Q (sonnet): Find patients IDs and samples in colorectal cancer that harbor the V600V alteration in BRAF ▶ read_guide { "uri": "cbioportal://common-pitfalls#16" } ◀ result ### 16. 🚨 SILENT QUERY SUBSTITUTION ("did you mean...") When the user's wording differs from canonical terminology (e.g. "V600V" looks like "V600E" with a typo, or "point mutation" sounds like "missense"), it is forbidden to silently rewrite the question and answer the rewritten version. Doing so produces an answer that looks confident but is for a different question — the user cannot tell what was changed. #### ❌ Wrong: silently substitute > User: *"Find patients in colorectal cancer with the V600V alteration in BRAF"* > Agent: *(internally treats this as V600E)* "I found 412 samples with BRAF V600E in colorectal studies..." > User: *"What is the most prevalent TP53 mutation in uterine cancer that is not a point mutation?"* > Agent: *(internally treats "point mutation" = "missense", silently excludes only missense)* "The most prevalent non-missense TP53 mutation is..." #### ✅ Correct: answer the literal question, flag any normalization For an unusual-looking variant the user may have typed deliberately: - Query for what was asked, literally. - If 0 rows come back, **explain *why* zero is the expected answer** before suggesting a likely-intended alternative. For synonymous variants (e.g. BRAF V600V, TP53 R175R), the explanation is: *cBioPortal's mutation tables filter out synonymous (silent) variants in most studies, so 0 hits means "filtered upstream", not "no such variant exists in any patient"*. Then ask: *"Did you mean V600E (the canonical activating variant)? Or would you like me to look for V600V in the studies that do retain synonymous calls?"* - If the wording is ambiguous (e.g. "point mutation"), ask the user which definition they meant before querying — do not pick one silently. #### Mutation-type terminology mapping (use this to disambiguate) | User says | Canonical definition | `mutation_type` filter | |---|---|---| | "point mutation" | Any SNV (single-nucleotide variant) — includes missense, nonsense, synonymous, splice-site SNVs | `mutation_type IN ('Missense_Mutation','Nonsense_Mutation','Silent','Splice_Site')` — **but ask the user to confirm scope first** | | "missense" | Single amino-acid substitution that changes the protein | `mutation_type = 'Missense_Mutation'` | | "nonsense" / "stop-gain" | Premature stop codon | `mutation_type = 'Nonsense_Mutation'` | | "synonymous" / "silent" | Nucleotide change with no amino-acid change | `mutation_type = 'Silent'` (**often filtered out of public datasets** — see below) | | "splice site" | Mutation in canonical splice acceptor/donor | `mutation_type = 'Splice_Site'` | | "frameshift" | Indel changing reading frame | `mutation_type IN ('Frame_Shift_Ins','Frame_Shift_Del')` | | "indel" / "in-frame" | In-frame insertion or deletion | `mutation_type IN ('In_Frame_Ins','In_Frame_Del')` | | "truncating" | Anything that disrupts the protein early | `mutation_type IN ('Nonsense_Mutation','Frame_Shift_Ins','Frame_Shift_Del','Splice_Site','Nonstop_Mutation')` | **Synonymous-variant filter. … (785 more chars) ▶ read_guide { "uri": "cbioportal://mutation-frequency-guide" } ▶ search_oncotree { "search_term": "colorectal" } ◀ result # Mutation Frequency Analysis Guide ## IMPORTANT: Reporting Mutation Frequencies - **ALWAYS report frequencies as percentages**, not raw counts: `frequency = (altered_samples / total_profiled_samples) × 100` - For quick frequency lookups, **prefer the TCGA Pan-Cancer Atlas study first**, then offer to expand to other studies - When reporting across multiple studies, show **ranges** (e.g., "TP53 is mutated in 30–60% of samples") rather than a single average - **NEVER** sum mutation events across studies to compute an aggregate frequency — this can exceed 100% due to double-counting - Warn users that samples may overlap across cohorts (e.g., MSK studies may share patients) - **Choose and state the counting unit**: use patient-level frequencies for prevalence/rate questions unless the user explicitly asks for samples; use sample-level frequencies when the user asks about samples. - **For "across cancer types" questions**, jump to the [Cross-Cancer-Type Mutation Frequency](#cross-cancer-type-mutation-frequency) section below — there is one correct recipe and several common wrong ones. ## Counting Unit: Samples vs Patients Before answering any mutation count or frequency question, decide whether the unit is samples or patients and state that choice in the answer. | User wording | Counting unit | |--------------|---------------| | "prevalence", "rate", "fraction of patients", "patients with", "how common is" | Patient-level: `COUNT(DISTINCT patient_unique_id)` | | "samples", "specimens", "biopsies", sample-level cohort composition | Sample-level: `COUNT(DISTINCT sample_unique_id)` | | Ambiguous | Ask, or default to patient-level for prevalence/rate language and say so | ### Cross-study sample-count caveat When an answer touches more than one study and reports a sample count, prepend a one-line caveat: > Sample IDs are unique within cBioPortal study prefixes, not guaranteed biological-sample identifiers across studies; overlapping cohorts can count the same patient/sample more than once. Prefer one of these safer approaches: - Use a shipped `cancer_study_query_preferences` cohort such as `pan_cancer_tcga` or `all_studies_non_redundant`. - Restrict to one named study. - Aggregate by `patient_unique_id` when the biological question is patient prevalence. ## STOP rule: a frequency above 100% means your query is wrong If your query returns a frequency over 100%, **do not try to debug or explain the data inconsistency to the user**. The cause is always one of these query bugs: - Summing mutation events instead of `COUNT(DISTINCT sample_unique_id)` for the numerator - Using a study-wide sample count as the denominator instead of the gene-specific profiled count - Cross-study aggregation where the same biological sample appears under multiple `sample_unique_id` values (e.g., MSK-IMPACT and MSK-CHORD share patients) - **Joining the profiled CTE through `gene_panel` / `gene_panel_list` without a WES branch.** `gene_panel_id = 'WES'` is *not* a row in … (36467 more chars) ◀ result {"result":[{"code":"COADREAD","name":"Colorectal Adenocarcinoma","score":60,"mainType":"Colorectal Cancer","tissue":"Bowel","hierarchy":"BOWEL > COADREAD"},{"code":"CAIS","name":"Colon Adenocarcinoma In Situ","score":45,"mainType":"Colorectal Cancer","tissue":"Bowel","hierarchy":"BOWEL > COADREAD > CAIS"},{"code":"CMC","name":"Medullary Carcinoma of the Colon","score":45,"mainType":"Colorectal Cancer","tissue":"Bowel","hierarchy":"BOWEL > CMC"},{"code":"COAD","name":"Colon Adenocarcinoma","score":45,"mainType":"Colorectal Cancer","tissue":"Bowel","hierarchy":"BOWEL > COADREAD > COAD"},{"code":"MACR","name":"Mucinous Adenocarcinoma of the Colon and Rectum","score":45,"mainType":"Colorectal Cancer","tissue":"Bowel","hierarchy":"BOWEL > COADREAD > MACR"},{"code":"READ","name":"Rectal Adenocarcinoma","score":45,"mainType":"Colorectal Cancer","tissue":"Bowel","hierarchy":"BOWEL > COADREAD > READ"},{"code":"SRCCR","name":"Signet Ring Cell Adenocarcinoma of the Colon and Rectum","score":45,"mainType":"Colorectal Cancer","tissue":"Bowel","hierarchy":"BOWEL > COADREAD > SRCCR"}]} ▶ resolve_and_route { "studyKeywords": [ "colorectal", "TCGA" ] } ◀ result {"success":true,"message":"Found 3 matching studies. Pick the best match (prefer TCGA → prefer PanCancer Atlas), use its metadata to call the appropriate navigation tool(s) and generate URLs immediately. Present other studies as alternatives.","data":{"totalCount":3,"studiesWithMetadata":[{"studyId":"coadread_tcga","name":"Colorectal Adenocarcinoma (TCGA, Firehose Legacy)","sampleCount":640,"studyViewUrl":"https://www.cbioportal.org/study?id=coadread_tcga","metadata":{"clinicalAttributeIds":["AGE","AJCC_METASTASIS_PATHOLOGIC_PM","AJCC_NODES_PATHOLOGIC_PN","AJCC_PATHOLOGIC_TUMOR_STAGE","AJCC_STAGING_EDITION","AJCC_TUMOR_PATHOLOGIC_PT","BRAF_GENE_ANALYSIS_INDICATOR","BRAF_GENE_ANALYSIS_RESULT","CANCER_TYPE","CANCER_TYPE_DETAILED","CLINICAL_STAGE","CLIN_M_STAGE","CLIN_N_STAGE","CLIN_T_STAGE","DAYS_TO_COLLECTION","DAYS_TO_INITIAL_PATHOLOGIC_DIAGNOSIS","DAYS_TO_PATIENT_PROGRESSION_FREE","DAYS_TO_SPECIMEN_COLLECTION","DAYS_TO_TUMOR_PROGRESSION","DFS_MONTHS","DFS_STATUS","DISEASE_CODE","ETHNICITY","EXTRANODAL_INVOLVEMENT","FORM_COMPLETION_DATE","FRACTION_GENOME_ALTERED","HEIGHT","HISTOLOGICAL_DIAGNOSIS","HISTORY_NEOADJUVANT_TRTYN","HISTORY_OTHER_MALIGNANCY","ICD_10","ICD_O_3_HISTOLOGY","ICD_O_3_SITE","INFORMED_CONSENT_VERIFIED","INITIAL_PATHOLOGIC_DIAGNOSIS_METHOD","INITIAL_PATHOLOGIC_DX_YEAR","IS_FFPE","KRAS_GENE_ANALYSIS_INDICATOR","KRAS_MUTATION","LONGEST_DIMENSION","LYMPHOVASCULAR_INVASION_INDICATOR","LYMPH_NODES_EXAMINED","LYMPH_NODES_EXAMINED_HE_COUNT","LYMPH_NODES_EXAMINED_IHC_COUNT","LYMPH_NODE_EXAMINED_COUNT","METHOD_OF_SAMPLE_PROCUREMENT","MUTATION_COUNT","NEW_TUMOR_EVENT_AFTER_INITIAL_TREATMENT","OCT_EMBEDDED","ONCOTREE_CODE","OS_MONTHS","OS_STATUS","OTHER_METHOD_OF_SAMPLE_PROCUREMENT","OTHER_PATIENT_ID","OTHER_SAMPLE_ID","PATHOLOGY_REPORT_FILE_NAME","PATHOLOGY_REPORT_UUID","PERINEURAL_INVASION","PHARMACEUTICAL_TX_ADJUVANT","PRIMARY_SITE_PATIENT","PROJECT_CODE","PROSPECTIVE_COLLECTION","RACE","RADIATION_TREATMENT_ADJUVANT","RESIDUAL_TUMOR","RETROSPECTIVE_COLLECTION","SAMPLE_COUNT","SAMPLE_INITIAL_WEIGHT","SAMPLE_TYPE","SAMPLE_TYPE_ID","SEX","SHORTEST_DIMENSION","SITE_OF_TUMOR_TISSUE","SOMATIC_STATUS","SPECIMEN_CURRENT_WEIGHT","SPECIMEN_FREEZING_METHOD","SPECIMEN_SECOND_LONGEST_DIMENSION","STAGE_OTHER","TIME_BETWEEN_CLAMPING_AND_FREEZING","TIME_BETWEEN_EXCISION_AND_FREEZING","TISSUE_SOURCE_SITE","TMB_NONSYNONYMOUS","TREATMENT_OUTCOME_FIRST_COURSE","TUMOR_STATUS","VASCULAR_INVASION_INDICATOR","VIAL_NUMBER","WEIGHT"],"molecularProfileIds":["coadread_tcga_gistic","coadread_tcga_linear_CNA","coadread_tcga_methylation_hm450","coadread_tcga_mrna","coadread_tcga_mrna_median_Zscores","coadread_tcga_mrna_median_all_sample_Zscores","coadread_tcga_mutations","coadread_tcga_protein_quantification","coadread_tcga_protein_quantification_zscores","coadread_tcga_rna_seq_v2_mrna","coadread_tcga_rna_seq_v2_mrna_median_Zscores","coadread_tcga_rna_seq_v2_mrna_median_all_sample_Zscores","coadread_tcga_rppa","coadread_tcga_rppa_Zscores"],"heatmapProfileIds":["coadr … (4966 more chars) ▶ clickhouse_run_select_query { "query": "SELECT cancer_study_identifier, count() AS n FROM genomic_event_derived WHERE variant_type = 'mutation' AND hugo_gene_symbol = 'BRAF' AND mutation_variant = 'V600V' GROUP BY cancer_study_identifier ORDER BY n DESC" } ◀ result {"rows":[]} ▶ clickhouse_run_select_query { "query": "SELECT mutation_variant, mutation_type, count() AS n FROM genomic_event_derived WHERE variant_type = 'mutation' AND hugo_gene_symbol = 'BRAF' AND mutation_variant LIKE 'V600%' GROUP BY mutation_variant, mutation_type ORDER BY n DESC" } ◀ result {"rows":[{"mutation_variant":"V600E","mutation_type":"Missense_Mutation","n":8529},{"mutation_variant":"V600K","mutation_type":"Missense_Mutation","n":348},{"mutation_variant":"V600M","mutation_type":"Missense_Mutation","n":81},{"mutation_variant":"V600R","mutation_type":"Missense_Mutation","n":61},{"mutation_variant":"V600_K601delinsE","mutation_type":"In_Frame_Del","n":49},{"mutation_variant":"V600D","mutation_type":"Missense_Mutation","n":12},{"mutation_variant":"V600G","mutation_type":"Missense_Mutation","n":9},{"mutation_variant":"V600_W604delinsDG","mutation_type":"In_Frame_Del","n":6},{"mutation_variant":"V600_W604delinsE","mutation_type":"In_Frame_Del","n":6},{"mutation_variant":"V600L","mutation_type":"Missense_Mutation","n":5},{"mutation_variant":"V600_K601delinsEI","mutation_type":"Missense_Mutation","n":4},{"mutation_variant":"V600_S602delinsDT","mutation_type":"In_Frame_Del","n":4},{"mutation_variant":"V600_R603del","mutation_type":"In_Frame_Del","n":3},{"mutation_variant":"V600_S602del","mutation_type":"In_Frame_Del","n":2},{"mutation_variant":"V600_W604delinsR","mutation_type":"In_Frame_Del","n":2},{"mutation_variant":"V600del","mutation_type":"In_Frame_Del","n":1}]} ═ answer (success) This confirms it clearly across the entire database (all studies, not just colorectal): every codon-600 variant of BRAF present is a missense or in-frame change (V600E dominates with 8,529 hits, plus K/M/R/D/G/L variants and several in-frame deletions). **V600V — the synonymous/silent variant — has zero occurrences anywhere.** ## Why this is expected, not a data gap BRAF V600V (a synonymous mutation — the codon changes but still encodes valine) is filtered out during the standard MAF-processing pipeline that essentially all cBioPortal studies (including every Colorectal Adenocarcinoma cohort — TCGA PanCancer Atlas, Firehose Legacy, and Nature 2012) use upstream of ingestion. Silent/synonymous variants aren't loaded because they don't alter the protein and aren't considered clinically actionable. So "0 patients/samples" reflects **filtered upstream**, not "this variant doesn't exist in any tumor." ## Next step — please confirm which you'd like 1. **BRAF V600E** — the canonical activating driver mutation, present in colorectal cancer (notably associated with the aggressive, often MSI/serrated-pathway subtype). This is very likely what was intended. 2. **Search specifically for synonymous calls** in a study that might retain `Silent` mutation type records for BRAF (uncommon, but I can check). If you'd like option 1, here are the colorectal studies I'd query (patient/sample level, mutation frequency with proper profiled-sample denominators): - [Colorectal Adenocarcinoma (TCGA, PanCancer Atlas)](https://www.cbioportal.org/study?id=coadread_tcga_pan_can_atlas_2018) — 594 samples - [Colorectal Adenocarcinoma (TCGA, Firehose Legacy)](https://www.cbioportal.org/study?id=coadread_tcga) — 640 samples - [Colorectal Adenocarcinoma (TCGA, Nature 2012)](https://www.cbioportal.org/study?id=coadread_tcga_pub) — 276 samples Let me know which you'd like — V600E, or a check for silent-variant retention — and I'll pull the exact patient IDs and sample IDs plus a direct link to the Mutations tab.