Q (sonnet-5.5): Find patients IDs and samples in colorectal cancer that harbor the V600V alteration in BRAF ▶ read_guide { "uri": "cbioportal://common-pitfalls#16" } ▶ read_guide { "uri": "cbioportal://mutation-frequency-guide" } ◀ result ### 16. 🚨 SILENT QUERY SUBSTITUTION ("did you mean...") When the user's wording differs from canonical terminology (e.g. "V600V" looks like "V600E" with a typo, or "point mutation" sounds like "missense"), it is forbidden to silently rewrite the question and answer the rewritten version. Doing so produces an answer that looks confident but is for a different question — the user cannot tell what was changed. #### ❌ Wrong: silently substitute > User: *"Find patients in colorectal cancer with the V600V alteration in BRAF"* > Agent: *(internally treats this as V600E)* "I found 412 samples with BRAF V600E in colorectal studies..." > User: *"What is the most prevalent TP53 mutation in uterine cancer that is not a point mutation?"* > Agent: *(internally treats "point mutation" = "missense", silently excludes only missense)* "The most prevalent non-missense TP53 mutation is..." #### ✅ Correct: answer the literal question, flag any normalization For an unusual-looking variant the user may have typed deliberately: - Query for what was asked, literally. - If 0 rows come back, **explain *why* zero is the expected answer** before suggesting a likely-intended alternative. For synonymous variants (e.g. BRAF V600V, TP53 R175R), the explanation is: *cBioPortal's mutation tables filter out synonymous (silent) variants in most studies, so 0 hits means "filtered upstream", not "no such variant exists in any patient"*. Then ask: *"Did you mean V600E (the canonical activating variant)? Or would you like me to look for V600V in the studies that do retain synonymous calls?"* - If the wording is ambiguous (e.g. "point mutation"), ask the user which definition they meant before querying — do not pick one silently. #### Mutation-type terminology mapping (use this to disambiguate) | User says | Canonical definition | `mutation_type` filter | |---|---|---| | "point mutation" | Any SNV (single-nucleotide variant) — includes missense, nonsense, synonymous, splice-site SNVs | `mutation_type IN ('Missense_Mutation','Nonsense_Mutation','Silent','Splice_Site')` — **but ask the user to confirm scope first** | | "missense" | Single amino-acid substitution that changes the protein | `mutation_type = 'Missense_Mutation'` | | "nonsense" / "stop-gain" | Premature stop codon | `mutation_type = 'Nonsense_Mutation'` | | "synonymous" / "silent" | Nucleotide change with no amino-acid change | `mutation_type = 'Silent'` (**often filtered out of public datasets** — see below) | | "splice site" | Mutation in canonical splice acceptor/donor | `mutation_type = 'Splice_Site'` | | "frameshift" | Indel changing reading frame | `mutation_type IN ('Frame_Shift_Ins','Frame_Shift_Del')` | | "indel" / "in-frame" | In-frame insertion or deletion | `mutation_type IN ('In_Frame_Ins','In_Frame_Del')` | | "truncating" | Anything that disrupts the protein early | `mutation_type IN ('Nonsense_Mutation','Frame_Shift_Ins','Frame_Shift_Del','Splice_Site','Nonstop_Mutation')` | **Synonymous-variant filter. … (785 more chars) ▶ search_oncotree { "search_term": "colorectal" } ◀ result # Mutation Frequency Analysis Guide ## IMPORTANT: Reporting Mutation Frequencies - **ALWAYS report frequencies as percentages**, not raw counts: `frequency = (altered_samples / total_profiled_samples) × 100` - For quick frequency lookups, **prefer the TCGA Pan-Cancer Atlas study first**, then offer to expand to other studies - When reporting across multiple studies, show **ranges** (e.g., "TP53 is mutated in 30–60% of samples") rather than a single average - **NEVER** sum mutation events across studies to compute an aggregate frequency — this can exceed 100% due to double-counting - Warn users that samples may overlap across cohorts (e.g., MSK studies may share patients) - **Choose and state the counting unit**: use patient-level frequencies for prevalence/rate questions unless the user explicitly asks for samples; use sample-level frequencies when the user asks about samples. - **For "across cancer types" questions**, jump to the [Cross-Cancer-Type Mutation Frequency](#cross-cancer-type-mutation-frequency) section below — there is one correct recipe and several common wrong ones. ## Counting Unit: Samples vs Patients Before answering any mutation count or frequency question, decide whether the unit is samples or patients and state that choice in the answer. | User wording | Counting unit | |--------------|---------------| | "prevalence", "rate", "fraction of patients", "patients with", "how common is" | Patient-level: `COUNT(DISTINCT patient_unique_id)` | | "samples", "specimens", "biopsies", sample-level cohort composition | Sample-level: `COUNT(DISTINCT sample_unique_id)` | | Ambiguous | Ask, or default to patient-level for prevalence/rate language and say so | ### Cross-study sample-count caveat When an answer touches more than one study and reports a sample count, prepend a one-line caveat: > Sample IDs are unique within cBioPortal study prefixes, not guaranteed biological-sample identifiers across studies; overlapping cohorts can count the same patient/sample more than once. Prefer one of these safer approaches: - Use a shipped `cancer_study_query_preferences` cohort such as `pan_cancer_tcga` or `all_studies_non_redundant`. - Restrict to one named study. - Aggregate by `patient_unique_id` when the biological question is patient prevalence. ## STOP rule: a frequency above 100% means your query is wrong If your query returns a frequency over 100%, **do not try to debug or explain the data inconsistency to the user**. The cause is always one of these query bugs: - Summing mutation events instead of `COUNT(DISTINCT sample_unique_id)` for the numerator - Using a study-wide sample count as the denominator instead of the gene-specific profiled count - Cross-study aggregation where the same biological sample appears under multiple `sample_unique_id` values (e.g., MSK-IMPACT and MSK-CHORD share patients) - **Joining the profiled CTE through `gene_panel` / `gene_panel_list` without a WES branch.** `gene_panel_id = 'WES'` is *not* a row in … (36467 more chars) ◀ result {"result":[{"code":"COADREAD","name":"Colorectal Adenocarcinoma","score":60,"mainType":"Colorectal Cancer","tissue":"Bowel","hierarchy":"BOWEL > COADREAD"},{"code":"CAIS","name":"Colon Adenocarcinoma In Situ","score":45,"mainType":"Colorectal Cancer","tissue":"Bowel","hierarchy":"BOWEL > COADREAD > CAIS"},{"code":"CMC","name":"Medullary Carcinoma of the Colon","score":45,"mainType":"Colorectal Cancer","tissue":"Bowel","hierarchy":"BOWEL > CMC"},{"code":"COAD","name":"Colon Adenocarcinoma","score":45,"mainType":"Colorectal Cancer","tissue":"Bowel","hierarchy":"BOWEL > COADREAD > COAD"},{"code":"MACR","name":"Mucinous Adenocarcinoma of the Colon and Rectum","score":45,"mainType":"Colorectal Cancer","tissue":"Bowel","hierarchy":"BOWEL > COADREAD > MACR"},{"code":"READ","name":"Rectal Adenocarcinoma","score":45,"mainType":"Colorectal Cancer","tissue":"Bowel","hierarchy":"BOWEL > COADREAD > READ"},{"code":"SRCCR","name":"Signet Ring Cell Adenocarcinoma of the Colon and Rectum","score":45,"mainType":"Colorectal Cancer","tissue":"Bowel","hierarchy":"BOWEL > COADREAD > SRCCR"}]} ▶ clickhouse_list_table_columns { "table": "genomic_event_derived" } ◀ result {"columns":[{"name":"sample_unique_id","type":"String"},{"name":"hugo_gene_symbol","type":"String","comment":"HUGO gene symbol (e.g., TP53, KRAS, BRAF). Use for gene-specific queries."},{"name":"entrez_gene_id","type":"Int32"},{"name":"gene_panel_stable_id","type":"LowCardinality(String)"},{"name":"cancer_study_identifier","type":"LowCardinality(String)"},{"name":"genetic_profile_stable_id","type":"LowCardinality(String)"},{"name":"variant_type","type":"LowCardinality(String)","comment":"Type of genomic event: \"mutation\" for SNVs/indels, \"cna\" for copy number alterations, \"structural_variant\" for SVs. Always filter by this."},{"name":"mutation_variant","type":"String","comment":"Protein change notation (e.g., p.V600E, p.R175H). Use for specific variant queries. \"NA\" for non-mutation events."},{"name":"mutation_type","type":"LowCardinality(String)"},{"name":"mutation_status","type":"LowCardinality(String)","comment":"For mutations: Somatic, Germline, UNKNOWN, or UNCALLED. Filter mutation_status != \"UNCALLED\" to exclude uncertain calls. Include all other statuses."},{"name":"driver_filter","type":"LowCardinality(String)"},{"name":"driver_filter_annotation","type":"String"},{"name":"driver_tiers_filter","type":"LowCardinality(String)"},{"name":"driver_tiers_filter_annotation","type":"String"},{"name":"cna_alteration","type":"Nullable(Int8)","comment":"Copy number alteration: only 2 = amplification (AMP) and -2 = deep deletion (HOMDEL) are stored. Shallow deletion (-1), diploid (0) and gain (1) are not in this table; query genetic_alteration_derived WHERE profile_type = 'gistic' (alteration_value is a String, e.g. '-1'). NULL for non-CNA events."},{"name":"cna_cytoband","type":"String"},{"name":"sv_event_info","type":"String"},{"name":"patient_unique_id","type":"String"},{"name":"off_panel","type":"Bool","comment":"Boolean: 1 = mutation outside gene panel coverage (off-panel), 0 = within panel (on-panel). Filter off_panel = 0 for reliable frequency calculations."}]} ▶ clickhouse_run_select_query { "query": "SELECT g.cancer_study_identifier, g.patient_unique_id, g.sample_unique_id, g.mutation_variant, g.mutation_type FROM genomic_event_derived g WHERE g.variant_type='mutation' AND g.hugo_gene_symbol='BRAF' AND (g.mutation_variant LIKE '%V600V%' OR g.mutation_type IN ('Silent')) LIMIT 100" } ▶ clickhouse_run_select_query { "query": "SELECT mutation_type, count() c FROM genomic_event_derived WHERE variant_type='mutation' AND hugo_gene_symbol='BRAF' GROUP BY mutation_type" } ◀ result {"rows":[]} ◀ result {"rows":[{"mutation_type":"Frame_Shift_Ins","c":140},{"mutation_type":"frameshift_deletion","c":1},{"mutation_type":"Nonsense_Mutation","c":219},{"mutation_type":"In_Frame_Ins","c":108},{"mutation_type":"Missense_Mutation","c":15842},{"mutation_type":"Splice_Site","c":125},{"mutation_type":"In_Frame_Del","c":284},{"mutation_type":"Frame_Shift_Del","c":229},{"mutation_type":"Splice_Region","c":45}]} ═ answer (success) I found no patients or samples with BRAF V600V in any colorectal study, or in any other study in the database. I searched every BRAF mutation record for the exact V600V protein change and for any silent (synonymous) mutation, and got zero hits. V600V is a synonymous change. It swaps the nucleotides at codon 600 but leaves valine as valine. cBioPortal studies almost always remove silent variants during data import. Across all studies, BRAF has no silent mutation records at all. The types present are missense, nonsense, frameshift, in-frame, and splice. So zero hits means the variant was filtered out upstream. It doesn't show that no patient carries it. I searched literally for V600V and did not swap in another variant. Did you mean **V600E**, the canonical activating BRAF variant? If so, I can list the colorectal patients and samples with V600E. I'd need to know which colorectal study or studies to use. Options include the TCGA colorectal PanCancer Atlas study or MSK cohorts, and I can ask you to pick or cover all of them.