Q (haiku): show me point mutations in EGFR in lung cancer ▶ read_guide { "uri": "cbioportal://mutation-frequency-guide" } ▶ read_guide { "uri": "cbioportal://common-pitfalls#16" } ◀ result # Mutation Frequency Analysis Guide ## IMPORTANT: Reporting Mutation Frequencies - **ALWAYS report frequencies as percentages**, not raw counts: `frequency = (altered_samples / total_profiled_samples) × 100` - For quick frequency lookups, **prefer the TCGA Pan-Cancer Atlas study first**, then offer to expand to other studies - When reporting across multiple studies, show **ranges** (e.g., "TP53 is mutated in 30–60% of samples") rather than a single average - **NEVER** sum mutation events across studies to compute an aggregate frequency — this can exceed 100% due to double-counting - Warn users that samples may overlap across cohorts (e.g., MSK studies may share patients) - **Choose and state the counting unit**: use patient-level frequencies for prevalence/rate questions unless the user explicitly asks for samples; use sample-level frequencies when the user asks about samples. - **For "across cancer types" questions**, jump to the [Cross-Cancer-Type Mutation Frequency](#cross-cancer-type-mutation-frequency) section below — there is one correct recipe and several common wrong ones. ## Counting Unit: Samples vs Patients Before answering any mutation count or frequency question, decide whether the unit is samples or patients and state that choice in the answer. | User wording | Counting unit | |--------------|---------------| | "prevalence", "rate", "fraction of patients", "patients with", "how common is" | Patient-level: `COUNT(DISTINCT patient_unique_id)` | | "samples", "specimens", "biopsies", sample-level cohort composition | Sample-level: `COUNT(DISTINCT sample_unique_id)` | | Ambiguous | Ask, or default to patient-level for prevalence/rate language and say so | ### Cross-study sample-count caveat When an answer touches more than one study and reports a sample count, prepend a one-line caveat: > Sample IDs are unique within cBioPortal study prefixes, not guaranteed biological-sample identifiers across studies; overlapping cohorts can count the same patient/sample more than once. Prefer one of these safer approaches: - Use a shipped `cancer_study_query_preferences` cohort such as `pan_cancer_tcga` or `all_studies_non_redundant`. - Restrict to one named study. - Aggregate by `patient_unique_id` when the biological question is patient prevalence. ## STOP rule: a frequency above 100% means your query is wrong If your query returns a frequency over 100%, **do not try to debug or explain the data inconsistency to the user**. The cause is always one of these query bugs: - Summing mutation events instead of `COUNT(DISTINCT sample_unique_id)` for the numerator - Using a study-wide sample count as the denominator instead of the gene-specific profiled count - Cross-study aggregation where the same biological sample appears under multiple `sample_unique_id` values (e.g., MSK-IMPACT and MSK-CHORD share patients) - **Joining the profiled CTE through `gene_panel` / `gene_panel_list` without a WES branch.** `gene_panel_id = 'WES'` is *not* a row in … (36467 more chars) ▶ resolve_and_route { "studyKeywords": [ "lung", "adenocarcinoma" ] } ◀ result ### 16. 🚨 SILENT QUERY SUBSTITUTION ("did you mean...") When the user's wording differs from canonical terminology (e.g. "V600V" looks like "V600E" with a typo, or "point mutation" sounds like "missense"), it is forbidden to silently rewrite the question and answer the rewritten version. Doing so produces an answer that looks confident but is for a different question — the user cannot tell what was changed. #### ❌ Wrong: silently substitute > User: *"Find patients in colorectal cancer with the V600V alteration in BRAF"* > Agent: *(internally treats this as V600E)* "I found 412 samples with BRAF V600E in colorectal studies..." > User: *"What is the most prevalent TP53 mutation in uterine cancer that is not a point mutation?"* > Agent: *(internally treats "point mutation" = "missense", silently excludes only missense)* "The most prevalent non-missense TP53 mutation is..." #### ✅ Correct: answer the literal question, flag any normalization For an unusual-looking variant the user may have typed deliberately: - Query for what was asked, literally. - If 0 rows come back, **explain *why* zero is the expected answer** before suggesting a likely-intended alternative. For synonymous variants (e.g. BRAF V600V, TP53 R175R), the explanation is: *cBioPortal's mutation tables filter out synonymous (silent) variants in most studies, so 0 hits means "filtered upstream", not "no such variant exists in any patient"*. Then ask: *"Did you mean V600E (the canonical activating variant)? Or would you like me to look for V600V in the studies that do retain synonymous calls?"* - If the wording is ambiguous (e.g. "point mutation"), ask the user which definition they meant before querying — do not pick one silently. #### Mutation-type terminology mapping (use this to disambiguate) | User says | Canonical definition | `mutation_type` filter | |---|---|---| | "point mutation" | Any SNV (single-nucleotide variant) — includes missense, nonsense, synonymous, splice-site SNVs | `mutation_type IN ('Missense_Mutation','Nonsense_Mutation','Silent','Splice_Site')` — **but ask the user to confirm scope first** | | "missense" | Single amino-acid substitution that changes the protein | `mutation_type = 'Missense_Mutation'` | | "nonsense" / "stop-gain" | Premature stop codon | `mutation_type = 'Nonsense_Mutation'` | | "synonymous" / "silent" | Nucleotide change with no amino-acid change | `mutation_type = 'Silent'` (**often filtered out of public datasets** — see below) | | "splice site" | Mutation in canonical splice acceptor/donor | `mutation_type = 'Splice_Site'` | | "frameshift" | Indel changing reading frame | `mutation_type IN ('Frame_Shift_Ins','Frame_Shift_Del')` | | "indel" / "in-frame" | In-frame insertion or deletion | `mutation_type IN ('In_Frame_Ins','In_Frame_Del')` | | "truncating" | Anything that disrupts the protein early | `mutation_type IN ('Nonsense_Mutation','Frame_Shift_Ins','Frame_Shift_Del','Splice_Site','Nonstop_Mutation')` | **Synonymous-variant filter. … (785 more chars) ◀ result {"success":true,"message":"Found 18 matching studies (top 5 with full metadata, rest with basic info). Pick the best match (prefer TCGA → prefer PanCancer Atlas), use its metadata to call the appropriate navigation tool(s) and generate URLs immediately. Present other studies as alternatives.","data":{"totalCount":18,"studiesWithMetadata":[{"studyId":"msk_met_2021","name":"MSK MetTropism (MSK, Cell 2021)","sampleCount":25775,"studyViewUrl":"https://www.cbioportal.org/study?id=msk_met_2021","metadata":{"clinicalAttributeIds":["AGE_AT_DEATH","AGE_AT_EVIDENCE_OF_METS","AGE_AT_LAST_CONTACT","AGE_AT_SEQUENCING","AGE_AT_SURGERY","CANCER_TYPE","CANCER_TYPE_DETAILED","DMETS_DX_ADRENAL_GLAND","DMETS_DX_BILIARY_TRACT","DMETS_DX_BLADDER_UT","DMETS_DX_BONE","DMETS_DX_BOWEL","DMETS_DX_BREAST","DMETS_DX_CNS_BRAIN","DMETS_DX_DIST_LN","DMETS_DX_FEMALE_GENITAL","DMETS_DX_HEAD_NECK","DMETS_DX_INTRA_ABDOMINAL","DMETS_DX_KIDNEY","DMETS_DX_LIVER","DMETS_DX_LUNG","DMETS_DX_MALE_GENITAL","DMETS_DX_MEDIASTINUM","DMETS_DX_OVARY","DMETS_DX_PLEURA","DMETS_DX_PNS","DMETS_DX_SKIN","DMETS_DX_UNSPECIFIED","FGA","FRACTION_GENOME_ALTERED","GENE_PANEL","IS_DIST_MET_MAPPED","METASTATIC_SITE","MET_COUNT","MET_SITE_COUNT","MSI_SCORE","MSI_TYPE","MUTATION_COUNT","ONCOTREE_CODE","ORGAN_SYSTEM","OS_MONTHS","OS_STATUS","PRIMARY_SITE","RACE","SAMPLE_COUNT","SAMPLE_COVERAGE","SAMPLE_TYPE","SEX","SUBTYPE","SUBTYPE_ABBREVIATION","TMB_NONSYNONYMOUS","TUMOR_PURITY"],"molecularProfileIds":["msk_met_2021_cna","msk_met_2021_mutations","msk_met_2021_structural_variants"],"availableComparisonTabs":["overlap","clinical","survival","alterations","mutations"]}},{"studyId":"luad_mskcc_2023_met_organotropism","name":"Lung Adenocarcinoma Met Organotropism (MSK, Cancer Cell 2023)","sampleCount":2653,"studyViewUrl":"https://www.cbioportal.org/study?id=luad_mskcc_2023_met_organotropism","metadata":{"clinicalAttributeIds":["ADJUVANT","ADJUVANT_CHEMOTHERAPY","ADJUVANT_IMMUNOTHERAPY","ADJUVANT_TARGETED","ADJUVANT_THERAPY","ADJUVANT_XRT","ADRENAL_MONTHS","ADRENAL_STATUS","AGE_AT_DOS_BX","BONE_MONTHS","BONE_STATUS","CANCER_TYPE","CANCER_TYPE_DETAILED","CELL_CYCLE","CIGARETTE_HX","CNS_MONTHS","CNS_STATUS","CSTAGE","DEATH","EVER_MET_SITE_ADRENAL","EVER_MET_SITE_BONE","EVER_MET_SITE_CNS","EVER_MET_SITE_LIVER_BILIARY_TRACT","EVER_MET_SITE_LN","EVER_MET_SITE_LUNG","EVER_MET_SITE_PLEURA","FGA","FRACTION_GENOME_ALTERED","FU_2YRS","GENE_PANEL","GROUP_NO","HAD_SURGERY","HIPPO","IMPACT_METASTATIC_LESION","IMPACT_PRIMARY_GROUP","INSTITUTE","IN_MATCHED","IS_WGD","LIVER_MONTHS","LIVER_STATUS","LN_MONTHS","LN_STATUS","LUNG_MONTHS","LUNG_STATUS","METASTATIC_BURDEN","METASTATIC_SITE","MONTHS_FROM_MATCHED_PRIM","MSI_SCORE","MSI_TYPE","MUTATION_COUNT","MYC_PATH","NEOADJUVANT","NEOADJUVANT_CHEMOTHERAPY","NEOADJUVANT_IMMUNOTHERAPY","NEOADJUVANT_TARGETED","NEOADJUVANT_XRT","NOTCH","NRF2","ONCOTREE_CODE","OS_MONTHS","OS_STATUS","PI3K","PLEURA_MONTHS","PLEURA_STATUS","PLOIDY","POST_SAMPLE_CHEMOTHERAPY","POST_SAMPLE_IMMUNOTHERAPY","POS … (9438 more chars) ═ answer (success) Good — I've loaded the guides on mutation terminology and point mutations. The key insight: **"point mutation" is ambiguous** — it can mean any SNV (including missense, nonsense, synonymous, splice), and I need to clarify with you first rather than assume. Also, **synonymous variants are filtered out of most cBioPortal studies**, which is important context. Now, I found multiple lung adenocarcinoma studies. The **TCGA PanCancer Atlas version** (`luad_tcga_pan_can_atlas_2018`, 566 samples) is the canonical lung adenocarcinoma cohort for cross-study consistency. I'll use that for the query. Let me clarify your question first: **By "point mutations," do you mean:** 1. **All SNVs** (single nucleotide variants, including missense, nonsense, splice site)? 2. **Missense mutations only** (amino acid changes)? 3. **Non-synonymous variants** (anything that changes the protein — missense, nonsense, frameshift)? Once you clarify, I'll retrieve the specific EGFR variants and link you to the mutations visualization in cBioPortal. In the meantime, here's what I know about EGFR in lung adenocarcinoma from the guides: - EGFR is one of the most frequently mutated genes in LUAD (~40% in TCGA) - Common driver variants include **exon 19 deletions**, **L858R (exon 21)**, and other activating mutations - These are the therapeutic targets for EGFR inhibitors like erlotinib and gefitinib **Which definition of "point mutations" would you like me to use?**