Q (sonnet): What are the top 10 most frequently mutated genes across all cancer types in the MSK-CHORD Study? ▶ read_guide { "uri": "cbioportal://mutation-frequency-guide" } ▶ get_study_guide { "study_id": "msk_chord_2024" } ◀ result # Mutation Frequency Analysis Guide ## IMPORTANT: Reporting Mutation Frequencies - **ALWAYS report frequencies as percentages**, not raw counts: `frequency = (altered_samples / total_profiled_samples) × 100` - For quick frequency lookups, **prefer the TCGA Pan-Cancer Atlas study first**, then offer to expand to other studies - When reporting across multiple studies, show **ranges** (e.g., "TP53 is mutated in 30–60% of samples") rather than a single average - **NEVER** sum mutation events across studies to compute an aggregate frequency — this can exceed 100% due to double-counting - Warn users that samples may overlap across cohorts (e.g., MSK studies may share patients) - **Choose and state the counting unit**: use patient-level frequencies for prevalence/rate questions unless the user explicitly asks for samples; use sample-level frequencies when the user asks about samples. - **For "across cancer types" questions**, jump to the [Cross-Cancer-Type Mutation Frequency](#cross-cancer-type-mutation-frequency) section below — there is one correct recipe and several common wrong ones. ## Counting Unit: Samples vs Patients Before answering any mutation count or frequency question, decide whether the unit is samples or patients and state that choice in the answer. | User wording | Counting unit | |--------------|---------------| | "prevalence", "rate", "fraction of patients", "patients with", "how common is" | Patient-level: `COUNT(DISTINCT patient_unique_id)` | | "samples", "specimens", "biopsies", sample-level cohort composition | Sample-level: `COUNT(DISTINCT sample_unique_id)` | | Ambiguous | Ask, or default to patient-level for prevalence/rate language and say so | ### Cross-study sample-count caveat When an answer touches more than one study and reports a sample count, prepend a one-line caveat: > Sample IDs are unique within cBioPortal study prefixes, not guaranteed biological-sample identifiers across studies; overlapping cohorts can count the same patient/sample more than once. Prefer one of these safer approaches: - Use a shipped `cancer_study_query_preferences` cohort such as `pan_cancer_tcga` or `all_studies_non_redundant`. - Restrict to one named study. - Aggregate by `patient_unique_id` when the biological question is patient prevalence. ## STOP rule: a frequency above 100% means your query is wrong If your query returns a frequency over 100%, **do not try to debug or explain the data inconsistency to the user**. The cause is always one of these query bugs: - Summing mutation events instead of `COUNT(DISTINCT sample_unique_id)` for the numerator - Using a study-wide sample count as the denominator instead of the gene-specific profiled count - Cross-study aggregation where the same biological sample appears under multiple `sample_unique_id` values (e.g., MSK-IMPACT and MSK-CHORD share patients) - **Joining the profiled CTE through `gene_panel` / `gene_panel_list` without a WES branch.** `gene_panel_id = 'WES'` is *not* a row in … (35166 more chars) ◀ result # MSK-CHORD (MSK, Nature 2024) **Study ID:** `msk_chord_2024` ## Overview Targeted sequencing via MSK-IMPACT panels. Clinical annotations include some derived from natural language processing (denoted NLP). **Exactly five cancer types** (`CANCER_TYPE`, patients): Non-Small Cell Lung Cancer 7,809, Colorectal Cancer 5,543, Breast Cancer 5,368, Prostate Cancer 3,211, Pancreatic Cancer 3,109. There is **no melanoma** or any other cancer type; say so up front if asked, instead of substituting another type. **No therapy-response variable.** There is no RECIST, objective response, or best-response attribute or event. For treatment-outcome questions (e.g. immunotherapy response), say this first; the only proxies are `OS_MONTHS`/`OS_STATUS`, or NLP radiology progression events (`Diagnosis` events with `SUBTYPE = 'Progression'`, key `PROGRESSION` = Y/N/Indeterminate), in patients with `Treatment` events of the relevant `SUBTYPE` (e.g. `Immuno`: 3,341 patients). Hand off the comparison to cBioPortal group comparison / survival. **Nearly one sample per patient: 24,950 patients / 25,040 samples.** Only 90 patients have more than one sample, and all 90 have samples from two different cancer types (second primaries); only 26 have both a `Primary` and a `Metastasis` sample. There is no meaningful same-patient (paired) primary-vs-metastasis cohort. For "same patient" / paired questions, say this up front, then offer the **unpaired** comparison of all `Primary` vs `Metastasis` samples (`SAMPLE_TYPE`), labelled as unpaired. ```sql SELECT countIf(n > 1) AS multi_sample_patients, -- 90 countIf(has_p AND has_m) AS primary_and_met -- 26 FROM (SELECT patient_unique_id, count() AS n, has(groupArray(attribute_value), 'Primary') AS has_p, has(groupArray(attribute_value), 'Metastasis') AS has_m FROM clinical_data_derived WHERE cancer_study_identifier = 'msk_chord_2024' AND attribute_name = 'SAMPLE_TYPE' GROUP BY patient_unique_id); ``` ## Gene Panels This study uses multiple MSK-IMPACT panel versions: - **IMPACT341**: Earlier version, 341 genes - **IMPACT410**: 410 genes - **IMPACT468**: 468 genes - **IMPACT505**: Latest version, 505 genes **Important:** Different samples may have different gene coverage. Always use gene-specific denominators when calculating mutation frequencies. ## Clinical Attributes - Semantic Guide ### Cancer Classification | Attribute | Description | Values | |-----------|-------------|--------| | `CANCER_TYPE` | Broad cancer category | e.g., "Non-Small Cell Lung Cancer", "Breast Cancer" | | `CANCER_TYPE_DETAILED` | Specific subtype | e.g., "Lung Adenocarcinoma", "Invasive Ductal Carcinoma" | | `ONCOTREE_CODE` | OncoTree classification code | Standardized cancer type codes | ### Sample Information | Attribute | Description | Values | |-----------|-------------|--------| | `SAMPLE_TYPE` | Sample origin | Primary, Metastasis, Local Recurrence, Unknown | | `SAMPLE_CLASS` | Sample … (5439 more chars) ▶ clickhouse_run_select_query { "query": "SELECT hugo_gene_symbol, total_mutation_events AS mutations, altered_samples, profiled_samples, frequency_pct FROM top_mutated_genes_in_study(study = 'msk_chord_2024', top_n = 10)" } ◀ result {"rows":[{"hugo_gene_symbol":"TP53","mutations":13876,"altered_samples":13124,"profiled_samples":25040,"frequency_pct":52.4},{"hugo_gene_symbol":"KRAS","mutations":7231,"altered_samples":7128,"profiled_samples":25040,"frequency_pct":28.5},{"hugo_gene_symbol":"APC","mutations":7187,"altered_samples":4777,"profiled_samples":25040,"frequency_pct":19.1},{"hugo_gene_symbol":"PIK3CA","mutations":4178,"altered_samples":3708,"profiled_samples":25040,"frequency_pct":14.8},{"hugo_gene_symbol":"EGFR","mutations":2556,"altered_samples":2159,"profiled_samples":25040,"frequency_pct":8.6},{"hugo_gene_symbol":"ARID1A","mutations":2205,"altered_samples":1843,"profiled_samples":25040,"frequency_pct":7.4},{"hugo_gene_symbol":"SMAD4","mutations":1925,"altered_samples":1796,"profiled_samples":25040,"frequency_pct":7.2},{"hugo_gene_symbol":"KMT2D","mutations":2390,"altered_samples":1783,"profiled_samples":25040,"frequency_pct":7.1},{"hugo_gene_symbol":"KMT2C","mutations":1979,"altered_samples":1629,"profiled_samples":25040,"frequency_pct":6.5},{"hugo_gene_symbol":"ATM","mutations":1681,"altered_samples":1388,"profiled_samples":25040,"frequency_pct":5.5}]} ▶ resolve_and_route { "studyIds": [ "msk_chord_2024" ] } ◀ result {"success":true,"message":"Found 1 study. Use the metadata to call the appropriate navigation tool(s).","data":{"totalCount":1,"studiesWithMetadata":[{"studyId":"msk_chord_2024","name":"MSK-CHORD (MSK, Nature 2024)","sampleCount":25040,"studyViewUrl":"https://www.cbioportal.org/study?id=msk_chord_2024","metadata":{"clinicalAttributeIds":["ADRENAL_GLANDS","BONE","CANCER_TYPE","CANCER_TYPE_DETAILED","CLINICAL_GROUP","CLINICAL_SUMMARY","CNS_BRAIN","CURRENT_AGE_DEID","DIAGNOSIS_DESCRIPTION","ETHNICITY","FRACTION_GENOME_ALTERED","GENDER","GENE_PANEL","GLEASON_FIRST_REPORTED","GLEASON_HIGHEST_REPORTED","GLEASON_SAMPLE_LEVEL","HER2","HISTORY_OF_PDL1","HR","ICD_O_HISTOLOGY_DESCRIPTION","INTRA_ABDOMINAL","LIVER","LUNG","LYMPH_NODES","METASTATIC_SITE","MSI_COMMENT","MSI_SCORE","MSI_TYPE","MUTATION_COUNT","NUM_ICDO_DX","ONCOTREE_CODE","OS_MONTHS","OS_STATUS","OTHER","PATHOLOGICAL_GROUP","PDL1_POSITIVE","PLEURA","PRIMARY_SITE","PRIOR_MED_TO_MSK","RACE","REPRODUCTIVE_ORGANS","SAMPLE_CLASS","SAMPLE_COUNT","SAMPLE_COVERAGE","SAMPLE_TYPE","SMOKING_PREDICTIONS_3_CLASSES","SOMATIC_STATUS","STAGE_HIGHEST_RECORDED","TMB_NONSYNONYMOUS","TUMOR_PURITY"],"molecularProfileIds":["msk_chord_2024_cna","msk_chord_2024_mutations","msk_chord_2024_structural_variants"],"availableComparisonTabs":["overlap","clinical","survival","alterations","mutations"]}}]}} ▶ navigate_to_results_view { "studyIds": [ "msk_chord_2024" ], "genes": [ "TP53", "KRAS", "APC", "PIK3CA", "EGFR", "ARID1A", "SMAD4", "KMT2D", "KMT2C", "ATM" ], "tab": "oncoprint" } ◀ result {"success":true,"message":"Navigating to https://www.cbioportal.org/results/oncoprint?cancer_study_list=msk_chord_2024&gene_list=TP53%0AKRAS%0AAPC%0APIK3CA%0AEGFR%0AARID1A%0ASMAD4%0AKMT2D%0AKMT2C%0AATM&case_set_id=msk_chord_2024_all","url":"https://www.cbioportal.org/results/oncoprint?cancer_study_list=msk_chord_2024&gene_list=TP53%0AKRAS%0AAPC%0APIK3CA%0AEGFR%0AARID1A%0ASMAD4%0AKMT2D%0AKMT2C%0AATM&case_set_id=msk_chord_2024_all","data":{"studyIds":["msk_chord_2024"],"studies":[{"studyId":"msk_chord_2024","name":"MSK-CHORD (MSK, Nature 2024)","sampleCount":25040}],"genes":["TP53","KRAS","APC","PIK3CA","EGFR","ARID1A","SMAD4","KMT2D","KMT2C","ATM"],"caseSetId":"msk_chord_2024_all","pageDescription":"Visual summary of alterations in queried genes across samples."}} ═ answer (success) ## Top 10 Most Frequently Mutated Genes — MSK-CHORD (`msk_chord_2024`) Sample-level frequencies, using gene-specific profiled denominators (25,040 total profiled samples for all genes; MSK-IMPACT panel-aware, WES-inclusive): | Rank | Gene | Altered Samples | Profiled Samples | Frequency | |------|------|-----------------|-------------------|-----------| | 1 | **TP53** | 13,124 | 25,040 | 52.4% | | 2 | **KRAS** | 7,128 | 25,040 | 28.5% | | 3 | **APC** | 4,777 | 25,040 | 19.1% | | 4 | **PIK3CA** | 3,708 | 25,040 | 14.8% | | 5 | **EGFR** | 2,159 | 25,040 | 8.6% | | 6 | **ARID1A** | 1,843 | 25,040 | 7.4% | | 7 | **SMAD4** | 1,796 | 25,040 | 7.2% | | 8 | **KMT2D** | 1,783 | 25,040 | 7.1% | | 9 | **KMT2C** | 1,629 | 25,040 | 6.5% | | 10 | **ATM** | 1,388 | 25,040 | 5.5% | **Note:** MSK-CHORD spans five cancer types (NSCLC, colorectal, breast, prostate, pancreatic cancer), each with different underlying mutation profiles. Several of these genes (APC, KRAS) are strongly driven by colorectal cancer representation, and EGFR by NSCLC. If you want the breakdown by individual cancer type, I can run that next. [View OncoPrint for Top 10 Genes in MSK-CHORD](https://www.cbioportal.org/results/oncoprint?cancer_study_list=msk_chord_2024&gene_list=TP53%0AKRAS%0AAPC%0APIK3CA%0AEGFR%0AARID1A%0ASMAD4%0AKMT2D%0AKMT2C%0AATM&case_set_id=msk_chord_2024_all) Visual summary of alterations in queried genes across samples (25,040 samples, MSK-CHORD).