Q (sonnet): How does the mutation landscape differ between primary and metastatic samples from the same patients in the MSK-CHORD Study? ▶ read_guide { "uri": "cbioportal://mutation-frequency-guide" } ▶ read_guide { "uri": "cbioportal://sample-filtering-guide" } ◀ result # Mutation Frequency Analysis Guide ## IMPORTANT: Reporting Mutation Frequencies - **ALWAYS report frequencies as percentages**, not raw counts: `frequency = (altered_samples / total_profiled_samples) × 100` - For quick frequency lookups, **prefer the TCGA Pan-Cancer Atlas study first**, then offer to expand to other studies - When reporting across multiple studies, show **ranges** (e.g., "TP53 is mutated in 30–60% of samples") rather than a single average - **NEVER** sum mutation events across studies to compute an aggregate frequency — this can exceed 100% due to double-counting - Warn users that samples may overlap across cohorts (e.g., MSK studies may share patients) - **Choose and state the counting unit**: use patient-level frequencies for prevalence/rate questions unless the user explicitly asks for samples; use sample-level frequencies when the user asks about samples. - **For "across cancer types" questions**, jump to the [Cross-Cancer-Type Mutation Frequency](#cross-cancer-type-mutation-frequency) section below — there is one correct recipe and several common wrong ones. ## Counting Unit: Samples vs Patients Before answering any mutation count or frequency question, decide whether the unit is samples or patients and state that choice in the answer. | User wording | Counting unit | |--------------|---------------| | "prevalence", "rate", "fraction of patients", "patients with", "how common is" | Patient-level: `COUNT(DISTINCT patient_unique_id)` | | "samples", "specimens", "biopsies", sample-level cohort composition | Sample-level: `COUNT(DISTINCT sample_unique_id)` | | Ambiguous | Ask, or default to patient-level for prevalence/rate language and say so | ### Cross-study sample-count caveat When an answer touches more than one study and reports a sample count, prepend a one-line caveat: > Sample IDs are unique within cBioPortal study prefixes, not guaranteed biological-sample identifiers across studies; overlapping cohorts can count the same patient/sample more than once. Prefer one of these safer approaches: - Use a shipped `cancer_study_query_preferences` cohort such as `pan_cancer_tcga` or `all_studies_non_redundant`. - Restrict to one named study. - Aggregate by `patient_unique_id` when the biological question is patient prevalence. ## STOP rule: a frequency above 100% means your query is wrong If your query returns a frequency over 100%, **do not try to debug or explain the data inconsistency to the user**. The cause is always one of these query bugs: - Summing mutation events instead of `COUNT(DISTINCT sample_unique_id)` for the numerator - Using a study-wide sample count as the denominator instead of the gene-specific profiled count - Cross-study aggregation where the same biological sample appears under multiple `sample_unique_id` values (e.g., MSK-IMPACT and MSK-CHORD share patients) - **Joining the profiled CTE through `gene_panel` / `gene_panel_list` without a WES branch.** `gene_panel_id = 'WES'` is *not* a row in … (35166 more chars) ▶ get_study_guide { "study_id": "msk_chord_2024" } ◀ result # Sample and Study Filtering Guide ## Overview Proper filtering is essential for meaningful cBioPortal analysis. This guide covers filtering by studies, sample types, and other criteria. ## Study-Level Filtering ### 1. Always Filter by Study Every query should specify a study to ensure consistent results: ```sql -- Always include study filtering SELECT * FROM your_table WHERE cancer_study_identifier = 'your_study_id' -- Additional filters... ``` ### 2. Find Available Studies ```sql -- Discover available studies SELECT cancer_study_identifier, name, description, type_of_cancer_id FROM cancer_study ORDER BY cancer_study_identifier; ``` ### 3. Study Information ```sql -- Get detailed study information SELECT cs.cancer_study_identifier, cs.name as study_name, cs.description, cs.sample_count, COUNT(DISTINCT p.internal_id) as patient_count FROM cancer_study cs LEFT JOIN patient p ON cs.cancer_study_id = p.cancer_study_id WHERE cs.cancer_study_identifier = 'your_study_id' GROUP BY cs.cancer_study_identifier, cs.name, cs.description, cs.sample_count; ``` ### 4. Find Studies by Available Data Types Use this when the user asks *"which studies have mutation and copy-number data for X"*, *"studies with expression for Y"*, *"is there any imaging data"*, or any *"studies with Z data"* question. `cancer_study` has one precomputed sample count per data type — filter on them. No join and no schema exploration needed. These are the same counts as the portal's study list and its "Data type" filter. **Canonical query — studies with ALL requested data types for a cancer type:** ```sql SELECT cancer_study_identifier, name, sample_count, mutation_sample_count, cna_sample_count FROM cancer_study WHERE type_of_cancer_id = 'luad' -- ← OncoTree code from search_oncotree AND mutation_sample_count > 0 -- ← one condition per required data type AND cna_sample_count > 0 ORDER BY sample_count DESC; ``` **Data type → column** (a count of 0 means the study doesn't have that data): | Data type in the user's question | Column (portal "Data type" filter label) | |---|---| | samples in the study | `sample_count` | | mutation / mutations | `mutation_sample_count` ("Mutations") | | copy-number / CNA / amplification / deletion | `cna_sample_count` ("CNA") | | structural variant / fusion | `structural_variant_sample_count` | | RNA / mRNA / gene expression (any platform) | `mrna_expression_sample_count` — use this for "which studies have expression data" | | specifically the portal's "RNA-Seq" filter | `rna_seq_sample_count` ("RNA-Seq") — counts only the `_rna_seq_v2_mrna` list; 0 for many GDC and iAtlas studies that do have RNA-Seq | | specifically the portal's "RNA (microarray)" filter | `mrna_microarray_sample_count` ("RNA (microarray)") — counts only the `_mrna` list | | microRNA | `mirna_sample_count` ("miRNA") | | protein / RPPA | `rppa_sample_count` ("RPPA") | | mas … (7083 more chars) ◀ result # MSK-CHORD (MSK, Nature 2024) **Study ID:** `msk_chord_2024` ## Overview Targeted sequencing via MSK-IMPACT panels. Clinical annotations include some derived from natural language processing (denoted NLP). **Exactly five cancer types** (`CANCER_TYPE`, patients): Non-Small Cell Lung Cancer 7,809, Colorectal Cancer 5,543, Breast Cancer 5,368, Prostate Cancer 3,211, Pancreatic Cancer 3,109. There is **no melanoma** or any other cancer type; say so up front if asked, instead of substituting another type. **No therapy-response variable.** There is no RECIST, objective response, or best-response attribute or event. For treatment-outcome questions (e.g. immunotherapy response), say this first; the only proxies are `OS_MONTHS`/`OS_STATUS`, or NLP radiology progression events (`Diagnosis` events with `SUBTYPE = 'Progression'`, key `PROGRESSION` = Y/N/Indeterminate), in patients with `Treatment` events of the relevant `SUBTYPE` (e.g. `Immuno`: 3,341 patients). Hand off the comparison to cBioPortal group comparison / survival. **Nearly one sample per patient: 24,950 patients / 25,040 samples.** Only 90 patients have more than one sample, and all 90 have samples from two different cancer types (second primaries); only 26 have both a `Primary` and a `Metastasis` sample. There is no meaningful same-patient (paired) primary-vs-metastasis cohort. For "same patient" / paired questions, say this up front, then offer the **unpaired** comparison of all `Primary` vs `Metastasis` samples (`SAMPLE_TYPE`), labelled as unpaired. ```sql SELECT countIf(n > 1) AS multi_sample_patients, -- 90 countIf(has_p AND has_m) AS primary_and_met -- 26 FROM (SELECT patient_unique_id, count() AS n, has(groupArray(attribute_value), 'Primary') AS has_p, has(groupArray(attribute_value), 'Metastasis') AS has_m FROM clinical_data_derived WHERE cancer_study_identifier = 'msk_chord_2024' AND attribute_name = 'SAMPLE_TYPE' GROUP BY patient_unique_id); ``` ## Gene Panels This study uses multiple MSK-IMPACT panel versions: - **IMPACT341**: Earlier version, 341 genes - **IMPACT410**: 410 genes - **IMPACT468**: 468 genes - **IMPACT505**: Latest version, 505 genes **Important:** Different samples may have different gene coverage. Always use gene-specific denominators when calculating mutation frequencies. ## Clinical Attributes - Semantic Guide ### Cancer Classification | Attribute | Description | Values | |-----------|-------------|--------| | `CANCER_TYPE` | Broad cancer category | e.g., "Non-Small Cell Lung Cancer", "Breast Cancer" | | `CANCER_TYPE_DETAILED` | Specific subtype | e.g., "Lung Adenocarcinoma", "Invasive Ductal Carcinoma" | | `ONCOTREE_CODE` | OncoTree classification code | Standardized cancer type codes | ### Sample Information | Attribute | Description | Values | |-----------|-------------|--------| | `SAMPLE_TYPE` | Sample origin | Primary, Metastasis, Local Recurrence, Unknown | | `SAMPLE_CLASS` | Sample … (5439 more chars) ▶ clickhouse_run_select_query { "query": "SELECT countIf(n > 1) AS multi_sample_patients, countIf(has_p AND has_m) AS primary_and_met FROM (SELECT patient_unique_id, count() AS n, has(groupArray(attribute_value), 'Primary') AS has_p, has(groupArray(attribute_value), 'Metastasis') AS has_m FROM clinical_data_derived WHERE cancer_study_identifier = 'msk_chord_2024' AND attribute_name = 'SAMPLE_TYPE' GROUP BY patient_unique_id)" } ◀ result {"rows":[{"multi_sample_patients":90,"primary_and_met":26}]} ▶ clickhouse_list_table_columns { "table": "mutation_panel_gene_coverage" } ▶ clickhouse_list_table_columns { "table": "mutation_wes_coverage" } ◀ result {"columns":[{"name":"sample_unique_id","type":"String"},{"name":"cancer_study_identifier","type":"LowCardinality(String)"},{"name":"hugo_gene_symbol","type":"String"},{"name":"gene_panel_id","type":"LowCardinality(String)"}]} ◀ result {"columns":[{"name":"sample_unique_id","type":"String"},{"name":"cancer_study_identifier","type":"LowCardinality(String)"}]} ▶ clickhouse_run_select_query { "query": " WITH sample_type AS ( SELECT sample_unique_id, attribute_value AS sample_type FROM clinical_data_derived WHERE cancer_study_identifier = 'msk_chord_2024' AND attribute_name = 'SAMPLE_TYPE' AND attribute_value IN ('Primary', 'Metastasis') ), altered AS ( SELECT st.sample_type, ged.hugo_gene_symbol, COUNT(DISTINCT ged.sample_unique_id) AS altered_samples FROM genomic_event_derived ged JOIN sample_type st USING (sample_unique_id) WHERE ged.cancer_study_identifier = 'msk_chord_2024' AND ged.variant_type = 'mutation' AND ged.mutation_status != 'UNCALLED' AND ged.off_panel = 0 GROUP BY st.sample_type, ged.hugo_gene_symbol ), profiled_per_gene AS ( SELECT p.sample_unique_id, p.hugo_gene_symbol FROM mutation_panel_gene_coverage p WHERE p.cancer_study_identifier = 'msk_chord_2024' ), profiled AS ( SELECT st.sample_type, pg.hugo_gene_symbol, COUNT(DISTINCT pg.sample_unique_id) AS profiled_samples FROM profiled_per_gene pg JOIN sample_type st USING (sample_unique_id) GROUP BY st.sample_type, pg.hugo_gene_symbol ) SELECT a.hugo_gene_symbol, maxIf(a.altered_samples, a.sample_type='Primary') AS primary_altered, maxIf(p.profiled_samples, p.sample_type='Primary') AS primary_profiled, maxIf(a.altered_samples, a.sample_type='Metastasis') AS met_altered, maxIf(p.profiled_samples, p.sample_type='Metastasis') AS met_profiled FROM altered a JOIN profiled p ON a.sample_type = p.sample_type AND a.hugo_gene_symbol = p.hugo_gene_symbol GROUP BY a.hugo_gene_symbol HAVING primary_altered + met_altered > 500 ORDER BY (primary_altered + met_altered) DESC LIMIT 25 " } ◀ result {"rows":[{"hugo_gene_symbol":"TP53","primary_altered":8104,"primary_profiled":15928,"met_altered":4890,"met_profiled":8878},{"hugo_gene_symbol":"KRAS","primary_altered":4751,"primary_profiled":15928,"met_altered":2326,"met_profiled":8878},{"hugo_gene_symbol":"APC","primary_altered":3345,"primary_profiled":15928,"met_altered":1400,"met_profiled":8878},{"hugo_gene_symbol":"PIK3CA","primary_altered":2351,"primary_profiled":15928,"met_altered":1327,"met_profiled":8878},{"hugo_gene_symbol":"EGFR","primary_altered":1338,"primary_profiled":15928,"met_altered":779,"met_profiled":8878},{"hugo_gene_symbol":"ARID1A","primary_altered":1199,"primary_profiled":15928,"met_altered":634,"met_profiled":8878},{"hugo_gene_symbol":"SMAD4","primary_altered":1142,"primary_profiled":15928,"met_altered":643,"met_profiled":8878},{"hugo_gene_symbol":"KMT2D","primary_altered":1203,"primary_profiled":15928,"met_altered":569,"met_profiled":8878},{"hugo_gene_symbol":"KMT2C","primary_altered":1059,"primary_profiled":15928,"met_altered":552,"met_profiled":8878},{"hugo_gene_symbol":"ATM","primary_altered":890,"primary_profiled":15928,"met_altered":483,"met_profiled":8878},{"hugo_gene_symbol":"CDKN2A","primary_altered":846,"primary_profiled":15928,"met_altered":483,"met_profiled":8878},{"hugo_gene_symbol":"FAT1","primary_altered":818,"primary_profiled":15928,"met_altered":421,"met_profiled":8878},{"hugo_gene_symbol":"ZFHX3","primary_altered":828,"primary_profiled":15394,"met_altered":382,"met_profiled":8393},{"hugo_gene_symbol":"NF1","primary_altered":778,"primary_profiled":15928,"met_altered":424,"met_profiled":8878},{"hugo_gene_symbol":"PTEN","primary_altered":782,"primary_profiled":15928,"met_altered":407,"met_profiled":8878},{"hugo_gene_symbol":"STK11","primary_altered":755,"primary_profiled":15928,"met_altered":428,"met_profiled":8878},{"hugo_gene_symbol":"PTPRT","primary_altered":780,"primary_profiled":15928,"met_altered":401,"met_profiled":8878},{"hugo_gene_symbol":"KEAP1","primary_altered":702,"primary_profiled":15928,"met_altered":475,"met_profiled":8878},{"hugo_gene_symbol":"BRAF","primary_altered":826,"primary_profiled":15928,"met_altered":307,"met_profiled":8878},{"hugo_gene_symbol":"FBXW7","primary_altered":876,"primary_profiled":15928,"met_altered":238,"met_profiled":8878},{"hugo_gene_symbol":"GATA3","primary_altered":634,"primary_profiled":15928,"met_altered":458,"met_profiled":8878},{"hugo_gene_symbol":"SMARCA4","primary_altered":692,"primary_profiled":15928,"met_altered":380,"met_profiled":8878},{"hugo_gene_symbol":"PTPRD","primary_altered":687,"primary_profiled":15928,"met_altered":381,"met_profiled":8878},{"hugo_gene_symbol":"RBM10","primary_altered":732,"primary_profiled":15928,"met_altered":323,"met_profiled":8878},{"hugo_gene_symbol":"CDH1","primary_altered":554,"primary_profiled":15928,"met_altered":406,"met_profiled":8878}]} ▶ clickhouse_run_select_query { "query": " WITH sample_type AS ( SELECT sample_unique_id, attribute_value AS sample_type FROM clinical_data_derived WHERE cancer_study_identifier = 'msk_chord_2024' AND attribute_name = 'SAMPLE_TYPE' AND attribute_value IN ('Primary', 'Metastasis') ), tmb AS ( SELECT sample_unique_id, toFloat64OrNull(attribute_value) AS tmb FROM clinical_data_derived WHERE cancer_study_identifier = 'msk_chord_2024' AND attribute_name = 'TMB_NONSYNONYMOUS' ) SELECT st.sample_type, count() AS n_samples, round(avg(t.tmb), 2) AS mean_tmb, round(quantile(0.5)(t.tmb), 2) AS median_tmb FROM sample_type st JOIN tmb t USING (sample_unique_id) WHERE t.tmb IS NOT NULL GROUP BY st.sample_type " } ◀ result {"rows":[{"sample_type":"Primary","n_samples":15928,"mean_tmb":7.77,"median_tmb":4.1},{"sample_type":"Metastasis","n_samples":8878,"mean_tmb":6.35,"median_tmb":4.32}]} ▶ clickhouse_run_select_query { "query": " WITH sample_type AS ( SELECT sample_unique_id, attribute_value AS sample_type FROM clinical_data_derived WHERE cancer_study_identifier = 'msk_chord_2024' AND attribute_name = 'SAMPLE_TYPE' AND attribute_value IN ('Primary', 'Metastasis') ), altered AS ( SELECT st.sample_type, ged.hugo_gene_symbol, COUNT(DISTINCT ged.sample_unique_id) AS altered_samples FROM genomic_event_derived ged JOIN sample_type st USING (sample_unique_id) WHERE ged.cancer_study_identifier = 'msk_chord_2024' AND ged.variant_type = 'mutation' AND ged.mutation_status != 'UNCALLED' AND ged.off_panel = 0 AND ged.hugo_gene_symbol IN ('TP53','KRAS','APC','PIK3CA','EGFR','BRAF','FBXW7','ARID1A','STK11','KEAP1','SMAD4','CDKN2A','ATM','GATA3','NF1','PTEN') GROUP BY st.sample_type, ged.hugo_gene_symbol ), profiled_per_gene AS ( SELECT p.sample_unique_id, p.hugo_gene_symbol FROM mutation_panel_gene_coverage p WHERE p.cancer_study_identifier = 'msk_chord_2024' AND p.hugo_gene_symbol IN ('TP53','KRAS','APC','PIK3CA','EGFR','BRAF','FBXW7','ARID1A','STK11','KEAP1','SMAD4','CDKN2A','ATM','GATA3','NF1','PTEN') ), profiled AS ( SELECT st.sample_type, pg.hugo_gene_symbol, COUNT(DISTINCT pg.sample_unique_id) AS profiled_samples FROM profiled_per_gene pg JOIN sample_type st USING (sample_unique_id) GROUP BY st.sample_type, pg.hugo_gene_symbol ) SELECT a.hugo_gene_symbol, maxIf(a.altered_samples, a.sample_type='Primary') AS primary_altered, maxIf(p.profiled_samples, p.sample_type='Primary') AS primary_profiled, round(maxIf(a.altered_samples, a.sample_type='Primary')*100.0/maxIf(p.profiled_samples, p.sample_type='Primary'),1) AS primary_pct, maxIf(a.altered_samples, a.sample_type='Metastasis') AS met_altered, maxIf(p.profiled_samples, p.sample_type='Metastasis') AS met_profiled, round(maxIf(a.altered_samples, a.sample_type='Metastasis')*100.0/maxIf(p.profiled_samples, p.sample_type='Metastasis'),1) AS met_pct FROM altered a JOIN profiled p ON a.sample_type = p.sample_type AND a.hugo_gene_symbol = p.hugo_gene_symbol GROUP BY a.hugo_gene_symbol ORDER BY primary_pct DESC " } ◀ result {"rows":[{"hugo_gene_symbol":"TP53","primary_altered":8104,"primary_profiled":15928,"primary_pct":50.9,"met_altered":4890,"met_profiled":8878,"met_pct":55.1},{"hugo_gene_symbol":"KRAS","primary_altered":4751,"primary_profiled":15928,"primary_pct":29.8,"met_altered":2326,"met_profiled":8878,"met_pct":26.2},{"hugo_gene_symbol":"APC","primary_altered":3345,"primary_profiled":15928,"primary_pct":21,"met_altered":1400,"met_profiled":8878,"met_pct":15.8},{"hugo_gene_symbol":"PIK3CA","primary_altered":2351,"primary_profiled":15928,"primary_pct":14.8,"met_altered":1327,"met_profiled":8878,"met_pct":14.9},{"hugo_gene_symbol":"EGFR","primary_altered":1338,"primary_profiled":15928,"primary_pct":8.4,"met_altered":779,"met_profiled":8878,"met_pct":8.8},{"hugo_gene_symbol":"ARID1A","primary_altered":1199,"primary_profiled":15928,"primary_pct":7.5,"met_altered":634,"met_profiled":8878,"met_pct":7.1},{"hugo_gene_symbol":"SMAD4","primary_altered":1142,"primary_profiled":15928,"primary_pct":7.2,"met_altered":643,"met_profiled":8878,"met_pct":7.2},{"hugo_gene_symbol":"ATM","primary_altered":890,"primary_profiled":15928,"primary_pct":5.6,"met_altered":483,"met_profiled":8878,"met_pct":5.4},{"hugo_gene_symbol":"FBXW7","primary_altered":876,"primary_profiled":15928,"primary_pct":5.5,"met_altered":238,"met_profiled":8878,"met_pct":2.7},{"hugo_gene_symbol":"CDKN2A","primary_altered":846,"primary_profiled":15928,"primary_pct":5.3,"met_altered":483,"met_profiled":8878,"met_pct":5.4},{"hugo_gene_symbol":"BRAF","primary_altered":826,"primary_profiled":15928,"primary_pct":5.2,"met_altered":307,"met_profiled":8878,"met_pct":3.5},{"hugo_gene_symbol":"NF1","primary_altered":778,"primary_profiled":15928,"primary_pct":4.9,"met_altered":424,"met_profiled":8878,"met_pct":4.8},{"hugo_gene_symbol":"PTEN","primary_altered":782,"primary_profiled":15928,"primary_pct":4.9,"met_altered":407,"met_profiled":8878,"met_pct":4.6},{"hugo_gene_symbol":"STK11","primary_altered":755,"primary_profiled":15928,"primary_pct":4.7,"met_altered":428,"met_profiled":8878,"met_pct":4.8},{"hugo_gene_symbol":"KEAP1","primary_altered":702,"primary_profiled":15928,"primary_pct":4.4,"met_altered":475,"met_profiled":8878,"met_pct":5.4},{"hugo_gene_symbol":"GATA3","primary_altered":634,"primary_profiled":15928,"primary_pct":4,"met_altered":458,"met_profiled":8878,"met_pct":5.2}]} ▶ resolve_and_route { "studyIds": [ "msk_chord_2024" ] } ◀ result {"success":true,"message":"Found 1 study. Use the metadata to call the appropriate navigation tool(s).","data":{"totalCount":1,"studiesWithMetadata":[{"studyId":"msk_chord_2024","name":"MSK-CHORD (MSK, Nature 2024)","sampleCount":25040,"studyViewUrl":"https://www.cbioportal.org/study?id=msk_chord_2024","metadata":{"clinicalAttributeIds":["ADRENAL_GLANDS","BONE","CANCER_TYPE","CANCER_TYPE_DETAILED","CLINICAL_GROUP","CLINICAL_SUMMARY","CNS_BRAIN","CURRENT_AGE_DEID","DIAGNOSIS_DESCRIPTION","ETHNICITY","FRACTION_GENOME_ALTERED","GENDER","GENE_PANEL","GLEASON_FIRST_REPORTED","GLEASON_HIGHEST_REPORTED","GLEASON_SAMPLE_LEVEL","HER2","HISTORY_OF_PDL1","HR","ICD_O_HISTOLOGY_DESCRIPTION","INTRA_ABDOMINAL","LIVER","LUNG","LYMPH_NODES","METASTATIC_SITE","MSI_COMMENT","MSI_SCORE","MSI_TYPE","MUTATION_COUNT","NUM_ICDO_DX","ONCOTREE_CODE","OS_MONTHS","OS_STATUS","OTHER","PATHOLOGICAL_GROUP","PDL1_POSITIVE","PLEURA","PRIMARY_SITE","PRIOR_MED_TO_MSK","RACE","REPRODUCTIVE_ORGANS","SAMPLE_CLASS","SAMPLE_COUNT","SAMPLE_COVERAGE","SAMPLE_TYPE","SMOKING_PREDICTIONS_3_CLASSES","SOMATIC_STATUS","STAGE_HIGHEST_RECORDED","TMB_NONSYNONYMOUS","TUMOR_PURITY"],"molecularProfileIds":["msk_chord_2024_cna","msk_chord_2024_mutations","msk_chord_2024_structural_variants"],"availableComparisonTabs":["overlap","clinical","survival","alterations","mutations"]}}]}} ▶ get_studyviewfilter_options { "studyId": "msk_chord_2024", "attributeIds": [ "SAMPLE_TYPE" ] } ◀ result {"success":true,"message":"Successfully retrieved StudyView filter options","data":{"studyId":"msk_chord_2024","attributes":[{"attributeId":"SAMPLE_TYPE","displayName":"Sample Type","description":"The type assigned to a sample (i.e., normal, primary, met, recurrence). Source: MPath","datatype":"STRING","values":["Local Recurrence","Unknown","Primary","Metastasis"]}]}} ▶ navigate_to_group_comparison { "studyIds": [ "msk_chord_2024" ], "groups": [ { "name": "Primary", "studyViewFilter": { "clinicalDataFilters": [ { "attributeId": "SAMPLE_TYPE", "values": [ { "value": "Primary" } ] } ] } }, { "name": "Metastasis", "studyViewFilter": { "clinicalDataFilters": [ { "attributeId": "SAMPLE_TYPE", "values": [ { "value": "Metastasis" } ] } ] } } ], "tab": "alterations" } ◀ result {"success":true,"message":"Navigating to https://www.cbioportal.org/comparison/alterations?comparisonId=6ab49e90c2115c492d884a5c","url":"https://www.cbioportal.org/comparison/alterations?comparisonId=6ab49e90c2115c492d884a5c","data":{"description":"Group comparison (2 custom groups)","studies":["msk_chord_2024"],"totalGroups":2,"groups":[{"name":"Primary","sampleCount":15928},{"name":"Metastasis","sampleCount":8878}],"studyViewUrl":"https://www.cbioportal.org/study?id=msk_chord_2024","groupUrls":[{"groupName":"Primary","url":"https://www.cbioportal.org/study?id=msk_chord_2024#filterJson=%7B%22clinicalDataFilters%22%3A%5B%7B%22attributeId%22%3A%22SAMPLE_TYPE%22%2C%22values%22%3A%5B%7B%22value%22%3A%22Primary%22%7D%5D%7D%5D%7D"},{"groupName":"Metastasis","url":"https://www.cbioportal.org/study?id=msk_chord_2024#filterJson=%7B%22clinicalDataFilters%22%3A%5B%7B%22attributeId%22%3A%22SAMPLE_TYPE%22%2C%22values%22%3A%5B%7B%22value%22%3A%22Metastasis%22%7D%5D%7D%5D%7D"}],"pageDescription":"Comparison of genomic alteration rates across Primary (15928) vs Metastasis (8878)."}} ▶ navigate_to_study_view { "studyIds": [ "msk_chord_2024" ], "filterJson": { "clinicalDataFilters": [ { "attributeId": "SAMPLE_TYPE", "values": [ { "value": "Primary" }, { "value": "Metastasis" } ] } ] } } ◀ result {"success":true,"message":"Navigating to https://www.cbioportal.org/study/summary?id=msk_chord_2024#filterJson=%7B%22clinicalDataFilters%22%3A%5B%7B%22attributeId%22%3A%22SAMPLE_TYPE%22%2C%22values%22%3A%5B%7B%22value%22%3A%22Primary%22%7D%2C%7B%22value%22%3A%22Metastasis%22%7D%5D%7D%5D%7D","url":"https://www.cbioportal.org/study/summary?id=msk_chord_2024#filterJson=%7B%22clinicalDataFilters%22%3A%5B%7B%22attributeId%22%3A%22SAMPLE_TYPE%22%2C%22values%22%3A%5B%7B%22value%22%3A%22Primary%22%7D%2C%7B%22value%22%3A%22Metastasis%22%7D%5D%7D%5D%7D","data":{"studyIds":["msk_chord_2024"],"studies":[{"studyId":"msk_chord_2024","name":"MSK-CHORD (MSK, Nature 2024)","sampleCount":25040}],"tab":"summary","hasFilters":true,"hasPlotsConfig":false,"pageDescription":"Overview of all study data displayed as interactive charts summarizing available clinical and genomic data. Click within a chart to filter. Additional charts may be available in the Charts menu."}} ═ answer (success) ## Important caveat before the numbers MSK-CHORD contains 24,950 patients but only 25,040 samples — nearly one sample per patient. Of the 90 patients who have more than one sample, only **26** have both a Primary and a Metastasis sample. There is no meaningful same-patient (longitudinally paired) primary-vs-metastasis cohort in this study — 26 pairs is too small, and it's not a representative or randomly-selected subset (these are patients who happened to be re-biopsied). What follows instead is an **unpaired** comparison: all 15,928 Primary samples vs. all 8,878 Metastasis samples (largely different patients, pooled across all five cancer types in MSK-CHORD: NSCLC, colorectal, breast, prostate, pancreatic). Because cancer-type mix differs between the two sample-type groups, some of the differences below partly reflect which cancer types get biopsied at metastasis (e.g., more prostate/pancreatic met biopsies) rather than a pure primary-vs-met biology effect. ### Top mutated genes, Primary vs. Metastasis samples (% of samples profiled for each gene) | Gene | Primary (n=15,928 profiled) | Metastasis (n=8,878 profiled) | |---|---|---| | TP53 | 50.9% (8,104) | 55.1% (4,890) | | KRAS | 29.8% (4,751) | 26.2% (2,326) | | APC | 21.0% (3,345) | 15.8% (1,400) | | PIK3CA | 14.8% (2,351) | 14.9% (1,327) | | EGFR | 8.4% (1,338) | 8.8% (779) | | ARID1A | 7.5% (1,199) | 7.1% (634) | | SMAD4 | 7.2% (1,142) | 7.2% (643) | | FBXW7 | 5.5% (876) | 2.7% (238) | | ATM | 5.6% (890) | 5.4% (483) | | CDKN2A | 5.3% (846) | 5.4% (483) | | BRAF | 5.2% (826) | 3.5% (307) | | NF1 | 4.9% (778) | 4.8% (424) | | PTEN | 4.9% (782) | 4.6% (407) | | STK11 | 4.7% (755) | 4.8% (428) | | KEAP1 | 4.4% (702) | 5.4% (475) | | GATA3 | 4.0% (634) | 5.2% (458) | **Notable patterns:** TP53 is slightly more frequent in metastatic samples; APC, FBXW7, and BRAF are notably less frequent in metastases (all three are enriched in colorectal cancer, which may reflect fewer colorectal met biopsies vs. primaries). KEAP1 and GATA3 are modestly higher in metastases. These are descriptive percentage-point differences only — no statistical test (e.g., Fisher's exact) has been run to establish significance; use the Alterations comparison tab linked below for that. **Tumor mutational burden (TMB, mutations/Mb):** Primary mean 7.77 (median 4.10, n=15,928); Metastasis mean 6.35 (median 4.32, n=8,878). Means are pulled up by hypermutated outliers in both groups; medians are similar. ### Explore further [Alteration frequency comparison: Primary vs. Metastasis](https://www.cbioportal.org/comparison/alterations?comparisonId=6ab49e90c2115c492d884a5c) — Comparison of genomic alteration rates across Primary (15,928 samples) vs Metastasis (8,878 samples). This tab can compute proper enrichment statistics per gene. [Study cohort filtered to Primary + Metastasis samples](https://www.cbioportal.org/study/summary?id=msk_chord_2024#filterJson=%7B%22clinicalDataFilters%22%3A%5B%7B%22attributeId%22%3A%22SAMPLE_TYPE%22%2C%22values%22%3A%5B%7B%22value%22%3A%22Primary%22%7D%2C%7B%22value%22%3A%22Metastasis%22%7D%5D%7D%5D%7D) — Overview of all study data displayed as interactive charts. If you'd like, I can re-run this restricted to a single cancer type (e.g., only NSCLC or only colorectal Primary vs. Metastasis) to remove the cancer-type-composition confound, or restrict to a specific gene of interest.