Q (sonnet): Can you think of any flaws in the methodology used in the MSK-CHORD paper (Jee et al., Nature 2024)? ▶ list_studies { "search": "MSK CHORD" } ◀ result {"result":[{"cancer_study_identifier":"msk_chord_2024","name":"MSK-CHORD (MSK, Nature 2024)","type_of_cancer_id":"mixed","sample_count":25040,"has_guide":true,"url":"https://www.cbioportal.org/study/summary?id=msk_chord_2024"}]} ▶ get_study_guide { "study_id": "msk_chord_2024" } ◀ result # MSK-CHORD (MSK, Nature 2024) **Study ID:** `msk_chord_2024` ## Overview Targeted sequencing via MSK-IMPACT panels. Clinical annotations include some derived from natural language processing (denoted NLP). **Exactly five cancer types** (`CANCER_TYPE`, patients): Non-Small Cell Lung Cancer 7,809, Colorectal Cancer 5,543, Breast Cancer 5,368, Prostate Cancer 3,211, Pancreatic Cancer 3,109. There is **no melanoma** or any other cancer type; say so up front if asked, instead of substituting another type. **No therapy-response variable.** There is no RECIST, objective response, or best-response attribute or event. For treatment-outcome questions (e.g. immunotherapy response), say this first; the only proxies are `OS_MONTHS`/`OS_STATUS`, or NLP radiology progression events (`Diagnosis` events with `SUBTYPE = 'Progression'`, key `PROGRESSION` = Y/N/Indeterminate), in patients with `Treatment` events of the relevant `SUBTYPE` (e.g. `Immuno`: 3,341 patients). Hand off the comparison to cBioPortal group comparison / survival. **Nearly one sample per patient: 24,950 patients / 25,040 samples.** Only 90 patients have more than one sample, and all 90 have samples from two different cancer types (second primaries); only 26 have both a `Primary` and a `Metastasis` sample. There is no meaningful same-patient (paired) primary-vs-metastasis cohort. For "same patient" / paired questions, say this up front, then offer the **unpaired** comparison of all `Primary` vs `Metastasis` samples (`SAMPLE_TYPE`), labelled as unpaired. ```sql SELECT countIf(n > 1) AS multi_sample_patients, -- 90 countIf(has_p AND has_m) AS primary_and_met -- 26 FROM (SELECT patient_unique_id, count() AS n, has(groupArray(attribute_value), 'Primary') AS has_p, has(groupArray(attribute_value), 'Metastasis') AS has_m FROM clinical_data_derived WHERE cancer_study_identifier = 'msk_chord_2024' AND attribute_name = 'SAMPLE_TYPE' GROUP BY patient_unique_id); ``` ## Gene Panels This study uses multiple MSK-IMPACT panel versions: - **IMPACT341**: Earlier version, 341 genes - **IMPACT410**: 410 genes - **IMPACT468**: 468 genes - **IMPACT505**: Latest version, 505 genes **Important:** Different samples may have different gene coverage. Always use gene-specific denominators when calculating mutation frequencies. ## Clinical Attributes - Semantic Guide ### Cancer Classification | Attribute | Description | Values | |-----------|-------------|--------| | `CANCER_TYPE` | Broad cancer category | e.g., "Non-Small Cell Lung Cancer", "Breast Cancer" | | `CANCER_TYPE_DETAILED` | Specific subtype | e.g., "Lung Adenocarcinoma", "Invasive Ductal Carcinoma" | | `ONCOTREE_CODE` | OncoTree classification code | Standardized cancer type codes | ### Sample Information | Attribute | Description | Values | |-----------|-------------|--------| | `SAMPLE_TYPE` | Sample origin | Primary, Metastasis, Local Recurrence, Unknown | | `SAMPLE_CLASS` | Sample … (5439 more chars) ═ answer (success) This is a general methodological question rather than a data query, so I'll answer it directly. To be clear about sourcing: the study-specific facts below (cohort structure, panel versions, NLP-derived fields, treatment-data limitations) come from the MSK-CHORD dataset as represented in cBioPortal; the broader methodological critique is general knowledge about retrospective real-world genomic cohort design, not a finding from any database query or literature review I performed. ## Potential methodological limitations in MSK-CHORD (Jee et al., *Nature* 2024) **1. Single-institution, tertiary-referral cohort → selection bias.** All patients were seen at MSK, a major academic cancer center. Patients referred to MSK are systematically enriched for advanced/refractory disease, rare or complex presentations, clinical-trial eligibility, and higher socioeconomic/insurance access. Prevalence estimates for mutations, biomarkers (e.g., PD-L1, MSI), and outcomes may not generalize to community oncology populations. **2. Sequencing-panel eligibility bias, not a population-based sample.** Patients were included because they underwent MSK-IMPACT testing, which itself is more often ordered for advanced, metastatic, or treatment-refractory tumors where genomic-guided therapy is being considered. Early-stage, low-risk, or patients who died/transferred care before testing was ordered are underrepresented — this is a form of immortal-time and indication bias that affects both mutation frequency estimates and survival analyses. **3. Heterogeneous, evolving gene panels (IMPACT341 → IMPACT505).** Because the panel expanded over the study period, gene coverage differs by sample/testing date. Any mutation-frequency comparison across time or across cohorts must use panel-specific denominators; pooling frequencies without accounting for panel version can bias apparent trends over calendar time (a version of testing-artifact confounding with real biological trends). **4. Reliance on NLP-derived clinical annotations.** A meaningful share of clinical variables (treatment start/stop, radiologic progression calls, some clinical/pathological groupings) were extracted via natural-language processing of unstructured notes rather than structured, manually curated fields. NLP extraction carries a nontrivial error rate, and confidence is only partially quantified (`INFERRED_TX_PROB`). This affects treatment-timing and "progression" endpoints in particular. **5. No formal response/RECIST endpoint.** The resource lacks a curated objective-response or best-overall-response variable. Treatment-outcome inferences rely on overall survival or on NLP-flagged "progression" events from radiology reports — a much noisier and less validated endpoint than RECIST-based response, especially for retrospective comparisons of drug efficacy. **6. Retrospective design with immortal time / guarantee-time bias risk.** Because sequencing, treatment, and outcome data are all assembled after the fact, any analysis linking "having a mutation" or "receiving a drug" to survival is vulnerable to immortal-time bias unless carefully time-anchored (patients must have survived long enough to be sequenced/treated). This is a generic pitfall of real-world genomic-outcomes cohorts, not something I've confirmed the authors mishandled — but it is a standard critique to raise when reviewing this class of paper. **7. Left/right censoring and incomplete treatment capture.** Prior therapies given outside MSK are captured separately (`PRIOR_MED_TO_MSK`) and may be less complete/reliable than in-house treatment records; patients who received care elsewhere before or after their MSK episode can have incomplete longitudinal treatment and outcome histories. **8. Near-absence of a paired primary–metastasis sub-cohort.** Only a small number of patients contributed both a primary and a metastatic sample (on the order of two dozen), which is far too few for a well-powered *paired* primary-vs-metastasis genomic comparison — any such comparisons in the paper are necessarily unpaired, comparing different patients' primary vs. metastatic samples rather than tracking clonal evolution within the same patient. **9. Five cancer types only, unevenly sized.** The cohort covers exactly five tumor types (NSCLC, colorectal, breast, prostate, pancreatic) with sample sizes ranging roughly 3,100–7,800 patients — conclusions about "pan-cancer" trends should be read as limited to these five relatively common solid tumors, not representative of rarer cancers. **10. Data-use and reproducibility constraints.** The resource is released under a restrictive non-commercial, no-derivatives license, and full external validation/reproduction by outside groups is correspondingly limited compared to fully open datasets. --- If you'd like, I can pull specific numbers from the MSK-CHORD data in cBioPortal to illustrate any of these points empirically — for example, the panel-version distribution over time, the primary-vs-metastasis sample imbalance, or the fraction of treatment records that are NLP-inferred vs. structured. I can also open the [MSK-CHORD Study View](https://www.cbioportal.org/study/summary?id=msk_chord_2024) so you can explore the cohort composition directly. Just let me know which angle you'd like quantified.