Q (sonnet-5.5): Can you think of any flaws in the methodology used in the MSK-CHORD paper (Jee et al., Nature 2024)? ═ answer (success) I haven't re-read the paper for this answer. What follows comes from my recollection of the study design, not from cBioPortal data, so check each point against the Methods and Supplement. In brief, MSK-CHORD links MSK-IMPACT sequencing with NLP-derived annotations from notes and radiology reports for about 25,000 patients with lung, breast, colorectal, prostate, or pancreatic cancer. It then builds survival models on that combined data. **Cohort and selection** - **Single institution, sequenced patients only.** Patients who get MSK-IMPACT are not a random sample. They are often those with enough disease burden, fitness, or survival time to be tested. This limits how far the results generalize, and it shapes the outcome distributions. - **Left truncation and immortal time bias.** Sequencing happens some time after diagnosis. If survival is counted from diagnosis, patients who died early can't enter the cohort. The authors handled this, as I recall, but the choice of time origin and the correction are worth checking. - **Limited ancestry and referral diversity.** A tertiary cancer center draws a particular referral population. **NLP-derived labels** - **Label noise.** Progression, metastatic sites, and treatment context are inferred from free text and radiology reports. Errors in these labels can propagate into the models. It matters how well the models did against manual annotation, and whether that varied by cancer type, note type, or era. - **Report-language dependence.** Progression inferred from report phrasing is not the same as RECIST or clinical progression. Templates and wording change over time. - **Care outside MSK.** Imaging, treatments, or events at other institutions may be missing. That can look like non-progression or censoring when it isn't. - **Transportability.** The NLP models were trained on MSK's documentation style. They may not work at other centers without retraining. **Outcomes and modeling** - **Confounding by indication.** Treatment and metastatic-site features are entangled with prognosis and treatment choice. Associations shouldn't be read causally. - **Time-varying features.** It has to be clear which features were known at the prediction time point. Any information from after that point would leak into the model. - **Validation.** Cross-validation or hold-out on the same institution's data overstates performance compared with external or prospective validation. Check whether temporal validation was done, and how the comparison against staging or other baselines was set up. - **Follow-up and censoring.** If death ascertainment is incomplete, or patients are lost to follow-up unevenly, the survival estimates are biased. **Genomic findings** - **Multiple testing and subgroup size.** Many gene and cancer-type combinations were tested. Findings like the SETD2 association in lung adenocarcinoma need independent replication, and effects in small subgroups are fragile. - **Heterogeneous genomic data.** Panel versions changed over time, sample type varies (primary vs. metastasis), and the sample may have been taken before or after treatment. All of these affect what a mutation call means for prognosis. - **Only five cancer types.** Conclusions about "cancer outcomes" in general go beyond what was tested. **Overall.** The main concerns are selection bias, NLP label quality, and internal-only validation. The paper's own limitations and supplementary validation sections are the best place to see how much of this the authors addressed. I can also look at how the MSK-CHORD cohort is represented in cBioPortal, such as sample counts, cancer types, and available clinical fields, if that would help you test any of these concerns.