Q (sonnet): Can you think of any flaws in the methodology used in the MSK-CHORD paper (Jee et al., Nature 2024)? ═ answer (success) I should note upfront: this question asks for scholarly critique of a published paper's methodology, not a data query or navigation task. I'll draw on what I know about MSK-CHORD (Jee et al., *Nature* 2024, "Automated real-world data integration improves cancer outcome prediction") and can supplement with what's actually in the cBioPortal database if that's useful, but I want to be upfront that a literature critique is somewhat outside the direct query/navigate scope of this tool — I'll answer based on general knowledge of the study design rather than pulling database evidence, since this isn't really a database question. With that caveat, here are some methodological considerations worth flagging about MSK-CHORD: **1. Single-institution referral bias** MSK-CHORD is derived entirely from patients treated at Memorial Sloan Kettering, a tertiary academic cancer center. Patients referred there skew toward higher socioeconomic status, later-line/refractory disease, unusual or complex cases, and greater access to clinical trials and MSK-IMPACT sequencing. Findings on treatment patterns and outcomes may not generalize to community oncology populations, and this limits external validity for real-world outcome inference. **2. NLP/automated curation vs. manual chart review — misclassification risk** The core methodological contribution is automated extraction of outcomes (e.g., treatment response, progression, death) from unstructured EHR text using NLP pipelines, replacing labor-intensive manual abstraction. While validated against manually curated subsets, any NLP-derived label carries residual misclassification error. Errors are unlikely to be purely random — they may correlate with note-writing style by service, tumor type, or era, potentially introducing differential misclassification that biases downstream association estimates in non-obvious directions. **3. Immortal time bias / left truncation in genomic-linked cohorts** Because MSK-IMPACT sequencing often occurs after a variable delay from diagnosis (sometimes only in patients who survive long enough to be sequenced, or after progression on first-line therapy), analyses linking mutation status to survival risk conditioning on a landmark that isn't uniformly defined relative to diagnosis — a classic setup for immortal time bias unless carefully handled with time-dependent covariates or landmark analysis. **4. Survivorship/sequencing-selection bias** Only patients who received MSK-IMPACT panel sequencing are included. Sequencing was not randomly ordered — it's more likely for patients with sufficient tissue, longer survival to accrue a sample, enrollment in precision oncology programs, or specific clinical suspicion. This creates a cohort that is not representative of the full disease population, potentially inflating apparent associations between certain mutations and outcomes. **5. Retrospective observational design — confounding by indication** Any treatment-outcome association is subject to confounding by indication (sicker patients get certain regimens), and while the paper uses this cohort partly to demonstrate feasibility of the pipeline, any clinical inference drawn from it (e.g., "drug X associated with better OS") cannot support causal claims without adjustment for known and unknown confounders — the paper is best read as a resource/methods paper, and readers sometimes over-extend its associative findings toward causal or clinically prescriptive interpretations. **6. Heterogeneous follow-up and informative censoring** Real-world follow-up is irregular — patients seen less frequently (or who transfer care elsewhere) generate informative censoring that can bias survival curves, unlike the fixed-schedule visits in clinical trials. **7. Validation cohort size and generalizability of NLP models across cancer types** NLP models trained/validated on the specific cancer types included (the paper focuses on a subset of common cancers) may not perform equivalently across all histologies, especially rarer tumor types with different note structures, vocabularies, or documentation practices — model performance metrics reported for the validation set may not hold uniformly. **8. Lack of race/ethnicity and social determinant granularity** Like many EHR-derived real-world datasets, structured demographic fields (race, ethnicity, insurance status) are often incomplete or inconsistently coded, limiting equity-focused analyses of the outcome predictions the pipeline generates. If you'd like, I can pull actual cohort composition, sample counts, cancer type distribution, or sequencing timing data for the MSK-CHORD dataset from cBioPortal (it is deposited there as `msk_chord_2024`) to ground any of these points empirically — for example, checking time from diagnosis to sequencing, or demographic completeness. Just let me know which angle you'd like quantified.