Q (sonnet): generate a oncoprint of kras, stk11, keap1, tp53 alterations stratified by smoking status in lung adenocarcinoma from the mskcc clincogenomic sequencing cohort. Please only select oncogenic somatic alterations and remove samples with unknown smoking status ▶ resolve_and_route { "studyKeywords": [ "MSK", "clinical sequencing cohort", "lung" ] } ◀ result {"success":false,"error":"No matching studies found","details":{"searchTerms":["MSK","clinical sequencing cohort","lung"]}} ▶ list_studies { "search": "MSK-IMPACT lung" } ▶ list_studies { "search": "clinical sequencing cohort lung" } ◀ result {"result":[{"cancer_study_identifier":"luad_mskcc_2020","name":"Lung Adenocarcinoma (MSK, J Thorac Oncol 2020)","type_of_cancer_id":"luad","sample_count":604,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=luad_mskcc_2020"},{"cancer_study_identifier":"bm_nsclc_mskcc_2023","name":"Non-Small Cell Lung Cancer Brain Metastasis (MSK, Nat Commun 2023)","type_of_cancer_id":"nsclc","sample_count":322,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=bm_nsclc_mskcc_2023"},{"cancer_study_identifier":"lung_msk_pdx","name":"Thoracic PDX (MSK, Provisional)","type_of_cancer_id":"lung","sample_count":318,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=lung_msk_pdx"},{"cancer_study_identifier":"nsclc_pd1_msk_2018","name":"Non-Small Cell Lung Cancer (MSK, J Clin Oncol 2018)","type_of_cancer_id":"nsclc","sample_count":240,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=nsclc_pd1_msk_2018"},{"cancer_study_identifier":"luad_mskimpact_2021","name":"Lung Adenocarcinoma (MSK, 2021)","type_of_cancer_id":"luad","sample_count":186,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=luad_mskimpact_2021"},{"cancer_study_identifier":"alk_msk_2026","name":"Non-Small Cell Lung Cancer (MSK, 2026)","type_of_cancer_id":"nsclc","sample_count":90,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=alk_msk_2026"},{"cancer_study_identifier":"asclc_msk_2024","name":"Atypical Small Cell Lung Cancer (MSK, Cancer Discov 2024)","type_of_cancer_id":"lung","sample_count":31,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=asclc_msk_2024"}]} ◀ result {"result":[]} ▶ list_studies { "search": "MSK-IMPACT Clinical Sequencing Cohort" } ▶ list_studies { "search": "MSK 2021 pan cancer" } ◀ result {"result":[{"cancer_study_identifier":"msk_impact_2017","name":"MSK-IMPACT Clinical Sequencing Cohort (MSK, Nat Med 2017)","type_of_cancer_id":"mixed","sample_count":10945,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=msk_impact_2017"},{"cancer_study_identifier":"msk_impact_50k_2026","name":"MSK-IMPACT 50K Clinical Sequencing Cohort (MSK, Cancer Cell 2026)","type_of_cancer_id":"mixed","sample_count":54331,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=msk_impact_50k_2026"},{"cancer_study_identifier":"pancan_mimsi_msk_2024","name":"Mixed Tumors - MiMSI Cohort (MSK, Nat Commun 2024)","type_of_cancer_id":"mixed","sample_count":5033,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=pancan_mimsi_msk_2024"},{"cancer_study_identifier":"brca_msk_erbb2_2026","name":"HER2-Low and HER2-Null Metastatic Breast Cancers treated with T-DXd (MSK, Clinical Cancer Research 2026)","type_of_cancer_id":"breast","sample_count":272,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=brca_msk_erbb2_2026"},{"cancer_study_identifier":"braf_msk_archer_2024","name":"BRAF Fusions - ARCHER Clinical Sequencing Cohort (MSK, Clin Cancer Res 2024)","type_of_cancer_id":"mixed","sample_count":107,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=braf_msk_archer_2024"},{"cancer_study_identifier":"braf_msk_impact_2024","name":"BRAF Fusions - IMPACT Clinical Sequencing Cohort (MSK, Clin Cancer Res 2024)","type_of_cancer_id":"mixed","sample_count":105,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=braf_msk_impact_2024"}]} ◀ result {"result":[{"cancer_study_identifier":"msk_met_2021","name":"MSK MetTropism (MSK, Cell 2021)","type_of_cancer_id":"mixed","sample_count":25775,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=msk_met_2021"}]} ▶ read_guide { "uri": "cbioportal://mutation-frequency-guide" } ◀ result # Mutation Frequency Analysis Guide ## IMPORTANT: Reporting Mutation Frequencies - **ALWAYS report frequencies as percentages**, not raw counts: `frequency = (altered_samples / total_profiled_samples) × 100` - For quick frequency lookups, **prefer the TCGA Pan-Cancer Atlas study first**, then offer to expand to other studies - When reporting across multiple studies, show **ranges** (e.g., "TP53 is mutated in 30–60% of samples") rather than a single average - **NEVER** sum mutation events across studies to compute an aggregate frequency — this can exceed 100% due to double-counting - Warn users that samples may overlap across cohorts (e.g., MSK studies may share patients) - **Choose and state the counting unit**: use patient-level frequencies for prevalence/rate questions unless the user explicitly asks for samples; use sample-level frequencies when the user asks about samples. - **For "across cancer types" questions**, jump to the [Cross-Cancer-Type Mutation Frequency](#cross-cancer-type-mutation-frequency) section below — there is one correct recipe and several common wrong ones. ## Counting Unit: Samples vs Patients Before answering any mutation count or frequency question, decide whether the unit is samples or patients and state that choice in the answer. | User wording | Counting unit | |--------------|---------------| | "prevalence", "rate", "fraction of patients", "patients with", "how common is" | Patient-level: `COUNT(DISTINCT patient_unique_id)` | | "samples", "specimens", "biopsies", sample-level cohort composition | Sample-level: `COUNT(DISTINCT sample_unique_id)` | | Ambiguous | Ask, or default to patient-level for prevalence/rate language and say so | ### Cross-study sample-count caveat When an answer touches more than one study and reports a sample count, prepend a one-line caveat: > Sample IDs are unique within cBioPortal study prefixes, not guaranteed biological-sample identifiers across studies; overlapping cohorts can count the same patient/sample more than once. Prefer one of these safer approaches: - Use a shipped `cancer_study_query_preferences` cohort such as `pan_cancer_tcga` or `all_studies_non_redundant`. - Restrict to one named study. - Aggregate by `patient_unique_id` when the biological question is patient prevalence. ## STOP rule: a frequency above 100% means your query is wrong If your query returns a frequency over 100%, **do not try to debug or explain the data inconsistency to the user**. The cause is always one of these query bugs: - Summing mutation events instead of `COUNT(DISTINCT sample_unique_id)` for the numerator - Using a study-wide sample count as the denominator instead of the gene-specific profiled count - Cross-study aggregation where the same biological sample appears under multiple `sample_unique_id` values (e.g., MSK-IMPACT and MSK-CHORD share patients) - **Joining the profiled CTE through `gene_panel` / `gene_panel_list` without a WES branch.** `gene_panel_id = 'WES'` is *not* a row in … (35166 more chars) ▶ read_guide { "uri": "cbioportal://sample-filtering-guide" } ▶ get_study_guide { "study_id": "msk_impact_2017" } ◀ result # Sample and Study Filtering Guide ## Overview Proper filtering is essential for meaningful cBioPortal analysis. This guide covers filtering by studies, sample types, and other criteria. ## Study-Level Filtering ### 1. Always Filter by Study Every query should specify a study to ensure consistent results: ```sql -- Always include study filtering SELECT * FROM your_table WHERE cancer_study_identifier = 'your_study_id' -- Additional filters... ``` ### 2. Find Available Studies ```sql -- Discover available studies SELECT cancer_study_identifier, name, description, type_of_cancer_id FROM cancer_study ORDER BY cancer_study_identifier; ``` ### 3. Study Information ```sql -- Get detailed study information SELECT cs.cancer_study_identifier, cs.name as study_name, cs.description, cs.sample_count, COUNT(DISTINCT p.internal_id) as patient_count FROM cancer_study cs LEFT JOIN patient p ON cs.cancer_study_id = p.cancer_study_id WHERE cs.cancer_study_identifier = 'your_study_id' GROUP BY cs.cancer_study_identifier, cs.name, cs.description, cs.sample_count; ``` ### 4. Find Studies by Available Data Types Use this when the user asks *"which studies have mutation and copy-number data for X"*, *"studies with expression for Y"*, *"is there any imaging data"*, or any *"studies with Z data"* question. `cancer_study` has one precomputed sample count per data type — filter on them. No join and no schema exploration needed. These are the same counts as the portal's study list and its "Data type" filter. **Canonical query — studies with ALL requested data types for a cancer type:** ```sql SELECT cancer_study_identifier, name, sample_count, mutation_sample_count, cna_sample_count FROM cancer_study WHERE type_of_cancer_id = 'luad' -- ← OncoTree code from search_oncotree AND mutation_sample_count > 0 -- ← one condition per required data type AND cna_sample_count > 0 ORDER BY sample_count DESC; ``` **Data type → column** (a count of 0 means the study doesn't have that data): | Data type in the user's question | Column (portal "Data type" filter label) | |---|---| | samples in the study | `sample_count` | | mutation / mutations | `mutation_sample_count` ("Mutations") | | copy-number / CNA / amplification / deletion | `cna_sample_count` ("CNA") | | structural variant / fusion | `structural_variant_sample_count` | | RNA / mRNA / gene expression (any platform) | `mrna_expression_sample_count` — use this for "which studies have expression data" | | specifically the portal's "RNA-Seq" filter | `rna_seq_sample_count` ("RNA-Seq") — counts only the `_rna_seq_v2_mrna` list; 0 for many GDC and iAtlas studies that do have RNA-Seq | | specifically the portal's "RNA (microarray)" filter | `mrna_microarray_sample_count` ("RNA (microarray)") — counts only the `_mrna` list | | microRNA | `mirna_sample_count` ("miRNA") | | protein / RPPA | `rppa_sample_count` ("RPPA") | | mas … (7083 more chars) ▶ get_study_guide { "study_id": "msk_impact_50k_2026" } ◀ result # Study Guide: MSK-IMPACT Clinical Sequencing Cohort (MSK, Nat Med 2017) **Study ID:** `msk_impact_2017` **Cancer Type:** mixed **Description:** Targeted sequencing of 10,000 clinical cases using the MSK-IMPACT assay ## Cohort Statistics - **Patients:** 10,336 - **Samples:** 10,946 ## Available Data Types - **COPY_NUMBER_ALTERATION**: Copy Number Alterations (MSK-IMPACT) - **MUTATION_EXTENDED**: Mutations (MSK-IMPACT) - **STRUCTURAL_VARIANT**: Structural variants ## Gene Panels - **IMPACT410**: 8,136 samples - **IMPACT341**: 2,809 samples ## Available Clinical Attributes | Attribute | Samples with Data | |-----------|------------------| | CANCER_TYPE_DETAILED | 10,945 | | PRIMARY_SITE | 10,945 | | MUTATION_COUNT | 10,945 | | FRACTION_GENOME_ALTERED | 10,945 | | TMB_NONSYNONYMOUS | 10,945 | | SAMPLE_CLASS | 10,945 | | SPECIMEN_TYPE | 10,945 | | SPECIMEN_PRESERVATION_TYPE | 10,945 | | CANCER_TYPE | 10,945 | | DNA_INPUT | 10,945 | | SOMATIC_STATUS | 10,945 | | SAMPLE_TYPE | 10,945 | | ONCOTREE_CODE | 10,945 | | SAMPLE_COLLECTION_SOURCE | 10,945 | | MATCHED_STATUS | 10,945 | | SAMPLE_COVERAGE | 10,945 | | TUMOR_PURITY | 10,945 | | METASTATIC_SITE | 10,945 | | OS_MONTHS | 1 | | VITAL_STATUS | 1 | ## Top Mutated Genes | Gene | Altered Samples | |------|----------------| | TP53 | 4,538 | | KRAS | 1,643 | | TERT | 1,460 | | PIK3CA | 1,355 | | APC | 1,121 | | ARID1A | 875 | | KMT2D | 851 | | PTEN | 665 | | KMT2C | 642 | | EGFR | 636 | ## Sample Types - **Primary**: 6,213 samples - **Metastasis**: 4,732 samples ## Query Tips for msk_impact_2017 ```sql -- Get all samples in this study SELECT DISTINCT sample_unique_id, patient_unique_id FROM clinical_data_derived WHERE cancer_study_identifier = 'msk_impact_2017'; -- Get mutations for a specific gene SELECT sample_unique_id, hugo_gene_symbol, mutation_variant, mutation_type FROM genomic_event_derived WHERE cancer_study_identifier = 'msk_impact_2017' AND hugo_gene_symbol = 'TP53' AND variant_type = 'mutation'; -- Get clinical data for specific attributes SELECT sample_unique_id, attribute_name, attribute_value FROM clinical_data_derived WHERE cancer_study_identifier = 'msk_impact_2017' AND attribute_name IN ('CANCER_TYPE', 'SAMPLE_TYPE', 'OS_MONTHS'); ``` ◀ result # Study Guide: MSK-IMPACT 50K Clinical Sequencing Cohort (MSK, Cancer Cell 2026) **Study ID:** `msk_impact_50k_2026` **Cancer Type:** mixed **Description:** MSK-50K (Memorial Sloan Kettering - 50,000 patients) is a pan-cancer clinical sequencing cohort of 54,331 tumors and their matched normal samples, profiled using MSK-IMPACT. The dataset includes somatic mutations, gene-level copy number alterations, structural variants in select genes, genome-wide allele-specific copy number, cancer cell fractions of mutations, derived mutational signatures, whole genome doubling status, and somatic/germline HLA Class I Genotype data. This data is available under the Creative Commons BY-NC-ND 4.0 license. For commercial use, please contact datarequests@mskcc.org. ## Cohort Statistics - **Patients:** 48,179 - **Samples:** 54,332 ## Available Data Types - **COPY_NUMBER_ALTERATION**: Copy Number Alterations - **GENERIC_ASSAY**: Arm-level copy-number from FACETS - **MUTATION_EXTENDED**: Mutations - **STRUCTURAL_VARIANT**: Structural Variants - **GENERIC_ASSAY**: Mutational Signature Contribution Scores (COSMIC v2) - **GENERIC_ASSAY**: HLA Loss of Heterozygosity ## Gene Panels - **WES** (Whole Exome): 54,330 samples — all genes profiled - **IMPACT468**: 34,549 samples - **IMPACT410**: 9,366 samples - **IMPACT505**: 7,817 samples - **IMPACT341**: 2,599 samples ## Available Clinical Attributes | Attribute | Samples with Data | |-----------|------------------| | PURITY_ESTIMATE_FROM_MUTATIONS | 54,331 | | FRACTION_GENOME_ALTERED | 54,331 | | CANCER_TYPE_DETAILED | 54,331 | | TMB_SCORE | 54,331 | | HLA_C2_LOH | 54,331 | | HLA_A2_LOH | 54,331 | | HLA_C1_LOH | 54,331 | | TUMOR_PURITY | 54,331 | | HLA_A1_LOH | 54,331 | | METASTATIC_SITE | 54,331 | | HLA_B1_LOH | 54,331 | | SOMATIC_STATUS | 54,331 | | MSI_SCORE | 54,331 | | GENE_PANEL | 54,331 | | ONCOTREE_CODE | 54,331 | | DISEASE_STATUS | 54,331 | | FACETS_PURITY | 54,331 | | FACETS_QC | 54,331 | | FACETS_PLOIDY | 54,331 | | FACETS_WGD | 54,331 | ## Top Mutated Genes | Gene | Altered Samples | |------|----------------| | TP53 | 24,541 | | KRAS | 8,917 | | PIK3CA | 7,636 | | TERT | 7,409 | | APC | 6,032 | | ARID1A | 5,343 | | KMT2D | 4,744 | | PTEN | 4,096 | | KMT2C | 3,632 | | EGFR | 3,290 | ## Sample Types - **Primary**: 32,878 samples - **Metastasis**: 20,830 samples - **Unknown**: 378 samples - **Local Recurrence**: 245 samples ## Query Tips for msk_impact_50k_2026 ```sql -- Get all samples in this study SELECT DISTINCT sample_unique_id, patient_unique_id FROM clinical_data_derived WHERE cancer_study_identifier = 'msk_impact_50k_2026'; -- Get mutations for a specific gene SELECT sample_unique_id, hugo_gene_symbol, mutation_variant, mutation_type FROM genomic_event_derived WHERE cancer_study_identifier = 'msk_impact_50k_2026' AND hugo_gene_symbol = 'TP53' AND variant_type = 'mutation'; -- Get clinical data for specific attributes SELECT sample_unique_id, … (189 more chars) ▶ clickhouse_run_select_query { "query": "SELECT cancer_study_identifier, attribute_name, COUNT(DISTINCT sample_unique_id) AS n FROM clinical_data_derived WHERE cancer_study_identifier IN ('msk_impact_2017','msk_impact_50k_2026') AND (lower(attribute_name) LIKE '%smok%') GROUP BY cancer_study_identifier, attribute_name ORDER BY cancer_study_identifier, attribute_name" } ◀ result {"rows":[{"cancer_study_identifier":"msk_impact_2017","attribute_name":"SMOKING_HISTORY","n":1}]} ▶ clickhouse_run_select_query { "query": "SELECT cs.cancer_study_identifier, cs.name, COUNT(DISTINCT cd.sample_unique_id) AS n FROM clinical_data_derived cd JOIN cancer_study cs ON cs.cancer_study_identifier = cd.cancer_study_identifier WHERE lower(cd.attribute_name) LIKE '%smok%' GROUP BY cs.cancer_study_identifier, cs.name ORDER BY n DESC LIMIT 30" } ◀ result {"rows":[{"cs.cancer_study_identifier":"hcc_clca_2024","name":"Hepatocellular Carcinoma (CLCA, Nature 2024)","n":494},{"cs.cancer_study_identifier":"bm_nsclc_mskcc_2023","name":"Non-Small Cell Lung Cancer Brain Metastasis (MSK, Nat Commun 2023)","n":322},{"cs.cancer_study_identifier":"lusc_cptac_2021","name":"Lung Squamous Cell Carcinoma (CPTAC, Cell 2021)","n":109},{"cs.cancer_study_identifier":"crc_sysucc_2022","name":"Colorectal Cancer- ChangKang Project (SYSUCC, Nat Commun 2022)","n":1},{"cs.cancer_study_identifier":"bladder_msk_2024","name":"Bladder Cancer (MSK, Eur Urol 2024)","n":1},{"cs.cancer_study_identifier":"luad_mskcc_2015","name":"Lung Adenocarcinoma (MSK, Science 2015)","n":1},{"cs.cancer_study_identifier":"lusc_tcga_pub","name":"Lung Squamous Cell Carcinoma (TCGA, Nature 2012)","n":1},{"cs.cancer_study_identifier":"sclc_ucologne_2015","name":"Small Cell Lung Cancer (U Cologne, Nature 2015)","n":1},{"cs.cancer_study_identifier":"kich_tcga_pub","name":"Kidney Chromophobe (TCGA, Cancer Cell 2014)","n":1},{"cs.cancer_study_identifier":"panet_msk_2018","name":"Metastatic Pancreatic Neuroendocrine Tumor (MSK, JCO Precis Oncol 2018)","n":1},{"cs.cancer_study_identifier":"hnsc_tcga_gdc","name":"Head and Neck Squamous Cell Carcinoma (TCGA GDC, 2025)","n":1},{"cs.cancer_study_identifier":"kich_tcga","name":"Kidney Chromophobe (TCGA, Firehose Legacy)","n":1},{"cs.cancer_study_identifier":"luad_tcga_pub","name":"Lung Adenocarcinoma (TCGA, Nature 2014)","n":1},{"cs.cancer_study_identifier":"lung_msk_mind_2020","name":"Lung Adenocarcinoma (MSK Mind,Nature Cancer 2022)","n":1},{"cs.cancer_study_identifier":"esca_broad","name":"Esophageal Adenocarcinoma (DFCI, Nat Genet 2013)","n":1},{"cs.cancer_study_identifier":"mixed_allen_2018","name":"MSS Mixed Solid Tumors (Broad/Dana-Farber, Nat Genet 2018)","n":1},{"cs.cancer_study_identifier":"blca_mskcc_solit_2014","name":"Bladder Cancer (MSK, Eur Urol 2014)","n":1},{"cs.cancer_study_identifier":"msk_chord_2024","name":"MSK-CHORD (MSK, Nature 2024)","n":1},{"cs.cancer_study_identifier":"chrcc_tcga_gdc","name":"Chromophobe Renal Cell Carcinoma (TCGA GDC, 2025)","n":1},{"cs.cancer_study_identifier":"alk_msk_2026","name":"Non-Small Cell Lung Cancer (MSK, 2026)","n":1},{"cs.cancer_study_identifier":"stes_tcga_pub","name":"Esophageal Carcinoma (TCGA, Nature 2017)","n":1},{"cs.cancer_study_identifier":"mbn_msk_2024","name":"Diffuse Large B-Cell Lymphoma (MSK, Haematologica 2024)","n":1},{"cs.cancer_study_identifier":"lusc_tcga_gdc","name":"Lung Squamous Cell Carcinoma (TCGA GDC, 2025)","n":1},{"cs.cancer_study_identifier":"luad_cptac_2020","name":"Lung Adenocarcinoma (CPTAC, Cell 2020)","n":1},{"cs.cancer_study_identifier":"paad_cptac_2021","name":"Pancreatic Ductal Adenocarcinoma (CPTAC, Cell 2021)","n":1},{"cs.cancer_study_identifier":"crc_eo_2020","name":"Colorectal Cancer (MSK, JNCI 2021)","n":1},{"cs.cancer_study_identifier":"kirp_tcga","name":"Kidney Renal Papillary Cell Carcinoma (TCGA, Firehose Legacy … (358 more chars) ▶ clickhouse_run_select_query { "query": "SELECT cancer_study_identifier, attribute_name, COUNT(DISTINCT sample_unique_id) AS n FROM clinical_data_derived WHERE cancer_study_identifier IN ('luad_mskcc_2020','luad_mskimpact_2021','luad_mskcc_2015','lung_msk_mind_2020','bm_nsclc_mskcc_2023') AND (lower(attribute_name) LIKE '%smok%' OR attribute_name IN ('SMOKING_HISTORY','SMOKING_STATUS','SMOKING_PACK_YEARS')) GROUP BY cancer_study_identifier, attribute_name ORDER BY cancer_study_identifier, attribute_name" } ◀ result {"rows":[{"cancer_study_identifier":"bm_nsclc_mskcc_2023","attribute_name":"SMOKING_STATUS","n":322},{"cancer_study_identifier":"luad_mskcc_2015","attribute_name":"MOLECULAR_SMOKING_SIGNATURE","n":1},{"cancer_study_identifier":"luad_mskcc_2015","attribute_name":"SMOKING_HISTORY","n":1},{"cancer_study_identifier":"luad_mskcc_2015","attribute_name":"SMOKING_PACK_YEARS","n":1},{"cancer_study_identifier":"luad_mskcc_2020","attribute_name":"SMOKING_HISTORY","n":1},{"cancer_study_identifier":"luad_mskcc_2020","attribute_name":"SMOKING_PACK_YEARS","n":1},{"cancer_study_identifier":"lung_msk_mind_2020","attribute_name":"SMOKING_STATUS","n":1}]} ═ answer (success) None of the dedicated LUAD studies have usable smoking data (n=1 is a data artifact, not real coverage) except `bm_nsclc_mskcc_2023` (brain metastasis NSCLC, not the "clinicogenomic sequencing cohort" the user named, and not pure LUAD). Given the user explicitly named the MSKCC clinicogenomic **sequencing cohort**, this points to `msk_impact_2017` (the study literally titled "MSK-IMPACT Clinical Sequencing Cohort") — but its smoking data coverage is essentially absent (1 sample). I need to flag this rather than silently substitute a different cohort or proceed with a broken filter. I looked into the MSK-IMPACT Clinical Sequencing Cohort (`msk_impact_2017`) — the study whose name matches "MSKCC Clinical Sequencing Cohort" — and it does **not** contain usable smoking status data: only 1 of its 10,945 samples has a `SMOKING_HISTORY` value recorded. The newer `msk_impact_50k_2026` cohort has no smoking attribute at all. I also checked the other dedicated MSK lung adenocarcinoma studies (`luad_mskcc_2015`, `luad_mskcc_2020`, `luad_mskimpact_2021`) — same problem, only a single sample each has any smoking field populated, which isn't real coverage. The one MSK lung cohort that *does* have smoking status broadly annotated is: - **Non-Small Cell Lung Cancer Brain Metastasis (MSK, Nat Commun 2023)** — `bm_nsclc_mskcc_2023`, 322 samples with `SMOKING_STATUS`. That's a brain-metastasis-focused NSCLC cohort, not the general "clinicogenomic sequencing cohort" you asked about, and it's NSCLC broadly rather than pure adenocarcinoma. **Before I build the OncoPrint, please clarify which you'd like:** 1. Proceed with `bm_nsclc_mskcc_2023` (NSCLC brain-met cohort), optionally filtered to adenocarcinoma histology if that attribute is present, since it's the only MSK lung cohort with real smoking-status coverage; or 2. Confirm you specifically meant `msk_impact_2017`, in which case the smoking stratification cannot be performed (data isn't there) and I can only show the unstratified oncoprint; or 3. Point me to a different specific study name/ID you had in mind. I don't want to silently substitute a cohort or drop the smoking stratification without your confirmation — let me know how you'd like to proceed.