Q (sonnet-5.5): generate a oncoprint of kras, stk11, keap1, tp53 alterations stratified by smoking status in lung adenocarcinoma from the mskcc clincogenomic sequencing cohort. Please only select oncogenic somatic alterations and remove samples with unknown smoking status ▶ list_studies { "search": "MSK-CHORD" } ▶ list_studies { "search": "clinical sequencing" } ◀ result {"result":[{"cancer_study_identifier":"msk_chord_2024","name":"MSK-CHORD (MSK, Nature 2024)","type_of_cancer_id":"mixed","sample_count":25040,"has_guide":true,"url":"https://www.cbioportal.org/study/summary?id=msk_chord_2024"}]} ▶ read_guide { "uri": "cbioportal://sample-filtering-guide" } ◀ result {"result":[{"cancer_study_identifier":"msk_impact_50k_2026","name":"MSK-IMPACT 50K Clinical Sequencing Cohort (MSK, Cancer Cell 2026)","type_of_cancer_id":"mixed","sample_count":54331,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=msk_impact_50k_2026"},{"cancer_study_identifier":"msk_impact_2017","name":"MSK-IMPACT Clinical Sequencing Cohort (MSK, Nat Med 2017)","type_of_cancer_id":"mixed","sample_count":10945,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=msk_impact_2017"},{"cancer_study_identifier":"braf_msk_archer_2024","name":"BRAF Fusions - ARCHER Clinical Sequencing Cohort (MSK, Clin Cancer Res 2024)","type_of_cancer_id":"mixed","sample_count":107,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=braf_msk_archer_2024"},{"cancer_study_identifier":"braf_msk_impact_2024","name":"BRAF Fusions - IMPACT Clinical Sequencing Cohort (MSK, Clin Cancer Res 2024)","type_of_cancer_id":"mixed","sample_count":105,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=braf_msk_impact_2024"},{"cancer_study_identifier":"msk_chord_2024","name":"MSK-CHORD (MSK, Nature 2024)","type_of_cancer_id":"mixed","sample_count":25040,"has_guide":true,"url":"https://www.cbioportal.org/study/summary?id=msk_chord_2024"},{"cancer_study_identifier":"crc_msk_2026","name":"Colorectal Adenocarcinoma (MSK, 2026)","type_of_cancer_id":"read","sample_count":7237,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=crc_msk_2026"},{"cancer_study_identifier":"pancan_mimsi_msk_2024","name":"Mixed Tumors - MiMSI Cohort (MSK, Nat Commun 2024)","type_of_cancer_id":"mixed","sample_count":5033,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=pancan_mimsi_msk_2024"},{"cancer_study_identifier":"breast_msk_2025","name":"Breast Cancer (MSK, Nat Genet 2025)","type_of_cancer_id":"brca","sample_count":3879,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=breast_msk_2025"},{"cancer_study_identifier":"brca_mbcproject_2022","name":"The Metastatic Breast Cancer Project (Provisional, December 2021)","type_of_cancer_id":"brca","sample_count":379,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=brca_mbcproject_2022"},{"cancer_study_identifier":"angs_painter_2025","name":"The Angiosarcoma Project - Count Me In (Provisional, April 2025)","type_of_cancer_id":"angs","sample_count":328,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=angs_painter_2025"},{"cancer_study_identifier":"brca_msk_erbb2_2026","name":"HER2-Low and HER2-Null Metastatic Breast Cancers treated with T-DXd (MSK, Clinical Cancer Research 2026)","type_of_cancer_id":"breast","sample_count":272,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=brca_msk_erbb2_2026"},{"cancer_study_identifier":"rcc_iatlas_immotion150_2018","name":"Renal Cell Carcinoma (IMmotion150 Clinical Trial, Nat Med. 2018) - iAtlas Harmonized","type_of_cancer_id":"rcc","sample_count":26 … (2065 more chars) ◀ result # Sample and Study Filtering Guide ## Overview Proper filtering is essential for meaningful cBioPortal analysis. This guide covers filtering by studies, sample types, and other criteria. ## Study-Level Filtering ### 1. Always Filter by Study Every query should specify a study to ensure consistent results: ```sql -- Always include study filtering SELECT * FROM your_table WHERE cancer_study_identifier = 'your_study_id' -- Additional filters... ``` ### 2. Find Available Studies ```sql -- Discover available studies SELECT cancer_study_identifier, name, description, type_of_cancer_id FROM cancer_study ORDER BY cancer_study_identifier; ``` ### 3. Study Information ```sql -- Get detailed study information SELECT cs.cancer_study_identifier, cs.name as study_name, cs.description, cs.sample_count, COUNT(DISTINCT p.internal_id) as patient_count FROM cancer_study cs LEFT JOIN patient p ON cs.cancer_study_id = p.cancer_study_id WHERE cs.cancer_study_identifier = 'your_study_id' GROUP BY cs.cancer_study_identifier, cs.name, cs.description, cs.sample_count; ``` ### 4. Find Studies by Available Data Types Use this when the user asks *"which studies have mutation and copy-number data for X"*, *"studies with expression for Y"*, *"is there any imaging data"*, or any *"studies with Z data"* question. `cancer_study` has one precomputed sample count per data type — filter on them. No join and no schema exploration needed. These are the same counts as the portal's study list and its "Data type" filter. **Canonical query — studies with ALL requested data types for a cancer type:** ```sql SELECT cancer_study_identifier, name, sample_count, mutation_sample_count, cna_sample_count FROM cancer_study WHERE type_of_cancer_id = 'luad' -- ← OncoTree code from search_oncotree AND mutation_sample_count > 0 -- ← one condition per required data type AND cna_sample_count > 0 ORDER BY sample_count DESC; ``` **Data type → column** (a count of 0 means the study doesn't have that data): | Data type in the user's question | Column (portal "Data type" filter label) | |---|---| | samples in the study | `sample_count` | | mutation / mutations | `mutation_sample_count` ("Mutations") | | copy-number / CNA / amplification / deletion | `cna_sample_count` ("CNA") | | structural variant / fusion | `structural_variant_sample_count` | | RNA / mRNA / gene expression (any platform) | `mrna_expression_sample_count` — use this for "which studies have expression data" | | specifically the portal's "RNA-Seq" filter | `rna_seq_sample_count` ("RNA-Seq") — counts only the `_rna_seq_v2_mrna` list; 0 for many GDC and iAtlas studies that do have RNA-Seq | | specifically the portal's "RNA (microarray)" filter | `mrna_microarray_sample_count` ("RNA (microarray)") — counts only the `_mrna` list | | microRNA | `mirna_sample_count` ("miRNA") | | protein / RPPA | `rppa_sample_count` ("RPPA") | | mas … (7866 more chars) ═ answer (success) "MSKCC clinicogenomic sequencing cohort" matches more than one study, so I need to know which you mean before I build the OncoPrint: - **MSK-CHORD** (`msk_chord_2024`, 25,040 samples). This is MSK's clinicogenomic dataset from Nature 2024, and it includes smoking history. - **MSK-IMPACT 50K Clinical Sequencing Cohort** (`msk_impact_50k_2026`, 54,331 samples). This is from Cancer Cell 2026. - **MSK-IMPACT Clinical Sequencing Cohort** (`msk_impact_2017`, 10,945 samples). This is from Nat Med 2017. Which one should I use? I can also run it on more than one, or all three. Smoking status may not be recorded in every one of these studies. I'll check that once you pick.