Q (haiku): What are the most frequently altered genes in KRAS wild-type lung adenocarcinoma patients? ▶ read_guide { "uri": "cbioportal://mutation-frequency-guide" } ▶ read_guide { "uri": "cbioportal://sample-filtering-guide" } ◀ result # Mutation Frequency Analysis Guide ## IMPORTANT: Reporting Mutation Frequencies - **ALWAYS report frequencies as percentages**, not raw counts: `frequency = (altered_samples / total_profiled_samples) × 100` - For quick frequency lookups, **prefer the TCGA Pan-Cancer Atlas study first**, then offer to expand to other studies - When reporting across multiple studies, show **ranges** (e.g., "TP53 is mutated in 30–60% of samples") rather than a single average - **NEVER** sum mutation events across studies to compute an aggregate frequency — this can exceed 100% due to double-counting - Warn users that samples may overlap across cohorts (e.g., MSK studies may share patients) - **Choose and state the counting unit**: use patient-level frequencies for prevalence/rate questions unless the user explicitly asks for samples; use sample-level frequencies when the user asks about samples. - **For "across cancer types" questions**, jump to the [Cross-Cancer-Type Mutation Frequency](#cross-cancer-type-mutation-frequency) section below — there is one correct recipe and several common wrong ones. ## Counting Unit: Samples vs Patients Before answering any mutation count or frequency question, decide whether the unit is samples or patients and state that choice in the answer. | User wording | Counting unit | |--------------|---------------| | "prevalence", "rate", "fraction of patients", "patients with", "how common is" | Patient-level: `COUNT(DISTINCT patient_unique_id)` | | "samples", "specimens", "biopsies", sample-level cohort composition | Sample-level: `COUNT(DISTINCT sample_unique_id)` | | Ambiguous | Ask, or default to patient-level for prevalence/rate language and say so | ### Cross-study sample-count caveat When an answer touches more than one study and reports a sample count, prepend a one-line caveat: > Sample IDs are unique within cBioPortal study prefixes, not guaranteed biological-sample identifiers across studies; overlapping cohorts can count the same patient/sample more than once. Prefer one of these safer approaches: - Use a shipped `cancer_study_query_preferences` cohort such as `pan_cancer_tcga` or `all_studies_non_redundant`. - Restrict to one named study. - Aggregate by `patient_unique_id` when the biological question is patient prevalence. ## STOP rule: a frequency above 100% means your query is wrong If your query returns a frequency over 100%, **do not try to debug or explain the data inconsistency to the user**. The cause is always one of these query bugs: - Summing mutation events instead of `COUNT(DISTINCT sample_unique_id)` for the numerator - Using a study-wide sample count as the denominator instead of the gene-specific profiled count - Cross-study aggregation where the same biological sample appears under multiple `sample_unique_id` values (e.g., MSK-IMPACT and MSK-CHORD share patients) - **Joining the profiled CTE through `gene_panel` / `gene_panel_list` without a WES branch.** `gene_panel_id = 'WES'` is *not* a row in … (35166 more chars) ◀ result # Sample and Study Filtering Guide ## Overview Proper filtering is essential for meaningful cBioPortal analysis. This guide covers filtering by studies, sample types, and other criteria. ## Study-Level Filtering ### 1. Always Filter by Study Every query should specify a study to ensure consistent results: ```sql -- Always include study filtering SELECT * FROM your_table WHERE cancer_study_identifier = 'your_study_id' -- Additional filters... ``` ### 2. Find Available Studies ```sql -- Discover available studies SELECT cancer_study_identifier, name, description, type_of_cancer_id FROM cancer_study ORDER BY cancer_study_identifier; ``` ### 3. Study Information ```sql -- Get detailed study information SELECT cs.cancer_study_identifier, cs.name as study_name, cs.description, cs.sample_count, COUNT(DISTINCT p.internal_id) as patient_count FROM cancer_study cs LEFT JOIN patient p ON cs.cancer_study_id = p.cancer_study_id WHERE cs.cancer_study_identifier = 'your_study_id' GROUP BY cs.cancer_study_identifier, cs.name, cs.description, cs.sample_count; ``` ### 4. Find Studies by Available Data Types Use this when the user asks *"which studies have mutation and copy-number data for X"*, *"studies with expression for Y"*, *"is there any imaging data"*, or any *"studies with Z data"* question. `cancer_study` has one precomputed sample count per data type — filter on them. No join and no schema exploration needed. These are the same counts as the portal's study list and its "Data type" filter. **Canonical query — studies with ALL requested data types for a cancer type:** ```sql SELECT cancer_study_identifier, name, sample_count, mutation_sample_count, cna_sample_count FROM cancer_study WHERE type_of_cancer_id = 'luad' -- ← OncoTree code from search_oncotree AND mutation_sample_count > 0 -- ← one condition per required data type AND cna_sample_count > 0 ORDER BY sample_count DESC; ``` **Data type → column** (a count of 0 means the study doesn't have that data): | Data type in the user's question | Column (portal "Data type" filter label) | |---|---| | samples in the study | `sample_count` | | mutation / mutations | `mutation_sample_count` ("Mutations") | | copy-number / CNA / amplification / deletion | `cna_sample_count` ("CNA") | | structural variant / fusion | `structural_variant_sample_count` | | RNA / mRNA / gene expression (any platform) | `mrna_expression_sample_count` — use this for "which studies have expression data" | | specifically the portal's "RNA-Seq" filter | `rna_seq_sample_count` ("RNA-Seq") — counts only the `_rna_seq_v2_mrna` list; 0 for many GDC and iAtlas studies that do have RNA-Seq | | specifically the portal's "RNA (microarray)" filter | `mrna_microarray_sample_count` ("RNA (microarray)") — counts only the `_mrna` list | | microRNA | `mirna_sample_count` ("miRNA") | | protein / RPPA | `rppa_sample_count` ("RPPA") | | mas … (7083 more chars) ▶ search_oncotree { "search_term": "lung adenocarcinoma" } ◀ result {"result":[{"code":"LUAD","name":"Lung Adenocarcinoma","score":80,"mainType":"Non-Small Cell Lung Cancer","tissue":"Lung","hierarchy":"LUNG > NSCLC > LUAD"},{"code":"LAIS","name":"Lung Adenocarcinoma In Situ","score":60,"mainType":"Non-Small Cell Lung Cancer","tissue":"Lung","hierarchy":"LUNG > LAIS"}]} ▶ list_studies { "search": "lung adenocarcinoma", "limit": 10 } ◀ result {"result":[{"cancer_study_identifier":"msk_met_2021","name":"MSK MetTropism (MSK, Cell 2021)","type_of_cancer_id":"mixed","sample_count":25775,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=msk_met_2021"},{"cancer_study_identifier":"luad_mskcc_2023_met_organotropism","name":"Lung Adenocarcinoma Met Organotropism (MSK, Cancer Cell 2023)","type_of_cancer_id":"luad","sample_count":2653,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=luad_mskcc_2023_met_organotropism"},{"cancer_study_identifier":"lung_msk_2017","name":"Non-Small Cell Cancer (MSK, Cancer Discov 2017)","type_of_cancer_id":"luad","sample_count":915,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=lung_msk_2017"},{"cancer_study_identifier":"luad_mskcc_2020","name":"Lung Adenocarcinoma (MSK, J Thorac Oncol 2020)","type_of_cancer_id":"luad","sample_count":604,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=luad_mskcc_2020"},{"cancer_study_identifier":"luad_tcga","name":"Lung Adenocarcinoma (TCGA, Firehose Legacy)","type_of_cancer_id":"luad","sample_count":586,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=luad_tcga"},{"cancer_study_identifier":"luad_tcga_gdc","name":"Lung Adenocarcinoma (TCGA GDC, 2025)","type_of_cancer_id":"luad","sample_count":571,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=luad_tcga_gdc"},{"cancer_study_identifier":"luad_tcga_pan_can_atlas_2018","name":"Lung Adenocarcinoma (TCGA, PanCancer Atlas)","type_of_cancer_id":"luad","sample_count":566,"has_guide":true,"url":"https://www.cbioportal.org/study/summary?id=luad_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"luad_msk_npjpo_2021","name":"Lung Adenocarcinoma (MSK, NPJ Precision Oncology 2021)","type_of_cancer_id":"luad","sample_count":426,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=luad_msk_npjpo_2021"},{"cancer_study_identifier":"luad_oncosg_2020","name":"Lung Adenocarcinoma (OncoSG, Nat Genet 2020)","type_of_cancer_id":"luad","sample_count":305,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=luad_oncosg_2020"},{"cancer_study_identifier":"lung_msk_mind_2020","name":"Lung Adenocarcinoma (MSK Mind,Nature Cancer 2022)","type_of_cancer_id":"lung","sample_count":247,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=lung_msk_mind_2020"}]} ▶ clickhouse_run_select_query { "query": " -- Get KRAS wild-type samples first WITH kras_wt AS ( SELECT DISTINCT s.sample_unique_id, s.cancer_study_identifier FROM genomic_event_derived s -- Left join to find KRAS mutations LEFT JOIN genomic_event_derived kras_mut ON s.sample_unique_id = kras_mut.sample_unique_id AND kras_mut.cancer_study_identifier = s.cancer_study_identifier AND kras_mut.hugo_gene_symbol = 'KRAS' AND kras_mut.variant_type = 'mutation' AND kras_mut.mutation_status != 'UNCALLED' WHERE s.cancer_study_identifier = 'luad_tcga_pan_can_atlas_2018' -- Get samples profiled for KRAS AND EXISTS ( SELECT 1 FROM sample_to_gene_panel_derived sgp WHERE sgp.sample_unique_id = s.sample_unique_id AND sgp.cancer_study_identifier = s.cancer_study_identifier AND sgp.alteration_type = 'MUTATION_EXTENDED' AND EXISTS ( SELECT 1 FROM gene_panel gp JOIN gene_panel_list gpl ON gp.internal_id = gpl.internal_id JOIN gene g ON gpl.gene_id = g.entrez_gene_id WHERE gp.stable_id = sgp.gene_panel_id AND g.hugo_gene_symbol = 'KRAS' ) ) -- Only include samples with no KRAS mutation (wild-type) AND kras_mut.sample_unique_id IS NULL ), -- Get altered gene counts in KRAS WT samples altered_genes AS ( SELECT ged.hugo_gene_symbol, COUNT(DISTINCT ged.sample_unique_id) AS altered_samples, COUNT(*) AS total_events FROM genomic_event_derived ged JOIN kras_wt wt ON ged.sample_unique_id = wt.sample_unique_id AND ged.cancer_study_identifier = wt.cancer_study_identifier WHERE ged.variant_type = 'mutation' AND ged.mutation_status != 'UNCALLED' AND ged.off_panel = 0 GROUP BY ged.hugo_gene_symbol ), -- Get profiled counts for each gene in KRAS WT cohort profiled_genes AS ( SELECT g.hugo_gene_symbol, COUNT(DISTINCT sgp.sample_unique_id) AS profiled_samples FROM sample_to_gene_panel_derived sgp JOIN gene_panel gp ON gp.stable_id = sgp.gene_panel_id JOIN gene_panel_list gpl ON gp.internal_id = gpl.internal_id JOIN gene g ON gpl.gene_id = g.entrez_gene_id JOIN kras_wt wt ON sgp.sample_unique_id = wt.sample_unique_id AND sgp.cancer_study_identifier = wt.cancer_study_identifier WHERE sgp.alteration_type = 'MUTATION_EXTENDED' GROUP BY g.hugo_gene_symbol ) SELECT a.hugo_gene_symbol, a.altered_samples, p.profiled_samples, ROUND(a.altered_samples * 100.0 / p.profiled_samples, 1) AS frequency_pct, a.total_events FROM altered_genes a JOIN profiled_genes p ON a.hugo_gene_symbol = p.hugo_gene_symbol ORDER BY frequency_pct DESC LIMIT 30; ", "max_rows": 100 } ◀ result {"error_message":"Query execution failed: HTTPDriver for https://dl96orhu96.us-east-1.aws.clickhouse.cloud:8443 received ClickHouse error code 1\n Code: 1. DB::Exception: Resolved identifier 's.sample_unique_id' in parent scope to expression 's.sample_unique_id' with correlated columns 'sample_unique_id' (Enable 'allow_experimental_correlated_subqueries' setting to allow correlated subqueries execution). In scope (SELECT _CAST(1, 'UInt64') FROM (SELECT 1 FROM cbioportal_public_librechat_blue.sample_to_gene_panel_derived AS sgp WHERE (sgp.sample_unique_id = s.sample_unique_id) AND (sgp.cancer_study_identifier = s.cancer_study_identifier) AND (sgp.alteration_type = 'MUTATION_EXTENDED') AND exists((SELECT 1 FROM gene_panel AS gp ALL INNER JOIN gene_panel_list AS gpl ON gp.internal_id = gpl.internal_id ALL INNER JOIN gene AS g ON gpl.gene_id = g.entrez_gene_id WHERE (gp.stable_id = sgp.gene_panel_id) AND (g.hugo_gene_symbol = 'KRAS')))) LIMIT _CAST(1, 'UInt64')). (UNSUPPORTED_METHOD) (version 26.4.1.2359 (official build))\n"} ▶ clickhouse_run_select_query { "query": " -- First, identify KRAS mutations in TCGA LUAD WITH kras_mutant_samples AS ( SELECT DISTINCT sample_unique_id FROM genomic_event_derived WHERE cancer_study_identifier = 'luad_tcga_pan_can_atlas_2018' AND hugo_gene_symbol = 'KRAS' AND variant_type = 'mutation' AND mutation_status != 'UNCALLED' ), -- Get all samples profiled for mutations in LUAD all_profiled_samples AS ( SELECT DISTINCT sample_unique_id FROM sample_to_gene_panel_derived WHERE cancer_study_identifier = 'luad_tcga_pan_can_atlas_2018' AND alteration_type = 'MUTATION_EXTENDED' ), -- KRAS wild-type = profiled for mutations but NOT in kras_mutant list kras_wt_samples AS ( SELECT sample_unique_id FROM all_profiled_samples WHERE sample_unique_id NOT IN (SELECT sample_unique_id FROM kras_mutant_samples) ), -- Get altered gene counts in KRAS WT samples altered_genes AS ( SELECT ged.hugo_gene_symbol, COUNT(DISTINCT ged.sample_unique_id) AS altered_samples, COUNT(*) AS total_events FROM genomic_event_derived ged WHERE cancer_study_identifier = 'luad_tcga_pan_can_atlas_2018' AND ged.sample_unique_id IN (SELECT sample_unique_id FROM kras_wt_samples) AND ged.variant_type = 'mutation' AND ged.mutation_status != 'UNCALLED' AND ged.off_panel = 0 GROUP BY ged.hugo_gene_symbol ), -- Get profiled counts for each gene in KRAS WT cohort using panel coverage profiled_genes AS ( SELECT g.hugo_gene_symbol, COUNT(DISTINCT sgp.sample_unique_id) AS profiled_samples FROM sample_to_gene_panel_derived sgp JOIN gene_panel gp ON gp.stable_id = sgp.gene_panel_id JOIN gene_panel_list gpl ON gp.internal_id = gpl.internal_id JOIN gene g ON gpl.gene_id = g.entrez_gene_id WHERE sgp.cancer_study_identifier = 'luad_tcga_pan_can_atlas_2018' AND sgp.sample_unique_id IN (SELECT sample_unique_id FROM kras_wt_samples) AND sgp.alteration_type = 'MUTATION_EXTENDED' GROUP BY g.hugo_gene_symbol ) SELECT a.hugo_gene_symbol, a.altered_samples, p.profiled_samples, ROUND(a.altered_samples * 100.0 / p.profiled_samples, 1) AS frequency_pct, a.total_events FROM altered_genes a JOIN profiled_genes p ON a.hugo_gene_symbol = p.hugo_gene_symbol ORDER BY frequency_pct DESC LIMIT 30; ", "max_rows": 100 } ◀ result {"rows":[]} ▶ clickhouse_run_select_query { "query": " -- Check basic data availability SELECT COUNT(DISTINCT sample_unique_id) as total_samples, COUNT(DISTINCT CASE WHEN variant_type = 'mutation' THEN sample_unique_id END) as mutation_samples, COUNT(*) as total_events FROM genomic_event_derived WHERE cancer_study_identifier = 'luad_tcga_pan_can_atlas_2018'; ", "max_rows": 10 } ◀ result {"rows":[{"total_samples":566,"mutation_samples":561,"total_events":342717}]} ▶ clickhouse_run_select_query { "query": " -- Check KRAS mutations SELECT COUNT(DISTINCT sample_unique_id) as kras_mutant_samples, COUNT(*) as kras_events FROM genomic_event_derived WHERE cancer_study_identifier = 'luad_tcga_pan_can_atlas_2018' AND hugo_gene_symbol = 'KRAS' AND variant_type = 'mutation' AND mutation_status != 'UNCALLED'; ", "max_rows": 10 } ◀ result {"rows":[{"kras_mutant_samples":168,"kras_events":172}]} ▶ clickhouse_run_select_query { "query": " -- Get top mutated genes in KRAS wild-type samples -- First, get all KRAS mutant samples WITH kras_mutant AS ( SELECT DISTINCT sample_unique_id FROM genomic_event_derived WHERE cancer_study_identifier = 'luad_tcga_pan_can_atlas_2018' AND hugo_gene_symbol = 'KRAS' AND variant_type = 'mutation' AND mutation_status != 'UNCALLED' ), -- Get gene frequencies in KRAS WT cohort gene_freqs AS ( SELECT hugo_gene_symbol, COUNT(DISTINCT sample_unique_id) as altered_samples, COUNT(*) as total_mutation_events FROM genomic_event_derived WHERE cancer_study_identifier = 'luad_tcga_pan_can_atlas_2018' AND variant_type = 'mutation' AND mutation_status != 'UNCALLED' AND off_panel = 0 AND sample_unique_id NOT IN (SELECT sample_unique_id FROM kras_mutant) GROUP BY hugo_gene_symbol ) SELECT gene_freqs.hugo_gene_symbol, gene_freqs.altered_samples, gene_freqs.total_mutation_events, (SELECT COUNT(DISTINCT sample_unique_id) FROM genomic_event_derived WHERE cancer_study_identifier = 'luad_tcga_pan_can_atlas_2018' AND variant_type = 'mutation' AND sample_unique_id NOT IN (SELECT sample_unique_id FROM kras_mutant)) AS wt_cohort_size FROM gene_freqs ORDER BY gene_freqs.altered_samples DESC LIMIT 30; ", "max_rows": 100 } ◀ result {"rows":[{"hugo_gene_symbol":"TP53","altered_samples":233,"total_mutation_events":247,"wt_cohort_size":393},{"hugo_gene_symbol":"TTN","altered_samples":190,"total_mutation_events":501,"wt_cohort_size":393},{"hugo_gene_symbol":"MUC16","altered_samples":158,"total_mutation_events":284,"wt_cohort_size":393},{"hugo_gene_symbol":"CSMD3","altered_samples":157,"total_mutation_events":276,"wt_cohort_size":393},{"hugo_gene_symbol":"RYR2","altered_samples":147,"total_mutation_events":289,"wt_cohort_size":393},{"hugo_gene_symbol":"LRP1B","altered_samples":132,"total_mutation_events":227,"wt_cohort_size":393},{"hugo_gene_symbol":"ZFHX4","altered_samples":122,"total_mutation_events":185,"wt_cohort_size":393},{"hugo_gene_symbol":"USH2A","altered_samples":120,"total_mutation_events":202,"wt_cohort_size":393},{"hugo_gene_symbol":"XIRP2","altered_samples":107,"total_mutation_events":143,"wt_cohort_size":393},{"hugo_gene_symbol":"FLG","altered_samples":105,"total_mutation_events":161,"wt_cohort_size":393},{"hugo_gene_symbol":"SPTA1","altered_samples":104,"total_mutation_events":143,"wt_cohort_size":393},{"hugo_gene_symbol":"COL11A1","altered_samples":89,"total_mutation_events":115,"wt_cohort_size":393},{"hugo_gene_symbol":"NAV3","altered_samples":88,"total_mutation_events":116,"wt_cohort_size":393},{"hugo_gene_symbol":"PCDH15","altered_samples":88,"total_mutation_events":132,"wt_cohort_size":393},{"hugo_gene_symbol":"FAT3","altered_samples":85,"total_mutation_events":123,"wt_cohort_size":393},{"hugo_gene_symbol":"MUC17","altered_samples":83,"total_mutation_events":118,"wt_cohort_size":393},{"hugo_gene_symbol":"ZNF536","altered_samples":79,"total_mutation_events":111,"wt_cohort_size":393},{"hugo_gene_symbol":"PCLO","altered_samples":77,"total_mutation_events":102,"wt_cohort_size":393},{"hugo_gene_symbol":"CSMD1","altered_samples":77,"total_mutation_events":106,"wt_cohort_size":393},{"hugo_gene_symbol":"TNR","altered_samples":76,"total_mutation_events":95,"wt_cohort_size":393},{"hugo_gene_symbol":"RYR3","altered_samples":75,"total_mutation_events":105,"wt_cohort_size":393},{"hugo_gene_symbol":"ANK2","altered_samples":73,"total_mutation_events":93,"wt_cohort_size":393},{"hugo_gene_symbol":"DNAH9","altered_samples":73,"total_mutation_events":94,"wt_cohort_size":393},{"hugo_gene_symbol":"ZNF804A","altered_samples":69,"total_mutation_events":89,"wt_cohort_size":393},{"hugo_gene_symbol":"APOB","altered_samples":69,"total_mutation_events":86,"wt_cohort_size":393},{"hugo_gene_symbol":"EGFR","altered_samples":69,"total_mutation_events":85,"wt_cohort_size":393},{"hugo_gene_symbol":"SI","altered_samples":69,"total_mutation_events":91,"wt_cohort_size":393},{"hugo_gene_symbol":"ADAMTS12","altered_samples":68,"total_mutation_events":88,"wt_cohort_size":393},{"hugo_gene_symbol":"ADGRG4","altered_samples":67,"total_mutation_events":78,"wt_cohort_size":393},{"hugo_gene_symbol":"CDH10","altered_samples":67,"total_mutation_events":83,"wt_cohort_size":393}]} ▶ clickhouse_run_select_query { "query": " -- Get combined mutation + CNA alterations in KRAS wild-type samples WITH kras_mutant AS ( SELECT DISTINCT sample_unique_id FROM genomic_event_derived WHERE cancer_study_identifier = 'luad_tcga_pan_can_atlas_2018' AND hugo_gene_symbol = 'KRAS' AND variant_type = 'mutation' AND mutation_status != 'UNCALLED' ), -- KRAS wild-type samples profiled for mutations wt_samples AS ( SELECT DISTINCT sample_unique_id FROM genomic_event_derived WHERE cancer_study_identifier = 'luad_tcga_pan_can_atlas_2018' AND variant_type = 'mutation' AND sample_unique_id NOT IN (SELECT sample_unique_id FROM kras_mutant) ), -- Combined alterations (mutations + CNA) combined_alterations AS ( SELECT hugo_gene_symbol, COUNT(DISTINCT sample_unique_id) as altered_samples, COUNT(*) as total_events, SUM(CASE WHEN variant_type = 'mutation' THEN 1 ELSE 0 END) as mutation_events, SUM(CASE WHEN variant_type = 'cna' THEN 1 ELSE 0 END) as cna_events FROM genomic_event_derived WHERE cancer_study_identifier = 'luad_tcga_pan_can_atlas_2018' AND sample_unique_id IN (SELECT sample_unique_id FROM wt_samples) AND ((variant_type = 'mutation' AND mutation_status != 'UNCALLED' AND off_panel = 0) OR (variant_type = 'cna' AND cna_alteration IN (2, -2))) GROUP BY hugo_gene_symbol ) SELECT hugo_gene_symbol, altered_samples, total_events, mutation_events, cna_events, ROUND(altered_samples * 100.0 / 393, 1) AS frequency_pct FROM combined_alterations ORDER BY altered_samples DESC LIMIT 30; ", "max_rows": 100 } ◀ result {"rows":[{"hugo_gene_symbol":"TP53","altered_samples":233,"total_events":251,"mutation_events":247,"cna_events":4,"frequency_pct":59.3},{"hugo_gene_symbol":"TTN","altered_samples":190,"total_events":504,"mutation_events":501,"cna_events":3,"frequency_pct":48.3},{"hugo_gene_symbol":"CSMD3","altered_samples":164,"total_events":287,"mutation_events":276,"cna_events":11,"frequency_pct":41.7},{"hugo_gene_symbol":"MUC16","altered_samples":159,"total_events":287,"mutation_events":284,"cna_events":3,"frequency_pct":40.5},{"hugo_gene_symbol":"RYR2","altered_samples":152,"total_events":299,"mutation_events":289,"cna_events":10,"frequency_pct":38.7},{"hugo_gene_symbol":"LRP1B","altered_samples":139,"total_events":241,"mutation_events":227,"cna_events":14,"frequency_pct":35.4},{"hugo_gene_symbol":"ZFHX4","altered_samples":133,"total_events":200,"mutation_events":185,"cna_events":15,"frequency_pct":33.8},{"hugo_gene_symbol":"FLG","altered_samples":129,"total_events":192,"mutation_events":161,"cna_events":31,"frequency_pct":32.8},{"hugo_gene_symbol":"USH2A","altered_samples":127,"total_events":210,"mutation_events":202,"cna_events":8,"frequency_pct":32.3},{"hugo_gene_symbol":"SPTA1","altered_samples":120,"total_events":163,"mutation_events":143,"cna_events":20,"frequency_pct":30.5},{"hugo_gene_symbol":"XIRP2","altered_samples":108,"total_events":145,"mutation_events":143,"cna_events":2,"frequency_pct":27.5},{"hugo_gene_symbol":"CSMD1","altered_samples":95,"total_events":131,"mutation_events":106,"cna_events":25,"frequency_pct":24.2},{"hugo_gene_symbol":"NAV3","altered_samples":94,"total_events":122,"mutation_events":116,"cna_events":6,"frequency_pct":23.9},{"hugo_gene_symbol":"COL11A1","altered_samples":94,"total_events":120,"mutation_events":115,"cna_events":5,"frequency_pct":23.9},{"hugo_gene_symbol":"ADAMTS12","altered_samples":93,"total_events":124,"mutation_events":88,"cna_events":36,"frequency_pct":23.7},{"hugo_gene_symbol":"PCDH15","altered_samples":91,"total_events":137,"mutation_events":132,"cna_events":5,"frequency_pct":23.2},{"hugo_gene_symbol":"ZNF536","altered_samples":91,"total_events":124,"mutation_events":111,"cna_events":13,"frequency_pct":23.2},{"hugo_gene_symbol":"PTPRD","altered_samples":91,"total_events":116,"mutation_events":84,"cna_events":32,"frequency_pct":23.2},{"hugo_gene_symbol":"FAT3","altered_samples":89,"total_events":128,"mutation_events":123,"cna_events":5,"frequency_pct":22.6},{"hugo_gene_symbol":"MUC17","altered_samples":88,"total_events":126,"mutation_events":118,"cna_events":8,"frequency_pct":22.4},{"hugo_gene_symbol":"CDH10","altered_samples":88,"total_events":116,"mutation_events":83,"cna_events":33,"frequency_pct":22.4},{"hugo_gene_symbol":"DNAH5","altered_samples":84,"total_events":109,"mutation_events":63,"cna_events":46,"frequency_pct":21.4},{"hugo_gene_symbol":"TNR","altered_samples":84,"total_events":104,"mutation_events":95,"cna_events":9,"frequency_pct":21.4},{"hugo_gene_symbol":"CDKN2A","altered_samples":84,"tota … (830 more chars) ▶ resolve_and_route { "studyIds": [ "luad_tcga_pan_can_atlas_2018" ] } ◀ result {"success":true,"message":"Found 1 study. Use the metadata to call the appropriate navigation tool(s).","data":{"totalCount":1,"studiesWithMetadata":[{"studyId":"luad_tcga_pan_can_atlas_2018","name":"Lung Adenocarcinoma (TCGA, PanCancer Atlas)","sampleCount":566,"studyViewUrl":"https://www.cbioportal.org/study?id=luad_tcga_pan_can_atlas_2018","metadata":{"clinicalAttributeIds":["AGE","AJCC_PATHOLOGIC_TUMOR_STAGE","AJCC_STAGING_EDITION","ANEUPLOIDY_SCORE","BUFFA_HYPOXIA_SCORE","CANCER_TYPE","CANCER_TYPE_ACRONYM","CANCER_TYPE_DETAILED","DAYS_LAST_FOLLOWUP","DAYS_TO_BIRTH","DAYS_TO_INITIAL_PATHOLOGIC_DIAGNOSIS","DFS_MONTHS","DFS_STATUS","DSS_MONTHS","DSS_STATUS","ETHNICITY","FORM_COMPLETION_DATE","FRACTION_GENOME_ALTERED","GENETIC_ANCESTRY_LABEL","GRADE","HISTORY_NEOADJUVANT_TRTYN","ICD_10","ICD_O_3_HISTOLOGY","ICD_O_3_SITE","INFORMED_CONSENT_VERIFIED","IN_PANCANPATHWAYS_FREEZE","MSI_SCORE_MANTIS","MSI_SENSOR_SCORE","MUTATION_COUNT","NEW_TUMOR_EVENT_AFTER_INITIAL_TREATMENT","ONCOTREE_CODE","OS_MONTHS","OS_STATUS","OTHER_PATIENT_ID","PATH_M_STAGE","PATH_N_STAGE","PATH_T_STAGE","PERSON_NEOPLASM_CANCER_STATUS","PFS_MONTHS","PFS_STATUS","PRIMARY_LYMPH_NODE_PRESENTATION_ASSESSMENT","PRIOR_DX","RACE","RADIATION_THERAPY","RAGNUM_HYPOXIA_SCORE","SAMPLE_COUNT","SAMPLE_TYPE","SEX","SOMATIC_STATUS","SUBTYPE","TBL_SCORE","TISSUE_PROSPECTIVE_COLLECTION_INDICATOR","TISSUE_RETROSPECTIVE_COLLECTION_INDICATOR","TISSUE_SOURCE_SITE","TISSUE_SOURCE_SITE_CODE","TMB_NONSYNONYMOUS","TUMOR_TISSUE_SITE","TUMOR_TYPE","WEIGHT","WINTER_HYPOXIA_SCORE"],"molecularProfileIds":["luad_tcga_pan_can_atlas_2018_gistic","luad_tcga_pan_can_atlas_2018_log2CNA","luad_tcga_pan_can_atlas_2018_mutations","luad_tcga_pan_can_atlas_2018_rna_seq_v2_mrna","luad_tcga_pan_can_atlas_2018_rna_seq_v2_mrna_median_Zscores","luad_tcga_pan_can_atlas_2018_rna_seq_v2_mrna_median_all_sample_Zscores","luad_tcga_pan_can_atlas_2018_rna_seq_v2_mrna_median_all_sample_ref_normal_Zscores","luad_tcga_pan_can_atlas_2018_rppa","luad_tcga_pan_can_atlas_2018_rppa_Zscores","luad_tcga_pan_can_atlas_2018_structural_variants"],"genericAssayProfiles":["luad_tcga_pan_can_atlas_2018_armlevel_cna","luad_tcga_pan_can_atlas_2018_genetic_ancestry","luad_tcga_pan_can_atlas_2018_methylation_hm27_hm450_merge","luad_tcga_pan_can_atlas_2018_methylation_hm450"],"heatmapProfileIds":["luad_tcga_pan_can_atlas_2018_rna_seq_v2_mrna_median_Zscores","luad_tcga_pan_can_atlas_2018_rna_seq_v2_mrna_median_all_sample_Zscores","luad_tcga_pan_can_atlas_2018_rna_seq_v2_mrna_median_all_sample_ref_normal_Zscores","luad_tcga_pan_can_atlas_2018_rppa_Zscores"],"availableComparisonTabs":["overlap","clinical","survival","alterations","mutations","mrna","protein","generic_assay_armlevel_cna","generic_assay_genetic_ancestry","generic_assay_methylation"]}}]}} ▶ get_studyviewfilter_options { "studyId": "luad_tcga_pan_can_atlas_2018", "geneSpecificQueries": [ { "hugoGeneSymbol": "KRAS", "profileType": "mutations" } ] } ◀ result {"success":true,"message":"Successfully retrieved StudyView filter options","data":{"studyId":"luad_tcga_pan_can_atlas_2018","geneSpecificCounts":[{"hugoGeneSymbol":"KRAS","profileType":"mutations","counts":[{"value":"Missense_Mutation","label":"Missense Mutation","count":171},{"value":"Nonsense_Mutation","label":"Nonsense Mutation","count":1}]}]}} ▶ navigate_to_study_view { "studyIds": [ "luad_tcga_pan_can_atlas_2018" ], "filterJson": { "mutationDataFilters": [ { "hugoGeneSymbol": "KRAS", "profileType": "mutations", "categorization": "MUTATED", "values": [ [ { "value": "NOT_MUTATED" } ] ] } ] } } ◀ result {"success":true,"message":"Navigating to https://www.cbioportal.org/study/summary?id=luad_tcga_pan_can_atlas_2018#filterJson=%7B%22mutationDataFilters%22%3A%5B%7B%22categorization%22%3A%22MUTATED%22%2C%22hugoGeneSymbol%22%3A%22KRAS%22%2C%22profileType%22%3A%22mutations%22%2C%22values%22%3A%5B%5B%7B%22value%22%3A%22NOT_MUTATED%22%7D%5D%5D%7D%5D%7D","url":"https://www.cbioportal.org/study/summary?id=luad_tcga_pan_can_atlas_2018#filterJson=%7B%22mutationDataFilters%22%3A%5B%7B%22categorization%22%3A%22MUTATED%22%2C%22hugoGeneSymbol%22%3A%22KRAS%22%2C%22profileType%22%3A%22mutations%22%2C%22values%22%3A%5B%5B%7B%22value%22%3A%22NOT_MUTATED%22%7D%5D%5D%7D%5D%7D","data":{"studyIds":["luad_tcga_pan_can_atlas_2018"],"studies":[{"studyId":"luad_tcga_pan_can_atlas_2018","name":"Lung Adenocarcinoma (TCGA, PanCancer Atlas)","sampleCount":566}],"tab":"summary","hasFilters":true,"hasPlotsConfig":false,"pageDescription":"Overview of all study data displayed as interactive charts summarizing available clinical and genomic data. Click within a chart to filter. Additional charts may be available in the Charts menu."}} ═ answer (success) Perfect! Now let me present the results comprehensively. ## Most Frequently Altered Genes in KRAS Wild-Type Lung Adenocarcinoma I analyzed **393 KRAS wild-type samples** from the TCGA Lung Adenocarcinoma PanCancer Atlas study (566 total samples; 168 KRAS-mutant samples excluded). ### Top 20 Most Frequently Altered Genes (Mutation + CNA) | Gene | Altered Samples | Frequency | Alterations | |------|-----------------|-----------|------------| | **TP53** | 233 | 59.3% | 247 mutations, 4 CNAs | | **TTN** | 190 | 48.3% | 501 mutations, 3 CNAs | | **CSMD3** | 164 | 41.7% | 276 mutations, 11 CNAs | | **MUC16** | 159 | 40.5% | 284 mutations, 3 CNAs | | **RYR2** | 152 | 38.7% | 289 mutations, 10 CNAs | | **LRP1B** | 139 | 35.4% | 227 mutations, 14 CNAs | | **ZFHX4** | 133 | 33.8% | 185 mutations, 15 CNAs | | **FLG** | 129 | 32.8% | 161 mutations, 31 CNAs | | **USH2A** | 127 | 32.3% | 202 mutations, 8 CNAs | | **SPTA1** | 120 | 30.5% | 143 mutations, 20 CNAs | | **XIRP2** | 108 | 27.5% | 143 mutations, 2 CNAs | | **CSMD1** | 95 | 24.2% | 106 mutations, 25 CNAs | | **NAV3** | 94 | 23.9% | 116 mutations, 6 CNAs | | **COL11A1** | 94 | 23.9% | 115 mutations, 5 CNAs | | **ADAMTS12** | 93 | 23.7% | 88 mutations, 36 CNAs | | **PCDH15** | 91 | 23.2% | 132 mutations, 5 CNAs | | **ZNF536** | 91 | 23.2% | 111 mutations, 13 CNAs | | **PTPRD** | 91 | 23.2% | 84 mutations, 32 CNAs | | **FAT3** | 89 | 22.6% | 123 mutations, 5 CNAs | | **MUC17** | 88 | 22.4% | 118 mutations, 8 CNAs | ### Key Observations 1. **TP53 dominates** at 59.3% frequency — the most consistently altered gene in KRAS wild-type adenocarcinomas. 2. **TTN is also highly frequent** (48.3%), typical of lung adenocarcinoma where large genes accumulate passenger mutations due to the hypermutated phenotype. 3. **Cancer-driver genes enriched:** Beyond TP53, clinically relevant genes include: - **EGFR** (69 samples, 17.6% frequency among all mutations) — key driver in KRAS-WT adenocarcinomas - **CDKN2A** (84 samples, 21.4%) — predominantly copy-number deletions (64 of 87 events) 4. **Many genes are large/genomic passengers** (TTN, MUC16, CSMD3, RYR2, LRP1B, etc.), as expected in a hypermutated cancer type. 5. **Copy-number alterations** are less common than mutations overall but contribute significantly for tumor suppressors like CDKN2A (64 homozygous deletions out of 87 total alterations). ### Explore the Data [View KRAS wild-type LUAD cohort in StudyView](https://www.cbioportal.org/study/summary?id=luad_tcga_pan_can_atlas_2018#filterJson=%7B%22mutationDataFilters%22%3A%5B%7B%22categorization%22%3A%22MUTATED%22%2C%22hugoGeneSymbol%22%3A%22KRAS%22%2C%22profileType%22%3A%22mutations%22%2C%22values%22%3A%5B%5B%7B%22value%22%3A%22NOT_MUTATED%22%7D%5D%5D%7D%5D%7D) This link opens the TCGA LUAD study pre-filtered to KRAS wild-type samples, where you can explore mutation and copy-number distributions interactively.