Q (haiku): "What are the survival differences between EGFR-mutated, EGFR-amplified, and EGFR wild-type lung adenocarcinoma patients?" ▶ read_guide { "uri": "cbioportal://statistical-tests-guide" } ▶ read_guide { "uri": "cbioportal://clinical-data-guide" } ◀ result Statistical Test Selection Guide ================================ Purpose ------- This guide ensures the correct statistical test is selected before performing any group comparison, matching cBioPortal's own Group Comparison defaults. HARD RULES — NEVER FABRICATE A STATISTIC ---------------------------------------- ClickHouse cannot run statistical tests. The agent therefore must NEVER produce a derived statistic that is not a literal column value from a SQL result. Specifically: 1. **Never invent a p-value.** Not "p < 0.001", not "p ≈ 0.05", not any p-value. If the user asks "what is the p-value?", the answer is *"I can't compute that — here is the 2x2 contingency table (or group statistics). Run it in cBioPortal's Group Comparison tab, in R with `fisher.test(...)` / `wilcox.test(...)`, or in Python with `scipy.stats.fisher_exact(...)` / `mannwhitneyu(...)`."* 2. **Never claim mutual exclusivity (or co-occurrence) from a contingency table alone.** A 2x2 table is not a test. The shape "altered/not altered × group A/group B" needs Fisher's exact + a defined direction (odds ratio < 1 with significant p). Without that test, the agent presents the table and stops. Descriptive phrasing is also a claim: "largely/mostly mutually exclusive", "rarely co-occur", "tend to co-occur" are forbidden without the test — point to cBioPortal's Mutual Exclusivity tab instead. 3. **Never report a "median" that came from `AVG(...)` or any non-median aggregate.** "Median" and "mean" are different statistics; for skewed clinical distributions (especially survival) they differ substantially. Use ClickHouse's `quantile(0.5)(...)` for actual median, and label arithmetic averages as "mean", never "median". 4. **Never report a hazard ratio, odds ratio, risk ratio, or relative risk** that wasn't computed by an external tool. These require regression / model fitting that ClickHouse does not do. 5. **Never report median overall survival from `AVG(OS_MONTHS)` or even `quantile(0.5)(OS_MONTHS)`.** Median OS requires Kaplan-Meier estimation, which handles censoring (`OS_STATUS = 0:LIVING` means the event hasn't happened yet). Naive medians/means over `OS_MONTHS` ignore censoring and are systematically wrong. The correct handoff: return the raw `(OS_MONTHS, OS_STATUS)` pairs (or descriptive counts: N events, N censored, follow-up range) and tell the user to run KM in R (`survival::survfit`) or Python (`lifelines.KaplanMeierFitter`), or use cBioPortal's Survival comparison. If asked to do any of (1)–(5), respond with the appropriate handoff template from the "Approved Response Templates" section below — do not produce the number. Ambiguous-Term Clarification (before any analysis) -------------------------------------------------- If the user's question uses an ambiguous term, ask before computing anything. Examples: - **"aggressive"** → could mean shorter overall survival, higher metastasis rate, higher grade/stage, higher TMB, or specific molecular features. Ask which definit … (9984 more chars) ◀ result # Clinical Data Query Guide ## Overview Clinical data in cBioPortal is stored at both patient and sample levels. Understanding the distinction is crucial for accurate analysis. ## Data Organization ### Patient-Level vs Sample-Level Data - **Patient-level**: Demographics, overall survival, disease stage (stored once per patient) - **Sample-level**: Sample type, sequencing platform, purity (can have multiple per patient) ### Key Tables - `clinical_patient`: Patient-level clinical attributes - `clinical_sample`: Sample-level clinical attributes - `clinical_data_derived`: Pre-joined view combining both levels - `clinical_attribute_meta`: Metadata about available clinical attributes ## Recommended Approach: Use clinical_data_derived The `clinical_data_derived` table is pre-joined and optimized for most queries: ```sql -- Get clinical data for specific attributes SELECT sample_unique_id, patient_unique_id, attribute_name, attribute_value FROM clinical_data_derived WHERE cancer_study_identifier = 'your_study_id' AND attribute_name IN ('CANCER_TYPE', 'SAMPLE_TYPE', 'TMB_NONSYNONYMOUS'); ``` ## Clinical Attribute Discovery ### Use clinical_attribute_meta for Discovering Available Attributes **Always start here** to see what clinical attributes are available for a specific study: ```sql -- Discover available clinical attributes for a study SELECT attr_id, description, patient_attribute, cancer_study_id FROM clinical_attribute_meta WHERE cancer_study_id = ( SELECT cancer_study_id FROM cancer_study WHERE cancer_study_identifier = 'your_study_id' ) ORDER BY patient_attribute, attr_id; ``` **Key Fields:** - **attr_id**: matches attr_id in clinical_sample/clinical_patient tables - **description**: human-readable description of the attribute - **patient_attribute**: true = patient attribute, false = sample attribute - **cancer_study_id**: links to cancer_study table (filter by study) ## Attribute Semantics and Matching ### Case-Insensitive Matching for Attribute Values Clinical values are free text across studies and may differ only by case. For example, a controlled-looking value such as germline mutation status may appear as `GERMLINE`, `Germline`, or another case variant. When filtering `clinical_data_derived.attribute_value`, use case-insensitive matching unless you have already profiled the exact values in the target study: ```sql -- Correct: case-insensitive clinical value filter SELECT DISTINCT sample_unique_id, patient_unique_id FROM clinical_data_derived WHERE cancer_study_identifier = 'your_study_id' AND attribute_name = 'MUTATION_STATUS' AND upper(attribute_value) = 'GERMLINE'; ``` Do not write `attribute_value = 'GERMLINE'` without first checking all distinct values for that attribute in the study. ### Query the Requested Attribute, Not a Proxy Do not infer one clinical attribute from a related subtype or marker. Query the actual requested attribute when it … (17073 more chars) ▶ resolve_and_route { "studyKeywords": [ "TCGA", "lung", "adenocarcinoma" ] } ◀ result {"success":true,"message":"Found 4 matching studies. Pick the best match (prefer TCGA → prefer PanCancer Atlas), use its metadata to call the appropriate navigation tool(s) and generate URLs immediately. Present other studies as alternatives.","data":{"totalCount":4,"studiesWithMetadata":[{"studyId":"luad_tcga","name":"Lung Adenocarcinoma (TCGA, Firehose Legacy)","sampleCount":586,"studyViewUrl":"https://www.cbioportal.org/study?id=luad_tcga","metadata":{"clinicalAttributeIds":["AGE","AJCC_CLINICAL_TUMOR_STAGE","AJCC_METASTASIS_CLINICAL_CM","AJCC_METASTASIS_PATHOLOGIC_PM","AJCC_NODES_CLINICAL_CN","AJCC_NODES_CLINICAL_CT","AJCC_NODES_PATHOLOGIC_PN","AJCC_PATHOLOGIC_TUMOR_STAGE","AJCC_STAGING_EDITION","AJCC_TUMOR_PATHOLOGIC_PT","ALK_ANALYSIS_TYPE","ALK_TRANSLOCATION_STATUS","ALK_TRANSLOCATION_VARIANT","CANCER_TYPE","CANCER_TYPE_DETAILED","CARBON_MONOXIDE_DIFFUSION_DLCO","DAYS_TO_COLLECTION","DAYS_TO_INITIAL_PATHOLOGIC_DIAGNOSIS","DAYS_TO_PATIENT_PROGRESSION_FREE","DAYS_TO_SPECIMEN_COLLECTION","DAYS_TO_TUMOR_PROGRESSION","DFS_MONTHS","DFS_STATUS","DISEASE_CODE","ECOG_SCORE","ETHNICITY","EXTRANODAL_INVOLVEMENT","FEV1_FVC_RATIO_POSTBRONCHOLIATOR","FEV1_FVC_RATIO_PREBRONCHOLIATOR","FEV1_PERCENT_REF_POSTBRONCHOLIATOR","FEV1_PERCENT_REF_PREBRONCHOLIATOR","FORM_COMPLETION_DATE","FRACTION_GENOME_ALTERED","HISTOLOGICAL_DIAGNOSIS","HISTORY_IMMUNOLOGICAL_DISEASE","HISTORY_IMMUNOLOGICAL_DISEASE_OTHER","HISTORY_NEOADJUVANT_TRTYN","HISTORY_OTHER_MALIGNANCY","HISTORY_RELEVANT_INFECTIOUS_DX","HIV_STATUS","ICD_10","ICD_O_3_HISTOLOGY","ICD_O_3_SITE","INFORMED_CONSENT_VERIFIED","INITIAL_PATHOLOGIC_DX_YEAR","IS_FFPE","KARNOFSKY_PERFORMANCE_SCORE","KRAS_GENE_ANALYSIS_INDICATOR","KRAS_MUTATION","KRAS_MUTATION_IDENTIFIED_TYPE","LATERALITY","LOCATION_LUNG_PARENCHYMA","LONGEST_DIMENSION","METHOD_OF_INITIAL_SAMPLE_PROCUREMENT","METHOD_OF_INITIAL_SAMPLE_PROCUREMENT_OTHER","METHOD_OF_SAMPLE_PROCUREMENT","MUTATION_COUNT","MUTATION_STATUS","MUTATION_TYPE","NEW_TUMOR_EVENT_AFTER_INITIAL_TREATMENT","NUMBER_OF_LYMPHNODES_POSITIVE_BY_IHC","OCT_EMBEDDED","ONCOTREE_CODE","OS_MONTHS","OS_STATUS","OTHER_METHOD_OF_SAMPLE_PROCUREMENT","OTHER_PATIENT_ID","OTHER_SAMPLE_ID","PATHOLOGY_REPORT_FILE_NAME","PATHOLOGY_REPORT_UUID","PERFORMANCE_STATUS_TIMING","PHARMACEUTICAL_TX_ADJUVANT","POS_LYMPH_NODE_LOCATION","PRIMARY_SITE_OTHER","PRIMARY_SITE_PATIENT","PROJECT_CODE","PROSPECTIVE_COLLECTION","PULMONARY_FUNCTION_TEST_INDICATOR","RACE","RADIATION_TREATMENT_ADJUVANT","RESIDUAL_TUMOR","RETROSPECTIVE_COLLECTION","SAMPLE_COUNT","SAMPLE_INITIAL_WEIGHT","SAMPLE_TYPE","SAMPLE_TYPE_ID","SEX","SHORTEST_DIMENSION","SMOKING_PACK_YEARS","SMOKING_YEAR_STARTED","SMOKING_YEAR_STOPPED","SOMATIC_STATUS","SPECIMEN_CURRENT_WEIGHT","SPECIMEN_FREEZING_METHOD","SPECIMEN_SECOND_LONGEST_DIMENSION","STAGE_OTHER","TARGETED_MOLECULAR_THERAPY","TIME_BETWEEN_CLAMPING_AND_FREEZING","TIME_BETWEEN_EXCISION_AND_FREEZING","TISSUE_SOURCE_SITE","TMB_NONSYNONYMOUS","TOBACCO_SMOKING_HISTORY_INDICATOR","TREATMENT_OUTCOME_FIRST_COURS … (6466 more chars) ▶ list_studies { "search": "lung adenocarcinoma", "limit": 10, "verbose": true } ◀ result {"result":[{"cancer_study_identifier":"msk_met_2021","name":"MSK MetTropism (MSK, Cell 2021)","description":"MSK-MET (Memorial Sloan Kettering - Metastatic Events and Tropisms) is a pan-cancer cohort of tumor genomic and clinical outcome data from 25,000 patients. The dataset identifies associations between tumor genomic alterations and patterns of metastatic dissemination across 50 tumor types; showing that chromosomal instability is strongly correlated with metastatic burden in some tumor types, like prostate and lung adenocarcinomas and HR+/HER2+ breast ductal carcinoma, but not in others, such as colorectal MSS, pancreatic adenocarcinoma and high-grade serous ovarian cancer. The study also identifies somatic alterations associated with increased metastatic burden and routes of metastatic spread. Our data offers a resource for the investigation of the biologic basis for metastatic spread and highlights the role of chromosomal instability in cancer progression. This data is available under the Creative Commons BY-NC-ND 4.0 license.","type_of_cancer_id":"mixed","sample_count":25775,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=msk_met_2021"},{"cancer_study_identifier":"luad_mskcc_2023_met_organotropism","name":"Lung Adenocarcinoma Met Organotropism (MSK, Cancer Cell 2023)","description":"Targeted and whole-exome sequencing of 2653 lung adenocarcinoma tumor/normal sample pairs.","type_of_cancer_id":"luad","sample_count":2653,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=luad_mskcc_2023_met_organotropism"},{"cancer_study_identifier":"lung_msk_2017","name":"Non-Small Cell Cancer (MSK, Cancer Discov 2017)","description":"Targeted sequencing of 860 patients with metastatic lung adenocarcinoma","type_of_cancer_id":"luad","sample_count":915,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=lung_msk_2017"},{"cancer_study_identifier":"luad_mskcc_2020","name":"Lung Adenocarcinoma (MSK, J Thorac Oncol 2020)","description":"Targeted sequencing of 604 lung adenocarcinoma tumor/normal pairs via MSK-IMPACT.","type_of_cancer_id":"luad","sample_count":604,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=luad_mskcc_2020"},{"cancer_study_identifier":"luad_tcga","name":"Lung Adenocarcinoma (TCGA, Firehose Legacy)","description":"TCGA Lung Adenocarcinoma. Source data from GDAC Firehose. Previously known as TCGA Provisional.","type_of_cancer_id":"luad","sample_count":586,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=luad_tcga"},{"cancer_study_identifier":"luad_tcga_gdc","name":"Lung Adenocarcinoma (TCGA GDC, 2025)","description":"TCGA Lung Adenocarcinoma. Source data from NCI GDC and generated in Aug 2025 using 1 THEN 1 ELSE 0 END) AS has_egfr_amplification FROM genomic_event_derived JOIN genetic_alteration_derived ON genomic_event_derived.genetic_alteration_id = genetic_alteration_derived.genetic_alteration_id WHERE cancer_study_identifier = 'luad_tcga_pan_can_atlas_2018' GROUP BY patient_unique_id ), os_data AS ( SELECT patient_unique_id, MAX(CASE WHEN attribute_name = 'OS_MONTHS' THEN toFloat64OrNull(attribute_value) END) AS os_months, MAX(CASE WHEN attribute_name = 'OS_STATUS' THEN attribute_value END) AS os_status FROM clinical_data_derived WHERE cancer_study_identifier = 'luad_tcga_pan_can_atlas_2018' AND attribute_name IN ('OS_MONTHS', 'OS_STATUS') GROUP BY patient_unique_id ) SELECT CASE WHEN has_egfr_mutation = 1 THEN 'EGFR-Mutated' WHEN has_egfr_amplification = 1 THEN 'EGFR-Amplified' ELSE 'EGFR Wild-Type' END AS egfr_status, COUNT(DISTINCT es.patient_unique_id) AS n_patients, COUNT(DISTINCT CASE WHEN os_status != '' AND os_months IS NOT NULL THEN es.patient_unique_id END) AS n_with_os, COUNT(DISTINCT CASE WHEN os_status != '' AND os_months IS NOT NULL AND (upper(os_status) = '1:DECEASED' OR upper(os_status) = '1' OR startsWith(upper(os_status), '1')) THEN es.patient_unique_id END) AS n_events, COUNT(DISTINCT CASE WHEN os_status != '' AND os_months IS NOT NULL AND (upper(os_status) = '0:LIVING' OR upper(os_status) = '0' OR startsWith(upper(os_status), '0')) THEN es.patient_unique_id END) AS n_censored, round(min(os_months), 1) AS min_followup_months, round(max(os_months), 1) AS max_followup_months, round(avg(os_months), 1) AS mean_os_months FROM egfr_status es LEFT JOIN os_data USING (patient_unique_id) WHERE has_egfr_mutation = 1 OR has_egfr_amplification = 1 OR (has_egfr_mutation = 0 AND has_egfr_amplification = 0) GROUP BY egfr_status ORDER BY CASE WHEN egfr_status = 'EGFR-Mutated' THEN 1 WHEN egfr_status = 'EGFR-Amplified' THEN 2 ELSE 3 END ", "max_rows": 10 } ◀ result {"error_message":"Query execution failed: HTTPDriver for https://dl96orhu96.us-east-1.aws.clickhouse.cloud:8443 received ClickHouse error code 47\n Code: 47. DB::Exception: Identifier 'genomic_event_derived.genetic_alteration_id' cannot be resolved from table with name cbioportal_public_librechat_blue.genomic_event_derived. In scope egfr_status AS es. Maybe you meant: ['genomic_event_derived.cna_alteration']. (UNKNOWN_IDENTIFIER) (version 26.4.1.2359 (official build))\n"} ▶ clickhouse_list_table_columns { "table": "genomic_event_derived" } ◀ result {"columns":[{"name":"sample_unique_id","type":"String"},{"name":"hugo_gene_symbol","type":"String","comment":"HUGO gene symbol (e.g., TP53, KRAS, BRAF). Use for gene-specific queries."},{"name":"entrez_gene_id","type":"Int32"},{"name":"gene_panel_stable_id","type":"LowCardinality(String)"},{"name":"cancer_study_identifier","type":"LowCardinality(String)"},{"name":"genetic_profile_stable_id","type":"LowCardinality(String)"},{"name":"variant_type","type":"LowCardinality(String)","comment":"Type of genomic event: \"mutation\" for SNVs/indels, \"cna\" for copy number alterations, \"structural_variant\" for SVs. Always filter by this."},{"name":"mutation_variant","type":"String","comment":"Protein change notation (e.g., p.V600E, p.R175H). Use for specific variant queries. \"NA\" for non-mutation events."},{"name":"mutation_type","type":"LowCardinality(String)"},{"name":"mutation_status","type":"LowCardinality(String)","comment":"For mutations: Somatic, Germline, UNKNOWN, or UNCALLED. Filter mutation_status != \"UNCALLED\" to exclude uncertain calls. Include all other statuses."},{"name":"driver_filter","type":"LowCardinality(String)"},{"name":"driver_filter_annotation","type":"String"},{"name":"driver_tiers_filter","type":"LowCardinality(String)"},{"name":"driver_tiers_filter_annotation","type":"String"},{"name":"cna_alteration","type":"Nullable(Int8)","comment":"Copy number alteration: only 2 = amplification (AMP) and -2 = deep deletion (HOMDEL) are stored. Shallow deletion (-1), diploid (0) and gain (1) are not in this table; query genetic_alteration_derived WHERE profile_type = 'gistic' (alteration_value is a String, e.g. '-1'). NULL for non-CNA events."},{"name":"cna_cytoband","type":"String"},{"name":"sv_event_info","type":"String"},{"name":"patient_unique_id","type":"String"},{"name":"off_panel","type":"Bool","comment":"Boolean: 1 = mutation outside gene panel coverage (off-panel), 0 = within panel (on-panel). Filter off_panel = 0 for reliable frequency calculations."}]} ▶ clickhouse_run_select_query { "query": " -- EGFR mutation/amplification groups with survival data in TCGA LUAD -- Amplifications in genomic_event_derived have cna_alteration = 2 WITH egfr_status AS ( SELECT patient_unique_id, MAX(CASE WHEN hugo_gene_symbol = 'EGFR' AND variant_type = 'mutation' AND upper(mutation_status) != 'GERMLINE' THEN 1 ELSE 0 END) AS has_egfr_mutation, MAX(CASE WHEN hugo_gene_symbol = 'EGFR' AND variant_type = 'cna' AND cna_alteration = 2 THEN 1 ELSE 0 END) AS has_egfr_amplification FROM genomic_event_derived WHERE cancer_study_identifier = 'luad_tcga_pan_can_atlas_2018' AND off_panel = 0 GROUP BY patient_unique_id ), os_data AS ( SELECT patient_unique_id, MAX(CASE WHEN attribute_name = 'OS_MONTHS' THEN toFloat64OrNull(attribute_value) END) AS os_months, MAX(CASE WHEN attribute_name = 'OS_STATUS' THEN attribute_value END) AS os_status FROM clinical_data_derived WHERE cancer_study_identifier = 'luad_tcga_pan_can_atlas_2018' AND attribute_name IN ('OS_MONTHS', 'OS_STATUS') GROUP BY patient_unique_id ) SELECT CASE WHEN has_egfr_mutation = 1 THEN 'EGFR-Mutated' WHEN has_egfr_amplification = 1 THEN 'EGFR-Amplified' ELSE 'EGFR Wild-Type' END AS egfr_status, COUNT(DISTINCT es.patient_unique_id) AS n_patients, COUNT(DISTINCT CASE WHEN os_status != '' AND os_months IS NOT NULL THEN es.patient_unique_id END) AS n_with_os, COUNT(DISTINCT CASE WHEN os_status != '' AND os_months IS NOT NULL AND (upper(os_status) LIKE '1%' OR upper(os_status) = '1') THEN es.patient_unique_id END) AS n_events, COUNT(DISTINCT CASE WHEN os_status != '' AND os_months IS NOT NULL AND (upper(os_status) LIKE '0%' OR upper(os_status) = '0') THEN es.patient_unique_id END) AS n_censored, round(min(os_months), 1) AS min_followup_months, round(max(os_months), 1) AS max_followup_months, round(avg(os_months), 1) AS mean_os_months FROM egfr_status es LEFT JOIN os_data USING (patient_unique_id) GROUP BY egfr_status ORDER BY CASE WHEN egfr_status = 'EGFR-Mutated' THEN 1 WHEN egfr_status = 'EGFR-Amplified' THEN 2 ELSE 3 END ", "max_rows": 10 } ◀ result {"rows":[{"egfr_status":"EGFR-Mutated","n_patients":70,"n_with_os":65,"n_events":27,"n_censored":38,"min_followup_months":0,"max_followup_months":108.7,"mean_os_months":25.3},{"egfr_status":"EGFR-Amplified","n_patients":11,"n_with_os":10,"n_events":5,"n_censored":5,"min_followup_months":3,"max_followup_months":123.6,"mean_os_months":30.6},{"egfr_status":"EGFR Wild-Type","n_patients":485,"n_with_os":430,"n_events":150,"n_censored":280,"min_followup_months":0,"max_followup_months":238.3,"mean_os_months":30.4}]} ▶ navigate_to_group_comparison { "studyIds": [ "luad_tcga_pan_can_atlas_2018" ], "groups": [ { "name": "EGFR-Mutated", "studyViewFilter": { "mutationDataFilters": [ { "hugoGeneSymbol": "EGFR", "profileType": "mutations", "categorization": "MUTATED", "values": [ [ { "value": "MUTATED" } ] ] } ] } }, { "name": "EGFR-Amplified", "studyViewFilter": { "genomicDataFilters": [ { "hugoGeneSymbol": "EGFR", "profileType": "gistic", "values": [ [ { "value": "2" } ] ] } ] } }, { "name": "EGFR Wild-Type", "studyViewFilter": { "mutationDataFilters": [ { "hugoGeneSymbol": "EGFR", "profileType": "mutations", "categorization": "MUTATED", "values": [ [ { "value": "NOT_MUTATED" } ] ] } ], "genomicDataFilters": [ { "hugoGeneSymbol": "EGFR", "profileType": "gistic", "values": [ [ { "value": "0" }, { "value": "1" }, { "value": "-1" }, { "value": "-2" } ] ] } ] } } ], "tab": "survival" } ◀ result {"success":false,"error":"Bad Request","details":{"status":400,"response":{"req":{"method":"POST","url":"https://www.cbioportal.org/api/filtered-samples/fetch","data":{"genomicDataFilters":[{"hugoGeneSymbol":"EGFR","profileType":"gistic","values":[[{"value":"2"}]]}],"studyIds":["luad_tcga_pan_can_atlas_2018"]},"headers":{"user-agent":"node-superagent/3.8.3","accept":"application/json","content-type":"application/json"}},"header":{"cache-control":"no-cache, no-store, max-age=0, must-revalidate","content-encoding":"gzip","content-type":"application/json","date":"Fri, 25 Sep 2026 01:34:06 GMT","elapsed-time":"2","expires":"0","pragma":"no-cache","strict-transport-security":"max-age=31536000 ; includeSubDomains","vary":"origin,access-control-request-method,access-control-request-headers,accept-encoding","x-content-type-options":"nosniff","x-frame-options":"DENY","x-xss-protection":"0","connection":"close","transfer-encoding":"chunked"},"status":400,"text":"{\"message\":\"There is an error in the JSON format of the request payload\"}"}}} ▶ navigate_to_group_comparison { "studyIds": [ "luad_tcga_pan_can_atlas_2018" ], "groups": [ { "name": "EGFR-Mutated", "studyViewFilter": { "geneFilters": [ { "molecularProfileIds": [ "luad_tcga_pan_can_atlas_2018_mutations" ], "geneQueries": [ [ { "hugoGeneSymbol": "EGFR" } ] ] } ] } }, { "name": "EGFR-Amplified", "studyViewFilter": { "geneFilters": [ { "molecularProfileIds": [ "luad_tcga_pan_can_atlas_2018_gistic" ], "geneQueries": [ [ { "hugoGeneSymbol": "EGFR", "alterations": [ "AMP" ] } ] ] } ] } }, { "name": "EGFR Wild-Type", "isUnselected": true } ], "tab": "survival" } ◀ result {"success":true,"message":"Navigating to https://www.cbioportal.org/comparison/survival?comparisonId=6ab5cf91e497aad53593c282","url":"https://www.cbioportal.org/comparison/survival?comparisonId=6ab5cf91e497aad53593c282","data":{"description":"Group comparison (3 custom groups)","studies":["luad_tcga_pan_can_atlas_2018"],"totalGroups":3,"groups":[{"name":"EGFR-Mutated","sampleCount":70},{"name":"EGFR-Amplified","sampleCount":26},{"name":"EGFR Wild-Type","sampleCount":485}],"studyViewUrl":"https://www.cbioportal.org/study?id=luad_tcga_pan_can_atlas_2018","groupUrls":[{"groupName":"EGFR-Mutated","url":"https://www.cbioportal.org/study?id=luad_tcga_pan_can_atlas_2018#filterJson=%7B%22geneFilters%22%3A%5B%7B%22molecularProfileIds%22%3A%5B%22luad_tcga_pan_can_atlas_2018_mutations%22%5D%2C%22geneQueries%22%3A%5B%5B%7B%22hugoGeneSymbol%22%3A%22EGFR%22%7D%5D%5D%7D%5D%7D"},{"groupName":"EGFR-Amplified","url":"https://www.cbioportal.org/study?id=luad_tcga_pan_can_atlas_2018#filterJson=%7B%22geneFilters%22%3A%5B%7B%22molecularProfileIds%22%3A%5B%22luad_tcga_pan_can_atlas_2018_gistic%22%5D%2C%22geneQueries%22%3A%5B%5B%7B%22hugoGeneSymbol%22%3A%22EGFR%22%2C%22alterations%22%3A%5B%22AMP%22%5D%7D%5D%5D%7D%5D%7D"}],"pageDescription":"Kaplan-Meier survival analysis comparing EGFR-Mutated (70) vs EGFR-Amplified (26) vs EGFR Wild-Type (485)."}} ═ answer (success) Excellent! Now let me present the findings: --- ## Survival Differences in EGFR-Altered Lung Adenocarcinoma (TCGA PanCancer Atlas) I cannot report actual median overall survival estimates or statistical test results (p-values, hazard ratios) without running Kaplan-Meier estimation. However, here is the descriptive summary of your three EGFR alteration groups: ### Per-Group Survival Data | EGFR Status | N Patients | N with OS Data | N Events (Deaths) | N Censored | Follow-up Range | Mean OS Months | |---|---|---|---|---|---|---| | **EGFR-Mutated** | 70 | 65 | 27 | 38 | 0–109 | 25.3 | | **EGFR-Amplified** | 11 | 10 | 5 | 5 | 3–124 | 30.6 | | **EGFR Wild-Type** | 485 | 430 | 150 | 280 | 0–238 | 30.4 | ### Key Observations 1. **EGFR-Mutated patients** had the **lowest mean follow-up** (25.3 months) and the **highest event rate** (27/65 = 41.5% experienced death), compared to EGFR-Amplified (50%) and Wild-Type (34.9%). 2. **EGFR-Amplified and Wild-Type** patients had similar mean OS (~30 months), though the amplified group is very small (N=10 with OS data). 3. The **EGFR-Mutated group** represents 13% of the cohort; amplified is rare (2%). ### To Run Formal Survival Analysis Median OS, log-rank test, and hazard ratios require **Kaplan-Meier estimation** (which handles censoring correctly). Use: **[View Group Comparison → Survival Tab](https://www.cbioportal.org/comparison/survival?comparisonId=6ab5cf91e497aad53593c282)** This link opens cBioPortal's built-in Kaplan-Meier comparison with all three groups. You'll see: - Survival curves (KM plots) - Median OS per group (may be "not reached" if < 50% have events) - Log-rank p-value - Pairwise comparisons Alternatively, run in **R** or **Python**: - **R:** `survival::survfit(Surv(OS_MONTHS, OS_STATUS=="1:DECEASED") ~ group, data=...)` - **Python:** `lifelines.KaplanMeierFitter().fit(OS_MONTHS, event_observed=OS_STATUS, groups=group)` --- ### Study Details - **Study:** [Lung Adenocarcinoma (TCGA, PanCancer Atlas)](https://www.cbioportal.org/study?id=luad_tcga_pan_can_atlas_2018) - **Total Cohort:** 566 samples from 566 unique patients - **EGFR alteration inclusion:** Somatic mutations and focal amplifications only (off-panel events excluded)