Q (sonnet): In the “Pan-cancer analysis of whole genomes” study what is the survival difference and the corresponding statistical significance between patients with mutations in both TP53 and KRAS versus patients with only a KRAS mutation? ▶ read_guide { "uri": "cbioportal://statistical-tests-guide" } ▶ list_studies { "search": "Pan-cancer analysis of whole genomes" } ◀ result Statistical Test Selection Guide ================================ Purpose ------- This guide ensures the correct statistical test is selected before performing any group comparison, matching cBioPortal's own Group Comparison defaults. HARD RULES — NEVER FABRICATE A STATISTIC ---------------------------------------- ClickHouse cannot run statistical tests. The agent therefore must NEVER produce a derived statistic that is not a literal column value from a SQL result. Specifically: 1. **Never invent a p-value.** Not "p < 0.001", not "p ≈ 0.05", not any p-value. If the user asks "what is the p-value?", the answer is *"I can't compute that — here is the 2x2 contingency table (or group statistics). Run it in cBioPortal's Group Comparison tab, in R with `fisher.test(...)` / `wilcox.test(...)`, or in Python with `scipy.stats.fisher_exact(...)` / `mannwhitneyu(...)`."* 2. **Never claim mutual exclusivity (or co-occurrence) from a contingency table alone.** A 2x2 table is not a test. The shape "altered/not altered × group A/group B" needs Fisher's exact + a defined direction (odds ratio < 1 with significant p). Without that test, the agent presents the table and stops. Descriptive phrasing is also a claim: "largely/mostly mutually exclusive", "rarely co-occur", "tend to co-occur" are forbidden without the test — point to cBioPortal's Mutual Exclusivity tab instead. 3. **Never report a "median" that came from `AVG(...)` or any non-median aggregate.** "Median" and "mean" are different statistics; for skewed clinical distributions (especially survival) they differ substantially. Use ClickHouse's `quantile(0.5)(...)` for actual median, and label arithmetic averages as "mean", never "median". 4. **Never report a hazard ratio, odds ratio, risk ratio, or relative risk** that wasn't computed by an external tool. These require regression / model fitting that ClickHouse does not do. 5. **Never report median overall survival from `AVG(OS_MONTHS)` or even `quantile(0.5)(OS_MONTHS)`.** Median OS requires Kaplan-Meier estimation, which handles censoring (`OS_STATUS = 0:LIVING` means the event hasn't happened yet). Naive medians/means over `OS_MONTHS` ignore censoring and are systematically wrong. The correct handoff: return the raw `(OS_MONTHS, OS_STATUS)` pairs (or descriptive counts: N events, N censored, follow-up range) and tell the user to run KM in R (`survival::survfit`) or Python (`lifelines.KaplanMeierFitter`), or use cBioPortal's Survival comparison. If asked to do any of (1)–(5), respond with the appropriate handoff template from the "Approved Response Templates" section below — do not produce the number. Ambiguous-Term Clarification (before any analysis) -------------------------------------------------- If the user's question uses an ambiguous term, ask before computing anything. Examples: - **"aggressive"** → could mean shorter overall survival, higher metastasis rate, higher grade/stage, higher TMB, or specific molecular features. Ask which definit … (9984 more chars) ◀ result {"result":[{"cancer_study_identifier":"pancan_pcawg_2020","name":"Pan-cancer analysis of whole genomes (ICGC/TCGA, Nature 2020)","type_of_cancer_id":"mixed","sample_count":2922,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=pancan_pcawg_2020"}]} ▶ read_guide { "uri": "cbioportal://mutation-frequency-guide" } ◀ result # Mutation Frequency Analysis Guide ## IMPORTANT: Reporting Mutation Frequencies - **ALWAYS report frequencies as percentages**, not raw counts: `frequency = (altered_samples / total_profiled_samples) × 100` - For quick frequency lookups, **prefer the TCGA Pan-Cancer Atlas study first**, then offer to expand to other studies - When reporting across multiple studies, show **ranges** (e.g., "TP53 is mutated in 30–60% of samples") rather than a single average - **NEVER** sum mutation events across studies to compute an aggregate frequency — this can exceed 100% due to double-counting - Warn users that samples may overlap across cohorts (e.g., MSK studies may share patients) - **Choose and state the counting unit**: use patient-level frequencies for prevalence/rate questions unless the user explicitly asks for samples; use sample-level frequencies when the user asks about samples. - **For "across cancer types" questions**, jump to the [Cross-Cancer-Type Mutation Frequency](#cross-cancer-type-mutation-frequency) section below — there is one correct recipe and several common wrong ones. ## Counting Unit: Samples vs Patients Before answering any mutation count or frequency question, decide whether the unit is samples or patients and state that choice in the answer. | User wording | Counting unit | |--------------|---------------| | "prevalence", "rate", "fraction of patients", "patients with", "how common is" | Patient-level: `COUNT(DISTINCT patient_unique_id)` | | "samples", "specimens", "biopsies", sample-level cohort composition | Sample-level: `COUNT(DISTINCT sample_unique_id)` | | Ambiguous | Ask, or default to patient-level for prevalence/rate language and say so | ### Cross-study sample-count caveat When an answer touches more than one study and reports a sample count, prepend a one-line caveat: > Sample IDs are unique within cBioPortal study prefixes, not guaranteed biological-sample identifiers across studies; overlapping cohorts can count the same patient/sample more than once. Prefer one of these safer approaches: - Use a shipped `cancer_study_query_preferences` cohort such as `pan_cancer_tcga` or `all_studies_non_redundant`. - Restrict to one named study. - Aggregate by `patient_unique_id` when the biological question is patient prevalence. ## STOP rule: a frequency above 100% means your query is wrong If your query returns a frequency over 100%, **do not try to debug or explain the data inconsistency to the user**. The cause is always one of these query bugs: - Summing mutation events instead of `COUNT(DISTINCT sample_unique_id)` for the numerator - Using a study-wide sample count as the denominator instead of the gene-specific profiled count - Cross-study aggregation where the same biological sample appears under multiple `sample_unique_id` values (e.g., MSK-IMPACT and MSK-CHORD share patients) - **Joining the profiled CTE through `gene_panel` / `gene_panel_list` without a WES branch.** `gene_panel_id = 'WES'` is *not* a row in … (36467 more chars) ▶ clickhouse_list_table_columns { "table": "genomic_event_derived" } ▶ clickhouse_list_table_columns { "table": "clinical_data_derived" } ◀ result {"columns":[{"name":"sample_unique_id","type":"String"},{"name":"hugo_gene_symbol","type":"String","comment":"HUGO gene symbol (e.g., TP53, KRAS, BRAF). Use for gene-specific queries."},{"name":"entrez_gene_id","type":"Int32"},{"name":"gene_panel_stable_id","type":"LowCardinality(String)"},{"name":"cancer_study_identifier","type":"LowCardinality(String)"},{"name":"genetic_profile_stable_id","type":"LowCardinality(String)"},{"name":"variant_type","type":"LowCardinality(String)","comment":"Type of genomic event: \"mutation\" for SNVs/indels, \"cna\" for copy number alterations, \"structural_variant\" for SVs. Always filter by this."},{"name":"mutation_variant","type":"String","comment":"Protein change notation (e.g., p.V600E, p.R175H). Use for specific variant queries. \"NA\" for non-mutation events."},{"name":"mutation_type","type":"LowCardinality(String)"},{"name":"mutation_status","type":"LowCardinality(String)","comment":"For mutations: Somatic, Germline, UNKNOWN, or UNCALLED. Filter mutation_status != \"UNCALLED\" to exclude uncertain calls. Include all other statuses."},{"name":"driver_filter","type":"LowCardinality(String)"},{"name":"driver_filter_annotation","type":"String"},{"name":"driver_tiers_filter","type":"LowCardinality(String)"},{"name":"driver_tiers_filter_annotation","type":"String"},{"name":"cna_alteration","type":"Nullable(Int8)","comment":"Copy number alteration: only 2 = amplification (AMP) and -2 = deep deletion (HOMDEL) are stored. Shallow deletion (-1), diploid (0) and gain (1) are not in this table; query genetic_alteration_derived WHERE profile_type = 'gistic' (alteration_value is a String, e.g. '-1'). NULL for non-CNA events."},{"name":"cna_cytoband","type":"String"},{"name":"sv_event_info","type":"String"},{"name":"patient_unique_id","type":"String"},{"name":"off_panel","type":"Bool","comment":"Boolean: 1 = mutation outside gene panel coverage (off-panel), 0 = within panel (on-panel). Filter off_panel = 0 for reliable frequency calculations."}]} ◀ result {"columns":[{"name":"internal_id","type":"Int32"},{"name":"sample_unique_id","type":"String","comment":"Globally unique sample ID: cancer_study_identifier + \"_\" + sample.stable_id. Empty for patient-level attributes. Use this for sample filtering and joins."},{"name":"patient_unique_id","type":"String","comment":"Globally unique patient ID: cancer_study_identifier + \"_\" + patient.stable_id. Present for both sample and patient-level attributes."},{"name":"attribute_name","type":"LowCardinality(String)","comment":"Clinical attribute name (e.g., SAMPLE_TYPE, CANCER_TYPE, AGE, OS_MONTHS). Use with attribute_value for filtering. AGE may be floored or capped for de-identification (e.g. all children recorded as 18, or everyone 89+ recorded as 89 or 90): before age statistics check for a pile-up at the min/max, and if present compute age from DAYS_TO_BIRTH (-days / 365.25)."},{"name":"attribute_value","type":"String","comment":"Value of the clinical attribute (String). For SAMPLE_TYPE: Primary, Metastasis, Local Recurrence, Unknown. Missing values are empty strings, so use toFloat64OrNull(attribute_value) for numeric comparisons — CAST fails on them."},{"name":"cancer_study_identifier","type":"LowCardinality(String)"},{"name":"type","type":"LowCardinality(String)","comment":"Data level: \"sample\" for sample-level attributes (e.g., SAMPLE_TYPE), \"patient\" for patient-level attributes (e.g., AGE, OS_MONTHS)."}]} ▶ clickhouse_run_select_query { "query": " WITH mut AS ( SELECT DISTINCT patient_unique_id, hugo_gene_symbol FROM genomic_event_derived WHERE cancer_study_identifier = 'pancan_pcawg_2020' AND variant_type = 'mutation' AND mutation_status != 'UNCALLED' AND hugo_gene_symbol IN ('TP53', 'KRAS') ), tp53_patients AS (SELECT DISTINCT patient_unique_id FROM mut WHERE hugo_gene_symbol = 'TP53'), kras_patients AS (SELECT DISTINCT patient_unique_id FROM mut WHERE hugo_gene_symbol = 'KRAS'), both_patients AS (SELECT patient_unique_id FROM tp53_patients WHERE patient_unique_id IN (SELECT patient_unique_id FROM kras_patients)), kras_only_patients AS (SELECT patient_unique_id FROM kras_patients WHERE patient_unique_id NOT IN (SELECT patient_unique_id FROM tp53_patients)) SELECT 'both' AS grp, count() AS n FROM both_patients UNION ALL SELECT 'kras_only' AS grp, count() AS n FROM kras_only_patients " } ◀ result {"rows":[{"grp":"both","n":187},{"grp":"kras_only","n":85}]} ▶ clickhouse_run_select_query { "query": " WITH mut AS ( SELECT DISTINCT patient_unique_id, hugo_gene_symbol FROM genomic_event_derived WHERE cancer_study_identifier = 'pancan_pcawg_2020' AND variant_type = 'mutation' AND mutation_status != 'UNCALLED' AND hugo_gene_symbol IN ('TP53', 'KRAS') ), tp53_patients AS (SELECT DISTINCT patient_unique_id FROM mut WHERE hugo_gene_symbol = 'TP53'), kras_patients AS (SELECT DISTINCT patient_unique_id FROM mut WHERE hugo_gene_symbol = 'KRAS'), both_patients AS (SELECT patient_unique_id FROM tp53_patients WHERE patient_unique_id IN (SELECT patient_unique_id FROM kras_patients)), kras_only_patients AS (SELECT patient_unique_id FROM kras_patients WHERE patient_unique_id NOT IN (SELECT patient_unique_id FROM tp53_patients)), os_status AS ( SELECT patient_unique_id, attribute_value AS os_status FROM clinical_data_derived WHERE cancer_study_identifier = 'pancan_pcawg_2020' AND attribute_name = 'OS_STATUS' ), os_months AS ( SELECT patient_unique_id, toFloat64OrNull(attribute_value) AS os_months FROM clinical_data_derived WHERE cancer_study_identifier = 'pancan_pcawg_2020' AND attribute_name = 'OS_MONTHS' ) SELECT 'both' AS grp, count() AS n_patients, countIf(s.os_status = '1:DECEASED') AS n_events, countIf(s.os_status = '0:LIVING') AS n_censored, min(m.os_months) AS min_followup, max(m.os_months) AS max_followup FROM both_patients b LEFT JOIN os_status s USING (patient_unique_id) LEFT JOIN os_months m USING (patient_unique_id) UNION ALL SELECT 'kras_only' AS grp, count() AS n_patients, countIf(s.os_status = '1:DECEASED') AS n_events, countIf(s.os_status = '0:LIVING') AS n_censored, min(m.os_months) AS min_followup, max(m.os_months) AS max_followup FROM kras_only_patients k LEFT JOIN os_status s USING (patient_unique_id) LEFT JOIN os_months m USING (patient_unique_id) " } ◀ result {"rows":[{"grp":"both","n_patients":187,"n_events":122,"n_censored":65,"min_followup":33.51973684,"max_followup":33.51973684},{"grp":"kras_only","n_patients":85,"n_events":41,"n_censored":44,"min_followup":0,"max_followup":29.60526316}]} ▶ clickhouse_run_select_query { "query": " SELECT attribute_name, type, count(), countIf(attribute_value != '') AS non_empty FROM clinical_data_derived WHERE cancer_study_identifier = 'pancan_pcawg_2020' AND attribute_name IN ('OS_MONTHS','OS_STATUS') GROUP BY attribute_name, type " } ◀ result {"rows":[{"attribute_name":"OS_STATUS","type":"patient","count()":2658,"non_empty":2458},{"attribute_name":"OS_MONTHS","type":"patient","count()":2658,"non_empty":282}]} ▶ clickhouse_run_select_query { "query": " WITH mut AS ( SELECT DISTINCT patient_unique_id, hugo_gene_symbol FROM genomic_event_derived WHERE cancer_study_identifier = 'pancan_pcawg_2020' AND variant_type = 'mutation' AND mutation_status != 'UNCALLED' AND hugo_gene_symbol IN ('TP53', 'KRAS') ), tp53_patients AS (SELECT DISTINCT patient_unique_id FROM mut WHERE hugo_gene_symbol = 'TP53'), kras_patients AS (SELECT DISTINCT patient_unique_id FROM mut WHERE hugo_gene_symbol = 'KRAS'), both_patients AS (SELECT patient_unique_id FROM tp53_patients WHERE patient_unique_id IN (SELECT patient_unique_id FROM kras_patients)), kras_only_patients AS (SELECT patient_unique_id FROM kras_patients WHERE patient_unique_id NOT IN (SELECT patient_unique_id FROM tp53_patients)), os_status AS ( SELECT patient_unique_id, attribute_value AS os_status FROM clinical_data_derived WHERE cancer_study_identifier = 'pancan_pcawg_2020' AND attribute_name = 'OS_STATUS' ), os_months AS ( SELECT patient_unique_id, toFloat64OrNull(attribute_value) AS os_months FROM clinical_data_derived WHERE cancer_study_identifier = 'pancan_pcawg_2020' AND attribute_name = 'OS_MONTHS' AND attribute_value != '' ) SELECT 'both' AS grp, (SELECT count() FROM both_patients) AS n_patients, countIf(s.os_status = '1:DECEASED') AS n_events, countIf(s.os_status = '0:LIVING') AS n_censored, countIf(s.os_status = '') AS n_missing_status, count(m.os_months) AS n_with_followup_time, min(m.os_months) AS min_followup, max(m.os_months) AS max_followup FROM both_patients b LEFT JOIN os_status s USING (patient_unique_id) LEFT JOIN os_months m USING (patient_unique_id) UNION ALL SELECT 'kras_only' AS grp, (SELECT count() FROM kras_only_patients) AS n_patients, countIf(s.os_status = '1:DECEASED') AS n_events, countIf(s.os_status = '0:LIVING') AS n_censored, countIf(s.os_status = '') AS n_missing_status, count(m.os_months) AS n_with_followup_time, min(m.os_months) AS min_followup, max(m.os_months) AS max_followup FROM kras_only_patients k LEFT JOIN os_status s USING (patient_unique_id) LEFT JOIN os_months m USING (patient_unique_id) " } ◀ result {"rows":[{"grp":"kras_only","n_patients":85,"n_events":41,"n_censored":44,"n_missing_status":0,"n_with_followup_time":3,"min_followup":0,"max_followup":29.60526316},{"grp":"both","n_patients":187,"n_events":122,"n_censored":65,"n_missing_status":0,"n_with_followup_time":1,"min_followup":33.51973684,"max_followup":33.51973684}]} ▶ resolve_and_route { "studyIds": [ "pancan_pcawg_2020" ] } ◀ result {"success":true,"message":"Found 1 study. Use the metadata to call the appropriate navigation tool(s).","data":{"totalCount":1,"studiesWithMetadata":[{"studyId":"pancan_pcawg_2020","name":"Pan-cancer analysis of whole genomes (ICGC/TCGA, Nature 2020)","sampleCount":2922,"studyViewUrl":"https://www.cbioportal.org/study?id=pancan_pcawg_2020","metadata":{"clinicalAttributeIds":["AGE","ALCOHOL","ALCOHOL_HISTORY_INTENSITY","ANCESTRY_PRIMARY","CANCER_TYPE","CANCER_TYPE_DETAILED","CELLULARITY","FIRST THERAPY_RESPONSE","FIRST_THERAPY","GRADE","HISTOLOGY","HISTOLOGY_ABBREVIATION","HISTOLOGY_TIER1","HISTOLOGY_TIER2","HISTOLOGY_TIER3","HISTOLOGY_TIER4","ICD_10","ICGC_SAMPLE_ID","MUTATION_COUNT","ONCOTREE_CODE","ORGAN_SYSTEM","OS_MONTHS","OS_STATUS","PLOIDY","PROJECT_CODE","PURITY","PURITY_CONFUGURATION","SAMPLE_CLASS","SAMPLE_COUNT","SAMPLE_TYPE","SEQUENCING_TYPE","SEX","STAGE","TBL_SCORE","TMB_NONSYNONYMOUS","TOBACCO_SMOKING_HISTORY_INDICATOR","TOBACCO_SMOKING_INTENSITY","TUMOR_SAMPLE_HISTOLOGY_CODE","WGD"],"molecularProfileIds":["pancan_pcawg_2020_cna","pancan_pcawg_2020_mirna","pancan_pcawg_2020_mirna_median_Zscores","pancan_pcawg_2020_mrna_seq_fpkm_capture","pancan_pcawg_2020_mrna_seq_fpkm_capture_all_sample_Zscores","pancan_pcawg_2020_mutations"],"genericAssayProfiles":["pancan_pcawg_2020_mutational_signatures_contribution_DBS","pancan_pcawg_2020_mutational_signatures_contribution_ID","pancan_pcawg_2020_mutational_signatures_contribution_SBS","pancan_pcawg_2020_mutational_signatures_counts_DBS","pancan_pcawg_2020_mutational_signatures_counts_ID","pancan_pcawg_2020_mutational_signatures_counts_SBS"],"heatmapProfileIds":["pancan_pcawg_2020_mirna_median_Zscores","pancan_pcawg_2020_mrna_seq_fpkm_capture_all_sample_Zscores"],"availableComparisonTabs":["overlap","clinical","survival","alterations","mutations","mrna","generic_assay_mutational_signature"]}}]}} ▶ navigate_to_group_comparison { "studyIds": [ "pancan_pcawg_2020" ], "tab": "survival", "groups": [ { "name": "TP53+KRAS mutant", "studyViewFilter": { "geneFilters": [ { "molecularProfileIds": [ "pancan_pcawg_2020_mutations" ], "geneQueries": [ [ { "hugoGeneSymbol": "TP53" } ], [ { "hugoGeneSymbol": "KRAS" } ] ] } ] } }, { "name": "KRAS-only mutant", "studyViewFilter": { "mutationDataFilters": [ { "hugoGeneSymbol": "KRAS", "profileType": "mutations", "categorization": "MUTATED", "values": [ [ { "value": "true" } ] ] }, { "hugoGeneSymbol": "TP53", "profileType": "mutations", "categorization": "MUTATED", "values": [ [ { "value": "false" } ] ] } ] } } ] } ◀ result {"success":false,"error":"Internal Server Error","details":{"status":500,"response":{"req":{"method":"POST","url":"https://www.cbioportal.org/api/filtered-samples/fetch","data":{"geneFilters":[{"molecularProfileIds":["pancan_pcawg_2020_mutations"],"geneQueries":[[{"hugoGeneSymbol":"TP53"}],[{"hugoGeneSymbol":"KRAS"}]]}],"studyIds":["pancan_pcawg_2020"]},"headers":{"user-agent":"node-superagent/3.8.3","accept":"application/json","content-type":"application/json"}},"header":{"cache-control":"no-cache, no-store, max-age=0, must-revalidate","content-encoding":"gzip","content-type":"application/json","date":"Sat, 26 Sep 2026 16:46:01 GMT","elapsed-time":"0","expires":"0","pragma":"no-cache","strict-transport-security":"max-age=31536000 ; includeSubDomains","vary":"origin,access-control-request-method,access-control-request-headers,accept-encoding","x-content-type-options":"nosniff","x-frame-options":"DENY","x-xss-protection":"0","connection":"close","transfer-encoding":"chunked"},"status":500,"text":"{\"timestamp\":1790441161310,\"status\":500,\"error\":\"Internal Server Error\",\"path\":\"/api/filtered-samples/fetch\"}"}}} ▶ get_studyviewfilter_options { "studyId": "pancan_pcawg_2020", "geneSpecificQueries": [ { "hugoGeneSymbol": "KRAS", "profileType": "mutations" }, { "hugoGeneSymbol": "TP53", "profileType": "mutations" } ] } ◀ result {"success":true,"message":"Successfully retrieved StudyView filter options","data":{"studyId":"pancan_pcawg_2020","geneSpecificCounts":[{"hugoGeneSymbol":"TP53","profileType":"mutations","counts":[{"value":"Missense_Mutation","label":"Missense Mutation","count":599},{"value":"Frame_Shift_Del","label":"Frame Shift Del","count":87},{"value":"In_Frame_Ins","label":"In Frame Ins","count":6},{"value":"Splice_Site","label":"Splice Site","count":69},{"value":"Frame_Shift_Ins","label":"Frame Shift Ins","count":36},{"value":"Splice_Region","label":"Splice Region","count":5},{"value":"In_Frame_Del","label":"In Frame Del","count":19},{"value":"Nonsense_Mutation","label":"Nonsense Mutation","count":117}]},{"hugoGeneSymbol":"KRAS","profileType":"mutations","counts":[{"value":"Missense_Mutation","label":"Missense Mutation","count":278}]}]}} ▶ navigate_to_group_comparison { "studyIds": [ "pancan_pcawg_2020" ], "tab": "survival", "groups": [ { "name": "TP53+KRAS mutant", "studyViewFilter": { "geneFilters": [ { "molecularProfileIds": [ "pancan_pcawg_2020_mutations" ], "geneQueries": [ [ { "hugoGeneSymbol": "TP53" }, { "hugoGeneSymbol": "KRAS" } ] ] } ] } }, { "name": "KRAS-only mutant", "studyViewFilter": { "mutationDataFilters": [ { "hugoGeneSymbol": "KRAS", "profileType": "mutations", "categorization": "MUTATED", "values": [ [ { "value": "true" } ] ] }, { "hugoGeneSymbol": "TP53", "profileType": "mutations", "categorization": "MUTATED", "values": [ [ { "value": "false" } ] ] } ] } } ] } ◀ result {"success":false,"error":"Internal Server Error","details":{"status":500,"response":{"req":{"method":"POST","url":"https://www.cbioportal.org/api/filtered-samples/fetch","data":{"mutationDataFilters":[{"hugoGeneSymbol":"KRAS","profileType":"mutations","categorization":"MUTATED","values":[[{"value":"true"}]]},{"hugoGeneSymbol":"TP53","profileType":"mutations","categorization":"MUTATED","values":[[{"value":"false"}]]}],"studyIds":["pancan_pcawg_2020"]},"headers":{"user-agent":"node-superagent/3.8.3","accept":"application/json","content-type":"application/json"}},"header":{"cache-control":"no-cache, no-store, max-age=0, must-revalidate","content-encoding":"gzip","content-type":"application/json","date":"Sat, 26 Sep 2026 16:46:36 GMT","elapsed-time":"209","expires":"0","pragma":"no-cache","strict-transport-security":"max-age=31536000 ; includeSubDomains","vary":"origin,access-control-request-method,access-control-request-headers,accept-encoding","x-content-type-options":"nosniff","x-frame-options":"DENY","x-xss-protection":"0","connection":"close","transfer-encoding":"chunked"},"status":500,"text":"{\"message\":\"SQL exception. If you are a maintainer of this instance, see logs for details.\"}"}}} ▶ navigate_to_study_view { "studyIds": [ "pancan_pcawg_2020" ], "filterJson": { "mutationDataFilters": [ { "hugoGeneSymbol": "KRAS", "profileType": "mutations", "categorization": "MUTATED", "values": [ [ { "value": "Mutated" } ] ] }, { "hugoGeneSymbol": "TP53", "profileType": "mutations", "categorization": "MUTATED", "values": [ [ { "value": "Not mutated" } ] ] } ] } } ◀ result {"success":true,"message":"Navigating to https://www.cbioportal.org/study/summary?id=pancan_pcawg_2020#filterJson=%7B%22mutationDataFilters%22%3A%5B%7B%22categorization%22%3A%22MUTATED%22%2C%22hugoGeneSymbol%22%3A%22KRAS%22%2C%22profileType%22%3A%22mutations%22%2C%22values%22%3A%5B%5B%7B%22value%22%3A%22Mutated%22%7D%5D%5D%7D%2C%7B%22categorization%22%3A%22MUTATED%22%2C%22hugoGeneSymbol%22%3A%22TP53%22%2C%22profileType%22%3A%22mutations%22%2C%22values%22%3A%5B%5B%7B%22value%22%3A%22Not%20mutated%22%7D%5D%5D%7D%5D%7D","url":"https://www.cbioportal.org/study/summary?id=pancan_pcawg_2020#filterJson=%7B%22mutationDataFilters%22%3A%5B%7B%22categorization%22%3A%22MUTATED%22%2C%22hugoGeneSymbol%22%3A%22KRAS%22%2C%22profileType%22%3A%22mutations%22%2C%22values%22%3A%5B%5B%7B%22value%22%3A%22Mutated%22%7D%5D%5D%7D%2C%7B%22categorization%22%3A%22MUTATED%22%2C%22hugoGeneSymbol%22%3A%22TP53%22%2C%22profileType%22%3A%22mutations%22%2C%22values%22%3A%5B%5B%7B%22value%22%3A%22Not%20mutated%22%7D%5D%5D%7D%5D%7D","data":{"studyIds":["pancan_pcawg_2020"],"studies":[{"studyId":"pancan_pcawg_2020","name":"Pan-cancer analysis of whole genomes (ICGC/TCGA, Nature 2020)","sampleCount":2922}],"tab":"summary","hasFilters":true,"hasPlotsConfig":false,"pageDescription":"Overview of all study data displayed as interactive charts summarizing available clinical and genomic data. Click within a chart to filter. Additional charts may be available in the Charts menu."}} ▶ navigate_to_group_comparison { "studyIds": [ "pancan_pcawg_2020" ], "tab": "survival", "groups": [ { "name": "TP53+KRAS mutant", "studyViewFilter": { "geneFilters": [ { "molecularProfileIds": [ "pancan_pcawg_2020_mutations" ], "geneQueries": [ [ { "hugoGeneSymbol": "TP53" }, { "hugoGeneSymbol": "KRAS" } ] ] } ] } }, { "name": "KRAS-only mutant", "studyViewFilter": { "mutationDataFilters": [ { "hugoGeneSymbol": "KRAS", "profileType": "mutations", "categorization": "MUTATED", "values": [ [ { "value": "Mutated" } ] ] }, { "hugoGeneSymbol": "TP53", "profileType": "mutations", "categorization": "MUTATED", "values": [ [ { "value": "Not mutated" } ] ] } ] } } ] } ◀ result {"success":false,"error":"Internal Server Error","details":{"status":500,"response":{"req":{"method":"POST","url":"https://www.cbioportal.org/api/filtered-samples/fetch","data":{"geneFilters":[{"molecularProfileIds":["pancan_pcawg_2020_mutations"],"geneQueries":[[{"hugoGeneSymbol":"TP53"},{"hugoGeneSymbol":"KRAS"}]]}],"studyIds":["pancan_pcawg_2020"]},"headers":{"user-agent":"node-superagent/3.8.3","accept":"application/json","content-type":"application/json"}},"header":{"cache-control":"no-cache, no-store, max-age=0, must-revalidate","content-encoding":"gzip","content-type":"application/json","date":"Sat, 26 Sep 2026 16:47:15 GMT","elapsed-time":"1","expires":"0","pragma":"no-cache","strict-transport-security":"max-age=31536000 ; includeSubDomains","vary":"origin,access-control-request-method,access-control-request-headers,accept-encoding","x-content-type-options":"nosniff","x-frame-options":"DENY","x-xss-protection":"0","connection":"close","transfer-encoding":"chunked"},"status":500,"text":"{\"timestamp\":1790441235467,\"status\":500,\"error\":\"Internal Server Error\",\"path\":\"/api/filtered-samples/fetch\"}"}}} ▶ navigate_to_group_comparison { "studyIds": [ "pancan_pcawg_2020" ], "tab": "survival", "groups": [ { "name": "TP53+KRAS mutant", "studyViewFilter": { "mutationDataFilters": [ { "hugoGeneSymbol": "KRAS", "profileType": "mutations", "categorization": "MUTATED", "values": [ [ { "value": "Mutated" } ] ] }, { "hugoGeneSymbol": "TP53", "profileType": "mutations", "categorization": "MUTATED", "values": [ [ { "value": "Mutated" } ] ] } ] } }, { "name": "KRAS-only mutant", "studyViewFilter": { "mutationDataFilters": [ { "hugoGeneSymbol": "KRAS", "profileType": "mutations", "categorization": "MUTATED", "values": [ [ { "value": "Mutated" } ] ] }, { "hugoGeneSymbol": "TP53", "profileType": "mutations", "categorization": "MUTATED", "values": [ [ { "value": "Not mutated" } ] ] } ] } } ] } ◀ result {"success":false,"error":"Internal Server Error","details":{"status":500,"response":{"req":{"method":"POST","url":"https://www.cbioportal.org/api/filtered-samples/fetch","data":{"mutationDataFilters":[{"hugoGeneSymbol":"KRAS","profileType":"mutations","categorization":"MUTATED","values":[[{"value":"Mutated"}]]},{"hugoGeneSymbol":"TP53","profileType":"mutations","categorization":"MUTATED","values":[[{"value":"Not mutated"}]]}],"studyIds":["pancan_pcawg_2020"]},"headers":{"user-agent":"node-superagent/3.8.3","accept":"application/json","content-type":"application/json"}},"header":{"cache-control":"no-cache, no-store, max-age=0, must-revalidate","content-encoding":"gzip","content-type":"application/json","date":"Sat, 26 Sep 2026 16:47:19 GMT","elapsed-time":"24","expires":"0","pragma":"no-cache","strict-transport-security":"max-age=31536000 ; includeSubDomains","vary":"origin,access-control-request-method,access-control-request-headers,accept-encoding","x-content-type-options":"nosniff","x-frame-options":"DENY","x-xss-protection":"0","connection":"close","transfer-encoding":"chunked"},"status":500,"text":"{\"message\":\"SQL exception. If you are a maintainer of this instance, see logs for details.\"}"}}} ▶ navigate_to_study_view { "studyIds": [ "pancan_pcawg_2020" ], "filterJson": { "mutationDataFilters": [ { "hugoGeneSymbol": "KRAS", "profileType": "mutations", "categorization": "MUTATED", "values": [ [ { "value": "Mutated" } ] ] }, { "hugoGeneSymbol": "TP53", "profileType": "mutations", "categorization": "MUTATED", "values": [ [ { "value": "Mutated" } ] ] } ] } } ◀ result {"success":true,"message":"Navigating to https://www.cbioportal.org/study/summary?id=pancan_pcawg_2020#filterJson=%7B%22mutationDataFilters%22%3A%5B%7B%22categorization%22%3A%22MUTATED%22%2C%22hugoGeneSymbol%22%3A%22KRAS%22%2C%22profileType%22%3A%22mutations%22%2C%22values%22%3A%5B%5B%7B%22value%22%3A%22Mutated%22%7D%5D%5D%7D%2C%7B%22categorization%22%3A%22MUTATED%22%2C%22hugoGeneSymbol%22%3A%22TP53%22%2C%22profileType%22%3A%22mutations%22%2C%22values%22%3A%5B%5B%7B%22value%22%3A%22Mutated%22%7D%5D%5D%7D%5D%7D","url":"https://www.cbioportal.org/study/summary?id=pancan_pcawg_2020#filterJson=%7B%22mutationDataFilters%22%3A%5B%7B%22categorization%22%3A%22MUTATED%22%2C%22hugoGeneSymbol%22%3A%22KRAS%22%2C%22profileType%22%3A%22mutations%22%2C%22values%22%3A%5B%5B%7B%22value%22%3A%22Mutated%22%7D%5D%5D%7D%2C%7B%22categorization%22%3A%22MUTATED%22%2C%22hugoGeneSymbol%22%3A%22TP53%22%2C%22profileType%22%3A%22mutations%22%2C%22values%22%3A%5B%5B%7B%22value%22%3A%22Mutated%22%7D%5D%5D%7D%5D%7D","data":{"studyIds":["pancan_pcawg_2020"],"studies":[{"studyId":"pancan_pcawg_2020","name":"Pan-cancer analysis of whole genomes (ICGC/TCGA, Nature 2020)","sampleCount":2922}],"tab":"summary","hasFilters":true,"hasPlotsConfig":false,"pageDescription":"Overview of all study data displayed as interactive charts summarizing available clinical and genomic data. Click within a chart to filter. Additional charts may be available in the Charts menu."}} ▶ navigate_to_study_view { "studyIds": [ "pancan_pcawg_2020" ], "filterJson": { "mutationDataFilters": [ { "hugoGeneSymbol": "KRAS", "profileType": "mutations", "categorization": "MUTATED", "values": [ [ { "value": "Mutated" } ] ] }, { "hugoGeneSymbol": "TP53", "profileType": "mutations", "categorization": "MUTATED", "values": [ [ { "value": "Not mutated" } ] ] } ] } } ◀ result {"success":true,"message":"Navigating to https://www.cbioportal.org/study/summary?id=pancan_pcawg_2020#filterJson=%7B%22mutationDataFilters%22%3A%5B%7B%22categorization%22%3A%22MUTATED%22%2C%22hugoGeneSymbol%22%3A%22KRAS%22%2C%22profileType%22%3A%22mutations%22%2C%22values%22%3A%5B%5B%7B%22value%22%3A%22Mutated%22%7D%5D%5D%7D%2C%7B%22categorization%22%3A%22MUTATED%22%2C%22hugoGeneSymbol%22%3A%22TP53%22%2C%22profileType%22%3A%22mutations%22%2C%22values%22%3A%5B%5B%7B%22value%22%3A%22Not%20mutated%22%7D%5D%5D%7D%5D%7D","url":"https://www.cbioportal.org/study/summary?id=pancan_pcawg_2020#filterJson=%7B%22mutationDataFilters%22%3A%5B%7B%22categorization%22%3A%22MUTATED%22%2C%22hugoGeneSymbol%22%3A%22KRAS%22%2C%22profileType%22%3A%22mutations%22%2C%22values%22%3A%5B%5B%7B%22value%22%3A%22Mutated%22%7D%5D%5D%7D%2C%7B%22categorization%22%3A%22MUTATED%22%2C%22hugoGeneSymbol%22%3A%22TP53%22%2C%22profileType%22%3A%22mutations%22%2C%22values%22%3A%5B%5B%7B%22value%22%3A%22Not%20mutated%22%7D%5D%5D%7D%5D%7D","data":{"studyIds":["pancan_pcawg_2020"],"studies":[{"studyId":"pancan_pcawg_2020","name":"Pan-cancer analysis of whole genomes (ICGC/TCGA, Nature 2020)","sampleCount":2922}],"tab":"summary","hasFilters":true,"hasPlotsConfig":false,"pageDescription":"Overview of all study data displayed as interactive charts summarizing available clinical and genomic data. Click within a chart to filter. Additional charts may be available in the Charts menu."}} ═ answer (success) ## Results — Pan-cancer analysis of whole genomes (PCAWG, Nature 2020) **Cohort:** `pancan_pcawg_2020` (2,922 samples). Counting unit: patients. | Group | Patients | Deceased (events) | Living (censored) | |---|---|---|---| | TP53-mutant **and** KRAS-mutant | 187 | 122 | 65 | | KRAS-mutant only (TP53 wild-type) | 85 | 41 | 44 | **A critical data limitation:** I cannot compute a Kaplan-Meier survival difference or a log-rank p-value for this comparison from the database directly. While vital status (`OS_STATUS`, deceased/living) is fully recorded for every patient, the actual **follow-up time** field (`OS_MONTHS`) is populated for only 4 of these 272 patients combined (essentially empty for both groups) in this study's clinical data. Without follow-up time, there is no way to build a survival curve — event/censoring counts alone don't tell you *when* those events happened, so a valid KM estimate or log-rank test cannot be constructed, by me or by the portal, from what's stored here. Given that, I did not attempt to fabricate a median survival or significance value (per policy, I never derive p-values, hazard ratios, or medians from raw survival columns — those require Kaplan-Meier/log-rank/Cox regression tools, and here the underlying time data isn't even usable). **What you can check directly:** the cBioPortal Group Comparison Survival tab would normally run this exact KM + log-rank comparison automatically, but my attempts to launch it programmatically for these two custom mutation-based groups hit a server-side error on cBioPortal.org (500 Internal Server Error on the group-filtering endpoint) — this is a backend issue, not a data-availability one on my end for the group-builder call itself. You can still build the comparison manually in the portal: 1. [KRAS+TP53 co-mutant cohort (StudyView, pre-filtered)](https://www.cbioportal.org/study/summary?id=pancan_pcawg_2020#filterJson=%7B%22mutationDataFilters%22%3A%5B%7B%22categorization%22%3A%22MUTATED%22%2C%22hugoGeneSymbol%22%3A%22KRAS%22%2C%22profileType%22%3A%22mutations%22%2C%22values%22%3A%5B%5B%7B%22value%22%3A%22Mutated%22%7D%5D%5D%7D%2C%7B%22categorization%22%3A%22MUTATED%22%2C%22hugoGeneSymbol%22%3A%22TP53%22%2C%22profileType%22%3A%22mutations%22%2C%22values%22%3A%5B%5B%7B%22value%22%3A%22Mutated%22%7D%5D%5D%7D%5D%7D) 2. [KRAS-mutant, TP53 wild-type cohort (StudyView, pre-filtered)](https://www.cbioportal.org/study/summary?id=pancan_pcawg_2020#filterJson=%7B%22mutationDataFilters%22%3A%5B%7B%22categorization%22%3A%22MUTATED%22%2C%22hugoGeneSymbol%22%3A%22KRAS%22%2C%22profileType%22%3A%22mutations%22%2C%22values%22%3A%5B%5B%7B%22value%22%3A%22Mutated%22%7D%5D%5D%7D%2C%7B%22categorization%22%3A%22MUTATED%22%2C%22hugoGeneSymbol%22%3A%22TP53%22%2C%22profileType%22%3A%22mutations%22%2C%22values%22%3A%5B%5B%7B%22value%22%3A%22Not%20mutated%22%7D%5D%5D%7D%5D%7D) From either page, select all filtered cases and use "Compare" → "Survival" to attempt the KM plot — but given what the database itself shows (follow-up time essentially absent for both groups), I'd expect the portal to display "not enough data" or a survival curve that is not meaningful/reliable for this particular study, rather than a robust comparison. If survival is central to your question, this PCAWG cohort likely isn't a good candidate for it; TCGA PanCancer Atlas studies for the same tumor types generally have much better `OS_MONTHS` completeness if you'd like me to check specific cancer subtypes there instead.