Q (haiku): Is KRAS G12C more aggressive than G12D? ▶ read_guide { "uri": "cbioportal://statistical-tests-guide" } ▶ read_guide { "uri": "cbioportal://common-pitfalls" } ◀ result Statistical Test Selection Guide ================================ Purpose ------- This guide ensures the correct statistical test is selected before performing any group comparison, matching cBioPortal's own Group Comparison defaults. HARD RULES — NEVER FABRICATE A STATISTIC ---------------------------------------- ClickHouse cannot run statistical tests. The agent therefore must NEVER produce a derived statistic that is not a literal column value from a SQL result. Specifically: 1. **Never invent a p-value.** Not "p < 0.001", not "p ≈ 0.05", not any p-value. If the user asks "what is the p-value?", the answer is *"I can't compute that — here is the 2x2 contingency table (or group statistics). Run it in cBioPortal's Group Comparison tab, in R with `fisher.test(...)` / `wilcox.test(...)`, or in Python with `scipy.stats.fisher_exact(...)` / `mannwhitneyu(...)`."* 2. **Never claim mutual exclusivity (or co-occurrence) from a contingency table alone.** A 2x2 table is not a test. The shape "altered/not altered × group A/group B" needs Fisher's exact + a defined direction (odds ratio < 1 with significant p). Without that test, the agent presents the table and stops. Descriptive phrasing is also a claim: "largely/mostly mutually exclusive", "rarely co-occur", "tend to co-occur" are forbidden without the test — point to cBioPortal's Mutual Exclusivity tab instead. 3. **Never report a "median" that came from `AVG(...)` or any non-median aggregate.** "Median" and "mean" are different statistics; for skewed clinical distributions (especially survival) they differ substantially. Use ClickHouse's `quantile(0.5)(...)` for actual median, and label arithmetic averages as "mean", never "median". 4. **Never report a hazard ratio, odds ratio, risk ratio, or relative risk** that wasn't computed by an external tool. These require regression / model fitting that ClickHouse does not do. 5. **Never report median overall survival from `AVG(OS_MONTHS)` or even `quantile(0.5)(OS_MONTHS)`.** Median OS requires Kaplan-Meier estimation, which handles censoring (`OS_STATUS = 0:LIVING` means the event hasn't happened yet). Naive medians/means over `OS_MONTHS` ignore censoring and are systematically wrong. The correct handoff: return the raw `(OS_MONTHS, OS_STATUS)` pairs (or descriptive counts: N events, N censored, follow-up range) and tell the user to run KM in R (`survival::survfit`) or Python (`lifelines.KaplanMeierFitter`), or use cBioPortal's Survival comparison. If asked to do any of (1)–(5), respond with the appropriate handoff template from the "Approved Response Templates" section below — do not produce the number. Ambiguous-Term Clarification (before any analysis) -------------------------------------------------- If the user's question uses an ambiguous term, ask before computing anything. Examples: - **"aggressive"** → could mean shorter overall survival, higher metastasis rate, higher grade/stage, higher TMB, or specific molecular features. Ask which definit … (9984 more chars) ◀ result # Common Query Pitfalls Guide ## Overview This guide highlights frequent mistakes when analyzing cBioPortal data and provides solutions to avoid them. ## Critical Pitfalls ### 1. 🚨 CRITICAL MUTATION FREQUENCY ERRORS #### ❌ WRONG: Using study-wide totals for gene frequencies ```sql -- INCORRECT - This gives wrong frequencies! SELECT hugo_gene_symbol, COUNT(DISTINCT sample_unique_id) as altered_samples, (SELECT COUNT(DISTINCT sample_unique_id) FROM genomic_event_derived WHERE cancer_study_identifier = 'your_study_id') as total_samples FROM genomic_event_derived WHERE variant_type = 'mutation' AND cancer_study_identifier = 'your_study_id' GROUP BY hugo_gene_symbol; ``` **Problem**: Different genes have different profiling coverage - you can't use study-wide totals! #### ❌ WRONG: Not using gene-specific profiling denominators ```sql -- INCORRECT - Missing gene-specific denominators SELECT hugo_gene_symbol, COUNT(DISTINCT sample_unique_id) as altered_samples FROM genomic_event_derived WHERE variant_type = 'mutation' GROUP BY hugo_gene_symbol; -- Missing: WHERE ARE THE DENOMINATORS FOR EACH GENE? ``` #### ❌ WRONG: Skipping individual gene profiling queries **Problem**: Failing to run separate profiling queries for EACH gene in results. **Each gene has different coverage**: TP53 might be profiled in 25,040 samples, MUC16 in 23,000, etc. #### ✅ CORRECT: Complete gene-specific workflow ```sql -- STEP 1: Get altered counts per gene SELECT hugo_gene_symbol, entrez_gene_id, COUNT(DISTINCT CASE WHEN off_panel = 0 THEN sample_unique_id END) AS numberOfAlteredSamplesOnPanel, COUNT(*) AS totalMutationEvents FROM genomic_event_derived WHERE variant_type = 'mutation' AND mutation_status != 'UNCALLED' GROUP BY entrez_gene_id, hugo_gene_symbol ORDER BY numberOfAlteredSamplesOnPanel DESC; -- STEP 2: FOR EACH GENE, run this profiling query: SELECT COUNT(DISTINCT stgp.sample_unique_id) AS numberOfProfiledSamples FROM sample_to_gene_panel_derived stgp JOIN gene_panel gp ON stgp.gene_panel_id = gp.stable_id JOIN gene_panel_list gpl ON gp.internal_id = gpl.internal_id JOIN gene g ON gpl.gene_id = g.entrez_gene_id WHERE stgp.alteration_type = 'MUTATION_EXTENDED' AND g.hugo_gene_symbol = 'TP53' -- Replace with each gene from Step 1 AND stgp.cancer_study_identifier = 'your_study_id'; -- STEP 3: Calculate frequency = numberOfAlteredSamplesOnPanel / numberOfProfiledSamples * 100 ``` #### 🚨 WORKFLOW REQUIREMENTS VIOLATIONS: - **Missing denominator columns**: Must show gene-specific profiled samples per row - **Wrong table format**: Should be | Gene | # Mutations | # Samples | Profiled Samples | Sample % | - **Not replacing gene symbols**: Must replace 'TP53' with actual gene for each query - **Using study totals**: Never use genomic_event_derived for total sample counts ### 2. 🚨 OFF-PANEL MUTATION INCLUSION #### ❌ Wrong: Including off-panel mutations ```sql -- INCORRECT - Includes mutations outside gene panels SELEC … (38012 more chars) ═ answer (success) Thank you for this question. I need to **clarify what you mean by "aggressive"** before I can provide accurate data. "Aggressive" could mean several things: - **Overall survival (OS)** — which variant has shorter median survival? - **Progression-free survival (PFS)** or **disease-free survival (DFS)**? - **Grade or stage at diagnosis** — are G12C tumors higher stage? - **Mutation frequency** — which is more common in advanced disease? - **Clinical outcome** — response to therapy, time to recurrence? Once you clarify, I'll pull the raw data from cBioPortal and point you to the right statistical test or visualization. For **survival comparison** specifically, I would need to: 1. Filter samples by KRAS G12C vs G12D mutation status 2. Extract their OS_MONTHS and OS_STATUS values (to account for censoring) 3. Hand off to Kaplan-Meier estimation (cBioPortal's Group Comparison → Survival tab, or R/Python) **I cannot produce a "median OS" or p-value myself** — ClickHouse cannot run statistical tests. But I can build the raw survival data and show you how to compute it. **Which cancer type(s)** are you interested in? (e.g., lung adenocarcinoma, pancreatic cancer, colorectal cancer)