Q (sonnet-5.5): In the TCGA Breast Cancer study if a patient has both a PIK3CA mutation and a PTEN deletion what is their probability of having a luminal A subtype? Is this probability different from a patient with only a PIK3CA mutation? ▶ read_guide { "uri": "cbioportal://statistical-tests-guide" } ▶ list_studies { "search": "TCGA breast" } ◀ result Statistical Test Selection Guide ================================ Purpose ------- This guide ensures the correct statistical test is selected before performing any group comparison, matching cBioPortal's own Group Comparison defaults. HARD RULES — NEVER FABRICATE A STATISTIC ---------------------------------------- ClickHouse cannot run statistical tests. The agent therefore must NEVER produce a derived statistic that is not a literal column value from a SQL result. Specifically: 1. **Never invent a p-value.** Not "p < 0.001", not "p ≈ 0.05", not any p-value. If the user asks "what is the p-value?", the answer is *"I can't compute that — here is the 2x2 contingency table (or group statistics). Run it in cBioPortal's Group Comparison tab, in R with `fisher.test(...)` / `wilcox.test(...)`, or in Python with `scipy.stats.fisher_exact(...)` / `mannwhitneyu(...)`."* 2. **Never claim mutual exclusivity (or co-occurrence) from a contingency table alone.** A 2x2 table is not a test. The shape "altered/not altered × group A/group B" needs Fisher's exact + a defined direction (odds ratio < 1 with significant p). Without that test, the agent presents the table and stops. Descriptive phrasing is also a claim: "largely/mostly mutually exclusive", "rarely co-occur", "tend to co-occur" are forbidden without the test — point to cBioPortal's Mutual Exclusivity tab instead. 3. **Never report a "median" that came from `AVG(...)` or any non-median aggregate.** "Median" and "mean" are different statistics; for skewed clinical distributions (especially survival) they differ substantially. Use ClickHouse's `quantile(0.5)(...)` for actual median, and label arithmetic averages as "mean", never "median". 4. **Never report a hazard ratio, odds ratio, risk ratio, or relative risk** that wasn't computed by an external tool. These require regression / model fitting that ClickHouse does not do. 5. **Never report median overall survival from `AVG(OS_MONTHS)` or even `quantile(0.5)(OS_MONTHS)`.** Median OS requires Kaplan-Meier estimation, which handles censoring (`OS_STATUS = 0:LIVING` means the event hasn't happened yet). Naive medians/means over `OS_MONTHS` ignore censoring and are systematically wrong. The correct handoff: return the raw `(OS_MONTHS, OS_STATUS)` pairs (or descriptive counts: N events, N censored, follow-up range) and tell the user to run KM in R (`survival::survfit`) or Python (`lifelines.KaplanMeierFitter`), or use cBioPortal's Survival comparison. If asked to do any of (1)–(5), respond with the appropriate handoff template from the "Approved Response Templates" section below — do not produce the number. Ambiguous-Term Clarification (before any analysis) -------------------------------------------------- If the user's question uses an ambiguous term, ask before computing anything. Examples: - **"aggressive"** → could mean shorter overall survival, higher metastasis rate, higher grade/stage, higher TMB, or specific molecular features. Ask which definit … (9984 more chars) ◀ result {"result":[{"cancer_study_identifier":"brca_tcga","name":"Breast Invasive Carcinoma (TCGA, Firehose Legacy)","type_of_cancer_id":"brca","sample_count":1108,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=brca_tcga"},{"cancer_study_identifier":"brca_tcga_gdc","name":"Invasive Breast Carcinoma (TCGA GDC, 2025)","type_of_cancer_id":"brca","sample_count":1102,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=brca_tcga_gdc"},{"cancer_study_identifier":"brca_tcga_pan_can_atlas_2018","name":"Breast Invasive Carcinoma (TCGA, PanCancer Atlas)","type_of_cancer_id":"brca","sample_count":1084,"has_guide":true,"url":"https://www.cbioportal.org/study/summary?id=brca_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"brca_tcga_pub","name":"Breast Invasive Carcinoma (TCGA, Nature 2012)","type_of_cancer_id":"brca","sample_count":825,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=brca_tcga_pub"},{"cancer_study_identifier":"brca_tcga_pub2015","name":"Breast Invasive Carcinoma (TCGA, Cell 2015)","type_of_cancer_id":"brca","sample_count":818,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=brca_tcga_pub2015"},{"cancer_study_identifier":"mixed_msk_tcga_2021","name":"RAD51B Associated Mixed Cancers (MSK, NPJ Breast Cancer 2021)","type_of_cancer_id":"mixed","sample_count":17,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=mixed_msk_tcga_2021"}]} ═ answer (success) "TCGA Breast Cancer study" matches several studies in cBioPortal, and the counts will differ between them. Which one do you want? | Study | ID | Samples | |---|---|---| | [Breast Invasive Carcinoma (TCGA, PanCancer Atlas)](https://www.cbioportal.org/study/summary?id=brca_tcga_pan_can_atlas_2018) | `brca_tcga_pan_can_atlas_2018` | 1,084 | | [Breast Invasive Carcinoma (TCGA, Firehose Legacy)](https://www.cbioportal.org/study/summary?id=brca_tcga) | `brca_tcga` | 1,108 | | [Invasive Breast Carcinoma (TCGA GDC, 2025)](https://www.cbioportal.org/study/summary?id=brca_tcga_gdc) | `brca_tcga_gdc` | 1,102 | | [Breast Invasive Carcinoma (TCGA, Nature 2012)](https://www.cbioportal.org/study/summary?id=brca_tcga_pub) | `brca_tcga_pub` | 825 | | [Breast Invasive Carcinoma (TCGA, Cell 2015)](https://www.cbioportal.org/study/summary?id=brca_tcga_pub2015) | `brca_tcga_pub2015` | 818 | I'd suggest the **PanCancer Atlas** study. It is the standard TCGA reference cohort, and it has the largest sample set with a curated study guide. I can also run the analysis on all of them and report each separately. I need two more definitions from you: 1. **PTEN "deletion":** should I count only deep (homozygous) deletions, which is GISTIC = −2? Or should I include shallow/heterozygous loss (−1) as well? I'd default to deep deletions only. 2. **"Luminal A":** I'll use the PAM50 subtype label in the study's clinical data (`Luminal A` vs. other subtypes). Tell me if you'd rather define it another way, for example by receptor status. Once you confirm, I'll build the counts for these groups: - PIK3CA-mutated and PTEN-deleted - PIK3CA-mutated only, with no PTEN deep deletion - PIK3CA-mutated, regardless of PTEN status For each group I'll give the number of Luminal A patients out of the total, which is the proportion you asked about. I can't compute a p-value or odds ratio myself. I'll give you the 2×2 table and a link to run the test in cBioPortal Group Comparison, or R/Python code to run it yourself.