Q (haiku): In TCGA colorectal cancer, do patients with high PTK2 expression have different survival from patients with low PTK2 expression? In Colorectal Adenocarcinoma (TCGA, PanCancer Atlas), the Kaplan-Meier curves for patients with PTK2 mRNA above vs below the median look slightly apart, with the high-PTK2 group doing a little worse early on. You can build this comparison in cBioPortal by querying PTK2 with an expression threshold (e.g. `PTK2: EXP>0`) and opening the Comparison → Survival tab. Continue this conversation: reply to the user's new message. is the difference in survival significant? ▶ read_guide { "uri": "cbioportal://statistical-tests-guide" } ◀ result Statistical Test Selection Guide ================================ Purpose ------- This guide ensures the correct statistical test is selected before performing any group comparison, matching cBioPortal's own Group Comparison defaults. HARD RULES — NEVER FABRICATE A STATISTIC ---------------------------------------- ClickHouse cannot run statistical tests. The agent therefore must NEVER produce a derived statistic that is not a literal column value from a SQL result. Specifically: 1. **Never invent a p-value.** Not "p < 0.001", not "p ≈ 0.05", not any p-value. If the user asks "what is the p-value?", the answer is *"I can't compute that — here is the 2x2 contingency table (or group statistics). Run it in cBioPortal's Group Comparison tab, in R with `fisher.test(...)` / `wilcox.test(...)`, or in Python with `scipy.stats.fisher_exact(...)` / `mannwhitneyu(...)`."* 2. **Never claim mutual exclusivity (or co-occurrence) from a contingency table alone.** A 2x2 table is not a test. The shape "altered/not altered × group A/group B" needs Fisher's exact + a defined direction (odds ratio < 1 with significant p). Without that test, the agent presents the table and stops. Descriptive phrasing is also a claim: "largely/mostly mutually exclusive", "rarely co-occur", "tend to co-occur" are forbidden without the test — point to cBioPortal's Mutual Exclusivity tab instead. 3. **Never report a "median" that came from `AVG(...)` or any non-median aggregate.** "Median" and "mean" are different statistics; for skewed clinical distributions (especially survival) they differ substantially. Use ClickHouse's `quantile(0.5)(...)` for actual median, and label arithmetic averages as "mean", never "median". 4. **Never report a hazard ratio, odds ratio, risk ratio, or relative risk** that wasn't computed by an external tool. These require regression / model fitting that ClickHouse does not do. 5. **Never report median overall survival from `AVG(OS_MONTHS)` or even `quantile(0.5)(OS_MONTHS)`.** Median OS requires Kaplan-Meier estimation, which handles censoring (`OS_STATUS = 0:LIVING` means the event hasn't happened yet). Naive medians/means over `OS_MONTHS` ignore censoring and are systematically wrong. The correct handoff: return the raw `(OS_MONTHS, OS_STATUS)` pairs (or descriptive counts: N events, N censored, follow-up range) and tell the user to run KM in R (`survival::survfit`) or Python (`lifelines.KaplanMeierFitter`), or use cBioPortal's Survival comparison. If asked to do any of (1)–(5), respond with the appropriate handoff template from the "Approved Response Templates" section below — do not produce the number. Ambiguous-Term Clarification (before any analysis) -------------------------------------------------- If the user's question uses an ambiguous term, ask before computing anything. Examples: - **"aggressive"** → could mean shorter overall survival, higher metastasis rate, higher grade/stage, higher TMB, or specific molecular features. Ask which definit … (9984 more chars) ═ answer (success) Based on the statistical-tests guide, **I cannot compute the p-value for you** — survival comparisons require a Kaplan-Meier test with log-rank, which ClickHouse cannot perform. However, I can point you to where to get the answer: **Use cBioPortal's Group Comparison → Survival tab** (recommended): 1. Go to the [TCGA Colorectal Adenocarcinoma study](https://www.cbioportal.org/study/summary?id=coadread_tcga_pan_can_atlas_2018) 2. Filter or compare PTK2 expression high (above median) vs. low (below median) 3. Open the **Survival** tab — it will show Kaplan-Meier curves with the log-rank p-value and median OS per group **Alternatively, compute it yourself in R or Python:** - **R:** Use `survival::survfit()` with `Surv(OS_MONTHS, OS_STATUS=="1:DECEASED")` to fit KM curves, then `survdiff()` for the log-rank test - **Python:** Use `lifelines.KaplanMeierFitter` and `logrank_test()` from `lifelines.statistics` The rough visual impression from my earlier response (high-PTK2 slightly worse early on) suggests a possible difference, but without the actual log-rank test statistic and p-value, you cannot say whether that difference is statistically significant or just random variation in the sample.