Q (sonnet-5.5): What is the median survival time in the Pediatric Neuroblastoma study from TARGET? ▶ list_studies { "search": "neuroblastoma TARGET" } ▶ read_guide { "uri": "cbioportal://statistical-tests-guide" } ◀ result {"result":[{"cancer_study_identifier":"nbl_target_2018_pub","name":"Pediatric Neuroblastoma (TARGET, 2018)","type_of_cancer_id":"nbl","sample_count":1089,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=nbl_target_2018_pub"},{"cancer_study_identifier":"nbl_target_gdc","name":"Neuroblastoma (TARGET GDC, 2025)","type_of_cancer_id":"nbl","sample_count":533,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=nbl_target_gdc"},{"cancer_study_identifier":"nbl_msk_2023","name":"Pediatric Neuroblastoma (MSK, Nat Genet 2023)","type_of_cancer_id":"nbl","sample_count":223,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=nbl_msk_2023"}]} ◀ result Statistical Test Selection Guide ================================ Purpose ------- This guide ensures the correct statistical test is selected before performing any group comparison, matching cBioPortal's own Group Comparison defaults. HARD RULES — NEVER FABRICATE A STATISTIC ---------------------------------------- ClickHouse cannot run statistical tests. The agent therefore must NEVER produce a derived statistic that is not a literal column value from a SQL result. Specifically: 1. **Never invent a p-value.** Not "p < 0.001", not "p ≈ 0.05", not any p-value. If the user asks "what is the p-value?", the answer is *"I can't compute that — here is the 2x2 contingency table (or group statistics). Run it in cBioPortal's Group Comparison tab, in R with `fisher.test(...)` / `wilcox.test(...)`, or in Python with `scipy.stats.fisher_exact(...)` / `mannwhitneyu(...)`."* 2. **Never claim mutual exclusivity (or co-occurrence) from a contingency table alone.** A 2x2 table is not a test. The shape "altered/not altered × group A/group B" needs Fisher's exact + a defined direction (odds ratio < 1 with significant p). Without that test, the agent presents the table and stops. Descriptive phrasing is also a claim: "largely/mostly mutually exclusive", "rarely co-occur", "tend to co-occur" are forbidden without the test — point to cBioPortal's Mutual Exclusivity tab instead. 3. **Never report a "median" that came from `AVG(...)` or any non-median aggregate.** "Median" and "mean" are different statistics; for skewed clinical distributions (especially survival) they differ substantially. Use ClickHouse's `quantile(0.5)(...)` for actual median, and label arithmetic averages as "mean", never "median". 4. **Never report a hazard ratio, odds ratio, risk ratio, or relative risk** that wasn't computed by an external tool. These require regression / model fitting that ClickHouse does not do. 5. **Never report median overall survival from `AVG(OS_MONTHS)` or even `quantile(0.5)(OS_MONTHS)`.** Median OS requires Kaplan-Meier estimation, which handles censoring (`OS_STATUS = 0:LIVING` means the event hasn't happened yet). Naive medians/means over `OS_MONTHS` ignore censoring and are systematically wrong. The correct handoff: return the raw `(OS_MONTHS, OS_STATUS)` pairs (or descriptive counts: N events, N censored, follow-up range) and tell the user to run KM in R (`survival::survfit`) or Python (`lifelines.KaplanMeierFitter`), or use cBioPortal's Survival comparison. If asked to do any of (1)–(5), respond with the appropriate handoff template from the "Approved Response Templates" section below — do not produce the number. Ambiguous-Term Clarification (before any analysis) -------------------------------------------------- If the user's question uses an ambiguous term, ask before computing anything. Examples: - **"aggressive"** → could mean shorter overall survival, higher metastasis rate, higher grade/stage, higher TMB, or specific molecular features. Ask which definit … (9984 more chars) ▶ read_guide { "uri": "cbioportal://clinical-data-guide" } ▶ clickhouse_list_table_columns { "table": "clinical_data_derived" } ◀ result # Clinical Data Query Guide ## Overview Clinical data in cBioPortal is stored at both patient and sample levels. Understanding the distinction is crucial for accurate analysis. ## Data Organization ### Patient-Level vs Sample-Level Data - **Patient-level**: Demographics, overall survival, disease stage (stored once per patient) - **Sample-level**: Sample type, sequencing platform, purity (can have multiple per patient) ### Key Tables - `clinical_patient`: Patient-level clinical attributes - `clinical_sample`: Sample-level clinical attributes - `clinical_data_derived`: Pre-joined view combining both levels - `clinical_attribute_meta`: Metadata about available clinical attributes ## Recommended Approach: Use clinical_data_derived The `clinical_data_derived` table is pre-joined and optimized for most queries: ```sql -- Get clinical data for specific attributes SELECT sample_unique_id, patient_unique_id, attribute_name, attribute_value FROM clinical_data_derived WHERE cancer_study_identifier = 'your_study_id' AND attribute_name IN ('CANCER_TYPE', 'SAMPLE_TYPE', 'TMB_NONSYNONYMOUS'); ``` ## Clinical Attribute Discovery ### Use clinical_attribute_meta for Discovering Available Attributes **Always start here** to see what clinical attributes are available for a specific study: ```sql -- Discover available clinical attributes for a study SELECT attr_id, description, patient_attribute, cancer_study_id FROM clinical_attribute_meta WHERE cancer_study_id = ( SELECT cancer_study_id FROM cancer_study WHERE cancer_study_identifier = 'your_study_id' ) ORDER BY patient_attribute, attr_id; ``` **Key Fields:** - **attr_id**: matches attr_id in clinical_sample/clinical_patient tables - **description**: human-readable description of the attribute - **patient_attribute**: true = patient attribute, false = sample attribute - **cancer_study_id**: links to cancer_study table (filter by study) ## Attribute Semantics and Matching ### Case-Insensitive Matching for Attribute Values Clinical values are free text across studies and may differ only by case. For example, a controlled-looking value such as germline mutation status may appear as `GERMLINE`, `Germline`, or another case variant. When filtering `clinical_data_derived.attribute_value`, use case-insensitive matching unless you have already profiled the exact values in the target study: ```sql -- Correct: case-insensitive clinical value filter SELECT DISTINCT sample_unique_id, patient_unique_id FROM clinical_data_derived WHERE cancer_study_identifier = 'your_study_id' AND attribute_name = 'MUTATION_STATUS' AND upper(attribute_value) = 'GERMLINE'; ``` Do not write `attribute_value = 'GERMLINE'` without first checking all distinct values for that attribute in the study. ### Query the Requested Attribute, Not a Proxy Do not infer one clinical attribute from a related subtype or marker. Query the actual requested attribute when it … (17073 more chars) ◀ result {"columns":[{"name":"internal_id","type":"Int32"},{"name":"sample_unique_id","type":"String","comment":"Globally unique sample ID: cancer_study_identifier + \"_\" + sample.stable_id. Empty for patient-level attributes. Use this for sample filtering and joins."},{"name":"patient_unique_id","type":"String","comment":"Globally unique patient ID: cancer_study_identifier + \"_\" + patient.stable_id. Present for both sample and patient-level attributes."},{"name":"attribute_name","type":"LowCardinality(String)","comment":"Clinical attribute name (e.g., SAMPLE_TYPE, CANCER_TYPE, AGE, OS_MONTHS). Use with attribute_value for filtering. AGE may be floored or capped for de-identification (e.g. all children recorded as 18, or everyone 89+ recorded as 89 or 90): before age statistics check for a pile-up at the min/max, and if present compute age from DAYS_TO_BIRTH (-days / 365.25)."},{"name":"attribute_value","type":"String","comment":"Value of the clinical attribute (String). For SAMPLE_TYPE: Primary, Metastasis, Local Recurrence, Unknown. Missing values are empty strings, so use toFloat64OrNull(attribute_value) for numeric comparisons — CAST fails on them."},{"name":"cancer_study_identifier","type":"LowCardinality(String)"},{"name":"type","type":"LowCardinality(String)","comment":"Data level: \"sample\" for sample-level attributes (e.g., SAMPLE_TYPE), \"patient\" for patient-level attributes (e.g., AGE, OS_MONTHS)."}]} ▶ clickhouse_run_select_query { "query": "WITH os AS ( SELECT patient_unique_id, maxIf(toFloat64OrNull(attribute_value), attribute_name='OS_MONTHS') AS os_months, maxIf(attribute_value, attribute_name='OS_STATUS') AS os_status FROM clinical_data_derived WHERE cancer_study_identifier='nbl_target_2018_pub' AND attribute_name IN ('OS_MONTHS','OS_STATUS') GROUP BY patient_unique_id) SELECT count() n, countIf(startsWith(os_status,'1')) events, countIf(startsWith(os_status,'0')) censored, min(os_months) mn, max(os_months) mx FROM os WHERE os_months IS NOT NULL AND os_status != ''" } ▶ resolve_and_route { "studyIds": [ "nbl_target_2018_pub" ] } ◀ result {"rows":[{"n":1072,"events":397,"censored":675,"mn":1,"mx":183}]} ◀ result {"success":true,"message":"Found 1 study. Use the metadata to call the appropriate navigation tool(s).","data":{"totalCount":1,"studiesWithMetadata":[{"studyId":"nbl_target_2018_pub","name":"Pediatric Neuroblastoma (TARGET, 2018)","sampleCount":1089,"studyViewUrl":"https://www.cbioportal.org/study?id=nbl_target_2018_pub","metadata":{"clinicalAttributeIds":["AGE","AGE_IN_DAYS","ANALYSIS_COHORT","CANCER_TYPE","CANCER_TYPE_DETAILED","CNS_SITE_OF_RELAPSE","DIAGNOSIS","EFSCENS","EFS_TIME","ETHNICITY","FIRST_EVENT","FRACTION_GENOME_ALTERED","GRADE","ICDO","ICDO_SNOMED_DESCRIPTION","INSS_STAGE","LAST_FOLLOWUP_YEAR","MKI","MUTATION_COUNT","MYCN","ONCOTREE_CODE","OS_DAYS","OS_MONTHS","OS_STATUS","PERCENTAGE_NECROSIS","PERCENTAGE_TUMOR_PURITY","PERCENT_TUMOR_CELLS_RELAPSE","PERCENT_TUMOR_VS_STROMA","PLOIDY","PROTOCOL","RACE","RELAPSE_PERCENT_NECROSIS","RELAPSE_PERCENT_STROMA","RISK_GROUP","SAMPLE_COUNT","SCENS","SEX","SNOMED","SOMATIC_STATUS","STIME","TMB_NONSYNONYMOUS","TUMOR_PLOIDY","TUMOR_SAMPLE_HISTOLOGY","TUMOR_TISSUE_SITE","YEAR_OF_DIAGNOSIS"],"molecularProfileIds":["nbl_target_2018_pub_gistic","nbl_target_2018_pub_mrna","nbl_target_2018_pub_mrna_median_Zscores","nbl_target_2018_pub_mrna_median_all_sample_Zscores","nbl_target_2018_pub_mutations","nbl_target_2018_pub_rna_seq_mrna","nbl_target_2018_pub_rna_seq_mrna_median_Zscores","nbl_target_2018_pub_rna_seq_mrna_median_all_sample_Zscores"],"heatmapProfileIds":["nbl_target_2018_pub_mrna_median_Zscores","nbl_target_2018_pub_mrna_median_all_sample_Zscores","nbl_target_2018_pub_rna_seq_mrna_median_Zscores","nbl_target_2018_pub_rna_seq_mrna_median_all_sample_Zscores"],"availableComparisonTabs":["overlap","clinical","survival","alterations","mutations","mrna"]}}]}} ═ answer (success) I can't give you a median survival time for this cohort. Median survival has to come from a Kaplan-Meier estimate, because many patients are still alive at last follow-up and their survival times are censored. A plain median or average of the follow-up times would be wrong, so I didn't calculate one. I used **Pediatric Neuroblastoma (TARGET, 2018)**. If you meant the newer TARGET GDC neuroblastoma study, tell me and I'll pull the same summary for it. **Overall survival summary (patients with survival data)** | | Patients | |---|---| | Total | 1,072 | | Deceased (events) | 397 (37%) | | Living at last follow-up (censored) | 675 (63%) | | Follow-up range | 1 to 183 months | Fewer than half of the patients had an event, so the Kaplan-Meier median is probably **not reached** in this cohort. The survival curve may never fall to 50%. To get the actual estimate, you can: - Run Kaplan-Meier in R (`survival::survfit`) or Python (`lifelines.KaplanMeierFitter`), using `OS_MONTHS` and `OS_STATUS`. - Use the Survival tab in cBioPortal Group Comparison, splitting the cohort by an attribute such as `RISK_GROUP` or `INSS_STAGE`. [View the Pediatric Neuroblastoma (TARGET, 2018) study](https://www.cbioportal.org/study?id=nbl_target_2018_pub)