Q (sonnet): In the Pediatric Neuroblastoma study from TARGET what is the survival difference and the corresponding statistical significance between patients who were older than four when diagnosed and the younger ones? ▶ read_guide { "uri": "cbioportal://statistical-tests-guide" } ▶ list_studies { "search": "neuroblastoma TARGET" } ◀ result Statistical Test Selection Guide ================================ Purpose ------- This guide ensures the correct statistical test is selected before performing any group comparison, matching cBioPortal's own Group Comparison defaults. HARD RULES — NEVER FABRICATE A STATISTIC ---------------------------------------- ClickHouse cannot run statistical tests. The agent therefore must NEVER produce a derived statistic that is not a literal column value from a SQL result. Specifically: 1. **Never invent a p-value.** Not "p < 0.001", not "p ≈ 0.05", not any p-value. If the user asks "what is the p-value?", the answer is *"I can't compute that — here is the 2x2 contingency table (or group statistics). Run it in cBioPortal's Group Comparison tab, in R with `fisher.test(...)` / `wilcox.test(...)`, or in Python with `scipy.stats.fisher_exact(...)` / `mannwhitneyu(...)`."* 2. **Never claim mutual exclusivity (or co-occurrence) from a contingency table alone.** A 2x2 table is not a test. The shape "altered/not altered × group A/group B" needs Fisher's exact + a defined direction (odds ratio < 1 with significant p). Without that test, the agent presents the table and stops. Descriptive phrasing is also a claim: "largely/mostly mutually exclusive", "rarely co-occur", "tend to co-occur" are forbidden without the test — point to cBioPortal's Mutual Exclusivity tab instead. 3. **Never report a "median" that came from `AVG(...)` or any non-median aggregate.** "Median" and "mean" are different statistics; for skewed clinical distributions (especially survival) they differ substantially. Use ClickHouse's `quantile(0.5)(...)` for actual median, and label arithmetic averages as "mean", never "median". 4. **Never report a hazard ratio, odds ratio, risk ratio, or relative risk** that wasn't computed by an external tool. These require regression / model fitting that ClickHouse does not do. 5. **Never report median overall survival from `AVG(OS_MONTHS)` or even `quantile(0.5)(OS_MONTHS)`.** Median OS requires Kaplan-Meier estimation, which handles censoring (`OS_STATUS = 0:LIVING` means the event hasn't happened yet). Naive medians/means over `OS_MONTHS` ignore censoring and are systematically wrong. The correct handoff: return the raw `(OS_MONTHS, OS_STATUS)` pairs (or descriptive counts: N events, N censored, follow-up range) and tell the user to run KM in R (`survival::survfit`) or Python (`lifelines.KaplanMeierFitter`), or use cBioPortal's Survival comparison. If asked to do any of (1)–(5), respond with the appropriate handoff template from the "Approved Response Templates" section below — do not produce the number. Ambiguous-Term Clarification (before any analysis) -------------------------------------------------- If the user's question uses an ambiguous term, ask before computing anything. Examples: - **"aggressive"** → could mean shorter overall survival, higher metastasis rate, higher grade/stage, higher TMB, or specific molecular features. Ask which definit … (9984 more chars) ◀ result {"result":[{"cancer_study_identifier":"nbl_target_2018_pub","name":"Pediatric Neuroblastoma (TARGET, 2018)","type_of_cancer_id":"nbl","sample_count":1089,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=nbl_target_2018_pub"},{"cancer_study_identifier":"nbl_target_gdc","name":"Neuroblastoma (TARGET GDC, 2025)","type_of_cancer_id":"nbl","sample_count":533,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=nbl_target_gdc"},{"cancer_study_identifier":"nbl_msk_2023","name":"Pediatric Neuroblastoma (MSK, Nat Genet 2023)","type_of_cancer_id":"nbl","sample_count":223,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=nbl_msk_2023"}]} ▶ read_guide { "uri": "cbioportal://clinical-data-guide" } ▶ clickhouse_list_table_columns { "table": "clinical_data_derived" } ◀ result # Clinical Data Query Guide ## Overview Clinical data in cBioPortal is stored at both patient and sample levels. Understanding the distinction is crucial for accurate analysis. ## Data Organization ### Patient-Level vs Sample-Level Data - **Patient-level**: Demographics, overall survival, disease stage (stored once per patient) - **Sample-level**: Sample type, sequencing platform, purity (can have multiple per patient) ### Key Tables - `clinical_patient`: Patient-level clinical attributes - `clinical_sample`: Sample-level clinical attributes - `clinical_data_derived`: Pre-joined view combining both levels - `clinical_attribute_meta`: Metadata about available clinical attributes ## Recommended Approach: Use clinical_data_derived The `clinical_data_derived` table is pre-joined and optimized for most queries: ```sql -- Get clinical data for specific attributes SELECT sample_unique_id, patient_unique_id, attribute_name, attribute_value FROM clinical_data_derived WHERE cancer_study_identifier = 'your_study_id' AND attribute_name IN ('CANCER_TYPE', 'SAMPLE_TYPE', 'TMB_NONSYNONYMOUS'); ``` ## Clinical Attribute Discovery ### Use clinical_attribute_meta for Discovering Available Attributes **Always start here** to see what clinical attributes are available for a specific study: ```sql -- Discover available clinical attributes for a study SELECT attr_id, description, patient_attribute, cancer_study_id FROM clinical_attribute_meta WHERE cancer_study_id = ( SELECT cancer_study_id FROM cancer_study WHERE cancer_study_identifier = 'your_study_id' ) ORDER BY patient_attribute, attr_id; ``` **Key Fields:** - **attr_id**: matches attr_id in clinical_sample/clinical_patient tables - **description**: human-readable description of the attribute - **patient_attribute**: true = patient attribute, false = sample attribute - **cancer_study_id**: links to cancer_study table (filter by study) ## Attribute Semantics and Matching ### Case-Insensitive Matching for Attribute Values Clinical values are free text across studies and may differ only by case. For example, a controlled-looking value such as germline mutation status may appear as `GERMLINE`, `Germline`, or another case variant. When filtering `clinical_data_derived.attribute_value`, use case-insensitive matching unless you have already profiled the exact values in the target study: ```sql -- Correct: case-insensitive clinical value filter SELECT DISTINCT sample_unique_id, patient_unique_id FROM clinical_data_derived WHERE cancer_study_identifier = 'your_study_id' AND attribute_name = 'MUTATION_STATUS' AND upper(attribute_value) = 'GERMLINE'; ``` Do not write `attribute_value = 'GERMLINE'` without first checking all distinct values for that attribute in the study. ### Query the Requested Attribute, Not a Proxy Do not infer one clinical attribute from a related subtype or marker. Query the actual requested attribute when it … (17073 more chars) ◀ result {"columns":[{"name":"internal_id","type":"Int32"},{"name":"sample_unique_id","type":"String","comment":"Globally unique sample ID: cancer_study_identifier + \"_\" + sample.stable_id. Empty for patient-level attributes. Use this for sample filtering and joins."},{"name":"patient_unique_id","type":"String","comment":"Globally unique patient ID: cancer_study_identifier + \"_\" + patient.stable_id. Present for both sample and patient-level attributes."},{"name":"attribute_name","type":"LowCardinality(String)","comment":"Clinical attribute name (e.g., SAMPLE_TYPE, CANCER_TYPE, AGE, OS_MONTHS). Use with attribute_value for filtering. AGE may be floored or capped for de-identification (e.g. all children recorded as 18, or everyone 89+ recorded as 89 or 90): before age statistics check for a pile-up at the min/max, and if present compute age from DAYS_TO_BIRTH (-days / 365.25)."},{"name":"attribute_value","type":"String","comment":"Value of the clinical attribute (String). For SAMPLE_TYPE: Primary, Metastasis, Local Recurrence, Unknown. Missing values are empty strings, so use toFloat64OrNull(attribute_value) for numeric comparisons — CAST fails on them."},{"name":"cancer_study_identifier","type":"LowCardinality(String)"},{"name":"type","type":"LowCardinality(String)","comment":"Data level: \"sample\" for sample-level attributes (e.g., SAMPLE_TYPE), \"patient\" for patient-level attributes (e.g., AGE, OS_MONTHS)."}]} ▶ clickhouse_run_select_query { "query": "SELECT attribute_name, count(), min(toFloat64OrNull(attribute_value)), max(toFloat64OrNull(attribute_value)) FROM clinical_data_derived WHERE cancer_study_identifier = 'nbl_target_2018_pub' AND attribute_name IN ('AGE', 'OS_MONTHS', 'OS_STATUS', 'DAYS_TO_BIRTH', 'AGE_AT_DIAGNOSIS', 'AGE_MONTHS') GROUP BY attribute_name" } ◀ result {"rows":[{"attribute_name":"AGE","count()":1076,"min(toFloat64OrNull(attribute_value))":0,"max(toFloat64OrNull(attribute_value))":21},{"attribute_name":"OS_STATUS","count()":1076},{"attribute_name":"OS_MONTHS","count()":1076,"min(toFloat64OrNull(attribute_value))":1,"max(toFloat64OrNull(attribute_value))":183}]} ▶ clickhouse_run_select_query { "query": "SELECT toFloat64OrNull(attribute_value) as age, count() as n FROM clinical_data_derived WHERE cancer_study_identifier = 'nbl_target_2018_pub' AND attribute_name = 'AGE' GROUP BY age ORDER BY age" } ◀ result {"rows":[{"age":0,"n":1},{"age":1,"n":226},{"age":2,"n":208},{"age":3,"n":222},{"age":4,"n":168},{"age":5,"n":126},{"age":6,"n":46},{"age":7,"n":26},{"age":8,"n":10},{"age":9,"n":10},{"age":10,"n":9},{"age":11,"n":4},{"age":12,"n":4},{"age":13,"n":3},{"age":14,"n":4},{"age":15,"n":2},{"age":16,"n":2},{"age":17,"n":1},{"age":18,"n":1},{"age":19,"n":1},{"age":21,"n":2}]} ▶ clickhouse_run_select_query { "query": "SELECT DISTINCT attribute_value FROM clinical_data_derived WHERE cancer_study_identifier = 'nbl_target_2018_pub' AND attribute_name = 'OS_STATUS'" } ◀ result {"rows":[{"attribute_value":"0:LIVING"},{"attribute_value":"1:DECEASED"},{}]} ▶ clickhouse_run_select_query { "query": "WITH patient_data AS ( SELECT patient_unique_id, maxIf(toFloat64OrNull(attribute_value), attribute_name = 'AGE') AS age, maxIf(toFloat64OrNull(attribute_value), attribute_name = 'OS_MONTHS') AS os_months, maxIf(attribute_value, attribute_name = 'OS_STATUS') AS os_status FROM clinical_data_derived WHERE cancer_study_identifier = 'nbl_target_2018_pub' AND attribute_name IN ('AGE', 'OS_MONTHS', 'OS_STATUS') GROUP BY patient_unique_id ) SELECT if(age > 4, 'Age > 4', 'Age <= 4') AS group_name, count() AS n_patients, countIf(startsWith(os_status, '1')) AS n_events, countIf(startsWith(os_status, '0')) AS n_censored, round(min(os_months),1) AS min_followup_months, round(max(os_months),1) AS max_followup_months, round(avg(os_months),1) AS mean_os_months FROM patient_data WHERE os_months IS NOT NULL AND os_status != '' AND age IS NOT NULL GROUP BY group_name ORDER BY group_name" } ◀ result {"rows":[{"group_name":"Age <= 4","n_patients":822,"n_events":276,"n_censored":546,"min_followup_months":1,"max_followup_months":183,"mean_os_months":58.2},{"group_name":"Age > 4","n_patients":250,"n_events":121,"n_censored":129,"min_followup_months":1,"max_followup_months":183,"mean_os_months":55.9}]} ▶ resolve_and_route { "studyIds": [ "nbl_target_2018_pub" ] } ◀ result {"success":true,"message":"Found 1 study. Use the metadata to call the appropriate navigation tool(s).","data":{"totalCount":1,"studiesWithMetadata":[{"studyId":"nbl_target_2018_pub","name":"Pediatric Neuroblastoma (TARGET, 2018)","sampleCount":1089,"studyViewUrl":"https://www.cbioportal.org/study?id=nbl_target_2018_pub","metadata":{"clinicalAttributeIds":["AGE","AGE_IN_DAYS","ANALYSIS_COHORT","CANCER_TYPE","CANCER_TYPE_DETAILED","CNS_SITE_OF_RELAPSE","DIAGNOSIS","EFSCENS","EFS_TIME","ETHNICITY","FIRST_EVENT","FRACTION_GENOME_ALTERED","GRADE","ICDO","ICDO_SNOMED_DESCRIPTION","INSS_STAGE","LAST_FOLLOWUP_YEAR","MKI","MUTATION_COUNT","MYCN","ONCOTREE_CODE","OS_DAYS","OS_MONTHS","OS_STATUS","PERCENTAGE_NECROSIS","PERCENTAGE_TUMOR_PURITY","PERCENT_TUMOR_CELLS_RELAPSE","PERCENT_TUMOR_VS_STROMA","PLOIDY","PROTOCOL","RACE","RELAPSE_PERCENT_NECROSIS","RELAPSE_PERCENT_STROMA","RISK_GROUP","SAMPLE_COUNT","SCENS","SEX","SNOMED","SOMATIC_STATUS","STIME","TMB_NONSYNONYMOUS","TUMOR_PLOIDY","TUMOR_SAMPLE_HISTOLOGY","TUMOR_TISSUE_SITE","YEAR_OF_DIAGNOSIS"],"molecularProfileIds":["nbl_target_2018_pub_gistic","nbl_target_2018_pub_mrna","nbl_target_2018_pub_mrna_median_Zscores","nbl_target_2018_pub_mrna_median_all_sample_Zscores","nbl_target_2018_pub_mutations","nbl_target_2018_pub_rna_seq_mrna","nbl_target_2018_pub_rna_seq_mrna_median_Zscores","nbl_target_2018_pub_rna_seq_mrna_median_all_sample_Zscores"],"heatmapProfileIds":["nbl_target_2018_pub_mrna_median_Zscores","nbl_target_2018_pub_mrna_median_all_sample_Zscores","nbl_target_2018_pub_rna_seq_mrna_median_Zscores","nbl_target_2018_pub_rna_seq_mrna_median_all_sample_Zscores"],"availableComparisonTabs":["overlap","clinical","survival","alterations","mutations","mrna"]}}]}} ▶ get_studyviewfilter_options { "studyId": "nbl_target_2018_pub", "attributeIds": [ "AGE" ] } ◀ result {"success":true,"message":"Successfully retrieved StudyView filter options","data":{"studyId":"nbl_target_2018_pub","attributes":[{"attributeId":"AGE","displayName":"Diagnosis Age","description":"Age at which a condition or disease was first diagnosed.","datatype":"NUMBER","continuous":true,"bins":[{"end":2,"count":435},{"start":3,"end":3,"count":222},{"start":4,"end":4,"count":168},{"start":4,"count":251}]}]}} ▶ navigate_to_group_comparison { "studyIds": [ "nbl_target_2018_pub" ], "tab": "survival", "groups": [ { "name": "Age <= 4", "studyViewFilter": { "clinicalDataFilters": [ { "attributeId": "AGE", "values": [ { "start": 0, "end": 4 } ] } ] } }, { "name": "Age > 4", "studyViewFilter": { "clinicalDataFilters": [ { "attributeId": "AGE", "values": [ { "start": 4.0000001 } ] } ] } } ] } ◀ result {"success":true,"message":"Navigating to https://www.cbioportal.org/comparison/survival?comparisonId=6ab5d0c0e497aad53593c28a","url":"https://www.cbioportal.org/comparison/survival?comparisonId=6ab5d0c0e497aad53593c28a","data":{"description":"Group comparison (2 custom groups)","studies":["nbl_target_2018_pub"],"totalGroups":2,"groups":[{"name":"Age <= 4","sampleCount":835},{"name":"Age > 4","sampleCount":253}],"studyViewUrl":"https://www.cbioportal.org/study?id=nbl_target_2018_pub","groupUrls":[{"groupName":"Age <= 4","url":"https://www.cbioportal.org/study?id=nbl_target_2018_pub#filterJson=%7B%22clinicalDataFilters%22%3A%5B%7B%22attributeId%22%3A%22AGE%22%2C%22values%22%3A%5B%7B%22start%22%3A0%2C%22end%22%3A4%7D%5D%7D%5D%7D"},{"groupName":"Age > 4","url":"https://www.cbioportal.org/study?id=nbl_target_2018_pub#filterJson=%7B%22clinicalDataFilters%22%3A%5B%7B%22attributeId%22%3A%22AGE%22%2C%22values%22%3A%5B%7B%22start%22%3A4.0000001%7D%5D%7D%5D%7D"}],"pageDescription":"Kaplan-Meier survival analysis comparing Age <= 4 (835) vs Age > 4 (253)."}} ═ answer (success) ## Pediatric Neuroblastoma (TARGET, 2018) — Age at Diagnosis and Overall Survival **Study:** [Pediatric Neuroblastoma (TARGET, 2018)](https://www.cbioportal.org/study/summary?id=nbl_target_2018_pub) — `nbl_target_2018_pub`, 1,089 samples. (Note: TARGET also has a separate, non-identical "Neuroblastoma (TARGET GDC, 2025)" cohort — let me know if you'd like that one instead or in addition.) I split patients at age 4 years (a standard neuroblastoma risk-stratification cutoff), using diagnosis age recorded per patient (patient-level count, n = 1,076 with age and survival data). ### Descriptive summary (patients, not samples) | Group | N patients | Events (deceased) | Censored (living) | Follow-up range (months) | |---|---|---|---|---| | Age ≤ 4 | 822 | 276 (33.6%) | 546 | 1 – 183 | | Age > 4 | 250 | 121 (48.4%) | 129 | 1 – 183 | A larger share of the older-at-diagnosis group had a death event recorded (48.4% vs 33.6%), suggestive of worse survival in that group — but this is a raw event-rate comparison, not a survival-curve or significance estimate. Since fewer than half of the ≤4 group experienced the event, the Kaplan–Meier median survival time is likely **not reached** for that group. **I cannot compute the actual survival difference or a p-value here** — that requires Kaplan-Meier estimation (to properly handle patients still alive/censored) plus a log-rank test, which the underlying database cannot run. The link below opens cBioPortal's own Survival Comparison tool for this exact split, which will render the KM curves and report the log-rank p-value directly: [Compare Overall Survival: Age ≤4 vs Age >4 in Pediatric Neuroblastoma (TARGET, 2018)](https://www.cbioportal.org/comparison/survival?comparisonId=6ab5d0c0e497aad53593c28a) - Group "Age ≤ 4": 835 samples - Group "Age > 4": 253 samples You can also run it yourself in R (`survival::survfit(Surv(OS_MONTHS, OS_STATUS=="1:DECEASED") ~ age_group) `+ `survdiff()` for the log-rank p-value) or Python (`lifelines.KaplanMeierFitter` + `lifelines.statistics.logrank_test`) using the raw `OS_MONTHS`/`OS_STATUS`/`AGE` fields from this study.