Q (sonnet): What fraction of patients were older than five when diagnosed according to the Pediatric Neuroblastoma study from TARGET? ▶ list_studies { "search": "neuroblastoma TARGET" } ▶ read_guide { "uri": "cbioportal://clinical-data-guide" } ◀ result {"result":[{"cancer_study_identifier":"nbl_target_2018_pub","name":"Pediatric Neuroblastoma (TARGET, 2018)","type_of_cancer_id":"nbl","sample_count":1089,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=nbl_target_2018_pub"},{"cancer_study_identifier":"nbl_target_gdc","name":"Neuroblastoma (TARGET GDC, 2025)","type_of_cancer_id":"nbl","sample_count":533,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=nbl_target_gdc"},{"cancer_study_identifier":"nbl_msk_2023","name":"Pediatric Neuroblastoma (MSK, Nat Genet 2023)","type_of_cancer_id":"nbl","sample_count":223,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=nbl_msk_2023"}]} ◀ result # Clinical Data Query Guide ## Overview Clinical data in cBioPortal is stored at both patient and sample levels. Understanding the distinction is crucial for accurate analysis. ## Data Organization ### Patient-Level vs Sample-Level Data - **Patient-level**: Demographics, overall survival, disease stage (stored once per patient) - **Sample-level**: Sample type, sequencing platform, purity (can have multiple per patient) ### Key Tables - `clinical_patient`: Patient-level clinical attributes - `clinical_sample`: Sample-level clinical attributes - `clinical_data_derived`: Pre-joined view combining both levels - `clinical_attribute_meta`: Metadata about available clinical attributes ## Recommended Approach: Use clinical_data_derived The `clinical_data_derived` table is pre-joined and optimized for most queries: ```sql -- Get clinical data for specific attributes SELECT sample_unique_id, patient_unique_id, attribute_name, attribute_value FROM clinical_data_derived WHERE cancer_study_identifier = 'your_study_id' AND attribute_name IN ('CANCER_TYPE', 'SAMPLE_TYPE', 'TMB_NONSYNONYMOUS'); ``` ## Clinical Attribute Discovery ### Use clinical_attribute_meta for Discovering Available Attributes **Always start here** to see what clinical attributes are available for a specific study: ```sql -- Discover available clinical attributes for a study SELECT attr_id, description, patient_attribute, cancer_study_id FROM clinical_attribute_meta WHERE cancer_study_id = ( SELECT cancer_study_id FROM cancer_study WHERE cancer_study_identifier = 'your_study_id' ) ORDER BY patient_attribute, attr_id; ``` **Key Fields:** - **attr_id**: matches attr_id in clinical_sample/clinical_patient tables - **description**: human-readable description of the attribute - **patient_attribute**: true = patient attribute, false = sample attribute - **cancer_study_id**: links to cancer_study table (filter by study) ## Attribute Semantics and Matching ### Case-Insensitive Matching for Attribute Values Clinical values are free text across studies and may differ only by case. For example, a controlled-looking value such as germline mutation status may appear as `GERMLINE`, `Germline`, or another case variant. When filtering `clinical_data_derived.attribute_value`, use case-insensitive matching unless you have already profiled the exact values in the target study: ```sql -- Correct: case-insensitive clinical value filter SELECT DISTINCT sample_unique_id, patient_unique_id FROM clinical_data_derived WHERE cancer_study_identifier = 'your_study_id' AND attribute_name = 'MUTATION_STATUS' AND upper(attribute_value) = 'GERMLINE'; ``` Do not write `attribute_value = 'GERMLINE'` without first checking all distinct values for that attribute in the study. ### Query the Requested Attribute, Not a Proxy Do not infer one clinical attribute from a related subtype or marker. Query the actual requested attribute when it … (17073 more chars) ▶ clickhouse_run_select_query { "query": "SELECT attr_id, description, patient_attribute FROM clinical_attribute_meta WHERE cancer_study_id = (SELECT cancer_study_id FROM cancer_study WHERE cancer_study_identifier = 'nbl_target_2018_pub') AND (upper(attr_id) LIKE '%AGE%' OR upper(attr_id) LIKE '%DAYS%')" } ◀ result {"rows":[{"attr_id":"AGE","description":"Age at which a condition or disease was first diagnosed.","patient_attribute":1},{"attr_id":"AGE_IN_DAYS","description":"Age at which a condition or disease was first diagnosed.","patient_attribute":1},{"attr_id":"INSS_STAGE","description":"Staging according to the International Neuroblastoma Staging System","patient_attribute":1},{"attr_id":"OS_DAYS","description":"Overall patient survival days","patient_attribute":1},{"attr_id":"PERCENTAGE_NECROSIS","description":"Percent necrosis.","patient_attribute":1},{"attr_id":"PERCENTAGE_TUMOR_PURITY","description":"Percent tumor purity.","patient_attribute":1}]} ▶ clickhouse_run_select_query { "query": " SELECT arrayMin(ages) AS min_age, arrayMax(ages) AS max_age, countEqual(ages, min_age) AS at_min, countEqual(ages, max_age) AS at_max, length(ages) AS patients FROM ( SELECT groupArray(toFloat64OrNull(attribute_value)) AS ages FROM clinical_data_derived WHERE cancer_study_identifier = 'nbl_target_2018_pub' AND attribute_name = 'AGE' AND toFloat64OrNull(attribute_value) IS NOT NULL ) " } ◀ result {"rows":[{"min_age":0,"max_age":21,"at_min":1,"at_max":2,"patients":1076}]} ▶ clickhouse_run_select_query { "query": " SELECT countIf(days > 1826.25) AS over_5, countIf(days <= 1826.25) AS five_or_under, count() AS total_with_age FROM ( SELECT patient_unique_id, toFloat64OrNull(attribute_value) AS days FROM clinical_data_derived WHERE cancer_study_identifier = 'nbl_target_2018_pub' AND attribute_name = 'AGE_IN_DAYS' AND toFloat64OrNull(attribute_value) IS NOT NULL ) " } ◀ result {"rows":[{"over_5":125,"five_or_under":951,"total_with_age":1076}]} ═ answer (success) `AGE` is not floored/capped in a problematic way here (min 0, max 21, no large pile-up), and `AGE_IN_DAYS` gives an exact age at diagnosis for 1,076 of the study's patients. Using 5 years = 1826.25 days as the cutoff: **Pediatric Neuroblastoma (TARGET, 2018)** — [nbl_target_2018_pub](https://www.cbioportal.org/study/summary?id=nbl_target_2018_pub) - Patients with age at diagnosis recorded: **1,076** - Older than 5 years at diagnosis: **125 patients (11.6%)** - 5 years old or younger: 951 patients (88.4%) So roughly **1 in 9 patients (~12%)** were older than five at diagnosis in this cohort — consistent with neuroblastoma being predominantly a disease of infancy/early childhood, with older age at diagnosis being a known adverse prognostic factor in the literature (this last point is general background, not a cBioPortal data field). You can explore this age distribution directly in the cohort view: [View age distribution in Pediatric Neuroblastoma (TARGET, 2018) Study View](https://www.cbioportal.org/study/summary?id=nbl_target_2018_pub) Note: there is also a newer, separate TARGET GDC neuroblastoma cohort (`nbl_target_gdc`, 533 samples) — let me know if you'd like that one checked as well, since "TARGET" could refer to either release.