Q (haiku): calculate median expression of ceacam5, itgb6, muc2, tpgb and muc1 mRNA in pancreatic cancer ▶ read_guide { "uri": "cbioportal://clinical-data-guide" } ▶ resolve_and_route { "studyKeywords": [ "pancreatic", "cancer" ] } ◀ result # Clinical Data Query Guide ## Overview Clinical data in cBioPortal is stored at both patient and sample levels. Understanding the distinction is crucial for accurate analysis. ## Data Organization ### Patient-Level vs Sample-Level Data - **Patient-level**: Demographics, overall survival, disease stage (stored once per patient) - **Sample-level**: Sample type, sequencing platform, purity (can have multiple per patient) ### Key Tables - `clinical_patient`: Patient-level clinical attributes - `clinical_sample`: Sample-level clinical attributes - `clinical_data_derived`: Pre-joined view combining both levels - `clinical_attribute_meta`: Metadata about available clinical attributes ## Recommended Approach: Use clinical_data_derived The `clinical_data_derived` table is pre-joined and optimized for most queries: ```sql -- Get clinical data for specific attributes SELECT sample_unique_id, patient_unique_id, attribute_name, attribute_value FROM clinical_data_derived WHERE cancer_study_identifier = 'your_study_id' AND attribute_name IN ('CANCER_TYPE', 'SAMPLE_TYPE', 'TMB_NONSYNONYMOUS'); ``` ## Clinical Attribute Discovery ### Use clinical_attribute_meta for Discovering Available Attributes **Always start here** to see what clinical attributes are available for a specific study: ```sql -- Discover available clinical attributes for a study SELECT attr_id, description, patient_attribute, cancer_study_id FROM clinical_attribute_meta WHERE cancer_study_id = ( SELECT cancer_study_id FROM cancer_study WHERE cancer_study_identifier = 'your_study_id' ) ORDER BY patient_attribute, attr_id; ``` **Key Fields:** - **attr_id**: matches attr_id in clinical_sample/clinical_patient tables - **description**: human-readable description of the attribute - **patient_attribute**: true = patient attribute, false = sample attribute - **cancer_study_id**: links to cancer_study table (filter by study) ## Attribute Semantics and Matching ### Case-Insensitive Matching for Attribute Values Clinical values are free text across studies and may differ only by case. For example, a controlled-looking value such as germline mutation status may appear as `GERMLINE`, `Germline`, or another case variant. When filtering `clinical_data_derived.attribute_value`, use case-insensitive matching unless you have already profiled the exact values in the target study: ```sql -- Correct: case-insensitive clinical value filter SELECT DISTINCT sample_unique_id, patient_unique_id FROM clinical_data_derived WHERE cancer_study_identifier = 'your_study_id' AND attribute_name = 'MUTATION_STATUS' AND upper(attribute_value) = 'GERMLINE'; ``` Do not write `attribute_value = 'GERMLINE'` without first checking all distinct values for that attribute in the study. ### Query the Requested Attribute, Not a Proxy Do not infer one clinical attribute from a related subtype or marker. Query the actual requested attribute when it … (17073 more chars) ◀ result {"success":true,"message":"Found 9 matching studies (top 5 with full metadata, rest with basic info). Pick the best match (prefer TCGA → prefer PanCancer Atlas), use its metadata to call the appropriate navigation tool(s) and generate URLs immediately. Present other studies as alternatives.","data":{"totalCount":9,"studiesWithMetadata":[{"studyId":"msk_met_2021","name":"MSK MetTropism (MSK, Cell 2021)","sampleCount":25775,"studyViewUrl":"https://www.cbioportal.org/study?id=msk_met_2021","metadata":{"clinicalAttributeIds":["AGE_AT_DEATH","AGE_AT_EVIDENCE_OF_METS","AGE_AT_LAST_CONTACT","AGE_AT_SEQUENCING","AGE_AT_SURGERY","CANCER_TYPE","CANCER_TYPE_DETAILED","DMETS_DX_ADRENAL_GLAND","DMETS_DX_BILIARY_TRACT","DMETS_DX_BLADDER_UT","DMETS_DX_BONE","DMETS_DX_BOWEL","DMETS_DX_BREAST","DMETS_DX_CNS_BRAIN","DMETS_DX_DIST_LN","DMETS_DX_FEMALE_GENITAL","DMETS_DX_HEAD_NECK","DMETS_DX_INTRA_ABDOMINAL","DMETS_DX_KIDNEY","DMETS_DX_LIVER","DMETS_DX_LUNG","DMETS_DX_MALE_GENITAL","DMETS_DX_MEDIASTINUM","DMETS_DX_OVARY","DMETS_DX_PLEURA","DMETS_DX_PNS","DMETS_DX_SKIN","DMETS_DX_UNSPECIFIED","FGA","FRACTION_GENOME_ALTERED","GENE_PANEL","IS_DIST_MET_MAPPED","METASTATIC_SITE","MET_COUNT","MET_SITE_COUNT","MSI_SCORE","MSI_TYPE","MUTATION_COUNT","ONCOTREE_CODE","ORGAN_SYSTEM","OS_MONTHS","OS_STATUS","PRIMARY_SITE","RACE","SAMPLE_COUNT","SAMPLE_COVERAGE","SAMPLE_TYPE","SEX","SUBTYPE","SUBTYPE_ABBREVIATION","TMB_NONSYNONYMOUS","TUMOR_PURITY"],"molecularProfileIds":["msk_met_2021_cna","msk_met_2021_mutations","msk_met_2021_structural_variants"],"availableComparisonTabs":["overlap","clinical","survival","alterations","mutations"]}},{"studyId":"pancreas_ctdna_msk_2025","name":"Pancreatic Cancer cfDNA (MSK, J Natl Cancer Inst 2025)","sampleCount":412,"studyViewUrl":"https://www.cbioportal.org/study?id=pancreas_ctdna_msk_2025","metadata":{"clinicalAttributeIds":["ADRENAL_GLANDS","BMI_FIRST_RESULT","BONE","CANCER_TYPE","CANCER_TYPE_DETAILED","CNS_BRAIN","ECOG_KPS","ETHNICITY","GENDER","GENE_PANEL","GLEASON_FIRST_REPORTED","GLEASON_HIGHEST_REPORTED","HISTORY_OF_D_MMR","HISTORY_OF_PDL1","INSTITUTE","INTRA_ABDOMINAL","LIVER","LUNG","LYMPH_NODES","METASTATIC_SITE","MUTATION_COUNT","NEW_LOC_ADVANCED_UNTREATED","NEW_LOC_UNTREATED","NEW_MET_DX_UNTREATED","NUM_ICDO_DX","ONCOTREE_CODE","OS_MONTHS","OS_STATUS","OTHER","PLEURA","PRIMARY_SITE","PRIOR_MED_TO_MSK","RACE","REPRODUCTIVE_ORGANS","SAMPLE_CLASS","SAMPLE_COUNT","SAMPLE_TYPE","SOMATIC_STATUS","STAGE_AT_DX","STAGE_HIGHEST_RECORDED","TOBACCO","YOST_INDEX_IMPUTED_MEDIAN"],"molecularProfileIds":["pancreas_ctdna_msk_2025_gistic","pancreas_ctdna_msk_2025_mutations","pancreas_ctdna_msk_2025_structural_variants"],"availableComparisonTabs":["overlap","clinical","survival","alterations","mutations"]}},{"studyId":"pancreas_msk_2024","name":"Pancreatic Cancer (MSK, Cancer Cell 2024)","sampleCount":395,"studyViewUrl":"https://www.cbioportal.org/study?id=pancreas_msk_2024","metadata":{"clinicalAttributeIds":["CANCER_TYPE","CANCER_TYPE_DETAILED", … (5518 more chars) ▶ clickhouse_list_table_columns { "table": "genomic_event_derived" } ▶ clickhouse_list_tables {} ◀ result {"columns":[{"name":"sample_unique_id","type":"String"},{"name":"hugo_gene_symbol","type":"String","comment":"HUGO gene symbol (e.g., TP53, KRAS, BRAF). Use for gene-specific queries."},{"name":"entrez_gene_id","type":"Int32"},{"name":"gene_panel_stable_id","type":"LowCardinality(String)"},{"name":"cancer_study_identifier","type":"LowCardinality(String)"},{"name":"genetic_profile_stable_id","type":"LowCardinality(String)"},{"name":"variant_type","type":"LowCardinality(String)","comment":"Type of genomic event: \"mutation\" for SNVs/indels, \"cna\" for copy number alterations, \"structural_variant\" for SVs. Always filter by this."},{"name":"mutation_variant","type":"String","comment":"Protein change notation (e.g., p.V600E, p.R175H). Use for specific variant queries. \"NA\" for non-mutation events."},{"name":"mutation_type","type":"LowCardinality(String)"},{"name":"mutation_status","type":"LowCardinality(String)","comment":"For mutations: Somatic, Germline, UNKNOWN, or UNCALLED. Filter mutation_status != \"UNCALLED\" to exclude uncertain calls. Include all other statuses."},{"name":"driver_filter","type":"LowCardinality(String)"},{"name":"driver_filter_annotation","type":"String"},{"name":"driver_tiers_filter","type":"LowCardinality(String)"},{"name":"driver_tiers_filter_annotation","type":"String"},{"name":"cna_alteration","type":"Nullable(Int8)","comment":"Copy number alteration: only 2 = amplification (AMP) and -2 = deep deletion (HOMDEL) are stored. Shallow deletion (-1), diploid (0) and gain (1) are not in this table; query genetic_alteration_derived WHERE profile_type = 'gistic' (alteration_value is a String, e.g. '-1'). NULL for non-CNA events."},{"name":"cna_cytoband","type":"String"},{"name":"sv_event_info","type":"String"},{"name":"patient_unique_id","type":"String"},{"name":"off_panel","type":"Bool","comment":"Boolean: 1 = mutation outside gene panel coverage (off-panel), 0 = within panel (on-panel). Filter off_panel = 0 for reliable frequency calculations."}]} ◀ result {"tables":[{"name":"allele_specific_copy_number"},{"name":"alteration_driver_annotation"},{"name":"cancer_study"},{"name":"cancer_study_query_preferences"},{"name":"cancer_study_tags"},{"name":"cbioportal_sequence_state"},{"name":"clinical_attribute_counts"},{"name":"clinical_attribute_meta"},{"name":"clinical_data_derived"},{"name":"clinical_event"},{"name":"clinical_event_data"},{"name":"clinical_event_data_derived"},{"name":"clinical_event_derived"},{"name":"clinical_patient"},{"name":"clinical_sample"},{"name":"cna_event"},{"name":"cna_panel_gene_coverage"},{"name":"cna_wes_coverage"},{"name":"co_altered_genes_in_study"},{"name":"copy_number_seg"},{"name":"copy_number_seg_file"},{"name":"gene"},{"name":"gene_alias"},{"name":"gene_alteration_frequency_by_cancer_type"},{"name":"gene_cna_distribution_in_study"},{"name":"gene_mutation_frequency_by_cancer_type"},{"name":"gene_mutation_frequency_in_studies"},{"name":"gene_mutation_frequency_in_study"},{"name":"gene_mutation_variants_in_study"},{"name":"gene_pair_coexpression"},{"name":"gene_panel"},{"name":"gene_panel_list"},{"name":"gene_panel_to_gene_derived"},{"name":"generic_assay_data_derived"},{"name":"generic_assay_meta_derived"},{"name":"generic_assay_profile_entity_derived"},{"name":"generic_entity_properties"},{"name":"geneset"},{"name":"geneset_gene"},{"name":"geneset_hierarchy_leaf"},{"name":"geneset_hierarchy_node"},{"name":"genetic_alteration"},{"name":"genetic_alteration_derived"},{"name":"genetic_entity"},{"name":"genetic_profile"},{"name":"genetic_profile_link"},{"name":"genetic_profile_samples"},{"name":"genomic_event_derived"},{"name":"gistic"},{"name":"gistic_to_gene"},{"name":"info"},{"name":"mut_sig"},{"name":"mutation"},{"name":"mutation_count_by_keyword"},{"name":"mutation_derived"},{"name":"mutation_event"},{"name":"mutation_panel_gene_coverage"},{"name":"mutation_wes_coverage"},{"name":"patient"},{"name":"reference_genome"},{"name":"reference_genome_gene"},{"name":"resource_definition"},{"name":"resource_patient"},{"name":"resource_sample"},{"name":"resource_study"},{"name":"sample"},{"name":"sample_cna_event"},{"name":"sample_derived"},{"name":"sample_list"},{"name":"sample_list_list"},{"name":"sample_profile"},{"name":"sample_to_gene_panel_derived"},{"name":"structural_variant"},{"name":"sv_panel_gene_coverage"},{"name":"sv_wes_coverage"},{"name":"top_cna_genes_in_study"},{"name":"top_mutated_genes_in_cohort"},{"name":"top_mutated_genes_in_study"},{"name":"top_sv_genes_in_study"},{"name":"treatment_counts_in_study"},{"name":"treatment_regimens_in_study"},{"name":"type_of_cancer"}]} ▶ clickhouse_list_table_columns { "table": "genetic_alteration_derived" } ◀ result {"columns":[{"name":"sample_unique_id","type":"String"},{"name":"cancer_study_identifier","type":"LowCardinality(String)"},{"name":"hugo_gene_symbol","type":"String"},{"name":"profile_type","type":"LowCardinality(String)"},{"name":"alteration_value","type":"Nullable(String)"}]} ▶ clickhouse_run_select_query { "query": " SELECT hugo_gene_symbol, profile_type, COUNT(DISTINCT sample_unique_id) AS samples_with_data, round(quantile(0.5)(toFloat64OrNull(alteration_value)), 2) AS median_expression, round(avg(toFloat64OrNull(alteration_value)), 2) AS mean_expression, round(min(toFloat64OrNull(alteration_value)), 2) AS min_expression, round(max(toFloat64OrNull(alteration_value)), 2) AS max_expression FROM genetic_alteration_derived WHERE cancer_study_identifier = 'paad_tcga_pan_can_atlas_2018' AND hugo_gene_symbol IN ('CEACAM5', 'ITGB6', 'MUC2', 'TPGB', 'MUC1') AND profile_type LIKE '%mrna%' AND alteration_value IS NOT NULL AND alteration_value != '' GROUP BY hugo_gene_symbol, profile_type ORDER BY hugo_gene_symbol ", "max_rows": 100 } ◀ result {"rows":[{"hugo_gene_symbol":"CEACAM5","profile_type":"rna_seq_v2_mrna","samples_with_data":177,"median_expression":10049.7,"mean_expression":23357.69,"min_expression":0,"max_expression":218618},{"hugo_gene_symbol":"CEACAM5","profile_type":"rna_seq_v2_mrna_median_Zscores","samples_with_data":177,"median_expression":-0.32,"mean_expression":0.09,"min_expression":-0.64,"max_expression":6.21},{"hugo_gene_symbol":"CEACAM5","profile_type":"rna_seq_v2_mrna_median_all_sample_Zscores","samples_with_data":177,"median_expression":0.25,"mean_expression":-0.06,"min_expression":-3.69,"max_expression":1.57},{"hugo_gene_symbol":"ITGB6","profile_type":"rna_seq_v2_mrna_median_Zscores","samples_with_data":177,"median_expression":-0.28,"mean_expression":0.03,"min_expression":-1.05,"max_expression":4.53},{"hugo_gene_symbol":"ITGB6","profile_type":"rna_seq_v2_mrna","samples_with_data":177,"median_expression":3740.15,"mean_expression":5204.67,"min_expression":0,"max_expression":26923.3},{"hugo_gene_symbol":"ITGB6","profile_type":"rna_seq_v2_mrna_median_all_sample_Zscores","samples_with_data":177,"median_expression":0.1,"mean_expression":-0.11,"min_expression":-6.26,"max_expression":1.62},{"hugo_gene_symbol":"MUC1","profile_type":"rna_seq_v2_mrna_median_Zscores","samples_with_data":177,"median_expression":-0.02,"mean_expression":0.31,"min_expression":-1.29,"max_expression":4.9},{"hugo_gene_symbol":"MUC1","profile_type":"rna_seq_v2_mrna","samples_with_data":177,"median_expression":14799,"mean_expression":18639.05,"min_expression":30.86,"max_expression":72349.5},{"hugo_gene_symbol":"MUC1","profile_type":"rna_seq_v2_mrna_median_all_sample_Zscores","samples_with_data":177,"median_expression":0.2,"mean_expression":0,"min_expression":-4.43,"max_expression":1.39},{"hugo_gene_symbol":"MUC2","profile_type":"rna_seq_v2_mrna","samples_with_data":177,"median_expression":6.53,"mean_expression":7785.51,"min_expression":0,"max_expression":507191},{"hugo_gene_symbol":"MUC2","profile_type":"rna_seq_v2_mrna_median_Zscores","samples_with_data":177,"median_expression":-0.27,"mean_expression":4.15,"min_expression":-0.28,"max_expression":287.71},{"hugo_gene_symbol":"MUC2","profile_type":"rna_seq_v2_mrna_median_all_sample_Zscores","samples_with_data":177,"median_expression":-0.45,"mean_expression":-0.14,"min_expression":-1.22,"max_expression":3.83}]} ▶ clickhouse_run_select_query { "query": " SELECT DISTINCT hugo_gene_symbol FROM genetic_alteration_derived WHERE cancer_study_identifier = 'paad_tcga_pan_can_atlas_2018' AND (upper(hugo_gene_symbol) LIKE '%TPGB%' OR upper(hugo_gene_symbol) LIKE '%TPBG%') ORDER BY hugo_gene_symbol ", "max_rows": 20 } ◀ result {"rows":[{"hugo_gene_symbol":"TPBG"},{"hugo_gene_symbol":"TPBGL"}]} ▶ clickhouse_run_select_query { "query": " SELECT hugo_gene_symbol, profile_type, COUNT(DISTINCT sample_unique_id) AS samples_with_data, round(quantile(0.5)(toFloat64OrNull(alteration_value)), 2) AS median_expression, round(avg(toFloat64OrNull(alteration_value)), 2) AS mean_expression, round(min(toFloat64OrNull(alteration_value)), 2) AS min_expression, round(max(toFloat64OrNull(alteration_value)), 2) AS max_expression FROM genetic_alteration_derived WHERE cancer_study_identifier = 'paad_tcga_pan_can_atlas_2018' AND hugo_gene_symbol = 'TPBG' AND profile_type LIKE '%mrna%' AND alteration_value IS NOT NULL AND alteration_value != '' GROUP BY hugo_gene_symbol, profile_type ORDER BY hugo_gene_symbol ", "max_rows": 100 } ◀ result {"rows":[{"hugo_gene_symbol":"TPBG","profile_type":"rna_seq_v2_mrna_median_all_sample_Zscores","samples_with_data":177,"median_expression":0.14,"mean_expression":0,"min_expression":-6.56,"max_expression":1.74},{"hugo_gene_symbol":"TPBG","profile_type":"rna_seq_v2_mrna_median_Zscores","samples_with_data":177,"median_expression":-0.03,"mean_expression":0.06,"min_expression":-2.21,"max_expression":3.79},{"hugo_gene_symbol":"TPBG","profile_type":"rna_seq_v2_mrna","samples_with_data":177,"median_expression":856.76,"mean_expression":894.89,"min_expression":11.71,"max_expression":2344.44}]} ▶ navigate_to_results_view { "studyIds": [ "paad_tcga_pan_can_atlas_2018" ], "genes": [ "CEACAM5", "ITGB6", "MUC2", "TPBG", "MUC1" ], "tab": "oncoprint", "oncoprintHeatmapTracks": [ { "molecularProfileId": "paad_tcga_pan_can_atlas_2018_rna_seq_v2_mrna_median_Zscores", "entities": [ "CEACAM5", "ITGB6", "MUC2", "TPBG", "MUC1" ] } ] } ◀ result {"success":true,"message":"Navigating to https://www.cbioportal.org/results/oncoprint?cancer_study_list=paad_tcga_pan_can_atlas_2018&gene_list=CEACAM5%0AITGB6%0AMUC2%0ATPBG%0AMUC1&case_set_id=paad_tcga_pan_can_atlas_2018_all&heatmap_track_groups=paad_tcga_pan_can_atlas_2018_rna_seq_v2_mrna_median_Zscores%2CCEACAM5%2CITGB6%2CMUC2%2CTPBG%2CMUC1","url":"https://www.cbioportal.org/results/oncoprint?cancer_study_list=paad_tcga_pan_can_atlas_2018&gene_list=CEACAM5%0AITGB6%0AMUC2%0ATPBG%0AMUC1&case_set_id=paad_tcga_pan_can_atlas_2018_all&heatmap_track_groups=paad_tcga_pan_can_atlas_2018_rna_seq_v2_mrna_median_Zscores%2CCEACAM5%2CITGB6%2CMUC2%2CTPBG%2CMUC1","data":{"studyIds":["paad_tcga_pan_can_atlas_2018"],"studies":[{"studyId":"paad_tcga_pan_can_atlas_2018","name":"Pancreatic Adenocarcinoma (TCGA, PanCancer Atlas)","sampleCount":184}],"genes":["CEACAM5","ITGB6","MUC2","TPBG","MUC1"],"caseSetId":"paad_tcga_pan_can_atlas_2018_all","oncoprintHeatmapTracks":[{"molecularProfileId":"paad_tcga_pan_can_atlas_2018_rna_seq_v2_mrna_median_Zscores","entities":["CEACAM5","ITGB6","MUC2","TPBG","MUC1"]}],"pageDescription":"Visual summary of alterations in queried genes across samples."}} ═ answer (success) ## Summary: Median mRNA Expression in Pancreatic Adenocarcinoma (TCGA) **Study:** Pancreatic Adenocarcinoma (TCGA, PanCancer Atlas) — 184 samples (177 with expression data) ### Raw mRNA Expression Levels (rna_seq_v2_mrna) | Gene | Median | Mean | Min | Max | |------|--------|------|-----|-----| | **MUC1** | **14,799** | 18,639 | 30.86 | 72,350 | | **CEACAM5** | **10,050** | 23,358 | 0 | 218,618 | | **ITGB6** | **3,740** | 5,205 | 0 | 26,923 | | **TPBG** | **857** | 895 | 11.71 | 2,344 | | **MUC2** | **6.53** | 7,786 | 0 | 507,191 | **Note:** MUC2 has the lowest median expression but high outlier values, indicating a small subset of samples with very high expression. ### Normalized Expression (Z-scores, median Zscores) For comparison across samples, the normalized z-score values show: | Gene | Median Z-score | Mean Z-score | |------|-----------------|--------------| | CEACAM5 | −0.32 | 0.09 | | ITGB6 | −0.28 | 0.03 | | MUC1 | −0.02 | 0.31 | | TPBG | −0.03 | 0.06 | | MUC2 | −0.27 | 4.15 | **Interpretation:** MUC1, ITGB6, and CEACAM5 are the most abundantly expressed genes. MUC2 expression is typically low but shows extreme heterogeneity (likely driven by mucinous vs ductal subtypes). --- ### Visualize in cBioPortal [View mRNA heatmap with all 5 genes](https://www.cbioportal.org/results/oncoprint?cancer_study_list=paad_tcga_pan_can_atlas_2018&gene_list=CEACAM5%0AITGB6%0AMUC2%0ATPBG%0AMUC1&case_set_id=paad_tcga_pan_can_atlas_2018_all&heatmap_track_groups=paad_tcga_pan_can_atlas_2018_rna_seq_v2_mrna_median_Zscores%2CCEACAM5%2CITGB6%2CMUC2%2CTPBG%2CMUC1)