Q (haiku): calculate median expression of ceacam5, itgb6, muc2, tpgb and muc1 mRNA in pancreatic cancer ▶ read_guide { "uri": "cbioportal://gene-expression-guide" } ▶ list_studies { "search": "pancreatic cancer", "limit": 20 } ◀ result # Gene Expression Analysis Guide This guide covers continuous-value genomic data: gene **expression**, **copy number** values, **methylation**, and related profile types. Mutation/CNA/SV *frequency* analysis lives in `cbioportal://mutation-frequency-guide`. ## Where this data lives Continuous per-sample-per-gene values are stored in `genetic_alteration_derived`: | Column | Description | |---|---| | `sample_unique_id` | `_` | | `cancer_study_identifier` | study scope | | `hugo_gene_symbol` | gene | | `profile_type` | which assay/normalization (see below) | | `alteration_value` | the actual value — stored as Nullable(String); cast with `toFloat64OrNull` | `alteration_value` is a string because the same column hosts many different value scales. The `''` and `'NA'` sentinels mean "missing"; always filter them out and use `toFloat64OrNull(alteration_value) IS NOT NULL` for downstream math. ## Discovering profile types for a study Different studies expose different profile types depending on what assays were run and how the data was normalized. Always check what a specific study supports before picking one: ```sql SELECT DISTINCT profile_type FROM genetic_alteration_derived WHERE cancer_study_identifier = 'brca_metabric' ORDER BY profile_type; ``` Common values across the public portal: | Family | Profile types | |---|---| | mRNA expression | `mrna`, `mrna_median_Zscores`, `mrna_seq_v2_rsem`, `mrna_seq_v2_rsem_Zscores`, `mrna_seq_cpm`, `mrna_seq_fpkm`, `mrna_U133`, `mrna_outliers` | | Copy number (continuous) | `cna`, `linear_CNA`, `log2CNA`, `cna_consensus`, `cna_rae`, `gistic` | | Methylation | `methylation_hm27`, `methylation_hm450`, `methylation_epic`, `methylation_promoters_rrbs` | | miRNA | `mirna`, `mirna_median_Zscores` | | Protein | `protein_quantification`, `protein_level`, `RPPA` | **Z-score vs raw choice.** When the user asks "is X correlated with Y", either works for Spearman (rank-based) — Pearson would care. Default to the non-Z-score variant if both exist, and call out which one in the response. ## Canonical recipe — Spearman correlation between two genes ```sql SELECT * FROM gene_pair_coexpression( study = 'brca_metabric', gene_a = 'TP53', gene_b = 'MYC', profile_type = 'mrna' ); ``` Returns one row: `(gene_a, gene_b, profile_type, spearman_correlation, num_samples)`. - `spearman_correlation` in [−1, 1]; `NULL` when fewer than 3 valid paired samples. - Mirrors cbioportal-backend's `ClickhouseCoExpressionMapper.getCoExpressions`, simplified to a pair lookup (the backend computes one ref gene vs ALL other genes for the coexpression page; here the agent asks about a specific pair). ### Verified examples | Study | gene_a | gene_b | profile_type | spearman | n | |---|---|---|---|---|---| | `brca_metabric` | TP53 | MYC | `mrna` | 0.118 | 1980 | | `brca_metabric` | ESR1 | PGR | `mrna` | 0.487 | 1980 | ESR1↔PGR is the textbook breast-cancer estrogen-receptor coregula … (2785 more chars) ◀ result {"result":[{"cancer_study_identifier":"pancreas_ctdna_msk_2025","name":"Pancreatic Cancer cfDNA (MSK, J Natl Cancer Inst 2025)","type_of_cancer_id":"paad","sample_count":412,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=pancreas_ctdna_msk_2025"},{"cancer_study_identifier":"pancreas_msk_2024","name":"Pancreatic Cancer (MSK, Cancer Cell 2024)","type_of_cancer_id":"pancreas","sample_count":395,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=pancreas_msk_2024"},{"cancer_study_identifier":"pancreas_cptac_gdc","name":"Pancreatic Cancer (CPTAC GDC, 2025)","type_of_cancer_id":"pancreas","sample_count":183,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=pancreas_cptac_gdc"},{"cancer_study_identifier":"paad_cptac_2021","name":"Pancreatic Ductal Adenocarcinoma (CPTAC, Cell 2021)","type_of_cancer_id":"paad","sample_count":140,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=paad_cptac_2021"},{"cancer_study_identifier":"paad_utsw_2015","name":"Pancreatic Cancer (UTSW, Nat Commun 2015)","type_of_cancer_id":"paad","sample_count":109,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=paad_utsw_2015"},{"cancer_study_identifier":"paad_iatlas_prince_2022","name":"Metastatic Pancreatic Adenocarcinoma (PRINCE Trial, Nat Med. 2022) - iAtlas Harmonized","type_of_cancer_id":"paad","sample_count":93,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=paad_iatlas_prince_2022"},{"cancer_study_identifier":"msk_met_2021","name":"MSK MetTropism (MSK, Cell 2021)","type_of_cancer_id":"mixed","sample_count":25775,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=msk_met_2021"},{"cancer_study_identifier":"paad_tcga_gdc","name":"Pancreatic Adenocarcinoma (TCGA GDC, 2025)","type_of_cancer_id":"paad","sample_count":186,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=paad_tcga_gdc"},{"cancer_study_identifier":"paad_tcga_pan_can_atlas_2018","name":"Pancreatic Adenocarcinoma (TCGA, PanCancer Atlas)","type_of_cancer_id":"paad","sample_count":184,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=paad_tcga_pan_can_atlas_2018"}]} ▶ clickhouse_run_select_query { "query": " SELECT DISTINCT profile_type FROM genetic_alteration_derived WHERE cancer_study_identifier = 'paad_tcga_pan_can_atlas_2018' ORDER BY profile_type " } ◀ result {"rows":[{"profile_type":"gistic"},{"profile_type":"log2CNA"},{"profile_type":"rna_seq_v2_mrna"},{"profile_type":"rna_seq_v2_mrna_median_Zscores"},{"profile_type":"rna_seq_v2_mrna_median_all_sample_Zscores"},{"profile_type":"rppa"},{"profile_type":"rppa_Zscores"}]} ▶ clickhouse_run_select_query { "query": " SELECT hugo_gene_symbol, quantile(0.5)(toFloat64OrNull(alteration_value)) AS median_expression, count() AS num_samples FROM genetic_alteration_derived WHERE cancer_study_identifier = 'paad_tcga_pan_can_atlas_2018' AND profile_type = 'rna_seq_v2_mrna' AND hugo_gene_symbol IN ('CEACAM5', 'ITGB6', 'MUC2', 'TPGB', 'MUC1') AND alteration_value NOT IN ('', 'NA') AND toFloat64OrNull(alteration_value) IS NOT NULL GROUP BY hugo_gene_symbol ORDER BY hugo_gene_symbol ", "max_rows": 100 } ◀ result {"rows":[{"hugo_gene_symbol":"CEACAM5","median_expression":10049.7,"num_samples":177},{"hugo_gene_symbol":"ITGB6","median_expression":3740.15,"num_samples":177},{"hugo_gene_symbol":"MUC1","median_expression":14799,"num_samples":177},{"hugo_gene_symbol":"MUC2","median_expression":6.5339,"num_samples":177}]} ═ answer (success) I found data for four of the five genes. **TPGB** was not found in the database—it may not be present in this study or the gene symbol may be incorrect. Here are the median mRNA expression values for pancreatic cancer (TCGA, PanCancer Atlas): | Gene | Median Expression (RSEM) | N Samples | |---|---|---| | **CEACAM5** | 10,049.7 | 177 | | **ITGB6** | 3,740.2 | 177 | | **MUC1** | 14,799.0 | 177 | | **MUC2** | 6.53 | 177 | | **TPGB** | *Not found* | — | **Key findings:** - **MUC1** shows the highest median expression (~14.8k RSEM reads) - **CEACAM5** is also highly expressed (~10k RSEM reads) - **ITGB6** has moderate expression (~3.7k RSEM reads) - **MUC2** is the lowest among detected genes (~6.5 RSEM reads) - All 177 tumor samples in this TCGA cohort have measurements for the first four genes **About TPGB:** This gene symbol doesn't appear in the TCGA pancreatic data. You may want to verify the gene name — could it be **TPGB** (trefoil factor family member 3) or a different symbol? I can search other pancreatic cancer cohorts if needed. Would you like me to query additional pancreatic cancer studies, search for alternative gene symbols, or look at Z-score normalized expression instead?