Q (haiku): are there cdkn2a het losses in gbm tcga study? ▶ read_guide { "uri": "cbioportal://gene-expression-guide" } ▶ search_oncotree { "search_term": "GBM" } ◀ result # Gene Expression Analysis Guide This guide covers continuous-value genomic data: gene **expression**, **copy number** values, **methylation**, and related profile types. Mutation/CNA/SV *frequency* analysis lives in `cbioportal://mutation-frequency-guide`. ## Where this data lives Continuous per-sample-per-gene values are stored in `genetic_alteration_derived`: | Column | Description | |---|---| | `sample_unique_id` | `_` | | `cancer_study_identifier` | study scope | | `hugo_gene_symbol` | gene | | `profile_type` | which assay/normalization (see below) | | `alteration_value` | the actual value — stored as Nullable(String); cast with `toFloat64OrNull` | `alteration_value` is a string because the same column hosts many different value scales. The `''` and `'NA'` sentinels mean "missing"; always filter them out and use `toFloat64OrNull(alteration_value) IS NOT NULL` for downstream math. ## Discovering profile types for a study Different studies expose different profile types depending on what assays were run and how the data was normalized. Always check what a specific study supports before picking one: ```sql SELECT DISTINCT profile_type FROM genetic_alteration_derived WHERE cancer_study_identifier = 'brca_metabric' ORDER BY profile_type; ``` Common values across the public portal: | Family | Profile types | |---|---| | mRNA expression | `rna_seq_v2_mrna`, `rna_seq_v2_mrna_median_Zscores`, `rna_seq_v2_mrna_median_all_sample_Zscores` (TCGA PanCancer Atlas), `mrna`, `mrna_median_Zscores`, `mrna_seq_v2_rsem`, `mrna_seq_v2_rsem_Zscores`, `mrna_seq_cpm`, `mrna_seq_fpkm`, `mrna_U133`, `mrna_outliers` | | Copy number (continuous) | `cna`, `linear_CNA`, `log2CNA`, `cna_consensus`, `cna_rae`, `gistic` | | Methylation | `methylation_hm27`, `methylation_hm450`, `methylation_epic`, `methylation_promoters_rrbs` | | miRNA | `mirna`, `mirna_median_Zscores` | | Protein | `protein_quantification`, `protein_level`, `RPPA` | **Z-score vs raw choice.** When the user asks "is X correlated with Y", either works for Spearman (rank-based) — Pearson would care. Default to the non-Z-score variant if both exist, and call out which one in the response. ## Canonical recipe — Spearman correlation between two genes ```sql SELECT * FROM gene_pair_coexpression( study = 'brca_metabric', gene_a = 'TP53', gene_b = 'MYC', profile_type = 'mrna' ); ``` Returns one row: `(gene_a, gene_b, profile_type, spearman_correlation, num_samples)`. - `spearman_correlation` in [−1, 1]; `NULL` when fewer than 3 valid paired samples. - Mirrors cbioportal-backend's `ClickhouseCoExpressionMapper.getCoExpressions`, simplified to a pair lookup (the backend computes one ref gene vs ALL other genes for the coexpression page; here the agent asks about a specific pair). ### Verified examples | Study | gene_a | gene_b | profile_type | spearman | n | |---|---|---|---|---|---| | `brca_metabric` | TP53 | MYC | `mrna` | 0.118 | 1980 | | … (4931 more chars) ▶ list_studies { "search": "GBM TCGA" } ◀ result {"result":[{"code":"GB","name":"Glioblastoma, IDH-Wildtype","score":90,"mainType":"Glioma","tissue":"CNS/Brain","hierarchy":"BRAIN > DIFG > ADIFG > GB","replacedCodes":["GBM"]},{"code":"SCGBM","name":"Small Cell Glioblastoma","score":55,"mainType":"Glioma","tissue":"CNS/Brain","hierarchy":"BRAIN > DIFG > ADIFG > GB > SCGBM"}]} ◀ result {"result":[{"cancer_study_identifier":"lgggbm_tcga_pub","name":"Merged Cohort of LGG and GBM (TCGA, Cell 2016)","type_of_cancer_id":"difg","sample_count":1122,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=lgggbm_tcga_pub"},{"cancer_study_identifier":"gbm_tcga","name":"Glioblastoma Multiforme (TCGA, Firehose Legacy)","type_of_cancer_id":"difg","sample_count":619,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=gbm_tcga"},{"cancer_study_identifier":"gbm_tcga_gdc","name":"Glioblastoma Multiforme (TCGA GDC, 2025)","type_of_cancer_id":"difg","sample_count":611,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=gbm_tcga_gdc"},{"cancer_study_identifier":"gbm_tcga_pan_can_atlas_2018","name":"Glioblastoma Multiforme (TCGA, PanCancer Atlas)","type_of_cancer_id":"difg","sample_count":592,"has_guide":true,"url":"https://www.cbioportal.org/study/summary?id=gbm_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"gbm_tcga_pub2013","name":"Glioblastoma (TCGA, Cell 2013)","type_of_cancer_id":"difg","sample_count":577,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=gbm_tcga_pub2013"},{"cancer_study_identifier":"gbm_tcga_pub","name":"Glioblastoma (TCGA, Nature 2008)","type_of_cancer_id":"difg","sample_count":206,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=gbm_tcga_pub"}]} ▶ clickhouse_list_table_columns { "table": "genetic_alteration_derived" } ▶ list_studies { "search": "gbm_tcga", "verbose": true } ◀ result {"columns":[{"name":"sample_unique_id","type":"String"},{"name":"cancer_study_identifier","type":"LowCardinality(String)"},{"name":"hugo_gene_symbol","type":"String"},{"name":"profile_type","type":"LowCardinality(String)"},{"name":"alteration_value","type":"Nullable(String)"}]} ◀ result {"result":[{"cancer_study_identifier":"lgggbm_tcga_pub","name":"Merged Cohort of LGG and GBM (TCGA, Cell 2016)","description":"Whole exome sequencing of 1,122 LGG and GBM tumor/normal pairs.","type_of_cancer_id":"difg","sample_count":1122,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=lgggbm_tcga_pub"},{"cancer_study_identifier":"gbm_tcga","name":"Glioblastoma Multiforme (TCGA, Firehose Legacy)","description":"TCGA Glioblastoma Multiforme. Source data from GDAC Firehose. Previously known as TCGA Provisional.","type_of_cancer_id":"difg","sample_count":619,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=gbm_tcga"},{"cancer_study_identifier":"gbm_tcga_gdc","name":"Glioblastoma Multiforme (TCGA GDC, 2025)","description":"TCGA Glioblastoma Multiforme. Source data from NCI GDC and generated in Aug 2025 using Cancer Data Aggregator.","type_of_cancer_id":"difg","sample_count":611,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=gbm_tcga_gdc"},{"cancer_study_identifier":"gbm_tcga_pan_can_atlas_2018","name":"Glioblastoma Multiforme (TCGA, PanCancer Atlas)","description":"Glioblastoma Multiforme TCGA PanCancer data. The original data is here. The publications are here.","type_of_cancer_id":"difg","sample_count":592,"has_guide":true,"url":"https://www.cbioportal.org/study/summary?id=gbm_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"gbm_tcga_pub2013","name":"Glioblastoma (TCGA, Cell 2013)","description":"Whole-exome and/or whole-genome sequencing of 291 of the 577 glioblastoma tumor/normal pairs. The Cancer Genome Atlas (TCGA) Glioblastoma Project.","type_of_cancer_id":"difg","sample_count":577,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=gbm_tcga_pub2013"},{"cancer_study_identifier":"gbm_tcga_pub","name":"Glioblastoma (TCGA, Nature 2008)","description":"Targeted sequencing in 91 of the 206 primary glioblastoma tumors (143 with matched normals) from the Cancer Genome Atlas (TCGA) Glioblastoma Project.","type_of_cancer_id":"difg","sample_count":206,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=gbm_tcga_pub"}]} ▶ clickhouse_run_select_query { "query": " SELECT DISTINCT profile_type FROM genetic_alteration_derived WHERE cancer_study_identifier = 'gbm_tcga' AND hugo_gene_symbol = 'CDKN2A' ORDER BY profile_type " } ◀ result {"rows":[{"profile_type":"gistic"},{"profile_type":"linear_CNA"},{"profile_type":"methylation_hm27"},{"profile_type":"methylation_hm450"},{"profile_type":"mrna"},{"profile_type":"mrna_U133"},{"profile_type":"mrna_U133_Zscores"},{"profile_type":"mrna_U133_all_sample_Zscores"},{"profile_type":"mrna_median_Zscores"},{"profile_type":"mrna_median_all_sample_Zscores"},{"profile_type":"rna_seq_v2_mrna"},{"profile_type":"rna_seq_v2_mrna_median_Zscores"},{"profile_type":"rna_seq_v2_mrna_median_all_sample_Zscores"}]} ▶ clickhouse_run_select_query { "query": " SELECT count(DISTINCT sample_unique_id) AS hetloss_samples, (SELECT count(DISTINCT sample_unique_id) FROM genetic_alteration_derived WHERE cancer_study_identifier = 'gbm_tcga' AND hugo_gene_symbol = 'CDKN2A' AND profile_type = 'gistic' AND alteration_value NOT IN ('', 'NA')) AS cna_profiled_samples, round(count(DISTINCT sample_unique_id) * 100.0 / (SELECT count(DISTINCT sample_unique_id) FROM genetic_alteration_derived WHERE cancer_study_identifier = 'gbm_tcga' AND hugo_gene_symbol = 'CDKN2A' AND profile_type = 'gistic' AND alteration_value NOT IN ('', 'NA')), 2) AS hetloss_percent FROM genetic_alteration_derived WHERE cancer_study_identifier = 'gbm_tcga' AND hugo_gene_symbol = 'CDKN2A' AND profile_type = 'gistic' AND alteration_value = '-1' " } ◀ result {"rows":[{"hetloss_samples":107,"cna_profiled_samples":577,"hetloss_percent":18.54}]} ═ answer (success) Excellent! Yes, there are **CDKN2A heterozygous losses in the GBM TCGA study**. Here are the results: ## CDKN2A Heterozygous Loss (HETLOSS) in GBM TCGA | Metric | Value | |--------|-------| | **Samples with HETLOSS** | 107 | | **Total CNA-profiled samples** | 577 | | **Frequency** | 18.54% | **Study:** [Glioblastoma Multiforme (TCGA, Firehose Legacy)](https://www.cbioportal.org/study/summary?id=gbm_tcga) CDKN2A heterozygous loss is found in approximately 1 in 5 GBM tumors in this cohort, based on GISTIC discrete copy-number calls. You can explore this in cBioPortal by filtering the study for CDKN2A losses or viewing the copy-number segment details.