Q (sonnet): show me EGFR expression across cancer types ▶ read_guide { "uri": "cbioportal://gene-expression-guide" } ◀ result # Gene Expression Analysis Guide This guide covers continuous-value genomic data: gene **expression**, **copy number** values, **methylation**, and related profile types. Mutation/CNA/SV *frequency* analysis lives in `cbioportal://mutation-frequency-guide`. ## Where this data lives Continuous per-sample-per-gene values are stored in `genetic_alteration_derived`: | Column | Description | |---|---| | `sample_unique_id` | `_` | | `cancer_study_identifier` | study scope | | `hugo_gene_symbol` | gene | | `profile_type` | which assay/normalization (see below) | | `alteration_value` | the actual value — stored as Nullable(String); cast with `toFloat64OrNull` | `alteration_value` is a string because the same column hosts many different value scales. The `''` and `'NA'` sentinels mean "missing"; always filter them out and use `toFloat64OrNull(alteration_value) IS NOT NULL` for downstream math. ## Discovering profile types for a study Different studies expose different profile types depending on what assays were run and how the data was normalized. Always check what a specific study supports before picking one: ```sql SELECT DISTINCT profile_type FROM genetic_alteration_derived WHERE cancer_study_identifier = 'brca_metabric' ORDER BY profile_type; ``` Common values across the public portal: | Family | Profile types | |---|---| | mRNA expression | `rna_seq_v2_mrna`, `rna_seq_v2_mrna_median_Zscores`, `rna_seq_v2_mrna_median_all_sample_Zscores` (TCGA PanCancer Atlas), `mrna`, `mrna_median_Zscores`, `mrna_seq_v2_rsem`, `mrna_seq_v2_rsem_Zscores`, `mrna_seq_cpm`, `mrna_seq_fpkm`, `mrna_U133`, `mrna_outliers` | | Copy number (continuous) | `cna`, `linear_CNA`, `log2CNA`, `cna_consensus`, `cna_rae`, `gistic` | | Methylation | `methylation_hm27`, `methylation_hm450`, `methylation_epic`, `methylation_promoters_rrbs` | | miRNA | `mirna`, `mirna_median_Zscores` | | Protein | `protein_quantification`, `protein_level`, `RPPA` | **Z-score vs raw choice.** When the user asks "is X correlated with Y", either works for Spearman (rank-based) — Pearson would care. Default to the non-Z-score variant if both exist, and call out which one in the response. ## Canonical recipe — Spearman correlation between two genes ```sql SELECT * FROM gene_pair_coexpression( study = 'brca_metabric', gene_a = 'TP53', gene_b = 'MYC', profile_type = 'mrna' ); ``` Returns one row: `(gene_a, gene_b, profile_type, spearman_correlation, num_samples)`. - `spearman_correlation` in [−1, 1]; `NULL` when fewer than 3 valid paired samples. - Mirrors cbioportal-backend's `ClickhouseCoExpressionMapper.getCoExpressions`, simplified to a pair lookup (the backend computes one ref gene vs ALL other genes for the coexpression page; here the agent asks about a specific pair). ### Verified examples | Study | gene_a | gene_b | profile_type | spearman | n | |---|---|---|---|---|---| | `brca_metabric` | TP53 | MYC | `mrna` | 0.118 | 1980 | | … (4931 more chars) ▶ clickhouse_run_select_query { "query": "WITH cohort AS ( SELECT cancer_study_identifier FROM cancer_study_query_preferences WHERE preference_name = 'pan_cancer_tcga' ), ct AS ( SELECT sample_unique_id, attribute_value AS cancer_type FROM clinical_data_derived WHERE attribute_name = 'CANCER_TYPE' AND cancer_study_identifier IN (SELECT cancer_study_identifier FROM cohort) ) SELECT ct.cancer_type, count() AS samples, round(median(log2(toFloat64OrNull(g.alteration_value) + 1)), 2) AS median_log2_rsem, round(quantile(0.9)(log2(toFloat64OrNull(g.alteration_value) + 1)), 2) AS p90_log2_rsem FROM genetic_alteration_derived g JOIN ct USING (sample_unique_id) WHERE g.cancer_study_identifier IN (SELECT cancer_study_identifier FROM cohort) AND g.hugo_gene_symbol = 'EGFR' AND g.profile_type = 'rna_seq_v2_mrna' AND g.alteration_value NOT IN ('', 'NA') AND toFloat64OrNull(g.alteration_value) IS NOT NULL GROUP BY ct.cancer_type ORDER BY p90_log2_rsem DESC;", "max_rows": 50 } ◀ result {"rows":[{"cancer_type":"Glioblastoma","samples":160,"median_log2_rsem":11.93,"p90_log2_rsem":15},{"cancer_type":"Glioma","samples":514,"median_log2_rsem":11.24,"p90_log2_rsem":13.05},{"cancer_type":"Head and Neck Cancer","samples":515,"median_log2_rsem":11.5,"p90_log2_rsem":12.98},{"cancer_type":"Renal Clear Cell Carcinoma","samples":510,"median_log2_rsem":11.28,"p90_log2_rsem":12.54},{"cancer_type":"Non-Small Cell Lung Cancer","samples":994,"median_log2_rsem":10.49,"p90_log2_rsem":12.05},{"cancer_type":"Esophagogastric Cancer","samples":593,"median_log2_rsem":9.6,"p90_log2_rsem":11.7},{"cancer_type":"Bladder Cancer","samples":407,"median_log2_rsem":9.27,"p90_log2_rsem":11.45},{"cancer_type":"Pleural Mesothelioma","samples":87,"median_log2_rsem":9.75,"p90_log2_rsem":11.33},{"cancer_type":"Cervical Cancer","samples":294,"median_log2_rsem":9.71,"p90_log2_rsem":11.32},{"cancer_type":"Hepatobiliary Cancer","samples":366,"median_log2_rsem":9.67,"p90_log2_rsem":11.22},{"cancer_type":"Renal Non-Clear Cell Carcinoma","samples":348,"median_log2_rsem":9.41,"p90_log2_rsem":11.03},{"cancer_type":"Thyroid Cancer","samples":498,"median_log2_rsem":9.61,"p90_log2_rsem":10.96},{"cancer_type":"Thymic Epithelial Tumor","samples":119,"median_log2_rsem":8.58,"p90_log2_rsem":10.93},{"cancer_type":"Cholangiocarcinoma","samples":36,"median_log2_rsem":9.47,"p90_log2_rsem":10.76},{"cancer_type":"Prostate Cancer","samples":493,"median_log2_rsem":9.5,"p90_log2_rsem":10.62},{"cancer_type":"Sarcoma","samples":253,"median_log2_rsem":8.62,"p90_log2_rsem":10.54},{"cancer_type":"Pancreatic Cancer","samples":177,"median_log2_rsem":9.36,"p90_log2_rsem":10.35},{"cancer_type":"Colorectal Cancer","samples":592,"median_log2_rsem":9.05,"p90_log2_rsem":9.95},{"cancer_type":"Breast Cancer","samples":1082,"median_log2_rsem":7.39,"p90_log2_rsem":9.76},{"cancer_type":"Endometrial Cancer","samples":584,"median_log2_rsem":8.11,"p90_log2_rsem":9.62},{"cancer_type":"Adrenocortical Carcinoma","samples":78,"median_log2_rsem":7.66,"p90_log2_rsem":9.59},{"cancer_type":"Non-Seminomatous Germ Cell Tumor","samples":86,"median_log2_rsem":7.24,"p90_log2_rsem":9.54},{"cancer_type":"Ovarian Epithelial Tumor","samples":300,"median_log2_rsem":7.76,"p90_log2_rsem":9.02},{"cancer_type":"Melanoma","samples":443,"median_log2_rsem":4.85,"p90_log2_rsem":7.96},{"cancer_type":"Pheochromocytoma","samples":147,"median_log2_rsem":4.9,"p90_log2_rsem":7.82},{"cancer_type":"Seminoma","samples":63,"median_log2_rsem":6.28,"p90_log2_rsem":7.74},{"cancer_type":"Miscellaneous Neuroepithelial Tumor","samples":31,"median_log2_rsem":5.43,"p90_log2_rsem":7.67},{"cancer_type":"Mature B-Cell Neoplasms","samples":48,"median_log2_rsem":4.25,"p90_log2_rsem":6.74},{"cancer_type":"Ocular Melanoma","samples":80,"median_log2_rsem":4.94,"p90_log2_rsem":6.69},{"cancer_type":"Leukemia","samples":173,"median_log2_rsem":-0.1,"p90_log2_rsem":2.67}]} ▶ clickhouse_run_select_query { "query": "SELECT cancer_study_identifier FROM cancer_study_query_preferences WHERE preference_name = 'pan_cancer_tcga' ORDER BY cancer_study_identifier", "max_rows": 50 } ◀ result {"rows":[{"cancer_study_identifier":"acc_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"blca_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"brca_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"cesc_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"chol_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"coadread_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"dlbc_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"esca_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"gbm_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"hnsc_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"kich_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"kirc_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"kirp_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"laml_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"lgg_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"lihc_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"luad_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"lusc_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"meso_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"ov_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"paad_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"pcpg_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"prad_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"sarc_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"skcm_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"stad_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"tgct_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"thca_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"thym_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"ucec_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"ucs_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"uvm_tcga_pan_can_atlas_2018"}]} ▶ resolve_and_route { "studyIds": [ "acc_tcga_pan_can_atlas_2018", "blca_tcga_pan_can_atlas_2018", "brca_tcga_pan_can_atlas_2018", "cesc_tcga_pan_can_atlas_2018", "chol_tcga_pan_can_atlas_2018", "coadread_tcga_pan_can_atlas_2018", "dlbc_tcga_pan_can_atlas_2018", "esca_tcga_pan_can_atlas_2018", "gbm_tcga_pan_can_atlas_2018", "hnsc_tcga_pan_can_atlas_2018", "kich_tcga_pan_can_atlas_2018", "kirc_tcga_pan_can_atlas_2018", "kirp_tcga_pan_can_atlas_2018", "laml_tcga_pan_can_atlas_2018", "lgg_tcga_pan_can_atlas_2018", "lihc_tcga_pan_can_atlas_2018", "luad_tcga_pan_can_atlas_2018", "lusc_tcga_pan_can_atlas_2018", "meso_tcga_pan_can_atlas_2018", "ov_tcga_pan_can_atlas_2018", "paad_tcga_pan_can_atlas_2018", "pcpg_tcga_pan_can_atlas_2018", "prad_tcga_pan_can_atlas_2018", "sarc_tcga_pan_can_atlas_2018", "skcm_tcga_pan_can_atlas_2018", "stad_tcga_pan_can_atlas_2018", "tgct_tcga_pan_can_atlas_2018", "thca_tcga_pan_can_atlas_2018", "thym_tcga_pan_can_atlas_2018", "ucec_tcga_pan_can_atlas_2018", "ucs_tcga_pan_can_atlas_2018", "uvm_tcga_pan_can_atlas_2018" ] } ◀ result Error: result (81,891 characters across 1 line) exceeds maximum allowed tokens. Output has been saved to /home/ino/.claude-cbioportal-team/projects/-tmp-mcp-qa-claude-5mhnjm0m/117bf706-6bcb-4a57-8a30-9e7431e6e832/tool-results/mcp-navigator-resolve_and_route-1790442041946.txt. Format: Plain text - For targeted searches (find a string): use grep on the file directly. - For analysis or summarization that requires reading the full content: the file's lines are too long for Read's offset/limit. Slice by character range via Bash instead — e.g. python3 -c 'print(open("/home/ino/.claude-cbioportal-team/projects/-tmp-mcp-qa-claude-5mhnjm0m/117bf706-6bcb-4a57-8a30-9e7431e6e832/tool-results/mcp-navigator-resolve_and_route-1790442041946.txt").read()[A:B])' in ~80,000-char spans until you have read 100% of it. - If the Agent tool is available, do this inside a subagent so the full output stays out of your main context. Give it the instruction above verbatim, and be explicit about what it must return — e.g. "Slice /home/ino/.claude-cbioportal-team/projects/-tmp-mcp-qa-claude-5mhnjm0m/117bf706-6bcb-4a57-8a30-9e7431e6e832/tool-results/mcp-navigator-resolve_and_route-1790442041946.txt in ~80,000-char spans via python (read()[A:B]) until you have read all 81,891 characters, then summarize and quote any key findings verbatim." A vague "summarize this" may lose detail. ▶ resolve_and_route { "studyIds": [ "gbm_tcga_pan_can_atlas_2018", "luad_tcga_pan_can_atlas_2018" ] } ◀ result {"success":true,"message":"Found 2 matching studies. Pick the best match (prefer TCGA → prefer PanCancer Atlas), use its metadata to call the appropriate navigation tool(s) and generate URLs immediately. Present other studies as alternatives.","data":{"totalCount":2,"studiesWithMetadata":[{"studyId":"gbm_tcga_pan_can_atlas_2018","name":"Glioblastoma Multiforme (TCGA, PanCancer Atlas)","sampleCount":592,"studyViewUrl":"https://www.cbioportal.org/study?id=gbm_tcga_pan_can_atlas_2018","metadata":{"clinicalAttributeIds":["AGE","AJCC_PATHOLOGIC_TUMOR_STAGE","AJCC_STAGING_EDITION","ANEUPLOIDY_SCORE","BUFFA_HYPOXIA_SCORE","CANCER_TYPE","CANCER_TYPE_ACRONYM","CANCER_TYPE_DETAILED","DAYS_LAST_FOLLOWUP","DAYS_TO_BIRTH","DAYS_TO_INITIAL_PATHOLOGIC_DIAGNOSIS","DFS_MONTHS","DFS_STATUS","DSS_MONTHS","DSS_STATUS","ETHNICITY","FORM_COMPLETION_DATE","FRACTION_GENOME_ALTERED","GENETIC_ANCESTRY_LABEL","GRADE","HISTORY_NEOADJUVANT_TRTYN","ICD_10","ICD_O_3_HISTOLOGY","ICD_O_3_SITE","INFORMED_CONSENT_VERIFIED","IN_PANCANPATHWAYS_FREEZE","MSI_SCORE_MANTIS","MSI_SENSOR_SCORE","MUTATION_COUNT","NEW_TUMOR_EVENT_AFTER_INITIAL_TREATMENT","ONCOTREE_CODE","OS_MONTHS","OS_STATUS","OTHER_PATIENT_ID","PATH_M_STAGE","PATH_N_STAGE","PATH_T_STAGE","PERSON_NEOPLASM_CANCER_STATUS","PFS_MONTHS","PFS_STATUS","PRIMARY_LYMPH_NODE_PRESENTATION_ASSESSMENT","PRIOR_DX","RACE","RADIATION_THERAPY","RAGNUM_HYPOXIA_SCORE","SAMPLE_COUNT","SAMPLE_TYPE","SEX","SOMATIC_STATUS","SUBTYPE","TBL_SCORE","TISSUE_PROSPECTIVE_COLLECTION_INDICATOR","TISSUE_RETROSPECTIVE_COLLECTION_INDICATOR","TISSUE_SOURCE_SITE","TISSUE_SOURCE_SITE_CODE","TMB_NONSYNONYMOUS","TUMOR_TISSUE_SITE","TUMOR_TYPE","WEIGHT","WINTER_HYPOXIA_SCORE"],"molecularProfileIds":["gbm_tcga_pan_can_atlas_2018_gistic","gbm_tcga_pan_can_atlas_2018_log2CNA","gbm_tcga_pan_can_atlas_2018_mutations","gbm_tcga_pan_can_atlas_2018_rna_seq_v2_mrna","gbm_tcga_pan_can_atlas_2018_rna_seq_v2_mrna_median_Zscores","gbm_tcga_pan_can_atlas_2018_rna_seq_v2_mrna_median_all_sample_Zscores","gbm_tcga_pan_can_atlas_2018_rppa","gbm_tcga_pan_can_atlas_2018_rppa_Zscores","gbm_tcga_pan_can_atlas_2018_structural_variants"],"genericAssayProfiles":["gbm_tcga_pan_can_atlas_2018_armlevel_cna","gbm_tcga_pan_can_atlas_2018_genetic_ancestry","gbm_tcga_pan_can_atlas_2018_methylation_hm27_hm450_merge","gbm_tcga_pan_can_atlas_2018_methylation_hm450"],"heatmapProfileIds":["gbm_tcga_pan_can_atlas_2018_rna_seq_v2_mrna_median_Zscores","gbm_tcga_pan_can_atlas_2018_rna_seq_v2_mrna_median_all_sample_Zscores","gbm_tcga_pan_can_atlas_2018_rppa_Zscores"],"availableComparisonTabs":["overlap","clinical","survival","alterations","mutations","mrna","protein","generic_assay_armlevel_cna","generic_assay_genetic_ancestry","generic_assay_methylation"]}},{"studyId":"luad_tcga_pan_can_atlas_2018","name":"Lung Adenocarcinoma (TCGA, PanCancer Atlas)","sampleCount":566,"studyViewUrl":"https://www.cbioportal.org/study?id=luad_tcga_pan_can_atlas_2018","metadata":{"clinicalAttributeIds":["AGE","AJCC_PATHOLO … (2394 more chars) ═ answer (success) cBioPortal's built-in **"Cancer Types Summary"** view within a single-gene query is actually the best fit for a true cross-cancer-type expression comparison, but that requires all 31 studies loaded, which is too large a metadata payload for me to resolve in one shot here. I'll give you the query results directly (below) and link you to the Glioblastoma study's Plots tab as an example entry point — but for the full pan-cancer boxplot, the fastest path is the public cBioPortal query page itself. ## EGFR mRNA expression across TCGA PanCancer Atlas cancer types Data: RNA-Seq (RSEM), `rna_seq_v2_mrna` profile, TCGA PanCancer Atlas cohorts (32 studies, one non-overlapping sample set per study). Values shown as log2(RSEM+1); higher = more expression. | Cancer type | Samples | Median log2(RSEM+1) | 90th pctile | |---|---|---|---| | Glioblastoma | 160 | 11.93 | 15.00 | | Glioma (LGG) | 514 | 11.24 | 13.05 | | Head and Neck Cancer | 515 | 11.50 | 12.98 | | Renal Clear Cell Carcinoma | 510 | 11.28 | 12.54 | | Non-Small Cell Lung Cancer | 994 | 10.49 | 12.05 | | Esophagogastric Cancer | 593 | 9.60 | 11.70 | | Bladder Cancer | 407 | 9.27 | 11.45 | | Pleural Mesothelioma | 87 | 9.75 | 11.33 | | Cervical Cancer | 294 | 9.71 | 11.32 | | Hepatobiliary Cancer | 366 | 9.67 | 11.22 | | Renal Non-Clear Cell Carcinoma | 348 | 9.41 | 11.03 | | Thyroid Cancer | 498 | 9.61 | 10.96 | | Thymic Epithelial Tumor | 119 | 8.58 | 10.93 | | Cholangiocarcinoma | 36 | 9.47 | 10.76 | | Prostate Cancer | 493 | 9.50 | 10.62 | | Sarcoma | 253 | 8.62 | 10.54 | | Pancreatic Cancer | 177 | 9.36 | 10.35 | | Colorectal Cancer | 592 | 9.05 | 9.95 | | Breast Cancer | 1,082 | 7.39 | 9.76 | | Endometrial Cancer | 584 | 8.11 | 9.62 | | Adrenocortical Carcinoma | 78 | 7.66 | 9.59 | | Non-Seminomatous Germ Cell Tumor | 86 | 7.24 | 9.54 | | Ovarian Epithelial Tumor | 300 | 7.76 | 9.02 | | Melanoma | 443 | 4.85 | 7.96 | | Pheochromocytoma | 147 | 4.90 | 7.82 | | Seminoma | 63 | 6.28 | 7.74 | | Miscellaneous Neuroepithelial Tumor | 31 | 5.43 | 7.67 | | Mature B-Cell Neoplasms | 48 | 4.25 | 6.74 | | Ocular Melanoma | 80 | 4.94 | 6.69 | | Leukemia (AML) | 173 | -0.10 | 2.67 | **Key observations:** - **Glioblastoma and lower-grade glioma** show the highest EGFR expression (median and tail), consistent with EGFR's known role as a glioma driver. - **Head and neck, renal clear cell, and NSCLC** also show high expression. - **Leukemia, ocular melanoma, and B-cell neoplasms** show near-baseline expression — expected, as EGFR is not typically expressed in hematologic or ocular melanocytic lineages. - Values are RSEM-based (not Z-scores), so they're directly comparable across these studies since they share a common normalization pipeline (unlike Z-scores, which are study-centered and not cross-comparable). Note: sample counts here reflect RNA-seq–profiled tumor samples per study; because these are 32 distinct TCGA cohorts, there's no cross-study double-counting to worry about here (each patient belongs to exactly one cancer type cohort). ### Explore interactively Two example entry points into the same RNA-seq profile (glioblastoma, the top-ranked cancer type, and lung adenocarcinoma, a common EGFR-relevant cancer): - [EGFR expression — Glioblastoma (TCGA PanCancer Atlas) Plots](https://www.cbioportal.org/results/plots?cancer_study_list=gbm_tcga_pan_can_atlas_2018&Z_SCORE_THRESHOLD=2.0&RPPA_SCORE_THRESHOLD=2.0&data_priority=0&profileFilter=rna_seq_v2_mrna&case_set_id=gbm_tcga_pan_can_atlas_2018_all&gene_list=EGFR&geneset_list=%20&tab_index=tab_visualize&plots_horz_selection=%7B%22dataType%22%3A%22mrna%22%2C%22selectedDataSourceOption%22%3A%22rna_seq_v2_mrna%22%7D&plots_vert_selection=%7B%7D) - [Glioblastoma study overview](https://www.cbioportal.org/study?id=gbm_tcga_pan_can_atlas_2018) For the full cross-cancer-type boxplot (all 32 cohorts side-by-side), the fastest path is [cBioPortal's Query page](https://www.cbioportal.org/query) — select the "TCGA PanCancer Atlas Studies" quick-select, query gene `EGFR`, choose the mRNA (RNA-Seq RSEM) profile, and open the **Plots** tab with "Cancer Type" as the grouping variable; it will render exactly the boxplot summarized in the table above.