Q (sonnet-5.5): How does ERBB2 mRNA expression vary across different cancer types in TCGA Pan-Cancer Atlas studies? ▶ read_guide { "uri": "cbioportal://gene-expression-guide" } ◀ result # Gene Expression Analysis Guide This guide covers continuous-value genomic data: gene **expression**, **copy number** values, **methylation**, and related profile types. Mutation/CNA/SV *frequency* analysis lives in `cbioportal://mutation-frequency-guide`. ## Where this data lives Continuous per-sample-per-gene values are stored in `genetic_alteration_derived`: | Column | Description | |---|---| | `sample_unique_id` | `_` | | `cancer_study_identifier` | study scope | | `hugo_gene_symbol` | gene | | `profile_type` | which assay/normalization (see below) | | `alteration_value` | the actual value — stored as Nullable(String); cast with `toFloat64OrNull` | `alteration_value` is a string because the same column hosts many different value scales. The `''` and `'NA'` sentinels mean "missing"; always filter them out and use `toFloat64OrNull(alteration_value) IS NOT NULL` for downstream math. ## Discovering profile types for a study Different studies expose different profile types depending on what assays were run and how the data was normalized. Always check what a specific study supports before picking one: ```sql SELECT DISTINCT profile_type FROM genetic_alteration_derived WHERE cancer_study_identifier = 'brca_metabric' ORDER BY profile_type; ``` Common values across the public portal: | Family | Profile types | |---|---| | mRNA expression | `rna_seq_v2_mrna`, `rna_seq_v2_mrna_median_Zscores`, `rna_seq_v2_mrna_median_all_sample_Zscores` (TCGA PanCancer Atlas), `mrna`, `mrna_median_Zscores`, `mrna_seq_v2_rsem`, `mrna_seq_v2_rsem_Zscores`, `mrna_seq_cpm`, `mrna_seq_fpkm`, `mrna_U133`, `mrna_outliers` | | Copy number (continuous) | `cna`, `linear_CNA`, `log2CNA`, `cna_consensus`, `cna_rae`, `gistic` | | Methylation | `methylation_hm27`, `methylation_hm450`, `methylation_epic`, `methylation_promoters_rrbs` | | miRNA | `mirna`, `mirna_median_Zscores` | | Protein | `protein_quantification`, `protein_level`, `RPPA` | **Z-score vs raw choice.** When the user asks "is X correlated with Y", either works for Spearman (rank-based) — Pearson would care. Default to the non-Z-score variant if both exist, and call out which one in the response. ## Canonical recipe — Spearman correlation between two genes ```sql SELECT * FROM gene_pair_coexpression( study = 'brca_metabric', gene_a = 'TP53', gene_b = 'MYC', profile_type = 'mrna' ); ``` Returns one row: `(gene_a, gene_b, profile_type, spearman_correlation, num_samples)`. - `spearman_correlation` in [−1, 1]; `NULL` when fewer than 3 valid paired samples. - Mirrors cbioportal-backend's `ClickhouseCoExpressionMapper.getCoExpressions`, simplified to a pair lookup (the backend computes one ref gene vs ALL other genes for the coexpression page; here the agent asks about a specific pair). ### Verified examples | Study | gene_a | gene_b | profile_type | spearman | n | |---|---|---|---|---|---| | `brca_metabric` | TP53 | MYC | `mrna` | 0.118 | 1980 | | … (4931 more chars) ▶ clickhouse_run_select_query { "query": "WITH cohort AS ( SELECT cancer_study_identifier FROM cancer_study_query_preferences WHERE preference_name = 'pan_cancer_tcga' ), ct AS ( SELECT sample_unique_id, attribute_value AS cancer_type FROM clinical_data_derived WHERE attribute_name = 'CANCER_TYPE' AND cancer_study_identifier IN (SELECT cancer_study_identifier FROM cohort) ) SELECT ct.cancer_type, count() AS samples, round(median(log2(toFloat64OrNull(g.alteration_value) + 1)), 2) AS median_log2_rsem, round(quantile(0.9)(log2(toFloat64OrNull(g.alteration_value) + 1)), 2) AS p90_log2_rsem FROM genetic_alteration_derived g JOIN ct USING (sample_unique_id) WHERE g.cancer_study_identifier IN (SELECT cancer_study_identifier FROM cohort) AND g.hugo_gene_symbol = 'ERBB2' AND g.profile_type = 'rna_seq_v2_mrna' AND g.alteration_value NOT IN ('', 'NA') AND toFloat64OrNull(g.alteration_value) IS NOT NULL GROUP BY ct.cancer_type ORDER BY p90_log2_rsem DESC", "max_rows": 100 } ▶ clickhouse_run_select_query { "query": "SELECT cancer_study_identifier FROM cancer_study_query_preferences WHERE preference_name = 'pan_cancer_tcga' ORDER BY cancer_study_identifier" } ◀ result {"rows":[{"cancer_type":"Breast Cancer","samples":1082,"median_log2_rsem":12.78,"p90_log2_rsem":14.75},{"cancer_type":"Bladder Cancer","samples":407,"median_log2_rsem":12.88,"p90_log2_rsem":14.28},{"cancer_type":"Esophagogastric Cancer","samples":593,"median_log2_rsem":12.02,"p90_log2_rsem":13.75},{"cancer_type":"Renal Non-Clear Cell Carcinoma","samples":348,"median_log2_rsem":12.63,"p90_log2_rsem":13.38},{"cancer_type":"Endometrial Cancer","samples":584,"median_log2_rsem":12.24,"p90_log2_rsem":13.36},{"cancer_type":"Thyroid Cancer","samples":498,"median_log2_rsem":12.71,"p90_log2_rsem":13.28},{"cancer_type":"Non-Small Cell Lung Cancer","samples":994,"median_log2_rsem":11.96,"p90_log2_rsem":13.23},{"cancer_type":"Cervical Cancer","samples":294,"median_log2_rsem":12.15,"p90_log2_rsem":13.22},{"cancer_type":"Prostate Cancer","samples":493,"median_log2_rsem":12.43,"p90_log2_rsem":13.03},{"cancer_type":"Pancreatic Cancer","samples":177,"median_log2_rsem":12.29,"p90_log2_rsem":12.99},{"cancer_type":"Cholangiocarcinoma","samples":36,"median_log2_rsem":12.42,"p90_log2_rsem":12.96},{"cancer_type":"Colorectal Cancer","samples":592,"median_log2_rsem":12.2,"p90_log2_rsem":12.9},{"cancer_type":"Ovarian Epithelial Tumor","samples":300,"median_log2_rsem":11.98,"p90_log2_rsem":12.8},{"cancer_type":"Non-Seminomatous Germ Cell Tumor","samples":86,"median_log2_rsem":11.57,"p90_log2_rsem":12.63},{"cancer_type":"Head and Neck Cancer","samples":515,"median_log2_rsem":11.69,"p90_log2_rsem":12.46},{"cancer_type":"Hepatobiliary Cancer","samples":366,"median_log2_rsem":11.12,"p90_log2_rsem":12.42},{"cancer_type":"Pleural Mesothelioma","samples":87,"median_log2_rsem":11.4,"p90_log2_rsem":12.34},{"cancer_type":"Renal Clear Cell Carcinoma","samples":510,"median_log2_rsem":11.32,"p90_log2_rsem":12.12},{"cancer_type":"Ocular Melanoma","samples":80,"median_log2_rsem":10.73,"p90_log2_rsem":11.76},{"cancer_type":"Sarcoma","samples":253,"median_log2_rsem":10.43,"p90_log2_rsem":11.54},{"cancer_type":"Thymic Epithelial Tumor","samples":119,"median_log2_rsem":9.64,"p90_log2_rsem":11.48},{"cancer_type":"Melanoma","samples":443,"median_log2_rsem":10.39,"p90_log2_rsem":11.39},{"cancer_type":"Glioblastoma","samples":160,"median_log2_rsem":10.18,"p90_log2_rsem":11.21},{"cancer_type":"Seminoma","samples":63,"median_log2_rsem":10.01,"p90_log2_rsem":10.57},{"cancer_type":"Glioma","samples":514,"median_log2_rsem":9.32,"p90_log2_rsem":10.37},{"cancer_type":"Miscellaneous Neuroepithelial Tumor","samples":31,"median_log2_rsem":8.51,"p90_log2_rsem":9.48},{"cancer_type":"Adrenocortical Carcinoma","samples":78,"median_log2_rsem":7.38,"p90_log2_rsem":9.13},{"cancer_type":"Mature B-Cell Neoplasms","samples":48,"median_log2_rsem":7.87,"p90_log2_rsem":8.78},{"cancer_type":"Pheochromocytoma","samples":147,"median_log2_rsem":7.75,"p90_log2_rsem":8.71},{"cancer_type":"Leukemia","samples":173,"median_log2_rsem":6.14,"p90_log2_rsem":7.34}]} ◀ result {"rows":[{"cancer_study_identifier":"acc_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"blca_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"brca_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"cesc_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"chol_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"coadread_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"dlbc_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"esca_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"gbm_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"hnsc_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"kich_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"kirc_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"kirp_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"laml_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"lgg_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"lihc_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"luad_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"lusc_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"meso_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"ov_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"paad_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"pcpg_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"prad_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"sarc_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"skcm_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"stad_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"tgct_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"thca_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"thym_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"ucec_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"ucs_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"uvm_tcga_pan_can_atlas_2018"}]} ▶ resolve_and_route { "studyIds": [ "acc_tcga_pan_can_atlas_2018", "blca_tcga_pan_can_atlas_2018", "brca_tcga_pan_can_atlas_2018", "cesc_tcga_pan_can_atlas_2018", "chol_tcga_pan_can_atlas_2018", "coadread_tcga_pan_can_atlas_2018", "dlbc_tcga_pan_can_atlas_2018", "esca_tcga_pan_can_atlas_2018", "gbm_tcga_pan_can_atlas_2018", "hnsc_tcga_pan_can_atlas_2018", "kich_tcga_pan_can_atlas_2018", "kirc_tcga_pan_can_atlas_2018", "kirp_tcga_pan_can_atlas_2018", "laml_tcga_pan_can_atlas_2018", "lgg_tcga_pan_can_atlas_2018", "lihc_tcga_pan_can_atlas_2018", "luad_tcga_pan_can_atlas_2018", "lusc_tcga_pan_can_atlas_2018", "meso_tcga_pan_can_atlas_2018", "ov_tcga_pan_can_atlas_2018", "paad_tcga_pan_can_atlas_2018", "pcpg_tcga_pan_can_atlas_2018", "prad_tcga_pan_can_atlas_2018", "sarc_tcga_pan_can_atlas_2018", "skcm_tcga_pan_can_atlas_2018", "stad_tcga_pan_can_atlas_2018", "tgct_tcga_pan_can_atlas_2018", "thca_tcga_pan_can_atlas_2018", "thym_tcga_pan_can_atlas_2018", "ucec_tcga_pan_can_atlas_2018", "ucs_tcga_pan_can_atlas_2018", "uvm_tcga_pan_can_atlas_2018" ] } ◀ result Error: result (81,891 characters across 1 line) exceeds maximum allowed tokens. Output has been saved to /home/ino/.claude-cbioportal-team/projects/-tmp-mcp-qa-claude-pqyf1b1k/413c3a0a-c576-4208-a446-0ace5d0abb15/tool-results/mcp-navigator-resolve_and_route-1790656969538.txt. Format: Plain text - For targeted searches (find a string): use grep on the file directly. - For analysis or summarization that requires reading the full content: the file's lines are too long for Read's offset/limit. Slice by character range via Bash instead — e.g. python3 -c 'print(open("/home/ino/.claude-cbioportal-team/projects/-tmp-mcp-qa-claude-pqyf1b1k/413c3a0a-c576-4208-a446-0ace5d0abb15/tool-results/mcp-navigator-resolve_and_route-1790656969538.txt").read()[A:B])' in ~80,000-char spans until you have read 100% of it. - If the Agent tool is available, do this inside a subagent so the full output stays out of your main context. Give it the instruction above verbatim, and be explicit about what it must return — e.g. "Slice /home/ino/.claude-cbioportal-team/projects/-tmp-mcp-qa-claude-pqyf1b1k/413c3a0a-c576-4208-a446-0ace5d0abb15/tool-results/mcp-navigator-resolve_and_route-1790656969538.txt in ~80,000-char spans via python (read()[A:B]) until you have read all 81,891 characters, then summarize and quote any key findings verbatim." A vague "summarize this" may lose detail. ▶ navigate_to_results_view { "studyIds": [ "acc_tcga_pan_can_atlas_2018", "blca_tcga_pan_can_atlas_2018", "brca_tcga_pan_can_atlas_2018", "cesc_tcga_pan_can_atlas_2018", "chol_tcga_pan_can_atlas_2018", "coadread_tcga_pan_can_atlas_2018", "dlbc_tcga_pan_can_atlas_2018", "esca_tcga_pan_can_atlas_2018", "gbm_tcga_pan_can_atlas_2018", "hnsc_tcga_pan_can_atlas_2018", "kich_tcga_pan_can_atlas_2018", "kirc_tcga_pan_can_atlas_2018", "kirp_tcga_pan_can_atlas_2018", "laml_tcga_pan_can_atlas_2018", "lgg_tcga_pan_can_atlas_2018", "lihc_tcga_pan_can_atlas_2018", "luad_tcga_pan_can_atlas_2018", "lusc_tcga_pan_can_atlas_2018", "meso_tcga_pan_can_atlas_2018", "ov_tcga_pan_can_atlas_2018", "paad_tcga_pan_can_atlas_2018", "pcpg_tcga_pan_can_atlas_2018", "prad_tcga_pan_can_atlas_2018", "sarc_tcga_pan_can_atlas_2018", "skcm_tcga_pan_can_atlas_2018", "stad_tcga_pan_can_atlas_2018", "tgct_tcga_pan_can_atlas_2018", "thca_tcga_pan_can_atlas_2018", "thym_tcga_pan_can_atlas_2018", "ucec_tcga_pan_can_atlas_2018", "ucs_tcga_pan_can_atlas_2018", "uvm_tcga_pan_can_atlas_2018" ], "genes": [ "ERBB2" ], "tab": "plots", "plotsHorzSelection": { "dataType": "clinical_attribute", "selectedGeneOption": "ERBB2" }, "plotsVertSelection": { "dataType": "MRNA_EXPRESSION", "selectedGeneOption": "ERBB2", "logScale": "true" } } ◀ result {"success":true,"message":"Navigating to https://www.cbioportal.org/results/plots?cancer_study_list=acc_tcga_pan_can_atlas_2018%2Cblca_tcga_pan_can_atlas_2018%2Cbrca_tcga_pan_can_atlas_2018%2Ccesc_tcga_pan_can_atlas_2018%2Cchol_tcga_pan_can_atlas_2018%2Ccoadread_tcga_pan_can_atlas_2018%2Cdlbc_tcga_pan_can_atlas_2018%2Cesca_tcga_pan_can_atlas_2018%2Cgbm_tcga_pan_can_atlas_2018%2Chnsc_tcga_pan_can_atlas_2018%2Ckich_tcga_pan_can_atlas_2018%2Ckirc_tcga_pan_can_atlas_2018%2Ckirp_tcga_pan_can_atlas_2018%2Claml_tcga_pan_can_atlas_2018%2Clgg_tcga_pan_can_atlas_2018%2Clihc_tcga_pan_can_atlas_2018%2Cluad_tcga_pan_can_atlas_2018%2Clusc_tcga_pan_can_atlas_2018%2Cmeso_tcga_pan_can_atlas_2018%2Cov_tcga_pan_can_atlas_2018%2Cpaad_tcga_pan_can_atlas_2018%2Cpcpg_tcga_pan_can_atlas_2018%2Cprad_tcga_pan_can_atlas_2018%2Csarc_tcga_pan_can_atlas_2018%2Cskcm_tcga_pan_can_atlas_2018%2Cstad_tcga_pan_can_atlas_2018%2Ctgct_tcga_pan_can_atlas_2018%2Cthca_tcga_pan_can_atlas_2018%2Cthym_tcga_pan_can_atlas_2018%2Cucec_tcga_pan_can_atlas_2018%2Cucs_tcga_pan_can_atlas_2018%2Cuvm_tcga_pan_can_atlas_2018&gene_list=ERBB2&case_set_id=all&plots_horz_selection=%7B%22selectedGeneOption%22%3A%222064%22%2C%22dataType%22%3A%22clinical_attribute%22%7D&plots_vert_selection=%7B%22selectedGeneOption%22%3A%222064%22%2C%22dataType%22%3A%22MRNA_EXPRESSION%22%2C%22logScale%22%3A%22true%22%7D","url":"https://www.cbioportal.org/results/plots?cancer_study_list=acc_tcga_pan_can_atlas_2018%2Cblca_tcga_pan_can_atlas_2018%2Cbrca_tcga_pan_can_atlas_2018%2Ccesc_tcga_pan_can_atlas_2018%2Cchol_tcga_pan_can_atlas_2018%2Ccoadread_tcga_pan_can_atlas_2018%2Cdlbc_tcga_pan_can_atlas_2018%2Cesca_tcga_pan_can_atlas_2018%2Cgbm_tcga_pan_can_atlas_2018%2Chnsc_tcga_pan_can_atlas_2018%2Ckich_tcga_pan_can_atlas_2018%2Ckirc_tcga_pan_can_atlas_2018%2Ckirp_tcga_pan_can_atlas_2018%2Claml_tcga_pan_can_atlas_2018%2Clgg_tcga_pan_can_atlas_2018%2Clihc_tcga_pan_can_atlas_2018%2Cluad_tcga_pan_can_atlas_2018%2Clusc_tcga_pan_can_atlas_2018%2Cmeso_tcga_pan_can_atlas_2018%2Cov_tcga_pan_can_atlas_2018%2Cpaad_tcga_pan_can_atlas_2018%2Cpcpg_tcga_pan_can_atlas_2018%2Cprad_tcga_pan_can_atlas_2018%2Csarc_tcga_pan_can_atlas_2018%2Cskcm_tcga_pan_can_atlas_2018%2Cstad_tcga_pan_can_atlas_2018%2Ctgct_tcga_pan_can_atlas_2018%2Cthca_tcga_pan_can_atlas_2018%2Cthym_tcga_pan_can_atlas_2018%2Cucec_tcga_pan_can_atlas_2018%2Cucs_tcga_pan_can_atlas_2018%2Cuvm_tcga_pan_can_atlas_2018&gene_list=ERBB2&case_set_id=all&plots_horz_selection=%7B%22selectedGeneOption%22%3A%222064%22%2C%22dataType%22%3A%22clinical_attribute%22%7D&plots_vert_selection=%7B%22selectedGeneOption%22%3A%222064%22%2C%22dataType%22%3A%22MRNA_EXPRESSION%22%2C%22logScale%22%3A%22true%22%7D","data":{"studyIds":["acc_tcga_pan_can_atlas_2018","blca_tcga_pan_can_atlas_2018","brca_tcga_pan_can_atlas_2018","cesc_tcga_pan_can_atlas_2018","chol_tcga_pan_can_atlas_2018","coadread_tcga_pan_can_atlas_2018","dlbc_tcga_pan_can_atlas_2018","esca_tcga_pan_can_atlas_2018","gbm_tcga_pan_can_atlas_2018","hn … (4667 more chars) ═ answer (success) **ERBB2 mRNA expression across the 32 TCGA PanCancer Atlas studies** The values are RSEM (`rna_seq_v2_mrna`), shown as log2(RSEM+1). This data is batch-normalized across the PanCancer Atlas studies, so the values can be compared between cancer types. Z-scores can't be used for this comparison because each study is centered on its own samples. The unit is samples, with about 10,000 tumors in total. Cancer types are ranked by the 90th percentile (p90) because amplification-driven expression shows up in the upper tail, not the median. | Cancer type | Samples | Median | p90 | |---|---|---|---| | Breast Cancer | 1,082 | 12.78 | **14.75** | | Bladder Cancer | 407 | 12.88 | 14.28 | | Esophagogastric Cancer | 593 | 12.02 | 13.75 | | Renal Non-Clear Cell Carcinoma | 348 | 12.63 | 13.38 | | Endometrial Cancer | 584 | 12.24 | 13.36 | | Thyroid Cancer | 498 | 12.71 | 13.28 | | Non-Small Cell Lung Cancer | 994 | 11.96 | 13.23 | | Cervical Cancer | 294 | 12.15 | 13.22 | | Prostate Cancer | 493 | 12.43 | 13.03 | | Pancreatic Cancer | 177 | 12.29 | 12.99 | | Cholangiocarcinoma | 36 | 12.42 | 12.96 | | Colorectal Cancer | 592 | 12.20 | 12.90 | | Ovarian Epithelial Tumor | 300 | 11.98 | 12.80 | | Non-Seminomatous Germ Cell Tumor | 86 | 11.57 | 12.63 | | Head and Neck Cancer | 515 | 11.69 | 12.46 | | Hepatobiliary Cancer | 366 | 11.12 | 12.42 | | Pleural Mesothelioma | 87 | 11.40 | 12.34 | | Renal Clear Cell Carcinoma | 510 | 11.32 | 12.12 | | Ocular Melanoma | 80 | 10.73 | 11.76 | | Sarcoma | 253 | 10.43 | 11.54 | | Thymic Epithelial Tumor | 119 | 9.64 | 11.48 | | Melanoma | 443 | 10.39 | 11.39 | | Glioblastoma | 160 | 10.18 | 11.21 | | Seminoma | 63 | 10.01 | 10.57 | | Glioma | 514 | 9.32 | 10.37 | | Miscellaneous Neuroepithelial Tumor | 31 | 8.51 | 9.48 | | Adrenocortical Carcinoma | 78 | 7.38 | 9.13 | | Mature B-Cell Neoplasms | 48 | 7.87 | 8.78 | | Pheochromocytoma | 147 | 7.75 | 8.71 | | Leukemia | 173 | 6.14 | 7.34 | **Main patterns** - **Highest tail:** Breast cancer has the highest p90 (14.75). Its median (12.78) is similar to many other epithelial cancers, so the high p90 means a subset of tumors have very high expression. - **Highest median:** Bladder has the highest median (12.88), just above breast (12.78). Its p90 (14.28) is second to breast. - **Next tier:** Esophagogastric, endometrial, lung, cervical and renal non-clear cell cancers have p90 values of about 13.2–13.8. This is consistent with a smaller high-expressing subset in these cancers. - **Intermediate:** Thyroid, prostate, pancreatic, colorectal, ovarian and head and neck cancers have medians of about 11.7–12.7 and a narrower spread. - **Lowest:** Non-epithelial and hematologic or neuroendocrine cancers have the lowest expression. Glioma, adrenocortical carcinoma, pheochromocytoma and B-cell neoplasms have medians of about 7.4–9.3, and leukemia is lowest at 6.14. These are descriptive summaries. I didn't run any statistical tests between cancer types. I haven't linked the high-expression tail to ERBB2 amplification in this analysis. **Explore the full distributions:** [ERBB2 mRNA expression by cancer type (Plots tab, all 32 TCGA PanCancer Atlas studies)](https://www.cbioportal.org/results/plots?cancer_study_list=acc_tcga_pan_can_atlas_2018%2Cblca_tcga_pan_can_atlas_2018%2Cbrca_tcga_pan_can_atlas_2018%2Ccesc_tcga_pan_can_atlas_2018%2Cchol_tcga_pan_can_atlas_2018%2Ccoadread_tcga_pan_can_atlas_2018%2Cdlbc_tcga_pan_can_atlas_2018%2Cesca_tcga_pan_can_atlas_2018%2Cgbm_tcga_pan_can_atlas_2018%2Chnsc_tcga_pan_can_atlas_2018%2Ckich_tcga_pan_can_atlas_2018%2Ckirc_tcga_pan_can_atlas_2018%2Ckirp_tcga_pan_can_atlas_2018%2Claml_tcga_pan_can_atlas_2018%2Clgg_tcga_pan_can_atlas_2018%2Clihc_tcga_pan_can_atlas_2018%2Cluad_tcga_pan_can_atlas_2018%2Clusc_tcga_pan_can_atlas_2018%2Cmeso_tcga_pan_can_atlas_2018%2Cov_tcga_pan_can_atlas_2018%2Cpaad_tcga_pan_can_atlas_2018%2Cpcpg_tcga_pan_can_atlas_2018%2Cprad_tcga_pan_can_atlas_2018%2Csarc_tcga_pan_can_atlas_2018%2Cskcm_tcga_pan_can_atlas_2018%2Cstad_tcga_pan_can_atlas_2018%2Ctgct_tcga_pan_can_atlas_2018%2Cthca_tcga_pan_can_atlas_2018%2Cthym_tcga_pan_can_atlas_2018%2Cucec_tcga_pan_can_atlas_2018%2Cucs_tcga_pan_can_atlas_2018%2Cuvm_tcga_pan_can_atlas_2018&gene_list=ERBB2&case_set_id=all&plots_horz_selection=%7B%22selectedGeneOption%22%3A%222064%22%2C%22dataType%22%3A%22clinical_attribute%22%7D&plots_vert_selection=%7B%22selectedGeneOption%22%3A%222064%22%2C%22dataType%22%3A%22MRNA_EXPRESSION%22%2C%22logScale%22%3A%22true%22%7D) The plot opens with ERBB2 mRNA expression on the y-axis, and you set the x-axis to a clinical attribute such as cancer type. Some studies group differently from the cancer-type labels above. For example, the glioma group combines LGG samples and the colorectal group combines COAD and READ. Overlapping cohorts can also double-count samples.