Q (haiku): What is the correlation coefficient between EGFR copy number and EGFR mRNA expression in ovarian cancer? ▶ read_guide { "uri": "cbioportal://gene-expression-guide" } ◀ result # Gene Expression Analysis Guide This guide covers continuous-value genomic data: gene **expression**, **copy number** values, **methylation**, and related profile types. Mutation/CNA/SV *frequency* analysis lives in `cbioportal://mutation-frequency-guide`. ## Where this data lives Continuous per-sample-per-gene values are stored in `genetic_alteration_derived`: | Column | Description | |---|---| | `sample_unique_id` | `_` | | `cancer_study_identifier` | study scope | | `hugo_gene_symbol` | gene | | `profile_type` | which assay/normalization (see below) | | `alteration_value` | the actual value — stored as Nullable(String); cast with `toFloat64OrNull` | `alteration_value` is a string because the same column hosts many different value scales. The `''` and `'NA'` sentinels mean "missing"; always filter them out and use `toFloat64OrNull(alteration_value) IS NOT NULL` for downstream math. ## Discovering profile types for a study Different studies expose different profile types depending on what assays were run and how the data was normalized. Always check what a specific study supports before picking one: ```sql SELECT DISTINCT profile_type FROM genetic_alteration_derived WHERE cancer_study_identifier = 'brca_metabric' ORDER BY profile_type; ``` Common values across the public portal: | Family | Profile types | |---|---| | mRNA expression | `rna_seq_v2_mrna`, `rna_seq_v2_mrna_median_Zscores`, `rna_seq_v2_mrna_median_all_sample_Zscores` (TCGA PanCancer Atlas), `mrna`, `mrna_median_Zscores`, `mrna_seq_v2_rsem`, `mrna_seq_v2_rsem_Zscores`, `mrna_seq_cpm`, `mrna_seq_fpkm`, `mrna_U133`, `mrna_outliers` | | Copy number (continuous) | `cna`, `linear_CNA`, `log2CNA`, `cna_consensus`, `cna_rae`, `gistic` | | Methylation | `methylation_hm27`, `methylation_hm450`, `methylation_epic`, `methylation_promoters_rrbs` | | miRNA | `mirna`, `mirna_median_Zscores` | | Protein | `protein_quantification`, `protein_level`, `RPPA` | **Z-score vs raw choice.** When the user asks "is X correlated with Y", either works for Spearman (rank-based) — Pearson would care. Default to the non-Z-score variant if both exist, and call out which one in the response. ## Canonical recipe — Spearman correlation between two genes ```sql SELECT * FROM gene_pair_coexpression( study = 'brca_metabric', gene_a = 'TP53', gene_b = 'MYC', profile_type = 'mrna' ); ``` Returns one row: `(gene_a, gene_b, profile_type, spearman_correlation, num_samples)`. - `spearman_correlation` in [−1, 1]; `NULL` when fewer than 3 valid paired samples. - Mirrors cbioportal-backend's `ClickhouseCoExpressionMapper.getCoExpressions`, simplified to a pair lookup (the backend computes one ref gene vs ALL other genes for the coexpression page; here the agent asks about a specific pair). ### Verified examples | Study | gene_a | gene_b | profile_type | spearman | n | |---|---|---|---|---|---| | `brca_metabric` | TP53 | MYC | `mrna` | 0.118 | 1980 | | … (4931 more chars) ▶ list_studies { "search": "ovarian cancer", "verbose": true, "limit": 20 } ◀ result {"result":[{"cancer_study_identifier":"msk_met_2021","name":"MSK MetTropism (MSK, Cell 2021)","description":"MSK-MET (Memorial Sloan Kettering - Metastatic Events and Tropisms) is a pan-cancer cohort of tumor genomic and clinical outcome data from 25,000 patients. The dataset identifies associations between tumor genomic alterations and patterns of metastatic dissemination across 50 tumor types; showing that chromosomal instability is strongly correlated with metastatic burden in some tumor types, like prostate and lung adenocarcinomas and HR+/HER2+ breast ductal carcinoma, but not in others, such as colorectal MSS, pancreatic adenocarcinoma and high-grade serous ovarian cancer. The study also identifies somatic alterations associated with increased metastatic burden and routes of metastatic spread. Our data offers a resource for the investigation of the biologic basis for metastatic spread and highlights the role of chromosomal instability in cancer progression. This data is available under the Creative Commons BY-NC-ND 4.0 license.","type_of_cancer_id":"mixed","sample_count":25775,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=msk_met_2021"},{"cancer_study_identifier":"hgsoc_tcga_gdc","name":"High-Grade Serous Ovarian Cancer (TCGA GDC, 2025)","description":"TCGA High-Grade Serous Ovarian Cancer. Source data from NCI GDC and generated in Aug 2025 using Cancer Data Aggregator.","type_of_cancer_id":"hgsoc","sample_count":604,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=hgsoc_tcga_gdc"},{"cancer_study_identifier":"ovary_geomx_gray_foundation_2024","name":"Ovarian Cancer (Gray Foundation, Cancer Discov 2024)","description":"Microregion transcriptomic data (GeoMx) and multiplex imaging (CyCIF) for High-Grade Serous ovarian Cancer Cases (HGSOC). This resource was supported by the Gray Foundation.","type_of_cancer_id":"ovary","sample_count":567,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=ovary_geomx_gray_foundation_2024"},{"cancer_study_identifier":"ov_tcga_pub","name":"Ovarian Serous Cystadenocarcinoma (TCGA, Nature 2011)","description":"Whole exome sequencing of 489 high-grade serous ovarian adenocarcinoma tumor/normal pairs. The Cancer Genome Atlas (TCGA) Serous Ovarian Cancer project. 489 cases.
Raw data via the TCGA Data Portal.","type_of_cancer_id":"hgsoc","sample_count":489,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=ov_tcga_pub"},{"cancer_study_identifier":"lgsoc_mapk_msk_2022","name":"Low-Grade Serous Ovarian Cancer (MSK, Clin Cancer Res 2022)","description":"Characterization of the somatic mutational landscape of MAPK pathway in 119 low-grade serous ovarian carcino … (3592 more chars) ▶ get_study_guide { "study_id": "ov_tcga_pan_can_atlas_2018" } ◀ result # Ovarian Serous Cystadenocarcinoma (TCGA, PanCancer Atlas) **Study ID:** `ov_tcga_pan_can_atlas_2018` See `_tcga_pancan_template.md` for common TCGA clinical attributes. ## Study-Specific Attributes ### Histology | Attribute | Description | Values (samples) | |-----------|-------------|--------| | `GRADE` | Histologic grade | G3 400, G2 65, GX 7, G1 5, GB 2, G4 1, blank 105 | | `CANCER_TYPE_DETAILED` | Histology | Serous Ovarian Cancer (all 585) | ### Not available in this study - **Platinum sensitivity and residual disease**: no treatment-response or surgical-outcome attributes. Survival (`OS_*`, `PFS_*`, `DFS_*`, `DSS_*`) is the only outcome data. - **BRCA / HRD status**: no clinical status attributes. Derive BRCA1/BRCA2 status from mutation (somatic calls only) and CNA data; there is no HRD score. - `AJCC_PATHOLOGIC_TUMOR_STAGE` is blank for all samples. `SUBTYPE` is `OV` (177) or blank (408) and carries no molecular subtype. ## Key Genes | Gene | Frequency | Notes | |------|-----------|-------| | TP53 | ~96% | Near-universal in high-grade serous | | BRCA1 | ~10% germline | PARP inhibitor sensitivity | | BRCA2 | ~6% germline | PARP inhibitor sensitivity | | NF1 | ~4% | RAS pathway | | RB1 | ~2% | Cell cycle | | CDK12 | ~3% | DNA repair | ## Notes - High-grade serous ovarian cancer (HGSOC) is genomically distinct from other ovarian subtypes - TP53 is mutated in virtually all HGSOC - BRCA1/2 mutations (germline or somatic) predict PARP inhibitor response - HRD score predicts platinum and PARP sensitivity even without BRCA mutation - Copy number alterations more prominent than mutations (except TP53) ▶ clickhouse_run_select_query { "query": " SELECT DISTINCT profile_type FROM genetic_alteration_derived WHERE cancer_study_identifier = 'ov_tcga_pan_can_atlas_2018' AND hugo_gene_symbol = 'EGFR' ORDER BY profile_type ", "max_rows": 50 } ◀ result {"rows":[{"profile_type":"gistic"},{"profile_type":"log2CNA"},{"profile_type":"protein_quantification"},{"profile_type":"protein_quantification_zscores"},{"profile_type":"rna_seq_v2_mrna"},{"profile_type":"rna_seq_v2_mrna_median_Zscores"},{"profile_type":"rna_seq_v2_mrna_median_all_sample_Zscores"},{"profile_type":"rppa"},{"profile_type":"rppa_Zscores"}]} ▶ clickhouse_run_select_query { "query": " WITH cna_data AS ( SELECT sample_unique_id, toFloat64OrNull(alteration_value) AS cn_value FROM genetic_alteration_derived WHERE cancer_study_identifier = 'ov_tcga_pan_can_atlas_2018' AND hugo_gene_symbol = 'EGFR' AND profile_type = 'log2CNA' AND alteration_value NOT IN ('', 'NA') AND toFloat64OrNull(alteration_value) IS NOT NULL ), mrna_data AS ( SELECT sample_unique_id, toFloat64OrNull(alteration_value) AS mrna_value FROM genetic_alteration_derived WHERE cancer_study_identifier = 'ov_tcga_pan_can_atlas_2018' AND hugo_gene_symbol = 'EGFR' AND profile_type = 'rna_seq_v2_mrna' AND alteration_value NOT IN ('', 'NA') AND toFloat64OrNull(alteration_value) IS NOT NULL ) SELECT round(spearmanRho(c.cn_value, m.mrna_value), 4) AS spearman_correlation, count() AS num_samples FROM cna_data c INNER JOIN mrna_data m USING (sample_unique_id) ", "max_rows": 10 } ◀ result {"error_message":"Query execution failed: HTTPDriver for https://dl96orhu96.us-east-1.aws.clickhouse.cloud:8443 received ClickHouse error code 46\n Code: 46. DB::Exception: Function with name `spearmanRho` does not exist. In scope WITH cna_data AS (SELECT sample_unique_id, toFloat64OrNull(alteration_value) AS cn_value FROM genetic_alteration_derived WHERE (cancer_study_identifier = 'ov_tcga_pan_can_atlas_2018') AND (hugo_gene_symbol = 'EGFR') AND (profile_type = 'log2CNA') AND (alteration_value NOT IN ('', 'NA')) AND (toFloat64OrNull(alteration_value) IS NOT NULL)), mrna_data AS (SELECT sample_unique_id, toFloat64OrNull(alteration_value) AS mrna_value FROM genetic_alteration_derived WHERE (cancer_study_identifier = 'ov_tcga_pan_can_atlas_2018') AND (hugo_gene_symbol = 'EGFR') AND (profile_type = 'rna_seq_v2_mrna') AND (alteration_value NOT IN ('', 'NA')) AND (toFloat64OrNull(alteration_value) IS NOT NULL)) SELECT round(spearmanRho(c.cn_value, m.mrna_value), 4) AS spearman_correlation, count() AS num_samples FROM cna_data AS c INNER JOIN mrna_data AS m USING (sample_unique_id). (UNKNOWN_FUNCTION) (version 26.4.1.2359 (official build))\n"} ▶ clickhouse_run_select_query { "query": " WITH cna_data AS ( SELECT sample_unique_id, toFloat64OrNull(alteration_value) AS cn_value FROM genetic_alteration_derived WHERE cancer_study_identifier = 'ov_tcga_pan_can_atlas_2018' AND hugo_gene_symbol = 'EGFR' AND profile_type = 'log2CNA' AND alteration_value NOT IN ('', 'NA') AND toFloat64OrNull(alteration_value) IS NOT NULL ), mrna_data AS ( SELECT sample_unique_id, toFloat64OrNull(alteration_value) AS mrna_value FROM genetic_alteration_derived WHERE cancer_study_identifier = 'ov_tcga_pan_can_atlas_2018' AND hugo_gene_symbol = 'EGFR' AND profile_type = 'rna_seq_v2_mrna' AND alteration_value NOT IN ('', 'NA') AND toFloat64OrNull(alteration_value) IS NOT NULL ) SELECT round(rankCorr(c.cn_value, m.mrna_value), 4) AS spearman_correlation, count() AS num_samples FROM cna_data c INNER JOIN mrna_data m USING (sample_unique_id) ", "max_rows": 10 } ◀ result {"rows":[{"spearman_correlation":0.3635,"num_samples":295}]} ▶ clickhouse_run_select_query { "query": " -- Verify using the canonical gene_pair_coexpression view for EGFR CN vs mRNA SELECT * FROM gene_pair_coexpression( study = 'ov_tcga_pan_can_atlas_2018', gene_a = 'EGFR', gene_b = 'EGFR', profile_type_a = 'log2CNA', profile_type_b = 'rna_seq_v2_mrna' ) ", "max_rows": 10 } ◀ result {"error_message":"Query execution failed: HTTPDriver for https://dl96orhu96.us-east-1.aws.clickhouse.cloud:8443 received ClickHouse error code 456\n Code: 456. DB::Exception: Substitution `profile_type` is not set. (UNKNOWN_QUERY_PARAMETER) (version 26.4.1.2359 (official build))\n"} ▶ clickhouse_run_select_query { "query": " -- Summary statistics for EGFR copy number and mRNA expression WITH cna_stats AS ( SELECT count() AS cna_samples, round(avg(toFloat64OrNull(alteration_value)), 3) AS cna_mean, round(median(toFloat64OrNull(alteration_value)), 3) AS cna_median, round(stddevSamp(toFloat64OrNull(alteration_value)), 3) AS cna_stdev, round(min(toFloat64OrNull(alteration_value)), 3) AS cna_min, round(max(toFloat64OrNull(alteration_value)), 3) AS cna_max FROM genetic_alteration_derived WHERE cancer_study_identifier = 'ov_tcga_pan_can_atlas_2018' AND hugo_gene_symbol = 'EGFR' AND profile_type = 'log2CNA' AND alteration_value NOT IN ('', 'NA') AND toFloat64OrNull(alteration_value) IS NOT NULL ), mrna_stats AS ( SELECT count() AS mrna_samples, round(avg(toFloat64OrNull(alteration_value)), 3) AS mrna_mean, round(median(toFloat64OrNull(alteration_value)), 3) AS mrna_median, round(stddevSamp(toFloat64OrNull(alteration_value)), 3) AS mrna_stdev, round(min(toFloat64OrNull(alteration_value)), 3) AS mrna_min, round(max(toFloat64OrNull(alteration_value)), 3) AS mrna_max FROM genetic_alteration_derived WHERE cancer_study_identifier = 'ov_tcga_pan_can_atlas_2018' AND hugo_gene_symbol = 'EGFR' AND profile_type = 'rna_seq_v2_mrna' AND alteration_value NOT IN ('', 'NA') AND toFloat64OrNull(alteration_value) IS NOT NULL ) SELECT * FROM cna_stats, mrna_stats ", "max_rows": 10 } ◀ result {"rows":[{"cna_samples":572,"cna_mean":0.111,"cna_median":0.02,"cna_stdev":0.418,"cna_min":-1,"cna_max":2.26,"mrna_samples":300,"mrna_mean":280.701,"mrna_median":215.533,"mrna_stdev":247.735,"mrna_min":5.947,"mrna_max":2504.067}]} ▶ resolve_and_route { "studyIds": [ "ov_tcga_pan_can_atlas_2018" ] } ◀ result {"success":true,"message":"Found 1 study. Use the metadata to call the appropriate navigation tool(s).","data":{"totalCount":1,"studiesWithMetadata":[{"studyId":"ov_tcga_pan_can_atlas_2018","name":"Ovarian Serous Cystadenocarcinoma (TCGA, PanCancer Atlas)","sampleCount":585,"studyViewUrl":"https://www.cbioportal.org/study?id=ov_tcga_pan_can_atlas_2018","metadata":{"clinicalAttributeIds":["AGE","AJCC_PATHOLOGIC_TUMOR_STAGE","AJCC_STAGING_EDITION","ANEUPLOIDY_SCORE","BUFFA_HYPOXIA_SCORE","CANCER_TYPE","CANCER_TYPE_ACRONYM","CANCER_TYPE_DETAILED","DAYS_LAST_FOLLOWUP","DAYS_TO_BIRTH","DAYS_TO_INITIAL_PATHOLOGIC_DIAGNOSIS","DFS_MONTHS","DFS_STATUS","DSS_MONTHS","DSS_STATUS","ETHNICITY","FORM_COMPLETION_DATE","FRACTION_GENOME_ALTERED","GENETIC_ANCESTRY_LABEL","GRADE","HISTORY_NEOADJUVANT_TRTYN","ICD_10","ICD_O_3_HISTOLOGY","ICD_O_3_SITE","INFORMED_CONSENT_VERIFIED","IN_PANCANPATHWAYS_FREEZE","MSI_SCORE_MANTIS","MSI_SENSOR_SCORE","MUTATION_COUNT","NEW_TUMOR_EVENT_AFTER_INITIAL_TREATMENT","ONCOTREE_CODE","OS_MONTHS","OS_STATUS","OTHER_PATIENT_ID","PATH_M_STAGE","PATH_N_STAGE","PATH_T_STAGE","PERSON_NEOPLASM_CANCER_STATUS","PFS_MONTHS","PFS_STATUS","PRIMARY_LYMPH_NODE_PRESENTATION_ASSESSMENT","PRIOR_DX","RACE","RADIATION_THERAPY","RAGNUM_HYPOXIA_SCORE","SAMPLE_COUNT","SAMPLE_TYPE","SEX","SOMATIC_STATUS","SUBTYPE","TBL_SCORE","TISSUE_PROSPECTIVE_COLLECTION_INDICATOR","TISSUE_RETROSPECTIVE_COLLECTION_INDICATOR","TISSUE_SOURCE_SITE","TISSUE_SOURCE_SITE_CODE","TMB_NONSYNONYMOUS","TUMOR_TISSUE_SITE","TUMOR_TYPE","WEIGHT","WINTER_HYPOXIA_SCORE"],"molecularProfileIds":["ov_tcga_pan_can_atlas_2018_gistic","ov_tcga_pan_can_atlas_2018_log2CNA","ov_tcga_pan_can_atlas_2018_mutations","ov_tcga_pan_can_atlas_2018_protein_quantification","ov_tcga_pan_can_atlas_2018_protein_quantification_zscores","ov_tcga_pan_can_atlas_2018_rna_seq_v2_mrna","ov_tcga_pan_can_atlas_2018_rna_seq_v2_mrna_median_Zscores","ov_tcga_pan_can_atlas_2018_rna_seq_v2_mrna_median_all_sample_Zscores","ov_tcga_pan_can_atlas_2018_rppa","ov_tcga_pan_can_atlas_2018_rppa_Zscores","ov_tcga_pan_can_atlas_2018_structural_variants"],"genericAssayProfiles":["ov_tcga_pan_can_atlas_2018_armlevel_cna","ov_tcga_pan_can_atlas_2018_genetic_ancestry","ov_tcga_pan_can_atlas_2018_methylation_hm27_hm450_merge","ov_tcga_pan_can_atlas_2018_methylation_hm450","ov_tcga_pan_can_atlas_2018_phosphoprotein_quantification"],"heatmapProfileIds":["ov_tcga_pan_can_atlas_2018_protein_quantification_zscores","ov_tcga_pan_can_atlas_2018_rna_seq_v2_mrna_median_Zscores","ov_tcga_pan_can_atlas_2018_rna_seq_v2_mrna_median_all_sample_Zscores","ov_tcga_pan_can_atlas_2018_rppa_Zscores"],"availableComparisonTabs":["overlap","clinical","survival","alterations","mutations","mrna","protein","generic_assay_armlevel_cna","generic_assay_genetic_ancestry","generic_assay_methylation","generic_assay_phosphosite_quantification"]}}]}} ▶ navigate_to_results_view { "studyIds": [ "ov_tcga_pan_can_atlas_2018" ], "genes": [ "EGFR" ], "tab": "plots", "plotsHorzSelection": { "selectedGeneOption": "EGFR", "dataType": "CNA", "selectedDataSourceOption": "log2CNA" }, "plotsVertSelection": { "selectedGeneOption": "EGFR", "dataType": "Expression", "selectedDataSourceOption": "rna_seq_v2_mrna" } } ◀ result {"success":true,"message":"Navigating to https://www.cbioportal.org/results/plots?cancer_study_list=ov_tcga_pan_can_atlas_2018&gene_list=EGFR&case_set_id=ov_tcga_pan_can_atlas_2018_all&plots_horz_selection=%7B%22selectedGeneOption%22%3A%221956%22%2C%22dataType%22%3A%22CNA%22%2C%22selectedDataSourceOption%22%3A%22log2CNA%22%7D&plots_vert_selection=%7B%22selectedGeneOption%22%3A%221956%22%2C%22dataType%22%3A%22Expression%22%2C%22selectedDataSourceOption%22%3A%22rna_seq_v2_mrna%22%7D","url":"https://www.cbioportal.org/results/plots?cancer_study_list=ov_tcga_pan_can_atlas_2018&gene_list=EGFR&case_set_id=ov_tcga_pan_can_atlas_2018_all&plots_horz_selection=%7B%22selectedGeneOption%22%3A%221956%22%2C%22dataType%22%3A%22CNA%22%2C%22selectedDataSourceOption%22%3A%22log2CNA%22%7D&plots_vert_selection=%7B%22selectedGeneOption%22%3A%221956%22%2C%22dataType%22%3A%22Expression%22%2C%22selectedDataSourceOption%22%3A%22rna_seq_v2_mrna%22%7D","data":{"studyIds":["ov_tcga_pan_can_atlas_2018"],"studies":[{"studyId":"ov_tcga_pan_can_atlas_2018","name":"Ovarian Serous Cystadenocarcinoma (TCGA, PanCancer Atlas)","sampleCount":585}],"genes":["EGFR"],"caseSetId":"ov_tcga_pan_can_atlas_2018_all","pageDescription":"Scatter/box plots: EGFR expression (y-axis) by EGFR cna (x-axis)."}} ═ answer (success) ## Results **EGFR copy number vs. mRNA expression correlation in ovarian cancer:** ### Spearman Correlation Coefficient: **0.3635** (295 samples with paired data) This is a **moderate positive correlation**, indicating that EGFR copy number gains are associated with higher mRNA expression in high-grade serous ovarian cancer (HGSOC). #### Summary Statistics | Metric | EGFR Copy Number (log₂) | EGFR mRNA (RSEM) | |--------|------------------------|------------------| | Samples | 572 | 300 | | Mean | 0.111 | 280.7 | | Median | 0.02 | 215.5 | | Std Dev | 0.418 | 247.7 | | Range | −1.0 to 2.26 | 5.9 to 2,504 | **Interpretation:** - A Spearman rho of 0.36 suggests a meaningful dose-response relationship: tumors with EGFR amplification tend to have elevated mRNA expression. - This is consistent with the expected gene dosage effect in cancer—more DNA copies → higher transcription. - The moderate (not strong) correlation reflects biological variability: not all copy number changes translate to proportional expression changes due to regulation at transcription and post-transcription levels. ### Visualize the Data [View EGFR copy number vs. mRNA scatter plot](https://www.cbioportal.org/results/plots?cancer_study_list=ov_tcga_pan_can_atlas_2018&gene_list=EGFR&case_set_id=ov_tcga_pan_can_atlas_2018_all&plots_horz_selection=%7B%22selectedGeneOption%22%3A%221956%22%2C%22dataType%22%3A%22CNA%22%2C%22selectedDataSourceOption%22%3A%22log2CNA%22%7D&plots_vert_selection=%7B%22selectedGeneOption%22%3A%221956%22%2C%22dataType%22%3A%22Expression%22%2C%22selectedDataSourceOption%22%3A%22rna_seq_v2_mrna%22%7D) This will show the scatter plot with EGFR copy number on the x-axis and EGFR mRNA expression on the y-axis in the TCGA Ovarian Serous Cystadenocarcinoma PanCancer Atlas (585 total samples).