Q (sonnet-5.5): Can you analyze thyroid hormone gene expression by diagnosis in the Pediatric Brain Tumor Atlas (PBTA, Provisional) cohort? ▶ list_studies { "search": "pediatric brain tumor atlas" } ▶ read_guide { "uri": "cbioportal://gene-expression-guide" } ◀ result {"result":[{"note":"No studies matched 'pediatric brain tumor atlas'. Before saying the study isn't available, read cbioportal://study-resolution-guide: names can omit centers or consortia (HTAN centers appear only as atlas codes in study ids, e.g. hta9 = OHSU), try fewer words, or it may be on another cBioPortal instance."}]} ◀ result # Gene Expression Analysis Guide This guide covers continuous-value genomic data: gene **expression**, **copy number** values, **methylation**, and related profile types. Mutation/CNA/SV *frequency* analysis lives in `cbioportal://mutation-frequency-guide`. ## Where this data lives Continuous per-sample-per-gene values are stored in `genetic_alteration_derived`: | Column | Description | |---|---| | `sample_unique_id` | `_` | | `cancer_study_identifier` | study scope | | `hugo_gene_symbol` | gene | | `profile_type` | which assay/normalization (see below) | | `alteration_value` | the actual value — stored as Nullable(String); cast with `toFloat64OrNull` | `alteration_value` is a string because the same column hosts many different value scales. The `''` and `'NA'` sentinels mean "missing"; always filter them out and use `toFloat64OrNull(alteration_value) IS NOT NULL` for downstream math. ## Discovering profile types for a study Different studies expose different profile types depending on what assays were run and how the data was normalized. Always check what a specific study supports before picking one: ```sql SELECT DISTINCT profile_type FROM genetic_alteration_derived WHERE cancer_study_identifier = 'brca_metabric' ORDER BY profile_type; ``` Common values across the public portal: | Family | Profile types | |---|---| | mRNA expression | `rna_seq_v2_mrna`, `rna_seq_v2_mrna_median_Zscores`, `rna_seq_v2_mrna_median_all_sample_Zscores` (TCGA PanCancer Atlas), `mrna`, `mrna_median_Zscores`, `mrna_seq_v2_rsem`, `mrna_seq_v2_rsem_Zscores`, `mrna_seq_cpm`, `mrna_seq_fpkm`, `mrna_U133`, `mrna_outliers` | | Copy number (continuous) | `cna`, `linear_CNA`, `log2CNA`, `cna_consensus`, `cna_rae`, `gistic` | | Methylation | `methylation_hm27`, `methylation_hm450`, `methylation_epic`, `methylation_promoters_rrbs` | | miRNA | `mirna`, `mirna_median_Zscores` | | Protein | `protein_quantification`, `protein_level`, `RPPA` | **Z-score vs raw choice.** When the user asks "is X correlated with Y", either works for Spearman (rank-based) — Pearson would care. Default to the non-Z-score variant if both exist, and call out which one in the response. ## Canonical recipe — Spearman correlation between two genes ```sql SELECT * FROM gene_pair_coexpression( study = 'brca_metabric', gene_a = 'TP53', gene_b = 'MYC', profile_type = 'mrna' ); ``` Returns one row: `(gene_a, gene_b, profile_type, spearman_correlation, num_samples)`. - `spearman_correlation` in [−1, 1]; `NULL` when fewer than 3 valid paired samples. - Mirrors cbioportal-backend's `ClickhouseCoExpressionMapper.getCoExpressions`, simplified to a pair lookup (the backend computes one ref gene vs ALL other genes for the coexpression page; here the agent asks about a specific pair). ### Verified examples | Study | gene_a | gene_b | profile_type | spearman | n | |---|---|---|---|---|---| | `brca_metabric` | TP53 | MYC | `mrna` | 0.118 | 1980 | | … (4931 more chars) ▶ list_studies { "search": "PBTA" } ▶ read_guide { "uri": "cbioportal://study-resolution-guide" } ◀ result {"result":[{"note":"No studies matched 'PBTA'. Before saying the study isn't available, read cbioportal://study-resolution-guide: names can omit centers or consortia (HTAN centers appear only as atlas codes in study ids, e.g. hta9 = OHSU), try fewer words, or it may be on another cBioPortal instance."}]} ◀ result # Study Resolution Guide Use this guide when the user names a study, cohort, portal, or data source that may not exist in the connected cBioPortal deployment. ## Routing Triggers Read this guide when the user mentions: - PBTA, Pediatric Brain Tumor Atlas, pediatric cBioPortal, Kids First - GENIE, AACR GENIE, MSK private cohorts, institutional cohorts - "download study", "which study", "find cohort", "data from [portal]" - a named cohort that `list_studies(search=...)` does not find - HTAN, or an HTAN center (OHSU, MSK, Vanderbilt, ...) ## Core Rules 1. Resolve the requested study before substituting another study. 2. If the requested study is not in this deployment, say so explicitly. 3. Do not silently analyze a substitute cohort. 4. If the user agrees to a substitute, keep a one-line scope caveat when reporting numbers. ## Known External cBioPortal Instances These are not necessarily queryable from this MCP server, but they are useful redirects: | User wording | Likely external instance | Scope | |---|---|---| | PBTA, Pediatric Brain Tumor Atlas, pediatric brain tumors | https://pedcbioportal.kidsfirstdrc.org/ | Pediatric cancer studies, including pediatric brain tumor cohorts | | GENIE | https://genie.cbioportal.org/ | AACR GENIE data access, depending on release and permissions | | MSK private / institutional cohorts | private institutional cBioPortal deployments | Not queryable from public cBioPortal unless exported to the public database | ## HTAN Studies HTAN study ids carry the atlas code (`hta9`, `hta8`, ...), and the names don't always name the center, so `list_studies(search="OHSU HTAN")` finds nothing. Search `list_studies(search="HTAN")` and match the center by atlas code: | Atlas code | Center | Study | |---|---|---| | `hta8` | MSK | `crc_hta8_htan_2024` | | `hta9` | OHSU | `brca_hta9_htan_2022` ("Breast Cancer (HTAN, 2022)") | | `hta11` | Vanderbilt | `crc_hta11_htan_2021` | For HTAN imaging (Minerva, MxIF, H&E), continue with the external-resources-guide. ## Study Resolution Workflow 1. Call `list_studies(search=...)` with the user's exact study/cohort phrase and close variants. 2. If a cancer type is mentioned, call `search_oncotree(search_term)` before disease-level study discovery. 3. If no matching study is found, check known external instances above before declaring the study absent. 4. If proposing a substitute, describe why it is a substitute and how its scope differs. ## Substitute-Cohort Answer Pattern > I cannot query PBTA from this cBioPortal deployment. PBTA is typically accessed through pediatric cBioPortal at https://pedcbioportal.kidsfirstdrc.org/. I can analyze `[substitute_study_id]` here, but its results should not be interpreted as PBTA results. When reporting numbers from a substitute: > Scope note: these counts are from `[substitute_study_id]` in this deployment, not from the requested PBTA cohort. ## Do Not - Do not answer a PBTA question with `brain_cptac_2020` numbers without a scope wa … (169 more chars) ═ answer (success) I can't analyze the Pediatric Brain Tumor Atlas (PBTA, Provisional) cohort here. It isn't in the cBioPortal deployment I'm connected to. I searched for "pediatric brain tumor atlas" and "PBTA" and found no matching study. PBTA is usually hosted on the Kids First pediatric cBioPortal: https://pedcbioportal.kidsfirstdrc.org/. To look at thyroid hormone gene expression by diagnosis there, you could: - Open the PBTA study and use the **Plots** tab. Set one axis to mRNA expression for your gene of interest and the other axis to a clinical attribute such as cancer type or diagnosis. - Query the genes together in the **Results View** and compare expression across diagnosis groups with **Group Comparison**. I haven't substituted another cohort, because that wouldn't give PBTA results. Other pediatric or brain tumor studies are available here, for example CPTAC brain or TARGET. I can run the analysis on one of those if you'd like, and I'd label the results clearly as not being PBTA. To go ahead, I need two things from you: 1. **Genes:** Which genes count as "thyroid hormone" genes? Examples include the receptors (THRA, THRB), the deiodinases (DIO1, DIO2, DIO3) and the transporters (SLC16A2/MCT8, SLCO1C1). I can use this set or your own list. 2. **Cohort:** Which substitute study should I use, if any?