Q (sonnet): What is the current and future support for storing and analyzing germline variants in cBioPortal, compared to other alternatives? ▶ list_guides {} ◀ result {"result":[{"uri":"cbioportal://mutation-frequency-guide","description":"Comprehensive guide for calculating gene mutation frequencies with gene-specific profiling denominators"},{"uri":"cbioportal://clinical-data-guide","description":"Guide for querying clinical data including patient vs sample level considerations"},{"uri":"cbioportal://sample-filtering-guide","description":"Guide for filtering samples and studies in cBioPortal queries"},{"uri":"cbioportal://common-pitfalls","description":"Guide to avoid common mistakes when querying cBioPortal data. If you already know which numbered pitfall applies, fetch just that section via read_guide(\"cbioportal://common-pitfalls#\") (e.g. #16) instead of the full guide"},{"uri":"cbioportal://treatment-guide","description":"Guide for querying treatment/clinical event data including drug agents, timelines, and linking to genomic data"},{"uri":"cbioportal://faq-guide","description":"General cBioPortal FAQ: history, how to cite, data types, reference genome, abbreviations, GISTIC thresholds, API access"},{"uri":"cbioportal://statistical-tests-guide","description":"Statistical test selection guide — decision matrix for choosing Fisher's exact, Wilcoxon, chi-squared, t-test, ANOVA, etc. based on data type and group count"},{"uri":"cbioportal://gene-expression-guide","description":"Gene expression / copy-number / methylation analysis. Covers genetic_alteration_derived, profile_type discovery, and the gene_pair_coexpression view for Spearman correlation between two genes"},{"uri":"cbioportal://external-resources-guide","description":"Guide for finding external linked resources such as imaging, pathology, Minerva, HTAN, or other resource_* table links before declaring data unavailable"},{"uri":"cbioportal://gene-resolution-guide","description":"Guide for resolving ambiguous gene symbols, aliases, gene families, and shorthand such as CD3 before querying expression or alteration data"},{"uri":"cbioportal://study-resolution-guide","description":"Guide for resolving requested studies, avoiding silent substitute cohorts, and redirecting to known external cBioPortal instances when data is not in this deployment"},{"uri":"cbioportal://germline-guide","description":"Guide for querying germline variant data — storage columns, study discovery, query patterns, and somatic vs germline considerations"},{"uri":"cbioportal://study-guide/{study_id}","description":"Dynamic study-specific guide - use get_study_guide(study_id) tool to generate"}]} ▶ read_guide { "uri": "cbioportal://germline-guide" } ▶ read_guide { "uri": "cbioportal://faq-guide" } ◀ result Germline Variant Guide ====================== Overview -------- cBioPortal stores both somatic AND germline variant data. Many cBioPortal features work identically for both variant types. This guide explains how to query germline variants and what to be aware of. How Germline Data is Stored --------------------------- ### mutation_status is free text — always compare case-insensitively Each study loads its own spelling. Germline calls appear as `'Germline'`, `'GERMLINE'` and `'germline'`; somatic calls as `'Somatic'`, `'SOMATIC'` and `'somatic'`; unannotated mutations as `'NA'`, `'.'`, `'Unknown'`, `'UNKNOWN'`, `'__UNKNOWN__'` and others. Matching one spelling silently drops whole studies (`mutation_status = 'Germline'` misses `all_stjude_2013`, `aml_stjude_2024` and `pog570_bcgsc_2020`). - **Germline:** `upper(mutation_status) = 'GERMLINE'` - **Somatic:** `upper(mutation_status) = 'SOMATIC'` — but only when the user asks for somatic-only. Many studies label their somatic calls `'NA'` or `'UNKNOWN'`, so for ordinary mutation questions follow common-pitfalls #3 and exclude only `'UNCALLED'`. - When unsure, list the values first: `SELECT mutation_status, count() FROM genomic_event_derived WHERE cancer_study_identifier = '{study_id}' AND variant_type = 'mutation' GROUP BY mutation_status` ### Where the column lives - `genomic_event_derived.mutation_status` (preferred): mutations, and structural variants (from `sv_status`: `'SOMATIC'`, `'Somatic'`, `'GERMLINE'`) - `mutation_derived.mutationStatus`: the same values for mutations Identifying Studies with Germline Data -------------------------------------- Not all studies include germline data. Always check before querying: ```sql -- Find studies containing germline mutations SELECT cancer_study_identifier, COUNT(*) as germline_count FROM genomic_event_derived WHERE variant_type = 'mutation' AND upper(mutation_status) = 'GERMLINE' GROUP BY cancer_study_identifier ORDER BY germline_count DESC ``` Common Query Patterns --------------------- ### Count germline vs somatic mutations per gene in a study ```sql SELECT hugo_gene_symbol, upper(mutation_status) AS status, COUNT(*) as count FROM genomic_event_derived WHERE cancer_study_identifier = '{study_id}' AND variant_type = 'mutation' AND upper(mutation_status) IN ('GERMLINE', 'SOMATIC') GROUP BY hugo_gene_symbol, status ORDER BY count DESC LIMIT 20 ``` ### Find patients with germline mutations in a specific gene ```sql SELECT DISTINCT patient_unique_id, sample_unique_id, mutation_variant, mutation_type FROM genomic_event_derived WHERE hugo_gene_symbol = '{GENE}' AND upper(mutation_status) = 'GERMLINE' AND cancer_study_identifier = '{study_id}' AND variant_type = 'mutation' ``` ### Germline mutation frequency The numerator is samples with a germline mutation in the gene; the denominator is samples **profiled** for the gene — not samples that happen to have a mutation in it. Take the denominator from mutation-frequency-guide Step 2 (or … (2222 more chars) ◀ result # cBioPortal FAQ Guide Curated answers to frequently asked general questions about cBioPortal. Source: [cBioPortal FAQ](https://docs.cbioportal.org/user-guide/faq/). ## What is cBioPortal? cBioPortal for Cancer Genomics is an open-access, open-source resource for interactive exploration of multidimensional cancer genomics data sets. It was originally developed at Memorial Sloan Kettering Cancer Center (MSK) and is now maintained by a multi-institutional team. ## History - **2008**: cBioPortal first became available online. - **2012**: First major publication — Cerami et al., *Cancer Discovery*. - **2013**: Second major publication — Gao et al., *Science Signaling*. - **2023**: Third major publication — de Bruijn et al., *Cancer Research*. ## How to Cite cBioPortal When using cBioPortal in publications, cite these three papers: 1. Cerami et al. "The cBio Cancer Genomics Portal: An Open Platform for Exploring Multidimensional Cancer Genomics Data." *Cancer Discovery* 2, 401–404 (2012). doi:10.1158/2159-8290.CD-12-0095 2. Gao et al. "Integrative Analysis of Complex Cancer Genomics and Clinical Profiles Using the cBioPortal." *Science Signaling* 6, pl1 (2013). doi:10.1126/scisignal.2004088 3. de Bruijn et al. "Analysis and Visualization of Longitudinal Genomic and Clinical Data from the AACR Project GENIE Biopharma Collaborative in cBioPortal." *Cancer Research* 83, 3861–3867 (2023). doi:10.1158/0008-5472.CAN-23-0816 Also cite the specific study publication(s) whose data you used. ## What Data Types Does cBioPortal Contain? - **Mutations** (non-synonymous somatic mutations) - **DNA copy-number alterations** (CNA) from GISTIC or other methods - **mRNA expression** data (RNA-seq, microarray) - **Protein/phosphoprotein** levels (RPPA, mass spectrometry) - **DNA methylation** data - **Structural variants** (gene fusions, rearrangements) - **Clinical data** (demographics, diagnosis, treatment, outcomes) Note: Synonymous mutations are not included in cBioPortal. ## Data Types Usually Not Stored Directly cBioPortal generally does not store: - raw CT, MRI, pathology-slide, or histology image files - raw sequencing files such as BAM, CRAM, or FASTQ - polygenic risk scores as a standard data type - full external clinical-trial databases However, this deployment may contain links to external viewers or portals through `resource_*` tables. For imaging, pathology, Minerva, HTAN, or viewer questions, read `cbioportal://external-resources-guide` and check those tables before saying the data is absent. ## What the MCP Agent Can and Cannot Produce The MCP agent can: - query cBioPortal data and return tables in text/JSON/Markdown - summarize counts, frequencies, and available attributes - provide cBioPortal study links and DataHub download links when applicable - provide SQL snippets or handoff instructions for R/Python/cBioPortal tools The MCP agent should not promise to: - render Kaplan-Meier plots or other visual figures directly - export large … (6558 more chars) ▶ clickhouse_run_select_query { "query": "SELECT cancer_study_identifier, COUNT(*) as germline_count, COUNT(DISTINCT hugo_gene_symbol) as genes, COUNT(DISTINCT sample_unique_id) as samples FROM genomic_event_derived WHERE variant_type = 'mutation' AND upper(mutation_status) = 'GERMLINE' GROUP BY cancer_study_identifier ORDER BY germline_count DESC" } ◀ result {"rows":[{"cancer_study_identifier":"coadread_tcga_pub","germline_count":101,"genes":86,"samples":59},{"cancer_study_identifier":"lung_nci_2022","germline_count":100,"genes":74,"samples":72},{"cancer_study_identifier":"pog570_bcgsc_2020","germline_count":74,"genes":24,"samples":68},{"cancer_study_identifier":"coadread_tcga","germline_count":70,"genes":59,"samples":48},{"cancer_study_identifier":"hnsc_a5consortium_2025","germline_count":67,"genes":1,"samples":66},{"cancer_study_identifier":"brca_tcga_pub","germline_count":64,"genes":10,"samples":58},{"cancer_study_identifier":"kirc_tcga_pub","germline_count":57,"genes":35,"samples":52},{"cancer_study_identifier":"all_stjude_2013","germline_count":52,"genes":44,"samples":33},{"cancer_study_identifier":"ov_tcga_pub","germline_count":51,"genes":2,"samples":50},{"cancer_study_identifier":"prad_su2c_2019","germline_count":40,"genes":8,"samples":39},{"cancer_study_identifier":"ov_tcga","germline_count":20,"genes":17,"samples":17},{"cancer_study_identifier":"aml_stjude_2024","germline_count":19,"genes":11,"samples":19},{"cancer_study_identifier":"mrt_bcgsc_2016","germline_count":10,"genes":7,"samples":8},{"cancer_study_identifier":"prad_tcga_pub","germline_count":7,"genes":2,"samples":7},{"cancer_study_identifier":"nhl_bcgsc_2011","germline_count":3,"genes":3,"samples":3},{"cancer_study_identifier":"stmyec_wcm_2022","germline_count":2,"genes":1,"samples":2}]} ═ answer (success) ## Current germline variant support in cBioPortal **Storage model:** cBioPortal has no separate "germline" data type — germline calls are stored in the same mutation tables as somatic calls, distinguished only by a free-text `mutation_status` field (values like `'Germline'`, `'GERMLINE'`, `'germline'` depending on the study). There is no dedicated germline schema, no ACMG pathogenicity classification field, and no separate germline-specific UI panel. **Actual coverage in this deployment:** Only a small number of studies have any germline variants loaded — I found germline-flagged mutations in just **16 studies**, e.g.: | Study | Germline variants | Genes | Samples | |---|---|---|---| | coadread_tcga_pub | 101 | 86 | 59 | | lung_nci_2022 | 100 | 74 | 72 | | pog570_bcgsc_2020 | 74 | 24 | 68 | | coadread_tcga | 70 | 59 | 48 | | brca_tcga_pub | 64 | 10 | 58 | | ov_tcga_pub | 51 | 2 | 50 | | prad_su2c_2019 | 40 | 8 | 39 | These are tiny counts relative to total mutation records in each study — germline is clearly a secondary, incidental annotation rather than a systematically curated dataset (e.g., no dedicated hereditary-cancer-panel study with comprehensive BRCA1/2, Lynch syndrome gene, etc. germline calling appears at scale). **Functional support that does work for germline data** (per the germline-guide): mutation frequency queries, protein domain/Mutation Mapper visualization, co-occurrence analysis, clinical correlation, gene panel coverage checks, and OncoPrint — all work identically for germline records once filtered correctly. **What does NOT work / is not designed for germline:** - No OncoKB or driver-annotation support for germline pathogenicity (OncoKB annotations target somatic variants) - No ACMG variant classification (pathogenic/likely pathogenic/VUS) storage - Mutual exclusivity analysis and TMB calculations assume somatic variants - `mutation_status` is inconsistently populated/spelled across studies — many studies simply don't annotate status at all (`'NA'`, `'UNKNOWN'`), so germline variants may be silently present but unlabeled and unfindable by filtering ## Future support I don't have visibility into the cBioPortal project's roadmap beyond what's in these guides, and I don't want to speculate. The FAQ guide confirms germline *studies* (dedicated cohorts built around germline variants + clinical data) are technically supported via the standard DataHub import format today, but there's no indication in the documentation available to me of planned schema changes (e.g., a dedicated ACMG classification field or germline-specific analysis modules). For authoritative roadmap information, I'd suggest checking the [cBioPortal GitHub repository](https://github.com/cBioPortal) issues/discussions or the docs.cbioportal.org release notes directly, rather than relying on my knowledge cutoff. ## Comparison to alternatives This is a general knowledge question outside what I can verify against the database, but based on well-established, publicly documented platform designs (not specific to this deployment): - **ClinVar** — the standard public repository specifically for germline/somatic variant clinical significance, with structured ACMG classifications (pathogenic, likely pathogenic, VUS, likely benign, benign) as a first-class field. Purpose-built for variant interpretation, not cohort-level genomic visualization. - **GDC (Genomic Data Commons)** — hosts both somatic and germline (restricted-access) data with more rigorous provenance/pipeline tracking, but is a data repository, not an interactive exploration tool (this distinction is also noted in the FAQ guide above). - **Franklin, VarSome, InterVar** — dedicated germline variant interpretation tools with automated ACMG classification — a capability cBioPortal explicitly does not have. - **GeneDx/Invitae-style clinical LIMS** — built specifically for hereditary cancer panel reporting with formal classification workflows. **Bottom line:** cBioPortal treats germline variants as a loosely-supported subset of its somatic-first mutation model — usable for basic frequency/visualization tasks in the handful of studies that include them, but lacking the structured pathogenicity classification and systematic curation that dedicated germline variant databases (ClinVar) or clinical interpretation tools provide. Would you like me to pull up one of these germline-containing studies (e.g., `coadread_tcga_pub` or `pog570_bcgsc_2020`) in the cBioPortal web interface to explore the germline calls directly?