Q (sonnet): What is the current and future support for storing and analyzing germline variants in cBioPortal, compared to other alternatives? ▶ read_guide { "uri": "cbioportal://germline-guide" } ▶ read_guide { "uri": "cbioportal://faq-guide" } ◀ result Germline Variant Guide ====================== Overview -------- cBioPortal stores both somatic AND germline variant data. Many cBioPortal features work identically for both variant types. This guide explains how to query germline variants and what to be aware of. How Germline Data is Stored --------------------------- ### mutation_status is free text — always compare case-insensitively Each study loads its own spelling. Germline calls appear as `'Germline'`, `'GERMLINE'` and `'germline'`; somatic calls as `'Somatic'`, `'SOMATIC'` and `'somatic'`; unannotated mutations as `'NA'`, `'.'`, `'Unknown'`, `'UNKNOWN'`, `'__UNKNOWN__'` and others. Matching one spelling silently drops whole studies (`mutation_status = 'Germline'` misses `all_stjude_2013`, `aml_stjude_2024` and `pog570_bcgsc_2020`). - **Germline:** `upper(mutation_status) = 'GERMLINE'` - **Somatic:** `upper(mutation_status) = 'SOMATIC'` — but only when the user asks for somatic-only. Many studies label their somatic calls `'NA'` or `'UNKNOWN'`, so for ordinary mutation questions follow common-pitfalls #3 and exclude only `'UNCALLED'`. - When unsure, list the values first: `SELECT mutation_status, count() FROM genomic_event_derived WHERE cancer_study_identifier = '{study_id}' AND variant_type = 'mutation' GROUP BY mutation_status` ### Where the column lives - `genomic_event_derived.mutation_status` (preferred): mutations, and structural variants (from `sv_status`: `'SOMATIC'`, `'Somatic'`, `'GERMLINE'`) - `mutation_derived.mutationStatus`: the same values for mutations Identifying Studies with Germline Data -------------------------------------- Not all studies include germline data. Always check before querying: ```sql -- Find studies containing germline mutations SELECT cancer_study_identifier, COUNT(*) as germline_count FROM genomic_event_derived WHERE variant_type = 'mutation' AND upper(mutation_status) = 'GERMLINE' GROUP BY cancer_study_identifier ORDER BY germline_count DESC ``` Common Query Patterns --------------------- ### Count germline vs somatic mutations per gene in a study ```sql SELECT hugo_gene_symbol, upper(mutation_status) AS status, COUNT(*) as count FROM genomic_event_derived WHERE cancer_study_identifier = '{study_id}' AND variant_type = 'mutation' AND upper(mutation_status) IN ('GERMLINE', 'SOMATIC') GROUP BY hugo_gene_symbol, status ORDER BY count DESC LIMIT 20 ``` ### Find patients with germline mutations in a specific gene ```sql SELECT DISTINCT patient_unique_id, sample_unique_id, mutation_variant, mutation_type FROM genomic_event_derived WHERE hugo_gene_symbol = '{GENE}' AND upper(mutation_status) = 'GERMLINE' AND cancer_study_identifier = '{study_id}' AND variant_type = 'mutation' ``` ### Germline mutation frequency The numerator is samples with a germline mutation in the gene; the denominator is samples **profiled** for the gene — not samples that happen to have a mutation in it. Take the denominator from mutation-frequency-guide Step 2 (or … (2222 more chars) ◀ result # cBioPortal FAQ Guide Curated answers to frequently asked general questions about cBioPortal. Source: [cBioPortal FAQ](https://docs.cbioportal.org/user-guide/faq/). ## What is cBioPortal? cBioPortal for Cancer Genomics is an open-access, open-source resource for interactive exploration of multidimensional cancer genomics data sets. It was originally developed at Memorial Sloan Kettering Cancer Center (MSK) and is now maintained by a multi-institutional team. ## History - **2008**: cBioPortal first became available online. - **2012**: First major publication — Cerami et al., *Cancer Discovery*. - **2013**: Second major publication — Gao et al., *Science Signaling*. - **2023**: Third major publication — de Bruijn et al., *Cancer Research*. ## How to Cite cBioPortal When using cBioPortal in publications, cite these three papers: 1. Cerami et al. "The cBio Cancer Genomics Portal: An Open Platform for Exploring Multidimensional Cancer Genomics Data." *Cancer Discovery* 2, 401–404 (2012). doi:10.1158/2159-8290.CD-12-0095 2. Gao et al. "Integrative Analysis of Complex Cancer Genomics and Clinical Profiles Using the cBioPortal." *Science Signaling* 6, pl1 (2013). doi:10.1126/scisignal.2004088 3. de Bruijn et al. "Analysis and Visualization of Longitudinal Genomic and Clinical Data from the AACR Project GENIE Biopharma Collaborative in cBioPortal." *Cancer Research* 83, 3861–3867 (2023). doi:10.1158/0008-5472.CAN-23-0816 Also cite the specific study publication(s) whose data you used. ## What Data Types Does cBioPortal Contain? - **Mutations** (non-synonymous somatic mutations) - **DNA copy-number alterations** (CNA) from GISTIC or other methods - **mRNA expression** data (RNA-seq, microarray) - **Protein/phosphoprotein** levels (RPPA, mass spectrometry) - **DNA methylation** data - **Structural variants** (gene fusions, rearrangements) - **Clinical data** (demographics, diagnosis, treatment, outcomes) Note: Synonymous mutations are not included in cBioPortal. ## Data Types Usually Not Stored Directly cBioPortal generally does not store: - raw CT, MRI, pathology-slide, or histology image files - raw sequencing files such as BAM, CRAM, or FASTQ - polygenic risk scores as a standard data type - full external clinical-trial databases However, this deployment may contain links to external viewers or portals through `resource_*` tables. For imaging, pathology, Minerva, HTAN, or viewer questions, read `cbioportal://external-resources-guide` and check those tables before saying the data is absent. ## What the MCP Agent Can and Cannot Produce The MCP agent can: - query cBioPortal data and return tables in text/JSON/Markdown - summarize counts, frequencies, and available attributes - provide cBioPortal study links and DataHub download links when applicable - provide SQL snippets or handoff instructions for R/Python/cBioPortal tools The MCP agent should not promise to: - render Kaplan-Meier plots or other visual figures directly - export large … (6558 more chars) ═ answer (success) This question is about cBioPortal's platform architecture and roadmap rather than a specific data query, so let me give you a direct answer based on what's documented, and flag where I'm stepping outside cBioPortal data into general knowledge. ## Current germline support in cBioPortal **Data model:** cBioPortal does not have a separate "germline track" — germline and somatic variants share the same mutation data schema. A variant's provenance is captured only via a free-text `mutation_status` field (values like `Germline`, `GERMLINE`, `Somatic`, `SOMATIC`, or unannotated `NA`/`Unknown`, spelled inconsistently study by study). This means: - **What works well for germline data:** loading germline calls alongside clinical annotation, viewing them in OncoPrint/Mutations tabs, protein-domain (lollipop) views, basic clinical correlation, and gene-panel coverage checks — because these features just consume the generic mutation table. - **What is not germline-aware:** driver annotation (`driver_filter`), OncoKB oncogenicity calls, mutual exclusivity analysis, and TMB calculations are all built and curated for **somatic** variants. Running them on a germline-heavy study will not produce meaningful driver/oncogenicity output. - **No ACMG classification field.** Pathogenic / likely pathogenic / VUS status is not a standard cBioPortal column — if a study wants to show it, it has to be loaded as a custom clinical attribute, not as structured variant metadata. - **Study-level, not systematic.** A given cohort either includes germline calls (e.g., some hereditary-cancer or pan-cancer cohorts have `mutation_status = 'Germline'` rows) or excludes them entirely during curation — there's no portal-wide standard for which cancer studies carry germline data, and public instances may intentionally strip germline variants for privacy reasons before public release. - **Reference/annotation pipeline (Genome Nexus + VEP)** is variant-agnostic, so annotation itself isn't a blocker — the gap is downstream interpretation tooling (driver calls, actionability, population-frequency filtering) that a germline-focused platform would need. ## Future support I don't have visibility into an official cBioPortal product roadmap through this database or the documentation available to me, so I can't confirm specific planned features (e.g., ACMG tiering, ClinVar integration, population-frequency filters, or a dedicated germline module). If you need the authoritative roadmap, I'd point you to the [cBioPortal GitHub repository](https://github.com/cBioPortal) issues/discussions or the [cBioPortal user forum](https://groups.google.com/forum/#!forum/cbioportal), where the development team posts and discusses upcoming features — I don't want to speculate here. ## Comparison to alternatives (general knowledge, not from cBioPortal data) This part draws on general knowledge of the field, not cBioPortal's own data: - **ClinVar** is the standard public repository specifically for germline variant pathogenicity classification (ACMG-tiered), but it has no cohort-level cancer genomics visualization, no OncoPrint-style cross-sample views, and no linkage to expression/CNA/clinical outcome data. - **GeneDx, Invitae/Labcorp genetics, and clinical LIMS/reporting systems** handle ACMG classification and clinical reporting workflows for germline testing but are not open cohort-exploration research platforms. - **PROMPT/PHS-style hereditary cancer registries and institutional biobank portals** (e.g., MSK's own IMPACT-associated germline pipelines) often build purpose-specific germline interpretation layers (ACMG tiers, penetrance, family history linkage) on top of or alongside cBioPortal-style somatic tools, because cBioPortal itself doesn't natively provide that interpretation layer. - **GDC/TCGA germline resource** (Huang et al. 2018 PanCanAtlas germline paper) computed pathogenicity but distributed it as supplementary tables/portals rather than through interactive cBioPortal-native tooling. **Bottom line:** cBioPortal can *store and display* germline variants reasonably well wherever a study chooses to load them (same schema, same visualization tabs), but it does not provide germline-specific interpretation (ACMG tiering, population-frequency filtering, hereditary-risk-specific driver calls) — that layer currently has to come from an external clinical genetics tool or a custom-curated clinical attribute layered on top. If you're working on a specific germline cohort in cBioPortal, I can check whether that study currently has germline calls loaded and how they're annotated.