Q (sonnet): What is the current and future support for storing and analyzing germline variants in cBioPortal, compared to other alternatives? ▶ read_guide { "uri": "cbioportal://germline-guide" } ▶ read_guide { "uri": "cbioportal://faq-guide" } ◀ result Germline Variant Guide ====================== Overview -------- cBioPortal stores both somatic AND germline variant data. Many cBioPortal features work identically for both variant types. This guide explains how to query germline variants and what to be aware of. How Germline Data is Stored --------------------------- ### mutation_status is free text — always compare case-insensitively Each study loads its own spelling. Germline calls appear as `'Germline'`, `'GERMLINE'` and `'germline'`; somatic calls as `'Somatic'`, `'SOMATIC'` and `'somatic'`; unannotated mutations as `'NA'`, `'.'`, `'Unknown'`, `'UNKNOWN'`, `'__UNKNOWN__'` and others. Matching one spelling silently drops whole studies (`mutation_status = 'Germline'` misses `all_stjude_2013`, `aml_stjude_2024` and `pog570_bcgsc_2020`). - **Germline:** `upper(mutation_status) = 'GERMLINE'` - **Somatic:** `upper(mutation_status) = 'SOMATIC'` — but only when the user asks for somatic-only. Many studies label their somatic calls `'NA'` or `'UNKNOWN'`, so for ordinary mutation questions follow common-pitfalls #3 and exclude only `'UNCALLED'`. - When unsure, list the values first: `SELECT mutation_status, count() FROM genomic_event_derived WHERE cancer_study_identifier = '{study_id}' AND variant_type = 'mutation' GROUP BY mutation_status` ### Where the column lives - `genomic_event_derived.mutation_status` (preferred): mutations, and structural variants (from `sv_status`: `'SOMATIC'`, `'Somatic'`, `'GERMLINE'`) - `mutation_derived.mutationStatus`: the same values for mutations Identifying Studies with Germline Data -------------------------------------- Not all studies include germline data. Always check before querying: ```sql -- Find studies containing germline mutations SELECT cancer_study_identifier, COUNT(*) as germline_count FROM genomic_event_derived WHERE variant_type = 'mutation' AND upper(mutation_status) = 'GERMLINE' GROUP BY cancer_study_identifier ORDER BY germline_count DESC ``` Common Query Patterns --------------------- ### Count germline vs somatic mutations per gene in a study ```sql SELECT hugo_gene_symbol, upper(mutation_status) AS status, COUNT(*) as count FROM genomic_event_derived WHERE cancer_study_identifier = '{study_id}' AND variant_type = 'mutation' AND upper(mutation_status) IN ('GERMLINE', 'SOMATIC') GROUP BY hugo_gene_symbol, status ORDER BY count DESC LIMIT 20 ``` ### Find patients with germline mutations in a specific gene ```sql SELECT DISTINCT patient_unique_id, sample_unique_id, mutation_variant, mutation_type FROM genomic_event_derived WHERE hugo_gene_symbol = '{GENE}' AND upper(mutation_status) = 'GERMLINE' AND cancer_study_identifier = '{study_id}' AND variant_type = 'mutation' ``` ### Germline mutation frequency The numerator is samples with a germline mutation in the gene; the denominator is samples **profiled** for the gene — not samples that happen to have a mutation in it. Take the denominator from mutation-frequency-guide Step 2 (or … (2222 more chars) ◀ result # cBioPortal FAQ Guide Curated answers to frequently asked general questions about cBioPortal. Source: [cBioPortal FAQ](https://docs.cbioportal.org/user-guide/faq/). ## What is cBioPortal? cBioPortal for Cancer Genomics is an open-access, open-source resource for interactive exploration of multidimensional cancer genomics data sets. It was originally developed at Memorial Sloan Kettering Cancer Center (MSK) and is now maintained by a multi-institutional team. ## History - **2008**: cBioPortal first became available online. - **2012**: First major publication — Cerami et al., *Cancer Discovery*. - **2013**: Second major publication — Gao et al., *Science Signaling*. - **2023**: Third major publication — de Bruijn et al., *Cancer Research*. ## How to Cite cBioPortal When using cBioPortal in publications, cite these three papers: 1. Cerami et al. "The cBio Cancer Genomics Portal: An Open Platform for Exploring Multidimensional Cancer Genomics Data." *Cancer Discovery* 2, 401–404 (2012). doi:10.1158/2159-8290.CD-12-0095 2. Gao et al. "Integrative Analysis of Complex Cancer Genomics and Clinical Profiles Using the cBioPortal." *Science Signaling* 6, pl1 (2013). doi:10.1126/scisignal.2004088 3. de Bruijn et al. "Analysis and Visualization of Longitudinal Genomic and Clinical Data from the AACR Project GENIE Biopharma Collaborative in cBioPortal." *Cancer Research* 83, 3861–3867 (2023). doi:10.1158/0008-5472.CAN-23-0816 Also cite the specific study publication(s) whose data you used. ## What Data Types Does cBioPortal Contain? - **Mutations** (non-synonymous somatic mutations) - **DNA copy-number alterations** (CNA) from GISTIC or other methods - **mRNA expression** data (RNA-seq, microarray) - **Protein/phosphoprotein** levels (RPPA, mass spectrometry) - **DNA methylation** data - **Structural variants** (gene fusions, rearrangements) - **Clinical data** (demographics, diagnosis, treatment, outcomes) Note: Synonymous mutations are not included in cBioPortal. ## What Cancer Types Are in the Database? Answer from the studies actually loaded, not from the OncoTree ontology (OncoTree lists hundreds of codes, most with no study here). One query: ```sql SELECT tc.type_of_cancer_id, tc.name, tc.tissue, count() AS studies FROM cancer_study cs JOIN type_of_cancer tc ON cs.type_of_cancer_id = tc.type_of_cancer_id GROUP BY tc.type_of_cancer_id, tc.name, tc.tissue ORDER BY studies DESC; ``` Report how many distinct cancer types there are across how many studies, grouped by `tissue`. `mixed` ("Mixed Cancer Types") marks multi-cancer studies such as MSK-IMPACT and MSK-CHORD; their per-sample cancer types are in `clinical_data_derived` (`CANCER_TYPE`). ## Data Types Usually Not Stored Directly cBioPortal generally does not store: - raw CT, MRI, pathology-slide, or histology image files - raw sequencing files such as BAM, CRAM, or FASTQ - polygenic risk scores as a standard data type - full external clinical-trial databases However, this deployment … (7371 more chars) ▶ clickhouse_run_select_query { "query": "SELECT cancer_study_identifier, COUNT(*) as germline_count FROM genomic_event_derived WHERE variant_type = 'mutation' AND upper(mutation_status) = 'GERMLINE' GROUP BY cancer_study_identifier ORDER BY germline_count DESC LIMIT 15", "max_rows": 15 } ◀ result {"rows":[{"cancer_study_identifier":"coadread_tcga_pub","germline_count":101},{"cancer_study_identifier":"lung_nci_2022","germline_count":100},{"cancer_study_identifier":"pog570_bcgsc_2020","germline_count":74},{"cancer_study_identifier":"coadread_tcga","germline_count":70},{"cancer_study_identifier":"hnsc_a5consortium_2025","germline_count":67},{"cancer_study_identifier":"brca_tcga_pub","germline_count":64},{"cancer_study_identifier":"kirc_tcga_pub","germline_count":57},{"cancer_study_identifier":"all_stjude_2013","germline_count":52},{"cancer_study_identifier":"ov_tcga_pub","germline_count":51},{"cancer_study_identifier":"prad_su2c_2019","germline_count":40},{"cancer_study_identifier":"ov_tcga","germline_count":20},{"cancer_study_identifier":"aml_stjude_2024","germline_count":19},{"cancer_study_identifier":"mrt_bcgsc_2016","germline_count":10},{"cancer_study_identifier":"prad_tcga_pub","germline_count":7},{"cancer_study_identifier":"nhl_bcgsc_2011","germline_count":3}]} ═ answer (success) ## Current support for germline variants in cBioPortal **cBioPortal is not somatic-only** — it stores germline variant calls in the same underlying mutation tables as somatic calls, distinguished by a `mutation_status` field (`GERMLINE` vs. `SOMATIC`, spelling varies per study). Confirming this against the live database, at least 15 studies currently carry germline-flagged variant records, including well-known cohorts such as `coadread_tcga_pub` (101 germline calls), `brca_tcga_pub` (64), `ov_tcga_pub` (51), `pog570_bcgsc_2020` (74), and pediatric cohorts like `all_stjude_2013` and `aml_stjude_2024`. Because germline and somatic mutations share the same data model, most of the standard cBioPortal machinery already works on germline calls without modification: - Study View, OncoPrint, mutation lollipop/diagram plots, and clinical correlation/plots - Group comparison and gene-panel coverage checks - Standard REST API / DataHub import formats (a study can be built purely around germline variants plus clinical annotations) **Current limitations:** - Germline coverage is inconsistent — most studies were curated with somatic variants as the primary focus, and many either exclude germline calls during processing or leave `mutation_status` unannotated (`NA`/`UNKNOWN`), which is not the same as confirming a variant is somatic. - cBioPortal does not store a dedicated ACMG-style pathogenicity classification (pathogenic/likely pathogenic/VUS/benign) for germline variants as a standard field. - Features built specifically around presumed-somatic biology — OncoKB driver annotation, mutual exclusivity analysis, and TMB calculation — are not designed for or validated on germline variant sets. - Some public-facing deployments may deliberately exclude germline data for privacy reasons, since germline findings are more sensitive (incidental/hereditary risk implications) than somatic tumor profiling. ## Future support I don't have access to cBioPortal's product roadmap or unreleased development plans, and I'm not able to speak to unannounced future features — that would be outside what the data and documentation available to me can confirm. For authoritative roadmap information, the cBioPortal GitHub repository (issues/discussions), the docs.cbioportal.org site, or the project's user forum would be the right sources. ## How this compares to alternatives (general knowledge, not from cBioPortal data) The following is general background, not sourced from cBioPortal's database or documentation, so treat it as context rather than a cBioPortal-verified claim: - **ClinVar** is the standard public repository purpose-built for germline (and some somatic) variant pathogenicity classification with ACMG-style curation — it is the go-to source for clinical-grade variant interpretation, which cBioPortal does not attempt to replicate. - **GDC (Genomic Data Commons)** hosts raw and controlled-access germline VCFs (e.g., from TCGA), but is a data repository, not an interactive visual analysis tool. - Dedicated hereditary-cancer / germline-focused platforms and clinical LIMS/variant-interpretation software (e.g., institutional pipelines built around ACMG guidelines) are typically stronger on classification and clinical actionability workflows than cBioPortal, which prioritizes cohort-level exploratory visualization across somatic and germline data together. In short: cBioPortal's strength is letting you explore germline variants side-by-side with somatic mutations, copy number, expression, and clinical/outcome data in one interactive cohort view — it is not designed to be a clinical-grade germline classification database like ClinVar, nor a raw-data repository like GDC. If you'd like, I can pull up the germline variant data for a specific gene or study (e.g., BRCA1/BRCA2 in one of the cohorts above) directly in cBioPortal's Study View or OncoPrint — just let me know which.