Q (sonnet-5.5): What is the current and future support for storing and analyzing germline variants in cBioPortal, compared to other alternatives? ▶ read_guide { "uri": "cbioportal://germline-guide" } ▶ read_guide { "uri": "cbioportal://faq-guide" } ◀ result Germline Variant Guide ====================== Overview -------- cBioPortal stores both somatic AND germline variant data. Many cBioPortal features work identically for both variant types. This guide explains how to query germline variants and what to be aware of. How Germline Data is Stored --------------------------- ### mutation_status is free text — always compare case-insensitively Each study loads its own spelling. Germline calls appear as `'Germline'`, `'GERMLINE'` and `'germline'`; somatic calls as `'Somatic'`, `'SOMATIC'` and `'somatic'`; unannotated mutations as `'NA'`, `'.'`, `'Unknown'`, `'UNKNOWN'`, `'__UNKNOWN__'` and others. Matching one spelling silently drops whole studies (`mutation_status = 'Germline'` misses `all_stjude_2013`, `aml_stjude_2024` and `pog570_bcgsc_2020`). - **Germline:** `upper(mutation_status) = 'GERMLINE'` - **Somatic:** `upper(mutation_status) = 'SOMATIC'` — but only when the user asks for somatic-only. Many studies label their somatic calls `'NA'` or `'UNKNOWN'`, so for ordinary mutation questions follow common-pitfalls #3 and exclude only `'UNCALLED'`. - When unsure, list the values first: `SELECT mutation_status, count() FROM genomic_event_derived WHERE cancer_study_identifier = '{study_id}' AND variant_type = 'mutation' GROUP BY mutation_status` ### Where the column lives - `genomic_event_derived.mutation_status` (preferred): mutations, and structural variants (from `sv_status`: `'SOMATIC'`, `'Somatic'`, `'GERMLINE'`) - `mutation_derived.mutationStatus`: the same values for mutations Identifying Studies with Germline Data -------------------------------------- Not all studies include germline data. Always check before querying: ```sql -- Find studies containing germline mutations SELECT cancer_study_identifier, COUNT(*) as germline_count FROM genomic_event_derived WHERE variant_type = 'mutation' AND upper(mutation_status) = 'GERMLINE' GROUP BY cancer_study_identifier ORDER BY germline_count DESC ``` Common Query Patterns --------------------- ### Count germline vs somatic mutations per gene in a study ```sql SELECT hugo_gene_symbol, upper(mutation_status) AS status, COUNT(*) as count FROM genomic_event_derived WHERE cancer_study_identifier = '{study_id}' AND variant_type = 'mutation' AND upper(mutation_status) IN ('GERMLINE', 'SOMATIC') GROUP BY hugo_gene_symbol, status ORDER BY count DESC LIMIT 20 ``` ### Find patients with germline mutations in a specific gene ```sql SELECT DISTINCT patient_unique_id, sample_unique_id, mutation_variant, mutation_type FROM genomic_event_derived WHERE hugo_gene_symbol = '{GENE}' AND upper(mutation_status) = 'GERMLINE' AND cancer_study_identifier = '{study_id}' AND variant_type = 'mutation' ``` ### Germline mutation frequency The numerator is samples with a germline mutation in the gene; the denominator is samples **profiled** for the gene — not samples that happen to have a mutation in it. Take the denominator from mutation-frequency-guide Step 2 (or … (2222 more chars) ◀ result # cBioPortal FAQ Guide Curated answers to frequently asked general questions about cBioPortal. Source: [cBioPortal FAQ](https://docs.cbioportal.org/user-guide/faq/). ## What is cBioPortal? cBioPortal for Cancer Genomics is an open-access, open-source resource for interactive exploration of multidimensional cancer genomics data sets. It was originally developed at Memorial Sloan Kettering Cancer Center (MSK) and is now maintained by a multi-institutional team. ## History - **2008**: cBioPortal first became available online. - **2012**: First major publication — Cerami et al., *Cancer Discovery*. - **2013**: Second major publication — Gao et al., *Science Signaling*. - **2023**: Third major publication — de Bruijn et al., *Cancer Research*. ## How to Cite cBioPortal When using cBioPortal in publications, cite these three papers: 1. Cerami et al. "The cBio Cancer Genomics Portal: An Open Platform for Exploring Multidimensional Cancer Genomics Data." *Cancer Discovery* 2, 401–404 (2012). doi:10.1158/2159-8290.CD-12-0095 2. Gao et al. "Integrative Analysis of Complex Cancer Genomics and Clinical Profiles Using the cBioPortal." *Science Signaling* 6, pl1 (2013). doi:10.1126/scisignal.2004088 3. de Bruijn et al. "Analysis and Visualization of Longitudinal Genomic and Clinical Data from the AACR Project GENIE Biopharma Collaborative in cBioPortal." *Cancer Research* 83, 3861–3867 (2023). doi:10.1158/0008-5472.CAN-23-0816 Also cite the specific study publication(s) whose data you used. ## What Data Types Does cBioPortal Contain? - **Mutations** (non-synonymous somatic mutations) - **DNA copy-number alterations** (CNA) from GISTIC or other methods - **mRNA expression** data (RNA-seq, microarray) - **Protein/phosphoprotein** levels (RPPA, mass spectrometry) - **DNA methylation** data - **Structural variants** (gene fusions, rearrangements) - **Clinical data** (demographics, diagnosis, treatment, outcomes) Note: Synonymous mutations are not included in cBioPortal. ## What Cancer Types Are in the Database? Answer from the studies actually loaded, not from the OncoTree ontology (OncoTree lists hundreds of codes, most with no study here). One query: ```sql SELECT tc.type_of_cancer_id, tc.name, tc.tissue, count() AS studies FROM cancer_study cs JOIN type_of_cancer tc ON cs.type_of_cancer_id = tc.type_of_cancer_id GROUP BY tc.type_of_cancer_id, tc.name, tc.tissue ORDER BY studies DESC; ``` Report how many distinct cancer types there are across how many studies, grouped by `tissue`. `mixed` ("Mixed Cancer Types") marks multi-cancer studies such as MSK-IMPACT and MSK-CHORD; their per-sample cancer types are in `clinical_data_derived` (`CANCER_TYPE`). ## Data Types Usually Not Stored Directly cBioPortal generally does not store: - raw CT, MRI, pathology-slide, or histology image files - raw sequencing files such as BAM, CRAM, or FASTQ - polygenic risk scores as a standard data type - full external clinical-trial databases However, this deployment … (7371 more chars) ▶ clickhouse_run_select_query { "query": "SELECT cancer_study_identifier, count() AS germline_count, uniqExact(sample_unique_id) AS samples FROM genomic_event_derived WHERE variant_type='mutation' AND upper(mutation_status)='GERMLINE' GROUP BY cancer_study_identifier ORDER BY germline_count DESC LIMIT 15" } ◀ result {"rows":[{"cancer_study_identifier":"coadread_tcga_pub","germline_count":101,"samples":59},{"cancer_study_identifier":"lung_nci_2022","germline_count":100,"samples":72},{"cancer_study_identifier":"pog570_bcgsc_2020","germline_count":74,"samples":68},{"cancer_study_identifier":"coadread_tcga","germline_count":70,"samples":48},{"cancer_study_identifier":"hnsc_a5consortium_2025","germline_count":67,"samples":66},{"cancer_study_identifier":"brca_tcga_pub","germline_count":64,"samples":58},{"cancer_study_identifier":"kirc_tcga_pub","germline_count":57,"samples":52},{"cancer_study_identifier":"all_stjude_2013","germline_count":52,"samples":33},{"cancer_study_identifier":"ov_tcga_pub","germline_count":51,"samples":50},{"cancer_study_identifier":"prad_su2c_2019","germline_count":40,"samples":39},{"cancer_study_identifier":"ov_tcga","germline_count":20,"samples":17},{"cancer_study_identifier":"aml_stjude_2024","germline_count":19,"samples":19},{"cancer_study_identifier":"mrt_bcgsc_2016","germline_count":10,"samples":8},{"cancer_study_identifier":"prad_tcga_pub","germline_count":7,"samples":7},{"cancer_study_identifier":"nhl_bcgsc_2011","germline_count":3,"samples":3}]} ═ answer (success) ## Germline variants in cBioPortal **Current support** - **Storage:** cBioPortal loads germline variants in the same mutation format as somatic ones. Each variant carries a mutation status of "Germline", either alongside somatic calls or in germline-only studies. - **Analysis:** Most features work the same for both variant types. These include the study summary, clinical correlations and plots, group comparison, OncoPrint, mutation diagrams and protein-domain views, and gene-panel coverage. - **Where it appears in this database:** Germline calls are present in only a modest number of studies, and mostly as small counts. The largest are: - `coadread_tcga_pub` (101 variants in 59 samples) - `lung_nci_2022` (100 variants in 72 samples) - `pog570_bcgsc_2020` (74 variants in 68 samples) - `coadread_tcga` (70 variants in 48 samples) - `hnsc_a5consortium_2025` (67 variants in 66 samples) - `brca_tcga_pub`, `kirc_tcga_pub`, `all_stjude_2013`, `ov_tcga_pub` and `prad_su2c_2019` - **Limitations:** - Many studies remove germline variants before loading, so absence of germline calls often means the data were filtered out. It doesn't mean the variants weren't present. - Many studies don't label mutation status at all, so those calls are "unknown", not somatic or germline. - Status labels are free text ("Germline", "GERMLINE", "germline"), so filters must match all spellings. - There is no standard field for ACMG pathogenicity classification. - Driver and OncoKB annotations, TMB and mutual-exclusivity analyses are designed for somatic variants. - Public instances may withhold germline data for privacy reasons. **Future support** I can't answer this from the database or from cBioPortal's documentation. I don't have the project roadmap, so I won't guess at planned features. For current plans, check the cBioPortal GitHub issues and releases or the docs at https://docs.cbioportal.org. Or ask the team through the cBioPortal community channels. **Comparison with alternatives** The database contains nothing about other tools, so this comparison is general background and not from cBioPortal data. Please check it against each tool's documentation. - **cBioPortal** is good for exploratory, cohort-level visualization. It handles germline and somatic variants side by side and links them to clinical data. It isn't built for formal germline classification, such as ACMG rules, or for handling raw variant files. - **Dedicated germline resources such as ClinVar, gnomAD and pipelines built around VCF files** are stronger for pathogenicity curation, population allele frequencies and raw-data processing. - **The GDC** is a repository for downloading the full data, including raw files. cBioPortal imports some GDC data and adds a friendlier way to explore it. If you want to load your own germline cohort, you can build a study from germline variants plus clinical attributes using the DataHub formats. To try it without submitting anything, you can run a private instance or use the standalone Mutation Mapper and OncoPrint tools. I didn't open any navigation links, because this question is about general support and not a specific study or gene. I can pull the germline variant counts for a particular gene or study if that would help.