Q (haiku): what's the best way to put clinical data into cbioportal? ▶ read_guide { "uri": "cbioportal://faq-guide" } ◀ result # cBioPortal FAQ Guide Curated answers to frequently asked general questions about cBioPortal. Source: [cBioPortal FAQ](https://docs.cbioportal.org/user-guide/faq/). ## What is cBioPortal? cBioPortal for Cancer Genomics is an open-access, open-source resource for interactive exploration of multidimensional cancer genomics data sets. It was originally developed at Memorial Sloan Kettering Cancer Center (MSK) and is now maintained by a multi-institutional team. ## History - **2008**: cBioPortal first became available online. - **2012**: First major publication — Cerami et al., *Cancer Discovery*. - **2013**: Second major publication — Gao et al., *Science Signaling*. - **2023**: Third major publication — de Bruijn et al., *Cancer Research*. ## How to Cite cBioPortal When using cBioPortal in publications, cite these three papers: 1. Cerami et al. "The cBio Cancer Genomics Portal: An Open Platform for Exploring Multidimensional Cancer Genomics Data." *Cancer Discovery* 2, 401–404 (2012). doi:10.1158/2159-8290.CD-12-0095 2. Gao et al. "Integrative Analysis of Complex Cancer Genomics and Clinical Profiles Using the cBioPortal." *Science Signaling* 6, pl1 (2013). doi:10.1126/scisignal.2004088 3. de Bruijn et al. "Analysis and Visualization of Longitudinal Genomic and Clinical Data from the AACR Project GENIE Biopharma Collaborative in cBioPortal." *Cancer Research* 83, 3861–3867 (2023). doi:10.1158/0008-5472.CAN-23-0816 Also cite the specific study publication(s) whose data you used. ## What Data Types Does cBioPortal Contain? - **Mutations** (non-synonymous somatic mutations) - **DNA copy-number alterations** (CNA) from GISTIC or other methods - **mRNA expression** data (RNA-seq, microarray) - **Protein/phosphoprotein** levels (RPPA, mass spectrometry) - **DNA methylation** data - **Structural variants** (gene fusions, rearrangements) - **Clinical data** (demographics, diagnosis, treatment, outcomes) Note: Synonymous mutations are not included in cBioPortal. ## What Cancer Types Are in the Database? Answer from the studies actually loaded, not from the OncoTree ontology (OncoTree lists hundreds of codes, most with no study here). One query: ```sql SELECT tc.type_of_cancer_id, tc.name, tc.tissue, count() AS studies FROM cancer_study cs JOIN type_of_cancer tc ON cs.type_of_cancer_id = tc.type_of_cancer_id GROUP BY tc.type_of_cancer_id, tc.name, tc.tissue ORDER BY studies DESC; ``` Report how many distinct cancer types there are across how many studies, grouped by `tissue`. `mixed` ("Mixed Cancer Types") marks multi-cancer studies such as MSK-IMPACT and MSK-CHORD; their per-sample cancer types are in `clinical_data_derived` (`CANCER_TYPE`). ## Data Types Usually Not Stored Directly cBioPortal generally does not store: - raw CT, MRI, pathology-slide, or histology image files - raw sequencing files such as BAM, CRAM, or FASTQ - polygenic risk scores as a standard data type - full external clinical-trial databases However, this deployment … (7371 more chars) ═ answer (success) Based on the guide, here's the best way to put clinical data into cBioPortal: ## Before You Start First, clarify your goal—are you trying to: 1. **Submit data to the public cBioPortal** — follow the official data curation process 2. **Run or populate a private cBioPortal instance** — use the standard study package formats 3. **Visualize your own data quickly** — use standalone tools ## Official Submission Process For the **public cBioPortal**, follow these steps: 1. **Prepare your data** using the cBioPortal datahub formats: - **Clinical data file** — tab-delimited, with required columns for patient and sample identifiers, plus your clinical attributes (age, sex, tumor grade, treatment, outcomes, etc.) - **Sample metadata** — links samples to patients - **Molecular data files** — mutations, copy number, expression (if applicable) - **Study metadata** — study name, description, cancer type, publication info 2. **Follow cBioPortal data standards:** - Use consistent identifiers for patients and samples - Format dates and numerical values consistently (the guide notes that `AGE` may be floored for de-identification) - Empty clinical values should be handled per cBioPortal conventions - Specify your reference genome (hg19/GRCh37 vs. hg38/GRCh38) 3. **Submit through the cBioPortal curation pipeline:** - Contact the cBioPortal team at the official documentation site: [cBioPortal Data Format Guide](https://docs.cbioportal.org/) - Provide study metadata, clinical files, and molecular data in the standardized formats ## Private Instance If you're running your own cBioPortal instance, create a **study package** with: - Clinical and sample metadata files (TSV format) - Molecular profile files (if applicable) - Study metadata configuration (XML or JSON) ## Quick Visualization Without Full Submission For rapid prototyping without full cBioPortal setup: - Use **cBioPortal standalone tools** like Mutation Mapper or OncoPrint for visualizing subsets - Export your processed data to CSV and load into R/Python for custom analysis **I'd recommend starting with the [cBioPortal documentation](https://docs.cbioportal.org/) on data formats—they have detailed templates for clinical data, file naming conventions, and validation scripts to ensure your data meets requirements before submission.** Is this for a public submission, or are you running a private instance?