Q (sonnet-5.5): what's the best way to put clinical data into cbioportal? ▶ read_guide { "uri": "cbioportal://faq-guide" } ◀ result # cBioPortal FAQ Guide Curated answers to frequently asked general questions about cBioPortal. Source: [cBioPortal FAQ](https://docs.cbioportal.org/user-guide/faq/). ## What is cBioPortal? cBioPortal for Cancer Genomics is an open-access, open-source resource for interactive exploration of multidimensional cancer genomics data sets. It was originally developed at Memorial Sloan Kettering Cancer Center (MSK) and is now maintained by a multi-institutional team. ## History - **2008**: cBioPortal first became available online. - **2012**: First major publication — Cerami et al., *Cancer Discovery*. - **2013**: Second major publication — Gao et al., *Science Signaling*. - **2023**: Third major publication — de Bruijn et al., *Cancer Research*. ## How to Cite cBioPortal When using cBioPortal in publications, cite these three papers: 1. Cerami et al. "The cBio Cancer Genomics Portal: An Open Platform for Exploring Multidimensional Cancer Genomics Data." *Cancer Discovery* 2, 401–404 (2012). doi:10.1158/2159-8290.CD-12-0095 2. Gao et al. "Integrative Analysis of Complex Cancer Genomics and Clinical Profiles Using the cBioPortal." *Science Signaling* 6, pl1 (2013). doi:10.1126/scisignal.2004088 3. de Bruijn et al. "Analysis and Visualization of Longitudinal Genomic and Clinical Data from the AACR Project GENIE Biopharma Collaborative in cBioPortal." *Cancer Research* 83, 3861–3867 (2023). doi:10.1158/0008-5472.CAN-23-0816 Also cite the specific study publication(s) whose data you used. ## What Data Types Does cBioPortal Contain? - **Mutations** (non-synonymous somatic mutations) - **DNA copy-number alterations** (CNA) from GISTIC or other methods - **mRNA expression** data (RNA-seq, microarray) - **Protein/phosphoprotein** levels (RPPA, mass spectrometry) - **DNA methylation** data - **Structural variants** (gene fusions, rearrangements) - **Clinical data** (demographics, diagnosis, treatment, outcomes) Note: Synonymous mutations are not included in cBioPortal. ## What Cancer Types Are in the Database? Answer from the studies actually loaded, not from the OncoTree ontology (OncoTree lists hundreds of codes, most with no study here). One query: ```sql SELECT tc.type_of_cancer_id, tc.name, tc.tissue, count() AS studies FROM cancer_study cs JOIN type_of_cancer tc ON cs.type_of_cancer_id = tc.type_of_cancer_id GROUP BY tc.type_of_cancer_id, tc.name, tc.tissue ORDER BY studies DESC; ``` Report how many distinct cancer types there are across how many studies, grouped by `tissue`. `mixed` ("Mixed Cancer Types") marks multi-cancer studies such as MSK-IMPACT and MSK-CHORD; their per-sample cancer types are in `clinical_data_derived` (`CANCER_TYPE`). ## Data Types Usually Not Stored Directly cBioPortal generally does not store: - raw CT, MRI, pathology-slide, or histology image files - raw sequencing files such as BAM, CRAM, or FASTQ - polygenic risk scores as a standard data type - full external clinical-trial databases However, this deployment … (7371 more chars) ═ answer (success) The best route depends on where you want the data to end up. Which of these fits? 1. **Submit to the public cBioPortal (cbioportal.org).** Follow the cBioPortal data curation and DataHub process. Your data has to be in the standard study format. 2. **Load into a private or local cBioPortal instance.** You build a study package and import it. 3. **Just visualize your own data without a full study.** Standalone tools such as Mutation Mapper and OncoPrint may be enough, depending on the data. ## Building a study package (options 1 and 2) Clinical data goes in tab-delimited files inside a study folder: - **Study metadata (`meta_study.txt`).** This holds the study ID, cancer type and description. - **Patient clinical data.** This is `data_clinical_patient.txt` with a matching `meta_clinical_patient.txt`. It has one row per patient, keyed by `PATIENT_ID`. It usually holds demographics and outcomes such as `OS_STATUS`, `OS_MONTHS`, `SEX` and `AGE`. - **Sample clinical data.** This is `data_clinical_sample.txt` with a matching `meta_clinical_sample.txt`. It has one row per sample, keyed by `PATIENT_ID` and `SAMPLE_ID`. It usually holds `CANCER_TYPE`, `ONCOTREE_CODE`, `SAMPLE_TYPE` and similar fields. - **Header rows.** Each clinical file starts with four `#` lines: display name, description, data type (`STRING`, `NUMBER` or `BOOLEAN`), and priority. The actual column names come on the fifth line. - **Timeline data (optional).** Treatments, diagnoses and other events go in a separate timeline file if you need them. - **Case lists.** You also need these, for example `cases_all.txt`. ## Practical tips - Keep `PATIENT_ID` and `SAMPLE_ID` consistent across every file. Any mismatch will cause import errors. - Use the standard attribute names where they exist, such as `OS_STATUS` and `OS_MONTHS`. The portal recognizes these and builds survival plots from them. - Write missing values as `NA` or leave the cell empty. Don't use placeholder text. - Run the validator (`validateData.py`) before importing. It catches most formatting problems. - Import with the cBioPortal importer script for a private instance. For the public portal, the curation team handles this step. - If you want clinical data attached to a study with germline or somatic variants, load it in the same package. The official documentation covers the exact file formats, including the `#` header lines, and the importer and validator usage. It's in the File Formats and Data Loading sections at [docs.cbioportal.org](https://docs.cbioportal.org). The formats can change between versions, so check them against your instance's version. Tell me which option you're going for, and what data you have (for example patient-level only, or with samples and treatments). I can then give you a more specific file layout.