Q (sonnet): what's the best way to put clinical data into cbioportal? ═ answer (success) The way you get clinical data into cBioPortal depends on whether you're loading it into your own local/institutional instance or contributing to the public portal. Here's the standard approach: ## Format cBioPortal expects clinical data as **tab-delimited text files**, split into two levels: 1. **Patient-level clinical data** (`data_clinical_patient.txt`) — one row per patient, attributes like SEX, AGE, OS_STATUS, OS_MONTHS, etc. 2. **Sample-level clinical data** (`data_clinical_sample.txt`) — one row per sample, attributes like SAMPLE_ID, PATIENT_ID, CANCER_TYPE, SAMPLE_TYPE, etc. Each file requires a specific header block (5 lines) before the actual column headers: ``` #Patient Identifier Sex ... #Identifier to uniquely specify a patient. ... #STRING ... #1 ... #PATIENT_ID SEX ... PATIENT_ID SEX ... patient1 Male ... ``` The five `#`-prefixed lines define: display name, description, datatype, priority, and are followed by the actual column name row. ## Steps to load it 1. **Prepare the files** following the format above — `PATIENT_ID` and `SAMPLE_ID` must match identifiers used in your mutation/CNA/expression files. 2. **Create/edit a `meta_clinical_patient.txt`** and `meta_clinical_sample.txt` file pointing to the data files, specifying `cancer_study_identifier` and `genetic_alteration_type: CLINICAL`. 3. Place these alongside your other data files (mutations, CNA, etc.) in a single study directory, with a top-level `meta_study.txt`. 4. **Import using the cBioPortal `metaImport.py` / `importer` scripts** (via the `cbioportal-docker-compose` or command-line importer) — this is done on a *local instance* you control, not on the public cbioportal.org site. 5. Validate first with the built-in **validator script** (`validateData.py`) to catch formatting errors before import. ## Key resources - Full data format specs and file templates: [cBioPortal File Formats documentation](https://docs.cbioportal.org/file-formats/) - Example study structure: the [datahub GitHub repo](https://github.com/cBioPortal/datahub) has many real `data_clinical_patient.txt` / `data_clinical_sample.txt` examples you can pattern-match against. - If you want your study included in the **public** cbioportal.org instance, the process is to submit it via the datahub repo (pull request) rather than direct upload — the cBioPortal team reviews and imports it. A few practical tips: - Missing values should be encoded as `NA`. - `OS_STATUS` should be formatted as `"0:LIVING"` / `"1:DECEASED"` and `OS_MONTHS` as a number, if you want survival analysis (Kaplan-Meier plots) to work correctly downstream. - Attribute names are case-sensitive and some (like `OS_STATUS`, `OS_MONTHS`, `SEX`, `CANCER_TYPE`) are "special" and drive specific portal features — matching these exact names is worthwhile even for custom attributes you want the survival/comparison tools to recognize. Would you like me to pull up a real example study's clinical files (via the database) so you can see the exact structure in practice, or are you setting up your own local cBioPortal instance and want more detail on the import scripts?