Q (haiku): what's the best way to put clinical data into cbioportal? ▶ list_guides {} ◀ result {"result":[{"uri":"cbioportal://mutation-frequency-guide","description":"Comprehensive guide for calculating gene mutation frequencies with gene-specific profiling denominators"},{"uri":"cbioportal://clinical-data-guide","description":"Guide for querying clinical data including patient vs sample level considerations"},{"uri":"cbioportal://sample-filtering-guide","description":"Guide for filtering samples and studies in cBioPortal queries"},{"uri":"cbioportal://common-pitfalls","description":"Guide to avoid common mistakes when querying cBioPortal data. If you already know which numbered pitfall applies, fetch just that section via read_guide(\"cbioportal://common-pitfalls#\") (e.g. #16) instead of the full guide"},{"uri":"cbioportal://treatment-guide","description":"Guide for querying treatment/clinical event data including drug agents, timelines, and linking to genomic data"},{"uri":"cbioportal://faq-guide","description":"General cBioPortal FAQ: history, how to cite, data types, reference genome, abbreviations, GISTIC thresholds, API access"},{"uri":"cbioportal://statistical-tests-guide","description":"Statistical test selection guide — decision matrix for choosing Fisher's exact, Wilcoxon, chi-squared, t-test, ANOVA, etc. based on data type and group count"},{"uri":"cbioportal://gene-expression-guide","description":"Gene expression / copy-number / methylation analysis. Covers genetic_alteration_derived, profile_type discovery, and the gene_pair_coexpression view for Spearman correlation between two genes"},{"uri":"cbioportal://external-resources-guide","description":"Guide for finding external linked resources such as imaging, pathology, Minerva, HTAN, or other resource_* table links before declaring data unavailable"},{"uri":"cbioportal://gene-resolution-guide","description":"Guide for resolving ambiguous gene symbols, aliases, gene families, and shorthand such as CD3 before querying expression or alteration data"},{"uri":"cbioportal://study-resolution-guide","description":"Guide for resolving requested studies, avoiding silent substitute cohorts, and redirecting to known external cBioPortal instances when data is not in this deployment"},{"uri":"cbioportal://germline-guide","description":"Guide for querying germline variant data — storage columns, study discovery, query patterns, and somatic vs germline considerations"},{"uri":"cbioportal://study-guide/{study_id}","description":"Dynamic study-specific guide - use get_study_guide(study_id) tool to generate"}]} ═ answer (success) Your question is about **data ingestion into cBioPortal** — how to contribute or upload clinical data. This is **not a query about data already in the database**, so it's out of scope for my query/navigation tools, which assume data is already loaded. For submitting clinical data to cBioPortal, here are the official resources: --- ## **Best approaches:** ### 1. **Official cBioPortal Documentation** The canonical guide is at **[cBioPortal Data Loading Guide](https://docs.cbioportal.org/file-formats-and-conventions/data-loading/)** — covers: - Clinical attributes format (`.txt` files with `PATIENT_ID`, `SAMPLE_ID`, attribute columns) - Sample/patient data distinctions - Validation and loading via the cBioPortal command-line tools ### 2. **If you're contributing to an existing cBioPortal deployment** (e.g., our local MSK instance): Contact the cBioPortal admin/data curator for your deployment. Data governance and file formats may differ by institution. (Your email suggests MSK — check with your data governance team.) ### 3. **If you're running your own cBioPortal instance** - Fork/clone the [cBioPortal GitHub repository](https://github.com/cBioPortal/cbioportal) - Use the **`scripts/importer/` tools** to validate and load data - Follow the [file format specs](https://docs.cbioportal.org/file-formats-and-conventions/) ### 4. **For quick validation/exploration** before formal loading Use my query tools to inspect what's already in the database — this can help you understand the schema and format conventions used locally. --- **What would help most?** - Are you loading data into an **existing deployment** (like MSK's cBioPortal)? - Are you **creating a new study/project** for cBioPortal? - Do you need to **query existing clinical data** to understand the format?