Q (sonnet): Can you find a study that I may emulate that has a data_clinical_outcomes.txt file and associated meta file to interogate ▶ list_guides {} ◀ result {"result":[{"uri":"cbioportal://mutation-frequency-guide","description":"Comprehensive guide for calculating gene mutation frequencies with gene-specific profiling denominators"},{"uri":"cbioportal://clinical-data-guide","description":"Guide for querying clinical data including patient vs sample level considerations"},{"uri":"cbioportal://sample-filtering-guide","description":"Guide for filtering samples and studies in cBioPortal queries"},{"uri":"cbioportal://common-pitfalls","description":"Guide to avoid common mistakes when querying cBioPortal data. If you already know which numbered pitfall applies, fetch just that section via read_guide(\"cbioportal://common-pitfalls#\") (e.g. #16) instead of the full guide"},{"uri":"cbioportal://treatment-guide","description":"Guide for querying treatment/clinical event data including drug agents, timelines, and linking to genomic data"},{"uri":"cbioportal://faq-guide","description":"General cBioPortal FAQ: history, how to cite, data types, reference genome, abbreviations, GISTIC thresholds, API access"},{"uri":"cbioportal://statistical-tests-guide","description":"Statistical test selection guide — decision matrix for choosing Fisher's exact, Wilcoxon, chi-squared, t-test, ANOVA, etc. based on data type and group count"},{"uri":"cbioportal://gene-expression-guide","description":"Gene expression / copy-number / methylation analysis. Covers genetic_alteration_derived, profile_type discovery, and the gene_pair_coexpression view for Spearman correlation between two genes"},{"uri":"cbioportal://external-resources-guide","description":"Guide for finding external linked resources such as imaging, pathology, Minerva, HTAN, or other resource_* table links before declaring data unavailable"},{"uri":"cbioportal://gene-resolution-guide","description":"Guide for resolving ambiguous gene symbols, aliases, gene families, and shorthand such as CD3 before querying expression or alteration data"},{"uri":"cbioportal://study-resolution-guide","description":"Guide for resolving requested studies, avoiding silent substitute cohorts, and redirecting to known external cBioPortal instances when data is not in this deployment"},{"uri":"cbioportal://germline-guide","description":"Guide for querying germline variant data — storage columns, study discovery, query patterns, and somatic vs germline considerations"},{"uri":"cbioportal://study-guide/{study_id}","description":"Dynamic study-specific guide - use get_study_guide(study_id) tool to generate"}]} ▶ read_guide { "uri": "cbioportal://faq-guide" } ▶ read_guide { "uri": "cbioportal://clinical-data-guide" } ◀ result # cBioPortal FAQ Guide Curated answers to frequently asked general questions about cBioPortal. Source: [cBioPortal FAQ](https://docs.cbioportal.org/user-guide/faq/). ## What is cBioPortal? cBioPortal for Cancer Genomics is an open-access, open-source resource for interactive exploration of multidimensional cancer genomics data sets. It was originally developed at Memorial Sloan Kettering Cancer Center (MSK) and is now maintained by a multi-institutional team. ## History - **2008**: cBioPortal first became available online. - **2012**: First major publication — Cerami et al., *Cancer Discovery*. - **2013**: Second major publication — Gao et al., *Science Signaling*. - **2023**: Third major publication — de Bruijn et al., *Cancer Research*. ## How to Cite cBioPortal When using cBioPortal in publications, cite these three papers: 1. Cerami et al. "The cBio Cancer Genomics Portal: An Open Platform for Exploring Multidimensional Cancer Genomics Data." *Cancer Discovery* 2, 401–404 (2012). doi:10.1158/2159-8290.CD-12-0095 2. Gao et al. "Integrative Analysis of Complex Cancer Genomics and Clinical Profiles Using the cBioPortal." *Science Signaling* 6, pl1 (2013). doi:10.1126/scisignal.2004088 3. de Bruijn et al. "Analysis and Visualization of Longitudinal Genomic and Clinical Data from the AACR Project GENIE Biopharma Collaborative in cBioPortal." *Cancer Research* 83, 3861–3867 (2023). doi:10.1158/0008-5472.CAN-23-0816 Also cite the specific study publication(s) whose data you used. ## What Data Types Does cBioPortal Contain? - **Mutations** (non-synonymous somatic mutations) - **DNA copy-number alterations** (CNA) from GISTIC or other methods - **mRNA expression** data (RNA-seq, microarray) - **Protein/phosphoprotein** levels (RPPA, mass spectrometry) - **DNA methylation** data - **Structural variants** (gene fusions, rearrangements) - **Clinical data** (demographics, diagnosis, treatment, outcomes) Note: Synonymous mutations are not included in cBioPortal. ## Data Types Usually Not Stored Directly cBioPortal generally does not store: - raw CT, MRI, pathology-slide, or histology image files - raw sequencing files such as BAM, CRAM, or FASTQ - polygenic risk scores as a standard data type - full external clinical-trial databases However, this deployment may contain links to external viewers or portals through `resource_*` tables. For imaging, pathology, Minerva, HTAN, or viewer questions, read `cbioportal://external-resources-guide` and check those tables before saying the data is absent. ## What the MCP Agent Can and Cannot Produce The MCP agent can: - query cBioPortal data and return tables in text/JSON/Markdown - summarize counts, frequencies, and available attributes - provide cBioPortal study links and DataHub download links when applicable - provide SQL snippets or handoff instructions for R/Python/cBioPortal tools The MCP agent should not promise to: - render Kaplan-Meier plots or other visual figures directly - export large … (6558 more chars) ◀ result # Clinical Data Query Guide ## Overview Clinical data in cBioPortal is stored at both patient and sample levels. Understanding the distinction is crucial for accurate analysis. ## Data Organization ### Patient-Level vs Sample-Level Data - **Patient-level**: Demographics, overall survival, disease stage (stored once per patient) - **Sample-level**: Sample type, sequencing platform, purity (can have multiple per patient) ### Key Tables - `clinical_patient`: Patient-level clinical attributes - `clinical_sample`: Sample-level clinical attributes - `clinical_data_derived`: Pre-joined view combining both levels - `clinical_attribute_meta`: Metadata about available clinical attributes ## Recommended Approach: Use clinical_data_derived The `clinical_data_derived` table is pre-joined and optimized for most queries: ```sql -- Get clinical data for specific attributes SELECT sample_unique_id, patient_unique_id, attribute_name, attribute_value FROM clinical_data_derived WHERE cancer_study_identifier = 'your_study_id' AND attribute_name IN ('CANCER_TYPE', 'SAMPLE_TYPE', 'TMB_NONSYNONYMOUS'); ``` ## Clinical Attribute Discovery ### Use clinical_attribute_meta for Discovering Available Attributes **Always start here** to see what clinical attributes are available for a specific study: ```sql -- Discover available clinical attributes for a study SELECT attr_id, description, patient_attribute, cancer_study_id FROM clinical_attribute_meta WHERE cancer_study_id = ( SELECT cancer_study_id FROM cancer_study WHERE cancer_study_identifier = 'your_study_id' ) ORDER BY patient_attribute, attr_id; ``` **Key Fields:** - **attr_id**: matches attr_id in clinical_sample/clinical_patient tables - **description**: human-readable description of the attribute - **patient_attribute**: true = patient attribute, false = sample attribute - **cancer_study_id**: links to cancer_study table (filter by study) ## Attribute Semantics and Matching ### Case-Insensitive Matching for Attribute Values Clinical values are free text across studies and may differ only by case. For example, a controlled-looking value such as germline mutation status may appear as `GERMLINE`, `Germline`, or another case variant. When filtering `clinical_data_derived.attribute_value`, use case-insensitive matching unless you have already profiled the exact values in the target study: ```sql -- Correct: case-insensitive clinical value filter SELECT DISTINCT sample_unique_id, patient_unique_id FROM clinical_data_derived WHERE cancer_study_identifier = 'your_study_id' AND attribute_name = 'MUTATION_STATUS' AND upper(attribute_value) = 'GERMLINE'; ``` Do not write `attribute_value = 'GERMLINE'` without first checking all distinct values for that attribute in the study. ### Query the Requested Attribute, Not a Proxy Do not infer one clinical attribute from a related subtype or marker. Query the actual requested attribute when it … (17073 more chars) ═ answer (success) I need to flag an important limitation before going further: the ClickHouse database I query only contains the *parsed* clinical attributes and events — it does not preserve the original file names from a study's source data package (e.g., `data_clinical_outcomes.txt`, `meta_clinical_outcomes.txt`). Those file names are a DataHub/study-package convention, not something stored as queryable metadata in this database, so I can't search for "which study has a file literally named `data_clinical_outcomes.txt`." That said, I can point you toward what's actually a good match for your intent: a study with a **separate, non-standard clinical file capturing outcomes/timeline data beyond the standard patient/sample files** — which is exactly what a `data_clinical_outcomes.txt`-style supplementary file is used for in DataHub packages (it's an optional pattern some contributors use alongside `data_clinical_patient.txt`/`data_clinical_sample.txt`, loaded the same way via a `meta_clinical_*.txt` with `datatype: CLINICAL`). A couple of options depending on what "emulate" means to you: 1. **If you want a real GitHub example of this exact file pattern** — that's outside what I can verify from the database; I dond recommend browsing the [cBioPortal datahub repository](https://github.com/cBioPortal/datahub/tree/master/public) directly (e.g., search the repo for `data_clinical_outcomes.txt`) since GitHub code search will show you the literal files and their paired meta files. 2. **If you want a study in this database with rich, detailed outcome/survival attributes** (multiple survival endpoints — OS, DFS, PFS, recurrence, treatment response — that would map well onto an "outcomes" file structure), I can find one for you. Could you clarify which you're after — or would you like me to identify a study here with the richest set of outcome-related clinical attributes as a modeling template?