Q (sonnet): Which cancer types have the highest aneuploidy scores and how does this correlate with mutation burden across TCGA Pan-Cancer studies? ▶ list_guides {} ◀ result {"result":[{"uri":"cbioportal://mutation-frequency-guide","description":"Comprehensive guide for calculating gene mutation frequencies with gene-specific profiling denominators"},{"uri":"cbioportal://clinical-data-guide","description":"Guide for querying clinical data including patient vs sample level considerations"},{"uri":"cbioportal://sample-filtering-guide","description":"Guide for filtering samples and studies in cBioPortal queries"},{"uri":"cbioportal://common-pitfalls","description":"Guide to avoid common mistakes when querying cBioPortal data. If you already know which numbered pitfall applies, fetch just that section via read_guide(\"cbioportal://common-pitfalls#\") (e.g. #16) instead of the full guide"},{"uri":"cbioportal://treatment-guide","description":"Guide for querying treatment/clinical event data including drug agents, timelines, and linking to genomic data"},{"uri":"cbioportal://faq-guide","description":"General cBioPortal FAQ: history, how to cite, data types, reference genome, abbreviations, GISTIC thresholds, API access"},{"uri":"cbioportal://statistical-tests-guide","description":"Statistical test selection guide — decision matrix for choosing Fisher's exact, Wilcoxon, chi-squared, t-test, ANOVA, etc. based on data type and group count"},{"uri":"cbioportal://gene-expression-guide","description":"Gene expression / copy-number / methylation analysis. Covers genetic_alteration_derived, profile_type discovery, and the gene_pair_coexpression view for Spearman correlation between two genes"},{"uri":"cbioportal://external-resources-guide","description":"Guide for finding external linked resources such as imaging, pathology, Minerva, HTAN, or other resource_* table links before declaring data unavailable"},{"uri":"cbioportal://gene-resolution-guide","description":"Guide for resolving ambiguous gene symbols, aliases, gene families, and shorthand such as CD3 before querying expression or alteration data"},{"uri":"cbioportal://study-resolution-guide","description":"Guide for resolving requested studies, avoiding silent substitute cohorts, and redirecting to known external cBioPortal instances when data is not in this deployment"},{"uri":"cbioportal://germline-guide","description":"Guide for querying germline variant data — storage columns, study discovery, query patterns, and somatic vs germline considerations"},{"uri":"cbioportal://study-guide/{study_id}","description":"Dynamic study-specific guide - use get_study_guide(study_id) tool to generate"}]} ▶ read_guide { "uri": "cbioportal://clinical-data-guide" } ◀ result # Clinical Data Query Guide ## Overview Clinical data in cBioPortal is stored at both patient and sample levels. Understanding the distinction is crucial for accurate analysis. ## Data Organization ### Patient-Level vs Sample-Level Data - **Patient-level**: Demographics, overall survival, disease stage (stored once per patient) - **Sample-level**: Sample type, sequencing platform, purity (can have multiple per patient) ### Key Tables - `clinical_patient`: Patient-level clinical attributes - `clinical_sample`: Sample-level clinical attributes - `clinical_data_derived`: Pre-joined view combining both levels - `clinical_attribute_meta`: Metadata about available clinical attributes ## Recommended Approach: Use clinical_data_derived The `clinical_data_derived` table is pre-joined and optimized for most queries: ```sql -- Get clinical data for specific attributes SELECT sample_unique_id, patient_unique_id, attribute_name, attribute_value FROM clinical_data_derived WHERE cancer_study_identifier = 'your_study_id' AND attribute_name IN ('CANCER_TYPE', 'SAMPLE_TYPE', 'TMB_NONSYNONYMOUS'); ``` ## Clinical Attribute Discovery ### Use clinical_attribute_meta for Discovering Available Attributes **Always start here** to see what clinical attributes are available for a specific study: ```sql -- Discover available clinical attributes for a study SELECT attr_id, description, patient_attribute, cancer_study_id FROM clinical_attribute_meta WHERE cancer_study_id = ( SELECT cancer_study_id FROM cancer_study WHERE cancer_study_identifier = 'your_study_id' ) ORDER BY patient_attribute, attr_id; ``` **Key Fields:** - **attr_id**: matches attr_id in clinical_sample/clinical_patient tables - **description**: human-readable description of the attribute - **patient_attribute**: true = patient attribute, false = sample attribute - **cancer_study_id**: links to cancer_study table (filter by study) ## Attribute Semantics and Matching ### Case-Insensitive Matching for Attribute Values Clinical values are free text across studies and may differ only by case. For example, a controlled-looking value such as germline mutation status may appear as `GERMLINE`, `Germline`, or another case variant. When filtering `clinical_data_derived.attribute_value`, use case-insensitive matching unless you have already profiled the exact values in the target study: ```sql -- Correct: case-insensitive clinical value filter SELECT DISTINCT sample_unique_id, patient_unique_id FROM clinical_data_derived WHERE cancer_study_identifier = 'your_study_id' AND attribute_name = 'MUTATION_STATUS' AND upper(attribute_value) = 'GERMLINE'; ``` Do not write `attribute_value = 'GERMLINE'` without first checking all distinct values for that attribute in the study. ### Query the Requested Attribute, Not a Proxy Do not infer one clinical attribute from a related subtype or marker. Query the actual requested attribute when it … (17073 more chars) ▶ read_guide { "uri": "cbioportal://sample-filtering-guide" } ▶ read_guide { "uri": "cbioportal://statistical-tests-guide" } ◀ result # Sample and Study Filtering Guide ## Overview Proper filtering is essential for meaningful cBioPortal analysis. This guide covers filtering by studies, sample types, and other criteria. ## Study-Level Filtering ### 1. Always Filter by Study Every query should specify a study to ensure consistent results: ```sql -- Always include study filtering SELECT * FROM your_table WHERE cancer_study_identifier = 'your_study_id' -- Additional filters... ``` ### 2. Find Available Studies ```sql -- Discover available studies SELECT cancer_study_identifier, name, description, type_of_cancer_id FROM cancer_study ORDER BY cancer_study_identifier; ``` ### 3. Study Information ```sql -- Get detailed study information SELECT cs.cancer_study_identifier, cs.name as study_name, cs.description, cs.sample_count, COUNT(DISTINCT p.internal_id) as patient_count FROM cancer_study cs LEFT JOIN patient p ON cs.cancer_study_id = p.cancer_study_id WHERE cs.cancer_study_identifier = 'your_study_id' GROUP BY cs.cancer_study_identifier, cs.name, cs.description, cs.sample_count; ``` ### 4. Find Studies by Available Data Types Use this when the user asks *"which studies have mutation and copy-number data for X"*, *"studies with expression for Y"*, *"is there any imaging data"*, or any *"studies with Z data"* question. `cancer_study` has one precomputed sample count per data type — filter on them. No join and no schema exploration needed. These are the same counts as the portal's study list and its "Data type" filter. **Canonical query — studies with ALL requested data types for a cancer type:** ```sql SELECT cancer_study_identifier, name, sample_count, mutation_sample_count, cna_sample_count FROM cancer_study WHERE type_of_cancer_id = 'luad' -- ← OncoTree code from search_oncotree AND mutation_sample_count > 0 -- ← one condition per required data type AND cna_sample_count > 0 ORDER BY sample_count DESC; ``` **Data type → column** (a count of 0 means the study doesn't have that data): | Data type in the user's question | Column (portal "Data type" filter label) | |---|---| | samples in the study | `sample_count` | | mutation / mutations | `mutation_sample_count` ("Mutations") | | copy-number / CNA / amplification / deletion | `cna_sample_count` ("CNA") | | structural variant / fusion | `structural_variant_sample_count` | | RNA / mRNA / gene expression (any platform) | `mrna_expression_sample_count` — use this for "which studies have expression data" | | specifically the portal's "RNA-Seq" filter | `rna_seq_sample_count` ("RNA-Seq") — counts only the `_rna_seq_v2_mrna` list; 0 for many GDC and iAtlas studies that do have RNA-Seq | | specifically the portal's "RNA (microarray)" filter | `mrna_microarray_sample_count` ("RNA (microarray)") — counts only the `_mrna` list | | microRNA | `mirna_sample_count` ("miRNA") | | protein / RPPA | `rppa_sample_count` ("RPPA") | | mas … (7866 more chars) ◀ result Statistical Test Selection Guide ================================ Purpose ------- This guide ensures the correct statistical test is selected before performing any group comparison, matching cBioPortal's own Group Comparison defaults. HARD RULES — NEVER FABRICATE A STATISTIC ---------------------------------------- ClickHouse cannot run statistical tests. The agent therefore must NEVER produce a derived statistic that is not a literal column value from a SQL result. Specifically: 1. **Never invent a p-value.** Not "p < 0.001", not "p ≈ 0.05", not any p-value. If the user asks "what is the p-value?", the answer is *"I can't compute that — here is the 2x2 contingency table (or group statistics). Run it in cBioPortal's Group Comparison tab, in R with `fisher.test(...)` / `wilcox.test(...)`, or in Python with `scipy.stats.fisher_exact(...)` / `mannwhitneyu(...)`."* 2. **Never claim mutual exclusivity (or co-occurrence) from a contingency table alone.** A 2x2 table is not a test. The shape "altered/not altered × group A/group B" needs Fisher's exact + a defined direction (odds ratio < 1 with significant p). Without that test, the agent presents the table and stops. Descriptive phrasing is also a claim: "largely/mostly mutually exclusive", "rarely co-occur", "tend to co-occur" are forbidden without the test — point to cBioPortal's Mutual Exclusivity tab instead. 3. **Never report a "median" that came from `AVG(...)` or any non-median aggregate.** "Median" and "mean" are different statistics; for skewed clinical distributions (especially survival) they differ substantially. Use ClickHouse's `quantile(0.5)(...)` for actual median, and label arithmetic averages as "mean", never "median". 4. **Never report a hazard ratio, odds ratio, risk ratio, or relative risk** that wasn't computed by an external tool. These require regression / model fitting that ClickHouse does not do. 5. **Never report median overall survival from `AVG(OS_MONTHS)` or even `quantile(0.5)(OS_MONTHS)`.** Median OS requires Kaplan-Meier estimation, which handles censoring (`OS_STATUS = 0:LIVING` means the event hasn't happened yet). Naive medians/means over `OS_MONTHS` ignore censoring and are systematically wrong. The correct handoff: return the raw `(OS_MONTHS, OS_STATUS)` pairs (or descriptive counts: N events, N censored, follow-up range) and tell the user to run KM in R (`survival::survfit`) or Python (`lifelines.KaplanMeierFitter`), or use cBioPortal's Survival comparison. If asked to do any of (1)–(5), respond with the appropriate handoff template from the "Approved Response Templates" section below — do not produce the number. Ambiguous-Term Clarification (before any analysis) -------------------------------------------------- If the user's question uses an ambiguous term, ask before computing anything. Examples: - **"aggressive"** → could mean shorter overall survival, higher metastasis rate, higher grade/stage, higher TMB, or specific molecular features. Ask which definit … (9984 more chars) ▶ clickhouse_run_select_query { "query": "SELECT DISTINCT attribute_name FROM clinical_data_derived WHERE upper(attribute_name) LIKE '%ANEUPLOIDY%' LIMIT 20" } ◀ result {"rows":[{"attribute_name":"ANEUPLOIDY"},{"attribute_name":"ANEUPLOIDY_SCORE"},{"attribute_name":"PERCENT_ANEUPLOIDY"},{"attribute_name":"ANEUPLOIDY_AMP_SCORE"},{"attribute_name":"ANEUPLOIDY_DEL_SCORE"}]} ▶ clickhouse_run_select_query { "query": "SELECT cancer_study_identifier FROM cancer_study_query_preferences WHERE preference_name = 'pan_cancer_tcga' ORDER BY cancer_study_identifier" } ◀ result {"rows":[{"cancer_study_identifier":"acc_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"blca_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"brca_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"cesc_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"chol_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"coadread_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"dlbc_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"esca_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"gbm_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"hnsc_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"kich_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"kirc_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"kirp_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"laml_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"lgg_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"lihc_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"luad_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"lusc_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"meso_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"ov_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"paad_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"pcpg_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"prad_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"sarc_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"skcm_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"stad_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"tgct_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"thca_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"thym_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"ucec_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"ucs_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"uvm_tcga_pan_can_atlas_2018"}]} ▶ clickhouse_run_select_query { "query": " WITH studies AS ( SELECT cancer_study_identifier FROM cancer_study_query_preferences WHERE preference_name = 'pan_cancer_tcga' ) SELECT cancer_study_identifier AS study, count() AS n, round(avg(aneuploidy),2) AS mean_aneuploidy, round(quantile(0.5)(aneuploidy),2) AS median_aneuploidy, round(avg(tmb),2) AS mean_tmb, round(quantile(0.5)(tmb),2) AS median_tmb FROM ( SELECT sample_unique_id, cancer_study_identifier, maxIf(toFloat64OrNull(attribute_value), attribute_name = 'ANEUPLOIDY_SCORE') AS aneuploidy, maxIf(toFloat64OrNull(attribute_value), attribute_name = 'TMB_NONSYNONYMOUS') AS tmb FROM clinical_data_derived WHERE cancer_study_identifier IN (SELECT cancer_study_identifier FROM studies) AND attribute_name IN ('ANEUPLOIDY_SCORE','TMB_NONSYNONYMOUS') GROUP BY sample_unique_id, cancer_study_identifier ) WHERE aneuploidy IS NOT NULL GROUP BY cancer_study_identifier ORDER BY mean_aneuploidy DESC ", "max_rows": 40 } ◀ result {"rows":[{"study":"tgct_tcga_pan_can_atlas_2018","n":149,"mean_aneuploidy":18.83,"median_aneuploidy":19,"mean_tmb":0.49,"median_tmb":0.4},{"study":"acc_tcga_pan_can_atlas_2018","n":89,"mean_aneuploidy":18.29,"median_aneuploidy":17,"mean_tmb":2.83,"median_tmb":0.9},{"study":"ucs_tcga_pan_can_atlas_2018","n":56,"mean_aneuploidy":17.21,"median_aneuploidy":19,"mean_tmb":4.32,"median_tmb":1.55},{"study":"lusc_tcga_pan_can_atlas_2018","n":482,"mean_aneuploidy":16.12,"median_aneuploidy":16,"mean_tmb":9.51,"median_tmb":7.73},{"study":"luad_tcga_pan_can_atlas_2018","n":501,"mean_aneuploidy":15.04,"median_aneuploidy":16,"mean_tmb":10.24,"median_tmb":6.87},{"study":"kich_tcga_pan_can_atlas_2018","n":65,"mean_aneuploidy":14.86,"median_aneuploidy":14,"mean_tmb":1.09,"median_tmb":0.67},{"study":"esca_tcga_pan_can_atlas_2018","n":162,"mean_aneuploidy":14.57,"median_aneuploidy":14.5,"mean_tmb":5.08,"median_tmb":3.65},{"study":"ov_tcga_pan_can_atlas_2018","n":552,"mean_aneuploidy":13.98,"median_aneuploidy":13,"mean_tmb":2.26,"median_tmb":1.9},{"study":"blca_tcga_pan_can_atlas_2018","n":401,"mean_aneuploidy":13.55,"median_aneuploidy":14,"mean_tmb":8.57,"median_tmb":5.9},{"study":"skcm_tcga_pan_can_atlas_2018","n":434,"mean_aneuploidy":12.99,"median_aneuploidy":11,"mean_tmb":23.85,"median_tmb":15.3},{"study":"coadread_tcga_pan_can_atlas_2018","n":586,"mean_aneuploidy":12.3,"median_aneuploidy":12,"mean_tmb":13.02,"median_tmb":3.5},{"study":"brca_tcga_pan_can_atlas_2018","n":1041,"mean_aneuploidy":12.12,"median_aneuploidy":11,"mean_tmb":2.76,"median_tmb":1.33},{"study":"hnsc_tcga_pan_can_atlas_2018","n":510,"mean_aneuploidy":11.33,"median_aneuploidy":10,"mean_tmb":5.35,"median_tmb":3.58},{"study":"stad_tcga_pan_can_atlas_2018","n":427,"mean_aneuploidy":11.18,"median_aneuploidy":10,"mean_tmb":12.87,"median_tmb":4},{"study":"sarc_tcga_pan_can_atlas_2018","n":244,"mean_aneuploidy":10.86,"median_aneuploidy":9.5,"mean_tmb":2.38,"median_tmb":1.33},{"study":"lihc_tcga_pan_can_atlas_2018","n":361,"mean_aneuploidy":10.48,"median_aneuploidy":8,"mean_tmb":3.54,"median_tmb":2.83},{"study":"chol_tcga_pan_can_atlas_2018","n":36,"mean_aneuploidy":10.42,"median_aneuploidy":9,"mean_tmb":2.32,"median_tmb":1.45},{"study":"cesc_tcga_pan_can_atlas_2018","n":294,"mean_aneuploidy":10.1,"median_aneuploidy":8,"mean_tmb":6.42,"median_tmb":2.83},{"study":"paad_tcga_pan_can_atlas_2018","n":165,"mean_aneuploidy":9.28,"median_aneuploidy":7,"mean_tmb":4.25,"median_tmb":1.2},{"study":"kirp_tcga_pan_can_atlas_2018","n":280,"mean_aneuploidy":8.24,"median_aneuploidy":8,"mean_tmb":2.75,"median_tmb":2.23},{"study":"gbm_tcga_pan_can_atlas_2018","n":566,"mean_aneuploidy":8.2,"median_aneuploidy":7,"mean_tmb":4.03,"median_tmb":1.73},{"study":"meso_tcga_pan_can_atlas_2018","n":81,"mean_aneuploidy":8.16,"median_aneuploidy":6,"mean_tmb":1.03,"median_tmb":0.93},{"study":"kirc_tcga_pan_can_atlas_2018","n":483,"mean_aneuploidy":7.28,"median_aneuploidy":4,"mean_tmb":1.8,"median_tmb":1.7},{"study":"dlbc_tcga_pan_ca … (1115 more chars) ▶ clickhouse_run_select_query { "query": " WITH studies AS ( SELECT cancer_study_identifier FROM cancer_study_query_preferences WHERE preference_name = 'pan_cancer_tcga' ), per_sample AS ( SELECT sample_unique_id, cancer_study_identifier, maxIf(toFloat64OrNull(attribute_value), attribute_name = 'ANEUPLOIDY_SCORE') AS aneuploidy, maxIf(toFloat64OrNull(attribute_value), attribute_name = 'TMB_NONSYNONYMOUS') AS tmb FROM clinical_data_derived WHERE cancer_study_identifier IN (SELECT cancer_study_identifier FROM studies) AND attribute_name IN ('ANEUPLOIDY_SCORE','TMB_NONSYNONYMOUS') GROUP BY sample_unique_id, cancer_study_identifier ) SELECT count() AS n_samples, round(corr(aneuploidy, tmb), 3) AS pearson_r_sample_level, round(corr(aneuploidy, log(tmb+0.01)), 3) AS pearson_r_sample_level_log_tmb FROM per_sample WHERE aneuploidy IS NOT NULL AND tmb IS NOT NULL " } ◀ result {"rows":[{"n_samples":9975,"pearson_r_sample_level":-0.046,"pearson_r_sample_level_log_tmb":0.17}]} ▶ clickhouse_run_select_query { "query": " WITH studies AS ( SELECT cancer_study_identifier FROM cancer_study_query_preferences WHERE preference_name = 'pan_cancer_tcga' ), per_sample AS ( SELECT sample_unique_id, cancer_study_identifier, maxIf(toFloat64OrNull(attribute_value), attribute_name = 'ANEUPLOIDY_SCORE') AS aneuploidy, maxIf(toFloat64OrNull(attribute_value), attribute_name = 'TMB_NONSYNONYMOUS') AS tmb FROM clinical_data_derived WHERE cancer_study_identifier IN (SELECT cancer_study_identifier FROM studies) AND attribute_name IN ('ANEUPLOIDY_SCORE','TMB_NONSYNONYMOUS') GROUP BY sample_unique_id, cancer_study_identifier ), per_study AS ( SELECT cancer_study_identifier, avg(aneuploidy) AS mean_an, avg(tmb) AS mean_tmb FROM per_sample WHERE aneuploidy IS NOT NULL AND tmb IS NOT NULL GROUP BY cancer_study_identifier ) SELECT count() AS n_studies, round(corr(mean_an, mean_tmb),3) AS pearson_r_study_level FROM per_study " } ◀ result {"rows":[{"n_studies":32,"pearson_r_study_level":0.123}]} ▶ resolve_and_route { "studyIds": [ "acc_tcga_pan_can_atlas_2018", "blca_tcga_pan_can_atlas_2018", "brca_tcga_pan_can_atlas_2018", "cesc_tcga_pan_can_atlas_2018", "chol_tcga_pan_can_atlas_2018", "coadread_tcga_pan_can_atlas_2018", "dlbc_tcga_pan_can_atlas_2018", "esca_tcga_pan_can_atlas_2018", "gbm_tcga_pan_can_atlas_2018", "hnsc_tcga_pan_can_atlas_2018", "kich_tcga_pan_can_atlas_2018", "kirc_tcga_pan_can_atlas_2018", "kirp_tcga_pan_can_atlas_2018", "laml_tcga_pan_can_atlas_2018", "lgg_tcga_pan_can_atlas_2018", "lihc_tcga_pan_can_atlas_2018", "luad_tcga_pan_can_atlas_2018", "lusc_tcga_pan_can_atlas_2018", "meso_tcga_pan_can_atlas_2018", "ov_tcga_pan_can_atlas_2018", "paad_tcga_pan_can_atlas_2018", "pcpg_tcga_pan_can_atlas_2018", "prad_tcga_pan_can_atlas_2018", "sarc_tcga_pan_can_atlas_2018", "skcm_tcga_pan_can_atlas_2018", "stad_tcga_pan_can_atlas_2018", "tgct_tcga_pan_can_atlas_2018", "thca_tcga_pan_can_atlas_2018", "thym_tcga_pan_can_atlas_2018", "ucec_tcga_pan_can_atlas_2018", "ucs_tcga_pan_can_atlas_2018", "uvm_tcga_pan_can_atlas_2018" ] } ◀ result {"success":false,"error":"Invalid study ID(s): ucs_tcga_pan_can_atlas_2018","details":{"invalidIds":["ucs_tcga_pan_can_atlas_2018"],"providedIds":["acc_tcga_pan_can_atlas_2018","blca_tcga_pan_can_atlas_2018","brca_tcga_pan_can_atlas_2018","cesc_tcga_pan_can_atlas_2018","chol_tcga_pan_can_atlas_2018","coadread_tcga_pan_can_atlas_2018","dlbc_tcga_pan_can_atlas_2018","esca_tcga_pan_can_atlas_2018","gbm_tcga_pan_can_atlas_2018","hnsc_tcga_pan_can_atlas_2018","kich_tcga_pan_can_atlas_2018","kirc_tcga_pan_can_atlas_2018","kirp_tcga_pan_can_atlas_2018","laml_tcga_pan_can_atlas_2018","lgg_tcga_pan_can_atlas_2018","lihc_tcga_pan_can_atlas_2018","luad_tcga_pan_can_atlas_2018","lusc_tcga_pan_can_atlas_2018","meso_tcga_pan_can_atlas_2018","ov_tcga_pan_can_atlas_2018","paad_tcga_pan_can_atlas_2018","pcpg_tcga_pan_can_atlas_2018","prad_tcga_pan_can_atlas_2018","sarc_tcga_pan_can_atlas_2018","skcm_tcga_pan_can_atlas_2018","stad_tcga_pan_can_atlas_2018","tgct_tcga_pan_can_atlas_2018","thca_tcga_pan_can_atlas_2018","thym_tcga_pan_can_atlas_2018","ucec_tcga_pan_can_atlas_2018","ucs_tcga_pan_can_atlas_2018","uvm_tcga_pan_can_atlas_2018"]}} ▶ resolve_and_route { "studyIds": [ "acc_tcga_pan_can_atlas_2018", "blca_tcga_pan_can_atlas_2018", "brca_tcga_pan_can_atlas_2018", "cesc_tcga_pan_can_atlas_2018", "chol_tcga_pan_can_atlas_2018", "coadread_tcga_pan_can_atlas_2018", "dlbc_tcga_pan_can_atlas_2018", "esca_tcga_pan_can_atlas_2018", "gbm_tcga_pan_can_atlas_2018", "hnsc_tcga_pan_can_atlas_2018", "kich_tcga_pan_can_atlas_2018", "kirc_tcga_pan_can_atlas_2018", "kirp_tcga_pan_can_atlas_2018", "laml_tcga_pan_can_atlas_2018", "lgg_tcga_pan_can_atlas_2018", "lihc_tcga_pan_can_atlas_2018", "luad_tcga_pan_can_atlas_2018", "lusc_tcga_pan_can_atlas_2018", "meso_tcga_pan_can_atlas_2018", "ov_tcga_pan_can_atlas_2018", "paad_tcga_pan_can_atlas_2018", "pcpg_tcga_pan_can_atlas_2018", "prad_tcga_pan_can_atlas_2018", "sarc_tcga_pan_can_atlas_2018", "skcm_tcga_pan_can_atlas_2018", "stad_tcga_pan_can_atlas_2018", "tgct_tcga_pan_can_atlas_2018", "thca_tcga_pan_can_atlas_2018", "thym_tcga_pan_can_atlas_2018", "ucec_tcga_pan_can_atlas_2018", "uvm_tcga_pan_can_atlas_2018" ] } ◀ result {"success":false,"error":"Internal Server Error","details":{"status":500,"response":{"req":{"method":"POST","url":"https://www.cbioportal.org/api/molecular-profiles/fetch","data":{"studyIds":["esca_tcga_pan_can_atlas_2018"]},"headers":{"user-agent":"node-superagent/3.8.3","accept":"application/json","content-type":"application/json"}},"header":{"cache-control":"no-cache, no-store, max-age=0, must-revalidate","content-encoding":"gzip","content-type":"application/json","date":"Sat, 26 Sep 2026 16:47:21 GMT","elapsed-time":"0","expires":"0","pragma":"no-cache","strict-transport-security":"max-age=31536000 ; includeSubDomains","vary":"origin,access-control-request-method,access-control-request-headers,accept-encoding","x-content-type-options":"nosniff","x-frame-options":"DENY","x-xss-protection":"0","connection":"close","transfer-encoding":"chunked"},"status":500,"text":"{\"timestamp\":1790441241431,\"status\":500,\"error\":\"Internal Server Error\",\"path\":\"/api/molecular-profiles/fetch\"}"}}} ▶ resolve_and_route { "studyIds": [ "tgct_tcga_pan_can_atlas_2018", "acc_tcga_pan_can_atlas_2018", "ucec_tcga_pan_can_atlas_2018", "thca_tcga_pan_can_atlas_2018" ] } ◀ result {"success":true,"message":"Found 4 matching studies. Pick the best match (prefer TCGA → prefer PanCancer Atlas), use its metadata to call the appropriate navigation tool(s) and generate URLs immediately. Present other studies as alternatives.","data":{"totalCount":4,"studiesWithMetadata":[{"studyId":"tgct_tcga_pan_can_atlas_2018","name":"Testicular Germ Cell Tumors (TCGA, PanCancer Atlas)","sampleCount":149,"studyViewUrl":"https://www.cbioportal.org/study?id=tgct_tcga_pan_can_atlas_2018","metadata":{"clinicalAttributeIds":["AGE","AJCC_PATHOLOGIC_TUMOR_STAGE","AJCC_STAGING_EDITION","ANEUPLOIDY_SCORE","CANCER_TYPE","CANCER_TYPE_ACRONYM","CANCER_TYPE_DETAILED","DAYS_LAST_FOLLOWUP","DAYS_TO_BIRTH","DAYS_TO_INITIAL_PATHOLOGIC_DIAGNOSIS","DFS_MONTHS","DFS_STATUS","DSS_MONTHS","DSS_STATUS","ETHNICITY","FORM_COMPLETION_DATE","FRACTION_GENOME_ALTERED","GENETIC_ANCESTRY_LABEL","GRADE","HISTORY_NEOADJUVANT_TRTYN","ICD_10","ICD_O_3_HISTOLOGY","ICD_O_3_SITE","INFORMED_CONSENT_VERIFIED","IN_PANCANPATHWAYS_FREEZE","MSI_SCORE_MANTIS","MSI_SENSOR_SCORE","MUTATION_COUNT","NEW_TUMOR_EVENT_AFTER_INITIAL_TREATMENT","ONCOTREE_CODE","OS_MONTHS","OS_STATUS","OTHER_PATIENT_ID","PATH_M_STAGE","PATH_N_STAGE","PATH_T_STAGE","PERSON_NEOPLASM_CANCER_STATUS","PFS_MONTHS","PFS_STATUS","PRIMARY_LYMPH_NODE_PRESENTATION_ASSESSMENT","PRIOR_DX","RACE","RADIATION_THERAPY","SAMPLE_COUNT","SAMPLE_TYPE","SEX","SOMATIC_STATUS","SUBTYPE","TBL_SCORE","TISSUE_PROSPECTIVE_COLLECTION_INDICATOR","TISSUE_RETROSPECTIVE_COLLECTION_INDICATOR","TISSUE_SOURCE_SITE","TISSUE_SOURCE_SITE_CODE","TMB_NONSYNONYMOUS","TUMOR_TISSUE_SITE","TUMOR_TYPE","WEIGHT"],"molecularProfileIds":["tgct_tcga_pan_can_atlas_2018_gistic","tgct_tcga_pan_can_atlas_2018_log2CNA","tgct_tcga_pan_can_atlas_2018_mutations","tgct_tcga_pan_can_atlas_2018_rna_seq_v2_mrna","tgct_tcga_pan_can_atlas_2018_rna_seq_v2_mrna_median_Zscores","tgct_tcga_pan_can_atlas_2018_rna_seq_v2_mrna_median_all_sample_Zscores","tgct_tcga_pan_can_atlas_2018_rppa","tgct_tcga_pan_can_atlas_2018_rppa_Zscores","tgct_tcga_pan_can_atlas_2018_structural_variants"],"genericAssayProfiles":["tgct_tcga_pan_can_atlas_2018_armlevel_cna","tgct_tcga_pan_can_atlas_2018_genetic_ancestry","tgct_tcga_pan_can_atlas_2018_methylation_hm27_hm450_merge","tgct_tcga_pan_can_atlas_2018_methylation_hm450"],"heatmapProfileIds":["tgct_tcga_pan_can_atlas_2018_rna_seq_v2_mrna_median_Zscores","tgct_tcga_pan_can_atlas_2018_rna_seq_v2_mrna_median_all_sample_Zscores","tgct_tcga_pan_can_atlas_2018_rppa_Zscores"],"availableComparisonTabs":["overlap","clinical","survival","alterations","mutations","mrna","protein","generic_assay_armlevel_cna","generic_assay_genetic_ancestry","generic_assay_methylation"]}},{"studyId":"acc_tcga_pan_can_atlas_2018","name":"Adrenocortical Carcinoma (TCGA, PanCancer Atlas)","sampleCount":92,"studyViewUrl":"https://www.cbioportal.org/study?id=acc_tcga_pan_can_atlas_2018","metadata":{"clinicalAttributeIds":["AGE","AJCC_PATHOLOGIC_TUMOR_STAGE","AJCC_STAGING_EDITION","ANE … (7397 more chars) ▶ navigate_to_study_view { "studyIds": [ "tgct_tcga_pan_can_atlas_2018" ], "tab": "plots", "plotsHorzSelection": { "selectedDataSourceOption": "ANEUPLOIDY_SCORE", "dataType": "clinical_attribute" }, "plotsVertSelection": { "selectedDataSourceOption": "TMB_NONSYNONYMOUS", "dataType": "clinical_attribute" } } ◀ result {"success":true,"message":"Navigating to https://www.cbioportal.org/study/plots?id=tgct_tcga_pan_can_atlas_2018&plots_horz_selection=%7B%22dataType%22%3A%22clinical_attribute%22%2C%22selectedDataSourceOption%22%3A%22ANEUPLOIDY_SCORE%22%7D&plots_vert_selection=%7B%22dataType%22%3A%22clinical_attribute%22%2C%22selectedDataSourceOption%22%3A%22TMB_NONSYNONYMOUS%22%7D","url":"https://www.cbioportal.org/study/plots?id=tgct_tcga_pan_can_atlas_2018&plots_horz_selection=%7B%22dataType%22%3A%22clinical_attribute%22%2C%22selectedDataSourceOption%22%3A%22ANEUPLOIDY_SCORE%22%7D&plots_vert_selection=%7B%22dataType%22%3A%22clinical_attribute%22%2C%22selectedDataSourceOption%22%3A%22TMB_NONSYNONYMOUS%22%7D","data":{"studyIds":["tgct_tcga_pan_can_atlas_2018"],"studies":[{"studyId":"tgct_tcga_pan_can_atlas_2018","name":"Testicular Germ Cell Tumors (TCGA, PanCancer Atlas)","sampleCount":149}],"tab":"plots","hasFilters":false,"hasPlotsConfig":true,"pageDescription":"Scatter/box plots: tmb nonsynonymous (y-axis) by aneuploidy score (x-axis)."}} ▶ navigate_to_study_view { "studyIds": [ "thca_tcga_pan_can_atlas_2018" ] } ◀ result {"success":true,"message":"Navigating to https://www.cbioportal.org/study/summary?id=thca_tcga_pan_can_atlas_2018","url":"https://www.cbioportal.org/study/summary?id=thca_tcga_pan_can_atlas_2018","data":{"studyIds":["thca_tcga_pan_can_atlas_2018"],"studies":[{"studyId":"thca_tcga_pan_can_atlas_2018","name":"Thyroid Carcinoma (TCGA, PanCancer Atlas)","sampleCount":500}],"tab":"summary","hasFilters":false,"hasPlotsConfig":false,"pageDescription":"Overview of all study data displayed as interactive charts summarizing available clinical and genomic data. Click within a chart to filter. Additional charts may be available in the Charts menu."}} ═ answer (success) ## Aneuploidy scores across TCGA PanCancer Atlas studies (32 cohorts, tumor samples) **Highest aneuploidy** (mean SCNA-based Aneuploidy Score, scale 0–39): | Rank | Cancer type (study) | N samples | Mean score | Median score | |---|---|---|---|---| | 1 | Testicular Germ Cell Tumors (`tgct`) | 149 | **18.8** | 19 | | 2 | Adrenocortical Carcinoma (`acc`) | 89 | **18.3** | 17 | | 3 | Uterine Carcinosarcoma (`ucs`) | 56 | **17.2** | 19 | | 4 | Lung Squamous Cell Carcinoma (`lusc`) | 482 | 16.1 | 16 | | 5 | Lung Adenocarcinoma (`luad`) | 501 | 15.0 | 16 | | 6 | Chromophobe Kidney (`kich`) | 65 | 14.9 | 14 | | 7 | Esophageal Carcinoma (`esca`) | 162 | 14.6 | 14.5 | | 8 | Ovarian Serous Carcinoma (`ov`) | 552 | 14.0 | 13 | **Lowest aneuploidy:** Thyroid Carcinoma (`thca`, mean 0.9), Acute Myeloid Leukemia (`laml`, 1.6), Prostate Adenocarcinoma (`prad`, 2.5), Thymoma (`thym`, 3.8), Low-Grade Glioma (`lgg`, 3.8) — these are largely genomically "quiet," near-diploid tumor types. ## Correlation with mutation burden (TMB, mutations/Mb) I computed Pearson correlation directly in the database (a literal aggregate, not an inference): - **Sample-level, pooled across all 9,975 samples in the 32 studies:** r = **−0.046** (essentially no linear relationship; using log-transformed TMB to reduce skew from hypermutators: r = **0.17**, still weak). - **Study-level, across the 32 cancer-type means:** r = **0.12** (weak positive). **Interpretation of the pattern (descriptive, not a formal test):** aneuploidy and point-mutation burden are largely **decoupled** across TCGA cancer types. The highest-aneuploidy cohorts (TGCT, ACC) have *low* TMB (means 0.5 and 2.8 mutations/Mb), while the highest-TMB cohorts — SKCM (mean 23.9), UCEC (mean 35.6, driven by MSI/POLE hypermutators) — have only moderate aneuploidy (13.0 and 6.0). This suggests chromosomal instability (copy-number-driven genomic instability) and mutational-signature-driven point-mutation burden are largely independent axes of genomic instability in most solid tumors, rather than two readouts of the same underlying process. This is a descriptive pattern from the correlation coefficients above, not a hypothesis test — I have not run a formal significance test on this correlation. Because TMB is heavily right-skewed by hypermutator outliers (MSI-high UCEC/COAD, POLE-mutant tumors), the **mean** TMB and **median** TMB rank cohorts very differently (see full table below) — always check both.
Full 32-study table (mean/median aneuploidy and TMB) Available on request — all 32 PanCancer Atlas cohorts were queried; ranked list above shows the top and bottom of the aneuploidy distribution. TMB means/medians per study are also on file if you want the complete ranking.
--- ### Explore interactively - [TGCT: Aneuploidy Score vs. TMB scatter plot](https://www.cbioportal.org/study/plots?id=tgct_tcga_pan_can_atlas_2018&plots_horz_selection=%7B%22dataType%22%3A%22clinical_attribute%22%2C%22selectedDataSourceOption%22%3A%22ANEUPLOIDY_SCORE%22%7D&plots_vert_selection=%7B%22dataType%22%3A%22clinical_attribute%22%2C%22selectedDataSourceOption%22%3A%22TMB_NONSYNONYMOUS%22%7D) — the highest-aneuploidy cohort, showing genomically unstable but mutationally quiet tumors. - [Thyroid Carcinoma study overview](https://www.cbioportal.org/study/summary?id=thca_tcga_pan_can_atlas_2018) — the lowest-aneuploidy cohort, for contrast. Note: a 32-study batch request to the portal's cross-cohort plots view hit a transient server error, so I linked two representative single-cohort views instead of one combined multi-study plot. If you'd like, I can retry the full 32-cohort comparison or build a custom Group Comparison (e.g., top-tertile vs. bottom-tertile aneuploidy cancer types) on request.