Q (sonnet): What percentage of endometrial cancer samples have hypermutation (>5000 mutations) and how does this correlate with histological subtype? ▶ read_guide { "uri": "cbioportal://mutation-frequency-guide" } ▶ read_guide { "uri": "cbioportal://clinical-data-guide" } ▶ list_studies { "search": "endometrial" } ◀ result # Mutation Frequency Analysis Guide ## IMPORTANT: Reporting Mutation Frequencies - **ALWAYS report frequencies as percentages**, not raw counts: `frequency = (altered_samples / total_profiled_samples) × 100` - For quick frequency lookups, **prefer the TCGA Pan-Cancer Atlas study first**, then offer to expand to other studies - When reporting across multiple studies, show **ranges** (e.g., "TP53 is mutated in 30–60% of samples") rather than a single average - **NEVER** sum mutation events across studies to compute an aggregate frequency — this can exceed 100% due to double-counting - Warn users that samples may overlap across cohorts (e.g., MSK studies may share patients) - **Choose and state the counting unit**: use patient-level frequencies for prevalence/rate questions unless the user explicitly asks for samples; use sample-level frequencies when the user asks about samples. - **For "across cancer types" questions**, jump to the [Cross-Cancer-Type Mutation Frequency](#cross-cancer-type-mutation-frequency) section below — there is one correct recipe and several common wrong ones. ## Counting Unit: Samples vs Patients Before answering any mutation count or frequency question, decide whether the unit is samples or patients and state that choice in the answer. | User wording | Counting unit | |--------------|---------------| | "prevalence", "rate", "fraction of patients", "patients with", "how common is" | Patient-level: `COUNT(DISTINCT patient_unique_id)` | | "samples", "specimens", "biopsies", sample-level cohort composition | Sample-level: `COUNT(DISTINCT sample_unique_id)` | | Ambiguous | Ask, or default to patient-level for prevalence/rate language and say so | ### Cross-study sample-count caveat When an answer touches more than one study and reports a sample count, prepend a one-line caveat: > Sample IDs are unique within cBioPortal study prefixes, not guaranteed biological-sample identifiers across studies; overlapping cohorts can count the same patient/sample more than once. Prefer one of these safer approaches: - Use a shipped `cancer_study_query_preferences` cohort such as `pan_cancer_tcga` or `all_studies_non_redundant`. - Restrict to one named study. - Aggregate by `patient_unique_id` when the biological question is patient prevalence. ## STOP rule: a frequency above 100% means your query is wrong If your query returns a frequency over 100%, **do not try to debug or explain the data inconsistency to the user**. The cause is always one of these query bugs: - Summing mutation events instead of `COUNT(DISTINCT sample_unique_id)` for the numerator - Using a study-wide sample count as the denominator instead of the gene-specific profiled count - Cross-study aggregation where the same biological sample appears under multiple `sample_unique_id` values (e.g., MSK-IMPACT and MSK-CHORD share patients) - **Joining the profiled CTE through `gene_panel` / `gene_panel_list` without a WES branch.** `gene_panel_id = 'WES'` is *not* a row in … (35166 more chars) ◀ result # Clinical Data Query Guide ## Overview Clinical data in cBioPortal is stored at both patient and sample levels. Understanding the distinction is crucial for accurate analysis. ## Data Organization ### Patient-Level vs Sample-Level Data - **Patient-level**: Demographics, overall survival, disease stage (stored once per patient) - **Sample-level**: Sample type, sequencing platform, purity (can have multiple per patient) ### Key Tables - `clinical_patient`: Patient-level clinical attributes - `clinical_sample`: Sample-level clinical attributes - `clinical_data_derived`: Pre-joined view combining both levels - `clinical_attribute_meta`: Metadata about available clinical attributes ## Recommended Approach: Use clinical_data_derived The `clinical_data_derived` table is pre-joined and optimized for most queries: ```sql -- Get clinical data for specific attributes SELECT sample_unique_id, patient_unique_id, attribute_name, attribute_value FROM clinical_data_derived WHERE cancer_study_identifier = 'your_study_id' AND attribute_name IN ('CANCER_TYPE', 'SAMPLE_TYPE', 'TMB_NONSYNONYMOUS'); ``` ## Clinical Attribute Discovery ### Use clinical_attribute_meta for Discovering Available Attributes **Always start here** to see what clinical attributes are available for a specific study: ```sql -- Discover available clinical attributes for a study SELECT attr_id, description, patient_attribute, cancer_study_id FROM clinical_attribute_meta WHERE cancer_study_id = ( SELECT cancer_study_id FROM cancer_study WHERE cancer_study_identifier = 'your_study_id' ) ORDER BY patient_attribute, attr_id; ``` **Key Fields:** - **attr_id**: matches attr_id in clinical_sample/clinical_patient tables - **description**: human-readable description of the attribute - **patient_attribute**: true = patient attribute, false = sample attribute - **cancer_study_id**: links to cancer_study table (filter by study) ## Attribute Semantics and Matching ### Case-Insensitive Matching for Attribute Values Clinical values are free text across studies and may differ only by case. For example, a controlled-looking value such as germline mutation status may appear as `GERMLINE`, `Germline`, or another case variant. When filtering `clinical_data_derived.attribute_value`, use case-insensitive matching unless you have already profiled the exact values in the target study: ```sql -- Correct: case-insensitive clinical value filter SELECT DISTINCT sample_unique_id, patient_unique_id FROM clinical_data_derived WHERE cancer_study_identifier = 'your_study_id' AND attribute_name = 'MUTATION_STATUS' AND upper(attribute_value) = 'GERMLINE'; ``` Do not write `attribute_value = 'GERMLINE'` without first checking all distinct values for that attribute in the study. ### Query the Requested Attribute, Not a Proxy Do not infer one clinical attribute from a related subtype or marker. Query the actual requested attribute when it … (17073 more chars) ◀ result {"result":[{"cancer_study_identifier":"ucec_ancestry_cds_msk_2023","name":"Endometrial Cancer (MSK, Cancer Discovery 2023)","type_of_cancer_id":"ucec","sample_count":1882,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=ucec_ancestry_cds_msk_2023"},{"cancer_study_identifier":"ucec_tcga","name":"Uterine Corpus Endometrial Carcinoma (TCGA, Firehose Legacy)","type_of_cancer_id":"ucec","sample_count":549,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=ucec_tcga"},{"cancer_study_identifier":"ucec_tcga_gdc","name":"Endometrial Carcinoma (TCGA GDC, 2025)","type_of_cancer_id":"ucec","sample_count":547,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=ucec_tcga_gdc"},{"cancer_study_identifier":"ucec_tcga_pan_can_atlas_2018","name":"Uterine Corpus Endometrial Carcinoma (TCGA, PanCancer Atlas)","type_of_cancer_id":"ucec","sample_count":529,"has_guide":true,"url":"https://www.cbioportal.org/study/summary?id=ucec_tcga_pan_can_atlas_2018"},{"cancer_study_identifier":"ucec_tcga_pub","name":"Uterine Corpus Endometrial Carcinoma (TCGA, Nature 2013)","type_of_cancer_id":"ucec","sample_count":373,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=ucec_tcga_pub"},{"cancer_study_identifier":"ucec_msk_2018","name":"Endometrial Cancer (MSK, 2018)","type_of_cancer_id":"ucec","sample_count":197,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=ucec_msk_2018"},{"cancer_study_identifier":"ucec_ccr_msk_2022","name":"Endometrial Carcinoma MSI (MSK, Clin Cancer Res 2022)","type_of_cancer_id":"ucec","sample_count":181,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=ucec_ccr_msk_2022"},{"cancer_study_identifier":"ucec_cptac_2020","name":"Endometrial Carcinoma (CPTAC, Cell 2020)","type_of_cancer_id":"ucec","sample_count":81,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=ucec_cptac_2020"},{"cancer_study_identifier":"ucs_msk_2024","name":"Endometrial Carcinoma (MSK, Molecular Oncology 2024)","type_of_cancer_id":"ucec","sample_count":69,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=ucs_msk_2024"},{"cancer_study_identifier":"ucec_ccr_cfdna_msk_2022","name":"Endometrial Carcinoma cfDNA (MSK, Clin Cancer Res 2022)","type_of_cancer_id":"ucec","sample_count":44,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=ucec_ccr_cfdna_msk_2022"},{"cancer_study_identifier":"ucec_msk_2024","name":"Endometrial and Ovarian Cancer (MSK, Nature Medicine 2024)","type_of_cancer_id":"mixed","sample_count":33,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=ucec_msk_2024"},{"cancer_study_identifier":"uec_msk_2024","name":"Endometrial cancer HR-positive (MSK, 2024)","type_of_cancer_id":"ucec","sample_count":24,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=uec_msk_2024"},{"cancer_study_identifier":"uccc_nih_2017","name":"Uterine Clear Cell Carcinoma (NIH, Cancer 2017)","type_of_cancer_id … (447 more chars) ▶ get_study_guide { "study_id": "ucec_tcga_pan_can_atlas_2018" } ◀ result # Uterine Corpus Endometrial Carcinoma (TCGA, PanCancer Atlas) **Study ID:** `ucec_tcga_pan_can_atlas_2018` See `_tcga_pancan_template.md` for common TCGA clinical attributes. ## Study-Specific Attributes ### Molecular Classification (TCGA) | Attribute | Description | Values (patients) | |-----------|-------------|--------| | `SUBTYPE` | TCGA molecular subtype | `UCEC_CN_HIGH` 163, `UCEC_MSI` 148, `UCEC_CN_LOW` 147, `UCEC_POLE` 49, blank 22 | ### Histology There is no `HISTOLOGICAL_TYPE` attribute; use `CANCER_TYPE_DETAILED`, plus `GRADE` for grade. | Attribute | Values (samples) | |-----------|--------| | `CANCER_TYPE_DETAILED` | Uterine Endometrioid Carcinoma 399, Uterine Serous Carcinoma/Uterine Papillary Serous Carcinoma 109, Uterine Mixed Endometrial Carcinoma 21 | | `GRADE` | G3 302, G2 119, G1 97, High Grade 11 | `AJCC_PATHOLOGIC_TUMOR_STAGE` is blank for all samples. ## Molecular Subtypes | Subtype | Characteristics | Prognosis | |---------|-----------------|-----------| | **POLE** | Ultra-hypermutated (>100 mut/Mb), POLE exonuclease mutations | Excellent | | **MSI** | Hypermutated, microsatellite instability, MLH1 silencing | Intermediate | | **CN-low** | Microsatellite stable, few copy number alterations | Intermediate | | **CN-high** | Serous-like, TP53 mutations, extensive CNA | Poor | ## Key Genes | Gene | Frequency | Subtype Association | |------|-----------|---------------------| | PTEN | ~65% | CN-low, MSI | | PIK3CA | ~50% | All subtypes | | PIK3R1 | ~30% | CN-low | | ARID1A | ~35% | MSI | | TP53 | ~25% | CN-high (>90% in this subtype) | | KRAS | ~20% | Various | | CTNNB1 | ~20% | CN-low | | POLE | ~7% | Defines POLE subtype | ## Notes - TCGA molecular classification has prognostic value and guides treatment - POLE and MSI subtypes are hypermutated but have different mechanisms - TP53 mutations with CN-high pattern indicate serous-like behavior regardless of histology - POLE mutations should be in exonuclease domain (proofreading) to be driver - MSI-H tumors respond well to immunotherapy ▶ clickhouse_run_select_query { "query": "SELECT attr_id, description, patient_attribute FROM clinical_attribute_meta WHERE cancer_study_id = (SELECT cancer_study_id FROM cancer_study WHERE cancer_study_identifier = 'ucec_tcga_pan_can_atlas_2018') AND (upper(attr_id) LIKE '%HISTOLOG%' OR upper(attr_id) LIKE '%SUBTYPE%' OR upper(attr_id) LIKE '%MUTATION_COUNT%' OR upper(attr_id) LIKE '%TMB%')" } ◀ result {"rows":[{"attr_id":"ICD_O_3_HISTOLOGY","description":"The third edition of the International Classification of Diseases for Oncology, published in 2000, used principally in tumor and cancer registries for coding the site (topography) and the histology (morphology) of neoplasms._The study of the structure of the cells and their arrangement to constitute tissues and, finally, the association among these to form organs. In pathology, the microscopic process of identifying normal and abnormal morphologic characteristics in tissues, by employing various cytochemical and immunocytochemical stains._A system of numbered categories for representation of data.","patient_attribute":1},{"attr_id":"MUTATION_COUNT","description":"Mutation Count","patient_attribute":0},{"attr_id":"SUBTYPE","description":"Subtype","patient_attribute":1},{"attr_id":"TMB_NONSYNONYMOUS","description":"TMB (nonsynonymous)","patient_attribute":0}]} ▶ clickhouse_run_select_query { "query": " WITH sample_data AS ( SELECT sample_unique_id, maxIf(toFloat64OrNull(attribute_value), attribute_name = 'MUTATION_COUNT') AS mutation_count, maxIf(attribute_value, attribute_name = 'CANCER_TYPE_DETAILED') AS histology, maxIf(attribute_value, attribute_name = 'SUBTYPE') AS molecular_subtype FROM clinical_data_derived WHERE cancer_study_identifier = 'ucec_tcga_pan_can_atlas_2018' AND attribute_name IN ('MUTATION_COUNT', 'CANCER_TYPE_DETAILED', 'SUBTYPE') GROUP BY sample_unique_id ) SELECT 'ALL' AS histology, count() AS n_samples, countIf(mutation_count > 5000) AS hypermutated, round(countIf(mutation_count > 5000) * 100.0 / count(), 1) AS pct_hypermutated FROM sample_data WHERE mutation_count IS NOT NULL UNION ALL SELECT histology, count() AS n_samples, countIf(mutation_count > 5000) AS hypermutated, round(countIf(mutation_count > 5000) * 100.0 / count(), 1) AS pct_hypermutated FROM sample_data WHERE mutation_count IS NOT NULL AND histology != '' GROUP BY histology ORDER BY n_samples DESC " } ◀ result {"rows":[{"histology":"Uterine Endometrioid Carcinoma","n_samples":388,"hypermutated":34,"pct_hypermutated":8.8},{"histology":"Uterine Serous Carcinoma/Uterine Papillary Serous Carcinoma","n_samples":108,"hypermutated":2,"pct_hypermutated":1.9},{"histology":"Uterine Mixed Endometrial Carcinoma","n_samples":21,"hypermutated":2,"pct_hypermutated":9.5},{"histology":"ALL","n_samples":517,"hypermutated":38,"pct_hypermutated":7.4}]} ▶ clickhouse_run_select_query { "query": " WITH sample_data AS ( SELECT sample_unique_id, maxIf(toFloat64OrNull(attribute_value), attribute_name = 'MUTATION_COUNT') AS mutation_count, maxIf(attribute_value, attribute_name = 'SUBTYPE') AS molecular_subtype FROM clinical_data_derived WHERE cancer_study_identifier = 'ucec_tcga_pan_can_atlas_2018' AND attribute_name IN ('MUTATION_COUNT', 'SUBTYPE') GROUP BY sample_unique_id ) SELECT molecular_subtype, count() AS n_samples, countIf(mutation_count > 5000) AS hypermutated, round(countIf(mutation_count > 5000) * 100.0 / count(), 1) AS pct_hypermutated, round(avg(mutation_count),0) AS mean_mutation_count, round(quantile(0.5)(mutation_count),0) AS median_mutation_count FROM sample_data WHERE mutation_count IS NOT NULL AND molecular_subtype != '' GROUP BY molecular_subtype ORDER BY n_samples DESC " } ◀ result {"rows":[]} ▶ clickhouse_run_select_query { "query": "SELECT attribute_value, count() FROM clinical_data_derived WHERE cancer_study_identifier = 'ucec_tcga_pan_can_atlas_2018' AND attribute_name = 'SUBTYPE' GROUP BY attribute_value" } ▶ clickhouse_list_table_columns { "table": "clinical_data_derived" } ◀ result {"rows":[{"count()":22},{"attribute_value":"UCEC_MSI","count()":148},{"attribute_value":"UCEC_POLE","count()":49},{"attribute_value":"UCEC_CN_HIGH","count()":163},{"attribute_value":"UCEC_CN_LOW","count()":147}]} ◀ result {"columns":[{"name":"internal_id","type":"Int32"},{"name":"sample_unique_id","type":"String","comment":"Globally unique sample ID: cancer_study_identifier + \"_\" + sample.stable_id. Empty for patient-level attributes. Use this for sample filtering and joins."},{"name":"patient_unique_id","type":"String","comment":"Globally unique patient ID: cancer_study_identifier + \"_\" + patient.stable_id. Present for both sample and patient-level attributes."},{"name":"attribute_name","type":"LowCardinality(String)","comment":"Clinical attribute name (e.g., SAMPLE_TYPE, CANCER_TYPE, AGE, OS_MONTHS). Use with attribute_value for filtering. AGE may be floored or capped for de-identification (e.g. all children recorded as 18, or everyone 89+ recorded as 89 or 90): before age statistics check for a pile-up at the min/max, and if present compute age from DAYS_TO_BIRTH (-days / 365.25)."},{"name":"attribute_value","type":"String","comment":"Value of the clinical attribute (String). For SAMPLE_TYPE: Primary, Metastasis, Local Recurrence, Unknown. Missing values are empty strings, so use toFloat64OrNull(attribute_value) for numeric comparisons — CAST fails on them."},{"name":"cancer_study_identifier","type":"LowCardinality(String)"},{"name":"type","type":"LowCardinality(String)","comment":"Data level: \"sample\" for sample-level attributes (e.g., SAMPLE_TYPE), \"patient\" for patient-level attributes (e.g., AGE, OS_MONTHS)."}]} ▶ clickhouse_run_select_query { "query": " WITH mut AS ( SELECT sample_unique_id, patient_unique_id, toFloat64OrNull(attribute_value) AS mutation_count FROM clinical_data_derived WHERE cancer_study_identifier = 'ucec_tcga_pan_can_atlas_2018' AND attribute_name = 'MUTATION_COUNT' ), subtype AS ( SELECT patient_unique_id, attribute_value AS molecular_subtype FROM clinical_data_derived WHERE cancer_study_identifier = 'ucec_tcga_pan_can_atlas_2018' AND attribute_name = 'SUBTYPE' AND attribute_value != '' ) SELECT s.molecular_subtype, count() AS n_samples, countIf(m.mutation_count > 5000) AS hypermutated, round(countIf(m.mutation_count > 5000) * 100.0 / count(), 1) AS pct_hypermutated, round(avg(m.mutation_count),0) AS mean_mutation_count, round(quantile(0.5)(m.mutation_count),0) AS median_mutation_count FROM mut m JOIN subtype s USING (patient_unique_id) WHERE m.mutation_count IS NOT NULL GROUP BY s.molecular_subtype ORDER BY n_samples DESC " } ◀ result {"rows":[{"molecular_subtype":"UCEC_CN_HIGH","n_samples":163,"hypermutated":0,"pct_hypermutated":0,"mean_mutation_count":71,"median_mutation_count":54},{"molecular_subtype":"UCEC_MSI","n_samples":148,"hypermutated":6,"pct_hypermutated":4.1,"mean_mutation_count":1020,"median_mutation_count":578},{"molecular_subtype":"UCEC_CN_LOW","n_samples":147,"hypermutated":0,"pct_hypermutated":0,"mean_mutation_count":77,"median_mutation_count":50},{"molecular_subtype":"UCEC_POLE","n_samples":49,"hypermutated":32,"pct_hypermutated":65.3,"mean_mutation_count":7427,"median_mutation_count":7385}]} ▶ resolve_and_route { "studyIds": [ "ucec_tcga_pan_can_atlas_2018" ] } ◀ result {"success":true,"message":"Found 1 study. Use the metadata to call the appropriate navigation tool(s).","data":{"totalCount":1,"studiesWithMetadata":[{"studyId":"ucec_tcga_pan_can_atlas_2018","name":"Uterine Corpus Endometrial Carcinoma (TCGA, PanCancer Atlas)","sampleCount":529,"studyViewUrl":"https://www.cbioportal.org/study?id=ucec_tcga_pan_can_atlas_2018","metadata":{"clinicalAttributeIds":["AGE","AJCC_PATHOLOGIC_TUMOR_STAGE","AJCC_STAGING_EDITION","ANEUPLOIDY_SCORE","BUFFA_HYPOXIA_SCORE","CANCER_TYPE","CANCER_TYPE_ACRONYM","CANCER_TYPE_DETAILED","DAYS_LAST_FOLLOWUP","DAYS_TO_BIRTH","DAYS_TO_INITIAL_PATHOLOGIC_DIAGNOSIS","DFS_MONTHS","DFS_STATUS","DSS_MONTHS","DSS_STATUS","ETHNICITY","FORM_COMPLETION_DATE","FRACTION_GENOME_ALTERED","GENETIC_ANCESTRY_LABEL","GRADE","HISTORY_NEOADJUVANT_TRTYN","ICD_10","ICD_O_3_HISTOLOGY","ICD_O_3_SITE","INFORMED_CONSENT_VERIFIED","IN_PANCANPATHWAYS_FREEZE","MSI_SCORE_MANTIS","MSI_SENSOR_SCORE","MUTATION_COUNT","NEW_TUMOR_EVENT_AFTER_INITIAL_TREATMENT","ONCOTREE_CODE","OS_MONTHS","OS_STATUS","OTHER_PATIENT_ID","PATH_M_STAGE","PATH_N_STAGE","PATH_T_STAGE","PERSON_NEOPLASM_CANCER_STATUS","PFS_MONTHS","PFS_STATUS","PRIMARY_LYMPH_NODE_PRESENTATION_ASSESSMENT","PRIOR_DX","RACE","RADIATION_THERAPY","RAGNUM_HYPOXIA_SCORE","SAMPLE_COUNT","SAMPLE_TYPE","SEX","SOMATIC_STATUS","SUBTYPE","TBL_SCORE","TISSUE_PROSPECTIVE_COLLECTION_INDICATOR","TISSUE_RETROSPECTIVE_COLLECTION_INDICATOR","TISSUE_SOURCE_SITE","TISSUE_SOURCE_SITE_CODE","TMB_NONSYNONYMOUS","TUMOR_TISSUE_SITE","TUMOR_TYPE","WEIGHT","WINTER_HYPOXIA_SCORE"],"molecularProfileIds":["ucec_tcga_pan_can_atlas_2018_gistic","ucec_tcga_pan_can_atlas_2018_log2CNA","ucec_tcga_pan_can_atlas_2018_mutations","ucec_tcga_pan_can_atlas_2018_rna_seq_v2_mrna","ucec_tcga_pan_can_atlas_2018_rna_seq_v2_mrna_median_Zscores","ucec_tcga_pan_can_atlas_2018_rna_seq_v2_mrna_median_all_sample_Zscores","ucec_tcga_pan_can_atlas_2018_rna_seq_v2_mrna_median_all_sample_ref_normal_Zscores","ucec_tcga_pan_can_atlas_2018_rppa","ucec_tcga_pan_can_atlas_2018_rppa_Zscores","ucec_tcga_pan_can_atlas_2018_structural_variants"],"genericAssayProfiles":["ucec_tcga_pan_can_atlas_2018_armlevel_cna","ucec_tcga_pan_can_atlas_2018_genetic_ancestry","ucec_tcga_pan_can_atlas_2018_methylation_hm27_hm450_merge","ucec_tcga_pan_can_atlas_2018_methylation_hm450"],"heatmapProfileIds":["ucec_tcga_pan_can_atlas_2018_rna_seq_v2_mrna_median_Zscores","ucec_tcga_pan_can_atlas_2018_rna_seq_v2_mrna_median_all_sample_Zscores","ucec_tcga_pan_can_atlas_2018_rna_seq_v2_mrna_median_all_sample_ref_normal_Zscores","ucec_tcga_pan_can_atlas_2018_rppa_Zscores"],"availableComparisonTabs":["overlap","clinical","survival","alterations","mutations","mrna","protein","generic_assay_armlevel_cna","generic_assay_genetic_ancestry","generic_assay_methylation"]}}]}} ▶ get_studyviewfilter_options { "studyId": "ucec_tcga_pan_can_atlas_2018", "attributeIds": [ "SUBTYPE", "CANCER_TYPE_DETAILED" ] } ◀ result {"success":true,"message":"Successfully retrieved StudyView filter options","data":{"studyId":"ucec_tcga_pan_can_atlas_2018","attributes":[{"attributeId":"CANCER_TYPE_DETAILED","displayName":"Cancer Type Detailed","description":"Cancer Type Detailed","datatype":"STRING","values":["Uterine Serous Carcinoma/Uterine Papillary Serous Carcinoma","Uterine Mixed Endometrial Carcinoma","Uterine Endometrioid Carcinoma"]},{"attributeId":"SUBTYPE","displayName":"Subtype","description":"Subtype","datatype":"STRING","values":["UCEC_MSI","UCEC_CN_HIGH","UCEC_POLE","NA","UCEC_CN_LOW"]}]}} ▶ navigate_to_study_view { "studyIds": [ "ucec_tcga_pan_can_atlas_2018" ], "filterJson": { "clinicalDataFilters": [ { "attributeId": "MUTATION_COUNT", "values": [ { "start": 5000 } ] } ] } } ◀ result {"success":true,"message":"Navigating to https://www.cbioportal.org/study/summary?id=ucec_tcga_pan_can_atlas_2018#filterJson=%7B%22clinicalDataFilters%22%3A%5B%7B%22attributeId%22%3A%22MUTATION_COUNT%22%2C%22values%22%3A%5B%7B%22start%22%3A5000%7D%5D%7D%5D%7D","url":"https://www.cbioportal.org/study/summary?id=ucec_tcga_pan_can_atlas_2018#filterJson=%7B%22clinicalDataFilters%22%3A%5B%7B%22attributeId%22%3A%22MUTATION_COUNT%22%2C%22values%22%3A%5B%7B%22start%22%3A5000%7D%5D%7D%5D%7D","data":{"studyIds":["ucec_tcga_pan_can_atlas_2018"],"studies":[{"studyId":"ucec_tcga_pan_can_atlas_2018","name":"Uterine Corpus Endometrial Carcinoma (TCGA, PanCancer Atlas)","sampleCount":529}],"tab":"summary","hasFilters":true,"hasPlotsConfig":false,"pageDescription":"Overview of all study data displayed as interactive charts summarizing available clinical and genomic data. Click within a chart to filter. Additional charts may be available in the Charts menu."}} ▶ navigate_to_group_comparison { "studyIds": [ "ucec_tcga_pan_can_atlas_2018" ], "groups": [ { "name": "POLE", "studyViewFilter": { "clinicalDataFilters": [ { "attributeId": "SUBTYPE", "values": [ { "value": "UCEC_POLE" } ] } ] } }, { "name": "MSI", "studyViewFilter": { "clinicalDataFilters": [ { "attributeId": "SUBTYPE", "values": [ { "value": "UCEC_MSI" } ] } ] } }, { "name": "CN-High", "studyViewFilter": { "clinicalDataFilters": [ { "attributeId": "SUBTYPE", "values": [ { "value": "UCEC_CN_HIGH" } ] } ] } }, { "name": "CN-Low", "studyViewFilter": { "clinicalDataFilters": [ { "attributeId": "SUBTYPE", "values": [ { "value": "UCEC_CN_LOW" } ] } ] } } ], "tab": "clinical" } ◀ result {"success":true,"message":"Navigating to https://www.cbioportal.org/comparison/clinical?comparisonId=6ab4a04cc2115c492d884a68","url":"https://www.cbioportal.org/comparison/clinical?comparisonId=6ab4a04cc2115c492d884a68","data":{"description":"Group comparison (4 custom groups)","studies":["ucec_tcga_pan_can_atlas_2018"],"totalGroups":4,"groups":[{"name":"POLE","sampleCount":49},{"name":"MSI","sampleCount":148},{"name":"CN-High","sampleCount":163},{"name":"CN-Low","sampleCount":147}],"studyViewUrl":"https://www.cbioportal.org/study?id=ucec_tcga_pan_can_atlas_2018","groupUrls":[{"groupName":"POLE","url":"https://www.cbioportal.org/study?id=ucec_tcga_pan_can_atlas_2018#filterJson=%7B%22clinicalDataFilters%22%3A%5B%7B%22attributeId%22%3A%22SUBTYPE%22%2C%22values%22%3A%5B%7B%22value%22%3A%22UCEC_POLE%22%7D%5D%7D%5D%7D"},{"groupName":"MSI","url":"https://www.cbioportal.org/study?id=ucec_tcga_pan_can_atlas_2018#filterJson=%7B%22clinicalDataFilters%22%3A%5B%7B%22attributeId%22%3A%22SUBTYPE%22%2C%22values%22%3A%5B%7B%22value%22%3A%22UCEC_MSI%22%7D%5D%7D%5D%7D"},{"groupName":"CN-High","url":"https://www.cbioportal.org/study?id=ucec_tcga_pan_can_atlas_2018#filterJson=%7B%22clinicalDataFilters%22%3A%5B%7B%22attributeId%22%3A%22SUBTYPE%22%2C%22values%22%3A%5B%7B%22value%22%3A%22UCEC_CN_HIGH%22%7D%5D%7D%5D%7D"},{"groupName":"CN-Low","url":"https://www.cbioportal.org/study?id=ucec_tcga_pan_can_atlas_2018#filterJson=%7B%22clinicalDataFilters%22%3A%5B%7B%22attributeId%22%3A%22SUBTYPE%22%2C%22values%22%3A%5B%7B%22value%22%3A%22UCEC_CN_LOW%22%7D%5D%7D%5D%7D"}],"pageDescription":"Clinical attribute distributions compared across POLE (49) vs MSI (148) vs CN-High (163) vs CN-Low (147)."}} ═ answer (success) ## Endometrial Cancer Hypermutation (>5,000 mutations) **Cohort:** [Uterine Corpus Endometrial Carcinoma (TCGA, PanCancer Atlas)](https://www.cbioportal.org/study/summary?id=ucec_tcga_pan_can_atlas_2018) — 529 samples. Note there are several other endometrial cohorts in cBioPortal (e.g. MSK 2018, CPTAC 2020); let me know if you'd like those included too. Counting unit: samples (mutation count is a sample-level attribute). ### Overall | | N (profiled) | Hypermutated (>5,000 mut.) | % | |---|---|---|---| | **All samples** | 517 | 38 | **7.4%** | ### By histological subtype (`CANCER_TYPE_DETAILED`) | Histology | N | Hypermutated | % | |---|---|---|---| | Uterine Endometrioid Carcinoma | 388 | 34 | 8.8% | | Uterine Serous Carcinoma/Papillary Serous Carcinoma | 108 | 2 | 1.9% | | Uterine Mixed Endometrial Carcinoma | 21 | 2 | 9.5% | ### By TCGA molecular subtype (`SUBTYPE`) — the stronger correlate | Subtype | N | Hypermutated | % | Mean mut. count | Median mut. count | |---|---|---|---|---|---| | **POLE** (ultramutated) | 49 | 32 | **65.3%** | 7,427 | 7,385 | | **MSI** (hypermutated) | 148 | 6 | 4.1% | 1,020 | 578 | | CN-high (serous-like) | 163 | 0 | 0% | 71 | 54 | | CN-low | 147 | 0 | 0% | 77 | 50 | **Interpretation:** Hypermutation in endometrial cancer is driven almost entirely by molecular subtype rather than histology per se. The POLE-ultramutated subtype accounts for nearly all samples exceeding 5,000 mutations (65% of POLE-subtype samples cross this threshold), consistent with POLE exonuclease-domain mutations causing extreme mutator phenotypes. MSI-high tumors are moderately mutated (mean ~1,000) but rarely exceed 5,000. Endometrioid histology is enriched for hypermutation mainly because it's the histology in which POLE and MSI tumors predominantly arise — serous carcinomas are almost never hypermutated by this threshold, consistent with their CN-high/TP53-driven biology. This descriptive breakdown is not a statistical test; a Kruskal-Wallis or Fisher's exact test would be needed to formally establish subtype/histology association with hypermutation status. ### Explore interactively - [Study View filtered to hypermutated samples (>5,000 mutations)](https://www.cbioportal.org/study/summary?id=ucec_tcga_pan_can_atlas_2018#filterJson=%7B%22clinicalDataFilters%22%3A%5B%7B%22attributeId%22%3A%22MUTATION_COUNT%22%2C%22values%22%3A%5B%7B%22start%22%3A5000%7D%5D%7D%5D%7D) — cross-tabulate with histology/grade/survival using the chart panel. - [Group Comparison: POLE vs MSI vs CN-High vs CN-Low](https://www.cbioportal.org/comparison/clinical?comparisonId=6ab4a04cc2115c492d884a68) — clinical/mutation distribution comparison across the four TCGA molecular subtypes (switch to the Mutations or Survival tab for further analysis).