Q (sonnet): How many samples are there that have any of these mutations in SEPHS1: p.Arg371Trp, p.Arg371Gln, p.Arg371Gly? ▶ read_guide { "uri": "cbioportal://mutation-frequency-guide" } ▶ clickhouse_list_table_columns { "table": "genetic_alteration_derived" } ◀ result # Mutation Frequency Analysis Guide ## IMPORTANT: Reporting Mutation Frequencies - **ALWAYS report frequencies as percentages**, not raw counts: `frequency = (altered_samples / total_profiled_samples) × 100` - For quick frequency lookups, **prefer the TCGA Pan-Cancer Atlas study first**, then offer to expand to other studies - When reporting across multiple studies, show **ranges** (e.g., "TP53 is mutated in 30–60% of samples") rather than a single average - **NEVER** sum mutation events across studies to compute an aggregate frequency — this can exceed 100% due to double-counting - Warn users that samples may overlap across cohorts (e.g., MSK studies may share patients) - **Choose and state the counting unit**: use patient-level frequencies for prevalence/rate questions unless the user explicitly asks for samples; use sample-level frequencies when the user asks about samples. - **For "across cancer types" questions**, jump to the [Cross-Cancer-Type Mutation Frequency](#cross-cancer-type-mutation-frequency) section below — there is one correct recipe and several common wrong ones. ## Counting Unit: Samples vs Patients Before answering any mutation count or frequency question, decide whether the unit is samples or patients and state that choice in the answer. | User wording | Counting unit | |--------------|---------------| | "prevalence", "rate", "fraction of patients", "patients with", "how common is" | Patient-level: `COUNT(DISTINCT patient_unique_id)` | | "samples", "specimens", "biopsies", sample-level cohort composition | Sample-level: `COUNT(DISTINCT sample_unique_id)` | | Ambiguous | Ask, or default to patient-level for prevalence/rate language and say so | ### Cross-study sample-count caveat When an answer touches more than one study and reports a sample count, prepend a one-line caveat: > Sample IDs are unique within cBioPortal study prefixes, not guaranteed biological-sample identifiers across studies; overlapping cohorts can count the same patient/sample more than once. Prefer one of these safer approaches: - Use a shipped `cancer_study_query_preferences` cohort such as `pan_cancer_tcga` or `all_studies_non_redundant`. - Restrict to one named study. - Aggregate by `patient_unique_id` when the biological question is patient prevalence. ## STOP rule: a frequency above 100% means your query is wrong If your query returns a frequency over 100%, **do not try to debug or explain the data inconsistency to the user**. The cause is always one of these query bugs: - Summing mutation events instead of `COUNT(DISTINCT sample_unique_id)` for the numerator - Using a study-wide sample count as the denominator instead of the gene-specific profiled count - Cross-study aggregation where the same biological sample appears under multiple `sample_unique_id` values (e.g., MSK-IMPACT and MSK-CHORD share patients) - **Joining the profiled CTE through `gene_panel` / `gene_panel_list` without a WES branch.** `gene_panel_id = 'WES'` is *not* a row in … (35166 more chars) ◀ result {"columns":[{"name":"sample_unique_id","type":"String"},{"name":"cancer_study_identifier","type":"LowCardinality(String)"},{"name":"hugo_gene_symbol","type":"String"},{"name":"profile_type","type":"LowCardinality(String)"},{"name":"alteration_value","type":"Nullable(String)"}]} ▶ clickhouse_run_select_query { "query": "SELECT DISTINCT mutation_variant, cancer_study_identifier FROM genomic_event_derived WHERE hugo_gene_symbol = 'SEPHS1' AND variant_type = 'mutation' ORDER BY mutation_variant" } ◀ result {"rows":[{"mutation_variant":"*251*","cancer_study_identifier":"normal_skin_keratinocytes_2024"},{"mutation_variant":"A103V","cancer_study_identifier":"coad_cptac_2019"},{"mutation_variant":"A103V","cancer_study_identifier":"coad_cptac_gdc"},{"mutation_variant":"A103V","cancer_study_identifier":"coadread_tcga_pan_can_atlas_2018"},{"mutation_variant":"A103V","cancer_study_identifier":"pan_pdo_wcm_2026"},{"mutation_variant":"A103V","cancer_study_identifier":"uec_cptac_gdc"},{"mutation_variant":"A151T","cancer_study_identifier":"ccle_broad_2019"},{"mutation_variant":"A151T","cancer_study_identifier":"ucec_tcga_pan_can_atlas_2018"},{"mutation_variant":"A152T","cancer_study_identifier":"ccle_broad_2019"},{"mutation_variant":"A152T","cancer_study_identifier":"ccle_broad_2025"},{"mutation_variant":"A155S","cancer_study_identifier":"luad_cptac_2020"},{"mutation_variant":"A155S","cancer_study_identifier":"luad_tcga_pan_can_atlas_2018"},{"mutation_variant":"A176P","cancer_study_identifier":"ccle_broad_2019"},{"mutation_variant":"A191V","cancer_study_identifier":"lihc_amc_prv"},{"mutation_variant":"A208T","cancer_study_identifier":"stad_tcga"},{"mutation_variant":"A208T","cancer_study_identifier":"stad_tcga_pan_can_atlas_2018"},{"mutation_variant":"A208T","cancer_study_identifier":"stad_tcga_pub"},{"mutation_variant":"A208T","cancer_study_identifier":"stes_tcga_pub"},{"mutation_variant":"A236T","cancer_study_identifier":"coadread_dfci_2016"},{"mutation_variant":"A240V","cancer_study_identifier":"crc_sysucc_2022"},{"mutation_variant":"A245V","cancer_study_identifier":"cesc_tcga_gdc"},{"mutation_variant":"A245V","cancer_study_identifier":"cesc_tcga_pan_can_atlas_2018"},{"mutation_variant":"A245V","cancer_study_identifier":"coadread_genentech"},{"mutation_variant":"A245V","cancer_study_identifier":"crc_sysucc_2022"},{"mutation_variant":"A245V","cancer_study_identifier":"difg_glass_2019"},{"mutation_variant":"A245V","cancer_study_identifier":"difg_glass"},{"mutation_variant":"A245V","cancer_study_identifier":"stad_tcga_pub"},{"mutation_variant":"A245V","cancer_study_identifier":"ucec_cptac_2020"},{"mutation_variant":"A245V","cancer_study_identifier":"uec_cptac_gdc"},{"mutation_variant":"A260T","cancer_study_identifier":"cscc_hgsc_bcm_2014"},{"mutation_variant":"A262S","cancer_study_identifier":"hcc_inserm_fr_2015"},{"mutation_variant":"A262T","cancer_study_identifier":"pancan_pdmr_2025"},{"mutation_variant":"A262T","cancer_study_identifier":"soft_tissue_tcga_gdc"},{"mutation_variant":"A262T","cancer_study_identifier":"npc_nusingapore"},{"mutation_variant":"A262T","cancer_study_identifier":"sarc_tcga"},{"mutation_variant":"A262T","cancer_study_identifier":"sarc_tcga_pan_can_atlas_2018"},{"mutation_variant":"A263T","cancer_study_identifier":"uec_cptac_gdc"},{"mutation_variant":"A275V","cancer_study_identifier":"ucec_tcga_gdc"},{"mutation_variant":"A275V","cancer_study_identifier":"ucec_tcga_pan_can_atlas_2018"},{"mutation_variant":"A275V","cancer_study_identifier … (4817 more chars) ▶ clickhouse_run_select_query { "query": "SELECT mutation_variant, cancer_study_identifier, COUNT(DISTINCT sample_unique_id) AS samples FROM genomic_event_derived WHERE hugo_gene_symbol = 'SEPHS1' AND variant_type = 'mutation' AND mutation_variant IN ('R371W', 'R371Q', 'R371G') GROUP BY mutation_variant, cancer_study_identifier ORDER BY mutation_variant, cancer_study_identifier" } ◀ result {"rows":[{"mutation_variant":"R371Q","cancer_study_identifier":"coad_cptac_gdc","samples":1},{"mutation_variant":"R371Q","cancer_study_identifier":"coad_tcga_gdc","samples":1},{"mutation_variant":"R371Q","cancer_study_identifier":"coadread_tcga","samples":2},{"mutation_variant":"R371Q","cancer_study_identifier":"coadread_tcga_pan_can_atlas_2018","samples":1},{"mutation_variant":"R371Q","cancer_study_identifier":"coadread_tcga_pub","samples":2},{"mutation_variant":"R371Q","cancer_study_identifier":"ohnca_cptac_gdc","samples":1},{"mutation_variant":"R371Q","cancer_study_identifier":"pancan_pcawg_2020","samples":1},{"mutation_variant":"R371W","cancer_study_identifier":"mel_dfci_2019","samples":1},{"mutation_variant":"R371W","cancer_study_identifier":"mel_iatlas_liu_2019","samples":1}]} ▶ clickhouse_run_select_query { "query": "SELECT COUNT(DISTINCT sample_unique_id) AS total_samples FROM genomic_event_derived WHERE hugo_gene_symbol = 'SEPHS1' AND variant_type = 'mutation' AND mutation_variant IN ('R371W', 'R371Q', 'R371G')" } ◀ result {"rows":[{"total_samples":11}]} ═ answer (success) No R371G (p.Arg371Gly) instances were found in the database — only R371Q and R371W occur. ## Result **11 samples** across the entire cBioPortal database carry any of the SEPHS1 p.Arg371Trp (R371W), p.Arg371Gln (R371Q), or p.Arg371Gly (R371G) mutations: | Variant | Study | Samples | |---|---|---| | R371Q | `coad_cptac_gdc` | 1 | | R371Q | `coad_tcga_gdc` | 1 | | R371Q | `coadread_tcga` | 2 | | R371Q | `coadread_tcga_pan_can_atlas_2018` | 1 | | R371Q | `coadread_tcga_pub` | 2 | | R371Q | `ohnca_cptac_gdc` | 1 | | R371Q | `pancan_pcawg_2020` | 1 | | R371W | `mel_dfci_2019` | 1 | | R371W | `mel_iatlas_liu_2019` | 1 | | R371G | — | 0 (not observed in any study) | **Important caveat:** These 11 rows span 9 different studies, several of which are overlapping cohorts of the same underlying TCGA colorectal data (`coad_tcga_gdc`, `coad_cptac_gdc`, `coadread_tcga`, `coadread_tcga_pub`, `coadread_tcga_pan_can_atlas_2018`) — the same physical patients/samples are very likely represented multiple times under different study-prefixed sample IDs. So 11 is a raw cross-study sample count, **not** 11 unique patients. If you want a non-redundant count, I'd recommend restricting to one TCGA colorectal cohort (e.g., the PanCancer Atlas version, which shows 1 sample) plus the two independent melanoma studies (2 samples) — giving 3 distinct samples across non-overlapping cohorts, plus the PCAWG and CPTAC/OHNCA cases if you want to include those too. Would you like me to consolidate to a de-duplicated patient-level count for a specific set of cohorts, or would you like a link to view these mutations directly in the portal (e.g., via the Mutations tab for a chosen study)?