Q (sonnet): "What are the top 5 frequently altered genes in the Osteosarcoma study from TARGET for mutations, copy numbers and SVs combined?" ▶ list_studies { "search": "Osteosarcoma TARGET" } ▶ read_guide { "uri": "cbioportal://mutation-frequency-guide" } ◀ result {"result":[{"cancer_study_identifier":"os_target_gdc","name":"Osteosarcoma (TARGET GDC, 2025)","type_of_cancer_id":"os","sample_count":159,"has_guide":true,"url":"https://www.cbioportal.org/study/summary?id=os_target_gdc"}]} ◀ result # Mutation Frequency Analysis Guide ## IMPORTANT: Reporting Mutation Frequencies - **ALWAYS report frequencies as percentages**, not raw counts: `frequency = (altered_samples / total_profiled_samples) × 100` - For quick frequency lookups, **prefer the TCGA Pan-Cancer Atlas study first**, then offer to expand to other studies - When reporting across multiple studies, show **ranges** (e.g., "TP53 is mutated in 30–60% of samples") rather than a single average - **NEVER** sum mutation events across studies to compute an aggregate frequency — this can exceed 100% due to double-counting - Warn users that samples may overlap across cohorts (e.g., MSK studies may share patients) - **Choose and state the counting unit**: use patient-level frequencies for prevalence/rate questions unless the user explicitly asks for samples; use sample-level frequencies when the user asks about samples. - **For "across cancer types" questions**, jump to the [Cross-Cancer-Type Mutation Frequency](#cross-cancer-type-mutation-frequency) section below — there is one correct recipe and several common wrong ones. ## Counting Unit: Samples vs Patients Before answering any mutation count or frequency question, decide whether the unit is samples or patients and state that choice in the answer. | User wording | Counting unit | |--------------|---------------| | "prevalence", "rate", "fraction of patients", "patients with", "how common is" | Patient-level: `COUNT(DISTINCT patient_unique_id)` | | "samples", "specimens", "biopsies", sample-level cohort composition | Sample-level: `COUNT(DISTINCT sample_unique_id)` | | Ambiguous | Ask, or default to patient-level for prevalence/rate language and say so | ### Cross-study sample-count caveat When an answer touches more than one study and reports a sample count, prepend a one-line caveat: > Sample IDs are unique within cBioPortal study prefixes, not guaranteed biological-sample identifiers across studies; overlapping cohorts can count the same patient/sample more than once. Prefer one of these safer approaches: - Use a shipped `cancer_study_query_preferences` cohort such as `pan_cancer_tcga` or `all_studies_non_redundant`. - Restrict to one named study. - Aggregate by `patient_unique_id` when the biological question is patient prevalence. ## STOP rule: a frequency above 100% means your query is wrong If your query returns a frequency over 100%, **do not try to debug or explain the data inconsistency to the user**. The cause is always one of these query bugs: - Summing mutation events instead of `COUNT(DISTINCT sample_unique_id)` for the numerator - Using a study-wide sample count as the denominator instead of the gene-specific profiled count - Cross-study aggregation where the same biological sample appears under multiple `sample_unique_id` values (e.g., MSK-IMPACT and MSK-CHORD share patients) - **Joining the profiled CTE through `gene_panel` / `gene_panel_list` without a WES branch.** `gene_panel_id = 'WES'` is *not* a row in … (35166 more chars) ▶ clickhouse_run_select_query { "query": "SELECT hugo_gene_symbol, total_mutation_events AS mutations, altered_samples, profiled_samples, frequency_pct FROM top_mutated_genes_in_study(study = 'os_target_gdc', top_n = 20)" } ▶ clickhouse_run_select_query { "query": "SELECT * FROM top_cna_genes_in_study(study = 'os_target_gdc', top_n = 20)" } ◀ result {"rows":[{"hugo_gene_symbol":"TP53","mutations":33,"altered_samples":32,"profiled_samples":143,"frequency_pct":22.4},{"hugo_gene_symbol":"MUC16","mutations":17,"altered_samples":16,"profiled_samples":143,"frequency_pct":11.2},{"hugo_gene_symbol":"TTN","mutations":18,"altered_samples":16,"profiled_samples":143,"frequency_pct":11.2},{"hugo_gene_symbol":"ATRX","mutations":11,"altered_samples":11,"profiled_samples":143,"frequency_pct":7.7},{"hugo_gene_symbol":"DNAH9","mutations":10,"altered_samples":10,"profiled_samples":143,"frequency_pct":7},{"hugo_gene_symbol":"RB1","mutations":8,"altered_samples":8,"profiled_samples":143,"frequency_pct":5.6},{"hugo_gene_symbol":"CNTNAP5","mutations":7,"altered_samples":7,"profiled_samples":143,"frequency_pct":4.9},{"hugo_gene_symbol":"DMD","mutations":7,"altered_samples":7,"profiled_samples":143,"frequency_pct":4.9},{"hugo_gene_symbol":"MUC4","mutations":11,"altered_samples":7,"profiled_samples":143,"frequency_pct":4.9},{"hugo_gene_symbol":"PCLO","mutations":7,"altered_samples":7,"profiled_samples":143,"frequency_pct":4.9},{"hugo_gene_symbol":"CSMD3","mutations":7,"altered_samples":6,"profiled_samples":143,"frequency_pct":4.2},{"hugo_gene_symbol":"ADGRV1","mutations":5,"altered_samples":5,"profiled_samples":143,"frequency_pct":3.5},{"hugo_gene_symbol":"DNAH3","mutations":5,"altered_samples":5,"profiled_samples":143,"frequency_pct":3.5},{"hugo_gene_symbol":"RYR2","mutations":6,"altered_samples":5,"profiled_samples":143,"frequency_pct":3.5},{"hugo_gene_symbol":"ALMS1","mutations":4,"altered_samples":4,"profiled_samples":143,"frequency_pct":2.8},{"hugo_gene_symbol":"CMYA5","mutations":4,"altered_samples":4,"profiled_samples":143,"frequency_pct":2.8},{"hugo_gene_symbol":"CSMD1","mutations":4,"altered_samples":4,"profiled_samples":143,"frequency_pct":2.8},{"hugo_gene_symbol":"CSMD2","mutations":4,"altered_samples":4,"profiled_samples":143,"frequency_pct":2.8},{"hugo_gene_symbol":"DNAI4","mutations":4,"altered_samples":4,"profiled_samples":143,"frequency_pct":2.8},{"hugo_gene_symbol":"GRM5","mutations":4,"altered_samples":4,"profiled_samples":143,"frequency_pct":2.8}]} ▶ clickhouse_run_select_query { "query": "SELECT * FROM top_sv_genes_in_study(study = 'os_target_gdc', top_n = 20)" } ◀ result {"rows":[{"hugo_gene_symbol":"LINC00901","cytoband":"3q13.31","cna_type":"HOMDEL","altered_samples":32,"profiled_samples":81,"frequency_pct":39.5},{"hugo_gene_symbol":"RN7SL442P","cytoband":"17p11.2","cna_type":"AMP","altered_samples":31,"profiled_samples":81,"frequency_pct":38.3},{"hugo_gene_symbol":"ALKBH5","cytoband":"17p11.2","cna_type":"AMP","altered_samples":30,"profiled_samples":81,"frequency_pct":37},{"hugo_gene_symbol":"ATPAF2","cytoband":"17p11.2","cna_type":"AMP","altered_samples":30,"profiled_samples":81,"frequency_pct":37},{"hugo_gene_symbol":"BRI3P3","cytoband":"17p11.2","cna_type":"AMP","altered_samples":30,"profiled_samples":81,"frequency_pct":37},{"hugo_gene_symbol":"CARS1P2","cytoband":"8q23.3","cna_type":"AMP","altered_samples":30,"profiled_samples":81,"frequency_pct":37},{"hugo_gene_symbol":"DRC3","cytoband":"17p11.2","cna_type":"AMP","altered_samples":30,"profiled_samples":81,"frequency_pct":37},{"hugo_gene_symbol":"DRG2","cytoband":"17p11.2","cna_type":"AMP","altered_samples":30,"profiled_samples":81,"frequency_pct":37},{"hugo_gene_symbol":"FLII","cytoband":"17p11.2","cna_type":"AMP","altered_samples":30,"profiled_samples":81,"frequency_pct":37},{"hugo_gene_symbol":"GID4","cytoband":"17p11.2","cna_type":"AMP","altered_samples":30,"profiled_samples":81,"frequency_pct":37},{"hugo_gene_symbol":"LLGL1","cytoband":"17p11.2","cna_type":"AMP","altered_samples":30,"profiled_samples":81,"frequency_pct":37},{"hugo_gene_symbol":"MIEF2","cytoband":"17p11.2","cna_type":"AMP","altered_samples":30,"profiled_samples":81,"frequency_pct":37},{"hugo_gene_symbol":"MYO15A","cytoband":"17p11.2","cna_type":"AMP","altered_samples":30,"profiled_samples":81,"frequency_pct":37},{"hugo_gene_symbol":"NCOR1","cytoband":"17p12-p11.2","cna_type":"AMP","altered_samples":30,"profiled_samples":81,"frequency_pct":37},{"hugo_gene_symbol":"RNU5E-8P","cytoband":"3q13.31","cna_type":"HOMDEL","altered_samples":30,"profiled_samples":81,"frequency_pct":37},{"hugo_gene_symbol":"RNU6-314P","cytoband":"17p11.2","cna_type":"AMP","altered_samples":30,"profiled_samples":81,"frequency_pct":37},{"hugo_gene_symbol":"RNU6-862P","cytoband":"17p11.2","cna_type":"AMP","altered_samples":30,"profiled_samples":81,"frequency_pct":37},{"hugo_gene_symbol":"RPL22P21","cytoband":"17p12","cna_type":"AMP","altered_samples":30,"profiled_samples":81,"frequency_pct":37},{"hugo_gene_symbol":"RPLP1P11","cytoband":"17p12","cna_type":"AMP","altered_samples":30,"profiled_samples":81,"frequency_pct":37},{"hugo_gene_symbol":"TRPS1-AS1","cytoband":"8q23.3","cna_type":"AMP","altered_samples":30,"profiled_samples":81,"frequency_pct":37}]} ◀ result {"rows":[]} ▶ clickhouse_run_select_query { "query": " WITH mut AS ( SELECT hugo_gene_symbol, COUNT(DISTINCT sample_unique_id) AS mut_samples FROM genomic_event_derived WHERE cancer_study_identifier = 'os_target_gdc' AND variant_type = 'mutation' AND mutation_status != 'UNCALLED' AND off_panel = 0 GROUP BY hugo_gene_symbol ), cna AS ( SELECT hugo_gene_symbol, COUNT(DISTINCT sample_unique_id) AS cna_samples FROM genomic_event_derived WHERE cancer_study_identifier = 'os_target_gdc' AND variant_type = 'cna' AND cna_alteration IN (2, -2) GROUP BY hugo_gene_symbol ), sv AS ( SELECT hugo_gene_symbol, COUNT(DISTINCT sample_unique_id) AS sv_samples FROM genomic_event_derived WHERE cancer_study_identifier = 'os_target_gdc' AND variant_type = 'structural_variant' GROUP BY hugo_gene_symbol ), combined AS ( SELECT coalesce(mut.hugo_gene_symbol, cna.hugo_gene_symbol, sv.hugo_gene_symbol) AS gene, coalesce(mut_samples,0) AS mut_samples, coalesce(cna_samples,0) AS cna_samples, coalesce(sv_samples,0) AS sv_samples FROM mut FULL OUTER JOIN cna ON mut.hugo_gene_symbol = cna.hugo_gene_symbol FULL OUTER JOIN sv ON coalesce(mut.hugo_gene_symbol, cna.hugo_gene_symbol) = sv.hugo_gene_symbol ) SELECT gene, mut_samples, cna_samples, sv_samples, (mut_samples + cna_samples + sv_samples) AS total_altered_events_approx FROM combined ORDER BY total_altered_events_approx DESC LIMIT 20 " } ◀ result {"rows":[{"gene":"TP53","mut_samples":32,"cna_samples":6,"sv_samples":0,"total_altered_events_approx":38},{"gene":"MYO15A","mut_samples":2,"cna_samples":30,"sv_samples":0,"total_altered_events_approx":32},{"gene":"CSMD3","mut_samples":6,"cna_samples":26,"sv_samples":0,"total_altered_events_approx":32},{"mut_samples":0,"cna_samples":32,"sv_samples":0,"total_altered_events_approx":32},{"gene":"FLII","mut_samples":1,"cna_samples":30,"sv_samples":0,"total_altered_events_approx":31},{"gene":"ZNF286A","mut_samples":1,"cna_samples":30,"sv_samples":0,"total_altered_events_approx":31},{"gene":"CCDC144A","mut_samples":2,"cna_samples":29,"sv_samples":0,"total_altered_events_approx":31},{"gene":"TRIM16","mut_samples":1,"cna_samples":30,"sv_samples":0,"total_altered_events_approx":31},{"mut_samples":0,"cna_samples":31,"sv_samples":0,"total_altered_events_approx":31},{"gene":"RAI1","mut_samples":1,"cna_samples":29,"sv_samples":0,"total_altered_events_approx":30},{"gene":"DNAH9","mut_samples":10,"cna_samples":20,"sv_samples":0,"total_altered_events_approx":30},{"mut_samples":0,"cna_samples":30,"sv_samples":0,"total_altered_events_approx":30},{"mut_samples":0,"cna_samples":30,"sv_samples":0,"total_altered_events_approx":30},{"mut_samples":0,"cna_samples":30,"sv_samples":0,"total_altered_events_approx":30},{"mut_samples":0,"cna_samples":30,"sv_samples":0,"total_altered_events_approx":30},{"mut_samples":0,"cna_samples":30,"sv_samples":0,"total_altered_events_approx":30},{"mut_samples":0,"cna_samples":30,"sv_samples":0,"total_altered_events_approx":30},{"mut_samples":0,"cna_samples":30,"sv_samples":0,"total_altered_events_approx":30},{"mut_samples":0,"cna_samples":30,"sv_samples":0,"total_altered_events_approx":30},{"mut_samples":0,"cna_samples":30,"sv_samples":0,"total_altered_events_approx":30}]} ▶ clickhouse_run_select_query { "query": " WITH mut_profiled AS ( SELECT DISTINCT sample_unique_id FROM sample_to_gene_panel_derived WHERE cancer_study_identifier = 'os_target_gdc' AND alteration_type = 'MUTATION_EXTENDED' ), cna_profiled AS ( SELECT DISTINCT sample_unique_id FROM sample_to_gene_panel_derived WHERE cancer_study_identifier = 'os_target_gdc' AND alteration_type = 'COPY_NUMBER_ALTERATION' ), denom AS ( SELECT count() AS n FROM ( SELECT sample_unique_id FROM mut_profiled UNION DISTINCT SELECT sample_unique_id FROM cna_profiled ) ), altered AS ( SELECT hugo_gene_symbol, COUNT(DISTINCT sample_unique_id) AS altered_samples FROM genomic_event_derived WHERE cancer_study_identifier = 'os_target_gdc' AND ( (variant_type = 'mutation' AND mutation_status != 'UNCALLED' AND off_panel = 0) OR (variant_type = 'cna' AND cna_alteration IN (2, -2)) ) AND hugo_gene_symbol NOT LIKE '%P' AND hugo_gene_symbol NOT LIKE 'RNU%' AND hugo_gene_symbol NOT LIKE 'RN7%' AND hugo_gene_symbol NOT LIKE 'LINC%' GROUP BY hugo_gene_symbol ) SELECT hugo_gene_symbol, altered_samples, (SELECT n FROM denom) AS profiled_samples, ROUND(altered_samples * 100.0 / (SELECT n FROM denom), 1) AS frequency_pct FROM altered ORDER BY altered_samples DESC LIMIT 15 " } ◀ result {"rows":[{"hugo_gene_symbol":"TP53","altered_samples":38,"profiled_samples":158,"frequency_pct":24.1},{"hugo_gene_symbol":"CSMD3","altered_samples":31,"profiled_samples":158,"frequency_pct":19.6},{"hugo_gene_symbol":"MYO15A","altered_samples":31,"profiled_samples":158,"frequency_pct":19.6},{"hugo_gene_symbol":"TRIM16","altered_samples":31,"profiled_samples":158,"frequency_pct":19.6},{"hugo_gene_symbol":"FLII","altered_samples":30,"profiled_samples":158,"frequency_pct":19},{"hugo_gene_symbol":"RNA5SP436","altered_samples":30,"profiled_samples":158,"frequency_pct":19},{"hugo_gene_symbol":"UBE2SP1","altered_samples":30,"profiled_samples":158,"frequency_pct":19},{"hugo_gene_symbol":"MIEF2","altered_samples":30,"profiled_samples":158,"frequency_pct":19},{"hugo_gene_symbol":"ZNF286A","altered_samples":30,"profiled_samples":158,"frequency_pct":19},{"hugo_gene_symbol":"ZNF286A-TBC1D26","altered_samples":30,"profiled_samples":158,"frequency_pct":19},{"hugo_gene_symbol":"ATPAF2","altered_samples":30,"profiled_samples":158,"frequency_pct":19},{"hugo_gene_symbol":"CARS1P2","altered_samples":30,"profiled_samples":158,"frequency_pct":19},{"hugo_gene_symbol":"BRI3P3","altered_samples":30,"profiled_samples":158,"frequency_pct":19},{"hugo_gene_symbol":"MEIS3P1","altered_samples":30,"profiled_samples":158,"frequency_pct":19},{"hugo_gene_symbol":"RAI1","altered_samples":30,"profiled_samples":158,"frequency_pct":19}]} ▶ clickhouse_run_select_query { "query": "SELECT resource_sample_counts, mutation_sample_count, cna_sample_count, sample_count FROM cancer_study WHERE cancer_study_identifier = 'os_target_gdc'" } ◀ result {"rows":[{"resource_sample_counts":{},"mutation_sample_count":143,"cna_sample_count":81,"sample_count":159}]} ▶ clickhouse_run_select_query { "query": "SELECT DISTINCT genetic_profile_id, genetic_alteration_type, datatype, name FROM genetic_profile WHERE cancer_study_identifier = 'os_target_gdc'" } ◀ result {"error_message":"Query execution failed: HTTPDriver for https://dl96orhu96.us-east-1.aws.clickhouse.cloud:8443 received ClickHouse error code 47\n Code: 47. DB::Exception: Unknown expression or function identifier `cancer_study_identifier` in scope SELECT DISTINCT genetic_profile_id, genetic_alteration_type, datatype, name FROM genetic_profile WHERE cancer_study_identifier = 'os_target_gdc'. (UNKNOWN_IDENTIFIER) (version 26.4.1.2359 (official build))\n"} ▶ resolve_and_route { "studyIds": [ "os_target_gdc" ] } ◀ result {"success":true,"message":"Found 1 study. Use the metadata to call the appropriate navigation tool(s).","data":{"totalCount":1,"studiesWithMetadata":[{"studyId":"os_target_gdc","name":"Osteosarcoma (TARGET GDC, 2025)","sampleCount":159,"studyViewUrl":"https://www.cbioportal.org/study?id=os_target_gdc","metadata":{"clinicalAttributeIds":["AGE","CANCER_TYPE","CANCER_TYPE_DETAILED","DAYS_TO_BIRTH","DAYS_TO_DEATH","DISEASE_TYPE","ETHNICITY","MUTATION_COUNT","ONCOTREE_CODE","OS_MONTHS","OS_STATUS","PRIMARY_DIAGNOSIS","PRIMARY_SITE_PATIENT","PROJECT_ID","RACE","SAMPLE_COUNT","SEX","SPECIMEN_TYPE","TMB_NONSYNONYMOUS","VITAL_STATUS"],"molecularProfileIds":["os_target_gdc_cna","os_target_gdc_mrna_seq_fpkm","os_target_gdc_mrna_seq_fpkm_Zscores","os_target_gdc_mrna_seq_read_counts_Zscores","os_target_gdc_mrna_seq_tpm","os_target_gdc_mrna_seq_tpm_Zscores","os_target_gdc_mutations","os_target_gdc_rna_seq_mrna"],"heatmapProfileIds":["os_target_gdc_mrna_seq_fpkm_Zscores","os_target_gdc_mrna_seq_read_counts_Zscores","os_target_gdc_mrna_seq_tpm_Zscores"],"availableComparisonTabs":["overlap","clinical","survival","alterations","mutations","mrna"]}}]}} ▶ navigate_to_results_view { "studyIds": [ "os_target_gdc" ], "genes": [ "TP53", "CSMD3", "MYO15A", "TRIM16", "FLII" ], "tab": "oncoprint" } ◀ result {"success":true,"message":"Navigating to https://www.cbioportal.org/results/oncoprint?cancer_study_list=os_target_gdc&gene_list=TP53%0ACSMD3%0AMYO15A%0ATRIM16%0AFLII&case_set_id=os_target_gdc_all","url":"https://www.cbioportal.org/results/oncoprint?cancer_study_list=os_target_gdc&gene_list=TP53%0ACSMD3%0AMYO15A%0ATRIM16%0AFLII&case_set_id=os_target_gdc_all","data":{"studyIds":["os_target_gdc"],"studies":[{"studyId":"os_target_gdc","name":"Osteosarcoma (TARGET GDC, 2025)","sampleCount":159}],"genes":["TP53","CSMD3","MYO15A","TRIM16","FLII"],"caseSetId":"os_target_gdc_all","pageDescription":"Visual summary of alterations in queried genes across samples."}} ▶ navigate_to_study_view { "studyIds": [ "os_target_gdc" ] } ◀ result {"success":true,"message":"Navigating to https://www.cbioportal.org/study/summary?id=os_target_gdc","url":"https://www.cbioportal.org/study/summary?id=os_target_gdc","data":{"studyIds":["os_target_gdc"],"studies":[{"studyId":"os_target_gdc","name":"Osteosarcoma (TARGET GDC, 2025)","sampleCount":159}],"tab":"summary","hasFilters":false,"hasPlotsConfig":false,"pageDescription":"Overview of all study data displayed as interactive charts summarizing available clinical and genomic data. Click within a chart to filter. Additional charts may be available in the Charts menu."}} ═ answer (success) ## Links - **[OncoPrint of top 5 combined-altered genes](https://www.cbioportal.org/results/oncoprint?cancer_study_list=os_target_gdc&gene_list=TP53%0ACSMD3%0AMYO15A%0ATRIM16%0AFLII&case_set_id=os_target_gdc_all)** — TP53, CSMD3, MYO15A, TRIM16, FLII. Visual summary of alterations in these genes across all 159 samples. - **[Study overview (StudyView)](https://www.cbioportal.org/study/summary?id=os_target_gdc)** — Osteosarcoma (TARGET GDC, 2025), 159 patients/samples. Explore full clinical and genomic charts here, including the built-in "CNA Genes" and "Mutated Genes" tables if you want to inspect the raw amplicon-driven CNA list yourself. **Bottom line:** TP53 is the only gene in this cohort that is recurrently altered on independent, gene-specific grounds (24.1%, mostly point mutation). The next four ranks (CSMD3, MYO15A, TRIM16, FLII) are almost entirely driven by a shared broad copy-number segment (17p11.2/17p12) rather than distinct gene-level selection — worth keeping in mind before treating this as a "top 5 driver genes" list in the oncogenic sense.