Q (sonnet-5.5): "What are the top 5 frequently altered genes in the Osteosarcoma study from TARGET for mutations, copy numbers and SVs combined?" ▶ list_studies { "search": "TARGET osteosarcoma" } ▶ read_guide { "uri": "cbioportal://mutation-frequency-guide" } ◀ result {"result":[{"cancer_study_identifier":"os_target_gdc","name":"Osteosarcoma (TARGET GDC, 2025)","type_of_cancer_id":"os","sample_count":159,"has_guide":true,"url":"https://www.cbioportal.org/study/summary?id=os_target_gdc"}]} ◀ result # Mutation Frequency Analysis Guide ## IMPORTANT: Reporting Mutation Frequencies - **ALWAYS report frequencies as percentages**, not raw counts: `frequency = (altered_samples / total_profiled_samples) × 100` - For quick frequency lookups, **prefer the TCGA Pan-Cancer Atlas study first**, then offer to expand to other studies - When reporting across multiple studies, show **ranges** (e.g., "TP53 is mutated in 30–60% of samples") rather than a single average - **NEVER** sum mutation events across studies to compute an aggregate frequency — this can exceed 100% due to double-counting - Warn users that samples may overlap across cohorts (e.g., MSK studies may share patients) - **Choose and state the counting unit**: use patient-level frequencies for prevalence/rate questions unless the user explicitly asks for samples; use sample-level frequencies when the user asks about samples. - **For "across cancer types" questions**, jump to the [Cross-Cancer-Type Mutation Frequency](#cross-cancer-type-mutation-frequency) section below — there is one correct recipe and several common wrong ones. ## Counting Unit: Samples vs Patients Before answering any mutation count or frequency question, decide whether the unit is samples or patients and state that choice in the answer. | User wording | Counting unit | |--------------|---------------| | "prevalence", "rate", "fraction of patients", "patients with", "how common is" | Patient-level: `COUNT(DISTINCT patient_unique_id)` | | "samples", "specimens", "biopsies", sample-level cohort composition | Sample-level: `COUNT(DISTINCT sample_unique_id)` | | Ambiguous | Ask, or default to patient-level for prevalence/rate language and say so | ### Cross-study sample-count caveat When an answer touches more than one study and reports a sample count, prepend a one-line caveat: > Sample IDs are unique within cBioPortal study prefixes, not guaranteed biological-sample identifiers across studies; overlapping cohorts can count the same patient/sample more than once. Prefer one of these safer approaches: - Use a shipped `cancer_study_query_preferences` cohort such as `pan_cancer_tcga` or `all_studies_non_redundant`. - Restrict to one named study. - Aggregate by `patient_unique_id` when the biological question is patient prevalence. ## STOP rule: a frequency above 100% means your query is wrong If your query returns a frequency over 100%, **do not try to debug or explain the data inconsistency to the user**. The cause is always one of these query bugs: - Summing mutation events instead of `COUNT(DISTINCT sample_unique_id)` for the numerator - Using a study-wide sample count as the denominator instead of the gene-specific profiled count - Cross-study aggregation where the same biological sample appears under multiple `sample_unique_id` values (e.g., MSK-IMPACT and MSK-CHORD share patients) - **Joining the profiled CTE through `gene_panel` / `gene_panel_list` without a WES branch.** `gene_panel_id = 'WES'` is *not* a row in … (36467 more chars) ▶ get_study_guide { "study_id": "os_target_gdc" } ▶ clickhouse_run_select_query { "query": "SELECT sl.stable_id, count() n FROM sample_list sl JOIN sample_list_list sll ON sl.list_id=sll.list_id WHERE sl.stable_id LIKE 'os_target_gdc%' GROUP BY sl.stable_id" } ◀ result # Osteosarcoma (TARGET GDC, 2025) **Study ID:** `os_target_gdc` ## Overview Pediatric osteosarcoma study from the TARGET (Therapeutically Applicable Research to Generate Effective Treatments) initiative. Whole exome sequencing data. ## Gene Panel - **WES** (Whole Exome Sequencing): all coding genes profiled - **143 of the 160 samples are profiled for mutations.** Use 143 as the mutation-frequency denominator (`sample_to_gene_panel_derived`, `alteration_type = 'MUTATION_EXTENDED'`), not the study's sample count — e.g. TP53 is mutated in 32/143 = 22.4%. ## Patients vs Samples 383 patients have clinical data, but only 153 of them have a sample (159 samples). Patient-level questions (age, sex, survival) use all patients with a value; genomic questions use the 143 mutation-profiled samples. ## Clinical Attributes - Semantic Guide ### Patient Demographics | Attribute | Description | Notes | |-----------|-------------|-------| | `AGE` | Age at diagnosis, **floored at 18** | Every patient younger than 18 is recorded as 18 (241 of 293). **Don't use it for age statistics** — use `DAYS_TO_BIRTH` | | `DAYS_TO_BIRTH` | Days from birth to diagnosis, negative | Age at diagnosis in years = `-DAYS_TO_BIRTH / 365.25`. 293 patients have a value; 90 are empty | | `SEX` | Patient sex | Male 172, Female 133, 78 empty | | `RACE`, `ETHNICITY` | Race, ethnicity | | ### Disease Characteristics | Attribute | Description | Notes | |-----------|-------------|-------| | `CANCER_TYPE_DETAILED` | Cancer type | Osteosarcoma for every sample | | `PRIMARY_SITE_PATIENT` | Primary site | "Appendicular Skeleton" for every patient with a value — no finer location (femur, tibia) | No histological subtype, tumor necrosis / chemotherapy response, or event-free survival attributes exist in this study. ### Clinical Outcomes | Attribute | Description | Notes | |-----------|-------------|-------| | `OS_MONTHS` | Overall survival in months | Time from diagnosis | | `OS_STATUS` | Overall survival status | 0:LIVING 171, 1:DECEASED 105, 107 empty | ## Age at Diagnosis Compute age from `DAYS_TO_BIRTH`, not `AGE`. A median from `AGE` comes out as 18 because every child is recorded as 18; the real median is about 15 years. ```sql SELECT count() AS patients, round(median(-toFloat64OrNull(attribute_value) / 365.25), 1) AS median_age_years, round(min(-toFloat64OrNull(attribute_value) / 365.25), 1) AS min_age_years, round(max(-toFloat64OrNull(attribute_value) / 365.25), 1) AS max_age_years FROM clinical_data_derived WHERE cancer_study_identifier = 'os_target_gdc' AND attribute_name = 'DAYS_TO_BIRTH' AND toFloat64OrNull(attribute_value) IS NOT NULL; -- 293 patients, median 15.2, range 3.6-87.1 ``` When reporting, say the age comes from `DAYS_TO_BIRTH` and that `AGE` is floored at 18. ## Notes & Caveats - This is a pediatric cancer cohort; age distribution is younger than adult studies - WES coverage means no gene panel filtering needed for frequency calculations - Part … (66 more chars) ◀ result {"rows":[{"stable_id":"os_target_gdc_cna","n":81},{"stable_id":"os_target_gdc_cnaseq","n":66},{"stable_id":"os_target_gdc_all","n":159},{"stable_id":"os_target_gdc_tpm","n":88},{"stable_id":"os_target_gdc_sequenced","n":143}]} ▶ clickhouse_run_select_query { "query": "SELECT hugo_gene_symbol, uniqExact(sample_unique_id) altered, uniqExactIf(sample_unique_id, variant_type='mutation') mut, uniqExactIf(sample_unique_id, variant_type='cna' AND cna_alteration=2) amp, uniqExactIf(sample_unique_id, variant_type='cna' AND cna_alteration=-2) del, uniqExactIf(sample_unique_id, variant_type='structural_variant') sv FROM genomic_event_derived WHERE cancer_study_identifier='os_target_gdc' AND hugo_gene_symbol!='' AND ((variant_type='mutation' AND mutation_status!='UNCALLED') OR (variant_type='cna' AND cna_alteration IN (2,-2)) OR variant_type='structural_variant') GROUP BY hugo_gene_symbol ORDER BY altered DESC, hugo_gene_symbol LIMIT 8; SELECT count(DISTINCT sample_unique_id) FROM sample_to_gene_panel_derived WHERE cancer_study_identifier='os_target_gdc'" } ▶ clickhouse_run_select_query { "query": "SELECT alteration_type, uniqExact(sample_unique_id) n FROM sample_to_gene_panel_derived WHERE cancer_study_identifier='os_target_gdc' GROUP BY alteration_type" } ◀ result {"error_message":"Query execution failed: HTTPDriver for https://dl96orhu96.us-east-1.aws.clickhouse.cloud:8443 received ClickHouse error code 62\n Code: 62. DB::Exception: Syntax error: failed at position 684 (end of query) (line 9, col 74): ;. Expected one of: token, DoubleColon, OR, AND, IS NOT DISTINCT FROM, IS DISTINCT FROM, IS NULL, IS NOT NULL, BETWEEN, NOT BETWEEN, LIKE, ILIKE, NOT LIKE, NOT ILIKE, REGEXP, IN, NOT IN, GLOBAL IN, GLOBAL NOT IN, MOD, DIV, alias, AS, Comma, OFFSET, WITH TIES, BY, FETCH, LIMIT, SETTINGS, UNION, EXCEPT, INTERSECT. (SYNTAX_ERROR) (version 26.4.1.2509 (official build))\n"} ◀ result {"rows":[{"alteration_type":"MUTATION_EXTENDED","n":143},{"alteration_type":"COPY_NUMBER_ALTERATION","n":81},{"alteration_type":"MRNA_EXPRESSION","n":88}]} ▶ clickhouse_run_select_query { "query": "SELECT hugo_gene_symbol, uniqExact(sample_unique_id) altered, uniqExactIf(sample_unique_id, variant_type='mutation') mut, uniqExactIf(sample_unique_id, variant_type='cna' AND cna_alteration=2) amp, uniqExactIf(sample_unique_id, variant_type='cna' AND cna_alteration=-2) del, uniqExactIf(sample_unique_id, variant_type='structural_variant') sv FROM genomic_event_derived WHERE cancer_study_identifier='os_target_gdc' AND hugo_gene_symbol!='' AND ((variant_type='mutation' AND mutation_status!='UNCALLED') OR (variant_type='cna' AND cna_alteration IN (2,-2)) OR variant_type='structural_variant') GROUP BY hugo_gene_symbol ORDER BY altered DESC, hugo_gene_symbol LIMIT 8" } ▶ clickhouse_run_select_query { "query": "SELECT uniqExact(sample_unique_id) n FROM sample_to_gene_panel_derived WHERE cancer_study_identifier='os_target_gdc' AND alteration_type IN ('MUTATION_EXTENDED','COPY_NUMBER_ALTERATION')" } ◀ result {"rows":[{"hugo_gene_symbol":"TP53","altered":38,"mut":32,"amp":1,"del":5,"sv":0},{"hugo_gene_symbol":"LINC00901","altered":32,"mut":0,"amp":0,"del":32,"sv":0},{"hugo_gene_symbol":"CSMD3","altered":31,"mut":6,"amp":26,"del":0,"sv":0},{"hugo_gene_symbol":"MYO15A","altered":31,"mut":2,"amp":30,"del":0,"sv":0},{"hugo_gene_symbol":"RN7SL442P","altered":31,"mut":0,"amp":31,"del":0,"sv":0},{"hugo_gene_symbol":"TRIM16","altered":31,"mut":1,"amp":29,"del":1,"sv":0},{"hugo_gene_symbol":"ALKBH5","altered":30,"mut":0,"amp":30,"del":0,"sv":0},{"hugo_gene_symbol":"ATPAF2","altered":30,"mut":0,"amp":30,"del":0,"sv":0}]} ◀ result {"rows":[{"n":158}]} ▶ resolve_and_route { "studyIds": [ "os_target_gdc" ] } ◀ result {"success":true,"message":"Found 1 study. Use the metadata to call the appropriate navigation tool(s).","data":{"totalCount":1,"studiesWithMetadata":[{"studyId":"os_target_gdc","name":"Osteosarcoma (TARGET GDC, 2025)","sampleCount":159,"studyViewUrl":"https://www.cbioportal.org/study?id=os_target_gdc","metadata":{"clinicalAttributeIds":["AGE","CANCER_TYPE","CANCER_TYPE_DETAILED","DAYS_TO_BIRTH","DAYS_TO_DEATH","DISEASE_TYPE","ETHNICITY","MUTATION_COUNT","ONCOTREE_CODE","OS_MONTHS","OS_STATUS","PRIMARY_DIAGNOSIS","PRIMARY_SITE_PATIENT","PROJECT_ID","RACE","SAMPLE_COUNT","SEX","SPECIMEN_TYPE","TMB_NONSYNONYMOUS","VITAL_STATUS"],"molecularProfileIds":["os_target_gdc_cna","os_target_gdc_mrna_seq_fpkm","os_target_gdc_mrna_seq_fpkm_Zscores","os_target_gdc_mrna_seq_read_counts_Zscores","os_target_gdc_mrna_seq_tpm","os_target_gdc_mrna_seq_tpm_Zscores","os_target_gdc_mutations","os_target_gdc_rna_seq_mrna"],"heatmapProfileIds":["os_target_gdc_mrna_seq_fpkm_Zscores","os_target_gdc_mrna_seq_read_counts_Zscores","os_target_gdc_mrna_seq_tpm_Zscores"],"availableComparisonTabs":["overlap","clinical","survival","alterations","mutations","mrna"]}}]}} ▶ navigate_to_results_view { "studyIds": [ "os_target_gdc" ], "genes": [ "TP53", "LINC00901", "CSMD3", "MYO15A", "RN7SL442P", "TRIM16" ] } ◀ result {"success":true,"message":"Navigating to https://www.cbioportal.org/results/oncoprint?cancer_study_list=os_target_gdc&gene_list=TP53%0ALINC00901%0ACSMD3%0AMYO15A%0ARN7SL442P%0ATRIM16&case_set_id=os_target_gdc_all","url":"https://www.cbioportal.org/results/oncoprint?cancer_study_list=os_target_gdc&gene_list=TP53%0ALINC00901%0ACSMD3%0AMYO15A%0ARN7SL442P%0ATRIM16&case_set_id=os_target_gdc_all","data":{"studyIds":["os_target_gdc"],"studies":[{"studyId":"os_target_gdc","name":"Osteosarcoma (TARGET GDC, 2025)","sampleCount":159}],"genes":["TP53","LINC00901","CSMD3","MYO15A","RN7SL442P","TRIM16"],"caseSetId":"os_target_gdc_all","pageDescription":"Visual summary of alterations in queried genes across samples."}} ═ answer (success) **Most frequently altered genes in Osteosarcoma (TARGET GDC, 2025)** I counted samples with a mutation, an amplification, or a homozygous deletion in each gene. Denominators are approximate: 158 samples have mutation or copy-number data, and 143 have mutation data. | Rank | Gene | Altered samples (of 158) | % | Mutation | Amp | Deep del | |---|---|---|---|---|---|---| | 1 | TP53 | 38 | 24.1% | 32 | 1 | 5 | | 2 | LINC00901 | 32 | 20.3% | 0 | 0 | 32 | | 3 (tie) | CSMD3 | 31 | 19.6% | 6 | 26 | 0 | | 3 (tie) | MYO15A | 31 | 19.6% | 2 | 30 | 0 | | 3 (tie) | RN7SL442P | 31 | 19.6% | 0 | 31 | 0 | | 3 (tie) | TRIM16 | 31 | 19.6% | 1 | 29 | 1 | - **Ties:** Four genes tie at 31 samples for ranks 3–6, so there is no unique top 5. I listed all four. - **Structural variants:** The study has no structural variant data, so SVs add nothing to these counts. - **What drives the ranking:** Apart from TP53, which is mostly mutated, the top genes are driven by copy-number events. These are probably large amplified or deleted segments that cover neighbouring genes, so it is not clear that each gene is an individual driver. - **Denominators:** The denominator is samples profiled for mutations or copy number. Some samples were profiled for only one of the two, so the percentages are approximate. TP53 mutation alone is 32/143 (22.4%) among mutation-profiled samples. [View these genes in the OncoPrint](https://www.cbioportal.org/results/oncoprint?cancer_study_list=os_target_gdc&gene_list=TP53%0ALINC00901%0ACSMD3%0AMYO15A%0ARN7SL442P%0ATRIM16&case_set_id=os_target_gdc_all) Visual summary of alterations in queried genes across samples. The OncoPrint uses all 159 samples in the study as its denominator, so its percentages will be slightly lower.