Q (sonnet-5.5): List the top 20 mutated genes in study nbl_msk_2023. ▶ read_guide { "uri": "cbioportal://mutation-frequency-guide" } ▶ list_studies { "search": "nbl_msk_2023" } ◀ result # Mutation Frequency Analysis Guide ## IMPORTANT: Reporting Mutation Frequencies - **ALWAYS report frequencies as percentages**, not raw counts: `frequency = (altered_samples / total_profiled_samples) × 100` - For quick frequency lookups, **prefer the TCGA Pan-Cancer Atlas study first**, then offer to expand to other studies - When reporting across multiple studies, show **ranges** (e.g., "TP53 is mutated in 30–60% of samples") rather than a single average - **NEVER** sum mutation events across studies to compute an aggregate frequency — this can exceed 100% due to double-counting - Warn users that samples may overlap across cohorts (e.g., MSK studies may share patients) - **Choose and state the counting unit**: use patient-level frequencies for prevalence/rate questions unless the user explicitly asks for samples; use sample-level frequencies when the user asks about samples. - **For "across cancer types" questions**, jump to the [Cross-Cancer-Type Mutation Frequency](#cross-cancer-type-mutation-frequency) section below — there is one correct recipe and several common wrong ones. ## Counting Unit: Samples vs Patients Before answering any mutation count or frequency question, decide whether the unit is samples or patients and state that choice in the answer. | User wording | Counting unit | |--------------|---------------| | "prevalence", "rate", "fraction of patients", "patients with", "how common is" | Patient-level: `COUNT(DISTINCT patient_unique_id)` | | "samples", "specimens", "biopsies", sample-level cohort composition | Sample-level: `COUNT(DISTINCT sample_unique_id)` | | Ambiguous | Ask, or default to patient-level for prevalence/rate language and say so | ### Cross-study sample-count caveat When an answer touches more than one study and reports a sample count, prepend a one-line caveat: > Sample IDs are unique within cBioPortal study prefixes, not guaranteed biological-sample identifiers across studies; overlapping cohorts can count the same patient/sample more than once. Prefer one of these safer approaches: - Use a shipped `cancer_study_query_preferences` cohort such as `pan_cancer_tcga` or `all_studies_non_redundant`. - Restrict to one named study. - Aggregate by `patient_unique_id` when the biological question is patient prevalence. ## STOP rule: a frequency above 100% means your query is wrong If your query returns a frequency over 100%, **do not try to debug or explain the data inconsistency to the user**. The cause is always one of these query bugs: - Summing mutation events instead of `COUNT(DISTINCT sample_unique_id)` for the numerator - Using a study-wide sample count as the denominator instead of the gene-specific profiled count - Cross-study aggregation where the same biological sample appears under multiple `sample_unique_id` values (e.g., MSK-IMPACT and MSK-CHORD share patients) - **Joining the profiled CTE through `gene_panel` / `gene_panel_list` without a WES branch.** `gene_panel_id = 'WES'` is *not* a row in … (36467 more chars) ◀ result {"result":[{"cancer_study_identifier":"nbl_msk_2023","name":"Pediatric Neuroblastoma (MSK, Nat Genet 2023)","type_of_cancer_id":"nbl","sample_count":223,"has_guide":false,"url":"https://www.cbioportal.org/study/summary?id=nbl_msk_2023"}]} ▶ clickhouse_run_select_query { "query": "SELECT hugo_gene_symbol, total_mutation_events, altered_samples, profiled_samples, frequency_pct FROM top_mutated_genes_in_study(study = 'nbl_msk_2023', top_n = 20)" } ▶ resolve_and_route { "studyIds": [ "nbl_msk_2023" ] } ◀ result {"rows":[{"hugo_gene_symbol":"ALK","total_mutation_events":34,"altered_samples":33,"profiled_samples":223,"frequency_pct":14.8},{"hugo_gene_symbol":"ATRX","total_mutation_events":24,"altered_samples":24,"profiled_samples":223,"frequency_pct":10.8},{"hugo_gene_symbol":"SMARCA4","total_mutation_events":9,"altered_samples":9,"profiled_samples":223,"frequency_pct":4},{"hugo_gene_symbol":"BRAF","total_mutation_events":8,"altered_samples":8,"profiled_samples":223,"frequency_pct":3.6},{"hugo_gene_symbol":"BRCA2","total_mutation_events":8,"altered_samples":8,"profiled_samples":223,"frequency_pct":3.6},{"hugo_gene_symbol":"TERT","total_mutation_events":7,"altered_samples":7,"profiled_samples":223,"frequency_pct":3.1},{"hugo_gene_symbol":"CREBBP","total_mutation_events":8,"altered_samples":6,"profiled_samples":223,"frequency_pct":2.7},{"hugo_gene_symbol":"FAT1","total_mutation_events":6,"altered_samples":6,"profiled_samples":223,"frequency_pct":2.7},{"hugo_gene_symbol":"KMT2D","total_mutation_events":9,"altered_samples":6,"profiled_samples":223,"frequency_pct":2.7},{"hugo_gene_symbol":"KRAS","total_mutation_events":6,"altered_samples":6,"profiled_samples":223,"frequency_pct":2.7},{"hugo_gene_symbol":"PTPN11","total_mutation_events":6,"altered_samples":6,"profiled_samples":223,"frequency_pct":2.7},{"hugo_gene_symbol":"ROS1","total_mutation_events":6,"altered_samples":6,"profiled_samples":223,"frequency_pct":2.7},{"hugo_gene_symbol":"CD79B","total_mutation_events":5,"altered_samples":5,"profiled_samples":223,"frequency_pct":2.2},{"hugo_gene_symbol":"IRS2","total_mutation_events":5,"altered_samples":5,"profiled_samples":223,"frequency_pct":2.2},{"hugo_gene_symbol":"KMT2B","total_mutation_events":5,"altered_samples":5,"profiled_samples":168,"frequency_pct":3},{"hugo_gene_symbol":"NCOR1","total_mutation_events":5,"altered_samples":5,"profiled_samples":223,"frequency_pct":2.2},{"hugo_gene_symbol":"NF1","total_mutation_events":6,"altered_samples":5,"profiled_samples":223,"frequency_pct":2.2},{"hugo_gene_symbol":"PDGFRA","total_mutation_events":5,"altered_samples":5,"profiled_samples":223,"frequency_pct":2.2},{"hugo_gene_symbol":"ARID1A","total_mutation_events":4,"altered_samples":4,"profiled_samples":223,"frequency_pct":1.8},{"hugo_gene_symbol":"AXL","total_mutation_events":4,"altered_samples":4,"profiled_samples":223,"frequency_pct":1.8}]} ◀ result {"success":true,"message":"Found 1 study. Use the metadata to call the appropriate navigation tool(s).","data":{"totalCount":1,"studiesWithMetadata":[{"studyId":"nbl_msk_2023","name":"Pediatric Neuroblastoma (MSK, Nat Genet 2023)","sampleCount":223,"studyViewUrl":"https://www.cbioportal.org/study?id=nbl_msk_2023","metadata":{"clinicalAttributeIds":["AGE_AT_DIAGNOSIS","AGE_AT_SEQ_REPORTED_YEARS","CANCER_TYPE","CANCER_TYPE_DETAILED","DISEASE_SUBTYPE","ETHNICITY","FRACTION_GENOME_ALTERED","GENE_PANEL","INSS_STAGE","INSTITUTE","MEDIAN_COVERAGE","METASTATIC_SITE","MSI_COMMENT","MSI_SCORE","MSI_TYPE","MUTATION_COUNT","ONCOTREE_CODE","OS_MONTHS","OS_STATUS","PRIMARY_SITE","RACE","SAMPLE_CLASS","SAMPLE_CODE","SAMPLE_COUNT","SAMPLE_COVERAGE","SAMPLE_TYPE","SEX","SOMATIC_STATUS","TIME_AT_LAST_FOLLOWUP_MONTHS","TMB_NONSYNONYMOUS","TUMOR_PLOIDY","TUMOR_PURITY","TUMOR_SITE","TUMOR_SITE_CATEGORY","TUMOR_SITE_DETAIL"],"molecularProfileIds":["nbl_msk_2023_cna","nbl_msk_2023_mutations","nbl_msk_2023_structural_variants"],"availableComparisonTabs":["overlap","clinical","survival","alterations","mutations"]}}]}} ▶ navigate_to_study_view { "studyIds": [ "nbl_msk_2023" ] } ◀ result {"success":true,"message":"Navigating to https://www.cbioportal.org/study/summary?id=nbl_msk_2023","url":"https://www.cbioportal.org/study/summary?id=nbl_msk_2023","data":{"studyIds":["nbl_msk_2023"],"studies":[{"studyId":"nbl_msk_2023","name":"Pediatric Neuroblastoma (MSK, Nat Genet 2023)","sampleCount":223}],"tab":"summary","hasFilters":false,"hasPlotsConfig":false,"pageDescription":"Overview of all study data displayed as interactive charts summarizing available clinical and genomic data. Click within a chart to filter. Additional charts may be available in the Charts menu."}} ═ answer (success) These are the 20 most frequently mutated genes in Pediatric Neuroblastoma (MSK, Nat Genet 2023). The counting unit is samples. There are 223 samples in the study, and all but one gene were profiled in all 223. Genes are ranked by the number of samples with a mutation. | # | Gene | Mutations | Mutated samples | Profiled samples | % | |---|------|-----------|-----------------|------------------|-----| | 1 | ALK | 34 | 33 | 223 | 14.8% | | 2 | ATRX | 24 | 24 | 223 | 10.8% | | 3 | SMARCA4 | 9 | 9 | 223 | 4.0% | | 4 | BRAF | 8 | 8 | 223 | 3.6% | | 5 | BRCA2 | 8 | 8 | 223 | 3.6% | | 6 | TERT | 7 | 7 | 223 | 3.1% | | 7 | CREBBP | 8 | 6 | 223 | 2.7% | | 8 | FAT1 | 6 | 6 | 223 | 2.7% | | 9 | KMT2D | 9 | 6 | 223 | 2.7% | | 10 | KRAS | 6 | 6 | 223 | 2.7% | | 11 | PTPN11 | 6 | 6 | 223 | 2.7% | | 12 | ROS1 | 6 | 6 | 223 | 2.7% | | 13 | CD79B | 5 | 5 | 223 | 2.2% | | 14 | IRS2 | 5 | 5 | 223 | 2.2% | | 15 | KMT2B | 5 | 5 | 168 | 3.0% | | 16 | NCOR1 | 5 | 5 | 223 | 2.2% | | 17 | NF1 | 6 | 5 | 223 | 2.2% | | 18 | PDGFRA | 5 | 5 | 223 | 2.2% | | 19 | ARID1A | 4 | 4 | 223 | 1.8% | | 20 | AXL | 4 | 4 | 223 | 1.8% | - **Ties:** Many genes have the same count, so ranks 4–20 are partly arbitrary. Ties are ordered alphabetically, and other genes with 4 mutated samples fall just outside the top 20. - **KMT2B:** Only 168 samples were profiled for this gene, so its percentage uses a smaller denominator. It has 5 mutated samples, which is why it ranks below genes with 6 despite the higher percentage. - **Driver status:** This ranking is by mutation frequency only. It doesn't say which mutations are drivers or oncogenic. [Open the study in cBioPortal](https://www.cbioportal.org/study/summary?id=nbl_msk_2023)