Q (haiku): List the top 20 mutated genes in study nbl_msk_2023. ▶ read_guide { "uri": "cbioportal://mutation-frequency-guide" } ▶ get_study_guide { "study_id": "nbl_msk_2023" } ◀ result # Mutation Frequency Analysis Guide ## IMPORTANT: Reporting Mutation Frequencies - **ALWAYS report frequencies as percentages**, not raw counts: `frequency = (altered_samples / total_profiled_samples) × 100` - For quick frequency lookups, **prefer the TCGA Pan-Cancer Atlas study first**, then offer to expand to other studies - When reporting across multiple studies, show **ranges** (e.g., "TP53 is mutated in 30–60% of samples") rather than a single average - **NEVER** sum mutation events across studies to compute an aggregate frequency — this can exceed 100% due to double-counting - Warn users that samples may overlap across cohorts (e.g., MSK studies may share patients) - **Choose and state the counting unit**: use patient-level frequencies for prevalence/rate questions unless the user explicitly asks for samples; use sample-level frequencies when the user asks about samples. - **For "across cancer types" questions**, jump to the [Cross-Cancer-Type Mutation Frequency](#cross-cancer-type-mutation-frequency) section below — there is one correct recipe and several common wrong ones. ## Counting Unit: Samples vs Patients Before answering any mutation count or frequency question, decide whether the unit is samples or patients and state that choice in the answer. | User wording | Counting unit | |--------------|---------------| | "prevalence", "rate", "fraction of patients", "patients with", "how common is" | Patient-level: `COUNT(DISTINCT patient_unique_id)` | | "samples", "specimens", "biopsies", sample-level cohort composition | Sample-level: `COUNT(DISTINCT sample_unique_id)` | | Ambiguous | Ask, or default to patient-level for prevalence/rate language and say so | ### Cross-study sample-count caveat When an answer touches more than one study and reports a sample count, prepend a one-line caveat: > Sample IDs are unique within cBioPortal study prefixes, not guaranteed biological-sample identifiers across studies; overlapping cohorts can count the same patient/sample more than once. Prefer one of these safer approaches: - Use a shipped `cancer_study_query_preferences` cohort such as `pan_cancer_tcga` or `all_studies_non_redundant`. - Restrict to one named study. - Aggregate by `patient_unique_id` when the biological question is patient prevalence. ## STOP rule: a frequency above 100% means your query is wrong If your query returns a frequency over 100%, **do not try to debug or explain the data inconsistency to the user**. The cause is always one of these query bugs: - Summing mutation events instead of `COUNT(DISTINCT sample_unique_id)` for the numerator - Using a study-wide sample count as the denominator instead of the gene-specific profiled count - Cross-study aggregation where the same biological sample appears under multiple `sample_unique_id` values (e.g., MSK-IMPACT and MSK-CHORD share patients) - **Joining the profiled CTE through `gene_panel` / `gene_panel_list` without a WES branch.** `gene_panel_id = 'WES'` is *not* a row in … (35166 more chars) ◀ result # Study Guide: Pediatric Neuroblastoma (MSK, Nat Genet 2023) **Study ID:** `nbl_msk_2023` **Cancer Type:** nbl **Description:** Targeted sequencing of 223 pediatric neuroblastoma tumor/normal pairs via MSK-IMPACT. ## Cohort Statistics - **Patients:** 352 - **Samples:** 224 ## Available Data Types - **COPY_NUMBER_ALTERATION**: Putative copy-number alterations from GISTIC - **MUTATION_EXTENDED**: Mutations - **STRUCTURAL_VARIANT**: Structural variants ## Gene Panels - **IMPACT468**: 165 samples - **IMPACT410**: 52 samples - **IMPACT341**: 3 samples - **IMPACT505**: 2 samples - **IMPACT-HEME-400**: 1 samples ## Available Clinical Attributes | Attribute | Samples with Data | |-----------|------------------| | SAMPLE_TYPE | 223 | | CANCER_TYPE_DETAILED | 223 | | MSI_COMMENT | 223 | | AGE_AT_SEQ_REPORTED_YEARS | 223 | | PRIMARY_SITE | 223 | | MUTATION_COUNT | 223 | | SOMATIC_STATUS | 223 | | FRACTION_GENOME_ALTERED | 223 | | SAMPLE_CODE | 223 | | TUMOR_PLOIDY | 223 | | TMB_NONSYNONYMOUS | 223 | | SAMPLE_CLASS | 223 | | MEDIAN_COVERAGE | 223 | | CANCER_TYPE | 223 | | INSTITUTE | 223 | | MSI_TYPE | 223 | | SAMPLE_COVERAGE | 223 | | METASTATIC_SITE | 223 | | GENE_PANEL | 223 | | TUMOR_PURITY | 223 | ## Top Mutated Genes | Gene | Altered Samples | |------|----------------| | ALK | 33 | | ATRX | 24 | | SMARCA4 | 9 | | BRAF | 8 | | BRCA2 | 8 | | TERT | 7 | | KMT2D | 6 | | CREBBP | 6 | | KRAS | 6 | | PTPN11 | 6 | ## Sample Types - **Primary**: 103 samples - **Recurrence**: 75 samples - **Metastasis**: 44 samples - **Unknown**: 1 samples ## Query Tips for nbl_msk_2023 ```sql -- Get all samples in this study SELECT DISTINCT sample_unique_id, patient_unique_id FROM clinical_data_derived WHERE cancer_study_identifier = 'nbl_msk_2023'; -- Get mutations for a specific gene SELECT sample_unique_id, hugo_gene_symbol, mutation_variant, mutation_type FROM genomic_event_derived WHERE cancer_study_identifier = 'nbl_msk_2023' AND hugo_gene_symbol = 'TP53' AND variant_type = 'mutation'; -- Get clinical data for specific attributes SELECT sample_unique_id, attribute_name, attribute_value FROM clinical_data_derived WHERE cancer_study_identifier = 'nbl_msk_2023' AND attribute_name IN ('CANCER_TYPE', 'SAMPLE_TYPE', 'OS_MONTHS'); ``` ▶ clickhouse_run_select_query { "query": " SELECT hugo_gene_symbol, altered_samples, profiled_samples, frequency_pct, total_mutation_events FROM top_mutated_genes_in_study(study = 'nbl_msk_2023', top_n = 20) ORDER BY altered_samples DESC ", "max_rows": 20 } ◀ result {"rows":[{"hugo_gene_symbol":"ALK","altered_samples":33,"profiled_samples":223,"frequency_pct":14.8,"total_mutation_events":34},{"hugo_gene_symbol":"ATRX","altered_samples":24,"profiled_samples":223,"frequency_pct":10.8,"total_mutation_events":24},{"hugo_gene_symbol":"SMARCA4","altered_samples":9,"profiled_samples":223,"frequency_pct":4,"total_mutation_events":9},{"hugo_gene_symbol":"BRAF","altered_samples":8,"profiled_samples":223,"frequency_pct":3.6,"total_mutation_events":8},{"hugo_gene_symbol":"BRCA2","altered_samples":8,"profiled_samples":223,"frequency_pct":3.6,"total_mutation_events":8},{"hugo_gene_symbol":"TERT","altered_samples":7,"profiled_samples":223,"frequency_pct":3.1,"total_mutation_events":7},{"hugo_gene_symbol":"CREBBP","altered_samples":6,"profiled_samples":223,"frequency_pct":2.7,"total_mutation_events":8},{"hugo_gene_symbol":"FAT1","altered_samples":6,"profiled_samples":223,"frequency_pct":2.7,"total_mutation_events":6},{"hugo_gene_symbol":"KMT2D","altered_samples":6,"profiled_samples":223,"frequency_pct":2.7,"total_mutation_events":9},{"hugo_gene_symbol":"KRAS","altered_samples":6,"profiled_samples":223,"frequency_pct":2.7,"total_mutation_events":6},{"hugo_gene_symbol":"PTPN11","altered_samples":6,"profiled_samples":223,"frequency_pct":2.7,"total_mutation_events":6},{"hugo_gene_symbol":"ROS1","altered_samples":6,"profiled_samples":223,"frequency_pct":2.7,"total_mutation_events":6},{"hugo_gene_symbol":"CD79B","altered_samples":5,"profiled_samples":223,"frequency_pct":2.2,"total_mutation_events":5},{"hugo_gene_symbol":"IRS2","altered_samples":5,"profiled_samples":223,"frequency_pct":2.2,"total_mutation_events":5},{"hugo_gene_symbol":"KMT2B","altered_samples":5,"profiled_samples":168,"frequency_pct":3,"total_mutation_events":5},{"hugo_gene_symbol":"NCOR1","altered_samples":5,"profiled_samples":223,"frequency_pct":2.2,"total_mutation_events":5},{"hugo_gene_symbol":"NF1","altered_samples":5,"profiled_samples":223,"frequency_pct":2.2,"total_mutation_events":6},{"hugo_gene_symbol":"PDGFRA","altered_samples":5,"profiled_samples":223,"frequency_pct":2.2,"total_mutation_events":5},{"hugo_gene_symbol":"ARID1A","altered_samples":4,"profiled_samples":223,"frequency_pct":1.8,"total_mutation_events":4},{"hugo_gene_symbol":"AXL","altered_samples":4,"profiled_samples":223,"frequency_pct":1.8,"total_mutation_events":4}]} ═ answer (success) Perfect! Here are the **top 20 mutated genes in the Pediatric Neuroblastoma (MSK, 2023) study**: | Gene | Altered Samples | Sample % | Profiled Samples | Total Mutations | |------|-----------------|----------|------------------|-----------------| | ALK | 33 | 14.8% | 223 | 34 | | ATRX | 24 | 10.8% | 223 | 24 | | SMARCA4 | 9 | 4.0% | 223 | 9 | | BRAF | 8 | 3.6% | 223 | 8 | | BRCA2 | 8 | 3.6% | 223 | 8 | | TERT | 7 | 3.1% | 223 | 7 | | CREBBP | 6 | 2.7% | 223 | 8 | | FAT1 | 6 | 2.7% | 223 | 6 | | KMT2D | 6 | 2.7% | 223 | 9 | | KRAS | 6 | 2.7% | 223 | 6 | | PTPN11 | 6 | 2.7% | 223 | 6 | | ROS1 | 6 | 2.7% | 223 | 6 | | CD79B | 5 | 2.2% | 223 | 5 | | IRS2 | 5 | 2.2% | 223 | 5 | | KMT2B | 5 | 3.0% | 168 | 5 | | NCOR1 | 5 | 2.2% | 223 | 5 | | NF1 | 5 | 2.2% | 223 | 6 | | PDGFRA | 5 | 2.2% | 223 | 5 | | ARID1A | 4 | 1.8% | 223 | 4 | | AXL | 4 | 1.8% | 223 | 4 | **Key findings:** - **ALK** is the most frequently mutated gene in this cohort (33/223 samples, 14.8%), reflecting its established importance in neuroblastoma pathogenesis. - **ATRX** is the second most common (24 samples, 10.8%), another known NBL driver. - The cohort includes 224 primary and relapsed/metastatic tumors profiled via MSK-IMPACT gene panels. - Note: **KMT2B** was only profiled in 168 samples (not all 223), likely on a subset of the gene panels; its frequency of 3.0% reflects this smaller profiled denominator.