Q (sonnet): How many unique patients have both primary and metastatic samples in the MSK-CHORD Study? ▶ read_guide { "uri": "cbioportal://sample-filtering-guide" } ◀ result # Sample and Study Filtering Guide ## Overview Proper filtering is essential for meaningful cBioPortal analysis. This guide covers filtering by studies, sample types, and other criteria. ## Study-Level Filtering ### 1. Always Filter by Study Every query should specify a study to ensure consistent results: ```sql -- Always include study filtering SELECT * FROM your_table WHERE cancer_study_identifier = 'your_study_id' -- Additional filters... ``` ### 2. Find Available Studies ```sql -- Discover available studies SELECT cancer_study_identifier, name, description, type_of_cancer_id FROM cancer_study ORDER BY cancer_study_identifier; ``` ### 3. Study Information ```sql -- Get detailed study information SELECT cs.cancer_study_identifier, cs.name as study_name, cs.description, cs.sample_count, COUNT(DISTINCT p.internal_id) as patient_count FROM cancer_study cs LEFT JOIN patient p ON cs.cancer_study_id = p.cancer_study_id WHERE cs.cancer_study_identifier = 'your_study_id' GROUP BY cs.cancer_study_identifier, cs.name, cs.description, cs.sample_count; ``` ### 4. Find Studies by Available Data Types Use this when the user asks *"which studies have mutation and copy-number data for X"*, *"studies with expression for Y"*, *"is there any imaging data"*, or any *"studies with Z data"* question. `cancer_study` has one precomputed sample count per data type — filter on them. No join and no schema exploration needed. These are the same counts as the portal's study list and its "Data type" filter. **Canonical query — studies with ALL requested data types for a cancer type:** ```sql SELECT cancer_study_identifier, name, sample_count, mutation_sample_count, cna_sample_count FROM cancer_study WHERE type_of_cancer_id = 'luad' -- ← OncoTree code from search_oncotree AND mutation_sample_count > 0 -- ← one condition per required data type AND cna_sample_count > 0 ORDER BY sample_count DESC; ``` **Data type → column** (a count of 0 means the study doesn't have that data): | Data type in the user's question | Column (portal "Data type" filter label) | |---|---| | samples in the study | `sample_count` | | mutation / mutations | `mutation_sample_count` ("Mutations") | | copy-number / CNA / amplification / deletion | `cna_sample_count` ("CNA") | | structural variant / fusion | `structural_variant_sample_count` | | RNA / mRNA / gene expression (any platform) | `mrna_expression_sample_count` — use this for "which studies have expression data" | | specifically the portal's "RNA-Seq" filter | `rna_seq_sample_count` ("RNA-Seq") — counts only the `_rna_seq_v2_mrna` list; 0 for many GDC and iAtlas studies that do have RNA-Seq | | specifically the portal's "RNA (microarray)" filter | `mrna_microarray_sample_count` ("RNA (microarray)") — counts only the `_mrna` list | | microRNA | `mirna_sample_count` ("miRNA") | | protein / RPPA | `rppa_sample_count` ("RPPA") | | mas … (7083 more chars) ▶ list_studies { "search": "MSK-CHORD" } ◀ result {"result":[{"cancer_study_identifier":"msk_chord_2024","name":"MSK-CHORD (MSK, Nature 2024)","type_of_cancer_id":"mixed","sample_count":25040,"has_guide":true,"url":"https://www.cbioportal.org/study/summary?id=msk_chord_2024"}]} ▶ clickhouse_run_select_query { "query": "SELECT attribute_value, COUNT(DISTINCT sample_unique_id) AS n_samples FROM clinical_data_derived WHERE cancer_study_identifier = 'msk_chord_2024' AND attribute_name = 'SAMPLE_TYPE' GROUP BY attribute_value ORDER BY n_samples DESC" } ◀ result {"rows":[{"attribute_value":"Primary","n_samples":15928},{"attribute_value":"Metastasis","n_samples":8878},{"attribute_value":"Unknown","n_samples":136},{"attribute_value":"Local Recurrence","n_samples":98}]} ▶ clickhouse_run_select_query { "query": "WITH primary_pts AS ( SELECT DISTINCT patient_unique_id FROM clinical_data_derived WHERE cancer_study_identifier = 'msk_chord_2024' AND attribute_name = 'SAMPLE_TYPE' AND attribute_value = 'Primary' ), met_pts AS ( SELECT DISTINCT patient_unique_id FROM clinical_data_derived WHERE cancer_study_identifier = 'msk_chord_2024' AND attribute_name = 'SAMPLE_TYPE' AND attribute_value = 'Metastasis' ) SELECT COUNT(*) AS patients_with_both FROM primary_pts WHERE patient_unique_id IN (SELECT patient_unique_id FROM met_pts)" } ◀ result {"rows":[{"patients_with_both":26}]} ▶ clickhouse_run_select_query { "query": "SELECT COUNT(DISTINCT patient_unique_id) AS total_patients FROM clinical_data_derived WHERE cancer_study_identifier = 'msk_chord_2024'" } ◀ result {"rows":[{"total_patients":24950}]} ▶ resolve_and_route { "studyIds": [ "msk_chord_2024" ] } ◀ result {"success":true,"message":"Found 1 study. Use the metadata to call the appropriate navigation tool(s).","data":{"totalCount":1,"studiesWithMetadata":[{"studyId":"msk_chord_2024","name":"MSK-CHORD (MSK, Nature 2024)","sampleCount":25040,"studyViewUrl":"https://www.cbioportal.org/study?id=msk_chord_2024","metadata":{"clinicalAttributeIds":["ADRENAL_GLANDS","BONE","CANCER_TYPE","CANCER_TYPE_DETAILED","CLINICAL_GROUP","CLINICAL_SUMMARY","CNS_BRAIN","CURRENT_AGE_DEID","DIAGNOSIS_DESCRIPTION","ETHNICITY","FRACTION_GENOME_ALTERED","GENDER","GENE_PANEL","GLEASON_FIRST_REPORTED","GLEASON_HIGHEST_REPORTED","GLEASON_SAMPLE_LEVEL","HER2","HISTORY_OF_PDL1","HR","ICD_O_HISTOLOGY_DESCRIPTION","INTRA_ABDOMINAL","LIVER","LUNG","LYMPH_NODES","METASTATIC_SITE","MSI_COMMENT","MSI_SCORE","MSI_TYPE","MUTATION_COUNT","NUM_ICDO_DX","ONCOTREE_CODE","OS_MONTHS","OS_STATUS","OTHER","PATHOLOGICAL_GROUP","PDL1_POSITIVE","PLEURA","PRIMARY_SITE","PRIOR_MED_TO_MSK","RACE","REPRODUCTIVE_ORGANS","SAMPLE_CLASS","SAMPLE_COUNT","SAMPLE_COVERAGE","SAMPLE_TYPE","SMOKING_PREDICTIONS_3_CLASSES","SOMATIC_STATUS","STAGE_HIGHEST_RECORDED","TMB_NONSYNONYMOUS","TUMOR_PURITY"],"molecularProfileIds":["msk_chord_2024_cna","msk_chord_2024_mutations","msk_chord_2024_structural_variants"],"availableComparisonTabs":["overlap","clinical","survival","alterations","mutations"]}}]}} ▶ get_studyviewfilter_options { "studyId": "msk_chord_2024", "attributeIds": [ "SAMPLE_TYPE" ] } ◀ result {"success":true,"message":"Successfully retrieved StudyView filter options","data":{"studyId":"msk_chord_2024","attributes":[{"attributeId":"SAMPLE_TYPE","displayName":"Sample Type","description":"The type assigned to a sample (i.e., normal, primary, met, recurrence). Source: MPath","datatype":"STRING","values":["Local Recurrence","Unknown","Primary","Metastasis"]}]}} ═ answer (success) The Study View doesn't support an "AND" (patient has sample-type A AND sample-type B) filter directly — its clinical filters are sample/patient-attribute filters applied independently, not a per-patient intersection across sample types. So a single filterJson can't reproduce "patients with both a primary and a met sample" exactly. I'll link to the Study View filtered to Metastasis samples (the rarer criterion) as a starting point, and note the 26-patient result comes directly from the database query. ## Summary In the **MSK-CHORD** cohort ([MSK, *Nature* 2024](https://www.cbioportal.org/study/summary?id=msk_chord_2024)), which has 25,040 samples from 24,950 patients: - **26 unique patients** have at least one **Primary** tumor sample *and* at least one **Metastasis** sample on record. - For context, sample-type breakdown: 15,928 Primary, 8,878 Metastasis, 136 Unknown, 98 Local Recurrence samples. This cohort is structured largely as one profiled sample per patient (24,950 patients vs 25,040 samples), so patients with paired primary+metastatic samples are a small minority. Because cBioPortal's Study View filters apply per-attribute rather than as a per-patient "has both" intersection, I can't generate a single filtered link that reproduces exactly these 26 patients. You can browse Metastasis-only samples here and cross-reference manually: [MSK-CHORD — Metastasis samples](https://www.cbioportal.org/study?id=msk_chord_2024) (Use the Sample Type chart on that page and select "Metastasis," then cross-check patient IDs against the Primary-sample list if you'd like to inspect these 26 patients individually.)