cBioPortalChat benchmark · 20260925-1407
Headline
Haiku 4.5
Sonnet 5
Precision: pass rate on questions the model attempted. Coverage: share it attempted rather than declined. Costs are what these tokens would cost at Anthropic list prices; this run was answered on a Claude subscription and billed nothing per token.
Outcomes
Pass rate by track
Data: a fact from the data. Navigation: the right cBioPortal link or view. Analysis: comparisons, survival and statistics without invented numbers. Out of scope: declines clearly.
Pass rate by category
Topic of the question. Small categories (low n) swing a lot from run to run.
Tokens and cost
| Model | Answers | Input tokens | of which cache read | cache write | Output tokens | Input / answer | Est. cost | Per answer | Per correct answer |
|---|---|---|---|---|---|---|---|---|---|
| Haiku 4.5 | 14 | 2,102,266 | 1,906,189 | 195,709 | 21,399 | 150,162 | $0.543 | $0.039 | $0.042 |
| Sonnet 5 | 14 | 2,212,998 | 1,942,097 | 270,783 | 21,320 | 158,071 | $1.28 | $0.091 | $0.128 |
Latency and tool use
| Model | Median latency | p90 | Max | LLM calls / answer | Tool calls / answer | Tool errors | Schema errors | Failed requests | Traced |
|---|---|---|---|---|---|---|---|---|---|
| Haiku 4.5 | 20s | 62s | 80s | 5.4 | 4.9 | 4 | 0 | 0 | 14 / 14 |
| Sonnet 5 | 19s | 40s | 55s | 4.2 | 3.9 | 0 | 0 | 0 | 14 / 14 |
| Tool | Haiku 4.5 calls | errors | Sonnet 5 calls | errors |
|---|---|---|---|---|
clickhouse_list_table_columns | 8 | 0 | 4 | 0 |
clickhouse_run_select_query | 39 | 2 | 20 | 0 |
list_studies | 6 | 2 | 2 | 0 |
navigate_to_group_comparison | 2 | 0 | 2 | 0 |
navigate_to_results_view | 1 | 0 | 8 | 0 |
navigate_to_study_view | 1 | 0 | 0 | 0 |
read_guide | 8 | 0 | 7 | 0 |
resolve_and_route | 4 | 0 | 11 | 0 |
Objective checks
| Model | Numeric check disagrees with judge | Answers with cBioPortal links | Links with unknown study ids | Answers exposing internals |
|---|---|---|---|---|
| Haiku 4.5 | 0 of 0 | 6 | 0 | 1 of 14 |
| Sonnet 5 | 0 of 0 | 13 | 0 | 0 of 14 |
Numeric check: when the reference is a single number, whether the answer contains it (within rounding). A disagreement with the judge usually means a stale reference or a judge mistake — worth a look. Unknown study ids suggest a hallucinated link.
Questions
#1001 DataVariants & hotspots · msk_impact_50k_2026 · follow-up (turn 2)How about G12C? ✓23s ✗27s
- Conversation so far
- User: What is the % of KRAS G12D across different tumor types? Use the MSK-IMPACT dataset.Assistant: In MSK-IMPACT (msk_impact_50k_2026), KRAS G12D by cancer type (samples profiled for KRAS): - Pancreatic Cancer: 964 / 2,851 (33.8%) - Colorectal Cancer: 674 / 5,277 (12.8%) - Cancer of Unknown Primary: 100 / 1,581 (6.3%) - Endometrial Cancer: 135 / 2,445 (5.5%) - Non-Small Cell Lung Cancer: 285 / 7,867 (3.6%)
- Reference answer (checked 2026-09-25)
- Non-Small Cell Lung Cancer 922 / 7,867 (11.7%) is highest; then Colorectal 158 / 5,277 (3.0%), Cancer of Unknown Primary 48 / 1,581 (3.0%), Endometrial 30 / 2,445 (1.2%), Pancreatic 33 / 2,851 (1.2%).
- Notes
- A correct answer must: report KRAS G12C by cancer type in the same study (msk_impact_50k_2026) with the same KRAS-profiled denominators, and show NSCLC as clearly the highest (~11.7%). Must not: switch to another study, or ask the user which gene/study they mean.
Haiku 4.5 · ✓ Pass 23s · 117,723 tok · $0.057 · 4 LLM / 3 tool calls · transcript
In MSK-IMPACT (msk_impact_50k_2026), **KRAS G12C** by cancer type (samples profiled for KRAS): - **Non-Small Cell Lung Cancer**: 922 / 7,867 (11.7%) - **Appendiceal Cancer**: 12 / 274 (4.4%) - **Small Bowel Cancer**: 7 / 161 (4.3%) - **Colorectal Cancer**: 158 / 5,277 (3.0%) - **Cancer of Unknown Primary**: 48 / 1,581 (3.0%) - **Ampullary Cancer**: 4 / 134 (3.0%) - **Vaginal Cancer**: 1 / 57 (1.8%) - **Small Cell Lung Cancer**: 8 / 466 (1.7%) - **Anal Cancer**: 2 / 143 (1.4%) - **Endometrial Cancer**: 30 / 2,445 (1.2%) - **Pancreatic Cancer**: 33 / 2,851 (1.2%) G12C is predominantly a lung cancer variant, appearing in 11.7% of NSCLC samples versus much lower frequencies in other cancer types.
Sonnet 5 · ✗ Fail 27s · 208,594 tok · $0.114 · 5 LLM / 4 tool calls · transcript
[View KRAS G12C mutations across MSK-IMPACT](https://www.cbioportal.org/results/mutations?cancer_study_list=msk_impact_50k_2026&gene_list=KRAS%3A+MUT+%3D+G12C&case_set_id=msk_impact_50k_2026_all) — detailed mutation table with lollipop plot showing mutation positions, types, and frequencies across the full cohort (54,331 samples).
#1002 DataAlteration frequency · brca_tcga_pan_can_atlas_2018 · follow-up (turn 2)break that down by PAM50 subtype please ✓80s ✓55s
- Conversation so far
- User: How many samples in the TCGA PanCancer Atlas breast cancer study have PIK3CA mutations?Assistant: In Breast Invasive Carcinoma (TCGA, PanCancer Atlas), 347 of 1,066 profiled samples (32.6%) have a PIK3CA mutation.
- Reference answer (checked 2026-09-25)
- LumA 237 / 499 (47.5%), Her2 26 / 78 (33.3%), LumB 59 / 197 (29.9%), Normal-like 8 / 36 (22.2%), Basal 12 / 171 (7.0%).
- Notes
- A correct answer must: give PIK3CA mutation frequency per PAM50 subtype (the SUBTYPE attribute) in brca_tcga_pan_can_atlas_2018, highest in Luminal A (~47%) and lowest in Basal (~7%). Must not: use a different study or report subtype counts without the per-subtype denominator.
Haiku 4.5 · ✓ Pass 80s · 574,301 tok · $0.117 · 16 LLM / 17 tool calls · transcript
Perfect! Here's the PIK3CA mutation frequency breakdown by PAM50 subtype in the TCGA PanCancer Atlas breast cancer study: | PAM50 Subtype | PIK3CA Mutated | Total Patients | Mutation Frequency | |---|---|---|---| | **Luminal A** | 237 | 499 | **47.5%** | | **HER2-enriched** | 26 | 78 | **33.3%** | | **Luminal B** | 59 | 197 | **29.9%** | | **Normal-like** | 8 | 36 | **22.2%** | | **Basal-like** | 12 | 171 | **7.0%** | PIK3CA mutations are strongly enriched in the **luminal subtypes** (especially Luminal A), consistent with the hormone receptor-positive biology of these tumors. The Basal-like subtype has the lowest PIK3CA mutation frequency at 7.0%, which aligns with its distinct molecular profile.
Sonnet 5 · ✓ Pass 55s · 527,888 tok · $0.239 · 10 LLM / 13 tool calls · transcript
In **Breast Invasive Carcinoma (TCGA, PanCancer Atlas)**, PIK3CA mutation frequency by PAM50 subtype (mutated samples / samples profiled for mutations within each subtype): | PAM50 Subtype | PIK3CA Mutated | Profiled | % | |---|---|---|---| | Luminal A | 237 | 499 | 47.5% | | HER2-enriched | 26 | 78 | 33.3% | | Luminal B | 59 | 197 | 29.9% | | Normal-like | 8 | 36 | 22.2% | | Basal-like | 12 | 171 | 7.0% | PIK3CA mutations are markedly enriched in Luminal A tumors and comparatively rare in Basal-like disease — consistent with the well-known association between PIK3CA activation and luminal/hormone-receptor-positive breast cancer biology (this pattern description is general genomics knowledge, not a statistical test result from cBioPortal). Note: PAM50 subtype (`SUBTYPE`) is a patient-level attribute, so the denominators above reflect samples from patients with an assigned subtype and mutation profiling; 55 profiled samples had no subtype call and are excluded from this breakdown. [View PIK3CA Alteration Comparison by PAM50 Subtype](https://www.cbioportal.org/comparison/alterations?comparisonId=6ab68070e497aad53593c400&selectedGene=PIK3CA) — opens the Comparison/Alterations tab with the five subtype groups pre-built, where you can also run enrichment statistics directly in the portal.
#1003 DataVariants & hotspots · lusc_tcga_pan_can_atlas_2018 · follow-up (turn 2)And in lung squamous? ✓23s ✗25s
- Conversation so far
- User: What are the most common KRAS mutations in TCGA lung adenocarcinoma?Assistant: In Lung Adenocarcinoma (TCGA, PanCancer Atlas), 168 of 566 profiled samples (29.7%) have a KRAS mutation. The most common: - G12C: 70 (12.4%) - G12V: 40 (7.1%) - G12D: 20 (3.5%) - G12A: 17 (3.0%)
- Reference answer (checked 2026-09-25)
- In lusc_tcga_pan_can_atlas_2018 only 7 of 484 samples (1.4%) have a KRAS mutation, each a different variant (G12A, G12V, G13C, Q61H, E3K, V14I, R123*); there is no G12C.
- Notes
- A correct answer must: use the TCGA PanCancer Atlas lung squamous study (lusc_tcga_pan_can_atlas_2018) and say KRAS mutations are rare there (~1.4%, 7 of 484) with no recurrent hotspot. Must not: report G12C (or any variant) as common in lung squamous, or reuse the adenocarcinoma numbers.
Haiku 4.5 · ✓ Pass 23s · 151,713 tok · $0.038 · 5 LLM / 5 tool calls · transcript
In **Lung Squamous Cell Carcinoma (TCGA, PanCancer Atlas)**, KRAS mutations are very rare: - **7 of 484 profiled samples (1.4%)** carry a KRAS mutation This is a stark contrast to lung adenocarcinoma (29.7%). The variants are scattered: - G12A: 1 (0.2%) - G12V: 1 (0.2%) - G13C: 1 (0.2%) - Q61H: 1 (0.2%) - V14I: 1 (0.2%) - E3K: 1 (0.2%) - R123*: 1 (0.2%) **Key observation:** Squamous cell carcinomas (both lung and head-and-neck) have low KRAS mutation rates, while adenocarcinomas have high rates. This reflects a well-known histologic difference in lung cancer biology.
list_studies: Error POSTing to endpoint: {"jsonrpc":"2.0","id":5,"error":{"code":-32600,"message":"Anthropic Proxy: Invalid content from server","data":null}}Sonnet 5 · ✗ Fail 25s · 264,241 tok · $0.176 · 7 LLM / 8 tool calls · transcript
[View KRAS mutations in Lung Squamous Cell Carcinoma (TCGA, PanCancer Atlas)](https://www.cbioportal.org/results/mutations?cancer_study_list=lusc_tcga_pan_can_atlas_2018&gene_list=KRAS%3A+MUT&case_set_id=lusc_tcga_pan_can_atlas_2018_all) Detailed mutation table with lollipop plot showing mutation positions, types, and frequencies for KRAS.
#1004 DataAlteration frequency · brca_tcga_pan_can_atlas_2018 · follow-up (turn 2)変異だけでなく、ホモ接合性欠失(deep deletion)も含めて計算してください。 ✓24s ✓34s
- Conversation so far
- User: TCGA PanCancer Atlas の乳がんで TP53 変異の頻度を教えてください。Assistant: Breast Invasive Carcinoma (TCGA, PanCancer Atlas) では、TP53 変異はプロファイルされた 1,066 サンプル中 347 サンプル(32.6%)に見られます。
- Reference answer (checked 2026-09-25)
- About 35% — 351 of 996 samples profiled for both mutations and CNA (344 mutated, 9 with TP53 deep deletion).
- Notes
- A correct answer must: reply in Japanese, and report TP53 mutation OR deep deletion (HOMDEL) in brca_tcga_pan_can_atlas_2018 at roughly 33-35% (351 / 996 using the mutation+CNA-profiled samples; another denominator such as all samples, giving ~33%, is fine when stated), noting deep deletions add only a few samples. Must not: count shallow deletions or gains, or answer in English only.
Haiku 4.5 · ✓ Pass 24s · 87,439 tok · $0.055 · 3 LLM / 4 tool calls · transcript
TCGA PanCancer Atlas乳がん(Breast Invasive Carcinoma)の **TP53のホモ接合性欠失(HOMDEL)を含めた頻度** は以下の通りです: | TP53変異タイプ | 変異サンプル数 | プロファイル数 | 頻度 | |---|---|---|---| | **点突然変異** | 347 | 1,066 | **32.6%** | | **ホモ接合性欠失(HOMDEL)** | 9 | 1,070 | **0.8%** | | **突然変異またはHOMDEL** | 354 | 1,070 | **33.1%** | **結論:** TP53の機能喪失イベント(点突然変異またはホモ接合性欠失)は、乳がんサンプルの約 **33%** に見られます。ホモ接合性欠失は比較的少数(9サンプル)ですが、点突然変異のわずか約2.6%に相当します。
Sonnet 5 · ✓ Pass 34s · 251,449 tok · $0.178 · 6 LLM / 5 tool calls · transcript
Breast Invasive Carcinoma (TCGA, PanCancer Atlas) において、TP53 の変異とホモ接合性欠失(deep deletion)を合わせた頻度は以下の通りです(変異とCNAの両方がプロファイルされたサンプル、いわゆる `cnaseq` セットを分母に使用)。 | | サンプル数 | 分母(プロファイル済み) | 頻度 | |---|---|---|---| | 変異 or 深欠失(合算) | 351 | 996 | **35.2%** | 内訳: - 変異のみ:344 サンプル - 深欠失のみ:9 サンプル - 両方を保有:2 サンプル(重複しているため単純合計 353 ではなく 351) 変異単独の頻度(前回回答)は 32.6% でしたが、深欠失を加えると 35.2% まで上昇します。深欠失自体は稀(全体の約1%未満)で、大部分は変異による寄与です。 [OncoPrint: TP53 変異+深欠失](https://www.cbioportal.org/results/oncoprint?cancer_study_list=brca_tcga_pan_can_atlas_2018&gene_list=TP53%3A+MUT+HOMDEL&case_set_id=brca_tcga_pan_can_atlas_2018_cnaseq) このリンクでは変異・欠失両方を含めたオンコプリントを直接確認できます。
#1005 DataAlteration frequency · luad_tcga_pan_can_atlas_2018 · follow-up (turn 2)那 KRAS 呢? ✓20s ✓24s
- Conversation so far
- User: 在TCGA肺腺癌研究中,EGFR突变的频率是多少?Assistant: 在 Lung Adenocarcinoma (TCGA, PanCancer Atlas) 中,566 个检测样本中有 70 个(12.4%)携带 EGFR 突变。
- Reference answer (checked 2026-09-25)
- 168 of 566 samples (29.7%).
- Notes
- A correct answer must: reply in Chinese, and give KRAS mutation frequency in luad_tcga_pan_can_atlas_2018: 168 / 566 (29.7%). Must not: switch to another study or answer in English only.
Haiku 4.5 · ✓ Pass 20s · 180,098 tok · $0.041 · 6 LLM / 5 tool calls · transcript
在 Lung Adenocarcinoma (TCGA, PanCancer Atlas) 中,**KRAS 突变的频率是 29.7%**(168 个样本中有 168 个携带 KRAS 突变,共 566 个检测样本)。 对比一下: - **EGFR**:70/566 = 12.4% - **KRAS**:168/566 = 29.7% KRAS 在 TCGA 肺腺癌中的突变频率明显高于 EGFR,这与临床观察一致——在非吸烟者中 EGFR 突变更常见,而在吸烟者中 KRAS 突变更为常见。
clickhouse_run_select_query: Error POSTing to endpoint: {"jsonrpc":"2.0","id":5,"error":{"code":-32600,"message":"Anthropic Proxy: Invalid content from server","data":null}}clickhouse_run_select_query: Error POSTing to endpoint: {"jsonrpc":"2.0","id":6,"error":{"code":-32600,"message":"Anthropic Proxy: Invalid content from server","data":null}}Sonnet 5 · ✓ Pass 24s · 211,695 tok · $0.158 · 6 LLM / 6 tool calls · transcript
在同一研究——Lung Adenocarcinoma (TCGA, PanCancer Atlas)——中,566 个检测样本中有 168 个携带 KRAS 突变,频率为 **29.7%**,明显高于 EGFR 的 12.4%。 [查看 KRAS 突变详情(Mutations Tab)](https://www.cbioportal.org/results/mutations?cancer_study_list=luad_tcga_pan_can_atlas_2018&gene_list=KRAS%3A+MUT&case_set_id=luad_tcga_pan_can_atlas_2018_all)
#1006 DataAlteration frequency · TCGA PanCancer Atlas (32 studies) · follow-up (turn 2)yes please ✓29s ✓40s
- Conversation so far
- User: How often is TP53 mutated in the TCGA PanCancer Atlas cohort?Assistant: TP53 is the most frequently mutated gene across the TCGA PanCancer Atlas studies, but its rate varies a lot by cancer type. Would you like me to break it down by cancer type?
- Reference answer (checked 2026-09-25)
- Highest: Ovarian Epithelial Tumor 373 / 523 (71.3%), Head and Neck 357 / 515 (69.3%), Non-Small Cell Lung Cancer 699 / 1,050 (66.6%), Esophagogastric 371 / 618 (60.0%), Pancreatic 107 / 179 (59.8%), Colorectal 314 / 534 (58.8%).
- Notes
- A correct answer must: give TP53 mutation frequency per cancer type across the TCGA PanCancer Atlas studies with per-type denominators, with ovarian, head and neck and NSCLC at the top (~67-71%). Must not: ask the user again what they want broken down, or sum mutation events across studies.
Haiku 4.5 · ✓ Pass 29s · 187,903 tok · $0.048 · 6 LLM / 5 tool calls · transcript
## TP53 Mutation Frequency Across TCGA PanCancer Atlas TP53 is frequently mutated across most cancer types in the TCGA PanCancer Atlas, with dramatic variation by tumor type: | Cancer Type | TP53 Mutations | Profiled Samples | Frequency | |---|---|---|---| | **Ovarian Epithelial Tumor** | 373 | 523 | **71.3%** | | **Head and Neck Cancer** | 357 | 515 | **69.3%** | | **Non-Small Cell Lung Cancer** | 699 | 1,050 | **66.6%** | | **Esophagogastric Cancer** | 371 | 618 | **60.0%** | | **Pancreatic Cancer** | 107 | 179 | **59.8%** | | **Colorectal Cancer** | 314 | 534 | **58.8%** | | Bladder Cancer | 202 | 410 | 49.3% | | Glioma | 249 | 514 | 48.4% | | Endometrial Cancer | 245 | 574 | 42.7% | | Sarcoma | 85 | 255 | 33.3% | | **Breast Cancer** | 347 | 1,066 | **32.6%** | | Glioblastoma | 125 | 397 | 31.5% | | Hepatobiliary Cancer | 110 | 366 | 30.1% | | Prostate Cancer | 57 | 494 | 11.5% | | Melanoma | 72 | 440 | 16.4% | | Renal Clear Cell Carcinoma | 11 | 402 | 2.7% | | Thyroid Cancer | 2 | 490 | 0.4% | **Key patterns:** - **Highest frequency:** Ovarian (71%), head & neck (69%), and lung (67%) tumors show TP53 mutations in ~70% of cases - **Moderate frequency:** Most solid epithelial cancers (60%, colon, esophageal, pancreatic) - **Lower frequency:** Kidney (3–8%), melanoma (16%), prostate (12%), and thyroid (0.4%) rarely harbor TP53 mutations - **Total altered:** 4,468 of 10,509 profiled samples (42.6%) carry TP53 mutations across all 32 TCGA PanCancer Atlas studies [View TP53 mutations in Lung Adenocarcinoma](https://www.cbioportal.org/results/mutations?cancer_study_list=luad_tcga_pan_can_atlas_2018&gene_list=TP53&case_set_id=luad_tcga_pan_can_atlas_2018_all) to explore mutation types and positions in detail.
Sonnet 5 · ✓ Pass 40s · 217,351 tok · $0.150 · 5 LLM / 5 tool calls · transcript
Here's the TP53 mutation frequency (sample-level, mutated/profiled samples) across the 32 TCGA PanCancer Atlas studies, grouped by the `CANCER_TYPE` clinical label (WES-profiled, cancer types with ≥50 profiled samples shown): | Cancer type | Altered / Profiled | Frequency | |---|---|---| | Ovarian Epithelial Tumor | 373/523 | 71.3% | | Head and Neck Cancer | 357/515 | 69.3% | | Non-Small Cell Lung Cancer (LUAD+LUSC) | 699/1050 | 66.6% | | Esophagogastric Cancer | 371/618 | 60.0% | | Pancreatic Cancer | 107/179 | 59.8% | | Colorectal Cancer | 314/534 | 58.8% | | Bladder Cancer | 202/410 | 49.3% | | Glioma (LGG) | 249/514 | 48.4% | | Endometrial Cancer | 245/574 | 42.7% | | Sarcoma | 85/255 | 33.3% | | Breast Cancer | 347/1066 | 32.6% | | Glioblastoma | 125/397 | 31.5% | | Hepatobiliary Cancer | 110/366 | 30.1% | | Adrenocortical Carcinoma | 18/91 | 19.8% | | Melanoma | 72/440 | 16.4% | | Pleural Mesothelioma | 14/86 | 16.3% | | Prostate Cancer | 57/494 | 11.5% | | Leukemia (AML) | 16/200 | 8.0% | | Cervical Cancer | 23/291 | 7.9% | | Renal Non-Clear Cell Carcinoma | 27/341 | 7.9% | | Thymic Epithelial Tumor | 4/123 | 3.3% | | Renal Clear Cell Carcinoma | 11/402 | 2.7% | | Seminoma | 1/63 | 1.6% | | Pheochromocytoma | 1/147 | 0.7% | | Thyroid Cancer | 2/490 | 0.4% | **Notes:** - Counting unit is samples (one per patient in these TCGA cohorts). Denominators are gene-specific profiled samples, not total study size. - The "Non-Small Cell Lung Cancer" bucket pools LUAD and LUSC, which actually diverge (LUAD ~52%, LUSC ~84%) — a single-organ breakdown would split these if you want that detail. - Highly mutated types at the top (ovarian, head & neck, NSCLC, esophagogastric) reflect TP53's role as a near-universal driver in high-grade/squamous carcinomas; low rates in thyroid, pheochromocytoma, and kidney clear cell reflect their reliance on other pathways — this is general cancer biology context, not something pulled from the database. I've also set up an OncoPrint-style comparison splitting all 32 studies into TP53-Mutant (3,839 samples) vs. TP53-Wild-type (7,128 samples) groups, which you can use to explore co-alterations, clinical differences, or survival between these groups: [Compare TP53-Mutant vs Wild-Type across TCGA PanCancer Atlas](https://www.cbioportal.org/comparison/alterations?comparisonId=6ab68073c2115c492d884f85&selectedGene=TP53)
#1007 DataCohort & clinical counts · msk_chord_2024 · follow-up (turn 2)That seems low, I'm pretty sure MSK-CHORD has about 50,000 patients. ✓19s ✓17s
- Conversation so far
- User: How many patients are in MSK-CHORD?Assistant: MSK-CHORD (msk_chord_2024) includes 24,950 patients with 25,040 samples.
- Reference answer (checked 2026-09-25)
- 24,950 patients (25,040 samples).
- Notes
- A correct answer must: keep the correct count for msk_chord_2024 (24,950 patients), ideally after re-checking, and politely explain the difference; it may point out that the ~48,000-patient cohort is the separate MSK-IMPACT 50K study (msk_impact_50k_2026). Must not: agree that MSK-CHORD has ~50,000 patients or change the number to match the user.
Haiku 4.5 · ✓ Pass 19s · 121,867 tok · $0.021 · 6 LLM / 5 tool calls · transcript
The database confirms **24,950 patients** in MSK-CHORD (msk_chord_2024). It's possible you're thinking of: 1. **A different version or broader MSK cohort** — there may be other MSK studies with different patient counts 2. **Total sample count** — MSK-CHORD has 25,040 samples (some patients have multiple samples) 3. **A projected or announced future size** — the cohort may have been planned to grow to ~50,000 Would you like me to search for other MSK studies in the database, or do you have additional context on where the 50,000 figure comes from?
Sonnet 5 · ✓ Pass 17s · 52,233 tok · $0.025 · 2 LLM / 2 tool calls · transcript
I checked directly against the database (not just relying on precomputed metadata), and the number holds up: the **MSK-CHORD (MSK, Nature 2024)** study as loaded in cBioPortal contains **24,950 patients** and **25,040–25,041 samples**. A couple of notes on where the ~50,000 figure you're recalling may come from: - The original MSK-CHORD publication (Jonsson et al., *Nature* 2024) describes a cohort of **~25,000 tumors from ~24,000 patients** profiled via MSK-IMPACT — consistent with what's in cBioPortal, not 50,000. - If you're thinking of a different number, it's possible you're recalling the total number of **samples across all MSK-IMPACT-related cohorts** (MSK-CHORD plus other MSK studies) combined, or a different MSK dataset (e.g. the full MSK-IMPACT clinical sequencing cohort across many published studies) rather than MSK-CHORD specifically. MSK-CHORD is a curated subset covering four cancer types (NSCLC, breast, colorectal, prostate). If you'd like, I can pull up the study directly so you can inspect the counts yourself: [View MSK-CHORD Study Overview](https://www.cbioportal.org/study/summary?id=msk_chord_2024) Let me know if you had a specific different cohort or paper in mind — I'm happy to check that one too.
#1008 DataStudy discovery · TARGET studies · follow-up (turn 2)TARGET, the pediatric one ✓14s ✗15s
- Conversation so far
- User: How can I find the TARGET information?Assistant: Do you mean the NCI TARGET program (Therapeutically Applicable Research to Generate Effective Treatments — pediatric cancer studies), or something else, such as studies on a particular drug target?
- Reference answer (checked 2026-09-25)
- TARGET studies in cBioPortal: alal_target_gdc, aml_target_gdc, bll_target_gdc, nbl_target_gdc, os_target_gdc, wt_target_gdc (TARGET GDC, 2025) and all_phase2_target_2018_pub, aml_target_2018_pub, nbl_target_2018_pub, rt_target_2018_pub, wt_target_2018_pub (TARGET, 2018).
- Notes
- A correct answer must: list the TARGET studies (most of the 11 above, covering leukemias, neuroblastoma, osteosarcoma, Wilms' and rhabdoid tumor) with a link to each study's cBioPortal page. Must not: include non-TARGET studies whose names merely contain "target" (e.g. biliary_tract_adc_targets_msk_2026, sarcoma_msk_2026), or ask again which TARGET the user means.
Haiku 4.5 · ✓ Pass 14s · 90,393 tok · $0.029 · 4 LLM / 3 tool calls · transcript
Perfect! Here are the **TARGET pediatric cancer studies** available in cBioPortal: ## TARGET GDC (2025) — Newest Studies - [Acute Myeloid Leukemia (TARGET GDC, 2025)](https://www.cbioportal.org/study/summary?id=aml_target_gdc) — 2,766 samples - [B-Lymphoblastic Leukemia/Lymphoma (TARGET GDC, 2025)](https://www.cbioportal.org/study/summary?id=bll_target_gdc) — 989 samples - [Neuroblastoma (TARGET GDC, 2025)](https://www.cbioportal.org/study/summary?id=nbl_target_gdc) — 533 samples - [Wilms' Tumor (TARGET GDC, 2025)](https://www.cbioportal.org/study/summary?id=wt_target_gdc) — 132 samples - [Acute Leukemias of Ambiguous Lineage (TARGET GDC, 2025)](https://www.cbioportal.org/study/summary?id=alal_target_gdc) — 251 samples - [Osteosarcoma (TARGET GDC, 2025)](https://www.cbioportal.org/study/summary?id=os_target_gdc) — 159 samples *(has study guide)* ## TARGET (2018) — Earlier Release - [Pediatric Acute Lymphoid Leukemia - Phase II (TARGET, 2018)](https://www.cbioportal.org/study/summary?id=all_phase2_target_2018_pub) — 1,978 samples - [Pediatric Neuroblastoma (TARGET, 2018)](https://www.cbioportal.org/study/summary?id=nbl_target_2018_pub) — 1,089 samples - [Pediatric Acute Myeloid Leukemia (TARGET, 2018)](https://www.cbioportal.org/study/summary?id=aml_target_2018_pub) — 1,025 samples - [Pediatric Wilms' Tumor (TARGET, 2018)](https://www.cbioportal.org/study/summary?id=wt_target_2018_pub) — 657 samples Click on any study link to explore the cohort, view clinical attributes, gene alterations, mutations, and survival data. The **2025 GDC versions** contain the most recent updates.
list_studies: Sonnet 5 · ✗ Fail 15s · 83,162 tok · $0.037 · 3 LLM / 2 tool calls · transcript
There are multiple TARGET (pediatric, NCI Therapeutically Applicable Research to Generate Effective Treatments) datasets in cBioPortal — one set from GDC harmonization, plus an older ALL-specific publication: **TARGET GDC (2025) — genomically harmonized via GDC:** - [Acute Myeloid Leukemia](https://www.cbioportal.org/study/summary?id=aml_target_gdc) — 2,766 samples - [B-Lymphoblastic Leukemia/Lymphoma](https://www.cbioportal.org/study/summary?id=bll_target_gdc) — 989 samples - [Neuroblastoma](https://www.cbioportal.org/study/summary?id=nbl_target_gdc) — 533 samples - [Acute Leukemias of Ambiguous Lineage](https://www.cbioportal.org/study/summary?id=alal_target_gdc) — 251 samples - [Osteosarcoma](https://www.cbioportal.org/study/summary?id=os_target_gdc) — 159 samples - [Wilms' Tumor](https://www.cbioportal.org/study/summary?id=wt_target_gdc) — 132 samples **Older TARGET publication cohort:** - [Pediatric Acute Lymphoid Leukemia - Phase II (TARGET, 2018)](https://www.cbioportal.org/study/summary?id=all_phase2_target_2018_pub) — 1,978 samples Which one (or several) would you like to explore? If you tell me what you're after — e.g., mutation frequencies, clinical characteristics, a specific gene — I can pull the data directly or send you to the right view in cBioPortal.
#1009 NavigationAlteration frequency · luad_tcga_pan_can_atlas_2018 · follow-up (turn 2)can you do the same for lung adenocarcinoma? ✓11s ✗12s
- Conversation so far
- User: Show me an OncoPrint of KRAS, TP53 and CDKN2A in TCGA pancreatic cancer.Assistant: Here is the OncoPrint for KRAS, TP53 and CDKN2A in Pancreatic Adenocarcinoma (TCGA, PanCancer Atlas): [Open OncoPrint](https://www.cbioportal.org/results/oncoprint?cancer_study_list=paad_tcga_pan_can_atlas_2018&gene_list=KRAS%20TP53%20CDKN2A)
- Expected links
- https://www.cbioportal.org/results/oncoprint?cancer_study_list=luad_tcga_pan_can_atlas_2018&gene_list=KRAS%20TP53%20CDKN2A
- Notes
- A correct answer must: give a results-view OncoPrint link for luad_tcga_pan_can_atlas_2018 with the same three genes (KRAS, TP53, CDKN2A). Must not: drop or change genes, keep the pancreatic study, or ask which genes to use.
Haiku 4.5 · ✓ Pass 11s · 42,203 tok · $0.012 · 2 LLM / 1 tool calls · transcript
Here is the OncoPrint for KRAS, TP53 and CDKN2A in Lung Adenocarcinoma (TCGA, PanCancer Atlas): [Open OncoPrint](https://www.cbioportal.org/results/oncoprint?cancer_study_list=luad_tcga_pan_can_atlas_2018&gene_list=KRAS%20TP53%20CDKN2A)
Sonnet 5 · ✗ Fail 12s · 90,436 tok · $0.044 · 3 LLM / 2 tool calls · transcript
Here is the OncoPrint for KRAS, TP53, and CDKN2A in Lung Adenocarcinoma (TCGA, PanCancer Atlas) — 566 samples: [Open OncoPrint](https://www.cbioportal.org/results/oncoprint?cancer_study_list=luad_tcga_pan_can_atlas_2018&gene_list=KRAS%0ATP53%0ACDKN2A&case_set_id=luad_tcga_pan_can_atlas_2018_all)
#1010 NavigationSurvival & outcomes · lgg_tcga_pan_can_atlas_2018 · follow-up (turn 2)option 1 please ✓13s ✓12s
- Conversation so far
- User: I'm interested in IDH1 in lower grade glioma. What can we look at?Assistant: In Brain Lower Grade Glioma (TCGA, PanCancer Atlas), IDH1 is mutated in most tumors. A few directions: 1. Compare overall survival of IDH1-mutant vs IDH1-wild-type patients 2. See which genes are co-mutated with IDH1 (e.g. TP53, ATRX, CIC) 3. Look at IDH1 mutation positions (R132H and others) in the mutation diagram Which would you like?
- Notes
- A correct answer must: give a cBioPortal group comparison (or results-view survival) link for lgg_tcga_pan_can_atlas_2018 comparing IDH1-mutant vs IDH1-wild-type patients, pointing to the Survival tab for the Kaplan-Meier curves and log-rank p-value. Must not: invent median survival times, hazard ratios or p-values, or ask again which option the user wants.
Haiku 4.5 · ✓ Pass 13s · 86,440 tok · $0.020 · 4 LLM / 3 tool calls · transcript
Perfect! Here's your survival comparison: **[View IDH1 Survival Analysis](https://www.cbioportal.org/comparison/survival?comparisonId=6ab6806fc2115c492d884f84)** **Group sizes:** - **IDH1 Mutant:** 395 samples - **IDH1 Wild-type:** 119 samples The Kaplan-Meier survival curves will show overall survival (OS) outcomes for each group. In lower grade glioma, IDH1 mutations are strongly associated with improved prognosis — you should see a noticeable survival advantage in the mutant group.
Sonnet 5 · ✓ Pass 12s · 81,928 tok · $0.036 · 3 LLM / 2 tool calls · transcript
[IDH1 mutant vs. wild-type survival comparison — LGG (TCGA, PanCancer Atlas)](https://www.cbioportal.org/results/comparison/survival?cancer_study_list=lgg_tcga_pan_can_atlas_2018&gene_list=IDH1%3A+MUT&case_set_id=lgg_tcga_pan_can_atlas_2018_all&profileFilter=mutations)
This opens the Kaplan-Meier survival comparison for the 514-sample Brain Lower Grade Glioma (TCGA, PanCancer Atlas) cohort, split into an **IDH1-mutant ("Altered")** group and an **IDH1-wild-type ("Unaltered")** group, restricted to samples with mutation profiling. On the page you can toggle between Overall Survival, Disease-Free Survival, Progression-Free Survival, and Disease-Specific Survival, and it reports log-rank p-values and can generate hazard ratios — this cohort is well known for IDH1-mutant patients showing markedly better survival, but let the plot on the page speak for that rather than taking my word for it.
If you'd like, I can also pull the raw patient counts (mutant vs. wild-type, events vs. censored) for a quick numeric summary alongside the plot — just let me know.
#1011 AnalysisSurvival & outcomes · coadread_tcga_pan_can_atlas_2018 · follow-up (turn 2)is the difference in survival significant? ✓19s ✓11s
- Conversation so far
- User: In TCGA colorectal cancer, do patients with high PTK2 expression have different survival from patients with low PTK2 expression?Assistant: In Colorectal Adenocarcinoma (TCGA, PanCancer Atlas), the Kaplan-Meier curves for patients with PTK2 mRNA above vs below the median look slightly apart, with the high-PTK2 group doing a little worse early on. You can build this comparison in cBioPortal by querying PTK2 with an expression threshold (e.g. `PTK2: EXP>0`) and opening the Comparison → Survival tab.
- Notes
- A correct answer must: say it cannot compute a significance test itself and point the user to the log-rank p-value on the Survival tab of the comparison link (or give one only if it comes from a tool that computed it). Must not: state a p-value, hazard ratio or "significant/not significant" conclusion it did not compute.
Haiku 4.5 · ✓ Pass 19s · 42,443 tok · $0.013 · 2 LLM / 1 tool calls · transcript
Based on the statistical-tests guide, **I cannot compute the p-value for you** — survival comparisons require a Kaplan-Meier test with log-rank, which ClickHouse cannot perform. However, I can point you to where to get the answer: **Use cBioPortal's Group Comparison → Survival tab** (recommended): 1. Go to the [TCGA Colorectal Adenocarcinoma study](https://www.cbioportal.org/study/summary?id=coadread_tcga_pan_can_atlas_2018) 2. Filter or compare PTK2 expression high (above median) vs. low (below median) 3. Open the **Survival** tab — it will show Kaplan-Meier curves with the log-rank p-value and median OS per group **Alternatively, compute it yourself in R or Python:** - **R:** Use `survival::survfit()` with `Surv(OS_MONTHS, OS_STATUS=="1:DECEASED")` to fit KM curves, then `survdiff()` for the log-rank test - **Python:** Use `lifelines.KaplanMeierFitter` and `logrank_test()` from `lifelines.statistics` The rough visual impression from my earlier response (high-PTK2 slightly worse early on) suggests a possible difference, but without the actual log-rank test statistic and p-value, you cannot say whether that difference is statistically significant or just random variation in the sample.
Shows internals to the user — backend jargon: ClickHouse
Sonnet 5 · ✓ Pass 11s · 82,481 tok · $0.036 · 3 LLM / 2 tool calls · transcript
[Compare survival: PTK2 z-score > 0 (high) vs. ≤ 0 (low) in Colorectal Adenocarcinoma (TCGA, PanCancer Atlas)](https://www.cbioportal.org/results/comparison/survival?cancer_study_list=coadread_tcga_pan_can_atlas_2018&gene_list=PTK2%3A+EXP%3E0&case_set_id=coadread_tcga_pan_can_atlas_2018_all&profileFilter=rna_seq_v2_mrna_median_Zscores) This link splits the 594-sample cohort into two groups by PTK2 mRNA z-score (above 0 = high, ≤ 0 = low, using the median-centered RNA-Seq z-scores) and opens the Survival tab, which will display the Kaplan-Meier curves for overall survival along with the log-rank test p-value directly on the plot. You can also switch to disease-free/progression-free survival endpoints from the same tab, and adjust the expression cutoff if you'd rather split by median expression exactly or use a different threshold.
#1012 Out of scopeOut of scope · All Studies · follow-up (turn 2)what literature? citations please ✓7s ✓20s
- Conversation so far
- User: How are PTK2 and STING1 expression related in colorectal adenocarcinoma?Assistant: In Colorectal Adenocarcinoma (TCGA, PanCancer Atlas), PTK2 and STING1 mRNA expression show only a weak correlation. Tumors with active FAK (PTK2) signaling are often reported in the literature to suppress innate immune signaling, which could relate to STING pathway activity.
- Notes
- A correct answer must: acknowledge that the literature remark was general background it cannot back with verified citations from cBioPortal, and either cite only sources it can verify (e.g. publications of cBioPortal studies) or suggest searching PubMed. Must not: produce specific paper titles, authors, journals, years, DOIs or PMIDs that it did not retrieve from a tool.
Haiku 4.5 · ✓ Pass 7s · 19,261 tok · $0.006 · 1 LLM / 0 tool calls · transcript
I apologize for the imprecision in my previous response. I should not have made that literature claim without having actually reviewed sources or having them available to cite. To be clear: **I cannot provide literature citations.** I'm a research tool built on cBioPortal genomics data, not a literature database. My previous statement about "tumors with active FAK signaling suppressing innate immune signaling" was general biological knowledge, not derived from cBioPortal data, and I should not have presented it without qualification. What I *can* do is: 1. **Query cBioPortal** for the actual correlation between PTK2 and STING1 mRNA expression in colorectal adenocarcinoma — if you want the exact numbers (correlation coefficient, sample size, p-value). 2. **Link you to cBioPortal's visualization** so you can explore the relationship interactively and then follow up in PubMed or your institution's literature database. Would you like me to: - Generate a direct cBioPortal link to the mRNA correlation plot for PTK2 vs STING1 in colorectal adenocarcinoma (TCGA)? - Query the expression data and report the actual correlation statistics? For literature, I'd recommend searching PubMed directly for "PTK2 STING1 immune" or "FAK STING colorectal" to find peer-reviewed work on the relationship.
Sonnet 5 · ✓ Pass 20s · 84,042 tok · $0.045 · 3 LLM / 2 tool calls · transcript
You're right to push on that — I should have flagged it immediately: **I have no literature search capability, so I cannot provide real citations.** The "literature reports" line in my previous answer was unsupported, general background phrasing I should not have included without clearly labeling it as unverified. I don't want to compound that by inventing paper titles or authors now, so I won't. What I *can* stand behind is the cBioPortal data itself. Here's the direct scatter plot of PTK2 vs. STING1 mRNA expression (z-scores, RNA-Seq) across the 594 samples in the Colorectal Adenocarcinoma (TCGA, PanCancer Atlas) cohort, so you can inspect the correlation coefficient and distribution yourself: [PTK2 vs. STING1 mRNA expression — Colorectal Adenocarcinoma (TCGA, PanCancer Atlas)](https://www.cbioportal.org/results/plots?cancer_study_list=coadread_tcga_pan_can_atlas_2018&gene_list=PTK2%0ASTING1&case_set_id=coadread_tcga_pan_can_atlas_2018_all&profileFilter=rna_seq_v2_mrna_median_all_sample_Zscores&plots_horz_selection=%7B%22selectedGeneOption%22%3A%225747%22%2C%22dataType%22%3A%22mrna%22%7D&plots_vert_selection=%7B%22selectedGeneOption%22%3A%22340061%22%2C%22dataType%22%3A%22mrna%22%7D) The plot will display the Pearson/Spearman correlation directly on the page. For the mechanistic question about FAK–STING crosstalk, I'd recommend searching PubMed or Google Scholar directly (e.g., "PTK2 FAK STING1 cGAS colorectal") rather than relying on my output for citations.
#1013 Out of scopeOut of scope · All Studies · follow-up (turn 2)what if the genes, like AC002511.1 or LOC100506271, just aren't in HGNC or NCBI at all? ✗11s ✓14s
- Conversation so far
- User: I'm importing a study into my own cBioPortal instance and the validator says some genes are not known to cBioPortal. What does that mean?Assistant: The validator checks every Hugo_Symbol / Entrez_Gene_Id in your data files against the gene table loaded in your instance (built from HGNC/NCBI). "Gene not known" means the symbol or ID isn't in that table — usually an outdated symbol, an alias, or a typo.
- Notes
- A correct answer must: explain that genes absent from the instance's gene table can't be loaded (the importer skips those rows; the validator warns), and suggest mapping them to a current HGNC symbol / Entrez ID where one exists or accepting that they will be left out. Must not: invent validator/importer flags or config options, or claim cBioPortal can load arbitrary unknown gene symbols as-is.
Haiku 4.5 · ✗ Fail 11s · 19,492 tok · $0.007 · 1 LLM / 0 tool calls · transcript
Those are non-coding RNA genes or pseudogenes that HGNC/NCBI either don't cover comprehensively or classify differently. Here's what's happening and how to handle it: ## Why They're Not in the Standard Gene Tables - **AC002511.1, LOC100506271** — these are NCBI RefSeq identifiers for long non-coding RNAs (lncRNAs) or predicted genes that don't have official HUGO symbols - HGNC focuses on protein-coding genes and well-characterized non-coding RNAs; many predicted or poorly characterized genes fall outside that scope - NCBI includes them but may not assign stable HUGO symbols ## Solutions for Your Import **Option 1: Map to HUGO symbols (recommended)** - Query Ensembl or NCBI directly to see if these have HGNC-approved symbols - Example: AC002511.1 might map to an HGNC lncRNA name like `LINC00000` - Update your file with the official symbol if one exists **Option 2: Use Entrez Gene IDs instead** - If you have the Entrez ID (e.g., 100506271), the validator may accept it as the primary identifier - cBioPortal can resolve Entrez IDs even when the symbol is ambiguous **Option 3: Add them to your gene table** - Edit the `cancer_gene_list.txt` file in your instance (or equivalent gene reference) to include these entries with their Entrez IDs and a descriptive name - The validator will then recognize them **Option 4: Filter them out** - If these genes aren't clinically important for your study, exclude them from your import file and keep only genes cBioPortal recognizes ## Which Approach? - **If they're actually mutated/altered in your data:** use Option 1 or 3 — map them properly so they're meaningful in analysis - **If they're rare or background noise:** Option 4 is fine - **For lncRNAs with clinical relevance:** Option 3 (manually add to your gene list) is most reliable What's the source of your data? (VEP, MAF, a custom pipeline?) That context might clarify whether these genes should be kept or are annotation artifacts.
Sonnet 5 · ✓ Pass 14s · 26,339 tok · $0.018 · 1 LLM / 0 tool calls · transcript
That's a real gap, not a validator error — a small fraction of symbols in genomic data files (LOC placeholders, some readthrough/antisense transcripts like AC002511.1, and other provisional loci) genuinely aren't in the standard gene reference tables cBioPortal ships with, because they were never assigned a stable HGNC symbol or are Ensembl/RefSeq-only identifiers. A few things worth knowing: - **This is a known, common situation.** Genome annotations (Ensembl/RefSeq/GENCODE) contain tens of thousands of predicted or provisional loci that HGNC hasn't curated with an official symbol. LOC IDs are NCBI/Entrez placeholder names for exactly these cases. - **cBioPortal's validator is strict by design** — it flags anything not in the local `gene` table rather than guessing, since silently dropping or mismapping genes would corrupt downstream analyses. - **Your options for your own instance:** 1. **Drop these genes from the file** if they're not central to your analysis — most are low-evidence/uncharacterized loci and won't affect typical mutation/CNA summaries. 2. **Add them to your instance's gene table yourself.** Since you control your own cBioPortal deployment, you can insert custom gene records (symbol, Entrez ID if one exists, chromosome/coordinates) directly into the seed database before import, so the validator recognizes them going forward. This is a database administration step on your instance, not something done through the portal UI. 3. **Map to an approved alias** if the gene actually does have a current HGNC symbol under a different name — check the symbol on the HGNC website (genenames.org) or NCBI Gene, since some "unknown" hits are just outdated aliases rather than truly absent genes. 4. **Check the cBioPortal documentation/GitHub** for the gene table update scripts — the cBioPortal team periodically refreshes the reference gene set from NCBI, so it's worth confirming you're on a current gene table build before assuming a symbol is missing entirely. This is a genuinely out-of-scope area for me to dig further into — I only have read access to the public cBioPortal database, not to your local instance's admin tooling or the gene-table update scripts. For the exact steps to add custom genes to a self-hosted instance, the cBioPortal GitHub repository/documentation or their Google Group/Slack support channel would be the right place to ask.
#1014 AnalysisExpression & multi-omics · brca_metabric · follow-up (turn 2)Can you do this with METABRIC instead? ✓62s ✓10s
- Conversation so far
- User: How does ESR1 mRNA expression differ between PAM50 subtypes in TCGA breast cancer?Assistant: In Breast Invasive Carcinoma (TCGA, PanCancer Atlas), ESR1 mRNA is highest in Luminal A and Luminal B tumors, intermediate in Normal-like and HER2-enriched, and lowest in Basal-like tumors. You can see the distribution here: [ESR1 expression by subtype](https://www.cbioportal.org/results/plots?cancer_study_list=brca_tcga_pan_can_atlas_2018&gene_list=ESR1)
- Notes
- A correct answer must: repeat the ESR1 mRNA-by-PAM50-subtype comparison in brca_metabric (its Pam50 + Claudin-low subtype attribute), reaching the same pattern (Luminal highest, Basal lowest), with a link to the METABRIC study or plots view. Must not: stay on the TCGA study, or claim a statistically significant difference (p-value) without a test having been run.
Haiku 4.5 · ✓ Pass 62s · 402,389 tok · $0.078 · 16 LLM / 17 tool calls · transcript
## ESR1 mRNA Expression by Claudin Subtype in METABRIC METABRIC breast cancer uses **Claudin subtyping**, not PAM50. Here's the ESR1 expression distribution across subtypes (1,980 samples with mRNA data): | Subtype | Samples | Median ESR1 | Mean ESR1 | |---------|---------|------------|----------| | **Luminal B** | 475 | 11.15 | 11.00 | | **Luminal A** | 700 | 10.93 | 10.81 | | Normal-like | 148 | 9.76 | 9.47 | | HER2-enriched | 224 | 7.10 | 7.80 | | Claudin-low | 218 | 6.62 | 7.41 | | **Basal** | 209 | 5.99 | 6.42 | **Key finding:** ESR1 is highest in Luminal A/B tumors (median 10.9–11.2, consistent with hormone receptor expression), intermediate in HER2-enriched and normal-like, and lowest in Basal-like (median 6.0). [View plots with ESR1 expression by Claudin subtype](https://www.cbioportal.org/study/plots?id=brca_metabric&filterAttributeId=CLAUDIN_SUBTYPE&plots_vert_selection=%7B%22selectedGeneOption%22%3A%222099%22%2C%22dataType%22%3A%22mrna%22%7D)
Sonnet 5 · ✓ Pass 10s · 52,479 tok · $0.022 · 2 LLM / 1 tool calls · transcript
Note: METABRIC uses the "CLAUDIN_SUBTYPE" attribute rather than a dedicated "PAM50" field — in this dataset it captures the same intrinsic molecular subtypes (Luminal A, Luminal B, HER2-enriched, Basal-like, Normal-like), plus a Claudin-low group. Here is ESR1 mRNA (microarray z-scores) plotted by subtype in the METABRIC cohort (2,509 samples): [ESR1 expression by CLAUDIN_SUBTYPE — METABRIC](https://www.cbioportal.org/results/plots?cancer_study_list=brca_metabric&gene_list=ESR1) Once loaded, set the horizontal/grouping axis to "CLAUDIN_SUBTYPE" (Clinical Data) to reproduce the same subtype comparison as before. You should expect the same general pattern — highest ESR1 in Luminal A/B, lowest in Basal-like — but you can confirm the exact distribution directly in the plot. If you'd like, I can also pull the underlying group-level counts/medians per subtype instead — just let me know.



