cBioPortalChat benchmark runs

Each run asks every question and grades the answers pass/fail. The runner is either LibreChat (the deployed agent through its Agents API) or Claude Code (headless, with the agent's prompt and the same MCP servers) — compare runs with the same runner. Each test set has its own table — compare runs within a table, not across; the test sets page describes each set and lists its questions. Newest first.

Main benchmark

Single questions covering data lookups, analysis, navigation links and out-of-scope requests. Questions in this set · input/questions.yaml

changed marks a setup value that differs from the run below it in this table (hover for the previous value).

RunTargetSetupQuestionsModelPass rate DataNavigationAnalysisOut of scope Cost / correctMedian latencyFailed requests
20261002-2344 beta
Claude Code 2.1.288 changed
cbioportal-mcp 2.13.3 68b96a0
navigator 1.0.0 b4ff15f
cBioPortal v7.1.2 / DB 3.0.0 / hgnc_v7_2025.10.7
146 (146 graded) Haiku 4.5 64% 68%35%74%75% $0.09524s0
Sonnet 5 62% 55%46%83%58% $0.22630s0
20261002-0306 beta
Claude Code 2.1.287 changed
cbioportal-mcp 2.13.3 68b96a0 changed
navigator 1.0.0 b4ff15f
cBioPortal v7.1.2 / DB 3.0.0 / hgnc_v7_2025.10.7
146 (146 graded) Haiku 4.5 56% 69%31%57%42% $0.10625s0
Sonnet 5 70% 74%46%80%58% $0.18727s0
20260929-0440 beta
Claude Code 2.1.284 changed
navigator 1.0.0 b4ff15f
cBioPortal v7.1.2 / DB 3.0.0 / hgnc_v7_2025.10.7
146 (145 graded) Sonnet 5.5 72% 87%46%69%67% $0.14822s0
20260926-1634 beta
Claude Code 2.1.283 changed
cbioportal-mcp 07ffe44 changed
navigator 1.0.0 b4ff15f changed
cBioPortal v7.1.2 / DB 3.0.0 / hgnc_v7_2025.10.7
146 (145 graded) Haiku 4.5 59% 77%31%49%67% $0.10434s0
Sonnet 5 74% 82%50%76%83% $0.18936s0
20260925-0129 beta
Claude Code 2.1.282
prompt 5dbe6fa4ced1 changed
cbioportal-mcp a06ebc2
navigator 1.0.0 4135b03
cBioPortal v7.1.2 / DB 3.0.0 / hgnc_v7_2025.10.7
146 (145 graded) Haiku 4.5 53% 77%27%33%58% $0.12329s0
Sonnet 5 65% 74%46%62%67% $0.20534s0
20260924-2354 prod
Claude Code 2.1.282 changed
cbioportal-mcp a06ebc2
navigator 1.0.0 4135b03
cBioPortal v7.1.2 / DB 3.0.0 / hgnc_v7_2025.10.7
146 (145 graded) Haiku 4.5 48% 66%38%31%33% $0.14231s0
Sonnet 5 57% 63%50%51%67% $0.25539s0
20260923-1919 beta
LibreChat
146 (145 graded) Haiku 4.5 54% 66%46%40%67% $0.17730s0
Sonnet 5 64% 79%69%44%50% $0.45172s1

Multi-turn follow-ups

Follow-up messages in a scripted conversation, modeled on real traffic. Smaller set; scores are not comparable with the main benchmark. Questions in this set · input/questions-multiturn.yaml

changed marks a setup value that differs from the run below it in this table (hover for the previous value).

RunTargetSetupQuestionsModelPass rate DataNavigationAnalysisOut of scope Cost / correctMedian latencyFailed requests
20260929-0459 beta
Claude Code 2.1.284 changed
navigator 1.0.0 b4ff15f
cBioPortal v7.1.2 / DB 3.0.0 / hgnc_v7_2025.10.7
14 (14 graded) Sonnet 5.5 93% 100%100%100%50% $0.09019s0
20260926-1702 beta
Claude Code 2.1.283 changed
cbioportal-mcp 07ffe44 changed
navigator 1.0.0 b4ff15f changed
cBioPortal v7.1.2 / DB 3.0.0 / hgnc_v7_2025.10.7
14 (14 graded) Haiku 4.5 79% 62%100%100%100% $0.04018s0
Sonnet 5 93% 100%100%50%100% $0.09122s0
20260925-1407 beta
Claude Code 2.1.282
cbioportal-mcp a06ebc2
navigator 1.0.0 4135b03
cBioPortal v7.1.2 / DB 3.0.0 / hgnc_v7_2025.10.7
14 (14 graded) Haiku 4.5 93% 100%100%100%50% $0.04220s0
Sonnet 5 71% 62%50%100%100% $0.12819s0

Subset questions

Statistics over a subset of samples or patients that per-study precomputed tables don't cover. Scores are not comparable with the main benchmark. Questions in this set · input/questions-subset.yaml

changed marks a setup value that differs from the run below it in this table (hover for the previous value).

RunTargetSetupQuestionsModelPass rate DataNavigationAnalysisOut of scope Cost / correctMedian latencyFailed requests
20261002-1859 beta
Claude Code 2.1.287
cbioportal-mcp 68b96a0
navigator 1.0.0 b4ff15f
cBioPortal v7.1.2 / DB 3.0.0 / hgnc_v7_2025.10.7
26 (26 graded) Haiku 4.5 54% 53%–57%– $0.16744s0
Sonnet 5 69% 79%–43%– $0.30240s0

Comparisons and checks

Not single benchmark runs: summaries that compare published runs, and checks that measure the database behind the MCP server directly without asking the agent any questions. The kind column says which. Newest first.

PageDateKindSummary
Beta vs prod database MCP — 2026-10-022026-10-02Run comparison Same agent and models, beta vs prod database MCP, 146 questions × 1 repeat (Claude Code runner). Pass rate prod→beta: Haiku 4.5 64.4%→56.2%, Sonnet 5 62.3%→69.9% — within ±8 pt run-to-run noise. Prod subset 20261002-2359 pending (still being graded).
Live ClickHouse latency check — 2026-09-292026-09-29DB latency check Read-only timings of the SQL in open cbioportal-mcp PRs #151, #154 and #160 (not merged or deployed) on the current public DB. #160 saves 0.24–0.60 s per call (median 0.29 s, identical results); #154 fallback SQL takes 0.08–1.35 s vs an estimated ~3–4 ms lookup; #151 schema calls cost 1.4–2.6 ms. DB share of answer latency not measured.