Beta vs prod database MCP — 2026-10-02
Read these before the numbers.
- One repeat each. Run-to-run noise on this set is about ±8 points of pass rate, so the observed pass-rate deltas (Haiku −8.2 pt, Sonnet +7.5 pt) are within noise. A single repeat cannot establish a regression or improvement.
- Local Claude Code runner on the Anthropic API, not LibreChat on Bedrock. The handoff router, fast path, tool-round cap and per-agent thinking settings of the deployed chat are not tested, and latencies are not comparable to the 2026-09-23 LibreChat run.
- Both runs use the single shared agent (prompt
5dbe6fa4ced1), whose instructions (stored in LibreChat's MongoDB) still say to read the matching guide and check table columns before any SQL. - The only intended difference is the database MCP: beta =
cbioportal/mcp:betaon the beta_green database; prod =cbioportal/mcp:latest. Both were reached through a port-forward (shown in the reports as “beta MCP (port-forward)” / “prod MCP (port-forward)”). The navigator MCP is the public one in both runs. - Judge:
claude-code:claude-sonnet-4-6for every grade in all three runs.
Headline: prod → beta
| Haiku 4.5 | Sonnet 5 | |||||
|---|---|---|---|---|---|---|
| Metric | prod | beta | Δ | prod | beta | Δ |
| Pass rate (graded answers) | 64.4% | 56.2% | −8.2 pt | 62.3% | 69.9% | +7.5 pt |
| Precision (passes / attempted) | 65.3% | 56.9% | −8.3 pt | 63.2% | 71.8% | +8.6 pt |
| Median latency | 24.3 s | 24.9 s | +0.6 s | 30.1 s | 27.5 s | −2.6 s |
| p90 latency | 69.9 s | 67.0 s | −2.8 s | 56.9 s | 58.9 s | +2.0 s |
| Answers under 10 s | 8.9% | 11.6% | +2.7 pt | 2.7% | 4.1% | +1.4 pt |
| Median LLM calls / answer | 6 | 6 | ±0 | 6 | 5 | −1 |
| Questions better / worse / same on beta | 15 better · 27 worse · 104 same | 26 better · 15 worse · 105 same | ||||
146 questions × 1 repeat per model, all graded, no failed requests. Green/red follow the compare's own verdict (better/worse for beta); with one repeat, ±8 pt of pass rate is within noise. “Better/worse” per question means its pass/fail outcome changed between the runs. Full per-question, per-category and per-track detail: Haiku 4.5 compare · Sonnet 5 compare.
The compare pages also list Tool rounds / answer as 0.00 for both runs. That metric is not populated for the Claude Code runner (it comes from LibreChat/Langfuse traces), so it is not a result.
26-question subset
Pending. Only the beta subset run is published here. Prod subset run 20261002-2359 is still being graded. There is no prod counterpart or subset comparison published yet.
| Beta, 26 questions | Answers | Pass rate | Precision | Median latency | p90 latency | Median LLM calls |
|---|---|---|---|---|---|---|
| Haiku 4.5 | 26 | 53.8% | 60.9% | 43.7 s | 85.8 s | 10.5 |
| Sonnet 5 | 26 | 69.2% | 72.0% | 40.0 s | 79.7 s | 6 |
Beta subset report (20261002-1859). The subset is a different question set (input/questions-subset.yaml, statistics over a subset of samples or patients that per-study precomputed tables and purpose-built tools don’t cover), so these numbers are not comparable with the 146-question table above. This run used a different setup from the main runs: a Pro login with no built-in plugins allowed, and the server-instructions probe failed at start-up, so the run doesn't record the beta MCP's instructions hash.
Setup and recorded metadata
- Runs: prod 20261002-2344 (started 2026-10-02 23:44 UTC) and beta 20261002-0306 (started 2026-10-02 03:06 UTC); question file
input/questions.yaml; models Haiku 4.5 and Sonnet 5; 1 repeat. - Database MCP instructions as served: prod 19,962 characters, beta 38,583 characters. Both servers report version 2.13.3.
- Recorded image metadata is wrong for the prod run. Its
run.jsonlistscbioportal/mcp:betawith the same digest as the beta run, because the version probe reads the beta Kubernetes deployment whatever MCP the run actually used. The different served instructions above (and beta-only tools appearing only in the beta run) show the two runs did reach different servers. - Claude Code: 2.1.288 (prod run), 2.1.287 (beta run), 2.1.287 (beta subset).
- Both main runs list target
beta: for the Claude Code runner that only selects which LibreChat agent's prompt is used. It is the same shared agent in both.
Reports and data
- Run reports: prod 20261002-2344 · beta 20261002-0306 · beta subset 20261002-1859. Each folder also has
run.json(every answer, trace and grade),summary.json, referenced screenshots and the shared agent prompt linked by the index. Raw transcript exports are omitted. - Compares (prod = A, beta = B): Haiku 4.5 HTML · Markdown · JSON; Sonnet 5 HTML · Markdown · JSON.
- Prod subset run and subset compares: pending — prod subset
20261002-2359is still being graded. - Published copies are scrubbed: local file paths, the ClickHouse host and patient/sample identifiers that tool results returned have been replaced by placeholders such as
[sample id]. Screenshots containing identifiers or contact details are omitted.