← All runs and checks

Beta vs prod database MCP — 2026-10-02

Run comparisonThe same 146 questions, asked of the same agent with the same models, once against the beta database MCP and once against the prod database MCP. Prod is the baseline; every delta reads prod → beta.

Read these before the numbers.

Headline: prod → beta

Haiku 4.5Sonnet 5
MetricprodbetaΔprodbetaΔ
Pass rate (graded answers)64.4%56.2%−8.2 pt62.3%69.9%+7.5 pt
Precision (passes / attempted)65.3%56.9%−8.3 pt63.2%71.8%+8.6 pt
Median latency24.3 s24.9 s+0.6 s30.1 s27.5 s−2.6 s
p90 latency69.9 s67.0 s−2.8 s56.9 s58.9 s+2.0 s
Answers under 10 s8.9%11.6%+2.7 pt2.7%4.1%+1.4 pt
Median LLM calls / answer66±065−1
Questions better / worse / same on beta15 better · 27 worse · 104 same26 better · 15 worse · 105 same

146 questions × 1 repeat per model, all graded, no failed requests. Green/red follow the compare's own verdict (better/worse for beta); with one repeat, ±8 pt of pass rate is within noise. “Better/worse” per question means its pass/fail outcome changed between the runs. Full per-question, per-category and per-track detail: Haiku 4.5 compare · Sonnet 5 compare.

The compare pages also list Tool rounds / answer as 0.00 for both runs. That metric is not populated for the Claude Code runner (it comes from LibreChat/Langfuse traces), so it is not a result.

26-question subset

Pending. Only the beta subset run is published here. Prod subset run 20261002-2359 is still being graded. There is no prod counterpart or subset comparison published yet.

Beta, 26 questionsAnswersPass ratePrecisionMedian latencyp90 latencyMedian LLM calls
Haiku 4.52653.8%60.9%43.7 s85.8 s10.5
Sonnet 52669.2%72.0%40.0 s79.7 s6

Beta subset report (20261002-1859). The subset is a different question set (input/questions-subset.yaml, statistics over a subset of samples or patients that per-study precomputed tables and purpose-built tools don’t cover), so these numbers are not comparable with the 146-question table above. This run used a different setup from the main runs: a Pro login with no built-in plugins allowed, and the server-instructions probe failed at start-up, so the run doesn't record the beta MCP's instructions hash.

Setup and recorded metadata

Reports and data