Fifteen questions with independently verified answers, put to four models eleven ways — answering alone, with live web search, and with Pipeworx — priced at what a customer pays.
More accurate and cheaper
What each arm costs, and what it is made of
| Arm | Correct | Agent tokens | Pipeworx requests | Total | Per answer | Pipeworx is … cheaper |
|---|---|---|---|---|---|---|
| Fable 5 alone | 2/15 | $0.2061 | — | $0.2061 | $0.0137 | 3× |
| Opus 5 alone | 2/15 | $0.1594 | — | $0.1594 | $0.0106 | 2× |
| GPT-5 alone | 1/14 | $0.2051 | — | $0.2051 | $0.0137 | 3× |
| DS Flash alone | 2/15 | $0.0012 | — | $0.0012 | $0.00008 | 0× |
| Fable 5 + web search | 6/15 | $17.98 | — | $17.98 | $1.20 | 228× |
| Opus 5 + web search | 8/15 | $7.07 | — | $7.07 | $0.4710 | 90× |
| GPT-5 + web search | 10/15 | $2.42 | — | $2.42 | $0.1617 | 31× |
| Fable 5 + Pipeworx | 13/15 | $1.07 | 17 × $0.005 = $0.0850 | $1.15 | $0.0769 | 15× |
| Opus 5 + Pipeworx | 12/15 | $0.8165 | 20 × $0.005 = $0.1000 | $0.9165 | $0.0611 | 12× |
| GPT-5 + Pipeworx | 12/12 | $0.7126 | 35 × $0.005 = $0.1750 | $0.8876 | $0.0592 | 11× |
| DS Flash + Pipeworx | 13/15 | $0.0039 | 15 × $0.005 = $0.0750 | $0.0789 | $0.0053 | 1× |
What changes when you connect Pipeworx
| Model | Alone | + Web search | + Pipeworx | × vs Pipeworx + DS Flash | |||
|---|---|---|---|---|---|---|---|
| Correct | Cost | Correct | Cost | Correct | Cost | alone / search / Pipeworx | |
| Fable 5 $10 / $50/Mtok | 2/15 | $0.2061 | 6/15 | $17.98 | 13/15 | $1.15 | 3× / 228× / 15× |
| Opus 5 $5 / $25/Mtok | 2/15 | $0.1594 | 8/15 | $7.07 | 12/15 | $0.9165 | 2× / 90× / 12× |
| GPT-5 $1.25 / $10/Mtok | 1/14 | $0.2051 | 10/15 | $2.42 | 12/12 | $0.8876 | 3× / 31× / 11× |
| DeepSeek V4 Flash $0.09 / $0.19 | 2/15 | $0.0012 | — | — | 13/15 | $0.0789 | 0.0× / — / 1× |
The bill for fifteen questions
Who got each question right
| Question | Answering alone | With web search | With Pipeworx | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Fable 5 | Opus 5 | GPT-5 | DS Flash | Fable 5 | Opus 5 | GPT-5 | Fable 5 | Opus 5 | GPT-5 | DS Flash | |
| 1 Apple's most recent 10-Q — date and periodmovedanswer: Filed 2026-05-01 → 2026-07-31 (new filing) | ✗refused | ✗refused | ✗refused | ✗ | ✗ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| 2 Most recent US unemployment rate (U-3)answer: 4.2%, June 2026 | ✗ | ✗refused | ✗refused | ✗ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| 3 NVIDIA revenue, most recent FY, to the millionanswer: $215,938 million | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| 4 CPI-U index level, seasonally adjustedanswer: 332.568, June 2026 | ✗refused | ✗refused | ✗refused | ✗refused | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| 12 Stripe's annual revenue from its 10-Ktrapanswer: No 10-K exists — Stripe is private | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| P1 Recruiting Phase 3 semaglutide studiesanswer: 29 | ✗refused | ✗refused | ✗refused | ✗refused | ✗refused | ✗refused | ✗ | ✓ | ✓ | ✓ | ✓ |
| P2 Semaglutide FDA shortage statusanswer: Tablets listed, To Be Discontinued | ✗refused | ✗ | ✗refused | ✗refused | ✗ | ✗ | ✗ | ✓ | ✓ | ✓ | ✓ |
| P3 More recruiting Phase 3 obesity trials — Novo or Lilly?answer: Eli Lilly: 12 vs 10 | ✗refused | ✗refused | ✗refused | ✗ | ✗refused | ✗ | ✗ | ✓ | ✓ | —no answer | ✗refused |
| P4 Who markets generic semaglutidetrapanswer: Nobody — Apotex tentative only | ✓ | ✓ | —no answer | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| P5 Eli Lilly 8-K filings in the last 90 daysanswer: 2; most recent 2026-05-20 | ✗refused | ✗refused | ✗refused | ✗refused | ✓ | ✓ | ✓ | ✗ | ✗ | ✓ | ✓ |
| V1 M5.0+ earthquakes worldwide, last 7 daysmovedanswer: 62 → 65 (live count) | ✗refused | ✗refused | ✗refused | ✗refused | ✗refused | ✗ | ✗refused | ✓ | ✓ | —no answer | ✓ |
| V2 Current temperature at SFO airport (KSFO)answer: moves hourly | ✗refused | ✗refused | ✗refused | ✗refused | ✗ | ✗ | ✓ | ✗ | ✗ | ✓ | ✓ |
| V3 Orbital launches in the last 7 daysanswer: 5 | ✗refused | ✗refused | ✗refused | ✗refused | ✗ | ✗ | ✗ | ✓ | ✗ | —no answer | ✗refused |
| V4 ECB euro reference rate for the US dollarmovedanswer: 1.1476 → 1.1485 (new rate) | ✗refused | ✗refused | ✗refused | ✗refused | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| V5 San Francisco Giants win-loss recordmovedanswer: 46-62 → 47-62 (they won) | ✗refused | ✗refused | ✗refused | ✗refused | ✗ | ✗ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Correct | 2/15 | 2/15 | 1/14 | 2/15 | 6/15 | 8/15 | 10/15 | 13/15 | 12/15 | 12/12 | 13/15 |
Three things the numbers say
1 · The answer is not in the model
Answering alone, the three frontier models managed 2, 2 and 1 of 15. Most attempts were refusals — “I don't have real-time access”, “check EDGAR”, “check the BLS website”. Where they did answer they were confidently wrong: every model alone missed NVIDIA's revenue, quoting last year's $130,497M against an actual $215,938M.
2 · Web search fixes some of it, expensively
Search lifted them to 6, 8 and 10 of 15 — real improvement — at $17.98, $7.07 and $2.42 for fifteen questions. Fable 5 spent 480,155 input tokens on a single launch-count question and got it wrong. And on the FDA shortage question all three search arms were more confidently wrong than without it: the web repeats the 2025 “shortage resolved” story, while the current FDA record is a database row nobody wrote an article about.
3 · Pipeworx makes them cheaper and better at once
Connected as an MCP tool, the same models reach 13, 12 and 12 correct while the bill falls to $1.15, $0.9165 and $0.8876. One call against a record replaces a dozen searches and the transcript they drag in.
And then the model stops mattering much. DeepSeek V4 Flash — a model roughly a hundred times cheaper per token than Fable 5 — matches it exactly once both are grounded: 13/15 against 13/15, for $0.0789 against $1.15.
The benchmark's own answers went stale overnight
Four of the fifteen true answers changed in 24 hours
Between the two runs Apple filed a new 10-Q (2026-07-31, period ending 2026-06-27), the ECB published a new reference rate (1.1485, up from 1.1476), the Giants won a game (47-62, from 46-62), and the rolling earthquake count moved from 62 to 65.
Every Pipeworx-connected arm tracked all four. Nothing was retrained, reindexed or re-cached — they asked, and got today's record. This was not designed into the experiment; it happened while it ran, and it is the clearest statement of the problem the product solves.
Where we lost, stated plainly
Our airport-code resolver silently answered about a different place
Asked for the temperature at KSFO, Fable 5 and Opus 5 with Pipeworx both reported 58.4°F against an actual 64.4°F. Both explained why, unprompted: the airport code did not resolve, so they fell back to a city lookup for “San Francisco”. GPT-5 got it right by reaching a METAR instead.
The models behaved well — they degraded, said so, and were still wrong. Answering about a nearby city without saying so is the same silent-substitution class as the stale-season default we found and fixed mid-benchmark.
Two more, unfixed
Comparison questions get one call, not two. Asked which sponsor runs more Phase 3 obesity trials, the cheap grounded arm returned only Novo's count. The frontier models with Pipeworx solved it by calling twice on their own initiative — which is exactly the gap: the orchestration, not the data.
“Launches that occurred” routes to a news feed rather than launch records. The model caught it out loud: “the total: 4 in the feed is the number of articles, not launches.”
If you take five lines to a slide
- Frontier models answering alone got 1–2 of 15 live-data questions. The answer is not in the model.
- Web search helped and cost up to $17.98 for fifteen questions.
- Pipeworx made every model cheaper and more accurate at once — Fable 5 16× cheaper than search, and better.
- Grounded, the cheapest model matches the most expensive: 13/15 for $0.0789 vs 13/15 for $1.15.
- Four true answers changed overnight. Only the grounded arms noticed.
Caveats
Fifteen questions is enough to see cost gaps of this size and not enough for a precise accuracy claim. The experiment-1 arms were graded by us and then re-graded blind by an independent model with arm labels stripped and answer order shuffled — 91.5% agreement, and every disagreement was resolved in the judge's favour or by re-running. The MCP arms have not yet had that treatment. The question set is built around live data on purpose: it measures the freshness gap, not general intelligence. Four questions decay within hours and are graded against ground truth captured the same day. GPT-5 with Pipeworx hit the harness's five-hop tool-loop cap on three questions and returned no answer, so its denominator is 12.