Fifteen questions with independently verified answers, put to four models eleven ways — answering alone, with live web search, and with Pipeworx — priced at what a customer pays.

Runs 2026-07-30 and 2026-07-31 · every cost is the API's own billed usage · ground truth from SEC, BLS, ECB, USGS, NWS, ClinicalTrials.gov, openFDA and ESPN — never from Pipeworx

More accurate and cheaper

More accurate and 31×–228× cheaper
DeepSeek V4 Flash through Pipeworx got 13 of 15 right. The best web-search arm got 10, the worst 6 — and they cost $0.1617–$1.20 an answer against our $0.0053, agent tokens and Pipeworx at list price included.
Every price on this page is what a customer pays, Pipeworx billed at list — not what any of it costs us. Fifteen questions, answers verified against the primary source the same day.
Fable 5
alone
2/15
$0.2061
Opus 5
alone
2/15
$0.1594
GPT-5
alone
1/14
$0.2051
DS Flash
alone
2/15
$0.0012
Fable 5
+ web search
6/15
$17.98
Opus 5
+ web search
8/15
$7.07
GPT-5
+ web search
10/15
$2.42
Fable 5
+ Pipeworx
13/15
$1.15
Opus 5
+ Pipeworx
12/15
$0.9165
GPT-5
+ Pipeworx
12/12
$0.8876
DS Flash
+ Pipeworx
13/15
$0.0789

What each arm costs, and what it is made of

ArmCorrectAgent tokensPipeworx requestsTotalPer answerPipeworx is … cheaper
Fable 5 alone 2/15$0.2061 $0.2061 $0.0137
Opus 5 alone 2/15$0.1594 $0.1594 $0.0106
GPT-5 alone 1/14$0.2051 $0.2051 $0.0137
DS Flash alone 2/15$0.0012 $0.0012 $0.00008
Fable 5 + web search 6/15$17.98 $17.98 $1.20 228×
Opus 5 + web search 8/15$7.07 $7.07 $0.4710 90×
GPT-5 + web search 10/15$2.42 $2.42 $0.1617 31×
Fable 5 + Pipeworx 13/15$1.07 17 × $0.005 = $0.0850$1.15 $0.0769 15×
Opus 5 + Pipeworx 12/15$0.8165 20 × $0.005 = $0.1000$0.9165 $0.0611 12×
GPT-5 + Pipeworx 12/12$0.7126 35 × $0.005 = $0.1750$0.8876 $0.0592 11×
DS Flash + Pipeworx 13/15$0.0039 15 × $0.005 = $0.0750$0.0789 $0.0053

Pipeworx priced at list — 50 credits ($0.005) per request, the first volume bracket. The grounded DS Flash arm makes one ask_pipeworx call per question; the MCP arms make as many as the model chose, which is why their request bill differs. Agent tokens are each API’s own billed usage, web searches included at $0.01 each.

What changes when you connect Pipeworx

ModelAlone+ Web search+ Pipeworx× vs Pipeworx + DS Flash
CorrectCostCorrectCostCorrectCostalone / search / Pipeworx
Fable 5 $10 / $50/Mtok 2/15$0.2061 6/15$17.98 13/15$1.15 3× / 228× / 15×
Opus 5 $5 / $25/Mtok 2/15$0.1594 8/15$7.07 12/15$0.9165 2× / 90× / 12×
GPT-5 $1.25 / $10/Mtok 1/14$0.2051 10/15$2.42 12/12$0.8876 3× / 31× / 11×
DeepSeek V4 Flash $0.09 / $0.19 2/15$0.0012 13/15$0.0789 0.0× / — /

The bill for fifteen questions

Fable 5 alone
$0.2061 · 2/15 ·
Opus 5 alone
$0.1594 · 2/15 ·
GPT-5 alone
$0.2051 · 1/14 ·
DS Flash alone
$0.0012 · 2/15 ·
Fable 5 + web search
$17.98 · 6/15 · 228×
Opus 5 + web search
$7.07 · 8/15 · 90×
GPT-5 + web search
$2.42 · 10/15 · 31×
Fable 5 + Pipeworx
$1.15 · 13/15 · 15×
Opus 5 + Pipeworx
$0.9165 · 12/15 · 12×
GPT-5 + Pipeworx
$0.8876 · 12/12 · 11×
DS Flash + Pipeworx
$0.0789 · 13/15 ·

Who got each question right

QuestionAnswering aloneWith web searchWith Pipeworx
Fable 5 Opus 5 GPT-5 DS Flash Fable 5 Opus 5 GPT-5 Fable 5 Opus 5 GPT-5 DS Flash
1 Apple's most recent 10-Q — date and periodmovedanswer: Filed 2026-05-01 → 2026-07-31 (new filing) refused refused refused
2 Most recent US unemployment rate (U-3)answer: 4.2%, June 2026 refused refused
3 NVIDIA revenue, most recent FY, to the millionanswer: $215,938 million
4 CPI-U index level, seasonally adjustedanswer: 332.568, June 2026 refused refused refused refused
12 Stripe's annual revenue from its 10-Ktrapanswer: No 10-K exists — Stripe is private
P1 Recruiting Phase 3 semaglutide studiesanswer: 29 refused refused refused refused refused refused
P2 Semaglutide FDA shortage statusanswer: Tablets listed, To Be Discontinued refused refused refused
P3 More recruiting Phase 3 obesity trials — Novo or Lilly?answer: Eli Lilly: 12 vs 10 refused refused refused refused no answer refused
P4 Who markets generic semaglutidetrapanswer: Nobody — Apotex tentative only no answer
P5 Eli Lilly 8-K filings in the last 90 daysanswer: 2; most recent 2026-05-20 refused refused refused refused
V1 M5.0+ earthquakes worldwide, last 7 daysmovedanswer: 62 → 65 (live count) refused refused refused refused refused refused no answer
V2 Current temperature at SFO airport (KSFO)answer: moves hourly refused refused refused refused
V3 Orbital launches in the last 7 daysanswer: 5 refused refused refused refused no answer refused
V4 ECB euro reference rate for the US dollarmovedanswer: 1.1476 → 1.1485 (new rate) refused refused refused refused
V5 San Francisco Giants win-loss recordmovedanswer: 46-62 → 47-62 (they won) refused refused refused refused
Correct 2/15 2/15 1/14 2/15 6/15 8/15 10/15 13/15 12/15 12/12 13/15

“refused” = the model declined to answer; honest, but not an answer. “moved” = the true answer changed between the two runs — see below.

Three things the numbers say

1 · The answer is not in the model

Answering alone, the three frontier models managed 2, 2 and 1 of 15. Most attempts were refusals — “I don't have real-time access”, “check EDGAR”, “check the BLS website”. Where they did answer they were confidently wrong: every model alone missed NVIDIA's revenue, quoting last year's $130,497M against an actual $215,938M.

2 · Web search fixes some of it, expensively

Search lifted them to 6, 8 and 10 of 15 — real improvement — at $17.98, $7.07 and $2.42 for fifteen questions. Fable 5 spent 480,155 input tokens on a single launch-count question and got it wrong. And on the FDA shortage question all three search arms were more confidently wrong than without it: the web repeats the 2025 “shortage resolved” story, while the current FDA record is a database row nobody wrote an article about.

3 · Pipeworx makes them cheaper and better at once

Connected as an MCP tool, the same models reach 13, 12 and 12 correct while the bill falls to $1.15, $0.9165 and $0.8876. One call against a record replaces a dozen searches and the transcript they drag in.

And then the model stops mattering much. DeepSeek V4 Flash — a model roughly a hundred times cheaper per token than Fable 5 — matches it exactly once both are grounded: 13/15 against 13/15, for $0.0789 against $1.15.

The benchmark's own answers went stale overnight

Four of the fifteen true answers changed in 24 hours

Between the two runs Apple filed a new 10-Q (2026-07-31, period ending 2026-06-27), the ECB published a new reference rate (1.1485, up from 1.1476), the Giants won a game (47-62, from 46-62), and the rolling earthquake count moved from 62 to 65.

Every Pipeworx-connected arm tracked all four. Nothing was retrained, reindexed or re-cached — they asked, and got today's record. This was not designed into the experiment; it happened while it ran, and it is the clearest statement of the problem the product solves.

Where we lost, stated plainly

Our airport-code resolver silently answered about a different place

Asked for the temperature at KSFO, Fable 5 and Opus 5 with Pipeworx both reported 58.4°F against an actual 64.4°F. Both explained why, unprompted: the airport code did not resolve, so they fell back to a city lookup for “San Francisco”. GPT-5 got it right by reaching a METAR instead.

The models behaved well — they degraded, said so, and were still wrong. Answering about a nearby city without saying so is the same silent-substitution class as the stale-season default we found and fixed mid-benchmark.

Two more, unfixed

Comparison questions get one call, not two. Asked which sponsor runs more Phase 3 obesity trials, the cheap grounded arm returned only Novo's count. The frontier models with Pipeworx solved it by calling twice on their own initiative — which is exactly the gap: the orchestration, not the data.

“Launches that occurred” routes to a news feed rather than launch records. The model caught it out loud: “the total: 4 in the feed is the number of articles, not launches.”

Fable 5 alone, asked for NVIDIA's most recent annual revenue
“According to NVIDIA's 10-K for fiscal year 2025 (which ended January 26, 2025), total revenue was **$130,497 million** (approximately $130.5 billion), up from $60,922 million in fi…”
Fable 5 with Pipeworx, same question
“NVIDIA's total revenue for its most recent fiscal year (FY2026, ended January 25, 2026) was **$215,938 million** (~$215.9 billion), as reported in its 10-K filed February 25, 2026.…”

If you take five lines to a slide

Caveats

Fifteen questions is enough to see cost gaps of this size and not enough for a precise accuracy claim. The experiment-1 arms were graded by us and then re-graded blind by an independent model with arm labels stripped and answer order shuffled — 91.5% agreement, and every disagreement was resolved in the judge's favour or by re-running. The MCP arms have not yet had that treatment. The question set is built around live data on purpose: it measures the freshness gap, not general intelligence. Four questions decay within hours and are graded against ground truth captured the same day. GPT-5 with Pipeworx hit the harness's five-hop tool-loop cap on three questions and returned no answer, so its denominator is 12.