Graded and priced — 20 questions, 4 models, 3 ways

Twenty questions with answers verified against the primary source, put to four models three ways — answering alone, with web search, and grounded through Pipeworx — priced at what a customer pays.

Run 2026-08-25 · every cost is the API's own billed usage, Pipeworx billed at list · ground truth from SEC EDGAR, FDA, ClinicalTrials.gov, US case law, the US Code, eCFR, USGS, ECB, USDA, UN Comtrade and EPA — captured from the source the same day, never from a model

Frontier models: 100% correct with Pipeworx, 63% with web search
Fable 5 + web search
60% correct
$13.00
Fable 5 + Pipeworx
100% correct
$1.26
Kimi K2 + Pipeworx
95% correct
$0.14
Twenty questions each, same questions, same blind judge. Cost is the full bill — model tokens, search fees, Pipeworx requests at list.
Fable 5 and Opus 5 went 40/40 grounded through Pipeworx and 25/40 with the strongest web search either can run — server-side search that fetches whole pages, not a results list. Alone, without either, they managed 15%. Grounded cost $2.04 for those 40 answers; searching cost $19.23.
Kimi K2 — open weights, $0.141 for all twenty questions — answered 19 of 20, one behind the frontier at a fraction of the price. The grounding is doing the work, not the model.
Correct, wrong, and declined are three different outcomes and this page keeps them apart. Across all 80 graded answers: alone 11% / 3% / 86%; with web search 49% / 15% / 36%; grounded 96% / 0% / 4%. Nothing the grounded arms returned was wrong; what they missed, they declined. A model that declines is behaving well — one that invents a figure is the risk this measures.
Fable 5
alone
3/20
$0.332
Opus 5
alone
3/20
$0.226
GPT-5
alone
1/20
$0.147
Kimi K2
alone
2/20
$0.0069
Fable 5
+ web search
12/20
$13.00
Opus 5
+ web search
13/20
$6.22
GPT-5
+ web search
6/20
$1.14
Kimi K2
+ web search
8/20
$0.183
Fable 5
+ Pipeworx
20/20
$1.26
Opus 5
+ Pipeworx
20/20
$0.788
GPT-5
+ Pipeworx
18/20
$0.588
Kimi K2
+ Pipeworx
19/20
$0.141

Search makes a model answer. It does not make it right.

Search converts silence into speech, not into accuracy.

Alone, the models declined 86% of the time and stated a wrong value on 3%. With web search, refusals fell to 36% — and wrong answers went to 15%, the same order of magnitude. The model becomes willing to answer. It does not become right.

Grounded: 0% wrong, 4% declined.

What each arm cost, and what it is made of

ArmCorrectAgent tokensPipeworx requests TotalPer answerPer correct answer
Fable 5alone 3/20$0.332 $0.332$0.017 $0.111
Opus 5alone 3/20$0.226 $0.226$0.011 $0.075
GPT-5alone 1/20$0.147 $0.147$0.0074 $0.147
Kimi K2alone 2/20$0.0069 $0.0069$0.0003 $0.0035
Fable 5+ web search 12/20$13.00 $13.00$0.650 $1.08
Opus 5+ web search 13/20$6.22 $6.22$0.311 $0.479
GPT-5+ web search 6/20$1.14 $1.14$0.057 $0.190
Kimi K2+ web search 8/20$0.183 $0.183$0.0091 $0.023
Fable 5+ Pipeworx 20/20$1.14 24 × $0.005 = $0.120 $1.26$0.063 $0.063
Opus 5+ Pipeworx 20/20$0.623 33 × $0.005 = $0.165 $0.788$0.039 $0.039
GPT-5+ Pipeworx 18/20$0.353 47 × $0.005 = $0.235 $0.588$0.029 $0.033
Kimi K2+ Pipeworx 19/20$0.026 23 × $0.005 = $0.115 $0.141$0.0071 $0.0074

Pipeworx billed at list — $0.005 per request, the government & open-data category rate. Agent tokens are each model's own billed usage through OpenRouter, web-search fees included. Grounded arms make as many requests as the model chose, which is why their request bill differs. Rate card pulled 2026-08-25.

Cost per correct answer

Fable 5 + web search
$1.08 · 12/20
Opus 5 + web search
$0.479 · 13/20
GPT-5 + web search
$0.190 · 6/20
GPT-5 alone
$0.147 · 1/20
Fable 5 alone
$0.111 · 3/20
Opus 5 alone
$0.075 · 3/20
Fable 5 + Pipeworx
$0.063 · 20/20
Opus 5 + Pipeworx
$0.039 · 20/20
GPT-5 + Pipeworx
$0.033 · 18/20
Kimi K2 + web search
$0.023 · 8/20
Kimi K2 + Pipeworx
$0.0074 · 19/20
Kimi K2 alone
$0.0035 · 2/20

Who got each question right

Questionalone+ web search+ Pipeworx
FableOpusGPT-5KimiFableOpusGPT-5KimiFableOpusGPT-5Kimi
1 What was the exact date of Apple's most recent 10-Q filing with the SEC, and what period did it cover? finance · Filed 2026-07-31, covering the quarter ended 2026-06-27 (accession 0000320193-26-000020)
2 What is the most recent monthly US unemployment rate (U-3) print, and for which reference month? finance · 4.1% for July 2026 (reference month 2026-07)
3 What was NVIDIA's total revenue in its most recent fiscal year, to the nearest million, as reported in its 10-K? finance · $215,938,000,000 for fiscal year 2026 (ended 2026-01-25)
4 What is the current level of the CPI-U index (not the percent change), seasonally adjusted, for the most recent month? finance · 332.813 for July 2026 (CPIAUCSL, index 1982-1984=100)
5 What was Stripe's most recent annual revenue as reported in its 10-K? finance · There is no such filing. Stripe is privately held and files no 10-K with the SEC.
6 How many currently recruiting Phase 3 interventional studies of semaglutide are there, and what conditions do they cover? pharma · 26 by the registry condition index, 33 by free-text match (ClinicalTrials.gov, RECRUITING + PHASE3, 2026-08-24
7 Is any semaglutide product currently on the FDA drug shortage list, and what is its status? pharma · Not on the CURRENT shortage list (status=Current returns zero). Present in the same FDA database under a diffe
8 How many FAERS adverse event reports for semaglutide list ischaemic optic neuropathy as a reaction? pharma · 631 reports
9 Which company markets generic semaglutide in the United States, and when was it approved? pharma · None. No generic semaglutide is approved in the US; every approved semaglutide product (Ozempic, Wegovy, Rybel
10 For the drug Eliquis, what is the expiration date of Orange Book patent number 6,967,208? pharma · 2026-11-21
11 How many magnitude 5.0 or greater earthquakes occurred worldwide in the last 7 days? other · 51 (USGS, captured 2026-08-24)
12 What is the most recent ECB euro foreign exchange reference rate for the US dollar, and for what date? other · 1.1664 on 2026-08-24
13 What is the San Francisco Giants' current win-loss record this season? other · 52-78 (.400), fourth in the NL West, 27.5 games back — 2026 season as of 2026-08-24
14 What is the current average US retail price for regular gasoline? other · $4.049 per gallon, week ending 2026-08-17
15 How many later court decisions have cited Miranda v. Arizona? legal · 58,288 citing references
16 How many later court decisions have cited Roe v. Wade? legal · 5,581 citing references
17 What did the court decide in Tesla, Inc. v. Charge Fusion Technologies, LLC, and when? legal · Decided 2026-03-31 by the US Court of Appeals for the Federal Circuit, case 24-2015 (docket 68904841), NONPREC
18 What data privacy bills are moving in the California legislature? legal · 173 matching bills. SB361 (data brokers: data collection and deletion) was chaptered by the Secretary of State
19 How many bills mentioning artificial intelligence has Congress introduced? legal · 1,671 bills
20 How many federal disaster declarations has FEMA issued in 2026? other · 114 declarations in fiscal year 2026 (one row per declaration per state), most recent FM-5673-AR, Pine Tree Ro

✓ correct · ✗ confidently wrong · – declined. Blind-judged: the judge sees one answer, the question and the ground truth, never which model or arm produced it.

Additional testing — the cheap tier, grounded

Ten more models, small and flash-class, all through Pipeworx on the same twenty questions. The question here is narrower: not whether grounding beats search, but whether a model that costs a few cents per million tokens can use a good tool well enough to matter.

Model + PipeworxCorrectWrongDeclined Cost, 20 questionsPer correct answer
Kimi K219/20 01 $0.141$0.0074
GLM-4.7 Flash19/20 01 $0.188$0.0099
Qwen3.7 Flash18/20 02 $0.114$0.0064
Kimi K2 Thinking18/20 02 $0.192$0.011
Mistral Small 3.217/20 03 $0.108$0.0064
Seed 1.6 Flash17/20 03 $0.113$0.0066
GPT-4.1 Nano15/20 41 $0.113$0.0075
Nova Lite15/20 32 $0.113$0.0076
Gemini 2.5 Flash Lite14/20 15 $0.098$0.0070
GPT-5 Nano13/20 07 $0.115$0.0089

Kimi K2 beat GPT-5, at a quarter of the price.

19/20 against 18/20, $0.141 against $0.588. Qwen3.7 Flash matched GPT-5 at $0.098. Grounding is doing most of the work these questions need, and it does not require an expensive model to do it.

But cheap is not uniformly safe.

Every wrong answer in this tier came from two models — GPT-4.1 Nano and Nova Lite invented values instead of declining. The rest either answered correctly or said they could not. Reasoning did not help either: Kimi K2 Thinking scored below plain Kimi K2 and cost more, because the answer was one tool call away and thinking about it longer is not the same as looking it up.

Additional testing — Perplexity, the search product itself

Every search arm above is a general model with a search tool attached. Perplexity is the other thing: a company whose entire product is answering questions from the live web, marketed specifically on finance and health coverage. Their models expose no tool calling, so there is no grounded arm and no alone arm to run — asking them plainly is running their search product. Same twenty questions, same blind judge, cost including their per-request search fee.

Perplexity modelCorrectWrongDeclined Cost, 20 questionsPer correct answer
Sonar Pro11/20 45 $0.161$0.015
Pro Search11/20 36 $0.430$0.039
Sonar9/20 65 $0.104$0.012
Reasoning Pro9/20 29 $0.125$0.014

Where they are strong, and where they are not

DomainSonar ProPro SearchSonarReasoning Pro
Finance5/55/54/55/5
Pharma3/53/53/53/5
Legal1/51/50/50/5
Other2/52/52/51/5

The finance claim holds up.

95% on the finance block — better than any non-grounded arm on this page, frontier models included. A filing date, a fiscal-year revenue figure, a CIK: these are published on pages a crawler can reach, and Perplexity reaches them.

Legal is 10% — and the misses are inventions, not declines.

2 correct out of 20 across the four models. What makes that worse than a low score is the shape of it: 10 of the legal answers were confidently wrong rather than declined. 2 questions were answered wrongly by all four models — How many later court decisions have cited Miranda v. Arizona? and What did the court decide in Tesla, Inc. v. Charge Fusion Technologies, LLC, and when?. A citation count and a 2026 docket are public record and free; they are simply not sitting on a page that answers the question, so the search product fills the gap with something plausible. That is the failure mode this whole page is about, and it shows up worst in the vendor built entirely around search.

What this does not show