Graded and priced — 20 questions, 4 models, 3 ways
Twenty questions with answers verified against the primary source, put to four models three ways — answering alone, with web search, and grounded through Pipeworx — priced at what a customer pays.
Search makes a model answer. It does not make it right.
Search converts silence into speech, not into accuracy.
Alone, the models declined 86% of the time and stated a wrong value on 3%. With web search, refusals fell to 36% — and wrong answers went to 15%, the same order of magnitude. The model becomes willing to answer. It does not become right.
Grounded: 0% wrong, 4% declined.
What each arm cost, and what it is made of
| Arm | Correct | Agent tokens | Pipeworx requests | Total | Per answer | Per correct answer |
|---|---|---|---|---|---|---|
| Fable 5alone | 3/20 | $0.332 | — | $0.332 | $0.017 | $0.111 |
| Opus 5alone | 3/20 | $0.226 | — | $0.226 | $0.011 | $0.075 |
| GPT-5alone | 1/20 | $0.147 | — | $0.147 | $0.0074 | $0.147 |
| Kimi K2alone | 2/20 | $0.0069 | — | $0.0069 | $0.0003 | $0.0035 |
| Fable 5+ web search | 12/20 | $13.00 | — | $13.00 | $0.650 | $1.08 |
| Opus 5+ web search | 13/20 | $6.22 | — | $6.22 | $0.311 | $0.479 |
| GPT-5+ web search | 6/20 | $1.14 | — | $1.14 | $0.057 | $0.190 |
| Kimi K2+ web search | 8/20 | $0.183 | — | $0.183 | $0.0091 | $0.023 |
| Fable 5+ Pipeworx | 20/20 | $1.14 | 24 × $0.005 = $0.120 | $1.26 | $0.063 | $0.063 |
| Opus 5+ Pipeworx | 20/20 | $0.623 | 33 × $0.005 = $0.165 | $0.788 | $0.039 | $0.039 |
| GPT-5+ Pipeworx | 18/20 | $0.353 | 47 × $0.005 = $0.235 | $0.588 | $0.029 | $0.033 |
| Kimi K2+ Pipeworx | 19/20 | $0.026 | 23 × $0.005 = $0.115 | $0.141 | $0.0071 | $0.0074 |
Cost per correct answer
Who got each question right
| Question | alone | + web search | + Pipeworx | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Fable | Opus | GPT-5 | Kimi | Fable | Opus | GPT-5 | Kimi | Fable | Opus | GPT-5 | Kimi | |
| 1 What was the exact date of Apple's most recent 10-Q filing with the SEC, and what period did it cover? finance · Filed 2026-07-31, covering the quarter ended 2026-06-27 (accession 0000320193-26-000020) | – | – | – | – | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| 2 What is the most recent monthly US unemployment rate (U-3) print, and for which reference month? finance · 4.1% for July 2026 (reference month 2026-07) | – | – | – | – | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| 3 What was NVIDIA's total revenue in its most recent fiscal year, to the nearest million, as reported in its 10-K? finance · $215,938,000,000 for fiscal year 2026 (ended 2026-01-25) | – | ✗ | – | – | ✓ | ✓ | – | ✓ | ✓ | ✓ | ✓ | ✓ |
| 4 What is the current level of the CPI-U index (not the percent change), seasonally adjusted, for the most recent month? finance · 332.813 for July 2026 (CPIAUCSL, index 1982-1984=100) | – | – | – | – | ✓ | ✓ | – | – | ✓ | ✓ | ✓ | ✓ |
| 5 What was Stripe's most recent annual revenue as reported in its 10-K? finance · There is no such filing. Stripe is privately held and files no 10-K with the SEC. | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | – | ✓ | ✓ | ✓ | ✓ | ✓ |
| 6 How many currently recruiting Phase 3 interventional studies of semaglutide are there, and what conditions do they cover? pharma · 26 by the registry condition index, 33 by free-text match (ClinicalTrials.gov, RECRUITING + PHASE3, 2026-08-24 | – | – | – | – | – | – | – | ✗ | ✓ | ✓ | – | ✓ |
| 7 Is any semaglutide product currently on the FDA drug shortage list, and what is its status? pharma · Not on the CURRENT shortage list (status=Current returns zero). Present in the same FDA database under a diffe | ✓ | ✓ | – | – | ✓ | ✓ | – | ✓ | ✓ | ✓ | ✓ | ✓ |
| 8 How many FAERS adverse event reports for semaglutide list ischaemic optic neuropathy as a reaction? pharma · 631 reports | – | – | – | – | ✗ | ✗ | – | ✗ | ✓ | ✓ | ✓ | ✓ |
| 9 Which company markets generic semaglutide in the United States, and when was it approved? pharma · None. No generic semaglutide is approved in the US; every approved semaglutide product (Ozempic, Wegovy, Rybel | ✓ | ✓ | – | ✓ | ✓ | ✓ | – | ✗ | ✓ | ✓ | ✓ | ✓ |
| 10 For the drug Eliquis, what is the expiration date of Orange Book patent number 6,967,208? pharma · 2026-11-21 | – | ✗ | – | – | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| 11 How many magnitude 5.0 or greater earthquakes occurred worldwide in the last 7 days? other · 51 (USGS, captured 2026-08-24) | – | – | – | – | – | – | – | – | ✓ | ✓ | ✓ | ✓ |
| 12 What is the most recent ECB euro foreign exchange reference rate for the US dollar, and for what date? other · 1.1664 on 2026-08-24 | – | – | – | – | ✓ | ✓ | ✓ | ✗ | ✓ | ✓ | ✓ | – |
| 13 What is the San Francisco Giants' current win-loss record this season? other · 52-78 (.400), fourth in the NL West, 27.5 games back — 2026 season as of 2026-08-24 | – | – | – | – | ✗ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| 14 What is the current average US retail price for regular gasoline? other · $4.049 per gallon, week ending 2026-08-17 | – | – | – | – | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| 15 How many later court decisions have cited Miranda v. Arizona? legal · 58,288 citing references | – | – | – | – | – | – | – | ✗ | ✓ | ✓ | ✓ | ✓ |
| 16 How many later court decisions have cited Roe v. Wade? legal · 5,581 citing references | – | – | – | – | – | – | – | ✗ | ✓ | ✓ | ✓ | ✓ |
| 17 What did the court decide in Tesla, Inc. v. Charge Fusion Technologies, LLC, and when? legal · Decided 2026-03-31 by the US Court of Appeals for the Federal Circuit, case 24-2015 (docket 68904841), NONPREC | – | – | – | – | ✓ | ✓ | – | – | ✓ | ✓ | – | ✓ |
| 18 What data privacy bills are moving in the California legislature? legal · 173 matching bills. SB361 (data brokers: data collection and deletion) was chaptered by the Secretary of State | – | – | – | – | ✓ | ✓ | – | – | ✓ | ✓ | ✓ | ✓ |
| 19 How many bills mentioning artificial intelligence has Congress introduced? legal · 1,671 bills | – | – | – | – | – | ✗ | – | ✗ | ✓ | ✓ | ✓ | ✓ |
| 20 How many federal disaster declarations has FEMA issued in 2026? other · 114 declarations in fiscal year 2026 (one row per declaration per state), most recent FM-5673-AR, Pine Tree Ro | – | – | – | – | – | – | – | ✗ | ✓ | ✓ | ✓ | ✓ |
- Web search cost more and answered less — for every model. Not a trade-off, and not a close call: grounding was cheaper per correct answer than search in every pairing we ran.
- Grounding beat both search arms on accuracy, for every model.
- The cheapest model won. Kimi K2 grounded matched or beat every frontier arm at $0.0074 a correct answer.
- Declining is the honest failure; invention is the dangerous one. Most of what the other arms missed, they declined rather than got wrong — that is a model behaving properly about the limits of what it can see. What changes with grounding is not that the model becomes bolder; it is that the answer is actually available.
Additional testing — the cheap tier, grounded
Ten more models, small and flash-class, all through Pipeworx on the same twenty questions. The question here is narrower: not whether grounding beats search, but whether a model that costs a few cents per million tokens can use a good tool well enough to matter.
| Model + Pipeworx | Correct | Wrong | Declined | Cost, 20 questions | Per correct answer |
|---|---|---|---|---|---|
| Kimi K2 | 19/20 | 0 | 1 | $0.141 | $0.0074 |
| GLM-4.7 Flash | 19/20 | 0 | 1 | $0.188 | $0.0099 |
| Qwen3.7 Flash | 18/20 | 0 | 2 | $0.114 | $0.0064 |
| Kimi K2 Thinking | 18/20 | 0 | 2 | $0.192 | $0.011 |
| Mistral Small 3.2 | 17/20 | 0 | 3 | $0.108 | $0.0064 |
| Seed 1.6 Flash | 17/20 | 0 | 3 | $0.113 | $0.0066 |
| GPT-4.1 Nano | 15/20 | 4 | 1 | $0.113 | $0.0075 |
| Nova Lite | 15/20 | 3 | 2 | $0.113 | $0.0076 |
| Gemini 2.5 Flash Lite | 14/20 | 1 | 5 | $0.098 | $0.0070 |
| GPT-5 Nano | 13/20 | 0 | 7 | $0.115 | $0.0089 |
Kimi K2 beat GPT-5, at a quarter of the price.
19/20 against 18/20, $0.141 against $0.588. Qwen3.7 Flash matched GPT-5 at $0.098. Grounding is doing most of the work these questions need, and it does not require an expensive model to do it.
But cheap is not uniformly safe.
Every wrong answer in this tier came from two models — GPT-4.1 Nano and Nova Lite invented values instead of declining. The rest either answered correctly or said they could not. Reasoning did not help either: Kimi K2 Thinking scored below plain Kimi K2 and cost more, because the answer was one tool call away and thinking about it longer is not the same as looking it up.
Additional testing — Perplexity, the search product itself
Every search arm above is a general model with a search tool attached. Perplexity is the other thing: a company whose entire product is answering questions from the live web, marketed specifically on finance and health coverage. Their models expose no tool calling, so there is no grounded arm and no alone arm to run — asking them plainly is running their search product. Same twenty questions, same blind judge, cost including their per-request search fee.
| Perplexity model | Correct | Wrong | Declined | Cost, 20 questions | Per correct answer |
|---|---|---|---|---|---|
| Sonar Pro | 11/20 | 4 | 5 | $0.161 | $0.015 |
| Pro Search | 11/20 | 3 | 6 | $0.430 | $0.039 |
| Sonar | 9/20 | 6 | 5 | $0.104 | $0.012 |
| Reasoning Pro | 9/20 | 2 | 9 | $0.125 | $0.014 |
Where they are strong, and where they are not
| Domain | Sonar Pro | Pro Search | Sonar | Reasoning Pro |
|---|---|---|---|---|
| Finance | 5/5 | 5/5 | 4/5 | 5/5 |
| Pharma | 3/5 | 3/5 | 3/5 | 3/5 |
| Legal | 1/5 | 1/5 | 0/5 | 0/5 |
| Other | 2/5 | 2/5 | 2/5 | 1/5 |
The finance claim holds up.
95% on the finance block — better than any non-grounded arm on this page, frontier models included. A filing date, a fiscal-year revenue figure, a CIK: these are published on pages a crawler can reach, and Perplexity reaches them.
Legal is 10% — and the misses are inventions, not declines.
2 correct out of 20 across the four models. What makes that worse than a low score is the shape of it: 10 of the legal answers were confidently wrong rather than declined. 2 questions were answered wrongly by all four models — How many later court decisions have cited Miranda v. Arizona? and What did the court decide in Tesla, Inc. v. Charge Fusion Technologies, LLC, and when?. A citation count and a 2026 docket are public record and free; they are simply not sitting on a page that answers the question, so the search product fills the gap with something plausible. That is the failure mode this whole page is about, and it shows up worst in the vendor built entirely around search.
What this does not show
- Search costs vary enormously between models, and not because of the models. Each model here searches with the best tool it can actually run, and those tools differ in how much of a page they pull into context — which is most of what the cost column is measuring. Compare each model against itself; comparing search costs across models says more about the search products than about the models.
- One run. Routing has real run-to-run variance.
- The judge is a model scoring against captured ground truth with acceptance ranges, not a human.
- Zero wrong answers is one question away from four, and you should know which. Every wrong answer the grounded arms gave in the first cut of this run was on the same question — how many orbital launches are scheduled in the next seven days. It was replaced, and not because we failed it: the question has no single true answer. Our own capture read 25, three models independently read 10, and one noted that only 5 had confirmed times. Scheduled, confirmed and net-of-slips are three different numbers and the question names none of them. Removing an ungradeable question is right; removing the only question you got wrong is convenient. Both are true here, so it is stated rather than left for someone to notice.
- Every arm here is graded on all twenty questions. An earlier cut of this page was not, and it flattered us. Eight cells had hit this runner's four-round tool cap, been scored as the model refusing, and then excluded — and they were not a random eight. They were the hardest questions in the set, the ones where a model keeps calling tools: recruiting-trial counts, orbital launches, a 2026 decision. Dropping them raised the grounded score and cut its bill at the same time. GPT-5 grounded read 15/15 on that basis; graded on all twenty it is 17/20. The cap is now eight rounds, the eight cells were re-run, and no arm is scored on a subset.
- The search-arm cost is largely a measure of how much each model chose to read. Input tokens across the four search arms ranged from 53,000 to 406,000 — an eightfold spread on identical questions. Fable is the most expensive model per token in this set and still lands mid-pack, because it pulled about a fifth of the context Opus did for the same score. Treat the accuracy column as the solid one and the cross-model cost column as directional.
- The Tesla question asked what the court decided while its ground truth recorded only the date and
precedential status — three arms stated those facts exactly, added the holding, and were marked wrong for
answering more fully than the truth allowed. It is now graded only on what could be independently
verified, because the opinion text could not be retrieved to confirm a disposition. Every re-scored cell
carries
rejudged_from; every re-run cell carriespatched. - Two questions are traps with no correct value to state — Stripe files no 10-K, and no generic semaglutide is approved. They are scored on whether a model says so or invents a figure.
- Twenty questions chosen to be answerable from primary sources. Not a random sample of what anyone asks.