@pipeworx/ntsb-investigations

Connect: https://pipeworx.io/mcp — every tool in the catalog, including @pipeworx/ntsb-investigations’s. Install: one-click buttons

Connect to just the @pipeworx/ntsb-investigations pack

https://gateway.pipeworx.io/ntsb-investigations/mcp — only @pipeworx/ntsb-investigations’s own tools, nothing else in the catalog.

No MCP client? Skip the connection: POST https://gateway.pipeworx.io/v1/tools/search_packs {"query":"..."} to find a tool below, GET /v1/tools/<name> for its schema, POST the same URL with arguments for the data — see For AI agents.

Tools: 1

NTSB aviation accident/incident investigations, searchable by aircraft registration (tail number) or by make/model. Returns NTSB number, event date, location, injury level, phase/occurrence, probable cause, findings, and the NTSB docket link for each match.

Tools

  • ntsb_search_investigations({ registration?, make?, model?, year?, state?, limit? }) — requires registration or make. model narrows within a make (matches “172”, “172S”, “172N”, … by prefix). year filters by event year, and is also how a registration match is narrowed when an N-number has been reassigned to a different aircraft over time (see below). Returns { found, data_as_of, source, count, query, results[] }, or { found: false, data_as_of, source, query } with no matches.

Auth

Keyless. No upstream account, no API key.

Data sources

  • https://data.ntsb.gov/avdata — the NTSB’s own eADMS aviation accident database (avall.zip, a Microsoft Access .mdb bulk file, confirmed live 2026-10-07 at ~96 MB zipped / ~561 MB unzipped). This is the full, current dataset — the SAME data CAROL’s aviation search queries, published as the bulk file NTSB itself ships for exactly this use.
  • https://data.ntsb.gov/Docket/?NTSBNumber=NNN — the NTSB docket page for a given ntsb_number, where the final report and supporting documents are published. Confirmed live (HTTP 200) for a real case pulled from this same dataset.

Why not CAROL’s own search, or the newer developer API

CAROL (data.ntsb.gov/carol-main-public), NTSB’s modern search UI, was probed first. Its JSON endpoints (api/Query/Main, api/Query/FileExport) are undocumented, and every QueryGroups/QueryRules request shape reverse-engineered from a public third-party CAROL proxy and from NTSB’s own CAROL-Guide.pdf returned HTTP 500 (“An unknown exception occured.” / “An error has occurred.”) — the frontend ships no devtools-free schema and the server gives no hint why the shape it wants differs from what a real browser session sends (likely a server-side session/anti-forgery requirement CAROL’s own JS sets up before the first query).

NTSB also runs a newer REST API at developer.ntsb.gov (Aviation / Safety Recommendations / Data Dictionary APIs, via an Azure API Management gateway). That portal requires signing up for a developer account and a subscription key before any call succeeds — an authentication wall, which per this repo’s standing rule is never a free pass to build around without a separate key decision, unlike the plain public bulk file this pack actually reads.

Refreshing the data

node mcps/ntsb-investigations/scripts/refresh.mjs [--local-mdb <path>] [--dry-run] [--r2-mode auto|api|wrangler] [--shard-max-bytes <n>]

Downloads avall.zip, unzips it, runs mdb-export (mdbtools — brew install mdbtools / apt install mdbtools) on five tables (aircraft, events, narratives, Findings, Events_Sequence, eADMSPUB_DataDictionary), joins them into one flat case-per-aircraft list, then shards it into the shared pipeworx-datasets R2 bucket (same bucket as mcps/indiana-code, mcps/illinois-code, mcps/crs-reports — see scripts/lib/statute-store.mjs) rather than writing one object:

  • ntsb/by-make/<slug>.json — one array per manufacturer. A make whose flat shard would exceed ~4 MB (only CESSNA, at this pull) is split further into ntsb/by-make/<slug>/<year>.json.
  • Manufacturers with under 5 rows (most of the ~4,174 distinct raw acft_make strings — OCR/data-entry variants, not real distinct makers) are pooled into 16 ntsb/by-make/_other/<bucket>.json shards by a stable hash, rather than one tiny file each.
  • ntsb/by-registration.json — normalised registration -> the (make, year) pairs it appears under, so a registration query knows which shard(s) to read without a full scan.
  • ntsb/manifest.json — data_as_of, source, record_count, and for every make the exact shard shape (flat / by_year + its years / bucket), so the read path never has to probe R2 to find out which key a make lives in.

This replaced an earlier single-object design (ntsb/aviation-cases.json, ~27 MB) that re-fetched and parsed the whole dataset on every call with no cache — two concurrent calls held two full parses in one isolate at once. src/index.ts now reads the manifest plus exactly the shard(s) one query needs (one, for a make/model query; the few a reused registration’s distinct (make, year) pairs resolve to, for a registration query) and never reads the old monolith, which this script deletes from the bucket once the shards are up. No module-scope cache, on purpose — see the header comment in src/index.ts. No Supabase migration; storage is R2 only. --local-mdb re-uses an already-downloaded avall.mdb instead of re-pulling 96 MB. data_as_of in every response is the date the refresh last ran, not a per-record field from the upstream.

Registration (N-number) reuse

Tail numbers are reassigned by the FAA to different aircraft over time. ntsb_search_investigations({ registration }) does not silently collapse this: every historical match for that tail number comes back, each with its own event_date, make, and model. Pass year to narrow to one incarnation when more than one comes back for the same registration.

A broad make may ask you to narrow it

A handful of common manufacturers (CESSNA, at this pull) have enough rows that their data is split into one shard per year. A query for one of those makes with neither model nor year would have to read every one of that make’s year-shards to answer — not a full dataset scan, but not the one or two small objects every other query costs either. Past a small fixed number of year-shards, the tool refuses with a message asking for year (or model, which does not change which shards get read but narrows the point of asking) instead of silently doing the larger read. A make/model/year combination, or an uncommon make on its own, is unaffected.

Tools

Tools

Regenerated from source · build October 8, 2026