@pipeworx/company-domain
Connect: https://pipeworx.io/mcp — every tool in the catalog, including @pipeworx/company-domain’s. Install: one-click buttons
Connect to just the @pipeworx/company-domain pack
https://gateway.pipeworx.io/company-domain/mcp — only @pipeworx/company-domain’s own tools, nothing else in the catalog.
No MCP client? Skip the connection: POST https://gateway.pipeworx.io/v1/tools/search_packs {"query":"..."} to find a tool below, GET /v1/tools/<name> for its schema, POST the same URL with arguments for the data — see For AI agents.
Tools: 2
Resolve a company NAME to its website domain, with evidence attributed per source — never a single blended confidence score.
Tools
resolve_domain(name, city?, state?, country?)— ranked candidate domains for one company name, each carryingmatched | inferred | not_foundstatus per source pluschecked(false means the source could not be reached — distinct from a confirmed negative). Sources: Wikidata (P856 official-website claim), GLEIF LEI registry, SEC EDGAR (US filers), UK Companies House (BYOK), Frenchrecherche-entreprises, a DNS-alive check, a homepage<title>/og:site_namematch, and an optional BYOK SERP fallback (serper/brave).resolve_domain_batch(rows[])— the same, for up to 100 rows in one call.
Auth
Keyless by default. Wikidata, GLEIF, SEC EDGAR, DNS and the homepage fetch all need no key.
Two sources are BYOK, passed as extra arguments on resolve_domain /
per-row on resolve_domain_batch:
_chApiKey— a UK Companies House API key (free at https://developer.company-information.service.gov.uk). Only called whencountryis GB/UK. Companies House has no website field in its own data (verified live 2026-10-08), so this only ever adds legal-entity confirmation for a GB/UK name, never a domain._serpApiKey(+ optional_serpProvider:"serper"default, or"brave") — a long-tail fallback search when none of the registries above carry a domain. The response disclosesserp.used: trueand the provider whenever it actually ran; with no key it is skipped (used: false) and nothing is called.
No new platform key was added for this pack (standing rule D4) — the only paid dependency (SERP) is caller-funded.
What each source actually carries — read this before trusting a not_found
Verified live on 2026-10-08, probing real records (Apple, Tesla, Microsoft,
Nvidia, Alphabet, Reddit and more for SEC; a live GLEIF record for Apple Inc.;
Companies House’s and recherche-entreprises’s own documented schemas):
| Source | Can it carry a domain? | What it’s actually used for |
|---|---|---|
| Wikidata (P856) | Yes | The only registry source that can directly produce a domain. A company can carry several P856 claims (regional sites); every call keeps all of them as separate candidates rather than guessing which is canonical — the one that also matches a guessed slug wins ranking by corroboration. |
| GLEIF LEI registry | No — no website field exists in the schema | Legal-entity confirmation only: LEI, jurisdiction, registration status, and whether the match is exact. |
SEC EDGAR (data.sec.gov/submissions) | Schema-present (website), but empty for the large majority of filers (Apple, Tesla, Microsoft, Nvidia, Alphabet, Reddit, Marinus Pharmaceuticals and more were all probed empty) | CIK + filed name confirmation; a matched domain on the rare filer that does populate it. |
| UK Companies House | No website field | GB/UK legal-entity confirmation (BYOK). |
French recherche-entreprises | No website field | FR legal-entity confirmation (SIREN, registered name). |
| Guessed domain (name → slug → TLD) | n/a — always inferred | The fallback that catches names none of the registries carry. Strips US legal-form suffixes (companyNameKey) and EU ones the shared helper doesn’t (SE, BV, SARL, HOLDING(S), GROUP, …), collapses dotted abbreviations ("S.A." → SA, "p.l.c." → PLC) before stripping, and adds an initials guess ("International Business Machines Corporation" → ibm). |
DNS-alive (@pipeworx/dns) | n/a — verification | Confirms a candidate domain actually resolves. A real NXDOMAIN is not_found; a DNS transport failure is checked: false with the error named — never collapsed into the same shape. |
Homepage title/og:site_name (@pipeworx/web-fetch) | n/a — verification | Confirms a candidate’s homepage text overlaps the company name. A bot-wall or auth refusal is reported checked: true, status: not_found with the refusal named (e.g. upstream_refused) — this is common and expected for large companies (Oracle, Adobe, ExxonMobil, Pfizer, Home Depot, Zapier and Grammarly all 403/bot-wall an unauthenticated fetch), not a pack defect. |
| SERP (BYOK) | Yes, when a key is supplied | Long-tail fallback; always inferred, never matched — a search result is never treated as self-attested the way a registry’s own field is. |
So a not_found on gleif / sec_edgar / companies_house / entreprises_fr
is expected on nearly every call — it means the source confirmed (or
didn’t find) the entity, not that it failed to look for a website it never
had.
Ranking — never a single blended score
Every candidate carries the full per-source evidence object (every field in the table above, every call). Ranking is ordinal, not a confidence number:
- Wikidata-sourced domain, DNS-alive
- Any-sourced domain, DNS-alive and homepage-title-matched
- DNS-alive only
- DNS-dead
Ties break on support_count — the plain integer count of how many
independent things (origins + DNS + homepage) corroborate that exact domain.
A verified: true boolean on each candidate means DNS-alive AND (wikidata- matched OR homepage-matched) — read the full evidence object before trusting
a candidate the response did NOT mark verified; see the homepage bot-wall note
above for why an unverified rank-1 is often still correct.
Measured precision — labelled set of 110 companies
eval/labelled-set.json (110 rows: 30 public US / SEC-listed, 20 private US,
20 UK, 20 EU, 20 smaller/private-software “SMB”) with each row’s ground-truth
domain set from the company’s own well-documented primary site — the same
fact a Wikipedia infobox or a stock-listing cover page states, not derived
from this pack. eval/run-eval.mjs calls resolve_domain live, for real,
over the network (same code path the gateway executes) and compares the
rank-1 candidate.
Reproduce with node mcps/company-domain/eval/run-eval.mjs.
n: 110
precision_at_1: 0.6818 (top-ranked candidate is the right domain)
rank1_verified_rate: 0.4455 (top-ranked AND independently confirmed by this pack's own evidence)
not_found_rate: 0.2727 (right domain never appeared among candidates at all)
by_label (hits / verified_hits / true_miss out of total):
public_us 20/30 hits · 12 verified · 8 true_miss
private_us 9/20 hits · 7 verified · 10 true_miss
uk 17/20 hits · 12 verified · 3 true_miss
eu 13/20 hits · 9 verified · 5 true_miss
smb 16/20 hits · 9 verified · 4 true_miss
Reading the numbers: precision_at_1 is whether the top-ranked candidate is
the right domain — the Claygent-comparable number. rank1_verified_rate is
the (strictly smaller) share where the rank-1 domain is ALSO independently
confirmed by this pack’s own evidence (wikidata or an unblocked homepage
read); the gap between the two is overwhelmingly homepage bot-walls on large,
well-known sites, not wrong answers — see the misses list this script prints
and the table above. not_found_rate is the share where the correct domain
never appeared among the candidates at all (hardest cases: marketing
abbreviations with no mechanical derivation from the legal name — “P&G” for
The Procter & Gamble Company, “JNJ” for Johnson & Johnson — and a few
registered-vs-trading-name mismatches like Oprah Winfrey’s “Harpo, Inc.” vs
oprah.com).
Pricing
10 credits per resolve_domain call ($0.001 at the standing peg, 1 credit =
$0.0001 — shared/src/overage.ts), keyless. Compare Claygent’s quoted
$0.002–$0.006 per row for the same “name → domain” lookup (the demand
evidence this pack was built from, Jul 27 GTM thread) — roughly 2–6x cheaper
per row, with per-source evidence Claygent does not expose.
Data sources
- https://www.wikidata.org/w/api.php (via
@pipeworx/wikidata) — P856 official-website claim. - https://api.gleif.org/api/v1 (via
@pipeworx/gleif) — LEI registry, identity confirmation only. - https://data.sec.gov/submissions/CIKxxxxxxxxxx.json — fetched directly
(not exposed by any
edgar/sectool today); CIK resolved via the sharedresolveSecEntityhelper theedgarpack itself uses. - https://api.company-information.service.gov.uk (via
@pipeworx/companies-house) — BYOK, GB/UK only. - https://recherche-entreprises.api.gouv.fr (via
@pipeworx/entreprises-fr) — FR only. @pipeworx/dns(Google DNS-over-HTTPS) — A-record alive check.@pipeworx/web-fetch— SSRF-safe homepage fetch (shared/src/ssrf.ts); this pack reads<title>andog:site_namefrom the raw HTML itself.@pipeworx/serper/@pipeworx/brave-search— BYOK long-tail fallback.
No data is baked or mirrored; every source above is called live, per request.
Tools
- resolve_domain — Ranked candidate website domains for a company name, with per-source evidence.
- resolve_domain_batch — resolve_domain for up to 100 company rows in one call.
Tools
resolve_domain— Ranked candidate website domains for a company name, with per-source evidence.resolve_domain_batch— resolve_domain for up to 100 company rows in one call.