@pipeworx/web-fetch
Connect: https://pipeworx.io/mcp — every tool in the catalog, including @pipeworx/web-fetch’s. Install: one-click buttons
Connect to just the @pipeworx/web-fetch pack
https://gateway.pipeworx.io/web-fetch/mcp — only @pipeworx/web-fetch’s own tools, nothing else in the catalog.
No MCP client? Skip the connection: POST https://gateway.pipeworx.io/v1/tools/search_packs {"query":"..."} to find a tool below, GET /v1/tools/<name> for its schema, POST the same URL with arguments for the data — see For AI agents.
Tools: 1
Fetch the exact contents of one public URL, live from its origin server, and get back the bytes (or text decoded from them) with a SHA-256 hash of exactly what the origin sent, plus the redirect chain and the response headers.
Tools
-
fetch_url(url, format?, max_bytes?, offset?)— one GET of one public http(s) URL. Returnsstatus,url_final,redirects(present only when the URL redirected),headers(content-type, content-length, last-modified, etag, content-language, date),content_family(text|pdf|archive|binary),bytes_total,sha256over the complete body, and a window of the body:format: "auto"(default) — readable text for HTML and PDF, the raw body for everything else.format: "raw"— the body itself: decoded text for text types, base64 for binary (zip, xlsx, gzip, PDF).sha256_returnedhashes exactly the bytes in the window.format: "text"— extracted text (text_extraction:html|pdf|none|failed). PDF extraction warnings are passed through inwarnings.max_bytes(default 100,000; at most 1,000,000 for text, 512,000 for base64) andoffsetpage through a large body; follownext_offset.
Text is decoded from the bytes, never guessed: the charset comes from a BOM, then the
Content-Typeheader, then the page’s own<meta charset>or XML declaration, then UTF-8 if the bytes are valid UTF-8, else windows-1252.charset_sourcesays which rule decided anddecode_replacementscounts any characters lost in decoding (0 means none).
Checking “verbatim” yourself
sha256 is taken over the entity body after transfer compression is removed —
the same bytes curl --compressed -s <url> | shasum -a 256 writes. Two caveats:
a dynamic page can legitimately hash differently a second later, and some
sites serve different content by location, so the answer is verbatim as
served to this fetch (colo, when known, names the edge location).
Refusals — always loud, never an empty body
error | when |
|---|---|
blocked_url | rule: scheme (not http/https), port (not 80/443), userinfo (user:pass@), private_address, dns_private (the name resolves to a private address), denylisted_host, too_many_redirects, invalid_redirect. Checked on the URL and every redirect hop. |
upstream_refused | 401 / 402 / 403 / 407 / 429 / 451, or a bot-wall page (reason: "bot_wall"). Carries status, a short body_summary, and a hint naming archive_evidence, which is never run automatically. |
upstream_down | 5xx, timeout, connection or DNS failure (phase). |
upstream_status | any other non-2xx, e.g. 404. |
unsupported_content_type | images, audio, video, fonts and executables (by declared type or by magic bytes). |
too_large | an offset beyond the 10 MB read limit on an origin without byte ranges. |
budget_exhausted | the per-site or hourly limit below; carries retry_after_s. |
invalid_argument | a malformed argument. |
Limits
- One GET, no cookies, no login, no caller headers, no JavaScript execution.
likely_needs_js: trueflags a page that renders its content with scripts. - 10 MB read per call; a larger body returns the first window with
truncated: true,body_complete: falseandsha256: null. - 25 s per call, at most 4 redirects.
- Per site (registrable domain), across all callers: 30 requests a minute and
1,000 a day (sec.gov: 5 a minute). A 200 response is reused for 10 minutes
(
served_from: "cache",cache_age_s). - Identifies itself honestly as
Pipeworx/1.0 (pipeworx.io); on a 403 it retries once asPipeworx/1.0 (+https://pipeworx.io; [email protected])and reports which it used inuser_agent_used.
Auth
Keyless.
Data sources
- Whatever public URL the caller names. Nothing is fetched except that URL and its redirects, plus a DNS-over-HTTPS lookup at https://cloudflare-dns.com/dns-query to refuse names that resolve to private addresses.
Tools
- fetch_url — Fetch the exact contents of ONE public web URL live from its origin server and return them verbatim with a SHA-256 hash of the bytes received.
Tools
fetch_url— Fetch the exact contents of ONE public web URL live from its origin server and return them verbatim with a SHA-256 hash of the bytes received.