@pipeworx/web-fetch

Connect: https://pipeworx.io/mcp — every tool in the catalog, including @pipeworx/web-fetch’s. Install: one-click buttons

Connect to just the @pipeworx/web-fetch pack

https://gateway.pipeworx.io/web-fetch/mcp — only @pipeworx/web-fetch’s own tools, nothing else in the catalog.

No MCP client? Skip the connection: POST https://gateway.pipeworx.io/v1/tools/search_packs {"query":"..."} to find a tool below, GET /v1/tools/<name> for its schema, POST the same URL with arguments for the data — see For AI agents.

Tools: 1

Fetch the exact contents of one public URL, live from its origin server, and get back the bytes (or text decoded from them) with a SHA-256 hash of exactly what the origin sent, plus the redirect chain and the response headers.

Tools

  • fetch_url(url, format?, max_bytes?, offset?) — one GET of one public http(s) URL. Returns status, url_final, redirects (present only when the URL redirected), headers (content-type, content-length, last-modified, etag, content-language, date), content_family (text | pdf | archive | binary), bytes_total, sha256 over the complete body, and a window of the body:

    • format: "auto" (default) — readable text for HTML and PDF, the raw body for everything else.
    • format: "raw" — the body itself: decoded text for text types, base64 for binary (zip, xlsx, gzip, PDF). sha256_returned hashes exactly the bytes in the window.
    • format: "text" — extracted text (text_extraction: html | pdf | none | failed). PDF extraction warnings are passed through in warnings.
    • max_bytes (default 100,000; at most 1,000,000 for text, 512,000 for base64) and offset page through a large body; follow next_offset.

    Text is decoded from the bytes, never guessed: the charset comes from a BOM, then the Content-Type header, then the page’s own <meta charset> or XML declaration, then UTF-8 if the bytes are valid UTF-8, else windows-1252. charset_source says which rule decided and decode_replacements counts any characters lost in decoding (0 means none).

Checking “verbatim” yourself

sha256 is taken over the entity body after transfer compression is removed — the same bytes curl --compressed -s <url> | shasum -a 256 writes. Two caveats: a dynamic page can legitimately hash differently a second later, and some sites serve different content by location, so the answer is verbatim as served to this fetch (colo, when known, names the edge location).

Refusals — always loud, never an empty body

errorwhen
blocked_urlrule: scheme (not http/https), port (not 80/443), userinfo (user:pass@), private_address, dns_private (the name resolves to a private address), denylisted_host, too_many_redirects, invalid_redirect. Checked on the URL and every redirect hop.
upstream_refused401 / 402 / 403 / 407 / 429 / 451, or a bot-wall page (reason: "bot_wall"). Carries status, a short body_summary, and a hint naming archive_evidence, which is never run automatically.
upstream_down5xx, timeout, connection or DNS failure (phase).
upstream_statusany other non-2xx, e.g. 404.
unsupported_content_typeimages, audio, video, fonts and executables (by declared type or by magic bytes).
too_largean offset beyond the 10 MB read limit on an origin without byte ranges.
budget_exhaustedthe per-site or hourly limit below; carries retry_after_s.
invalid_argumenta malformed argument.

Limits

  • One GET, no cookies, no login, no caller headers, no JavaScript execution. likely_needs_js: true flags a page that renders its content with scripts.
  • 10 MB read per call; a larger body returns the first window with truncated: true, body_complete: false and sha256: null.
  • 25 s per call, at most 4 redirects.
  • Per site (registrable domain), across all callers: 30 requests a minute and 1,000 a day (sec.gov: 5 a minute). A 200 response is reused for 10 minutes (served_from: "cache", cache_age_s).
  • Identifies itself honestly as Pipeworx/1.0 (pipeworx.io); on a 403 it retries once as Pipeworx/1.0 (+https://pipeworx.io; [email protected]) and reports which it used in user_agent_used.

Auth

Keyless.

Data sources

  • Whatever public URL the caller names. Nothing is fetched except that URL and its redirects, plus a DNS-over-HTTPS lookup at https://cloudflare-dns.com/dns-query to refuse names that resolve to private addresses.

Tools

  • fetch_url — Fetch the exact contents of ONE public web URL live from its origin server and return them verbatim with a SHA-256 hash of the bytes received.

Tools

  • fetch_url — Fetch the exact contents of ONE public web URL live from its origin server and return them verbatim with a SHA-256 hash of the bytes received.

Regenerated from source · build October 8, 2026