@pipeworx/arquivo-pt

Connect: https://pipeworx.io/mcp — every tool in the catalog, including @pipeworx/arquivo-pt’s. Install: one-click buttons

Connect to just the @pipeworx/arquivo-pt pack

https://gateway.pipeworx.io/arquivo-pt/mcp — only @pipeworx/arquivo-pt’s own tools, nothing else in the catalog.

No MCP client? Skip the connection: POST https://gateway.pipeworx.io/v1/tools/search_packs {"query":"..."} to find a tool below, GET /v1/tools/<name> for its schema, POST the same URL with arguments for the data — see For AI agents.

Tools: 3

Full-text search of archived web pages in Arquivo.pt, the Portuguese web archive run by FCCN (captures since 1996). Where the Internet Archive’s Wayback Machine (pack wayback) needs the URL, Arquivo.pt indexes the text of what it captured, so a caller can find archived pages by phrase, then read a capture’s extracted text or list every preserved version of a URL.

Tools

  • arquivo_search_pages(query, from?, to?, site?, type?, limit?, offset?, per_site?) — full-text hits for terms or a quoted phrase: title, original URL, capture date, decoded snippet, archived replay URL, extracted-text and screenshot links, plus estimated_total and next_offset. Zero hits answer found:false with reason:"no_match" and a hint.
  • arquivo_url_history(url, from?, to?, limit?, offset?) — every preserved capture of a domain, host or full URL, newest first, with capture timestamp, crawl HTTP status, MIME type, size, digest and replay links. Zero captures answer found:false with reason:"no_captures".
  • arquivo_page_text(url, timestamp, max_chars?) — the plain text Arquivo.pt extracted from one capture (original_url + captured_at_ts of a hit). A capture that does not exist answers found:false, reason:"capture_not_found".

Every response carries source (the exact upstream URL) and data_as_of. A transport failure, a non-JSON body or a changed response shape throws a loud error naming Arquivo.pt and the HTTP status — never an empty list.

Auth

Keyless.

Coverage, honestly

  • Centred on the Portuguese web (.pt sites and Portuguese-language pages), but international pages linked from them are captured too. Probed 2026-10-08: “climate change” ≈ 52.8M estimated hits (ipcc.ch, eea.europa.eu), “quantum computing” ≈ 1.08M (microsoft.com, en.wikipedia.org), “Federal Reserve interest rates” ≈ 1.9M (federalreserve.gov, vox.com). English queries work; the top hits skew to pages Portuguese sites cite.
  • The full-text index lags the crawl by years. Probed 2026-10-08, a search restricted to from=2025 returns 0 hits for any query while arquivo_url_history lists captures from mid-2026. Use the URL history for anything recent.
  • The pack sends to only when you pass one. The upstream default (the previous calendar year) already covers the whole index, and forcing to=<now> made the same query return zero items from a live Worker while estimated_nr_results stayed at 3.8M (probed twice, 2026-10-08).
  • Hits are deduplicated to 2 per site by default (per_site); raise it to see more captures of one site.

Data sources

Tools

  • arquivo_search_pages — Full-text search of archived web pages in Arquivo.pt, the Portuguese web archive (captures since 1996), when the URL is unknown. Finds pages whose TEXT contains the given terms or a quoted phrase and
  • arquivo_url_history — List every preserved capture of a specific URL in Arquivo.pt, the Portuguese web archive, newest first — the version history of a page or domain from 1996 to the current crawl. Each capture carries th
  • arquivo_page_text — Read the extracted plain text of one archived page capture held by Arquivo.pt, the Portuguese web archive — the page content with HTML stripped, as the archive indexed it. Takes the original URL and t

Tools

  • arquivo_page_text — Read the extracted plain text of one archived page capture held by Arquivo.pt, the Portuguese web archive — the page content with HTML stripped, as the archive indexed it. Takes the original URL and t
  • arquivo_search_pages — Full-text search of archived web pages in Arquivo.pt, the Portuguese web archive (captures since 1996), when the URL is unknown. Finds pages whose TEXT contains the given terms or a quoted phrase and
  • arquivo_url_history — List every preserved capture of a specific URL in Arquivo.pt, the Portuguese web archive, newest first — the version history of a page or domain from 1996 to the current crawl. Each capture carries th

Regenerated from source · build October 8, 2026