@pipeworx/arquivo-pt
Connect: https://pipeworx.io/mcp — every tool in the catalog, including @pipeworx/arquivo-pt’s. Install: one-click buttons
Connect to just the @pipeworx/arquivo-pt pack
https://gateway.pipeworx.io/arquivo-pt/mcp — only @pipeworx/arquivo-pt’s own tools, nothing else in the catalog.
No MCP client? Skip the connection: POST https://gateway.pipeworx.io/v1/tools/search_packs {"query":"..."} to find a tool below, GET /v1/tools/<name> for its schema, POST the same URL with arguments for the data — see For AI agents.
Tools: 3
Full-text search of archived web pages in Arquivo.pt, the Portuguese web
archive run by FCCN (captures since 1996). Where the Internet Archive’s
Wayback Machine (pack wayback) needs the URL, Arquivo.pt indexes the text
of what it captured, so a caller can find archived pages by phrase, then read
a capture’s extracted text or list every preserved version of a URL.
Tools
arquivo_search_pages(query, from?, to?, site?, type?, limit?, offset?, per_site?)— full-text hits for terms or a quoted phrase: title, original URL, capture date, decoded snippet, archived replay URL, extracted-text and screenshot links, plusestimated_totalandnext_offset. Zero hits answerfound:falsewithreason:"no_match"and a hint.arquivo_url_history(url, from?, to?, limit?, offset?)— every preserved capture of a domain, host or full URL, newest first, with capture timestamp, crawl HTTP status, MIME type, size, digest and replay links. Zero captures answerfound:falsewithreason:"no_captures".arquivo_page_text(url, timestamp, max_chars?)— the plain text Arquivo.pt extracted from one capture (original_url+captured_at_tsof a hit). A capture that does not exist answersfound:false, reason:"capture_not_found".
Every response carries source (the exact upstream URL) and data_as_of.
A transport failure, a non-JSON body or a changed response shape throws a loud
error naming Arquivo.pt and the HTTP status — never an empty list.
Auth
Keyless.
Coverage, honestly
- Centred on the Portuguese web (
.ptsites and Portuguese-language pages), but international pages linked from them are captured too. Probed 2026-10-08: “climate change” ≈ 52.8M estimated hits (ipcc.ch, eea.europa.eu), “quantum computing” ≈ 1.08M (microsoft.com, en.wikipedia.org), “Federal Reserve interest rates” ≈ 1.9M (federalreserve.gov, vox.com). English queries work; the top hits skew to pages Portuguese sites cite. - The full-text index lags the crawl by years. Probed 2026-10-08, a
search restricted to
from=2025returns 0 hits for any query whilearquivo_url_historylists captures from mid-2026. Use the URL history for anything recent. - The pack sends
toonly when you pass one. The upstream default (the previous calendar year) already covers the whole index, and forcingto=<now>made the same query return zero items from a live Worker whileestimated_nr_resultsstayed at 3.8M (probed twice, 2026-10-08). - Hits are deduplicated to 2 per site by default (
per_site); raise it to see more captures of one site.
Data sources
- https://arquivo.pt/textsearch?q=… — full-text search (
q,from,to,siteSearch,type,maxItems≤ 500,offset,dedupValue,dedupField). - https://arquivo.pt/textsearch?versionHistory=… — URL version history.
- https://arquivo.pt/textextracted?m=… — extracted text
of one capture. The
mvalue is the original URL followed by/and the 14-digit timestamp, so a URL ending in/yields//before the stamp — that is correct, do not “fix” it. - API reference: https://github.com/arquivo/pwa-technologies/wiki/Arquivo.pt-API.
A URL passed as
qis an HTTP 400 upstream; the pack refuses it first and points atarquivo_url_history. - Snippets come back HTML-escaped with Latin-1 named entities
(
Inteligência); the pack decodes them.
Tools
- arquivo_search_pages — Full-text search of archived web pages in Arquivo.pt, the Portuguese web archive (captures since 1996), when the URL is unknown. Finds pages whose TEXT contains the given terms or a quoted phrase and
- arquivo_url_history — List every preserved capture of a specific URL in Arquivo.pt, the Portuguese web archive, newest first — the version history of a page or domain from 1996 to the current crawl. Each capture carries th
- arquivo_page_text — Read the extracted plain text of one archived page capture held by Arquivo.pt, the Portuguese web archive — the page content with HTML stripped, as the archive indexed it. Takes the original URL and t
Tools
arquivo_page_text— Read the extracted plain text of one archived page capture held by Arquivo.pt, the Portuguese web archive — the page content with HTML stripped, as the archive indexed it. Takes the original URL and tarquivo_search_pages— Full-text search of archived web pages in Arquivo.pt, the Portuguese web archive (captures since 1996), when the URL is unknown. Finds pages whose TEXT contains the given terms or a quoted phrase andarquivo_url_history— List every preserved capture of a specific URL in Arquivo.pt, the Portuguese web archive, newest first — the version history of a page or domain from 1996 to the current crawl. Each capture carries th