arquivo_search_pages
Pack: arquivo-pt · Connect: https://pipeworx.io/mcp (see Connect below for a single-pack URL)
No MCP client? Call it directly: GET https://gateway.pipeworx.io/v1/tools/arquivo_search_pages for the schema, then POST the same URL with its arguments for the data.
Full-text search of archived web pages in Arquivo.pt, the Portuguese web archive (captures since 1996), when the URL is unknown. Finds pages whose TEXT contains the given terms or a quoted phrase and returns, per hit, the page title, original URL, capture date, a text snippet with the matching terms, the archived (replay) URL, and links to the extracted text and a screenshot. Coverage: centred on the Portuguese web (.pt sites and Portuguese-language pages) but international English-language sites that Portuguese pages link to are captured too — probes for “climate change”, “quantum computing” and “Federal Reserve interest rates” each return millions of estimated hits led by ipcc.ch, microsoft.com and federalreserve.gov. The full-text index lags the crawl: it covers captures up to about 2020, so for newer captures of a known URL use arquivo_url_history. Supports a date window (YYYY or YYYYMMDD), restricting to one site, restricting to a file type (pdf, html, doc), exclusion with a leading minus (Albert -Einstein), and paging by offset. Do not pass a URL as the query; use arquivo_url_history for that.
Parameters
| Name | Type | Required | Description |
|---|---|---|---|
query | string | yes | Search terms. Quote a phrase for an exact match (“inteligência artificial”); prefix a term with - to exclude it. Portuguese or English both work. Must not be a URL. |
from | string | no | Earliest capture date, as YYYY, YYYYMMDD or YYYYMMDDhhmmss. Default 1996. |
to | string | no | Latest capture date, as YYYY, YYYYMMDD or YYYYMMDDhhmmss. Omit it unless you need a window: the full-text index ends in 2020. |
site | string | no | Restrict hits to one site, e.g. “publico.pt” or “http://www.publico.pt”. |
type | string | no | Restrict to a file type by MIME subtype: pdf, html, doc, xls, ppt, rtf, ps. |
limit | number | no | Hits to return (default 10, max ${MAX_ITEMS}). |
offset | number | no | Hit offset for paging (default 0). The response carries next_offset when more hits exist. |
per_site | number | no | Maximum hits per site, to spread results across sources (default 2 — the upstream default). Raise it to see more captures of the same site. |
Example call
Arguments
{
"query": "\"inteligência artificial\"",
"limit": 5
}
curl
curl -X POST https://gateway.pipeworx.io/arquivo-pt/mcp \
-H 'Content-Type: application/json' \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"arquivo_search_pages","arguments":{"query":"\"inteligência artificial\"","limit":5}}}'
TypeScript (@pipeworx/sdk)
import { Pipeworx } from '@pipeworx/sdk';
const pipeworx = new Pipeworx();
const result = await pipeworx.call('arquivo_search_pages', {
"query": "\"inteligência artificial\"",
"limit": 5
});
Connect
Add this to your MCP client config — every tool in the catalog, including this one — or use one-click install buttons:
{
"mcpServers": {
"pipeworx": {
"url": "https://pipeworx.io/mcp"
}
}
}
Connect to just the arquivo-pt pack
{
"mcpServers": {
"arquivo-pt": {
"url": "https://gateway.pipeworx.io/arquivo-pt/mcp"
}
}
}
See Getting Started for client-specific install steps.