arquivo_search_pages

Pack: arquivo-pt · Connect: https://pipeworx.io/mcp (see Connect below for a single-pack URL)

No MCP client? Call it directly: GET https://gateway.pipeworx.io/v1/tools/arquivo_search_pages for the schema, then POST the same URL with its arguments for the data.

Full-text search of archived web pages in Arquivo.pt, the Portuguese web archive (captures since 1996), when the URL is unknown. Finds pages whose TEXT contains the given terms or a quoted phrase and returns, per hit, the page title, original URL, capture date, a text snippet with the matching terms, the archived (replay) URL, and links to the extracted text and a screenshot. Coverage: centred on the Portuguese web (.pt sites and Portuguese-language pages) but international English-language sites that Portuguese pages link to are captured too — probes for “climate change”, “quantum computing” and “Federal Reserve interest rates” each return millions of estimated hits led by ipcc.ch, microsoft.com and federalreserve.gov. The full-text index lags the crawl: it covers captures up to about 2020, so for newer captures of a known URL use arquivo_url_history. Supports a date window (YYYY or YYYYMMDD), restricting to one site, restricting to a file type (pdf, html, doc), exclusion with a leading minus (Albert -Einstein), and paging by offset. Do not pass a URL as the query; use arquivo_url_history for that.

Parameters

NameTypeRequiredDescription
querystringyesSearch terms. Quote a phrase for an exact match (“inteligência artificial”); prefix a term with - to exclude it. Portuguese or English both work. Must not be a URL.
fromstringnoEarliest capture date, as YYYY, YYYYMMDD or YYYYMMDDhhmmss. Default 1996.
tostringnoLatest capture date, as YYYY, YYYYMMDD or YYYYMMDDhhmmss. Omit it unless you need a window: the full-text index ends in 2020.
sitestringnoRestrict hits to one site, e.g. “publico.pt” or “http://www.publico.pt”.
typestringnoRestrict to a file type by MIME subtype: pdf, html, doc, xls, ppt, rtf, ps.
limitnumbernoHits to return (default 10, max ${MAX_ITEMS}).
offsetnumbernoHit offset for paging (default 0). The response carries next_offset when more hits exist.
per_sitenumbernoMaximum hits per site, to spread results across sources (default 2 — the upstream default). Raise it to see more captures of the same site.

Example call

Arguments

{
  "query": "\"inteligência artificial\"",
  "limit": 5
}

curl

curl -X POST https://gateway.pipeworx.io/arquivo-pt/mcp \
  -H 'Content-Type: application/json' \
  -d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"arquivo_search_pages","arguments":{"query":"\"inteligência artificial\"","limit":5}}}'

TypeScript (@pipeworx/sdk)

import { Pipeworx } from '@pipeworx/sdk';
const pipeworx = new Pipeworx();

const result = await pipeworx.call('arquivo_search_pages', {
  "query": "\"inteligência artificial\"",
  "limit": 5
});

Connect

Add this to your MCP client config — every tool in the catalog, including this one — or use one-click install buttons:

{
  "mcpServers": {
    "pipeworx": {
      "url": "https://pipeworx.io/mcp"
    }
  }
}
Connect to just the arquivo-pt pack
{
  "mcpServers": {
    "arquivo-pt": {
      "url": "https://gateway.pipeworx.io/arquivo-pt/mcp"
    }
  }
}

See Getting Started for client-specific install steps.

Regenerated from source · build October 8, 2026