querido-diario-mcp-server
This server provides read-only, local-first MCP tools to discover Brazilian municipalities and search their official gazettes via the Querido Diário API.
search_cities: Find municipalities by partial name (case-insensitive), optionally filtered by Brazilian state code, and resolve their 7-digit IBGE territory IDs.
get_city: Fetch detailed metadata for a single municipality using its exact territory ID (name, state, availability date, publication URLs).
search_gazettes: Full-text search over gazette content with flexible query syntax (required/excluded terms, exact phrases), filter by city, date range, pagination (max 50 per call), and sorting by relevance or date. Returns structured results with excerpts, URLs, and metadata.
No API key or telemetry: Runs locally via stdio; only makes read-only HTTPS requests to the public Querido Diário API.
Typed, validated I/O: All inputs are validated; outputs are structured Pydantic models, with explicit error handling for upstream failures.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@querido-diario-mcp-serverFind mentions of 'licitação' in Belo Horizonte gazettes since January 2024."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Querido Diário MCP Server
A local-first Model Context Protocol server that gives AI clients structured, read-only access to Brazilian municipal official gazettes indexed by Querido Diário.
No API key · Runs locally · Read-only · No telemetry · Open source
The package is published on PyPI and the MCP Registry.
What it does
Querido Diário provides searchable municipal official gazettes through a public API. This server exposes a small, typed MCP interface on top of that API so an MCP client can:
resolve Brazilian municipalities and their IBGE territory IDs;
search indexed official gazettes by text, municipality and date range;
receive structured results that are easier for agents to inspect and combine with other tools.
The server runs as a local stdio subprocess and only performs read-only HTTPS requests to the public Querido Diário API.
Related MCP server: Prefeitura PE Recife: CPOM
Quick start
Requires Python 3.12+ and uv.
No repository clone is required:
uvx querido-diario-mcp-serverThis starts the MCP server on stdio. It is intended to be launched by an MCP client rather than used as an interactive terminal application.
Client configuration
All supported clients use the same command:
uvx querido-diario-mcp-serverClaude Desktop / Claude Code
{
"mcpServers": {
"querido-diario": {
"command": "uvx",
"args": ["querido-diario-mcp-server"]
}
}
}Use this in claude_desktop_config.json for Claude Desktop or a project-level .mcp.json for Claude Code.
Cursor
{
"mcpServers": {
"querido-diario": {
"command": "uvx",
"args": ["querido-diario-mcp-server"]
}
}
}Add it to .cursor/mcp.json or your global Cursor MCP configuration.
Codex CLI
[mcp_servers.querido-diario]
command = "uvx"
args = ["querido-diario-mcp-server"]Other MCP clients that support local stdio servers can launch the same command.
Tools
The server intentionally exposes a small, read-only tool surface.
Tool | Purpose |
| Search Brazilian municipalities by name and resolve the 7-digit IBGE territory ID. Supports an optional state filter. |
| Fetch one municipality by its exact 7-digit IBGE territory ID. |
| Full-text search over indexed gazettes with municipality, publication date, pagination and sorting filters. |
search_gazettes uses the simple-query-string syntax supported by the upstream search API. Examples include +required, -excluded and "exact phrase".
Example
A user can ask:
Find mentions of ACME Ltda in Porto Alegre gazettes
from January through July 2026.An MCP client can then execute:
1. search_cities(city_name="Porto Alegre")
-> territory_id: "4314902"
2. search_gazettes(
query='"ACME Ltda"',
territory_ids=["4314902"],
published_since="2026-01-01",
published_until="2026-07-31",
)The result is returned as structured tool output for the client to summarize, filter or combine with other tools.
Architecture
MCP client
|
| stdio / MCP
v
server.py
|
| validated tool arguments
v
client.py
|
| typed read-only HTTPS requests
v
Querido Diário public APIThe implementation keeps the protocol boundary separate from the HTTP integration:
src/querido_diario_mcp_server/
__init__.py package version and CLI entry point
config.py environment-driven configuration
errors.py integration error hierarchy
models.py typed Pydantic domain models
client.py async HTTP client with no MCP dependency
server.py MCP lifespan, validation and tool definitionsA single httpx.AsyncClient connection pool is created for the server lifecycle and reused across tool calls. The HTTP client can be tested independently of MCP, while integration tests exercise the real MCP server in process.
Design principles
Read-only by design. The integration only calls a fixed set of upstream GET endpoints.
Local-first. There is no hosted middleware, proprietary backend or account requirement.
Typed boundaries. API responses and MCP outputs use explicit Pydantic models.
Small tool surface. The server exposes only the operations required to discover cities and search gazettes.
Explicit failures. Upstream failures are converted to concise MCP tool errors rather than raw tracebacks or HTML responses.
No arbitrary URL fetching. Tool callers cannot turn the server into a generic HTTP or SSRF primitive.
No telemetry. The server does not collect usage data.
Security
The server never automatically follows gazette URLs returned by the upstream API and does not accept arbitrary target URLs from tool callers. Territory IDs, dates, pagination and sort options are validated before use.
For vulnerability reporting and the supported security scope, see SECURITY.md.
Configuration
Variable | Default | Purpose |
|
| Base URL for the Querido Diário API. Can be overridden for local or staging environments. |
Normal users do not need to set any environment variables.
Development
Clone the repository only if you want to contribute or work on the implementation:
git clone https://github.com/lucaspmgomess/querido-diario-mcp-server.git
cd querido-diario-mcp-server
uv syncRun the same checks enforced by CI:
uv run ruff check .
uv run ruff format --check .
uv run pyright
uv run pytest --covRun the server from the checkout:
uv run querido-diario-mcp-serverInspect the MCP tools interactively:
uv run mcp dev src/querido_diario_mcp_server/server.py:mcpTest strategy
Unit tests mock the HTTP boundary with httpx.MockTransport, so the automated suite does not depend on the public API.
Coverage includes:
successful city and gazette searches;
query serialization and repeated territory IDs;
date ranges, pagination and sorting;
empty result sets;
4xx and 5xx upstream failures;
malformed responses;
connection errors and timeouts;
MCP tool discovery, schemas and structured outputs;
validation errors and upstream failures at the MCP boundary.
Manual smoke-test scripts are available for validating the real production API and the full MCP-to-API path:
uv run python scripts/smoke_test_api.py
uv run python scripts/smoke_test_mcp.pyProject status
The project is currently in beta. The MCP tool surface is intentionally small and changes are kept conservative. Release history is documented in CHANGELOG.md.
The following are intentionally outside the current scope:
arbitrary URL fetching;
automatic PDF or full-text download;
OCR;
write operations;
crawling and background jobs;
built-in LLM summarization;
a web interface.
Relationship to Querido Diário
This is a community-built, unofficial project.
Querido Diário is maintained by Open Knowledge Brasil and its community. This repository is not an official Open Knowledge Brasil project, is not endorsed or maintained by the organization, and only consumes the project's public API.
Relevant upstream repositories:
Additional notes about upstream endpoint history are kept in docs/upstream-api-history.md.
Contributing
Issues and pull requests are welcome. See CONTRIBUTING.md for the development workflow, project scope and review checklist.
License
Available Tools
3 toolsget_cityA
Retrieve one Brazilian municipality by its 7-digit IBGE territory ID.
Use this when the exact territory_id is already known (from `search_cities`, or
supplied directly) and you need that city's details: name, state, Querido Diário
availability date, and known official publication URLs. If you only have a city
name, call `search_cities` instead.
| Name | Required | Description | Default |
|---|---|---|---|
| territory_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| level | No | |
| state_code | Yes | |
| territory_id | Yes | |
| territory_name | Yes | |
| publication_urls | No | |
| availability_date | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. 'Retrieve' signals a read-only lookup, and the description discloses what the result contains: name, state, availability date, and official publication URLs. It does not discuss not-found behavior, but for a simple single-ID lookup this is acceptable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short, front-loaded with the primary action, and every sentence is informative. The usage guidance and alternative routing are placed after the purpose without padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple: one required parameter, no nested objects, an output schema exists, and the description names the only two sibling tools. It covers what the tool returns, when to use it, and when not to use it. Nothing essential is missing for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema only lists territory_id as a plain string with no description. The description compensates by explaining it is a 7-digit IBGE territory ID and that it typically comes from search_cities. This gives the agent enough semantic grounding to select and format the parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Retrieve one Brazilian municipality by its 7-digit IBGE territory ID.' The phrase 'one' and the ID-based lookup clearly distinguish it from the fuzzy/name-based sibling search_cities.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use this tool: when the exact territory_id is already known. It also names the alternative and its trigger: 'If you only have a city name, call search_cities instead.' This is direct routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_citiesA
Find Brazilian municipalities indexed by Querido Diário and resolve their 7-digit IBGE territory IDs.
Call this FIRST whenever a request names a city but you don't already have its
numeric territory ID — both `get_city` and `search_gazettes` require that ID, not
a city name. For example, to answer "find gazettes about ACME in Porto Alegre",
first call search_cities(city_name="Porto Alegre") to get its territory_id, then
pass that ID to search_gazettes.
Matching is by partial, case-insensitive name similarity (as implemented by the
upstream API), not exact string equality, so it tolerates minor spelling
variation. `state_code` is an optional two-letter Brazilian state/UF filter
(e.g. "RS", "SP") applied after the upstream search, useful for disambiguating
cities that share a name across different states.
| Name | Required | Description | Default |
|---|---|---|---|
| city_name | Yes | ||
| state_code | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| cities | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It clearly discloses the matching semantics (partial, case-insensitive similarity rather than exact equality) and the post-filtering behavior of state_code. It does not go into return shape or pagination, but an output schema exists and the core behavior is well explained.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose, immediately followed by usage context and parameter semantics. Every paragraph earns its place, including the concrete example that clarifies the intended call flow without excessive verbosity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter search tool with no annotations, the description covers purpose, when to call it, how to call it, parameter behavior, and the relationship to sibling tools. An output schema exists, so the lack of return-format detail is acceptable. Nothing essential for correct tool invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description fully compensates. It explains city_name as a partial, case-insensitive name query and state_code as an optional two-letter Brazilian state/UF filter applied after the upstream search, including examples like 'RS' and 'SP' and its disambiguation purpose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Find') and a concrete resource ('Brazilian municipalities indexed by Querido Diário'), and explains the tool's unique role: resolving 7-digit IBGE territory IDs. It explicitly differentiates from siblings by noting that both get_city and search_gazettes require the territory ID that this tool provides.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use guidance: 'Call this FIRST whenever a request names a city but you don't already have its numeric territory ID.' It names the sibling tools and gives a concrete example workflow, effectively telling the agent when to use this tool versus alternatives and what to do after.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_gazettesA
Search the text of published Brazilian municipal official gazettes.
Each result is one gazette (one city's official publication, one edition, one
date) with a short excerpt showing where `query` matched. This searches gazette
*content*, not city metadata — resolve a city name to its territory_id with
`search_cities` first.
`query` uses OpenSearch's "simple query string syntax": bare words are OR'd
together, a leading `+` requires a term, a leading `-` excludes it, and double
quotes match an exact phrase — e.g. '"João da Silva"' for an exact name, or
'+licitação +pregão' to require both terms together. An empty query returns
matching gazettes' metadata without excerpts.
`territory_ids` restricts the search to specific cities by their 7-digit IBGE ID
(from `search_cities` or `get_city`); omit it to search across all available
cities. `published_since` / `published_until` are inclusive ISO dates
(YYYY-MM-DD) bounding the gazette's publication date — leave both unset to search
the full available history. `size` is capped at 50 results per call; use `offset`
to page through more. `sort_by` defaults to relevance; use "descending_date" or
"ascending_date" to sort chronologically instead.
Example: to find mentions of "Empresa X" in Porto Alegre gazettes published
between January and July 2026, once territory_id "4314902" is known, call
search_gazettes(query='"Empresa X"', territory_ids=["4314902"],
published_since="2026-01-01", published_until="2026-07-31").
| Name | Required | Description | Default |
|---|---|---|---|
| size | No | ||
| query | No | ||
| offset | No | ||
| sort_by | No | relevance | |
| territory_ids | No | ||
| published_since | No | ||
| published_until | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| gazettes | Yes | |
| total_gazettes | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so thoroughly: each result is one gazette with an excerpt, empty queries return metadata without excerpts, dates are inclusive, size is capped at 50, and sort defaults to relevance. This goes well beyond the schema in explaining actual behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but justified by seven undocumented parameters and no annotations. It is front-loaded with the core semantics, organized into focused paragraphs per concern, and closes with a complete example, with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity, optional parameters, and lack of annotations/schema descriptions, the description is complete: it covers result units, matching behavior, all filters, defaults, limits, pagination, and a realistic example. The presence of an output schema means return-value details are already covered elsewhere.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain every parameter, and it does: query syntax (OR, +, -, quoted phrases), territory_ids as 7-digit IBGE IDs, inclusive ISO dates, size/offset pagination, and sort_by values. The included example ties the parameters together concretely.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Search the text of published Brazilian municipal official gazettes.' It then explicitly distinguishes gazette content from city metadata and points to search_cities, making the purpose unmistakable and differentiating it from sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear workflow: resolve a city name to territory_id with search_cities first, or omit territory_ids to search all cities. It also explains query syntax, date filtering, pagination, and sort options, so an agent knows exactly when and how to invoke the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.0- First observed
get_city - First observed
search_cities - First observed
search_gazettes
TDQS
Scored across 3 tools
Each tool has a clearly distinct role: search_cities resolves city names to IDs, get_city retrieves metadata by ID, and search_gazettes searches gazette content. No two tools overlap in purpose or output.
All tool names follow a consistent verb_noun snake_case pattern: search_cities, get_city, search_gazettes. The singular/plural difference for 'city' is natural and does not create confusion.
Three tools is a well-scoped set for a focused gazette search server. Each tool is necessary and covers a distinct step in the core workflow: resolve city, inspect city, search gazettes.
The core search workflow is fully covered: city name resolution, city detail lookup, and gazette full-text search with filtering. Minor gaps exist (e.g., no direct single-gazette fetch endpoint or full-text retrieval beyond excerpts), but they do not break the primary use case.
Maintenance
Related MCP Connectors
Searches MUNICIPAL official gazettes from thousands of city halls (Querido Diário / Open Knowledge B
Brazilian legal stack in one MCP: lawsuits, court publications, case law, tenders, certificates.
MCP server for Brazilian Federal Senate open data (legislative, administrative, e-Cidadania).
Discover, resolve, and query official Brazilian economic data with semantic search and provenance.
Related MCP Servers
- AlicenseAqualityBmaintenanceExposes official IBGE data as MCP tools, including Brazilian localities, SIDRA statistical aggregates, and population indicators.11MIT
- AlicenseNot gradedqualityCmaintenanceEnables querying official data from the Prefeitura de Recife (Recife City Hall) through a hosted, read-only MCP server. Supports MCP over HTTP with usage-based prepaid billing.MIT
- AlicenseNot gradedqualityCmaintenanceProvides read-only access to Brazilian municipality codes from IBGE, an official government source, via MCP. Works with any MCP-compatible client to query municipality data using natural language.MIT
- AlicenseNot gradedqualityCmaintenanceEnables querying the Brazilian government's transparency portal (Portal da Transparência) to search official public spending and transparency data. Read-only access via natural language from any MCP client, with prepaid pay-per-query usage.MIT