CensoSenso MCP
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@CensoSenso MCPCompare a população de São Paulo e Rio de Janeiro no Censo 2022"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
CensoSenso MCP
Servidor MCP independente para consulta, comparação e análise de dados oficiais do IBGE, com 24 ferramentas, procedência reproduzível, validação de contratos e transporte remoto via Cloudflare Workers.
Endpoint MCP remoto: https://censosenso.poderdapalavra.org/mcp
Protótipo público 1. A versão 0.7.0 é experimental, somente leitura e disponibilizada sem SLA. O canal oficial inicial é o endpoint remoto acima. O pacote npm ainda não integra este lançamento.
O CensoSenso MCP é uma implementação de Model Context Protocol voltada ao uso analítico de dados públicos brasileiros. A superfície atual reúne ferramentas de localidades, Censo Demográfico, SIDRA, indicadores, Cidades@, saúde, saneamento, economia, classificações, malhas administrativas, vizinhança municipal, recortes territoriais, notícias e calendário.
O princípio central é simples: uma resposta útil não deve trazer apenas um número. Sempre que o contrato da ferramenta permitir, o resultado preserva fonte, URL reproduzível, período ou vintage, instante de extração, indicação de derivação e atributos de qualidade.
Sumário
Related MCP server: ibge-data-mcp
Política de versionamento
A sequência pública usa versões pré-1.0: 0.x.0 para acréscimo funcional e 0.x.y para correções ou melhorias de funcionalidades existentes. A família 0.9.x é reservada ao beta e 1.0.0 ao primeiro lançamento estável. Consulte docs/VERSIONAMENTO.md.
Versão e estado
versão do servidor: 0.7.0;
linguagem principal: TypeScript;
runtime: Node.js 22+;
protocolo: Model Context Protocol;
transporte local: STDIO;
transporte remoto: Streamable HTTP;
hospedagem remota: Cloudflare Workers;
endpoint:
https://censosenso.poderdapalavra.org/mcp;health:
https://censosenso.poderdapalavra.org/health;status:
https://censosenso.poderdapalavra.org/status;licença do software: MIT;
fonte principal dos dados: IBGE;
natureza deste lançamento: protótipo público experimental, sem SLA.
A superfície atual possui 24 ferramentas: 22 componentes ibge_* e duas ferramentas genéricas de pesquisa assistida, search e fetch.
Documentação técnica
O deploy remoto foi homologado com a sequência MCP real:
initialize;emissão/captura de
mcp-session-id;notifications/initialized;tools/list;tools/call;consulta efetiva a fonte oficial;
retorno estruturado com procedência.
Validação dos 645 municípios paulistas
A evidência completa desta execução está em docs/VALIDACAO_SP_645.md. O resultado municipal público pode ser consultado em validation/sp/0.6.0/municipios-sp-645.csv.
A versão 0.6.0 possui um gate reproduzível específico para o universo municipal do Estado de São Paulo.
Execução validada em 5 de outubro de 2026:
Métrica | Resultado |
Municípios retornados pela API oficial | 645 |
Validação estrutural | aprovada |
Divergências globais | 0 |
Divergências municipais | 0 |
Amostra determinística de chamadas reais | 21 |
Amostra aprovada | 21/21 |
Falhas na amostra | 0 |
SHA-256 do snapshot validado |
|
O gate consulta:
https://servicodados.ibge.gov.br/api/v1/localidades/estados/35/municipios?orderBy=nomeO que é verificado nos 645 registros
Para cada município:
código IBGE com exatamente 7 dígitos;
prefixo
35, correspondente ao Estado de São Paulo;unicidade dos códigos;
nome não vazio;
unicidade dos nomes normalizados em pt-BR;
código da UF igual a 35, quando informado pela origem;
sigla da UF igual a SP, quando informada;
região igual a Sudeste, quando informada.
No conjunto:
total obrigatório de 645 registros;
inexistência de códigos duplicados;
inexistência de qualquer registro marcado como divergente.
Como a amostra de 21 municípios é construída
Além da validação estrutural 645/645, o gate cria uma amostra determinística para chamar as ferramentas reais do servidor.
A amostra começa por:
município de São Paulo;
Guararema;
primeiro e último código na ordenação;
primeiro e último nome na ordenação alfabética.
Depois inclui municípios nos quantis de código:
10%;
25%;
50%;
75%;
90%.
Também inclui:
ao menos um município de cada região geográfica intermediária encontrada;
um nome contendo caracteres acentuados;
um nome composto com espaço.
As duplicidades são removidas por código e o resultado é ordenado deterministicamente.
Cada município da amostra é validado simultaneamente por:
ibge_geocodigo(codigo=...)
ibge_localidade(codigo=...)O gate compara código, nome e UF retornados pelas ferramentas com o registro oficial usado como referência.
Artefatos produzidos pelo gate
A execução gera:
snapshot-candidato.json;meta.json;validacao-645.jsonl;amostra.json;amostra-resultados.json;resumo.json.
O snapshot é serializado de forma determinística e recebe SHA-256. A promoção para snapshot versionado só é permitida se não houver divergência estrutural nem falha na amostra real.
Comando:
npm run gate:sp-645Arquitetura
flowchart LR
A[Cliente MCP remoto] -->|Streamable HTTP| B[Cloudflare Worker]
A2[Cliente MCP local] -->|STDIO| C[Servidor local]
B --> D[buildServer / registerAll]
C --> D
D --> E{Ferramenta}
E --> F[Localidades / Cidades / CNAE / Países]
E --> G[SIDRA / Censo / Indicadores / Saúde]
E --> H[Malhas / Vizinhos]
E --> I[Recortes temáticos]
E --> J[search / fetch]
F --> K[APIs oficiais IBGE]
G --> K
H --> K
I --> L[Snapshots oficiais versionados]
J --> K
D --> M[Validação + cache + métricas]
K --> N[Adapter de domínio]
L --> N
N --> O[Procedência]
O --> P[content + structuredContent + _meta]A mesma lógica de registro das ferramentas é usada pelo servidor STDIO e pelo Worker HTTP. O Worker acrescenta controles de transporte, segurança e operação, mas não mantém uma implementação paralela das regras de domínio.
Árvore funcional
src/
server.ts registro MCP e instruções
cache.ts cache TTL e metadata de extração
retry.ts retry, timeout e erros upstream
provenance.ts adapter de procedência IBGE
stats.ts adapter estatístico SIDRA
pagination.ts guarda de cursor
validation.ts validações compartilhadas
tools/ ferramentas de domínio
data/recortes/ snapshots territoriais versionados
worker/
src/ transporte Cloudflare
tests/ testes do Worker
scripts/ smoke e testes HTTP
wrangler.jsonc configuração de deploy
scripts/
gate-sp-645.mjs validação dos 645 municípios de SP
normalizar_recortes.py ETL dos recortes territoriais
smoke-mcp.mjs smoke STDIO/HTTP
run_integration_gate.ps1 contratos reais
tests/ testes unitários e de integração
evals/ avaliação de seleção de ferramentas
baselines/ superfície MCP versionada
docs/ arquitetura, auditorias e metodologiaFerramentas disponíveis
Para schemas, parâmetros, decisões, fontes e limitações ferramenta por ferramenta, consulte docs/FERRAMENTAS.md. Para a visão consolidada de origem, cache e derivação, consulte docs/MATRIZ_FONTES_E_PROCESSOS.md.
Domínio | Ferramentas |
Localidades |
|
SIDRA |
|
Censo e indicadores |
|
Geografia |
|
Economia/classificações |
|
Saúde e condições de vida |
|
Nomes |
|
Países |
|
Informação corrente |
|
Pesquisa assistida |
|
Seleção recomendada
Pergunta | Ferramenta |
Panorama de um município |
|
Censo por tema |
|
Série econômica/social conhecida |
|
Comparar localidades |
|
Consultar tabela SIDRA |
|
Descobrir tabela |
|
Inspecionar dimensões/classificações |
|
Resolver nome/código IBGE |
|
Geometria administrativa |
|
Mapa coroplético municipal |
|
Vizinhança municipal |
|
Recorte territorial especial |
|
Fluxo de uma chamada
flowchart TD
A[tools/call] --> B[Schema Zod estrito]
B -->|inválido| C[Erro explícito]
B -->|válido| D[Handler]
D --> E{Origem}
E -->|API| F[cachedFetch]
E -->|snapshot| G[JSON versionado]
F --> H[fetchWithRetry]
H --> I[Resposta oficial]
G --> I
I --> J[Adapter de domínio]
J --> K{estatísticas?}
K -->|sim| L[mcp-stats]
K -->|não| M[resultado de domínio]
L --> M
M --> N[provenienciaIbge]
N --> O[toMcpResult]
O --> P[content]
O --> Q[structuredContent]
O --> R[_meta]A política do projeto é recusar ambiguidades relevantes em vez de escolher silenciosamente uma interpretação plausível.
Algoritmos e regras de decisão
1. Validação territorial
isValidIbgeCode interpreta o tamanho do código:
1 dígito → região;
2 dígitos → UF;
7 dígitos → município;
9 dígitos → distrito.
Para municípios e distritos, os dois primeiros dígitos precisam corresponder a uma UF oficial conhecida.
2. Normalização de UF
Entradas como:
SP
São Paulo
35são resolvidas por um único resolver compartilhado, evitando lógicas divergentes entre ferramentas.
3. Resolução em ibge_geocodigo
Quando recebe codigo, a ferramenta remove caracteres não numéricos e decide o nível territorial pelo comprimento.
Quando recebe nome:
verifica correspondência com UF;
verifica correspondência com região;
consulta municípios;
restringe por UF, quando informada;
usa busca textual por inclusão do nome;
limita a 20 correspondências;
se houver exatamente uma correspondência, retorna imediatamente a hierarquia detalhada;
se houver várias, retorna lista e exige escolha explícita.
A ferramenta não usa correspondência fuzzy probabilística para inventar uma localidade.
4. Datas
Entradas aceitas:
DD/MM/AAAA;DD-MM-AAAA;AAAA-MM-DD.
A ordem mês-dia-ano fornecida pelo usuário não é aceita por ser ambígua no contexto brasileiro.
Quando uma API do IBGE exige MM-DD-AAAA, a conversão é feita somente depois da validação da entrada.
5. Períodos SIDRA
São aceitos, entre outros:
last;first;all;last N;ano
AAAA;intervalo
AAAA-AAAA;mês
AAAAMM;trimestre no formato usado pelo SIDRA;
múltiplos períodos separados por vírgula.
6. Paginação MCP
As listas MCP atuais são entregues em página única e não emitem nextCursor.
Por isso, se o cliente envia cursor em:
tools/list;resources/list;resources/templates/list;prompts/list;
o servidor responde JSON-RPC -32602 Invalid params, em vez de ignorar um cursor inválido.
7. Algoritmos delegados
Alguns algoritmos não são reimplementados localmente:
estatística genérica →
@sbissoli/mcp-stats;ranking do índice
search→@sbissoli/mcp-search;modelo canônico de procedência →
@sbissoli/mcp-provenance;contiguidade topológica →
@turf/boolean-touches;centróide →
@turf/centroid;distância geodésica →
@turf/distance.
O CensoSenso implementa os adapters, regras de decisão, seleção de colunas, tratamento de erros, fontes, schemas e contratos ao redor desses algoritmos.
Extração, cache e tolerância a falhas
Cache
O cache é em memória por processo/isolate.
TTL padrão:
Classe | TTL |
| 24 horas |
| 1 hora |
| 15 minutos |
| 1 minuto |
A chave de cache é construída com parâmetros:
parâmetros
undefinedsão removidos;nomes são ordenados lexicograficamente;
a representação
k=vé concatenada à base.
Isso evita chaves distintas para a mesma consulta apenas por diferença na ordem dos parâmetros.
Data real de extração
Cada entrada de cache preserva a data do fetch real.
Em cache hit:
o valor pode ser servido novamente;
served_from_cache=true;retrieved_atcontinua sendo o instante original de obtenção upstream.
Isso evita declarar como “extraído agora” um dado que foi obtido anteriormente.
Retry
Padrão:
máximo: 4 retries além da tentativa inicial;
atraso inicial: 2.000 ms;
multiplicador: 2;
atraso máximo: 16.000 ms.
Fórmula:
delay = min(initialDelay × multiplier^(attempt-1), maxDelay)Status HTTP retryable:
429, 500, 502, 503, 504Também são considerados transitórios erros de rede como:
ECONNREFUSED;ECONNRESET;ETIMEDOUT;ENOTFOUND;EAI_AGAIN;conexão recusada/resetada;
timeout.
Cada tentativa recebe AbortSignal próprio e timeout configurável.
Erros da fonte
Quando uma API devolve erro:
o código HTTP é preservado;
o corpo textual é lido;
páginas HTML de borda são descartadas;
JSON de erro é inspecionado por
message,mensagem,erro,erroroudetail;espaços são normalizados;
o texto é truncado em 300 caracteres.
A intenção é devolver ao cliente a razão fornecida pela fonte — por exemplo, “nível territorial incompatível” — em vez de reduzir tudo a “HTTP 400”.
Procedência e auditabilidade
O CensoSenso usa um adapter sobre @sbissoli/mcp-provenance.
Contexto:
namespace:
io.github.hilaliskandar.censosenso;locale:
pt-BR;timezone: Brasília, UTC-03:00;
modo padrão:
concise.
A procedência pode ser emitida em três canais:
structuredContent.provenance;structuredContent.attribution;_metanamespaced.
O Markdown legível também pode receber rodapé compacto.
Fontes registradas
Entre as fontes oficiais usadas:
API de Localidades;
SIDRA;
API de Agregados;
API de Nomes;
API de Malhas;
GeoFTP;
API de Notícias;
Projeções de População;
CNAE;
Calendário;
Países;
Pesquisas/Cidades@.
Derivação
Regra:
filtrar, paginar ou reserializar dado bruto → não é derivação estatística;
calcular média, mediana, percentis, rankings ou agregações → resultado marcado como derivado e acompanhado de nota metodológica.
SIDRA e estatísticas
Fluxo recomendado
ibge_sidra_tabelas
↓
ibge_sidra_metadados
↓
ibge_sidraWrappers especializados simplificam temas frequentes:
ibge_censo;ibge_indicadores;ibge_datasaude;ibge_comparar;ibge_cidades.
Estatísticas sobre o conjunto completo
Quando estatisticas=true, o cálculo ocorre antes da paginação/truncagem.
São produzidos:
n;
soma;
mínimo;
máximo;
média;
mediana;
desvio-padrão;
percentis;
top;
bottom.
Marcadores SIDRA excluídos da distribuição:
-
..
...
XValores não numéricos também ficam fora de n e são contabilizados em registrosSemValor.
Resolução de agruparPor
A coluna solicitada é resolvida nesta ordem:
correspondência exata sem diferença de caixa/acento;
aliases conhecidos;
correspondência parcial bidirecional;
se houver mais de uma candidata, a operação é recusada como ambígua.
Exemplos de aliases:
UF,estado→Unidade da Federação;cidade→Município;região→Grande Região;indicador→Variável.
Quando SIDRA publica simultaneamente “Unidade da Federação” e “Unidade da Federação (Código)”, o rótulo legível é preferido porque os grupos são semanticamente equivalentes.
Mistura de variáveis
Se uma consulta traz várias variáveis e o usuário não escolhe agrupamento, o servidor agrupa automaticamente por Variável para evitar calcular uma única distribuição sobre unidades potencialmente incompatíveis.
Geografia, vizinhança e recortes territoriais
Malhas administrativas
ibge_malhas consulta a API oficial de Malhas para geometrias administrativas.
Mapas coropléticos
ibge_mapa combina indicadores municipais auditados do SIDRA com a malha municipal oficial do IBGE e produz SVG vetorial para recortes explícitos de até 50 municípios de uma mesma UF. A primeira versão oferece classificação por quantis ou intervalos iguais e registra o produto como derivado.
Vizinhança municipal
Sem raio:
geometria do município
↓
candidatos territoriais
↓
@turf/boolean-touches
↓
municípios topologicamente contíguosNeste modo, “vizinho” significa tocar a fronteira do município.
Com raio:
geometria
↓
centroid
↓
distance
↓
filtro por kmEsse modo é aproximação por distância entre centróides e é declarado como tal.
Recortes temáticos
A auditoria do projeto encontrou bloqueio HTTP 403 no WFS temático para clientes automatizados. Para evitar que uma origem instável fizesse parte do caminho crítico, os atributos e composições foram migrados para snapshots oficiais versionados.
Fluxo:
fonte oficial IBGE/GeoFTP
↓
inspeção de arquivo
↓
validação de estrutura/vintage/encoding
↓
normalização
↓
JSON versionado
↓
recortes-gerados.ts
↓
ibge_malhas_temaRecortes atualmente versionados:
Amazônia Legal 2024;
Biomas 2025;
Semiárido 2022;
Zona Costeira 2021;
Faixa de Fronteira 2024;
Regiões Metropolitanas 2025;
RIDEs 2025.
A ferramenta retorna composição e atributos. Geometria administrativa continua em ibge_malhas.
Pesquisa assistida search/fetch
O índice de pesquisa é construído na primeira chamada e agrega:
catálogo SIDRA;
municípios;
indicadores conhecidos;
temas do Censo;
indicadores de saúde;
recortes territoriais.
A construção consulta em paralelo:
catálogo oficial de agregados;
lista nivelada de municípios.
Depois acrescenta os dicionários auditados locais.
O índice possui TTL equivalente a CACHE_TTL.STATIC — 24 horas.
Chamadas simultâneas durante a primeira construção compartilham a mesma Promise, evitando reconstruções concorrentes.
Se a construção falha, o resultado falho não é mantido: a próxima chamada tenta novamente.
search
O ranking genérico é fornecido por @sbissoli/mcp-search. O CensoSenso define:
entradas;
títulos;
palavras-chave;
URLs;
vocabulário de consulta;
fontes;
limite de resultados.
fetch
O ID do resultado determina qual ferramenta real renderiza o documento:
tabela SIDRA → metadados SIDRA;
município → hierarquia + população;
tema Censo → catálogo auditado;
saúde → indicador auditado;
recorte → snapshot versionado;
indicador geral → ferramenta de indicadores.
Assim, fetch não inventa um documento independente: ele recompõe a informação usando os mesmos adapters e contratos das ferramentas de domínio.
Validação, testes e gates
O projeto separa testes determinísticos de contratos que dependem de fontes externas.
Gate completo
npm run gate:deployEncadeia:
checagem do gerador de homologação;
typecheck;
lint;
EOL;
Prettier;
build;
suíte Vitest;
cobertura;
npm cido Worker;typecheck do Worker;
testes do Worker.
Gate dos 645 municípios
npm run gate:sp-645Integração real
Em PowerShell:
.\scripts\run_integration_gate.ps1Smoke
STDIO:
node scripts/smoke-mcp.mjs --stdioRemoto:
node scripts/smoke-mcp.mjs https://censosenso.poderdapalavra.orgBaseline da superfície MCP
node scripts/dump-surface.mjs --stdioBaseline atual:
baselines/surface-stdio-0.7.0.jsonMudança de ferramentas, schemas, resources ou prompts deve produzir diff explícito no baseline.
Instalação local
A árvore homologada do protótipo 0.7.0 está publicada neste repositório. O CI público usa apenas runners hospedados pelo GitHub e não realiza deploy de produção.
Pré-requisitos:
Git;
Node.js 22 ou superior;
npm compatível com o lockfile;
Python somente se for necessário regenerar snapshots territoriais.
Clone:
git clone https://github.com/hilaliskandar/hilaliskandar-censosenso-mcp.git
cd hilaliskandar-censosenso-mcp
npm ciBuild:
npm run buildExecução STDIO:
node dist/index.jsOu:
npm startConfiguração genérica de um cliente MCP local:
{
"mcpServers": {
"censosenso": {
"command": "node",
"args": ["/caminho/para/hilaliskandar-censosenso-mcp/dist/index.js"]
}
}
}A forma exata do arquivo de configuração varia entre clientes MCP.
Uso remoto
Clientes compatíveis com MCP Streamable HTTP devem apontar para:
https://censosenso.poderdapalavra.org/mcpO protótipo público 1:
não exige conta CensoSenso;
não exige OAuth;
não exige API key;
é somente leitura.
Exemplos de intenção
Qual era a população de Campinas no Censo 2022?
Liste os municípios do Espírito Santo.
Compare a população de Campinas, Jundiaí e Sorocaba.
Quais municípios fazem fronteira com Guararema?
Qual é o índice de envelhecimento de Campinas?
Liste os municípios da Amazônia Legal.Exemplos de chamadas
ibge_geocodigo(nome="Campinas", uf="SP")
ibge_censo(ano="2022", tema="estrutura_etaria", nivel_territorial="6", localidades="3509502")
ibge_vizinhos(municipio="3518305")
ibge_malhas_tema(tema="amazonia_legal")Desenvolvimento e reprodução
Worker local:
cd worker
npm ci
npm run devA equivalência funcional de uma reprodução exige, no mínimo:
npm ci;npm run gate:deploy;negociação MCP via STDIO;
negociação MCP via Worker;
tools/listcom superfície esperada;ao menos uma chamada real ao IBGE;
procedência nos canais esperados;
integridade dos snapshots;
ausência de mudança inexplicada no baseline;
resposta válida de
/health,/statuse/mcp.
CI/CD
Fluxo operacional:
push/merge
↓
GitHub
↓
gates
↓
Cloudflare Workers Build
↓
wrangler deploy
↓
health/status/MCPSegurança, privacidade e limites operacionais
Modelo de segurança
ferramentas de dados são read-only;
schemas de entrada são estritos;
Host/Origin são validados no transporte;
HTTPS via Cloudflare;
Bearer opcional existe no código;
OAuth não está habilitado;
/metricsé privado por padrão.
Sem METRICS_API_KEY, /metrics responde:
HTTP 404 Not FoundEsse comportamento foi confirmado no domínio de produção antes da abertura pública.
Rate limiting
O Worker usa token bucket em memória por cliente/IP.
Configuração atual:
burst: 20 tokens;
refill: 5 tokens/s;
máximo: 1.000 buckets por isolate;
evicção FIFO quando o mapa atinge o limite.
Importante: esse limite é por isolate. Não constitui cota global exata.
Telemetria
O projeto pode manter:
contadores agregados no Durable Object
USAGE;logs de método, path, status e duração;
metadata de deploy;
cache em memória.
O código não foi projetado para persistir nos contadores de uso:
valores de argumentos;
conteúdo das perguntas;
datasets retornados.
Leia PRIVACY.md e SECURITY.md.
Limitações
sem SLA;
disponibilidade depende de APIs oficiais externas;
o WFS temático permanece fora do runtime principal;
rate limit não é global;
OAuth ainda não está habilitado;
pacote npm ainda não publicado;
o protótipo pode mudar de contrato em versões posteriores.
Dados, licença e atribuições
Software
Licença: MIT.
Copyright:
Copyright (c) 2026 Carlos Alexandre Gomes <hilaliskandar@gmail.com>Código de origem
Partes deste projeto derivam de:
IBGE Brasil MCP / ibge-br-mcp
Autor original: Sidney da Silva Pereira Bissoli
Repositório: https://github.com/SidneyBissoli/ibge-br-mcp
Licença: MIT.
O aviso MIT original é preservado em THIRD_PARTY_LICENSES.md.
A linhagem e as responsabilidades estão descritas em NOTICE.md.
Dados
O software consulta dados públicos oficiais do IBGE.
A licença MIT deste repositório cobre o software e não substitui os termos aplicáveis aos dados e serviços oficiais consultados.
O CensoSenso:
não é produto oficial do IBGE;
não representa o IBGE;
não implica endosso do IBGE.
Status do protótipo e próximos passos
Concluído:
versão 0.7.0;
24 ferramentas;
transporte STDIO;
Streamable HTTP;
domínio próprio;
homologação remota;
procedência estruturada;
snapshots territoriais;
CI/gates;
hardening de
/metrics;validação dos 645 municípios paulistas;
documentação de segurança e privacidade;
licença e atribuições públicas.
Antes da divulgação ampla:
publicar a árvore homologada completa neste repositório;
executar novamente o gate completo na árvore pública;
executar smoke remoto após a migração;
criar tag/release pública;
iniciar com um grupo externo reduzido;
monitorar erros, uso e rate limiting;
ampliar a divulgação somente após janela inicial estável.
Itens deliberadamente posteriores:
publicação npm;
OAuth obrigatório;
rate limit global rígido;
eventual retorno do WFS temático ao runtime.
Contato
Carlos Alexandre Gomes
GitHub: @hilaliskandar
E-mail: hilaliskandar@gmail.com
Para vulnerabilidades, consulte SECURITY.md. Para contribuições, consulte CONTRIBUTING.md.
Available Tools
23 toolsfetchDocumento para Deep ResearchARead-onlyIdempotent
Returns the full document for an id obtained from search, as { id, title, text, url, metadata }: text is the readable content (Markdown) and url the canonical public page to cite.
Companion of search in the OpenAI Deep Research contract, over the IBGE (Brazilian official statistics: SIDRA tables, municipalities, known indicators) catalog. Only ids returned by search are valid; an unknown id returns an error.
The ibge_* tools (ibge_sidra, ibge_cidades, ibge_indicadores, ibge_comparar…) remain the tools for data queries.
Behavior: read-only and idempotent — a live GET against the public source when the document needs it.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Identificador de um documento devolvido por `search` |
Output Schema
| Name | Required | Description |
|---|---|---|
| id | Yes | Identificador único do documento no servidor; é o que `fetch` recebe |
| url | Yes | URL pública canônica do documento — a citação do ChatGPT depende dela |
| text | Yes | Conteúdo integral do documento, legível (Markdown) |
| title | Yes | Título legível do documento |
| metadata | No | Pares chave/valor adicionais sobre o documento (tipo, fonte, período…) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description nonetheless adds real behavioral context beyond them: the 'live GET against the public source when the document needs it' (a network fetch with possible lazy retrieval) and the error-on-unknown-id behavior. It does restate 'read-only and idempotent', which duplicates annotations, keeping it just under a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and return shape, then companion context, then behavior — a sensible ordering with little waste. Slightly dense and includes minor redundancy with the annotations ('read-only and idempotent').
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter fetch with full annotations and an output schema, the description covers purpose, provenance constraint, error behavior, and sibling routing. Nothing needed to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single `id` param is documented, so the baseline is 3. The description still adds value by clarifying provenance ('id obtained from `search`') and the failure mode for invalid ids, which the schema does not express.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Returns the full document') and resource (a document by id from `search`), and names the return shape { id, title, text, url, metadata }. It explicitly positions itself against `search` and the `ibge_*` siblings, so an agent can distinguish it without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit preconditions ('only ids returned by `search` are valid; an unknown id returns an error') and routes the agent away from the `ibge_*` tools for data queries. Both when-to-use and when-not-to-use are covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ibge_calendarioCalendário de divulgaçõesARead-onlyIdempotent
Queries IBGE release and collection calendar.
Features:
List upcoming survey releases
Filter by product (IPCA, PNAD, GDP, etc.)
Filter by period
Distinguish releases from field collections
Event types:
Release: Publication of survey results
Collection: Field research period
Examples:
Upcoming releases: (no parameters)
IPCA releases: produto="IPCA"
2024 calendar: de="01/01/2024", ate="31/12/2024"
Field collections: tipo="coleta"
Use a different tool when:
Already-published news and releases → ibge_noticias
Behavior: read-only and idempotent — a live GET against the public IBGE Calendário API. Returns a Markdown list.
| Name | Required | Description | Default |
|---|---|---|---|
| de | No | Data inicial no formato DD/MM/AAAA (ex: '01/01/2024') | |
| ate | No | Data final no formato DD/MM/AAAA (ex: '31/12/2024') | |
| tipo | No | Tipo de evento: 'divulgacao' (publicações), 'coleta' (pesquisas de campo), ou 'todos' | divulgacao |
| pagina | No | Número da página (padrão: 1) | |
| produto | No | Filtrar por produto/pesquisa (ex: 'IPCA', 'PNAD', 'PIB') | |
| quantidade | No | Quantidade de resultados por página (padrão: 20) |
Output Schema
| Name | Required | Description |
|---|---|---|
| total | Yes | Total de eventos disponíveis para os critérios |
| pagina | No | Página atual retornada |
| eventos | Yes | Lista de eventos do calendário (divulgações/coletas) |
| produto | No | Filtro de produto aplicado, quando informado |
| provenance | Yes | Bloco de proveniência (contrato v1.0): fonte, URL, período, extração e licença |
| attribution | Yes | URLs canônicas das fontes desta resposta (lista de atribuição) |
| totalPaginas | No | Total de páginas disponíveis |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so 'read-only and idempotent' is largely redundant. The description still adds genuine context beyond annotations: it is a live GET against the public IBGE Calendário API and returns a Markdown list.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded one-line purpose, then tightly grouped Features / Event types / Examples / Routing / Behavior sections. The 'read-only and idempotent' clause duplicates the annotations and could be trimmed, but overall structure is clean and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-param read-only query tool with a 100%-documented schema, full annotations and an output schema, the description supplies everything else an agent needs: purpose, event semantics, examples, sibling routing, and API/return context. Nothing material is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% (baseline 3), but the description adds meaning the schema does not: it interprets the 'tipo' enum values as Release vs Collection events and shows realistic values for produto (IPCA, PNAD, GDP) and de/ate formatting. The example block converts abstract params into concrete usage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Queries IBGE release and collection calendar') and immediately names the sibling it must not be confused with (ibge_noticias). An agent can distinguish this calendar-query tool from the news tool without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit 'Use a different tool when' routing to ibge_noticias for already-published content, plus four concrete invocation examples (no params, produto, date range, tipo). When-to-use and when-not-to-use are both covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ibge_censoCenso DemográficoARead-onlyIdempotent
Queries IBGE Demographic Census data (1970-2022).
Simplified tool to access census data without knowing SIDRA table codes.
Available years: 1970, 1980, 1991, 2000, 2010, 2022
Available themes:
populacao: Resident population
alfabetizacao: Literacy rate
domicilios: Housing characteristics
idade_sexo: Age pyramid
religiao: Religion distribution
cor_raca: Race/color
rendimento: Monthly income
educacao: Education level
trabalho: Employment
Examples:
Population 2022: ano="2022", tema="populacao"
Historical series: ano="todos", tema="populacao"
Literacy 2010 by state: ano="2010", tema="alfabetizacao", nivel_territorial="3"
List tables: tema="listar"
Statistics mode: for largest/smallest/mean/median/distribution/ranking questions over census data ("which municipality had the largest 2022 population?") use estatisticas=true — full distribution + top/bottom computed over ALL rows before truncation; agruparPor="" ranks groups by descending sum. In this mode campos/formato are ignored and registros comes empty.
Use a different tool when:
One municipality's current panel (estimate, HDI, GDP) → ibge_cidades
Current estimates or non-census indicator time series → ibge_indicadores
Comparing/ranking localities → ibge_comparar
An arbitrary SIDRA table → ibge_sidra
Behavior: read-only and idempotent — a live GET against the public IBGE SIDRA API. Returns Markdown plus a typed structuredContent payload.
| Name | Required | Description | Default |
|---|---|---|---|
| ano | No | Ano do censo (1970, 1980, 1991, 2000, 2010, 2022) ou 'todos' para série histórica | |
| tema | No | Tema dos dados: - populacao: População residente - alfabetizacao: Taxa de alfabetização - domicilios: Características dos domicílios - idade_sexo: Distribuição por sexo e idade para pirâmide etária - estrutura_etaria: Índice de envelhecimento, idade mediana e razão de sexo - religiao: Distribuição por religião - cor_raca: Cor ou raça - rendimento: Rendimento mensal - migracao: Migração - educacao: Nível de instrução - trabalho: Ocupação e trabalho - indigenas: População indígena - quilombolas: População quilombola - saneamento: Abastecimento de água; para esgotamento sanitário use ibge_datasaude(saneamento_esgoto) - deficiencia: Pessoas com deficiência - nupcialidade: Estado civil - fecundidade: Taxa de fecundidade - listar: Lista tabelas disponíveis | populacao |
| topN | No | Tamanho das listas top/bottom quando estatisticas=true sem agruparPor (padrão: 10, máx: 100) | |
| campos | No | Selecionar apenas algumas colunas por rótulo, separadas por vírgula (ex: 'Valor,Ano'). Reduz o volume da resposta. | |
| formato | No | Formato de saída | tabela |
| agruparPor | No | Com estatisticas=true, agrupa pela coluna informada (rótulo, ex: 'Unidade da Federação', 'Ano') e ranqueia os grupos por soma decrescente (grupos[0] = maior total), cada grupo com sua mini-distribuição. Nome curto ('UF', 'estado', 'cidade', 'região') e rótulo parcial ('Federação') são resolvidos, e a resposta diz em `aviso` por qual coluna agrupou; rótulo que casa com duas colunas é recusado em vez de escolhido | |
| localidades | No | Códigos das localidades ou 'all' | all |
| estatisticas | No | Computa estatísticas (mínimo/máximo/média/mediana/desvio-padrão/percentis) sobre TODOS os registros da consulta, antes da paginação, + ranking top/bottom. Use para 'qual o maior/menor', 'média', 'mediana', 'distribuição', 'ranking'. Quando true, ignora pagina, campos e formato | |
| nivel_territorial | No | Nível territorial (código N): 1=Brasil, 2=Região, 3=UF, 6=Município | 1 |
Output Schema
| Name | Required | Description |
|---|---|---|
| ano | No | Ano(s) de referência |
| tema | No | Tema do censo consultado |
| tabela | No | Tabela SIDRA de origem |
| colunas | Yes | Rótulos das colunas, na ordem |
| descricao | No | Descrição da tabela |
| registros | Yes | Registros: cada um mapeia rótulo da coluna -> valor |
| provenance | Yes | Bloco de proveniência (contrato v1.0): fonte, URL, período, extração e licença |
| attribution | Yes | URLs canônicas das fontes desta resposta (lista de atribuição) |
| estatisticas | No | Bloco estatístico presente quando estatisticas=true (registros vem vazio nesse modo) |
| totalRegistros | Yes | Total de registros de dados |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/idempotentHint/openWorldHint, so the safety profile is covered. The description still adds real context beyond them: it is a live GET against the public IBGE SIDRA API, it returns Markdown plus a typed structuredContent payload, and statistics mode has side effects (campos/formato ignored, registros left empty) that an agent must anticipate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose, then organized under labeled sections (years, themes, examples, statistics mode, alternatives, behavior). It is long and repeats the year list already present in the enum, but the structure keeps it scannable and each section carries distinct routing or mode information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with an output schema and full annotation coverage, the definition supplies the routing, mode semantics, and example calls an agent needs. The only shortfall is that the in-description theme list is a subset of the schema enum, so an agent skimming only the description could under-estimate available themes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so the baseline is 3, but the description goes further with concrete call examples that show parameter interaction (ano="todos" for a historical series, nivel_territorial="3" for literacy by state, tema="listar" to enumerate tables) and explains that agruparPor ranks groups by descending sum. The description's theme list covers only 9 of the 18 enum values, leaving the rest to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Queries IBGE Demographic Census data 1970-2022') plus the scoping constraint that it works without knowing SIDRA table codes, which is exactly the distinction from ibge_sidra. An agent can tell it apart from siblings without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Contains an explicit 'Use a different tool when' block that routes to four alternatives (ibge_cidades, ibge_indicadores, ibge_comparar, ibge_sidra) with the condition selecting each. It also names the exact question shape ('which municipality had the largest 2022 population?') that triggers estatisticas=true.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ibge_cidadesPanorama municipal (Cidades@)ARead-onlyIdempotent
Queries municipal indicators from IBGE (similar to Cidades@ portal).
Features:
General overview of a municipality (population, HDI, GDP, etc.)
Query specific indicators
Historical indicator data over years
List available surveys and indicators
Available indicators: populacao, area, densidade, pib_per_capita, idh, escolarizacao, mortalidade, salario_medio, receitas, despesas
Examples:
São Paulo overview: tipo="panorama", municipio="3550308"
Population history: tipo="historico", municipio="3550308", indicador="populacao"
View surveys: tipo="pesquisas"
Available indicators: tipo="indicador"
This tool is the panel for a SINGLE municipality (Cidades@). Use a different tool when:
Census themes / historical series → ibge_censo
Comparing multiple municipalities → ibge_comparar
A macro indicator time series → ibge_indicadores
Behavior: read-only and idempotent — uses the public IBGE Pesquisas (Cidades@) API and Localidades lookup; panorama may combine several indicator calls and explicitly reports partial upstream gaps. Returns Markdown plus a typed structuredContent payload.
| Name | Required | Description | Default |
|---|---|---|---|
| uf | No | Código ou sigla da UF para filtrar (ex: 35 ou SP) | |
| tipo | No | Tipo de consulta: panorama (resumo geral), indicador (específico), pesquisas (listar), historico | panorama |
| pesquisa | No | ID da pesquisa para filtrar indicadores | |
| indicador | No | ID do indicador ou nome para busca | |
| municipio | No | Código IBGE do município (7 dígitos) |
Output Schema
| Name | Required | Description |
|---|---|---|
| nome | No | Nome do município/indicador |
| tipo | Yes | Tipo de consulta (panorama, indicador, pesquisas, historico) |
| avisos | No | Avisos de indisponibilidade parcial ou limitações da resposta |
| municipio | No | Código IBGE do município |
| provenance | Yes | Bloco de proveniência (contrato v1.0): fonte, URL, período, extração e licença |
| attribution | Yes | URLs canônicas das fontes desta resposta (lista de atribuição) |
| indicadores | Yes | Indicadores retornados (vazio para respostas de catálogo) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, so the safety profile is covered. The description adds real context beyond that: it names the upstream API (public IBGE Pesquisas/Cidades@ plus Localidades lookup), warns that panorama may fan out into several indicator calls and report partial upstream gaps, and states the dual Markdown + structuredContent return. Useful, though it stops short of detailing rate limits or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then clearly sectioned into features, indicator values, examples, sibling routing, and behavior. Slightly long, but every block carries routing or parameter information and nothing is repeated filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return format need not be explained in depth, and the description still notes Markdown plus typed structuredContent. With full schema coverage, explicit disambiguation from all relevant siblings, and documented behavior, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, so the baseline is 3, but the description goes further: it enumerates valid indicator identifiers (populacao, area, idh, etc.) and the example block shows how tipo/municipio/indicador combine for each query mode, which the schema's short field descriptions do not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (queries) and resource (municipal indicators from IBGE/Cidades@) and immediately scopes it to a SINGLE municipality. It explicitly names the four sibling tools it is not (ibge_censo, ibge_comparar, ibge_indicadores), so an agent can route without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides concrete call examples per tipo value and an explicit 'use a different tool when' block mapping each alternative scenario to the correct sibling. When-to-use and when-not-to-use are both spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ibge_cnaeClassificação CNAEARead-onlyIdempotent
Queries CNAE (National Classification of Economic Activities) from IBGE.
CNAE is the official classification for economic activities in Brazil.
Hierarchical structure:
Section (letter A-U): 21 main categories
Division (2 digits): 87 divisions
Group (3 digits): 285 groups
Class (4-5 digits): 673 classes
Subclass (7 digits): 1,332 subclasses
Features:
Search by CNAE code
Search by activity description
List by hierarchical level
Show complete hierarchy
Examples:
Search software: busca="software"
Specific code: codigo="6201-5/01"
View section: codigo="J"
List divisions: nivel="divisoes"
Behavior: read-only and idempotent — a live GET against the public IBGE CNAE API. Returns Markdown.
| Name | Required | Description | Default |
|---|---|---|---|
| busca | No | Termo para buscar na descrição das atividades (ex: 'software', 'restaurante', 'comércio') | |
| nivel | No | Nível hierárquico para listar (padrão: mostra todos os níveis relevantes) | |
| codigo | No | Código CNAE para buscar (seção, divisão, grupo, classe ou subclasse). Exemplos: - Seção: "A" (agricultura) - Divisão: "01" (agricultura e pecuária) - Grupo: "01.1" (produção de lavouras) - Classe: "01.11" (cultivo de cereais) - Subclasse: "0111-3/01" (cultivo de arroz) | |
| limite | No | Número máximo de resultados (padrão: 20) |
Output Schema
| Name | Required | Description |
|---|---|---|
| modo | Yes | Modo de resposta que gerou os dados |
| busca | No | Presente no modo de busca por termo |
| lista | No | Presente no modo de listagem por nível |
| codigo | No | Presente no modo de consulta por código |
| provenance | Yes | Bloco de proveniência (contrato v1.0): fonte, URL, período, extração e licença |
| attribution | Yes | URLs canônicas das fontes desta resposta (lista de atribuição) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, non-destructive, open-world behavior, so the bar is lower. The description still adds real context by specifying it is 'a live GET against the public IBGE CNAE API' and 'Returns Markdown,' which goes beyond the annotations and informs the agent about the external source and output medium.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and uses clear headers for hierarchy, features, examples, and behavior, making it scannable. It is slightly verbose with the level-by-level counts (21/87/285/673/1,332) that are informative but not strictly necessary for invocation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With annotations covering the safety profile, an output schema present, and 100% schema description coverage, the description need not explain return values. It rounds out the picture with hierarchy structure, feature list, worked examples, and source/format behavior, leaving nothing essential for a correct call unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the codigo field already documents format and examples in the schema, so the schema carries the semantic load. The description's examples for busca, codigo, and nivel add marginal reinforcement, but it never addresses the limite parameter; baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Queries) and resource (CNAE - National Classification of Economic Activities from IBGE), and clarifies the domain as Brazilian economic-activity classification. This cleanly distinguishes it from domain-adjacent siblings like ibge_cidades, ibge_estados, and ibge_pesquisas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'Features' and 'Examples' sections give concrete invocation scenarios (search by term, by code, list by level, show hierarchy), which effectively communicates when to use the tool. It lacks explicit when-not guidance or named alternatives among the many IBGE siblings, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ibge_compararComparação entre localidadesARead-onlyIdempotent
Compares data between localities (municipalities or states).
Available indicators:
populacao: Current population estimate
populacao_censo: Census 2022 population
pib: GDP per capita
area: Territorial area (km²)
densidade: Population density (inhab/km²)
alfabetizacao: Literacy rate
domicilios: Number of households
Features:
Compare up to 10 localities at once
Calculate statistics (max, min, average, variation)
Generate ranked output
Accept municipality codes (7 digits) or state codes (2 digits)
Examples:
Compare capitals: localidades="3550308,3304557,4106902", indicador="populacao"
Compare states: localidades="35,33,41", indicador="pib"
Area ranking: localidades="3550308,3304557", formato="ranking"
List indicators: indicador="listar"
Use a different tool when:
You want a time series or one known indicator rather than comparing a fixed set of localities → ibge_indicadores
You need census themes or historical census data → ibge_censo
You know the exact SIDRA table and need arbitrary dimensions → ibge_sidra
Use this tool ONLY to rank/compare 2–10 localities on one indicator. For a single locality, use ibge_cidades (municipal panel), ibge_censo, or ibge_sidra.
Behavior: read-only and idempotent — a live GET against the public IBGE APIs (SIDRA and Localidades). Returns Markdown plus a typed structuredContent payload.
| Name | Required | Description | Default |
|---|---|---|---|
| formato | No | Formato de saída: tabela, json ou ranking (ordenado) | tabela |
| indicador | No | Indicador para comparação: - populacao: Estimativa populacional atual - populacao_censo: População do Censo 2022 - pib: PIB a preços correntes (Mil Reais) - area: Área territorial (km²) - densidade: Densidade demográfica (hab/km²) - alfabetizacao: Taxa de alfabetização - domicilios: Número de domicílios - listar: Lista indicadores disponíveis | populacao |
| localidades | Yes | Códigos IBGE das localidades separados por vírgula (ex: "3550308,3304557,4106902"). Use 7 dígitos para municípios, 2 dígitos para UFs. |
Output Schema
| Name | Required | Description |
|---|---|---|
| nome | No | Nome do indicador |
| tabela | No | Tabela SIDRA de origem |
| formato | No | Formato solicitado |
| indicador | No | Indicador comparado |
| provenance | Yes | Bloco de proveniência (contrato v1.0): fonte, URL, período, extração e licença |
| attribution | Yes | URLs canônicas das fontes desta resposta (lista de atribuição) |
| localidades | Yes | Localidades comparadas, com o valor do indicador |
| estatisticas | No | Estatísticas agregadas (quando há ao menos 2 valores positivos) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, openWorld and non-destructive. The description adds behavioral context beyond them by disclosing it performs a live GET against the public IBGE APIs (SIDRA and Localidades) and returns Markdown plus a typed structuredContent payload. It does not cover error/rate-limit behavior, but with safety fully annotated this is solid added value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded, then useful sections (indicators, features, examples, exclusions, behavior) in a scannable structure. It runs long and duplicates the schema's indicator definitions, but each sentence is on-topic and aids selection.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With annotations covering the safety profile, 100% schema coverage, an output schema handling return values, and explicit sibling routing, the definition supplies everything an agent needs to select and invoke it correctly. No material gap remains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds a real constraint absent from the schema ('Compare up to 10 localities at once') and concrete code-format examples (7-digit municipalities, 2-digit states) that clarify how to populate 'localidades'. The indicator list largely duplicates the schema enum.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Compares data between localities (municipalities or states)') and immediately names the scope of the operation. It differentiates itself from siblings by explicitly stating it compares a fixed set of localities rather than producing time series or single-locality panels.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an explicit 'Use a different tool when:' block naming three alternatives (ibge_indicadores, ibge_censo, ibge_sidra) with the conditions that select each, plus a 'use ONLY to rank/compare 2–10 localities' constraint and single-locality fallbacks (ibge_cidades, ibge_censo, ibge_sidra). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ibge_datasaudeIndicadores de saúdeARead-onlyIdempotent
Queries Brazil health indicators, served through IBGE's SIDRA (some originally produced by DataSUS, e.g. mortality and births).
Mortality and Birth:
mortalidade_infantil: Infant mortality rate
nascidos_vivos: Live births by location
obitos: Deaths by residence
Demographic Indicators:
esperanca_vida: Life expectancy at birth
fecundidade: Fertility rate
Sanitation:
saneamento_agua: Water supply
saneamento_esgoto: Sewage system
Health Coverage:
plano_saude: Health insurance coverage
autoavaliacao_saude: Self-rated health status
Territorial levels: 1=Brazil, 2=Region, 3=State, 6=Municipality
Use this tool for health, mortality, fertility, sanitation and health-coverage indicators. Use ibge_indicadores for general economic/social series such as GDP, prices, labor and population estimates.
Examples:
Infant mortality: indicador="mortalidade_infantil"
Life expectancy by state: indicador="esperanca_vida", nivel_territorial="3"
Deaths in SP: indicador="obitos", nivel_territorial="3", localidade="35"
List indicators: indicador="listar"
Statistics mode: for largest/smallest/mean/median/distribution/ranking questions ("which state has the highest infant mortality?", "median life expectancy across states") use estatisticas=true — full distribution + top/bottom over ALL rows before truncation; agruparPor="" ranks groups by descending sum. In this mode campos/formato are ignored and registros comes empty.
Use a different tool when:
A single municipality's general panel (which also includes infant mortality) → ibge_cidades
Population/demographic counts (not health-specific) → ibge_censo or ibge_sidra
Behavior: read-only and idempotent — a live GET against the public IBGE SIDRA API. Returns Markdown plus a typed structuredContent payload.
| Name | Required | Description | Default |
|---|---|---|---|
| topN | No | Tamanho das listas top/bottom quando estatisticas=true sem agruparPor (padrão: 10, máx: 100) | |
| campos | No | Selecionar apenas algumas colunas por rótulo, separadas por vírgula (ex: 'Valor,Ano'). Reduz o volume da resposta. | |
| formato | No | Formato de saída | tabela |
| periodo | No | Período: 'last', 'all', ou ano específico | last |
| indicador | Yes | Indicador de saúde. Disponíveis: - mortalidade_infantil: Taxa de mortalidade infantil - esperanca_vida: Esperança de vida ao nascer - nascidos_vivos: Nascidos vivos - obitos: Óbitos por local de residência - fecundidade: Taxa de fecundidade - saneamento_agua: Abastecimento de água - saneamento_esgoto: Esgotamento sanitário - plano_saude: Cobertura de plano de saúde - autoavaliacao_saude: Autoavaliação de saúde boa ou muito boa - listar: Lista indicadores disponíveis | |
| agruparPor | No | Com estatisticas=true, agrupa pela coluna informada (rótulo, ex: 'Unidade da Federação', 'Ano') e ranqueia os grupos por soma decrescente (grupos[0] = maior total), cada grupo com sua mini-distribuição. Nome curto ('UF', 'estado', 'cidade', 'região') e rótulo parcial ('Federação') são resolvidos, e a resposta diz em `aviso` por qual coluna agrupou; rótulo que casa com duas colunas é recusado em vez de escolhido | |
| localidade | No | Código da localidade ou 'all' | all |
| estatisticas | No | Computa estatísticas (mínimo/máximo/média/mediana/desvio-padrão/percentis) sobre TODOS os registros da consulta, antes da paginação, + ranking top/bottom. Use para 'qual o maior/menor', 'média', 'mediana', 'distribuição', 'ranking'. Quando true, ignora pagina, campos e formato | |
| nivel_territorial | No | Nível territorial (código N): 1=Brasil, 2=Região, 3=UF, 6=Município | 1 |
Output Schema
| Name | Required | Description |
|---|---|---|
| nome | No | Nome do indicador |
| fonte | No | Fonte do dado |
| colunas | Yes | Rótulos das colunas, na ordem |
| indicador | No | Chave do indicador de saúde consultado |
| registros | Yes | Registros: cada um mapeia rótulo da coluna -> valor |
| provenance | Yes | Bloco de proveniência (contrato v1.0): fonte, URL, período, extração e licença |
| attribution | Yes | URLs canônicas das fontes desta resposta (lista de atribuição) |
| estatisticas | No | Bloco estatístico presente quando estatisticas=true (registros vem vazio nesse modo) |
| totalRegistros | Yes | Total de registros de dados |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, open-world and non-destructive, so the safety profile is covered; the description still adds that it is a live GET against the public SIDRA API and returns Markdown plus structuredContent. The statistics-mode side effects (campos/formato ignored, registros empty, ranking computed over ALL rows before truncation) are genuine behavioral disclosure. It stops short of auth/rate-limit or error behavior, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well front-loaded and sectioned, but it duplicates content already present at full fidelity in the schema: the nine-indicator list with Portuguese glosses, and the territorial-level legend (1=Brazil ... 6=Municipality). Roughly a third of the text is restatement rather than added value; the material that does earn its place is the sibling routing and the examples.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter open-world query tool with an output schema, annotations and 100% schema coverage, the description supplies everything else an agent needs: indicator taxonomy, mode semantics, territorial levels, sibling disambiguation and concrete examples. Return-format explanation is a bonus since the output schema already covers it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds operational meaning the schema lacks: examples tie localidade to a territorial level ('Deaths in SP: indicador="obitos", nivel_territorial="3", localidade="35"') and explain the statistics/ranking mode's intent. It does not, however, add format details for periodo or campos beyond what the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb and resource ('Queries Brazil health indicators, served through IBGE's SIDRA') and enumerates the indicator families it covers. It explicitly separates itself from ibge_indicadores ('general economic/social series such as GDP, prices, labor and population estimates'), so an agent can route correctly without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('Use this tool for health, mortality, fertility, sanitation and health-coverage indicators') and a dedicated 'Use a different tool when' block naming ibge_cidades, ibge_censo and ibge_sidra with the exact condition that selects each. Worked examples further pin down invocation for three distinct query shapes.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ibge_estadosEstados do BrasilARead-onlyIdempotent
Lists all Brazilian states from IBGE.
Features:
Lists all 27 states (26 states + Federal District)
Filter by region (North, Northeast, Southeast, South, Central-West)
Sort by ID, name, or abbreviation
Examples:
List all states: (no parameters)
Northeast states: regiao="NE"
Sorted by abbreviation: ordenar="sigla"
Use a different tool when:
Municipalities of a state → ibge_municipios
Details/hierarchy of one locality by code → ibge_localidade
Behavior: read-only and idempotent — a live GET against the public IBGE Localidades API. Returns a Markdown table.
| Name | Required | Description | Default |
|---|---|---|---|
| regiao | No | Filtrar por região: N (Norte), NE (Nordeste), SE (Sudeste), S (Sul), CO (Centro-Oeste) | |
| ordenar | No | Campo para ordenação dos resultados | nome |
Output Schema
| Name | Required | Description |
|---|---|---|
| total | Yes | Total de estados retornados |
| estados | Yes | Lista de estados |
| provenance | Yes | Bloco de proveniência (contrato v1.0): fonte, URL, período, extração e licença |
| attribution | Yes | URLs canônicas das fontes desta resposta (lista de atribuição) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds genuinely new context — it is a live GET against the public IBGE Localidades API returning a Markdown table — though the return format is partly redundant given an output schema exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the one-line purpose, then uses labeled Features/Examples/routing sections that are scannable and waste-free. Every sentence either defines scope, demonstrates a call, or routes to a sibling.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With annotations covering safety, a full input schema with enums, and an output schema covering return values, the description needs only purpose, examples and routing — all of which are present. Nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and both params are enum-constrained with self-documenting values, so the baseline is 3. The description goes beyond the schema by showing the actual wire values in context (regiao="NE", ordenar="sigla"), which reduces the chance of an agent guessing at an invalid code.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource ('Lists all Brazilian states from IBGE') and immediately scopes it with the exact count (27) and the two filter/sort axes. The 'Use a different tool when' section names the sibling tools it is not, so an agent can separate it from ibge_municipios and ibge_localidade without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the alternatives (ibge_municipios for municipalities, ibge_localidade for single-locality details) and the condition that selects each, plus concrete invocation examples for no-param, region-filtered, and sorted calls. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ibge_geocodigoCódigos geográficos do IBGEARead-onlyIdempotent
Decodes IBGE codes or searches codes by locality name.
Features:
Decode region, state, municipality, or district codes
Search IBGE code by name
Show complete geographic hierarchy
Return related codes
Code structure:
1 digit: Region (1=North, 2=Northeast, 3=Southeast, 4=South, 5=Central-West)
2 digits: State (11-53)
7 digits: Municipality
9 digits: District
Examples:
Decode municipality: codigo="3550308"
Decode state: codigo="35"
Search by name: nome="São Paulo"
Municipality in state: nome="Campinas", uf="SP"
This tool decodes a code's structure and resolves name→code at any level. Use a different tool when:
You only need to list/search municipalities → ibge_municipios
You want the full detailed record of one locality → ibge_localidade
Behavior: read-only and idempotent — a live GET against the public IBGE Localidades API. Returns Markdown.
| Name | Required | Description | Default |
|---|---|---|---|
| uf | No | Estado por sigla (SP), nome (São Paulo) ou código IBGE (35) para restringir a busca por nome de município | |
| nome | No | Nome da localidade para encontrar o código IBGE (estado ou município) | |
| codigo | No | Código IBGE para decodificar. Formatos aceitos: - 1 dígito: Região (1-5) - 2 dígitos: UF (11-53) - 7 dígitos: Município - 9 dígitos: Distrito |
Output Schema
| Name | Required | Description |
|---|---|---|
| nome | No | Nome da localidade resolvida |
| tipo | Yes | Tipo do resultado: localidade decodificada (regiao/uf/municipio/distrito) ou lista de municípios encontrados (lista) |
| sigla | No | Sigla da região ou UF, quando aplicável |
| total | No | Quantidade de municípios encontrados na busca por nome (apenas tipo lista) |
| codigo | No | Código IBGE da localidade resolvida (ausente em resultados do tipo lista) |
| regiao | No | Nome da região à qual a UF pertence (apenas tipo uf) |
| estados | No | Estados pertencentes à região (apenas tipo regiao) |
| matches | No | Municípios encontrados na busca por nome (apenas tipo lista) |
| hierarquia | No | Hierarquia geográfica completa, da região ao município/distrito (tipo municipio/distrito) |
| provenance | Yes | Bloco de proveniência (contrato v1.0): fonte, URL, período, extração e licença |
| attribution | Yes | URLs canônicas das fontes desta resposta (lista de atribuição) |
| codigoSidra | No | Código SIDRA de 6 dígitos do município (apenas tipo municipio) |
| regiaoCodigo | No | Código IBGE da região à qual a UF pertence (apenas tipo uf) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly, idempotent, openWorld and non-destructive. The description adds genuinely new context beyond them: it is a live GET against the public IBGE Localidades API and returns Markdown. It does not discuss rate limits or error behavior, so it stops short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then cleanly sectioned into features, code structure, examples and alternatives, which suits the tool's branching input modes. The code-structure bullets partially duplicate the parameter schema, costing some efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained. Between the level-decoding rules, the four examples, the named alternatives and the read-only/API/Markdown behavior note, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, so the baseline is 3, but the description adds meaning the schema lacks — the digit-length-to-level mapping (1/2/7/9 digits) and worked examples such as nome="Campinas", uf="SP" that show how nome and uf combine. It stops short of documenting mutual exclusivity between codigo and nome.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific dual verb (decodes / searches) against a specific resource (IBGE geographic codes), and the 'Use a different tool when' block explicitly names ibge_municipios and ibge_localidade, so an agent can distinguish it from siblings without reading schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing is provided: it says when to use this tool (decode a code at any level, resolve name→code) and when not to, naming two concrete alternatives plus the condition that selects each. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ibge_indicadoresIndicadores econômicos e sociaisARead-onlyIdempotent
Queries IBGE economic and social indicators.
Available indicators:
Economic:
pib: GDP at current prices
pib_variacao: GDP variation (%)
pib_per_capita: GDP per capita
industria: Industrial production
comercio: Retail sales
servicos: Services volume
Prices:
ipca: Monthly IPCA
ipca_acumulado: 12-month IPCA
inpc: Monthly INPC
Labor:
desemprego: Unemployment rate
ocupacao: Employed people
rendimento: Average income
informalidade: Informality rate
Population:
populacao: Population estimate
densidade: Population density
Examples:
GDP: indicador="pib"
IPCA last 12 months: indicador="ipca", periodos="last 12"
Unemployment by state: indicador="desemprego", nivel_territorial="3"
List indicators: indicador="listar"
Statistics mode: for largest/smallest/mean/median/distribution/ranking questions ("which state has the highest unemployment?", "median GDP per capita across states") use estatisticas=true — full distribution + top/bottom over ALL rows before truncation; agruparPor="" (e.g. "Unidade da Federação", "Trimestre") ranks groups by descending sum. In this mode campos/formato are ignored and registros comes empty.
Use a different tool when:
Comparing/ranking localities → ibge_comparar
Census themes → ibge_censo
One municipality's panel → ibge_cidades
You know the exact SIDRA table or need arbitrary variables/classifications → ibge_sidra
Behavior: read-only and idempotent — a live GET against the public IBGE SIDRA API. Returns Markdown plus a typed structuredContent payload.
| Name | Required | Description | Default |
|---|---|---|---|
| topN | No | Tamanho das listas top/bottom quando estatisticas=true sem agruparPor (padrão: 10, máx: 100) | |
| campos | No | Selecionar apenas algumas colunas por rótulo, separadas por vírgula (ex: 'Valor,Ano'). Reduz o volume da resposta. | |
| formato | No | Formato de saída | tabela |
| periodos | No | Períodos (ex: '2023', 'last', 'last 4') | last |
| categoria | No | Filtrar por categoria de indicadores | |
| indicador | No | Nome do indicador (ex: "pib", "ipca", "desemprego", "populacao"). Use "listar" para ver todos os indicadores disponíveis. | |
| agruparPor | No | Com estatisticas=true, agrupa pela coluna informada (rótulo, ex: 'Unidade da Federação', 'Ano') e ranqueia os grupos por soma decrescente (grupos[0] = maior total), cada grupo com sua mini-distribuição. Nome curto ('UF', 'estado', 'cidade', 'região') e rótulo parcial ('Federação') são resolvidos, e a resposta diz em `aviso` por qual coluna agrupou; rótulo que casa com duas colunas é recusado em vez de escolhido | |
| localidades | No | Códigos das localidades ou 'all' | all |
| estatisticas | No | Computa estatísticas (mínimo/máximo/média/mediana/desvio-padrão/percentis) sobre TODOS os registros da consulta, antes da paginação, + ranking top/bottom. Use para 'qual o maior/menor', 'média', 'mediana', 'distribuição', 'ranking'. Quando true, ignora pagina, campos e formato | |
| nivel_territorial | No | Nível territorial (código N): 1=Brasil, 2=Região, 3=UF, 6=Município, 7=Região Metropolitana, 8=Mesorregião, 9=Microrregião, 14=RIDE | 1 |
Output Schema
| Name | Required | Description |
|---|---|---|
| nome | No | Nome do indicador |
| tabela | No | Tabela SIDRA de origem |
| colunas | Yes | Rótulos das colunas, na ordem |
| indicador | No | Chave do indicador consultado |
| registros | Yes | Registros: cada um mapeia rótulo da coluna -> valor |
| provenance | Yes | Bloco de proveniência (contrato v1.0): fonte, URL, período, extração e licença |
| attribution | Yes | URLs canônicas das fontes desta resposta (lista de atribuição) |
| estatisticas | No | Bloco estatístico presente quando estatisticas=true (registros vem vazio nesse modo) |
| totalRegistros | Yes | Total de registros de dados |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so safety is covered. The description still adds value: it discloses the live GET against the public SIDRA API, the Markdown + structuredContent return shape, and the surprising mode side effect that campos/formato are ignored and registros comes back empty when estatisticas=true.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Long but well-structured with category headings, bullets and examples, and the core purpose is front-loaded. Every section (indicator list, examples, statistics mode, sibling routing, behavior) earns its place for a 10-parameter tool, though it is denser than strictly necessary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers indicator vocabulary, examples, the non-obvious statistics mode, and sibling routing, and an output schema exists so return values need not be detailed. Nothing an agent needs to invoke this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description goes further by explaining statistics-mode interactions (ignores pagina/campos/formato), agruparPor group-ranking and label-resolution behavior, and by supplying concrete example values for indicador, periodos and nivel_territorial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Queries') and resource ('IBGE economic and social indicators'), then enumerates the exact indicators available grouped by category. It explicitly differentiates from siblings (ibge_sidra, ibge_comparar, ibge_censo, ibge_cidades) so an agent can route without opening other schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Contains an explicit 'Use a different tool when:' block naming each alternative and its selecting condition, plus worked examples mapping parameters to questions. The statistics-mode guidance states exactly when to set estatisticas=true.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ibge_localidadeDetalhes de localidadeARead-onlyIdempotent
Returns details of a specific locality by IBGE code.
Features:
State information (2-digit code)
Municipality information (7-digit code)
District information (9-digit code)
Complete hierarchy (region, mesoregion, microregion)
Examples:
São Paulo state: codigo=35
São Paulo city: codigo=3550308
District: codigo=355030805
This tool returns the full record of ONE locality you already have the code for. Use a different tool when:
You have a name and need the code → ibge_municipios (municipalities) or ibge_geocodigo (any level)
You want to decompose/understand a code's structure → ibge_geocodigo
Behavior: read-only and idempotent — a live GET against the public IBGE Localidades API. Returns a Markdown record.
| Name | Required | Description | Default |
|---|---|---|---|
| tipo | No | Tipo da localidade. Se não informado, será inferido pelo tamanho do código. | |
| codigo | Yes | Código IBGE da localidade (estado: 2 dígitos, município: 7 dígitos, distrito: 9 dígitos) |
Output Schema
| Name | Required | Description |
|---|---|---|
| id | Yes | Código IBGE da localidade |
| nome | Yes | Nome da localidade |
| tipo | Yes | Tipo da localidade retornada |
| sigla | No | Sigla da UF (apenas para estados) |
| estado | No | Estado da localidade (município ou distrito) |
| regiao | No | Região do estado (apenas para estados) |
| municipio | No | Município ao qual o distrito pertence (apenas para distritos) |
| provenance | Yes | Bloco de proveniência (contrato v1.0): fonte, URL, período, extração e licença |
| attribution | Yes | URLs canônicas das fontes desta resposta (lista de atribuição) |
| mesorregiao | No | Mesorregião do município |
| microrregiao | No | Microrregião do município |
| regiaoImediata | No | Região imediata do município |
| regiaoIntermediaria | No | Região intermediária do município |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint, so the description does not need to restate safety. It adds genuine value beyond them by naming the backing source ('live GET against the public IBGE Localidades API') and the return format ('Markdown record'). It does not mention any rate limits or auth requirements, which are minor for a public API.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose, then organizes the remainder into labeled Features, Examples, When-to-use, and Behavior blocks. Every section is short and actionable; the concrete code examples partially duplicate the schema's digit-length note, but they remain useful for disambiguation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be enumerated; the description still notes the Markdown format. Combined with the hierarchy explanation, the code examples, and the sibling routing rules, an agent has everything needed to call this correctly for a three-level locality lookup.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already documents both 'codigo' digit lengths and the 'tipo' enum with inference from code length, so the baseline is 3. The description's concrete examples (São Paulo state=35, city=3550308, district=355030805) add grounded meaning by mapping real entities onto the abstract code lengths, marginally exceeding the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Returns details of a specific locality by IBGE code') and immediately clarifies scope with the exact code lengths for state, municipality, and district levels. It explicitly distinguishes itself from siblings by name ('Use a different tool when... ibge_municipios or ibge_geocodigo'), so an agent can route correctly without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an explicit selection condition ('returns the full record of ONE locality you already have the code for') and pairs it with two named alternatives for the opposite cases (having a name, wanting to decompose a code). This is textbook when-to-use / when-not-to-use guidance with named siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ibge_malhasMalhas geográficasARead-onlyIdempotent
Gets geographic meshes (maps) from IBGE in GeoJSON, TopoJSON, or SVG format.
Features:
Meshes for Brazil, regions, states, municipalities
Different resolution levels (internal divisions)
Different quality levels
Formats: GeoJSON (data), TopoJSON (compact), SVG (image)
Locality types:
"BR" or "1" = Entire Brazil
State abbreviation (e.g., "SP", "RJ")
State code (e.g., "35" for SP)
Municipality code (7 digits)
Resolution (internal divisions):
0 = Outline only
2 = States
5 = Municipalities
Examples:
Brazil with states: localidade="BR", resolucao="2"
São Paulo with municipalities: localidade="SP", resolucao="5"
SVG format: localidade="BR", formato="svg"
Use a different tool when:
Thematic meshes (biomes, Legal Amazon, semi-arid, metropolitan regions) → ibge_malhas_tema
Behavior: read-only and idempotent — a live GET against the public IBGE Malhas API. Returns the mesh in the requested format (GeoJSON, TopoJSON, or SVG).
| Name | Required | Description | Default |
|---|---|---|---|
| tipo | No | Tipo de divisão territorial | |
| formato | No | Formato de saída (padrão: geojson) | geojson |
| qualidade | No | Qualidade do traçado: 'minima', 'intermediaria' ou 'maxima' (padrão). Os números 1–4 do IBGE antigo continuam aceitos e são traduzidos. | maxima |
| resolucao | No | Divisões internas a desenhar dentro da malha pedida: 0 = Sem divisões internas (só o contorno) 1 = Macrorregiões (apenas quando localidade=BR) 2 = Unidades da Federação (BR ou uma região) 3 = Mesorregiões 4 = Microrregiões 5 = Municípios Cada nível aceita só as divisões menores que ele: município aceita nenhuma, UF aceita 3, 4 e 5. | 0 |
| localidade | Yes | Código IBGE ou sigla da localidade (ex: 'BR', 'SP', '35', '3550308') | |
| intrarregiao | No | Divisão interna pelo nome, alternativa a resolucao: 'regiao', 'UF', 'regiao-intermediaria', 'regiao-imediata', 'mesorregiao', 'microrregiao' ou 'municipio'. Quando informado, prevalece sobre resolucao. |
Output Schema
| Name | Required | Description |
|---|---|---|
| url | No | URL para download da malha completa |
| tipo | No | Tipo de divisão territorial, quando informado |
| formato | Yes | Formato de saída solicitado (geojson, topojson ou svg) |
| qualidade | No | Qualidade do traçado solicitada |
| resolucao | No | Resolução/divisões internas solicitada |
| localidade | Yes | Código IBGE ou sigla da localidade consultada |
| provenance | Yes | Bloco de proveniência (contrato v1.0): fonte, URL, período, extração e licença |
| attribution | Yes | URLs canônicas das fontes desta resposta (lista de atribuição) |
| intrarregiao | No | Divisão interna desenhada dentro da malha (vocabulário da API v3) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered; the description's 'read-only and idempotent' sentence largely repeats them. What it does add is the nature of the call — a live GET against the public IBGE Malhas API — which tells the agent data is fetched remotely rather than cached locally.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well front-loaded: purpose first, then features, locality types, resolution, examples, and sibling routing, all in scannable bullets. It loses a point because the resolution and format enumerations restate schema fields that already carry the same information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description need not explain return shapes, and annotations carry the safety profile. Given 6 parameters, one required, the description supplies locality coding, resolution levels, format choices, precedence-free guidance, and sibling routing — nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3 and the schema already documents resolucao, qualidade, formato and intrarregiao. The description goes further by giving worked mappings ('BR' or '1' = Brazil, 'SP'/'35' = São Paulo, 7-digit municipality code) and example parameter combinations, which reduce ambiguity when constructing a call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Gets geographic meshes (maps) from IBGE') and enumerates the scope it covers (Brazil, regions, states, municipalities; GeoJSON/TopoJSON/SVG). It also names the sibling it is not — ibge_malhas_tema for thematic meshes — so an agent can separate the two without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Contains an explicit 'Use a different tool when' block that routes thematic meshes (biomes, Legal Amazon, semi-arid, metropolitan regions) to ibge_malhas_tema, plus three concrete invocation examples. The condition that selects the alternative and the condition that selects this tool are both stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ibge_malhas_temaMalhas temáticasARead-onlyIdempotent
Lists what a THEMATIC territorial recorte of Brazil contains: how many records, with which codes and names, from versioned official IBGE/GeoFTP snapshots.
Available recortes:
biomas: the six continental biomes
amazonia_legal: Legal Amazon boundary
semiarido: semi-arid area
costeiro: coastal municipalities
fronteira: border-strip municipalities
metropolitana: metropolitan regions
ride: Integrated Development Regions
listar: the catalogue itself, without querying the source
Filtering with codigo: biomas uses the biome code; every municipality-composition recorte (amazonia_legal, semiarido, costeiro, fronteira, metropolitana and ride) accepts the 7-digit IBGE municipality code. Ask without codigo to list records.
THEMATIC GEOMETRY IS NOT PART OF THIS TOOL'S CONTRACT. Use ibge_malhas for supported administrative geometry.
Use a different tool when:
Administrative meshes WITH geometry (country/region/state/municipality outlines) → ibge_malhas
Behavior: read-only and idempotent. Attributes/listings come exclusively from versioned official IBGE/GeoFTP snapshots audited by the laboratory. Returns Markdown plus a typed structuredContent payload.
| Name | Required | Description | Default |
|---|---|---|---|
| tema | Yes | Recorte temático do território: - biomas: os seis biomas continentais - amazonia_legal: limite da Amazônia Legal - semiarido: área do semiárido - costeiro: municípios da zona costeira - fronteira: municípios da faixa de fronteira - metropolitana: regiões metropolitanas - ride: Regiões Integradas de Desenvolvimento - listar: lista os recortes disponíveis, sem consultar a fonte | |
| codigo | No | Filtra registros do recorte. Em biomas, usa cd_bioma (ex. "1"). Nos recortes compostos por municípios — amazonia_legal, semiarido, costeiro, fronteira, metropolitana e ride — usa o código IBGE municipal de 7 dígitos. | |
| limite | No | Quantas feições trazer (padrão 50, máx. 600). O total do recorte vem sempre, mesmo quando o limite corta a lista. |
Output Schema
| Name | Required | Description |
|---|---|---|
| tema | Yes | Recorte solicitado (ou 'listar') |
| temas | No | Lista de recortes disponíveis (somente no modo 'listar') |
| codigo | No | Código usado como filtro, quando informado |
| versao | No | Versão/vintage da fonte temática consultada |
| feicoes | No | Total de registros retornáveis na fonte; consulte tipo_registro para interpretar a unidade |
| registros | No | Atributos de cada feição (sem geometria) |
| provenance | Yes | Bloco de proveniência (contrato v1.0): fonte, URL, período, extração e licença |
| attribution | Yes | URLs canônicas das fontes desta resposta (lista de atribuição) |
| fonte_dados | No | URL oficial da fonte de atributos/listagem usada |
| tipo_registro | No | Unidade dos registros retornados: feição geográfica do recorte ou município componente |
| feicoes_retornadas | No | Quantas vieram nesta resposta |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly/idempotent/destructive annotations, the description discloses provenance (versioned official IBGE/GeoFTP snapshots audited by the laboratory), the return format (Markdown plus typed structuredContent), and a hard contract boundary ('THEMATIC GEOMETRY IS NOT PART OF THIS TOOL'S CONTRACT'). These are behavioral facts the annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then cleanly sectioned into recortes, filtering, contract boundary, and routing. Slightly redundant in stating the geometry exclusion twice (once as a contract note, once under 'Use a different tool when'), which costs it the top score.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only listing tool with an output schema, full parameter coverage, and rich annotations, the description supplies everything an agent needs: scope, enumeration, filter semantics, sibling routing, and return format. Nothing material is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already spells out the enum values, the cd_bioma vs. 7-digit municipality code rule, and the limite default/max with 'total always returned'. The description's parameter guidance largely restates the schema rather than adding syntax or edge-case meaning, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Lists what a THEMATIC territorial recorte of Brazil contains'), enumerates all eight recortes, and explicitly distinguishes itself from the sibling ibge_malhas. An agent can identify the tool's scope and output shape without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing is given: 'Use a different tool when: Administrative meshes WITH geometry → ibge_malhas', plus the inverse rule that thematic geometry is out of contract. It also tells the agent how to get a plain listing ('Ask without codigo to list records') and that 'listar' bypasses the source entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ibge_municipiosMunicípios do BrasilARead-onlyIdempotent
Lists Brazilian municipalities from IBGE.
Features:
List municipalities by state (using state abbreviation)
List all municipalities in Brazil (5,570 municipalities)
Search by municipality name
Returns 7-digit IBGE code
Examples:
São Paulo municipalities: uf="SP"
Search by name: busca="Campinas"
MG municipalities containing "Belo": uf="MG", busca="Belo"
Use a different tool when:
Resolve/decode a code at any level (region, state, district), not just municipalities → ibge_geocodigo
Full details/hierarchy of one locality by code → ibge_localidade
Neighboring municipalities → ibge_vizinhos
Behavior: read-only and idempotent — a live GET against the public IBGE Localidades API. Returns a Markdown table.
| Name | Required | Description | Default |
|---|---|---|---|
| uf | No | Estado por sigla (SP), nome (São Paulo) ou código IBGE (35). Se não informado, retorna todos os municípios do Brasil. | |
| busca | No | Termo para buscar no nome do município | |
| limite | No | Número máximo de resultados (padrão: 100, máximo: 5570) |
Output Schema
| Name | Required | Description |
|---|---|---|
| uf | No | UF informada no filtro (como recebida na entrada) |
| busca | No | Termo de busca aplicado ao nome do município |
| total | Yes | Total de municípios encontrados antes do limite |
| municipios | Yes | Lista de municípios retornados (após filtro e limite) |
| provenance | Yes | Bloco de proveniência (contrato v1.0): fonte, URL, período, extração e licença |
| attribution | Yes | URLs canônicas das fontes desta resposta (lista de atribuição) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false. The description restates 'read-only and idempotent' and adds the live GET source and Markdown table output, which is useful context but largely redundant with annotations. Output schema exists, so no need to explain returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose, then features, examples, and exclusions in labeled sections. Every line earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Complete for a read-only filtered-list tool: purpose, modes, examples, boundaries, behavior, and output shape all covered. An agent has everything needed to select and invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds value beyond the schema: concreteness of the 7-digit code result, the 5,570 total, and worked examples showing uf setup with busca combined. Marginal lift above schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Lists Brazilian municipalities from IBGE') and specifies the scope precisely (5,570 municipalities, 7-digit codes). It distinguishes itself from siblings ibge_geocodigo, ibge_localidade, and ibge_vizinhos with explicit routing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit 'Use a different tool when' section names three alternatives (geocodigo, localidade, vizinhos) and the exact conditions selecting each. Examples show the three main invocation modes. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ibge_nomesFrequência e ranking de nomesARead-onlyIdempotent
Queries name frequency and rankings in Brazil (IBGE).
Features:
Name frequency (tipo='frequencia'):
Birth frequency by decade
Multiple names separated by comma
Filter by sex and locality
Name ranking (tipo='ranking'):
Most popular names
Filter by decade, sex, and locality
Available decades: 1930-2010
Examples:
Frequency of "Maria": tipo="frequencia", nomes="Maria"
Compare names: tipo="frequencia", nomes="João,José,Pedro"
2000s ranking: tipo="ranking", decada=2000
Female names: tipo="ranking", sexo="F"
Behavior: read-only and idempotent — a live GET against the public IBGE Nomes (Censo) API. Returns a Markdown table.
| Name | Required | Description | Default |
|---|---|---|---|
| sexo | No | Filtrar por sexo: M (masculino) ou F (feminino) | |
| tipo | Yes | Tipo de consulta: 'frequencia' para buscar nomes específicos ou 'ranking' para ver os mais populares | |
| nomes | No | Para tipo='frequencia': Nome ou nomes separados por vírgula | |
| decada | No | Para tipo='ranking': Década do ranking (ex: 1990, 2000, 2010) | |
| limite | No | Para tipo='ranking': Número de nomes (padrão: 20) | |
| localidade | No | Código IBGE da localidade (UF: 2 dígitos, Município: 7 dígitos) |
Output Schema
| Name | Required | Description |
|---|---|---|
| tipo | Yes | Tipo da consulta realizada |
| ranking | No | Resultado do ranking (presente quando tipo='ranking') |
| frequencia | No | Resultados de frequência (presente quando tipo='frequencia') |
| provenance | Yes | Bloco de proveniência (contrato v1.0): fonte, URL, período, extração e licença |
| attribution | Yes | URLs canônicas das fontes desta resposta (lista de atribuição) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, non-destructive, open-world, so the redundant 'read-only and idempotent' line earns nothing. However, the description adds genuinely new context: the live GET against the public IBGE Nomes (Censo) API, the Markdown-table return shape, and the 1930–2010 decade constraint that the schema does not state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded one-line summary followed by numbered features, examples, and a behavior note — well organized and scannable. Minor waste from the read-only/idempotent sentence duplicating annotations and some repeated param hints, but nothing egregious.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be detailed beyond the Markdown-table note. Both query modes, their required inputs, and the decade range are covered, making the definition sufficient for correct invocation; only minor edge cases (e.g., locality code validation) are left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema description coverage, the schema already documents each parameter and even tags them by tipo ('Para tipo=frequencia', 'Para tipo=ranking'), including comma-separated names and the limit default. The description largely restates those bindings and the decade example list, adding only the 1930–2010 range, so the baseline 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Queries name frequency and rankings in Brazil (IBGE)') and enumerates the two distinct query modes. The resource is unique among the ibge_* siblings, so an agent can immediately tell this tool apart from geography/indicator/calendar tools without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The feature breakdown plus four worked examples make the when-to-use for each mode (tipo='frequencia' vs tipo='ranking') unambiguous. It does not name or exclude any sibling alternative, and gives no explicit when-not-to-use, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ibge_noticiasNotícias do IBGEARead-onlyIdempotent
Searches and lists already-published IBGE news articles and press releases.
Use this to find recent IBGE publications or announcements about a survey or topic — when an indicator was released, or news mentioning a term like "censo". Results are sorted newest-first; with no parameters it returns the 10 most recent items.
Parameters:
busca: free-text term to match (e.g. "PIB", "censo")
tipo: "release" (official publication of survey results) or "noticia" (general news); omit for both
de / ate: date range, format DD/MM/AAAA (e.g. de="01/01/2024", ate="31/12/2024")
destaque: true to return only featured items
quantidade: how many to return (default 10, max 100); pagina: page number to page through more
Each item returns: title, type (release/news), publication date, editoria (section), related products/surveys, a featured flag, a plain-text summary, and a link to the full article. The header reports the total count and current page.
Examples:
Latest 10 news: (no parameters)
Search census: busca="censo"
2024 news: de="01/01/2024", ate="31/12/2024"
Releases only: tipo="release"
Use a different tool when:
Scheduled/upcoming release dates (not yet published) → ibge_calendario
Behavior: read-only and idempotent — a live GET against the public IBGE Notícias API. Returns a Markdown list.
| Name | Required | Description | Default |
|---|---|---|---|
| de | No | Data inicial no formato DD/MM/AAAA (ex: 01/01/2024) | |
| ate | No | Data final no formato DD/MM/AAAA (ex: 31/12/2024) | |
| tipo | No | Tipo de publicação: 'release' ou 'noticia' | |
| busca | No | Termo para buscar nas notícias | |
| pagina | No | Número da página para paginação | |
| destaque | No | Filtrar apenas notícias em destaque | |
| quantidade | No | Quantidade de notícias a retornar (padrão: 10, máximo: 100) |
Output Schema
| Name | Required | Description |
|---|---|---|
| busca | No | Termo de busca aplicado, se houver |
| total | Yes | Total de notícias encontradas na consulta |
| pagina | Yes | Página atual |
| noticias | Yes | Lista de notícias retornadas |
| provenance | Yes | Bloco de proveniência (contrato v1.0): fonte, URL, período, extração e licença |
| attribution | Yes | URLs canônicas das fontes desta resposta (lista de atribuição) |
| totalPaginas | Yes | Número total de páginas |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnly, idempotent, non-destructive, openWorld), but the description adds real behavior: newest-first sorting, default 10 with no params, max 100, pagination via pagina, and a live GET returning Markdown. It slightly restates 'read-only and idempotent' already in annotations, but the added operational detail is substantive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then parameters, examples, and routing in clearly delimited sections. Slightly long with some redundancy (return-field list overlaps the output schema), but every section is scannable and earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Complete for a zero-required-param search tool: coverage of all 7 params, default/pagination behavior, enum meaning, invocation examples, and sibling routing. The output schema exists, yet the description still conveniently summarizes returned fields without it being necessary.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description genuinely extends meaning: it clarifies that tipo='release' is an official survey publication vs 'noticia' general news and that omitting returns both, and adds the DD/MM/AAAA format with examples beyond the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource: 'Searches and lists already-published IBGE news articles and press releases.' It also distinguishes itself from the sibling ibge_calendario by scoping to already-published content, so an agent can tell them apart without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use it ('find recent IBGE publications or announcements... when an indicator was released') and names the alternative with the disambiguating condition: scheduled/upcoming releases route to ibge_calendario. Includes concrete invocation examples.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ibge_paisesDados de paísesARead-onlyIdempotent
Queries international country data via IBGE.
Features:
List all countries (following UN M49 methodology)
Country details (area, languages, currency, location)
Search countries by name
Filter by region/continent
Available regions: americas, europa, africa, asia, oceania
Country codes: Use ISO-ALPHA-2 (e.g., BR, US, AR, PT, JP)
Examples:
List all: tipo="listar"
Brazil details: tipo="detalhes", pais="BR"
Search: tipo="buscar", busca="Argentina"
Americas countries: tipo="listar", regiao="americas"
Available indicators: tipo="indicadores"
Behavior: read-only and idempotent — queries the public IBGE Países API when data is requested and returns Markdown plus typed structuredContent with provenance.
| Name | Required | Description | Default |
|---|---|---|---|
| pais | No | Código ISO-ALPHA-2 do país (ex: BR, US, AR) ou código M49 | |
| tipo | No | Tipo de consulta: listar (todos), detalhes (de um país), indicadores, buscar | listar |
| busca | No | Termo de busca para filtrar países pelo nome | |
| regiao | No | Filtrar por região/continente: americas, europa, africa, asia, oceania | |
| indicadores | No | IDs dos indicadores separados por | (ex: 77819|77820), usados apenas com tipo=detalhes |
Output Schema
| Name | Required | Description |
|---|---|---|
| pais | No | Detalhes de um país específico (modo detalhes) |
| tipo | Yes | Modo de consulta que originou este resultado |
| busca | No | Termo de busca aplicado, se houver |
| total | No | Total de países encontrados (modos listar/buscar) |
| paises | No | Lista de países (modos listar/buscar). Limitada aos 50 primeiros na exibição |
| regiao | No | Filtro de região/continente aplicado, se houver |
| provenance | Yes | Bloco de proveniência (contrato v1.0): fonte, URL, período, extração e licença |
| attribution | Yes | URLs canônicas das fontes desta resposta (lista de atribuição) |
| indicadores | No | Indicadores disponíveis para consulta de países (modo indicadores) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/idempotentHint/openWorldHint, so the safety profile is covered. The description adds value beyond that by naming the backing source (public IBGE Países API) and the return shape (Markdown plus typed structuredContent with provenance).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then cleanly sectioned into Features, regions, codes, examples, and behavior. It is somewhat long and repeats region values already present in the schema, but no sentence is wasted and the structure is easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description needn't detail returns, and it still sketches the response format. Combined with annotations covering safety and a fully-documented 5-parameter schema, nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description goes further by pairing parameters in worked examples (e.g., tipo="detalhes" with pais="BR", tipo="listar" with regiao="americas"), clarifying valid combinations. The region and ISO-ALPHA-2 details largely restate the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Queries international country data via IBGE") and enumerates the four query modes. The 'países/countries' resource is clearly distinct from the state, municipality, and locality siblings, so an agent can route without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The Examples block shows concrete invocation patterns for every mode (listar, detalhes, buscar, region filter, indicadores), giving clear context for how to use the tool. It stops short of stating when-not-to-use or naming a sibling alternative, so it lacks explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ibge_pesquisasPesquisas do IBGEARead-onlyIdempotent
Lists available IBGE surveys and their tables.
Features:
List all IBGE surveys (Census, PNAD, GDP, etc.)
Search by name or code
Show details and tables of a specific survey
Categorize surveys by theme
Main surveys:
Census: Demographic, Agricultural, MUNIC
PNAD Contínua: Employment, income, education
National Accounts: GDP, investments
Economic Surveys: Industry, Commerce, Services
Price Indices: IPCA, INPC
Examples:
List all: (no parameters)
Search population: busca="população"
PNAD details: detalhes="pnad"
This lists surveys, not data. To find table codes use ibge_sidra_tabelas; to query data use ibge_sidra (or a wrapper: ibge_censo, ibge_indicadores, ibge_comparar, ibge_cidades).
Behavior: read-only and idempotent — a live GET against the public IBGE SIDRA/Pesquisas API. Returns a Markdown list.
| Name | Required | Description | Default |
|---|---|---|---|
| busca | No | Termo para buscar no nome ou ID da pesquisa | |
| detalhes | No | Código da pesquisa para ver detalhes e tabelas disponíveis |
Output Schema
| Name | Required | Description |
|---|---|---|
| modo | Yes | Modo de consulta que originou este resultado: lista de pesquisas ou detalhes de uma |
| busca | No | Termo de busca aplicado, se houver (modo lista) |
| total | No | Total de pesquisas encontradas (modo lista) |
| pesquisa | No | Detalhes de uma pesquisa específica (modo detalhes) |
| pesquisas | No | Lista de pesquisas (modo lista) |
| provenance | Yes | Bloco de proveniência (contrato v1.0): fonte, URL, período, extração e licença |
| attribution | Yes | URLs canônicas das fontes desta resposta (lista de atribuição) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the description's restatement of 'read-only and idempotent' earns no credit. It does add that this is a live GET against the public IBGE SIDRA/Pesquisas API, which clarifies the network dependency, but says nothing about rate limits, failures, or empty-result behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and the sibling routing sentence, then structured into Features/Main surveys/Examples/Behavior. The 'Main surveys' catalog is domain-flavored filler that lengthens the definition without changing how the tool is called, keeping this short of a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value detail is unnecessary. Between purpose, explicit sibling alternatives, parameter examples, and the domain framing of survey families, an agent has everything needed to select and invoke this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description goes beyond that by showing the intended shape of each argument in context (busca="população" for name search, detalhes="pnad" for a survey's tables), which clarifies that detalhes takes a survey code rather than an arbitrary string.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Lists available IBGE surveys and their tables') and immediately disambiguates from siblings with 'This lists surveys, not data.' The agent knows exactly what this tool returns versus ibge_sidra and ibge_sidra_tabelas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes to alternatives: ibge_sidra_tabelas for table codes, ibge_sidra or the named wrappers for data. The Examples block maps each parameter mode (none, busca, detalhes) to a concrete intent, so when-to-use is fully specified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ibge_sidraConsulta de tabelas SIDRAARead-onlyIdempotent
Queries SIDRA tables (IBGE's Automatic Recovery System).
SIDRA contains data from IBGE surveys like Census, PNAD, GDP, etc.
Common tables:
6579: Population estimates (annual)
9514: Census 2022 population
200: Census population (1970-2010)
4714: Population, territorial area and density (Census 2022)
4099: Unemployment rate (PNAD Contínua, quarterly)
5436: Average real income (PNAD Contínua, quarterly)
6706: GDP at current prices
5938: GDP per capita
Territorial levels:
1: Brazil
2: Region (North, Northeast, etc.)
3: State (UF)
6: Municipality
7: Metropolitan Region
Examples:
Brazil population 2023: tabela="6579", periodos="2023"
Population by state: tabela="6579", nivel_territorial="3"
Census 2022 by municipality: tabela="9514", nivel_territorial="6", localidades="3550308"
Statistics mode: for largest/smallest/mean/median/distribution/ranking questions ("which municipality has the largest population?", "median GDP by state") use estatisticas=true — it computes min/max/mean/median/std-dev/labeled percentiles over ALL data rows BEFORE pagination and returns top/bottom rankings (default 10, cap 100 via topN), so one call answers what would otherwise require paging thousands of records. With agruparPor="" (e.g. "Unidade da Federação", "Ano") it ranks groups by descending sum, each with its own mini-distribution. Queries mixing several variables auto-group by "Variável" (units differ). SIDRA absence markers ("-", "..", "...", "X") are excluded from n. In this mode pagina/campos/formato are ignored and registros comes empty. Very large queries are refused by the source (since 2026-09-16 SIDRA tables are read through the Aggregates API, whose ceiling is lower than SIDRA's old 100,000-value cap: all municipalities × 12 yearly periods fails, × 8 works) — narrow periodos (e.g. "last 4") or raise nivel_territorial.
ibge_sidra is the low-level engine. Prefer a friendlier wrapper when it fits:
Census themes (1970–2022) → ibge_censo
Economic/social time series → ibge_indicadores
Rank/compare 2–10 localities → ibge_comparar
One municipality's panel → ibge_cidades Use ibge_sidra_tabelas and ibge_sidra_metadados to find a table code and its structure before querying.
Behavior: read-only and idempotent — a live GET against the public IBGE SIDRA API. Returns Markdown plus a typed structuredContent payload.
| Name | Required | Description | Default |
|---|---|---|---|
| topN | No | Tamanho das listas top/bottom quando estatisticas=true sem agruparPor (padrão: 10, máx: 100) | |
| campos | No | Selecionar apenas algumas colunas por rótulo, separadas por vírgula (ex: 'Valor,Ano'). Reduz o volume da resposta. Omitir traz todas. | |
| pagina | No | Página de resultados (100 registros por página) | |
| tabela | Yes | Código da tabela SIDRA (ex: 6579 para estimativas de população, 9514 para censo 2022) | |
| formato | No | Formato de saída: 'json' para dados brutos ou 'tabela' para formato legível | tabela |
| periodos | No | Períodos: 'last' para último, 'all' para todos, ou anos específicos (ex: 2020,2021,2022) | last |
| variaveis | No | IDs das variáveis separados por vírgula, ou 'allxp' para todas | allxp |
| agruparPor | No | Com estatisticas=true, agrupa pela coluna informada (rótulo, ex: 'Unidade da Federação', 'Ano') e ranqueia os grupos por soma decrescente (grupos[0] = maior total), cada grupo com sua mini-distribuição. Nome curto ('UF', 'estado', 'cidade', 'região') e rótulo parcial ('Federação') são resolvidos, e a resposta diz em `aviso` por qual coluna agrupou; rótulo que casa com duas colunas é recusado em vez de escolhido | |
| localidades | No | Códigos das localidades separados por vírgula, ou 'all' para todas | all |
| estatisticas | No | Computa estatísticas (mínimo/máximo/média/mediana/desvio-padrão/percentis) sobre TODOS os registros da consulta, antes da paginação, + ranking top/bottom. Use para 'qual o maior/menor', 'média', 'mediana', 'distribuição', 'ranking'. Quando true, ignora pagina, campos e formato | |
| classificacoes | No | Classificações no formato 'id[categorias]' (ex: '2[6794]' para sexo masculino) | |
| nivel_territorial | No | Nível territorial (código N): 1=Brasil, 2=Região, 3=UF, 6=Município, 7=Região Metropolitana, 8=Mesorregião, 9=Microrregião, 10=Distrito, 11=Subdistrito, 13=RM/RIDE, 14=RIDE, 15=Aglomeração Urbana, 17=Região Geográfica Imediata, 18=Região Geográfica Intermediária, 105=Macrorregião de Saúde, 106=Região de Saúde, 114=Aglomerado Subnormal, 127=Amazônia Legal, 128=Semiárido | 1 |
Output Schema
| Name | Required | Description |
|---|---|---|
| nome | Yes | Nome da tabela (quando conhecido) |
| tabela | Yes | Código da tabela SIDRA consultada |
| colunas | Yes | Rótulos das colunas, na ordem |
| paginacao | Yes | Metadados de paginação para continuação |
| registros | Yes | Registros da página atual: cada um mapeia rótulo da coluna -> valor |
| provenance | Yes | Bloco de proveniência (contrato v1.0): fonte, URL, período, extração e licença |
| attribution | Yes | URLs canônicas das fontes desta resposta (lista de atribuição) |
| estatisticas | No | Bloco estatístico presente quando estatisticas=true (registros vem vazio nesse modo) |
| totalRegistros | Yes | Total de registros de dados disponíveis (todas as páginas) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, but the description adds substantial behavioral context beyond those: it specifies it is a 'live GET against the public IBGE SIDRA API', mentions the return format ('Markdown plus a typed structuredContent payload'), explains that statistics mode ignores pagina/campos/formato, and discloses that SIDRA absence markers are excluded and that very large queries are refused due to the Aggregates API ceiling. This is rich behavioral disclosure with no contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but well-structured: it front-loads the core purpose, then uses labeled sections for common tables, territorial levels, examples, statistics mode, and sibling guidance. Every sentence earns its place—even the long statistics section is necessary for a complex feature. It is slightly verbose (e.g., repeating the Aggregates API ceiling note) but remains organized and scannable, so a 4 is appropriate rather than a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 12 parameters, 1 required, and an output schema present, this description is nearly exhaustive. It covers common usage patterns, edge cases (statistics mode ignoring some params, large query limits), how to discover table codes and metadata, and how to disambiguate from sibling tools. The output schema handles return-value details, so nothing essential is missing for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description goes far beyond: it provides a curated list of common table codes with meanings, explains territorial level codes, gives concrete examples for each parameter, details the statistics mode and agruparPor resolution behavior (including refusal for ambiguous labels), and explains that topN caps at 100. This adds practical semantics that the schema alone does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb+resource statement: 'Queries SIDRA tables (IBGE's Automatic Recovery System).' It further distinguishes itself from siblings by explicitly naming ibge_censo, ibge_indicadores, ibge_comparar, and ibge_cidades as friendlier wrappers for specific use cases, and references ibge_sidra_tabelas and ibge_sidra_metadados as lookup tools. This leaves no ambiguity about what the tool does and how it differs from the others.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit guidance on when to use this low-level engine vs. alternative tools ('Prefer a friendlier wrapper when it fits: Census themes → ibge_censo, Economic/social time series → ibge_indicadores, Rank/compare 2–10 localities → ibge_comparar, One municipality's panel → ibge_cidades'). It also explains the statistics mode for ranking/distribution questions and warns about the Aggregates API ceiling, telling the user to narrow periods or raise territorial level. This is explicit when/when-not guidance with alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ibge_sidra_metadadosMetadados de tabela SIDRAARead-onlyIdempotent
Returns metadata for a specific SIDRA table.
Features:
General info (name, survey, subject, periodicity)
Available territorial levels
Variable list with units
Classifications and categories
Available periods
Use this tool to understand table structure BEFORE querying data with ibge_sidra.
Examples:
Population table metadata: tabela="6579"
Census 2022 metadata: tabela="9514"
PNAD unemployment: tabela="4714"
Use this after finding a table code (ibge_sidra_tabelas) and before querying with ibge_sidra.
Behavior: read-only and idempotent — a live GET against the public IBGE SIDRA API. Returns Markdown.
| Name | Required | Description | Default |
|---|---|---|---|
| tabela | Yes | Código da tabela/agregado SIDRA (ex: '6579', '9514', '4714') | |
| incluir_periodos | No | Incluir lista de períodos disponíveis (padrão: true) | |
| incluir_localidades | No | Incluir níveis territoriais disponíveis (padrão: false) |
Output Schema
| Name | Required | Description |
|---|---|---|
| url | No | URL da tabela no SIDRA |
| nome | Yes | Nome da tabela |
| codigo | Yes | Código da tabela/agregado SIDRA |
| assunto | No | Assunto/tema da tabela |
| periodos | No | Períodos disponíveis para a tabela (quando incluir_periodos) |
| pesquisa | No | Nome da pesquisa de origem |
| variaveis | No | Variáveis da tabela, com unidades e classificações/categorias |
| provenance | Yes | Bloco de proveniência (contrato v1.0): fonte, URL, período, extração e licença |
| attribution | Yes | URLs canônicas das fontes desta resposta (lista de atribuição) |
| periodicidade | No | Periodicidade da pesquisa |
| niveisTerritoriais | No | Níveis territoriais disponíveis para a tabela |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint, destructiveHint=false, but the description goes further by explaining the mechanism: 'a live GET against the public IBGE SIDRA API' and specifying the return format is Markdown. This adds context beyond the structured hints about how and where the call executes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose, followed by a scannable bullet list of returned fields, usage guidance, and concrete examples. The examples with real table codes are useful, though the list of features is somewhat verbose since the output schema likely covers some of it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given an output schema exists, the description doesn't need to explain return values in detail, yet it still provides helpful categories and example table codes. Combined with explicit routing to sibling tools and API behavior, an agent has everything needed to invoke and interpret this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all three parameters including defaults. The description lists metadata categories that map to the optional flags (periods, territorial levels) but does not add syntax or format details beyond the schema's own descriptions. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the specific verb and resource ('Returns metadata for a specific SIDRA table') and enumerates the metadata fields returned. It clearly distinguishes itself from the sibling data-query tool ibge_sidra and the table-search tool ibge_sidra_tabelas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing: 'Use this after finding a table code (ibge_sidra_tabelas) and before querying with ibge_sidra.' Both prerequisite and follow-up tools are named, and the 'BEFORE querying data' condition is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ibge_sidra_tabelasBusca de tabelas SIDRAARead-onlyIdempotent
Lists and searches available SIDRA tables.
Features:
List all SIDRA tables (aggregates)
Search by table name: every word must match (AND), accents and case ignored, and everyday Portuguese is resolved to the IBGE's own wording (renda→rendimento, desemprego→desocupação, cidade→município, gênero→sexo); when that happens the response says so in notas_vocabulario
Filter by survey (Census, PNAD, GDP, etc.)
Shows code and name of each table
SIDRA contains data from various surveys:
Demographic Census
PNAD Contínua (employment, income)
National Accounts (GDP)
Industrial Survey
Agricultural Survey
Examples:
List tables: (no parameters)
Search population tables: busca="população"
Census tables: pesquisa="censo"
This is step 1 of the SIDRA workflow: find a table code → ibge_sidra_metadados (structure) → ibge_sidra (query). For common data, a wrapper is usually easier: ibge_censo, ibge_indicadores, ibge_comparar, ibge_cidades.
Behavior: read-only and idempotent — a live GET against the public IBGE SIDRA API. Returns a Markdown table.
| Name | Required | Description | Default |
|---|---|---|---|
| busca | No | Termos para buscar no nome das tabelas/agregados (sem distinção de acento ou caixa; AND entre as palavras; a palavra de todo dia é traduzida para a do IBGE — renda→rendimento, desemprego→desocupação, cidade→município) | |
| limite | No | Número máximo de resultados (padrão: 20) | |
| pesquisa | No | Filtrar por código ou nome da pesquisa (ex: 'censo', 'pnad', 'pib') |
Output Schema
| Name | Required | Description |
|---|---|---|
| busca | No | Termo de busca aplicado, se houver |
| total | Yes | Total de tabelas que correspondem aos critérios |
| tabelas | Yes | Lista de tabelas SIDRA retornadas |
| pesquisa | No | Filtro de pesquisa aplicado, se houver |
| provenance | Yes | Bloco de proveniência (contrato v1.0): fonte, URL, período, extração e licença |
| attribution | Yes | URLs canônicas das fontes desta resposta (lista de atribuição) |
| notas_vocabulario | No | Quando a busca foi ampliada para a palavra que o IBGE usa (renda→rendimento), diz qual |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds valuable behavioral context: it performs a live GET against the public IBGE SIDRA API, returns a Markdown table, and explains vocabulary normalization (renda→rendimento, etc.) with a notas_vocabulario note. It could add rate-limit or pagination details, but it goes beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with clear sections (Features, Surveys, Examples, Workflow) and front-loads the core purpose. It is somewhat long, but every section earns its place by providing routing and behavioral context. A few redundant lines (e.g., listing surveys twice) could be trimmed, but it remains efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list/search tool with zero required parameters, an output schema, and full schema coverage, the description is complete. It explains the workflow position, the return format, the vocabulary normalization behavior, and the alternatives. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters. The description adds examples and explains the AND/accents/case behavior for busca, but it mostly restates what the schema says. Baseline 3 is appropriate because the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Lists and searches available SIDRA tables') and clearly distinguishes this tool from siblings by naming the SIDRA workflow and wrappers. It also enumerates features and examples, so an agent can tell it apart from ibge_sidra_metadados and ibge_sidra without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says this is step 1 of the SIDRA workflow, names the follow-up tools (ibge_sidra_metadados, ibge_sidra), and advises that wrappers (ibge_censo, ibge_indicadores, ibge_comparar, ibge_cidades) are usually easier for common data. This gives clear when-to-use and when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ibge_vizinhosMunicípios vizinhosARead-onlyIdempotent
Finds nearby/neighboring municipalities.
Features:
Search by IBGE code (7 digits) or municipality name
Without raio: returns municipalities that actually touch the reference municipality in the official municipal mesh
With raio: returns municipalities whose centroids fall within the requested radius in km
Optionally includes population data
Examples:
By code: municipio="3550308"
By name: municipio="Campinas", uf="SP"
With population: municipio="3550308", incluir_dados=true
Contiguity is topological (shared boundary); radius mode is an approximation based on centroid distance. For listing/searching municipalities, use ibge_municipios.
Behavior: read-only and idempotent — uses the public IBGE Localidades and Malhas v3 APIs; population enrichment, when requested, uses SIDRA. Returns a Markdown list.
| Name | Required | Description | Default |
|---|---|---|---|
| uf | No | Estado por sigla (SP), nome (São Paulo) ou código IBGE (35) — obrigatório se usar nome do município | |
| raio | No | Raio em km para buscar municípios próximos, calculado pela distância entre centróides municipais | |
| municipio | Yes | Código IBGE do município (7 dígitos) ou nome do município | |
| incluir_dados | No | Incluir dados populacionais dos vizinhos |
Output Schema
| Name | Required | Description |
|---|---|---|
| total | Yes | Quantidade de municípios encontrados |
| vizinhos | Yes | Lista de municípios contíguos ou próximos, conforme o critério espacial solicitado |
| municipio | Yes | Município de referência da consulta |
| provenance | Yes | Bloco de proveniência (contrato v1.0): fonte, URL, período, extração e licença |
| attribution | Yes | URLs canônicas das fontes desta resposta (lista de atribuição) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, openWorld, non-destructive), yet the description adds substantive traits: the exact upstream sources (IBGE Localidades, Malhas v3, SIDRA for population), the topological-vs-approximate nature of each mode, and the return format (Markdown list). This is real context an agent cannot get from annotations alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well front-loaded: purpose first, then features, examples, and behavior. Sections earn their place, though the raio/centroid explanation is repeated in both Features and Behavior, which is mild redundancy rather than waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists, so return-value documentation is not required, but the description still notes the Markdown-list format. Combined with mode semantics, data provenance, and examples, an agent has everything needed to call this 4-parameter tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3; the description goes beyond it by tying raio's presence to a mode switch and by showing concrete invocation shapes for municipio, uf, and incluir_dados. It does not add format details (e.g. acceptable sigla/nome variants) beyond what the schema already states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Finds nearby/neighboring municipalities') and immediately names the sibling it is not ('For listing/searching municipalities, use ibge_municipios'). An agent can distinguish it from the other 20+ ibge_* tools without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly contrasts the two operating modes (contiguous neighbors vs. radius by centroid distance) and states the selection condition, plus routes listing/searching to ibge_municipios. Worked examples for code, name+uf, and population enrichment remove remaining ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
searchBusca para Deep ResearchARead-onlyIdempotent
Searches the IBGE (Brazilian official statistics: SIDRA tables, municipalities, known indicators) catalog and returns up to 10 matching documents as { id, title, url }, ordered by relevance (an empty list means nothing matched).
This tool exists for the OpenAI Deep Research contract: ChatGPT deep research, company knowledge and research workflows over the Responses API require exactly the tools search and fetch. Pass one of the returned ids to fetch to read the document.
For direct questions and for data (values, series, rankings) prefer the ibge_* tools (ibge_sidra, ibge_cidades, ibge_indicadores, ibge_comparar…), which return the actual data with provenance — this is a catalog index, not a data query.
Query: natural language or keywords, Portuguese or English; accents and case are ignored.
Behavior: read-only and idempotent — the catalog comes from the public source and is cached in memory.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Termos de busca em linguagem natural ou palavras-chave (acentos e caixa são ignorados) |
Output Schema
| Name | Required | Description |
|---|---|---|
| results | Yes | Documentos encontrados, em ordem de relevância |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotent and non-destructive, so the description only needs to add context — and it does: results are cached in memory from the public source, an empty list signals no match, and the index is explicitly not a data source. It does not discuss ranking mechanics or catalog staleness, but the annotation burden is already met.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with what the tool does and returns, then guidelines, then query semantics, then behavior — a sensible ordering with no true filler. However, the accent/case sentence duplicates the schema and the Deep Research contract paragraph is longer than the routing information it carries.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need no expansion, and the description still summarizes cardinality and ordering. For a single-parameter search tool sitting inside a 22-tool catalog, everything an agent needs to select it, query it, and route its output is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single `query` param is already documented in the schema, including the accent/case insensitivity the description repeats. The description's only marginal addition is that Portuguese or English both work, which is a small gain over the structured field — baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — searches the IBGE catalog — and pins the exact return shape ({ id, title, url }, up to 10, relevance-ordered). It explicitly distinguishes itself from siblings by naming the ibge_* data tools and fetch as the downstream consumer of returned ids.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use (OpenAI Deep Research / Responses API contract requiring exactly `search` and `fetch`), when-not-to-use (direct questions and data/values/rankings should go to ibge_sidra, ibge_cidades, ibge_indicadores, ibge_comparar), and the follow-up action (pass a returned id to `fetch`).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
23 tool updates
v0.6.0- First observed
fetch - First observed
ibge_calendario - First observed
ibge_censo - First observed
ibge_cidades - First observed
ibge_cnae - First observed
ibge_comparar - First observed
ibge_datasaude - First observed
ibge_estados - First observed
ibge_geocodigo - First observed
ibge_indicadores - First observed
ibge_localidade - First observed
ibge_malhas - First observed
ibge_malhas_tema - First observed
ibge_municipios - First observed
ibge_nomes - First observed
ibge_noticias - First observed
ibge_paises - First observed
ibge_pesquisas - First observed
ibge_sidra - First observed
ibge_sidra_metadados - First observed
ibge_sidra_tabelas - First observed
ibge_vizinhos - First observed
search
TDQS
Scored across 23 tools
Several SIDRA-based statistical wrappers (ibge_censo, ibge_indicadores, ibge_datasaude, ibge_comparar, ibge_cidades) overlap in the indicators they expose, and population data can be obtained from multiple tools. The extensive cross-references in descriptions help mitigate confusion, but an agent could still misselect when a question spans themes (e.g., population by municipality).
21 of 23 tools follow a consistent `ibge_` + snake_case noun pattern (e.g., ibge_sidra, ibge_municipios, ibge_sidra_tabelas). The two exceptions, `search` and `fetch`, are mandated by the OpenAI Deep Research contract and explained, but they still break the otherwise uniform prefix convention.
23 tools is heavy, but they map to distinct IBGE API families (SIDRA, Localidades, Nomes, Notícias, Malhas, CNAE, Países, etc.). The overlap among SIDRA wrappers suggests some consolidation is possible, but the breadth of the domain makes the count reasonable rather than excessive.
The surface covers the major IBGE data domains comprehensively: geography, census, economic/social indicators, health, names, news, release calendar, CNAE, countries, and meshes. The low-level ibge_sidra tool provides arbitrary SIDRA query access, closing gaps for less common surveys; no obvious read-only operation is missing for the stated domains.
Maintenance
Related MCP Connectors
IBGE: geography, census, economy and health from the official APIs, with provenance. 23 tools.
Discover, resolve, and query official Brazilian economic data with semantic search and provenance.
Banco Central do Brasil (BCB): SGS series, Focus expectations, PTAX, stats + provenance. 17 tools.
Brazilian public data API for AI agents. BCB, IBGE, CVM, B3, compliance. x402 payments on Base.
Related MCP Servers
- AlicenseAqualityBmaintenanceExposes official IBGE data as MCP tools, including Brazilian localities, SIDRA statistical aggregates, and population indicators.11MIT
- FlicenseNot gradedqualityDmaintenanceEnables LLMs to access Brazilian IBGE statistical data (population, economy, agriculture) via natural language, with tools for querying aggregated data and metadata.-
- AlicenseAqualityDmaintenanceEnables AI agents to access Brazilian statistical, geographic, and economic data in real-time via IBGE public APIs.3211 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables agents to query 153 live Brazilian government data tools across federal and state sources, including economy, legislature, judiciary, elections, health, education, and more, reading directly from original APIs.MIT