revisor-notas
It is an MCP server that validates and annotates telemedicine clinical notes in SOAP (F/S/O/A/P) format entirely on your machine.
validar_nota_soap: validates a note using deterministic rules (required sections, CID format, plan numbering, footer, subjective fields) and, if available, a local LLM for semantic coherence between complaint, exam, and conduct, plus red flags not investigated.Returns structured problems with section, severity (
erro/aviso), description, suggestion, and origin (regraorsemantica).Falls back to rule-only validation when the local LLM is unavailable, indicating
checagem_semantica_feita: falseand the reason.sugerir_correcoes: returns the same note with inline<<ERRO: ...>>/<<AVISO: ...>>annotations, without rewriting the original text.Groups document-level issues in a
<<PROBLEMAS NO DOCUMENTO>>block and places each annotation only once at the end of the relevant section.Guarantees privacy: only localhost LLM calls are allowed, nothing is persisted, and no note content is logged.
Can be used directly via MCP clients like Claude Code, or through CLI commands for validation and annotation.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@revisor-notasconfira os campos obrigatórios e a coerência desta nota SOAP"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
revisor-notas-mcp
Servidor MCP que confere notas clínicas no formato SOAP (F/S/O/A/P): checa se os campos obrigatórios estão preenchidos e usa um LLM local para apontar incoerências entre queixa, exame e conduta. Roda inteiramente na sua máquina, o texto da nota não sai dela.
⚠️ Não substitui o julgamento clínico
Isto não é um software de apoio à decisão clínica. Não foi validado clinicamente, não foi avaliado por nenhum órgão regulador e não é um dispositivo médico.
A checagem semântica é probabilística. Ela é feita por um LLM lendo texto livre: erra para os dois lados. Ausência de avisos não atesta que a nota está correta, e um aviso pode ser falso positivo.
A responsabilidade pela nota continua sendo de quem assina. A ferramenta aponta o que pode ter faltado; ela não sabe do paciente nada além do que está escrito.
O projeto foi escrito em torno de um template específico de telemedicina. Se o seu formato for outro, as regras determinísticas precisam ser ajustadas.
Requisitos
O quê | Versão | Para quê |
Python | 3.12+ | runtime |
recente | dependências e venv | |
Um LLM local com API OpenAI-compatible | - | checagem semântica (opcional) |
Sem banco de dados, sem sync e sem timer systemd: este projeto não persiste nada e só roda quando o cliente MCP chama uma tool.
Related MCP server: Konsilium
Instalação
A distribuição é por clone + uv (não há pacote no PyPI).
git clone https://github.com/fabianofilho/revisor-notas-mcp.git
cd revisor-notas-mcp
uv sync
cp .env.example .envConfiguração
Variável | Padrão | Observação |
|
| llama.cpp. Ollama: |
|
| llama.cpp e LM Studio aceitam qualquer nome |
|
| curto de propósito: a checagem não pode travar a resposta |
O endpoint precisa ser localhost. Qualquer outro host é recusado com exceção, não com aviso, ver Privacidade.
uv run revisor-cli llm # confirma o LLM local
uv run revisor-cli validar nota.txt # ou cole a nota no stdin
uv run revisor-cli validar nota.txt --sem-llm # só as regras determinísticas
uv run revisor-cli anotar nota.txt # nota com anotações inlineLigando ao Claude Code
claude mcp add revisor-notas --scope user \
-e QWEN_ENDPOINT=http://127.0.0.1:8080/v1 \
-e QWEN_MODEL=local-model \
-- uv --directory /caminho/para/revisor-notas-mcp run revisor-notas-mcpUso
validar_nota_soap(texto_nota: str)
Lista os problemas encontrados, cada um com seção, severidade, sugestão e a origem
(regra ou semantica).
{
"total_erros": 0,
"total_avisos": 2,
"checagem_semantica_feita": true,
"problemas": [
{
"secao": "F, S, A, P",
"severidade": "aviso",
"descricao": "A queixa é dor torácica irradiando para o braço esquerdo, com hipertensão e tabagismo, o que sugere etiologia cardíaca. A hipótese final é dor muscular e a conduta é analgésico simples.",
"origem": "semantica"
},
{
"secao": "S",
"severidade": "aviso",
"descricao": "Sinais de alarme esperados que não aparecem como investigados: sudorese, dispneia, náuseas.",
"origem": "semantica"
}
]
}Esse exemplo é a saída real de uma nota sintética de teste. Note que total_erros é zero:
a nota está formalmente correta, as regras determinísticas passam limpo. O que a
camada semântica aponta é clínico, e é exatamente o que regra não pega.
sugerir_correcoes(texto_nota: str)
A mesma nota com os problemas de validar_nota_soap inseridos como linhas <<ERRO: ...>>
ou <<AVISO: ...>>. Não reescreve o texto original: nenhuma linha da nota é alterada,
e quem decide o que mudar é quem assina.
Cada problema aparece uma única vez, no fim da seção a que se refere. Um achado da camada semântica que cita várias seções vai para a primeira delas que existe na nota, com as seções entre colchetes, por exemplo
<<AVISO [F, A, P]: ...>>.O que não tem onde ficar (cabeçalho ausente, seção ausente ou vazia sem marcador, seção fora do padrão) vai para um bloco
<<PROBLEMAS NO DOCUMENTO>>no topo.total_anotacoesé o número de problemas, igual ao número de linhas<<...>>fora o título do bloco.
<<PROBLEMAS NO DOCUMENTO>>
<<ERRO [O]: Seção -O: ausente ou vazia. → Preencha a seção -O: conforme o template.>>
#TELEMEDICINA#
-F: Paciente refere cefaleia. Sem sinais de alarme.
...
-A: Cefaleia tensional (CID: R51).
-P:
1. Dipirona 500mg via oral se dor.
...
Atendimento realizado via telemedicina, não sendo possível aferição de sinais vitais.
<<AVISO: Plano sem os itens numerados [2, 4, 5]. → O template prevê os itens 1 a 5, sendo o 5 o do atestado.>>Saída real (sem LLM) de uma nota sintética de teste, com linhas omitidas.
As duas camadas
Regras determinísticas (rules/) rodam sempre, sem LLM: cabeçalho, as cinco seções,
CID em formato válido, itens numerados do plano, item de sinais de alarme, rodapé fixo, e
os campos do subjetivo (medicações em uso, antecedentes, alergia, hábitos).
Checagem semântica (llm/) olha coerência entre queixa, exame e conduta; e sinais de
alarme típicos do diagnóstico que não aparecem como investigados (negativa explícita conta
como investigado).
Se o LLM local estiver fora do ar (ou der timeout, ou devolver algo que não é JSON), a
validação por regras responde sozinha: checagem_semantica_feita vem false e
motivo_semantica_pulada diz o porquê. As regras nunca ficam bloqueadas pelo modelo.
Os prompts do LLM estão em src/revisor_notas_mcp/prompts/ e fazem parte do pacote.
Limitações conhecidas
O parser espera um template específico. Ele tolera variação de formatação (marcador
com ou sem hífen, caixa baixa, espaço sobrando), mas as regras de conteúdo (quais campos
são obrigatórios, qual rodapé, quais itens no plano) foram escritas para um template de
telemedicina. Adaptar para outro formato significa mexer em rules/checklist.py.
A checagem semântica depende de um modelo pequeno. Rodando local, ela custa alguns segundos e a qualidade varia com o modelo. Um modelo fraco vai gerar falso positivo.
A validação de CID é sintática, não semântica. Ela confere o formato (letra + dois dígitos, com subcategoria opcional), não se o código corresponde ao diagnóstico escrito.
Nada é cacheado. Cada chamada reprocessa a nota do zero, porque cachear significaria guardar o texto, e isso o projeto não faz.
Privacidade
Este é o ponto central do projeto, e ele está em código e em teste, não só aqui:
Nenhuma chamada de rede além de
localhost. O cliente do LLM recusa com exceção (EndpointNaoLocal) um endpoint que aponte para qualquer outro host. Configuração perigosa estoura em vez de degradar em silêncio.Nada é gravado. Sem banco, sem cache em disco, sem log do conteúdo. O processamento é em memória, por chamada. Há um teste que confere que nenhum trecho da nota aparece em log, inclusive quando a chamada ao LLM falha.
Fixtures de teste são sintéticas, nunca dado real, nem anonimizado.
Sem telemetria, sem analytics.
O que sai da sua máquina: nada. O que vai para o seu LLM local: o texto da nota, pelo
localhost.
Segurança
Veja SECURITY.md para reportar uma vulnerabilidade. As mudanças por versão estão no CHANGELOG.md.
Contribuindo
Veja CONTRIBUTING.md. A regra mais importante e sem exceção: nenhuma fixture pode conter dado real de paciente, nem anonimizado.
Licença e atribuição
Apache License 2.0: escolhida por o projeto tocar em dado de paciente, onde a cláusula explícita de patente é mais protetiva.
Construído no contexto do IA.med.
Available Tools
2 toolssugerir_correcoesA
Devolve a nota com anotações inline do que precisa ser ajustado.
Não reescreve o texto original: cada problema de validar_nota_soap vira uma linha <<ERRO: ...>> ou <<AVISO: ...>>, uma única vez, no fim da seção a que se refere. O que não tem seção na nota (cabeçalho, seção ausente) vai para um bloco <> no topo. Quem decide o que mudar é o médico.
Args: texto_nota: a nota completa a ser anotada.
| Name | Required | Description | Default |
|---|---|---|---|
| texto_nota | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| aviso | No | |
| nota_anotada | Yes | |
| total_anotacoes | Yes | Número de problemas; cada um aparece uma vez |
| motivo_semantica_pulada | No | |
| checagem_semantica_feita | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden — and it delivers. It discloses that the original text is never rewritten ('Não reescreve o texto original'), specifies the exact annotation format (<<ERRO: ...>> / <<AVISO: ...>>), guarantees each problem appears once ('uma única vez'), defines placement rules including the special-case block for header/missing-section issues, and clarifies that the doctor decides what to change. This is rich, non-obvious behavioral context an agent needs to invoke the tool correctly and interpret its output.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded in the first sentence, followed by dense behavioral details and a clean Args section. Every sentence earns its place, though the middle passage is information-dense and could be slightly better organized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema is present, so return-value documentation is handled externally, and the single parameter is fully explained in the description. The behavioral contract (format, placement, deduplication, non-rewriting guarantee) is well covered; the only notable gap is explicit routing guidance against the sibling tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the parameter meaning. The Args section explains 'texto_nota' as 'a nota completa a ser anotada', conveying that the input must be the full note and that it is the object being annotated — meaning the schema's bare title 'Texto Nota' does not provide. It stops short of format constraints (e.g., whether it must be a SOAP note), which is a minor gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Devolve a nota com anotações inline do que precisa ser ajustado' — returns the note with inline annotations of what needs adjustment. It also differentiates itself from its sibling validar_nota_soap by stating that each problem from that validation becomes an annotation line, making the boundary between the two tools clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The reference to 'cada problema de validar_nota_soap' implies this tool consumes validation problems and renders them as annotations, giving the agent context about when it is relevant. However, it never explicitly states when to choose this tool over validar_nota_soap or when not to use it, leaving the routing decision to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validar_nota_soapA
Valida uma nota clínica no formato #TELEMEDICINA# (F/S/O/A/P).
Checa campos obrigatórios, CID, numeração do plano e rodapé por regras determinísticas, e usa o LLM local para conferir coerência entre queixa, exame e conduta, além de sinais de alarme não investigados. A parte do LLM é probabilística e vem marcada como tal; se ele estiver fora do ar, as regras respondem sozinhas e motivo_semantica_pulada diz por que a semântica não rodou.
O texto processado não sai da máquina nem é gravado em lugar nenhum.
Args: texto_nota: a nota completa, como seria colada no prontuário.
| Name | Required | Description | Default |
|---|---|---|---|
| texto_nota | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| aviso | No | |
| problemas | Yes | |
| total_erros | Yes | |
| total_avisos | Yes | |
| motivo_semantica_pulada | No | |
| checagem_semantica_feita | Yes | False quando a semântica não rodou (motivo em motivo_semantica_pulada); as regras valeram assim mesmo |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully carries the behavioral burden. It discloses that LLM-based checks are probabilistic and flagged as such, that offline LLM triggers a rule-only fallback with motivo_semantica_pulada, and that processed text never leaves the machine or is persisted. These are essential behavioral traits beyond the bare input schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence carries substantive information: purpose, rule/LLM split, privacy, and parameter meaning. It is slightly long but not verbose — the detail about fallback and privacy is valuable, not filler. Front-loading the core purpose is a good structural choice.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter validation tool with an output schema, the description covers input semantics, processing modes (deterministic vs. LLM), failure behavior, and privacy. Since an output schema exists, return details are not required, and the mention of motivo_semantica_pulada aligns with what the schema likely exposes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains that texto_nota is the complete note 'como seria colada no prontuário', which conveys content expectations and the format (#TELEMEDICINA# F/S/O/A/P). A concrete example or length hint would be an extra improvement, but the given context is sufficient.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Valida') and resource (nota clínica no formato #TELEMEDICINA#), and details exactly what is checked (obrigatórios, CID, numeração do plano, rodapé, coerência). This clearly differentiates it from the sibling 'sugerir_correcoes' — validating is not suggesting corrections.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a clear use context: when a SOAP note needs validation against deterministic rules and semantic consistency. It does not explicitly exclude scenarios or point to the sibling as an alternative, but the scope of validation is concrete enough that an agent can infer when to call it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
- Changed
sugerir_correcoes2 fields changed- added
Output schema / properties / motivo_semantica_puladaAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Motivo Semantica Pulada" +} - added
Output schema / properties / total_anotacoes / descriptionAdded value: +"Número de problemas; cada um aparece uma vez"
- Changed
validar_nota_soap1 field changed- changed
Output schema / properties / checagem_semantica_feita / descriptionPrevious value: -"False quando o LLM local não respondeu; as regras valeram assim mesmo"New value: +"False quando a semântica não rodou (motivo em motivo_semantica_pulada); as regras valeram assim mesmo"
2 tool updates
v0.1.0- First observed
sugerir_correcoes - First observed
validar_nota_soap
TDQS
Scored across 2 tools
The two tools are clearly related but serve distinct output purposes: one produces a validation report, the other an annotated version of the note. Descriptions explicitly differentiate them, though an agent might briefly hesitate about which to call for a given user request.
Both tool names follow the same verb_noun pattern in Portuguese, using snake_case: validar_nota_soap and sugerir_correcoes. The convention is consistent and the verbs accurately describe the actions.
With only two tools, the server is on the smaller side, but they cover the two core operations for a note reviewer: validate and suggest corrections. The count feels reasonable rather than thin because each tool is substantial and has a clear purpose.
The domain is clinical note review, and the pair covers validation (including semantic checks) and actionable inline suggestions. There are no obvious dead ends, though a tool for a structured list of errors separately from annotation is already covered by validar_nota_soap.
Maintenance
Related MCP Connectors
Consent-gated tools that turn user health notes into non-diagnostic appointment-prep materials.
HealthGuard - 12-tool health/medical AI safety MCP: PII redaction, HIPAA, GDPR Art.9.
Medical RAG: semantic search for clinical guidelines, drug interactions, diagnoses & EHR data.
Medical RAG: semantic search for clinical guidelines, drug interactions, diagnoses & EHR data.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables secure natural language querying of a clinical SQL Server database using a local Ollama model, with strict SQL validation to protect sensitive data.-
- AlicenseNot gradedqualityCmaintenanceEnables privacy-first medical document analysis with multi-perspective AI review. Ingest documents, run consilium reviews, generate doctor letters, and search patient memory—all through natural language.Apache 2.0

soapnoteapi-mcpofficial
AlicenseAqualityDmaintenanceEnables AI agents to turn clinical transcripts or audio recordings into structured SOAP notes, ICD-10/CPT billing codes, patient summaries, and visit summaries via SOAPNoteAPI.646 npmMIT- FlicenseAqualityCmaintenanceEnables local analysis of unstructured documents (PDF, DOCX, PPTX, SVG, PNG) by extracting text and structure with citation anchors, and verifies summaries against source material before a human approves saving a report.9-