The Gastrologer — GI clinical tools for agents
Server Details
Physician-governed GI tools: claim verification, teaching retrieval, HRM and instrument scoring
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-03-26
- URL
TDQS
Scored across 15 tools
There is a dense cluster of claim-related tools (check_gi_claim, resolve_gi_claim, search_gi_verifications, get_gi_verification, get_evidence_bundle, verify_claim_hash) whose boundaries are only partially clarified by their verbose descriptions. The distinction between checking pasted text vs resolving a proposition vs searching vs fetching by id is real but subtle and easy to miscall. Other tools (read_module, list_modules, score_clinical_instrument, query_relations) are clearly distinct.
Names follow a mostly consistent snake_case verb_noun pattern (get_*, search_*, list_*, read_*, verify_*, resolve_*). The only irregularity is the inconsistent use of the 'gi_' infix (check_gi_claim, resolve_gi_claim, get_gi_verification) versus omitting it elsewhere (get_claims, get_evidence_bundle). This is a minor deviation that does not impede readability.
Fifteen tools is at the top of the well-scoped range but plausibly justified by the breadth of the clinical domain (search, fetch, verify, score, graph query, module reading). One tool (classify_hrm_metrics) is a quarantined legacy alias that arguably does not earn its place. Overall reasonable.
The surface covers the core lifecycle: discover (list/search), read (read_module, get_claims), verify (verify_claim_hash, get_gi_verification), resolve (check/resolve_gi_claim), and analyze (query_relations, score_clinical_instrument). Minor gaps exist, such as no explicit write/adjudication or bundle-diff tooling, but agents can work around them.
Available Tools
15 toolscheck_gi_claimARead-onlyIdempotentInspect
Fact-check pasted text (a social post, caption, or sentence): canonical wording or its full explicit negation can receive a scoped governed verdict; other matches return unresolved reviewed candidates with evidence. No video retrieval or individual medical advice.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotent, non-destructive, closed-world behavior. The description adds meaningful matching semantics: canonical wording or full explicit negation can get a scoped governed verdict, while other matches return unresolved reviewed candidates with evidence. It does not detail response structure or auth, but adds substantial context beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core purpose and followed by precise scope exclusions. Every clause earns its place without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read-only tool with rich annotations but no output schema, the description explains return behavior and scope exclusions adequately. It could mention the input length limit, but the schema provides maxLength.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains that the input is pasted text and gives examples, but it omits the maxLength 2000 constraint and other formatting details. It partially compensates for the missing schema description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
It states a specific verb (Fact-check) and resource (pasted text such as a social post, caption, or sentence), and describes the outcome. It does not name a sibling tool or distinguish itself from resolve_gi_claim, so sibling differentiation is missing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives implied usage context (pasted text) and scope exclusions (no video retrieval, no individual medical advice). However, it does not explicitly say when to use this versus resolve_gi_claim or other sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
classify_hrm_metricsBRead-onlyIdempotentInspect
Legacy HRM operation quarantined after rule discrepancies were reproduced against reviewed teaching. Returns needs_review without classification. Do not invoke to obtain a diagnosis; this compatibility name is omitted from new plugin products.
| Name | Required | Description | Default |
|---|---|---|---|
| metrics | Yes | ||
| position | No | ||
| provocative | No | ||
| manufacturer | No | ||
| secondary_testing | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnly/idempotent/non-destructive, and the description goes further by disclosing the return behavior ('Returns needs_review without classification') and the reason for quarantine (rule discrepancies reproduced). That is meaningful behavioral context beyond the structured hints, and it is consistent with them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences that front-load the status ('quarantined') and the practical consequence ('Returns needs_review'), then the prohibition. Dense phrasing like 'reproduced against reviewed teaching' is opaque but costs little space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a quarantined no-op with no output schema, the description covers the essentials an agent needs (don't call it, it won't classify). However, the nested metrics/position/manufacturer inputs are entirely unexplained, leaving a gap if an agent ever needed to reason about the shape of expected input.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Five parameters, including two nested objects and two enums, have 0% schema description coverage, and the description supplies no parameter meaning at all. The only implicit signal is that inputs are irrelevant because the tool never classifies, which is thin compensation for a schema this complex.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies this as a 'Legacy HRM operation' that is 'quarantined,' which tells the agent it is a disabled/deprecated entry point, but it never states what HRM metric classification actually is (e.g., high-resolution manometry analysis) or what a working version would return. The effective behavior — 'Returns needs_review without classification' — is clear, but the nominal purpose is left vague.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit prohibition: 'Do not invoke to obtain a diagnosis,' plus the deployment note that this name is omitted from new plugin products — strong when-not-to-use guidance. It stops short of naming a working alternative among the siblings (e.g., score_clinical_instrument), so it is clear but not fully routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_claimsBRead-onlyIdempotentInspect
Get the claim ledger (claim → evidence rows) for a module, when published.
| Name | Required | Description | Default |
|---|---|---|---|
| slug | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint=false and destructiveHint=false, so the safety and determinism profile is covered by structured data. The description's only added behavioral signal is the 'when published' gate, which hints at failure or empty results for unpublished modules but never spells out what happens. No mention of auth, rate limits or return complexity beyond the row shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One tight sentence with the resource front-loaded and the constraint ('when published') appended. Nothing is padded and there is no redundant restatement of the name or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the burden of describing the return value; it does so partially by naming 'claim → evidence rows'. Annotations cover the safety profile. What is missing is the behavior of the publication gate and the slug semantics, which for a single-param, no-output-schema tool leaves a real gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there is a single required parameter (slug) that the schema documents only via a regex pattern. The description's 'for a module' implicitly maps slug to a module identifier, but it never states that the slug is the module slug or what format it expects. With low coverage the description should compensate more than it does.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb (get) and resource (the claim ledger, sketched as claim → evidence rows) scoped to a module and gated on publication status. That is enough for an agent to know roughly what comes back. It does not, however, contrast this tool with close siblings like get_evidence_bundle or verify_claim_hash, so an agent still has to infer which ledger-style tool it wants.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'when published' implies a precondition for using the tool, which is genuine but terse guidance. There is no explicit when-to-use versus when-not-to-use, nor any naming of the alternatives (e.g. get_evidence_bundle) that an agent should pick instead. Usage is therefore inferred rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_corpus_stateARead-onlyIdempotentInspect
The current GASTROLOGER_CLINICAL_STATE: version, corpus/section/module Merkle roots, counts and hash scheme — pin an answer to an exact knowledge state.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is fully covered. The description adds the semantic content of the state (version, Merkle roots, counts, hash scheme), which is useful, but says nothing about caching, error behavior, or return format beyond a field list.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the resource and immediately lists the return fields. No wasted sentences, though the internal jargon name and em-dash clause make it slightly harder to parse than it needs to be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no parameters, the description carries the return-shape burden and does so by enumerating version, roots, counts, and hash scheme — enough for an agent to know what pinning to a knowledge state yields. It stops short of explaining how to use the hashes downstream, but is largely complete for a zero-arg read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is no parameter surface to document; baseline is 4. The description does not need to explain invocation arguments.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource ('GASTROLOGER_CLINICAL_STATE') and enumerates what it returns: version, corpus/section/module Merkle roots, counts and hash scheme. This is a clear read of the corpus-level state, distinguishable from siblings like get_claims or list_modules, though it does not explicitly name those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'pin an answer to an exact knowledge state' implies when to reach for it (to anchor/timestamp a result), but there is no explicit when/when-not statement or named alternative such as verify_claim_hash. Usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_evidence_bundleARead-onlyIdempotentInspect
Resolve up to 10 claims (by claim_ids or free-text query) into an evidence bundle: statements, verdicts, review dates, evidence pointers, canonical URLs, clinical-state pin. Optional locale (BCP-47) reports human-localized module availability; canonical identifiers are language-independent. Bulk export beyond 10 claims is the paid tier (x402) at /api/premium/evidence-pack.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | ||
| locale | No | ||
| claim_ids | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, closed-world behavior, so the safety profile is covered. The description adds non-annotation context: a hard 10-claim ceiling, the paid-tier escape hatch with its endpoint, and the distinction that locale affects only human-localized module availability while canonical identifiers stay language-independent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action and its output shape, then limits, then the locale nuance. Dense but every clause carries information; only the premium-tier sentence borders on tangential for an agent deciding whether to call this tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by enumerating the returned bundle fields (statements, verdicts, review dates, evidence pointers, canonical URLs, clinical-state pin). Combined with the 10-claim limit and premium-tier note, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry the load, and it does: it identifies claim_ids and query as alternative resolution inputs and explains locale as a BCP-47 code controlling human-localized module availability. It does not clarify whether query and claim_ids are mutually exclusive or how they interact, leaving one gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (resolve) and resource (claims into an evidence bundle), and enumerates the bundle contents (statements, verdicts, review dates, evidence pointers, canonical URLs, clinical-state pin). This clearly distinguishes it from siblings like get_claims and get_evidence_log, which do not bundle verdicts and review dates together.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the two accepted input modes (claim_ids or free-text query) and names the boundary condition for the alternative: bulk export beyond 10 claims goes to the paid tier at /api/premium/evidence-pack. It does not explicitly say when to prefer this over get_claims or get_evidence_log, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_evidence_logBRead-onlyIdempotentInspect
Dated evidence-review log for all modules.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and a closed world, so the safety profile is fully covered without the description. The description only adds the 'all modules' scope and 'dated' framing, which is modest extra context for a read tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence that front-loads the resource and scope with no filler. It is efficient, though the fragmentary phrasing leaves the operation implicit rather than stated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read tool with no output schema and annotations covering safety, the description is just barely adequate: it does not say what an entry contains, how results are ordered, or how this log relates to get_evidence_bundle or the module-scoped tools. Nothing is wrong, but it is thin.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; the schema correctly declares an empty object with additionalProperties false. Baseline 4 applies since parameter semantics are a non-issue here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource ('evidence-review log') and its scope ('for all modules'), and 'Dated' hints at ordering. However, it is a bare noun phrase with no verb, so the action (retrieve/list) is only implied, and it does nothing to separate it from near-siblings like get_evidence_bundle or get_claims.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no mention of when this log is preferable to get_evidence_bundle, and no prerequisites. The scope qualifier 'for all modules' weakly implies it is the global variant versus per-module siblings, but that must be inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_gi_verificationBRead-onlyIdempotentInspect
Get one canonical verification record (statement, verdict, review date, evidence module, machine URLs) by its slug or claim id.
| Name | Required | Description | Default |
|---|---|---|---|
| slug | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint, idempotentHint, destructiveHint=false), lowering the bar. The description adds that the record is 'canonical' and lists what it contains (statement, verdict, review date, evidence module, machine URLs), which is useful given no output schema, but says nothing about not-found behavior or what makes a record canonical.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the identifier-first construction and a parenthetical field list; nothing is wasted and the key action leads.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter lookup, the description compensates for the missing output schema by enumerating the returned record's components, and annotations cover the safety semantics. Only error/absence behavior and the slug-vs-claim-id ambiguity remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With one parameter and 0% schema description coverage, the description must carry the load. It clarifies the slug may also be a claim id, which adds meaning, but this conflicts slightly with the schema's slug-only pattern '^[A-Za-z0-9-]{3,110}$'. No format or validation guidance beyond that.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Get one canonical verification record') and enumerates the record's fields, which is concrete. The phrase 'one ... record ... by its slug' implicitly distinguishes it from search_gi_verifications' multi-result retrieval, though no sibling is named outright.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'by its slug or claim id' hints at identifier-based lookup, but there is no explicit when-to-use guidance, no conditions, and no named alternatives such as search_gi_verifications or resolve_gi_claim. The agent must infer the routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_modulesARead-onlyIdempotentInspect
List every clinician teaching module and patient chapter: title, canonical URL, evidence-review date.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and a closed world, so the safety profile is covered. The description adds the shape of the result (title, canonical URL, evidence-review date), which is genuinely useful, but says nothing about result volume, ordering, or pagination for a full-corpus listing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with zero filler, front-loaded with the verb and scope, and the returned fields appended as a compact list. Nothing is redundant or padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no input parameters and no output schema, the description carries the burden of describing what comes back, and it does list the three returned fields. It is nearly complete; only guidance on result size/ordering/pagination for a full-corpus listing is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so by rule the baseline is 4. The description correctly adds no parameter detail because there is none to add, and nothing in the schema is left unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (clinician teaching modules and patient chapters) and even enumerates the returned fields. The word 'every' signals an exhaustive enumeration, which implicitly separates it from read_module and search_teaching, but no sibling is named explicitly, keeping it just below the top score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: an agent can infer this is the tool for dumping the full catalog rather than fetching one module or searching. There is no explicit when-to-use statement, no exclusions, and no mention of search_teaching/read_module as alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
query_relationsARead-onlyIdempotentInspect
Query the evidence-backed clinical relationship graph (subject–predicate–object triples over diseases, medications, tests, findings, guidelines). Every relation cites governed claim IDs and/or evidence registry IDs. Filter by node text (subject/object), predicate (treated_by, evaluated_by, contraindicated_by, increases_risk_of, reduces_risk_of, causes, associated_with, suggests, argues_against, complicated_by, recommended_by, supported_by, conflicts_with, supersedes, requires_context_of), kind, or a claim_id. Also GET /api/tools/relations.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | filter subject_kind or object_kind (disease, medication, test, finding, guideline, risk_factor, mechanism, intervention) | |
| node | No | substring match against subject_id/object_id | |
| limit | No | ||
| claim_id | No | ||
| predicate | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds genuine non-annotation context: every returned relation cites governed claim IDs and/or evidence registry IDs, which tells the agent about provenance guarantees and how to trace results back to governance records.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences are front-loaded with the purpose and data model, then the filter options. The trailing 'Also GET /api/tools/relations' is a REST endpoint aside that adds little for an agent already invoking the tool, a minor bit of waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and five optional parameters, the description carries most of the burden and does cover the graph model, provenance citations, and filterable dimensions. It is still silent on result shape, ordering, and how the limit interacts with pagination, leaving small gaps for a query tool of this breadth.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 40%, so the description must compensate. It does so well for 'predicate' by enumerating all fifteen allowed values (treated_by, evaluated_by, contraindicated_by, etc.), which the schema leaves as a bare string with no enum. However, 'limit' (max 25) and 'claim_id' receive no explanation in the description, and 'node' is only loosely restated as subject/object text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource – 'Query the evidence-backed clinical relationship graph' – and immediately defines the data model as subject–predicate–object triples over diseases, medications, tests, findings, and guidelines. This clearly distinguishes it from siblings like get_claims or search_gi_verifications without needing to open a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description lists the filter axes (node text, predicate, kind, claim_id), which implies how to use the tool, but never states when to choose this over get_claims or get_evidence_bundle, which also deal with claims and evidence. Usage is inferable from the filter set rather than explicitly guided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_moduleARead-onlyIdempotentInspect
Read ONE module's full teaching text (markdown) with required citation. Not for bulk corpus assembly.
| Name | Required | Description | Default |
|---|---|---|---|
| slug | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive), so the bar is lower; the description adds real value by disclosing the return content (full markdown teaching text) and an output-side obligation (required citation), which annotations do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero waste, with the scope constraint ('ONE module') front-loaded before the exclusion clause.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description carries return-value burden and does so reasonably (markdown teaching text, citation required). Minor gaps remain: what the required citation looks like and error behavior for an unmatched slug.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% for the single slug parameter, but the description conveys its semantic role by emphasizing 'ONE module' (identifying a single record) and implicitly ruling out bulk selection. It adds no format or slug-value guidance beyond the schema's pattern.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Read) and resource (ONE module's full teaching text in markdown) and immediately scopes it against bulk usage. An agent can distinguish it from list_modules and search_teaching without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The exclusion 'Not for bulk corpus assembly' implies bulk work belongs elsewhere, but no sibling (e.g., search_teaching, list_modules) is named as the alternative, so routing remains partly inferential.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resolve_gi_claimARead-onlyIdempotentInspect
Resolve one supplied proposition against governed claims. Only preserved canonical wording or its full explicit negation is certified; lexical candidates and changed qualifiers remain unresolved. Does not search external evidence.
| Name | Required | Description | Default |
|---|---|---|---|
| claim | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish a safe, idempotent, closed-world read, so safety needs no restating. The description adds genuinely non-obvious behavior: only exact canonical wording or its full explicit negation is certified, while lexical candidates and changed qualifiers are left unresolved — this outcome rule is valuable context an agent cannot get from the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences: purpose first, then the certification rule, then the scope exclusion. No filler, no restatement of the tool name, and each sentence contributes a distinct constraint.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, read-only, idempotent tool with no output schema, the description covers what it does, which inputs resolve, and what it does not consult. The main remaining gap is the lack of any contrast with the sibling check_gi_claim, which an agent must resolve on its own.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single required 'claim' parameter has 0% schema description coverage, so the description must compensate. It adds that the input is a 'proposition' and that the preserved canonical wording form is what gets certified, but it never states the expected format, length, or normalization requirements of the claim string.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and scope: 'Resolve one supplied proposition against governed claims,' which tells the agent it is a resolution/certification operation over an existing governed corpus. It is somewhat jargon-heavy ('governed claims', 'certified') and never names or contrasts with the nearby sibling check_gi_claim, so sibling differentiation is left to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The line 'Does not search external evidence' is an explicit when-not boundary, implying this is a closed-corpus lookup rather than an evidence search. However, it names no alternative tool and gives no positive trigger for choosing resolve_gi_claim over check_gi_claim or get_gi_verification, so usage remains largely implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
score_clinical_instrumentARead-onlyIdempotentInspect
Score a governed clinical instrument from structured responses: IBS-SSS (ibs-sss, 5 slider items 0-100 + days×10, total 0-500) and GERD-Q (gerd-q, 6 choice items, total 0-18, ≥8 consistent with GERD). Returns total, percent of max, band label + interpretation, and the instrument citation. Every item is required; invalid/missing responses are rejected without defaulting or clamping. Educational scoring — not a diagnosis. Also POST /api/tools/score-instrument.
| Name | Required | Description | Default |
|---|---|---|---|
| responses | Yes | item_id → numeric value (see instrument spec at /tools/instruments.json) | |
| instrument | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so safety is covered. The description goes beyond them by disclosing strict validation behavior ('Every item is required; invalid/missing responses are rejected without defaulting or clamping') and the return shape (total, percent of max, band label + interpretation, citation) despite no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded: instrument specs, then return values, then validation rule, then the educational caveat. The trailing 'Also POST /api/tools/score-instrument' is redundant endpoint trivia that adds little for an agent, a minor structural blemish.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description supplies the return fields itself, covers both enum instruments with their scoring scales, and states the strict-rejection behavior. For a two-param, nested-object scoring tool this is complete enough to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50% (the 'responses' param), but the description compensates with per-instrument semantics: IBS-SSS uses 5 slider items 0-100 plus days×10 (total 0-500) and GERD-Q uses 6 choice items (total 0-18, ≥8 consistent with GERD). It also points to /tools/instruments.json for the item spec, though exact item_ids are left external.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Score) plus the exact resources (IBS-SSS, GERD-Q) and their scoring ranges, so the agent knows precisely what the tool produces. None of the sibling tools (claim verification, corpus/module reads, relation queries) overlap with clinical instrument scoring, making this easily distinguishable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for use: it scores the two named governed instruments and adds an explicit boundary ('Educational scoring — not a diagnosis'). However, it never states when NOT to use it or names an alternative for, e.g., free-form symptom assessment, so guidance is contextual but not fully routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_gi_verificationsBRead-onlyIdempotentInspect
Search canonical GI fact-check verifications by free text; returns matching governed claims with canonical /verify/ URLs.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world), so the bar is lower. The description still adds real value by disclosing what comes back — matching 'governed' claims carrying canonical /verify/ URLs — which matters because there is no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with the action front-loaded and no filler. The semicolon-joined return-value clause earns its place, though it slightly compresses two distinct ideas into one line.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter search tool with no output schema, explaining the result shape was the right call, but pagination/limit behavior, what makes a claim 'governed' or 'canonical', and ranking semantics are all left unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning. 'By free text' partially explains the required query, but the limit parameter (max 10) is never mentioned, leaving half the parameters undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search) and resource (canonical GI fact-check verifications) plus the matching mode (free text). This clearly separates it from a single-record fetch like the sibling get_gi_verification, though the description never names that sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'by free text' implies when to use this versus an exact-ID lookup, but no alternative tool is named and no condition or exclusion is given. Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_teachingARead-onlyIdempotentInspect
Search clinician teaching modules AND patient chapters by free-text query (title, description, and body text); returns matching modules with snippets.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint=false, and destructiveHint=false, so the safety profile is covered. The description usefully adds that matches are returned with snippets, but says nothing about pagination, result ordering, or why limit caps at 10.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the action and resource, then adds scope and return shape. No filler, nothing to trim.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and thin annotations, the description does enough to call the tool correctly: it names the corpus, the search fields, and the snippet return. Only the limit parameter and result-size behavior are left unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the load. It does clarify that the query string is matched against title, description, and body text, which is real added value, but the limit parameter (and its maximum of 10) is never explained or given a default.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (search) plus resource (clinician teaching modules AND patient chapters) and the fields searched (title, description, body text). It implicitly separates itself from list_modules/read_module by being query-driven, but never names a sibling to make the distinction explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: a free-text query is required, so the agent can infer this is for keyword lookup rather than browsing. No explicit when-to-use vs. list_modules, no exclusions, and no note on what happens when a query matches nothing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_claim_hashARead-onlyIdempotentInspect
Verify one clinical claim against the published corpus root via its Merkle inclusion proof (optionally against a bundle_hash you hold). Returns the claim content only when it is physician-activated or clinician-adjudicated.
| Name | Required | Description | Default |
|---|---|---|---|
| claim_id | Yes | ||
| bundle_hash | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds meaningful behavior beyond that: the optional alternative trust root (bundle_hash you hold) and the conditional disclosure rule that claim content is returned only for physician-activated or clinician-adjudicated claims.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, front-loaded with the core action and mechanism, with the return-condition caveat last. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description should explain return semantics. It discloses the conditional content-return rule but not what a verification result looks like when the proof fails or the claim is not activated, leaving an agent to guess the failure response shape.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the load. It clarifies that bundle_hash is optional and represents an alternative root of trust the caller holds, which is genuinely useful, but it says nothing about claim_id semantics or format beyond the schema pattern.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (verify), a specific resource (one clinical claim), and the mechanism (Merkle inclusion proof against the published corpus root). This is clearly distinguishable from siblings like get_claims, get_evidence_bundle, or check_gi_claim.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the usage context (verifying a claim's inclusion against the corpus root, optionally against a caller-held bundle_hash), but it never says when to prefer this over check_gi_claim or resolve_gi_claim, nor does it state preconditions or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
15 tool updates
- First observed
check_gi_claim - First observed
classify_hrm_metrics - First observed
get_claims - First observed
get_corpus_state - First observed
get_evidence_bundle - First observed
get_evidence_log - First observed
get_gi_verification - First observed
list_modules - First observed
query_relations - First observed
read_module - First observed
resolve_gi_claim - First observed
score_clinical_instrument - First observed
search_gi_verifications - First observed
search_teaching - First observed
verify_claim_hash
Related MCP Connectors
Free medical evidence tools: graded CliniAtlas answers, PubMed metadata and EU SmPC links.
- Eureka EHROAuthmd.eureka
The Eureka EHR as MCP tools: patients, schedule, notes, messages, billing, and orders.
FDA and CMS evidence for AI medical devices: 510(k), postmarket, reimbursement, and compliance.
Healthcare staffing simulator — ED, walk-in clinic, and appointment office DES tools.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceHealthcare AI Governance - MCP server providing AI-powered tools and automation by MEOK AI Labs3 npm47 PyPIMIT
- FlicenseNot gradedqualityBmaintenanceBridges Codex with MiMoCode as a coding agent for planning, implementation, and review via MCP tools.-
- AlicenseBqualityDmaintenanceMCP server for searching research grants across NSF (US), ERC (EU), and KRF/NRF (Korea) via a unified interface. NIH excluded—covered by existing connectors.39 npmMIT
- AlicenseBqualityDmaintenanceEnables medical evidence retrieval and analysis via AI-powered search and summarization tools.148 npmMIT
Glama MCP Gateway
Add one secure layer between your agents and this server.