ot-dossier-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@ot-dossier-mcpassemble a NOD2 IBD evidence dossier and validate it"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Open Targets evidence invariants
A reproducible target-dossier prototype for NOD2 and TNF in inflammatory bowel disease (IBD). It explores which scientific constraints can be enforced in code and which still require human judgment.
Built by Chris Lawrence, RN with AI coding assistance. Public Open Targets data only; no employer or client materials. Local research prototype, not a clinical recommendation system.
Review in five minutes
Read the evaluation findings.
Compare the original live NOD2 dossier with the corrected dossier and signed review.
Inspect the domain validator, thin MCP server, and four fault-injection tests.
Run the offline demo; no API key is required.
Related MCP server: opentargets-mcp
What this adds
Open Targets already has an official MCP server for access to its API. This project is a separate, task-specific server that replays frozen GraphQL responses. It does not call, replace, or claim superiority over the official MCP.
assemble_target_disease_evidence: creates a typed packet with disease scope, source-record provenance, and explicit missingness and retrieval limits.validate_dossier_references: checks packet-bound citations and selected structured assertions. It does not determine whether arbitrary prose is true.A bounded agent loop requires final validation and records tool calls, repairs, versions, and token use. A reusable skill guides synthesis. Markdown and JSON outputs support subsequent review.
For example, a Crohn disease record remains descendant evidence for selected IBD. Calling it direct IBD evidence in a structured assertion is rejected. Inferring what intervention to use from a LoF/risk label still requires scientific judgment.
Synthetic fault | Checked behavior |
Descendant evidence asserted as direct | Reject with |
Approval for an unrelated indication asserted for IBD | Reject with |
Empty safety data interpreted as a safe target | Preserve unknown state; reject structured safety assertion |
Nonexistent citation | Reject with |
Results and limits
The reviewed live NOD2 run completed with 2 generation requests, 2 MCP calls, 0 repairs, and 63,391 input / 1,341 output tokens. Structural validation passed, but human-assisted review identified three error-bearing claims: a mismatched source citation, an uncited named variant, and an incorrect numeric lower bound. A separate edited derivative was signed off by Chris Lawrence, RN on September 22, 2026, with partial factual coverage disclosed. It is not unassisted model success.
TNF passed scripted integration through the real MCP server. Live TNF scientific evaluation is not completed. Both targets use 300-row frozen evidence captures, not exhaustive or representative evidence samples. There is no held-out benchmark, official-MCP comparison, independent publication review, or specialist validation. See evaluation details and limitations.
Quick start — Python 3.12
From this repository directory:
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.lock
python -m pip install --no-deps --no-build-isolation -e .
python -m pytest -q
PYTHONPATH=src python -m ot_dossier.agent.loop --target NOD2 --out generated/demo-nod2-01Dependency installation needs network access once. Tests and this scripted
demo use frozen data and no model API. Expected demo result: two MCP calls and
zero repairs. Use a new output directory on each run; the agent refuses overwrite.
The loop launches its own stdio MCP server. Open generated/demo-nod2-01/dossier.md
and result.json afterward.
For deterministic source-observation drafts without the agent loop:
PYTHONPATH=src python -m ot_dossier.cli --target NOD2 --out generated
PYTHONPATH=src python -m ot_dossier.cli --target TNF --out generatedThe ordinary CLI overwrites its named draft outputs in the specified directory;
the agent's per-run directories are immutable by convention and protected from
overwrite by the host. Installed entry points are ot-dossier and
ot-dossier-mcp. See the beginner MCP walkthrough or
optional paid live-run guide. Keep API keys out of files and Git.
The supported runtime is an editable checkout: fixtures and skill files are loaded from the repository. Wheel-only deployment and remote hosting are not supported.
Architecture and reproducibility
Open Targets GraphQL → frozen cassettes + checksum manifest
↓
deterministic assembler
↓
typed EvidencePacket
↓
two stdio MCP tools ↔ bounded agent host
↓
required validation → JSON + Markdown
↓
attributed human/assistant reviewArchitecture explains the code boundaries. Packet IDs and record IDs are content-derived; provenance retains cassette names, response hashes, and JSON pointers. Frozen replay is reproducible, but a fresh API query need not return the same evidence. Verification records the checks; the pre-upload audit includes a clean installation and 84 passing tests.
Frozen data release 26.06 | NOD2 | TNF |
Captured evidence rows | 300 | 300 |
Upstream matching rows | 4,003 | 20,990 |
Direct / descendant rows | 31 / 269 | 64 / 236 |
Curated target safety records returned | 0 | 9 |
Clinical context is limited to drug-bearing rows in that evidence capture.
An empty result is not proof of safety or absence of drugs. Source-reported
APPROVAL stays with its exact indication; PHASE_4 is not converted to approval.
The source reference case contains five reviewed source facts and five prohibited inferences; its metadata distinguishes human source review from publication-level validation. The artifact index separates original outputs, scripted runs, review findings, and edited derivatives.
Scope and attribution
No UI, deployment, ranking, RAG/vector database, additional biomedical source, official-MCP comparator, or large benchmark. See scope decisions.
Code: Apache-2.0. Open Targets Platform data is marked CC0; retain upstream attribution and consult its licensing and citation guidance. Only queried fields are frozen; no paper full text is redistributed. Sources and NOTICE document attribution. No endorsement by Open Targets, Anthropic, or a data provider is implied.
Available Tools
2 toolsassemble_target_disease_evidenceC
Replay frozen evidence for NOD2/TNF and IBD. Returns scoped records and typed missingness.
| Name | Required | Description | Default |
|---|---|---|---|
| target | Yes | ||
| disease | No | IBD |
Output Schema
| Name | Required | Description |
|---|---|---|
| target | Yes | |
| records | Yes | |
| release | Yes | |
| endpoint | Yes | |
| sections | Yes | |
| packet_id | Yes | |
| snapshot_id | Yes | |
| retrieved_at | Yes | |
| disease_scope | Yes | |
| schema_version | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It usefully indicates that the tool replays frozen evidence and returns typed missingness, implying a deterministic read-oriented operation, but it does not disclose permissions, side effects, rate limits, or whether any state changes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two tightly written sentences with no filler. The action and scope are front-loaded, and the return behavior is stated immediately after, so every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained in depth. However, with no annotations and 0% schema description coverage, the description leaves key calling context missing: it does not explain parameter semantics, usage relative to the sibling, or the operational meaning of "scoped records" and "typed missingness."
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It offers concrete values (NOD2/TNF for target, IBD for disease), which adds some meaning beyond the bare schema, but it does not explicitly map these to the parameters or clarify whether they are examples, allowed values, or required formats.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names an action ("Replay") and a resource ("frozen evidence") scoped to NOD2/TNF and IBD, so the general intent is inferable. However, "replay frozen evidence" is jargon-heavy and does not clearly state that it assembles evidence for an arbitrary target/disease pair, nor does it distinguish this tool from validate_dossier_references.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool, when not to use it, or how it relates to the sibling validate_dossier_references. The description supplies context but no selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_dossier_referencesC
Check references and structured assertions against a previously assembled packet. Not entailment.
| Name | Required | Description | Default |
|---|---|---|---|
| dossier | Yes | ||
| packet_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| scope | No | |
| valid | Yes | |
| findings | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It clarifies only the negative scope ('Not entailment'); it does not say whether the tool is read-only or mutating, whether it can fail/repair the dossier, or what kind of result (pass/fail vs. diagnostics) it produces, beyond the existence of an output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences with zero filler; the scope-limiting clause is placed immediately after the action. It is arguably too compressed for a nested-schema tool, but every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool taking a complex nested Dossier with claims, record references, and section states (and a sibling assembler producing that input), the description is too thin. The output schema covers return values, but reference-matching behavior, validation failure semantics, and parameter meaning are left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the two parameters (packet_id, dossier with nested Claim/Datum structures) are entirely undocumented. The phrase 'previously assembled packet' hints that packet_id refers to prior output, but nothing explains the Dossier payload, the id/reference matching semantics, or the assertion enums.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Check) and resource (references and structured assertions) with a clear scoping phrase (against a previously assembled packet), and the 'Not entailment' clause narrows what kind of checking this is. It implicitly contrasts with the sibling assemble_target_disease_evidence (assembly vs. validation) but never names it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'against a previously assembled packet' implies this runs after assembly, and 'Not entailment' sets a scope boundary, so usage is implied rather than stated. No explicit when/when-not guidance or named alternative (e.g. the assembly sibling) is given, and no preconditions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.1.0- First observed
assemble_target_disease_evidence - First observed
validate_dossier_references
TDQS
Scored across 2 tools
The two tools have distinct verbs and roles: one assembles/replays evidence, the other validates references against an assembled packet. They are part of the same pipeline, which introduces mild sequencing ambiguity, but a careful reading of the descriptions clearly separates their purposes.
Both tools use snake_case with a leading verb (assemble_..., validate_...) and noun-phrase tails. The verb_noun pattern is consistent throughout, with no style mixing.
Only 2 tools is on the thin side, though the server's scope (a single frozen evidence dossier workflow) is narrow enough that this is borderline rather than clearly wrong. A few more operations would make it feel better scoped.
The surface covers an assemble-and-validate loop but appears locked to one hardcoded disease pair (NOD2/TNF and IBD) with no discovery, listing, or update operations, and no way to build or modify the underlying packet. Notable gaps remain for a general dossier workflow.
Related MCP Connectors
Authenticated public evidence search, verification, research jobs, exports, and webhooks.
Read-only game, setup, place, evidence and travel decision tools with explicit provenance.
Machine-readable entity discovery with provenance, trust and verified source evidence.
Auditable MCP server for PubMed, Europe PMC, ClinicalTrials.gov, and bioRxiv/medRxiv queries
Related MCP Servers
- FlicenseBqualityDmaintenanceUnofficial Model Context Protocol server for accessing Open Targets platform data for gene-drug-disease associations research.611-
- AlicenseBqualityAmaintenanceMCP server that exposes the Open Targets Platform GraphQL API as a set of tools for querying biomedical data such as targets, diseases, drugs, and genetic evidence.6871 PyPI19MIT
- AlicenseNot gradedqualityBmaintenanceEnables defining and verifying evidence contracts for claims in READMEs, releases, or product pages using constrained verifiers and generating hash-chained receipts and reports.12 npmMIT
- AlicenseNot gradedqualityAmaintenanceA local-first, deterministic, read-only MCP server that audits test suites for false-green tests, tautological assertions, and mock-contract drift, ensuring tests truly validate production code. It provides tools to detect test fidelity issues, verify mock drift, and synthesize strict mock contracts.1MIT