heor-agent-mcp
The HEORAgent MCP server is an AI-powered Health Economics and Outcomes Research (HEOR) platform that automates evidence generation, economic modeling, and HTA submission workflows.
Literature & Evidence
Literature search across 44 data sources (PubMed, ClinicalTrials.gov, NICE TAs, ICER, CADTH, PBAC, Embase, Cochrane, FDA, LATAM/APAC sources, etc.) with PRISMA-style audit trail
Abstract screening with PICO-based relevance scoring and study design classification
Risk of bias assessment using Cochrane RoB 2 (RCTs), ROBINS-I (observational), and AMSTAR-2 (systematic reviews) with GRADE domain summaries
Evidence network building, NMA feasibility assessment, and indirect treatment comparisons (Bucher, frequentist NMA with consistency checks)
Population-adjusted indirect comparisons (MAIC/STC) with feasibility assessment
Survival curve fitting for 5 parametric distributions per NICE DSU TSD 14
Economic Modeling
Cost-effectiveness modeling (Markov, PartSA, decision-tree) with PSA (up to 10,000 Monte Carlo iterations), OWSA tornado diagrams, CEAC, EVPI/EVPPI, and WTP assessment against NHS, US payer, and societal thresholds
Budget impact modeling (ISPOR-compliant, year-by-year output, treatment-displacement modeling)
Utility value set analysis with EQ-5D-3L/5L references and baseline-utility-aware ICER impact estimation
HTA & Regulatory
HTA dossier drafting for NICE, EMA, FDA, IQWiG, HAS, and EU JCA with structured GRADE tables and gap analysis
Pharmacovigilance classification into EMA GVP categories with ENCePP protocol templates, PASS/PAES classification, and RMP implications
ITC feasibility assessment evaluating the 3-assumption framework
Project & Knowledge Management
Persistent project workspaces with directory skeleton and metadata
Full-text knowledge search, file read/write across
raw/andwiki/trees, Obsidian-compatible markdown with wikilinksMulti-tool workflow orchestration (e.g., end-to-end MAIC pipelines, full dossier creation)
Link validation for HTTP citation URLs
Output & Integration
Output formats: Markdown, structured JSON, or DOCX
Full audit trail per tool call (source selection, queries, inclusions/exclusions, assumptions, warnings)
Available via Claude Desktop/Code, Claude.ai, any MCP-compatible host, and a ChatGPT Custom GPT adapter
Provides access to Embase and ScienceDirect databases (requires ELSEVIER_API_KEY) for comprehensive literature review across 41 data sources with audit trails.
Hosts source code repository for the HEORAgent MCP server at github.com/neptun2000/heor-agent-mcp for development and collaboration.
Searches academic literature through Google Scholar API (requires SERPAPI_KEY) as part of comprehensive literature review across 41 data sources with audit trails.
Runs the HEORAgent MCP server as a Node.js application with support for both stdio and HTTP transport modes.
Distributes the HEORAgent MCP server as an npm package for easy installation and version management.
Writes compiled evidence to project wiki in Obsidian-compatible markdown format with wikilinks for persistent knowledge management.
Searches 35M+ biomedical citations through NCBI E-utilities as part of comprehensive literature review across 41 data sources with PRISMA-style audit trails.
Hosts MCP server deployment for web UI backend, enabling tool execution without local setup when using the companion chat interface.
Hosts companion chat interface at web-michael-ns-projects.vercel.app for interacting with HEOR tools through Claude Opus with BYOK (Bring Your Own Key) functionality.
HEORAgent MCP Server
AI-powered Health Economics and Outcomes Research (HEOR) agent as a Model Context Protocol server.
Try it now → HEORAgent on ChatGPT (ChatGPT Plus / Team) · Web UI (Claude, BYOK) ·
npx heor-agent-mcpfor Claude Desktop / Claude Code
Automates literature review across 44 data sources, risk of bias assessment (RoB 2 / ROBINS-I / AMSTAR-2), EQ-5D value set impact estimation, state-of-the-art cost-effectiveness modelling, HTA dossier preparation for NICE / EMA / FDA / IQWiG / HAS / EU JCA, and a persistent project knowledge base — all callable as MCP tools from Claude.ai, Claude Code, and any MCP-compatible host.
Built for pharmaceutical, biotech, CRO, and medical affairs teams who need rigorous, auditable HEOR workflows without building infrastructure from scratch.
First 60 seconds
Verify your install works before wiring it into Claude / Cursor / Continue. Open two terminal tabs:
Tab 1 — start the server in HTTP mode:
MCP_HTTP_PORT=8080 npx heor-agent-mcp@latestYou should see:
HEORAgent MCP server running on HTTP port 8080Tab 2 — confirm it responds:
curl -s http://localhost:8080/healthExpected output:
{"status":"ok","server":"heor-agent-mcp","version":"1.10.2"}✅ If you see the JSON above, the npm package works on your machine. Any further issues are in your MCP client config (Claude Desktop / Cursor / Continue), not the server.
❌ If you see command not found, run node --version — you need Node ≥20. If you see a different error, file a quick issue at https://github.com/neptun2000/heor-agent-mcp/issues with the output.
Now stop Tab 1 (Ctrl+C) and pick your client below — you don't need the HTTP mode for the actual integration; Claude / Cursor / Continue all use stdio.
Related MCP server: FHIR MCP Server
Quick Start (per client)
Pick your MCP host:
Claude Code
claude mcp add heor-agent -- npx heor-agent-mcpThen restart Claude Code.
Claude Desktop / claude.ai Desktop
Edit your MCP config file (~/Library/Application Support/Claude/claude_desktop_config.json on macOS) and add:
{
"mcpServers": {
"heor-agent": {
"command": "npx",
"args": ["heor-agent-mcp"]
}
}
}Then restart Claude Desktop.
Cursor / Continue / Cline
Same config shape as Claude Desktop above; the file path differs by client:
Cursor:
Settings → MCP → Add new MCP serverContinue:
~/.continue/config.jsonunder themcpServerskeyCline:
Settings → MCP Servers → Edit MCP Settings
Hosted (no install)
ChatGPT (Plus / Team): HEORAgent on ChatGPT — type
/heorto use it; works on any conversation.Web UI (Claude, BYOK): web-michael-ns-projects.vercel.app — bring your Anthropic API key; runs the full v1.6.3 toolset.
Your first prompt
Once your MCP host is configured, paste any of these to verify end-to-end:
Run a literature search for semaglutide cost-effectiveness in T2D
using PubMed, NICE TAs, and ICER reports. Set runs=2.Run irb_review for an industry-funded interventional Phase 2 trial in
relapsed MM — multi-site US+EU, pseudonymized data, greater-than-minimal
risk. I need the review tier, GDPR/HIPAA DMP, SAE framework, and the
ready-to-paste cover letter.Run jca_pico_scope for osimertinib in EGFR-mutant 2L NSCLC across
DE/FR/IT/ES/NL. Then prepare an EU JCA dossier draft using the picos.The first prompt exercises literature_search + validate_links (free, no API keys needed). The second exercises irb_review (pure decision tree, instant). The third exercises jca_pico_scope → hta_dossier pipeline.
What's new
See CHANGELOG.md for full version history. Current: v1.23.0 (45 tools, 44 data sources).
v1.17.0–v1.23.0 — Living Evidence Intelligence (review → reimbursement)
A connected RWE + cross-deliverable layer (see docs/FEATURES.md for the full table):
RWE & real-world safety:
rwe.method_select(study-design selection),pv.comparative_safety(class-level FAERS-style AE ranking),evidence.triangulation(per-outcome RCT↔RWE concordance).Governed social listening:
rwe.social_listening_protocol+pv.social_listening_triage(GVP Module VI ICSR triage — no scraping).One source of truth:
evidence.claim_registry(author/auto-import a figure once),evidence.consistency_check(detect drift across dossier/publication/payer),publication.draft(reuse claims; CONSORT/STROBE/PRISMA/CHEERS + GPP2022/ICMJE).Living orchestration:
evidence.gap_analysis(iEGP),workflow.living_evidence(SLR → living KB → JCA/HTA runbook),hta.living_gvd(regenerate only the GVD sections whose figures changed).
v1.13.0 — AI Transparency Disclosure (ISPOR ELEVATE-GenAI aligned)
16 tools now accept an ai_disclosure_level parameter:
Value | Behaviour |
| No disclosure block appended |
| Model ID · tools called · data sources · date · human-review reminder |
| Standard block + ISPOR ELEVATE-GenAI full citation |
Default by tool tier: HTA/regulatory tools (hta_dossier, hta_workflow, jca_pico_scope, pv_classify, etc.) default to "submission"; analysis tools (risk_of_bias, cost_effectiveness_model, etc.) default to "standard". Pass ai_disclosure_level: "off" to suppress.
Environment-level default: set HEORAGENT_DISCLOSURE_LEVEL=off|standard|submission to override the built-in per-tool defaults globally.
Web UI persona defaults: payer and HTA-reviewer personas always use "submission"; analyst personas default to "standard" and switch to "off" for scratch / exploratory prompts.
v1.0.4 highlights (still in v1.6.3)
Pharmacovigilance + workflow orchestration:
pv_classifytool — classifies a planned study into its EMA pharmacovigilance regulatory category (PASS imposed/voluntary, PAES, RMP Annex 4, DUS, active surveillance registry, pregnancy registry, spontaneous reporting, ICH E2E plan). Returns the matching GVP module (V/VI/VIII/VIII Addendum I), ENCePP protocol template ID, RMP implications, FDA analogue, and submission obligations. Pure decision-tree per EMA GVP rev 4 + EU Regulation 1235/2010 Article 107a. <200ms response.hta_dossierPharmacovigilance Plan section — passpv_classificationfrompv_classifytohta_dossierand the dossier output now includes a PV Plan section between RoB and CEA. Without it, a one-line "PV plan not provided" note flags the gap so reviewers see what's missing.maic_workfloworchestrator (v1.0.6) — runs the full MAIC discovery+screening pipeline (ITC feasibility + parallel literature_search + screening + RoB + network) in one MCP call. Built for ChatGPT-5.3 surfaces where chaining 5+ tool calls in parallel is unreliable; works equally well from Claude.examplestool (v1.0.5) — pre-filled JSON inputs for heavy-schema tools (CEA, BIA, survival, MAIC, Bucher) plus amaic_workflow_recipemulti-step prompt template for ChatGPT users.CMS IRA awareness — when
pv_classifyis called with US jurisdiction, output explicitly notes that CMS IRA Medicare price-negotiation calculations exclude PV cost data — track those obligations in the regulatory budget, not the HEOR cost-effectiveness model.GRADE I²-based inconsistency, GRADE upgrading (Guyatt 2011), Bucher consistency check, EQ-5D 5L baseline-utility-aware impact (v1.0.4) — see CHANGELOG.md.
ChatGPT Custom GPT support (v1.0.4) — OpenAPI 3.1 adapter at
/api/openapilets you build a Custom GPT in 5 minutes. See ChatGPT Custom GPT below.Surface-tagged analytics (v1.0.4) — every
tool_callPostHog event carries asurfaceproperty (claude_anthropic_web/chatgpt_adapter/claude_desktop/direct_mcp).
See CHANGELOG.md for the full diff.
Tools (28)
Tool | Purpose |
| Search 44 data sources with a full PRISMA-style audit trail |
| PICO-based relevance scoring and study design classification |
| Cochrane RoB 2 / ROBINS-I / AMSTAR-2 with GRADE RoB domain summary |
| Build treatment comparison network and assess NMA feasibility |
| Bucher and frequentist NMA with automatic consistency check vs direct h2h evidence (NICE DSU TSD 18) |
| MAIC and STC for population-adjusted indirect comparisons |
| Fit 5 parametric distributions to KM data (NICE DSU TSD 14) |
| Assess the 3-assumption ITC framework and recommend Bucher / NMA / MAIC / STC / ML-NMR |
| Markov / PartSA / decision-tree CEA with PSA, OWSA, CEAC, EVPI, EVPPI; QALY + evLYG support |
| ISPOR-compliant BIA with year-by-year output and treatment-displacement modelling |
| Draft submissions for NICE, EMA, FDA, IQWiG, HAS, and EU JCA — GRADE table uses structured RoB when |
| EQ-5D-3L / 5L value-set reference + baseline-utility-aware Biz 2026 ICER impact estimator (UK 5L transition) |
| HTTP validation of citation URLs before presentation |
| Initialize a persistent project workspace |
| Full-text search across a project's raw/ and wiki/ trees |
| Read any file from a project's knowledge base |
| Write compiled evidence to the project wiki (Obsidian-compatible) |
literature_search
Searches across 44 sources in parallel. Every call returns a source selection table showing which of the 44 sources were used and why — essential for HTA audit trails.
Example call:
{
"query": "semaglutide cardiovascular outcomes type 2 diabetes",
"sources": ["pubmed", "clinicaltrials", "nice_ta", "cadth_reviews", "icer_reports"],
"max_results": 20,
"output_format": "text"
}cost_effectiveness_model
Multi-state Markov model (default) or Partitioned Survival Analysis (oncology), following ISPOR good practice and NICE reference case (3.5% discount rate, half-cycle correction). Includes:
PSA — 1,000–10,000 Monte Carlo iterations, probability cost-effective at WTP thresholds
OWSA — one-way sensitivity analysis with tornado summary
CEAC — cost-effectiveness acceptability curve
EVPI — expected value of perfect information
WTP assessment — verdict against NHS (£25–35K/QALY, updated April 2026), US payer ($100–150K), societal thresholds
Example call:
{
"intervention": "Semaglutide 1mg SC weekly",
"comparator": "Sitagliptin 100mg daily",
"indication": "Type 2 Diabetes Mellitus",
"time_horizon": "lifetime",
"perspective": "nhs",
"model_type": "markov",
"clinical_inputs": { "efficacy_delta": 0.5, "mortality_reduction": 0.15 },
"cost_inputs": { "drug_cost_annual": 3200, "comparator_cost_annual": 480 },
"utility_inputs": { "qaly_on_treatment": 0.82, "qaly_comparator": 0.76 },
"run_psa": true,
"output_format": "docx"
}hta_dossier_prep
Drafts submission-ready sections for six HTA frameworks with gap analysis:
Body | Country | Submission types |
NICE | UK | STA, MTA, early_access |
EMA | EU | STA, MTA |
FDA | US | STA, MTA |
IQWiG | Germany | STA, MTA |
HAS | France | STA, MTA |
JCA | EU (Reg. 2021/2282) | initial, renewal, variation (with PICOs) |
Accepts piped output from literature_search and cost_effectiveness_model.
risk_of_bias
Assesses risk of bias using the appropriate Cochrane instrument, auto-detected from study_type:
Study type | Instrument |
RCT | RoB 2 (5 domains: randomization, deviations, missing data, measurement, reporting) |
Observational | ROBINS-I (7 domains: confounding, selection, classification, deviations, missing data, measurement, reporting) |
Systematic review | AMSTAR-2 (16 items, critical vs non-critical) |
Returns a rob_results object you can pass directly to hta_dossier_prep — this replaces the heuristic RoB estimate in the GRADE table with structured domain judgments.
Example call:
{
"studies": [{ "id": "pmid_1", "study_type": "RCT", "title": "...", "abstract": "..." }],
"output_format": "json"
}Pipeline:
literature_search→screen_abstracts→risk_of_bias→hta_dossier_prep
Knowledge base tools
Projects live at ~/.heor-agent/projects/{project-id}/ with:
raw/literature/— auto-populated literature search resultsraw/models/— auto-populated model runsraw/dossiers/— auto-populated dossier draftsreports/— generated DOCX fileswiki/— manually curated, Obsidian-compatible markdown with[[wikilinks]]
Pass project: "project-id" to any tool and results are saved automatically.
Examples
Copy-paste prompts to try in Claude Code, Claude Desktop, or the web UI.
Single-tool examples
Literature search
Search the literature for tirzepatide cardiovascular outcomes in type 2 diabetes. Use PubMed, ClinicalTrials.gov, and NICE TAs.
Survival curve fitting
Fit survival curves to this OS data from KEYNOTE-189: time 0 survival 1.0, time 6 survival 0.88, time 12 survival 0.72, time 18 survival 0.60, time 24 survival 0.51, time 36 survival 0.38. Use months.
Budget impact
Estimate the 5-year NHS budget impact of semaglutide for obesity. 200,000 eligible patients, drug cost £1,200/year, comparator (orlistat) £250/year, uptake 15% year 1 to 40% year 5.
Cost-effectiveness model
Build a CE model for semaglutide vs sitagliptin in T2D, NHS perspective, lifetime horizon, with PSA.
Indirect comparison (Bucher)
I have two trials: SUSTAIN-1 showed semaglutide vs placebo HR 0.74 (0.58-0.95) for HbA1c, and AWARD-5 showed dulaglutide vs placebo HR 0.78 (0.65-0.93). Run a Bucher indirect comparison between semaglutide and dulaglutide.
MAIC (population-adjusted comparison)
Run a MAIC between SUSTAIN-7 (N=300, semaglutide vs placebo, HR 0.74, CI 0.58-0.95, age 56±10, BMI 33±5) and AWARD-11 (N=600, dulaglutide vs placebo, HR 0.78, CI 0.65-0.93, age 58±9, BMI 35±6). Adjust for age and BMI.
Multi-tool workflows
Abstract screening workflow
Search PubMed for pembrolizumab in NSCLC, then screen the results with population adults with NSCLC, intervention pembrolizumab, comparator chemotherapy, outcomes overall survival and PFS.
Evidence network + NMA feasibility
Search for GLP-1 receptor agonists in T2D using PubMed, build an evidence network from the results, and assess NMA feasibility.
CE model with scenarios
Build a CE model for dapagliflozin vs placebo in heart failure, NHS perspective, lifetime horizon, with PSA. Add scenarios: "20% price reduction" with drug cost 400, "10-year horizon" with time_horizon 10yr.
End-to-end HTA workflow
Full dossier preparation
Create a project for semaglutide in obesity targeting NICE and ICER. Search literature for evidence, screen the results for adults with obesity comparing semaglutide to placebo for weight loss outcomes, assess risk of bias on the screened studies, then draft a NICE STA dossier using the screened results and rob_results.
This single prompt exercises: project_create → literature_search → screen_abstracts → risk_of_bias → hta_dossier_prep (GRADE RoB from structured assessment).
Data Sources
44 sources across 10 categories. Every literature_search call includes a source selection table showing used/not-used status and reason for each.
PubMed — 35M+ biomedical citations (NCBI E-utilities)
ClinicalTrials.gov — NIH/NLM trial registry (CT.gov v2 API)
bioRxiv / medRxiv — Life sciences and medical preprints
ChEMBL — Drug bioactivity, mechanisms, ADMET (EMBL-EBI)
Wiley Online Library — Pharmacoeconomics, Health Economics, Journal of Medical Economics, Value in Health (CrossRef, ~77% abstract coverage, no key required)
WHO GHO — WHO Global Health Observatory
World Bank — Demographics, macroeconomics, health expenditure
OECD Health — OECD health statistics (expenditure, workforce, outcomes)
IHME GBD — Global Burden of Disease (DALYs, prevalence across 204 countries)
All of Us — NIH precision medicine cohort
FDA Orange Book — Drug approvals and therapeutic equivalence
FDA Purple Book — Licensed biologics and biosimilars
NICE TAs (UK) · CADTH (Canada) · ICER (US) · PBAC (Australia)
G-BA AMNOG (Germany) · IQWiG (Germany) · HAS (France)
AIFA (Italy) · TLV (Sweden) · INESSS (Quebec, Canada)
CMS NADAC (US drug acquisition costs)
PSSRU (UK unit costs) · NHS National Cost Collection · BNF (UK drug pricing)
PBS Schedule (Australia)
DATASUS · CONITEC · ANVISA (Brazil)
PAHO (Pan American regional) · IETS (Colombia) · FONASA (Chile)
HITAP (Thailand)
Source | Env variable |
Embase |
|
ScienceDirect |
|
Cochrane Library |
|
Citeline |
|
Pharmapendium |
|
Cortellis |
|
Google Scholar |
|
ISPOR — HEOR methodology and conference abstracts
OHE (Office of Health Economics) — EQ-5D value set research and HEOR methodology
EuroQol Group — EQ-5D instruments, value sets, and registry
Output Formats
All tools support output_format:
text(default) — Markdown with formatted tables and headingsjson— Structured objects for downstream toolsdocx— Microsoft Word document, saved to disk, path returned in response
DOCX files are saved to ~/.heor-agent/projects/{project}/reports/ (when a project is set) or ~/.heor-agent/reports/ (global). The tool response contains the absolute path — ready to attach to submissions or share with stakeholders.
Audit Trail
Every tool call returns a full audit record:
Source selection table — all 44 sources with used/not-used and reason
Sources queried — queries sent, response counts, status, latency
Inclusions / exclusions — counts with reasons
Methodology — PRISMA-style for literature, ISPOR/NICE for economics
Assumptions — every assumption logged with justification
Warnings — data quality flags, missing API keys, failed sources
Suitable for inclusion in HTA submission appendices.
Configuration
# Optional — enterprise data sources
ELSEVIER_API_KEY=... # Embase + ScienceDirect
COCHRANE_API_KEY=... # Cochrane Library
CITELINE_API_KEY=... # Citeline
PHARMAPENDIUM_API_KEY=... # Pharmapendium
CORTELLIS_API_KEY=... # Cortellis
SERPAPI_KEY=... # Google Scholar
# Optional — knowledge base location
HEOR_KB_ROOT=~/.heor-agent # Default
# Optional — localhost proxy for enterprise APIs behind corporate VPN
HEOR_PROXY_URL=http://localhost:8787
# Optional — hosted tier (future)
HEOR_API_KEY=...Web UI
A companion chat interface is available at:
https://web-michael-ns-projects.vercel.app
Chat with Claude Sonnet 4.6 + all 22 HEOR tools
BYOK (Bring Your Own Key) — paste your Anthropic API key in the settings; it stays in your browser's localStorage and is never stored on our servers
Markdown rendering with styled tables, tool call cards with live progress timers, and theme-aware mermaid network diagrams
12 example prompts covering literature search, CEA, BIA, NMA, ITC feasibility, RoB, EQ-5D 5L, EU JCA dossiers
Per-request MCP sessions (no cross-user session bleed)
The web UI calls the hosted MCP server on Railway for tool execution. No setup required — just add your API key and start querying.
Self-hosting the web UI
cd web
npm install
echo "ANTHROPIC_API_KEY=sk-ant-..." > .env.local # optional server-side fallback
npm run dev -- -p 3456Set MCP_SERVER_URL to point to your own MCP server instance (default: the public Railway deployment).
ChatGPT Custom GPT
🟢 Live: HEORAgent on ChatGPT →
Open in ChatGPT (Plus / Team / Enterprise account required), pick a conversation starter, and you're querying 44 HEOR data sources.
HEORAgent is also available as a ChatGPT Custom GPT — useful when you (or your team) prefer the ChatGPT interface or have a ChatGPT Plus/Team account but no Anthropic API access.
Behind the scenes, the web tier exposes an OpenAPI 3.1 adapter at /api/openapi, with one POST endpoint per tool at /api/v1/{tool_name}. ChatGPT speaks this contract natively.
What's different from the Anthropic surface
Web UI / MCP / Claude Desktop | ChatGPT Custom GPT | |
Streaming | yes (SSE) | no (45s single response) |
| up to 10,000 | capped to 1,000 (CEA) / 500 (BIA) |
| 1–5 | capped to 1 |
| up to 100 | capped to 30 |
Auth model | BYOK Anthropic | optional |
Surface label in PostHog |
|
|
The caps exist because ChatGPT Actions hard-fail at the 45-second response timeout. PSA, multi-run literature search, and full max_results would routinely exceed it. The web UI and MCP clients are unaffected.
Build a Custom GPT (ChatGPT Plus / Team required)
Visit chatgpt.com/gpts/editor and click Create.
Configure tab — fill in name (e.g., "HEORAgent"), description, and conversation starters. Paste the system prompt from
web/lib/claude.ts(or write your own — the tool descriptions are self-documenting).Actions → Create new action → Import from URL → paste:
https://web-michael-ns-projects.vercel.app/api/openapiChatGPT auto-imports all 17 endpoints with their schemas.
Authentication — choose None for the open public endpoint, or API Key with the
CHATGPT_ADAPTER_TOKENvalue if you've configured one (recommended for prod).Privacy policy URL — required by GPT Store. Use the web UI's privacy URL or your own.
Test in the playground (right pane), then Publish → "Anyone with the link" or "GPT Store".
Securing the adapter for production
By default the /api/v1/* endpoint is open. Two layers of protection are recommended for any public-facing GPT:
# 1. Token-gate the endpoint
cd web
vercel env add CHATGPT_ADAPTER_TOKEN production # generate a long random token
# Configure the same token in your Custom GPT under Authentication → API Key
# 2. Built-in rate limit
# 60 req/min per IP is enforced automatically (lib/rateLimit.ts).
# For multi-region/high-traffic prod, swap in @upstash/ratelimit + Vercel KV.Sample call (manual, no GPT needed)
curl -X POST https://web-michael-ns-projects.vercel.app/api/v1/utility_value_set \
-H "Content-Type: application/json" \
-d '{
"action": "estimate_impact",
"indication_type": "non_cancer_qol_only",
"baseline_utility": 0.85,
"base_icer": 30000
}'Returns the Biz 2026 baseline-utility-adjusted ICER projection (the new EQ-5D 5L impact estimator).
HTTP Transport
The server supports both stdio (default, for local MCP clients) and Streamable HTTP (for hosted deployment).
# Stdio mode (default — for Claude Code, Claude Desktop)
npx heor-agent-mcp
# HTTP mode — for hosted deployment, Smithery, web UI backend
npx heor-agent-mcp --http # port 8787
MCP_HTTP_PORT=3000 npx heor-agent-mcp # custom portHTTP endpoints:
POST/GET/DELETE /mcp— MCP Streamable HTTP protocolGET /health— health checkGET /.well-known/mcp/server-card.json— Smithery discovery
Development
git clone https://github.com/neptun2000/heor-agent-mcp
cd heor-agent-mcp
npm install
npm test # 401 tests across 84 suites
npm run build # Compile TypeScript to dist/
npm run dev # Run with tsx (no build step)Requires: Node.js ≥ 20.
Architecture
┌────────────────────────────────────────────┐
│ MCP Host (Claude.ai / Claude Code / etc.) │
└────────────────┬───────────────────────────┘
│ stdio
┌────────────────▼──────────────────────────┐
│ heor-agent-mcp server │
│ ┌──────────────────────────────────────┐ │
│ │ 17 MCP tools (Zod-validated) │ │
│ ├──────────────────────────────────────┤ │
│ │ DirectProvider (default) │ │
│ │ ├─ 44 source fetchers │ │
│ │ ├─ Audit builder + PRISMA trail │ │
│ │ ├─ Markov / PartSA economic models │ │
│ │ ├─ Markdown + DOCX formatters │ │
│ │ └─ Knowledge base (YAML + MD) │ │
│ └──────────────────────────────────────┘ │
└───────────────────────────────────────────┘
│
┌────────────┴─────────────┐
▼ ▼
┌────────────┐ ┌──────────────────┐
│ ~/.heor- │ │ External APIs │
│ agent/ │ │ (PubMed, NICE, │
│ projects/ │ │ ICER, CADTH, …) │
└────────────┘ └──────────────────┘License
MIT — see LICENSE.
Trust & Transparency
HEORAgent is a research and analysis tool — not a clinical decision-support system. It is not classified as high-risk under the EU AI Act because it does not drive individual diagnosis, treatment, or monitoring; it falls under limited-risk transparency obligations only. Every output is intended for review by a qualified HEOR/HTA/PV professional before any action is taken.
EU AI Pact signatory — committed to AI governance, high-risk system mapping, and AI literacy promotion (voluntary commitments per the European Commission, ahead of the AI Act's August 2026 deadline).
PRISMA-style audit trail on every tool call (sources queried, succeeded, failed, assumptions applied).
AI commentary explicitly labelled — domain claims (ICERs, trial results, regulatory decisions) come exclusively from tool outputs, never from training-data recall.
Methodology cited inline — ISPOR, NICE DSU TSDs, NICE PMG36, Cochrane Handbook, GRADE, EMA GVP, Cope 2014, Phillippo 2016, Biz 2026.
Full statement: /ai-transparency — risk classification, human oversight model, methodological references, and reporting channel.
Disclaimer
All outputs are preliminary and for research orientation only. Results require validation by a qualified health economist before use in any HTA submission, payer negotiation, regulatory filing, or clinical decision. This tool does not replace professional HEOR expertise.
Distribution
Channel | How to use | Who pays |
npm |
| User's Claude subscription |
Smithery | User's Claude subscription | |
Web UI | User's own Anthropic API key (BYOK) | |
Hosted MCP |
| Free (tool execution only) |
Links
Available Tools
7 toolscost_effectiveness_modelA
Build a cost-utility analysis (ICER, QALY, PSA, sensitivity analysis) for a drug vs comparator. Follows ISPOR good practice guidelines and NICE reference case. Includes probabilistic sensitivity analysis (PSA), one-way sensitivity, and cost-effectiveness acceptability curve (CEAC).
| Name | Required | Description | Default |
|---|---|---|---|
| intervention | Yes | Drug or treatment name | |
| comparator | Yes | Comparator (standard of care) | |
| indication | Yes | Disease or condition | |
| time_horizon | Yes | Modelling horizon: 'lifetime', '5yr', '10yr', or years as number | |
| perspective | Yes | ||
| model_type | No | Model type. Default: markov. Use 'partsa' for oncology. | |
| clinical_inputs | Yes | ||
| cost_inputs | Yes | ||
| utility_inputs | No | ||
| run_psa | No | Run probabilistic sensitivity analysis (default: true) | |
| psa_iterations | No | PSA iterations (default: 1000, max: 10000) | |
| output_format | No | ||
| project | No | Project ID for knowledge base persistence. When set, model run is saved to ~/.heor-agent/projects/{project}/raw/models/ |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden but lacks critical behavioral details. It mentions analysis types but doesn't disclose whether this is a read-only or write operation, what permissions are needed, whether it's computationally intensive, or what happens to the output (e.g., where results are stored). The mention of 'project' parameter saving to a directory hints at persistence but isn't fully explained.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured in two sentences: the first states the core purpose and scope, the second enumerates specific analysis components. Every phrase adds value without redundancy, making it front-loaded and easy to parse despite the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with 13 parameters, nested objects, no output schema, and no annotations, the description provides good high-level context but leaves gaps. It explains the analytical approach but doesn't cover behavioral aspects like computational requirements, error handling, or output details. The absence of annotations increases the burden on the description, which it partially meets but not fully.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds significant value beyond the 62% schema coverage by explaining the analytical framework (cost-utility analysis with specific components like ICER, QALY, PSA) and referencing guidelines. While it doesn't detail individual parameters, it provides crucial context about what the tool fundamentally does that the schema alone doesn't convey, though some parameter relationships remain implicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Build a cost-utility analysis') and the comprehensive scope of what it produces (ICER, QALY, PSA, sensitivity analysis, CEAC). It distinguishes itself from siblings by focusing on economic modeling rather than knowledge management or dossier preparation, with explicit mention of following ISPOR and NICE guidelines.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage through its reference to specific guidelines (ISPOR, NICE) and analysis types, suggesting it's for health economic evaluations. However, it doesn't explicitly state when to use this tool versus alternatives like 'hta_dossier_prep' or provide clear exclusions or prerequisites for usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hta_dossier_prepA
Structure evidence into HTA body-specific submission format (NICE STA, EMA, FDA, IQWiG, HAS, EU JCA). Produces draft sections with gap analysis. Accepts output from literature_search and cost_effectiveness_model.
| Name | Required | Description | Default |
|---|---|---|---|
| hta_body | Yes | HTA body. Use 'jca' for EU Joint Clinical Assessment (EUHTA Reg. 2021/2282). | |
| submission_type | Yes | Submission type. Use 'initial'/'renewal'/'variation' for JCA. | |
| drug_name | Yes | ||
| indication | Yes | ||
| evidence_summary | No | Text summary or JSON array from literature_search output | |
| model_results | No | JSON output from cost_effectiveness_model | |
| picos | No | JCA: list of PICOs from the scoping decision. If omitted, a default PICO is generated. | |
| output_format | No | ||
| project | No | Project ID for knowledge base persistence. When set, dossier draft is saved to ~/.heor-agent/projects/{project}/raw/dossiers/ |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It describes the tool's core function (structuring evidence into submission formats) and output (draft sections with gap analysis), but doesn't mention important behavioral aspects like whether this is a read-only or write operation, what permissions might be needed, how long processing takes, or error handling. The description adds value but leaves significant behavioral questions unanswered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is perfectly concise with two sentences that each earn their place. The first sentence establishes the core function and scope, while the second specifies input requirements and integration points. There's zero wasted language, and the most important information (what the tool does) comes first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with 9 parameters, no annotations, and no output schema, the description is adequate but has clear gaps. It explains the tool's purpose and input relationships well, but doesn't describe the output format in detail (beyond 'draft sections with gap analysis'), doesn't mention error conditions or limitations, and leaves behavioral aspects unspecified. Given the complexity, more completeness would be expected.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 67% schema description coverage, the description compensates well by providing crucial context about parameter relationships and tool integration. The statement 'accepts output from literature_search and cost_effectiveness_model' clarifies the semantics of evidence_summary and model_results parameters, which is valuable information not captured in the schema descriptions alone. However, it doesn't explain all parameters beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verbs ('structure evidence', 'produces draft sections with gap analysis') and resources ('HTA body-specific submission format'). It explicitly distinguishes from sibling tools by mentioning it 'accepts output from literature_search and cost_effectiveness_model', showing it operates downstream of those tools rather than duplicating their functions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool by specifying it processes outputs from two specific sibling tools (literature_search and cost_effectiveness_model). However, it doesn't explicitly state when NOT to use it or mention alternatives for similar formatting tasks, which prevents a perfect score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
knowledge_readA
Read a file from a project's raw/ or wiki/ tree. Path is relative to project root. Only raw/ and wiki/ subtrees accessible.
| Name | Required | Description | Default |
|---|---|---|---|
| project | Yes | Project ID | |
| path | Yes | Relative path (e.g. 'wiki/trials/sustain-6.md' or 'raw/literature/pubmed_12345.md') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It discloses the read-only nature implicitly through the verb 'Read' and specifies access limitations ('Only raw/ and wiki/ subtrees accessible'), but doesn't mention authentication requirements, rate limits, error conditions, or what happens with invalid paths. It adds some behavioral context but leaves important operational details unspecified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences with zero waste. The first states the core purpose, the second clarifies path semantics, and the third specifies access limitations. Every sentence earns its place by adding distinct, necessary information. The description is appropriately sized and front-loaded with the main functionality.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read operation with 2 parameters, 100% schema coverage, and no output schema, the description provides adequate but not complete context. It covers the what, where, and access limitations, but lacks information about return format, error handling, authentication needs, or performance characteristics. Given the simplicity of the tool and good schema coverage, this is minimally viable but has clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters fully. The description adds minimal value beyond the schema by clarifying that path is 'relative to project root' and providing example formats, but doesn't explain parameter interactions or constraints beyond what's in the schema descriptions. This meets the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Read a file') and resource ('from a project's raw/ or wiki/ tree'), with explicit scope limitations ('Only raw/ and wiki/ subtrees accessible'). It distinguishes from sibling tools like knowledge_write (write vs read) and knowledge_search (search vs direct read).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context about when to use this tool ('Read a file from a project's raw/ or wiki/ tree') and path requirements ('Path is relative to project root'), but doesn't explicitly state when NOT to use it or name specific alternatives like knowledge_search for broader searching. The sibling tool list shows knowledge_search exists, but no direct comparison is made.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
knowledge_searchA
Search a project's knowledge base (raw/ and wiki/) for text matches. Returns file paths with line numbers and snippets. Use this to find previously-retrieved literature, model runs, and compiled wiki content without re-querying external APIs.
| Name | Required | Description | Default |
|---|---|---|---|
| project | Yes | Project ID (must exist) | |
| query | Yes | Search query — multi-term searches match ANY term (OR) | |
| paths | No | Which subtrees to search. Default: both. | |
| max_results | No | Max matches to return (default 20, max 100) | |
| case_sensitive | No | Case-sensitive search (default false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses that the tool returns file paths with line numbers and snippets, which adds useful behavioral context. However, it lacks details on permissions, rate limits, or error handling, leaving gaps for a search tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose, followed by usage guidance, in two efficient sentences with zero wasted words, making it highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description provides basic purpose and usage but lacks details on return format, error cases, or operational constraints. For a search tool with 5 parameters, this is adequate but has clear gaps in completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents all parameters. The description adds no additional parameter semantics beyond what's in the schema, such as explaining search logic or result formatting, resulting in the baseline score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verbs ('search a project's knowledge base') and resources ('raw/ and wiki/'), and distinguishes it from siblings like knowledge_read, knowledge_write, and literature_search by specifying it searches previously-retrieved content without external API calls.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context on when to use this tool ('to find previously-retrieved literature, model runs, and compiled wiki content without re-querying external APIs'), which implicitly differentiates it from literature_search. However, it does not explicitly state when not to use it or name alternatives, keeping it from a perfect score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
knowledge_writeA
Write a file to the project's wiki/ tree. Path MUST start with 'wiki/' and end with '.md'. Use this to compile/organize evidence from raw/ files into a structured knowledge base. Supports Obsidian-style [[wikilinks]].
| Name | Required | Description | Default |
|---|---|---|---|
| project | Yes | Project ID | |
| path | Yes | Relative path starting with 'wiki/', ending with .md (e.g. 'wiki/trials/sustain-6.md') | |
| content | Yes | Markdown content. Can include YAML frontmatter and [[wikilinks]]. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It discloses that the tool writes files and supports Obsidian-style wikilinks, which is useful behavioral context. However, it doesn't mention important behavioral aspects like whether this overwrites existing files, what permissions are required, error conditions, or response format. For a write operation with zero annotation coverage, this leaves significant gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured in three sentences: first states the core action and constraints, second provides usage context, third adds feature support. Every sentence adds value with zero wasted words, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a write tool with 3 parameters, 100% schema coverage, but no annotations and no output schema, the description provides adequate purpose and usage context. However, it lacks important behavioral details about mutation effects (overwrite vs create), error handling, and response format that would be needed for complete understanding. The absence of output schema means the description should ideally address what the tool returns.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters. The description adds some value by clarifying the path format ('MUST start with wiki/ and end with .md') and content capabilities ('Can include YAML frontmatter and [[wikilinks]]'), but doesn't provide significant additional semantic context beyond what the schema descriptions offer.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Write a file'), target resource ('to the project's wiki/ tree'), and distinguishes it from sibling tools like knowledge_read and knowledge_search by focusing on writing rather than reading/searching. It provides concrete purpose beyond just the tool name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context about when to use this tool ('to compile/organize evidence from raw/ files into a structured knowledge base') and mentions path format requirements. However, it doesn't explicitly state when NOT to use it or name specific alternatives among siblings like knowledge_read for reading existing files.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
literature_searchA
Search PubMed, ClinicalTrials.gov, bioRxiv/medRxiv, ChEMBL, FDA Orange Book, FDA Purple Book, enterprise sources (Embase, ScienceDirect, Cochrane, Citeline, Pharmapendium, Cortellis), HTA cost reference sources (CMS NADAC, PSSRU, NHS National Cost Collection, BNF, PBS Schedule), LATAM sources (DATASUS, CONITEC, ANVISA, PAHO, IETS, FONASA), APAC sources (HITAP), and HTA appraisal/guidance sources (NICE TAs, CADTH CDR/pCODR, ICER, PBAC PSDs, G-BA AMNOG, HAS Transparency Committee, IQWiG, AIFA, TLV Sweden, INESSS Quebec) for evidence on a drug or indication. Returns structured results including HTA precedents and appraisal decisions with a full audit trail suitable for HTA submissions.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Research question (e.g. 'semaglutide type 2 diabetes cost-effectiveness') | |
| sources | No | Data sources to query. Default: pubmed, clinicaltrials, biorxiv, chembl (+ embase if ELSEVIER_API_KEY set). Use 'who_gho' and 'world_bank' for epidemiology and demographic data. Use 'oecd' for OECD health statistics (expenditure, hospital beds, physicians, life expectancy). Use 'ihme_gbd' for Global Burden of Disease estimates (DALYs, prevalence, mortality across 204 countries). Use 'orange_book' for FDA drug approvals and therapeutic equivalence. Use 'purple_book' for FDA-licensed biologics and biosimilars. Enterprise (require API key): 'cochrane' (COCHRANE_API_KEY), 'citeline' (CITELINE_API_KEY), 'pharmapendium' (PHARMAPENDIUM_API_KEY), 'cortellis' (CORTELLIS_API_KEY). HTA cost reference sources: 'cms_nadac' (US drug acquisition costs via CMS API), 'pssru' (UK unit costs, reference links), 'nhs_costs' (NHS National Cost Collection, reference links), 'bnf' (UK drug pricing, reference links), 'pbs_schedule' (Australia PBS/MBS pricing, reference links). LATAM sources (explicit request only): 'datasus' (Brazil SUS hospital/ambulatory data), 'conitec' (Brazil HTA reports), 'anvisa' (Brazil drug pricing/registry), 'paho' (Pan American regional health statistics), 'iets' (Colombia HTA reports), 'fonasa' (Chile public health insurance data). APAC sources (explicit request only): 'hitap' (Thailand HTA reports and methodology). HTA appraisal/precedent sources (explicit request only): 'nice_ta' (NICE Technology Appraisals, UK), 'cadth_reviews' (CADTH CDR/pCODR, Canada), 'icer_reports' (ICER evidence reports and HBPBs, US), 'pbac_psd' (PBAC Public Summary Documents, Australia), 'gba_decisions' (G-BA AMNOG benefit assessments, Germany), 'has_tc' (HAS Transparency Committee opinions, France), 'iqwig' (IQWiG systematic reviews and dossier assessments, Germany), 'aifa' (AIFA reimbursement decisions, Italy), 'tlv' (TLV value-based pricing decisions, Sweden), 'inesss' (INESSS drug evaluations, Quebec Canada). | |
| max_results | No | Maximum results to return (default: 20, max: 100) | |
| date_from | No | Exclude results before this date (ISO format: YYYY-MM-DD) | |
| output_format | No | Output format. 'docx' requires hosted tier. | |
| project | No | Project ID for knowledge base persistence. When set, results are saved to ~/.heor-agent/projects/{project}/raw/literature/ |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the tool returns structured results with audit trails suitable for HTA submissions, which adds useful context about output quality. However, it lacks details on rate limits, authentication needs for enterprise sources, or potential costs/limitations of the search operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is overly verbose with a long list of sources in the first sentence, making it difficult to parse quickly. While informative, it could be more front-loaded and structured for clarity, with some details better placed in the schema or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of the tool (6 parameters, no output schema, no annotations), the description is mostly complete by explaining the broad purpose, sources, and output suitability. However, it could better address behavioral aspects like authentication or limitations to fully compensate for the lack of annotations and output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description adds minimal value beyond the schema by implying the query parameter should target drug/indication evidence, but does not provide additional syntax or format details. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('search') and the comprehensive scope of resources (PubMed, ClinicalTrials.gov, bioRxiv/medRxiv, etc.), distinguishing it from sibling tools like cost_effectiveness_model or hta_dossier_prep. It explicitly mentions the purpose is to find evidence for drugs or indications with HTA suitability.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool (searching evidence for drugs/indications with HTA submissions), but does not explicitly state when not to use it or name specific alternatives among siblings. It implies usage for evidence gathering versus other tools focused on modeling or dossier preparation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
project_createA
Initialize a new HEOR project workspace with directory skeleton and project.yaml metadata. Idempotent — returns existing project if already created. Required before using the project parameter in other tools.
| Name | Required | Description | Default |
|---|---|---|---|
| project_id | Yes | Short identifier (alphanumeric + hyphens, e.g. 'semaglutide-t2d') | |
| drug | Yes | Drug or intervention name | |
| indication | Yes | Disease/condition being treated | |
| hta_targets | No | HTA bodies to target (optional) | |
| notes | No | Free-text project notes (optional) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively reveals key traits: the idempotent nature ('Idempotent — returns existing project if already created'), the prerequisite requirement for other tools, and the initialization of specific resources. However, it doesn't mention potential side effects like file system changes or error conditions, leaving some behavioral aspects uncovered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose in the first sentence, followed by critical behavioral notes. Every sentence earns its place by providing essential information without redundancy. The structure is logical and efficiently conveys key points in minimal text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (initialization with idempotency) and the absence of both annotations and an output schema, the description does a good job covering purpose, usage, and key behavior. However, it lacks details on what the tool returns (since no output schema exists) and doesn't fully address all potential side effects or error cases, leaving some gaps in completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description doesn't add any meaningful semantic context beyond what's in the schema—it doesn't explain relationships between parameters or provide usage examples. This meets the baseline expectation when schema coverage is complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Initialize a new HEOR project workspace'), the resources involved ('directory skeleton and project.yaml metadata'), and distinguishes this tool from siblings by explaining its prerequisite role for using the 'project' parameter in other tools. It goes beyond a simple tautology by detailing what initialization entails.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool ('Required before using the `project` parameter in other tools') and provides a clear alternative scenario ('returns existing project if already created'). This gives the agent precise guidance on timing and fallback behavior without misleading information.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v0.1.7- First observed
cost_effectiveness_model - First observed
hta_dossier_prep - First observed
knowledge_read - First observed
knowledge_search - First observed
knowledge_write - First observed
literature_search - First observed
project_create
TDQS
Scored across 7 tools
Each tool has a clearly distinct purpose with no overlap: cost-effectiveness modeling, HTA dossier preparation, knowledge base reading/searching/writing, literature searching, and project creation. The descriptions clearly differentiate their functions, making misselection unlikely.
Most tools follow a consistent verb_noun pattern (e.g., literature_search, project_create). The only minor deviation is 'knowledge_read/write/search' which uses a noun_verb format, but this is internally consistent within the knowledge tools and still readable.
With 7 tools, this is well-scoped for a Health Economics and Outcomes Research (HEOR) agent. Each tool earns its place by covering distinct aspects of the workflow: project setup, evidence gathering, analysis, and documentation.
The toolset provides complete coverage for the HEOR domain: project initialization, comprehensive literature searching, cost-effectiveness modeling, HTA dossier preparation, and knowledge management. There are no obvious gaps; it supports the full lifecycle from evidence collection to submission-ready outputs.
Maintenance
Related MCP Connectors
Hybrid human + AI expertise for faster, trusted answers and decisions via MCP Server.
Bioinformatics MCP for genomic variant interpretation, gene-disease evidence and literature.
NIH clinical trials and FDA adverse event reports. 4 MCP tools for health research.
AI-powered medical document management for cancer patients. Google Drive, Gmail, Calendar via MCP.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceAn advanced integrated MCP server platform that combines 600+ tools and multiple biomedical databases to enable comprehensive information retrieval across molecules, proteins, genes, and diseases for accelerating therapeutic research.38-
- FlicenseNot gradedqualityDmaintenanceA comprehensive MCP server that bridges AI applications with FHIR healthcare data systems, enabling patient data access, clinical data retrieval, and data quality assessment.4-
- AlicenseNot gradedqualityDmaintenanceAn MCP server with 60 tools connecting AI assistants to Czech healthcare databases (SUKL, MKN-10, NRPZS) and global biomedical sources (PubMed, ClinicalTrials.gov, OpenFDA).1MIT
- AlicenseNot gradedqualityDmaintenanceMCP server for clinical and pharmaceutical data, enabling search of ClinicalTrials.gov, PubMed, FDA, and ICH guidelines without API keys.19 npmMIT