Skip to main content
Glama

med-lit-mcp

med-lit-mcp is an MCP server for medical literature reviews. It takes a review through four deliberate stages, search → screening → fetch → wiki, and ends with an evidence-linked wiki you can open in Obsidian.

Your MCP client's model does the reading and writing (screening decisions, entity extraction, syntheses). med-lit-mcp searches the databases, stores everything, checks the model's work against the source text, and exports the wiki. Every screening decision, extracted mention and synthesis citation must be a verbatim quote or a reference to the stored text.

It works with Hermes, Claude Code and Claude Desktop on Linux and macOS. It needs no embedding endpoint, no model API key and no build step. The only required setting is NCBI_EMAIL.

Install

med-lit-mcp is on PyPI and runs through uv: uvx fetches the server and its dependencies on first start. The first start takes a while, so pre-warm the cache and check your settings once:

NCBI_EMAIL=you@example.org uvx med-lit-mcp --check

To try the latest unreleased code instead, replace med-lit-mcp with --from git+https://github.com/junhewk/med-lit-mcp med-lit-mcp in any command below.

Hermes

hermes mcp add med-lit --command uvx \
  --env NCBI_EMAIL=you@example.org \
  --connect-timeout 180 \
  --args med-lit-mcp
hermes mcp test med-lit

--env must come before --args, because --args takes everything after it. Hermes passes an MCP server only the variables given with --env (plus basics such as PATH and HOME), so add optional keys there too, for example --env NCBI_EMAIL=… NCBI_API_KEY=….

Claude Code

claude mcp add med-lit --scope user -e NCBI_EMAIL=you@example.org -- \
  uvx med-lit-mcp

Claude Desktop

Add the server to claude_desktop_config.json: ~/Library/Application Support/Claude/ on macOS, ~/.config/Claude/ on Linux. Desktop starts servers with a minimal PATH, so use the absolute path printed by which uvx.

{
  "mcpServers": {
    "med-lit": {
      "command": "/Users/you/.local/bin/uvx",
      "args": ["med-lit-mcp"],
      "env": {"NCBI_EMAIL": "you@example.org"}
    }
  }
}

Updating

uvx caches the installed version. To pick up a new release, run uvx med-lit-mcp@latest --check once, then restart the client.

Related MCP server: research-hub

Settings

All settings are environment variables passed in the MCP client configuration.

Variable

Purpose

NCBI_EMAIL

Required for PubMed/PMC search and for fetching. Also sent to Unpaywall, which requires a contact email. Use a real address.

NCBI_API_KEY

Optional; raises NCBI rate limits.

SEMANTIC_SCHOLAR_API_KEY (or S2_API_KEY)

Enables Semantic Scholar by default. Without a key it shares a public rate limit that usually answers HTTP 429, so it is skipped unless requested. Request a free key.

OPENALEX_API_KEY

Optional OpenAlex key.

MED_LIT_PROJECTS_DIR

Where new projects are created when no path is given. Default: ~/med-lit.

MED_LIT_STATE_DIR

Where the list of known projects is kept. Default: $XDG_DATA_HOME/med-lit-mcp or ~/.local/share/med-lit-mcp.

MED_LIT_STAGES

Optional, for advanced users. Exposes only the tools up to a stage: search, screening, fetch or wiki (default: all). For example, fetch hides the 11 wiki tools. Clients that load tools on demand rarely need this; it helps with clients that load every tool up front. Project, status and guide tools are always available.

Projects

Each review is a project: one self-contained folder holding the review's wiki and, hidden inside it, everything else.

~/med-lit/LLMs in shared decision making/      ← the project, and the wiki
├── index.md
├── log.md
├── entities/Shared decision making.md
├── sources/Guirgus 2026 - Assessing Artificial Intelligence in Patient Education.md
└── .med-lit/                                  ← hidden working data
    ├── project.json                           ← name, id, format version
    ├── med-lit.sqlite3                        ← this review's articles and knowledge graph
    └── runs/<run-id>/                         ← searches, screening decisions, fetched text
  • One topic, one project. Unrelated reviews never share entities or syntheses. One project can hold several searches, for example an update search a year later.

  • The folder is the unit. Move, copy, archive or back it up as a whole. If you move it, ask the agent to open_project the new location.

  • Where it lives is up to you: the default is ~/med-lit/<name>, or give a path when creating the project.

Run a study

Talk to your agent. The server's instructions tell it to start each stage only when you ask, and to confirm the search question with you first.

You: Start a new review on how large language models are used in shared decision making in healthcare.

Agent: create_project, reads guide("question"), drafts a PCC question, calls validate_question, and shows it to you.

You: Looks good, run it.

Agent: start_search with the approved question_id, which returns a run ID and candidate counts.

You: Screen it. Include studies of LLMs supporting patient–clinician decisions; exclude purely technical benchmarks.

Agent: set_screening_criteria, then repeats next_screening_batch and record_screening_decisions.

You: Include the uncertain one about decision aids: it describes an LLM-drafted aid.

Agent: review_article.

You: Fetch the included articles and build the wiki.

Agent: fetch_articles, then wiki_tasks, which hands each article and each batch of entity pages to a fresh subagent, then export_wiki.

Claude Code also exposes the guided prompts /mcp__med-lit__plan_search, screen_run and build_wiki.

How the stages work:

  • Question structure. Groups within a component are ANDed, and a group's synonyms are ORed. Give each separately required facet (for example technology and task) its own group; put only true synonyms in a group; leave out ambiguous bare acronyms such as LLM unless you approve them. PCC context is for the setting only. Groups can suggest candidate MeSH headings, which are checked against NCBI before use.

  • Date range. Without a from_date filter, a search covers only the last three years. validate_question states the effective start date, so ask for an earlier one if you need it.

  • Screening uses titles and abstracts only. Include and exclude decisions need a verbatim quote from the record. A decision whose quote cannot be verified is stored as uncertain, and uncertain articles wait for the researcher's own decision.

  • Changing criteria requires replace=true, starts a new revision, and re-screens every article. Earlier decisions stay in the history.

  • Fetch tries PMC full text (NCBI, then Europe PMC). For articles without a PMCID it asks Unpaywall for legal open-access copies of the DOI, preferring a PMC copy and otherwise extracting text from an open-access PDF. The abstract is the last resort, and abstract-only articles are always labelled as such. The source, license and version (published, accepted or submitted) of each full text are recorded. To retry abstract-only articles later, ask for a fetch with retry_abstract_only.

  • The wiki is built from paged article text of about 12,000 characters per page, without truncation.

    • Each entity mention and relationship must be quoted verbatim from its page.

    • Entities are matched lexically: case, plurals, spelling variants and acronyms defined in the text, such as "large language models (LLMs)". Only exact key or acronym matches merge automatically. Near matches are queued as possible duplicates for the agent or researcher to merge or keep distinct.

    • Syntheses must cite their sources as [uid].

    • Entities, roles and relationships follow the ontology described below.

    • Article pages are large, so wiki_tasks splits the work into self-contained tasks: one per article, then duplicate review, then batches of five entity pages. Clients that can delegate (Hermes subagents, Claude Code agents) give each task to a fresh subagent, which keeps the main conversation short. Other clients work through the same tasks in order.

Long steps are split into bounded calls. If a call reports running or remaining > 0, call it again with the same run ID. A failed stage is retried with the same tool and run ID.

If some sources fail while others succeed (for example PMC answering HTTP 500 while NCBI is degraded), the run completes with the working sources and lists the rest in source_failures. Later, resume_search with retry_failed_sources=true searches only the failed sources again and adds their new records to the same run. Those records then appear as pending in screening.

Entity types, roles and relationships

The wiki is a small knowledge graph. Its model follows the lightweight ontology of the Simple Graph Builder Obsidian plugin: a few fixed entity types, free-form relationship verbs, and a detail note on each relationship. That plugin's types are general-purpose, so med-lit-mcp uses medical types, each refining exactly one of the plugin's types. The full contract is in docs/ontology.md.

Type

Covers

Plugin type

CONDITION

Diseases, disorders, symptoms, diagnoses

CONCEPT

INTERVENTION

Treatments, drugs, procedures, surgeries, therapies, programs

METHOD

TECHNOLOGY

AI systems and models, software, devices, platforms

TOOL

METHOD

Study designs, research and analytic methods, instruments, scores

METHOD

GUIDELINE

Laws, regulations, clinical guidelines, reporting standards, policies

DOCUMENT

DATASET

Datasets, databases, registries, benchmarks

TOOL

CONCEPT

Ideas, principles, ethical, social or professional notions

CONCEPT

PERSON, ORGANIZATION, PLACE

People; institutions; countries, regions, care settings

same name

Roles. Each mention of an entity can record the PICO/PCC role it plays in that article's own study: population, intervention, comparator, outcome, concept or context. Roles belong to mentions rather than entities, because type 2 diabetes can be the population of one study and an outcome of another.

Relationships are directed, such as ChatGPT —evaluates→ patient education materials. Each has a free-form verb, an optional detail note, and a verbatim evidence quote from every article that asserts it. A relationship's strength is the number of articles supporting it.

Tools

Stage

Tools

Projects

create_project, list_projects, open_project

Search

validate_question, start_search, resume_search

Guidance

guide (rules for each stage, on demand)

Screening

set_screening_criteria, next_screening_batch, record_screening_decisions, review_article

Fetch

fetch_articles

Wiki

wiki_tasks, next_wiki_article, get_article_page, record_extraction, find_entities, list_duplicate_candidates, resolve_duplicates, merge_entities, next_synthesis, record_synthesis, export_wiki

Status

list_runs, get_run_status, list_articles

Tools that work on a whole project take a project name, which can be left out while only one project exists. Tools that work on one search take its run ID.

validate_question stores the checked question and returns a question_id; start_search takes that id, so a search always runs exactly the question and sources the researcher approved.

Tool descriptions are kept to a line each, and the detailed rules for each stage come from guide(topic) when the agent needs them. Hermes, Claude Code and Claude Desktop all load MCP tools on demand (Tool Search), finding them by name; the server instructions list every tool by stage for that reason.

The wiki in Obsidian

A project folder is plain Markdown with YAML front matter and relative links, so Obsidian can open it with no plugin: choose Open folder as vault and pick the project folder. Obsidian ignores the hidden .med-lit folder.

File

Contents

index.md

Synthesized entities grouped by type, with a one-line summary each

entities/<Name>.md

One page per entity: summary, key aspects, synthesis, relationships with evidence, and source articles with the entity's role in each. Entities without a synthesis yet get a short page listing their mentions.

sources/<Author Year - Title>.md

One page per fetched article: bibliographic details, abstract, the entities it mentions and their roles, and a notice when only the abstract was available

log.md

Articles added to the wiki, newest first

Page names are the names people read, so they appear as page titles in Obsidian. Each name is assigned once and stays stable; the entity or article id is kept in the front matter. Front matter becomes Obsidian properties (entity_type, sgb_type, aliases, sources, version, stale and so on), which you can search, sort and filter. aliases lets Obsidian resolve an acronym or spelling variant to the right page. Every relationship and citation is a link, so Obsidian's graph view connects entities to each other and to their source articles. Obsidian's graph draws links without labels, so relationship verbs appear on the pages rather than on the graph edges.

Things to know:

  • Don't edit generated pages. Every export rewrites the pages it generated, so keep your own notes in separate files and link to the generated pages. Export only writes or removes files carrying generator: med-lit-mcp in their front matter. It never overwrites any other file and lists any it skipped under not_overwritten.

  • Exports rewrite only changed files, so Git or a sync service sees real changes only.

  • Back up the whole folder, including .med-lit/. Avoid opening the same project from two machines through a file-sync service while it is in use: the database inside is a live SQLite file.

Simple Graph Builder. The ontology is designed so that a future version of the Simple Graph Builder plugin can import a project's graph into your own vault, with typed relationships and their evidence. Until then, if you open a project inside a vault where that plugin runs, add the project folder to its Analysis Exclusions. Otherwise it would re-extract entities from these pages with its own LLM, creating a second, unsourced copy of the graph.

Viewer (optional)

uvx --from med-lit-mcp med-lit-viewer --port 3000

The viewer lists the runs of every known project, with screening reasons and evidence, fetched text and each project's wiki pages, and lets researchers record manual screening decisions. It binds to 127.0.0.1. For remote access, put it behind an authenticated reverse proxy or a private network such as Tailscale.

Development

uv sync
uv run pytest
uvx ruff check src
npx @modelcontextprotocol/inspector uv run med-lit-mcp

To run a server from a checkout, use uv run --directory /path/to/med-lit-mcp med-lit-mcp as the MCP command. Unlike uvx --from <path>, it always runs the checkout's current code.

Windows is not supported (run locks use fcntl).

The database search engine in src/med_lit_mcp/medsearch/ (query compilation, MeSH resolution, PubMed/PMC/OpenAlex/Semantic Scholar/Scopus clients, deduplication and ranking) was merged from hermes-medical-search 0.2.1 (MIT) and is maintained here. For debugging it can run on its own: uv run python -m med_lit_mcp.medsearch --help.

License

med-lit-mcp is licensed under the PolyForm Noncommercial License 1.0.0. You may use, modify and share it for any noncommercial purpose, including research, teaching, and use by educational institutions, public research organizations and public health organizations. Commercial use needs a separate license; open an issue to ask. NOTICE records the MIT-licensed origin of the search engine.

Available Tools

26 tools
create_projectB

Create a project: one folder per review, holding its wiki and (hidden) its data.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesReview topic, e.g. 'LLMs in shared decision making'
pathNoFolder for the project (for example inside an Obsidian vault); default ~/med-lit/<name>

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and openWorldHint=false, so the write-and-local nature is known. The description usefully adds that a dedicated folder is created and that data is hidden, but omits key mutation behavior such as what happens on a name collision, idempotency, or required permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the core action leads. The parenthetical '(hidden)' is slightly cryptic but carries real behavioral information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be described, and annotations cover the safety profile. However, for a creation tool with a default-derived path, the description should say something about collision/overwrite behavior or naming constraints — that gap keeps it at minimum-viable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: both 'name' (review topic, 1-80 chars) and 'path' (default ~/med-lit/<name>) are fully documented in the schema. The description adds nothing about parameters, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Create) and resource (project) and adds structural context — 'one folder per review, holding its wiki and (hidden) its data' — which tells the agent the side effect and layout. It does not explicitly name sibling tools like open_project or list_projects, so differentiation is implied rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to create a new project versus opening an existing one (open_project) or listing them (list_projects). No prerequisites or conditions are given, so the agent must infer selection from the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

export_wikiB

Rewrite every Markdown page of a project's wiki (entities, sources, index, log) from its database.

ParametersJSON Schema
NameRequiredDescriptionDefault
projectNoProject name from list_projects; may be omitted when only one project exists

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and openWorldHint=false, so the agent knows this mutates state locally. The description usefully adds that it is a whole-wiki operation ('every Markdown page') across four page categories, but it never says the files are overwritten or where they are written, so the destructive/overwrite nature is only implied.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that states the scope and the page categories with no filler. It could be tightened because 'from its database' is slightly ambiguous about the direction of the operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. However, for a mutating whole-wiki export the description omits where pages are written and when to invoke it, leaving the operational context only partially covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 100% description coverage for the single 'project' parameter, including the note that it may be omitted when only one project exists. The description adds nothing about the parameter, so the schema does the heavy lifting and the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a concrete verb+resource: 'Rewrite every Markdown page of a project's wiki' and enumerates the affected page types (entities, sources, index, log). It is clear enough to distinguish an agent's intent, though it never contrasts itself with sibling wiki tools like wiki_tasks or next_wiki_article, and the name 'export' versus the verb 'rewrite' creates minor tension.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance: nothing says when to run this versus wiki_tasks, next_wiki_article, or get_article_page, nor any prerequisites such as needing a populated database or an open project. The agent must infer timing entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fetch_articlesA

Fetch full text of included articles (PMC, Unpaywall open access, else abstract); repeat while remaining > 0. Rules: guide("fetch").

ParametersJSON Schema
NameRequiredDescriptionDefault
uidsNoSpecific included articles; default is the next pending ones
run_idYesRun ID from start_search
max_itemsNo
retry_failedNoAlso retry articles whose fetch failed
retry_abstract_onlyNoAlso retry articles that only have an abstract

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and openWorldHint=true, so the external/state-mutating nature is partly covered. The description adds real value by disclosing the free-full-text-first fallback strategy and the loop-until-exhausted behavior, but it doesn't explain why it is non-read-only (that results are stored) or mention rate limits against the open-access sources.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences that lead with the main action and the source strategy, with no padding. The trailing "Rules: guide('fetch')" is slightly cryptic but is a useful pointer rather than waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described, and annotations carry the safety profile. The description supplies the pipeline context (source waterfall, repeated invocation until remaining hits zero) and routes to guide for rules, which is sufficient for an agent to call it correctly; only the run_id dependency chain is left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 80%, so the schema already documents uids, run_id, and the retry flags well. The description adds nothing about parameters (it never mentions run_id, max_items, or the retry semantics), so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Fetch full text of included articles") and even explains the source waterfall (PMC, Unpaywall open access, else abstract), so the agent knows exactly what content it returns. It does not, however, distinguish itself from the related read tools (get_article_page, list_articles), so sibling differentiation is missing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a workflow cue ("repeat while remaining > 0") and defers rules to a sibling via "guide('fetch')", which implies usage context. But it never states when this tool should be chosen over get_article_page or list_articles and gives no explicit exclusions or prerequisites beyond the run_id hint.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_entitiesB
Read-only

Search a project's wiki entities by name, alias, acronym, or spelling variant.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryYes
projectNoProject name from list_projects; may be omitted when only one project exists
entity_typeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds one genuinely useful behavioral fact — lookup is fuzzy across aliases, acronyms, and spelling variants — but says nothing about result capping or ordering despite a limit parameter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the verb and resource with zero filler. Nothing redundant or padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and annotations cover safety. However, for a 4-parameter search tool with a 10-value entity_type enum and a 1-50 limit, the description is too thin to let an agent invoke it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 25%: only 'project' is documented in the schema, and the description adds nothing about query, limit, or entity_type. Since the description must compensate for the coverage gap and instead repeats only the query notion, it leaves entity_type filtering and limit behavior entirely opaque.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (search) and resource (project's wiki entities) and even enumerates the match modes (name, alias, acronym, spelling variant), which is unusually informative. It is distinguishable from merge_entities and wiki_tasks, though it never names a sibling explicitly, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no exclusion criteria, and no mention of alternatives such as merge_entities or resolve_duplicates, which also operate on entities. The reader can infer this is the lookup entry point, but nothing routes them explicitly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_article_pageA
Read-only

Re-read one page of a fetched article (pages are numbered from 0).

ParametersJSON Schema
NameRequiredDescriptionDefault
uidYes
pageYes
run_idYesRun ID from start_search

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description contributes the genuinely useful 0-indexed paging convention, but says nothing about what a 'page' contains, whether it re-fetches from the network or reads from cache, or what the response looks like.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with zero waste; the core action comes first and the paging convention follows immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and three required parameters, the description should carry more weight. It covers the paging model adequately for a simple read, but leaves 'uid' and the nature of a page undefined, which an agent needs in order to call this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33% — only run_id is documented in the schema. The description compensates for exactly one gap by explaining that pages are numbered from 0, which is real added meaning for 'page', but 'uid' remains completely unexplained in both the schema and the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: re-reading one page of an already-fetched article. It distinguishes itself from fetch_articles and list_articles via the word 're-read,' but it never names a sibling or clarifies the relationship, so the differentiation is implied rather than explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Re-read' implies the usage context (after an article has already been fetched), which is useful but inferred rather than stated. There are no explicit when-to-use/when-not conditions and no naming of alternatives such as fetch_articles for a first read.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_run_statusA
Read-only

Show a run's progress in every stage and the available next steps.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYesRun ID from start_search

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered by structured data rather than the prose. The description adds a modest behavioral signal by indicating the response contains per-stage progress and next-step guidance, but says nothing about freshness, blocking behavior, or errors for unknown run IDs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence stating what is shown first (progress by stage) and then what else is returned (next steps). No filler, no restatement of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description does not need to enumerate return fields, and the read-only annotation covers risk. For a one-parameter status lookup it is nearly sufficient, missing only a pointer on how it relates to list_runs and resume_search.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single required parameter with 100% schema description coverage, so the schema already documents run_id, its format constraint, and its origin ("Run ID from start_search"). The description contributes nothing further about the parameter, which is the expected baseline when the schema does the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ("Show a run's progress") and scopes it to stages plus next steps, which is meaningfully more specific than the bare tool name. It does not, however, differentiate itself from siblings like list_runs or resume_search, so an agent still has to infer which run-oriented tool applies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the agent can guess this is called with a run_id to inspect progress, and the phrase "available next steps" hints the output guides what to do next. There is no explicit when-to-use, when-not-to-use, or named alternative such as list_runs for enumerating runs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

guideC
Read-only

Rules and how-to for a stage: workflow, question, screening, fetch, wiki, extraction, duplicates, synthesis or ontology.

ParametersJSON Schema
NameRequiredDescriptionDefault
topicYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the agent knows this is a safe, local, side-effect-free read. The description adds that the content is prescriptive rules/how-to rather than data, which is mild extra context. It doesn't describe the shape or breadth of the returned guidance, but the output schema covers return values, so a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with the purpose front-loaded followed by the enumerated stages. It is tight and wastes little, though restating the enum values is somewhat redundant with the schema. Efficient overall.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With annotations covering safety and an output schema defining the return, the main missing piece is usage context — when in the workflow to consult this guide relative to the sibling action tools. For a meta/reference tool in a large workflow suite, that gap keeps it from being complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but the single parameter carries an enum of nine values that the description reproduces verbatim. It confirms the valid topics but adds no per-topic meaning (e.g., what 'duplicates' or 'ontology' guidance contains). It partially compensates for the coverage gap by surfacing the options but adds nothing beyond the enum itself.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description conveys that this is a reference/guidance tool that returns rules and how-to material scoped to a named stage, and it enumerates those stages (workflow, question, screening, etc.). However, the verb is nominal ('Rules and how-to for a stage') rather than an explicit action, so the agent must infer that the tool retrieves documentation. It is understandable but not sharply stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this tool versus the many action siblings (validate_question, start_search, set_screening_criteria, etc.). It never says 'call this first to learn how a stage works' or gives any precondition or workflow context. The agent is left to infer usage entirely from the topic enum.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_articlesC
Read-only

List a run's articles with their screening, fetch, and wiki status.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
offsetNo
run_idYesRun ID from start_search
decisionNo
wiki_statusNo
fetch_statusNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds that results include screening/fetch/wiki status fields, but omits any pagination or filter behavior despite limit/offset existing in the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded, no filler. It is efficient, though arguably too terse for a six-parameter tool with filterable enums.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described, but with 6 parameters, low schema coverage, and undocumented filter enums, the description leaves the agent guessing about filterability and pagination of this list operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 17% – just run_id is documented. The description hints that screening (decision), fetch_status, and wiki_status appear in results but never states they can be used as filters, and limit/offset are wholly unexplained in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb (list), resource (articles), scope (a run's), and states the returned status dimensions (screening, fetch, wiki). It does not differentiate itself from fetch_articles, but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance and no named alternative (e.g., fetch_articles for full text). The agent must infer that this is a browsing/listing operation versus a fetching one.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_duplicate_candidatesB
Read-only

List entity pairs that may be the same thing, with an example mention of each.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
run_idNoOnly pairs involving this run's articles
projectNoProject name from list_projects; may be omitted when only one project exists

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds that each pair comes with an example mention, a small behavioral detail, but says nothing about scoring, confidence, ordering, or pagination behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence that front-loads the resource. It is appropriately sized, though it carries almost no additional information beyond the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and annotations cover safety. However, for a discovery tool that feeds merge/resolve siblings, the missing usage guidance and parameter explanation leave the definition merely adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 67% (limit is undocumented), and the description adds no parameter meaning at all. It never explains that run_id scopes results to a search run or that project may be omitted in single-project setups, leaving the two most consequential filters under-explained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('entity pairs that may be the same thing'), which is clearly the duplicate-candidate detection concept. It does not differentiate itself from the closely related sibling resolve_duplicates or merge_entities, but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus resolve_duplicates or merge_entities, no mention of prerequisites or workflow position. The agent must infer that this is a discovery step feeding the resolution tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_projectsA
Read-only

List known projects with their folders. available=false means the folder moved or is missing.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds genuine context beyond that by explaining how to interpret the returned `available` flag (folder moved or missing), which affects how an agent acts on results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero filler, with the core action front-loaded and the interpretive caveat second. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-param read-only list tool with an output schema, the essentials are covered: what is returned and the one non-obvious field semantic. Slightly thin on ordering/pagination and where projects come from, but nothing required to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline for a no-param tool applies. The `available` semantic it explains is an output field, not an input, so it does not inflate this dimension.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ("List known projects with their folders"), which cleanly separates it from write-oriented siblings like create_project and open_project. It does not explicitly name or contrast with those siblings, so it stays short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: an agent can infer this is the discovery step for projects/folders before calling open_project. There is no explicit when-to-use statement, no prerequisites, and no mention of alternatives among the 24 siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_runsB
Read-only

List a project's searches (runs), newest first.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
projectNoProject name from list_projects; may be omitted when only one project exists

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds the newest-first ordering, which is useful behavioral context, but says nothing about pagination or how limit interacts with results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with zero filler; the scope and ordering are conveyed compactly with no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and ordering is covered. However, for choosing among the many sibling tools the agent gets no guidance, and the limit parameter's semantics are left to inference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 50%: the project parameter is documented but limit has no description at all. The description offers no parameter meaning and does not compensate for the undocumented limit beyond what its name/default imply.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List), resource (a project's searches/runs), and result ordering (newest first), which lets an agent understand the operation immediately. It does not explicitly distinguish itself from related siblings like get_run_status or list_projects, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance, no conditions for choosing this over get_run_status, and no prerequisites stated. Usage is only implied by the phrase 'a project's searches'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

merge_entitiesC
Destructive

Merge entities into one, and/or rename or retype an entity.

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonYes
keep_idYes
projectNoProject name from list_projects; may be omitted when only one project exists
merge_idsYesEntities folded into keep_id; empty to only rename or retype
entity_typeNo
canonical_nameNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true and openWorldHint=false, so the safety profile is covered structurally; the description adds nothing about irreversibility, what happens to the folded entities' data, or permission requirements. For a destructive merge with no extra behavioral context, this is thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the primary action front-loaded and the optional capability trailing. It is not padded or repetitive, though it is arguably too terse given the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, 6-parameter mutation tool with a required reason field and partial schema coverage, the description omits the merge-vs-rename distinction mechanism, the required reason, and any warning about data loss. An output schema exists, so return values need not be explained, but the input-side gaps remain serious.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33%, leaving keep_id, reason, entity_type, and canonical_name undocumented. The description's 'rename or retype' weakly maps to canonical_name/entity_type but gives no syntax, no mention that reason is required or capped at 400 chars, and no note that merge_ids is capped at 10.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Merge entities into one') and adds the secondary capability of rename/retype. However, it does not distinguish itself from close siblings such as resolve_duplicates or list_duplicate_candidates, so an agent cannot tell which to pick from the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus resolve_duplicates or list_duplicate_candidates, and no prerequisites (such as being in the right project). The 'and/or' phrasing is the only usage hint, and it is implicit at best.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

next_screening_batchB

Get the next unscreened titles/abstracts with the criteria to judge them against.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYesRun ID from start_search
batch_sizeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and openWorldHint=false. The description says 'Get', which reads like a passive read, yet it never explains that calling it advances/consumes the screening queue, whether items can be re-fetched, or what happens if the batch is not recorded. For a state-advancing 'next' tool this is a notable disclosure gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler; the verb and scope come first. It is efficient, though very terse given the workflow it participates in.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. However, for a two-parameter, non-read-only queue tool, the description omits prerequisite sequencing, iteration behavior, and batch_size semantics, leaving the agent with a minimum-viable picture.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: run_id is documented only as 'Run ID from start_search', and batch_size carries no description at all in the schema. The description does not compensate by explaining batch_size limits/semantics or how run_id links to start_search, so the undocumented parameter remains opaque.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Get') and resource ('next unscreened titles/abstracts'), plus the payload detail that screening criteria are returned alongside. That distinguishes it reasonably well from fetch_articles/list_articles, though it doesn't name any sibling explicitly to route the agent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: mentioning 'unscreened' and 'criteria to judge them against' signals it belongs inside the screening loop, but there is no explicit when-to-use, no prerequisite (start_search/set_screening_criteria first), and no pointer to record_screening_decisions as the follow-up.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

next_synthesisC

Get the next entity whose wiki page needs writing, with its evidence from the project.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idNoOnly entities from this run's articles
projectNoProject name from list_projects; may be omitted when only one project exists
min_sourcesNoEntities need at least this many source articles

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are sparse (readOnlyHint=false, openWorldHint=false) and the description does not fill the gap. Critically, it does not say whether calling this advances a queue or reserves/locks the returned entity, despite readOnlyHint=false implying a state-affecting operation. No mention of auth, ordering, or what happens when no entity needs writing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the key payload front-loaded after the verb. No filler, though it is terse enough that a bit more procedural framing could have been added without bloat.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained in prose. However, for a workflow-step tool sitting among next_wiki_article and next_screening_batch, the description omits the state-advancing/queue behavior that an agent needs to call it safely, leaving a real gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so run_id, project, and min_sources are already documented in the schema, and the description adds nothing about their semantics. Baseline 3 is appropriate when structured fields carry the parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Get') and resource ('the next entity whose wiki page needs writing') plus the payload ('its evidence'), so the agent knows exactly what comes back. It is distinguishable from list_articles or find_entities, but it never contrasts itself with the closely named sibling next_wiki_article or next_screening_batch, so sibling differentiation is incomplete.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to call this versus alternatives such as next_wiki_article or a batch flow, nor when not to call it. The agent must infer that this is a step in the synthesis workflow and has no explicit cue about sequencing or preconditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

next_wiki_articleB

Get the next page of article text to extract wiki entities from (a JSON header plus the page).

ParametersJSON Schema
NameRequiredDescriptionDefault
uidNoServe only this article's pages (for a subagent that owns one article)
run_idYesRun ID from start_search
max_pagesNoLimit pages read for the next article

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, consistent with 'next page' implying a stateful cursor advance. The description adds a return-format hint (JSON header plus page), which is useful, but does not confirm that pages are consumed or what happens after the last page. Adequate but thin for a stateful iterator.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that states the action and the extraction context with no filler. Slightly dense parenthetical, but nothing wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and only two annotations, the description leaves open iteration termination, page-size defaults, and how the returned header relates to run state. It is enough to attempt a call but not enough to confidently drive the pagination loop.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so uid, run_id, and max_pages are already documented in the schema (including the subagent ownership case for uid and the run_id pattern). The description adds no parameter-level detail beyond the schema, which is the baseline 3 when the schema does the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb and resource: 'Get the next page of article text.' The purpose of feeding entity extraction is clear. It does not, however, differentiate itself from the sibling get_article_page, which an agent could plausibly confuse it with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies a pagination loop ('the next page') but never says when to call this versus get_article_page, what ends the iteration, or what run state is required. Usage must be inferred from the schema and the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

open_projectB

Register an existing project folder, for example after it was moved or copied from elsewhere.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesFolder of an existing project

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false and openWorldHint=false, and the word 'Register' is consistent with that state-changing profile, so there is no contradiction. The description adds the notion of re-registering a relocated folder but stays silent on error behavior (e.g. missing/invalid path) or what the registration actually changes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the action front-loaded and the qualifying scenario trailing it. Nothing is wasted, though the brevity is close to under-specification rather than optimal conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need no explanation, and one fully documented parameter plus annotations cover most of the surface. Still, for a state-mutating tool the description omits failure conditions and any distinction from create_project, which an agent needs to choose correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the single 'path' parameter is fully documented in the schema itself. The description only restates that the path is a folder of an existing project, adding no format, validation, or relative-vs-absolute detail beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Register) and resource (an existing project folder), which is more precise than the bare name. The phrase 'existing project folder' implicitly distinguishes it from create_project, though it never names that sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The clause 'for example after it was moved or copied from elsewhere' gives a motivating scenario, which implies when to use it. However, it gives no explicit when-not guidance and does not point to create_project as the alternative for new projects, leaving the agent to infer the routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_extractionB
Idempotent

Record the entities and relationships extracted from one page (re-recording replaces it). Rules: guide("extraction").

ParametersJSON Schema
NameRequiredDescriptionDefault
uidYes
pageYes
run_idYesRun ID from start_search
entitiesYes
relationshipsYes
content_sha256Yes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare idempotentHint=true and openWorldHint=false, so the safety profile is covered. The description adds genuine value by disclosing that re-recording a page replaces the prior record, which explains the overwrite semantics the agent must understand. It does not cover auth, rate limits, or partial-failure behavior, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences, front-loaded with the core action and the overwrite caveat, then the rules pointer. No waste, though the guide reference is terse enough to feel like a pointer rather than guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is unnecessary, and the overwrite behavior is stated. However, for a 6-required-parameter mutation tool with 17% schema coverage, the meaning of uid, content_sha256, and the run linkage is left undocumented, which is a meaningful gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Top-level schema description coverage is only 17%: uid, content_sha256, page, entities, and relationships carry no schema descriptions, and the description does not explain them either. It only implies a per-page scope; the run linkage and hashing parameters remain opaque.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Record') and resource (entities and relationships extracted from one page), and the parenthetical clarifies scope. It does not name a contrasting sibling like find_entities or merge_entities, so it is clear but not fully differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only usage directive is 'Rules: guide("extraction")', which routes the agent to a rules source rather than stating when to call this versus alternatives. It implies this is part of the extraction workflow but leaves conditions and prerequisites to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_screening_decisionsB

Record screening decisions; include/exclude need a verbatim quote as evidence.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYesRun ID from start_search
revisionYesrevision from next_screening_batch
decisionsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and openWorldHint=false, so safety profile is partly covered. The description adds a genuine behavioral constraint beyond the schema — include/exclude require a verbatim quote — which is useful. It omits mutation consequences, revision/conflict behavior, and the 25-item batch cap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the key constraint front-loaded. It is appropriately sized and wastes no words, though the terse phrasing leaves room for more structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. However, for a mutating batch tool in a multi-step screening pipeline, the description omits revision/conflict semantics, batch limits, and workflow ordering that an agent needs to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 67%, so the schema already documents run_id, revision, and the evidence field. The description reinforces the evidence rule but adds no meaning for run_id, revision, uid, or the decision enum that the schema does not already supply. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Record screening decisions'), which is a clear write operation distinct from read-oriented siblings like next_screening_batch or review_article. It does not, however, explicitly name or contrast itself against those siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use or when-not-to-use guidance. The description never references the workflow position (e.g., call after next_screening_batch, before the next batch/revision), nor does it mention alternatives. The agent must infer ordering from parameter names alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_synthesisB

Save an entity page written only from next_synthesis evidence. Rules: guide("synthesis").

ParametersJSON Schema
NameRequiredDescriptionDefault
projectNoProject name from list_projects; may be omitted when only one project exists
summaryYes2-3 sentence overview
entity_idYes
synthesisYesMarkdown article citing sources as [uid] and linking entities as [[Name]]
key_aspectsYes
input_digestYes
related_entitiesYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this is a write operation (readOnlyHint=false) in a closed world, so the safety profile is covered. The description adds a genuine behavioral constraint not in the annotations — the page must be synthesized only from next_synthesis evidence — but says nothing about validation failures, length limits, or overwrite/merge behavior for an existing entity page.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler, and the core action plus its source constraint is front-loaded. The trailing 'Rules: guide("synthesis")' fragment is terse and slightly cryptic but costs little space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, but this is a 7-parameter write tool with 6 required fields and the description conveys almost none of their intent or the constraints (maxLength/minLength, required evidence). For a complex mutation tool it is under-specified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 43%, so the description is expected to compensate for undocumented parameters, yet it explains none of them. Opaque required fields like entity_id and input_digest are left entirely unaddressed in the description, adding no meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Save) and resource (an entity page), and constrains the source material to next_synthesis evidence, which distinguishes it from the read-side sibling next_synthesis. It stops short of explicitly naming an alternative tool to use instead, so sibling differentiation is implied rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'written only from next_synthesis evidence' implies the workflow prerequisite (gather evidence via next_synthesis first) and the pointer 'guide("synthesis")' routes the agent to rules. However, no explicit when-not-to-use conditions or step sequencing are given, so usage must be partly inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resolve_duplicatesC
Destructive

Merge or keep distinct pairs of possible duplicate entities. Rules: guide("duplicates").

ParametersJSON Schema
NameRequiredDescriptionDefault
projectNoProject name from list_projects; may be omitted when only one project exists
decisionsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and openWorldHint=false, so the safety profile is partly covered. The description adds nothing about irreversibility of merges, what happens to the losing entity, or the 50-decision cap, which matters for a destructive write operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly worded sentences with the operation front-loaded and no filler. The 'guide("duplicates")' pointer is terse but functional rather than padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, batch-write tool with a nested decision object and no annotation-independent detail, the description is too thin. It omits the decision structure, the merge consequences, and any eligibility or ordering constraints, even though an output schema does exist.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50%. The phrase 'merge or keep distinct' loosely maps to the action enum, but the 'keep' field (entity vs candidate), entity_id/candidate_id, and reason are unexplained, so the description does not compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb pair ('merge or keep distinct') and resource ('pairs of possible duplicate entities'), so an agent knows exactly what operation this performs. However, it does not differentiate from siblings like merge_entities or list_duplicate_candidates, which also touch duplicate entities.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance and no comparison to alternatives such as merge_entities. The only pointer is 'Rules: guide("duplicates")', which tells the agent where rules live but not when this tool is the right choice versus its siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

review_articleC

Record the researcher's own include/exclude decision and reason for one article.

ParametersJSON Schema
NameRequiredDescriptionDefault
uidYes
reasonYes
run_idYesRun ID from start_search
decisionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and openWorldHint=false, consistent with the implied write, so there is no contradiction. Beyond that, the description discloses nothing about whether a decision can be overwritten, idempotency on re-review, permission requirements, or the effect of recording a decision on downstream workflow state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with the key noun (include/exclude decision) front-loaded and no filler. Nothing is wasted, though there is also nothing extra offered.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a fully-required four-parameter mutation tool with low schema coverage, the description is too thin: it omits the relationship to batch screening siblings, the consequence of the recorded decision, and any workflow prerequisites. The presence of an output schema excuses it from explaining return values, but not from the remaining gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is low at 25% (only run_id is documented in the schema). The description names the decision and reason fields and implies the article via 'one article', partially compensating, but adds no detail on the include/exclude enum, reason length limits, or the uid reference.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Record') and resource (an include/exclude decision plus reason for one article), so the agent knows exactly what operation this performs. It does not, however, distinguish itself from the sibling record_screening_decisions, which appears to cover the same conceptual space.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this versus next_screening_batch, record_screening_decisions, or set_screening_criteria, nor any stated precondition such as requiring an active search run. The agent must infer its place in the workflow from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_screening_criteriaC

Save the researcher's inclusion and exclusion criteria for screening. Rules: guide("screening").

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYesRun ID from start_search
excludeYesExclusion criteria in the researcher's words
includeYesInclusion criteria in the researcher's words
replaceNoRequired to change existing criteria; starts a new revision and re-screens everything

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false/openWorldHint=false, and 'Save' implies mutation, so the direction is consistent. However, the description omits the significant behavioral trait that supplying replace=true starts a new revision and re-screens everything (only present in the schema), and says nothing about overwrite semantics or required run state. Minimal value added beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Very short and front-loaded, with the purpose stated first. The second sentence ('Rules: guide("screening")') is cryptic internal shorthand that spends a sentence without conveying actionable meaning to an outside agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists and schema coverage is complete, so return values and parameter details are handled elsewhere. What is missing is the usage context (where this fits relative to start_search and record_screening_decisions) for a mutation tool whose replace flag is destructive to prior screening work.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so run_id, include, exclude and replace are all documented in the schema, including the consequential replace behavior. The description adds no parameter-level meaning of its own, which matches the baseline 3 when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: save inclusion/exclusion criteria for screening. This clearly separates it from screening-execution siblings like next_screening_batch and record_screening_decisions. It stops short of naming any sibling explicitly, so 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only usage signal is the opaque 'Rules: guide("screening")' line, which does not state when to call this versus alternatives such as record_screening_decisions, nor any prerequisites or sequencing. An agent can infer this is a setup step but gets no explicit when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_questionA
Read-only

Check a drafted PICO/PCC search question and return a question_id for start_search. Show the result to the researcher and wait for approval. Rules: guide("question").

ParametersJSON Schema
NameRequiredDescriptionDefault
sourcesNo
questionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds a genuinely non-obvious behavioral trait: it is a human-in-the-loop gate where the agent must present the result and pause for researcher approval before proceeding, plus a pointer to the governing ruleset guide('question').

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with purpose and output, then the approval step and rules pointer. No filler, though 'Rules: guide("question").' is telegraphic enough that its intent is slightly opaque.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described, and the HITL step is covered. However, for a tool whose input is a deeply nested ResearchQuestion (framework, components, filters), the description offers only a guide reference and no orientation on constructing or when to reject a question, leaving meaningful gaps at this complexity level.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Top-level schema description coverage is 0%, so the description must compensate for the 'sources' and 'question' parameters, but it says nothing about their format, defaults, or the framework/component structure beyond a terse 'guide("question")' pointer. The only parameter-relevant help is delegated to another tool rather than explained here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Check a drafted PICO/PCC search question') and names the downstream consumer ('return a question_id for start_search'), which distinguishes it from siblings like start_search. An agent knows exactly what this tool produces and where that artifact goes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The mention of returning a question_id 'for start_search' implies this runs before search, and 'Show the result to the researcher and wait for approval' adds a workflow step, but there is no explicit when-to-use clause or exclusion relative to siblings such as start_search or guide. Usage is inferred rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wiki_tasksB
Read-only

Start here to build or update the wiki: the current step as self-contained tasks, each sized for one fresh subagent. Rules: guide("wiki").

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYesRun ID from start_search

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered structurally. The description adds useful consumption context ('each sized for one fresh subagent') but says nothing about rate limits, task ordering, or what happens across steps; the cryptic 'Rules: guide("wiki")' adds little.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences, with the entry-point instruction front-loaded. The trailing 'Rules: guide("wiki")' is terse to the point of ambiguity but cheap in length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the one parameter is fully documented. However, for a workflow-orchestration tool competing with many wiki/synthesis siblings, the description leaves the relationship to next_wiki_article and next_synthesis unclear.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and there is a single run_id parameter already documented in the schema ('Run ID from start_search'). The description adds no further parameter meaning, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action (fetch the current step as self-contained tasks) tied to a resource (the wiki build/update workflow) and marks itself as the entry point. It is distinguishable from siblings like next_wiki_article or guide, though the phrasing 'the current step as self-contained tasks' is slightly abstract.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Start here to build or update the wiki' gives a clear entry-point cue, implying use at the beginning of a wiki workflow. It does not name alternatives (e.g., next_wiki_article, guide) or state when-not to use it, so guidance is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 26 tool updatesv0.1.1
    • First observedcreate_project
    • First observedexport_wiki
    • First observedfetch_articles
    • First observedfind_entities
    • First observedget_article_page
    • First observedget_run_status
    • First observedguide
    • First observedlist_articles
    • First observedlist_duplicate_candidates
    • First observedlist_projects
    • First observedlist_runs
    • First observedmerge_entities
    • First observednext_screening_batch
    • First observednext_synthesis
    • First observednext_wiki_article
    • First observedopen_project
    • First observedrecord_extraction
    • First observedrecord_screening_decisions
    • First observedrecord_synthesis
    • First observedresolve_duplicates
    • First observedresume_search
    • First observedreview_article
    • First observedset_screening_criteria
    • First observedstart_search
    • First observedvalidate_question
    • First observedwiki_tasks

TDQS

B3.2/5.0

Scored across 26 tools

Disambiguation4/5

Most tools target distinct workflow stages (project setup, search, screening, fetching, wiki extraction, synthesis), so an agent can usually tell them apart. A few boundaries are fuzzy, especially between record_screening_decisions and review_article, and between resolve_duplicates and merge_entities, but descriptions provide enough guidance.

Naming Consistency4/5

Nearly all tools use snake_case with a predictable verb-led pattern such as list_articles, create_project, record_extraction, and resolve_duplicates. Minor deviations like guide and wiki_tasks are readable but slightly break the strict verb_noun convention.

Tool Count3/5

26 tools is heavy for a single MCP server, and several stage-specific helpers could likely be consolidated. The complex systematic-review pipeline justifies more tools than a simple CRUD server, but the count is still at the borderline-heavy end.

Completeness4/5

The surface covers the core review lifecycle: projects, question validation, search, screening, fetching, extraction, wiki curation, synthesis, and export. Minor gaps like project deletion/renaming or criteria updates are workable around, but not every lifecycle operation is exposed.

Maintenance

ActivityNo data
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    F
    maintenance
    Enables creation of persistent, compounding knowledge bases using Karpathy's LLM Wiki pattern with LLM-maintained markdown wikis. Supports automated ingestion, cross-referencing, synthesis, and linting of sources as an alternative to traditional RAG systems.
    61
    MIT
  • A
    license
    B
    quality
    A
    maintenance
    AI-operable research workspace integrating Zotero, Obsidian, and NotebookLM. Search papers (arXiv/Semantic Scholar/PubMed/CrossRef), ingest into Zotero, sync per-paper notes to Obsidian, verify NotebookLM briefs. All three external tools optional.
    76
    150 PyPI
    58
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables traceable scholarly literature reviews using free APIs, generating reports where every claim links to evidence IDs.
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables LLM-assisted biomedical literature screening and structured extraction from PubMed alerts, PMIDs, DOIs, and GEO accessions, with full-text retrieval and multi-provider LLM support.
    98 PyPI
    2
    MIT