Lacuna Research MCP
Lacuna Research MCP is a read-only MCP server for exploring a knowledge graph of ML/AI research — papers, works, research directions, authors, venues, institutions, and AI-generated research proposals.
Search the corpus with
search_lacuna: query papers, works, directions, authors, venues, institutions, hypotheses, or resources, with filters for author, date range, venue, result limit/offset, and sort order.Tune ranking: choose
lexical/default(production lexical+semantic for paper relevance searches),semantic(embedding-based, papers only), orbm25_title_abstract; optionally weight specific lexicalfieldsliketitle^4,abstract.Explore novel research ideas with
get_hypothesis, returning either a compact proposal context or the full version history and signal counts.Browse research directions with
get_direction(compact context or raw cluster record) andget_direction_papersfor paginated, citation-ready paper lists.Read papers and works:
get_workgroups versions of the same research;get_paperreads a specific artifact in context, full, preview, blog, figures, concepts, or neighbors views (with configurable figure previews and linked code/resources).Inspect researchers via
get_author_context, plus dedicated paging through an author'sget_author_papers,get_author_directions, andget_author_neighbors.Map the landscape with
get_venue_context(optionally year-scoped) andget_institution_context/get_institution_authors.Ground answers in sources: start from compact agent-oriented views by default, or request
view="full"for raw upstream metadata when needed.
Provides access to machine learning research papers and related metadata from arXiv, enabling search and retrieval of papers within the Lacuna research map.
Provides access to scholarly publication records from dblp, enabling search and retrieval of papers, authors, and venues within the Lacuna research map.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Lacuna Research MCPSearch for recent papers on large language model alignment"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Lacuna Research MCP
🔬 Ground your coding agent in novel ideas, papers, and the ML landscape
Lacuna Research MCP gives AI researchers' coding agents:
Novel research proposals. Explore novel research ideas generated with Alien Science.
Research directions. Navigate concept clusters with linked papers, authors, and proposals.
Agent-ready literature. Search recent papers in markdown with source links.
Researcher intelligence. Trace authors, publications, directions, impact, and related researchers.
Landscape mapping. Compare venues, institutions, leading researchers, and publication activity.
Lacuna, built by Tiptree Systems, is a research map of machine learning: a heterogeneous knowledge graph linking papers, research directions, authors, venues, institutions, and generated research proposals, with a source trail from every derived object back to the exact paper and page that produced it. Its pipeline reconciles scholarly records from OpenAlex, OpenReview, DBLP, and arXiv; extracts concept elements from paper text and clusters them into research directions (Lacuna paper); and samples novel research proposals from those directions with Alien Science. The map spans more than 730,000 papers, 190,000+ author profiles, 38,000+ research directions, and 3,000+ research proposals built from over 15 million concept elements — and it grows continuously as new arXiv and other AI papers are ingested.
Install · First use · Tools · API reference · Configuration
Install
The easiest way to install Lacuna Research MCP is to ask your coding agent, such as Codex or Claude Code:
Install and configure the
lacuna-research-mcppackage from PyPI for this client.
For manual setup, the instructions below use uvx to run the latest tagged release from PyPI. Install uv first; Lacuna Research MCP requires Python 3.11 or newer.
Codex
Add the server with the Codex CLI:
codex mcp add lacuna-research -- uvx lacuna-research-mcpAlternatively, add the following to ~/.codex/config.toml (or to .codex/config.toml in a trusted project for project-only setup):
[mcp_servers.lacuna-research]
command = "uvx"
args = ["lacuna-research-mcp"]Run codex mcp list to verify the server is configured. The Codex app, CLI, and IDE extension share this configuration on the same machine.
Claude Code
Add the server for all of your projects with the Claude Code CLI:
claude mcp add --scope user lacuna-research -- uvx lacuna-research-mcpOmit --scope user to add it only to the current project. Alternatively, add the following under the top-level mcpServers object in ~/.claude.json:
{
"mcpServers": {
"lacuna-research": {
"type": "stdio",
"command": "uvx",
"args": ["lacuna-research-mcp"]
}
}
}Run claude mcp get lacuna-research to verify the server is configured.
Claude Desktop
Open Settings → Developer → Edit Config, then add the server under mcpServers in claude_desktop_config.json:
{
"mcpServers": {
"lacuna-research": {
"command": "uvx",
"args": ["lacuna-research-mcp"]
}
}
}Restart Claude Desktop after saving the file.
Other MCP clients
For any client that supports local stdio MCP servers, use this standard configuration:
{
"mcpServers": {
"lacuna-research": {
"command": "uvx",
"args": ["lacuna-research-mcp"]
}
}
}Standalone command
Install the MCP server as a persistent command:
uv tool install lacuna-research-mcp
lacuna-research-mcpRun it without installing a persistent command:
uvx lacuna-research-mcpIf this exits with a prerelease error (for example, "Python 3.14.0rc2 is a prerelease..."),
your uv is using an early Python 3.14 prerelease that the MCP SDK does not support.
Update uv (uv self update), install a released Python (uv python install 3.14),
or run uvx --python 3.13 lacuna-research-mcp.
Latest development version
PyPI contains tagged releases. To try the latest code from the main branch instead:
uvx --from git+https://github.com/tiptreesystems/lacuna-research-mcp.git lacuna-research-mcpLocal development
With uv:
git clone https://github.com/tiptreesystems/lacuna-research-mcp.git
cd lacuna-research-mcp
uv sync --extra dev
uv run lacuna-research-mcpWith pip:
cd <path-to-lacuna-research-mcp>
python3 -m venv .venv
. .venv/bin/activate
python -m pip install -U pip
python -m pip install -e .Related MCP server: scholar-memory
First use
After connecting the server, call:
search_lacuna(query="LLM jailbreak defense", search_type="hypothesis", limit=10)search_lacuna(query="methods for detecting prompt injection attacks", search_type="paper", limit=10)(production lexical+semantic paper ranking by default)get_hypothesis(hypothesis_id_or_url="bd35de182c2325ae")get_paper(artifact_id_or_url="art_79c57fbfec094f26b79c422cf08fed34")(defaults toview="context")get_direction(cluster_id_or_url=25108)(defaults toview="context")search_lacuna(query="ImageNet", search_type="resource", resource_kind="dataset")(datasets linked to papers; follow up withget_resource)
Scope
The corpus covers machine learning and AI research: papers, research directions, authors' research output, venues, institutions, generated research hypotheses, and research resources (GitHub repositories, Hugging Face datasets and models, Zenodo records) linked to papers. It does not contain biographies, news, or non-research web content. Agents should answer questions outside that scope from other sources.
What it exposes
search_lacunaUses Lacuna's public/api/v1/searchendpoint for directions, papers, authors, venues, institutions, hypotheses, and resources. Explicit paper searches (search_type="paper") use the server's production lexical+semantic ranker when the other ranking arguments remain at their defaults. Passsearch_type="hypothesis"for hypothesis search. Passsearch_type="resource"for code repositories, datasets, models, and demos linked to papers, optionally filtered withresource_kind(codebase,dataset,model,demo) andprovider(github,huggingface,zenodo); for example,resource_kind="dataset"finds benchmark datasets. Resource search rejectsdate_from/date_to,venue, year sorts, and semantic ranking (resources carry no dates, venues, or embeddings).get_hypothesis(hypothesis_id_or_url, view="context")Hypothesis/proposal.view="context"(default) is a compact single-fetch proposal context (summary, abstract, linked directions);view="full"returns the server's version record with version history and signal counts. Proposal bodies are inversions[].markdown.get_direction(cluster_id_or_url, view="context")Research direction/cluster.view="context"(default) requests the compact agent-oriented summary;view="full"returns the raw cluster record.get_direction_papers(cluster_id_or_url, page, limit, view="compact")Paginated papers attached to a direction.view="compact"(default) returns citation-ready rows (id, url, title, year, venue, a few authors, abstract snippet);view="full"returns the raw upstream paper records.get_paper(artifact_id_or_url, view="context", figure_limit=None, include_resources=True)Paper lookup with available versions; use a version’s artifact ID to read it. Linked resources (code repositories, datasets, models, and demos) are included by default in context and full views; passinclude_resources=Falseto omit them. Other views ignore this option.view="context"(default) requests the compact agent-oriented context; other views are"full","preview","blog","figures","concepts", or"neighbors". In context view,figure_limitcaps the figure preview (server default 3; pass 0 to suppress previews while keeping afigures_truncatedsignal).get_resource(resource_id_or_url)Resource lookup by id (from resource search orget_paper'sresourceslist) or Lacuna/resource/...URL: external URL, summary, facets (tasks, modalities, size, license), provider metrics, README/card excerpt, linked papers underpublications(withpaper_idand relationship such asdataset_for), related directions, and mentioned resources.Author tools:
get_author_context(…, view="context"),get_author_papers,get_author_directions, andget_author_neighbors. Start withget_author_context, which defaults to the compact agent-oriented view (capped papers plus a readableimpact_directionslist instead of rawimpact_clusterstelemetry). Use the dedicated papers, directions, and neighbors tools to page through those collections without repeating the author context.view="full"returns the server-bounded full-shape context (collections remain capped at 100). Passinclude_neighbors=trueto explicitly include similar authors; this may add significant server latency.Venue and institution tools:
get_venue_context(…, view="context"),get_institution_context(…, view="context"),get_institution_authors. Context tools default to compact (capped lists, duplicated blocks dropped; venue keeps a recent-activity slice that always includes the requestedyear). Useget_institution_authorsto page through an institution's complete author list.
Wrapped APIs
MCP tool | Lacuna API endpoint |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
For id-or-URL arguments, pass either the raw id returned by search_lacuna or the corresponding Lacuna page URL. The MCP normalizes Lacuna-relative url and *_url fields to absolute URLs, and it also absolutifies Lacuna links (including /resource/... pages) inside fields named summary_markdown, article_markdown, markdown, content, or description.
get_author_context is bounded server-side in both views. The default compact view returns a curated briefing; view="full" returns the larger complete shape with embedded collections capped at 100. Use get_author_papers and get_author_directions rather than trying to page embedded context collections.
Search requests are capped at 50 results per call, and direction-paper page
requests are capped at 100 results per call.
In search_lacuna, date_from and date_to are inclusive publication-date
bounds. Accepted formats are YYYY, YYYY-MM, and YYYY-MM-DD; for example,
date_from="2020", date_to="2022-03" includes papers from January 1, 2020
through March 31, 2022.
search_lacuna exposes these ranking profiles:
default/lexicalThe default profile for all searches. Withsearch_type="paper",sort="relevance", andfieldsunset, it uses the server's production lexical+semantic ranker with graceful fallback. The MCP's defaultsearch_type="all"uses the server's default lexical ranking instead.semanticUse for embedding-based paper retrieval. The semantic query omits the normal lexical ranking leg, but the server can still overlay exact-title lexical matches. Only supported forpaperandallsearches (only papers have semantic embeddings), and not withresource_kind/providerfilters.bm25_title_abstract/bm25Use for lexical matching constrained to title and abstract. Rejected forauthorandinstitutionsearches (those records have no title or abstract fields).
All searches use the server default unless ranking_profile is provided. The MCP rejects unsupported profile/type combinations because the server would otherwise fall back to substring search and silently ignore the requested ranking profile.
sort accepts relevance (default), year_desc, and year_asc. Year sorts cannot be combined with ranking_profile="semantic" — the server would silently ignore the sort — so the MCP rejects that combination; for recent-and-relevant queries, keep semantic ranking and constrain recency with date_from/date_to instead.
Environment variables
LACUNA_SITE_URLDefaults tohttps://lacuna.tiptreesystems.comLACUNA_MCP_TIMEOUTDefaults to30LACUNA_MCP_MAX_RETRIESDefaults to2; applies to timeouts, transport errors, and HTTP429,502,503, and504responses.LACUNA_MCP_USER_AGENTDefaults tolacuna-research-mcp/{package_version}LACUNA_MCP_BEARER_TOKENOptional bearer token sent asAuthorization: Bearer ...for private Lacuna deployments.LACUNA_MCP_LOG_LEVELDefaults toWARNING(one ofDEBUG,INFO,WARNING,ERROR,CRITICAL). The default keeps normal operation quiet; lower it only for debugging, sinceINFO/DEBUGlet the HTTP client log full request URLs — including the search query string — to stderr, which some MCP hosts retain.
Environment variables are read once when the MCP server is created, or on the first direct tool/API call if the module is imported without calling create_mcp().
Implementation layout
The server is a thin MCP adapter over Lacuna's HTTP API. The implementation is split by responsibility:
lacuna_research_mcp/server.pyMCPServer app creation, tool registration, lifespan cleanup, and thelacuna-research-mcpCLI entrypoint.lacuna_research_mcp/tools.pyMCP tool functions. Each tool normalizes its inputs, calls the matching Lacuna API endpoint through the shared client helpers, and returns JSON-compatible data.lacuna_research_mcp/client.pyRuntime HTTP access to Lacuna: the event-loop-boundhttpx.AsyncClient, retry handling, error wrapping, JSON parsing, URL normalization on responses, andapi_get/api_object/api_payload.lacuna_research_mcp/config.pyRuntime constants,RuntimeConfig, package user-agent construction, and environment parsing.lacuna_research_mcp/ids.pyHelpers that accept either raw ids or Lacuna page URLs and extract safe API path segments.lacuna_research_mcp/normalize.pyResponse post-processing for relative Lacuna URLs and markdown links.lacuna_research_mcp/errors.pyUser-facing exception type for Lacuna API access failures.
Notes
get_paperandget_directiondefault toview="context". These context views request Lacuna's compact agent-oriented payloads by default to keep MCP responses small. Paper context includessummary_markdownwhen available (otherwiseabstract), authors, and figures; direction context includessummary_markdown, capped papers/authors/related directions, and truncation markers. Useview="full"when you need the raw metadata, and the other paper views (preview,blog,figures,concepts,neighbors) when you want one isolated sub-resource.An explicit relevance-sorted paper search with no custom
fieldsdefaults to the server's production lexical+semantic ranker. Setranking_profile="semantic"for embedding-based retrieval (with a possible exact-title overlay) or"bm25_title_abstract"for title-and-abstract lexical matching.Most detail tools accept either the id returned by search or the corresponding Lacuna URL.
Relative Lacuna URLs in
url/*_urlresponse fields and fields namedsummary_markdown,article_markdown,markdown,content, ordescriptionare normalized to absolute URLs.Venue and institution keys are opaque hashes (for example
d7bf22905bd6), never human-readable names likeicml. Find the key withsearch_lacuna(search_type="venue")first, or pass a/venue/...page URL.
Citation
If you find our work helpful, feel free to cite the papers behind Lacuna's research-proposal generation and research map.
Research-proposal generation — Alien Science
@inproceedings{artiles2026alien,
title = {Alien Science: Sampling Coherent but Cognitively Unavailable Research Directions from Idea Atoms},
author = {Artiles, Alejandro H. and Weiss, Martin and Brinkmann, Levin and Goyal, Anirudh and Rahaman, Nasim},
booktitle = {ICLR 2026 Workshop on Post-AGI Science and Society},
year = {2026},
url = {https://openreview.net/forum?id=XZWkDET1ia}
}Research map — Lacuna
@misc{weiss2026lacunaresearchmapmachine,
title = {Lacuna: A Research Map for Machine Learning},
author = {Martin Weiss and Miles Q. Li and Alejandro H. Artiles and Yacine Mkhinini and Chris Pal and Hugo Larochelle and Nasim Rahaman},
year = {2026},
eprint = {2606.26246},
archivePrefix = {arXiv},
primaryClass = {cs.DL},
url = {https://arxiv.org/abs/2606.26246}
}License
MIT. See LICENSE.
Available Tools
13 toolsget_author_contextARead-onlyIdempotent
Fetch agent-oriented context for a Lacuna author.
Author profiles describe research output (papers, directions, impact). A
free-text affiliation field may be present but can be incomplete or
outdated and may not represent current employment; the corpus has no
biography or employment history, so do not infer those from this data.
view selects the response shape:
"context" (default, recommended): Lacuna's compact author context — capped readable papers and an
impact_directionslist (named research directions) in place of the raw numericimpact_clusterstelemetry, with the duplicated nested author record dropped server-side."full": the bounded full-shape author context (raw
impact_clusters, nested author record; server collections are capped at 100).
Set include_neighbors=True to include similar authors as named, linkable records. Neighbor computation may add significant server latency.
| Name | Required | Description | Default |
|---|---|---|---|
| view | No | context | |
| author_id_or_url | Yes | ||
| include_neighbors | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate read-only, idempotent, non-destructive behavior. The description adds valuable context: caveats about affiliation data being incomplete/outdated, no biography/employment history, latency implications of include_neighbors, and server-side collection caps (100). This significantly exceeds what annotations alone provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured: a lead sentence, a data-caveat paragraph, and a bulleted view explanation. Every section earns its place, though the affiliation caveat could be slightly more concise without losing meaning. It is appropriately sized for the complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is complete for this tool's complexity: it explains response shape variants, data reliability limits, server caps, and neighbor latency. Since an output schema exists, it need not describe return fields in detail. No major gaps are evident.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, so the description must compensate. It thoroughly explains the 'view' parameter with distinct response shapes and the 'include_neighbors' parameter with latency trade-offs. The 'author_id_or_url' parameter is self-explanatory but lacks format details, which is a minor gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool 'Fetch agent-oriented context for a Lacuna author' with a specific verb and resource. It distinguishes from sibling tools by describing a composite context (papers, directions, impact) rather than a single aspect like papers or directions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool, especially through the view parameter choices ('context' recommended vs 'full') and the include_neighbors option. However, it does not explicitly name alternative sibling tools for specific needs, though the context implies this is the go-to for a holistic author profile.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_author_directionsARead-onlyIdempotent
Fetch one page of an author's named research directions.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| offset | No | ||
| author_id_or_url | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already communicate read-only, idempotent, non-destructive behavior. The description adds the useful behavioral detail that only one page is returned at a time, and that directions are 'named', which is not captured by annotations. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence with the verb first and no unnecessary words. It earns its place entirely, achieving high clarity in minimal space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple paginated read-only tool with a full output schema and robust annotations, the description covers the core purpose and pagination. Missing details include parameter clarification and explicit sibling differentiation, but the basic usage is clear enough for straightforward selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description must compensate. It only hints at limit/offset via 'one page', adding some meaning, but does not explain author_id_or_url or the exact semantics of limit/offset beyond naming. This is minimal compensation for the lack of schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool fetches one page of an author's named research directions, using a specific verb and resource. It distinguishes itself from siblings like get_direction (a single direction) and get_author_papers (papers) by focusing on the author's directions with pagination.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The pagination hint 'one page' implies a use case for paging through results, but the description gives no explicit when-to-use vs alternatives, no exclusions, and does not mention sibling tools. Usage context is implied rather than directly stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_author_neighborsARead-onlyIdempotent
Fetch one ranked page of neighboring/similar Lacuna authors.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| offset | No | ||
| author_id_or_url | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, and non-destructive behavior. The description adds 'one ranked page' to indicate pagination and 'neighboring/similar' to clarify the concept, but it does not disclose additional behavioral traits such as ranking criteria or error handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence with the action verb first and no unnecessary words. It is concise, front-loaded, and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although the tool has an output schema and safety annotations, the description lacks details on what qualifies as a 'neighbor' or how ranking works. It also does not explain the required input parameter, leaving some contextual gaps for an agent to infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain any of the three parameters (author_id_or_url, limit, offset). While parameter names are somewhat self-explanatory, the description fails to compensate for the lack of schema descriptions, adding no extra meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Fetch') and the resource ('neighboring/similar Lacuna authors'), distinguishing it from sibling tools like get_author_papers or get_author_context. The phrase 'one ranked page' also signals pagination, making the purpose specific and clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for retrieving similar authors but does not explicitly state when to use it or contrast with alternatives. No exclusions or alternative tool mentions are provided, so guidance is limited to the implied context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_author_papersARead-onlyIdempotent
Fetch one page of an author's papers, ordered from newest to oldest.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| offset | No | ||
| author_id_or_url | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, non-destructive, and idempotent behavior. The description adds valuable behavioral context by specifying that results are paginated ('one page') and ordered newest-to-oldest, which is not captured in annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no unnecessary words. It efficiently conveys the action, resource, pagination scope, and ordering.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the core purpose and one behavioral detail, but it lacks explicit usage guidelines and parameter explanations. The presence of an output schema and clear annotations partially offsets these gaps, making it minimally viable but not fully comprehensive for a tool with 3 parameters and several siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description needed to compensate by explaining the parameters. It only implies pagination ('one page') without naming limit/offset or explaining their semantics, and it does not clarify the format or valid values for author_id_or_url beyond its name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'Fetch' and clearly identifies the resource as 'an author's papers', with pagination ('one page') and ordering ('newest to oldest'). This distinguishes it from sibling tools like get_paper (single paper) and get_direction_papers (papers for a direction).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies that the tool is for retrieving an author's papers, but it does not explicitly state when to use it over alternatives such as get_author_directions or get_direction_papers. There is no when-not-to-use guidance or mention of alternative tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_directionARead-onlyIdempotent
Fetch a Lacuna research direction/cluster.
Use a numeric cluster ID (for example 25005) or a direction URL ending in that number.
view selects the response shape (context typically contains the fields
full provides plus the agent-oriented summary content):
"context" (default, recommended): compact agent-oriented summary with summary_markdown, capped papers/authors/related_directions, and truncation markers.
"full": raw upstream cluster record only. Cheaper than context when you only need basic cluster metadata.
| Name | Required | Description | Default |
|---|---|---|---|
| view | No | context | |
| cluster_id_or_url | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive), so the bar is lower. The description earns credit by disclosing view behavior: context returns a capped agent-oriented summary with truncation markers, while full returns the raw upstream record and is cheaper for metadata-only needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first line, followed by input format and then a well-structured bulleted breakdown of the two views. Slightly verbose with line breaks, but every sentence carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values needn't be fully described, yet the description still usefully characterizes what each view returns. Input format and view semantics are covered; only sibling routing is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry the load, and it does. It explains cluster_id_or_url accepts either a numeric ID or a URL ending in that number, and it precisely defines both enum values of view, including the default and the trade-off between them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Fetch a Lacuna research direction/cluster.' The 'direction/cluster' wording disambiguates the terminology, but it doesn't explicitly differentiate itself from siblings like get_direction_papers or get_author_directions, so an agent still infers the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear guidance on the view parameter (context default/recommended vs full when only metadata is needed), but says nothing about when to choose get_direction over sibling tools, nor any 'when-not' conditions. Input-format guidance ('numeric cluster ID or URL') is not tool-selection guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_direction_papersARead-onlyIdempotent
Fetch paginated papers associated with a Lacuna research direction/cluster.
view selects the per-paper shape:
"compact" (default, recommended): citation-ready rows (id, url, title, year, venue, a few authors, abstract snippet). Drops the raw upstream info blob and levels.cluster internals that otherwise dominate the payload.
"full": the complete upstream paper records. Much larger; use only when you need the raw metadata (openalex/dblp/arxiv ids, etc.).
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | ||
| view | No | compact | |
| limit | No | ||
| cluster_id_or_url | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent, and the description adds valuable behavioral detail: payload shapes, size implications, and what the compact view discards (raw blob, cluster internals). This goes beyond the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, front-loaded with a clear opening sentence, and uses bulleted view details. No redundant text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only paginated fetch with an output schema, the description covers the two views and their tradeoffs, and mentions pagination. Minor gaps: no explicit discussion of default page/limit behavior, but those are in the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description explains the view parameter in depth but leaves cluster_id_or_url, page, and limit to their schema names. The required parameter's semantics are implied by the tool name but not elaborated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool fetches paginated papers for a Lacuna research direction/cluster, using a specific verb and resource. It distinguishes from siblings like get_paper (single paper) and get_author_papers (author-specific papers).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context for the view parameter, recommending compact by default and reserving full for raw metadata needs. It does not explicitly name alternative tools, but the purpose is unambiguous enough for an agent to infer when to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_hypothesisARead-onlyIdempotent
Fetch a generated novel ML/AI research proposal from Lacuna.
Use after search_lacuna(search_type="hypothesis") to inspect a proposal.
view selects the response shape:
"context" (default, recommended): compact single-fetch proposal context — summary_markdown, abstract, and linked directions, with the raw upstream record (whose markdown duplicates summary_markdown) dropped server-side.
"full": the server's version record, including version history and signal counts. Proposal bodies are in versions[].markdown. Use only when you need versions or signals.
| Name | Required | Description | Default |
|---|---|---|---|
| view | No | context | |
| hypothesis_id_or_url | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds behavioral detail beyond the annotations by explaining what each view returns, including that the 'context' view drops the raw upstream record server-side and that 'full' includes version history and signal counts. It discloses side effects (dropping duplicate markdown) and the content shape, with no contradiction to the readOnlyHint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured: a one-sentence purpose, a usage note, and a bulleted list of views. Every sentence adds value, and the front-loaded structure makes it easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has only two parameters and an output schema, the description covers the essential decision points: what each view returns, which is default/recommended, and the scenario for 'full'. It is complete for an agent to select and invoke the tool correctly, and the output schema covers return values.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The 'view' parameter is thoroughly explained with its two enum values, default, and recommendation. The 'hypothesis_id_or_url' parameter is not explicitly described beyond its name and the note that it is used after search_lacuna, which is somewhat implicit given the 0% schema description coverage. However, the name is self-explanatory and the usage guidance fills in the context, so it is only slightly below full compensation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear, specific verb+resource statement: 'Fetch a generated novel ML/AI research proposal from Lacuna.' It distinguishes the tool from siblings by focusing on proposals and explicitly ties it to search_lacuna, making its purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit when-to-use guidance: 'Use after search_lacuna(search_type="hypothesis") to inspect a proposal.' It also explains when to choose each view, recommending 'context' by default and reserving 'full' for cases needing versions or signals, which is strong usage direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_institution_authorsARead-onlyIdempotent
Fetch one page of authors affiliated with a Lacuna institution.
Results are ordered by paper count. Use authors_total, authors_returned, authors_offset, and authors_truncated to page through the complete list.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| offset | No | ||
| institution_key_or_url | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds meaningful behavioral context beyond annotations: it discloses that results are one page at a time, ordered by paper count, and how pagination works via output fields. This is sufficient transparency for a read-only tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three short sentences, front-loaded with the core purpose. Every sentence earns its place: purpose, ordering, and pagination instructions. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (3 params, no nested objects, output schema present), the description covers the essential points: what it fetches, ordering, and pagination. It could add a note on what institution_key_or_url refers to, but the parameter name and sibling tools provide enough context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does provide pagination-related context but does not explicitly explain the limit/offset parameters or the form of institution_key_or_url. The parameter names are fairly self-explanatory, and the output-pagination fields indirectly imply usage, but the description could be more explicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Fetch') and resource ('one page of authors affiliated with a Lacuna institution'), clearly distinguishing it from sibling tools like get_institution_context or get_author_papers. The scope ('one page') and ordering ('by paper count') are also stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear in-tool guidance: it tells the user how to page through results using authors_total, authors_returned, authors_offset, and authors_truncated. It does not explicitly mention alternatives/exclusions, but given no sibling tool serves the same purpose, the context is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_institution_contextARead-onlyIdempotent
Fetch agent-oriented context for a Lacuna institution.
view selects the response shape:
"context" (default, recommended): compact institution context — capped top authors with the duplicated institution block dropped server-side.
"full": the complete institution context shape (duplicated institution record and a bounded author list with explicit truncation metadata).
Use get_institution_authors to page through the complete author list.
| Name | Required | Description | Default |
|---|---|---|---|
| view | No | context | |
| institution_key_or_url | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnly/openWorld/idempotent, and the description adds meaningful behavioral context: the 'context' view returns a 'capped top authors' list with 'the duplicated institution block dropped server-side,' while 'full' includes 'a bounded author list with explicit truncation metadata.' This goes beyond just stating the operation and clarifies response shape differences.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded, with the purpose in the first sentence. Bullet points efficiently explain the two views, and the final sentence directs to a sibling tool. No filler or redundant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists and annotations cover safety, the description is largely complete: it explains the two response shapes, provides an alternative for full author listing, and gives a default recommendation. The only gap is the lack of detail on the institution identifier parameter, but overall the tool is well-specified for an agent to select and invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, so the description must compensate. The 'view' parameter is thoroughly explained with enum meanings and recommendations, but the required 'institution_key_or_url' parameter lacks any detail (e.g., format or example) beyond its self-descriptive name. The description mentions 'a Lacuna institution' but does not explicitly define what a key or URL looks like.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb + resource: 'Fetch agent-oriented context for a Lacuna institution,' clearly distinguishing it from sibling tools like get_author_context and get_venue_context. It further differentiates this tool from get_institution_authors by noting that the latter is for paging through the complete author list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly recommends the default 'context' view and explains when to use it vs. the 'full' view. It also provides an explicit alternative: 'Use get_institution_authors to page through the complete author list,' which tells the agent when NOT to use this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_paperARead-onlyIdempotent
Fetch a Lacuna paper by artifact id or paper URL.
include_resources adds linked resources (code repositories, datasets, models, and demos) in context and full views by default. Pass False to omit them; other views ignore this option. Pass a resource entry's id to get_resource for further details.
view selects the response shape. context requests Lacuna's compact
agent-oriented context by default, while the four single-field views
(blog/figures/concepts/neighbors) return isolated sub-resources:
"context" (default, recommended): agent-oriented summary with summary_markdown, authors, and a small figure preview. Start here for almost everything. Available versions are included when present; to read another version, use its artifact ID from the returned version list.
"full": raw upstream paper record. Cheaper than context when you only need basic metadata.
"preview": compact card with unique fields
excerpt,excerpt_kind,bookmarked. Use for citation-style display."blog": just the summary_markdown content, without the rest of the context envelope.
"figures": just the figures list.
"concepts": just the concepts list.
"neighbors": just the related-papers list.
figure_limit (context view only) caps the figure preview (server default 3).
Pass 0 to suppress figure previews while keeping a figures_truncated signal.
| Name | Required | Description | Default |
|---|---|---|---|
| view | No | context | |
| figure_limit | No | ||
| include_resources | No | ||
| artifact_id_or_url | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly/idempotent annotations, it discloses non-obvious behavior: include_resources defaults to True for context and full views, other views ignore it, figure_limit defaults to 3 server-side, and passing 0 suppresses previews while retaining a `figures_truncated` signal. It also explains that other versions require reading the returned version list's artifact ID.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the one-line purpose, then a bulleted view breakdown that is scannable. It is somewhat long, and 'other views ignore this option' plus the include_resources explanation repeat the same default, but nearly every sentence carries usable routing detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description need not explain return shapes, and it correctly focuses on view selection, defaults, and cross-tool handoff. For a 7-mode fetch tool, an agent has everything needed to pick a view and call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry all meaning, and it does: every enum value of `view` gets a definition and intended use, and both `include_resources` and `figure_limit` are explained with defaults and edge-case behavior. Only artifact_id_or_url is left to its name, which is self-evident.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Fetch a Lacuna paper') plus the accepted identifier forms (artifact id or paper URL). It also names 'get_resource' as the follow-up tool, so an agent can locate it against siblings like get_hypothesis or get_author_papers without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit routing: "context" is the default and recommended starting point "for almost everything," with per-view guidance on when to prefer full (cheaper, basic metadata), preview (citation display), and the single-field views. It also states the condition for switching views ('Pass a resource entry's id to get_resource').
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_resourceARead-onlyIdempotent
Fetch a Lacuna research resource (code repository, dataset, model, or demo).
Accepts a resource artifact ID (art_...) or a Lacuna resource URL containing that ID (/resource//art_...).
The response includes the external url, a summary, facets (tasks, modalities,
size, license, access), provider metrics, a README/card excerpt,
publications (linked papers with paper_id, title,
venue, year, and relationship such as dataset_for or code_for), related
research directions, and other resources mentioned by this one. Pass a
publication's paper_id to get_paper to read the paper and its reported
results.
| Name | Required | Description | Default |
|---|---|---|---|
| resource_id_or_url | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the burden is light; the description adds real behavioral context by enumerating the response payload (facets, provider metrics, README excerpt, linked publications with relationship types, related directions). It does not cover failure modes for malformed IDs or slugs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the one-line purpose, then input format, then response shape. Every sentence carries information, though the response-field enumeration is long and partially duplicated by the output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so describing return values is redundant-but-helpful rather than necessary. Input format and sibling routing are covered; missing only error/auth behavior, which is minor for a read-only lookup.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry the single parameter — and it does, specifying that it accepts an 'art_...' artifact ID or a Lacuna resource URL of the form /resource/<slug>/art_..., which is the key detail an agent needs to construct a valid call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (fetch) and resource (Lacuna research resource: code repository, dataset, model, or demo), which concretely bounds the tool and separates it from generic siblings like search_lacuna or get_paper.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the input form and routes the agent onward — 'Pass a publication's paper_id to get_paper' — which is genuine usage guidance. It does not, however, state when to choose this tool over search_lacuna or when-not conditions, so it stops short of full when/alternatives coverage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_venue_contextARead-onlyIdempotent
Fetch agent-oriented context for a Lacuna venue, optionally scoped to a year.
Venue keys are opaque hashes (e.g. "d7bf22905bd6"), never human-readable names like "icml". Find the key first via search_lacuna(search_type="venue") or pass a /venue/... page URL.
view selects the response shape:
"context" (default, recommended): compact venue context — capped top authors, non-placeholder top clusters, and a recent-activity slice (the requested
yearis always included), with the duplicated venue block and full year histogram dropped server-side."full": the complete venue context (full year histogram, all top authors/ clusters, duplicated venue record).
| Name | Required | Description | Default |
|---|---|---|---|
| view | No | context | |
| year | No | ||
| venue_key_or_url | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnly, idempotent, non-destructive), it discloses that venue keys are opaque hashes, that the context view drops duplicated venue blocks and full histogram, and that the requested year is always included. This adds meaningful behavioral detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and every sentence provides useful information, with the key format and view modes clearly explained. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers key sourcing, view differences, and year inclusion. The output schema presumably documents return fields, so it's not necessary to explain them. Minor ambiguity in 'agent-oriented context' and lack of explicit year range, but overall complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema coverage, the description compensates well: it explains venue_key_or_url format and how to find it, and details view options. Year gets partial treatment as 'optionally scoped to a year' but lacks explicit type/range, though schema provides the type.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it fetches agent-oriented context for a Lacuna venue, with a specific verb and resource. It distinguishes itself from sibling tools by focusing on venue context, and provides guidance on venue keys.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains how to obtain the venue key via search_lacuna and recommends the 'context' view, but lacks explicit when-not-to-use or alternative tool pointers beyond that.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_lacunaARead-onlyIdempotent
Search Lacuna's ML/AI corpus for papers, research directions, authors, venues, institutions, novel research hypotheses, and research resources (code repositories, datasets, models, and demos linked to papers).
For novel ML/AI research ideas, use search_type="hypothesis", then get_hypothesis on promising results.
search_type accepts all, cluster, paper, author, institution, venue, hypothesis, or resource. Use paper for literature and get_paper to read a result. Use other sources for biographies, news, and non-research web content.
Resources: search_type="resource" searches GitHub repositories, Hugging
Face datasets/models, and Zenodo records that Lacuna has linked to at least
one paper. To find benchmark or evaluation datasets, use
search_type="resource" with resource_kind="dataset" (for example query
"image classification benchmark" or a dataset name like "ImageNet").
resource_kind accepts codebase, dataset, model, or demo (one value or a
list); provider accepts github, huggingface, or zenodo. Both filters require
search_type "resource" or "all" (with either filter set, only resources are
returned). Each result has an external url, a Lacuna context_url, and
linked_paper_count; call get_resource on its id to list the linked papers
(with their relationship to the resource), then get_paper on those papers
for reported results and numbers. Resource search does not support
date_from/date_to, venue, year sorting, or semantic ranking.
ranking_profile accepts:
default / lexical (default): production ranking; relevance-sorted paper searches combine lexical and semantic retrieval when fields is unset.
semantic: conceptual paper retrieval; supported for paper and all.
bm25_title_abstract / bm25: lexical paper matching over those fields.
sort accepts relevance (default), year_desc, or year_asc. Semantic ranking cannot use year sorting; constrain recency with date_from/date_to instead. date_from and date_to are inclusive YYYY, YYYY-MM, or YYYY-MM-DD bounds. author_id_or_url constrains a paper search to one author. Pass an author ID or Lacuna author page URL; it requires search_type="paper".
fields optionally restricts and weights lexical fields, for example "title^4,abstract". Supported names are title, abstract, summary, concepts, name, top_names, and venue, plus description, topics, provider_key, and paper_titles for resources (a resource's title already includes its description, topics, README text, and linked paper titles). Fields must exist on the selected search_type, weights must be within 0 < weight <= 100, and fields cannot be combined with semantic ranking.
| Name | Required | Description | Default |
|---|---|---|---|
| sort | No | relevance | |
| limit | No | ||
| query | Yes | ||
| venue | No | ||
| fields | No | ||
| offset | No | ||
| date_to | No | ||
| provider | No | ||
| date_from | No | ||
| search_type | No | all | |
| resource_kind | No | ||
| ranking_profile | No | ||
| author_id_or_url | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, open-world, non-destructive, but the description adds substantial behavioral constraints beyond them: resource search does not support date_from/date_to, venue, year sorting, or semantic ranking; semantic ranking cannot use year sorting; fields cannot be combined with semantic ranking and weights must fall in 0 < w <= 100. These are exactly the limitations an agent needs to avoid erroring.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then workflow routing, then parameter behavior; each block earns its place given 13 parameters with zero schema documentation. It is dense and occasionally repetitive (resource filters are restated across several sentences), so it stops just short of maximally tight.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and annotations cover the safety profile; the description fills the remaining gaps (filter interdependencies, ranking restrictions, date formats, follow-up tools). An agent has everything needed to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden and largely does: it documents search_type values, resource_kind and provider enums and their dependency on search_type, ranking_profile options and their applicability, sort values, inclusive date bound formats, author_id_or_url semantics and its search_type=paper requirement, and fields syntax with weight bounds. Only limit, offset, query, and venue receive no explanation, which is a minor residual gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search) and resource (Lacuna's ML/AI corpus), and enumerates the entity types retrievable: papers, directions, authors, venues, institutions, hypotheses, and resources. It clearly separates itself from read-style siblings like get_paper and get_resource, which it names as follow-up steps.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing guidance: use search_type=hypothesis for novel ideas then get_hypothesis; use paper for literature then get_paper; use resource with resource_kind=dataset for benchmarks. It also names the conditions under which filters apply (resource_kind/provider require search_type resource or all) and the alternatives for biographies/news.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.5.1- Added
get_resource - Removed
get_work - Changed
search_lacuna2 fields changed- added
Input schema / properties / providerAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "items": { + "type": "string" + }, + "type": "array" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Provider" +} - added
Input schema / properties / resource_kindAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "items": { + "type": "string" + }, + "type": "array" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Resource Kind" +}
3 tool updates
v0.3.0- Changed
get_paper1 field changed- added
Input schema / properties / include_resourcesAdded value: +{ + "default": true, + "title": "Include Resources", + "type": "boolean" +}
- Added
get_work - Changed
search_lacuna1 field changed- added
Input schema / properties / author_id_or_urlAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Author Id Or Url" +}
12 tool updates
v0.2.0- First observed
get_author_context - First observed
get_author_directions - First observed
get_author_neighbors - First observed
get_author_papers - First observed
get_direction - First observed
get_direction_papers - First observed
get_hypothesis - First observed
get_institution_authors - First observed
get_institution_context - First observed
get_paper - First observed
get_venue_context - First observed
search_lacuna
TDQS
Scored across 13 tools
Most tools target clearly distinct entity+facet combinations, and descriptions explicitly distinguish context views from dedicated paginated fetches. However, get_author_context's include_neighbors option overlaps with get_author_neighbors, and context views include capped papers/authors that dedicated tools like get_author_papers and get_direction_papers also provide, creating minor boundary blur.
All tools follow a consistent snake_case verb_noun pattern: search_lacuna for search, and get_<entity>_<facet> or get_<entity>_context for fetches. There is no mixing of camelCase, different verb styles, or irregular naming.
13 tools map cleanly to the seven main entity types (paper, author, direction, venue, institution, resource, hypothesis) plus pagination and search needs. The set is neither bloated nor thin for a read-only research corpus.
Search covers all entity types, and dedicated fetch tools exist for papers, directions, authors, venues, institutions, resources, and hypotheses, with pagination where needed. Minor gaps remain, such as no non-context get_author or get_venue, no batch paper retrieval, but agents can work around these via search and context tools.
Maintenance
Related MCP Connectors
Search GitHub, npm, PyPI, StackOverflow, ArXiv from one MCP — built for coding agents.
Knowledge Network for AI Agents and creators: Search, rate, and review programming guides via MCP
Shared, peer-validated knowledge archive for AI agents — search, contribute, and validate via MCP
- WauldoOAuthcom.wauldo
Stateless agentic tools over MCP: concept extraction, long-context, knowledge graph, planning.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to search code by meaning, explore codebase structure, store and query knowledge with temporal facts, and read source code through a set of MCP tools.310 npm7MIT
- AlicenseNot gradedqualityDmaintenanceEnables scientific literature research through multi-agent search, analysis, and semantic memory, exposing 9 MCP tools for querying, storing, and retrieving research findings.1MIT
- AlicenseAqualityCmaintenanceEnables AI agents to search and retrieve academic papers, author profiles, and citation data from the Scopus database via MCP tools.7MIT
- AlicenseNot gradedqualityBmaintenanceEnables coding agents to explore a repository map via MCP, with tools for briefs, scoping, symbol lookup, module details, and freshness checks.8 npmMIT