Skip to main content
Glama

Server Details

Multi-LLM entity enrichment: schemas, single/batch enrichment, fusion, model benchmarks.

Ownership verified
Status
Healthy
Uptime
14.7% over 43 days
OAuth
Works in Glama
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL
Repository
TOT-Concept/mcp-server-entity-enricher
GitHub Stars
0
Server Listing
mcp-server-entity-enricher

TDQS

A4.4/5.0

Scored across 58 tools

Disambiguation4/5

Despite the large surface, most tools have clearly distinct roles: schema property add/update/move/nest, semantic concept CRUD/merge/probe, database sync lifecycle, and benchmarking are well-separated by their descriptions. Some overlap remains in the schema and semantic clusters (e.g., add vs update vs move property; analyze_sample vs analyze_schema vs classify_database_model), but detailed descriptions make selection tractable.

Naming Consistency5/5

Virtually all 58 tools follow a consistent snake_case verb_noun convention (list_schemas, create_database_sync, update_schema_property, delete_benchmark_scenario). Minor multi-word names like sync_records_to_database still fit the pattern. There is no mixing of case styles or vague verbs.

Tool Count2/5

58 tools is far above the sensible range even for a complex platform covering enrichment, schema editing, semantic IDs, database sync, benchmarks, and attachments. The surface is too large to navigate efficiently and many operations could be consolidated without losing capability.

Completeness5/5

The surface covers the full lifecycle of its domains: schema CRUD and granular property editing, semantic concept creation/probing/merging/migration, database sync registration, fetching, acknowledgement and teardown, benchmarking create/run/score/reference, attachment upload/delete, job control, and enrichment single/batch/retry/fusion. Only minor gaps exist (e.g., no list/get attachment), which agents can generally work around.

Available Tools

58 tools
ack_database_deltasAcknowledge database deltasA
Destructive
Inspect

Acknowledge every delta through up_to_id after successful application, releasing its lease. Acknowledgement can permanently purge delivered deltas and entity state according to registration options; never acknowledge unapplied or failed work. Returns acknowledged and purged counts. No LLM call. Apply/ack workflow: enricher://docs/database-sync.

ParametersJSON Schema
NameRequiredDescriptionDefault
up_to_idYesAcknowledge every delta with id <= this value.
database_idYesDatabase sync UUID.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Even though destructiveHint=true already flags danger, the description adds genuinely important behavioral context: acknowledgement can permanently purge deltas and entity state, it releases a lease, and it returns acknowledged/purged counts. It also states 'No LLM call', which helps an agent decide how to invoke it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, each earning its place: core action, side effects, return summary, and workflow reference. The most decision-relevant warning is front-loaded early.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the 100% schema coverage, existing output schema, and destructive annotation, the description covers everything needed for correct invocation: lease semantics, purge risk, success criteria, return counts, and a docs pointer. No critical gap remains.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters well. The description reinforces that up_to_id bounds the acknowledgement range and mentions returned counts, but it does not add much parameter-level meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb (acknowledge), a precise resource (database deltas up to up_to_id), and states the postcondition (releasing its lease). This clearly differentiates it from sibling tools like fetch_database_deltas and sync_records_to_database.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when to use the tool—after successful application of deltas—and gives a strong when-not: never acknowledge unapplied or failed work. It points to the apply/ack workflow docs, though it does not name sibling alternatives directly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_schema_propertyAdd schema propertyA
Destructive
Inspect

Add a property under the root (parent_path=''), an object path or '$defs.X'. Accepts a scalar, nested object or reference to an existing $defs entity or $enums vocabulary. Requires editor; no LLM call. Omitted nullable means required at enrichment. Adding inside $defs affects every usage site. The same validation and working-copy/publication rules as update_schema_property apply. Read the container with get_schema_part first. Property format: enricher://docs/schema-reference.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesNew property name (letters, digits, underscores).
schema_idYesUUID of the saved schema.
definitionYes{type|ref, description?, examples?, nullable?, flags?, properties?} — 'properties' nests the same shape per child for an inline object.
parent_pathNo'' = root, or an object path / '$defs.X'.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the annotations by disclosing that omitted nullable means required at enrichment, that adding inside $defs affects every usage site, and that the tool follows the same validation and working-copy/publication rules as update_schema_property. These are meaningful behavioral side effects an agent needs to know.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is information-dense but every sentence earns its place: primary operation, accepted values, prerequisites, side effects, and reference to documentation. It is front-loaded with the most important scoping information and contains no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with an output schema and rich annotations, the description covers prerequisites, permissions, side effects, and cross-tool references. It even points to the property format documentation. Nothing essential for invoking the tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already covers 100% of parameters, so the baseline is 3. The description adds extra semantics beyond the schema by explaining accepted definition forms (scalar, nested object, reference to $defs or $enums) and the nullable default behavior, which helps the agent construct valid input.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Add a property') and immediately scopes where it applies: root, object path, or '$defs.X'. It also distinguishes the operation from siblings like update_schema_property and move_schema_property by focusing on the act of adding, not modifying or moving.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear usage context: requires editor, no LLM call, read the container with get_schema_part first, and the same validation/publication rules as update_schema_property apply. It does not explicitly state exclusions or name alternatives for when to use update instead of add, but the purpose is clear enough for an agent to select it correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_semantic_conceptAdd semantic conceptAInspect

Add an identity concept at zero usage, or add text as an alias using alias_of. Requires editor; resolution may incur embedding/judge cost. Probe first; if the text already resolves to an incumbent, offer that concept instead of blindly retrying. Aliases resolving to another concept are refused. embedding_model selects a new type's space only. Returns concept details and link; manage existing aliases with update_concept_alias. See enricher://docs/semantic-ids.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesThe identity text to add.
alias_ofNosemantic_id of the concept this text is a surface form of; omit to add a standalone concept.
judge_floorNoSimilarity at or above which a candidate is put to the identity judge. Omit to use the organization default (Settings → Organization).
concept_typeYesConcept type the text belongs to.
embedding_modelNoComposite key (provider::model) to embed a NEW concept type under.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide readOnlyHint=false and destructiveHint=false, but the description adds meaningful behavioral detail beyond that: editor permission is required, embedding/judge costs may be incurred, resolution refusal behavior, and the fact that embedding_model only affects new concept types. This gives the agent important execution-time expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence earns its place: core purpose, prerequisites/costs, best-practice probing, refusal behavior, parameter scoping, return-value pointer, and pointer to documentation. It is front-loaded with the main purpose and remains compact despite covering multiple operational considerations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, an output schema exists, and annotations cover mutability and destructiveness, the description is complete. It addresses prerequisites, costs, edge cases, alternatives, and even points to further documentation. Nothing critical for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already covers all parameters at 100%, so the baseline is 3. The description adds value by explaining alias_of semantics, the refusal behavior for conflicting aliases, and the special scoping of embedding_model to new types only. Most parameters remain schema-described, but the added context justifies a score above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Add an identity concept at zero usage, or add text as an alias using alias_of.' It clearly distinguishes the two modes of this tool and implicitly separates it from sibling tools like delete_semantic_concepts and update_concept_alias.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use and when-not-to-use guidance: 'Probe first; if the text already resolves to an incumbent, offer that concept instead of blindly retrying,' and 'Aliases resolving to another concept are refused.' It also names the alternative for managing existing aliases: update_concept_alias.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

analyze_sampleAnalyze sampleAInspect

Analyze sample property ambiguity and relationship identity scoping before schema generation. Requires editor; billed analysis persists a record but does not modify the sample. Returns findings with competing interpretations and suggested_names, plus identity_scoping for sites mixing an entity's own facts with facts about its parent relationship. Names are judged in their parent context, including missing units, periods or ranges. Apply only approved corrections before create_schema_from_sample. Optional: generate_sample already checks its initial sample. How to interpret findings: enricher://docs/schema-from-sample.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoModel composite key. 'auto' (default) lets the server pick the org's default schema-generation model.auto
sample_jsonYesThe sample entity object to analyze.
protected_fieldsNoLeaf names you own (still flagged, but no rename is proposed for them).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already set readOnlyHint=false, openWorldHint=true, destructiveHint=false. The description adds valuable detail: it persists a record (so not read-only in the strict sense) but does not modify the sample, and it is billed. This clarifies the exact side effects beyond the annotations. There is no contradiction; the description enriches the behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is informative but a bit dense with several clauses and a doc link. It front-loads the purpose and then covers side effects, usage, and output. While not minimal, every sentence contributes value—no wasted words. It earns a 4 for being structured and efficient without being overly terse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (3 parameters, nested objects, output schema exists), the description covers prerequisites, side effects, usage context, and output semantics. It even points to a documentation resource for interpretation. The output schema is present, so return values are structured. Nothing critical is missing for an agent to correctly invoke and understand the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description does not add per-parameter detail beyond what the schema provides. It mentions 'protected_fields' indirectly via 'Leaf names you own,' but that is also in the schema. It does not explain the 'model' parameter or the exact structure of 'sample_json' beyond the schema. Thus it meets the baseline but adds no extra semantic value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-resource pair: 'Analyze sample property ambiguity and relationship identity scoping before schema generation.' This clearly distinguishes it from siblings like analyze_schema and generate_sample. It also states what it produces (findings with competing interpretations and suggested_names) and the scoping output, making the tool's purpose unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when to use the tool: 'before schema generation' and 'Apply only approved corrections before create_schema_from_sample.' It also notes a prerequisite ('Requires editor') and a side effect ('billed analysis persists a record but does not modify the sample'). It references an alternative: 'Optional: generate_sample already checks its initial sample,' which helps the agent decide whether this tool is needed. This is thorough guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

analyze_schemaAnalyze schemaAInspect

Analyze a saved schema's property ambiguity and relationship identity scoping, writing annotations to the schema. Requires editor; billed. Incremental by default; force=true rechecks all sites. Returns findings, suggested descriptions, identity-scoping annotations and pending unification proposals; it does not apply suggested structural changes. Fails with ambiguity_check_disabled if the feature is off. Use update_schema_property for approved descriptions or renames; review structural changes and publish them when linked. Interpretation and modeling guidance: enricher://docs/schema-reference.

ParametersJSON Schema
NameRequiredDescriptionDefault
forceNoRe-analyze every property, not just unannotated ones.
modelNoModel composite key. 'auto' (default) lets the server pick the org's default schema-generation model.auto
schema_idYesUUID of the saved schema.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description carries substantial behavioral context beyond the annotations (readOnlyHint=false, openWorldHint=true, destructiveHint=false). It discloses the mutation target ('writing annotations to the schema'), a negative behavior ('does not apply suggested structural changes'), the incremental-vs-force behavior, billing, the required role, and a specific error condition — all beyond what the annotations provide. No contradiction with annotations exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but front-loaded with the core purpose, and every subsequent sentence earns its place: prerequisite, billing, behavior control, return values, exclusions, error condition, sibling routing, and a docs pointer. It is longer than typical, but the tool's complexity (mutation, incremental modes, error states) justifies the length; a leaner pass was possible but the structure is effective.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Highly complete for a complex tool. An output schema exists (so return values need not be spelled out), yet the description still summarizes what it returns, states what it does not apply, covers the error case, prerequisites, and the alternative tool, and points to interpretation guidance docs. An agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds modest value by clarifying the default behavior of force ('Incremental by default; force=true rechecks all sites'), which the schema does not explicitly state — it only documents the default false value. This reinforces the schema rather than merely repeating it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb plus resource and action: 'Analyze a saved schema's property ambiguity and relationship identity scoping, writing annotations to the schema.' It then explicitly states what the tool does NOT do ('it does not apply suggested structural changes'), which distinguishes it from mutation-heavy siblings like update_schema, move_schema_property, and update_schema_property.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit routing to an alternative with a condition: 'Use update_schema_property for approved descriptions or renames; review structural changes and publish them when linked.' It also states a prerequisite ('Requires editor') and documents a failure mode ('Fails with ambiguity_check_disabled if the feature is off'), giving an agent clear when-to/when-not-to context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

answer_job_questionAnswer job questionAInspect

Resume a paused job with answers to the questions returned under pause. answers maps question IDs to {option_ids: [...], text: ...}; omitted questions use defaults. Relay consequential unanswered choices to the user unless those defaults were already authorized. Resuming may continue billed model work. Returns the next pause, terminal result or running status after wait_seconds; poll get_job_status if still running. Works with generate_sample clarification in both knowledge and source modes. See enricher://docs/documents.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYesPaused job ID.
answersNoMap of question id -> {option_ids: list[str], text: str | null}. Omit to resume with the planner's defaults.
wait_secondsNoHow long to wait for the job's next pause or completion before returning (0 = return immediately after resuming).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the annotations by disclosing that resuming 'may continue billed model work', that omitted answers use defaults, and that consequential unanswered choices should be relayed unless already authorized. It also explains the return behavior and polling fallback, which the annotations do not cover.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the core action. Each sentence adds necessary information: answers structure, default behavior, cost implication, return behavior, polling alternative, and compatibility. There is no filler or redundant restatement of schema fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, the description covers the input contract, default behavior, cost implications, return semantics, and follow-up polling. It also references enricher://docs/documents for further detail. Nothing critical appears missing for an agent to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema coverage is 100%, the description adds meaningful semantics: it defines the answers map structure as '{option_ids: [...], text: ...}', explains that omitted questions use defaults, and clarifies wait_seconds behavior by describing what is returned after the wait. This enriches the schema significantly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Resume a paused job with answers to the questions returned under pause.' This clearly differentiates the tool from siblings like cancel_job and get_job_status, and the title 'Answer job question' is expanded into a concrete action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description states when to use the tool: when a job is paused and questions were returned. It also provides routing guidance by saying to 'poll get_job_status if still running' and mentions compatibility with generate_sample clarification. It lacks an explicit when-not-to-use statement, but the context is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

assign_sync_hostAssign sync hostA
Destructive
Inspect

Assign or clear the host provisioning a database sync. Requires owner and a sync-enabled plan. host accepts a connected host ID/name; null unassigns it. Assignment may create the physical database and start synchronization automatically. Moving hosts revokes the old host's minted credential; an existing manual pairing is not evicted. Candidates come from create_database_sync or list_database_syncs. No LLM call. Managed and manual setup: enricher://docs/database-sync.

ParametersJSON Schema
NameRequiredDescriptionDefault
hostNoSync host id or name; null clears the assignment.
database_idYesDatabase sync UUID (from list_database_syncs).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare safety flags, so the description carries the burden of behavior. It discloses that assignment may create the physical database and start synchronization automatically, revokes the old host's minted credential on host move, and does not evict existing manual pairings. This is rich, non-obvious context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core operation and each subsequent sentence adds meaningful behavioral or prerequisite detail. It is dense but not bloated, and the doc link provides optional depth without cluttering the main guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutating tool with an output schema, the description covers prerequisites, side effects, parameter sourcing, and gotchas like credential revocation and manual pairing preservation. Nothing critical is missing for an agent to invoke this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema coverage is 100%, so the schema already documents both parameters. The description adds some nuance by explaining null means unassign and naming candidate sources, but most parameter meaning is already present in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb pair, 'Assign or clear,' and names the exact resource, 'the host provisioning a database sync.' It clearly distinguishes this from related tools like create_database_sync and list_database_syncs by specifying that candidates come from those tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit prerequisites (owner, sync-enabled plan), parameter semantics (null to unassign), and candidate sourcing from create_database_sync or list_database_syncs. It does not fully enumerate when to avoid this tool versus sibling assignment alternatives, but the guidance is otherwise clear and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cancel_jobCancel jobA
Idempotent
Inspect

Request cancellation of a pending, running or paused LLM job. In-flight model calls may finish and persist records; remaining work is skipped. Cancellation does not roll back records or database writes. Read get_job_status and list_records(job_id=...) afterward to inspect the outcome. No new LLM call is started by this tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYesJob ID to cancel.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds important behavioral details beyond annotations: in-flight model calls may finish and persist records, remaining work is skipped, cancellation does not roll back writes, and no new LLM call is started. These consequences are material and not derivable from the annotation hints alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense, purposeful sentences: action and state scope first, then behavioral consequences, then follow-up guidance. There is no filler, repetition of schema fields, or unnecessary background.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a single well-documented parameter and an output schema, the description covers action, effects, persistence semantics, side effects, and recommended follow-through. Nothing needed for correct invocation or interpretation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter, job_id, is already fully described in the schema as 'Job ID to cancel' with 100% coverage. The description references job_id in the follow-up example but adds no new semantic detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Request cancellation of a pending, running or paused LLM job.' It also distinguishes itself from read/delete-like operations by explicitly noting cancellation is not a rollback. This leaves no ambiguity about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear post-conditions and next steps: inspect outcome via get_job_status and list_records(job_id=...). It implies when to use the tool, though it does not explicitly enumerate alternatives or cases where cancellation is inappropriate. The context is strong enough for an agent to select it correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

classify_database_modelClassify database modelAInspect

Start a billed analysis proposing database keys, SQL types, indexes and relationship ownership on a linked schema. Requires editor. Registration already starts the initial pass; use this after relevant edits. Incremental scope covers new properties or changed JSON types/multilingual flags; an empty scope does not rerun unchanged fields. Returns job_id: poll get_job_status and inspect get_schema's working copy. Correct proposals with property tools, then publish_schema to ship changes. Entity-level indexes and ownership choices: enricher://docs/database-sync.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoModel composite key. 'auto' (default) lets the server pick the org's default schema-generation model, falling back to the model that generated the schema.auto
schema_idYesSaved schema UUID to classify.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses billing, editor permission, asynchronous job_id return, incremental scope behavior, and that proposals are written to get_schema's working copy until publishing. These behavioral traits go well beyond the annotations and do not contradict readOnlyHint, openWorldHint, or destructiveHint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence carries distinct value: purpose/cost, permission, timing, incremental behavior, response handling, and follow-up workflow. It is front-loaded with the core action and remains dense without being repetitive.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a billed, asynchronous analysis with downstream editing and publishing, the description covers cost, privileges, job-result handling, incremental rerun semantics, corrective workflow, and documentation for entity-level choices. Little is left for an agent to infer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema provides 100% description coverage for both parameters, so the schema already explains schema_id and model. The description does not add significant parameter-level meaning beyond that, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Start a billed analysis') and concrete deliverables ('proposing database keys, SQL types, indexes and relationship ownership on a linked schema'). It is clearly distinguishable from generic analysis siblings by naming the job workflow and what the analysis produces.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when to use it ('use this after relevant edits'), when it is unnecessary ('empty scope does not rerun unchanged fields'), and what prerequisites apply ('Requires editor'). It also sequences follow-up tools: poll get_job_status, inspect get_schema, correct with property tools, then publish_schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_benchmark_scenarioCreate benchmark scenarioAInspect

Create a reusable benchmark with a mandatory scoring judge. scenario_type='enrichment' needs schema_id and entity_data (the entity to enrich, as enrich_entity takes it — refused when it carries no value; never put it in description); 'sample_generation' needs sample_request; 'schema_generation' needs entity_samples (1..20 samples of one entity type). Enrichment and schema generation need a verified gold reference via set_benchmark_reference before running. Sample generation is rubric-scored and takes no reference. Requires owner and a plan with benchmarks; creating the scenario does not run the models. Returns the scenario and link. Next call run_benchmark when its reference requirements are satisfied. See enricher://docs/model-benchmark.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
languageNosample_generation: output language for names + values.en
strategyNoEnrichment: pinned strategy (no 'auto'): single_pass | expert_domains | multi_expertisesingle_pass
languagesNoEnrichment: defaults to ['en'].
schema_idNoUUID of the saved schema to enrich against (enrichment only, required there).
descriptionNoFree-text note shown in the Benchmarks tab; no model ever reads it. The entity to enrich goes in entity_data, never here.
entity_dataNoEnrichment (required there): the fixed entity input every model enriches — the same JSON enrich_entity takes. Read get_schema.input_contract first: identifying fields, preserve paths and keys for supplied array items. Refused when it carries no value.
repetitionsNoRun each model N times per run; keeps mean + consistency spread.
scenario_typeNoenrichment | sample_generation | schema_generation (immutable).enrichment
attachment_idsNoAttachment UUIDs included in every run.
entity_samplesNoschema_generation: the 1..20 fixed input samples (JSON objects of one entity type) every model converts to a schema. Several samples let the scoring read evidence the reference cannot state alone (nullable, types, identity).
sample_requestNosample_generation: the free-text sample request every model answers — the kind of entity, what the sample must contain, any budget or structural preference (same contract as generate_sample's request).
typical_objectNosample_generation: a specific instance to model (e.g. 'Serena Williams').
enable_web_searchNosample_generation: ground values with the model's web search.
naming_conventionNosample_generation: auto | snake_case | camelCase.auto
generate_semantic_idsNoschema_generation: add semantic_id properties to keyed objects.
scoring_judge_model_keyYesLLM judge composite key used to score results (required).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint=false annotation, the description discloses that the tool only creates and does not run models, requires an owner and plan, returns the scenario and link, and refuses entity_data that carries no value. It also surfaces the preconditions for later execution, giving a clear behavioral contract.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every clause carries routing or behavioral information; the mandatory judge and scenario-type dispatch are front-loaded, and no sentence is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 17-parameter tool with three scenario modes, the description provides type-specific requirements, reference preconditions, owner/plan prerequisites, a link to docs, and a follow-up call, while the output schema covers return values. Nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already high (94%), but the description adds cross-tool semantics: entity_data is 'the same JSON enrich_entity takes', sample_request follows 'same contract as generate_sample's request', entity_samples are '1..20 samples of one entity type', and the entity must 'never' go in description. This materially improves parameter understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb-resource pair: 'Create a reusable benchmark with a mandatory scoring judge.' It then enumerates the three scenario_type branches and references related tools, making it clearly distinct from execution tools like run_benchmark and update/delete siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives direct when-to-use rules per scenario_type ('needs schema_id and entity_data', 'needs sample_request', 'needs entity_samples'), states the gold-reference prerequisite via set_benchmark_reference, and names the next action (run_benchmark) once reference requirements are met. This is explicit routing with alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_database_syncCreate database syncAInspect

Register a saved schema for relational synchronization to PostgreSQL, MySQL or SQLite. Requires owner and a sync-enabled plan. Starts a billed classification job when available; returns database ID, classification_job_id or classification_skipped, stamped_keys and registration_notices. The webhook signing secret stays outside the MCP response and can be managed in the web app. The schema is initially unpublished: wait for classification, review keys/options/notices, then publish_schema before any data can sync. pk_strategy locks once the physical model ships. Owned child arrays replace previous membership, so omitted children are deleted on re-enrichment. purge_entity_state transfers custody to replicas; relay custody_warning. A connected host may provision automatically; otherwise get_database_setup_instructions starts browser-confirmed client pairing. Modeling, publication and delivery: enricher://docs/database-sync.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesName of the physical database (existing on the user's server, or created by `ee-database run --create-missing`); shared by every schema linked to this database sync. Also used (snake-cased) as the conventional replica database name by the ee-database CLI.
dialectNoSQL dialect the deltas are rendered in: postgres | mysql | sqlite.postgres
on_gapsNoAdmission gate — what is written when an enrichment has gaps (non-nullable fields unfilled). 'reject_entity': nothing — one gap anywhere refuses the whole enrichment. 'skip_children' (default): the entity without its incomplete children — an array item is dropped, a shared 1-1 reference is detached (child not written, parent's foreign key NULL); gaps with no such child above them still reject. 'accept_partial': everything — gaps land as NULLs, and under last-write-wins a later partial run erases what an earlier one filled.skip_children
schema_idYesSaved schema UUID to connect the database to.
pk_strategyNosurrogate (default): physical surrogate IDs with unique schema keys; natural: use schema keys as physical primary keys and restrict later re-keying. Locks once the physical model ships; decide before publication. See enricher://docs/database-sync.surrogate
target_hostNoSync host (id or name) that should provision and sync this registration automatically (managed ee-database mode). Omitted: auto-assigned when exactly one eligible host is connected; otherwise the response's connected_hosts lists the candidates — relay the choice to the user and call assign_sync_host, or fall back to get_database_setup_instructions for browser-confirmed manual pairing.
key_languageNoLanguage of multilingual identity tokens (ISO 639-1), shared by all schemas of this database. Usually defaults to schema language when needed; publication may request it after classification. Locked once chosen for the database.
purge_on_ackNoDelete delivered delta copies once acknowledged.
index_scalarsNonone: identity/feed indexes; keys: add natural keys and relation access paths; filterable (default): add dates, search intent, spatial/range roles and entity query indexes; all: every scalar. Later changes queue index migrations.filterable
notify_debounce_sNoQuiet period (seconds) before a delta-available notification fires: each new delta resets the timer, so a burst is announced once. Default 5 for MCP callers (agent workflows expect near-immediate reaction); the web app defaults to 30 to coalesce human-scale editing bursts.
propagate_not_nullNoMirror required schema fields as SQL NOT NULL. Omitted: on for strict gap policies, off for accept_partial. True with accept_partial is refused. Later tightening requires validating existing replica rows; violations can quarantine the migration.
purge_entity_stateNoDelete fully delivered entity state after every linked database acknowledges it. Transfers custody to replicas, makes server snapshots incomplete and limits duplicate checks. Relay custody_warning and obtain agreement before enabling. Default false.
purge_on_ack_delay_daysNoGrace period for purge_on_ack: keep acknowledged delta copies this many days (from acknowledgement) before the hourly purge deletes them. Omit to delete them at acknowledgement. Bounded by the plan's delta retention ceiling.
pattern_index_localized_keysNoAdd per-language pattern-match indexes on indexed localized keys/labels (default false). Useful for prefix autocomplete; increases write/index cost. Concrete SQL depends on the selected dialect.
purge_entity_state_delay_daysNoGrace period for purge_entity_state: a fully-delivered entity row is kept until it has gone this many days without an update, then the hourly purge deletes it. Omit to delete it as soon as every database of the schema acknowledged it.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Despite sparse annotations, the description discloses many non-obvious side effects: it starts a billed classification job, leaves the schema initially unpublished, locks pk_strategy once the physical model ships, replaces owned child arrays on re-enrichment, and transfers custody for purge_entity_state. It also warns that the webhook signing secret is not returned and instructs relaying custody_warning. This goes far beyond what the annotations convey and does not contradict them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but information-dense: each sentence introduces a distinct fact such as prerequisites, billing, return fields, unpublished state, key locks, child replacement, purge semantics, host provisioning, and a docs pointer. It is front-loaded with the core registration purpose before moving to follow-up behavior. No filler sentences are present, so the length is justified for a tool with this many parameters and side effects.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with this many parameters and external effects, the description covers all major decision points: plan requirements, billing, lifecycle, locks, child replacement, purge custody, host provisioning, and documentation. The output schema exists and the description still lists the primary return fields. It is complete enough for an agent to call the tool correctly and know what to do next.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already documents all 15 parameters with 100% coverage, so the baseline is 3 and the description does not need to repeat parameter details. The description adds lifecycle context around pk_strategy locking and purge_entity_state custody transfer, but these largely echo the schema's own parameter descriptions. It therefore adds marginal parameter-level value without needing compensation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific action and object: register a saved schema for relational synchronization to PostgreSQL, MySQL or SQLite. This makes the tool's scope unambiguous and distinguishes it from siblings like create_schema_from_sample, publish_schema, and get_database_setup_instructions by centering on registration of an existing schema. The billing and publication lifecycle details reinforce the purpose rather than obscure it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states hard prerequisites: requires owner and a sync-enabled plan. It also lays out the correct workflow: wait for classification, review keys/options/notices, then publish_schema before any data can sync. It names explicit fallbacks such as get_database_setup_instructions for browser-confirmed pairing when no host provisions automatically, so an agent knows when this tool applies versus the surrounding lifecycle tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_schema_from_sampleCreate schema from sampleAInspect

Generate and auto-save a schema from reviewed samples, returning schema_id, schema content and record links. Supply entity_samples (or samples_csv), sample_record_id, or both; explicit samples override stored JSON while attachment inheritance is preserved. Samples must describe one entity type in one language. Generation combines their fields and annotates relationships; it does not redesign the approved structure. Resolve consequential modeling choices and whether to generate semantic IDs before calling; semantic IDs need an organization embedding model and add cost. Requires editor; synchronous and billed. Review returned suggestions before applying edits with property tools. Sample review and canonicalizations: enricher://docs/schema-from-sample; schema format: enricher://docs/schema-reference.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoModel composite key, or auto (default) for the organization task selection. Explicit keys are discovered through list_models.auto
languageNoLanguage of schema type names/descriptions and annotations; defaults to the sample-key language. Sample property names are not translated. Separate from enrichment output languages.
samples_csvNoSamples as CSV text instead of entity_samples (never both): the first row is ALWAYS the header, each data row one sample. Delimiter ',', ';' or tab; headers become identifier keys ('Author Name' -> author_name); each column gets one type (integer, number, boolean or text; empty cell = null; decimal commas read in ';'/tab text). Over 20 rows, 20 are kept, covering every column. Read the result's csv_import and relay the kept rows and renamed headers.
attachment_idsNoUUIDs used as source context and persisted for regeneration. With sample_record_id, omit to inherit its linked attachments; pass [] to deliberately use none, or a non-empty list to override.
entity_samplesNo1..20 reviewed instances of one entity type. Required without sample_record_id. Explicit samples override stored JSON while keeping attachment inheritance. Fields are unioned; missing/null observations become nullable. Use consistent names across samples and array items. Single-language data only; set multilingual flags after generation with update_schema_property.
timeout_secondsNoWall-clock cap; returns a timeout error past this.
sample_record_idNoUUID of a successful sample_generation record. Its stored sample(s) are used when entity_samples is omitted, and its linked attachments are inherited when attachment_ids is omitted.
generate_semantic_idsNoAdd semantic IDs to eligible keyed objects. Requires an organization embedding model and adds resolution cost. Recommend for reusable identities without stable machine keys; obtain agreement before enabling unless already authorized. Defaults false. Modeling guidance: enricher://docs/schema-from-sample.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well past the annotations: discloses that it auto-saves, is synchronous and billed, requires editor role, that semantic IDs need an org embedding model and add cost, and that generation will not redesign the approved structure. The mutation ('auto-save') is consistent with readOnlyHint=false and destructiveHint=false. This is exactly the extra context annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, inputs, and constraints, and every sentence is actionable (inputs, override rules, cost, role, docs link). It is somewhat long and partially overlaps the schema parameter descriptions, but there is little true filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, and annotations cover the safety profile. The description still supplies the workflow, prerequisites, billing/cost caveats, and doc pointers an agent needs to call this correctly in one shot.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and each parameter already has a detailed inline description, so the schema carries most semantic load; baseline 3 applies. The description restates override/inheritance behavior and the single-entity/single-language constraint, but mostly mirrors what the schema descriptions already say rather than adding new meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource pair ('Generate and auto-save a schema from reviewed samples') and names the concrete outputs (schema_id, schema content, record links). It is clearly distinguishable from siblings like analyze_sample, save_schema, and create_benchmark_scenario. An agent knows exactly what this produces without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives real when-to-use guidance: supply entity_samples or samples_csv (never both), or sample_record_id, and resolve modeling/semantic-ID decisions before calling. It also directs the agent to the docs link for review/canonicalization. It stops short of naming explicit sibling alternatives (e.g. save_schema, add_schema_property) as substitutes, so it is clear context but not full routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_attachmentDelete attachmentA
DestructiveIdempotent
Inspect

Permanently delete an attachment in your organization, including its stored file. No LLM call. Existing records remain, but future calls or schema regeneration cannot reuse the deleted source. Delete only when that source is no longer needed; this is not a required post-enrichment step. Returns the deleted ID and filename.

ParametersJSON Schema
NameRequiredDescriptionDefault
attachment_idYesUUID of the attachment to delete.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses that deletion is permanent and includes the stored file, that existing records remain, and that the deleted source cannot be reused. This adds significant behavioral context beyond the destructiveHint annotation, which only indicates destructiveness. It also notes 'No LLM call', a useful operational detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise (three sentences) and front-loads the core action, then adds essential caveats about permanence and when not to use. Every sentence adds value without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple (one parameter) and has an output schema, so return values need no explanation. The description covers purpose, scope, conditions, and consequences thoroughly, making it complete for an agent to decide and invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already fully documents the single parameter attachment_id with type and description (coverage 100%). The tool description adds no additional meaning or context about the parameter itself, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'delete', the resource 'attachment', and the scope 'in your organization, including its stored file'. It distinguishes itself from sibling delete tools (delete_schema, delete_benchmark_scenario) by specifying the attachment domain, and it clarifies it is not a required post-enrichment step, making its purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage guidance: 'Delete only when that source is no longer needed' and explicitly states 'this is not a required post-enrichment step', preventing misuse. It also notes that future calls or schema regeneration cannot reuse the deleted source, which helps the agent decide when deletion is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_benchmark_scenarioDelete benchmark scenarioA
DestructiveIdempotent
Inspect

Delete a benchmark scenario and its stored results. Requires owner and a plan with benchmarks. No LLM call; inspect the scenario before deleting it.

ParametersJSON Schema
NameRequiredDescriptionDefault
scenario_idYesUUID of the scenario to delete.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds meaningful behavioral context beyond the annotations. It confirms that the tool deletes both the scenario and its stored results (specifying the destructive scope), notes that it requires ownership and a plan with benchmarks, clarifies that no LLM call is involved, and advises inspecting the scenario before deletion. These details enrich the understanding of side effects and prerequisites, and they do not contradict the annotations (readOnlyHint false, destructiveHint true, idempotentHint true).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, front-loaded with the core purpose ('Delete a benchmark scenario and its stored results') followed by essential usage and safety notes. Every sentence earns its place with no redundancy or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simplicity of the tool (one parameter, output schema present, annotations covering destructive and idempotent nature), the description covers all critical aspects: what is deleted, required ownership/plan, absence of LLM call, and a safety caution. Nothing essential is missing for an agent to use it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 100% description coverage for the single parameter scenario_id, already documenting it as the UUID of the scenario to delete. The description does not add parameter-specific information beyond that, so it meets the baseline but provides no extra semantic value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Delete' and the resource 'benchmark scenario and its stored results', which is specific and distinguishes it from sibling tools like get_benchmark_scenario or update_benchmark_scenario. The phrase 'and its stored results' adds clarity about the full scope of the operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides practical usage context by specifying prerequisites ('Requires owner and a plan with benchmarks') and a caution ('inspect the scenario before deleting it'). While it doesn't explicitly name alternatives or state when not to use it, the purpose is self-evident for a delete operation, and the guidance is sufficient for an agent to decide when to invoke it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_database_syncDelete database syncA
Destructive
Inspect

Delete a database registration and its queued deltas, stopping its feed. Requires owner and a sync-enabled plan; no LLM call. External replica tables remain untouched. Entity state and schema database flags remain by default; delete_entity_state and clear_database_model additionally remove data/model settings from schemas left with no registration. Obtain approval for those irreversible teardown options. See enricher://docs/database-sync.

ParametersJSON Schema
NameRequiredDescriptionDefault
database_idYesDatabase sync UUID (from list_database_syncs).
delete_entity_stateNoAlso hard-delete the stored entity state of the schemas left with no database (nothing writes to it anymore). Enrichment records are untouched. Irreversible.
clear_database_modelNoAlso clear the database model those schemas carry (database_key, db_type, index, unique_group, shared, ordered, db_name/db_name_absolute, the classification ledger and the key-language lock), returning them to plain enrichment schemas. Implies delete_entity_state. Irreversible — a later re-link re-classifies from scratch.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes beyond the destructiveHint annotation by detailing exactly what is deleted (registration, queued deltas, feed) and what is preserved (external replicas, entity state, schema flags by default). It explains the side effects of the optional flags and their irreversibility, and even warns to obtain approval. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is moderately long but each sentence contributes unique information: action, requirements, side effects, options, and a doc link. It is front-loaded with the core purpose and logically organized, though slightly dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive operation with three parameters, the description covers prerequisites, exact side effects, default behavior, and irreversible options. It also references documentation for deeper details. Combined with the output schema (not shown but present), the agent has enough to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema coverage is 100%, the description adds critical nuance: it explains that clear_database_model implies delete_entity_state, and clarifies the meaning of 'schemas left with no registration'. This is not present in the schema and is essential for correct parameter usage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Delete a database registration and its queued deltas') with clear scope, distinguishing it from sibling tools like delete_schema or delete_attachment. It also names prerequisites (owner, sync-enabled plan) which further clarifies its role.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear prerequisites (owner, sync-enabled plan, no LLM call) and what remains untouched (external replica tables). It doesn't explicitly name alternatives, but the purpose is distinct enough that an agent can infer when to use it. The note about obtaining approval for irreversible teardown options is useful guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_schemaDelete schemaA
DestructiveIdempotent
Inspect

Soft-delete a saved schema by UUID. Requires editor; no LLM call. Restoration and permanent deletion are available in the web app, not through this tool. Returns the deletion outcome.

ParametersJSON Schema
NameRequiredDescriptionDefault
schema_idYesUUID of the schema to delete.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the tool as destructive and idempotent, but the description adds substantial behavioral context beyond that: the operation is a soft delete, editor permission is required, no LLM call is made, and the tool cannot restore or permanently delete. This gives an agent an accurate model of side effects even though destructiveHint is present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, each earning its place: the operation and target, the requirements and constraints, and the return behavior. The most decision-relevant information is front-loaded, and there is no redundant filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with an output schema, the description is fully sufficient. It covers what the tool does, under what permissions, what side effects it has, and what it will not do. No important operational detail an agent needs before calling it is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% since schema_id is documented as 'UUID of the schema to delete.' The tool description repeats that the operation targets a saved schema by UUID but adds no new parameter-level meaning. Baseline 3 is appropriate because the schema carries the parameter documentation burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb plus resource: 'Soft-delete a saved schema by UUID.' It adds the soft-delete nuance, which distinguishes it from permanent deletion and from sibling tools that delete other resource types. The phrase 'Returns the deletion outcome' completes the picture without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear conditions of use: it requires editor permissions, it is not an LLM call, and it only supports soft deletion. It explicitly states that restoration and permanent deletion are not available through this tool and belong to the web app, which is a useful when-not. It does not name a specific sibling tool as an alternative, but for a schema delete operation that is not necessary.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_semantic_conceptsDelete semantic conceptsA
Destructive
Inspect

Delete concepts selected by ids, concept_types or unused_only. Defaults to impact_only=true: review affected records, schemas and replicas before obtaining approval. Execution requires editor; clearing whole types requires owner. Deletion breaks convergence with IDs already stored in replicas; future resolution may mint new IDs. An unused-only deletion still removes that vocabulary. No LLM call. Returns impact counts or deletion results. See enricher://docs/semantic-ids.

ParametersJSON Schema
NameRequiredDescriptionDefault
idsNoExplicit concept semantic_ids to delete.
impact_onlyNoTrue = report the blast radius only; False = delete.
unused_onlyNoOnly concepts no record references (usage 0).
concept_typesNoScope to these concept types; without ids and unused_only this clears the whole types (owner role).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the tool as destructive and not read-only, but the description adds substantial behavioral context: impact-only review, role-based permission requirements, convergence-breaking consequences, future ID minting, unused-only still removing vocabulary, and absence of LLM calls. This far exceeds what annotations alone convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence carries essential information: selection method, safety default, permissions, long-term side effects, execution behavior, return value, and reference docs. It is front-loaded with the core action and immediately provides decision-relevant constraints.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive tool with four parameters, annotations, and an output schema, the description covers all necessary decision points: what is deleted, how to limit impact, who may execute, what side effects occur, what is returned, and where to find further documentation. No critical operational context is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaningful relational semantics beyond the schema, such as how ids, concept_types, and unused_only interact, the owner-role requirement for clearing whole types, and the impact_only default. This lifts it above the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Delete concepts selected by ids, concept_types or unused_only.' It clearly identifies the object being deleted and the selection criteria, distinguishing it from sibling tools that delete schemas, attachments, or syncs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear operational context: impact_only defaults to true, approval is needed, editor vs owner role requirements are stated, and a docs link is provided. It does not explicitly name sibling alternatives or state when not to use this tool, but the selection criteria and role conditions make the intended usage clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

enrich_entityEnrich entityAInspect

Enrich one entity against exactly one of schema_id or target_schema, returning structured output, record_id, costs and any database outcome. Read get_schema's input_contract first (published version when linked); never invent preserve values. Models may be omitted for auto selection. Generation is billed. Two or more models fuse only when all succeed: check failed_models before reporting success. A failed leg prevents automatic fusion and database admission. classification_warning returns success=false, error_code and classification; bypass only after user confirmation. Database sync defaults on: report database.status and database_warning, including partial writes or overwrites; admission is not proof of replica delivery. Recovery and fusion: enricher://docs/enrichment-and-fusion. This tool does not expose web-search activation.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelsNoModel composite keys. Omit or pass ['auto'] for one automatically selected model; auto alone never fuses. Use list_models for explicit choices and model-count limits.
strategyNoauto (default — server picks from the schema) | single_pass (simple schemas, 1 LLM call) | expert_domains (medium schemas with clear domains) | multi_expertise (large multi-domain schemas, parallel per-expertise calls — best quality, higher cost)auto
languagesNoISO 639-1 codes; defaults to ['en'] server-side. The first language is the primary one used for all non-multilingual string fields; multilingual fields get one value per language.
schema_idNoUUID of a saved schema. Mutually exclusive with target_schema.
entity_dataYesEntity identifiers and supplied values. Read get_schema.input_contract first: preserve paths and keys for supplied array items are required. Identifying field names are guidance; arbitrary names are accepted.
database_syncNoWhether this run feeds the schema's entity layer and linked databases. Leave true for normal enrichments. Set false for the one-model recovery leg of a failed fusion run (see the recovery ladder above): the merge_records call that follows is what should write, so the intermediate single-model write is skipped. The record itself is still saved either way.
target_schemaNoInline JSON Schema document in the supported Entity Enricher dialect; prefer schema_id. Format: enricher://docs/schema-reference.
attachment_idsNoUUIDs of attachments (from upload_attachment) to provide as source material for this enrichment.
timeout_secondsNoWall-clock cap. Past it the call returns `enrichment_timeout` with the job_id and the job is cancelled — a leg still in flight may finish and persist a partial record, reachable via list_records(job_id=...) and recoverable with retry_expertises. Multi-model fusion runs and reasoning models routinely need more than the default; raise it or use start_batch_enrichment (async).
arbitration_modelNoOptional LLM model key used to resolve fusion conflicts when 2+ models are selected. Without this, conflicts are resolved by deterministic voting.
classification_modelNoOptional pre-flight classifier model key. When set, the entity is type-checked before enrichment to catch mismatches.
force_after_classification_warningNoSet to true to bypass a previous classification warning. Use only after explicit user confirmation.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Against sparse annotations (readOnlyHint=false, openWorldHint=true, destructiveHint=false), the description carries rich adjacent behavior: generation is billed, fusion occurs only when all models succeed (check failed_models), a failed leg blocks automatic fusion and database admission, database sync defaults on with partial-write/overwrite disclosure, admission does not prove replica delivery, and classification warnings require user confirmation to bypass. It also flags timeout behavior (partial records via job_id). No contradiction with annotations; readOnlyHint=false aligns with the write/database-sync semantics described.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The front-loaded purpose sentence hits the core contract immediately, and every subsequent sentence earns its place (billing, fusion semantics, database sync, classification warning, recovery doc links). It is a dense wall of text rather than structured paragraphs, but the complexity of a 12-parameter tool with fusion and database-sync semantics justifies the length and density.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the high complexity (12 params, nested objects, output schema, fusion and database-sync semantics), the description is remarkably complete: it covers single-vs-batch routing, billing, fusion failure modes, database partial writes and replica-delivery caveats, classification-warning bypass, timeout partial-record recoverability, and links to docs (enricher://docs/enrichment-and-fusion, schema-reference). The presence of an output schema relieves the need to explain return values, so nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3 and the schema properties already document every parameter. The description nonetheless adds meaning beyond the schema for key params: models ('Two or more models fuse only when all succeed'), database_sync (defaults on, partial writes, admission vs replica delivery), and force_after_classification_warning (bypass only after explicit user confirmation). This elevates it above the baseline, though it doesn't enumerate every parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Enrich one entity') with precise scoping ('against exactly one of schema_id or target_schema') and the output contract ('structured output, record_id, costs and any database outcome'). The single-entity scope visually distinguishes it from start_batch_enrichment, and references to get_schema, list_models, and retry_expertises orient it among siblings. An agent can tell exactly what this tool does and what it returns.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description prescribes a prerequisite workflow ('Read get_schema's input_contract first'), an exclusion rule ('never invent preserve values'), and pointers to alternatives (list_models for explicit model choice, start_batch_enrichment for async in the timeout context, retry_expertises for recovery). It does not, however, state an explicit 'use X instead when...' condition for the primary batch alternative in the headline text, leaving some routing inference to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fetch_database_deltasFetch database deltasAInspect

Read the next ordered window of SQL deltas and canonical payloads for a database sync. claim=false is replayable; claim=true leases the window for 120 seconds. Apply a claimed window transactionally, including schema deltas, before ack_database_deltas. limit is also bounded by the registration's page_limit. snapshot_required pauses delivery until the replica reapplies its snapshot. No LLM call. Leasing, cursors, quarantine and snapshot handling: enricher://docs/database-sync.

ParametersJSON Schema
NameRequiredDescriptionDefault
claimNoLease the window (requires ack).
limitNo
sinceNoCursor: return deltas with id > since.
database_idYesDatabase sync UUID (from list_database_syncs).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are minimal (readOnlyHint=false, destructiveHint=false), so the description carries the disclosure burden. It goes beyond them by detailing lease duration, replayability, transactional application, page_limit bounding, snapshot pause behavior, and the fact that no LLM call is made. This is strong behavioral transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence carries distinct information: scope, replay/lease semantics, transactional ordering, limit constraint, snapshot gating, and a pointer to detailed docs. It is front-loaded with the core read action and wastes no words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a sync-consumption tool with an output schema, the description covers the operational essentials: lease behavior, acknowledgment dependency, ordering, limit interaction, and snapshot handling. It also links to documentation for leasing, cursors, quarantine, and snapshot handling, so an agent has enough context to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 75% schema coverage, the schema already explains database_id, claim, and since. The description adds meaning by explaining that claim=false is replayable while claim=true leases the window, and that limit is also bounded by the registration's page_limit. This usefully supplements the schema rather than repeating it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action and resource: 'Read the next ordered window of SQL deltas and canonical payloads for a database sync.' This clearly distinguishes it from acknowledge-only or sync-management siblings such as ack_database_deltas and create_database_sync. The scope of the operation is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides actionable guidance on claim=false vs claim=true, the 120-second lease, and the requirement to apply a claimed window before calling ack_database_deltas. It does not explicitly enumerate when-not-to-use alternatives, but the coordination with ack_database_deltas and replayability semantics give clear context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_sampleGenerate sampleAInspect

Generate editable sample JSON from a free-text request for schema authoring. Without attachments, use model knowledge and optional web search; with attachments, extract from the sources only (search does not relax that rule). Each sample is one instance in one language; set sample_count separately from request. Attachments force one sample. Requires editor; generation is billed. Returns a job_id and may already be paused or complete: relay pause questions through answer_job_question, otherwise poll get_job_status. Review returned samples and warnings before create_schema_from_sample; do not silently change facts or structure. For relationship modeling, multiple documents or hybrid extraction plus research, read enricher://docs/schema-from-sample and enricher://docs/documents.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoAuto (default) chooses the organization task model with attachment/search capabilities. Explicit provider::model bypasses this capability matching; provider combinations or quota may still fail.auto
requestNoEntity type, desired fields, scope and size/depth budget. Required without attachments; optional source-mode instructions otherwise. Put the number of instances in sample_count, not in this text.
languageNoOutput language code for the generated field names AND values (e.g. 'en', 'fr'); an explicit code applies even when attachments are in another language. Omitted (default) → the generator follows the language the request is written in (its text and any typical object), else the attachment's, else English.
auto_answerNoOmit or false to pause for clarification; true authorizes standard interpretations and planner defaults without asking, in either mode.
sample_countNoNumber of same-type instances (1..20), default 1. Consider 3 varied instances when designing a schema. Attachments force 1. Inspect samples_note for under-delivery or the cap.
wait_secondsNoHow long to wait for the first pause or completion before returning (0 = return the job_id immediately).
attachment_idsNoUUIDs from upload_attachment. Providing any attachment switches the call into source mode: transcribe the document or describe visible photo attributes only, with an interactive planner.
typical_objectsNoUp to sample_count concrete instances to anchor knowledge mode (e.g. ['Sanofi', 'Pfizer']), one per generated sample in order — slots beyond len(typical_objects) are named by the model's own instance roster. In source mode the attachment remains authoritative and this is ignored.
enable_web_searchNoUse builtin search in knowledge mode (default false). Source mode remains source-only. With an explicit unsupported model this option is ignored. Hybrid tasks: enricher://docs/documents.
naming_conventionNoauto | snake_case | camelCaseauto

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description discloses that generation requires an editor, is billed, returns a job_id, and may already be paused or complete. It also warns that attachments force one sample, source mode is strict, and the agent must not silently change facts or structure. This is substantial behavioral context the annotations do not provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: purpose, mode rules, billing/editor requirement, job lifecycle, review obligation, and advanced doc pointers. It is front-loaded with the core purpose and then layers necessary workflow detail without repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 10 parameters and a complex asynchronous workflow, the description covers the full call path: modes, attachments, job_id return, pause/completion states, next-step tools, and warnings. An output schema exists, so return values need no further explanation, and advanced cases are routed to documentation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds cross-parameter meaning: sample_count must be set separately from request, each sample is one instance in one language, attachments force one sample, and typical_objects anchors knowledge mode. These clarifications help an agent use the parameters together correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Generate editable sample JSON from a free-text request for schema authoring.' It clearly distinguishes this from downstream siblings like create_schema_from_sample and get_job_status, so an agent knows exactly what this tool produces.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit mode-based guidance: without attachments use model knowledge and optional web search; with attachments extract from sources only and search does not relax that. It also names the exact alternatives for follow-up actions: answer_job_question for pauses, get_job_status for polling, and create_schema_from_sample after review, plus doc links for advanced relationship modeling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_benchmark_scenarioGet benchmark scenarioA
Read-only
Inspect

Read one benchmark scenario with per-model quality, cost and speed results. include_reference=true adds the reference, the fixed entity_data and the schema-generation entity_samples. Config changes can make existing results stale. No LLM call. For a ranked subset use get_benchmark_scenario_results. Read after run_benchmark completes.

ParametersJSON Schema
NameRequiredDescriptionDefault
scenario_idYesUUID of the scenario.
include_referenceNoInclude reference_output + entity_data (can be large).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint=true and destructiveHint=falsehistorically. The description adds meaningful behavioral context beyond this: 'No LLM call' clarifies cost/latency characteristics, while 'Config changes can make existing results stale' warns about data freshness. It also explains the side effects of include_reference=true, which is a behavioral trait not evident from annotations alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is four short sentences, each contributing essential information: what the tool returns, the parameter effect, a staleness warning, and a routing hint. It is front-loaded with the core purpose and has zero filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is a simple read with an output schema present, so return values need no elaboration. The description covers the main use case (after run_benchmark), the caveat (config changes), the parameter semantics, and the alternative tool. Nothing an agent needs to safely and effectively invoke it is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so a baseline of 3 is expected. However, the description adds value by explaining that include_reference=true adds not only reference_output and entity_data (already in the schema) but also 'the schema-generation entity_samples', which is not mentioned in the schema. This clarifies the parameter's full impact beyond the structured definition.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Read one benchmark scenario') and the resource with its content ('per-model quality, cost and speed results'). It also explicitly distinguishes from the sibling get_benchmark_scenario_results by noting that the sibling returns a 'ranked subset', which is a clear differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit usage guidance: 'Read after run_benchmark completes' and 'For a ranked subset use get_benchmark_scenario_results'. It also warns about staleness with 'Config changes can make existing results stale', helping agents decide when to refresh or trust the data.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_benchmark_scenario_resultsGet ranked benchmark scenario resultsA
Read-only
Inspect

Filter, rank and limit a scenario's per-model benchmark results. No LLM call. overall blends quality, speed and cost using organization task weights and is null if a component is missing. Status tags are independent: success does not exclude stale or stale_score results. Missing sort metrics come last in either direction. See enricher://docs/model-benchmark for interpretation.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoKeep only the top N after sorting (None = all).
statusNoKeep results carrying ANY of these tags: success | failed | stale (config_hash changed since this run, re-run it) | stale_score (reference/scoring config changed since scored, rescore it) | unscored (ran fine, never scored). Empty/None = every status.
sort_byNoMetric to sort by.overall
providersNoKeep only these provider names (empty/None = every provider).
model_keysNoKeep only these model composite keys (empty/None = every model).
sort_orderNoSort direction.desc
scenario_idYesUUID of the scenario.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint/destructiveHint annotations, the description discloses key behavioral nuances: 'overall' is a weighted blend that becomes null if any component is missing, status tags are independent (success does not exclude stale/stale_score), and missing sort metrics always sort last. These are critical for correctly interpreting results and are not visible in the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no wasted words. The opening verb phrase immediately states the action, followed by essential behavioral caveats and a pointer to docs. Each sentence contributes distinct information, and the most important constraints (status independence, null handling) are front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The presence of an output schema means return values need not be described. The description covers filtering, ranking, limiting, and the tricky semantics of nulls and status tags. It also provides a documentation link for deeper interpretation. Nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema description coverage, the baseline is 3. The description adds value by explaining the meaning of 'overall' (weighted blend) and the independence of status tags, which are not fully captured in the parameter descriptions. This goes beyond the schema's per-parameter docs, though it doesn't exhaustively cover every parameter nuance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (filter, rank, limit), a specific resource (a scenario's per-model benchmark results), and the scope (per-model). It distinguishes itself from siblings by focusing on retrieving and ranking results rather than creating or running benchmarks. The 'No LLM call' clarification reinforces its read-only purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies this is the tool for retrieving and ranking existing benchmark results, with explicit filtering and sorting semantics. It does not explicitly name alternatives or when-not-to-use conditions, but the context (read-only, results-focused) makes the use case unambiguous. A brief reference to the docs for interpretation adds some guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_database_setup_instructionsGet database setup instructionsA
Read-only
Inspect

Return non-secret install, browser-confirmed pairing and run instructions for an ee-database sync client. Requires owner and a sync-enabled plan; no LLM call and no credential is issued or exposed to the MCP client. Run pair_command on the intended replica host: the CLI opens verification_url, the user chooses the database and confirms, and the credential travels directly to the polling CLI. Its DSN stays on that host. Use manual pairing only when managed host provisioning is not already handling the registration. Setup: enricher://docs/database-sync.

ParametersJSON Schema
NameRequiredDescriptionDefault
database_idYesDatabase sync UUID (from list_database_syncs).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnly and non-destructive annotations, the description discloses meaningful behavior: no LLM call, no credential issued or exposed to the MCP client, the DSN stays on the host, and the pairing flow is browser-confirmed. This gives an agent an accurate mental model of side effects and security boundaries.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but efficient: purpose is front-loaded, followed by prerequisites, security-relevant behavior, the pairing flow, and a clear when-to-use caveat. Every sentence adds information an agent needs for correct invocation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has a single documented parameter, an output schema, and read-only annotations, the description covers the remaining operational context: required permissions, setup steps, credential handling, and the manual-pairing condition. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the single parameter 'database_id' is documented with its source ('from list_database_syncs'). The tool description does not add parameter-specific detail, but the schema already carries the necessary meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Return non-secret install, browser-confirmed pairing and run instructions for an ee-database sync client.' This goes well beyond the title and makes the tool's function unambiguous relative to related database-sync siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states preconditions ('Requires owner and a sync-enabled plan') and gives an explicit routing rule: 'Use manual pairing only when managed host provisioning is not already handling the registration.' It does not name sibling tools as alternatives, but the context is clear enough for an agent to decide when this tool is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_enum_candidatesGet enum candidatesA
Read-only
Inspect

List observed values outside each open enum's current vocabulary, with counts from recent enrichment records. No LLM call. Use the report to propose admitted members or rejected_values, or close a vocabulary only when it is exhaustive. This read does not edit the enum. An open enum allows other values; a closed one constrains output to its members. Named enums can be read with get_schema_part and changed through update_schema. See enricher://docs/schema-reference.

ParametersJSON Schema
NameRequiredDescriptionDefault
schema_idYesUUID of the saved schema.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds valuable behavioral context: it is a read that does not edit the enum, it makes no LLM call, and it reports counts from recent enrichment records. This goes beyond the annotations by explaining the operation's non-mutating nature and data source.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the core action and scope appear in the first sentence, followed by usage guidance and exclusions. Every sentence adds value, and the reference to enricher://docs/schema-reference is a useful pointer without bloating the text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read tool with a full output schema and annotations covering safety, the description is complete. It explains the purpose, the open/closed enum distinction, when to use it, and what it does not do. An agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents the single parameter (schema_id as UUID of the saved schema). The description does not add parameter-specific details beyond what the schema provides, but with full coverage the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('List') and resource ('observed values outside each open enum's current vocabulary'), and distinguishes it from related tools by noting it is a read that does not edit the enum. It also clarifies the distinction between open and closed enums, which helps an agent understand the tool's exact scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when to use this tool: to propose admitted members or rejected_values, or to close a vocabulary only when exhaustive. It also names alternatives (get_schema_part for reading named enums, update_schema for changing them) and states that this tool does not edit the enum, providing clear when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_job_statusGet job statusA
Read-only
Inspect

Read a job's status, progress and compact terminal summary with persisted record IDs. Status is pending, running, paused, completed, failed or cancelled. Relay pause questions with answer_job_question. include_result=true returns full terminal details; batch summaries count entities and database outcomes. A failed job is not usable output; inspect individual records for partial recovery. Unknown IDs may be invalid, expired or lost after restart: look for persisted outputs with list_records(job_id=...), without assuming success. events_after= adds the job's event log past that cursor (per-model completions, scoring progress, pauses) — the poll equivalent of the SSE stream; pass the returned last_seq next time. No LLM call. Polling, failure and recovery guidance: enricher://docs/enrichment-and-fusion.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYesJob ID returned by a start tool.
events_afterNoEvent-log cursor: 0 returns the job's events from the start, a previous call's last_seq returns only the newer ones (at most 100 per call; events_has_more says when to call again). Omit to skip the log.
include_resultNoInclude the full terminal result payload (can be large). Default returns a compact scalar summary per model.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Despite readOnlyHint and destructiveHint already indicating a safe read, the description adds substantial behavior: possible statuses, include_result behavior, failed-job semantics, unknown ID causes, event log semantics, and 'No LLM call'. It even notes the polling relationship to the SSE stream, which is valuable beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but each sentence carries operational value: scope, statuses, alternative tools, result modes, failure handling, unknown IDs, event cursor mechanics, and a documentation link. It is front-loaded with the core purpose and avoids filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a polling/status tool with an output schema, the description is complete. It covers all parameters, failure modes, recovery paths, alternative tools, and the polling loop. The presence of an output schema means return-value details are not required, and nothing essential is left out.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema by explaining what include_result returns ('full terminal details', 'batch summaries count entities and database outcomes') and how events_after works as a cursor ('pass the returned last_seq next time'). This is more than a restatement.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description starts with a clear action verb and resource: 'Read a job's status, progress and compact terminal summary'. It also differentiates itself from related tools by naming answer_job_question and list_records as separate paths, making the tool's scope unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly tells the agent when to use alternative tools: 'Relay pause questions with answer_job_question' and for unknown IDs, 'look for persisted outputs with list_records(job_id=...), without assuming success'. It also explains the polling pattern with events_after and points to docs for failure and recovery guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_recordGet recordA
Read-only
Inspect

Read one persisted record's structured_output, entity_input_data, validation errors, expertise verdicts and metrics. failed_expertises and partial identify incomplete work even when output exists; use retry_expertises only on a record with failed domains. Fusion metadata identifies its source models and arbitration method. database_sync, when present, reports admission and per-replica delivery state. Returns record_url and the related schema link. No LLM call. Interpretation and recovery: enricher://docs/enrichment-and-fusion.

ParametersJSON Schema
NameRequiredDescriptionDefault
record_idYesUUID of the record.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the bar is lower, but the description still adds substantial behavioral context: 'No LLM call' signals cost/latency profile, 'failed_expertises and partial identify incomplete work even when output exists' reveals that a returned record may still represent failed work, and it explains the semantics of fusion metadata and the conditional database_sync field. No contradiction with annotations — 'Read' and 'No LLM call' align with the read-only hint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-structured: core purpose first, then result-interpretation semantics, sibling routing, metadata meaning, and finally cost/docs pointers. Every sentence carries distinct information with no filler, though it is on the longer side and slightly front-loads detail before the simplest 'no LLM call' takeaway.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read tool with an output schema and safety annotations, the description covers everything an agent needs: what fields the record contains, how to detect incomplete/failed work, which sibling to route to for recovery, what fusion metadata means, the meaning of the conditional database_sync field, and a docs link for interpretation. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%: the single parameter record_id is already documented as 'UUID of the record.' The description reinforces this via 'one persisted record' but adds no parameter-specific format, validation, or edge-case detail beyond the schema. Baseline 3 is appropriate when the schema carries the full parameter documentation burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-resource pair ('Read one persisted record's...') and enumerates exactly which fields are returned (structured_output, entity_input_data, validation errors, expertise verdicts, metrics). The singular scope 'one persisted record' cleanly distinguishes it from list_records, and the focus on enrichment/verdict data separates it from get_schema and get_stats.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit routing guidance: 'use retry_expertises only on a record with failed domains', telling the agent when a sibling is appropriate and what precondition must be checked first. It also clarifies interpretation ('failed_expertises and partial identify incomplete work even when output exists') and points to recovery docs. It does not explicitly contrast with list_records or get_stats, but the read-one-record scope implies the boundary.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_schemaGet schemaA
Read-only
Inspect

Read a saved schema with its properties, annotations and input_contract. Before enrichment, use version='published' for a database-linked schema; default 'working' is for editing. identifying_keys guide entity naming; preserve values belong to the caller and are required; each supplied array item must carry its array_item_keys. Never invent caller-owned values. The response includes version, publish_state and schema_url. A requested published contract that does not exist returns not_found. For a small edit prefer get_schema_part. No LLM call. Contract details: enricher://docs/schema-reference.

ParametersJSON Schema
NameRequiredDescriptionDefault
versionNo'working' (default) — the editable copy updates/edits apply to; 'published' — the contract enrichment and database sync use (only meaningful for schemas linked to a database sync).working
schema_idYesUUID of the saved schema.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the read-only nature is covered. The description adds value by stating 'No LLM call', explaining the not_found behavior for a missing published contract, and describing response fields (version, publish_state, schema_url). It also discloses the semantics of identifying_keys and preserve values, which goes beyond the annotation baseline.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is longer than average but every sentence carries substantive content: purpose, version selection, entity-naming guidance, error behavior, and a pointer to documentation. It is front-loaded with the core purpose and ends with a contract reference. Some redundancy with the schema exists, but it is not padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the return format is already documented. The description covers version selection, error cases (not_found), the distinction from get_schema_part, and the special semantics of identifying_keys and preserve values. No critical information for a correct call is missing; the agent knows exactly when to call it and what to expect.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description enhances parameter understanding by clarifying the version semantics (working vs published and their respective use cases) and notes that schema_id is a UUID. It adds practical guidance on how to choose the version, which the schema only partially covers.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb-resource pair: 'Read a saved schema' and enumerates exactly what is read (properties, annotations, input_contract). It also implicitly distinguishes itself from get_schema_part by saying 'For a small edit prefer get_schema_part' and from list_schemas by focusing on a single saved schema. This makes the tool's purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit usage guidance is given: use version='published' before enrichment for database-linked schemas, default 'working' for editing, and prefer get_schema_part for small edits. These conditions directly tell an agent when to choose this tool over alternatives, satisfying the 'when/when-not' requirement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_schema_partGet schema partA
Read-only
Inspect

Read only the schema fragment needed for an edit. Omit path for the root/type index; use '$defs.X' or '$enums.X' for a definition, an object path for its subtree, or a leaf path for its property card and relations. Dot-separated paths use '[]' for array items. The response identifies shared definition usage, identity participation, database flags and whether entity state or linked databases exist. Editing a $defs property affects every usage site. Reads the working copy; no LLM call. Path examples: enricher://docs/schema-reference.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoOmit for the index; else a property/object path or '$defs.X'/'$enums.X'.
schema_idYesUUID of the saved schema.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is known. The description adds valuable behavioral context beyond annotations: it 'Reads the working copy; no LLM call', which informs cost and latency, and it details what the response identifies (shared definition usage, identity participation, etc.). It also warns that 'Editing a $defs property affects every usage site', a useful consequence for the editing workflow. This goes beyond the annotations and enhances transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-structured: it opens with the core purpose, then explains path syntax, response content, an editing caveat, and finally a note on working copy/no LLM. Each sentence adds value; there is minimal redundancy. While it is longer than minimal, it is organized front-loaded and earns its length through necessary usage details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the presence of an output schema (which defines the return format) and annotations covering read-only/destructive safety, the description completes the picture by explaining how to target different schema parts, what the response highlights, and the implication of editing shared definitions. It also clarifies the working copy and no-LLM behavior. Nothing an agent needs to call this tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so both parameters are documented. The description adds meaning beyond the schema by detailing path variants ('$defs.X', '$enums.X', object/leaf paths) and giving the array syntax '[]'. It also provides a concrete example (enricher://docs/schema-reference), which helps an agent construct valid paths. This exceeds the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Read only the schema fragment needed for an edit.' This clearly states the tool's function and its scoped purpose, distinguishing it from siblings like get_schema (which likely returns the full schema). The detailed path semantics further clarify exactly what can be fetched, leaving no ambiguity about the tool's role.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides strong contextual guidance: it's for reading a fragment 'needed for an edit', and it gives explicit path syntax (e.g., '$defs.X', '[]' for arrays). However, it does not explicitly name alternatives or state when NOT to use this tool (e.g., 'for the full schema, use get_schema instead'). The context implies the right usage but lacks an explicit exclusion or comparison.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_semantic_conceptGet semantic conceptA
Read-only
Inspect

Read one concept's aliases, identity source keys, linked records and nearest neighbors within its own type/model slice. Record links are capped; records_truncated signals omitted links. neighbors_limit=0 skips neighbors. Use returned alias IDs with update_concept_alias and similarities to assess a proposed merge. No LLM call. Never compare similarities across embedding spaces or concept types. See enricher://docs/semantic-ids.

ParametersJSON Schema
NameRequiredDescriptionDefault
concept_idYesThe concept's semantic_id (UUID).
neighbors_limitNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the tool as read-only and non-destructive, and the description adds meaningful behavioral context: 'Record links are capped; records_truncated signals omitted links,' 'neighbors_limit=0 skips neighbors,' and 'No LLM call.' These details go well beyond what annotations alone provide and set accurate expectations about output limits and deterministic behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence adds value: capabilities, output truncation, parameter edge case, downstream use, determinism, cross-space warning, and a docs pointer. Front-loading the main function makes it immediately scannable for an agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the small parameter count, strong annotations, and presence of an output schema, this description is complete. It covers what is returned, how results are limited, how neighbors can be skipped, how to use the results, and important constraints. An agent can invoke this tool correctly without needing additional clarification.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%, with concept_id documented but neighbors_limit lacking a description. The description compensates by explicitly explaining 'neighbors_limit=0 skips neighbors' and clarifies nearest-neighbor behavior with 'within its own type/model slice.' It does not add much for concept_id, but the schema already covers that parameter adequately.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Read one concept's aliases, identity source keys, linked records and nearest neighbors.' It also adds scope by saying 'within its own type/model slice,' which differentiates this tool from broader semantic list/merge operations. This is clearly more informative than the tool's title alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a concrete workflow: 'Use returned alias IDs with update_concept_alias and similarities to assess a proposed merge.' It also warns against comparing similarities across embedding spaces or concept types. It does not explicitly name sibling alternatives or when-not-to-use conditions, but the context is clear enough to guide correct selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_statsGet statisticsA
Read-only
Inspect

Read organization-wide record totals, success rate, tokens and cost summary. No LLM call. This tool has no job filter; use list_records(job_id=...) for a particular run and benchmark tools for comparative quality/cost/speed scores.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, covering safety. The description adds a useful behavioral trait: 'No LLM call,' indicating a cheap/fast operation. This goes beyond annotations but is not exhaustive; it doesn't mention edge cases, but given output schema and read-only annotations, it's sufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences. The first front-loads the core purpose, the second provides routing guidance. Zero filler or redundant phrasing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool with an output schema and annotations covering safety, the description covers purpose, usage boundaries, and behavioral cost. Nothing an agent needs to correctly invoke it is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, so the baseline is 4. The description adds context about scope ('organization-wide') but no parameter-specific details are needed. Schema coverage is trivially 100% with no properties.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb 'Read' and a clear resource: organization-wide record totals, success rate, tokens, and cost summary. Explicitly contrasts with sibling tools like list_records and benchmark tools, making its distinct scope obvious.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Directly tells the agent when to use this tool vs alternatives: use for organization-wide stats, use list_records(job_id=...) for a specific run, and benchmark tools for comparative scores. This is explicit and unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

import_semantic_conceptsImport semantic conceptsAInspect

Resolve 1..1000 texts against one concept type. mint=false returns exact/matched/would_mint outcomes without minting; mint=true creates misses and requires owner (reporting requires editor). Resolution can call embeddings and the identity judge even in report mode. Review would_mint rows before authorizing creation. A new type may be initialized with embedding_model; existing types cannot switch spaces through import. Returns per-text outcomes. See enricher://docs/semantic-ids.

ParametersJSON Schema
NameRequiredDescriptionDefault
mintNoFalse = report only; True = create the unmatched rows (owner role).
textsYesIdentity texts to resolve (1..1000).
judge_floorNoSimilarity at or above which a candidate is put to the identity judge. Omit to use the organization default (Settings → Organization).
concept_typeYesConcept type (slice) to resolve against.
embedding_modelNoComposite key (provider::model) seeding a NEW concept type's slice — an import may open the type it resolves into. Refused when that type's vocabulary already lives in another model.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond annotations, the description discloses key side effects and runtime behavior: minting creates misses, resolution can invoke embeddings and the identity judge even in report mode, and embedding_model can initialize a new type. This adds material context about permissions, side effects, and model coupling that annotations alone would not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: scope, modes, permissions, side-effect warnings, embedding restrictions, and return summary are all covered in six sentences. It is front-loaded with the core action and mode distinction.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with an output schema, the description need not explain return shapes; it already says 'returns per-text outcomes.' Combined with permission details, mode distinctions, embedding caveats, and a docs link, an agent has what it needs to select and invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaningful semantics: it explains the mint flag's outcome modes, describes embedding_model as a composite key used only for new types, and clarifies that a type's vocabulary cannot switch embedding spaces. This goes beyond the schema's field-level descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Resolve 1..1000 texts against one concept type.' It clearly distinguishes this bulk-resolution/minting tool from sibling tools like add_semantic_concept by emphasizing range, concept-type scoping, and optional minting behavior.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit mode guidance: mint=false is report-only, mint=true creates misses and requires owner, reporting requires editor. It also states a clear exclusion ('existing types cannot switch spaces through import') and advises reviewing would_mint rows before authorizing creation. It does not name alternative tools directly, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_benchmark_scenariosList benchmark scenariosA
Read-only
Inspect

List compact benchmark scenario summaries and total. Scenarios test enrichment, sample generation or schema generation. No LLM call. Read one with get_benchmark_scenario or create one with create_benchmark_scenario. Lifecycle and reference requirements: enricher://docs/model-benchmark.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds value by stating 'No LLM call' – a behavioral detail not in annotations – and provides a lifecycle reference ('enricher://docs/model-benchmark') that hints at operational context. This is beyond what structured fields provide, though not extensive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, with the core action front-loaded ('List compact benchmark scenario summaries and total'). It includes sibling routing and a documentation pointer without excess. Every sentence earns its place; nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with zero parameters and an output schema, the description covers purpose, usage, alternatives, and behavior. It even points to lifecycle documentation. An agent can confidently invoke this tool based on the description alone, with no missing critical details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the input schema is empty. The baseline for 0 params is 4. The description doesn't need to explain parameters; it instead mentions the output ('summaries and total'), which is appropriate. No redundancy with the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'List compact benchmark scenario summaries and total.' It clearly distinguishes from siblings by mentioning alternatives (get_benchmark_scenario and create_benchmark_scenario) and defines the scope (scenarios test enrichment, sample generation or schema generation). This is precise and immediately tells the agent what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides clear routing: 'Read one with get_benchmark_scenario or create one with create_benchmark_scenario,' which tells the agent when to use alternatives. It also notes 'No LLM call,' implying it's a lightweight operation. However, it doesn't explicitly state conditions like 'when you need detailed info, use get...' but the guidance is clear enough. The lifecycle reference adds context for when to consult documentation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_database_syncsList database syncsA
Read-only
Inspect

List a saved schema's database registrations, linked schemas, options and sync hosts. Returns pending and quarantined delta counts, projection_upgrade_pending and database links. No LLM call. Pending zero alone does not prove healthy delivery: check quarantine and migration blockers. Use list_entity_states for server-side rows and get_record for per-record delivery state. PostgreSQL, MySQL and SQLite delivery is performed by the user's sync client. See enricher://docs/database-sync.

ParametersJSON Schema
NameRequiredDescriptionDefault
schema_idYesSaved schema UUID.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, and the description adds meaningful behavioral context beyond that: 'No LLM call,' the caveat that pending zero is not proof of healthy delivery, the note that delivery is performed by the user's sync client for PostgreSQL/MySQL/SQLite, and a pointer to documentation. This gives the agent a realistic model of what the operation does and its limitations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact yet information-dense: it lists the resource scope, return contents, a critical caveat, routing to alternatives, delivery model, and a docs reference. Every sentence earns its place, and the most important scoping information is front-loaded before caveats and alternatives.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description does not need to explain return values in depth. It covers operational context, caveats, sibling routing, and delivery behavior. For a read-only list operation with one required parameter, nothing essential is missing for an agent to select and invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single parameter schema_id is already described as 'Saved schema UUID.' The description reinforces this by saying 'a saved schema's...' but does not add new semantic detail such as formats, defaults, or relationship to other identifiers. Baseline 3 is appropriate because the schema carries the parameter documentation burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'List a saved schema's database registrations, linked schemas, options and sync hosts.' It enumerates the exact returned data and distinguishes itself from sibling tools by naming list_entity_states and get_record as alternatives. An agent can clearly identify what this tool does and what it does not do.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit routing guidance: 'Use list_entity_states for server-side rows and get_record for per-record delivery state.' It also warns when not to rely on a single signal ('Pending zero alone does not prove healthy delivery') and directs to check quarantine and migration blockers. This is strong when-to-use and when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_entity_statesList entity statesA
Read-only
Inspect

Browse a schema's current merged entity rows, not per-run records. Requires editor; no LLM call. Returns identities, revision, last_record_id and optionally payload, with limit/offset pagination. Rejected entities have no row; purge_entity_state may also remove delivered rows. Server-side state does not prove the external replica has applied its deltas. Use get_record and list_database_syncs for delivery and rejection diagnostics. See enricher://docs/database-sync.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
offsetNo
schema_idYesSaved schema UUID.
entity_typeNoFilter to one entity type (the response lists the types present).
include_payloadNoInclude each entity's current merged payload. Set false for a compact identity-only listing (keys, revision, last_record_id).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint and destructiveHint annotations, the description adds that the tool requires editor permissions and incurs no LLM call cost. It also discloses that server-side state does not prove the external replica has applied deltas, and that purge_entity_state may have removed delivered rows. These are behavioral caveats not captured in annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each carrying distinct information: purpose, requirements/returns, and caveats/routing. The most important scoping ('not per-run records') is front-loaded. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers purpose, prerequisites, return contents, pagination, edge cases (rejected/purged rows), and explicitly routes to alternative tools for diagnostics. With an output schema present, it need not detail the response structure further. It is complete for an agent to decide when and how to invoke it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already documents schema_id, entity_type, and include_payload with descriptions, covering 60% of parameters. The description adds only that pagination uses limit/offset and that include_payload can be set false for a compact listing, which largely mirrors the schema. For the undocumented limit and offset, their purpose is evident from names and defaults, so the description adds limited additional value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool lists a schema's current merged entity rows, explicitly distinguishing it from per-run records. It names the returned fields (identities, revision, last_record_id, optional payload) and pagination, making the tool's purpose specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says to use get_record and list_database_syncs for delivery and rejection diagnostics, routing the agent away from this tool for those purposes. It also notes that rejected entities have no row and that purge_entity_state may remove rows, which informs when to expect data. The requirement for editor and no LLM call are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_modelsList modelsA
Read-only
Inspect

List available model keys, nominal capabilities, languages, strategies, auto-selected defaults and organization profile_limits. Use when choosing explicit models or checking plan limits; ordinary calls may omit models or use auto without fetching this large catalogue. Auto also accounts for attachment capabilities. is_available means a usable provider key exists, not that provider quota or every combination of tools and media will work. Missing capability flags mean unsupported on this discovery surface. default_models_web_search is a search-only preview, not attachment-specific. No LLM call. Model selection and costs: enricher://docs/enrichment-and-fusion.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint and destructiveHint annotations, the description adds important behavioral semantics: is_available does not guarantee quota or every tool/media combination, missing capability flags mean unsupported on this discovery surface, default_models_web_search is search-only, Auto accounts for attachment capabilities, and no LLM call is made. These caveats materially affect how an agent interprets results and costs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense and front-loaded with the core enumeration, then use cases, then disambiguation caveats. It is longer than many tool descriptions, but each sentence adds distinct value; only the trailing doc-link and model-selection aside are auxiliary, keeping it appropriately sized rather than padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read-only discovery tool, the description covers purpose, when to call, interpretation pitfalls, auto behavior, behavior (no LLM call), and a documentation pointer. With an output schema present, an agent has everything it needs to decide whether and how to invoke the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool accepts no parameters, so the schema has nothing to document; the description's field-level detail relates to output semantics rather than inputs. Per the 0-parameter baseline, no additional parameter meaning is required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description starts with a specific verb and resource — 'List available model keys, nominal capabilities, languages, strategies, auto-selected defaults and organization profile_limits' — making the tool's scope unmistakable. The detailed enumeration also distinguishes it from sibling list_* tools such as list_schemas and list_records without needing to name them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when to use the tool ('when choosing explicit models or checking plan limits') and when not to ('ordinary calls may omit models or use auto without fetching this large catalogue'). This gives an agent a clear decision rule for invoking it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_recordsList recordsA
Read-only
Inspect

List compact, paginated records in your organization, most recent first. Filter by job_id to retrieve persisted outputs of an asynchronous workflow; type, success, model and search further narrow the result. Both per-model enrichment and arbitration records may exist for one entity. Use get_record for full output and failures; use list_entity_states for current merged entity state instead of run history. No LLM call. Fetch every page when a job produces more than one page of records.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageNo
modelNoFilter by model composite key (e.g. 'anthropic::claude-sonnet-4-6').
job_idNoFilter to a single job (every model + expertise in one batch shares one job_id).
searchNoSubstring match against the entity's `name` field in the output.
successNoTrue = only successful records; False = only failed.
page_sizeNo
record_typeNoFilter by type: enrichment | classification | arbitration | sample_generation | schema_generation | schema_edit | schema_annotation | ambiguity_analysis | db_classification | benchmark_scoring | playground

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false. The description adds extra behavioral context: 'No LLM call' (performance), 'most recent first' (ordering), the note that both per-model enrichment and arbitration records may exist for one entity, and the pagination advice. These go beyond annotation coverage and add value without contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, each earning its place. The core purpose is front-loaded, then filters, then alternatives, then a practical pagination note. No fluff or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Complete for a read-only list tool: output schema exists, annotations cover safety, and the description covers usage alternatives, ordering, pagination, and the multi-record nuance. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 71%, so baseline is 3. The description lists the filterable parameters (job_id, type, success, model, search) but only summarizes them without adding new semantics; it does not mention page or page_size, though the schema already documents defaults and constraints. No significant value added beyond schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (list) and resource (records) with clear modifiers: 'compact, paginated records in your organization, most recent first.' It also differentiates from siblings by explicitly naming get_record and list_entity_states, making the tool's scope unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use guidance: 'Filter by job_id to retrieve persisted outputs of an asynchronous workflow' and directly points to alternatives: 'Use get_record for full output and failures; use list_entity_states for current merged entity state instead of run history.' Also advises on pagination: 'Fetch every page when a job produces more than one page of records.' No gaps.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_schemasList schemasA
Read-only
Inspect

List saved schemas in your organization, pinned first. Returns compact summaries with IDs, names and links; no LLM call. Use get_schema to inspect a chosen schema and its input contract, or get_schema_part for a targeted edit. For a database-linked schema, enrich against its published version; the working copy may contain unpublished changes.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, and the description adds meaningful behavioral context: pinned ordering, compact summary format with IDs/names/links, and the notable 'no LLM call' trait. It also warns about unpublished changes in working copies, which is valuable beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: the first states the core action and output, the second routes to sibling tools, and the third provides a caveat about published vs. working copies. The most important information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that the tool has no parameters, has an output schema, and is covered by safety annotations, the description is complete. It explains what is returned, how results are ordered, and the key caveat about database-linked schemas, leaving nothing essential for the agent to guess.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there is no parameter semantics to document. Per the baseline for 0-parameter tools, this is adequate; the description appropriately focuses on output and usage rather than parameter details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'List saved schemas in your organization, pinned first.' It also clarifies the output is compact summaries with IDs, names, and links, and distinguishes itself from get_schema and get_schema_part, so an agent can tell it apart from related tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly names alternatives and the conditions for choosing them: 'Use get_schema to inspect a chosen schema and its input contract, or get_schema_part for a targeted edit.' It also adds practical guidance about enriching database-linked schemas against the published version rather than the working copy.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_semantic_conceptsList semantic conceptsA
Read-only
Inspect

Browse organization concepts with aliases, usage counts and type/model facets. view='review' returns pairs escalated by the identity judge, not merely similar pairs. Returns a filtered page; use get_semantic_concept for details and neighbors. Similarities are comparable only within one concept_type/embedding_model slice. No LLM call. Vocabulary review and curation: enricher://docs/semantic-ids.

ParametersJSON Schema
NameRequiredDescriptionDefault
viewNo'concepts' = filtered concept page; 'review' = pairs the identity judge left for a person: escalated (unsure), found duplicated while answering another question, or separated at a similarity high enough to double-check.concepts
limitNo
offsetNo
searchNoSubstring match against concept texts.
sort_byNoref_count | text | concept_type | created_atref_count
sort_orderNodesc
concept_typesNoRestrict to these concept types (empty/None = every type).
min_ref_countNoOnly concepts used by at least this many records.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered. The description adds valuable behavioral context: 'No LLM call' (performance hint), the comparability constraint across concept_type/embedding_model slices, and the specific semantics of 'review' view. These go beyond what annotations state, so a 4 is warranted.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is five sentences, front-loaded with the core purpose and key constraints. It avoids fluff and includes a useful doc link. It is appropriately concise for the complexity, though it could be slightly more structured by separating the review-view note, but overall it's efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a listing tool with an output schema, the description covers the essential decisions: the two views, the alternative for detail, the comparability caveat, and the read-only nature. It does not explicitly mention pagination behavior, but the schema provides limits/offsets and the output schema covers return structure. The review-view distinction is important and well stated. Overall, nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 63%, so the description should compensate for undocumented parameters (limit, offset, sort_order). It does hint at filtering via 'usage counts' and 'type/model facets', which maps to min_ref_count and concept_types, but it does not clarify pagination or sort semantics. Since the schema already documents the key params (view, search, concept_types, min_ref_count, sort_by), the description adds only marginal value, so a baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Browse') and resource ('organization concepts') and enumerates the key facets returned (aliases, usage counts, type/model). It explicitly differentiates from the sibling get_semantic_concept by directing users there for details and neighbors, so an agent can immediately tell this is the list operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly names the alternative tool (get_semantic_concept) and the condition that selects it (when details/neighbors are needed). It also explains the two view modes, including the special 'review' behavior, which helps the agent choose the right view. The pointer to documentation adds further context for when this tool is appropriate for vocabulary review.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

merge_recordsMerge recordsA
Destructive
Inspect

Fuse two or more records of the same entity into a new arbitration record. Without arbitration_model use voting, median and union rules; with an arbiter, conflicts may incur LLM cost. Returns output, conflicts, fusion metadata and a new record ID. The merged result can feed linked databases and overwrite current entity values. Inspect database warnings and the actual fusion method; an arbiter failure can fall back to rules. Use this after separately recovered model runs, not to merge different entities. See enricher://docs/enrichment-and-fusion.

ParametersJSON Schema
NameRequiredDescriptionDefault
result_idsYesUUIDs of the enrichment records to merge (minimum 2).
attachment_idsNoAttachments passed to the arbitration LLM (ignored for rule-based merges).
arbitration_modelNoModel composite key for LLM arbitration (None = rule-based).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (destructiveHint=true, readOnlyHint=false), the description discloses that the merge can overwrite current entity values, may incur LLM cost, can fall back to rules on arbiter failure, and advises inspecting database warnings and the actual fusion method. This is rich, honest behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact yet information-dense. Every sentence contributes: purpose, arbitration modes, return contents, destructive side effects, fallback behavior, usage guidance, and a docs reference. It is front-loaded with the core purpose and avoids redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers prerequisites, exclusions, cost, failure modes, side effects, and return value categories. Since an output schema exists, detailed return structure is not required. For a destructive, potentially costly merge operation, this is complete enough for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaningful context for arbitration_model by explaining rule-based behavior and LLM cost implications, and clarifies that result_ids must refer to the same entity. This goes beyond the schema's bare parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Fuse two or more records of the same entity') and a clear resource ('records' into a 'new arbitration record'). It also distinguishes itself by emphasizing same-entity merging, which separates it from sibling tools like merge_semantic_concepts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit usage context: use after separately recovered model runs, not for merging different entities. It also explains when rule-based vs LLM arbitration applies. However, it does not name alternative sibling tools or explicitly say when not to use them, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

merge_semantic_conceptsMerge semantic conceptsA
Destructive
Inspect

Merge a loser concept into a winner. Defaults to impact_only=true: inspect counts and affected replicas before obtaining approval to execute. impact_only=false requires owner and rewrites aliases, entity identities and referencing payloads, queuing convergence to linked replicas. No LLM call. Similarity alone does not establish identity; inspect get_semantic_concept and the judge review evidence first. Returns impact or merge results. See enricher://docs/semantic-ids.

ParametersJSON Schema
NameRequiredDescriptionDefault
loser_idYessemantic_id of the concept folded into the winner.
winner_idYessemantic_id of the concept that survives.
impact_onlyNoTrue = report the merge's blast radius only; False = merge (owner role).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate destructiveHint=true and readOnlyHint=false, which the description complements by detailing side effects: impact_only=false rewrites aliases, entity identities, referencing payloads, and queues convergence. It also discloses 'No LLM call' and warns that similarity alone does not establish identity, adding important behavioral context beyond the structured annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and well-structured: it opens with the core action, then explains default behavior, conditional behavior, prerequisites, and output. Every sentence contributes essential information with no redundancy, making it easy for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, two-phase operation, the description covers all necessary aspects: default safety mode, owner requirement, consequences (alias/identity rewrites, convergence queueing), non-LLM nature, prerequisite evidence checks, and output type. It references documentation and related tools, ensuring an agent has enough context to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters. The description adds value by elaborating on impact_only's dual semantics (impact reporting vs. actual merge with owner requirement) and the consequences of each mode, going beyond the schema's brief descriptions. It doesn't add new info for winner_id/loser_id, but the mode explanation justifies a score above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action: 'Merge a loser concept into a winner.' It distinguishes two operational modes (impact_only true/false) and explicitly notes that similarity alone is insufficient, implying this is an identity-based merge rather than a fuzzy match. This separates it from sibling tools like merge_records or update_concept_alias.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit when-to-use guidance: defaults to impact_only=true for inspecting counts and affected replicas before approval, while impact_only=false requires owner role. It also advises inspecting get_semantic_concept and judge review evidence first, and points to enricher://docs/semantic-ids for further details. No ambiguity about prerequisites or modes.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

migrate_semantic_embeddingsMigrate semantic embedding modelAInspect

Inspect or migrate the organization's concept embedding space. action='status' is read-only; preview, start and cancel require owner. preview estimates cost and reports potential collisions; review them before authorizing start. target_model is required for preview/start; source_model and concept_types scope the move. Starting launches billed background re-embedding while enrichment continues, then switches the selected slices after coverage is complete. Cancel marks the transition cancelled; it does not guarantee interruption or rollback of in-flight work. Returns migration status or preview. See enricher://docs/semantic-ids.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesstatus | preview | start | cancel
source_modelNoEmbedding space to move; omitted = the org's current default.
target_modelNoComposite key (provider::model) to move to. Required for preview/start.
concept_typesNoRestrict the move to these concept types; empty = the whole space.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and openWorldHint=true, so the description carries the burden of behavioral disclosure. It discloses that start launches billed background re-embedding, continues enrichment, and switches slices after coverage, plus the limitation that cancel does not guarantee interruption. It also notes owner permissions and cost preview, which are not in annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence adds value: purpose, permission requirements, action-specific guidance, parameter requirements, and behavioral caveats. It front-loads the purpose and organizes details logically, making it efficient for an agent to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and a complex multi-action tool, the description covers permissions, parameter requirements, side effects, and limitations. It leaves no obvious gap for an agent to call it correctly, and it references documentation for deeper detail.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so parameters are documented, but the description adds crucial conditional semantics: target_model is required for preview/start, and source_model defaults to the current space when omitted. It also clarifies that concept_types scopes the move, which is not obvious from the schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool inspects or migrates the organization's concept embedding space, with explicit action types (status, preview, start, cancel). This distinguishes it from other tools by focusing on embedding migration specifically. No other sibling tool appears to cover this domain.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly maps actions to permission levels: status is read-only, while preview, start, and cancel require owner. It also advises running preview before start to review cost and collisions, and clarifies that cancel does not guarantee interruption or rollback. This is actionable guidance for an agent deciding which action to invoke.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

move_schema_propertyMove schema propertyA
Destructive
Inspect

Move one property into the root, an object path or '$defs.X', preserving its flags and expertise. Requires editor; no LLM call. Moving into a definition changes every usage site. An identity composition naming the property is re-spelled automatically when the destination stays within the same 1-1 closure (identity_rebound reports it); the move is refused when it would take the key out of that closure or leave a semantic ID with no identity material, and on collisions or recursive containment. Read source and destination with get_schema_part first. Structural changes on database-linked schemas require publish_schema. Paths and migration guidance: enricher://docs/schema-reference.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesProperty path, e.g. 'ceremonies[].ceremony_type'.
schema_idYesUUID of the saved schema.
new_parent_pathNo'' = root, or an object path / '$defs.X'.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the destructiveHint annotation by explaining side effects (changing every usage site), automatic re-spelling behavior, refusal conditions, collisions, recursive containment, and semantic ID constraints. No contradictions with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is information-dense but every sentence adds operational value: purpose, prerequisites, side effects, refusal conditions, and related tool requirements. It is front-loaded with the core action and structured logically.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, the description covers prerequisites, side effects, failure modes, related publishing requirements, and points to documentation. The output schema fills any remaining return-value expectations, making this complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so parameters are already documented. The description still adds meaning by explaining the destination semantics and the implications of moving into $defs.X, though it does not detail the path format beyond the schema's example.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Move one property') with precise destination options ('root, an object path or $defs.X') and key semantics ('preserving its flags and expertise'). This clearly differentiates it from sibling tools like add_schema_property and update_schema_property.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit operational guidance: requires editor, no LLM call, read source and destination with get_schema_part first, and publish_schema for database-linked schemas. It also covers when the move is refused, which is strong usage context beyond simple alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nest_schema_regionMaterialize entity regionA
Destructive
Inspect

Materialize an entity region from get_schema's x-entityMap. Ordinary flat members (e.g. product_id, product_name on an order line) move into a new object named after the region, regions hanging from it move along (nest them in turn inside the new object), pairing facts stay on the host, and the moved names shed the region's tokens (product_name → name) unless strip_prefix=false. Compact scalar occurrences (e.g. manufacturer_name or each therapeutic_classes[] item) are all converted to references to one shared $defs entity; host_path and strip_prefix do not apply to that form. host_path is '' for the root, an object path, 'path[]' for an array's items, or '$defs.X'. Defaults to dry_run=true: inspect the returned schema_content and notes, then persist with dry_run=false only after approval of the change. Requires editor; no LLM call. Database-linked structural edits still require publish_schema. Modeling consequences: enricher://docs/schema-reference.

ParametersJSON Schema
NameRequiredDescriptionDefault
dry_runNoReport the rewrite without persisting it (the default).
host_pathNoContainer holding the flat members ('' = root, object path, 'path[]', '$defs.X').
region_idYesEntity region id from x-entityMap.regions (get_schema).
schema_idYesUUID of the saved schema.
strip_prefixNoDrop the region's name tokens from the moved field names.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already signal destructiveHint=true, but the description goes far beyond that by explaining the dry-run default, the approval-before-persist workflow, the no-LLM-call constraint, and the publish_schema dependency. It also discloses nuanced transformation behavior such as compact scalar occurrences becoming shared $defs references and host_path/strip_prefix not applying to that form. No contradiction with annotations exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: behavior, special cases, workflow, constraints, and references are all covered without repetition. The most important operational detail (dry_run default and approval requirement) is placed near the end but is clearly highlighted, and the overall structure is logical.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex transformation tool with five parameters, an output schema, and destructive annotations, the description covers all critical aspects: what gets moved, how names change, edge cases, the safe execution workflow, permissions, and downstream dependencies. The output schema exists, so return-value details need not be repeated in the description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema coverage is 100%, the description adds substantial meaning beyond the schema: it enumerates valid host_path forms, clarifies when strip_prefix applies, explains that compact scalar occurrences are converted to shared references, and ties region_id to get_schema's x-entityMap. This materially helps an agent choose correct parameter values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Materialize') and resource ('an entity region from get_schema's x-entityMap'), and goes on to explain exactly what materialization means: moving flat members into a new object, nesting child regions, and stripping prefixes. This clearly distinguishes it from generic schema-editing siblings like update_schema or move_schema_property.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives strong usage context: it requires editor access, makes no LLM call, defaults to dry_run, and requires approval before persisting. It also notes that database-linked structural edits still require publish_schema, providing a clear when-not-to-use boundary. It does not name an explicit alternative tool for the same operation, but the context is sufficient for an agent to route correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

probe_semantic_conceptProbe semantic conceptA
Idempotent
Inspect

Preview identity resolution without adding a concept or increasing its usage. Requires editor; uncached resolution may call embeddings and the identity judge. Returns exact_hit, match or no_match, the matched concept and neighbors. Probe before adding; a matched incumbent may already represent the intended entity. embedding_model can select the space for a new concept type, not change an existing type's space. Inspect judge evidence as well as similarity. See enricher://docs/semantic-ids.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesThe identity text to resolve.
neighborsNo
judge_floorNoSimilarity at or above which a candidate is put to the identity judge. Omit to use the organization default (Settings → Organization).
concept_typeYesConcept type (slice) to resolve against.
embedding_modelNoComposite key (provider::model) to embed a NEW concept type under.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description discloses that uncached resolution may invoke embeddings and the identity judge, requires editor permissions, and asserts no concept addition or usage increase. This gives a clear side-effect/cost picture without contradicting the idempotentHint/openWorldHint annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: purpose, permission/cost, return values, usage decision, parameter caveat, and inspection advice. Key purpose is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema and rich parameter descriptions, this is complete for an agent to call correctly: it explains return semantics, when to use, side effects, and the one non-obvious parameter behavior. Minor omissions like judge_floor default behavior are already in the schema, so not significant.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already covers 80% of parameter semantics, and the description adds meaningful guidance for embedding_model ('select the space for a new concept type, not change an existing type's space'). It also clarifies judge evidence and matched-concept behavior, going beyond the schema's basic descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific action ('Preview identity resolution') and explicitly constrains it ('without adding a concept or increasing its usage'), distinguishing it from mutation siblings like add_semantic_concept. The output vocabulary and use of 'probe' make the tool's role unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Probe before adding; a matched incumbent may already represent the intended entity' is an explicit when-to-use instruction tied to the add workflow. The description also specifies the editor requirement and when embedding_model applies, so an agent can decide between this and add_semantic_concept.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

publish_schemaPublish schemaA
Destructive
Inspect

Publish a database-linked schema's working copy as the contract used by enrichment and replicas. A newly linked schema sends nothing until first publication; unlinked drafts cannot be published. Requires editor; no LLM call. Call validate_only=true to inspect the diff, blockers, warnings and per-database migration_sql. Transform migrations require user-approved confirm_transforms=true; cross-schema conflicts remain blockers. key_language may be required if classification revealed a multilingual key. Returns publication state; queued migrations apply asynchronously on replicas. See enricher://docs/database-sync for the review and delivery workflow.

ParametersJSON Schema
NameRequiredDescriptionDefault
schema_idYesUUID of the schema to publish.
key_languageNoISO 639-1 language for multilingual identity tokens. Normally adopted from the database lock. Supply when key_language_required requests a choice; review the suggested language with the user. Shared by all schemas of that database.
validate_onlyNoDry-run: return the diff + blockers without publishing.
confirm_transformsNoAcknowledge a transform migration (re-key, type change, or a column/table rename) run against the replicas' own data. Required when the preview says requires_confirm — ALWAYS show the transforms to the user and get their explicit approval before passing true.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark destructiveHint=true, and the description adds substantial behavioral context: queued migrations apply asynchronously on replicas, cross-schema conflicts remain blockers, editor permission is required, and no LLM call is made. This goes well beyond the annotation signals.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every clause earns its place: purpose, preconditions, permissions, validation mode, transform confirmation, language requirements, async behavior, and a docs pointer. It is front-loaded with the core purpose and does not waste words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex, destructive publish operation, the description covers prerequisites, permissions, validation paths, async replica behavior, and conflict handling. An output schema exists, so return-value details are not required in the description. A docs reference is included for the full workflow.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Even though schema description coverage is 100%, the description enriches each key parameter: validate_only returns diff/blockers/warnings/migration_sql, confirm_transforms is required for user-approved transform migrations, and key_language is needed when multilingual classification requires it. This adds real decision-making value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Publish a database-linked schema's working copy as the contract used by enrichment and replicas.' It clearly distinguishes publishing from editing or saving by noting that unlinked drafts cannot be published and that a newly linked schema sends nothing until first publication.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives strong usage context: it requires editor permission, supports a validate_only dry-run, and explains when confirm_transforms and key_language are needed. It does not explicitly name sibling tools like save_schema or update_schema as alternatives for draft editing, so it stops short of full alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resolve_unify_proposalResolve unify proposalA
Destructive
Inspect

Resolve one pending entity-type unification proposal from get_schema. action='accept' maps the proposed site onto the winning $def; field_map overrides proposed correspondences. Unmapped fields are retained as nullable; other usages of the winning definition are affected. action='dismiss' keeps the sites separate. Defaults to dry_run=true: inspect the returned schema_content and notes, then persist with dry_run=false only after approval of the change. Requires editor; no LLM call. Database-linked structural edits still require publish_schema. Modeling consequences: enricher://docs/schema-reference.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYes'accept' or 'dismiss'.
dry_runNoReport the rewrite without persisting it (the default).
field_mapNoaccept only: override of the loser→winner field correspondence.
schema_idYesUUID of the saved schema.
proposal_idYesProposal id from x-entityMap.proposals (get_schema).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true, and the description goes well beyond by disclosing the dry_run default and the required inspection-then-persist workflow, the side effect on 'other usages of the winning definition', the 'Requires editor' permission, and that it makes no LLM call. These are concrete behavioral details that materially affect invocation, exceeding what annotations alone convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-organized: it opens with the core action, explains the two modes, then dry_run, then requirements, and ends with a doc link. Every sentence carries useful information; nothing is fluff. It is longer than minimal but each clause earns its place, and the structure is logical.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (5 params, destructive, dry_run flow, side effects, permissions), the description covers all essential aspects: the workflow, safety defaults, required auth, side effects, and the publish_schema dependency. With an output schema present, return values are already documented. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% so each parameter already has a description. The tool description adds semantic context by explaining the meaning of action values ('accept' maps, 'dismiss' keeps separate), the role of field_map as an override, and the dry_run default behavior. This adds value beyond the schema but does not fully redefine the parameters; it's a solid enhancement.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Resolve one pending entity-type unification proposal from get_schema.' It clearly distinguishes the two actions (accept/dismiss) and their effects, making the tool's purpose unambiguous even among many siblings. It also names the source (get_schema) and the separate publish_schema step, which differentiates it from related tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the context (pending proposal from get_schema), the default dry_run behavior and the required approval flow, and explicitly notes that database-linked structural edits still require publish_schema — a clear when-not. However, it does not explicitly contrast with other schema-editing tools like update_schema or save_schema, though the proposal-specific nature makes the use case clear. This is strong but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

retry_expertisesRetry failed expertisesAInspect

Retry only an existing record's failed expertise domains, then update its output and attempt the run's fusion/synchronization. Billed for retried work. Requires failed_expertises on that record (get_record); a surviving successful sibling is not retryable and returns no_failed_expertises. Supply its entity_input_data and saved_schema_id; model optionally substitutes the failed model. Returns job_id: poll get_job_status, then re-read get_record. When the failed leg left no record, use a one-model enrich_entity with database_sync=false followed by merge_records instead. Recovery decisions: enricher://docs/enrichment-and-fusion.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoModel to retry with (provider::model from list_models). Defaults to the record's own model — pass a stronger one when a domain fails repeatedly on it (retrying the same model that just failed usually fails again). Only the failed domains are re-run and re-billed; the record stays attributed to its original model.
languagesNoISO 639-1 codes; defaults to ['en'].
record_idYesEnrichment record with failed expertises.
schema_idNoSaved schema UUID (get_record -> saved_schema_id). Or pass target_schema.
entity_dataYesThe record's original entity input (get_record -> entity_input_data).
target_schemaNoRaw schema dict when no saved schema exists.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses beyond annotations: it states billing ('Billed for retried work'), the condition of failed_expertises, the re-run and re-billing scope for failed domains only, and that the record stays attributed to its original model. It also tells the agent to poll get_job_status and re-read get_record. Annotations only set readOnlyHint=false and destructiveHint=false; the description adds substantial behavioral context without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but each sentence earns its place. It front-loads the core function, then presents prerequisites, parameter guidance, return value, and alternatives succinctly. There is no fluff or repetition; it is an efficient, well-structured text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (6 parameters, nested object, output schema), the description is complete: it specifies the required inputs, the output handling, the alternative path when no record exists, and links to further recovery docs. It covers all aspects an agent needs to invoke the tool correctly without guessing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description enriches parameter meaning significantly: it tells the agent to supply entity_input_data and saved_schema_id from get_record, explains the model parameter's default behavior and when to override it, and clarifies that schema_id can be replaced by target_schema. This goes beyond the schema's field definitions and guides correct invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Retry only an existing record's failed expertise domains') and clarifies that it operates on failed domains only, distinguishing it from sibling tools like enrich_entity and merge_records by explicitly mentioning those alternatives. It also names the output (job_id) and the follow-up steps.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when to use this tool (when a record has failed expertises) and when not ('a surviving successful sibling is not retryable and returns no_failed_expertises'). It names concrete alternatives for the no-record case ('use a one-model enrich_entity with database_sync=false followed by merge_records instead') and directs to a documentation link for recovery decisions. This leaves nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

revert_benchmark_reference_updatesRevert benchmark reference updatesA
Idempotent
Inspect

Undo automatic edits a scoring pass made to a scenario's reference. Every scoring pass folds what the scored models showed the reference should be (a candidate the judge found better, a rule the samples prove, a value the reference lacked) and writes it into the reference — an edit a pass already made is only replaced by stronger evidence (samples, a wrong verdict, more agreeing models), never by one more model's better verdict; the log is reference_meta.auto_applied on get_benchmark_scenario, each entry with its inverse patch. Pass the revision ids to undo: the inverse is applied, the entry is marked reverted and its (path, attribute) is pinned so no later pass re-applies it (a manual set_benchmark_reference lifts every pin). Scores never go stale from this. Requires owner and a plan with benchmarks. No LLM call.

ParametersJSON Schema
NameRequiredDescriptionDefault
scenario_idYesUUID of the scenario.
revision_idsYesIds of the reference_meta.auto_applied entries to undo (1..200).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the annotations by explaining exactly what happens on revert: the inverse patch is applied, the entry is marked reverted, the path/attribute is pinned to prevent re-application, and manual set_benchmark_reference lifts pins. It also states that scores never go stale and that no LLM call is made, which is valuable behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence contributes: purpose, background on how edits are made, the revert mechanism, pinning behavior, prerequisites, and side effects. It is front-loaded with the core purpose and avoids redundant restatement of the title or schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, the description covers what the tool does, where to find revision ids, what happens after reversion, how pins interact with future passes and manual sets, prerequisites, and the fact that scores remain valid. An output schema exists, so return-value details are not required here.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaning by explaining that revision_ids are entries from reference_meta.auto_applied on get_benchmark_scenario and that passing them triggers the inverse patch. This enriches the schema's bare parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Undo automatic edits a scoring pass made to a scenario's reference.' It also distinguishes this from manual reference changes by referencing set_benchmark_reference and automatic scoring-pass edits, making the tool's unique role clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies when to use the tool: when you need to undo automatic scoring-pass edits to a scenario reference. It points to get_benchmark_scenario for finding revision ids and mentions set_benchmark_reference as the manual counterpart, but it does not explicitly state exclusions or a direct 'use X instead' rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_benchmarkRun benchmarkAInspect

Start billed asynchronous execution and scoring of a benchmark. Requires owner, a benchmark-enabled plan and a judge; a verified reference is also required except for sample generation. Supply model_keys or providers; omitting both runs all active models with usable provider keys. Repetitions and judging increase cost. Returns job_id and total_models: poll get_job_status, then read get_benchmark_scenario_results. Re-running replaces each selected model's previous result. An organization's benchmark runs and scoring passes execute one at a time: a launch while another is in flight is queued (queue_position, status 'pending'), and a launch on a scenario whose run is still queued folds its models into that run (merged=true, job_id names the queued run). Reference setup and score interpretation: enricher://docs/model-benchmark.

ParametersJSON Schema
NameRequiredDescriptionDefault
providersNoProvider names (e.g. ['anthropic', 'mistral']) — runs every active model of those providers that has a valid key.
model_keysNoExplicit model composite keys. Overrides `providers`.
scenario_idYesUUID of the scenario to run.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description discloses billing, asynchronous execution, queueing behavior, model folding into queued runs, and the side effect that re-running replaces previous results. It also explains the merged=true and queue_position semantics. This is rich behavioral context that annotations alone would not provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: purpose, prerequisites, parameter behavior, return/next steps, side effects, queueing semantics, and a docs pointer. It is front-loaded with the core action and remains structured despite covering a complex tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, the presence of an output schema, and the rich annotations, the description covers prerequisites, side effects, queueing, cost, follow-up tools, and documentation. Nothing essential for invoking the tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds value by explaining the relationship between model_keys and providers, the behavior when both are omitted, and the cost impact of repetitions and judging. This goes beyond the schema's per-parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Start billed asynchronous execution and scoring of a benchmark.' It clearly distinguishes this from sibling tools like get_benchmark_scenario_results and get_job_status by framing it as the launch action, not a read or setup action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear operational context: prerequisites (owner, plan, judge, verified reference), model selection behavior, cost implications, and the expected follow-up polling/read workflow. It does not explicitly name alternatives or state when not to use this tool, but the guidance is strong enough for an agent to select it correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_schemaSave schemaAInspect

Save a directly authored schema and return its ID and link. Requires editor; no LLM call or generation charge. schema_content is a JSON Schema 2020-12 document in Entity Enricher's supported dialect: title, type='object', properties, optional $defs and x-* extension sections. The server validates it and makes colliding names unique. Entity definitions, enum vocabularies and localized fields have different projection rules. Unknown keywords are dropped, not rejected: read ignored_keywords (path, keyword, hint) and applied_repairs in the result — a property flag such as semantic_id placed on a $defs entity object lands there, with the level it is read at. Use create_schema_from_sample to derive a schema from data, or update_schema for an existing schema. A minimal valid example and supported annotations are in enricher://docs/schema-reference.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
tagsNoOptional tags.
is_pinnedNoPin it to the top of listings.
schema_contentYesFull schema document: title, type=object, properties and optional $defs/x-* sections. See enricher://docs/schema-reference for a minimal example.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the annotations by explaining that the server validates the schema, makes colliding names unique, drops unknown keywords rather than rejecting them, and returns ignored_keywords and applied_repairs. It also clarifies that there is no LLM call or generation charge. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: purpose, prerequisite, schema format, validation behavior, return diagnostics, and routing to alternatives are all covered without filler. The most decision-relevant facts are front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with four parameters and a nested object, the description covers prerequisites, validation behavior, return information, and points to a reference doc for a minimal example and supported annotations. An output schema exists, so the return shape does not need to be fully spelled out here.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, so the schema already documents name, tags, and is_pinned. The description adds substantial meaning for schema_content by specifying the JSON Schema 2020-12 dialect, supported sections, projection rules, and the behavior for unknown keywords. This compensates well for the nested object's complexity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Save a directly authored schema and return its ID and link.' It also differentiates from siblings by naming create_schema_from_sample and update_schema as alternative operations, so an agent can tell exactly what this tool is for.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly provides when-to-use guidance and alternatives: 'Use create_schema_from_sample to derive a schema from data, or update_schema for an existing schema.' It also states a prerequisite ('Requires editor'), which helps the agent decide whether it can invoke the tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_benchmark_referenceSet benchmark referenceAInspect

Save the gold reference for an enrichment or schema-generation benchmark. Only set reference_verified=true after checking its values against trusted evidence or obtaining human sign-off; a generated answer alone is not verification. Schema-generation references must be GeneratedJsonSchema objects. Sample-generation scenarios reject references because they are rubric-scored. Requires owner and a plan with benchmarks; no LLM call. A verified reference enables run_benchmark. See enricher://docs/model-benchmark.

ParametersJSON Schema
NameRequiredDescriptionDefault
sourceNoProvenance: generated | pasted | record (default 'pasted').
scenario_idYesUUID of the scenario.
reference_outputYesExpected entity JSON for enrichment, or a schema document for schema_generation. Sample-generation scenarios do not accept a reference.
source_record_idNoRecord UUID the reference was copied from (when source='record').
reference_verifiedNoExplicit sign-off that the reference is correct (required to run).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations are minimal (readOnlyHint=false, destructiveHint=false, openWorldHint=false), so the description carries the full burden of behavioral disclosure. It adds substantial context: no LLM call, prerequisite of owner and plan, the dependency on verification to enable run_benchmark, and the requirement for GeneratedJsonSchema objects in schema-generation. This goes well beyond the annotations and is accurate with no contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-organized paragraph. It front-loads the core purpose, then delivers critical caveats (verification, schema-generation type, sample-generation rejection), then prerequisites, dependency, and a doc link. Every sentence adds value; there is no redundancy or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (5 parameters, nested objects, output schema), the description covers all essential context: when to use it, constraints on reference types, verification requirements, prerequisites, and the downstream effect on run_benchmark. The link to documentation provides further detail. Nothing an agent needs to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already has 100% coverage, so the baseline is 3. The description adds meaning by clarifying the semantics of reference_verified (explicit sign-off, not a generated answer), the type restriction for schema-generation (GeneratedJsonSchema objects), and the rejection of references for sample-generation. This enriches the schema's parameter descriptions, particularly for reference_verified and reference_output, justifying a 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's action: saving the gold reference for enrichment or schema-generation benchmarks. It distinguishes itself from siblings by clarifying the scope (enrichment/schema-generation, not sample-generation) and by noting it enables run_benchmark. It also specifies prerequisites, making the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-not-to-use guidance: sample-generation scenarios reject references, and reference_verified should only be set true after verification or sign-off. It also states requirements (owner and plan with benchmarks). However, it does not explicitly name an alternative tool for setting references in sample-generation scenarios, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

start_batch_enrichmentStart batch enrichmentAInspect

Start billed asynchronous enrichment of an entity list against exactly one of schema_id or target_schema. Returns job_id and total. No fixed entity-count cap; live prompt quotas and credits can stop remaining work. Each entity follows the enrichment/fusion pipeline; every model must succeed for its automatic fusion and database admission. A confident classification mismatch skips that entity without enrichment, never pauses the batch. Poll get_job_status, then list_records(job_id=...). Attachments apply to every entity. Unlike enrich_entity, this tool exposes neither database_sync=false nor web-search activation. Input contracts, partial results and recovery: enricher://docs/batch-enrichment.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelsNoModel composite keys (call list_models to discover them). Optional: omit (or pass ['auto']) to let the server pick the org's default model — pinned per-task default if set, else the best blended benchmark score.
entitiesYesEntities to enrich (each a free-form dict naming the entity — the schema's identifying fields ideally, but any field names work).
strategyNoauto (default) | single_pass | expert_domains | multi_expertiseauto
languagesNoISO 639-1 codes; defaults to ['en'].
schema_idNoUUID of a saved schema. Mutually exclusive with target_schema.
target_schemaNoInline schema document in the supported Entity Enricher dialect. Prefer schema_id to link records. See enricher://docs/schema-reference.
attachment_idsNoAttachment UUIDs applied as source material to every entity.
arbitration_modelNoOptional LLM for auto-fusion conflict resolution (None = rule-based).
classification_modelNoOptional classifier key. Confident mismatches skip that entity; softer verdicts become prompt context. Batch classification never pauses.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate mutation (readOnlyHint=false) but the description adds substantial detail: it's billed, has no fixed entity-count cap but live quotas can stop it, each entity must pass all models for fusion/admission, and confident mismatches skip without pausing the batch. This goes well beyond the annotations and is highly useful.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-structured, leading with the core action and constraints, then returns, caveats, workflow, and sibling differentiation. Every sentence adds value, though it is on the longer side. Given the complexity (9 params, async behavior), this length is justified.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with this complexity, the description covers billing, async behavior, failure modes, model selection, workflow, attachments, differences from the sibling, and points to docs for detailed contracts. The output schema likely covers the return shape (job_id, total), so nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaningful context beyond schema fields: the model auto-selection logic ('pinned per-task default... else best blended benchmark score'), the classification_model behavior ('Confident mismatches skip that entity; softer verdicts become prompt context'), and the one-of constraint reiterated. It also clarifies attachments apply to every entity. This elevates it to a 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Start billed asynchronous enrichment of an entity list' and immediately constrains it to 'exactly one of schema_id or target_schema', which clearly distinguishes it from the single-entity enrich_entity sibling. It also states the return values (job_id and total) and the workflow (poll get_job_status then list_records).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly contrasts with enrich_entity ('Unlike enrich_entity, this tool exposes neither database_sync=false nor web-search activation') and gives a post-call workflow ('Poll get_job_status, then list_records(job_id=...).'). It also references documentation for input contracts and recovery, leaving no doubt about when to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sync_records_to_databaseInject records into the database syncA
Destructive
Inspect

Validate and inject stored or supplied enrichment output into the entity layer and linked syncs. May incur semantic-resolution cost. record_id alone reuses its output; adding structured_output creates a new derived record. Without record_id, supply structured_output and saved_schema_id. Each item uses the current published contract and admission gate. Only enrichment/arbitration records qualify; base records whose fusion is in the same request are skipped. Inspect each outcome for rejected or partial writes. Use after fixing a rejected output or an enrich_entity run with database_sync=false. See enricher://docs/enrichment-and-fusion.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemsYesEntities to inject. Each item: {record_id?, structured_output?, saved_schema_id?} — at least one of record_id / structured_output, and saved_schema_id required when there is no record_id.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnlyHint=false, destructiveHint=true), the description discloses additional behaviors: 'May incur semantic-resolution cost,' 'Each item uses the current published contract and admission gate,' and 'Inspect each outcome for rejected or partial writes.' These are significant operational traits not captured in annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-structured, front-loading the core action and then logically adding cost, parameter rules, qualifications, and usage guidance. Every sentence adds value with no fluff, making it efficient for an agent to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (one parameter with intricate logic) and that an output schema exists, the description covers all necessary aspects: what it does, when to use it, behavioral expectations, parameter semantics, and a pointer to docs. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

While the schema covers the items structure (100% coverage), the description adds crucial semantics: 'record_id alone reuses its output; adding structured_output creates a new derived record. Without record_id, supply structured_output and saved_schema_id.' This explains the interplay and required combinations beyond the schema's basic description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb and resource: 'Validate and inject stored or supplied enrichment output into the entity layer and linked syncs.' It distinguishes this from sibling tools like enrich_entity and create_database_sync by specifying it operates on enrichment output and injects into syncs. The purpose is unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use guidance: 'Use after fixing a rejected output or an enrich_entity run with database_sync=false.' It also clarifies what qualifies: 'Only enrichment/arbitration records qualify; base records whose fusion is in the same request are skipped.' This gives clear direction and differentiates from alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_benchmark_scenarioUpdate benchmark scenarioA
Destructive
Inspect

Edit a benchmark's test definition or scoring configuration. Requires owner and a plan with benchmarks; no model run. Only supplied fields change, but sample_params and schema_gen_params replace their parameter objects wholesale. The judge may be replaced, not cleared; scenario_type is immutable. Changed test definitions make previous results stale. Re-run affected models to refresh their scores. See enricher://docs/model-benchmark.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNo
strategyNo
languagesNo
schema_idNoNew saved-schema UUID.
descriptionNoFree-text note shown in the Benchmarks tab; no model ever reads it.
entity_dataNoEnrichment: replacement fixed entity input (same contract as enrich_entity's entity_data; refused when it carries no value).
repetitionsNo
scenario_idYesUUID of the scenario.
sample_paramsNosample_generation: full replacement task params object {request, typical_object, naming_convention, language, enable_web_search}.
entity_samplesNoschema_generation: replacement input samples (1..20, whole list).
scoring_sourceNoFeed this scenario's results into the live per-model scores of the options API: 'organization' (owner+), 'global' (system admin only, fallback for orgs without their own source), or 'off' to stop using it.
schema_gen_paramsNoschema_generation: full replacement task params object {generate_semantic_ids}.
scoring_judge_model_keyNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare destructiveHint=true and readOnlyHint=false, but the description adds substantial behavioral context: partial-update semantics ('Only supplied fields change'), wholesale replacement of sample_params and schema_gen_params, the replace-not-clear constraint on the judge, scenario_type immutability, authorization requirements, and the side effect that changed definitions make previous results stale. This goes well beyond what annotations convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Five dense sentences, no filler. The purpose is front-loaded, and every subsequent sentence adds a distinct piece of operational knowledge: prerequisites, partial-update semantics, replacement behavior, immutability constraints, and staleness consequences. The docs pointer is a single efficient clause.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 13 parameters, an output schema (so return values need not be described), and annotations covering the safety profile, the description covers the critical behavioral gotchas — replacement semantics, immutability, side effects, and authorization. It does not enumerate every parameter's interaction, but the schema's 62% coverage plus the description's focus on the riskiest behaviors makes this adequately complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema description coverage at 62%, the schema does not carry the full burden, and the description compensates by explaining the replacement behavior of sample_params/schema_gen_params and the replace-not-clear constraint on the scoring judge. It adds meaning the raw schema (which only labels them 'replacement' objects) does not fully convey, though some param-level richness is already present in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Edit a benchmark's test definition or scoring configuration.' This clearly distinguishes it from siblings like create_benchmark_scenario, get_benchmark_scenario, run_benchmark, and delete_benchmark_scenario, since it names exactly what aspect of the benchmark is being mutated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context: the prerequisite ('Requires owner and a plan with benchmarks'), an explicit when-not ('no model run', distinguishing it from run_benchmark), and follow-up guidance ('Re-run affected models to refresh their scores'). It stops short of naming sibling alternatives by name, which is the only reason it doesn't earn a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_concept_aliasUpdate concept surface formA
Destructive
Inspect

Remove or promote a concept alias using alias IDs from get_semantic_concept. Requires editor; no LLM call. action='remove' stops that surface form resolving to this concept; removing the last alias is refused. action='set_canonical' changes its displayed form. To add an alias use add_semantic_concept(alias_of=...). Returns the outcome. See enricher://docs/semantic-ids.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYes'remove' prunes the surface form; 'set_canonical' promotes it.
alias_idYesThe surface-form row's own id (UUID).
concept_idYesThe concept's semantic_id (UUID).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Even though annotations already flag readOnlyHint=false and destructiveHint=true, the description adds substantial behavioral context: the exact effect of each action, the guardrail on removing the last alias, the editor permission requirement, the note that no LLM call is made, and that the outcome is returned. This surpasses the annotations' binary flags.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: it opens with the core purpose, then provides necessary usage context, action behavior, alternatives, and result. Every sentence contributes distinct value—no filler or restating of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the three required parameters, an output schema, and annotations already covering safety, the description covers permissions, action semantics, a critical edge case (last alias removal), and points to further docs. An agent has everything needed to invoke the tool correctly and interpret its result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline 3 applies. The description goes further by explaining the semantic effect of the action enum values ('removing the last alias is refused', 'promotes its displayed form') and by telling the agent where to obtain alias_id and concept_id (from get_semantic_concept). It does not redefine the UUID parameters, but the schema already does.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb ('remove or promote') applied to a specific resource ('concept alias'), and names the source of IDs ('get_semantic_concept'). It differentiates between the two action modes and explicitly directs adding aliases to a sibling tool, so an agent can distinguish it from add_semantic_concept or delete_semantic_concepts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says when to use the remove action, when to use set_canonical, and gives a concrete alternative for adding an alias ('use add_semantic_concept(alias_of=...)'). It also states the prerequisite ('using alias IDs from get_semantic_concept') and a constraint ('removing the last alias is refused').

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_schemaUpdate schemaA
Destructive
Inspect

Edit a saved schema's metadata or replace its full schema_content without an LLM call. Requires editor. Only supplied values change. For one property, prefer update_schema_property, add_schema_property or move_schema_property. Replacements must use GeneratedJsonSchema; unknown keywords are dropped, not rejected, and reported in ignored_keywords / applied_repairs. On database-linked schemas this edits the working copy: neutral edits propagate automatically; structural edits take effect after publish_schema. Returns the updated schema and link. Contract and edit workflow: enricher://docs/schema-reference.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoNew name (must be unique).
tagsNoReplacement tag list.
is_pinnedNoPin or unpin.
schema_idYesUUID of the schema to update.
key_languageNoPre-set the key language for multilingual database keys ahead of a database link (ISO 639-1). Settable only while the schema has no linked database and no entity state; locked afterwards.
schema_contentNoFull replacement schema document, following get_schema.schema_content. See enricher://docs/schema-reference.
ambiguity_check_enabledNoEnable/disable the ambiguity check for this schema — the pass that flags properties whose name admits more than one meaning (gates analyze_schema and the generation post-pass).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate a destructive, non-read-only operation, but the description adds substantial behavior beyond that: replacements must use GeneratedJsonSchema, unknown keywords are dropped (not rejected) and reported in ignored_keywords / applied_repairs, and database-linked edits operate on a working copy with conditional propagation. It also discloses a docs reference for the contract and edit workflow. No statement contradicts the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence earns its place: purpose, auth, sibling routing, replacement invariants, linked-schema behavior, return value, and doc link. The most important scoping sentence is front-loaded. It is dense but not bloated, with no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex mutation tool with 7 parameters, a destructive annotation, and linked-schema workflow nuances, the description covers what an agent needs to call it correctly: the target, the authorization check, the exact replacement format, the unknown-keyword behavior, the publish/gating semantics, and the return shape. An output schema exists for return details, so the description need not duplicate that.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema coverage is 100%, so the baseline is 3. The description adds meaningful semantics beyond the field descriptions: it frames the operation as 'only supplied values change', imposes the GeneratedJsonSchema constraint on schema_content, and explains the working-copy/publish behavior that affects how parameters like schema_content behave on linked schemas. It does not enumerate each parameter, but the schema already does that, so the extra value is above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-resource pairing: 'Edit a saved schema's metadata or replace its full schema_content without an LLM call.' It clearly distinguishes itself from sibling tools by naming update_schema_property, add_schema_property, and move_schema_property as the alternatives for single-property edits. The scope is unambiguous and it does not rely on the title alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when to prefer alternative tools ('For one property, prefer update_schema_property, add_schema_property or move_schema_property') and conditions the sibling choices. It also gives concrete guidance for database-linked schemas, explaining that neutral edits propagate automatically while structural edits take effect after publish_schema. The requirement for an editor role and the pointer to the contract docs round out practical usage conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_schema_propertyUpdate schema propertyA
Destructive
Inspect

Edit or remove one property by path without replacing the full schema. Requires editor; no LLM call. Only supplied fields change; null clears an entry in flags. Server validation rejects or normalizes invalid combinations and reports applied_repairs. Editing inside $defs affects every usage site. Removing an identity-source member is refused until its identity is recomposed. Renames preserve identity references and record migration intent. Database-linked schemas remain working-copy edits until publish_schema; inspect migration_grade. For additions or relocation use add_schema_property or move_schema_property. Paths, flags and rename rules: enricher://docs/schema-reference.

ParametersJSON Schema
NameRequiredDescriptionDefault
refNo'#/$defs/X' or '#/$enums/X'; clears type.
pathYesProperty path, e.g. 'ceremonies[].ceremony_type'.
typeNoNew JSON type (string/number/integer/boolean/array/object); clears $ref.
flagsNoFlag updates (null clears): expertise, preserve, multilingual, language_discriminator, identifying, nullable, format, pattern, semantic_id, judge_floor, semantic_concept_type, semantic_embedding_model, database_key, db_type, db_type_length, index, unique_group, shared, ordered, db_name, db_name_absolute.
removeNoDelete the property instead.
examplesNo
new_nameNoRename the property.
schema_idYesUUID of the saved schema.
descriptionNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations by disclosing permissions ('Requires editor'), validation behavior ('rejects or normalizes invalid combinations and reports applied_repairs'), side effects ('Editing inside $defs affects every usage site'), restrictions ('Removing an identity-source member is refused'), and state implications ('remain working-copy edits until publish_schema'). Annotations only indicate destructiveHint=true; the description adds rich behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place. It front-loads the core action, then covers permissions, validation, side effects, restrictions, rename behavior, and routing to siblings. No fluff or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (9 params, mutation, side effects, restrictions) and the existence of an output schema, the description covers everything an agent needs: permissions, validation, side effects, restrictions, alternatives, and a pointer to detailed rules. Nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 78%, so the schema already documents most parameters. The description adds meaningful value by clarifying that null clears entries in flags, pointing to the external reference for path/flag/rename rules, and noting that ref clears type and type clears $ref. This goes beyond the schema's basic descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Edit or remove'), a resource ('one property by path'), and the key behavior (without replacing the full schema). It also distinguishes itself from siblings add_schema_property and move_schema_property, making its purpose unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names when to use this tool versus alternatives: 'For additions or relocation use add_schema_property or move_schema_property.' It also clarifies that database-linked schemas remain working-copy edits until publish_schema, giving clear context on when this is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upload_attachmentUpload attachmentAInspect

Upload base64 file bytes as reusable source material; returns id and requires_capability. Pass the ID in attachment_ids to sample/schema generation, enrichment or benchmarks. Supported format handling depends on server MIME policy: extracted text or model-readable binary. Prefer auto model selection for attachment capabilities. Uploading does not itself run an LLM. A generated sample with attachments is source-only; see enricher://docs/documents for formats, multiple-file behavior and research workflows. Retain attachments needed for later runs or regeneration.

ParametersJSON Schema
NameRequiredDescriptionDefault
filenameYesOriginal filename including extension (e.g. 'report.pdf').
media_typeNoOptional MIME hint; the server still sniffs the magic bytes.
content_base64YesThe file's bytes, base64-encoded (no data: prefix).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are minimal (readOnlyHint=false, destructiveHint=false, openWorldHint=false), so the description carries the behavioral burden. It adds useful behavior: MIME policy affects format handling, uploading does not run an LLM, and attachments are reusable source material. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is several sentences long but each sentence serves a purpose: core action, usage, MIME behavior, LLM clarification, source-only note, and retention advice. It is not overly verbose, though it could be tightened by merging related ideas. The most important info is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema (not shown but exists) and the description mentions return values (id, requires_capability). It covers formats, usage, and references for more details. It is sufficiently complete for an agent to call it correctly, though it could specify auth requirements if any.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with each parameter (filename, media_type, content_base64) documented. The description does not add extra parameter-level semantics beyond what's in the schema; it reinforces the base64 encoding and use of the returned ID. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Upload base64 file bytes as reusable source material') with a clear resource and outcome (returns id and requires_capability). It also explains how the ID is used downstream, distinguishing it from delete_attachment and other siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use the tool (for sample/schema generation, enrichment, benchmarks) and gives a pointer to docs for formats and workflows. It also notes 'Prefer auto model selection for attachment capabilities' and clarifies that uploading does not run an LLM. It does not explicitly name alternatives or exclusion conditions, but the usage context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool update
    • Changedcreate_schema_from_sample1 field changed
      • addedInput schema / properties / samples_csv
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Samples as CSV text instead of entity_samples (never both): the first row is ALWAYS the header, each data row one sample. Delimiter ',', ';' or tab; headers become identifier keys ('Author Name' -> author_name); each column gets one type (integer, number, boolean or text; empty cell = null; decimal commas read in ';'/tab text). Over 20 rows, 20 are kept, covering every column. Read the result's csv_import and relay the kept rows and renamed headers.",
        +  "title": "Samples Csv"
        +}
  2. 1 tool update
    • Changedlist_records1 field changed
      • changedInput schema / properties / record_type / description
        Previous value: -"Filter by type: enrichment | sample_generation | schema_generation | schema_edit | playground | classification | arbitration | ambiguity_analysis | db_classification | benchmark_scoring"New value: +"Filter by type: enrichment | classification | arbitration | sample_generation | schema_generation | schema_edit | schema_annotation | ambiguity_analysis | db_classification | benchmark_scoring | playground"
  3. 58 tool updates
    • First observedack_database_deltas
    • First observedadd_schema_property
    • First observedadd_semantic_concept
    • First observedanalyze_sample
    • First observedanalyze_schema
    • First observedanswer_job_question
    • First observedassign_sync_host
    • First observedcancel_job
    • First observedclassify_database_model
    • First observedcreate_benchmark_scenario
    • First observedcreate_database_sync
    • First observedcreate_schema_from_sample
    • First observeddelete_attachment
    • First observeddelete_benchmark_scenario
    • First observeddelete_database_sync
    • First observeddelete_schema
    • First observeddelete_semantic_concepts
    • First observedenrich_entity
    • First observedfetch_database_deltas
    • First observedgenerate_sample
    • First observedget_benchmark_scenario
    • First observedget_benchmark_scenario_results
    • First observedget_database_setup_instructions
    • First observedget_enum_candidates
    • First observedget_job_status
    • First observedget_record
    • First observedget_schema
    • First observedget_schema_part
    • First observedget_semantic_concept
    • First observedget_stats
    • First observedimport_semantic_concepts
    • First observedlist_benchmark_scenarios
    • First observedlist_database_syncs
    • First observedlist_entity_states
    • First observedlist_models
    • First observedlist_records
    • First observedlist_schemas
    • First observedlist_semantic_concepts
    • First observedmerge_records
    • First observedmerge_semantic_concepts
    • First observedmigrate_semantic_embeddings
    • First observedmove_schema_property
    • First observednest_schema_region
    • First observedprobe_semantic_concept
    • First observedpublish_schema
    • First observedresolve_unify_proposal
    • First observedretry_expertises
    • First observedrevert_benchmark_reference_updates
    • First observedrun_benchmark
    • First observedsave_schema
    • First observedset_benchmark_reference
    • First observedstart_batch_enrichment
    • First observedsync_records_to_database
    • First observedupdate_benchmark_scenario
    • First observedupdate_concept_alias
    • First observedupdate_schema
    • First observedupdate_schema_property
    • First observedupload_attachment

Related MCP Connectors

Related MCP Servers

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.