Pathwize
Server Details
Run data annotation and evaluation tasks with vetted domain experts: create, publish, get results.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-06-18
- URL
TDQS
Scored across 10 tools
Each tool targets a distinct action: task lifecycle (create/update/publish/close), reads (get_task/list_tasks/get_results), item ingestion (add_items), and pool/dashboard access. The only mild adjacency is add_items vs update_task (both touch task content), but descriptions clearly separate item ingestion from settings changes.
All ten tools follow a consistent snake_case verb_noun pattern (add_items, create_task, get_results, list_pools, publish_task, etc.). No mixed conventions or vague standalone verbs.
Ten tools is well-scoped for a labeling/evaluation task management server, and each tool maps to a necessary step in the create-add-publish-review workflow. Nothing feels redundant or padded.
The surface covers the full task lifecycle (create, update, publish, close, get, list), item ingestion, results retrieval, and pool listing. Minor gaps exist: no delete/archive task, no item removal, and no pool/team management, but core workflows have no dead ends.
Available Tools
10 toolsadd_itemsAdd items to an AI taskAIdempotentInspect
Add items by public URL; Pathwize copies the files into its own storage. reference: true adds them to the qualification set (optionally with gold_label for automatic scoring). Idempotent per id, up to 1000 per call.
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | ||
| taskId | Yes | ||
| reference | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare idempotentHint=true, which the description reinforces with 'Idempotent per id'. Beyond that it discloses real behavioral context: files are copied into Pathwize storage (a side effect worth knowing), a hard cap of 1000 per call, and the downstream effect of reference:true with gold_label. It still doesn't describe the return value or what a re-add of an existing id does.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact clauses, front-loaded with the core action and storage behavior, then the reference/gold_label semantics and the idempotency/limit. Dense and semicolon-driven but every clause carries information; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description should ideally say what a successful call returns or how partial failures are reported, and it doesn't. It covers the storage side effect, idempotency, and the batch limit adequately, but the response shape and taskId semantics remain gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden, and it only explains reference and gold_label (qualification set, automatic scoring). The required taskId and the per-item url/id/mime parameters are left entirely undocumented. Partial compensation for a low-coverage schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: adding items to a task by public URL, with the notable detail that files are copied into Pathwize's own storage. This is clearly distinguishable from create_task or publish_task. It stops short of explicitly naming which sibling to prefer in ambiguous cases, so it lands at a clear-but-undifferentiated 4.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the usage context (items must be reachable by public URL, and reference:true routes them into the qualification set), which is enough to infer when the tool applies. It does not state when NOT to use it, nor does it point to any alternative tool for other item-adding scenarios. Implied guidance only.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
close_taskClose or pause an AI taskAIdempotentInspect
closed ends the task, draft pauses it. Experts are not emailed.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | ||
| taskId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide only idempotentHint, so the description carries most of the behavioral burden. 'Experts are not emailed' is a genuinely useful side-effect disclosure absent from structured data, but the description says nothing about permissions required, whether closing is reversible, or effects on dependent tasks.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Eleven words, front-loaded with the state semantics and ending with the notification caveat; every clause earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter state-change tool with no output schema, the description covers enum semantics and one side effect, which is adequate but leaves reversible-ness and permission requirements unstated for a mutation operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does add real meaning for the 'status' enum values, which the schema lists without explanation, but the required 'taskId' identifier is never described.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description specifies what the two status values do ('closed ends the task, draft pauses it'), which is the substantive purpose beyond the tautological name. It does not differentiate itself from siblings like update_task or publish_task, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Choosing between 'closed' and 'draft' is implied by the definitions of each value, but there is no guidance on when to prefer close_task over update_task or publish_task, and no stated prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_taskCreate an AI task (draft)AInspect
Create a labeling or evaluation task as a draft; nobody is notified until publish_task. Typical order: list_pools, create_task with instructions, classes, rate, pool_ids and one example image per class, add_items, then publish_task after the user confirms.
| Name | Required | Description | Default |
|---|---|---|---|
| rate | No | Hourly pay for the expert in EUR | |
| type | No | Default data_labeling | |
| title | Yes | ||
| domain | No | ||
| classes | No | Class names experts choose from | |
| deadline | No | ||
| examples | No | Example images by public URL with the class they show. Labeling tasks need one per class before they can open. | |
| language | No | Language of the instructions | |
| pool_ids | No | Pools the task runs on (see list_pools) | |
| expertise | No | What experts need to know, one line | |
| allow_note | No | ||
| annotation | No | Default classification | |
| instructions | No | Instructions for the experts in Markdown. A table whose first column names the classes becomes the per-class briefing. | |
| members_only | No | Only members of these pools see the task, even on public pools | |
| multi_select | No | ||
| qualification | No | ||
| allow_cant_tell | No | Let experts skip an item without a class (default true) | |
| seconds_per_item | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it delivers the single most important behavioral fact: the task is a draft and nobody is notified until publish_task. It also discloses the labeling prerequisite (one example image per class before a task can open). Gaps remain on permissions/ownership and error behavior, so it stops short of 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the draft/no-notification fact and then the call sequence. Dense with actionable detail and no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 18-parameter mutation tool with no annotations and no output schema, the description covers the essential setup sequence and the draft semantics that an agent must know. It doesn't touch several consequential parameters (qualification, members_only, deadline), which is a real but minor gap given the schema documents them.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 61%, so a moderate baseline applies. The description names instructions, classes, rate, pool_ids and examples and adds the 'one example image per class' constraint for examples plus a pointer to list_pools for pool_ids, but the other documented parameters (qualification, deadline, annotation, language, etc.) get no added meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (create) and resource (labeling/evaluation task), and pins the scope with 'as a draft.' The parenthetical distinction from publish_task and update_task is clear enough that an agent can route without opening a sibling schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives the explicit workflow order: list_pools, create_task, add_items, publish_task after user confirmation. This names the alternatives to reach for and the condition ('after the user confirms') that gates the follow-up call.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_resultsGet labeled resultsCRead-onlyInspect
Finished items with final label, agreement, notes and flags. Page with starting_after (the last item id).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| since | No | ||
| taskId | Yes | ||
| starting_after | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=true already establishes the safe-read profile, so the description isn't burdened with that. It adds useful return-content context ('final label, agreement, notes and flags') and a pagination cue, but omits whether results exist only for closed tasks, how paging terminates, or any ordering/rate-limit behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler, and the returned-field summary leads before the paging note. It is efficient, though arguably under-specified rather than deliberately concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, 0% parameter coverage, and no annotation detail beyond readOnly, the description leaves too much open: three of four parameters are unexplained and the response shape is only hinted at by a field list. For a 4-parameter retrieval tool this is insufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 4 parameters, so the description must compensate. It explains only starting_after ('the last item id') and implies the paging pattern; limit, since, and taskId receive no clarification of format, units, or effect.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource precisely: 'Finished items with final label, agreement, notes and flags,' which tells an agent this returns completed/labeled results rather than task metadata. However, it never differentiates itself from siblings like get_task or list_tasks, and the verb is only implied by the tool name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this versus get_task, list_tasks, or the other read siblings, nor any prerequisite (e.g. that the task must be finished/closed). The only usage note is the pagination mechanism, which is a mechanical detail rather than a selection criterion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_taskGet an AI taskBRead-onlyInspect
One task with settings, classes, examples, qualification, reference item count and item counts by status.
| Name | Required | Description | Default |
|---|---|---|---|
| taskId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=true already signals a safe read, so the bar is lower. The description adds useful context by disclosing what is bundled into the response (settings, classes, examples, qualification, item counts), but says nothing about failure modes (unknown taskId) or auth requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with no filler, though it is a noun-phrase fragment rather than a front-loaded statement of action. Nothing is wasted, but it reads more like a field list than a structured definition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description reasonably compensates by enumerating the returned fields, which is good. For a one-parameter read-only tool this covers most needs, but parameter origin and error behavior remain undocumented.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single taskId parameter at 0% schema description coverage, so the description should carry meaning. It implies which task is returned but never mentions taskId, its UUID format, or where the id comes from; the parameter is self-evident, but compensation for the coverage gap is absent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description makes clear this returns a single task ("One task") and enumerates the payload contents (settings, classes, examples, qualification, counts), which distinguishes it from sibling list_tasks. It lacks an explicit verb, but the resource and scope are unambiguous for an agent to select it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no named alternative (e.g. list_tasks for many tasks, get_results for outcomes). The singular-vs-plural distinction with list_tasks must be inferred from the word "One".
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_poolsList expert poolsARead-onlyInspect
Pools a task can run on: active curated pools plus your own, with member counts.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=true already declaring the safety profile, the description adds real behavioral value: it discloses which pools are returned (active curated plus the user's own) and that member counts are included. This is meaningful scope information beyond the annotation, though ordering/pagination behavior is unstated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact clause that front-loads the resource and immediately qualifies scope and returned metadata. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only list with no output schema, the description covers what will be returned, which is the main thing an agent needs. Only minor details like result ordering are absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. There is nothing for the description to clarify on the parameter side.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (expert pools) and scope ('pools a task can run on', 'active curated pools plus your own, with member counts'). The listing verb is only implied via the name/title, and there is no sibling that needs distinguishing since none of the siblings deal with pools, so a 4 is appropriate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'pools a task can run on' implies the context (inspect available pools before assigning a task), but the description gives no explicit when-to-use or when-not-to-use guidance and names no alternatives. Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_tasksList AI tasksARead-onlyInspect
Your AI tasks, newest first, with status and item counts.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds genuinely useful behavior for a read tool with no output schema: the sort order and the fact that each entry carries status and item counts. It says nothing about volume, pagination, or whether completed/closed tasks are included.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the resource and then stacks ordering and payload information with zero filler. Nothing could be cut without losing signal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-argument, read-only list operation with no output schema, the description covers the essentials: what is returned, in what order, and with what fields. Only the absence of any scope/limit information (pagination, whether closed tasks appear) holds it back.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing to disambiguate and the schema is trivially complete; the baseline for a parameterless tool applies. No param details are needed or expected.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (list) plus resource (AI tasks) and adds the ordering ('newest first') and the fields returned, so an agent knows exactly what this call produces. It does not explicitly differentiate itself from the plural/singular sibling get_task, which keeps it from a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus get_task, get_results, or list_pools, and no exclusions or prerequisites are stated. The scope word 'Your' faintly implies per-user scope, but that is inference rather than instruction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
open_dashboardOpen the business dashboard, signed inAInspect
One-time link that opens the Pathwize business dashboard signed in as this business (optionally one task). Single use, one hour; give it only to the user.
| Name | Required | Description | Default |
|---|---|---|---|
| taskId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does substantial work: single-use semantics, a one-hour expiry, the signed-in identity context, and a security caveat about distribution. It omits what happens on reuse (error vs. silent failure) and any permission requirements, but this is unusually rich disclosure for a link-generation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core deliverable ("One-time link that opens..."), followed by tight fragments covering scope, lifetime, and distribution. No filler sentences; the telegraphic "Single use, one hour" is efficient rather than unclear.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, no-output-schema tool, the description covers what is produced (a link), how long it lives, that it is single-use, and who may receive it. Only minor gaps remain: whether a new link can be regenerated and what the agent receives back as a value.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% for the single taskId parameter, so the description must compensate, and it does: "optionally one task" conveys both the optionality (matching required=0) and the semantic effect of scoping the link to one task. It stops short of describing the expected value beyond the schema's uuid format, which the schema already supplies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb and resource: it produces a one-time signed-in link to the Pathwize business dashboard, optionally scoped to a task. This is clearly distinguishable from the task-CRUD siblings (create_task, get_task, close_task, etc.), though the description never explicitly contrasts itself with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Give it only to the user" is a genuine sharing constraint and "optionally one task" hints at the taskId use case, so usage is implied. There is no explicit when-to-use/when-not guidance or named alternative, leaving the agent to infer the context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
publish_taskPublish an AI taskAInspect
Open a task for experts. This emails every expert in the task's pools once and cannot be undone, so call it only after the user explicitly confirmed, with confirm: true. Fails with a list of anything missing.
| Name | Required | Description | Default |
|---|---|---|---|
| taskId | Yes | ||
| confirm | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so: it discloses the side effect (emails every expert in the task's pools), the count/scope (once), irreversibility ('cannot be undone'), the confirmation gate, and the failure behavior (returns a list of anything missing). This is exactly the behavioral context an agent needs before a destructive-ish broadcast.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with what the tool does and immediately followed by the irreversibility warning and the required precondition. No filler; every clause adds decision-relevant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations and no output schema, the description covers action, side effects, irreversibility, precondition, and error shape. It could say more about the resulting task state (e.g. what status it moves to, whether it can be re-published), but nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains confirm well, including why it exists (explicit user confirmation) beyond the schema's const: true. taskId is left unexplained, though its name and uuid format make its meaning self-evident.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Open a task for experts'), making the state-changing action clear. It does not explicitly contrast with close_task or create_task among siblings, so an agent must infer the lifecycle position, but the action itself is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a hard precondition: call only after the user explicitly confirmed, with confirm: true. That is strong usage guidance. It stops short of naming alternatives (e.g. create_task for drafts) or stating when not to call, so it is clear context without full routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_taskUpdate an AI taskBIdempotentInspect
Change a task's settings; only the fields you pass change, examples are added. Does not publish.
| Name | Required | Description | Default |
|---|---|---|---|
| rate | No | Hourly pay for the expert in EUR | |
| title | No | ||
| domain | No | ||
| taskId | Yes | ||
| classes | No | Class names experts choose from | |
| deadline | No | ||
| examples | No | Example images by public URL with the class they show. Labeling tasks need one per class before they can open. | |
| language | No | Language of the instructions | |
| pool_ids | No | Pools the task runs on (see list_pools) | |
| expertise | No | What experts need to know, one line | |
| allow_note | No | ||
| annotation | No | Default classification | |
| instructions | No | Instructions for the experts in Markdown. A table whose first column names the classes becomes the per-class briefing. | |
| members_only | No | Only members of these pools see the task, even on public pools | |
| multi_select | No | ||
| qualification | No | ||
| allow_cant_tell | No | Let experts skip an item without a class (default true) | |
| seconds_per_item | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The only annotation is idempotentHint=true, so the description carries most of the burden. It discloses two genuinely non-obvious behaviors: PATCH-style partial mutation and that 'examples are added' rather than replaced (a surprising additive semantic the schema does not state). It omits permissions, effect on already-published tasks, and whether fields can be cleared, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact clauses, front-loaded with the core operation and followed by the two most decision-relevant nuances. No wasted words, though the terseness edges toward under-specification for an 18-parameter mutation tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 18 inputs, two enums, a nested qualification object, no output schema, and only an idempotentHint annotation, this description is far too thin. It never addresses whether the task must be unpublished, what happens to a live task, or validation/limit constraints, so the agent has gaps for most non-trivial invocations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 18 parameters and only 56% schema description coverage, roughly half the inputs are undocumented in structured fields, and the description names none of them. Its only parameter-level contribution is the additive behavior of 'examples'; the many bare fields (title, domain, deadline, flags, nested qualification) get no added meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Change a task's settings') and pins down the key behavioral distinction from create_task/publish_task with 'Does not publish.' It doesn't explicitly name the sibling tools, so it stops short of a 5, but an agent can tell what this does and what it does not.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The partial-update rule ('only the fields you pass change') hints at how to use it, and 'Does not publish' implies publish_task is a separate follow-up step. But there is no explicit when-to-use/when-not statement, no mention of required preconditions (e.g. task state, ownership), and no direct reference to alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
10 tool updates
- First observed
add_items - First observed
close_task - First observed
create_task - First observed
get_results - First observed
get_task - First observed
list_pools - First observed
list_tasks - First observed
open_dashboard - First observed
publish_task - First observed
update_task
Related MCP Connectors
Delegate tasks to vetted human experts - research, writing, analysis, and data work.
Create surveys, feedback and classification campaigns. Collect human answers and export results.
API for AI agents to delegate tasks to real humans.
Connect your AI to human workers. Get paid to help AI.
Related MCP Servers
- AlicenseAqualityDmaintenanceReal human judgment as agent tools -- an AI agent can ask a question and get back a structured, schema-validated JSON answer from a real quality-scored human. 16 response types (yes/no, ratings, rankings, A/B tests, sentiment, image/video/audio review, voice/video/photo capture). Fully programmatic signup with a $5 free trial credit, no card required.762 npmMIT

humanforaiofficial
AlicenseAqualityBmaintenanceEnables AI agents to hire real human operators for tasks requiring physical presence, human perception, or judgment, such as verification, testing, data collection, and physical-world tasks.458 npm1MIT- AlicenseAqualityBmaintenanceEnables AI agents to recruit real humans for evaluation tasks like surveys, A/B tests, and ratings on text, images, audio, and video, returning aggregated results directly into the conversation.137MIT
- FlicenseNot gradedqualityDmaintenanceLets AI agents natively discover and hire human experts for tasks they can't do themselves, such as research, verification, and expert calls.-
Glama MCP Gateway
Add one secure layer between your agents and this server.