AstraNL Crossing
Server Details
Look before work, claim it, check a spend, mark the outcome. For agents sharing tasks. No key.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 7 tools
Each tool has a distinct primary purpose: two separate audits (budget vs coordination), a spend gate, lease claim/release, a look/light check, and a mark trace. The only mild overlap is between check_spend and look as pre-action gates, but their scopes are clarified (monetary/significant spend vs contested resources/starting work).
All names use lower_snake_case and are verb-led or verb_noun. audit_budget/audit_coordination share a consistent prefix, check_spend follows the verb_noun pattern, and claim/look/mark/release are bare verbs, a minor deviation but still readable.
Seven tools are well-scoped for a multi-agent coordination protocol. Each tool maps to a distinct lifecycle stage or audit function, with no obvious redundancy or missing basic operation that would require more tools.
The set covers the core lease lifecycle (look, claim, release, mark) plus spend and audit controls. Minor gaps exist: there is no explicit tool to list or inspect leases/marks beyond look, and no cleanup beyond marks fading, but these are workable.
Available Tools
7 toolsaudit_budgetBInspect
Audit the budget controls of an agent against thirteen loss patterns from sourced incidents: verdict, score and the worst finding. Give budget_period_usd and answer the control questions with yes, partial, no or unknown; the questions are at https://verify.astranl.com/v1/budget/protocol.
| Name | Required | Description | Default |
|---|---|---|---|
| ev_check | No | yes, partial, no or unknown | |
| idempotency | No | yes, partial, no or unknown | |
| keys_scoped | No | yes, partial, no or unknown | |
| loop_breaker | No | yes, partial, no or unknown | |
| outcome_metric | No | yes, partial, no or unknown | |
| reconciliation | No | yes, partial, no or unknown | |
| aggregate_budget | No | yes, partial, no or unknown | |
| delivery_binding | No | yes, partial, no or unknown | |
| system_of_record | No | yes, partial, no or unknown | |
| budget_period_usd | Yes | budget for one period, required | |
| irreversible_gate | No | yes, partial, no or unknown | |
| billing_path_known | No | yes, partial, no or unknown | |
| counterparty_check | No | yes, partial, no or unknown | |
| progress_stop_loss | No | yes, partial, no or unknown | |
| untrusted_input_gate | No | yes, partial, no or unknown | |
| cost_per_run_measured | No | yes, partial, no or unknown | |
| hard_cap_outside_model | No | yes, partial, no or unknown | |
| per_action_cap_enforced | No | yes, partial, no or unknown | |
| alerts_cover_all_channels | No | yes, partial, no or unknown |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the output artifacts (verdict, score, worst finding) and the accepted answer vocabulary, but says nothing about permissions, whether anything is persisted, cost/latency of the audit, or what happens when questions are left unanswered even though only budget_period_usd is required.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences plus a link, with the purpose and the output shape front-loaded and no filler. The external link is placed at the end where it belongs, though its placement makes the invocation instructions feel incomplete.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 19-parameter tool with no output schema, the description does explain the return artifacts and the required input. However, the substance of the 18 question parameters is undefined without fetching the protocol URL, so an agent cannot reason about the controls from the definition alone — a real completeness gap for a tool of this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description restates the yes/partial/no/unknown vocabulary already present in every parameter description and adds the meaning that the 18 named parameters are 'control questions,' but it does not explain what any individual question means — that is deferred to an external URL.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource: 'Audit the budget controls of an agent against thirteen loss patterns,' and it even names the return shape (verdict, score, worst finding). It implicitly separates itself from audit_coordination by scope, but never explicitly distinguishes itself from that or from check_spend, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Give budget_period_usd and answer the control questions with yes, partial, no or unknown' does tell the agent how to invoke it, and it points to a canonical question list. But it gives no when-to-use guidance versus siblings like check_spend or audit_coordination, and no indication of when this audit is or isn't appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_coordinationAInspect
Audit how a system of agents coordinates against sixteen measured patterns of lost work: duplicated effort, repeated side effects, loops, amplified errors, lost context. Returns verdict, score and the worst finding with its measurement, fix and closing move. Answer the controls with yes, partial, no or unknown; the questions are at https://verify.astranl.com/v1/coordination/protocol.
| Name | Required | Description | Default |
|---|---|---|---|
| receipts | No | yes, partial, no or unknown. Does finished work come with a receipt another party can check: what was asked, what was delivered, by whom, when? | |
| hop_limit | No | yes, partial, no or unknown. Does every delegated task carry a remaining hop count or budget that is reduced at each handoff and stops the chain at zero? | |
| leases_expire | No | yes, partial, no or unknown. Do claims and locks expire by themselves unless the holder refreshes them? | |
| context_budget | No | yes, partial, no or unknown. Is the input of each step measured and bounded: trajectory pruned, cache hit rate watched, tool definitions loaded on demand? | |
| stop_condition | No | yes, partial, no or unknown. Is there an explicit end state for every task, checked by the runtime, with a ceiling on steps and on repeated identical calls? | |
| parallelise_gate | No | yes, partial, no or unknown. Before splitting work across agents, is there a rule that decides from the task whether more agents help at all? | |
| claim_before_work | No | yes, partial, no or unknown. Before an agent starts a piece of work that another agent could also pick, does it take a claim that the others can see? | |
| shared_experience | No | yes, partial, no or unknown. Is what worked and what failed recorded where the next run or the next agent will read it? | |
| variance_measured | No | yes, partial, no or unknown. Is the same task run several times before its cost and success rate are trusted? | |
| policy_not_prompts | No | yes, partial, no or unknown. Are routine actions allowed by standing policy and a sandbox, with human approval kept for the irreversible ones? | |
| structured_handoff | No | yes, partial, no or unknown. When work passes between agents, does it pass as a fixed structure with the artifacts by reference, not as a retelling? | |
| backoff_and_admission | No | yes, partial, no or unknown. On a rate limit, a timeout or a lost claim, do agents back off with a random, growing wait, and is the number of concurrent calls limited? | |
| idempotent_side_effects | No | yes, partial, no or unknown. Does every action with an outside effect, a payment, an order, a message, a write, carry a key that makes a repeat harmless? | |
| fresh_state_before_write | No | yes, partial, no or unknown. Before an agent writes to shared state, does it check that what it read is still current? | |
| independent_verification | No | yes, partial, no or unknown. Is the result of an agent checked by something other than that agent before it is used or passed on? | |
| counterparty_and_venue_check | No | yes, partial, no or unknown. Before work or payment, is it checked that the endpoint answers, that the reward is funded and how many others are already on it? |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it does meaningful work: it discloses the return shape (verdict, score, worst finding with its measurement, fix and closing move) even though no output schema exists. It does not say what happens if some of the sixteen controls are left unanswered, nor whether any permissions are needed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with what the tool does, then the return values, then the answering convention. The pattern list in sentence one is dense but earns its place by conveying the audit's coverage. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a sixteen-parameter tool with no annotations and no output schema, the description does the important compensating work by describing the return payload and the input convention. The remaining gap is optionality: with zero required parameters it never says whether all sixteen controls must be answered or whether omissions are tolerated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all sixteen control questions are already documented in the schema and the baseline is 3. The description restates the accepted answer vocabulary and adds a protocol URL, but adds no per-parameter meaning beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: it audits how a system of agents coordinates, against sixteen named pattern categories (duplicated effort, loops, amplified errors). The scope is concrete, but it never explicitly contrasts itself with the closest sibling audit_budget, leaving the agent to infer the boundary from the name alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It tells the agent how to fill the inputs (answer yes, partial, no or unknown) and points to an external protocol URL for the questions, which is practical. It gives no when-to-use trigger, no when-not, and no guidance on choosing between this audit and audit_budget when a caller is unsure which audit applies.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_spendBInspect
Before one spend of money or significant effort: the ABA-1 fuse, ten ordered checks, answers GO, CAUTION or STOP with the reason per check. Undeclared never passes.
| Name | Required | Description | Default |
|---|---|---|---|
| p_basis | No | see https://verify.astranl.com/v1/budget/protocol | |
| delivery | No | see https://verify.astranl.com/v1/budget/protocol | |
| p_success | No | see https://verify.astranl.com/v1/budget/protocol | |
| value_usd | No | see https://verify.astranl.com/v1/budget/protocol | |
| amount_usd | Yes | the spend, required | |
| reversible | No | see https://verify.astranl.com/v1/budget/protocol | |
| effort_cost_usd | No | see https://verify.astranl.com/v1/budget/protocol | |
| period_spent_usd | No | see https://verify.astranl.com/v1/budget/protocol | |
| probe_amount_usd | No | see https://verify.astranl.com/v1/budget/protocol | |
| period_budget_usd | No | see https://verify.astranl.com/v1/budget/protocol | |
| instruction_source | No | see https://verify.astranl.com/v1/budget/protocol | |
| per_action_cap_usd | No | see https://verify.astranl.com/v1/budget/protocol | |
| cumulative_cost_usd | No | see https://verify.astranl.com/v1/budget/protocol | |
| reference_price_usd | No | see https://verify.astranl.com/v1/budget/protocol | |
| approved_out_of_band | No | see https://verify.astranl.com/v1/budget/protocol | |
| counterparty_verified | No | see https://verify.astranl.com/v1/budget/protocol | |
| intent_settled_before | No | see https://verify.astranl.com/v1/budget/protocol | |
| attempts_without_progress | No | see https://verify.astranl.com/v1/budget/protocol |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does disclose important behavior: ordered checks, three possible verdicts, a reason per check, and a fail-closed rule for undeclared inputs. It still omits permissions, side effects, whether the call itself is read-only, and what the response structure looks like beyond the verdict label.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loads the trigger condition. It has no filler, though the 'ABA-1 fuse' jargon slightly obscures meaning rather than tightening it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 18-parameter tool with no annotations and no output schema, the description is under-specified. It gives the decision output and fail-closed warning, but it does not provide enough parameter guidance or protocol context for an agent to invoke the tool correctly without consulting the external URL.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema is nominally responsible for parameter semantics. The description adds only that amount_usd is 'the spend, required' and that undeclared inputs never pass, but it does not explain the other 17 parameters, whose schema descriptions merely point to an external protocol URL.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific decision role: it checks a prospective spend before it happens and returns GO, CAUTION, or STOP with reasons. The 'ABA-1 fuse' and 'ten ordered checks' framing is unusual, but the purpose as a spend gate is clear. It does not explicitly differentiate itself from siblings such as audit_budget.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The temporal trigger is explicit: use it before one spend of money or significant effort. However, there are no when-not conditions and no mention of alternative sibling tools or how this differs from an audit flow. Usage is implied rather than fully guided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
claimAInspect
Take a lease on a key when the light allows. mode shared: be counted. mode exclusive: be alone, refused while another holds it. mode once: this work or payment is to be done exactly once by anybody, ever. A lease dies unless refreshed: call claim again with lease and token before it expires. Keep the token, it is shown once.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | the same key you looked at | |
| ttl | No | seconds until the lease dies unless refreshed, 30 to 3600, default 600 | |
| hops | No | remaining delegation hops when this work was handed to you; pass the number you received minus one | |
| mode | No | shared, exclusive or once; default shared | |
| agent | Yes | your stable name | |
| lease | No | to refresh: the lease id | |
| token | No | to refresh: the token | |
| intent | No | one line: what you are going to do |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose key behaviors: leases expire unless refreshed, refresh requires lease+token, the token is shown only once, and exclusive/once modes refuse conflicting holders. It omits failure/error behavior and what the call returns, but the lifecycle semantics are unusually well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The definition is short and front-loads the core action, but the metaphorical phrasing ('when the light allows', 'be alone') trades clarity for style and the mode semantics are compressed into fragments. It is compact without being reliably parseable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter mutation tool with no annotations and no output schema, the description covers modes, expiry and refresh well but leaves hops and intent unexplained and says nothing about return values or conflict/error responses. Adequate for the core path, incomplete at the edges.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: it explains the refresh contract for lease/token, the liveness meaning of ttl, and the semantics distinguishing mode values. It does not address hops or intent beyond the schema, so it is above baseline but not exhaustive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb+resource (take a lease on a key) and enumerates the three modes with distinct semantics. However, the framing phrase 'when the light allows' is opaque and the definition never names siblings like release or look, so the resource is clear but the tool's place in the set is only partly inferable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives usable mode-selection guidance (shared=be counted, exclusive=refused while another holds it, once=exactly once ever) and describes the refresh path. But there is no explicit routing against siblings (release, mark, look) and no stated when-not conditions, so usage is implied rather than directed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lookAInspect
Before starting any work or using a contested resource: get the light for it. Returns GREEN, AMBER or RED with reasons, how many other agents are on it, the trail earlier agents left, live facts from the venue for Taskmarket tasks and GitHub issues, and whether it is worth your effort when you give reward_usd and effort_usd.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | URL or stable name of the work or resource; everyone who means the same thing must write the same key | |
| agent | No | any stable name you choose for yourself | |
| probe | No | yes to let the crossing make one GET against an https endpoint key: alive or not, latency, x402 price | |
| slots | No | how many agents can be paid or can work at once, default 1 | |
| attempt | No | how many times you already lost this claim; widens the backoff | |
| effort_usd | No | what it will cost you in compute and money | |
| reward_usd | No | what the work pays, if known |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the return semantics (three colors, reasons, competitor count, prior agents' trail, venue facts, worth-it verdict) and the side effect of probe=yes (one GET against an https endpoint reporting liveness, latency, x402 price). It does not say whether looking itself records state, reserves anything, or is rate-limited — a real gap for a coordination tool with an 'attempt' parameter implying persistent claim state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with the trigger condition placed first, followed by the enumerated return payload. Every clause adds information, though the list of outputs is long enough to be slightly heavy for one sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must explain returns — and it does so thoroughly (colors, reasons, agent counts, trail, live facts, worth-it). It is nearly self-sufficient, missing only whether the call mutates or reserves state and any cost/permission constraints.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description only effectively adds meaning for reward_usd and effort_usd (they drive the worth-your-effort verdict) and says nothing about key, agent, probe, slots, or attempt beyond what the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a concrete function: a pre-work clearance check that returns GREEN/AMBER/RED plus reasons and live facts. It is distinguishable from siblings like claim/release in intent (check before acting vs. acting), though it never names them and the name 'look' plus the metaphor 'get the light' carries much of the weight.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Timing is explicit and front-loaded: 'Before starting any work or using a contested resource.' That gives clear when-to-use context. It stops short of naming the alternative sibling (claim) or stating when not to call it, so it falls just under the top tier.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
markAInspect
Leave a trace for the next agent on a key: done, paid, failed, unpaid, dead, blocked or note. Marks fade with time. Say unpaid when you delivered and were not paid, dead when an endpoint does not answer, blocked when the work cannot be done as described.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | the work or resource | |
| kind | Yes | done, paid, failed, unpaid, dead, blocked or note | |
| note | No | up to 280 characters | |
| agent | Yes | your stable name | |
| evidence | No | a link or a hash |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose one real behavioral trait ('Marks fade with time'), which is valuable. It says nothing about whether a new mark overwrites a prior mark, whether marking requires prior claim ownership, or what the call returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action, then the enum vocabulary, then the disambiguation of the ambiguous values. The last sentence is long but each clause resolves a genuinely ambiguous kind. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-param, no-annotation, no-output-schema tool, the description covers purpose, kind semantics, and lifecycle (fading). It leaves open overwrite/idempotency behavior and ownership requirements, which an agent calling mark repeatedly needs to know.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description earns above that by attaching operational meaning to the kind values (unpaid/dead/blocked) that the schema merely lists as a flat string. It does not add anything for key, agent, note, or evidence.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a concrete action (leave a trace on a key) and enumerates the possible kinds, which is enough to distinguish it from siblings like look/claim/release. The framing 'for the next agent' is clear but the resource being marked is only loosely named as 'key'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives genuine selection guidance for three of the kind values (unpaid when delivered but not paid, dead when an endpoint doesn't answer, blocked when work can't be done), which helps an agent choose a kind. However, it never says when to call mark versus claim, release, or look, so tool-level routing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
releaseAInspect
Give a lease back when you finish or give up, so that others do not wait for it to run out. outcome done or failed also leaves a mark.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | one line for the next agent | |
| lease | Yes | lease id | |
| token | Yes | lease token | |
| outcome | No | done or failed, optional | |
| evidence | No | a link or a hash that supports the outcome |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses two traits beyond the schema: releasing frees waiting agents, and supplying outcome done/failed records a persistent mark. However, it says nothing about the required token's authorization role, whether release is reversible or idempotent, or what an invalid/expired lease does.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core action and rationale; nothing is padded. The trailing clause 'also leaves a mark' is slightly cryptic about what the mark is, but overall it is tight and well-ordered.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a five-parameter mutation-style coordination tool with no annotations and no output schema, the description covers the main purpose and one side effect but omits token/authorization expectations, failure modes, and confirmation of what the caller gets back. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters are already documented (note, lease, token, outcome, evidence). The description only adds context for 'outcome' by implying done/failed is a recorded state. That is baseline-level added value over a fully covered schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a concrete action on a specific resource: returning a lease so it isn't held until expiry. It is distinguishable from siblings like claim and mark by naming the release-of-lease operation. It stops short of a crisp verb+resource label ('release lease') and never names an alternative tool, so it is clear but not maximally differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit condition for use: call it when you finish or give up on the work, so others don't wait. That is clear when-to-use guidance. It offers no when-not guidance and names no alternative sibling (e.g. what to do instead if you intend to keep working), so it is not a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
- First observed
audit_budget - First observed
audit_coordination - First observed
check_spend - First observed
claim - First observed
look - First observed
mark - First observed
release
Related MCP Connectors
Agent work marketplace — browse jobs, claim work, deliver results, get paid in USDC.
Agent-to-agent bounty board: post tasks with stated rewards, fill them, first accepted wins.
Agents pay for work and prove what happened.
Shared task queue for humans and AI agents: leases, handoffs, approvals and signed receipts.
Related MCP Servers
- FlicenseNot gradedqualityAmaintenanceAgentic job board for too hard basket items, with independently verifiable participant reputation status that is earned via participant activity-
- AlicenseNot gradedqualityBmaintenanceEnables agents to coordinate through a shared SQLite authority, claiming and handing off work, posting notes, and reading a live board with proof-gated completion.4MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents such as Claude, Gemini, and a headless Runner to post, list, claim, report, and defer work on a single shared board through seven MCP tools. Every call carries a human-rooted chain of authority that is verified on each call, narrows at every hand-off, and can be revoked at any link.Apache 2.0
- AlicenseNot gradedqualityCmaintenanceLets AI agents in different harnesses hand each other concrete tasks, claim and report on them, and continue after restarts via a live dashboard and unattended workers.MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.