MangoMe
MangoMe MCP server provides governed persistent operational memory and work-state management for long-lived multi-agent AI work.
Manage session recovery, workspace binding, health, and canonical restore without state creation.
Admit and govern work via intake, specs, projects, families, contracts, artifacts, and typed entity links.
Plan, start, update, claim done, and close slices under persisted plans.
Persist, attest, verify, gate, override, approve, and accept evidence/assurance.
Query canonical state: status, project overview, family views, context, and graph edges.
Compile/dispose bounded execution contexts and UAI/1 compact semantic transport packets.
Publish worker runtimes, check execution eligibility, authorize/complete delegations, and record model economics.
Discover/onboard via scopes, repository locations, Big Bang scans, filesystem scans/references, and reconcile candidates.
Build RB/1 reproduction bindings, check evidence freshness, refresh views, diagnose maintenance, and migrate schemas.
Persists canonical MangoMe work state — contracts, specifications, families, slices, plans, evidence, and approvals — in MongoDB as the production store.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@MangoMewhat's the status of the AVCOS contract family?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
MangoMe
MangoMe is not just a state machine
MangoMe is the canonical operational memory underneath long-running AI work.
It combines responsibilities that are usually scattered across chats, repositories, task trackers, reports and agent memory:
DOCUMENT STORE
+
WORK GRAPH
+
DURABLE WORK IDENTITY
+
STATE MACHINE
+
SPECIFICATION / NORMATIVE BASELINE HISTORY
+
EVIDENCE / PROVENANCE LEDGER
+
EXECUTION ECONOMICS
+
DETERMINISTIC CONTEXT COMPILER
+
COMPACT SEMANTIC TRANSPORTA worker is not a truth source. A worker can execute, observe, propose and claim completion. MangoMe persists what the work is, what is expected, what was observed, what was claimed, and what was independently verified.
Workers are ephemeral executors. MangoMe is the canonical operational record within its governed scope.
The central distinction is:
N = normative truth
What must be true?
O = observed truth
What is actually present?
J = worker judgment
What does the worker conclude from N and O?
V = independent verification
What has been independently demonstrated?MangoMe externalizes state so agents spend model capacity on the unresolved delta instead of repeatedly reconstructing history from prompts, filesystem paths or prior agent prose.
Externalize state. Localize uncertainty. Preserve agency. Verify independently.
Why MangoMe exists
Long-running AI work tends to fail in a predictable way: artifacts survive, but the shared operational understanding does not.
A Claude session ends. Codex continues. Another agent sees a different database. A worker says “done”. A specification has changed. A previous test result exists but is stale. Two agents touch related work. A coordinator reconstructs state from a repository instead of reading the canonical system.
MangoMe moves that problem out of the prompt and into durable, inspectable state.
A normal worker should not need to know or expose MangoMe internals to the user. The worker should use MangoMe to organize its execution, persist progress and recover safely.
v0.3.6 — bitemporal truth + hardened MongoDB trust boundary
v0.3.6 closes the two remaining core gaps identified after the v0.3.x architecture review. BTTM/1 adds non-destructive bitemporal truth maintenance (valid_time vs known_time) with support/assumption/dependency invalidation and PCH revalidation heating. MTB/1 adds an optional fail-closed production trust-boundary profile so canonical MongoDB credentials can live behind a dedicated MangoMe service identity instead of in worker-visible environment/configuration. See docs/bitemporal-truth-maintenance.md and docs/trust-boundary.md.
v0.3.5 — out-of-band control-plane self-maintenance
v0.3.5 closes a real control-plane deadlock found during live Codex operation. An explicit operator request to update/repair MangoMe itself must not require enter_work, a new WorkIdentity, or a self-approval persisted into the same MangoMe control plane being repaired. reconcile_assignment now returns the read-only CPM/1 disposition for explicit MangoMe self-maintenance and short-circuits canonical restore/admission.
CPM/1 is deliberately narrow: it covers only MangoMe source/package/runtime/service maintenance named by the current user. It never authorizes canonical MongoDB/schema mutation. If the requested target requires a database/schema migration without separate explicit authorization, the worker must stop with DATABASE_CHANGE_REQUIRED. Historical host memory and old eval/audit artifacts remain candidate-only hints; current user intent plus current repository/runtime observation outrank them for maintenance state. Casual status remarks are not work orders.
v0.3.4 — non-blocking bootstrap / lazy discovery hotfix
v0.3.4 fixes a real-world zero-touch failure found during a Codex/MCP run: an explicit session_bootstrap on an unbound workspace could fall through to full workspace attachment, which may perform filesystem inventory and candidate Big-Bang discovery. That made a read-only bootstrap capable of blocking for minutes.
Binding and discovery are now separate operations. Managed startup creates only a volatile READ_ONLY_FAST_PATH workspace binding. No filesystem scan, Big-Bang scan, Git archaeology, or canonical mutation occurs merely because MangoMe starts or a worker requests restore/reconciliation. Canonical restore is lazy at the recovery/effect boundary, while full discovery remains explicit through refresh/onboarding operations.
The operational invariants are:
MANGOME_MUST_NOT_BLOCK_COGNITION
BINDING_IS_NOT_DISCOVERY
BOOTSTRAP_PERFORMS_NO_FILESYSTEM_SCAN
BINDING_FAILURE_BLOCKS_EFFECTS_NOT_REASONINGworkspace_status(refresh=True) remains available when a host/operator explicitly wants a full attachment/discovery refresh.
v0.3.3 — zero-touch assignment reconciliation
v0.3.3 removes the worker-facing restore-first ritual from the normal managed workflow. A worker may inspect, search, reason, classify and form a tentative decomposition before MangoMe involvement. MangoMe becomes mandatory at the effect boundary, not at the first thought.
user request
↓
free local understanding / tentative decomposition
↓
reconcile_assignment
↓
existing WorkIdentity / baseline / unfinished state / authority
↓
governed productive effectThe governing rule is THINK FREELY, RECONCILE BEFORE EFFECT. Managed clients bind the workspace through a cheap read-only fast path when MangoMe initializes; canonical restore is resolved lazily when reconciliation/recovery or an effect gate actually needs it. Workers therefore do not call session_restore merely because a session started. session_restore remains the explicit recovery/status primitive, while reconcile_assignment is the normal worker-facing bridge from tentative judgment to governed work.
The effect boundary includes durable code/file/database mutation, commits, deployment, external side effects, canonical MangoMe mutations, normative promotion and assurance claims. Tentative reasoning is never promoted automatically, and existing work is still reused before new work is admitted.
v0.3.3 also adds FJD/1 Fast Judgment Decisions, a deliberately small provider-neutral typed-decision layer for cheap repeated judgments. It accepts strict BOOL, SCORE, and CHOICE outputs with explicit confidence and returns USE_SIGNAL, REVIEW, REVIEW_REQUIRED, or ESCALATE. FJD/1 performs no model inference itself; a host may use a heuristic, small local classifier, specialized decision engine, or stronger model. Every result remains WORKER_JUDGMENT: confidence can guide classification, triage, routing, activation, or prioritization, but can never create truth, Evidence, assurance, verification, acceptance, normative promotion, or mutation authority. See docs/fast-judgment.md.
v0.3.2 — scoped recursive audit and impact closure
v0.3.2 adds SRA/1, a bounded recursive audit layer for the normal engineering reality that a worker is competent inside a local assignment but does not and should not need to understand the entire system before acting.
initial scope
↓
inspect current frontier
↓
persist finding / evidence / assumptions / unknowns
↓
material impact?
no → frontier shrinks
yes → follow bounded confirmed relations
↓
repeat until fixpoint or explicit boundarySRA/1 separates knowledge, inspection, mutation, and assurance boundaries. The inspection frontier may grow when a finding has material impact; the mutation scope is frozen at audit start and never grows automatically. A closed audit therefore means only that the configured impact frontier has reached FIXPOINT_REACHED or BOUNDED_FIXPOINT. It never means the entire system is correct or that the audited object is VERIFIED.
The active audit frontier is supplied to PCH/1 as additional cognitive roots, so the worker receives the currently relevant impact region without loading the entire historical graph. See docs/scoped-recursive-audit.md.
v0.3.1 — persistent cognitive hygiene and thermal working sets
v0.3.1 adds a deterministic homeostasis layer above MangoMe's durable v0.3 work graph and below the existing ContextCompiler. It does not create another memory system. Canonical MongoDB state remains the historical record; PCH/1 computes only which objects should be cognitively resident for the present task.
Persistent canonical history H_t
↓
PCH/1 Cognitive Hygiene
graph reachability + task relevance
+ authority + freshness + support
↓
Bounded active working set A_t
HOT / WARM / COLD
↓
existing ContextCompiler
↓
worker / UAI context C_tThe central invariants are:
temperature != truth
temperature != assurance
COLD != deleted
|C_t| <= B_context
|A_t| <= B_active
H_t remains historically recoverableHOT, WARM, and COLD are buckets over a continuous task-relative temperature in [0,1]. Current canonical execution roots are pinned. Superseded or historical state normally cools, but a targeted historical query can reheat it for inspection/revalidation without restoring normative authority. The GC analogy is deliberately non-destructive: PCH/1 borrows reachability, generations and residency pressure, but "collection" only removes objects from the active worker set.
The full diagnostic surface is available through cognitive_hygiene; normal worker execution receives only the bounded selection summary through ContextCompiler and UAI/1.
v0.3.0 — durable WorkIdentity, progressive persistence, and playbook isolation
v0.3.0 changes MangoMe's primary architectural anchor. Contracts and Specifications remain important normative inputs, but they no longer define the identity of the work. Playbooks remain procedural inputs and are explicitly non-normative.
Playbooks may change.
Specs may evolve.
Workers and sessions may disappear.
The identity of the work and its assurance history must survive them all.
The governing model is:
Intent-first
↓
Identity-bound
↓
Playbook-guided
↓
Spec-governed
↓
Assurance-preservingMangoMe now separates three persistence levels:
VOLATILE
selected Playbook body / transient execution context
PROGRESSIVE
crash-recovery checkpoint / working cursor / non-normative selection state
CANONICAL
WorkIdentity / NormativeBaseline binding / Plan authority / Evidence / Claims / Assurance historyA prompt may admit a durable WorkIdentity, but it does not automatically become a Specification or Contract. Productive Plans are bound to the current immutable NormativeBaseline; if the effective Spec/Contract truth changes, an older Plan stops with BASELINE_DRIFT rather than silently inheriting the new target.
Playbooks are deliberately weaker than canonical state:
Playbook = HOW work is normally performed
Specification = WHAT must become true
MangoMe = WHAT survives: identity, persistence, provenance, execution binding, and independent assuranceThe hard recovery rule is therefore:
filesystem / Git / Playbook / prior agent prose
↓
may guide or corroborate
↓
MUST NOT reconstruct WorkIdentity, current normative truth, or assuranceFor managed MCP operation, first admission after STATE_NOT_FOUND may trust the current client-relayed user intent. Once canonical workspace state exists, minting additional durable work authority or a new WorkTurn belongs to the host/control-plane capability boundary. Recovered workers consume a bound turn; they do not self-authorize current-user intent. Direct library use remains a lower-level integration boundary and must be isolated from untrusted workers in high-assurance deployments.
The ContextCompiler and UAI projection now expose the persistence boundary explicitly so canonical identity/assurance survives context reduction while volatile Playbook material is trimmed first.
v0.2.2 — execution integrity hardening
v0.2.2 is a hardening release built on GitHub main commit:
10c7859826a0fa379e99ffbcada0321eb1c36560That commit was v0.1.9rc4.
The release addresses an observed failure mode: an agent received an audit assignment, but instead of performing the audit it generated a large new audit/planning artifact and presented that planning output as the result.
MangoMe must allow agents to organize work internally without allowing planning to replace execution.
The v0.2.2 execution rule is therefore:
USER ASSIGNMENT
↓
RESTORE CANONICAL STATE
↓
RELEVANT SLICES EXIST?
┌───────────────┴───────────────┐
YES NO
↓ ↓
REUSE EXISTING MATERIALIZE MINIMUM
SLICES INTERNAL SLICE(S)
└───────────────┬───────────────┘
↓
EXECUTE WORK
↓
OBSERVE REALITY
↓
PERSIST EVIDENCE
↓
RECONCILE EXPECTED/ACTUAL
↓
REPAIR / ADAPT IF NEEDED
↓
RE-OBSERVE
↓
CLAIM COMPLETION
↓
INDEPENDENT VERIFICATIONThe user should receive the result of the assignment, not MangoMe's internal decomposition.
Related MCP server: Geond Agent Protocol
Contract-first turn binding and immutable generations
MangoMe treats the current user turn and the recovered work state as different things. Restored ACTIVE work is durable context, not an instruction to continue it. For admitted contract work, an agent first binds the current turn to the intended logical Contract with one of six modes: QUERY, CONTINUE, EXECUTE, VERIFY, MODIFY, or CONTROL.
The recovery priority is therefore:
CURRENT USER TURN
↓
EFFECTIVE CONTRACT / SPECIFICATION
↓
OBSERVED ARTIFACTS / EVIDENCE
↓
UNRESOLVED DELTA
↓
DERIVED PLAN / SLICE STATEPlans and Slices remain durable and useful, but they are execution decomposition rather than normative truth. An active Slice never authorizes the current turn by itself.
Contract bodies can exist in more than one physical location. MangoMe keeps the logical Contract identity separate from those locations and can persist immutable canonical generations of the normative body in MongoDB. A local or Git/DMS copy may change, but that changed file is only an observation until an explicitly bound MODIFY turn promotes it.
Promotion is single-writer and compare-and-swap protected. A generation grant is bound to the Contract, actor, current turn, base generation, and base content hash. Only that grant may create the next immutable generation; concurrent or stale promotion fails closed instead of merging or overwriting silently.
This gives the intended asymmetry:
logical Contract durable identity
canonical generations immutable normative history
physical files mutable working representations / bindings
Plans and Slices derived execution stateThe key invariants are:
MangoMe context never substitutes for current user intent.
A changed contract file is not canonical truth until the authorized MODIFY turn promotes a new generation.
Slices are internal execution state
Slices are durable execution addresses, not normal user-facing deliverables.
For an admitted assignment:
if suitable Slices already exist, MangoMe reuses them;
it does not create replacement Slices merely because a new session or agent started;
if an admitted Family genuinely has no Slice, MangoMe may materialize the minimum internal Slice needed to execute the assignment;
internal Slice/Plan details are not normal human-facing output;
a user does not need to ask for, manage, name or approve ordinary internal Slices.
The new prepare_assignment surface exists for this execution path.
prepare_assignment
↓
reuse existing Slice(s)
OR
materialize minimum internal Slice
↓
create/bind Plan
↓
start executable target(s)It returns internal execution bindings for the client/worker, while the presentation policy explicitly marks Slice details as internal.
Planning is not execution
For audit, review and verification work, a checklist, plan, specification or newly authored contract is not a substitute for executing the requested work.
An audit-like assignment cannot reach DONE_CLAIMED merely because the worker produced planning prose. The Slice must have persisted non-claim Evidence demonstrating that observation actually occurred.
This closes the failure pattern:
AUDIT REQUEST
↓
WRITE AN AUDIT PLAN
↓
CLAIM SUCCESSThe valid path is:
AUDIT REQUEST
↓
OBSERVE
↓
EVIDENCE
↓
FINDINGS / DELTA
↓
OPTIONAL REPAIR
↓
RE-OBSERVE
↓
DONE_CLAIMEDObserve before repair
Audit/reconciliation work first establishes the delta between normative and observed state. Repair, adaptation or correction is a subsequent action where the assignment permits mutation.
A read-only audit remains read-only.
A repair-capable assignment may continue from findings into correction, but it must not silently rewrite normative Contract/Specification truth merely to make implementation appear compliant.
Zero-touch operation
Zero-touch is a user-interface property, not a weaker truth model.
user gives a normal task
↓
worker understands / inspects / forms tentative decomposition
↓
reconcile_assignment
↓
STATE_FOUND | STATE_PARTIAL | STATE_NOT_FOUND
↓
STATE_FOUND → reconcile with existing WorkIdentity/baseline/work
STATE_PARTIAL → bounded validation/backfill only
STATE_NOT_FOUND → never synthesize recovery state
↓
new work? → enter_work
admitted work? → bind current turn + prepare_assignment / begin_work
↓
Plan-bound internal Slice execution
↓
Evidence / claims / verificationThe user should not have to say:
“create a Slice”;
“start Big Bang”;
“call
begin_work”;“create a Plan”;
or otherwise operate MangoMe vocabulary manually.
The internal machinery remains strict even though the user-facing workflow is simple.
Restore is not creation
Recovery is identity-driven.
STATE_FOUND
STATE_PARTIAL
STATE_NOT_FOUNDSTATE_NOT_FOUND is a real state. It does not mean “create replacement state and call it restored”.
For admitted work:
Recovery follows identity. Discovery must never create or reconstruct admitted identity/state.
Filesystem paths, Git/worktree history, contract folders, evidence folders and previous-agent prose may validate a bounded unresolved delta. They are not substitutes for canonical MangoMe state.
GOAL_ACTIVE or an unfinished Family does not itself grant execution permission. Productive work follows canonical executable items and active plan bindings.
MangoMe is infrastructure
Agents may use MangoMe but must not modify MangoMe source, tests, packaging or managed configuration unless the explicit assignment targets MangoMe itself.
A blocked application task, failed verification, missing state or runtime defect is not implicit permission to self-edit the governance substrate.
Hard filesystem/process enforcement remains a host/runtime responsibility.
This distinction is important:
application task fails
≠
permission to modify MangoMeCanonical persistence and runtime binding
MongoDB remains MangoMe's production canonical store.
Production canonical persistence = MongoDBv0.2.2 strengthens runtime binding because a correct state model is useless if different agents silently connect to different databases.
Managed client configuration now carries the expected database identity:
MANGOME_DATABASE
MANGOME_EXPECTED_DATABASEIf the configured runtime database does not match the expected managed database, initialization fails closed with:
WRONG_MANGOME_DATABASEThis is intentionally separate from source/version binding:
MANGOME_EXPECTED_VERSION
MANGOME_EXPECTED_SOURCE_ROOT
MANGOME_EXPECTED_DATABASEThe practical invariant is:
GitHub/source identity
+
runtime package identity
+
managed MCP target
+
MongoDB database identity
=
one intended MangoMe runtimeDirect worker access to MongoDB remains outside MangoMe's service trust boundary. Production deployments should isolate canonical database write credentials from untrusted workers.
Schema evolution fails closed on the future
MangoMe persists a schema_version on canonical documents.
Older known schema versions can be upgraded deterministically by registered migrations.
v0.2.2 adds the opposite boundary as well: an older reader must not silently interpret a future schema it does not understand.
persisted schema <= reader schema
→ normal read / migration path
persisted schema > reader schema
→ UnsupportedSchemaVersionThis protects mixed-version agent environments from quietly treating newer state as if it were older compatible state.
Hard context envelope
Canonical truth is never truncated merely to fit an agent prompt.
MangoMe instead compiles a disposable execution projection.
v0.2.2 adds a deterministic hard byte envelope to ContextCompiler:
ContextCompiler(service).compile(
family_id,
slice_id,
max_bytes=...
)The compiler follows a deterministic reduction policy when the projection exceeds the envelope:
preserve normative and active execution state;
compact optional Evidence payload detail;
compact optional Contract storage detail;
omit older optional Evidence entries where required;
fail closed if mandatory state itself cannot fit.
The failure is explicit:
ContextBudgetExceededThe rule is:
Canonical truth is never compressed destructively. Only the execution projection is bounded.
MangoMe uses a byte envelope rather than pretending it can guarantee provider-specific token counts without the exact tokenizer. Provider-reported token counts continue to belong in Execution Receipts.
Work Graph query hardening
MangoMe does not introduce an in-memory graph truth store in v0.2.2.
Before adding a Change-Stream-driven DAG cache, the repository first removes avoidable broad reads.
Graph lookup now uses indexed edge predicates for incoming and outgoing relationships rather than reading the whole Edge collection and filtering in Python.
MongoDB indexes include the relevant directional access paths, including reverse to_id lookup and Evidence subject access.
The current decision is deliberate:
first: query/index correctness
then: measure
only then: consider an in-memory topology cacheA second live graph representation would add cache invalidation and runtime-consistency failure modes. That is not justified until profiling shows indexed MongoDB traversal to be a real bottleneck.
Core entities
MangoMe gives first-class identity to durable operational objects including:
Request
Project
Contract Family
Contract Contribution
Specification
Plan
Slice
Claim
Evidence
Artifact
Graph Edge
Approval
Model Profile
Execution Receipt
Materialized Status ViewEvery first-class entity has an immutable internal entity_id.
Human identifiers such as declared contract or Slice IDs may collide without silently overwriting history; ambiguity is surfaced rather than hidden.
Core invariants
1. Work identity is more durable than a session
A chat, prompt, branch or model invocation is not the identity of the work.
2. A prompt is not automatically Contract Truth
User intent may admit operational work. Durable normative contract evolution remains explicit and append-only.
3. Slices are durable internal execution addresses
Execution state and assurance state remain separate dimensions.
Execution examples:
PLANNED
STARTED
ACTIVE
PAUSED
BLOCKED
DONE_CLAIMED
CANCELLEDAssurance examples:
UNVERIFIED
PARTIAL
VERIFIED
ACCEPTED
REJECTEDA normal state is:
DONE_CLAIMED / UNVERIFIED4. Existing Slices are reused
A new worker/session does not justify duplicate work decomposition.
When an admitted Family already contains relevant Slices, assignment preparation reuses them. Missing internal work structure may be created only where necessary for execution.
5. Plan before mutate
Productive mutations remain bound to a persisted Plan.
submit_plan / prepare_assignment
↓
start_slice(plan_id=P)
↓
update_slice_progress(plan_id=P)
↓
claim_done(plan_id=P)6. DONE is a claim
worker: "done"
↓
DONE_CLAIMED
↓
independent Evidence / gates
↓
VERIFIED
↓
optional authorized acceptance
↓
ACCEPTEDDONE_CLAIMED != VERIFIED remains a core invariant.
7. Evidence is not automatically proof
Worker-authored Evidence can support work, but independent verification remains separately authorized.
Audit/review work in v0.2.2 must at minimum persist real observational Evidence before claiming completion; prose planning alone is insufficient.
8. Normative truth is not rewritten to fit implementation
If observed reality conflicts with the current Specification, that conflict is a finding. A worker may repair implementation where authorized or propose an amendment, but it may not silently redefine the requirement.
9. Parallelism is allowed; canonical writes are guarded
MangoMe does not globally serialize agents. It protects canonical document writes through revision / compare-and-swap semantics and surfaces collisions where useful.
Contract evolution
Contract contributions are append-only.
Supported relations include:
ADDS_TO
AMENDS
EXTENDS
REPAIRS
RECOVERS
SUPERSEDES
CONFLICTS_WITH
VALIDATES
IMPLEMENTS
PART_OF
EXPOSED_BY
RELATES_TOeffective_family_view resolves confirmed supersession and exposes conflict/suggestion state without silently merging ambiguous prose.
AV/1 — adversarial completion verification
MangoMe treats worker completion as a claim to inspect.
DONE_CLAIMED
↓
completion_review
↓
verifier observes / replays / inspects
↓
submit_verification_observation
↓
attested independent Evidence
↓
verify_slice
↓
VERIFIEDThe verifier path is intentionally separate from worker execution.
A worker may not self-upgrade its completion claim into verified truth.
RB/1 — reproducible Evidence binding
Where Evidence should remain reusable, MangoMe can bind it to concrete execution inputs through RB/1.
A reproduction binding can include:
command identity;
caller-supplied exit code;
relevant input hashes;
Git commit context;
output Artifact identity;
stdout/stderr hashes;
environment names without secret values;
a deterministic fingerprint.
Freshness can later become REUSABLE, STALE, UNKNOWN, UNBOUND or INADMISSIBLE, but freshness never creates VERIFIED by itself.
UAI/1 — compact semantic transport
MangoMe can compile canonical state into a compact transport representation for expensive workers.
The key rule remains:
MangoMe does not compress canonical truth. It compiles a disposable execution projection of that truth.
UAI/1 packets include a semantic hash, and UAI/1R worker results bind back to that hash. Decoded results never directly mutate canonical state; actions still flow through normal MangoMe mutation, Evidence and verification paths.
Assignment surfaces
reconcile_assignment
Read-only RAE/1 bridge from tentative worker understanding to governed work. It resolves canonical restore state and returns whether the worker should reconcile with existing work, perform bounded recovery/backfill, or admit genuinely new work. It never canonicalizes the worker's tentative decomposition.
enter_work
Zero-touch admission for genuinely new ordinary work.
prepare_assignment
v0.2.2 execution-preparation surface for already admitted work.
It:
classifies the assignment;
uses the current effective Specification;
reuses existing Slices;
materializes a minimal internal Slice only when none exist;
persists a Plan;
binds executable targets to that Plan;
exposes a presentation policy that keeps Slices internal.
It is specifically intended to prevent agents from replacing execution with a new user-facing planning artifact.
begin_work
Convenience composition for an already admitted Family/Specification when a specific bounded Slice is being started.
MCP tools
The v0.2.2 MCP surface includes the existing MangoMe tools plus the hardened assignment path.
Session / recovery
health
workspace_status
reconcile_assignment
session_restore
session_bootstrap
recovery_context
Intake / specification
intake_request
create_spec
enter_work
prepare_assignment
begin_work
Identity / registry
resolve
create_project
create_family
register_contract
import_contract_bundle
attach_artifact
link_entities
Planning / execution
submit_plan
start_slice
update_slice_progress
claim_done
close_plan
Evidence / assurance
submit_evidence
attest_evidence
completion_review
submit_verification_observation
set_gate
set_gate_controlled
verify_slice
request_override
approve_override
reject_override
list_approvals
accept_slice
Read / truth
status
project_overview
effective_family_view
read_context
compile_execution_context
graph
Compact semantic transport
compile_uai_context
expand_uai_context
decode_uai_result
render_uai_result
Runtime / delegation
publish_worker_runtime
execution_eligibility
authorize_delegation
complete_delegation
delegation_status
Fast judgment / decision gating
assess_fast_judgment
record_fast_judgment
fast_judgment_status
Economics
register_model
record_execution_receipt
model_stats
Discovery / maintenance
discovery_scopes
repository_locations
bigbang_scan
reconcile_bigbang
filesystem_scan
filesystem_references
build_reproduction_binding
evidence_freshness
refresh_views
maintenance_diagnose
migrate_schemaInstallation
Requirements:
Python 3.10+
MongoDB for canonical production persistence
python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
pytestManaged client setup:
mangome setup --client autoClaude Code can use private LOCAL scope by default; project scope remains explicit:
mangome setup --client claude-code --claude-scope projectFor deterministic local tests:
export MANGOME_BACKEND=memory
mangome-mcpFor canonical MongoDB persistence:
export MANGOME_BACKEND=mongo
export MANGOME_MONGODB_URI='mongodb://127.0.0.1:27017'
export MANGOME_DATABASE='mangome'
mangome-mcpManaged clients generated by MangoMe additionally bind the expected database identity.
The default MCP transport is stdio.
Streamable HTTP can be enabled through the existing MANGOME_MCP_TRANSPORT, host and port variables.
Agent Skill
Canonical Skill source:
skill/mangome/SKILL.mdThe project keeps mirrored Skill surfaces for packaging and supported clients. They must remain synchronized.
The v0.3.3 Skill explicitly teaches the normal worker path:
think freely and form only a tentative local decomposition;
reconcile with MangoMe before productive effect;
use
reconcile_assignmentinstead of a ritual restore-first step;resolve/reuse existing WorkIdentity, plans and evidence before creating new work;
keep Slice/Plan mechanics internal;
use
prepare_assignmentfor admitted assignments;planning is never a substitute for execution;
audits must persist actual observation Evidence;
observe the delta before repair;
do not mutate MangoMe unless MangoMe itself is the assignment target.
Typical lifecycle
Ordinary new work:
understand locally
→ reconcile_assignment
→ STATE_NOT_FOUND for genuinely new work
→ enter_work
→ internal Plan/Slice execution
→ Evidence
→ claim_done
→ independent verification where requiredExisting admitted work:
understand locally
→ reconcile_assignment
→ STATE_FOUND
→ bind current turn / prepare_assignment
→ reuse existing Slice(s)
→ execute
→ persist Evidence
→ reconcile
→ repair/adapt if authorized
→ claim_done
→ verifyAudit/review path:
reconcile_assignment
→ prepare_assignment
→ inspect actual system
→ persist observations
→ compare N vs O
→ findings
→ optional authorized repair
→ re-observe
→ DONE_CLAIMEDRepository structure
.
├── src/mangome/
│ ├── service.py # canonical domain operations / assignment preparation
│ ├── integrity.py # assurance and authority invariants
│ ├── hygiene.py # PCH/1 thermal working-set / cognitive homeostasis
│ ├── context.py # bounded execution context / hard envelope
│ ├── interlingua.py # UAI/1 compile/decode/render
│ ├── importer.py # Big-Bang discovery/reconciliation
│ ├── filesystem.py # deterministic filesystem inventory / evidence freshness
│ ├── operability.py # client bootstrap, identity/database binding, auto-attach
│ ├── maintenance.py # deterministic diagnostics/migrations
│ ├── schema.py # schema evolution / future-version fail-closed
│ ├── mcp_server.py # MCP v2 surface
│ └── storage/
│ ├── mongo.py # canonical MongoDB backend / indexes / OCC
│ └── memory.py # deterministic test backend
├── skill/mangome/
├── src/mangome/skill/
├── .github/skills/mangome/
├── .claude/skills/mangome/
├── docs/
├── examples/
├── tests/
├── pyproject.toml
└── server.pyRelease-layout integrity
v0.2.2 also hardens the repository/package boundary itself. The executable Python package lives under src/mangome/; package modules must not be duplicated into the repository root. Version metadata in pyproject.toml, src/mangome/__init__.py, and the MCP server must agree. Regression tests fail if shadow copies such as root-level runtime.py, service.py, mcp_server.py, operability.py, or duplicate root Skill files appear.
MangoMe setup/doctor code is not a host-network manager. The release guard also rejects source changes that add direct management of resolver/network/firewall surfaces such as /etc/resolv.conf, systemd-resolved, Netplan, iptables/nftables, or UFW to the MangoMe runtime. Host-network repair is outside MangoMe's normal execution scope.
Deliberate v0.2.2 non-goals
v0.2.2 does not:
replace MongoDB;
add an in-memory DAG as a second graph truth;
add MongoDB Change Stream cache synchronization;
let workers self-verify;
expose Slices as normal user-facing workflow;
create recovery state when
STATE_NOT_FOUND;automatically reinterpret a prompt as a durable Contract contribution;
let planning artifacts count as execution Evidence;
guarantee provider-specific token counts from byte budgeting;
claim that external model dispatch enforcement exists unless the host actually binds to MangoMe authorization.
Current status
v0.3.2 adds Scoped Recursive Audit/Impact Closure on top of the v0.3.1 Persistent Cognitive Hygiene layer.
The current execution architecture is:
ordinary user intent / recovered WorkIdentity
↓
current WorkTurn + immutable NormativeBaseline
↓
canonical Project / Family / Spec / Plan / Slice / Evidence graph
↓
SRA/1 bounded audit frontier (when auditing)
↓
PCH/1 thermal working-set selection
↓
existing ContextCompiler hard envelope
↓
UAI/1 or normal worker context
↓
execution / Evidence / DONE_CLAIMED
↓
independent verification
↓
VERIFIED / optional ACCEPTEDPCH/1 is deliberately projection-only. It does not delete history, alter assurance, infer correctness, or create a second source of truth. HOT/WARM/COLD express current cognitive residency; canonical state remains governed by MangoMe's existing truth, authority and verification paths.
Important current limits
SRA/1 closes bounded impact scope, not whole-system correctness. A
BOUNDED_FIXPOINTexplicitly means configured depth/object limits prevented further traversal.Audit scope expansion never grants additional mutation authority; external code/filesystem enforcement remains a host/runtime responsibility.
PCH/1 uses an explicit deterministic heuristic policy; its weights and thresholds are an inspectable baseline for evaluation, not a claim of optimal cognitive allocation.
The current graph-distance calculation is scoped to the Family execution context plus deterministic synthetic relations; it is not yet a host-wide graph navigator.
Freshness can consume explicit
CURRENT,STALE,SOURCE_CHANGED,ENVIRONMENT_CHANGED, orREVALIDATION_REQUIREDmarkers, but v0.3.3 does not yet implement a general bitemporal truth-maintenance engine.Temperature controls activation only. It cannot promote evidence, change normative authority, verify a Slice, or accept work.
FJD/1 confidence is advisory worker judgment only. The default 0.80/0.60 thresholds are inspectable operational defaults, not calibrated truth guarantees; high-impact decisions still require review.
MangoMe can only enforce writes that pass through MangoMe. Direct MongoDB/admin access remains outside the service trust boundary.
External model dispatch and verifier command execution remain host responsibilities.
Product boundaries
MangoMe deliberately does not:
treat every prompt as a contract;
infer verification from a worker's DONE statement;
automatically recover/replay a dead model session;
serialize all project work behind a global lock;
use an LLM for ordinary status calculation;
silently canonicalize ambiguous discovery;
infer implementation correctness from filesystem presence alone;
let compact UAI output directly mutate canonical state;
make an external reviewer a source of truth;
promise interchangeable production persistence semantics;
require Scrum, sprints or story points.
License
Apache License 2.0.
See LICENSE.
Available Tools
90 toolsaccept_sliceB
Move VERIFIED to ACCEPTED using a trusted approved ACCEPT_SLICE decision.
| Name | Required | Description | Default |
|---|---|---|---|
| turn_id | No | ||
| slice_id | Yes | ||
| accepted_by | No | ||
| approval_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses the state transition and the precondition of an approved decision, but it omits permissions, reversibility, side effects, and failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler. For a four-parameter mutation tool, however, it is arguably too terse to be fully appropriate, even though it is structurally efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. But with no annotations and 0% schema description coverage across four parameters, the description leaves critical invocation details—especially parameter roles and behavioral constraints—unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across four parameters. The description implies that an approved ACCEPT_SLICE decision relates to approval_id and that the state transition relates to slice_id, but it never names or explains turn_id or accepted_by, so parameter meaning remains mostly undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific operation: moving a slice from VERIFIED to ACCEPTED using an approved ACCEPT_SLICE decision. This is clear for the domain, though it does not explicitly mention the resource (slice) or differentiate from siblings like verify_slice or reject_override.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies when to use the tool by naming the required source state (VERIFIED) and the required decision type (trusted approved ACCEPT_SLICE). However, it offers no explicit when-not guidance or alternatives to consider, leaving routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
approve_overrideC
Approve using the owner runtime capability; actor strings alone are insufficient.
| Name | Required | Description | Default |
|---|---|---|---|
| decided_by | Yes | ||
| approval_id | Yes | ||
| decision_ref | No | ||
| approval_token | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It discloses a meaningful precondition: approval must originate from the owner runtime capability, and actor strings are not enough. It does not state side effects, reversibility, or what happens on approval, but it does add real operational context beyond the tool name.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is only one sentence with no filler, and the key verb 'Approve' is front-loaded. It is appropriately compact, though the brevity contributes to the gaps in parameter and usage detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an action tool with four parameters, no annotations, and multiple sibling approval/rejection tools, the description is too thin: it omits how approval_id, decision_ref, and approval_token are used, what 'owner runtime capability' means, and when to prefer this over reject_override or request_override. The presence of an output schema covers return values, but the invocation semantics remain under-specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description should compensate, but it explains none of the four parameters: approval_id, decided_by, decision_ref, or approval_token. The warning about actor strings being insufficient may relate to decided_by, but it does not clarify any parameter's meaning, format, or requirements.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Approve') and the tool name supplies the object ('override'), so the core action is identifiable. It also adds a distinguishing constraint: approval must use the owner runtime capability, not actor strings. However, it never explains what an override approval entails or names the closest sibling alternatives, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool by stating that actor strings alone are insufficient and that the owner runtime capability is required, which is a form of usage guidance. It does not explicitly say when to choose approve_override over reject_override, request_override, or list_approvals, and offers no exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
assess_fast_judgmentA
Validate and gate provider-neutral typed fast judgments without persisting them.
Use for cheap BOOL/SCORE/CHOICE signals such as routing, triage, relevance or activation. The result is worker judgment only and never grants truth, Evidence, assurance, verification or mutation authority.
| Name | Required | Description | Default |
|---|---|---|---|
| source | No | HOST | |
| purpose | Yes | ||
| decisions | Yes | ||
| model_ref | No | ||
| high_impact | No | ||
| review_confidence | No | ||
| use_signal_confidence | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does real work: it discloses the non-persistence guarantee and explicitly bounds the output's authority ('never grants truth, Evidence, assurance, verification or mutation authority'). It does not cover auth requirements or the meaning of high_impact/review_confidence, but the side-effect and authority boundaries are unusually well stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action and scope, followed by usage then the authority disclaimer. No filler; each sentence adds a distinct layer.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the conceptual boundaries are covered well. However, for a 7-parameter tool with no annotations and zero schema descriptions, the complete absence of parameter-level guidance leaves a real gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 7 parameters, so the description is the only source of parameter meaning — and it explains none of them. The BOOL/SCORE/CHOICE mention loosely hints at the shape of `decisions`, but purpose, source, model_ref, high_impact, review_confidence and use_signal_confidence are entirely undefined.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb pair (validate and gate) on a specific resource (provider-neutral typed fast judgments), and the qualifier 'without persisting them' sharply distinguishes it from the sibling record_fast_judgment. It stops short of naming that sibling, so the agent must infer the routing from the phrase alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use for cheap BOOL/SCORE/CHOICE signals such as routing, triage, relevance or activation' gives concrete context for when the tool applies. There is no explicit when-not guidance or named alternative (e.g. record_fast_judgment for persistence), which is the only gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
attach_artifactC
Register a physical artifact/reference without changing its external storage.
| Name | Required | Description | Default |
|---|---|---|---|
| checksum | No | ||
| metadata | No | ||
| belongs_to | No | ||
| logical_name | Yes | ||
| artifact_type | Yes | ||
| storage_system | Yes | ||
| physical_location | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses a key behavioral trait: the operation registers without changing external storage, which helps an agent understand it is metadata-only and not a file-moving or storage-mutating action. However, with no annotations, the description still leaves gaps around idempotency, error behavior, checksum validation, and permissions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, focused sentence with no filler; it front-loads the core purpose. It is concise and easy to parse, though its brevity does not allow room for sibling differentiation or usage guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given seven parameters, no annotations, no schema descriptions, and a large set of sibling tools, the description is too sparse for an agent to confidently select and correctly invoke the tool. The output schema reduces the need to describe return values, but key context about when to use this versus registers, models, or links is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the schema's lack of parameter explanations. The description does not define or clarify any of the seven parameters, leaving the agent to infer meaning from names like logical_name, physical_location, and belongs_to without details on formats, constraints, or relationships.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the action ('Register') and the resource ('a physical artifact/reference'), and adds a meaningful qualifier ('without changing its external storage'). However, it does not explicitly distinguish itself from sibling tools like register_contract or register_model, so the differentiation is left to inference from the resource type.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided about when to use this tool versus alternatives such as register_contract, register_model, link_entities, or record_execution_receipt. The description gives a general sense of the operation but no conditions, exclusions, or explicit scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
attest_evidenceC
Attest evidence using a runtime verifier/owner capability. The capability is never persisted.
| Name | Required | Description | Default |
|---|---|---|---|
| authority | No | VERIFIER | |
| attested_by | Yes | ||
| evidence_id | Yes | ||
| capability_token | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses one important trait—'The capability is never persisted'—but does not explain side effects, whether the attestation record is stored, what permissions are needed, or any failure modes. This is insufficient for a mutation-like operation with no annotation support.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler. The primary action is front-loaded, and the key non-persistence trait is stated immediately afterward. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although an output schema exists, the tool still lacks usage guidance, parameter semantics, and behavioral side-effect information. For a 4-parameter tool with no annotations and no schema descriptions, this definition is not complete enough for an agent to invoke it reliably.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only vaguely references the capability concept and persistence. It gives some context for capability_token and authority, but says nothing about evidence_id or attested_by, leaving most parameter semantics unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear action ('Attest evidence') and a specific resource, and adds a mechanism ('runtime verifier/owner capability'). It is not a tautology and is understandable on its own, though it does not explicitly differentiate itself from sibling evidence-related tools like submit_evidence or verify_slice.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus alternatives such as submit_evidence, verify_slice, or evidence_freshness. The description implies attestation is a distinct operation, but it never states conditions, prerequisites, or exclusions, leaving the agent to guess.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_contextC
Return the bounded current audit frontier, findings, assumptions, mutation boundary and PCH working-set support.
| Name | Required | Description | Default |
|---|---|---|---|
| audit_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full burden, and it discloses almost nothing: 'Return' implies a read, but there is no statement about auth/permission requirements, whether the audit_id must reference an open audit, freshness, or error behavior. The opaque 'PCH working-set support' phrase adds no usable behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler or redundancy, which is appropriately sized for a one-parameter read tool. It is efficient even if the phrasing is jargon-dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be spelled out, and it is only a one-parameter tool. Still, the description leaves the audit_id undefined and the enumerated concepts largely uninterpretable, so an agent has minimum-viable but incomplete footing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions the single required audit_id parameter, so it does not compensate for the gap. It offers no format, source, or constraint guidance for the identifier.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Return) and enumerates the resources surfaced (audit frontier, findings, assumptions, mutation boundary, PCH working-set), so the general purpose is discernible. However, the enumerated items are opaque domain jargon and the description does nothing to distinguish this from siblings like audit_status or audit_mutation_allowed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no indication of when to call this versus the many sibling audit tools (start_scoped_audit, audit_status, record_audit_finding, close_scoped_audit). No prerequisites, no alternatives, no triggering context are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_mutation_allowedC
Check the audit's frozen mutation boundary. This does not replace normal Plan/WorkTurn or host permissions.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | ||
| audit_id | Yes | ||
| entity_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses that this is not an authorization grant and does not replace host permissions, but it does not confirm whether the operation is read-only, whether it has side effects, or what authority is required to call it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very short, front-loaded, and contains no filler. Both sentences serve a purpose, though the extreme brevity contributes to the parameter and usage gaps noted elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. Still, for a tool with three undocumented parameters, no annotations, and a complex audit/permissions domain, the description is too thin: it omits parameter semantics and has only a bare disclaimer about usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are three parameters with 0% schema description coverage, and the description says nothing about audit_id, ref, or entity_id. The meaning and format of each parameter must be inferred entirely from the schema titles, which is insufficient.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('Check') and a domain resource ('the audit's frozen mutation boundary'), but 'frozen mutation boundary' is jargon that does not clearly explain what is returned or how it differs from audit_status/audit_context. It is more concrete than a tautology, but still leaves the core purpose partially inferential.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second sentence clarifies that this does not replace normal Plan/WorkTurn or host permissions, which gives a usage caveat. However, it provides no explicit when-to-use condition or comparison to any of the many sibling tools, leaving alternatives and invocation context unstated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
audit_statusC
Return SRA/1 progress and whether impact closure reached a fixpoint or configured boundary.
| Name | Required | Description | Default |
|---|---|---|---|
| audit_id | Yes | ||
| frontier_limit | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It implies a safe read by saying 'Return', but never explicitly states it is non-mutating, what permissions are needed, or how audit_id scoping affects the result. For a status tool with zero annotation coverage this is thin, though the output schema does cover the return shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no wasted words; the core action and subject lead. Density of domain terminology slightly reduces readability but there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values are covered, but the two undocumented parameters and the absence of any usage/routing guidance leave the definition incomplete for a tool sitting among many *_status siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description mentions neither audit_id (required) nor frontier_limit (with default 16). Nothing in the prose compensates, so an agent learns nothing about either parameter's meaning or the role of frontier_limit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Return) and resource (SRA/1 progress and impact-closure fixpoint/boundary status), which reads as a read-only status query. The jargon ('SRA/1', 'impact closure', 'configured boundary') is domain-specific and opaque, but the purpose is discernible and distinct from mutation siblings like start_scoped_audit/close_scoped_audit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to call this versus alternatives such as audit_context, audit_mutation_allowed, or the other *_status tools. The agent is left to infer that a progress/fixpoint report is what it wants, with no stated preconditions or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
authorize_delegationC
Persist one bounded delegation authorization from current runtime/cost/concurrency state.
This is a policy/coordination record, not the model dispatch itself. The host must
refuse dispatch when authorized is false. High-cost workers require explicit
owner approval and only one high-cost delegation may be active per family.
| Name | Required | Description | Default |
|---|---|---|---|
| purpose | Yes | ||
| task_key | Yes | ||
| family_id | Yes | ||
| worker_key | Yes | ||
| input_scope | No | ||
| cost_ceiling | No | STANDARD | |
| owner_approval_id | No | ||
| coordinator_actor_id | Yes | ||
| required_capabilities | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It does reveal important constraints: it is a persisted record, the host must refuse dispatch when 'authorized' is false, and high-cost workers require explicit owner approval with a limit of one active per family. However, it omits details about error handling, idempotency, or what happens when constraints are violated. The description gives some context but not a complete behavioral picture.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, consisting of two short paragraphs. The main purpose is front-loaded in the first sentence, and the second paragraph adds relevant context about constraints. There is no wasted wording, and it is appropriately sized for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (9 parameters, 5 required) and the absence of schema descriptions, the description is insufficient. It does not explain what each parameter does, how they relate to the runtime/cost/concurrency state mentioned, or what the output schema contains. The description provides some high-level policy context but leaves an agent without enough information to correctly construct a valid request. For a tool with this many parameters, a much more thorough explanation is needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate by explaining parameters. It does not. The only parameter-related hint is that high-cost workers require owner approval, which loosely maps to 'cost_ceiling' and 'owner_approval_id', but the description does not explicitly tie these to parameters. Parameters like 'family_id', 'worker_key', 'task_key', 'purpose', 'input_scope', and 'required_capabilities' are left entirely unexplained. The description adds minimal semantic value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: 'Persist one bounded delegation authorization from current runtime/cost/concurrency state.' It specifies the resource (delegation authorization) and the verb (persist), and clarifies that this is a policy/coordination record rather than the model dispatch itself, distinguishing it from dispatch-related siblings. However, it does not explicitly name any sibling tool, so it lacks the strongest form of differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not provide explicit guidance on when to use this tool versus alternatives. It implies its use for creating authorizations but does not contrast with other delegation-related tools like 'complete_delegation' or 'delegation_status'. The statement 'This is a policy/coordination record, not the model dispatch itself' hints at usage but is not a clear directive. No when-to-use or when-not-to-use conditions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
backfill_work_identityC
Explicitly bind historical canonical family state to WorkIdentity without inventing user intent.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | ||
| family_id | Yes | ||
| project_id | Yes | ||
| source_ref | No | ||
| controller_token | No | ||
| controller_actor_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It implies a mutation/backfill operation but never states required authorization (the controller_actor_id/controller_token suggest a control plane), idempotency, conflict behavior on existing bindings, or reversibility.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with no filler, and the action leads the clause. Its brevity is not the problem — the abstraction is — so it earns credit for structure without reaching a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, but for a six-parameter mutation with zero annotation and zero schema coverage, the description omits who may invoke it, what state is mutated, and how it relates to the many adjacent binding/family tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across six parameters, and the description only loosely gestures at 'family state'. It adds no meaning for family_id, project_id, controller_actor_id, controller_token, source_ref, or title, leaving the agent to guess at the required control-plane actors.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a recognizable action (binding state to a WorkIdentity) via a backfill, which is more than a tautology. However, 'historical canonical family state' is domain jargon that does not clearly distinguish it from siblings such as bind_work_turn, effective_family_view, or bind_contract_turn.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when this tool should be used versus alternatives, no preconditions, and no note on when not to use it. The clause 'without inventing user intent' gestures at a constraint but does not give an agent any actionable routing guidance relative to the many bind_* and family_* siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
begin_workC
Compose intake -> plan -> start for ordinary work without MangoMe jargon.
Existing callers may pass proposed_slice. Zero-touch callers can instead pass
slice_title and optionally slice_objective; slice_declared_id is optional and
MangoMe generates a stable AUTO-* id when omitted. Domain failures are returned
as structured error data so the worker can self-correct.
| Name | Required | Description | Default |
|---|---|---|---|
| intent | Yes | ||
| spec_id | No | ||
| turn_id | No | ||
| work_id | No | ||
| actor_id | Yes | ||
| estimate | No | ||
| family_id | Yes | ||
| slice_title | No | ||
| contract_ids | No | ||
| request_text | Yes | ||
| classification | No | EXISTING_CONTRACT_WORK | |
| expected_scope | No | ||
| proposed_slice | No | ||
| slice_objective | No | ||
| slice_declared_id | No | ||
| expected_artifacts | No | ||
| classification_source | No | MANGOME_ZERO_TOUCH | |
| acceptance_expectations | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses only that 'Domain failures are returned as structured error data' and that a stable AUTO-* id is generated when slice_declared_id is omitted. It does not state whether the tool performs writes, what state it creates, permissions required, idempotency, or other mutation-side effects, leaving major behavioral traits undocumented.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is reasonably concise at four sentences and front-loads the composite purpose before moving to parameter variants and failure behavior. Most sentences earn their place, though the phrase 'without MangoMe jargon' is vague and adds little operational value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex 18-parameter composite mutation with no annotations, the description is incomplete. Output schema exists, so return values need not be detailed, but the description leaves most parameter semantics, usage versus siblings, and behavioral side effects unexplained, so an agent would struggle to invoke the tool correctly without additional inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 18 parameters, so the description must compensate. It adds meaning for a few optional parameters — proposed_slice, slice_title, slice_objective, and slice_declared_id — but leaves required parameters (family_id, actor_id, request_text, intent) and many optionals (classification, expected_scope, expected_artifacts, acceptance_expectations, estimate, etc.) completely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a composite sequence — 'Compose intake -> plan -> start' — which gives a rough purpose, but the verb 'Compose' is vague and the phrase 'ordinary work without MangoMe jargon' does not clarify scope. It does not distinguish this wrapper from siblings like intake_request, submit_plan, or start_slice, so an agent cannot tell when this tool is preferable to calling those separately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives caller-specific parameter guidance — 'Existing callers may pass proposed_slice' versus 'Zero-touch callers can instead pass slice_title' — which is useful. However, it never says when to use this tool versus the many sibling tools such as intake_request, enter_work, or start_slice, nor does it provide any exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bigbang_scanC
Candidate-only onboarding discovery. Forbidden as a recovery/state reconstruction path for admitted work.
| Name | Required | Description | Default |
|---|---|---|---|
| roots | Yes | ||
| id_patterns | No | ||
| include_git | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description bears the full burden of disclosing behavior. It only states the purpose and a usage restriction; it says nothing about side effects, required permissions, rate limits, or what the tool actually does beyond 'discovery.' It doesn't even indicate whether it is read-only or mutating. This is a significant transparency gap for a tool with zero annotation support.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is exceptionally brief – one sentence with two clauses. It is concise and front-loads the core purpose, but it sacrifices necessary detail. It doesn't waste words, but it under-delivers on content. As a result, it is appropriately sized for what it says, but that size is too small for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given three parameters, zero schema descriptions, and no annotations, the description is far from complete. It does not explain how to use the tool, what inputs are required, or what the output represents (though an output schema exists, the description adds no context). The tool is sufficiently complex that the description should provide at least a minimal example or note about the 'roots' parameter, but it does not. The description leaves an agent with insufficient information to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, meaning the description must compensate for the lack of parameter documentation. However, the description does not mention any of the three parameters (roots, id_patterns, include_git) or their semantics. The agent must rely solely on the bare schema property names, which are insufficient. This is a critical failure to aid parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action: 'Candidate-only onboarding discovery' – a clear verb+resource with a defined scope (candidate onboarding). It also draws a boundary against recovery/state reconstruction, which distinguishes it from recovery-related siblings like session_restore and recovery_context. However, the phrasing is terse and doesn't elaborate on what 'discovery' entails, so it stops short of being fully self-explanatory.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear when-not-to-use: 'Forbidden as a recovery/state reconstruction path for admitted work.' It implies when to use (for candidate onboarding) but does not name alternative tools or explain what to use for recovery instead. The exclusion is useful but the guidance on alternatives is implicit rather than explicit, so a 3 is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bind_contract_turnC
Bind the CURRENT user turn to one canonical Contract. Restore state alone never authorizes this turn.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | Yes | ||
| actor_id | Yes | ||
| ttl_minutes | No | ||
| contract_ref | Yes | ||
| request_text | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It communicates the core binding action and the important rule that restore alone is insufficient, but it omits mutation effects, permission requirements, idempotency, and reversibility.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences and front-loads the core action before the critical constraint. Every sentence adds some signal, and there is no wasted wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex contract-binding tool with 5 undocumented parameters and no annotations, the description is too sparse. It gives the central purpose and one important rule, but lacks parameter semantics, operational context, and usage alternatives.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are 5 parameters with 0% schema description coverage, and the description does not mention or explain any of them. In particular, request_text, mode, actor_id, contract_ref, and ttl_minutes remain fully undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource pairing: binding the CURRENT user turn to one canonical Contract. It is clear enough to distinguish from generic state tools, but it does not explicitly differentiate itself from the sibling bind_work_turn or contract_state tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The sentence 'Restore state alone never authorizes this turn' implies that this tool is required after restoring state to authorize the current turn. However, it does not explicitly state when to use this tool versus alternatives, nor does it name sibling tools or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bind_work_turnC
Control-plane bind the current user turn to durable WorkIdentity. Workers cannot self-mint authority.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | Yes | ||
| actor_id | Yes | ||
| work_ref | Yes | ||
| request_text | Yes | ||
| controller_token | No | ||
| controller_actor_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it discloses one meaningful behavioral fact: this is the control-plane authority-granting path and workers cannot self-mint authority. However, it says nothing about side effects, idempotency, required privileges for the controller fields, failure modes, or what happens to prior bindings.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the primary action front-loaded and no filler. The second sentence earns its place as scoping context rather than repeating the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having an output schema, this is a complex authority-binding operation with 6 params, 5 required, no annotations, and 0% param coverage. The description omits prerequisites, mode values, token semantics, and error behavior, leaving it far too thin for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 6 parameters, and the description explains none of them. work_ref, mode, actor_id, controller_actor_id, and controller_token all go undocumented in both schema and prose, so an agent has no semantic guidance for any argument.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'bind the current user turn to durable WorkIdentity.' The action is identifiable, but the jargon ('Control-plane', 'WorkIdentity') and the absence of any reference to the near-identical sibling bind_contract_turn leave the agent without sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second sentence ('Workers cannot self-mint authority') is a rationale/constraint, not when-to-use guidance. There is no statement of preconditions, no when-not-to-use, and no routing toward or away from bind_contract_turn or enter_work/begin_work.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bitemporal_truth_statusC
Summarize current BTTM/1 support states for one WorkIdentity.
| Name | Required | Description | Default |
|---|---|---|---|
| work_ref | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it says almost nothing. 'Summarize' weakly implies a read, but there is no disclosure of permissions, side effects, mutation risk, caching/freshness semantics, or what a 'support state' means. For a bitemporal status tool with zero annotation coverage, this is a substantial gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the issue is under-specification rather than verbosity. It is as short as it can be while still naming an operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, but that is the only structural aid. With no annotations, no parameter documentation, and undefined terminology, an agent lacks enough context to call this confidently or to prefer it over the sibling truth tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and work_ref has no description in the schema. The description's 'one WorkIdentity' hints that the parameter is a scoping identifier, but it never states what form work_ref takes or how it relates to a WorkIdentity, so it only partially compensates for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
It names a verb ('Summarize') and a scope ('one WorkIdentity'), so the operation is identifiable, but the core resource phrase 'current BTTM/1 support states' is opaque jargon that an agent cannot decode without domain knowledge. It also offers nothing to distinguish it from the many truth-related siblings like truth_at, truth_assertion_status, or record_truth_assertion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as truth_at or truth_assertion_status, which sit in the same truth-assertion family. The agent is left to infer the selection criterion entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
build_reproduction_bindingB
Build an RB/1 reproduction binding from current files/Git metadata without executing the command or reading environment values.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | Yes | ||
| command | Yes | ||
| exit_code | Yes | ||
| git_commit | No | ||
| input_paths | Yes | ||
| stderr_sha256 | No | ||
| stdout_sha256 | No | ||
| max_hash_bytes | No | ||
| environment_names | No | ||
| output_artifact_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It does disclose two important behavioral traits: the tool does not execute the command and does not read environment values. However, it does not disclose side effects, whether state is mutated, or what a 'binding' actually produces beyond the output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that states the purpose first and appends the constraints. Every word earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter tool with zero schema coverage and no annotations, the description is incomplete. It never explains what an 'RB/1 reproduction binding' is for, what output the agent should expect (despite an output schema), or what distinguishes this build step from related reproduction/execution tools among the many siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 10 parameters, so the description must compensate. It loosely maps 'files' to input_paths and 'Git metadata' to git_commit, and 'environment values' to environment_names, but leaves command, cwd, exit_code, stdout_sha256, stderr_sha256, max_hash_bytes, and output_artifact_id entirely unexplained. This is insufficient compensation for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Build') and a specific resource ('RB/1 reproduction binding'), and states the data sources (current files/Git metadata). It also conveys key constraints (no command execution, no environment reads). It doesn't explicitly differentiate from siblings, but the 'RB/1' qualifier plus the non-execution constraint make it reasonably distinguishable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'without executing the command or reading environment values' implies an offline/metadata-only mode, which is a soft usage hint. However, there is no explicit guidance on when to choose this over alternative tools or when not to use it — no named siblings or conditions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
checkpoint_workC
Persist PROGRESSIVE crash-recovery state. It can only reference an existing canonical WorkIdentity.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | EXECUTION_CHECKPOINT | |
| payload | Yes | ||
| plan_id | No | ||
| turn_id | No | ||
| actor_id | Yes | ||
| work_ref | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, yet it only establishes that this is a write ("Persist"). It says nothing about idempotency, whether a new checkpoint supersedes prior ones, retention or overwrite behavior, required permissions, or any rate/ordering constraints — all material for a crash-recovery write.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, and the primary purpose statement is front-loaded ahead of the constraint. It is efficient, though the brevity shades into under-specification rather than disciplined editing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a six-parameter, nested-object mutation with zero annotation coverage and zero schema descriptions, the description is far too thin. It covers the work_ref prerequisite but leaves the agent without enough context to construct payload, kind, or the plan/turn linkage correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across six parameters, so the description must compensate and largely does not. The single constraint it adds (the reference must map to an existing canonical WorkIdentity) clarifies work_ref, but kind, payload, plan_id, turn_id, and actor_id remain entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a verb ("Persist") and a resource ("crash-recovery state"), so the general intent is inferable, but "PROGRESSIVE crash-recovery state" is jargon that never says concretely what is being written (a snapshot? a delta? a pointer?). It also does nothing to separate this from siblings such as begin_work, bind_work_turn, or backfill_work_identity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second sentence supplies a real precondition — the work reference must already exist as a canonical WorkIdentity — which implies when this tool is valid versus when it is not. However, no alternative tool is named for the otherwise-implied case, and there is no guidance on when a checkpoint should be taken relative to begin_work or bind_work_turn.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
claim_doneB
Record DONE_CLAIMED under the slice's active plan; this never implies verification.
| Name | Required | Description | Default |
|---|---|---|---|
| plan_id | Yes | ||
| summary | No | ||
| actor_id | Yes | ||
| slice_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description must carry the full behavioral burden. It does contribute a meaningful behavioral guarantee: recording DONE_CLAIMED 'never implies verification.' It also scopes the action to the 'active plan.' Still, it does not disclose side effects, failure conditions, idempotency, or what happens if no active plan exists, leaving notable transparency gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single semicolon-joined sentence that is front-loaded with the core action and carries an important caveat, with zero filler. It is concise in form, though the brevity contributes to missing behavioral and parameter detail. Still, as a structure it is efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter mutation-style tool with no annotations and zero schema description coverage, the description is not complete enough. It gives the core semantic intent and one critical exclusion, but an agent would still be left guessing about required actor context, optional summary usage, and relationship to other status-changing siblings like completion_review or submit_evidence.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description should compensate for all four parameters. It only implicitly references slice_id and plan_id via 'slice's active plan'; actor_id and summary are entirely unexplained. This is a partial but insufficient contribution to parameter clarity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Record') and names a concrete resource/action ('DONE_CLAIMED under the slice's active plan'), which goes far beyond the tool name. The explicit caveat 'never implies verification' immediately separates it from sibling tools like verify_slice, giving an agent a clear sense of what this tool is and is not.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it: when you want to record a done claim rather than perform verification. However, it never names an alternative tool or states a concrete condition such as 'use verify_slice when actual verification is needed.' The guidance is contextually suggestive but not explicit enough to earn a 4.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
close_planB
Close an unbound plan so it stops generating collision traffic.
| Name | Required | Description | Default |
|---|---|---|---|
| plan_id | Yes | ||
| actor_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It does reveal a key effect ('stops generating collision traffic'), but it does not disclose side effects, reversibility, permissions, or what 'close' actually does to the plan beyond stopping traffic.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no wasted words. It front-loads the action and object, then gives the purpose. Every part earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has two required parameters, no annotations, and no parameter documentation, yet the description does not explain actor_id or the meaning of 'unbound plan' and 'collision traffic'. While an output schema exists and return values need not be described, the missing parameter semantics and side-effect context make this incomplete for safe invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not mention plan_id or actor_id at all. The schema only provides field names and types, so the agent cannot infer what values are expected, why actor_id is required, or how the parameters relate to closing a plan.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Close'), a specific resource ('an unbound plan'), and a concrete outcome ('stops generating collision traffic'). It is clearly distinguishable from sibling tools like submit_plan and start_slice, even without explicit sibling comparison.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool: when an unbound plan is generating collision traffic and should be stopped. However, it does not explicitly state when not to use it or mention any alternative tools, leaving some inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
close_scoped_auditB
Close only after no pending frontier remains; closure never promotes underlying work to VERIFIED.
| Name | Required | Description | Default |
|---|---|---|---|
| summary | No | ||
| actor_id | Yes | ||
| audit_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it does disclose one non-obvious trait: closure does not promote underlying work to VERIFIED. It is silent on permissions/actor requirements, reversibility or idempotency of closing, and what happens to recorded findings.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the precondition front-loaded and a second clause that adds distinct information. No filler, though the terseness borders on underspecification.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, but for a mutation tool with zero annotations and fully undocumented parameters the description should cover required authority and the effect on audit state. It covers the precondition and one side-effect constraint and little else.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and none of the three parameters (audit_id, actor_id, optional summary) are explained in the description. The description does not compensate at all for the schema gap, leaving actor attribution and summary semantics entirely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The verb 'close' plus the tool name identifies the resource (a scoped audit), and the clause 'closure never promotes underlying work to VERIFIED' actively separates it from verification-oriented siblings like verify_slice and accept_slice. It stops short of a full plain-language statement of what closure does to the audit record itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Close only after no pending frontier remains' states an explicit gating precondition for invocation, which is real when-to-use guidance rather than vague implication. It does not, however, point at siblings such as audit_status or audit_mutation_allowed for checking the frontier before calling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cognitive_hygieneD
Compute the PCH/1 thermal map and bounded active working set without mutating canonical truth.
| Name | Required | Description | Default |
|---|---|---|---|
| slice_id | No | ||
| family_id | Yes | ||
| query_text | No | ||
| max_active_objects | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description bears the full burden. The clause 'without mutating canonical truth' is a real behavioral disclosure (implies a read-only/non-mutating computation), which merits more than a 1, but permissions, cost, failure modes, and side effects are entirely undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no padding or redundancy. The problem is under-specification rather than verbosity, and the sentence itself is efficiently constructed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a four-parameter tool with no annotations and zero schema description coverage, the description is far too thin. An output schema exists so return format is partly covered, but the inputs and invocation context remain completely unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Four parameters at 0% schema description coverage, and the description adds no meaning for any of them (slice_id, family_id, query_text, max_active_objects). It neither explains what family_id identifies nor what 'max_active_objects' bounds.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The verb 'Compute' is generic and the resources ('PCH/1 thermal map', 'bounded active working set') are undefined internal jargon, so an agent cannot infer what the tool actually produces or how it differs from any sibling. It is more than a bare tautology because a subject is named, but the naming conveys no actionable meaning.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool, when not to, or which sibling to prefer. The sibling list contains dozens of lifecycle/context tools, and nothing here routes the agent between them.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compile_execution_contextC
Produce a compact deterministic execution package after PCH/1 working-set selection.
| Name | Required | Description | Default |
|---|---|---|---|
| slice_id | No | ||
| family_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description bears full behavioral disclosure. It mentions 'deterministic' but does not state side effects, permissions, idempotency, or whether the tool mutates state. 'Produce' is ambiguous about read vs write behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no wasted words, but it is under-specified rather than concise. The front-loaded verb is fine, yet the rest is too terse to be useful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema and no annotations, but the description omits parameter meanings, behavioral traits, and any differentiation from the many siblings. It is not complete enough for an agent to confidently invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not mention either parameter. Required family_id and optional slice_id are left completely unexplained, so the description fails to compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a verb 'Produce' and a resource 'execution package' but 'PCH/1 working-set selection' and 'compact deterministic execution package' are undefined domain jargon. It does not clearly distinguish this tool from siblings like compile_uai_context or prepare_assignment.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Only implies a prerequisite ('after PCH/1 working-set selection') but gives no explicit when-to-use guidance, no alternatives, and no conditions for choosing this tool over related ones. The agent is left to infer usage entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compile_uai_contextC
Compile canonical MangoMe truth into a versioned compact UAI/1 execution packet with semantic hash and savings metrics.
| Name | Required | Description | Default |
|---|---|---|---|
| slice_id | No | ||
| family_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the full burden. It does disclose meaningful behavior: it produces a versioned, compact packet, includes a semantic hash, and reports savings metrics. However, it does not state whether compilation has side effects, whether it is idempotent, whether it requires existing state, or what canonical truth depends on.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that packs the purpose, input source, output format, versioning behavior, and computed metrics. There is no filler, and the core action is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema helps with return values, but the tool still lacks usage guidance and completely omits parameter semantics. For a compile operation with 2 parameters and no annotations, an agent cannot confidently invoke it correctly from this description alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not mention family_id or slice_id at all. The description entirely fails to explain what family_id identifies or how slice_id changes compilation, leaving an agent with two parameters and no semantic guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action (compile), a clear source resource (canonical MangoMe truth), and a concrete output (versioned compact UAI/1 execution packet with semantic hash and savings metrics). It is distinct enough from sibling tools like expand_uai_context, decode_uai_result, and compile_execution_context, though it never names them directly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives such as compile_execution_context, expand_uai_context, or read_context. There are no prerequisites, exclusions, or context cues that would help an agent decide this is the right compilation step.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
complete_delegationC
Checkpoint/close a bounded delegation so recovery can continue without rediscovery.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | COMPLETED | |
| artifact | No | ||
| delegation_id | Yes | ||
| missing_delta | No | ||
| next_dependency | No | ||
| coordinator_actor_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description is the only behavioral disclosure. It indicates a state-changing action and one consequence (recovery can continue), but does not say whether closing is reversible, what side effects occur, what happens if the delegation is already closed, or any authorization requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler; it communicates the action and purpose efficiently. It loses a point only because 'Checkpoint/close' is compact to the point of being ambiguous.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite the presence of an output schema, the description is too thin for a 6-parameter tool with no annotations. It fails to explain required parameters, optional outcome fields, or the lifecycle context needed to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description provides no meaning for any of the 6 parameters. Terms like status, artifact, missing_delta, and next_dependency are left entirely unexplained, so an agent cannot correctly populate them from this description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Checkpoint/close') applied to a specific resource ('a bounded delegation') and gives a clear rationale ('so recovery can continue without rediscovery'). It distinguishes the tool's role from siblings like delegation_status or authorize_delegation, though it does not name them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The rationale implies the tool is used at a delegation checkpoint during recovery, which is useful contextual guidance. However, it gives no explicit when-to-use versus alternatives, no exclusions, and no conditions under which a delegation should not be completed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
completion_reviewC
Build an AV/1 adversarial verification brief from persisted DONE claims, plan/spec obligations, Evidence and an optional observed change set.
| Name | Required | Description | Default |
|---|---|---|---|
| slice_id | Yes | ||
| changed_paths | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It says the tool 'builds' a brief, but does not state whether this is a read-only operation, whether it persists anything, requires special permissions, or has side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. It conveys the core purpose efficiently, though the unexplained 'AV/1' jargon slightly reduces clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The presence of an output schema reduces the need to describe return values, but the description still lacks usage context, parameter semantics, and behavioral transparency. For a 2-parameter tool with no annotations, this is incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It vaguely maps 'optional observed change set' to changed_paths, but slice_id is not explained at all, and the relationship between the described inputs and the parameters is left implicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Build') and resource ('AV/1 adversarial verification brief') and names the input sources. It is clear about what the tool produces, though it does not explicitly differentiate it from sibling tools like verify_slice.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is used after DONE claims and evidence have been persisted, and optionally with a changed set. However, it gives no explicit guidance about when to choose this tool over alternatives such as verify_slice or claim_done.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contract_stateC
Read one logical Contract and its immutable canonical generations. Physical paths are bindings, not identity.
| Name | Required | Description | Default |
|---|---|---|---|
| contract_ref | Yes | ||
| include_content | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so the description carries the full behavioral burden. It discloses that generations are immutable canonical and that physical paths are bindings not identity, but omits read-only confirmation, permission requirements, and error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded purpose, no redundant text. The second sentence adds domain context and is not padding, though it is somewhat abstract.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists, so return values need not be explained. However, with no annotations and 0% parameter description coverage, the description leaves include_content behavior and contract_ref format unaddressed, making it incomplete for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description adds no meaning for contract_ref or include_content. Parameter names alone must carry all semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb 'Read' and resource 'logical Contract' plus scope 'immutable canonical generations.' Distinguishes read from write siblings like register_contract, but does not name alternatives such as resolve or truth_at.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use or when-not-to-use guidance; no alternative tools named. Usage is only implied by the verb 'Read.'
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_familyB
Administrative durable-family creation; requires control-plane authority.
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | ||
| scope_ids | No | ||
| family_key | Yes | ||
| project_ids | No | ||
| controller_token | No | ||
| controller_actor_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose one meaningful behavioral trait — that this is an administrative operation requiring control-plane authority — which is genuine added context beyond the schema. However, it says nothing about mutation effects, idempotency, or what the controller_token/controller_actor_id parameters imply behaviorally.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded clause with no waste, but it is under-specified for a six-parameter mutation tool rather than truly concise. Brevity here comes at the cost of completeness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, but for a creation/mutation tool with no annotations, six undocumented parameters, and only an authority note, the description leaves too much unspecified for an agent to invoke it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Six parameters with 0% schema description coverage, and the description adds nothing about any of them. The required family_key/title and the optional scope_ids, project_ids, controller_token, and controller_actor_id are undocumented in both the schema and the description, so the description fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('creation' of a 'durable-family'), which is distinguishable from siblings like create_project and create_spec. It lacks explicit differentiation language, but the resource noun is unique enough for an agent to identify it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'requires control-plane authority' gives a prerequisite/context cue for when this tool is available, but no alternatives are named and no when-not-to-use guidance is given. Usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_projectC
Administrative identity creation; requires control-plane authority.
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | ||
| description | No | ||
| project_key | Yes | ||
| controller_token | No | ||
| controller_actor_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it does disclose one genuinely useful behavioral trait: the control-plane authority requirement. However it says nothing about side effects, uniqueness/conflict behavior on project_key, or reversibility, which matters for a creation tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single short clause, so there is no wasted padding and the key constraint is front-loaded. Brevity here reflects under-specification rather than disciplined conciseness, but structurally it is tight.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need no explanation, but the description remains thin for a 5-parameter mutation with no annotations and no parameter documentation. The authority note is the one substantive point; the rest of what an agent needs to call this correctly is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 5 parameters, and the description adds no parameter-level meaning whatsoever. Nothing explains project_key, controller_token, or controller_actor_id, so the description fails to compensate for the total coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The phrase 'Administrative identity creation' gestures at creation but names the resource as 'identity' rather than 'project', and never states plainly what the tool produces. It is more label than description, leaving the agent to infer the resource from the tool name. Sibling tools like create_family and create_spec show a naming convention that this description does not reinforce.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states a prerequisite ('requires control-plane authority') but gives no when-to-use, when-not-to-use, or alternative-tool guidance. With many creation-flavored siblings (create_spec, create_family, register_contract), the agent gets no help distinguishing this from them.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_specC
Append normative Specification truth. Work-bound families require a MODIFY WorkTurn.
| Name | Required | Description | Default |
|---|---|---|---|
| turn_id | No | ||
| work_id | No | ||
| actor_id | No | ||
| family_id | Yes | ||
| objective | Yes | ||
| constraints | No | ||
| contract_ids | No | ||
| deliverables | No | ||
| out_of_scope | No | ||
| controller_token | No | ||
| required_evidence | No | ||
| supersedes_spec_id | No | ||
| acceptance_criteria | No | ||
| controller_actor_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. 'Append' implies mutation and 'normative' implies authority, but the description does not disclose permissions, side effects, reversibility, conflict handling with supersedes_spec_id, or authentication needs for a 14-parameter write operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very short and front-loads the core purpose, with no wasted words. However, for a 14-parameter tool, two terse sentences are underspecified rather than appropriately sized, so it is only minimally adequate.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists, so return values need not be explained. But with no annotations, 0% schema description coverage, and 14 parameters, the description is missing almost all operational context: required permissions, parameter meanings, when to use it, and how work-bound families must supply a MODIFY WorkTurn.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 14 parameters, and the description does not explain any parameter. The phrase about 'MODIFY WorkTurn' loosely hints at turn/work binding but does not name or clarify turn_id, work_id, family_id, objective, constraints, or any other field.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Append') and resource ('normative Specification truth'), giving a clear high-level purpose. It does not differentiate this tool from siblings such as record_truth_assertion or create_family, but the core action is identifiable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides one conditional prerequisite: 'Work-bound families require a MODIFY WorkTurn.' However, it gives no guidance on when to use create_spec versus alternatives, no exclusions, and no workflow ordering beyond that single condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
decode_uai_resultA
Validate and expand a compact UAI/1R worker result. This never mutates canonical state.
| Name | Required | Description | Default |
|---|---|---|---|
| result_json | Yes | ||
| expected_context_hash | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states the key trait that it never mutates canonical state, which is important. However, it does not describe error handling, what happens on validation failure, or any other side effects. The disclosure is moderate, not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences: 'Validate and expand a compact UAI/1R worker result. This never mutates canonical state.' It front-loads the core purpose and adds a critical behavioral guarantee with zero filler. Perfectly structured for quick parsing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the main purpose and non-mutation, and an output schema exists to handle return values. However, it omits details about validation semantics and the specific role of expected_context_hash, leaving some gaps in completeness for a 2-parameter tool. It is adequate but not exhaustive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It implies that result_json is the compact worker result and mentions validation, which likely involves expected_context_hash. However, it does not explicitly define either parameter, leaving some ambiguity. It adds partial meaning but not enough to fully offset the 0% coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('validate and expand') and resource ('compact UAI/1R worker result'), and clarifies it never mutates canonical state. This clearly distinguishes it from mutation-heavy siblings like compile_uai_context or render_uai_result, making the tool's purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear context: it is used for validating and expanding a compact UAI/1R worker result. It does not explicitly mention alternatives or exclusions, but the context is specific enough that an agent can infer when to use this tool. It lacks explicit 'when-not-to-use' guidance, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delegation_statusB
Return persisted delegation/checkpoint state for one family.
| Name | Required | Description | Default |
|---|---|---|---|
| family_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It adds useful context by calling the state 'persisted' and 'checkpoint', implying a durable read-only query. However, it does not disclose auth requirements, error behavior, freshness, or whether the state reflects in-flight delegation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler. Every word adds meaning: action, resource, persistence qualifier, and scope.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple (one required parameter) and has an output schema, so not describing return values is acceptable. However, the missing parameter semantics and absence of any usage guidance leave a meaningful gap for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and family_id is only shown with its title. The description mentions 'one family' but does not explain what family_id is, how to obtain it, or what format/value constraints apply.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Return'), a specific resource ('persisted delegation/checkpoint state'), and a clear scope ('for one family'). This distinguishes it from mutation siblings like authorize_delegation and complete_delegation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance about when to use this tool versus alternatives. With many sibling status/state retrieval tools (status, workspace_status, effective_family_view), the agent is left to guess which one fits its situation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
discovery_scopesB
Return portable typed discovery scopes; this does not scan or admit content.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace_root | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the transparency burden. It usefully discloses a key negative behavior (no scanning or content admission) and 'Return' implies read-only retrieval, but it does not address side effects, permissions, or error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the verb and resource appear first, and the negative scope statement earns its place. There is no filler or redundant restating of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even with an output schema present, the description leaves key gaps: what 'discovery scopes' are, how workspace_root affects results, and when to choose this over the many sibling tools. It is too terse for a fully informed agent call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description never mentions workspace_root, its meaning, format, or effect on the returned scopes. The description adds no parameter-level value beyond the schema's name and default.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Return') and resource ('portable typed discovery scopes'), and explicitly contrasts with content scanning/admission. This distinguishes it from sibling tools like filesystem_scan and intake_request, though the term 'discovery scopes' remains domain-specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The clause 'this does not scan or admit content' implies exclusions and indirectly points toward alternatives like filesystem_scan or intake_request. However, it never states a positive use case, prerequisites, or explicit routing conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
effective_family_viewC
Return the effective append-only contract-family view, supersession, relations, and conflicts.
| Name | Required | Description | Default |
|---|---|---|---|
| family_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the behavioral disclosure burden. 'Return' weakly implies a read operation, and 'append-only' describes the view's nature rather than the operation's behavior. There is no explicit statement that the call is side-effect-free, no failure/side-effect context, and no explanation of how conflicts are resolved.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that leads with the verb and object, then lists the key output categories. Every word earns its place; there is no filler or redundant restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema covers return structure and the description names the major result categories, so the return side is adequately addressed. However, with no annotations and no usage or parameter guidance, the description is only minimally complete for an agent deciding when and how to invoke this tool among many sibling view/query tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has zero description coverage for the single parameter, and the description never directly explains family_id. The phrase 'contract-family view' is a weak clue connecting family_id to the family domain, but the description does not compensate for the missing schema coverage with format, provenance, or constraint details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Return') and names a precise resource ('effective append-only contract-family view'), then enumerates the main output categories: supersession, relations, and conflicts. It is clear, but it does not explicitly differentiate this from sibling read/view tools such as graph, resolve, or project_overview.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to use this tool versus alternatives. The agent is not told how this view differs from sibling tools like graph, resolve, or project_overview, so usage must be inferred from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
enter_workC
Enter governed work from ordinary user intent with no MangoMe identifiers required.
The managed workspace becomes a deterministic operational Project/Family + WorkIdentity. The current user request is admitted as an operational-intent baseline, not automatically as a Specification or Contract. Discovery candidates are never promoted automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| intent | No | ||
| actor_id | Yes | ||
| estimate | No | ||
| slice_title | No | ||
| request_text | Yes | ||
| expected_scope | No | ||
| slice_objective | No | ||
| controller_token | No | ||
| required_evidence | No | ||
| expected_artifacts | No | ||
| acceptance_criteria | No | ||
| controller_actor_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose real behavioral traits – the workspace becomes a deterministic Project/Family + WorkIdentity, the request is admitted as an operational-intent baseline only, and discovery candidates are never auto-promoted. However, it omits idempotency, permissions, reversibility, and escalation semantics for a mutation-style tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The three short paragraphs are reasonably front-loaded, but the middle and final sentences are dense internal jargon that restates governance philosophy rather than adding callable information. Size is acceptable; the wording is opaque rather than wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-parameter governance tool with no annotations, the description is far too abstract to guide correct invocation. The presence of an output schema relieves it of explaining return values, but it still fails to cover parameter intent, prerequisites, or how to choose it over near-identical siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are 12 parameters with 0% schema description coverage, and the description addresses none of them. Terms like 'actor_id', 'controller_token', 'required_evidence', 'expected_scope', and 'slice_objective' are entirely undocumented, so the agent has no basis for supplying anything beyond the two required fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses an abstract verb+resource ('Enter governed work') but never states concretely what the tool creates or returns. It is heavily laden with internal jargon (MangoMe identifiers, WorkIdentity, operational-intent baseline) and gives no way to tell it apart from close siblings like begin_work, start_slice, or intake_request.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is a faint usage hint ('from ordinary user intent with no MangoMe identifiers required') implying use when IDs are absent, but no explicit when/when-not and no alternatives are named. With dozens of similarly named sibling tools, this leaves the agent guessing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
evidence_freshnessA
Check whether attested PASS evidence remains current under legacy hash bindings or RB/1 reproduction bindings. Never reruns commands or changes assurance state.
| Name | Required | Description | Default |
|---|---|---|---|
| live_check | No | ||
| evidence_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it does so well by explicitly stating the tool never reruns commands and never changes assurance state. This effectively communicates a read-only, non-mutating safety profile, though it does not address auth requirements or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences deliver the core purpose and the critical negative behavior with no filler. The information is front-loaded and every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description gives a clear purpose and safety profile, and the output schema presumably covers return values. However, with no annotations and zero parameter documentation, the agent is left guessing about what evidence_id should contain and what live_check controls, which is a meaningful gap for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate by explaining the parameters. It does not mention evidence_id or live_check at all, leaving their semantics completely undefined beyond the raw schema field names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Check') and resource ('attested PASS evidence') and adds the precise scoping condition ('under legacy hash bindings or RB/1 reproduction bindings'). It also clarifies what the tool does not do, which helps distinguish it from verification and attestation siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for checking freshness without side effects ('Never reruns commands or changes assurance state'), giving clear context for when to use it. It does not explicitly name alternatives or exclusion conditions, so it stops short of a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
execution_eligibilityA
Evaluate the worker's CURRENT runtime capabilities/cost before dispatch or a protected action.
MangoMe returns an authorization decision but does not itself dispatch the model. The external orchestrator must enforce a negative decision at its real dispatch boundary and should re-check before capability-sensitive actions such as DEPLOY.
| Name | Required | Description | Default |
|---|---|---|---|
| family_id | No | ||
| worker_key | Yes | ||
| cost_ceiling | No | STANDARD | |
| delegation_key | No | ||
| owner_approval_id | No | ||
| required_capabilities | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that the tool returns an authorization decision but does not dispatch the model, which is critical behavioral information. However, it doesn't disclose details about authorization decision types, potential side effects (read-only vs mutation), or what happens on failure, which are gaps given the lack of annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded with the core purpose, followed by necessary clarifications about dispatch enforcement. The second paragraph is slightly verbose but adds critical context. Overall, it is structured well with minimal waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 6 parameters and an output schema, but no parameter descriptions in the schema. The description provides high-level context (authorization decision, enforcement boundary) but lacks detailed parameter semantics and behavioral specifics like return value structure, error handling, or side effects. Given the complexity and missing schema descriptions, the description is incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, but it does not explain any of the six parameters (worker_key, cost_ceiling, required_capabilities, etc.). The description only mentions 'runtime capabilities/cost' at a high level. With zero schema descriptions, this is a significant gap; the description could have outlined the meaning of cost_ceiling or required_capabilities.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool evaluates the worker's current runtime capabilities/cost before dispatch or a protected action, distinguishing it as an authorization check rather than a dispatch mechanism. It mentions 'MangoMe' and 'orchestrator' which could be more explicit but overall the purpose is clear and specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear guidance on when to use the tool (before dispatch or protected actions like DEPLOY) and explicitly states that the external orchestrator must enforce negative decisions, implying when not to rely on this tool for enforcement. It doesn't name specific sibling alternatives but the context is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
expand_uai_contextC
Round-trip a UAI/1 context packet back into its canonical semantic execution projection and verify its hash.
| Name | Required | Description | Default |
|---|---|---|---|
| wire | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the entire burden. It discloses that the tool 'verify its hash,' which is a behavioral trait, but provides no other details—such as side effects, error behavior, or whether it is read-only. The description is too sparse to give an agent confidence about the tool's behavior beyond the stated operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence that front-loads the verb and keeps essential information. It is appropriately sized for the operation, though it is dense and uses specialized jargon that may hinder comprehension. Efficiency is good, but clarity suffers slightly from terminology.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is an obscure, domain-specific operation with no usage context or differentiation from numerous siblings. The output schema exists but the description doesn't need to explain returns. However, it omits when to use this tool versus compile_uai_context or decode_uai_result, and lacks details on hash verification behavior or failure modes. An agent would struggle to decide when to invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It loosely ties the parameter `wire` to a 'UAI/1 context packet,' giving some meaning, but does not explain the format, constraints, or how the packet is structured. For a single parameter with no schema documentation, this is insufficient.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action: 'Round-trip a UAI/1 context packet back into its canonical semantic execution projection and verify its hash.' It identifies the resource (UAI/1 context packet) and the verb (round-trip, verify). It is clear what the tool does, though it does not explicitly differentiate from siblings like compile_uai_context or decode_uai_result, so it's not a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus alternatives. It does not mention any conditions, prerequisites, or exclusions. The description only states what it does, leaving the agent to infer appropriate usage from the domain context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fast_judgment_statusB
Read one persisted FJD/1 progressive judgment and its confidence/fallback disposition.
| Name | Required | Description | Default |
|---|---|---|---|
| judgment_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the behavioral burden. "Read" establishes this is a non-mutating lookup and "persisted" hints the record must already exist, but it says nothing about authorization needs, error behavior for unknown ids, or what the fallback disposition implies operationally.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the resource and its returned disposition are stated immediately. It is dense with domain jargon but not padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the description does say what is retrieved. However, with no annotations and no usage context, an agent still lacks the prerequisites and the distinction from the record/assess siblings to invoke this confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema gives 0% description coverage for judgment_id, so the description must compensate. It only implies the id identifies "one persisted ... judgment"; it never states the id format or where an agent obtains it, leaving the single required parameter under-specified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ("Read") and resource ("one persisted FJD/1 progressive judgment") plus what it returns ("confidence/fallback disposition"). This differentiates it from siblings like assess_fast_judgment and record_fast_judgment, which write rather than read, though the "FJD/1" jargon is unexplained.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no stated prerequisites (e.g., that a judgment must first be recorded), and no named alternative among the many sibling status tools. The word "persisted" weakly implies a prior write, but nothing is explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
filesystem_referencesC
Targeted validation of an already-canonical identity; never use lexical references to invent recovery state.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| declared_id | Yes | ||
| present_only | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses a behavioral constraint (avoid inventing recovery state) but does not state whether the operation is read-only, destructive, requires auth, or what side effects occur. The cryptic warning offers minimal transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very short and concise, but it is under-specified rather than efficiently structured. It is a single sentence with no front-loading of key information, and it omits critical details, so conciseness does not translate to clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has an output schema (which may cover returns), the description still fails to explain the tool's core purpose fully, its inputs beyond the schema, or when to invoke it. The cryptic phrasing leaves an agent with insufficient context to use it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It hints at 'declared_id' via 'declared identity' but does not explain limit or present_only. The description adds little meaning beyond the parameter names and fails to clarify their semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('validate') and a resource ('canonical identity'), but 'canonical identity' is ambiguous and not explicitly tied to filesystem references, despite the tool name. It is not a tautology, but it lacks the clarity needed to distinguish it from siblings like filesystem_scan.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a caution ('never use lexical references to invent recovery state') but gives no explicit guidance on when to use this tool versus alternatives. It does not mention conditions or scenarios that select this tool over others.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
filesystem_scanC
Broad inventory for onboarding/maintenance; forbidden as state reconstruction for admitted work.
| Name | Required | Description | Default |
|---|---|---|---|
| roots | Yes | ||
| max_depth | No | ||
| max_files | No | ||
| max_hash_bytes | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description bears full responsibility for disclosing safety and behavioral traits. It only says 'broad inventory', which implies a read operation but does not explicitly state read-only, performance implications, or what data is returned. Minimal disclosure, but not outright contradictory.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single sentence, which is concise, but it is under-specified and does not front-load the core action or key constraints. The brevity comes at the cost of usefulness, so it does not earn its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A complex tool with 4 parameters, 0% schema coverage, and no parameter descriptions in the text is completely inadequate. The description fails to explain what the tool returns, how parameters interact, or any operational context, making it nearly impossible to use correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description makes no mention of any of the four parameters (roots, max_depth, max_files, max_hash_bytes). Given schema description coverage is 0%, the agent has no clue what these mean or how to set them. This is a critical gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states 'broad inventory' without a clear verb or subject, remaining vague about what is inventoried (filesystem? files? directories?). It does not distinguish from siblings like filesystem_references or bigbang_scan, and the name alone carries most of the meaning.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context: use for onboarding/maintenance, and explicitly forbidden for state reconstruction of admitted work. This gives both when and when-not, though it names no alternative tool, so it stops short of full guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
graphA
Return validated confirmed/suggested incoming and outgoing graph edges for an entity.
| Name | Required | Description | Default |
|---|---|---|---|
| entity_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds useful behavioral nuance—edges are 'validated,' can be 'confirmed/suggested,' and are both 'incoming and outgoing'—which goes beyond simply saying 'returns graph edges.' However, because no annotations are provided, it does not disclose error behavior, permissions, or scalability/pagination characteristics, and 'validated' is left unexplained.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the action, and every word adds meaning ('validated,' 'confirmed/suggested,' 'incoming and outgoing'). There is no filler, repetition, or extraneous detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given one required parameter, a clear return focus, and an output schema, the description covers the essentials for a simple read query. It could mention usage context or define validated/confirmed/suggested more precisely, but those gaps are minor at this complexity level.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the parameter semantics burden. It confirms that entity_id identifies the entity whose edges are returned, but adds no detail on ID format or accepted values. This is acceptable for a single self-descriptive parameter, though the description barely adds meaning beyond the schema title.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Return') and resource ('graph edges') with differentiating details: 'validated confirmed/suggested' and both 'incoming and outgoing.' This clearly distinguishes it from generic sibling tools like status, link_entities, or resolve, so an agent can identify the tool's purpose from the description alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use graph versus alternatives, when not to use it, or what conditions make it the appropriate choice. The description only states what the tool does, leaving applicability entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
healthA
Return readiness even when backing-store initialization fails.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses a notable behavioral trait (returns readiness even if backing-store init fails), which is valuable. However, it doesn't mention side effects, whether it's read-only, or what happens under normal conditions, leaving some uncertainty.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that conveys the core purpose and a critical edge case without any wasted words. It is optimally concise for a zero-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no parameters and an existing output schema, the description covers the essential behavior (readiness even on failure). It doesn't elaborate on what 'readiness' means or typical usage context, but that is largely implied and the output schema fills in return details. Minor gap: no explicit statement that this is a health check for system liveness, though the name and content suggest it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so per the baseline rule it scores 4. The description adds nothing about parameters (correctly, as there are none), and the schema is trivial (empty object).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool returns readiness and specifies a key condition (even when backing-store initialization fails). It is a specific verb+resource and distinct enough from siblings like 'status' or 'workspace_status', though it doesn't explicitly differentiate from them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives. It doesn't state prerequisites, typical scenarios, or why one would call it over other status-like tools. The context is implied (health check) but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
import_contract_bundleB
Controller-authorized onboarding of historical contract/slice state.
This is explicit import/backfill, never restore. When exactly one project is supplied, the imported family is immediately bound to historical WorkIdentity without inventing user intent; productive execution still requires an admitted normative target + WorkTurn.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | BASE | |
| slices | Yes | ||
| actor_id | Yes | ||
| scope_ids | No | ||
| family_key | Yes | ||
| declared_id | Yes | ||
| project_ids | No | ||
| family_title | Yes | ||
| contract_title | Yes | ||
| storage_system | No | ||
| controller_token | No | ||
| physical_location | No | ||
| controller_actor_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose real behavior: controller authorization is required, single-project imports auto-bind to historical WorkIdentity, and productive execution needs an admitted normative target + WorkTurn. It still omits idempotency, reversibility, what existing state is overwritten, and permission specifics for a mutating import.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded and compact at three sentences with no obvious filler, but phrases like 'without inventing user intent' and 'admitted normative target' are opaque rather than informative. Density here trades clarity for brevity rather than earning every word.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex 13-parameter import with no annotations and 0% schema coverage, the description is far too thin; the output schema covers return values, but input semantics, preconditions, and side effects are largely absent. An agent could not reliably construct a correct call from this alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% across 13 parameters, so the description must compensate, and it only gestures at a few concepts (project, family, controller, slices) without mapping them to any parameter. Most of the 13 fields (kind, scope_ids, storage_system, physical_location, declared_id, contract_title, etc.) receive no semantic clarification from either source.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action on a specific resource: 'onboarding of historical contract/slice state', and explicitly frames it as 'import/backfill, never restore', which separates it from siblings like session_restore and backfill_work_identity. However, 'onboarding' is dense jargon and the sentence is not immediately parseable for an agent unfamiliar with the domain.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'This is explicit import/backfill, never restore' provides a clear when-not (not a restore operation), which is useful routing guidance. But there is no explicit statement of when to choose this over the many adjacent siblings (backfill_work_identity, register_contract, create_family), so selection guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
intake_requestB
Persist and categorize a new assignment before execution.
| Name | Required | Description | Default |
|---|---|---|---|
| family_id | No | ||
| source_ref | No | ||
| request_text | Yes | ||
| classification | No | ||
| classification_source | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden. It meaningfully states that the tool persists and categorizes rather than executes, which is a useful behavioral signal. It does not mention side effects, reversibility, or whether existing state is modified, but it is not misleading.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with the main action front-loaded and no filler, repetition of the tool name, or irrelevant detail. It is appropriately short, though brevity comes at the cost of parameter guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter tool with no annotations and no schema descriptions, the one-line description leaves key invocation details unstated, such as what constitutes a valid request and how classification is supplied or derived. The presence of an output schema mitigates return-value ambiguity, but input semantics remain under-specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description adds no explicit meaning for request_text, family_id, source_ref, classification, or classification_source. The verbs 'persist' and 'categorize' only faintly hint at request_text and classification, so the description fails to compensate for the absent schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb phrase ('Persist and categorize') and resource ('a new assignment') and situates the action 'before execution,' which separates it from execution-oriented siblings like begin_work and start_slice. It does not explicitly distinguish it among other creation tools such as create_project or register_contract, so differentiation is partial.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'before execution' gives a clear temporal usage context, implying this tool belongs to the intake stage rather than execution. However, it provides no explicit when-not-to-use guidance and names no alternatives, so routing must be inferred from sibling names.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
invalidate_truth_assertionA
Close an assertion's known-time interval and propagate REVALIDATION_REQUIRED to dependents.
Requires a canonical MODIFY work turn. Historical truth is preserved; invalidation never deletes evidence/history.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | ||
| turn_id | Yes | ||
| actor_id | Yes | ||
| known_at | No | ||
| assertion_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and does so well: it discloses the side effect (propagate REVALIDATION_REQUIRED to dependents), the prerequisite (canonical MODIFY work turn), and the preservation guarantee (historical truth preserved, evidence/history never deleted). This gives an agent essential operational context for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core action, followed by the prerequisite and preservation guarantee. Every sentence adds distinct value with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is complex, with five parameters, no annotations, and zero schema description coverage. The description supplies important behavioral and prerequisite context, but it omits parameter semantics and usage alternatives. Given the output schema exists, return-value explanation is not required, but the definition remains incomplete for a mutation tool of this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for five parameters, four of them required. It indirectly references assertion_id and turn_id, but provides no explanation of actor_id, reason, known_at, or the expected formats. Most parameter semantics are left to the bare schema, which has no property descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action on a specific resource: closing an assertion's known-time interval and propagating REVALIDATION_REQUIRED. This is more precise than a tautology and distinguishes it from sibling tools like record_truth_assertion and truth_at, though it does not explicitly name or contrast against those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states a prerequisite ('Requires a canonical MODIFY work turn') but gives no guidance on when to choose this tool over alternatives such as record_truth_assertion, truth_assertion_status, or bitemporal_truth_status. There are no when-not-use conditions or explicit alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
link_entitiesC
Create a typed relation; normative contract evolution requires a MODIFY WorkTurn.
| Name | Required | Description | Default |
|---|---|---|---|
| to_id | Yes | ||
| status | No | CONFIRMED | |
| from_id | Yes | ||
| to_type | Yes | ||
| turn_id | No | ||
| work_id | No | ||
| relation | Yes | ||
| from_type | Yes | ||
| confidence | No | ||
| source_actor_id | No | ||
| controller_token | No | ||
| controller_actor_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and largely fails: it never states whether the relation is idempotent, what authorization the controller_token/controller_actor_id imply, whether status defaults to CONFIRMED silently, or what side effects a relation edge has. The MODIFY WorkTurn remark hints at a workflow constraint but is too under-explained to act on.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is short, but this is under-specification rather than conciseness: a cryptic jargon fragment ("MODIFY WorkTurn") is front-loaded after the verb phrase and does not earn its place as helpful guidance for an agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, but for a 12-parameter, no-annotation, domain-heavy mutation tool the definition is far too thin. The authorization and workflow semantics implied by the tokens and the WorkTurn clause are left unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 12 parameters with 0% description coverage, and the description compensates only by implying that 'relation' is 'typed' (from_type/to_type). Nothing is said about status, confidence, turn_id, work_id, source_actor_id, or the controller authorization fields, leaving the majority of the surface undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening clause gives a specific verb and resource ("Create a typed relation"), which maps reasonably to a link_entities tool. However, it never differentiates itself from plausible siblings like graph, attach_artifact, or the various create_* tools, and the dangling semi-colon clause is opaque rather than clarifying.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only usage signal is the cryptic fragment "normative contract evolution requires a MODIFY WorkTurn," which implies a precondition but does not state when to call this tool, when not to, or which sibling to use instead. No explicit alternative routing is offered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_approvalsA
List approval requests, optionally filtered by status and subject.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | ||
| subject_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of behavioral disclosure. It states the read/list behavior and the optional filtering dimensions, but does not disclose possible status values, default behavior when filters are omitted, or pagination/return behavior beyond what the output schema already supplies.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, efficient sentence that front-loads the primary action and then specifies the optional filters. No filler or redundant wording is present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter read tool with an output schema, the description is minimally adequate. However, the absence of annotations and the lack of status-value guidance or filter behavior leave meaningful gaps in what an agent needs to call it correctly without external domain knowledge.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for parameter semantics. It only names 'status and subject' without explaining valid statuses, what subject_id refers to, or how the filters interact with each other, leaving the agent without semantic guidance beyond property names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the resource ('approval requests') and the specific action ('List'), with optional filtering by status and subject. The verb and resource set it apart from mutation-oriented siblings like approve_override and reject_override.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this tool is for retrieving approval requests and supports optional filtering, which gives reasonable context. However, it does not explicitly state when to use this tool versus alternatives or mention any exclusions or alternatives among the sibling workflow tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
maintenance_diagnoseA
Report stale plan candidates, collisions, approvals and unresolved graph work without auto-fixing it.
| Name | Required | Description | Default |
|---|---|---|---|
| stale_after_hours | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral disclosure burden. It clearly signals a non-mutating diagnostic action by saying 'Report... without auto-fixing it', which is the most important behavioral trait. It does not go deeper into side effects or execution characteristics, but the output schema helps cover the result side.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One tight sentence with a front-loaded verb, a precise list of report categories, and an explicit scope exclusion. Every part earns its place and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is low-complexity and has an output schema, so the description does not need to explain return values. However, it leaves the `stale_after_hours` threshold implicit and gives no hints about when to prefer this over sibling diagnostic/reporting tools. It is adequate but has clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions the `stale_after_hours` parameter or how it influences the report. The agent has to infer from the parameter name that it controls the staleness threshold, which is not explicitly connected to the tool's behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Report' and names specific reportable items: stale plan candidates, collisions, approvals, and unresolved graph work. It also explicitly excludes auto-fixing, which distinguishes it from any mutation-oriented sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'without auto-fixing it' implies a diagnostic-only usage context, but the description does not explicitly say when to use this tool versus related siblings like health, status, or list_approvals. No alternative tools or exclusion conditions are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migrate_schemaB
Dry-run or explicitly persist registered lazy schema migrations.
| Name | Required | Description | Default |
|---|---|---|---|
| dry_run | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the behavioral burden. It does disclose that one mode is non-mutating ('dry-run') and the other mutating ('explicitly persist'), which is key. However, it does not explain consequences of persisting, such as irreversibility, required permissions, or impact on the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very short and front-loads the core action. Every word contributes, though it borders on under-specification and uses jargon that could obscure meaning. It is efficient but not fully self-explanatory.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with an output schema, the description covers the essential call shape. Still, it omits context such as the need for migrations to be registered first, what persistence actually changes, and any safety or approval implications. It is minimally adequate but leaves gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It effectively explains the single boolean parameter: dry_run means perform a dry-run, while explicitly persisting corresponds to the non-dry-run path. This adds meaning beyond the bare boolean type and default.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies a clear action resource: it dry-runs or persists registered lazy schema migrations. It distinguishes the two modes, though it does not differentiate the tool from sibling tools and relies on niche terminology like 'lazy schema migrations'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus alternatives, and no prerequisites or exclusions are stated. The dry-run vs. persist phrasing implies two use modes, but it does not tell an agent when persistence is appropriate or what must be true before invoking it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
model_statsC
Return empirical cost/verified-outcome statistics.
| Name | Required | Description | Default |
|---|---|---|---|
| model_id | No | ||
| work_class | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. 'Return' implies a read-only operation, but the description does not state side effects, data freshness, access requirements, aggregation scope, or whether filters change behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler or redundancy. It is terse almost to a fault, but conciseness itself is handled well.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple and has an output schema, so the return type may be covered, but the description omits usage context, filter semantics, and behavioral details. An agent could call it correctly with no arguments, but would struggle to use it deliberately for a specific model or work class.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description should compensate, but it never mentions model_id or work_class. Only the property names and titles hint that they select a model and work class; the description does not clarify whether they are filters, how they combine, or what values are valid.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a clear verb ('Return') and a specific object ('empirical cost/verified-outcome statistics'), so an agent can tell it is a query tool for model outcome data. It does not explicitly contrast with look-alike siblings such as health or status, but the resource and data type are distinct enough.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call model_stats instead of health, workspace_status, status, or evidence_freshness, and no conditions or exclusions. An agent must infer the intended use from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prepare_assignmentC
Prepare internal execution for an admitted assignment without exposing MangoMe orchestration to the user.
Existing Slices are reused and never duplicated. If the admitted Family has no Slice, MangoMe materializes one minimal internal execution Slice. For audit/review assignments, the worker must still perform the audit and persist Evidence; planning or producing another contract/specification is not completion.
| Name | Required | Description | Default |
|---|---|---|---|
| intent | No | ||
| turn_id | No | ||
| work_id | No | ||
| actor_id | Yes | ||
| estimate | No | ||
| family_id | Yes | ||
| read_only | No | ||
| request_text | Yes | ||
| classification | No | ||
| expected_scope | No | ||
| expected_artifacts | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose meaningful behavior: existing Slices are reused and never duplicated, a minimal internal Slice is materialized only when the Family lacks one, and for audit/review it must persist Evidence (planning alone is not completion). That is useful mutation/idempotency context, but it omits permissions, side effects on the assignment, and what 'prepare' actually changes internally.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then supporting behavior. Three tight sentences with no obvious padding, though the audit/review sentence drifts toward edge-case policy rather than core function.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation. However, for an 11-parameter, high-complexity orchestration tool with zero annotation coverage, the description leaves the entire parameter surface undocumented and does not explain ownership/permission or lifecycle impact. Behavioral intent is covered but the call contract is not.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 11 parameters, and the description explains none of them. Key fields like family_id, actor_id, request_text, read_only, classification, expected_scope, and expected_artifacts carry no added meaning anywhere. This is a severe gap for a high-arity tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Prepare internal execution for an admitted assignment.' It also clarifies the notable behavior that Slice reuse happens and no duplication occurs, which helps distinguish it from peers that create/generate artifacts. It leans heavily on internal jargon ('MangoMe orchestration', 'Slice', 'Family') and names no sibling, so an agent gets the gist but must infer the domain model.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use versus a sibling (e.g., start_slice, enter_work, submit_plan). The only conditional is a behavioral expectation ('For audit/review assignments, the worker must still perform the audit and persist Evidence'), which is a completion rule, not selection guidance. No exclusions or alternatives are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
project_overviewC
Explain a complete project's current state across all known families.
| Name | Required | Description | Default |
|---|---|---|---|
| project_ref | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It does not mention whether the tool performs a read-only operation, whether it requires authentication or specific permissions, or whether it has performance implications. It also doesn't describe what happens if the project does not exist or what the output structure looks like.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very short and to the point, with no wasted words. The core purpose is stated in a single sentence. However, its brevity comes at the cost of missing critical information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, a single parameter, and no parameter description, the tool is not adequately described. There is no information about return values or edge cases. While the output schema exists (as indicated by has output schema), the description does not help an agent understand what to expect from a successful call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description says nothing about the 'project_ref' parameter beyond what is already in the input schema (type string). Since schema description coverage is 0%, the description adds no context about how to format the reference, what it represents, or how it should be obtained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'explain' and clearly references a complete project's current state across all known families. However, it does not distinguish this from sibling tools like 'graph', 'status', 'health', or 'workspace_status' which could also provide state or overview information, so some ambiguity remains.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives. There is no mention of when this tool is preferred over 'status' or 'health', nor any exclusions. An agent would have to infer usage from the name and description alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
promote_contract_generationA
Promote a changed contract working copy into a new immutable canonical generation.
This requires the single active generation grant created for the same MODIFY turn, actor and Contract. Merely observing a changed local file never updates canonical contract truth.
| Name | Required | Description | Default |
|---|---|---|---|
| content | Yes | ||
| turn_id | Yes | ||
| actor_id | Yes | ||
| grant_id | Yes | ||
| change_type | No | MODIFY | |
| contract_ref | Yes | ||
| source_checksum | No | ||
| source_storage_system | No | ||
| source_physical_location | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses that this is a mutating operation creating an immutable generation, that a pre-existing grant is required, and that file observation alone does not change canonical truth. It does not cover failure modes, conflicts with concurrent generations, reversibility, or permissions beyond the grant.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action before the grant precondition and the caveat about observation. No filler, though the 'Merely observing...' sentence is a mild restatement of the same point.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation. For a 9-parameter mutating tool with zero annotation coverage, however, the description leaves several parameters and the failure/conflict behavior undocumented, which is a meaningful gap for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 9 parameters, so the description must compensate. It does partially: the phrase 'the same MODIFY turn, actor and Contract' ties together turn_id, actor_id, contract_ref and change_type, and grant_id is named explicitly. But content, source_checksum, source_storage_system and source_physical_location get no meaning at all.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Promote') and resource ('contract working copy') and names the resulting artifact ('new immutable canonical generation'), which is far more precise than the bare tool name. It also distinguishes itself from the passive alternative ('merely observing a changed local file'), though it doesn't name a sibling tool that does the observing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete precondition: the single active generation grant for the same MODIFY turn, actor and Contract must exist. That tells the agent when the call can succeed. It stops short of naming alternatives or stating explicit exclusion cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
publish_worker_runtimeB
Publish a host-observed worker runtime snapshot through the router capability.
This is not a worker self-report. Model identity, runtime mode, capabilities and cost are separate facts; hosts should refresh the snapshot when provider/tool availability changes (for example a degraded/reserve mode).
| Name | Required | Description | Default |
|---|---|---|---|
| active | No | ||
| metadata | No | ||
| model_id | No | ||
| cost_class | No | STANDARD | |
| worker_key | Yes | ||
| owner_gated | No | ||
| capabilities | No | ||
| router_token | No | ||
| runtime_mode | No | NORMAL | |
| router_actor_id | Yes | ||
| max_parallel_tasks | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It does add meaningful behavior: the snapshot is host-observed, not self-reported, and should be refreshed on availability changes. However, it does not disclose side effects such as whether this is an upsert, whether prior snapshots are replaced, or what authorization/state implications 'through the router capability' entails.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is brief and front-loaded: the first sentence states the action and object, and the following sentences add essential nuance without fluff. Every sentence contributes meaningful context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite the output schema existing, the tool has 11 parameters, no annotations, and no parameter descriptions. The description gives a good high-level scenario but leaves crucial invocation details uncovered, such as what router_actor_id and router_token represent, what 'through the router capability' implies, and the effect of the active flag.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only maps a few concepts loosely: model identity, runtime mode, capabilities, and cost align with model_id, runtime_mode, capabilities, and cost_class. Required parameters like worker_key and router_actor_id, plus active, owner_gated, router_token, max_parallel_tasks, and metadata, remain semantically unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource ('Publish a host-observed worker runtime snapshot') and adds a key clarifying distinction: it is not a worker self-report. This makes the tool's purpose reasonably clear, though it does not explicitly distinguish itself from sibling tools beyond that contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage context: hosts should publish/refresh the snapshot when provider/tool availability changes, such as degraded/reserve mode. It also gives a clear when-not ('This is not a worker self-report'), but it does not name alternative sibling tools or explicitly describe when another tool should be used instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_contextC
Read all relevant family context. Reading is unrestricted.
| Name | Required | Description | Default |
|---|---|---|---|
| family_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It only states that reading is unrestricted, which is a minor behavioral trait. It omits side effects, permissions, data scope, or any potential non-obvious behavior such as filtering, caching, or aggregation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely short and front-loaded, with the core verb and object in the first sentence. No filler words exist. However, the brevity borders on under-specification, sacrificing necessary detail for compactness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value format is presumably covered. But the description does not explain the scope of 'relevant' context, how family_id filters the data, or what distinguishes this tool from close siblings. For a read tool with one parameter, this is inadequate for reliable selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the only parameter 'family_id' is not mentioned in the description at all. The description adds no meaning beyond the bare schema, leaving the agent without any understanding of what family_id should contain or how it affects the retrieved context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear action ('Read') and a resource ('all relevant family context'), which satisfies the basic requirement. However, it does not distinguish itself from sibling tools like recovery_context or effective_family_view, and 'all relevant' is vague and could apply to several other context-reading tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Reading is unrestricted' implies there are no access restrictions, which is a small usage signal. But there is no guidance on when to choose this tool over alternatives, no exclusions, and no context about prerequisites or typical invocation scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reconcile_assignmentC
Bridge free worker reasoning into governed MangoMe work without creating truth.
The worker may inspect/reason and form a tentative decomposition before this call.
This operation is read-only with respect to canonical work identity and normative
truth: it resolves the managed workspace's canonical restore state and tells the
worker which governed transition is appropriate before productive effect. For an
elliptical follow-up maintenance request, target may carry the target already
resolved from the current conversation (for example MangoMe) without searching
historical host memory.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | ||
| request_text | Yes | ||
| workspace_root | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses that the operation is 'read-only with respect to canonical work identity and normative truth' and that it resolves state before effect. But it stops short of clarifying the tool's own side effects (it is a 'reconcile'/'bridge' operation) or its permission/rate/return behavior, leaving meaningful behavioral gaps for a governance-critical tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences with little waste, but the front-loaded statement is abstract jargon before concrete information arrives. The elliptical-follow-up detail is useful but the overall prose is heavy and slow to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. However, for a complex governance operation with zero annotations and 0% parameter coverage, the description covers the conceptual 'why' but leaves the two undocumented parameters and the tool's own mutation profile unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains only `target` (carrying a resolved value like `MangoMe` without searching historical host memory). The required `request_text` and `workspace_root` are not addressed at all, leaving two of three parameters undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific role for the tool ('resolves the managed workspace's canonical restore state and tells the worker which governed transition is appropriate before productive effect'), which is more than a tautology. However, it is wrapped in dense domain jargon ('MangoMe work', 'canonical work identity', 'normative truth') and never cleanly distinguishes itself from sibling tools like prepare_assignment, begin_work, or submit_plan. An agent would struggle to tell precisely what class of action this performs versus its neighbors.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies usage ('The worker may inspect/reason and form a tentative decomposition before this call') and gives a specific case ('For an elliptical follow-up maintenance request, `target` may carry the target already resolved...'). But it never states explicit when-to-use / when-not-to-use or names an alternative sibling, leaving the routing decision to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reconcile_bigbangC
Reconcile candidate discovery only before work admission; never rebuild admitted state from discovered records.
| Name | Required | Description | Default |
|---|---|---|---|
| records | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the burden of behavioral disclosure. It states a critical safety constraint ('never rebuild admitted state from discovered records') which is valuable, but it does not disclose what the tool does with the records (e.g., mutates? reads? creates? side effects) or what happens on invocation. This is insufficient for a tool with such an opaque name.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very brief (one sentence) and front-loads the key constraint. It is concise and has no fluff, but the conciseness sacrifices necessary detail. Still, for what it says, it is efficient, so it earns a 4 for structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having an output schema, the description leaves too much unsaid. The tool name is cryptic, and the description does not clarify the purpose, the meaning of 'records', or the expected behavior. With one required parameter and no parameter details, the description is far from complete for an agent to use it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain the 'records' parameter at all. The only parameter is 'records' (an array of objects), but the description does not specify what these records are, their structure, or how they relate to 'candidate discovery'. Since coverage is 0%, the description should compensate but fails to do so.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description says 'Reconcile candidate discovery only before work admission; never rebuild admitted state from discovered records.' It is unclear what 'reconcile' does operationally — does it merge, compare, update? The verb is vague and there is no explicit resource beyond 'candidate discovery'. It hints at a distinction but does not state a clear purpose that would help an agent understand what action to take.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is a partial usage guideline: 'only before work admission' suggests a timing constraint, but there is no mention of when NOT to use it or what alternatives exist among siblings (e.g., bigbang_scan, discover_scopes). No exclusions are given, and an agent cannot infer which other tool to use in other contexts.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_audit_findingC
Persist one scoped audit finding; EXPAND follows the configured graph frontier without expanding mutation authority.
| Name | Required | Description | Default |
|---|---|---|---|
| impact | No | NONE | |
| summary | Yes | ||
| actor_id | Yes | ||
| audit_id | Yes | ||
| unknowns | No | ||
| subject_id | No | ||
| assumptions | No | ||
| subject_ref | No | ||
| affected_ids | No | ||
| evidence_ids | No | ||
| affected_refs | No | ||
| finding_class | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It hints at a safety boundary ('without expanding mutation authority'), which is a useful behavioral trait, but says nothing about persistence semantics, idempotency, deduplication, or what 'EXPAND' actually triggers.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact clauses with the core action front-loaded and no padding. The second clause is dense but not wasteful in length, even if its meaning is unclear.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values needn't be described, but with 12 undocumented parameters, no annotations, and an unexplained 'EXPAND' behavior, the definition leaves too much unspecified for an agent to invoke confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are 12 parameters with 0% schema description coverage, so the description must compensate and largely does not. Only 'EXPAND' vaguely gestures at a finding_class value; required params like audit_id, actor_id, summary, and finding_class are left entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (persist) and resource (audit finding) with a scope qualifier, which distinguishes it from siblings like start_scoped_audit and close_scoped_audit. However, the second clause introducing 'EXPAND' is opaque and doesn't clarify the tool's core purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance is provided, nor any indication of when to prefer this over start_scoped_audit or submit_verification_observation. The 'EXPAND' clause reads as cryptic jargon rather than actionable usage direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_execution_receiptC
Record model/agent cost and durable outcome economics, including optional UAI/1 transport telemetry.
| Name | Required | Description | Default |
|---|---|---|---|
| outcome | No | UNKNOWN | |
| actor_id | Yes | ||
| currency | No | EUR | |
| metadata | No | ||
| model_id | No | ||
| slice_id | Yes | ||
| family_id | Yes | ||
| human_cost | No | ||
| work_class | No | UNCLASSIFIED | |
| repair_cost | No | ||
| input_tokens | No | ||
| output_tokens | No | ||
| execution_cost | No | ||
| verification_cost | No | ||
| context_tokens_raw | No | ||
| interlingua_version | No | ||
| context_tokens_compiled | No | ||
| output_tokens_interlingua | No | ||
| context_tokens_interlingua | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden for behavioral disclosure. It implies a write operation but does not state side effects, reversibility, authentication needs, or what happens on successful record. Even the optional telemetry is not elaborated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no wasted words, and the core action is front-loaded. However, it is so terse that valuable context is omitted, so it is not efficiently informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (19 parameters, zero schema descriptions, no annotations) and the fact an output schema exists, the description is severely incomplete. It fails to explain required parameters, the meaning of 'durable outcome economics', or how the telemetry fields relate, leaving an agent unable to correctly populate the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It only vaguely references 'model/agent cost' and 'UAI/1 transport telemetry', but does not explain any of the 19 parameters, including required ones like family_id, slice_id, actor_id, or the many cost and token fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool records cost and outcome economics, and mentions optional telemetry. It names a specific verb and resource (recording receipt) and distinguishes it from general tools, though it does not explicitly differentiate from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. It does not mention prerequisites, typical scenarios, or situations where a different tool should be preferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_fast_judgmentC
Persist an FJD/1 result as PROGRESSIVE WORKER_JUDGMENT bound to WorkIdentity.
| Name | Required | Description | Default |
|---|---|---|---|
| source | No | HOST | |
| purpose | Yes | ||
| work_id | Yes | ||
| actor_id | Yes | ||
| decisions | Yes | ||
| model_ref | No | ||
| high_impact | No | ||
| context_refs | No | ||
| review_confidence | No | ||
| use_signal_confidence | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It indicates a write/persist operation and binding to WorkIdentity, but omits permissions, idempotency, side effects, handling of high_impact or confidence fields, and what happens on repeated calls.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The single sentence is front-loaded and free of filler, but it is under-specified for a 10-parameter mutation tool. Brevity here reflects missing information rather than efficient communication.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 10 parameters, no annotations, 0% schema description coverage, and a complex domain operation, the description is far too thin. The output schema reduces the need to explain return values, but usage, parameter meanings, and behavioral constraints are almost entirely absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 10 parameters, and the description does not explain required fields such as purpose, actor_id, or decisions, nor optional fields like source, model_ref, context_refs, review_confidence, or use_signal_confidence. It only faintly connects 'WorkIdentity' to work_id and 'FJD/1 result' to decisions, leaving nearly all parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Persist') and resource ('an FJD/1 result as PROGRESSIVE WORKER_JUDGMENT bound to WorkIdentity'), so the core operation is clear. However, it relies on domain jargon and does not differentiate itself from siblings like assess_fast_judgment or fast_judgment_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus alternatives. It implies persistence after an FJD/1 result exists, but never states required preconditions, sequence, or when-not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_truth_assertionB
Record one canonical BTTM/1 assertion under a VERIFY or MODIFY work turn.
valid_* describes represented-world validity; known_* is MangoMe transaction time. Worker prose/confidence is never sufficient grounding: an assertion needs Evidence or support/assumption/dependency links before BTTM can mark it SUPPORTED.
| Name | Required | Description | Default |
|---|---|---|---|
| value | Yes | ||
| turn_id | Yes | ||
| actor_id | Yes | ||
| valid_to | No | ||
| work_ref | Yes | ||
| predicate | Yes | ||
| depends_on | No | ||
| known_from | No | ||
| source_ref | No | ||
| subject_id | Yes | ||
| valid_from | No | ||
| contradicts | No | ||
| support_ids | No | ||
| evidence_ids | No | ||
| assertion_key | No | ||
| assumption_ids | No | ||
| supersedes_assertion_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the full burden. It usefully discloses bitemporal semantics (valid_* = represented-world validity, known_* = transaction time) and a genuine domain rule that prose/confidence alone cannot ground an assertion and Evidence or link IDs are required before it is marked SUPPORTED. It does not say what happens when that requirement is unmet (rejected, pending, error), nor anything about idempotency or permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, front-loaded with the action and scope, then the parameter-family semantics and the grounding rule. No filler, though the unexplained jargon reduces the value of the terseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 17-parameter mutation tool with no annotations and no schema descriptions, the description covers the central domain rule (grounding requirement) but omits usage routing, most parameter meanings, and failure behavior. An output schema exists, so return values need not be explained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 17 parameters, so the description must compensate. It clarifies the valid_*/known_* parameter families and ties evidence/support/assumption/dependency links to the grounding rule, but leaves work_ref, subject_id, predicate, actor_id, turn_id, assertion_key, contradicts and supersedes_assertion_id entirely unexplained, and 'value' has no type at all.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Record one canonical BTTM/1 assertion') and scopes it to a VERIFY or MODIFY work turn, which separates it from siblings like invalidate_truth_assertion or truth_assertion_status. The domain vocabulary (BTTM/1, MangoMe transaction time) is never defined, but the action itself is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use or when-not-to-use guidance, and no alternatives are named despite many truth-related siblings (truth_at, truth_assertion_status, invalidate_truth_assertion). The only contextual constraint is that the assertion sits under a VERIFY/MODIFY work turn, which is implied rather than explained.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recovery_contextC
Return compact authoritative recovery state; never derive admitted work state from path discovery.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace_root | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It adds some useful context by labeling the state 'authoritative' and warning against path-derived inference, but it does not disclose side effects, permissions, idempotency, or return behavior. This is too thin to be considered transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence with no filler. The primary action is front-loaded, and the behavioral caution is appended without bloating the length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although the tool signature is simple and an output schema exists, the description leaves gaps around when to use this tool among many recovery/context siblings and what workspace_root does. The cryptic 'never derive' clause is not enough to make the definition complete for agent selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 0%, and the description does not mention workspace_root at all, nor its role in scoping recovery state. The parameter's name is somewhat self-explanatory, but with such low coverage the description needed to compensate and did not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb and resource: 'Return compact authoritative recovery state.' The additional clause adds precision about what the tool does not do, which helps distinguish it from path-derived state inference. However, it does not explicitly name or differentiate among the many related sibling tools like read_context or session_restore.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance. The warning 'never derive admitted work state from path discovery' is a semantic constraint, not a usage guideline, and no alternatives are mentioned. An agent must infer when recovery_context is preferable to session_restore, workspace_status, or read_context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
refresh_viewsB
Refresh all deterministic family views without LLM use.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must fully disclose behavioral traits. It only states the action and constraint, but fails to explain side effects, whether the refresh is destructive, how it interacts with existing data, or what the response contains. The term 'refresh' implies mutation but consequences are unspecified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no fluff, front-loading the action and constraint. It is efficient, though it could add a bit more behavioral context without losing brevity. Overall, well-structured for its minimal content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless tool with an output schema, the description still needs to provide usage context and behavioral expectations. It lacks explanation of what constitutes a 'family view', what 'deterministic' means in this system, and what the refresh operation entails. Given the large sibling set, an agent would benefit from more context on when this refresh is appropriate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema fully covers inputs trivially. The description appropriately avoids redundant parameter details. Baseline of 4 applies for parameterless tools, and no additional information is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description specifies the action (refresh), the resource (all deterministic family views), and a key constraint (without LLM use). It clearly states what the tool does and distinguishes it from siblings like effective_family_view, which likely reads rather than refreshes. However, it does not explicitly name alternatives or contrast with them, so it misses a point for differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use: when needing to refresh deterministic family views without LLM involvement. It gives a condition but does not explicitly state when not to use it or mention alternative tools. An agent must infer that non-deterministic or LLM-dependent views belong to other tools, leaving room for ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
register_contractC
Append normative Contract truth; work-bound families require a MODIFY WorkTurn.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | BASE | |
| title | Yes | ||
| turn_id | No | ||
| work_id | No | ||
| actor_id | No | ||
| checksum | No | ||
| family_id | Yes | ||
| declared_id | Yes | ||
| storage_system | No | ||
| controller_token | No | ||
| physical_location | No | ||
| controller_actor_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses that work-bound families need a MODIFY WorkTurn (a genuine prerequisite), but is silent on permissions, reversibility, what 'append' mutates, or error behavior for a 12-parameter mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence led by the action verb, with the conditional clause placed after. No wasted words, though the extreme terseness borders on cryptic for a tool of this complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values needn't be described, but a 12-param, unannotated mutation tool with the MODIFY WorkTurn dependency only partially explained leaves substantial gaps. The description is under-specified relative to the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 12 parameters, so the description must compensate and largely does not. The words 'families' and 'WorkTurn' loosely gesture at family_id and turn_id/work_id, but declared_id, kind, checksum, controller_actor_id and the rest receive no semantic explanation anywhere.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"Append normative Contract truth" gives a specific verb (append/register) and resource (Contract), distinguishing it from read-only siblings like contract_state. The phrasing is domain-heavy jargon, but an agent can infer it creates/registers an authoritative contract record.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It supplies one real usage rule: work-bound families require a MODIFY WorkTurn. That is a useful conditional, but there is no guidance on when to choose this over sibling writers like bind_contract_turn, register_truth_assertion, or register_playbook.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
register_modelC
Register a model/access-path identity for empirical execution comparisons.
| Name | Required | Description | Default |
|---|---|---|---|
| metadata | No | ||
| provider | No | ||
| model_key | Yes | ||
| access_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It only says 'register,' which implies a write operation, but it does not state side effects, idempotency, permissions, reversibility, or what output to expect. This is a serious gap for an action that likely creates a persistent record.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence front-loaded with the verb and object. It wastes no words and is easy to parse. However, its brevity borders on under-specification, so it does not earn a 5 for structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 4 parameters, no annotations, and an implied mutation, this description is far too thin. It lacks usage context, parameter explanations, behavioral details, and any sense of what the tool accomplishes beyond a vague registration. An agent cannot confidently call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the 4 parameters (model_key, access_path, provider, metadata). The phrase 'model/access-path identity' only hints at two of them, and provider and metadata are left entirely unexplained. No parameter-level meaning is added.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('register') and names its resource ('model/access-path identity') with a purpose clause ('for empirical execution comparisons'). This distinguishes it from sibling registration tools like register_contract, though it does not explicitly name alternatives. It is clear but could be more concrete about what 'register' entails.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus alternatives, no prerequisites, and no exclusions. The phrase 'for empirical execution comparisons' hints at a use case, but it does not tell the agent when to prefer this over register_contract or other siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
register_playbookC
Register procedural Playbook metadata. Playbooks never become normative project truth.
| Name | Required | Description | Default |
|---|---|---|---|
| source | Yes | ||
| version | Yes | ||
| metadata | No | ||
| description | Yes | ||
| content_hash | No | ||
| playbook_key | Yes | ||
| applicability | No | ||
| required_capabilities | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden. It discloses one useful semantic trait (playbooks are non-normative), but says nothing about permissions, mutation effects, idempotency, or what happens on duplicate keys for what is clearly a registration/mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the core purpose front-loaded and zero wasted words. It is efficient, though arguably under-specified rather than over-long.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, but the description is inadequate for an 8-parameter mutation tool with no annotations and zero parameter documentation in the schema. Essential invocation context is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 8 parameters (4 required), and the description adds no meaning for any of them. For a tool with playbook_key, version, description, source, applicability, and required_capabilities, this leaves every parameter's semantics undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Register procedural Playbook metadata'), which is clear and distinguishable from sibling select_playbook. The second sentence adds a scope constraint, but there is no explicit sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this versus alternatives such as select_playbook, nor any prerequisites or exclusions. The note that playbooks 'never become normative project truth' hints at scope but does not tell the agent when to invoke the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reject_overrideD
Reject using the owner runtime capability.
| Name | Required | Description | Default |
|---|---|---|---|
| decided_by | Yes | ||
| approval_id | Yes | ||
| decision_ref | No | ||
| approval_token | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It does not state whether the action is destructive, requires authentication, has side effects, or what happens on success/failure. This is a critical gap for a mutating operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely short, but this is under-specification rather than conciseness. It lacks structure and does not earn its place by conveying useful information; it is essentially a tautology of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 4 parameters, a non-trivial purpose, and an output schema, the description is wholly inadequate. It provides no context on inputs, outputs, or expected behavior, making it impossible for an agent to use the tool correctly without additional information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, meaning the parameters (approval_id, decided_by, decision_ref, approval_token) have no descriptions. The description does not compensate by explaining any of these fields or their roles, leaving the agent to guess at their meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Reject using the owner runtime capability' is vague. It identifies the verb 'reject' but the resource and context are unclear – it does not specify what is being rejected (e.g., an override request) or what 'owner runtime capability' means. It fails to differentiate from sibling tools like approve_override or request_override.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool, what prerequisites exist, or how it differs from alternatives. There is no mention of appropriate scenarios or when to prefer this over approve_override or list_approvals.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
release_contract_generation_grantB
Release an unused contract-generation write grant owned by the current actor.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | ||
| actor_id | Yes | ||
| grant_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden. It does disclose two useful facts: the grant must be unused and must be owned by the current actor, which hints at failure conditions. However it says nothing about permissions, reversibility, or whether the release is destructive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the key scope constraint ('owned by the current actor') is placed at the end but the sentence is short enough that this is not a readability problem.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a state-mutating grant operation with zero annotation coverage and zero schema parameter descriptions, the description is too thin: it omits preconditions beyond 'unused', permission requirements, and error behavior. The existence of an output schema does excuse it from describing return values, but the input-side gaps remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for three parameters, so the description must compensate. It only clarifies that actor_id is constrained to the calling actor; grant_id and reason remain entirely undocumented, and 'reason' is not even hinted at.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description pairs a specific verb ('release') with a specific resource ('contract-generation write grant') and a scope qualifier ('owned by the current actor'). It is distinguishable from the nearby sibling promote_contract_generation, though it never names that sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The word 'unused' implies the trigger condition (a grant that was issued but never consumed), but the description never states when to call this versus promote_contract_generation, nor what happens if the grant has already been used. Usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
render_uai_resultB
Render a structured UAI/1R worker result into deterministic human-readable English or German.
| Name | Required | Description | Default |
|---|---|---|---|
| language | No | en | |
| result_json | Yes | ||
| expected_context_hash | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions 'deterministic' and the output language options, which are useful, but it does not disclose error behavior, the role of expected_context_hash, or any side effects. The description adds some behavioral context but leaves significant gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear sentence that directly states purpose and output languages. It is concise and front-loaded with the primary action, but it is so brief that it omits crucial operational details. It is appropriately sized but slightly under-specified.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although an output schema exists, the description lacks parameter semantics, usage guidance, and behavioral caveats, making it incomplete for an agent to correctly invoke the tool. The description does not explain the required context hash or how language selection works, so an agent would have to infer or experiment.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the parameters. It does not mention result_json, expected_context_hash, or the allowed language values beyond the default 'en'. The description provides no semantic information about the parameters, leaving agents without guidance on what to supply.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (render), the resource (structured UAI/1R worker result), and the outcome (deterministic human-readable English or German). It distinguishes this from siblings like decode_uai_result by emphasizing human-readable output, making the tool's purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use when a human-readable representation is needed, but it does not explicitly contrast with sibling tools such as decode_uai_result or compile_uai_context, nor does it state when not to use it. The usage context is implied but not explicitly guided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
repository_locationsC
Return observed physical Git checkout/worktree locations without inferring project truth.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace_root | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It discloses a key trait: it returns observed data and does not infer project truth. However, it does not mention read-only nature, permissions, side effects, or the return format (though an output schema exists). The single behavioral note adds value but is not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence that front-loads the primary action and includes a meaningful qualifier. It has no filler or redundancy, making it highly efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with one optional parameter and an output schema, which reduces the need for extensive description. However, the lack of parameter explanation and usage guidance leaves gaps. The core purpose is clear, but an agent might not know how to effectively call it without more context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description does not mention the 'workspace_root' parameter at all. Since schema description coverage is 0%, the description should compensate, but it fails to explain what this parameter controls or how it affects the returned locations. The agent is left to guess its semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: 'Return observed physical Git checkout/worktree locations'. It specifies the resource and adds a qualifier 'without inferring project truth' that clarifies the nature of the data. It does not explicitly name sibling tools, but the purpose is distinct enough from generic filesystem or status tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. The qualifier 'without inferring project truth' hints at a use case for raw data, but there is no explicit when-to-use or when-not-to-use, nor any mention of sibling tools or conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
request_overrideC
Create an explicit approval request; requesting approval does not grant it.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | ||
| subject_id | Yes | ||
| action_type | Yes | ||
| requested_by | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It does reveal a key behavioral trait (requesting does not grant approval) but provides no other side effects, authorization requirements, or process details. The single caveat is helpful but insufficient for a tool that creates an implicit record.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, efficient and front-loaded with the primary action. It includes a useful caveat without unnecessary elaboration. However, it is so brief that it falls short on essential information, though that's more an issue of completeness than conciseness itself.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a 4-parameter required schema and no annotations, the description is far from complete. It fails to explain the parameters, the outcome, or any considerations for calling the tool. An agent would lack critical information to correctly invoke it, especially without a clear output schema understanding.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it mentions none of the four required parameters (reason, subject_id, action_type, requested_by). The agent receives no guidance on what these parameters mean, their formats, or how to populate them correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb and resource: 'Create an explicit approval request'. It distinguishes itself from siblings like approve_override and reject_override by explicitly noting that requesting approval does not grant it, making its purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context by clarifying that requesting does not grant approval, which hints that this tool is for creating requests while approval/rejection are handled elsewhere. However, it doesn't explicitly say when to use this tool vs approve_override or reject_override, nor mention any prerequisites or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resolveC
Resolve a family, declared contract id, or slice id without semantic guessing.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the behavioral burden. It does not disclose whether resolution is read-only, what happens on missing or ambiguous IDs, or whether any side effects occur. 'Without semantic guessing' is a weak behavioral hint but does not convey operational semantics like error handling or response structure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler words. It is concise and directly addresses the core operation, though it is so terse that it sacrifices explanatory power.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although the tool has only one parameter and an output schema, the description leaves critical gaps: no usage context, no behavior guarantees, and minimal parameter semantics. An agent cannot confidently decide when to call it or what to expect beyond a vague 'resolve' action.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema describes 'query' only as a string with zero coverage. The description clarifies that the query can be a family, declared contract id, or slice id, but does not explain expected formats, examples, or how the tool distinguishes among the three. This partial compensation is insufficient for a 0%-coverage parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('resolve') and resource types ('family, declared contract id, or slice id'), which distinguishes it from sibling tools that create, register, or scan. The phrase 'without semantic guessing' adds a hint of exact matching, but does not name alternatives explicitly, so it stops short of full differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus siblings like read_context, status, or graph. It does not state prerequisites, exclusions, or scenarios where another tool would be more appropriate, leaving the agent to infer usage from the mnemonic verb alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
select_playbookC
Select a replaceable procedural Playbook for a WorkIdentity; selection is PROGRESSIVE/non-normative.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | ||
| turn_id | No | ||
| actor_id | Yes | ||
| work_ref | Yes | ||
| playbook_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden. It mentions 'replaceable' and 'PROGRESSIVE/non-normative', but does not say whether selection mutates state, replaces an existing playbook, requires permissions, or can be undone—leaving core behavioral effects unexplained.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One compact sentence that is front-loaded with the action and resource. The second clause is terse but adds a small behavioral qualifier; overall there is no padding, though the all-caps qualifier is cryptic rather than clarifying.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter selection tool with no annotations and no structured parameter descriptions, the description is too thin. An output schema may cover return values, but prerequisites, side effects, and parameter roles are left entirely to inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for 5 parameters, and the description explains none of them. It alludes to 'Playbook' and 'WorkIdentity' but does not map work_ref, playbook_id, actor_id, reason, or turn_id to meaning, formats, or required relationships.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Select') and a distinct resource ('procedural Playbook') scoped to a WorkIdentity. Sibling differentiation is implied by contrast with register_playbook, but not made explicit, and the phrasing is domain jargon rather than plain language.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use condition, prerequisites, or alternative-tool routing is given. The clause 'selection is PROGRESSIVE/non-normative' hints at a mode of selection but does not tell an agent when to choose this over register_playbook or other workflow tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_bootstrapB
Bounded read-only bootstrap: bind cheaply, then read canonical restore state.
Bootstrap never runs filesystem inventory, Big-Bang discovery, or repository archaeology. Those remain explicit onboarding/maintenance operations.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace_root | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it does disclose useful traits: read-only, bounded, and deliberately scoped away from heavy inventory work. It still omits permissions/auth needs, idempotency, and what 'bind' actually changes (if anything), so coverage is partial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, with the core scope statement front-loaded before the negative-scope caveat. The only drag is the jargon-heavy phrasing, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the negative scope clarifies the tool's boundary. But the lone undocumented parameter and the vague 'bind cheaply' framing leave an agent without enough to invoke this confidently in a large, dense sibling set.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one parameter, workspace_root, with 0% schema description coverage, and the description never mentions it. An agent gets no guidance on what the root should be, whether it is optional (schema default null), or what happens when omitted, so the description fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states it is a 'bounded read-only bootstrap' that 'bind[s] cheaply' and reads 'canonical restore state', giving a rough sense of a lightweight startup/restore operation. However, 'bind cheaply' and 'canonical restore state' are jargon that leaves the concrete output and action ambiguous, and while it distinguishes itself from discovery/archaeology work, it never clearly names what it produces. Purpose is vague-but-directional rather than specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives exclusion guidance by naming what bootstrap does not do (filesystem inventory, Big-Bang discovery, repository archaeology) and routing those to onboarding/maintenance. That implies a when-to-use boundary, but it never names a concrete sibling (e.g. session_restore, bigbang_scan, filesystem_scan) as the alternative, so the routing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_restoreB
Restore canonical session/work state without discovery or state creation.
| Name | Required | Description | Default |
|---|---|---|---|
| workspace_root | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose a meaningful behavioral trait: it will not create state or perform discovery, so an agent knows this is a non-mutating rehydration. It omits auth/permission requirements, what happens when no prior state exists, and any side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the core action and its boundary condition are stated up front. It is appropriately terse, if slightly elliptical.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema covers return values and the tool is low-complexity with one optional param, so the description needn't explain returns. Still, with no annotations and an undocumented param, it leaves gaps around the workspace_root effect and the failure mode when no canonical state exists.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter workspace_root has 0% schema description coverage and is never mentioned in the description, so its meaning, format, or effect on restore scope is left entirely to inference from the name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Restore) and resource (canonical session/work state), and the negation 'without discovery or state creation' implicitly differentiates it from creation-oriented siblings like session_bootstrap. It is clear what the tool does, though it does not name a sibling to route against.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'without discovery or state creation' implies the context of use (rehydrating an existing session rather than bootstrapping one), which is genuine implied guidance. However, it never explicitly names an alternative or states the prerequisite condition for choosing this over session_bootstrap or recovery_context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_gateB
Set a gate. PASS requires attested PASS evidence; WAIVED requires approved owner decision.
| Name | Required | Description | Default |
|---|---|---|---|
| status | Yes | ||
| gate_id | Yes | ||
| turn_id | No | ||
| actor_id | No | ||
| slice_id | Yes | ||
| approval_id | No | ||
| evidence_ids | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden, and it does disclose real prerequisites: PASS needs attested evidence and WAIVED needs an owner decision. It says nothing about mutation effects, permission requirements, idempotency, or failure behavior, leaving most behavioral questions unanswered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and then the status constraints; nothing is wasted. It is terse to the point of under-specification, but that is a completeness issue rather than a verbosity one.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter mutation tool with zero schema description coverage and no annotations, the description is far too thin: it documents neither the parameters nor the relationship to set_gate_controlled. The output schema spares it from explaining return values, but that is the only burden lifted.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 7 parameters, so the description must compensate and largely does not. "Attested evidence" loosely gestures at evidence_ids and "approved owner decision" at approval_id, but slice_id, gate_id, turn_id, and actor_id are entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"Set a gate" is a clear verb plus resource, and the second sentence narrows what setting a gate means by tying it to status semantics. It does not, however, distinguish itself from the sibling tool set_gate_controlled, which an agent might reasonably consider instead.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"PASS requires attested PASS evidence; WAIVED requires approved owner decision" gives conditional prerequisites for two statuses, which is genuine usage guidance. It never says when to prefer set_gate over set_gate_controlled, or what scenarios select each status, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_gate_controlledD
Compatibility alias for v0.1.1 controlled gate updates.
| Name | Required | Description | Default |
|---|---|---|---|
| status | Yes | ||
| gate_id | Yes | ||
| actor_id | Yes | ||
| slice_id | Yes | ||
| approval_id | No | ||
| evidence_ids | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure, but it only mentions 'controlled gate updates' without detailing side effects, required permissions, reversibility, or response behavior. The word 'updates' implies mutation, yet nothing else about the operation's behavior is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded with the compatibility-alias framing, so it is concise. However, the brevity is closer to under-specification than effective compression, since the single sentence omits nearly all operationally relevant content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a six-parameter tool with an output schema, but the description provides only an alias label and a vague reference to gate updates. Despite the output schema existing, the agent still lacks enough information about purpose, parameter semantics, or behavior to invoke the tool confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate by explaining key parameters like status, approval_id, evidence_ids, and actor_id. It does not mention any parameter semantics at all, leaving the agent to guess what values are valid and how the parameters relate to the update operation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies the tool only as a 'Compatibility alias for v0.1.1 controlled gate updates,' which does not directly state what action the tool performs or what effect it has. An agent must infer that it updates a gate's controlled status, and there is no explanation of what 'controlled' means or how this differs from the sibling set_gate tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'compatibility alias' provides a weak contextual signal that this is for backward compatibility with v0.1.1, but there is no explicit guidance on when to use this tool versus set_gate or any other alternative. No conditions, exclusions, or migration advice are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_scoped_auditC
Start SRA/1 bounded recursive audit from a local scope. Scope may expand; mutation authority never does.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | READ_ONLY | |
| plan_id | No | ||
| turn_id | No | ||
| actor_id | Yes | ||
| slice_id | No | ||
| family_id | Yes | ||
| max_depth | No | ||
| objective | Yes | ||
| audit_kind | No | GENERAL | |
| target_ids | No | ||
| assumptions | No | ||
| max_objects | No | ||
| target_refs | No | ||
| mutation_scope_ids | No | ||
| expansion_relations | No | ||
| mutation_scope_refs | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose one genuinely useful behavioral guarantee – 'Scope may expand; mutation authority never does' – which is non-obvious and not derivable from the schema. But it omits auth requirements, what state gets created, and the effect of the many mutation_scope_* and mode parameters.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no waste and the key behavioral caveat front-loaded in the second clause. However, given 16 undocumented parameters and no annotations, the extreme terseness is under-specification rather than effective conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need no explanation, but the description is far too thin for a 16-parameter, no-annotation mutation-capable tool. Required parameters (family_id, actor_id, objective) go unexplained and the mode default of READ_ONLY is never mentioned, leaving the agent unable to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 16 parameters with cryptic names like mode, audit_kind, max_depth, expansion_relations, and mutation_scope_ids/refs. The description explains none of them, so an agent has no way to know what values mode accepts or what mutation authority entails. This is a severe gap for a 16-param tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the verb 'Start' and the resource 'SRA/1 bounded recursive audit from a local scope,' which conveys that this begins an audit run. However, 'SRA/1' is undefined jargon and the description does nothing to distinguish it from siblings like audit_context, close_scoped_audit, or record_audit_finding, leaving the agent to infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use, when-not-to-use, or alternative-tool guidance. The 'local scope' phrasing hints at context but never states prerequisites or when a sibling like close_scoped_audit or audit_status would be preferred. The agent gets no routing help.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_sliceC
Start a slice and bind it to the actor's persisted active plan.
| Name | Required | Description | Default |
|---|---|---|---|
| plan_id | Yes | ||
| actor_id | Yes | ||
| slice_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description must disclose behavior itself. It does reveal one side effect—binding a slice to a persisted active plan—but it does not explain whether this mutates existing plan state, whether it is idempotent, what happens if no active plan exists, or what the operation returns. For a state-changing tool, this is a meaningful gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is brief and front-loaded with the core action, but 'start' + 'bind' is the minimum viable phrasing. It earns conciseness but the lack of any structured context, examples, or caveats makes it feel under-specified rather than efficiently complete.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given three required parameters TWZ, no annotations, and several closely related sibling tools, this description is too thin for reliable invocation. It omits the expected state of the plan, failure semantics, and how 'start' differs from 'accept', 'begin', or 'enter' work, so an agent would likely need additional tool exploration or trial and error.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only loosely implies the roles of slice_id, actor_id, and plan_id. It never explicitly maps each parameter to its meaning or explains how they relate, leaving the agent to infer which ID is the actor, the plan, or the slice.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies a specific action ('start a slice') and a distinct binding responsibility ('bind it to the actor's persisted active plan'), which helps separate this from siblings like verify_slice or accept_slice. It is not a tautology and names the primary resource and effect, though 'start' remains somewhat generic without more detail.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit guidance about when to use this tool over closely related siblings such as begin_work, enter_work, accept_slice, or verify_slice. It also does not state prerequisites like the existence of an active plan or whether the slice must already be accepted or verified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
statusC
Return deterministic materialized state for one family.
| Name | Required | Description | Default |
|---|---|---|---|
| family_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full behavioral burden. It states the result is 'deterministic' and 'materialized,' which adds some context, but it does not disclose whether the operation is read-only, requires authentication, or has side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no filler. It front-loads the action and scope, making it easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although an output schema exists and there is only one parameter, the description omits usage context, behavioral expectations, and definition of key terms like 'materialized state.' An agent would likely struggle to distinguish this from related family-state tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description only says 'for one family,' adding minimal meaning to family_id. It does not explain how to obtain or format the family_id, nor what 'materialized state' includes.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Return deterministic materialized state for one family.' It clearly identifies the operation and scope. However, it does not differentiate from overlapping siblings like effective_family_view or read_context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It does not mention exclusions, prerequisites, or relationships to sibling tools such as effective_family_view or workspace_status.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
submit_evidenceC
Persist evidence. New evidence is UNATTESTED until a trusted verifier/owner attests it.
| Name | Required | Description | Default |
|---|---|---|---|
| result | No | ||
| source | Yes | ||
| payload | No | ||
| actor_id | No | ||
| subject_id | Yes | ||
| artifact_id | No | ||
| evidence_type | Yes | ||
| evidence_class | No | CLAIM |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose a meaningful behavioral trait: new evidence remains UNATTESTED until a trusted verifier/owner attests it. However, it does not mention persistence side effects, authorization requirements, failure modes, or whether submission is idempotent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The definition is short and free of filler, and the attestation note is useful. However, it is so minimal that it approaches under-specification rather than representing a well-structured, appropriately detailed description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter mutation tool with no annotations and no parameter documentation, the description is incomplete. The attestation lifecycle fact is valuable, but the agent still lacks enough context to confidently select and correctly populate the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description names none of the eight parameters, including required fields like subject_id, evidence_type, and source. It also does not clarify ambiguous optional fields such as evidence_class, result, payload, actor_id, or artifact_id.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a clear verb and resource ('Persist evidence') and adds the important lifecycle fact that new evidence is UNATTESTED until attested. This helps distinguish it from the sibling attest_evidence tool, though it does not elaborate on what evidence submission entails.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance is given about when to use this tool versus alternatives like attest_evidence, verify_slice, or submit_verification_observation. The implied usage is 'when persisting evidence,' but no exclusion criteria or preferred conditions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
submit_planC
Record a plan. v0.3 work requires WorkIdentity + controller-minted EXECUTE/CONTINUE turn + current baseline.
| Name | Required | Description | Default |
|---|---|---|---|
| intent | Yes | ||
| spec_id | No | ||
| turn_id | No | ||
| work_id | No | ||
| actor_id | Yes | ||
| estimate | No | ||
| family_id | Yes | ||
| request_id | Yes | ||
| contract_ids | No | ||
| expected_scope | No | ||
| proposed_slices | Yes | ||
| expected_artifacts | No | ||
| normative_baseline_id | No | ||
| acceptance_expectations | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full behavioral burden. It discloses a version-specific prerequisite, which is useful context, but says nothing about mutation effects, idempotency, permissions, or what happens to existing state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very short and front-loads the core action before the prerequisite clause. There is no wasted text, though the second sentence is dense jargon that may not earn its place without further clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 14-parameter mutation tool with no annotations and 0% schema description coverage, the description is drastically incomplete. While the presence of an output schema removes the need to explain return values, the description omits almost everything an agent needs to call this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 14 parameters, and the description only hints at three of them ('WorkIdentity', 'EXECUTE/CONTINUE turn', 'current baseline'). The five required parameters (family_id, request_id, actor_id, intent, proposed_slices) receive no explanation at all.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Record') and resource ('a plan'), so an agent can tell it creates a plan record. However, it does not distinguish this from siblings like close_plan or begin_work, and the second sentence shifts to version-specific prerequisites rather than clarifying purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only guidance is a cryptic prerequisite clause: 'v0.3 work requires WorkIdentity + controller-minted EXECUTE/CONTINUE turn + current baseline.' It does not say when to use submit_plan versus alternatives, nor does it mention exclusions or sequencing with siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
submit_verification_observationC
Persist an independent AV/1 verifier observation. REPLAY requires an intact RB/1 binding; MangoMe does not execute the check itself.
| Name | Required | Description | Default |
|---|---|---|---|
| claim | Yes | ||
| source | Yes | ||
| status | Yes | ||
| details | No | ||
| turn_id | No | ||
| slice_id | Yes | ||
| artifact_id | No | ||
| reproduction | No | ||
| evidence_type | Yes | ||
| evidence_class | Yes | ||
| verifier_token | Yes | ||
| observation_type | Yes | ||
| verifier_actor_id | Yes | ||
| original_evidence_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose one genuinely useful behavioral trait — that MangoMe 'does not execute the check itself,' i.e. this only persists an observation rather than performing verification — plus a REPLAY/RB/1 precondition. But it omits auth requirements, whether persisted observations are immutable, failure behavior, and other traits expected of a write tool with no annotation support.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no padding, and the core action is front-loaded. The efficiency is undercut only by unexplained jargon that forces the reader to look elsewhere for meaning, but structurally it is tight.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be described, but the definition remains incomplete: 14 parameters, 9 required, with zero schema descriptions and nothing in the prose; no annotations; and heavy unexplained domain jargon (AV/1, RB/1, MangoMe, REPLAY). An agent lacks enough to invoke this reliably.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
14 parameters with 0% schema description coverage, and the description explains none of them. It doesn't clarify the roles of verifier_token, verifier_actor_id, observation_type, evidence_type, evidence_class, status, or the optional details/reproduction objects. With no annotations and no parameter prose, the description fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('Persist ... verifier observation'), so the basic action is identifiable. However, the AV/1 qualifier is unexplained jargon, and nothing distinguishes this from nearby siblings such as verify_slice, submit_evidence, attest_evidence, or record_truth_assertion. Purpose is clear-ish but not distinguishable within its family.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance and no named alternative. The line about REPLAY requiring an intact RB/1 binding hints at a precondition, but it is opaque and does not tell the agent which sibling to call instead or under what circumstances this tool applies.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
trust_boundary_statusA
Report MangoMe's MongoDB trust-boundary posture without exposing credentials.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It adds one important behavioral trait—'without exposing credentials'—which signals a security-conscious read operation. But it does not explicitly state read-only nature, required permissions, or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, front-loaded with the core purpose and a key constraint. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists, so return values need not be described. The description covers what is reported and a key security constraint, but for a status tool with no annotations it could specify read-only behavior or usage context. Still largely complete for its simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, and the schema is empty; the baseline for no params is 4. Schema coverage is 100%, so no additional parameter semantics are needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb 'Report' and resource 'MongoDB trust-boundary posture' are stated, making the tool's intent clear. However, it does not differentiate from sibling status tools like fast_judgment_status or audit_status, so an agent must infer its niche.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, no alternatives mentioned. The description only says what it does, leaving the agent to guess when this status is relevant versus the many other status tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
truth_assertion_statusC
Evaluate one assertion's supportability without changing truth or activation state.
| Name | Required | Description | Default |
|---|---|---|---|
| known_at | No | ||
| valid_at | No | ||
| assertion_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It clearly discloses one important trait: the operation does not change truth or activation state, so it is non-mutating. But it does not cover permissions, authentication requirements, error conditions, or other side-effect details that an agent might need.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler. It is concise and every clause contributes, though the brevity leaves little structure for usage or parameter context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. Still, for a tool with three parameters at 0% schema description coverage and many closely related truth/assertion siblings, the description is substantially incomplete: it omits parameter semantics, usage alternatives, and permission or side-effect context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for three parameters. The description only implies that an assertion is evaluated, giving no meaning for the required assertion_id or the optional known_at and valid_at temporal parameters. It therefore fails to compensate for the completely undocumented schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Evaluate') and resource ('one assertion's supportability'), and it clarifies the non-mutating scope by saying it does not change truth or activation state. It distinguishes from write siblings like record_truth_assertion and invalidate_truth_assertion, but does not differentiate from read-oriented siblings such as truth_at or bitemporal_truth_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'without changing truth or activation state' implies a read-only status check, which gives some contextual guidance. However, there is no explicit when-to-use guidance, no when-not-to-use guidance, and no alternatives named among the many truth-related sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
truth_atC
Query BTTM/1 independently by represented-world time and MangoMe-known time.
| Name | Required | Description | Default |
|---|---|---|---|
| known_at | No | ||
| valid_at | No | ||
| work_ref | Yes | ||
| include_unsupported | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, yet it says nothing about whether this is a read or write operation, what is returned, permission requirements, or the effect of include_unsupported. For a bitemporal query tool this is a substantial gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single front-loaded sentence with no filler, but it is terse to the point of obscurity — every word is jargon the agent must guess at, so brevity here costs more than it saves.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, but with 4 parameters, 0% schema coverage, no annotations, and unexplained domain jargon, the description is far too thin for a bitemporal query tool. Key context an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 4 parameters. The description alludes to the two time axes (valid_at / known_at) but never maps them to parameter names, and says nothing about work_ref or include_unsupported, leaving half the parameters undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a verb ('Query') and a resource ('BTTM/1'), and points at two temporal dimensions, but 'BTTM/1' and 'MangoMe' are opaque jargon that an agent cannot resolve without external knowledge. It does not distinguish itself from siblings like bitemporal_truth_status or truth_assertion_status, so the specific purpose remains ambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no exclusions, and no routing to alternatives among the numerous truth-related siblings. The agent must infer everything about when this tool is the right choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_slice_progressC
Persist slice progress; the same active plan that started the slice is mandatory.
| Name | Required | Description | Default |
|---|---|---|---|
| blocker | No | ||
| plan_id | Yes | ||
| actor_id | Yes | ||
| slice_id | Yes | ||
| total_steps | No | ||
| current_step | No | ||
| execution_state | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full disclosure burden. It adds one genuinely useful behavioral constraint (the plan_id must be the same active plan that started the slice, implying rejection on mismatch). However, it doesn't disclose mutation semantics explicitly, idempotency, or failure behavior beyond the implied plan-matching rule, leaving most of the burden unmet.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a compact two-clause sentence with no filler. The primary action is front-loaded and the constraint follows. The phrase 'is mandatory' is slightly awkward (mandatory what?), but the overall structure is efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter mutation tool with zero annotations and zero schema coverage, this description is thin. The output schema covers return values, but the description misses usage context, parameter semantics for 6 of 7 params, and behavioral details like what happens when the plan mismatch rule is violated. An agent would need substantial external knowledge to invoke this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for all 7 undocumented parameters. It clarifies plan_id's semantic constraint ('the same active plan that started the slice is mandatory') but says nothing about slice_id, actor_id, blocker, total_steps, current_step, or execution_state, leaving agents to infer from parameter names alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb+resource ('Persist slice progress') that clearly identifies the tool's action. It distinguishes reasonably from lifecycle siblings like start_slice, verify_slice, and claim_done by the nature of 'progress persistence', though it doesn't explicitly name them. It could be richer, but the core purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use or when-not-to-use guidance and names no alternatives despite a large sibling set (61 tools) containing closely related state/transition tools like claim_done, start_slice, and begin_work. The plan-matching constraint is a precondition, not usage routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_sliceC
Verify DONE_CLAIMED only when PASS gates (or gateless proof) include independent AV/1 observed PASS Evidence.
| Name | Required | Description | Default |
|---|---|---|---|
| turn_id | No | ||
| slice_id | Yes | ||
| evidence_ids | No | ||
| verifier_token | No | ||
| verifier_actor_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It indicates a conditional verification operation but does not describe permissions required, side effects, state changes, reversibility, or what happens on failure, leaving the agent with little beyond the high-level gate condition.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence with no wasted words, so it is concise. However, the structure is not front-loaded for clarity; the condition dominates and the core action is buried in jargon, making it hard to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given five parameters, zero schema description coverage, no annotations, and a complex verification domain, the description is not complete enough. The existence of an output schema means return values need not be explained, but input semantics and behavioral expectations remain largely absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for five parameters, and the description does not identify or explain slice_id, verifier_actor_id, evidence_ids, verifier_token, or turn_id. It only alludes to evidence in the abstract, which does not help an agent understand how to supply the required inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action ('Verify DONE_CLAIMED') and a precondition involving PASS gates and observed PASS Evidence. However, the purpose is expressed in dense domain jargon without plain-language explanation, and it does not differentiate this tool from siblings such as completion_review or accept_slice.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear, explicit condition for when the tool may be used: only when PASS gates (or gateless proof) include independent AV/1 observed PASS Evidence. It does not name alternative tools or describe when not to use it, but the gate condition itself acts as a strong usage constraint.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
work_contextC
Read identity-bound canonical/progressive/volatile context without filesystem reconstruction.
| Name | Required | Description | Default |
|---|---|---|---|
| work_ref | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It implies a read operation and rules out filesystem reconstruction, but says nothing about permissions, scope, identity binding semantics, or what 'volatile' context means for callers; an output schema exists so return values need not be described, but the operational profile is thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the verb first and no filler; every clause is doing something. The phrasing is dense to the point of obscurity, but it is not padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, a single undocumented parameter, and many closely named context-reading siblings, the description is not sufficient for reliable selection or invocation. The output schema covers return values, but the gaps around routing and parameter meaning remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one parameter, work_ref, with 0% schema description coverage, so the description must compensate and largely does not. 'Identity-bound' loosely hints that work_ref is an identity reference, but its expected format, source, and relationship to work identity are never stated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description has a specific verb ('Read') and a stated resource ('identity-bound canonical/progressive/volatile context'), but the resource is opaque jargon that an agent cannot map to a concrete artifact. 'Without filesystem reconstruction' gestures at a boundary, but the core purpose remains only partially intelligible.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit statement of when to call this versus alternatives such as read_context, recovery_context, or audit_context, all of which sit adjacent in the sibling list. The only hint is the negative clause 'without filesystem reconstruction', which implies an alternative without naming it or the condition that selects this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workspace_statusC
Return attachment state plus any durable workspace project already known to MangoMe.
| Name | Required | Description | Default |
|---|---|---|---|
| refresh | No | ||
| workspace_root | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It mentions 'already known to MangoMe' implying a cache/state check but does not disclose whether the call triggers side effects (refresh parameter suggests possible mutation) or what 'attachment state' means. The description is too brief to inform the agent about behavior like failure modes or caching semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the main functionality. It is appropriately concise, though it could be slightly longer to add usage context without bloat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema, which may describe the return value, but the description omits key operational details: when to call, what 'attachment' refers to, prerequisites, and side effects of the refresh parameter. Complexity is moderate, but the empty annotations and low schema coverage mean the description should be more substantive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain what the 'refresh' or 'workspace_root' parameters do. The description only mentions the output, not inputs. Since the schema provides parameter names but no explanatory descriptions, the description must compensate and fails to do so.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool returns 'attachment state plus any durable workspace project already known to MangoMe', which is a specific resource and outcome. However, the meaning of 'attachment state' and 'durable workspace project' is unclear without deeper domain knowledge, and it does not distinguish itself from siblings like 'status' or 'project_overview'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives. The description only says what it returns, not when an agent should invoke it. Given the large sibling set, an explicit use-case statement (e.g., 'call this to check current workspace status before starting work') is missing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
41 tool updates
v0.3.4- Changed
accept_slice1 field changed- added
Input schema / properties / turn_idAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Turn Id" +}
- Added
assess_fast_judgment - Added
audit_context - Added
audit_mutation_allowed - Added
audit_status - Added
backfill_work_identity - Changed
begin_work2 fields changed- added
Input schema / properties / turn_idAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Turn Id" +} - added
Input schema / properties / work_idAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Work Id" +}
- Added
bind_contract_turn - Added
bind_work_turn - Added
bitemporal_truth_status - Added
checkpoint_work - Added
close_scoped_audit - Added
cognitive_hygiene - Added
contract_state - Changed
create_family2 fields changed- added
Input schema / properties / controller_actor_idAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Controller Actor Id" +} - added
Input schema / properties / controller_tokenAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Controller Token" +}
- Changed
create_project2 fields changed- added
Input schema / properties / controller_actor_idAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Controller Actor Id" +} - added
Input schema / properties / controller_tokenAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Controller Token" +}
- Changed
create_spec5 fields changed- added
Input schema / properties / actor_idAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Actor Id" +} - added
Input schema / properties / controller_actor_idAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Controller Actor Id" +} - added
Input schema / properties / controller_tokenAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Controller Token" +} - added
Input schema / properties / turn_idAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Turn Id" +} - added
Input schema / properties / work_idAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Work Id" +}
- Changed
enter_work2 fields changed- added
Input schema / properties / controller_actor_idAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Controller Actor Id" +} - added
Input schema / properties / controller_tokenAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Controller Token" +}
- Added
fast_judgment_status - Changed
import_contract_bundle2 fields changed- added
Input schema / properties / controller_actor_idAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Controller Actor Id" +} - added
Input schema / properties / controller_tokenAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Controller Token" +}
- Added
invalidate_truth_assertion - Changed
link_entities4 fields changed- added
Input schema / properties / controller_actor_idAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Controller Actor Id" +} - added
Input schema / properties / controller_tokenAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Controller Token" +} - added
Input schema / properties / turn_idAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Turn Id" +} - added
Input schema / properties / work_idAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Work Id" +}
- Added
prepare_assignment - Added
promote_contract_generation - Added
reconcile_assignment - Added
record_audit_finding - Added
record_fast_judgment - Added
record_truth_assertion - Changed
register_contract4 fields changed- added
Input schema / properties / controller_actor_idAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Controller Actor Id" +} - added
Input schema / properties / controller_tokenAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Controller Token" +} - added
Input schema / properties / turn_idAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Turn Id" +} - added
Input schema / properties / work_idAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Work Id" +}
- Added
register_playbook - Added
release_contract_generation_grant - Added
select_playbook - Changed
set_gate1 field changed- added
Input schema / properties / turn_idAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Turn Id" +}
- Added
start_scoped_audit - Changed
submit_plan7 fields changed- added
Input schema / properties / normative_baseline_idAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Normative Baseline Id" +} - added
Input schema / properties / spec_id / anyOfAdded value: +[ + { + "type": "string" + }, + { + "type": "null" + } +] - added
Input schema / properties / spec_id / defaultAdded value: +null - removed
Input schema / properties / spec_id / typeRemoved value: -"string" - added
Input schema / properties / turn_idAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Turn Id" +} - added
Input schema / properties / work_idAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Work Id" +} - changed
Input schema / requiredPrevious value: -[ - "family_id", - "request_id", - "spec_id", - "actor_id", - "intent", - "proposed_slices" -]New value: +[ + "family_id", + "request_id", + "actor_id", + "intent", + "proposed_slices" +]
- Changed
submit_verification_observation1 field changed- added
Input schema / properties / turn_idAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Turn Id" +}
- Added
trust_boundary_status - Added
truth_assertion_status - Added
truth_at - Changed
verify_slice1 field changed- added
Input schema / properties / turn_idAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Turn Id" +}
- Added
work_context
40 tool updates
v0.1.2- Changed
accept_slice4 fields changed- added
Input schema / properties / accepted_by / anyOfAdded value: +[ + { + "type": "string" + }, + { + "type": "null" + } +] - added
Input schema / properties / accepted_by / defaultAdded value: +null - removed
Input schema / properties / accepted_by / typeRemoved value: -"string" - changed
Input schema / requiredPrevious value: -[ - "slice_id", - "approval_id", - "accepted_by" -]New value: +[ + "slice_id", + "approval_id" +]
- Changed
approve_override1 field changed- added
Input schema / properties / approval_tokenAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Approval Token" +}
- Added
attest_evidence - Added
authorize_delegation - Added
begin_work - Changed
bigbang_scan2 fields changed- added
Input schema / properties / id_patternsAdded value: +{ + "anyOf": [ + { + "items": { + "type": "string" + }, + "type": "array" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Id Patterns" +} - added
Input schema / properties / include_gitAdded value: +{ + "default": true, + "title": "Include Git", + "type": "boolean" +}
- Added
build_reproduction_binding - Changed
claim_done2 fields changed- added
Input schema / properties / plan_idAdded value: +{ + "title": "Plan Id", + "type": "string" +} - changed
Input schema / requiredPrevious value: -[ - "slice_id", - "actor_id" -]New value: +[ + "slice_id", + "actor_id", + "plan_id" +]
- Added
compile_uai_context - Added
complete_delegation - Added
completion_review - Added
decode_uai_result - Added
delegation_status - Added
discovery_scopes - Added
effective_family_view - Added
enter_work - Added
evidence_freshness - Added
execution_eligibility - Added
expand_uai_context - Added
filesystem_references - Added
filesystem_scan - Added
health - Added
maintenance_diagnose - Added
migrate_schema - Added
project_overview - Added
publish_worker_runtime - Added
reconcile_bigbang - Changed
record_execution_receipt3 fields changed- added
Input schema / properties / context_tokens_interlinguaAdded value: +{ + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Context Tokens Interlingua" +} - added
Input schema / properties / interlingua_versionAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Interlingua Version" +} - added
Input schema / properties / output_tokens_interlinguaAdded value: +{ + "anyOf": [ + { + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Output Tokens Interlingua" +}
- Added
recovery_context - Changed
reject_override1 field changed- added
Input schema / properties / approval_tokenAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Approval Token" +}
- Added
render_uai_result - Added
repository_locations - Added
session_bootstrap - Added
session_restore - Changed
set_gate2 fields changed- added
Input schema / properties / actor_idAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Actor Id" +} - added
Input schema / properties / approval_idAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Approval Id" +}
- Changed
submit_evidence1 field changed- added
Input schema / properties / evidence_classAdded value: +{ + "default": "CLAIM", + "title": "Evidence Class", + "type": "string" +}
- Added
submit_verification_observation - Changed
update_slice_progress2 fields changed- added
Input schema / properties / plan_idAdded value: +{ + "title": "Plan Id", + "type": "string" +} - changed
Input schema / requiredPrevious value: -[ - "slice_id", - "actor_id" -]New value: +[ + "slice_id", + "actor_id", + "plan_id" +]
- Changed
verify_slice1 field changed- added
Input schema / properties / verifier_tokenAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Verifier Token" +}
- Added
workspace_status
32 tool updates
v0.1.1- First observed
accept_slice - First observed
approve_override - First observed
attach_artifact - First observed
bigbang_scan - First observed
claim_done - First observed
close_plan - First observed
compile_execution_context - First observed
create_family - First observed
create_project - First observed
create_spec - First observed
graph - First observed
import_contract_bundle - First observed
intake_request - First observed
link_entities - First observed
list_approvals - First observed
model_stats - First observed
read_context - First observed
record_execution_receipt - First observed
refresh_views - First observed
register_contract - First observed
register_model - First observed
reject_override - First observed
request_override - First observed
resolve - First observed
set_gate - First observed
set_gate_controlled - First observed
start_slice - First observed
status - First observed
submit_evidence - First observed
submit_plan - First observed
update_slice_progress - First observed
verify_slice
TDQS
Scored across 90 tools
Many tools overlap or are indistinguishable: enter_work/begin_work/prepare_assignment/reconcile_assignment all touch work initiation; multiple status/context tools (status, project_overview, workspace_status, work_context, read_context, recovery_context, session_restore, session_bootstrap) have unclear boundaries. The agent will struggle to select the right tool without deep MangoMe-specific knowledge.
All lowercase snake_case, but patterns are inconsistent: bare verbs (resolve, health), noun phrases (fast_judgment_status, contract_state), verb_noun (create_spec), and aliases (set_gate_controlled). No predictable verb_noun convention.
90 tools is an extreme mismatch for any server; many are niche governance operations that could be consolidated or hidden behind higher-level tools. This count far exceeds the 3-15 range for a well-scoped set.
The surface broadly covers a complex governance lifecycle (identity, contracts, truth, verification, audit, delegation, discovery, migration), but gaps remain: no delete/list/search operations for many entities, and some flows depend on context tools rather than direct CRUD.
Maintenance
Related MCP Connectors
Cross-agent artifact workspace with provenance across Claude Code, Codex, Cursor, LangGraph.
Durable agent-to-agent handoffs and shared scratchpad for multi-agent workflows.
Shared task queue for humans and AI agents: leases, handoffs, approvals and signed receipts.
Machine-native research commons for agent evidence, discovery, rooms, and bounded research quests.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables AI coding agents to execute formal, stateful workflows with typed contracts, postcondition enforcement, and structured retry logic.1Apache 2.0
- AlicenseAqualityCmaintenanceLocal-first shared memory and coordination layer for AI coding agents, with repository evidence, reservations, handoffs, code graph context, and dashboard review backed by PostgreSQL/pgvector.303Apache 2.0
- AlicenseBqualityAmaintenanceAn append-only coordination memory for multi-agent and human work, backed by SQLite, with a local dashboard and acceptance contracts that enforce integrator review before work is considered accepted.431MIT
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to maintain persistent, inspectable understanding through typed, revisable updates, and to coordinate multi-agent work via shared graph-based stigmergy.81 npm1MIT