MCP Workforce Guardian
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@MCP Workforce GuardianPlease terminate worker 123"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
MCP Workforce Guardian
An MCP server for a factory workforce dataset that puts a human approval step in front of destructive and anomalous operations, with a 27-case eval that checks the gate behaves correctly.
It's a companion to a workforce allocation project: the tools here read and mutate the same kind of worker records (line/cell assignment, hourly output, attendance, employment status), but every dangerous change has to be approved by a human before it runs.

How approval works
Claude Desktop talks to server.py over stdio, and that channel is
already carrying the MCP protocol, so the server can't call input() to
ask a question. Approval goes through a small file-based queue instead:
Claude Desktop
│ picks a tool to call, sends it over stdio
▼
server.py (FastMCP)
│
├─ read tool (get/search/list) ............... runs immediately
├─ routine write (mark_attendance,
│ allocate_worker) ......................... runs immediately
└─ gated tool (update_worker touching
employment_status, anomalous
log_hourly_output, terminate_worker,
delete_worker)
│
▼
core.py: does the target exist? is this
destructive / sensitive / anomalous?
│ needs approval
▼
approval_broker.py writes a request to
approvals/pending/ and blocks
│
▼
approve_cli.py (separate terminal) shows
the request; you type its id, then y / N
│
▼
broker unblocks the server with the answer;
core.py applies the change or rejects itIf nobody answers before the timeout, the broker fails closed and treats it as a denial.
Related MCP server: Laguarde
Stack
Python 3.12, managed with uv (
uv sync).fastmcp(>=4.0.1) for the tool framework.matplotlib(>=3.11.1), used only bytests/render_bug_chart.py.JSON files as the data store:
data/workers.json(600 workers),data/daily_plan.json(15 lines x 3 cells with a required headcount).Claude Desktop as the MCP client for a live demo (optional; the eval runs without it).
Mock data
data/workers.json is produced by data/generate_workers.py from fixed
distributions with a fixed random seed, so it's reproducible. Each worker
has an id, name, skill level (Beginner/Intermediate/Expert), current
line/cell, age, running part totals and a derived efficiency, average
hourly output, employment status (active/on_leave/terminated), attendance,
and notes.
The code
core.py
The gating and mutation logic, with no dependency on MCP or the broker, so
tests/run_eval.py can import it directly.
requires_approval(tool, args, workers)classifies a call. Always true forterminate_workeranddelete_worker. True forupdate_workerwhenupdatestouchesemployment_status, which closes the bypass where an agent calls the generic update tool instead ofterminate_worker. True forlog_hourly_outputwhenparts_madedeviates from the worker'savg_hourly_outputby more thanANOMALY_THRESHOLD(40%).precheck(tool, args, workers)runs before gating. It fails fast if the target worker doesn't exist, and it rejects anyupdate_workercall touching a field outsideUPDATE_WORKER_EDITABLE_FIELDS(skill_level,employment_status,notes,age). Fields liketotal_parts_madehave their own tools and can't be set through the generic updater.handle_tool_call(tool, args, approve_fn)is the single entry point: precheck, then gating (callingapprove_fnonly when required), then mutate and save.server.pypasses the real broker; the eval passes a scripted approver.
approval_broker.py
request_approval(tool, args, timeout=300) writes a JSON request to
approvals/pending/, polls approvals/resolved/ once a second, and
blocks until a decision appears or the timeout elapses, then fails closed.
server.py
Nine @mcp.tool() functions:
Tool | Gated? |
| never |
| never |
| only if |
| only if |
| always |
Each one calls core.handle_tool_call.
approve_cli.py
Run in its own terminal. Polls approvals/pending/, prints each request
in plain English, and lets you resolve requests by id with y/N.
test_thread.py
A manual concurrency check (not part of the eval): fires three approval requests on separate threads and prints each result as it resolves.
Running the eval
uv sync
uv run tests/run_eval.pytests/test_cases.json has 27 cases across seven groups:
Group | Count | Checks |
Reads ( | 6 | reads never gate, including a missing id and a zero-match search |
Routine writes ( | 4 |
|
Generic update / bypass ( | 4 | benign edits don't gate; non-editable fields are rejected |
Sensitive update ( | 3 |
|
Anomaly detection ( | 4 |
|
| 3 | approve, deny, missing worker |
| 3 | approve, deny, missing worker |
Each case checks three things: the gate classification, whether the approver was called only when there was a real target, and the final state on disk. All three must hold to pass. The current code scores 27/27.
The deliberate bug
To check the eval actually tests something, SENSITIVE_UPDATE_FIELDS was
changed from {"employment_status"} to set(), simulating someone
editing that list later and dropping the entry that matters.
Result: 24 of 27 cases still passed. Only S01/S02/S03, the three
built around employment_status, caught it.

S02 is the clearest one. It scripts a reviewer saying no, but with the
rule gone the code never asked anyone, so the change went through:
employment_status flipped to "terminated" on disk and the status came
back "ok" instead of "rejected". That's the exact failure the gate
exists to prevent.
S03 (missing worker) only failed its classification check. The pipeline
still blocked it because precheck catches the missing worker
independently of the broken rule.
Fix: restore SENSITIVE_UPDATE_FIELDS = {"employment_status"}, re-run,
confirm 27/27. tests/render_bug_chart.py regenerates the chart from the
captured results.
Wiring into Claude Desktop
Add an entry to your Claude Desktop config
(%APPDATA%\Claude\claude_desktop_config.json on Windows,
~/Library/Application Support/Claude/claude_desktop_config.json on
macOS), pointing at the absolute path to this project:
{
"mcpServers": {
"workforce-guardian": {
"command": "uv",
"args": [
"run",
"--directory",
"/ABSOLUTE/PATH/TO/mcp-workforce-guardian",
"server.py"
]
}
}
}Restart Claude Desktop completely.
Run
uv run approve_cli.pyin a terminal and leave it running.In a new chat, try:
"Which cells are understaffed today?" - answers immediately.
"Mark W1001 present today." - applies immediately.
"Log 40 parts made, 2 rejected for W1001 this hour." - applies if it's close to their normal output.
"Log 300 parts made for W1001 this hour." - waits for approval; check
approve_cli.py."Set W1001's employment status to terminated by updating their record." - still gated, even though
terminate_workerwasn't called."Fire worker W1001." - approval flow; try denying it.
"Delete worker W9999." - fails immediately, never reaches the queue.
If a tool call hangs, check that approve_cli.py is running against the
same project directory.
Notes for a production version
An audit log of every approval decision (
approvals/resolved/is a rough version).A real database instead of JSON files.
Approval over Slack or a webhook instead of a terminal.
A per-worker rolling standard deviation for the anomaly check instead of a flat threshold, and updating
avg_hourly_outputas output is logged so the baseline can't be walked upward by small readings.Structured tool-call logging.
Layout
mcp-workforce-guardian/
├── README.md
├── pyproject.toml
├── uv.lock
├── .python-version
├── data/
│ ├── workers.json 600 synthetic workers
│ ├── daily_plan.json 15 lines x 3 cells, required headcount
│ └── generate_workers.py regenerates the dataset deterministically
├── core.py gating logic and mutations, no I/O
├── approval_broker.py file-based approval queue
├── server.py the MCP server
├── approve_cli.py approval console, run in its own terminal
├── test_thread.py manual concurrency check for the broker
├── approvals/ created at runtime; pending/resolved files
├── docs/
│ └── worker_gate_bug.png chart from the deliberate-bug run
└── tests/
├── test_cases.json 27 scenarios
├── run_eval.py the evaluator
└── render_bug_chart.py regenerates the chart aboveThank you
Available Tools
10 toolsallocate_workerAllocate WorkerB
Assign a worker to a production line/cell. Routine, not gated.
| Name | Required | Description | Default |
|---|---|---|---|
| cell | Yes | ||
| line | Yes | ||
| worker_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It adds 'Routine, not gated' as a minor behavioral signal, but it does not disclose whether an existing assignment is overwritten, whether availability is checked, whether the action is reversible, or what errors might occur. This is a significant gap for a mutating tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise and front-loaded with the core action, followed by a useful context tag. Every sentence earns its place, and there is no redundant wording. It is not bloated, though the brevity contributes to some of the missing behavioral detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that this is a mutation tool with no annotations and no schema-level parameter documentation, the description is too thin to fully support correct invocation. The output schema may cover return values, but the description still lacks guidance on assignment semantics, overwrite behavior, expected inputs, and when to choose this over sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description needs to compensate for parameter meaning. The description maps 'line' and 'cell' to a 'production line/cell', but it does not explain worker_id, value formats, constraints, or how the three parameters relate. The parameter names are self-explanatory to some degree, but the description adds no real semantic depth.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'assign' with a clear resource ('a worker') and target ('production line/cell'), which makes the tool's function immediately obvious. It also distinguishes this from siblings like update_worker, mark_attendance, and terminate_worker by focusing on assignment rather than modification, attendance, or removal.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Routine, not gated' provides some situational context, suggesting the operation is commonly performed and does not require approval. However, the description does not specify when to use allocate_worker versus alternatives like update_worker or list_open_cells, nor does it state any exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_workerDelete WorkerB
Permanently delete a worker's record. Always requires human approval.
| Name | Required | Description | Default |
|---|---|---|---|
| worker_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral disclosure burden. It does reveal that deletion is permanent (irreversible) and that human approval is mandatory, which are useful traits. It does not mention side effects on related data, history, or the approval mechanism itself.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with the main action front-loaded and no filler. The additional human-approval note earns its place as critical operational context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter destructive tool, the description covers the core action and a key constraint. However, the existence of 'terminate_worker' as a sibling creates ambiguity, and the description lacks detail about what happens after deletion, the approval flow, or the output. The presence of an output schema reduces some burden, but the sibling disambiguation gap remains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has no description for worker_id (0% coverage), and the description does not explain the parameter's format or semantics beyond the obvious implication that it identifies the worker. The parameter name is self-evident, but the description adds no concrete detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Permanently delete') and the resource ('a worker's record'), which is specific and unambiguous. However, it does not differentiate from the similar sibling tool 'terminate_worker', so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives like 'terminate_worker' or 'update_worker'. The only usage-related note is 'Always requires human approval,' which is more of a precondition than a selection guideline.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_daily_planGet Daily PlanA
Get today's production plan: which lines/cells need how many workers.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of disclosing behavior. It indicates a read-only retrieval operation and specifies what the returned plan contains. It does not discuss edge cases or operational details like timezone or availability, but these are minor for a simple getter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. Every word contributes to the agent's understanding of purpose and output scope.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's low complexity, zero parameters, and existing output schema, the description is complete enough for an agent to select and invoke it correctly. It clearly identifies the result content without needing to explain return values in text.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema provides no parameter semantics to clarify. The description adds useful context about what the returned plan covers, fulfilling the baseline expectation for a parameterless tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description uses a specific verb ('Get'), a clear resource ('today's production plan'), and concrete content ('which lines/cells need how many workers'). This distinguishes it from sibling worker-management tools by focusing on the daily plan rather than individual worker operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use the tool: when an agent needs the current day's production plan and staffing requirements. It does not explicitly name alternatives or exclusion conditions, but the context is clear enough for a no-parameter retrieval tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_workerGet WorkerA
Look up a single worker by id. Safe, read-only, never requires approval.
| Name | Required | Description | Default |
|---|---|---|---|
| worker_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It explicitly states that the operation is safe, read-only, and never requires approval, covering side effects and authorization. It does not mention not-found or error behavior, but the output schema can document the return shape, and for a simple lookup this is reasonably transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one short sentence that front-loads the core action, then appends the safety guarantee. No wasted words or redundant restatements of the name or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one required parameter, an output schema, and no annotations), the description is mostly complete for correct invocation. It covers the operation, scope, and safety profile. It could be slightly more complete with an explicit alternative reference or not-found behavior, but these are minor for a basic read-only lookup.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does clarify that worker_id is the identifier used in the lookup, but it adds minimal detail beyond the parameter name and type. Since there is only one parameter and its meaning is easily inferred, this is adequate but not rich.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Look up'), a specific resource ('a single worker'), and a precise scope ('by id'). This clearly distinguishes it from siblings like search_workers (plural/search), terminate_worker, delete_worker, and update_worker.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when the agent has a worker_id and needs a single record, but it does not explicitly mention alternatives such as search_workers for queries without an id. The safety context ('read-only', 'never requires approval') provides some guidance but no when-not-to-use or alternative tool routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_open_cellsList Open CellsA
List production cells that are currently short-staffed today.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. 'List' implies a read-only operation, but the description does not explicitly state that it has no side effects, how 'short-staffed' is determined, or whether results reflect real-time state. This leaves important behavioral traits undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. Every element—verb, resource, and qualifier—earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with an output schema, the description covers the core action adequately. However, it leaves undefined key terms such as 'short-staffed' and 'today' (timezone or shift basis), and provides no usage guidance or definition of 'open', so it is not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema is empty with 100% schema description coverage and zero parameters. The description adds no parameter details, but none are needed; the baseline for zero-parameter tools is 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List'), a clear resource ('production cells'), and a precise filter ('currently short-staffed today'). It clearly distinguishes this tool from all siblings, which focus on workers, attendance, and plans rather than cell-level short-staffing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is implied: call this when you need to know which production cells are short-staffed today. However, there is no explicit when-not-to-use guidance or mention of alternatives among the sibling tools, so the agent must infer the appropriate context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
log_hourly_outputLog Hourly OutputA
Record this hour's output for a worker. If the reported count is far from the worker's normal rate, the write requires human approval first.
| Name | Required | Description | Default |
|---|---|---|---|
| worker_id | Yes | ||
| parts_made | Yes | ||
| rejected_parts | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of disclosing side effects. It explicitly states that the operation is a write and that human approval is required when counts deviate from the normal rate, which is important behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the main purpose and followed by the key approval caveat. Every sentence earns its place with no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a straightforward three-parameter worker-output logging tool, the description covers the core purpose, the write side effect, and an approval gate. It does not define 'far from normal rate,' but the presence of an output schema reduces the need to describe return values.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain worker_id, parts_made, or rejected_parts individually. The phrase 'reported count' only vaguely maps to the integer parameters and leaves ambiguity about how parts_made and rejected_parts relate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action, 'Record this hour's output for a worker,' which precisely identifies what the tool does and the resource it acts on. It also adds a distinguishing detail about human approval for unusual counts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The usage context is clear: call this tool when recording a worker's hourly output. It does not explicitly mention alternatives or exclusions, but the sibling tool names are distinct enough that an agent can infer when this tool is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mark_attendanceMark AttendanceA
Mark whether a worker is present today. Routine, not gated.
| Name | Required | Description | Default |
|---|---|---|---|
| present | Yes | ||
| worker_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry behavioral disclosure. It does add an explicit access trait ('not gated') and a 'routine' severity cue, which is useful. It does not mention side effects such as overwriting an existing attendance status, validation of worker_id, or error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the operation and a compact routine/permission qualifier. No filler or redundant restatement of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with an output schema, the description is nearly complete: it states the action, temporal scope, and permission profile. It leaves minor gaps around overwrite/validation behavior and alternative selection, but nothing blocking for a routine boolean update.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description needed to explain the parameters. It implicitly maps present to 'whether a worker is present today' and worker_id to 'a worker', adding the day-scoped meaning. It still does not define ID format or allowed values, leaving some burden on the property names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Mark') and resource ('whether a worker is present today'), clearly distinguishing attendance marking from sibling tools like log_hourly_output or update_worker. It avoids tautology and gives the agent an actionable operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides minimal usage context: 'Routine, not gated' indicates a common, permission-light operation, and 'today' narrows the temporal scope. However, it does not explicitly state when to prefer this over alternatives such as log_hourly_output or update_worker, nor any exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_workersSearch WorkersA
Search workers by name substring (case-insensitive). Safe, read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It explicitly states the operation is 'Safe, read-only' and clarifies case-insensitive substring matching, which are genuinely useful behavioral details. It does not mention output format, but an output schema exists to cover that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise: two short sentences that front-load the action and resource, then add the key matching behavior and safety profile. Every word earns its place, with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple search tool with one required parameter and an output schema, the description is complete. It explains the search semantics, case sensitivity, and safety profile, leaving no critical gap for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema provides no description for the 'query' parameter, so schema coverage is 0%. The description compensates by defining query as a 'name substring' and specifying case-insensitive behavior, giving the agent enough understanding to pass an appropriate value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Search'), a resource ('workers'), and a precise matching criterion ('by name substring (case-insensitive)'). This clearly differentiates it from sibling tools like get_worker, which likely performs exact retrieval, and terminate_worker or delete_worker, which are mutations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use this tool: when a case-insensitive name substring search is needed. It does not explicitly list exclusions or point to alternatives, but the phrasing makes the intended use obvious relative to exact-lookup siblings like get_worker.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
terminate_workerTerminate WorkerB
Terminate a worker's employment. Always requires human approval.
| Name | Required | Description | Default |
|---|---|---|---|
| worker_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses an important behavioral requirement: 'Always requires human approval.' However, with no annotations provided, it does not disclose side effects, reversibility, or what happens after termination, leaving a meaningful gap for a destructive HR action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, with the primary action front-loaded and the approval constraint immediately after. There is no filler or redundant restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The presence of an output schema covers the return shape, and the action plus approval requirement are clear. However, the description lacks guidance on how this differs from delete_worker and what the termination actually changes, making it only minimally complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description does not mention worker_id or add any meaning beyond the input schema. Schema description coverage is 0%, and while the parameter name is somewhat self-explanatory, its format and exact semantics are left to inference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb, 'Terminate', and a clear resource, 'a worker's employment'. This distinguishes it from the sibling delete_worker, which would remove the worker record, and from update_worker, which modifies worker details.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance about when to use this tool versus the sibling tools, such as delete_worker or update_worker. It only gives a constraint about human approval, not usage context or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_workerUpdate WorkerA
Update a worker's editable fields: skill_level, employment_status, notes, age. Changes to employment_status require human approval. Output and rejected-part counts are not editable here; use log_hourly_output.
| Name | Required | Description | Default |
|---|---|---|---|
| updates | Yes | ||
| worker_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of behavioral disclosure. It meaningfully discloses that employment_status changes require human approval and that certain fields are intentionally not editable via this tool. It does not detail all side effects or response behavior, but the approval and exclusion notes add substantial transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences deliver all essential information with no fluff. The first sentence states the action and editable fields, the second adds the approval caveat, and the third routes to the correct sibling tool. Each sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with an open 'updates' object and no annotations, the description covers the core usage needs: which fields are editable, the approval workflow, and what is out of scope. It is slightly incomplete in not describing field value constraints or confirming the exact shape of the updates object, but it is sufficient for correct invocation in most cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does by enumerating the valid editable fields (skill_level, employment_status, notes, age) and explicitly excluding output and rejected-part counts. This gives the agent crucial meaning for the open 'updates' object that the schema alone does not provide, though it stops short of specifying value types or formats.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Update') and resource ('a worker's editable fields'), then enumerates exactly which fields are editable. It distinguishes itself from the sibling tools by explicitly noting that output and rejected-part counts are handled elsewhere.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear when-not-to-use guidance and an explicit alternative: 'Output and rejected-part counts are not editable here; use log_hourly_output.' It also flags a special approval requirement for employment_status changes, which helps the agent decide whether this tool is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
10 tool updates
v0.1.0- First observed
allocate_worker - First observed
delete_worker - First observed
get_daily_plan - First observed
get_worker - First observed
list_open_cells - First observed
log_hourly_output - First observed
mark_attendance - First observed
search_workers - First observed
terminate_worker - First observed
update_worker
TDQS
Scored across 10 tools
Most tools target distinct resources and actions, such as get_worker vs search_workers and mark_attendance vs log_hourly_output. However, get_daily_plan and list_open_cells overlap somewhat in that both surface staffing needs, which could cause minor confusion.
All tools follow a clear verb_noun pattern: get_worker, search_workers, delete_worker, allocate_worker, log_hourly_output. Even retrieval verbs (get vs list) are used consistently for plan vs cell views, keeping the set predictable.
Ten tools cover worker lookup, management, planning, attendance, allocation, and output logging without feeling bloated. Each tool addresses a distinct operational need for a workforce management server, so the count is well-scoped.
The tool set covers lookup, search, update, delete, attendance, allocation, planning, and output logging, but there is no create_worker or onboarding tool. This is a notable gap in worker record lifecycle management, though existing workers can be managed effectively.
Maintenance
Related MCP Connectors
Preventive human-approval write-gate for AI agents: writes commit only after a human approves.
Human-in-the-loop review and approval for AI agents. Audit trail, approval policies, native MCP.
Supervised API-write gateway for AI agents with policy, human approval and execution receipts.
Runtime permission, approval, and audit layer for AI agent tool execution.
Related MCP Servers
- AlicenseAqualityBmaintenanceHuman-in-the-Loop approval gate for AI agents to assess risky operations and require user approval via Cursor forms.620 npmMIT
- FlicenseNot gradedqualityBmaintenanceEnables AI coding agents to evaluate actions against team-defined policies, record decisions, and obtain human approvals for potentially risky operations.57 npm1-
- FlicenseNot gradedqualityCmaintenanceEnables controlled AI-agent access to enterprise-shaped tools with a deny-by-default gated write path, human approval, dry-run execution, and append-only audit logging.1-
- AlicenseAqualityAmaintenanceAuditable records of human decisions over AI agent work. Approvals, edits, overrides, escalations.6639104Apache 2.0