Skip to main content
Glama
Sarthak2426

MCP Workforce Guardian

by Sarthak2426

MCP Workforce Guardian

eval

An MCP server for a factory workforce dataset that puts a human approval step in front of destructive and anomalous operations, with a 27-case eval that checks the gate behaves correctly.

It's a companion to a workforce allocation project: the tools here read and mutate the same kind of worker records (line/cell assignment, hourly output, attendance, employment status), but every dangerous change has to be approved by a human before it runs.

Agent tries to delete a worker; a human denies it in the approval console; the agent reports the record was untouched

How approval works

Claude Desktop talks to server.py over stdio, and that channel is already carrying the MCP protocol, so the server can't call input() to ask a question. Approval goes through a small file-based queue instead:

Claude Desktop
   │  picks a tool to call, sends it over stdio
   ▼
server.py (FastMCP)
   │
   ├─ read tool (get/search/list) ............... runs immediately
   ├─ routine write (mark_attendance,
   │  allocate_worker) ......................... runs immediately
   └─ gated tool (update_worker touching
      employment_status, anomalous
      log_hourly_output, terminate_worker,
      delete_worker)
         │
         ▼
      core.py: does the target exist? is this
      destructive / sensitive / anomalous?
         │  needs approval
         ▼
      approval_broker.py writes a request to
      approvals/pending/ and blocks
         │
         ▼
      approve_cli.py (separate terminal) shows
      the request; you type its id, then y / N
         │
         ▼
      broker unblocks the server with the answer;
      core.py applies the change or rejects it

If nobody answers before the timeout, the broker fails closed and treats it as a denial.

Related MCP server: Laguarde

Stack

  • Python 3.12, managed with uv (uv sync).

  • fastmcp (>=4.0.1) for the tool framework.

  • matplotlib (>=3.11.1), used only by tests/render_bug_chart.py.

  • JSON files as the data store: data/workers.json (600 workers), data/daily_plan.json (15 lines x 3 cells with a required headcount).

  • Claude Desktop as the MCP client for a live demo (optional; the eval runs without it).

Mock data

data/workers.json is produced by data/generate_workers.py from fixed distributions with a fixed random seed, so it's reproducible. Each worker has an id, name, skill level (Beginner/Intermediate/Expert), current line/cell, age, running part totals and a derived efficiency, average hourly output, employment status (active/on_leave/terminated), attendance, and notes.

The code

core.py

The gating and mutation logic, with no dependency on MCP or the broker, so tests/run_eval.py can import it directly.

  • requires_approval(tool, args, workers) classifies a call. Always true for terminate_worker and delete_worker. True for update_worker when updates touches employment_status, which closes the bypass where an agent calls the generic update tool instead of terminate_worker. True for log_hourly_output when parts_made deviates from the worker's avg_hourly_output by more than ANOMALY_THRESHOLD (40%).

  • precheck(tool, args, workers) runs before gating. It fails fast if the target worker doesn't exist, and it rejects any update_worker call touching a field outside UPDATE_WORKER_EDITABLE_FIELDS (skill_level, employment_status, notes, age). Fields like total_parts_made have their own tools and can't be set through the generic updater.

  • handle_tool_call(tool, args, approve_fn) is the single entry point: precheck, then gating (calling approve_fn only when required), then mutate and save. server.py passes the real broker; the eval passes a scripted approver.

approval_broker.py

request_approval(tool, args, timeout=300) writes a JSON request to approvals/pending/, polls approvals/resolved/ once a second, and blocks until a decision appears or the timeout elapses, then fails closed.

server.py

Nine @mcp.tool() functions:

Tool

Gated?

get_worker, search_workers, get_daily_plan, list_open_cells

never

mark_attendance, allocate_worker

never

update_worker

only if updates touches employment_status

log_hourly_output

only if parts_made is >40% off baseline

terminate_worker, delete_worker

always

Each one calls core.handle_tool_call.

approve_cli.py

Run in its own terminal. Polls approvals/pending/, prints each request in plain English, and lets you resolve requests by id with y/N.

test_thread.py

A manual concurrency check (not part of the eval): fires three approval requests on separate threads and prints each result as it resolves.

Running the eval

uv sync
uv run tests/run_eval.py

tests/test_cases.json has 27 cases across seven groups:

Group

Count

Checks

Reads (R)

6

reads never gate, including a missing id and a zero-match search

Routine writes (W)

4

mark_attendance, allocate_worker never gate

Generic update / bypass (U)

4

benign edits don't gate; non-editable fields are rejected

Sensitive update (S)

3

employment_status via update_worker: approved, denied, missing worker

Anomaly detection (L)

4

log_hourly_output at baseline, a big jump approved and denied, missing worker

terminate_worker (T)

3

approve, deny, missing worker

delete_worker (D)

3

approve, deny, missing worker

Each case checks three things: the gate classification, whether the approver was called only when there was a real target, and the final state on disk. All three must hold to pass. The current code scores 27/27.

The deliberate bug

To check the eval actually tests something, SENSITIVE_UPDATE_FIELDS was changed from {"employment_status"} to set(), simulating someone editing that list later and dropping the entry that matters.

Result: 24 of 27 cases still passed. Only S01/S02/S03, the three built around employment_status, caught it.

27 eval cases with S01-S03 failing, everything else passing

S02 is the clearest one. It scripts a reviewer saying no, but with the rule gone the code never asked anyone, so the change went through: employment_status flipped to "terminated" on disk and the status came back "ok" instead of "rejected". That's the exact failure the gate exists to prevent.

S03 (missing worker) only failed its classification check. The pipeline still blocked it because precheck catches the missing worker independently of the broken rule.

Fix: restore SENSITIVE_UPDATE_FIELDS = {"employment_status"}, re-run, confirm 27/27. tests/render_bug_chart.py regenerates the chart from the captured results.

Wiring into Claude Desktop

Add an entry to your Claude Desktop config (%APPDATA%\Claude\claude_desktop_config.json on Windows, ~/Library/Application Support/Claude/claude_desktop_config.json on macOS), pointing at the absolute path to this project:

{
  "mcpServers": {
    "workforce-guardian": {
      "command": "uv",
      "args": [
        "run",
        "--directory",
        "/ABSOLUTE/PATH/TO/mcp-workforce-guardian",
        "server.py"
      ]
    }
  }
}
  1. Restart Claude Desktop completely.

  2. Run uv run approve_cli.py in a terminal and leave it running.

  3. In a new chat, try:

    • "Which cells are understaffed today?" - answers immediately.

    • "Mark W1001 present today." - applies immediately.

    • "Log 40 parts made, 2 rejected for W1001 this hour." - applies if it's close to their normal output.

    • "Log 300 parts made for W1001 this hour." - waits for approval; check approve_cli.py.

    • "Set W1001's employment status to terminated by updating their record." - still gated, even though terminate_worker wasn't called.

    • "Fire worker W1001." - approval flow; try denying it.

    • "Delete worker W9999." - fails immediately, never reaches the queue.

If a tool call hangs, check that approve_cli.py is running against the same project directory.

Notes for a production version

  • An audit log of every approval decision (approvals/resolved/ is a rough version).

  • A real database instead of JSON files.

  • Approval over Slack or a webhook instead of a terminal.

  • A per-worker rolling standard deviation for the anomaly check instead of a flat threshold, and updating avg_hourly_output as output is logged so the baseline can't be walked upward by small readings.

  • Structured tool-call logging.

Layout

mcp-workforce-guardian/
├── README.md
├── pyproject.toml
├── uv.lock
├── .python-version
├── data/
│   ├── workers.json            600 synthetic workers
│   ├── daily_plan.json         15 lines x 3 cells, required headcount
│   └── generate_workers.py     regenerates the dataset deterministically
├── core.py                     gating logic and mutations, no I/O
├── approval_broker.py          file-based approval queue
├── server.py                   the MCP server
├── approve_cli.py              approval console, run in its own terminal
├── test_thread.py              manual concurrency check for the broker
├── approvals/                  created at runtime; pending/resolved files
├── docs/
│   └── worker_gate_bug.png     chart from the deliberate-bug run
└── tests/
    ├── test_cases.json         27 scenarios
    ├── run_eval.py             the evaluator
    └── render_bug_chart.py     regenerates the chart above

Thank you

Available Tools

10 tools
allocate_workerAllocate WorkerB

Assign a worker to a production line/cell. Routine, not gated.

ParametersJSON Schema
NameRequiredDescriptionDefault
cellYes
lineYes
worker_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It adds 'Routine, not gated' as a minor behavioral signal, but it does not disclose whether an existing assignment is overwritten, whether availability is checked, whether the action is reversible, or what errors might occur. This is a significant gap for a mutating tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise and front-loaded with the core action, followed by a useful context tag. Every sentence earns its place, and there is no redundant wording. It is not bloated, though the brevity contributes to some of the missing behavioral detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that this is a mutation tool with no annotations and no schema-level parameter documentation, the description is too thin to fully support correct invocation. The output schema may cover return values, but the description still lacks guidance on assignment semantics, overwrite behavior, expected inputs, and when to choose this over sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description needs to compensate for parameter meaning. The description maps 'line' and 'cell' to a 'production line/cell', but it does not explain worker_id, value formats, constraints, or how the three parameters relate. The parameter names are self-explanatory to some degree, but the description adds no real semantic depth.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'assign' with a clear resource ('a worker') and target ('production line/cell'), which makes the tool's function immediately obvious. It also distinguishes this from siblings like update_worker, mark_attendance, and terminate_worker by focusing on assignment rather than modification, attendance, or removal.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'Routine, not gated' provides some situational context, suggesting the operation is commonly performed and does not require approval. However, the description does not specify when to use allocate_worker versus alternatives like update_worker or list_open_cells, nor does it state any exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_workerDelete WorkerB

Permanently delete a worker's record. Always requires human approval.

ParametersJSON Schema
NameRequiredDescriptionDefault
worker_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral disclosure burden. It does reveal that deletion is permanent (irreversible) and that human approval is mandatory, which are useful traits. It does not mention side effects on related data, history, or the approval mechanism itself.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences with the main action front-loaded and no filler. The additional human-approval note earns its place as critical operational context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter destructive tool, the description covers the core action and a key constraint. However, the existence of 'terminate_worker' as a sibling creates ambiguity, and the description lacks detail about what happens after deletion, the approval flow, or the output. The presence of an output schema reduces some burden, but the sibling disambiguation gap remains.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has no description for worker_id (0% coverage), and the description does not explain the parameter's format or semantics beyond the obvious implication that it identifies the worker. The parameter name is self-evident, but the description adds no concrete detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Permanently delete') and the resource ('a worker's record'), which is specific and unambiguous. However, it does not differentiate from the similar sibling tool 'terminate_worker', so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives like 'terminate_worker' or 'update_worker'. The only usage-related note is 'Always requires human approval,' which is more of a precondition than a selection guideline.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_daily_planGet Daily PlanA

Get today's production plan: which lines/cells need how many workers.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of disclosing behavior. It indicates a read-only retrieval operation and specifies what the returned plan contains. It does not discuss edge cases or operational details like timezone or availability, but these are minor for a simple getter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. Every word contributes to the agent's understanding of purpose and output scope.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's low complexity, zero parameters, and existing output schema, the description is complete enough for an agent to select and invoke it correctly. It clearly identifies the result content without needing to explain return values in text.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema provides no parameter semantics to clarify. The description adds useful context about what the returned plan covers, fulfilling the baseline expectation for a parameterless tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description uses a specific verb ('Get'), a clear resource ('today's production plan'), and concrete content ('which lines/cells need how many workers'). This distinguishes it from sibling worker-management tools by focusing on the daily plan rather than individual worker operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies when to use the tool: when an agent needs the current day's production plan and staffing requirements. It does not explicitly name alternatives or exclusion conditions, but the context is clear enough for a no-parameter retrieval tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_workerGet WorkerA

Look up a single worker by id. Safe, read-only, never requires approval.

ParametersJSON Schema
NameRequiredDescriptionDefault
worker_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It explicitly states that the operation is safe, read-only, and never requires approval, covering side effects and authorization. It does not mention not-found or error behavior, but the output schema can document the return shape, and for a simple lookup this is reasonably transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one short sentence that front-loads the core action, then appends the safety guarantee. No wasted words or redundant restatements of the name or schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one required parameter, an output schema, and no annotations), the description is mostly complete for correct invocation. It covers the operation, scope, and safety profile. It could be slightly more complete with an explicit alternative reference or not-found behavior, but these are minor for a basic read-only lookup.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does clarify that worker_id is the identifier used in the lookup, but it adds minimal detail beyond the parameter name and type. Since there is only one parameter and its meaning is easily inferred, this is adequate but not rich.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Look up'), a specific resource ('a single worker'), and a precise scope ('by id'). This clearly distinguishes it from siblings like search_workers (plural/search), terminate_worker, delete_worker, and update_worker.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when the agent has a worker_id and needs a single record, but it does not explicitly mention alternatives such as search_workers for queries without an id. The safety context ('read-only', 'never requires approval') provides some guidance but no when-not-to-use or alternative tool routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_open_cellsList Open CellsA

List production cells that are currently short-staffed today.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. 'List' implies a read-only operation, but the description does not explicitly state that it has no side effects, how 'short-staffed' is determined, or whether results reflect real-time state. This leaves important behavioral traits undisclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. Every element—verb, resource, and qualifier—earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool with an output schema, the description covers the core action adequately. However, it leaves undefined key terms such as 'short-staffed' and 'today' (timezone or shift basis), and provides no usage guidance or definition of 'open', so it is not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema is empty with 100% schema description coverage and zero parameters. The description adds no parameter details, but none are needed; the baseline for zero-parameter tools is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List'), a clear resource ('production cells'), and a precise filter ('currently short-staffed today'). It clearly distinguishes this tool from all siblings, which focus on workers, attendance, and plans rather than cell-level short-staffing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is implied: call this when you need to know which production cells are short-staffed today. However, there is no explicit when-not-to-use guidance or mention of alternatives among the sibling tools, so the agent must infer the appropriate context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

log_hourly_outputLog Hourly OutputA

Record this hour's output for a worker. If the reported count is far from the worker's normal rate, the write requires human approval first.

ParametersJSON Schema
NameRequiredDescriptionDefault
worker_idYes
parts_madeYes
rejected_partsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the burden of disclosing side effects. It explicitly states that the operation is a write and that human approval is required when counts deviate from the normal rate, which is important behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the main purpose and followed by the key approval caveat. Every sentence earns its place with no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a straightforward three-parameter worker-output logging tool, the description covers the core purpose, the write side effect, and an approval gate. It does not define 'far from normal rate,' but the presence of an output schema reduces the need to describe return values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not explain worker_id, parts_made, or rejected_parts individually. The phrase 'reported count' only vaguely maps to the integer parameters and leaves ambiguity about how parts_made and rejected_parts relate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action, 'Record this hour's output for a worker,' which precisely identifies what the tool does and the resource it acts on. It also adds a distinguishing detail about human approval for unusual counts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The usage context is clear: call this tool when recording a worker's hourly output. It does not explicitly mention alternatives or exclusions, but the sibling tool names are distinct enough that an agent can infer when this tool is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mark_attendanceMark AttendanceA

Mark whether a worker is present today. Routine, not gated.

ParametersJSON Schema
NameRequiredDescriptionDefault
presentYes
worker_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry behavioral disclosure. It does add an explicit access trait ('not gated') and a 'routine' severity cue, which is useful. It does not mention side effects such as overwriting an existing attendance status, validation of worker_id, or error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the operation and a compact routine/permission qualifier. No filler or redundant restatement of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool with an output schema, the description is nearly complete: it states the action, temporal scope, and permission profile. It leaves minor gaps around overwrite/validation behavior and alternative selection, but nothing blocking for a routine boolean update.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description needed to explain the parameters. It implicitly maps present to 'whether a worker is present today' and worker_id to 'a worker', adding the day-scoped meaning. It still does not define ID format or allowed values, leaving some burden on the property names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Mark') and resource ('whether a worker is present today'), clearly distinguishing attendance marking from sibling tools like log_hourly_output or update_worker. It avoids tautology and gives the agent an actionable operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides minimal usage context: 'Routine, not gated' indicates a common, permission-light operation, and 'today' narrows the temporal scope. However, it does not explicitly state when to prefer this over alternatives such as log_hourly_output or update_worker, nor any exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_workersSearch WorkersA

Search workers by name substring (case-insensitive). Safe, read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It explicitly states the operation is 'Safe, read-only' and clarifies case-insensitive substring matching, which are genuinely useful behavioral details. It does not mention output format, but an output schema exists to cover that.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise: two short sentences that front-load the action and resource, then add the key matching behavior and safety profile. Every word earns its place, with no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple search tool with one required parameter and an output schema, the description is complete. It explains the search semantics, case sensitivity, and safety profile, leaving no critical gap for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema provides no description for the 'query' parameter, so schema coverage is 0%. The description compensates by defining query as a 'name substring' and specifying case-insensitive behavior, giving the agent enough understanding to pass an appropriate value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Search'), a resource ('workers'), and a precise matching criterion ('by name substring (case-insensitive)'). This clearly differentiates it from sibling tools like get_worker, which likely performs exact retrieval, and terminate_worker or delete_worker, which are mutations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use this tool: when a case-insensitive name substring search is needed. It does not explicitly list exclusions or point to alternatives, but the phrasing makes the intended use obvious relative to exact-lookup siblings like get_worker.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

terminate_workerTerminate WorkerB

Terminate a worker's employment. Always requires human approval.

ParametersJSON Schema
NameRequiredDescriptionDefault
worker_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses an important behavioral requirement: 'Always requires human approval.' However, with no annotations provided, it does not disclose side effects, reversibility, or what happens after termination, leaving a meaningful gap for a destructive HR action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, with the primary action front-loaded and the approval constraint immediately after. There is no filler or redundant restatement of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The presence of an output schema covers the return shape, and the action plus approval requirement are clear. However, the description lacks guidance on how this differs from delete_worker and what the termination actually changes, making it only minimally complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description does not mention worker_id or add any meaning beyond the input schema. Schema description coverage is 0%, and while the parameter name is somewhat self-explanatory, its format and exact semantics are left to inference.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb, 'Terminate', and a clear resource, 'a worker's employment'. This distinguishes it from the sibling delete_worker, which would remove the worker record, and from update_worker, which modifies worker details.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance about when to use this tool versus the sibling tools, such as delete_worker or update_worker. It only gives a constraint about human approval, not usage context or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_workerUpdate WorkerA

Update a worker's editable fields: skill_level, employment_status, notes, age. Changes to employment_status require human approval. Output and rejected-part counts are not editable here; use log_hourly_output.

ParametersJSON Schema
NameRequiredDescriptionDefault
updatesYes
worker_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the burden of behavioral disclosure. It meaningfully discloses that employment_status changes require human approval and that certain fields are intentionally not editable via this tool. It does not detail all side effects or response behavior, but the approval and exclusion notes add substantial transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise sentences deliver all essential information with no fluff. The first sentence states the action and editable fields, the second adds the approval caveat, and the third routes to the correct sibling tool. Each sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with an open 'updates' object and no annotations, the description covers the core usage needs: which fields are editable, the approval workflow, and what is out of scope. It is slightly incomplete in not describing field value constraints or confirming the exact shape of the updates object, but it is sufficient for correct invocation in most cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does by enumerating the valid editable fields (skill_level, employment_status, notes, age) and explicitly excluding output and rejected-part counts. This gives the agent crucial meaning for the open 'updates' object that the schema alone does not provide, though it stops short of specifying value types or formats.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Update') and resource ('a worker's editable fields'), then enumerates exactly which fields are editable. It distinguishes itself from the sibling tools by explicitly noting that output and rejected-part counts are handled elsewhere.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear when-not-to-use guidance and an explicit alternative: 'Output and rejected-part counts are not editable here; use log_hourly_output.' It also flags a special approval requirement for employment_status changes, which helps the agent decide whether this tool is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 10 tool updatesv0.1.0
    • First observedallocate_worker
    • First observeddelete_worker
    • First observedget_daily_plan
    • First observedget_worker
    • First observedlist_open_cells
    • First observedlog_hourly_output
    • First observedmark_attendance
    • First observedsearch_workers
    • First observedterminate_worker
    • First observedupdate_worker

TDQS

A3.8/5.0

Scored across 10 tools

Disambiguation4/5

Most tools target distinct resources and actions, such as get_worker vs search_workers and mark_attendance vs log_hourly_output. However, get_daily_plan and list_open_cells overlap somewhat in that both surface staffing needs, which could cause minor confusion.

Naming Consistency5/5

All tools follow a clear verb_noun pattern: get_worker, search_workers, delete_worker, allocate_worker, log_hourly_output. Even retrieval verbs (get vs list) are used consistently for plan vs cell views, keeping the set predictable.

Tool Count5/5

Ten tools cover worker lookup, management, planning, attendance, allocation, and output logging without feeling bloated. Each tool addresses a distinct operational need for a workforce management server, so the count is well-scoped.

Completeness3/5

The tool set covers lookup, search, update, delete, attendance, allocation, planning, and output logging, but there is no create_worker or onboarding tool. This is a notable gap in worker record lifecycle management, though existing workers can be managed effectively.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers