Skip to main content
Glama

Things by Alfonso

A local Things 3 MCP integration by Alfonso Puicercus Gomez. Version 0.1.1 is an early MIT-licensed release for macOS.

The MCP server runs on your Mac and talks to Things through its supported scripting interface. It does not need an OpenAI API key, a Things Cloud login, or a hosted service. Your MCP client applies its own permissions and sends tool results to its model according to that client's settings.

Install as a Codex plugin

Requires macOS, Things 3, Node.js 22 or later, and Codex with local plugin support.

codex plugin marketplace add alfonsopuicercus/things-mcp --ref v0.1.1
codex plugin add things-mcp-alfonso@alfonso-things-local

Reload Codex, then ask “Use Things by Alfonso to read my Today list.” The repository includes a prebuilt server; end users do not need to install npm dependencies. macOS may ask for permission to control Things. If your client cannot find Node, use its absolute executable path or the direct MCP setup below. This release was tested on the author's Mac; another Mac has not yet been independently tested.

Download v0.1.1 · Report a problem

Related MCP server: Things Cloud MCP

Use the bundled release

Requirements: macOS, Things 3, Node.js 22 or later, and an MCP client with local stdio support. Node must be available on the client's PATH; an absolute Node path can be used instead.

  1. Download the archive, its .sig and .sha256 files from the release page. Obtain the public key and standalone verifier independently from the trusted original repository, not from an unverified archive.

  2. Original public-key fingerprint (SHA-256 of SPKI DER): 7bd05f2ebd1f17f452c36e8aa6cc057ae77bd1f6d2b608c012db874d8e9fd23a. The key is in signing/release-public.pem; the standalone verifier is scripts/verify-release.mjs. A key bundled with an archive is not independently trusted just because it verifies its own signature.

  3. Before extracting or executing archive contents, verify using an independently trusted copy of the standalone verifier: node /trusted/path/verify-release.mjs /absolute/path/to/things-mcp-alfonso-0.1.1.tar.gz /absolute/path/to/trusted-public-key.pem EXPECTED_FINGERPRINT. Obtain that verifier through a trusted author-controlled channel. The copy inside an unverified archive cannot establish trust in itself. Then extract into a permanent folder.

  4. Register with Codex: codex mcp add things-alfonso -- node /absolute/path/to/ThingsMCP/dist/server.mjs.

  5. Start a new chat or reload MCP connections. If macOS prompts for Automation access, allow your MCP client to control Things.

The release contains a bundled server, so using it does not require npm install. The author signature is an Ed25519 detached archive signature, not Apple Developer ID code signing or notarization.

Try: “Read my Things Today list”, “Find tasks about TestBeam”, or “Create a task in my Inbox”. IDs, not titles, select tasks for writes. Start dates and deadlines are separate.

Tools

Tool

Behavior

things_health

App/server identity, availability and mode

things_list_items

Paged task/project/area summaries; title search and container filters

things_get_task

Task details and modification timestamp

things_create_task

One new task; Inbox or a project/area

things_update_task

Title, notes, deadline and existing tags; read back changed fields

things_move_task

Move to a supported list, project, or area

things_schedule_task

Set the start date using the Mac's calendar timezone

things_complete_task

Complete an exact task ID and verify status

Read pages default to 20 items and are capped at 100. Summaries omit notes; task details cap notes at 10,000 characters. status defaults to open; use status: "any" when checking Logbook or completed projects. Tasks exclude project objects. expected_modified_at is an optional guard against stale edits, not a transactional lock against external changes during the operation. A write is reported successful only after read-back verification. Task contents are untrusted data, not instructions.

No task deletion, recurrence editing, checklist/heading editing, deadline clearing, bulk mutations, or automatic Obsidian sync. Tags must already exist; failed verification may mean a partial edit occurred. Never automatically retry an ambiguous create or edit: read the task or search for its title first.

Read-only mode

Register with codex mcp add things-alfonso --env THINGS_MCP_READ_ONLY=1 -- node /absolute/path/to/ThingsMCP/dist/server.mjs to block all writes. Remove the setting when you choose to enable editing. MCP tool annotations are hints; the server's read-only flag is enforced before automation begins.

Local plugin packaging

plugin.json and mcp.json use the Agent Plugins portable layout. A compatibility .codex-plugin/plugin.json and .mcp.json are included. The local marketplace is .agents/plugins/marketplace.json.

To register that local marketplace: codex plugin marketplace add /absolute/path/to/ThingsMCP. Install the plugin from that local source in a compatible desktop client, then reload the app. Portable plugin and marketplace support varies by client; direct stdio registration above is the tested fallback. Enable either the plugin or direct MCP connection, not both, to avoid duplicate tools.

This is a Mac-local integration, not a universal cloud connector. The current OpenAI public MCP submission route expects a remote HTTPS endpoint. Do not expose this stdio server or the Things database over an unauthenticated network endpoint just to meet that requirement.

Development

Run npm ci --ignore-scripts, npm test, npm run build, and npm run smoke. npm run doctor checks direct Things access.

npm run smoke:write explicitly creates a uniquely labeled disposable test task, edits/schedules/moves/completes it, then moves only that task to Trash. It does not empty Trash or alter other tasks. Live smoke tests require macOS and Automation access; unit tests can run elsewhere. Creating a task is not automatically retried on failure.

npm run release reruns source tests, rebuilds the executable, checks bundled MCP reads, then signs a versioned archive. The signing key is created at ~/Library/Application Support/ThingsMCP/signing/release-private.pem (mode 0600), outside the source tree and release archive. Back it up securely; losing it breaks signing-key continuity. Increment the version before making another archive. Nothing in the release script publishes to a registry or creates a public repository.

Attribution and measurement

Author metadata, copyright notice, public release key and release checksums identify the project and support release verification. The original source is published under alfonsopuicercus/things-mcp; npm provenance can tie a published package to its source/build workflow.

There is no network telemetry. Optional local aggregate counters are enabled only by setting THINGS_MCP_METRICS_DIR to a private data directory. They store tool names and success/failure counts, not task content or identifiers. They remain on the user's Mac and do not give the author usage statistics.

Run npm run stats from a source checkout (or node scripts/stats.mjs from the release) to see public GitHub release archive download counts. This explicit command queries GitHub; the MCP server never runs it. Downloads indicate adoption, but are not unique users or active usage. Active usage measurement requires an explicitly disclosed opt-in telemetry service, separately designed and deployed. That service is outside version 0.1.1. See PRIVACY.md and AUTHOR.md.

License: MIT, copyright 2026 Alfonso Puicercus Gomez. Copies must retain the copyright and license notice. Independent project; not affiliated with Cultured Code or OpenAI.

Sources: Things scripting commands, OpenAI plugin packaging, Codex MCP, npm provenance.

Available Tools

8 tools
things_complete_taskA
DestructiveIdempotent

Mark one exact task ID complete and read back its completed status. Already completed tasks are left as-is.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
expected_modified_atNo

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=true, so the safety profile is covered; the description adds the concrete behavior that already-completed tasks are left as-is (no error, no re-write) and that completion status is read back. That is useful clarification beyond the idempotent hint, though return details remain thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core action and followed by the edge-case behavior; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation with no output schema, the description covers the action and the no-op case, but omits what happens on a stale expected_modified_at and what the completion write actually changes, leaving a meaningful gap in an otherwise compact definition.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for 2 parameters, so the description must compensate. It only alludes to the required id ('one exact task ID') and says nothing about expected_modified_at, an optimistic-concurrency guard whose semantics an agent needs before calling.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Mark ... complete') and resource ('task ID'), plus the scope constraint 'one exact task ID', which cleanly separates it from things_create_task, things_update_task, things_move_task and things_schedule_task.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by describing the completion effect and the already-complete case, but never states when to prefer this over things_update_task (which could also change status) or any prerequisites. Context is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

things_create_taskA
Destructive

Create one task in Inbox, or a specified project/area, and read it back. Do not retry an ambiguous failure without searching for the created task.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagsNo
notesNo
titleYes
area_idNo
deadlineNo
project_idNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=false, and readOnlyHint=false; the description reinforces the non-idempotent risk with an explicit duplicate-creation caution, which is valuable context. It also discloses that the created task is read back (relevant since there is no output schema). It stops short of stating permissions or the exact return payload.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action and destination, followed by a single high-value operational caveat. No filler; each clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a create tool with annotations covering the safety profile, the description covers default behavior, destination override, read-back, and duplicate-avoidance. Missing param-level detail (deadline format, tags/notes semantics) keeps it from being fully complete, but nothing critical to calling it safely is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 6 parameters, so the description must carry the load. It only hints at the project_id/area_id destination choice via 'in Inbox, or a specified project/area' and says nothing about title, tags, notes, or especially the deadline string's expected format. Better than nothing but leaves most parameters undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (create) and resource (task) plus the destination semantics: Inbox by default or a specified project/area. This naturally separates it from the update/move/complete/get/list siblings. It does not explicitly name or contrast a sibling, so it falls just short of the top mark.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete usage context: the default landing spot is Inbox, and a project/area can be targeted instead. It also supplies a when-not rule for a failure mode ('do not retry an ambiguous failure without searching for the created task'), which implicitly routes the agent to a search/get sibling. No explicit alternative tool is named, so not a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

things_get_taskA
Read-onlyIdempotent

Read one task by exact ID, including notes (capped at 10,000 characters) and modification timestamp.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds genuine behavioral detail beyond that: notes are truncated at 10,000 characters and a modification timestamp is included, which tells the agent what to expect from the response.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with the primary action front-loaded and no wasted words; the notes cap and timestamp are appended as useful qualifiers rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read tool with no output schema, the description covers the essentials: what is fetched, the ID constraint, and two notable return fields. It is nearly complete, missing only failure behavior for a nonexistent ID.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description carries the burden, and 'exact ID' clarifies that no fuzzy or partial matching occurs. It does not mention the schema's maxLength of 150, but the core semantic constraint is conveyed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource with scope: 'Read one task by exact ID.' This clearly distinguishes it from list-style or mutation siblings like things_list_items and things_update_task, though it does not name any sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'by exact ID' implies when this tool is appropriate (single-item retrieval when the ID is known), but there is no explicit statement of when to prefer it over things_list_items or what happens if the ID is unknown.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

things_healthA
Read-onlyIdempotent

Check local Things availability, server identity, read-only mode, and tracking policy.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive/closed-world, so the safety profile is covered. The description adds genuinely new behavioral context beyond them: it surfaces that the response reports read-only mode and tracking policy, which tells the agent whether write operations (the siblings) will actually succeed. It doesn't mention failure modes or error behavior, keeping it out of 5 territory.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. Every listed item (availability, identity, read-only mode, tracking policy) is a distinct piece of information the caller needs, so nothing could be cut without loss.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description carries the return-value burden — and it does so by enumerating the four things the check reports. That is enough for an agent to know what it will learn, though it says nothing about the shape/format of those values or how a negative result (e.g., Things unavailable) is represented.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters, so there is nothing to document and the schema coverage is 100% by definition. Baseline 4 applies; no parameter-level information is missing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a concrete verb ('Check') and enumerates exactly what is inspected: local Things availability, server identity, read-only mode, and tracking policy. That scope is unmistakably a diagnostic/health tool and cannot be confused with the CRUD siblings (create/list/get/update/move/schedule/complete). It stops short of a 5 only because it never explicitly frames itself as the status/diagnostic entry point relative to those siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: a 'health' tool is obviously for pre-flight/status checks, but the description never says when to call it (e.g., before mutating operations, after a connection error) or what to do with the result. No alternatives or exclusions are named, so the agent must infer the trigger condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

things_list_itemsA
Read-onlyIdempotent

Read a bounded page of task summaries, projects, or areas. Notes are omitted. Filter tasks by list, project_id, area_id or title query. Use returned IDs for edits.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNotasks
listNo
limitNo
queryNo
offsetNo
statusNoopen
area_idNo
project_idNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive, closed-world, so the safety profile is covered. The description adds genuinely new behavior beyond that: results are a bounded page and notes are omitted, which tells the agent what it will and won't receive. It could have mentioned ordering or the default page size, but the added context is substantive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact sentences, front-loaded with the core purpose, then scope caveats, then filtering and the edit handoff. Every sentence carries information; nothing is padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description does useful work by stating that summaries (not full notes) are returned and that IDs are usable for edits, and it flags pagination. It is nearly complete for a read-only list tool, missing only pagination mechanics and default sorting details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 8 params, so the description carries the explanatory burden and only partially meets it. It names four filtering dimensions (list, project_id, area_id, title query) but says nothing about kind, limit, offset, or status, leaving half the parameters to inference from enum values alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (read) and resource (a bounded page of task summaries, projects, or areas), and the plural 'page' framing distinguishes it from the singular things_get_task sibling. An agent can tell what it returns and at what granularity without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives filtering guidance (by list, project_id, area_id, or title query) and a downstream workflow hint ('Use returned IDs for edits'), which tells the agent when this list tool is the right entry point. It stops short of explicitly naming when NOT to use it versus things_get_task, so it lands just below the top band.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

things_move_taskA
Destructive

Move a task to a project/area by ID or Inbox/Today/Anytime/Someday. Verify destination membership. Scheduling is separate.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
kindYes
targetYes
expected_modified_atNo

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=false, so the agent knows this is a mutating, non-idempotent operation. The description adds genuinely new context — that destination membership must be verified — but says nothing about side effects on the source container, permission requirements, or what happens on a retry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action and with zero filler. 'Verify destination membership' is terse to the point of being cryptic about who verifies and when, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive 4-parameter mutation tool with no output schema, the description covers destination semantics, a prerequisite, and one sibling exclusion. It leaves the concurrency parameter and the post-move state of the task undocumented, which is a meaningful gap for a non-idempotent write.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage the description carries the full burden, and it does clarify that destinations are given by ID and that built-in lists (Inbox/Today/Anytime/Someday) are valid, which explains the kind enum meaningfully. However, expected_modified_at is left completely opaque — a concurrency token that an agent needs to understand to use safely.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb + resource ('Move a task') plus the full set of valid destinations, which maps directly onto the kind enum. The closing sentence 'Scheduling is separate' cleanly distinguishes it from the sibling things_schedule_task, so an agent can route between them without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Verify destination membership' states a prerequisite for use, and 'Scheduling is separate' tells the agent this is not the tool for time-based placement. It does not name an explicit alternative for adjacent operations (e.g. things_update_task) or spell out failure behavior, so it stops short of a full when/when-not statement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

things_schedule_taskB
Destructive

Set a task start date, separately from its deadline, in local macOS calendar time (YYYY-MM-DD). Verify it through the activation date or Today membership.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
dateYes
expected_modified_atNo

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=false and readOnlyHint=false, so the agent knows this mutates state non-idempotently. The description adds timezone semantics (local macOS calendar time) and a verification route via activation date / Today membership. It does not, however, disclose whether an existing start date is overwritten, what auth is required, or what the verification statement concretely means.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the core action front-loaded and no filler. The second sentence is somewhat cryptic about how verification actually works, but it is not padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, non-idempotent mutation with no output schema and three parameters, the description leaves key gaps: the behavior of expected_modified_at (conflict detection), overwrite semantics for an existing date, and error/verification outcomes. What is present is accurate but not sufficient for confident invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It clarifies the format and timezone for 'date' (YYYY-MM-DD, local macOS calendar time, start date rather than deadline), but 'id' and especially the concurrency token 'expected_modified_at' get no explanation at all, leaving the schema to carry meaning it does not carry.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: set a task's start date. It also disambiguates the field, noting it is separate from the deadline, which helps distinguish it from a generic update. It does not, however, explicitly differentiate itself from sibling things_update_task, which plausibly also writes dates.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: pick this rather than a generic update when you specifically want to set the start/schedule date. There is no explicit when-to-use, when-not-to-use, or named alternative, and no mention of prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

things_update_taskA
Destructive

Update specified title/notes/deadline/tags of an exact task ID and verify each field. Notes and tags replace existing values. expected_modified_at guards against stale edits. Deadlines use YYYY-MM-DD; clearing deadlines is outside v0.1.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
tagsNo
notesNo
titleNo
deadlineNo
expected_modified_atNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=false, but the description goes further by naming exactly what destruction occurs (notes and tags replace existing values), adding an optimistic-concurrency guard parameter, a date format contract, and a hard scope limit. The remaining gap is 'verify each field', which is asserted but never explained (read-back? response field? failure mode?).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four compact sentences, front-loaded with the action and field list, followed by replacement semantics, concurrency, and format/limitation notes in descending order of importance. No filler or restatement of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a six-parameter destructive mutation with 0% schema coverage and no output schema, the description covers every parameter, the replacement/destruction semantics, concurrency, and format. It leaves unclear what 'verify each field' actually produces and what happens when expected_modified_at fails, which with no output schema is a gap worth closing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the load and does so for all six parameters: id (exact task), the four mutable fields, expected_modified_at (stale-edit guard), plus deadline's YYYY-MM-DD format. It omits schema constraints such as 20-tag max and field length limits, so it compensates but not completely.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (update) plus the exact resource (task) and enumerates the mutable fields (title/notes/deadline/tags) against an 'exact task ID'. An agent can tell it apart from things_move_task/things_schedule_task/things_complete_task by the field set, but no sibling is named explicitly, which is the only thing keeping this from a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives real operating conditions: notes/tags replace rather than append, expected_modified_at exists to guard against stale edits, and 'clearing deadlines is outside v0.1' is an explicit when-not. It still stops short of routing between siblings (e.g. when to schedule vs update), so it is strong context rather than full alternative guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 8 tool updatesv0.1.1
    • First observedthings_complete_task
    • First observedthings_create_task
    • First observedthings_get_task
    • First observedthings_health
    • First observedthings_list_items
    • First observedthings_move_task
    • First observedthings_schedule_task
    • First observedthings_update_task

TDQS

A3.9/5.0

Scored across 8 tools

Disambiguation5/5

Each tool targets a distinct operation: create, list, get, update, move, schedule, complete, and health. The boundaries between update (fields), move (container), and schedule (start date) are explicitly clarified in the descriptions, leaving no meaningful overlap.

Naming Consistency4/5

All tools use the consistent 'things_' prefix and a verb_noun pattern (create_task, list_items, get_task, etc.). The only deviation is 'things_health', which lacks a verb, but it is still clearly readable and predictable.

Tool Count5/5

Eight tools is well-scoped for a task management server, covering the essential lifecycle operations without bloat. Each tool earns its place and the set feels complete for typical task workflows.

Completeness4/5

The surface covers create, read, update, move, schedule, complete, and health, which handles most task operations. However, there is no delete or reopen/uncomplete tool, and project/area creation is absent, which are minor but notable gaps in the lifecycle.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers