Skip to main content
Glama

mustdo-mcp

日本語版はこちら (README.ja.md)

MCP server for MustDo — the iOS To-Do alarm that keeps ringing until you do it.

It lets Claude (Claude Code, Claude Desktop, claude.ai, the Claude iPhone app) and any other Model Context Protocol client read and write your MustDo To-Dos.

  • Your data stays in your iCloud. MustDo has no backend of its own. To-Dos live in the CloudKit private database of your Apple ID (container iCloud.jp.lightning.mustdo, zone MustDo). This server talks to Apple's CloudKit Web Services and reads/writes the same records as the iPhone app.

  • The developer never stores your To-Dos. Neither the local server in this repository nor the hosted relay (see below) keeps To-Do content on Lightning LLC servers.

  • After a write, CloudKit pushes a silent notification to your iPhone, so the app updates right away.

Two ways to use it

A. Hosted relay (recommended)

B. Run this repository locally

Works from

claude.ai, Claude iPhone app, any client that supports remote MCP + OAuth

Claude Code / Claude Desktop on a Mac (stdio)

Setup

Add a custom connector, sign in with your Apple ID

Node 24, build, configure, sign in

What Lightning LLC stores

Your CloudKit sign-in token only, encrypted with AWS KMS (see below)

Nothing

A. Hosted relay — https://ltng.jp/api/mustdo/mcp

  1. claude.ai → Settings → Connectors → Add custom connector → URL https://ltng.jp/api/mustdo/mcp (name it "MustDo").

  2. Click Connect. You will see a consent page on ltng.jp explaining what is stored, then Apple's sign-in page. Sign in with the same Apple ID you use in the MustDo app.

  3. Done. The same connector is available in the Claude iPhone app.

What the relay keeps, honestly:

  • When you sign in, Apple issues a CloudKit sign-in token (ckWebAuthToken). The relay stores this token encrypted with AWS KMS (AWS Tokyo region) so it can call CloudKit on your behalf on each request.

  • It also stores hashed OAuth access/refresh tokens for the connector itself.

  • It does not store or log your To-Do content, your Apple ID email, your password, or your raw iCloud user ID.

  • Disconnect: remove the connector in claude.ai and visit https://ltng.jp/api/mustdo/disconnect. After confirming with your Apple ID, the stored token is deleted immediately. It is also deleted automatically when Apple invalidates the sign-in (the tools then return RECONNECT_REQUIRED; just reconnect).

Full write-up: https://ltng.jp/mustdo/mcp.

B. Run locally (stdio)

Requirements:

  • macOS with Node.js 24 or newer

  • The MustDo app installed and synced to iCloud at least once (the app creates the zone and the Account record)

  • A CloudKit API Token for the MustDo container — see the next section

git clone https://github.com/lightning-llc-jpn/mustdo-mcp mustdo-mcp
cd mustdo-mcp
npm install
npm run build        # -> dist/index.js
npm test             # vitest; CloudKit is mocked

Register with Claude Code:

claude mcp add mustdo \
  -e MUSTDO_CK_API_TOKEN=<MUSTDO_CK_API_TOKEN> \
  -e MUSTDO_CK_ENV=production \
  -- node /path/to/mustdo-mcp/dist/index.js

Or put the settings in ~/.mustdo/config.json and register without -e:

{
  "apiToken": "<MUSTDO_CK_API_TOKEN>",
  "environment": "production"
}

Then ask Claude to run the sign_in tool once. The server opens Apple's sign-in page in your browser, listens on http://localhost:51234/callback, and saves the returned token to ~/.mustdo/auth.json (mode 0600).

About the CloudKit API Token

CloudKit Web Services needs two tokens on every request:

Token

What it is

Who has it

ckAPIToken

Identifies the container (iCloud.jp.lightning.mustdo). Created in CloudKit Dashboard by the container owner. Apple designs it for use "from a website or an embedded web view", i.e. it is a client-side token with a fixed Sign-in Callback URL.

Lightning LLC (the container belongs to the MustDo developer team). You cannot create one yourself — CloudKit Dashboard only lets a team create tokens for its own containers.

ckWebAuthToken

Your personal CloudKit session, issued by Apple when you sign in with your Apple ID. This is the credential that actually grants access to your private database.

Only you. Stored in ~/.mustdo/auth.json; never leaves your Mac except to Apple.

Because the API Token is container-wide and cannot be created by end users, Lightning LLC provides the value for local use. Use the value shown on https://ltng.jp/mustdo/mcp as MUSTDO_CK_API_TOKEN (the token for the production environment has its Sign-in Callback set to https://ltng.jp/api/mustdo/oauth/local-callback, which simply redirects back to http://localhost:51234/callback without storing anything). If the page does not show a token, local use is not currently offered — use the hosted relay.

Configuration

Environment variables win over ~/.mustdo/config.json.

env

config.json key

default

meaning

MUSTDO_CK_API_TOKEN

apiToken

(required)

CloudKit API Token for the MustDo container

MUSTDO_CK_ENV

environment

development

production for App Store data. development is only useful for the MustDo developers

MUSTDO_CK_CONTAINER

container

iCloud.jp.lightning.mustdo

leave as is

MUSTDO_CK_ZONE

zoneName

MustDo

leave as is

MUSTDO_CALLBACK_PORT

callbackPort

51234

port the sign-in callback listens on

MUSTDO_HOME

—

~/.mustdo

where config.json and auth.json live

MUSTDO_LOG_LEVEL

—

INFO

DEBUG / INFO / WARN / ERROR (JSON lines on stderr)

auth.json is per environment; switching MUSTDO_CK_ENV requires another sign_in. Apple expires the session after a while (CloudKit returns HTTP 421); tools then return NOT_SIGNED_IN and you run sign_in again.

Related MCP server: iCloud CalDAV MCP Connector

Tools

All tools return JSON. Dates in output are ISO 8601 (UTC). Dates in input may be:

  • YYYY-MM-DD — that day at the account's default time (see below)

  • YYYY-MM-DDTHH:mm — wall-clock time in the account's time zone

  • Full ISO 8601 with offset

Tool

Arguments

What it does

sign_in

waitSeconds? (5–300, default 90)

Local only. Opens Apple's sign-in page and stores ckWebAuthToken. Not present on the relay.

get_me

—

Account info: timeZone, today, weekday, defaultTime, defaultDueNext, trial/subscription dates, canAddTodo. Call this before doing date math.

list_todos

date?, from?, to? (YYYY-MM-DD, inclusive), status? (pending default / done / skipped / all), includeRepeating? (default true)

Lists To-Dos. Deleted ones are excluded. Repeating To-Dos are returned as templates under repeating (not expanded) with any per-day occurrences in range.

add_todo

title (1–200), due?, repeat?, notes? (≤2000), sound?

Creates a To-Do with source = "mcp". If due is omitted, it is tomorrow at the default time.

update_todo

id, title?, due?, repeat? (object or null), notes? (string or null), sound?, status?

Changes only the fields you pass. repeat replaces the whole rule; null or { "kind": "none" } removes it.

complete_todo

id, occurrenceDate?

One-off: status = done. Repeating: marks that day's occurrence done (default: today).

snooze_todo

id, until, occurrenceDate?

Snoozes until until. Repeating: only that day's occurrence.

delete_todo

id

Soft delete (deletedAt). The app purges it after 14 days.

Default time and omitted due

Each account has a default alarm time (Account.defaultTime, HH:mm, set in the app's settings; 09:00 if unset).

  • add_todo with no due → tomorrow (in the account's time zone) at the default time. Month/year boundaries and DST transitions follow the wall clock.

  • due: "2026-10-10" → that day at the default time.

  • get_me returns defaultTime and defaultDueNext so a client can tell the user when the alarm will ring.

Examples:

// "Remind me to buy milk" → tomorrow at the default time
{ "title": "Buy milk" }

// "Call the dentist on the 10th" → that day at the default time
{ "title": "Call the dentist", "due": "2026-10-10" }

// "Today at 3pm"
{ "title": "Submit report", "due": "2026-10-06T15:00" }

Repeat rules

Same vocabulary as the iOS Calendar app. Used as input to add_todo / update_todo and returned by list_todos.

{
  "kind": "none | daily | weekly | monthly | yearly",
  "interval": 1,                                   // 1 = every, 2 = every other … (1–99)
  "weekdays": [2, 4],                              // weekly only. 1 = Sun … 7 = Sat. Empty → weekday of `due`
  "monthly": { "mode": "dayOfMonth | weekdayOrdinal", "ordinal": 1, "weekday": 2 },
  "end": { "kind": "never | until | count", "until": "2026-12-31", "count": 10 }
}

You want

repeat

Every day

{ "kind": "daily" }

Every Mon & Wed

{ "kind": "weekly", "weekdays": [2, 4] }

Every other week

{ "kind": "weekly", "interval": 2 }

Every 3 days

{ "kind": "daily", "interval": 3 }

5th of every month

{ "kind": "monthly" } with due on the 5th (29–31 fall back to month end)

First Monday of every month

{ "kind": "monthly", "monthly": { "mode": "weekdayOrdinal", "ordinal": 1, "weekday": 2 } }

Last Friday of every month

{ "kind": "monthly", "monthly": { "mode": "weekdayOrdinal", "ordinal": -1, "weekday": 6 } }

Every year

{ "kind": "yearly" } (month/day of due; Feb 29 → Feb 28 in non-leap years)

10 times, then stop

{ "kind": "daily", "end": { "kind": "count", "count": 10 } }

Until Dec 31

{ "kind": "weekly", "weekdays": [2], "end": { "kind": "until", "until": "2026-12-31" } }

Omitted fields take defaults (interval 1, weekdays [], monthly.mode dayOfMonth, end.kind never). Ranges: interval 1–99, ordinal 1–5 or -1, weekday 1–7, count 1–999. Anything else → INVALID_ARGUMENT. The legacy shape { "kind", "weekdays", "until" } is still accepted.

Expansion happens in the app, not here. list_todos returns the template plus occurrences (per-day done / skipped / snooze). Range filtering only drops templates that definitely cannot fire in range (first occurrence after the range, until before the range, weekly with no matching weekday).

Errors

Failures come back with isError: true and a body of { "code": "...", "message": "...", "details"?: {...} }.

code

Meaning

NOT_SIGNED_IN

No valid ckWebAuthToken (local). Run sign_in.

RECONNECT_REQUIRED

Same, on the hosted relay. Reconnect the connector in claude.ai.

NOT_CONFIGURED

API Token missing or rejected by CloudKit (HTTP 401/403).

PAYMENT_REQUIRED

The account's trial has ended and there is no active subscription. Only add_todo is affected.

NOT_FOUND

No such To-Do, or it was deleted.

INVALID_ARGUMENT

Bad date, out-of-range repeat rule, empty title, etc.

CONFLICT

Another device changed the record twice in a row. Retry.

CLOUDKIT_ERROR

Any other CloudKit error (e.g. zone missing because the app has never synced).

SIGN_IN_TIMEOUT

The 10-minute sign-in listener expired.

INTERNAL

Unexpected error.

Security model

  • Access to your To-Dos is gated by your own Apple ID session (ckWebAuthToken), issued by Apple's sign-in page. No one — including the developer — can read your private database without it.

  • The API Token only identifies the container and fixes where Apple may redirect after sign-in. Apple positions it as a client-side token (it is normally embedded in CloudKit JS web pages). By itself it grants no access to any user's private data.

  • Locally, the session is stored in ~/.mustdo/auth.json (0600), logs go to stderr as JSON and never include tokens, and the sign-in listener binds to 127.0.0.1 / ::1 only.

  • On the relay, the session is encrypted with AWS KMS per user (envelope encryption with encryption context), OAuth tokens are stored only as peppered SHA-256 hashes, PKCE S256 is mandatory, refresh tokens rotate with reuse detection, and To-Do content is never written to storage or logs.

  • Writes use CloudKit recordChangeTag (optimistic locking) and retry once on conflict; deletes are soft.

  • Anything you find: see SECURITY.md.

Development

npm install
npm run build       # tsc → dist/
npm test            # vitest (CloudKit mocked in test/fakeCloudKit.ts)
npm run typecheck   # tsc --noEmit
MUSTDO_LOG_LEVEL=DEBUG node dist/index.js   # run the stdio server by hand

Layout:

src/
  index.ts      entry point (stdio)
  server.ts     wiring for stdio: sign_in + the 7 To-Do tools
  tools.ts      tool definitions (zod schemas) shared by stdio and the relay
  core.ts       public entry for the shared core (no auth/config/stdio)
  service.ts    tool logic: filtering, upserts, conflict retry
  cloudkit.ts   thin CloudKit Web Services client (query / lookup / modify, 421 handling)
  records.ts    CloudKit record <-> model conversion
  dates.ts      time-zone-aware date math using Intl only
  auth.ts       auth.json and the sign-in callback listener
  config.ts     env / ~/.mustdo/config.json
  model.ts      enums, allowlists, RepeatRule types (mirrors the Swift app)
  errors.ts     ToolError / NotSignedInError
  log.ts        JSON logs to stderr (setLogSink to redirect)
test/           vitest

dist/core.js (package exports) is the shared core consumed by the hosted relay: everything except auth.ts, config.ts, index.ts and server.ts. Keep it free of anything that touches the local file system or a browser.

License

MIT — Copyright (c) 2026 Lightning LLC. See LICENSE.

MustDo is a product of Lightning LLC. Apple, iCloud and CloudKit are trademarks of Apple Inc.

Available Tools

8 tools
add_todoTODO を追加A

TODO を追加する(source=mcp)。ひとことで TODO を言われたら title だけ渡す。due を省略すると明日の既定時刻(Account の timeZone での今日の翌日、Account.defaultTime、既定 09:00)になる。「明日 9 時」「今日 15 時」など利用者が時刻を言ったときだけ due を渡す(YYYY-MM-DDTHH:mm は Account の timeZone の時計、ISO8601 も可)。日付だけ言われたら due に YYYY-MM-DD を渡すと「その日の既定時刻」になる。繰り返しは repeat で(iOS カレンダーと同じ語彙: 毎日 / 毎週 月水 / 隔週 / 毎月 5 日 / 毎月 第 1 月曜 / 最終金曜 / 毎年 / 3 日ごと / 10 回で終了 / 12/31 まで)。初回は due(毎月 5 日なら due を 5 日に、毎年なら due の月日)。試用(30 日)も購読も切れていると isError + code=PAYMENT_REQUIRED。

ParametersJSON Schema
NameRequiredDescriptionDefault
dueNo鳴らす時刻。YYYY-MM-DDTHH:mm(Account の timeZone の時計)か、オフセット付き ISO8601。YYYY-MM-DD だけなら「その日の既定時刻(Account.defaultTime、既定 09:00)」。
notesNo
soundNoアラーム音。default, chime, marimba, beep, urgent, vibrateOnly, siren, klaxon, rapid, redalert, musicbox, harp, morning, bowl, droplet, breeze, horror, dread, ghost, lament, cello, rainy, fanfare, skip, sparkle
titleYes
repeatNo繰り返し(iOS カレンダーと同じ語彙)。省略した項目は既定値(interval 1・weekdays []・monthly.mode dayOfMonth・end.kind never)。範囲外は INVALID_ARGUMENT。例: 毎日 {kind:daily} / 毎週 月水 {kind:weekly, weekdays:[2,4]} / 隔週 {kind:weekly, interval:2} / 3 日ごと {kind:daily, interval:3} / 毎月 5 日 {kind:monthly}(due を 5 日に)/ 毎月 第 1 月曜 {kind:monthly, monthly:{mode:weekdayOrdinal, ordinal:1, weekday:2}} / 毎月 最終金曜 {kind:monthly, monthly:{mode:weekdayOrdinal, ordinal:-1, weekday:6}} / 毎年 {kind:yearly}(due の月日)/ 10 回で終了 {kind:daily, end:{kind:count, count:10}} / 12/31 まで {kind:weekly, weekdays:[2], end:{kind:until, until:"2026-12-31"}}

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses the default due behavior (next day at Account.defaultTime, default 09:00), timezone semantics, default repeat values, and a failure condition (trial/subscription expired → isError + code=PAYMENT_REQUIRED). It stops short of stating authentication requirements (implied by sign_in) or any rate limits, so it is not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded and no sentence is pure filler, but the definition is a single long run-on paragraph mixing due defaults, repeat vocabulary, and payment failure. Bullets or short sections would improve scanability for an agent parsing it, though the information density is justified by the nested repeat schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter tool with a nested repeat object and no output schema, the description covers the tricky invocation aspects (due defaults, recurrence encoding, payment-required error). It omits auth prerequisites and says nothing about the response, but with no output schema defined the latter is less critical. Adequate for correct invocation given the schema's own descriptions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 60%, and the description compensates heavily for the complex parameters: it explains due formats and the date-only default-time behavior, and translates natural-language recurrence phrases into repeat object shapes (毎週 月水 → weekdays [2,4], 隔週 → interval 2, etc.). It says nothing about notes or sound, which remain schema-documented only, but adds real meaning beyond the schema for the high-complexity fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource (「TODO を追加する」) and adds the creation source (source=mcp). It is trivially distinguishable from siblings like update_todo, complete_todo, snooze_todo, and delete_todo, which all operate on existing TODOs. No ambiguity about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Rich guidance on optional parameters: pass only title for a bare TODO; pass due only when the user states a time (「明日 9 時」「今日 15 時」); pass YYYY-MM-DD for date-only; how to express recurrence. It does not, however, route among sibling tools or state exclusions (e.g., when to use update_todo instead), so it falls short of explicit alternative-selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

complete_todoTODO を完了A

やった、にする。単発は status=done と completedAt。繰り返しは occurrenceDate(YYYY-MM-DD、省略で TODO の timeZone での今日)の回を Occurrence に done で記録する。

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
occurrenceDateNo繰り返しのどの回か。省略で今日

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full burden and does well: it discloses the exact mutations for each case (status=done + completedAt for one-off; a done Occurrence record keyed by occurrenceDate for recurring). It omits idempotency, error states (e.g., already-completed), and auth requirements, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The action ('やった、にする') is front-loaded, and the remaining clauses each earn their place by covering the two execution branches. It is dense but free of filler; slightly telegraphic phrasing is the only minor cost.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-parameter mutation tool with no annotations and no output schema, the definition covers the important complexity: the divergence between one-off and recurring completion, plus the optional date semantics. Error/edge behavior and return expectations are the only meaningful gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: occurrenceDate is documented, id is not. The description adds real value beyond the schema by specifying the timeZone-aware default ('today in the TODO's timeZone'), which the schema's terse '省略で今日' does not clarify. It adds nothing for the required id parameter, so a 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb+resource ('mark as done') and goes further by splitting the behavior into one-off vs. recurring cases. It implicitly separates this from update_todo via completion semantics, but never names or contrasts a sibling explicitly, so it lands at 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the completion semantics and the single-vs-recurring branch, which tells an agent which code path applies. However, there is no guidance on when to prefer complete_todo over update_todo or snooze_todo, and no stated preconditions or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_todoTODO を削除B

deletedAt を書く soft delete。アプリ側で 14 日後に物理削除される。

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden and does add real value: it discloses the soft-delete mechanism and the 14-day retention before physical purge. It omits other relevant behavior such as idempotency, what happens when deleting an already-deleted TODO, or any permission requirements, leaving meaningful gaps for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short, front-loaded sentences with zero filler; the core behavior (soft delete via deletedAt) leads and the follow-on cleanup detail follows immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple (one required param, no output schema, no annotations), and the description does cover the key lifecycle behavior. Still, it says nothing about the id parameter, return value, or failure modes, and with no annotations those gaps are not covered elsewhere.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one parameter (id) with 0% schema description coverage, so the description should compensate but never mentions the identifier at all. It does not clarify what the id refers to or its expected format, leaving the schema's bare string entirely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific operation (soft delete) and its mechanism (writing deletedAt) for the TODO resource. However, it never distinguishes itself from siblings like complete_todo, snooze_todo, or update_todo, which also mutate a TODO, so an agent must infer the difference from the tool name alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to choose this over complete_todo, snooze_todo, or update_todo, nor any prerequisite (e.g. ownership or existing state). The delete intent is only implied by the words 'soft delete'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_meアカウント情報と今日A

Account レコード(trialStartedAt / subscribedUntil / timeZone / defaultTime)と、その時間帯での今日の日付・曜日、due を省略したときに使われる「明日の既定時刻」(defaultDueNext、ISO8601)を返す。日付を扱う前に呼ぶ。

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose a genuinely non-obvious behavioral trait: defaultDueNext is a derived value tied to the account time zone and only matters as the fallback when due is omitted. It still does not state whether authentication (sibling sign_in) is required or what happens on failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler: the first front-loads exactly what is returned, the second front-loads the call precondition at the end where it is most actionable. Every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must explain return values and it does so field by field, including the derived ISO8601 value. For a zero-parameter read tool the definition gives the agent everything needed to call and interpret it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4 and there is no parameter syntax for the description to add. Nothing in the schema needs compensating for.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the exact resource (Account レコード) and enumerates the fields returned (trialStartedAt, subscribedUntil, timeZone, defaultTime, defaultDueNext), so an agent knows precisely what it gets. It does not explicitly contrast itself with siblings like list_todos or sign_in, but the resource is distinctive enough that confusion is unlikely.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The closing line "日付を扱う前に呼ぶ" gives an explicit precondition for calling the tool, which is real usage guidance rather than implication. It stops short of naming a when-not case or an alternative tool, so it is not a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_todosTODO 一覧A

TODO を一覧する。date(1 日)か from/to(両端含む、YYYY-MM-DD、Account の timeZone)で絞る。省略で全期間。status の既定は pending。削除済み(deletedAt あり)は含めない。繰り返しは展開せず repeating にテンプレートとして返し、範囲内の Occurrence(回ごとの done / skipped / スヌーズ)を添える。展開はアプリ側で MCP はしない。規則(repeat): kind × interval(1=毎、2=隔)。daily は interval 日ごと、weekly は weekdays(1=日…7=土。空なら dueAt の曜日)の曜日、monthly は dueAt の日(dayOfMonth。無い月は月末)か第 n 曜日(weekdayOrdinal、ordinal -1 = 最終)、yearly は dueAt の月日。終了は end(never / until = その日を含む / count = 開始からの回数)。初回は dueAt、各回の時刻は dueAt の時分(timeZone の時計)。

ParametersJSON Schema
NameRequiredDescriptionDefault
toNoこの日まで(含む)
dateNoこの 1 日だけ
fromNoこの日から(含む)
statusNo既定 pending。done / skipped は単発だけに効く
includeRepeatingNo既定 true。false で繰り返しテンプレートを省く

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does well: it discloses that deleted records (deletedAt present) are excluded, that recurring items are not expanded but returned as 'repeating' templates plus in-range occurrences, and that expansion is the app's job, not the MCP's. It stops short of stating the read-only nature or the response shape explicitly, but the read semantics are strongly implied by 一覧する.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose and filtering rules are front-loaded, and the dense recurrence-rule block at the end is justified because there is no output schema describing the 'repeating' template. Every clause is information-bearing, though the spec-dump style for recurrence could be tightened.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-required-param list tool with no output schema, the description covers filtering, defaults, exclusion rules, and the structure of recurring items and their occurrences. Given the complexity of the recurrence domain, it is largely complete, though the shape of a returned TODO item itself is left unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema: from/to are inclusive on both ends, dates are interpreted in the Account's timeZone, and status defaults to pending. These are semantic details the schema fields do not convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening 'TODO を一覧する' states a specific verb (list) and resource (TODO), immediately separable from the mutation siblings (add_todo, update_todo, complete_todo, delete_todo). It also defines the scope precisely with the date/from-to filtering semantics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly says when to use date (single day) vs. from/to (inclusive range) and what the omission default is (全期間). It also states the status default (pending) and that deleted items are excluded. It does not explicitly name sibling alternatives, but the filtering guidance is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sign_iniCloud にサインインA

Apple のサインイン画面をブラウザで開き、戻ってきた ckWebAuthToken を ~/.mustdo/auth.json に保存する。他のツールが NOT_SIGNED_IN を返したときに使う。ユーザーがブラウザ操作できるときだけ呼ぶこと。

ParametersJSON Schema
NameRequiredDescriptionDefault
waitSecondsNoブラウザが戻るまで待つ秒数(既定 90)

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the side effect (a browser flow), the persisted artifact and its exact path (~/.mustdo/auth.json), and the human-in-the-loop prerequisite. It is silent on failure/timeout behavior and whether an existing token is overwritten, which keeps it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences: action, trigger, precondition. The routing constraint is front-loaded and every clause earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-optional-parameter, no-output-schema tool, the description covers purpose, trigger, precondition, and side effect. Only minor gaps remain (failure/timeout handling, token-overwrite semantics), which are not essential for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and only one optional parameter exists, so the schema already fully documents waitSeconds (default 90, range 5-300). The description adds no parameter syntax or defaults beyond that, making the baseline 3 correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (opens Apple's sign-in page, persists ckWebAuthToken to ~/.mustdo/auth.json). This is unmistakably different from every sibling tool, all of which are todo operations, so an agent can route to it without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit trigger (when another tool returns NOT_SIGNED_IN) and an explicit exclusion (call only when the user can operate the browser). Both the when and the when-not are stated, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

snooze_todoTODO をスヌーズC

until(YYYY-MM-DDTHH:mm は TODO の timeZone、または ISO8601)まで鳴らさない。繰り返しは occurrenceDate の回だけ。

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
untilYesYYYY-MM-DDTHH:mm(Account の timeZone の時計)か、オフセット付き ISO8601
occurrenceDateNo繰り返しのどの回か。省略で今日

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, yet it only discloses the suppression effect and that recurring items are scoped to one occurrence. It omits whether auth is required, how it interacts with existing reminders, idempotency, and what value is returned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact clauses with no filler, and the cardinal constraint (until) is front-loaded. It is dense to the point of being slightly cryptic but wastes no words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations and no output schema, the description covers the core timing parameters but omits the id parameter, sibling routing, side effects, and return behavior. It is minimally sufficient but leaves clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%; the description restates the until format and clarifies recurring-occurrence behavior, adding some meaning beyond the schema. However, it leaves id undocumented and its timeZone wording ("TODO の timeZone") conflicts with the schema's "Account の timeZone", which is a minor semantic inconsistency rather than added clarity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description conveys a specific action and effect—suppressing firing until a given time ("〜まで鳴らさない")—which, with the name snooze_todo, is distinguishable from update_todo or complete_todo. It does not explicitly name or contrast a sibling, keeping it short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no stated when-to-use or when-not-to-use versus the siblings (e.g. update_todo, complete_todo). The only usage hint is the recurrence note ("繰り返しは occurrenceDate の回だけ"), which is narrow and does not help an agent choose this tool over alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_todoTODO を更新A

TODO の題名・時刻・繰り返し・メモ・音・状態を変える。渡した項目だけ変わる。repeat は丸ごと置き換え(省略した項目は既定値)。repeat: null か {kind: none} で繰り返しを外す。notes: null でメモを消す。繰り返し TODO の status は変えられない(complete_todo / snooze_todo を使う)。

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
dueNo鳴らす時刻。YYYY-MM-DDTHH:mm(Account の timeZone の時計)か、オフセット付き ISO8601。YYYY-MM-DD だけなら「その日の既定時刻(Account.defaultTime、既定 09:00)」。
notesNo
soundNoアラーム音。default, chime, marimba, beep, urgent, vibrateOnly, siren, klaxon, rapid, redalert, musicbox, harp, morning, bowl, droplet, breeze, horror, dread, ghost, lament, cello, rainy, fanfare, skip, sparkle
titleNo
repeatNo繰り返し(iOS カレンダーと同じ語彙)。省略した項目は既定値(interval 1・weekdays []・monthly.mode dayOfMonth・end.kind never)。範囲外は INVALID_ARGUMENT。例: 毎日 {kind:daily} / 毎週 月水 {kind:weekly, weekdays:[2,4]} / 隔週 {kind:weekly, interval:2} / 3 日ごと {kind:daily, interval:3} / 毎月 5 日 {kind:monthly}(due を 5 日に)/ 毎月 第 1 月曜 {kind:monthly, monthly:{mode:weekdayOrdinal, ordinal:1, weekday:2}} / 毎月 最終金曜 {kind:monthly, monthly:{mode:weekdayOrdinal, ordinal:-1, weekday:6}} / 毎年 {kind:yearly}(due の月日)/ 10 回で終了 {kind:daily, end:{kind:count, count:10}} / 12/31 まで {kind:weekly, weekdays:[2], end:{kind:until, until:"2026-12-31"}}
statusNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden well: it discloses partial mutation, wholesale replacement of repeat, repeat removal via null or {kind:none}, notes deletion via null, and the status restriction on repeating TODOs. It stops short of auth, error, or return behavior, but covers the main mutation semantics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded, moving from the general update operation to key semantic caveats without wasted sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter mutation tool with no annotations and no output schema, the description provides the crucial update semantics an agent needs. Minor gaps remain around id requirements and error/return behavior, but core invocation safety is covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 43%, so the description must compensate. It adds valuable semantics for repeat, notes, and status, but does not explain id, due-time formatting, or sound choices beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (変える) and resource (TODO), enumerates the mutable fields, and distinguishes the tool from complete_todo and snooze_todo for repeating TODO status operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly explains partial-update behavior (渡した項目だけ変わる) and routes the agent away from using this tool to change status on repeating TODOs, naming complete_todo and snooze_todo as alternatives. It does not cover every sibling relationship, but the relevant exclusion is explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 8 tool updatesv0.1.0
    • First observedadd_todo
    • First observedcomplete_todo
    • First observeddelete_todo
    • First observedget_me
    • First observedlist_todos
    • First observedsign_in
    • First observedsnooze_todo
    • First observedupdate_todo

TDQS

A3.8/5.0

Scored across 8 tools

Disambiguation4/5

Each tool targets a distinct operation: auth (sign_in, get_me), CRUD (add/update/delete/list_todo), and special status actions (complete_todo, snooze_todo). The only mild overlap is that update_todo can also change status, but the description explicitly redirects recurring occurrences to complete_todo/snooze_todo, keeping boundaries clear.

Naming Consistency5/5

Consistent snake_case verb_noun pattern throughout: sign_in, get_me, add_todo, update_todo, complete_todo, snooze_todo, delete_todo, list_todos. No mixed conventions or vague verbs.

Tool Count5/5

8 tools is well-scoped for a personal TODO server. Auth, account info, CRUD, and recurring-occurrence actions each earn their place without redundancy.

Completeness4/5

Covers the full TODO lifecycle including soft delete, completion, snooze, recurrence, and listing with filters. Minor gaps: no restore-undelete, no single-todo get, and no explicit unsnooze, but agents can work around these.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers