Skip to main content
Glama
Sealjay

mcp-hey

by Sealjay

mcp-hey

Sealjay/mcp-hey MCP server Bun TypeScript Python MCP License: MIT GitHub issues GitHub stars

A local Model Context Protocol (MCP) server that gives Claude read/write access to your Hey.com inbox via reverse-engineered web APIs.

mcp-hey has two moving parts: a Bun/TypeScript MCP server that exposes Hey tools over stdio, and a small Python helper that uses the system webview to capture session cookies at login. Everything runs locally — no cloud relay, no credentials stored, just session cookies on disk.

Warning — unofficial API. Hey.com does not publish a public API; mcp-hey reverse-engineers its web endpoints and pairs them with browser-identical HTTP requests. Things can break without notice. The current documented surface lives in docs/API.md.

Features

  • Read emails from Imbox, Feed, Paper Trail, Set Aside, Reply Later, Drafts, Trash, and Spam

  • Download attachments and parse calendar invites from emails

  • Send and reply to email threads

  • Search emails across boxes

  • Organise mail (set aside, reply later, screen in/out, bubble up)

  • Local SQLite cache for faster repeated reads and full-text search

  • Lightweight — around 30 MB idle memory

  • Browser-identical headers and TLS posture to avoid detection

  • Runs entirely on your machine; stdio transport with no network exposure

Related MCP server: Outlook MCP

Setup

Prerequisites

  • Bun 1.1 or later

  • Python 3.10 or later (plus UV if you want to follow the Python tooling in CLAUDE.md)

  • A Hey.com account

  • Platform: developed and tested on macOS and Linux. Windows users will likely need WSL — pywebview's Windows backend is not currently exercised.

Installation

  1. Clone this repository

    git clone https://github.com/Sealjay/mcp-hey.git
    cd mcp-hey
  2. Install dependencies

    bun install
    uv pip install -r auth/requirements.txt
  3. First run — authenticate

    bun run dev
    1. A system webview opens with Hey.com's login page. Log in normally.

    2. The helper captures session cookies to data/hey-cookies.json (permissions 600) and exits.

    3. Press Ctrl+C — your MCP client will launch its own server instance from here on.

    4. Subsequent runs reuse the stored session until it expires.

MCP client configuration

All clients below use the same command/args shape. On macOS, you'll almost certainly need the absolute path to bun — see macOS: bun PATH below.

Claude Code

The quickest route is the CLI:

claude mcp add --transport stdio hey --scope user -- bun run /absolute/path/to/mcp-hey/src/index.ts

The server is available immediately in the current session.

Alternatively, add to .mcp.json at your project root (or ~/.claude.json for a user-scoped server):

{
  "mcpServers": {
    "hey": {
      "type": "stdio",
      "command": "bun",
      "args": ["run", "/absolute/path/to/mcp-hey/src/index.ts"]
    }
  }
}

If you edit the file directly, restart the Claude Code session to pick it up.

Claude Desktop

Add to ~/Library/Application Support/Claude/claude_desktop_config.json (macOS):

{
  "mcpServers": {
    "hey": {
      "command": "bun",
      "args": ["run", "/absolute/path/to/mcp-hey/src/index.ts"]
    }
  }
}

Restart Claude Desktop. You should see hey listed as an available integration.

Cursor

Add to ~/.cursor/mcp.json:

{
  "mcpServers": {
    "hey": {
      "command": "bun",
      "args": ["run", "/absolute/path/to/mcp-hey/src/index.ts"]
    }
  }
}

Restart Cursor.

Docker

A Dockerfile is included for containerised deployments and Glama compatibility.

Build the image:

docker build -t mcp-hey .

Smoke-test the server (should return a JSON-RPC response listing available tools):

printf '{"jsonrpc":"2.0","id":1,"method":"tools/list"}\n' | docker run -i mcp-hey

Note: The Docker image runs the MCP server only. The Python auth helper and webview login are not available inside the container. You must provide pre-existing session cookies via a volume mount to data/hey-cookies.json for authenticated operations.

macOS: bun PATH

GUI apps (Claude Desktop, Cursor) and shells launched by Claude Code don't always inherit the PATH from your interactive terminal, so a Homebrew-installed bun may fail with spawn bun ENOENT or simply never connect. Fix by using the absolute path to bun in command:

  • Apple Silicon Homebrew — /opt/homebrew/bin/bun

  • Intel Homebrew — /usr/local/bin/bun

  • Manual install — run which bun in your terminal to find it

Example:

{
  "mcpServers": {
    "hey": {
      "command": "/opt/homebrew/bin/bun",
      "args": ["run", "/absolute/path/to/mcp-hey/src/index.ts"]
    }
  }
}

Architecture

Component

Description

MCP server

Bun/TypeScript, stdio transport, ~30 MB idle memory

Auth helper

Python/pywebview, spawns on-demand for login via system webview

Cache

Local SQLite store for messages, threads, and search index

Communication

File-based session sharing via data/hey-cookies.json

Data flow

  1. MCP client (Claude Code, Claude Desktop, Cursor, etc.) launches bun run src/index.ts over stdio.

  2. On startup the server validates data/hey-cookies.json. If missing or expired it spawns auth/hey-auth.py, which opens Hey in a system webview and writes fresh cookies.

  3. Tool calls hit Hey.com directly with browser-realistic headers; responses are parsed (HTML via node-html-parser) and cached in SQLite.

  4. Write operations fetch a fresh CSRF token before submitting.

Project structure

mcp-hey/
  src/
    index.ts           # MCP server entry point
    hey-client.ts      # HTTP client with cookie injection
    session.ts         # Session management and validation
    errors.ts          # Error classes and sanitisation
    cache/             # SQLite cache (db, schema, messages, search)
    tools/             # MCP tool implementations
      read.ts          # Reading and listing
      send.ts          # Send, reply, forward
      organise.ts      # Triage, labels, bubble up, etc.
      http-helpers.ts  # Shared CSRF retry and endpoint fallback
      attachments.ts   # Download attachments, parse calendar invites
    __tests__/         # Test suites
  auth/
    hey-auth.py        # Python auth helper (pywebview)
    requirements.txt
  data/
    hey-cookies.json   # Session storage (gitignored, chmod 600)
  docs/
    API.md             # Hey.com API surface documentation
    TOOLS.md           # MCP tool reference (34 tools)
    hey-features-doc.md  # Hey.com feature mapping

Available tools

34 tools grouped by function. See docs/TOOLS.md for parameters, return shapes, and error behaviour.

Category

Tools

Read

hey_list_emails (imbox, feed, paper_trail, trash, spam, drafts, sent), hey_imbox_summary, hey_list_set_aside, hey_list_reply_later, hey_list_screener, hey_read_email, hey_download_attachment, hey_get_calendar_invite

Labels & Collections

hey_list_labels, hey_list_label_emails, hey_label, hey_list_collections, hey_list_collection_emails, hey_collection

Send

hey_send_email, hey_reply, hey_forward

Triage

hey_set_aside, hey_unset_aside, hey_reply_later, hey_remove_reply_later, hey_move_to, hey_set_status, hey_mark_unseen, hey_mark_seen, hey_read_status, hey_thread_mute

Bubble up

hey_bubble_up, hey_bubble_up_if_no_reply, hey_pop_bubble

Screener

hey_screen, hey_screen_by_id

Search

hey_search

Cache

hey_cache_status

Privacy and security

  • No credentials are ever stored — only session cookies, written with 600 permissions.

  • Authentication happens entirely inside Hey's own login page (system webview).

  • All data stays on your machine. No telemetry is emitted by this project.

  • MCP uses stdio transport — the server never opens a network listener.

  • Session validity is checked on startup and before sensitive operations.

See SECURITY.md for how to report vulnerabilities.

Limitations

  • Prompt-injection risk: as with many MCP servers, this one is subject to the lethal trifecta. A malicious email arriving in your inbox could attempt to instruct Claude to exfiltrate other messages. Treat the tool surface accordingly and review risky actions before approving them.

  • Unofficial API: Hey.com's frontend can change without notice and break things. Expect occasional breakage and check docs/API.md for known deltas.

  • No real-time notifications: polling only.

  • Attachment uploads are not yet supported.

  • Single account per MCP server instance.

  • Account risk: aggressive or abnormal access patterns could in theory trigger Hey's anti-abuse systems. The server respects x-ratelimit headers and backs off exponentially, but there are no guarantees.

  • English UI only: the server parses Hey.com's HTML responses and matches English-language strings (e.g. "You ignored this thread", label names, button text). It will not work correctly if Hey.com is set to a non-English locale.

Troubleshooting

  • Auth webview does not open — confirm Python 3.10+ is on PATH and uv pip install -r auth/requirements.txt succeeded. On Linux ensure a webview backend is available (python -c "import webview" should not error).

  • 401/403 responses after weeks of use — your Hey session has expired. Delete data/hey-cookies.json and run bun run dev again to re-auth.

  • Rate limits (429) — the client respects x-ratelimit headers and backs off. If you see sustained 429s, reduce concurrent tool use or wait a few minutes.

  • MCP client can't launch the server — args must be an absolute path, not relative. If bun itself fails with spawn bun ENOENT, see macOS: bun PATH.

  • Cookie name changed — Hey has renamed session cookies before (e.g. _hey_session → session_token, see docs/API.md changelog). If auth silently fails after a Hey update, capture fresh cookies and compare.

Contributing

Contributions welcome via pull request. Please:

  • Use conventional commits (feat, fix, docs, refactor, test, perf, cicd, revert, WIP).

  • Run bun run format and bun run lint before pushing (powered by Biome).

  • Ensure bun test passes.

  • Update docs/API.md if you discover or change any Hey.com API behaviour.

See CLAUDE.md for the full development workflow.

Licence

MIT Licence — see LICENCE.

Available Tools

37 tools
hey_bubble_upA
Idempotent

Schedule an email thread to bubble up (reappear) at a specific time. Returns {success, error?}. Reversible via hey_pop_bubble. Requires the thread's topic_id (use topic_id from any list operation); posting IDs are not accepted and will 404.

ParametersJSON Schema
NameRequiredDescriptionDefault
dateNoDate in YYYY-MM-DD format. Required when slot is 'custom', ignored otherwise.
slotYesWhen to bubble up: now (immediately), today (evening), tomorrow (morning), weekend (Saturday), next_week (Monday), surprise_me (random), custom (specific date - requires 'date' parameter)
topic_idYesThe topic ID (thread ID) to schedule. Use the `topicId` field from hey_list_* responses (NOT `postingId` — that's a different ID and will 404).

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond annotations (idempotent, non-destructive), description adds return format {success, error?} and the critical constraint that posting IDs are not accepted. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with all critical information front-loaded. No redundant text, every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given annotations, schema completeness, and no output schema, the description covers all necessary aspects: purpose, parameters, constraints, return format, and reversibility.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, and description adds crucial context: topic_id must come from list operations, slot options are explained, and date parameter requirement for custom slot. The distinction between topic_id and postingId is essential and well clarified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Explicitly states the action (schedule an email thread to bubble up) and resource (email thread). Differentiates from sibling hey_pop_bubble by mentioning reversibility.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clearly states when to use (schedule a thread to reappear) and what not to use (posting IDs cause 404). Mentions reversible alternative hey_pop_bubble but does not contrast with hey_bubble_up_if_no_reply.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_bubble_up_if_no_replyA
Idempotent

Schedule an email thread to bubble up ONLY if there's no reply by a deadline date. Returns {success, error?}. The thread only reappears if no reply arrives. Reversible via hey_pop_bubble. Requires the thread's topic_id (use topic_id from any list operation); posting IDs are not accepted.

ParametersJSON Schema
NameRequiredDescriptionDefault
dateYesDeadline date in YYYY-MM-DD format. The email will bubble up only if no reply is received by this date.
topic_idYesThe topic ID (thread ID) to schedule. Use the `topicId` field from hey_list_* responses (NOT `postingId` — that's a different ID and will 404).

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate idempotent and non-destructive behavior. The description adds the return shape and reversibility, but does not elaborate on side effects or triggers. Adequate but not exceptional.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core purpose, no wasted words. Efficient and clear.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool, the description provides return value, reversibility, and ID clarification. No output schema exists, but the return shape is described. Complete for the complexity level.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Both parameters are fully described in the schema (100% coverage). The description adds crucial context about using the correct ID field ('topicId' vs 'postingId'), which goes beyond schema documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the conditional behavior ('bubble up ONLY if there's no reply by a deadline date'), which differentiates it from siblings like 'hey_bubble_up'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly mentions when to use (conditional bubbling) and provides a reversible alternative ('Reversible via hey_pop_bubble'). Also warns about the correct ID type. Could be improved by explicitly stating when not to use, but the alternative is implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_cache_statusA
Read-only

Get local cache statistics: message counts, cache age, and storage estimate. Read-only with no side effects. Pass folder to scope message/unread counts to one cached view. Use before deciding whether to pass force_refresh on other list tools.

ParametersJSON Schema
NameRequiredDescriptionDefault
folderNoOptional: scope stats to one cached view (imbox, feed, paper_trail, set_aside, reply_later). Note this differs from hey_list_emails' folder enum — trash, spam, and drafts aren't cache-tracked and silently return 0 here.

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, and the description confirms 'Read-only with no side effects'. It adds value by clarifying that the data is local cache state (not live server state), which affects staleness interpretation. It stops short of describing freshness guarantees or how the storage estimate is derived.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, each earning its place: what it returns, its safety profile, and how/when to use it, with the most decision-relevant guidance front-loaded after the payload description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description must convey what comes back, and it does (message counts, cache age, storage estimate). The optional parameter is explained and the pre-flight use case is fully covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaning the schema does not: it specifies that folder scopes 'message/unread counts' specifically, telling the agent which outputs are affected. The schema-handled caveats about untracked folders complement this.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb and resource ('Get local cache statistics') and enumerates the returned content: message counts, cache age, storage estimate. It also distinguishes itself from the sibling list tools by framing its role as a pre-flight check for force_refresh.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use it: 'Use before deciding whether to pass force_refresh on other list tools.' This routes the agent to the right decision point rather than leaving usage implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_collectionA
Idempotent

Add or remove an email thread from a collection. Returns {success, error?}. Reversible by calling again with the opposite action. Use hey_list_collections to discover collection IDs.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesWhether to add or remove the thread from the collection
topic_idYesThe topic/thread ID to add to or remove from the collection
collection_idYesThe collection ID (use hey_list_collections to see available collections)

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds real value beyond the annotations: the return shape ({success, error?}) and the fact that the operation is reversible by calling again with the opposite action. Annotations already cover the safety/freshness profile (destructiveHint=false, idempotentHint=true), so the description's contribution is the response contract and undo path.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences with zero filler; the core action is front-loaded, followed by return shape, then the discovery hint. Every sentence carries distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter mutation with no output schema, the description covers the operation, the return contract, reversibility, and the ID-discovery prerequisite. Combined with complete schema coverage and annotations, an agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so topic_id, collection_id, and the action enum are already fully documented in the schema. The description only repeats the hey_list_collections hint for collection_id and adds nothing new about parameter format or constraints, matching the baseline for fully covered schemas.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb pair (add/remove) and a specific resource (an email thread within a collection), which an agent can distinguish from siblings such as hey_list_collections or hey_list_collection_emails. No ambiguity about what the tool mutates.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent to hey_list_collections to discover valid collection IDs, which is the key prerequisite for calling this tool. It stops short of stating when not to use it (e.g., versus hey_set_status or label tools), so it falls just below the top band.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_delete_draftA
DestructiveIdempotent

Permanently delete a draft by ID. This does not move it to Trash — it is gone immediately, same as the trash icon in Hey's Drafts list. Returns {success, draftId, error?}. Irreversible; there is no hey_restore_draft. Use hey_list_emails(folder='drafts') first if you don't already have the draftId from hey_save_draft or hey_edit_draft.

ParametersJSON Schema
NameRequiredDescriptionDefault
draft_idYesThe draft's message ID, from hey_save_draft's draftId or hey_list_emails(folder='drafts')

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=true and idempotentHint=true, but the description adds what those hints cannot express: the delete bypasses Trash, is immediate and irreversible, and there is no restore counterpart. It also discloses the return shape ({success, draftId, error?}) despite no output schema existing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, front-loaded with the destructive semantic and irreversibility, then return shape, then prerequisite. No filler and nothing repeated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter destructive tool with no output schema, the description covers effect (immediate, non-Trash), reversibility (none), return value, and how to obtain the ID. Nothing an agent needs to invoke it safely is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the schema already names the source of draftId, so the description's provenance note is largely redundant with structured data. It adds no format, validation, or ID-shape detail beyond what the schema provides; baseline 3 applies for a single fully documented parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Permanently delete a draft by ID') and immediately scopes it against the sibling set by noting there is no hey_restore_draft. An agent can distinguish this from hey_save_draft / hey_edit_draft without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states the prerequisite path ('Use hey_list_emails(folder='drafts') first if you don't already have the draftId from hey_save_draft or hey_edit_draft'), and explicitly excludes the assumed undo path. Both when-to-use and when-not-to-rely-on-restore are covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_download_attachmentA

Download a single attachment from an email and save it to disk. First call hey_read_email to get the attachments[] array with IDs, filenames, and sizes, then call this tool with the attachment_id to save the file. Returns {local_path, filename, size, mime}.

ParametersJSON Schema
NameRequiredDescriptionDefault
email_idYesThe email's topic or entry ID (same ID used with hey_read_email — topic IDs are resolved automatically)
save_pathNoOptional path or directory to save into. Must be within ~/. Defaults to ~/Downloads/hey-attachments/<date>/<filename>. Trailing '/' is treated as a directory. Duplicate filenames are auto-numbered (invite-1.ics, invite-2.ics).
attachment_idYesThe attachment ID from hey_read_email's attachments array (format: part-N, e.g. 'part-1')

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, and openWorldHint=true (a filesystem-writing, non-idempotent operation). The description adds genuinely new behavior: the path must stay within ~, the default save location, and duplicate-filename auto-numbering (invite-1.ics, invite-2.ics), which is the concrete explanation for the non-idempotent hint. It does not mention permission or overwrite failure modes, so it falls short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with zero filler: purpose first, then the prerequisite workflow, then the return shape. Every sentence earns its place and the most important constraint (read email first) is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, but the description supplies the return shape ({local_path, filename, size, mime}), the prerequisite tool, and the save-path behavior. Combined with complete schema coverage and annotations, an agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters (including the save_path default and directory semantics) are already fully documented in the schema. The description only restates that attachment_id comes from hey_read_email's attachments[] array, which the schema also says, so it adds little beyond baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb+resource+effect: 'Download a single attachment from an email and save it to disk.' It is the only download/attachment tool among the siblings, so the scope is unambiguous without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit two-step workflow: call hey_read_email first to obtain the attachments[] array, then call this tool with the attachment_id. It names the prerequisite sibling and the exact input it provides, so an agent knows when and how to sequence the call.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_edit_draftA
Idempotent

Update an existing draft's recipients, subject, or body. Every field you pass replaces the draft's current value outright (no merging) — omit a field to leave it as Hey last saved it. Returns {success, draftId, error?}. Previous field values are not recoverable once overwritten, so re-fetch via hey_list_emails(folder='drafts') first if you need to preserve them. Use hey_save_draft to create a draft before calling this, and hey_delete_draft to remove one instead of blanking it out.

ParametersJSON Schema
NameRequiredDescriptionDefault
ccNoReplace the CC list with these addresses
toNoReplace the recipient list with these addresses
bodyNoReplace the body content (HTML supported)
subjectNoReplace the subject line
draft_idYesThe draft's message ID, from hey_save_draft's draftId or hey_list_emails(folder='drafts')

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial context beyond the annotations: outright (non-merging) replacement semantics, the omit-to-leave-unchanged rule, irreversibility of overwritten values, and the return shape {success, draftId, error?}. There is mild tension with destructiveHint=false given the explicit warning that previous values are not recoverable, but the target draft itself is not destroyed, so this reads as nuance rather than a hard contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose and the crucial no-merge/omit rule are front-loaded, followed by return shape, irreversibility, and routing. Every sentence carries information, though the description is dense and runs slightly long for a five-parameter update tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description supplies the return shape and the key behavioral rules (replacement, omission, irreversibility) an agent needs to invoke it safely. Nothing essential is missing for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description goes beyond the schema by clarifying that omitted fields retain their last-saved value and that passing a field replaces rather than merges — a semantic the schema does not state and that materially affects how the agent should call it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (update), resource (an existing draft), and the exact mutable fields (recipients, subject, body). It also explicitly distinguishes itself from the sibling create tool hey_save_draft and the removal tool hey_delete_draft, so an agent can route correctly without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states prerequisites ('Use hey_save_draft to create a draft before calling this') and a when-not condition ('use hey_delete_draft to remove one instead of blanking it out'). It also gives a concrete guidance action — re-fetch via hey_list_emails(folder='drafts') first if preservation matters.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_forwardA

Forward an existing email to new recipients immediately. The original thread remains unchanged. Returns {success, error?}. Use instead of hey_send_email when sharing existing content; use hey_reply for responding within a thread.

ParametersJSON Schema
NameRequiredDescriptionDefault
ccNoList of CC recipient email addresses
toYesList of recipient email addresses
bccNoList of BCC recipient email addresses
bodyNoOptional message to include above the forwarded content
entry_idYesThe entry ID of the email to forward

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Description discloses that forwarding is immediate and non-destructive (thread unchanged), and mentions the return shape {success, error?}, adding value beyond annotations. Could include more on authorization or error conditions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise sentences, each serving a purpose: action/effect, key behavior, and usage guidance. No extraneous information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers purpose, usage, and return type for a simple forwarding tool. Lacks details on error handling or behavior with invalid entry_id, but adequate given tool complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema coverage, the description adds little beyond the schema. It reinforces that body is optional but does not provide additional semantic or formatting guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool forwards an existing email immediately and notes the original thread remains unchanged, distinguishing it from siblings like hey_send_email and hey_reply.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs to use this tool instead of hey_send_email for sharing existing content and hey_reply for replying within a thread, providing clear when-to-use and alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_get_calendar_inviteA
Read-only

Extract and parse a calendar invite (.ics) from an email. First call hey_read_email — if calendar_invites[] is present, call this tool to get full details: title, start, end, location, attendees, organizer, description, and raw_ics. To save the .ics file to disk, use hey_download_attachment instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
email_idYesThe email's topic or entry ID (same ID used with hey_read_email — topic IDs are resolved automatically)
attachment_idNoOptional attachment ID when the email has multiple .ics parts

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

ReadOnlyHint annotation is consistent with description of extraction. Description adds context that it returns full details and requires prior call to hey_read_email, which is useful behavioral sequence.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no waste: first sentence states purpose, second provides usage guidance. Front-loaded with core action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, description lists all returned fields (title, start, end, etc.). Covers prerequisites, input, output, and relationships with siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema has 100% parameter description coverage; description adds workflow context for email_id (same as hey_read_email) and clarifies attachment_id purpose for multiple .ics parts.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the tool extracts and parses a calendar invite (.ics) from an email, with specific verb and resource. It distinguishes itself from siblings like hey_read_email and hey_download_attachment.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells when to use (after hey_read_email if calendar_invites present) and when not to use (use hey_download_attachment to save .ics). Provides clear workflow.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_imbox_summaryA
Read-only

Get a complete Imbox summary including screener count, bubbled up emails, and new emails. Use this for a comprehensive view of the inbox state.

ParametersJSON Schema
NameRequiredDescriptionDefault
force_refreshNoBypass cache and fetch fresh data (default: false)

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate read-only and open-world hints. Description adds that it includes specific counts and emails, but doesn't mention caching behavior or other traits beyond what annotations provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no unnecessary words. Front-loaded with the purpose and usage guidance. Very concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description provides a reasonable expectation of what the summary includes. Could be more explicit about the complete return structure, but adequate for a simple tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% for the single parameter (force_refresh). Description does not add any additional meaning or usage details about the parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states 'Get a complete Imbox summary' with specific included items (screener count, bubbled up emails, new emails). Differentiates from sibling tools by being a high-level overview, distinct from list or action tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States 'use this for a comprehensive view of the inbox state', providing context for when to use it. However, no explicit alternatives or when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_labelA
Idempotent

Add or remove a label on an email thread. Returns {success, error?}. Reversible by calling again with the opposite action. Use hey_list_labels to discover available label IDs.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesWhether to add or remove the label
label_idYesThe label ID to add or remove (use hey_list_labels to see available labels)
topic_idYesThe topic/thread ID to label or unlabel

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the mutation profile (readOnly=false, destructive=false, idempotent=true), but the description adds real value by disclosing the return shape ({success, error?}) despite no output schema, and by explaining reversibility. It doesn't cover auth requirements or rate limits, but the added context is substantive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with zero filler: the action is front-loaded, the return shape follows, and the discovery pointer closes it out. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-param mutation tool with full schema coverage but no output schema, the description supplies the missing return shape, reversibility semantics, and ID discovery path. Nothing an agent needs to invoke it correctly is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the enum, label_id, and topic_id are fully documented in the schema. The description's only param-related guidance (use hey_list_labels for IDs) duplicates the label_id schema description, so there is no net addition beyond structured fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb pair (add/remove) and resource (label on an email thread), which is unambiguous and distinct from siblings like hey_set_aside or hey_move_to. An agent can tell immediately what this tool does without reading further.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent to hey_list_labels to discover available label IDs before calling. It gives clear prerequisite context, though it doesn't contrast against similarly-scoped siblings like hey_set_status or explain when labeling is preferable to those alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_list_collection_emailsA
Read-only

List emails in a specific collection. Returns cached results unless force_refresh=true.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageNoPage number for pagination, 1-1000 (default: 1)
limitNoMaximum number of emails to return, 1-100 (default: 25)
collection_idYesThe collection ID to list emails from (use id from hey_list_collections)
force_refreshNoBypass cache and fetch fresh data (default: false)

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true. The description adds real behavioral value by disclosing the caching behavior ('Returns cached results unless force_refresh=true'), which the annotations do not convey. It stops short of stating auth needs, rate limits, or return shape, but this is useful added context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero filler; the primary purpose is front-loaded and the caching caveat follows. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only, collection-scoped list tool with full schema coverage and safety annotations, the description covers purpose and the one non-obvious behavior (caching). Pagination and return format are handled by the schema, so little is missing, though it could note the collection_id dependency more explicitly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (page, limit, collection_id, force_refresh) are already documented in the schema, including defaults and ranges. The description only restates the force_refresh semantics already present in the schema, adding no new parameter meaning beyond the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List emails in a specific collection'), which clearly identifies the operation. It separates itself from plain list_email indirectly via the 'collection' scope, but never names a sibling collection-vs-label distinction, so it falls short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance, no prerequisites, and no mention of alternatives such as hey_list_emails or hey_list_label_emails. The scoping is inferable from the tool name but the description offers no routing help.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_list_collectionsA
Read-only

List all collections in Hey.com. Returns array of {id, name}. Use the id with hey_collection or hey_list_collection_emails.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate readOnlyHint and openWorldHint. Description adds that it returns an array of {id, name} and lists all collections, but does not elaborate on behavior like pagination or edge cases. Still, the return format is useful context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core action 'List all collections,' and no wasted words. Every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no parameters, no output schema, and simple design, the description is complete. It specifies return format and how to use results with sibling tools, covering the agent's needs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

No parameters (0), so schema coverage is irrelevant per guidelines at 100% vacuous. Description adds meaning by explaining the tool's function and output, which compensates for lack of param info. Baseline 4 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it lists all collections in Hey.com, specifies the return format (array of {id, name}), and distinguishes from sibling tools like hey_collection (single collection) and hey_list_collection_emails (emails of a collection).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit guidance: 'Use the id with hey_collection or hey_list_collection_emails,' telling the agent when to switch to alternatives. Also, no parameters means no ambiguity in usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_list_emailsA
Read-only

List emails in a Hey.com folder/view. Returns cached results unless force_refresh=true. Each email includes id, topicId, postingId, entryId, from, subject, date, and unread status. For folder=sent, from is always "Me" and the recipient (unverified against a live Sent page) is under to/toEmail instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageNoPage number for pagination, 1-1000 (default: 1)
limitNoMaximum number of emails to return, 1-100 (default: 25)
folderYesThe folder/view to list emails from: imbox (important), feed (newsletters), paper_trail (receipts), trash, spam, drafts, or sent (emails you sent)
force_refreshNoBypass cache and fetch fresh data (default: false)

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so safety is covered. The description adds genuinely useful behavioral context beyond them: results are cached unless force_refresh=true, and it candidly flags that the sent-folder recipient field is unverified against a live Sent page — a real reliability caveat.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three front-loaded sentences: purpose, caching behavior, then returned fields. Each carries distinct information with no filler, though the return-field enumeration is somewhat list-like and could be trimmed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly compensates by enumerating the returned fields (id, topicId, postingId, entryId, from, subject, date, unread). Caching semantics and the sent-folder edge case round it out; only pagination/interaction with page and limit is left entirely to the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so parameter definitions are already well documented (baseline 3). The description goes further by explaining force_refresh's cache-bypass effect in context and by describing the sent-folder quirk where from is always "Me" and the recipient moves to to/toEmail, which is nowhere in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("List emails in a Hey.com folder/view") and scopes it to the folder/view axis, which distinguishes it from siblings like hey_list_label_emails and hey_list_collection_emails. It does not name those siblings explicitly, so the differentiation is inferable rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the folder/view framing, and the force_refresh sentence signals when a cache bypass is warranted. However, it never says when to prefer this over hey_search or the label/collection listing tools, so the agent must infer the selection rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_list_label_emailsA
Read-only

List emails with a specific label. Returns cached results unless force_refresh=true.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageNoPage number for pagination, 1-1000 (default: 1)
limitNoMaximum number of emails to return, 1-100 (default: 25)
label_idYesThe label ID to list emails from (use id from hey_list_labels)
force_refreshNoBypass cache and fetch fresh data (default: false)

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and openWorldHint, so safety is covered. The description adds genuinely new behavioral context the annotations do not carry: results come from a cache by default and are only re-fetched when force_refresh=true. It omits pagination behavior and return shape, but adds real value beyond the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero waste, with the core operation stated first and the caching caveat second. Nothing is padded or redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, open-world list tool with a fully documented schema and no output schema, the description covers the essential non-obvious behavior (caching). It is nearly complete; the only gap is what a result looks like and pagination limits, which the schema partially covers.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description contributes semantics the schema does not: it explains the consequence of force_refresh ('returns cached results unless...') rather than merely restating 'bypass cache'. label_id, page and limit are left to the schema, which documents them adequately.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb and resource with a clear scope qualifier ('emails with a specific label'), which distinguishes it from the broader hey_list_emails. It does not, however, explicitly name or contrast with the closest siblings (hey_list_emails, hey_list_collection_emails), so differentiation must be inferred from the scope phrase.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the scope ('with a specific label'), and the caching clause hints at when to set force_refresh, but the description never states when to prefer this over hey_list_emails, hey_list_collection_emails, or hey_search. No exclusions or prerequisites are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_list_labelsA
Read-only

List all labels/folders in Hey.com. Returns array of {id, name}. Use the id with hey_label or hey_list_label_emails.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description adds value beyond them by disclosing the return shape ('array of {id, name}'), which matters since there is no output schema; it omits ordering/pagination but those are minor for a label list.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences with zero waste: purpose, return shape, and next-step routing are front-loaded in that order. Nothing could be cut without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-param, side-effect-free listing tool with annotations covering safety, the description supplies the one thing structured data does not: the return shape. It stops short of noting whether results are cached or ordered, a minor gap for this complexity level.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. There is nothing for the description to clarify, and it correctly avoids inventing parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List all labels/folders in Hey.com') and goes further to declare the exact return shape ({id, name}) and the follow-up tools that consume the id. An agent can distinguish it from hey_list_label_emails or hey_list_collections without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description names two follow-up tools that consume the returned id, which implies a discover-then-act workflow, but it never says when to prefer this tool over the sibling list tools or state any precondition. Usage is inferred rather than specified.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_list_reply_laterA
Read-only

List emails currently in the Reply Later stack. Returns cached results (same shape as hey_list_emails) unless force_refresh=true. Use hey_remove_reply_later with the returned postingId to remove an item.

ParametersJSON Schema
NameRequiredDescriptionDefault
force_refreshNoBypass cache and fetch fresh data (default: false)

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and openWorldHint, so safety is covered. The description adds genuinely useful behavior: results are cached unless force_refresh=true, and the output shape matches hey_list_emails. This goes beyond the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, zero filler, with the core purpose front-loaded and the caching caveat and removal pointer following in priority order.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by stating the return shape matches hey_list_emails and that items carry a postingId. That is enough for an agent to act, though it could note pagination or empty-stack behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single force_refresh parameter is fully documented in the schema including its default. The description's mention that force_refresh bypasses the cache largely restates the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List emails currently in the Reply Later stack'), which is distinguishable from sibling list tools like hey_list_set_aside and hey_list_emails. An agent can identify the target stack without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the follow-up tool hey_remove_reply_later with the exact identifier (postingId) needed, which routes the agent correctly for removal. It does not explicitly say when not to use this versus hey_list_emails, so it falls short of the 5 bar.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_list_screenerA
Read-only

List senders waiting in the Screener for approval. Returns cached results (same shape as hey_list_emails, plus clearanceId) unless force_refresh=true. Use hey_screen_by_id with the returned clearanceId to approve or reject.

ParametersJSON Schema
NameRequiredDescriptionDefault
force_refreshNoBypass cache and fetch fresh data (default: false)

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnlyHint=true, openWorldHint=true), and the description adds a genuinely useful trait they do not: results are cached and only refreshed when force_refresh=true. That caching disclosure is real behavioral value beyond structured fields, though rate limits, pagination, or staleness windows are not addressed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the purpose, then the caching caveat, then the follow-up tool. Every sentence earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter, annotation-covered list tool with no output schema, the description compensates by describing the return shape ('same shape as hey_list_emails, plus clearanceId'), which is exactly what an agent needs to chain into hey_screen_by_id. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With a single parameter at 100% schema description coverage, the schema already documents force_refresh as a cache-bypass flag with a default of false. The description echoes that behavior ('unless force_refresh=true') but adds no format or edge-case detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List senders waiting in the Screener for approval') and immediately differentiates itself from the sibling hey_screen_by_id, which handles approval/rejection instead. An agent can tell what this tool does and what it does not do without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent to hey_screen_by_id with the returned clearanceId for the follow-up approve/reject step, which gives clear situational context. It stops short of stating when *not* to use this tool (e.g., vs hey_list_emails or hey_screen for other inbox views), so it is clear but not fully exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_list_set_asideA
Read-only

List emails currently in the Set Aside stack. Returns cached results (same shape as hey_list_emails) unless force_refresh=true. Use hey_unset_aside with the returned postingId to remove an item.

ParametersJSON Schema
NameRequiredDescriptionDefault
force_refreshNoBypass cache and fetch fresh data (default: false)

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint and openWorldHint; the description adds real behavioral detail the annotations do not cover: results come from a cache and force_refresh bypasses it, and the payload is the same shape as hey_list_emails. It omits cache scope/TTL, but that is a minor gap given the read-only safety profile is already declared.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, zero waste, front-loaded with the core purpose and followed by cache behavior and the follow-up action. Every sentence carries information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully tells the agent the result set shape matches hey_list_emails and how to act on the returned postingId. Caching behavior and the removal path are covered; only finer details like refresh cost or list size limits are absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% for the single force_refresh parameter, so the baseline is already covered. The description reinforces the semantics by explaining that true bypasses the cache, which adds practical meaning to the flag beyond the schema's terse wording.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List emails currently in the Set Aside stack') and places it against siblings by noting the result shape matches hey_list_emails and that removal is done via hey_unset_aside. An agent can distinguish this from hey_list_emails, hey_list_reply_later, and hey_list_collections without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear operational context: results are cached unless force_refresh=true, and the returned postingId feeds hey_unset_aside. It stops short of an explicit 'use this instead of X when...' routing rule, so it is strong but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_mark_seenA
Idempotent

Clear the orange 'New for you' tray dot — mirrors Hey's 'Mark all as seen' UI affordance. Returns {success, error?}. Pass a posting_id to clear just that one item (POST /postings/seen); omit posting_id to clear the entire Imbox tray in one call (POST /boxes/{imboxId}/observation). Reversible per-item via hey_mark_unseen. Use after triaging or skimming a batch to keep the tray tidy without re-marking threads unread.

ParametersJSON Schema
NameRequiredDescriptionDefault
posting_idNoOptional. The posting ID (from postingId field on list operations) to clear from the New for you tray. When omitted, the entire Imbox tray is marked seen in a single bulk request.

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds details beyond annotations: explains return format {success, error?}, HTTP methods, and that the action is reversible per-item. Consistent with idempotentHint and non-destructive nature.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Concise yet informative: each sentence adds value. Front-loaded with purpose, then specifics, then usage guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Fully describes both usage modes, return format, and reversibility. No output schema needed given simplicity and parameter count of 1.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Adds meaning beyond the schema by explaining that omitting posting_id clears entire tray and passing a specific ID clears one item. References postingId field from list operations for clarity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the tool clears the 'New for you' tray dot, specifying both per-item and bulk modes. It distinguishes from sibling tool hey_mark_unseen by mentioning reversibility.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly advises using after triaging or skimming a batch to keep tray tidy. Mentions reversal via hey_mark_unseen, providing clear alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_mark_unseenA
Idempotent

Mark a thread as unseen/unread to reset its read status. Returns {success, error?}. Reading the email via hey_read_email implicitly marks it as seen again. To clear the orange 'New for you' tray dot without re-marking unseen, use hey_mark_seen.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe topic/thread ID to mark as unseen

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate idempotentHint=true and readOnlyHint=false, but the description adds context about the return format ({success, error?}) and the side effect of reading the email implicitly marking it seen. This goes beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each one adding value. Front-loaded with the core action, then return type, then usage caveats. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple with one parameter and no output schema. The description fully covers purpose, usage guidelines, side effects, and return format, leaving no gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the description only restates what the schema already says ('The topic/thread ID to mark as unseen'). No additional meaning or clarification is provided.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's action: 'Mark a thread as unseen/unread to reset its read status.' It uses a specific verb and resource, and distinguishes itself from the sibling 'hey_mark_seen' by explaining the difference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit guidance on when to use this tool (to reset read status) and when not to (to clear the orange dot without re-marking unseen, use 'hey_mark_seen'). It also mentions the side effect of reading email marking it seen.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_move_toA
Idempotent

Move an email thread between Hey.com views: imbox, feed, or paper_trail. Returns {success, error?}. Use paper_trail for receipts/automated mail, feed for newsletters, imbox to restore. Reversible by moving to a different destination. Does not affect trash, spam, or screener — use hey_set_status or hey_screen for those.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe topic/thread ID to move (use topicId from list operations)
destinationYesTarget view: imbox (important mail), feed (newsletters/updates), paper_trail (receipts/automated)

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses return format ({success, error?}), reversibility, and scope boundaries beyond annotations. Annotations already indicate non-read-only, non-destructive, and idempotent; description adds context confirming these.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence provides essential information—action, return value, usage advice, reversibility, and exclusions. No redundancy; core purpose is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers primary use cases and return format. Minor gaps exist (e.g., error conditions for invalid id), but the presence of 'error?' in return implies error handling. With no output schema, this is sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema has 100% description coverage, so baseline is 3. Description adds value by explaining the meaning of each destination value and when to use them, going beyond the schema's enum labels.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the action (move email thread) and the specific views (imbox, feed, paper_trail). Differentiates from sibling tools by noting it does not affect trash, spam, or screener.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use each destination (paper_trail for receipts, feed for newsletters, imbox to restore) and provides alternatives for other actions (hey_set_status, hey_screen).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_pop_bubbleA
Idempotent

Pop (dismiss) a bubbled-up email thread so it sinks back into the Imbox. The thread is not deleted — it just stops being pinned at the top. Returns {success, error?}. Requires the thread's topic_id (use topic_id from any list operation); posting IDs are not accepted.

ParametersJSON Schema
NameRequiredDescriptionDefault
topic_idYesThe topic ID (thread ID) to pop/unbubble. Use the `topicId` field from hey_list_* responses (NOT `postingId` — that's a different ID and will 404).

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses idempotent behavior (consistent with annotations), non-destructive nature, return value {success, error?}, and error condition for incorrect ID types. Adds value beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Concise two-sentence description front-loaded with action and effect, no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 1-parameter tool without output schema, the description covers purpose, usage, prerequisites, return type, and error cases completely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema already covers 100% of parameters with detailed guidance on topic_id vs postingId. Description reinforces but adds minimal new information beyond schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool pops/dismisses a bubbled email thread, distinguishing it from sibling tools like hey_bubble_up by specifying the reverse action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit instructions on when to use, including the prerequisite of using topic_id (not posting_id) and clarifies that the thread is not deleted, just unpinned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_read_emailA
Read-only

Read an email thread's full content. Returns all messages in the thread via entries[] array (each with entryId, from, to, cc, date, body). Also returns attachments[] metadata and calendar_invites[] when present — use hey_download_attachment to save files to disk, or hey_get_calendar_invite to parse .ics details. Use format='html' (default) for rich content with thread entries, or format='text' for decoded RFC822 plain text of the first message.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe topic/thread ID or entry ID to read (use topicId from list operations for full threads)
formatNohtml (default): rich HTML with all thread entries. text: decoded plain text of the first message only
force_refreshNoBypass cache and fetch fresh data (default: false)

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations (readOnlyHint, openWorldHint) indicate safe, non-destructive operation. Description reinforces read-only nature and details output structure (entries[], attachments[], calendar_invites[]) and format behavior, adding value beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise sentences, front-loaded with main purpose, then details on output and alternative tools. No wasted words; each sentence adds distinct value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema, but description adequately explains return fields and links to sibling tools for next steps. With annotations, provides complete context for safe and effective use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but description adds context: explains id can be topicId or entryId, clarifies format default and behavior (html = full thread, text = first message only). Provides meaning beyond schema definitions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states it reads an email thread's full content, using specific verb 'Read' and resource 'email thread'. It distinguishes from siblings like hey_download_attachment and hey_get_calendar_invite by mentioning them for further actions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides guidance on when to use format='html' vs 'text', and references sibling tools for attachments and calendar invites. Implies using topicId from list operations, but does not explicitly state when not to use this tool (e.g., for single messages vs threads).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_read_statusA
Idempotent

Set the read/unread status of an email entry. Returns {success, error?}. Reversible by calling again with the opposite status. Operates on individual entries (use entryId), not whole threads. For marking an entire thread as unseen, use hey_mark_unseen instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe entry ID to update (use entryId from list operations)
statusYesTarget status: read or unread

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate the tool is not read-only, not destructive, and idempotent. The description adds the return format {success, error?}, reversibility, and per-entry scope, which are valuable but not extensive behavioral details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, no wasted words, front-loaded with primary action. Every sentence serves a purpose: main action, return value, reversibility, scope, and sibling reference.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool without output schema, the description covers usage, behavior, return value, and differentiation from siblings. No obvious gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema covers both parameters with descriptions, so baseline is 3. The description adds guidance to use entryId from list operations and explains that it operates on individual entries, providing context beyond schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the tool sets read/unread status of an email entry. Distinguishes itself from the sibling hey_mark_unseen by specifying it works on individual entries, not threads.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells when to use this tool (for individual entries) and when to use an alternative (hey_mark_unseen for entire threads). Also states the operation is reversible.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_remove_reply_laterA
Idempotent

Remove an email from Reply Later, marking it "Done" and moving it back to the Imbox. Returns {success, error?}. Requires the posting_id from hey_list_reply_later. Reversible by calling hey_reply_later again with the thread's topic_id.

ParametersJSON Schema
NameRequiredDescriptionDefault
posting_idYesThe posting ID to remove from Reply Later (use postingId from hey_list_reply_later)

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover safety flags (idempotentHint, destructiveHint=false, readOnlyHint=false), and the description adds substantive context beyond them: the "Done" status change, the move back to the Imbox, and the return shape {success, error?}. It also discloses reversibility and the exact recovery path, which is not derivable from the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, each carrying distinct information (action, side effects, return shape, prerequisite, reversal). The primary action is front-loaded and there is no redundant restatement of the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter mutation with no output schema, the description supplies everything needed: the source of the required id, the side effects, the return shape, and the undo path. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the parameter is documented, so the baseline is 3. The description adds value by stating where to obtain posting_id (hey_list_reply_later), giving the agent a concrete sourcing workflow rather than just a type.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (remove) and resource (email from Reply Later) and goes further by naming the observable side effects: marking the item "Done" and moving it back to the Imbox. This distinguishes it cleanly from siblings like hey_set_aside, hey_reply_later, and hey_unset_aside.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the prerequisite source for posting_id (hey_list_reply_later) and describes the inverse operation (hey_reply_later with topic_id), which is strong routing. It lacks an explicit when-not condition, which keeps it just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_replyA

Reply to an email thread, preserving threading on both ends. Returns {success, error?}. By default the reply goes to the other thread participants (your own address is automatically excluded); pass to to redirect to specific recipients — e.g. chasing your own thread without looping back to yourself, or redirecting off a mailing-list address. Prefer this over hey_send_email when responding to an existing thread; use hey_send_email only for new conversations.

ParametersJSON Schema
NameRequiredDescriptionDefault
ccNoOptional CC override. Only honoured when `to` is also provided.
toNoOptional override of the To: line. Replaces the auto-detected participants. Use to: (a) chase your own thread without looping back to yourself, (b) redirect a mailing-list reply to a specific person, or (c) generally target the reply at recipients other than the thread defaults.
bodyYesReply body content (HTML supported)
thread_idYesThe thread/topic ID to reply to

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare non-readOnly, non-idempotent, non-destructive, open-world. The description adds real value beyond them: the return shape ({success, error?}), the default recipient behavior (own address auto-excluded), and threading preservation. It doesn't mention auth needs or rate limits, but for this tool those are minor.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and return shape, then recipient defaults, then override cases, then sibling routing. Every sentence earns its place and nothing is repeated verbatim from the schema without purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by stating the return shape. Combined with the recipient-default and threading notes, an agent has everything needed to invoke this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaning the schema lacks: the default recipient set and the automatic exclusion of your own address, which explains why `to` exists at all. The detailed `to`/`cc` mechanics are largely mirrored from the schema, capping it below 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource ('Reply to an email thread') with a stated distinguishing trait ('preserving threading on both ends'). It explicitly names the sibling it is not (hey_send_email) and the condition that differentiates them, so an agent can disambiguate without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit routing guidance: 'Prefer this over hey_send_email when responding to an existing thread; use hey_send_email only for new conversations.' It also explains when to reach for the `to` override (chasing your own thread, redirecting off a mailing list).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_reply_laterA
Idempotent

Move an email thread to Reply Later. Reversible via hey_remove_reply_later (requires postingId from hey_list_reply_later). Returns {success, error?}. Use for emails you intend to respond to but not right now — if no reply is planned, use hey_set_aside instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe topic or entry ID to mark for reply later (use topicId or entryId from list operations)

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this is a mutation (readOnlyHint=false), idempotent, and non-destructive; the description adds genuinely new context: reversibility via hey_remove_reply_later, the postingId prerequisite from hey_list_reply_later, and the return shape {success, error?}. It stops short of describing side effects like removal from the Imbox, which keeps it below a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short clauses: action, reversal path, decision rule with alternative. Front-loaded and every sentence earns its place with no padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description supplies the return shape, the undo path, and the routing rule against a sibling — everything an agent needs to call and recover from this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single 'id' parameter is documented there (topicId or entryId from list operations), so the baseline applies. The description's mention of postingId pertains to the reversal tool rather than the input, adding no new meaning about the argument itself.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Move an email thread to Reply Later') and explicitly contrasts with siblings hey_set_aside and hey_remove_reply_later, so an agent can distinguish it without opening other schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('emails you intend to respond to but not right now') and when-not-to-use with a named alternative ('if no reply is planned, use hey_set_aside instead'), plus the reversal path. Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_save_draftA

Save a new draft email without sending it. Creates an entry in Hey's Drafts folder; recipients, subject, and body are all optional so a partially-written draft can be saved. Returns {success, draftId?, error?}. Sending isn't available via MCP yet — finish and send the draft from the Hey web/app UI, or use hey_edit_draft to keep revising it. Use hey_send_email/hey_reply instead when you want to send immediately rather than draft first.

ParametersJSON Schema
NameRequiredDescriptionDefault
ccNoOptional list of CC recipient email addresses
toNoOptional list of recipient email addresses
bodyNoOptional email body content (HTML supported)
subjectNoOptional email subject line

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, and idempotentHint=false, so the safety profile is partly covered. The description adds real value: it names the Drafts folder destination, discloses the return shape ({success, draftId?, error?}), and surfaces the key workflow limitation that sending isn't available via MCP.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then folder destination, optionality, return shape, and alternatives. Dense but every sentence carries distinct information; nothing is redundant with the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description supplies the return shape, the folder behavior, the optionality semantics, and the critical workflow limitation that sending must happen outside MCP. An agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the params are already documented (baseline 3). The description adds rationale beyond the schema: recipients, subject, and body are all optional specifically so a partially-written draft can be saved, which explains the intended usage pattern.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Save a new draft email') plus scope ('without sending it'). It's immediately distinguishable from siblings like hey_edit_draft, hey_send_email, and hey_reply, which it names directly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: use hey_edit_draft to keep revising, hey_send_email/hey_reply to send immediately, and finish in the web/app UI since MCP can't send. When-to-use, when-not, and alternatives are all stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_screenA
DestructiveIdempotent

Approve or reject a sender by email address. Approve: routes the sender's current and future emails into the chosen destination (defaults to imbox). Reject (a.k.a. screen out): blocks the sender from sending you further emails — works for both pending screener entries AND already-approved senders (falls back to the contact-page 'Screened Out' affordance via /contacts/{id}/clearance). Reject does NOT flag emails as spam; existing emails are left untouched. Reversible from the Hey UI by visiting the contact page; not yet exposed via MCP. Returns {success, error?}. Use hey_list_screener to see pending senders, or hey_screen_by_id for clearance IDs.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesapprove routes to `destination`; reject blocks future emails and does not affect existing mail.
destinationNoWhere future emails from this sender land when approved: imbox (default, important mail), feed (newsletters/updates), paper_trail (receipts/automated). Ignored when action is reject.
sender_emailYesThe sender's email address

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructive/idempotent/openWorld, and the description adds substantial non-redundant context: reject does NOT mark spam, existing emails are untouched, the operation falls back to the contact-page clearance route, and it is reversible in the Hey UI but not via MCP. It even discloses the return shape, which annotations cannot.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the verb pair, then layers approve/reject semantics economically in a few tight sentences. Slightly heavy on parenthetical implementation detail (the /contacts/{id}/clearance path), but nearly every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-param mutation with no output schema, the description covers action semantics, edge cases (already-approved senders), side effects, reversibility, and return shape. Nothing the agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already documents destination defaults and that it is ignored on reject; the description largely restates these. It adds the 'screen out' alias and the approve routing/default, but no syntax or format detail beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb pair and resource ('Approve or reject a sender by email address') and explicitly differentiates from siblings by naming hey_list_screener and hey_screen_by_id. An agent can pick this tool over those without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use routing ('Use hey_list_screener to see pending senders, or hey_screen_by_id for clearance IDs') and clarifies that reject applies to both pending and already-approved senders. The alternative tools and the condition selecting them are stated outright.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_screen_by_idA
DestructiveIdempotent

Approve or reject a first-time sender from the Screener by clearance ID. Approve: allows future emails from this sender into the chosen destination (defaults to imbox). Reject: blocks future emails from this sender via the screener. Reversible from the Hey UI's contact page (not yet via MCP). Returns {success, error?}. Use hey_list_screener to get clearance IDs; for senders that have already left the screener, use hey_screen by email instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesapprove routes to `destination`; reject blocks future emails and does not flag as spam.
destinationNoWhere future emails from this sender land when approved: imbox (default, important mail), feed (newsletters/updates), paper_trail (receipts/automated). Ignored when action is reject.
clearance_idYesThe clearance ID from hey_list_screener

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations (destructiveHint, idempotentHint) by spelling out the downstream effect of each action on future mail flow, the default destination, and the reversibility constraint ('reversible from the Hey UI's contact page, not yet via MCP'). It even discloses the return shape '{success, error?}' in the absence of an output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the action and resource, then branches cleanly into Approve/Reject effects, a reversibility note, the return contract, and sibling routing. Every sentence carries information; none is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description supplies the return contract, both action semantics, the destination default, reversibility limits, and sibling navigation. An agent has everything needed to invoke this destructive, idempotent tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already documents the destination default (imbox) and enum meanings, so the description largely restates what is structured. It adds only light value by tying destination to the approve branch and clarifying it is ignored on reject.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific compound action (approve/reject) on a precise resource (first-time sender in the Screener) keyed by clearance ID. It is immediately distinguishable from hey_list_screener and hey_screen by the ID-vs-email keying, which the description itself calls out.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: 'Use hey_list_screener to get clearance IDs' and 'for senders that have already left the screener, use hey_screen by email instead.' Both the prerequisite and the exclusion condition that selects a sibling are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_send_emailA

Send a new email immediately (no draft stage). Returns {success, error?}. Use for standalone outbound messages; use hey_reply for thread responses, or hey_forward to share existing emails with new recipients.

ParametersJSON Schema
NameRequiredDescriptionDefault
ccNoList of CC recipient email addresses
toYesList of recipient email addresses
bodyYesEmail body content (HTML supported)
subjectYesEmail subject line

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations indicate mutating operation (readOnlyHint=false) and description adds immediate sending and return shape, but lacks details on potential side effects or failure modes beyond a brief mention.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no unnecessary words; first sentence defines action and return, second provides direct usage comparison to siblings.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers purpose, usage alternatives, and return format briefly. Lacks details on constraints like attachment support or rate limits, but sufficient for a simple send email tool given sibling differentiation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with descriptions for all parameters; description adds no extra semantic meaning beyond the schema, meeting baseline expectations.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states 'Send a new email immediately' and differentiates from sibling tools hey_reply and hey_forward with specific usage contexts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use this tool vs alternatives: 'Use for standalone outbound messages; use hey_reply for thread responses, or hey_forward to share existing emails with new recipients.'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_set_asideA
Idempotent

Move an email thread to Set Aside for later. Reversible via hey_unset_aside (requires postingId from hey_list_set_aside). Returns {success, error?}. Use for emails you plan to revisit but don't need to reply to — for emails needing a reply, use hey_reply_later instead. Does not affect future emails from the sender.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe topic or entry ID to set aside (use topicId or entryId from list operations)

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare it non-read-only, idempotent, and non-destructive; the description adds the operational details that matter: the action is reversible via hey_unset_aside, which requires a postingId obtained from hey_list_set_aside, the response shape is {success, error?}, and it does not affect future emails from the sender.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four compact sentences with the core action, reversibility, return shape, and routing rule front-loaded; no filler and nothing repeated from structured fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter mutation with no output schema, the description supplies the return shape, the undo path, and the side-effect boundary, leaving nothing an agent needs to invoke it correctly unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single parameter is already documented there (topicId/entryId). The description adds no syntax or format detail for the id beyond what the schema provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Move) and resource (email thread) plus the destination state (Set Aside). It differentiates itself from siblings by naming hey_reply_later as the alternative for a different intent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('emails you plan to revisit but don't need to reply to'), when-not ('emails needing a reply'), and the alternative tool (hey_reply_later). Also names the reversal path (hey_unset_aside) and the prerequisite ID source.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_set_statusA
DestructiveIdempotent

Change an email thread's status. Returns {success, error?}. Trash and spam are reversible via restore and unspam actions respectively. DESTRUCTIVE: trash removes from Imbox, spam blocks the sender. Paper Trail bundles (postingId-only items with no thread) support action=trash only; spam/restore/unspam on a bundle return an explicit error pointing to the bundle's individual entries.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe topic/thread ID (use topicId from list operations). For Paper Trail bundles, pass the postingId — trash works via a posting-based fallback, other actions return an error.
actionYesThe status action: trash (move to Trash), restore (recover from Trash), spam (mark as spam and block sender), unspam (restore from spam folder)

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide destructiveHint and idempotentHint; description adds concrete details: 'trash removes from Imbox', 'spam blocks the sender', and error handling for bundles. No contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences with no unnecessary words. Front-loaded with purpose, then return format, then edge cases. Highly efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Addresses return format, destructive effects, reversibility, and special case (bundles). No output schema needed given simplicity. Complete for minimum viable understanding.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema covers both parameters with descriptions (100% coverage). Description adds context on id parameter usage (topicId vs postingId) and action effects, but enum already lists actions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states it changes an email thread's status, which is a specific verb+resource. It distinguishes from sibling tools like hey_label or hey_collection by focusing on status actions (trash, spam, etc.).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit context on reversibility (trash/spam), destructive effects, and bundle-specific limitations (only trash works on bundles). Could mention when to use this vs other status-changing tools like hey_screen, but the domain is different.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_thread_muteA
Idempotent

Mute or unmute a thread (called 'Ignore' in Hey.com's UI). Muting stops notifications for the thread but keeps it in its current view — the thread is not moved or deleted. Returns {success, error?}. Reversible by calling with the opposite action. To check if a thread is currently muted, use hey_read_email — the response includes a 'muted' field.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesmute to stop notifications, unmute to resume them
posting_idYesThe posting ID of the thread (use postingId from list operations)

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses key behaviors: reversible, keeps thread in current view, returns {success, error?}. Aligns with annotations (idempotentHint true, destructiveHint false). No contradictions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise sentences front-loading the core purpose, with no extraneous text. Every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple tool, annotations, and full schema coverage, the description fully satisfies the agent's needs. Mentions return format despite no output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema covers both parameters (100%). Description adds context for posting ID (use from list ops) and action enum values, going beyond schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states 'Mute or unmute a thread' with the Hey.com term 'Ignore' for clarity. Distinguishes from siblings like hey_read_email (which checks mute status) and hey_move_to (which moves threads).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains when to use (stop/resume notifications) and provides explicit guidance to check mute status via hey_read_email. Could add more on when not to use, but sufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hey_unset_asideA
Idempotent

Remove an email from Set Aside, moving it back to the Imbox or its original location. Returns {success, error?}. Requires the posting_id from hey_list_set_aside. Reversible by calling hey_set_aside again with the thread's topic_id.

ParametersJSON Schema
NameRequiredDescriptionDefault
posting_idYesThe posting ID to remove from Set Aside (use postingId from hey_list_set_aside)

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=true), and the description adds genuinely useful context beyond them: the return shape {success, error?} and the fact that the action is reversible via hey_set_aside. It does not mention failure modes or what happens if the posting_id is absent, which keeps it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the action and effect, followed by return shape and prerequisite. No filler and every sentence carries operational information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool with no output schema, the description covers the effect, the return shape, the parameter source, and reversibility. An agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single parameter's description already states it comes from hey_list_set_aside, so the description largely repeats rather than extends the schema. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (Remove) and resource (an email from Set Aside), and states the resulting state change (moved back to the Imbox or original location). It is clearly distinguishable from the sibling hey_set_aside, which performs the inverse operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the prerequisite (requires posting_id from hey_list_set_aside) and names the alternative operation that reverses it (hey_set_aside with topic_id). It does not spell out a when-not-to-use condition, but the context is unambiguous for a single-purpose unsave tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 11 tool updatesv0.5.0
    • Changedhey_cache_status1 field changed
      • changedInput schema / properties / folder / description
        Previous value: -"Optional folder to get specific stats for"New value: +"Optional: scope stats to one cached view (imbox, feed, paper_trail, set_aside, reply_later). Note this differs from hey_list_emails' folder enum — trash, spam, and drafts aren't cache-tracked and silently return 0 here."
    • Addedhey_delete_draft
    • Changedhey_download_attachment1 field changed
      • changedInput schema / properties / attachment_id / description
        Previous value: -"The attachment ID from hey_read_email's attachments array (e.g. 'part-1')"New value: +"The attachment ID from hey_read_email's attachments array (format: part-N, e.g. 'part-1')"
    • Addedhey_edit_draft
    • Changedhey_list_collection_emails3 fields changed
      • changedInput schema / properties / collection_id / description
        Previous value: -"The collection ID to list emails from"New value: +"The collection ID to list emails from (use id from hey_list_collections)"
      • changedInput schema / properties / limit / description
        Previous value: -"Maximum number of emails to return (default: 25)"New value: +"Maximum number of emails to return, 1-100 (default: 25)"
      • changedInput schema / properties / page / description
        Previous value: -"Page number for pagination (default: 1)"New value: +"Page number for pagination, 1-1000 (default: 1)"
    • Changedhey_list_emails4 fields changed
      • changedInput schema / properties / folder / description
        Previous value: -"The folder/view to list emails from: imbox (important), feed (newsletters), paper_trail (receipts), trash, spam, or drafts"New value: +"The folder/view to list emails from: imbox (important), feed (newsletters), paper_trail (receipts), trash, spam, drafts, or sent (emails you sent)"
      • changedInput schema / properties / folder / enum
        Previous value: -[
        -  "imbox",
        -  "feed",
        -  "paper_trail",
        -  "trash",
        -  "spam",
        -  "drafts"
        -]New value: +[
        +  "imbox",
        +  "feed",
        +  "paper_trail",
        +  "trash",
        +  "spam",
        +  "drafts",
        +  "sent"
        +]
      • changedInput schema / properties / limit / description
        Previous value: -"Maximum number of emails to return (default: 25)"New value: +"Maximum number of emails to return, 1-100 (default: 25)"
      • changedInput schema / properties / page / description
        Previous value: -"Page number for pagination (default: 1)"New value: +"Page number for pagination, 1-1000 (default: 1)"
    • Changedhey_list_label_emails3 fields changed
      • changedInput schema / properties / label_id / description
        Previous value: -"The label/folder ID to list emails from"New value: +"The label ID to list emails from (use id from hey_list_labels)"
      • changedInput schema / properties / limit / description
        Previous value: -"Maximum number of emails to return (default: 25)"New value: +"Maximum number of emails to return, 1-100 (default: 25)"
      • changedInput schema / properties / page / description
        Previous value: -"Page number for pagination (default: 1)"New value: +"Page number for pagination, 1-1000 (default: 1)"
    • Addedhey_save_draft
    • Changedhey_screen1 field changed
      • changedInput schema / properties / action / description
        Previous value: -"approve: allow this sender's emails through (routed to `destination`). reject: block future emails from this sender. Does not flag as spam, does not move existing emails. Reversible via the Hey UI's contact page."New value: +"approve routes to `destination`; reject blocks future emails and does not affect existing mail."
    • Changedhey_screen_by_id1 field changed
      • changedInput schema / properties / action / description
        Previous value: -"approve: allow this sender's emails through (routed to `destination`). reject: block future emails from this sender. Does not flag as spam."New value: +"approve routes to `destination`; reject blocks future emails and does not flag as spam."
    • Changedhey_search2 fields changed
      • changedInput schema / properties / limit / description
        Previous value: -"Maximum number of results (default: 25)"New value: +"Maximum number of results, 1-100 (default: 25)"
      • changedInput schema / properties / query / description
        Previous value: -"Search query"New value: +"Search query text, 1-500 characters after trimming."
  2. 33 tool updatesv0.4.1
    • Removedhey_add_label
    • Removedhey_add_to_collection
    • Changedhey_bubble_up3 fields changed
      • removedInput schema / properties / posting_id
        Removed value: -{
        -  "description": "The topic or posting ID to schedule (use topicId for 'now' slot)",
        -  "type": "string"
        -}
      • addedInput schema / properties / topic_id
        Added value: +{
        +  "description": "The topic ID (thread ID) to schedule. Use the `topicId` field from hey_list_* responses (NOT `postingId` — that's a different ID and will 404).",
        +  "type": "string"
        +}
      • changedInput schema / required
        Previous value: -[
        -  "posting_id",
        -  "slot"
        -]New value: +[
        +  "topic_id",
        +  "slot"
        +]
    • Changedhey_bubble_up_if_no_reply3 fields changed
      • removedInput schema / properties / posting_id
        Removed value: -{
        -  "description": "The topic or posting ID to schedule (use topicId preferred)",
        -  "type": "string"
        -}
      • addedInput schema / properties / topic_id
        Added value: +{
        +  "description": "The topic ID (thread ID) to schedule. Use the `topicId` field from hey_list_* responses (NOT `postingId` — that's a different ID and will 404).",
        +  "type": "string"
        +}
      • changedInput schema / required
        Previous value: -[
        -  "posting_id",
        -  "date"
        -]New value: +[
        +  "topic_id",
        +  "date"
        +]
    • Addedhey_collection
    • Changedhey_download_attachment1 field changed
      • changedInput schema / properties / save_path / description
        Previous value: -"Optional absolute path or directory to save into. Must be within the user's home directory (~). Defaults to ~/Downloads/hey-attachments/<email_id>/<filename>. Trailing '/' is treated as a directory."New value: +"Optional path or directory to save into. Must be within ~/. Defaults to ~/Downloads/hey-attachments/<date>/<filename>. Trailing '/' is treated as a directory. Duplicate filenames are auto-numbered (invite-1.ics, invite-2.ics)."
    • Removedhey_ignore_thread
    • Addedhey_label
    • Removedhey_list_drafts
    • Addedhey_list_emails
    • Removedhey_list_feed
    • Removedhey_list_imbox
    • Removedhey_list_paper_trail
    • Removedhey_list_spam
    • Removedhey_list_trash
    • Addedhey_mark_seen
    • Removedhey_not_spam
    • Changedhey_pop_bubble3 fields changed
      • removedInput schema / properties / posting_id
        Removed value: -{
        -  "description": "The topic or posting ID to pop/unbubble (use topicId preferred)",
        -  "type": "string"
        -}
      • addedInput schema / properties / topic_id
        Added value: +{
        +  "description": "The topic ID (thread ID) to pop/unbubble. Use the `topicId` field from hey_list_* responses (NOT `postingId` — that's a different ID and will 404).",
        +  "type": "string"
        +}
      • changedInput schema / required
        Previous value: -[
        -  "posting_id"
        -]New value: +[
        +  "topic_id"
        +]
    • Addedhey_read_status
    • Removedhey_remove_from_collection
    • Removedhey_remove_label
    • Changedhey_reply1 field changed
      • changedInput schema / properties / to / description
        Previous value: -"Optional override of the To: line. Use this when chasing a thread where you sent the most recent message, so the chase lands on the original recipient instead of looping back to your own address."New value: +"Optional override of the To: line. Replaces the auto-detected participants. Use to: (a) chase your own thread without looping back to yourself, (b) redirect a mailing-list reply to a specific person, or (c) generally target the reply at recipients other than the thread defaults."
    • Removedhey_restore
    • Addedhey_screen
    • Addedhey_screen_by_id
    • Removedhey_screen_in
    • Removedhey_screen_in_by_id
    • Removedhey_screen_out
    • Addedhey_set_status
    • Removedhey_spam
    • Addedhey_thread_mute
    • Removedhey_trash
    • Removedhey_unignore_thread
  3. 3 tool updatesv0.3.4
    • Changedhey_download_attachment2 fields changed
      • changedInput schema / properties / email_id / description
        Previous value: -"The email ID containing the attachment"New value: +"The email's topic or entry ID (same ID used with hey_read_email — topic IDs are resolved automatically)"
      • changedInput schema / properties / save_path / description
        Previous value: -"Optional absolute path or directory to save into. Defaults to ~/Downloads/hey-attachments/<email_id>/<filename>. Trailing '/' is treated as a directory."New value: +"Optional absolute path or directory to save into. Must be within the user's home directory (~). Defaults to ~/Downloads/hey-attachments/<email_id>/<filename>. Trailing '/' is treated as a directory."
    • Changedhey_get_calendar_invite1 field changed
      • changedInput schema / properties / email_id / description
        Previous value: -"The email ID containing the calendar invite"New value: +"The email's topic or entry ID (same ID used with hey_read_email — topic IDs are resolved automatically)"
    • Changedhey_read_email2 fields changed
      • changedInput schema / properties / format / description
        Previous value: -"Format to return (default: html)"New value: +"html (default): rich HTML with all thread entries. text: decoded plain text of the first message only"
      • changedInput schema / properties / id / description
        Previous value: -"The email ID to read"New value: +"The topic/thread ID or entry ID to read (use topicId from list operations for full threads)"
  4. 7 tool updatesv0.3.0
    • Changedhey_bubble_up1 field changed
      • changedInput schema / properties / posting_id / description
        Previous value: -"The posting ID to schedule"New value: +"The topic or posting ID to schedule (use topicId for 'now' slot)"
    • Changedhey_bubble_up_if_no_reply1 field changed
      • changedInput schema / properties / posting_id / description
        Previous value: -"The posting ID to schedule"New value: +"The topic or posting ID to schedule (use topicId preferred)"
    • Addedhey_move_to
    • Removedhey_move_to_paper_trail
    • Changedhey_pop_bubble1 field changed
      • changedInput schema / properties / posting_id / description
        Previous value: -"The posting ID to pop/unbubble"New value: +"The topic or posting ID to pop/unbubble (use topicId preferred)"
    • Changedhey_reply_later3 fields changed
      • removedInput schema / properties / entry_id
        Removed value: -{
        -  "description": "The entry ID to mark for reply later (use entryId from list operations)",
        -  "type": "string"
        -}
      • addedInput schema / properties / id
        Added value: +{
        +  "description": "The topic or entry ID to mark for reply later (use topicId or entryId from list operations)",
        +  "type": "string"
        +}
      • changedInput schema / required
        Previous value: -[
        -  "entry_id"
        -]New value: +[
        +  "id"
        +]
    • Changedhey_set_aside3 fields changed
      • removedInput schema / properties / entry_id
        Removed value: -{
        -  "description": "The entry ID to set aside (use entryId from list operations)",
        -  "type": "string"
        -}
      • addedInput schema / properties / id
        Added value: +{
        +  "description": "The topic or entry ID to set aside (use topicId or entryId from list operations)",
        +  "type": "string"
        +}
      • changedInput schema / required
        Previous value: -[
        -  "entry_id"
        -]New value: +[
        +  "id"
        +]
  5. 44 tool updatesv0.1.0
    • First observedhey_add_label
    • First observedhey_add_to_collection
    • First observedhey_bubble_up
    • First observedhey_bubble_up_if_no_reply
    • First observedhey_cache_status
    • First observedhey_download_attachment
    • First observedhey_forward
    • First observedhey_get_calendar_invite
    • First observedhey_ignore_thread
    • First observedhey_imbox_summary
    • First observedhey_list_collection_emails
    • First observedhey_list_collections
    • First observedhey_list_drafts
    • First observedhey_list_feed
    • First observedhey_list_imbox
    • First observedhey_list_label_emails
    • First observedhey_list_labels
    • First observedhey_list_paper_trail
    • First observedhey_list_reply_later
    • First observedhey_list_screener
    • First observedhey_list_set_aside
    • First observedhey_list_spam
    • First observedhey_list_trash
    • First observedhey_mark_unseen
    • First observedhey_move_to_paper_trail
    • First observedhey_not_spam
    • First observedhey_pop_bubble
    • First observedhey_read_email
    • First observedhey_remove_from_collection
    • First observedhey_remove_label
    • First observedhey_remove_reply_later
    • First observedhey_reply
    • First observedhey_reply_later
    • First observedhey_restore
    • First observedhey_screen_in
    • First observedhey_screen_in_by_id
    • First observedhey_screen_out
    • First observedhey_search
    • First observedhey_send_email
    • First observedhey_set_aside
    • First observedhey_spam
    • First observedhey_trash
    • First observedhey_unignore_thread
    • First observedhey_unset_aside

TDQS

A4/5.0

Scored across 37 tools

Disambiguation4/5

Most tools target clearly distinct actions, and descriptions explicitly guide choice between similar pairs (hey_screen vs hey_screen_by_id, hey_set_aside vs hey_reply_later, hey_bubble_up vs hey_bubble_up_if_no_reply). The main soft spot is the read-state trio (hey_mark_seen, hey_mark_unseen, hey_read_status), which is subtly differentiated (tray dot vs thread vs entry) and could trip up an agent, but the descriptions do address it.

Naming Consistency4/5

Every tool uses a consistent hey_ snake_case prefix, which makes the set scan cleanly. Minor deviation: some names are verb_noun (hey_list_emails, hey_send_email) while others are state-noun or ad hoc (hey_imbox_summary, hey_cache_status, hey_thread_mute), but the convention is stable and readable throughout.

Tool Count3/5

At 37 tools this is on the heavy side and pushes past the comfortable 3-15 band. The email-client domain justifies breadth, but several clusters could be consolidated (three read-state tools, three bubble-up tools, separate list/view tools) rather than expanded separately.

Completeness4/5

Coverage is broad: listing across views, search, read, reply/forward/send, drafts, labels, collections, screener, set-aside/reply-later, bubbling, muting, moving, and status changes. Explicit gaps are acknowledged (drafts can't be sent via MCP, screener reversals are UI-only), which are minor relative to the surface.

Maintenance

ActivityMaintained
ResponsivenessSlow

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables semantic search and AI-powered analysis of Outlook emails using RAG-based natural language queries and Vision AI for architectural documents, with specialized support for AEC workflows.
    MIT
  • A
    license
    A
    quality
    D
    maintenance
    Local MCP server for multi-account IMAP/SMTP email (iCloud + Gmail via app-specific passwords). Never marks mail read. Cross-folder search, idempotent sends, TLS verified.
    8
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    An MCP server that exposes a local notmuch email database to an LLM client such as Claude. It is read-first: searching, reading, and understanding mail is always available; writing anything (drafts, tags, exported files) requires an explicit opt-in flag and is confined to clearly bounded locations.
    13
    2
    MIT