Mailsac MCP Server
OfficialClick on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Mailsac MCP ServerSign up for a new account on localhost:3000 and confirm it by email."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Mailsac MCP server
The official Model Context Protocol server for Mailsac. It gives AI coding agents (Claude Code, Cursor, VS Code, Claude Desktop and other MCP clients) disposable test inboxes, so an agent can check the email your application sends: sign-up confirmations, password resets, magic links and one-time codes.
A typical conversation:
Sign up for a new account on http://localhost:3000 and confirm it by email.
The agent creates a unique address with Mailsac, fills in your sign-up form, waits for the confirmation email, follows the link, and tells you what it found. The same tools help it write and debug end-to-end tests for those flows.
Tools
Tool | What it does | Mailsac operations |
| A new unique address, e.g. | none |
| Waits until a matching email arrives (optionally by subject, sender or time), then returns the subject, text, links, the most likely confirm/reset/login link ( | 1 per check, plus 1 to read |
| Emails stored at an address, newest first. | 1 |
| One email as text (with links and codes), HTML, raw MIME or headers. | 1 |
| Delete one email, or all emails at an address. | 1 per call or message |
| Your account's private domains, to use with | 1 |
Related MCP server: agent-inbox
Set up
You need a Mailsac API key. The free plan works: create an account, then create a key.
Claude Code
claude mcp add mailsac -e MAILSAC_API_KEY=your-key -- npx -y @mailsac/mcpCursor (.cursor/mcp.json), Claude Desktop (claude_desktop_config.json) and other clients that use the
same format:
{
"mcpServers": {
"mailsac": {
"command": "npx",
"args": ["-y", "@mailsac/mcp"],
"env": { "MAILSAC_API_KEY": "your-key" }
}
}
}VS Code (.vscode/mcp.json):
{
"servers": {
"mailsac": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@mailsac/mcp"],
"env": { "MAILSAC_API_KEY": "your-key" }
}
}
}Settings
Variable | Default | Purpose |
| (required) | Your Mailsac API key |
|
| Domain for new test addresses. Set it to your private domain. |
|
| API base URL |
Public inboxes and private domains
Addresses at @mailsac.com are public: anyone who guesses an address can read its mail. They are ideal for
made-up test data and need no setup. For anything real, such as a staging environment that sends real
customer names, use a private domain: your own subdomain (for example test.example.com) or a zero-setup
yourteam.msdc.co subdomain from the domains page. Set MAILSAC_DOMAIN and every
new address uses it.
Operations and limits
Each API call uses one Mailsac operation. wait_for_email checks every 3 seconds
by default, so a message that arrives within a few seconds costs about 2 to 5 operations. The free plan
includes 1,500 operations a month; paid plans start at 25,000.
How Mailsac knows this is the MCP server
Requests from this server carry Mailsac-Client: mcp and a mailsac-mcp/<version> user agent. Mailsac
counts these in aggregate to see how many people use Mailsac through AI agents. Email content is never part of that
count.
Develop
npm install
npm test # builds, then runs the tests with a fake Mailsac API
MAILSAC_API_KEY=your-key node dist/index.js # runs the server over stdioLicense
MIT
Available Tools
6 toolscreate_test_addressCreate a test email addressA
Return a new, unique address that can receive email immediately, with nothing to set up. Use one address per test so tests never read each other's mail. Makes no API call.
| Name | Required | Description | Default |
|---|---|---|---|
| domain | No | Domain to use. Defaults to mailsac.com. Use a private domain for real or sensitive data. | |
| prefix | No | Readable start of the address, e.g. "signup" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations the description carries the full burden, and it delivers meaningful traits: the address is unique, immediately receivable, needs "nothing to set up," and notably "Makes no API call." That last point usefully tells the agent there is no network round-trip, rate limit, or auth wall. It omits any note on persistence, expiry, or cleanup expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core purpose and escalating to a usage rule and a behavioral caveat. No sentence is filler; even the terse "Makes no API call" earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-required-parameter generator with no output schema, the description covers purpose, the intended per-test usage pattern, and the key behavioral fact that no API call occurs. It could go further on what the returned address looks like or whether addresses expire, but nothing critical to correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already explains domain (with default and privacy advice) and prefix. The description adds nothing about either parameter, which is acceptable when the schema does the work. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Return a new, unique address that can receive email immediately") and implicitly separates itself from the consumer siblings (wait_for_email, read_email, delete_emails). It never names a sibling or explicitly contrasts its role, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete usage direction: "Use one address per test so tests never read each other's mail," which tells the agent the intended lifecycle pattern. There is no when-not guidance or explicit routing to alternatives, but the context for calling it is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_emailsDelete emailsADestructive
Delete one email (messageId) or every email at an address (omit messageId). Clean up after a test.
| Name | Required | Description | Default |
|---|---|---|---|
| address | Yes | ||
| messageId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
destructiveHint=true already tells the agent this is a destructive operation, and the description's omitting-messageId clause usefully adds that the bulk case wipes an entire address rather than one message. However it never states irreversibility, whether deletion can be undone, or any auth/ownership requirement — gaps that matter for a destructive tool with a low annotation bar.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, with the branching behavior front-loaded and the motivating use case last. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with no output schema, the description covers both the single and bulk modes and the intended cleanup scenario. The remaining gap is that a destructive operation is not flagged as irreversible and messageId's provenance is unstated, but the destructiveHint annotation partially compensates.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden, and it does convey the key semantics: messageId present means one message, absent means all mail at address. It does not explain the format/origin of messageId (e.g. obtained from list_emails or read_email) nor the address format, so coverage is good but not complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (delete) and resource (email/emails), and explicitly disambiguates its two operating modes: single message via messageId, or the whole mailbox at an address when messageId is omitted. Among siblings (list_emails, read_email, create_test_address) only this tool deletes, so its role is unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Clean up after a test" gives one intended context and hints at a paired workflow with wait_for_email/read_email, but there is no guidance on when not to use it, no warning about permanent loss, and no mention of constraints (e.g. only owned test addresses). The usage context is implied rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_domainsList private domainsA
List the private domains on this Mailsac account. Mail to a private domain is visible only to the account; pass one as domain to create_test_address.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, but it does add a genuine behavioral fact: mail to a private domain is visible only to the account. It still omits read-only confirmation, pagination behavior, and what the account scope actually means for results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, with the primary action front-loaded and the usage hint following immediately. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter list tool with no output schema, the description covers what is listed, the visibility semantics, and the follow-up call. It could state the return shape (e.g. a list of domain names) explicitly, but nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. The description goes slightly beyond by naming the downstream parameter slot (`domain` on create_test_address) where a returned value is consumed, which helps an agent chain calls.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource tied to the Mailsac account ('List the private domains on this Mailsac account'), which is unambiguous about what is returned. It does not differentiate against sibling list tools (list_emails, list_domains vs them) beyond naming create_test_address as a downstream consumer.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the use case by explaining that a listed domain can be passed as `domain` to create_test_address, which is a helpful workflow pointer. However, it never states when to call this tool versus alternatives such as list_emails, nor any prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_emailsList emails at an addressA
List the emails currently stored for an address, newest first (no bodies). One operation.
| Name | Required | Description | Default |
|---|---|---|---|
| address | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses ordering ('newest first'), the exclusion of bodies, and scope ('currently stored'), but says nothing about pagination, result limits, permissions, or what 'One operation' actually means — which reads as cryptic rather than informative.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the core action and result shape, with the ordering and body-exclusion constraints trailing compactly. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description does the important work of describing the return set (emails, newest first, bodies omitted). The main gaps are pagination/limits and the meaning of the trailing 'One operation', which are not explained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% for the single 'address' parameter; the schema only declares type=string/format=email. The description adds that the address selects the mailbox whose emails are listed, which is minimal but consistent with the schema's format hint.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource ('List the emails') scoped to 'an address', and it distinguishes itself from read_email by noting there are 'no bodies'. It does not name siblings explicitly, but the resource and result shape are unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the '(no bodies)' note hints that read_email is the tool for message content, and 'newest first' signals a listing/scan use case. There is no explicit when-to-use or when-not-to-use guidance against wait_for_email, delete_emails, or list_domains.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_emailRead an emailB
Read one email. format "text" (default) also returns links, actionLink and codes; "html" returns the HTML body; "raw" the full MIME source; "headers" the parsed headers.
| Name | Required | Description | Default |
|---|---|---|---|
| format | No | text | |
| address | Yes | ||
| messageId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden, and it does disclose what each format returns (links/actionLink/codes for text, HTML body, MIME source, parsed headers). It omits whether reading marks the message as read, permission requirements, and error behavior, so it is partially complete but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One tight sentence, front-loaded with the core purpose before enumerating format behaviors. The semicolon-dense format list is compressed but every clause adds information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description covers the main gap by explaining the return shape for each format. A simple read tool needs little more, though side effects and failure modes are unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it fully explains the format enum's four values and their outputs. address and messageId remain undocumented, though their names and required status make their meaning self-evident.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read one email'), with 'one' implying single-message scope versus the sibling list_emails. It does not name an alternative tool or clarify the boundary with wait_for_email, so the differentiation is only implied.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use or when-not-to-use guidance, and no sibling is named as an alternative. The reader can infer that a messageId and address are prerequisites from the required parameters, but the description itself offers no routing context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wait_for_emailWait for an emailA
Poll an address until a matching email arrives, then return its subject, sender, text, links, the most likely confirm/reset/login link (actionLink) and candidate one-time codes. Each poll uses one Mailsac operation.
| Name | Required | Description | Default |
|---|---|---|---|
| address | Yes | Address to watch | |
| fromContains | No | Only match emails whose sender contains this text | |
| receivedAfter | No | ISO timestamp; ignore emails received before it (useful when an address is reused) | |
| timeoutSeconds | No | Give up after this long | |
| subjectContains | No | Only match emails whose subject contains this text | |
| pollIntervalSeconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses the polling cost ('Each poll uses one Mailsac operation') and the return payload, but never states what happens when the timeout elapses (error, empty result, partial match) or whether the call blocks. That failure-mode gap is material for a wait-style tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence front-loads the core action and then lists the returned artifacts without padding. Slightly overloaded by packing the full return list and the polling-cost note together, but every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description does cover the returned fields, which is the main need. However, for a 6-parameter, timeout-driven tool with zero annotations, it omits timeout/failure semantics and concurrency behavior, leaving a real gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 83%, so the schema already documents address, fromContains, receivedAfter, subjectContains and timeoutSeconds. The description adds no filter syntax or matching semantics beyond 'matching email', so it does not extend the schema meaningfully.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Poll an address until a matching email arrives') and enumerates what is returned, making it clearly distinguishable from siblings like list_emails and read_email. The purpose is unambiguous without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the polling behavior — an agent infers this is for awaiting an inbound message rather than reading one already present — but no sibling is named and there is no explicit when-not guidance. Adequate but leaves the choice against read_email/list_emails to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v0.1.0- First observed
create_test_address - First observed
delete_emails - First observed
list_domains - First observed
list_emails - First observed
read_email - First observed
wait_for_email
TDQS
Scored across 6 tools
Most tools have clearly distinct purposes (create address, list, read, delete, list domains). wait_for_email overlaps somewhat with list_emails/read_email since all retrieve mail, but the descriptions clarify that wait polls for arrival while the others inspect existing mail.
All six tools follow a clean snake_case verb_noun pattern (create_test_address, wait_for_email, list_emails, read_email, delete_emails, list_domains). Naming is fully predictable.
Six tightly scoped tools are ideal for a test-email service, each mapping to a distinct step in the inbox testing lifecycle. Nothing feels redundant or missing count-wise.
Covers the full receiving workflow: mint an address, await mail, list/read messages, delete for cleanup, and discover private domains. Outbound sending and message search aren't covered, but these are minor for a receive-focused testing tool.
Maintenance
Related MCP Connectors
Disposable inboxes for AI agents: create, wait for delivery, and extract email content or links.
Disposable email inboxes for AI agents — read messages and verification codes.
Real email inboxes for AI agents: create addresses, send, wait for mail and verification codes.
Real email inboxes for AI agents: create inboxes, catch verification codes, extract OTPs, reply.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to create disposable email inboxes and automatically extract OTPs, magic links, and verification codes from incoming emails.10 npmMIT
- AlicenseAqualityDmaintenanceEnables AI agents to create temporary email addresses, receive confirmation emails, and extract verification links, automating sign-up and email verification workflows without manual intervention.623 npm63MIT
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to generate temporary email addresses, receive emails, and automatically extract OTP codes and links from incoming messages for automation and testing workflows.MIT
- AlicenseAqualityAmaintenanceProvides disposable email inboxes for AI agents to automatically receive and extract OTPs and magic links, enabling seamless email verification during autonomous workflows.353 npmMIT