MCP Headless Gmail Server
Provides tools for reading recent emails, retrieving full email content in chunks, and sending emails through Gmail using OAuth 2.0 authentication with automatic token refresh handling.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@MCP Headless Gmail Serversend an email to john@example.com about our meeting tomorrow"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
MCP Headless Gmail Server (NPM & Docker)
A MCP (Model Context Protocol) server that provides get, send Gmails without local credential or token setup.
Why MCP Headless Gmail Server?
Critical Advantages
Headless & Remote Operation: Unlike other MCP Gmail solutions that require running outside of docker and local file access, this server can run completely headless in remote environments with no browser no local file access.
Decoupled Architecture: Any client can complete the OAuth flow independently, then pass credentials as context to this MCP server, creating a complete separation between credential storage and server implementation.
Nice but not critical
Focused Functionality: In many use cases, especially for marketing applications, only Gmail access is needed without additional Google services like Calendar, making this focused implementation ideal.
Docker-Ready: Designed with containerization in mind for a well-isolated, environment-independent, one-click setup.
Reliable Dependencies: Built on the well-maintained google-api-python-client library.
Related MCP server: Gmail MCP
Features
Get most recent emails from Gmail with the first 1k characters of the body
Get full email body content in 1k chunks using offset parameter
Send emails through Gmail
Refresh access tokens separately
Automatic refresh token handling
Prerequisites
Python 3.10 or higher
Google API credentials (client ID, client secret, access token, and refresh token)
Installation
Installing via Smithery
To install mcp-headless-gmail for Claude Desktop automatically via Smithery:
npx -y @smithery/cli install @AbhinavBansal17/mcp-headless-gmail --client claudeManual Installation
# Clone the repository
git clone https://github.com/baryhuang/mcp-headless-gmail.git
cd mcp-headless-gmail
# Install dependencies
pip install -e .Docker
Building the Docker Image
# Build the Docker image
docker build -t mcp-headless-gmail .Usage with Claude Desktop
You can configure Claude Desktop to use the Docker image by adding the following to your Claude configuration:
docker
{
"mcpServers": {
"gmail": {
"command": "docker",
"args": [
"run",
"-i",
"--rm",
"buryhuang/mcp-headless-gmail:latest"
]
}
}
}npm version
{
"mcpServers": {
"gmail": {
"command": "npx",
"args": [
"@peakmojo/mcp-server-headless-gmail"
]
}
}
}Note: With this configuration, you'll need to provide your Google API credentials in the tool calls as shown in the Using the Tools section. Gmail credentials are not passed as environment variables to maintain separation between credential storage and server implementation.
Cross-Platform Publishing
To publish the Docker image for multiple platforms, you can use the docker buildx command. Follow these steps:
Create a new builder instance (if you haven't already):
docker buildx create --useBuild and push the image for multiple platforms:
docker buildx build --platform linux/amd64,linux/arm64,linux/arm/v7 -t buryhuang/mcp-headless-gmail:latest --push .Verify the image is available for the specified platforms:
docker buildx imagetools inspect buryhuang/mcp-headless-gmail:latest
Usage
The server provides Gmail functionality through MCP tools. Authentication handling is simplified with a dedicated token refresh tool.
Starting the Server
mcp-server-headless-gmailUsing the Tools
When using an MCP client like Claude, you have two main ways to handle authentication:
Refreshing Tokens (First Step or When Tokens Expire)
If you have both access and refresh tokens:
{
"google_access_token": "your_access_token",
"google_refresh_token": "your_refresh_token",
"google_client_id": "your_client_id",
"google_client_secret": "your_client_secret"
}If your access token has expired, you can refresh with just the refresh token:
{
"google_refresh_token": "your_refresh_token",
"google_client_id": "your_client_id",
"google_client_secret": "your_client_secret"
}This will return a new access token and its expiration time, which you can use for subsequent calls.
Getting Recent Emails
Retrieves recent emails with the first 1k characters of each email body:
{
"google_access_token": "your_access_token",
"max_results": 5,
"unread_only": false
}Response includes:
Email metadata (id, threadId, from, to, subject, date, etc.)
First 1000 characters of the email body
body_size_bytes: Total size of the email body in bytescontains_full_body: Boolean indicating if the entire body is included (true) or truncated (false)
Getting Full Email Body Content
For emails with bodies larger than 1k characters, you can retrieve the full content in chunks:
{
"google_access_token": "your_access_token",
"message_id": "message_id_from_get_recent_emails",
"offset": 0
}You can also get email content by thread ID:
{
"google_access_token": "your_access_token",
"thread_id": "thread_id_from_get_recent_emails",
"offset": 1000
}The response includes:
A 1k chunk of the email body starting from the specified offset
body_size_bytes: Total size of the email bodychunk_size: Size of the returned chunkcontains_full_body: Boolean indicating if the chunk contains the remainder of the body
To retrieve the entire email body of a long message, make sequential calls increasing the offset by 1000 each time until contains_full_body is true.
Sending an Email
{
"google_access_token": "your_access_token",
"to": "recipient@example.com",
"subject": "Hello from MCP Gmail",
"body": "This is a test email sent via MCP Gmail server",
"html_body": "<p>This is a <strong>test email</strong> sent via MCP Gmail server</p>"
}Token Refresh Workflow
Start by calling the
gmail_refresh_tokentool with either:Your full credentials (access token, refresh token, client ID, and client secret), or
Just your refresh token, client ID, and client secret if the access token has expired
Use the returned new access token for subsequent API calls.
If you get a response indicating token expiration, call the
gmail_refresh_tokentool again to get a new token.
This approach simplifies most API calls by not requiring client credentials for every operation, while still enabling token refresh when needed.
Obtaining Google API Credentials
To obtain the required Google API credentials, follow these steps:
Go to the Google Cloud Console
Create a new project
Enable the Gmail API
Configure OAuth consent screen
Create OAuth client ID credentials (select "Desktop app" as the application type)
Save the client ID and client secret
Use OAuth 2.0 to obtain access and refresh tokens with the following scopes:
https://www.googleapis.com/auth/gmail.readonly(for reading emails)https://www.googleapis.com/auth/gmail.send(for sending emails)
Token Refreshing
This server implements automatic token refreshing. When your access token expires, the Google API client will use the refresh token, client ID, and client secret to obtain a new access token without requiring user intervention.
Security Note
This server requires direct access to your Google API credentials. Always keep your tokens and credentials secure and never share them with untrusted parties.
License
See the LICENSE file for details.
Available Tools
3 toolsgmail_get_email_body_chunkB
Get a 1k character chunk of an email body starting from the specified offset
| Name | Required | Description | Default |
|---|---|---|---|
| google_access_token | No | Google OAuth2 access token | |
| message_id | No | ID of the message to retrieve | |
| thread_id | No | ID of the thread to retrieve (will get the first message if multiple exist) | |
| offset | No | Offset in characters to start from (default: 0) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the chunk size (1k characters) and offset behavior, but lacks details on error handling (e.g., invalid offsets), authentication requirements (implied by google_access_token but not explained), rate limits, or what happens if the email body is shorter than the offset. This leaves significant gaps for a tool that interacts with external data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core functionality. It wastes no words and directly communicates the tool's purpose without unnecessary elaboration, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of handling email data with multiple parameters and no output schema, the description is insufficient. It does not explain the return format (e.g., plain text, HTML), error cases, or how partial chunks are handled. With no annotations and incomplete behavioral details, it fails to provide enough context for reliable use by an AI agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters (google_access_token, message_id, thread_id, offset). The description adds minimal value by mentioning the offset parameter and default behavior, but does not provide additional context beyond what the schema offers, such as how message_id and thread_id interact or format requirements.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Get'), resource ('a 1k character chunk of an email body'), and scope ('starting from the specified offset'), distinguishing it from sibling tools like gmail_get_recent_emails (which lists emails) and gmail_send_email (which sends emails). It precisely defines what the tool does without being vague or tautological.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It does not mention prerequisites (e.g., needing a message_id or thread_id), nor does it explain scenarios where this tool is appropriate (e.g., for large email bodies) versus using other tools. Usage is implied but not explicitly stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gmail_get_recent_emailsA
Get the most recent emails from Gmail (returns metadata, snippets, and first 1k chars of body)
| Name | Required | Description | Default |
|---|---|---|---|
| google_access_token | No | Google OAuth2 access token | |
| max_results | No | Maximum number of emails to return (default: 10) | |
| unread_only | No | Whether to return only unread emails (default: False) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses the return format (metadata, snippets, and first 1k chars of body) and the scope ('most recent'), which are useful behavioral traits. However, it doesn't mention rate limits, pagination, error handling, or whether this is a read-only operation (though implied by 'Get'). The description doesn't contradict any annotations since none are provided.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core purpose and includes key details about the return data. Every part earns its place, with no redundant or vague language. It's appropriately sized for a tool with good schema coverage and no complex behavioral nuances to explain.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (3 parameters, no output schema, no annotations), the description is adequate but has gaps. It covers the purpose and return format, but lacks details on authentication requirements (beyond the parameter), error cases, or how 'most recent' is determined (e.g., sorting by date). With no output schema, it should ideally describe the response structure more fully, but the mention of metadata/snippets/body chars provides some context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters (google_access_token, max_results, unread_only) with their types and defaults. The description doesn't add any parameter-specific semantics beyond what's in the schema, such as explaining OAuth2 scopes or how 'most recent' interacts with max_results/unread_only. Baseline 3 is appropriate when schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Get the most recent emails') and resource ('from Gmail'), specifying what data is returned (metadata, snippets, and first 1k chars of body). It distinguishes from sibling tools like gmail_get_email_body_chunk (which gets specific body chunks) and gmail_send_email (which sends emails). However, it doesn't explicitly mention the sibling differentiation in the description text itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for retrieving recent emails with metadata and partial body content, but doesn't provide explicit guidance on when to use this tool versus alternatives. No when-not-to-use scenarios or prerequisites (like authentication needs) are mentioned, though the need for a google_access_token is clear from the schema.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gmail_send_emailC
Send an email via Gmail with optional file attachments
| Name | Required | Description | Default |
|---|---|---|---|
| google_access_token | No | Google OAuth2 access token | |
| to | Yes | Recipient email address | |
| subject | Yes | Email subject | |
| body | Yes | Email body content (plain text) | |
| html_body | No | Email body content in HTML format (optional) | |
| attachments | No | Optional list of file attachments |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It states this is a send operation (implying mutation/write) but doesn't mention authentication requirements (though the schema shows google_access_token parameter), rate limits, error conditions, what happens on success/failure, or whether emails are sent immediately or queued. The description adds minimal behavioral context beyond what's implied by 'send'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that states the core purpose and key feature (attachments) with zero wasted words. It's appropriately sized and front-loaded with the essential information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description is insufficient. It doesn't explain what happens after sending (success/failure indicators), authentication requirements, rate limits, or error handling. The 100% schema coverage helps with parameters, but behavioral aspects are largely undocumented given this is a write operation with potential side effects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 6 parameters thoroughly. The description adds value by mentioning 'optional file attachments' which helps contextualize the attachments parameter, but doesn't provide additional semantic context beyond what's in the schema descriptions. Baseline 3 is appropriate when schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('send an email') and resource ('via Gmail') with a specific feature mention ('optional file attachments'). It distinguishes from sibling tools (gmail_get_email_body_chunk, gmail_get_recent_emails) by being a write operation rather than a read operation, though it doesn't explicitly name the siblings for comparison.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives or when not to use it. It mentions 'optional file attachments' but doesn't explain when to use attachments versus inline content, nor does it reference the sibling tools for different email-related tasks.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
- First observed
gmail_get_email_body_chunk - First observed
gmail_get_recent_emails - First observed
gmail_send_email
TDQS
Scored across 3 tools
Each tool has a clearly distinct purpose: gmail_get_email_body_chunk retrieves specific content from an email, gmail_get_recent_emails fetches metadata and previews of multiple emails, and gmail_send_email handles outgoing messages. There is no overlap in functionality that could cause confusion.
All tools follow a consistent 'gmail_verb_noun' pattern with snake_case, making them predictable and easy to understand. The naming convention is uniform across all three tools.
With only 3 tools, the server feels thin for a Gmail integration, lacking essential operations like searching emails, managing labels, or deleting messages. While the tools cover basic read and send functions, the scope is limited compared to typical email management needs.
The toolset has significant gaps for a Gmail server: there is no way to search or filter emails beyond recent ones, no ability to update or delete emails, no label management, and no full email retrieval beyond chunks. This incomplete surface will likely cause agent failures in common email workflows.
Maintenance
Related MCP Connectors
Manage Gmail end-to-end: search, read, send, draft, label, and organize threads. Automate workflow…
Permissioned access to Gmail, Drive and Calendar via the user's own Google account
Never-stored live email: read, send, organize, schedule and auto-triage Gmail or any IMAP mailbox.
Manage Gmail messages, threads, labels, drafts, and settings from your workflows. Send and organiz…
Related MCP Servers
- AlicenseBqualityFmaintenanceA headless server that enables reading and sending Gmail emails through API calls without requiring local credentials or browser access, designed to run remotely in containerized environments.41155MIT
- AlicenseCqualityDmaintenanceEnables comprehensive Gmail management through the Gmail API, including sending/receiving emails, organizing labels and threads, managing drafts, and configuring account settings with secure OAuth2 authentication.646271MIT
- AlicenseNot gradedqualityBmaintenanceEnables reading, sending, archiving, and managing Gmail emails and labels through Google OAuth authentication, acting as an OAuth proxy to the Gmail API.8811MIT
- FlicenseNot gradedqualityNot gradedmaintenanceEnables interaction with Gmail through OAuth 2.0 authentication, allowing users to fetch unread emails and create draft replies that are properly threaded with original conversations.-