MCP Headless Gmail Server
MCP ヘッドレス Gmail サーバー
ローカル認証情報やトークンの設定なしで Gmail の取得、送信ができる MCP (Model Context Protocol) サーバー。
MCP ヘッドレス Gmail サーバーを選ぶ理由
重要な利点
ヘッドレス & リモート操作: Docker 外での実行とローカル ファイル アクセスを必要とする他の MCP Gmail ソリューションとは異なり、このサーバーはブラウザーもローカル ファイル アクセスもないリモート環境で完全にヘッドレスで実行できます。
分離されたアーキテクチャ: どのクライアントも OAuth フローを独立して完了し、資格情報をコンテキストとしてこの MCP サーバーに渡すことができるため、資格情報の保存とサーバーの実装が完全に分離されます。
いいけど批判的ではない
重点的な機能: 多くのユースケース、特にマーケティング アプリケーションでは、カレンダーなどの追加の Google サービスなしで Gmail へのアクセスのみが必要なため、この重点的な実装が理想的です。
Docker 対応: コンテナ化を考慮して設計されており、適切に分離された、環境に依存しない、ワンクリック セットアップが可能です。
信頼できる依存関係: 適切に管理された google-api-python-client ライブラリに基づいて構築されています。
Related MCP server: MCP Headless Gmail Server
特徴
Gmail から最新のメールを本文の最初の 1,000 文字で取得します
オフセットパラメータを使用して、1k チャンクでメール本文の全内容を取得します。
Gmail経由でメールを送信する
アクセストークンを個別に更新する
自動リフレッシュトークン処理
前提条件
Python 3.10以上
Google API 認証情報(クライアント ID、クライアント シークレット、アクセス トークン、リフレッシュ トークン)
インストール
# Clone the repository
git clone https://github.com/baryhuang/mcp-headless-gmail.git
cd mcp-headless-gmail
# Install dependencies
pip install -e .ドッカー
Dockerイメージの構築
# Build the Docker image
docker build -t mcp-headless-gmail .Claude Desktopでの使用
Claude 構成に以下を追加することで、Claude Desktop が Docker イメージを使用するように構成できます。
ドッカー
{
"mcpServers": {
"gmail": {
"command": "docker",
"args": [
"run",
"-i",
"--rm",
"buryhuang/mcp-headless-gmail:latest"
]
}
}
}npmバージョン
{
"mcpServers": {
"gmail": {
"command": "npx",
"args": [
"@peakmojo/mcp-server-headless-gmail"
]
}
}
}注: この設定では、 「ツールの使用」セクションに示されているように、ツール呼び出し時にGoogle API認証情報を提供する必要があります。認証情報の保存とサーバー実装を分離するため、Gmailの認証情報は環境変数として渡されません。
クロスプラットフォームパブリッシング
複数のプラットフォーム向けにDockerイメージを公開するには、 docker buildxコマンドを使用します。以下の手順に従ってください。
新しいビルダー インスタンスを作成します(まだ作成していない場合)。
docker buildx create --use複数のプラットフォーム用のイメージをビルドしてプッシュします。
docker buildx build --platform linux/amd64,linux/arm64,linux/arm/v7 -t buryhuang/mcp-headless-gmail:latest --push .指定されたプラットフォームでイメージが使用可能であることを確認します。
docker buildx imagetools inspect buryhuang/mcp-headless-gmail:latest
使用法
サーバーはMCPツールを通じてGmail機能を提供します。専用のトークン更新ツールにより、認証処理が簡素化されます。
サーバーの起動
mcp-server-headless-gmailツールの使用
Claude のような MCP クライアントを使用する場合、認証を処理する主な方法は 2 つあります。
トークンの更新(最初のステップまたはトークンの有効期限が切れたとき)
アクセス トークンとリフレッシュ トークンの両方がある場合:
{
"google_access_token": "your_access_token",
"google_refresh_token": "your_refresh_token",
"google_client_id": "your_client_id",
"google_client_secret": "your_client_secret"
}アクセス トークンの有効期限が切れている場合は、リフレッシュ トークンだけで更新できます。
{
"google_refresh_token": "your_refresh_token",
"google_client_id": "your_client_id",
"google_client_secret": "your_client_secret"
}これにより、新しいアクセス トークンとその有効期限が返され、後続の呼び出しで使用できるようになります。
最近のメールを取得する
各メール本文の最初の 1,000 文字を含む最近のメールを取得します。
{
"google_access_token": "your_access_token",
"max_results": 5,
"unread_only": false
}回答には以下が含まれます:
電子メールのメタデータ (ID、スレッド ID、送信元、送信先、件名、日付など)
メール本文の最初の1000文字
body_size_bytes: メール本文の合計サイズ(バイト単位)contains_full_body: 本文全体が含まれているか(true)、切り捨てられているか(false)を示すブール値
メール本文の全文を取得する
本文が 1,000 文字を超えるメールの場合は、完全なコンテンツをチャンク単位で取得できます。
{
"google_access_token": "your_access_token",
"message_id": "message_id_from_get_recent_emails",
"offset": 0
}スレッド ID でメールの内容を取得することもできます。
{
"google_access_token": "your_access_token",
"thread_id": "thread_id_from_get_recent_emails",
"offset": 1000
}応答には次のものが含まれます。
指定されたオフセットから始まるメール本文の1kチャンク
body_size_bytes: メール本文の合計サイズchunk_size: 返されるチャンクのサイズcontains_full_body: チャンクに本体の残りが含まれているかどうかを示すブール値
長いメッセージのメール本文全体を取得するには、 contains_full_body true になるまで、オフセットを 1000 ずつ増やしながら連続呼び出しを実行します。
メールを送信する
{
"google_access_token": "your_access_token",
"to": "recipient@example.com",
"subject": "Hello from MCP Gmail",
"body": "This is a test email sent via MCP Gmail server",
"html_body": "<p>This is a <strong>test email</strong> sent via MCP Gmail server</p>"
}トークン更新ワークフロー
まず、次のいずれかの方法で
gmail_refresh_tokenツールを呼び出します。完全な認証情報(アクセストークン、リフレッシュトークン、クライアントID、クライアントシークレット)、または
アクセストークンの有効期限が切れている場合は、リフレッシュトークン、クライアントID、クライアントシークレットのみ
返された新しいアクセス トークンを後続の API 呼び出しに使用します。
トークンの有効期限が切れたことを示す応答を受け取った場合は、
gmail_refresh_tokenツールを再度呼び出して新しいトークンを取得してください。
このアプローチでは、すべての操作でクライアント資格情報を要求せず、必要なときにトークンの更新も可能にすることで、ほとんどの API 呼び出しが簡素化されます。
Google API認証情報の取得
必要な Google API 認証情報を取得するには、次の手順に従います。
Google Cloud Consoleにアクセスします
新しいプロジェクトを作成する
Gmail APIを有効にする
OAuth同意画面を設定する
OAuth クライアント ID 資格情報を作成します (アプリケーションの種類として「デスクトップ アプリ」を選択します)
クライアントIDとクライアントシークレットを保存する
OAuth 2.0 を使用して、次のスコープのアクセス トークンと更新トークンを取得します。
https://www.googleapis.com/auth/gmail.readonly(メールを読むため)https://www.googleapis.com/auth/gmail.send(メール送信用)
トークンの更新
このサーバーはトークンの自動更新を実装しています。アクセストークンの有効期限が切れると、Google APIクライアントはリフレッシュトークン、クライアントID、クライアントシークレットを使用して、ユーザーの介入なしに新しいアクセストークンを取得します。
セキュリティに関する注意事項
このサーバーはGoogle API認証情報に直接アクセスする必要があります。トークンと認証情報は常に安全に保管し、信頼できない相手と共有しないでください。
ライセンス
詳細については、LICENSE ファイルを参照してください。
Available Tools
4 toolsgmail_get_email_body_chunkB
Get a 1k character chunk of an email body starting from the specified offset
| Name | Required | Description | Default |
|---|---|---|---|
| google_access_token | Yes | Google OAuth2 access token | |
| message_id | No | ID of the message to retrieve | |
| thread_id | No | ID of the thread to retrieve (will get the first message if multiple exist) | |
| offset | No | Offset in characters to start from (default: 0) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It mentions the 1k character chunking behavior, which is valuable, but doesn't address authentication needs (though implied by google_access_token parameter), error handling, rate limits, or what happens with invalid offsets/message_ids. For a tool with no annotation coverage, this leaves significant gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that immediately conveys the core functionality without any wasted words. It's appropriately sized and front-loaded with the essential information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (4 parameters, no output schema, no annotations), the description is minimally adequate. It explains the chunking behavior but lacks details about authentication requirements, error conditions, and how this tool relates to siblings. Without annotations or output schema, more behavioral context would be helpful for safe usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so parameters are well-documented in the schema. The description adds context about the 'offset' parameter (default: 0) and clarifies that thread_id retrieves the first message if multiple exist, providing some value beyond the schema. However, it doesn't explain parameter interactions or provide additional semantic context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Get'), resource ('email body chunk'), and key constraint ('1k character chunk starting from specified offset'), making the purpose immediately understandable. However, it doesn't explicitly differentiate from sibling tools like gmail_get_recent_emails, which retrieves multiple emails rather than a specific body chunk.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided about when to use this tool versus alternatives. The description doesn't mention prerequisites (like needing a message_id or thread_id), nor does it explain when this tool is appropriate compared to gmail_get_recent_emails for retrieving email content.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gmail_get_recent_emailsC
Get the most recent emails from Gmail (returns metadata, snippets, and first 1k chars of body)
| Name | Required | Description | Default |
|---|---|---|---|
| google_access_token | Yes | Google OAuth2 access token | |
| max_results | No | Maximum number of emails to return (default: 10) | |
| unread_only | No | Whether to return only unread emails (default: False) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It mentions what data is returned (metadata, snippets, first 1k chars of body) which is helpful, but doesn't cover important behavioral aspects like authentication requirements (beyond the parameter), rate limits, pagination behavior, error conditions, or whether this is a read-only operation. For a tool with no annotations, this leaves significant gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that communicates the core functionality and return format. It's appropriately sized for a straightforward retrieval tool, though it could potentially benefit from slightly more detail given the lack of annotations and output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of email retrieval (3 parameters, no annotations, no output schema), the description is insufficiently complete. It doesn't explain the return format in detail, doesn't mention authentication requirements beyond the parameter, and doesn't cover important behavioral aspects. For a tool with no annotations or output schema, the description should provide more context about what to expect from the operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already fully documents all three parameters. The description doesn't add any parameter-specific information beyond what's in the schema. The baseline of 3 is appropriate when the schema does all the parameter documentation work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Get the most recent emails from Gmail' specifies the verb (get) and resource (emails). It distinguishes from sibling 'gmail_get_email_body_chunk' by indicating it returns metadata, snippets, and partial body content, but doesn't explicitly differentiate from other siblings like 'gmail_send_email' or 'gmail_refresh_token' beyond the obvious functional difference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives is provided. The description doesn't mention when to use this versus 'gmail_get_email_body_chunk' for full body retrieval, or when to use 'gmail_refresh_token' for token management. Usage context is implied by the tool name but not explicitly stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gmail_refresh_tokenB
Refresh the access token using the refresh token and client credentials
| Name | Required | Description | Default |
|---|---|---|---|
| google_access_token | No | Google OAuth2 access token (optional if expired) | |
| google_refresh_token | Yes | Google OAuth2 refresh token | |
| google_client_id | Yes | Google OAuth2 client ID for token refresh | |
| google_client_secret | Yes | Google OAuth2 client secret for token refresh |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden but only states what the tool does at a high level. It doesn't disclose behavioral traits like whether this invalidates previous tokens, rate limits, error conditions, or what the refreshed token enables. For a security-sensitive operation with zero annotation coverage, this is inadequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose with zero waste. It's appropriately sized and front-loaded, with every word contributing to understanding the core functionality.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a security-critical token refresh operation with no annotations and no output schema, the description is insufficient. It doesn't explain what happens after refresh (e.g., token lifetime, scope preservation), error handling, or integration with sibling tools. Given the complexity and lack of structured data, more context is needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 4 parameters thoroughly. The description adds no parameter-specific semantics beyond what's in the schema (e.g., it doesn't explain relationships between parameters or provide usage examples). Baseline 3 is appropriate when schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Refresh') and resource ('access token'), specifying it uses refresh token and client credentials. It distinguishes from sibling tools (email-related operations) by focusing on authentication token management, though it doesn't explicitly name alternatives for token refresh.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when access tokens expire (via 'optional if expired' in schema), but doesn't explicitly state when to use this tool versus alternatives like initial authentication or other token management methods. No guidance on prerequisites or exclusions is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gmail_send_emailC
Send an email via Gmail
| Name | Required | Description | Default |
|---|---|---|---|
| google_access_token | Yes | Google OAuth2 access token | |
| to | Yes | Recipient email address | |
| subject | Yes | Email subject | |
| body | Yes | Email body content (plain text) | |
| html_body | No | Email body content in HTML format (optional) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states the action ('send an email') but lacks critical details: it doesn't mention authentication requirements (implied by the 'google_access_token' parameter but not explicitly stated), potential rate limits, error handling, or what happens upon success (e.g., whether it returns a confirmation). This leaves significant gaps for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with zero wasted words—'Send an email via Gmail' is front-loaded and directly conveys the core action. It's appropriately sized for the tool's complexity, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (a mutation tool with 5 parameters, no annotations, and no output schema), the description is incomplete. It fails to address key contextual aspects like authentication needs, behavioral traits (e.g., what 'send' entails operationally), or output expectations, leaving the agent with insufficient information for reliable use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, with all parameters well-documented in the schema (e.g., 'to' as recipient email, 'body' as plain text content). The description adds no additional meaning beyond the schema, such as explaining parameter interactions (e.g., 'body' vs. 'html_body') or constraints. This meets the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Send an email via Gmail' clearly states the verb ('send') and resource ('email via Gmail'), making the purpose immediately understandable. However, it doesn't differentiate from sibling tools like 'gmail_get_recent_emails' or 'gmail_refresh_token' beyond the obvious action distinction, which keeps it from a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., needing authentication via 'google_access_token'), nor does it clarify scenarios where other tools like 'gmail_get_recent_emails' might be more appropriate, leaving usage context entirely implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
- First observed
gmail_get_email_body_chunk - First observed
gmail_get_recent_emails - First observed
gmail_refresh_token - First observed
gmail_send_email
TDQS
Scored across 4 tools
Each tool has a clearly distinct purpose: retrieving email body chunks, listing recent emails, refreshing tokens, and sending emails. There is no overlap in functionality that would cause confusion or misselection.
All tools follow a consistent 'gmail_verb_noun' pattern with snake_case, making them predictable and easy to understand. The naming convention is uniform across all four tools.
With 4 tools, the count is reasonable for a Gmail server, though it feels slightly thin for covering all common email operations. It includes core functions but could benefit from additional tools like searching or managing drafts.
The tools cover basic email operations (read, list, send) and authentication, but there are notable gaps such as searching emails, managing labels, or handling attachments. This could limit agents in performing more complex email tasks.
Maintenance
Related MCP Connectors
Manage Gmail end-to-end: search, read, send, draft, label, and organize threads. Automate workflow…
Never-stored live email: read, send, organize, schedule and auto-triage Gmail or any IMAP mailbox.
Permissioned access to Gmail, Drive and Calendar via the user's own Google account
A MCP server for Gmail that lets you search, read, and draft emails and replies.
Related MCP Servers
- AlicenseNot gradedqualityNot gradedmaintenanceA Model Context Protocol server that enables applications to interact with Gmail through a clean API, supporting email searching, sending, reading, and label management.MIT
- AlicenseBqualityDmaintenanceEnables reading and sending Gmail messages in headless/remote environments without local credential storage, using OAuth tokens passed directly in API calls for containerized deployments.362 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables reading, sending, archiving, and managing Gmail emails and labels through Google OAuth authentication, acting as an OAuth proxy to the Gmail API.150 npm11MIT
- AlicenseNot gradedqualityDmaintenanceA Model Context Protocol server that exposes the Gmail API for integration with LLMs, enabling email management tasks such as reading, labeling, and searching emails.7MIT