grafana-unified-mcp
grafana-unified-mcp
一个MCP服务器,位于多个Grafana实例之前。包含标准Grafana MCP服务器的所有工具,外加一个额外参数——instance——用于指定针对哪个Grafana运行。
query_prometheus(instance="tenant-a", expr="up", datasourceUid="...")
search_dashboards(instance="tenant-b", query="login latency")为什么存在
上游的 grafana/mcp-grafana 在进程启动时一次性绑定 GRAFANA_URL。它每次请求都会读取 X-Grafana-Service-Account-Token,但URL是固定的——而曾经用于覆盖它的头部现在已被明确弃用。来自上游 validate_url.go:
已弃用:X-Grafana-URL 不再配置 Grafana 客户端。此中间件暂时保留以处理格式错误的头部。
因此,一个 mcp-grafana 进程只能与一个 Grafana 通信。十个 Grafana 意味着十个服务器、十个客户端配置条目,以及十组名称相同的工具,模型需要从中区分。
此服务器通过为每个实例运行一个上游子进程,并根据 instance 参数将每个调用路由到正确的实例来解决此问题。工具在运行时从实际二进制文件中发现,因此您将获得上游暴露的所有工具——目前是65个工具——这里没有针对每个工具的代码,当上游添加更多工具时也无需更新。
Related MCP server: mcphub
工作原理
┌──────────────────────────────────┐
Claude Code / routines / │ grafana-unified-mcp │
cloud sessions │ │
│ │ ┌────────────────────────────┐ │
│ streamable-HTTP │ │ bearer auth │ │
│ Authorization: Bearer … │ │ → Principal(instances, │ │
├──────────────────────────────►│ │ read-only|read-write) │ │
│ │ └────────────┬───────────────┘ │
│ │ │ │
│ │ ┌────────────▼───────────────┐ │
│ │ │ catalog: inject `instance` │ │
│ │ │ filter by caller's grant │ │
│ │ └────────────┬───────────────┘ │
│ │ │ route on │
│ │ │ instance=… │
│ │ ┌────────────▼───────────────┐ │
│ │ │ child pool (lazy, reaped) │ │
│ │ └──┬──────────┬──────────┬───┘ │
└───────────────────────────────┴─────┼──────────┼──────────┼──────┘
│ stdio │ stdio │ stdio
┌─────▼────┐ ┌───▼──────┐ ┌▼─────────┐
│mcp-grafana│ │mcp-grafana│ │mcp-grafana│
│ tenant-a │ │ tenant-b │ │ … │
└─────┬────┘ └───┬──────┘ └┬─────────┘
▼ ▼ ▼
tenant-a tenant-b …Grafana子进程在首次使用时启动,保持热状态,空闲时被回收(--idle-timeout,默认15分钟),如果死亡则会透明地重新生成。无法访问的 Grafana 只会影响其自身的实例。
安装
两部分:上游二进制文件和本包。
# 1. the upstream mcp-grafana binary (needs Go 1.26+; GOTOOLCHAIN=auto fetches it)
deploy/install-mcp-grafana.sh /usr/local/bin
# 2. this server
python3 -m venv /opt/grafana-unified-mcp/.venv
/opt/grafana-unified-mcp/.venv/bin/pip install 'grafana-unified-mcp[aws] @ .'如果您已有二进制文件,请使用 MCP_GRAFANA_BINARY=/path/to/mcp-grafana 或 --mcp-grafana-binary 指向它。
配置
端点
正是您期望的形式——实例名称对应上游环境变量:
{
"tenant-a": {
"GRAFANA_URL": "https://tenant-a.example.cloud/grafana",
"GRAFANA_SERVICE_ACCOUNT_TOKEN": "glsa_…"
},
"tenant-b": {
"GRAFANA_URL": "https://tenant-b.example.cloud/grafana",
"GRAFANA_SERVICE_ACCOUNT_TOKEN": "glsa_…",
"description": "Tenant B production"
}
}每个实例的可选键:GRAFANA_ORG_ID、GRAFANA_USERNAME / GRAFANA_PASSWORD、description、extra_env、extra_args。为了将机密信息排除在文档本身之外,请使用 GRAFANA_SERVICE_ACCOUNT_TOKEN_ENV(从此进程的环境中读取)或 GRAFANA_SERVICE_ACCOUNT_TOKEN_FILE(子进程读取的路径)。
认证
{
"clients": [
{
"name": "claude-routines",
"token_sha256": "3f786850e387550fdab836ed7e6dc881de23001b…",
"instances": ["tenant-a", "tenant-b"],
"scope": "read-only"
},
{
"name": "platform-oncall",
"token_sha256": "…",
"instances": ["*"],
"scope": "read-write"
}
]
}生成一个令牌及其哈希:
grafana-unified-mcp --hash-token # generates one
grafana-unified-mcp --hash-token 'my-existing-token'将 token 提供给客户端;将 token_sha256 放入文档。令牌通过 hmac.compare_digest 进行摘要比较,每次尝试都会检查每个客户端,因此匹配位置不会通过时序泄露。
每个调用者强制实施两项内容:
instances—— 调用者看到的instance枚举被限制为其授权范围,调用命名实例之外的实例会被拒绝,并返回与不存在的实例相同的消息,因此令牌无法枚举其无法访问的内容。scope——read-only调用者甚至看不到变更工具。这种划分来自上游自身的readOnlyHint注释(目前65个工具中有49个是只读的),而不是在此维护的列表,因此上游添加的工具无需代码更改即可分类。任何未注释的工具都被视为非只读。
为了双重保险,添加 --child-arg=--disable-write 以在源头上为所有调用者剥离写入工具。
无认证运行
--auth-mode none 为所有能够访问端口的调用者提供服务,只读。没有身份来限定实例范围,因此所有配置的实例保持可读——但没有任何内容可写,因为开放的端口不应能够重写仪表板或删除快照。这在三个层面强制执行:
发布的目录省略了所有变更工具;
即使客户端直接命名一个工具,授权检查也会拒绝它;
子进程以
--disable-write启动,因此上游也会拒绝它们。
第三层使其不仅仅是一个过滤器。上游将 grafana_api_request 替换为一个单独的仅GET注册——没有 body 参数,method 限定为 GET,运行时拒绝非GET请求——因此即使第一层和第二层存在错误,也无法变成写入操作。
stdio 则不同:本地调用者已经持有端点文档及其中的所有令牌,因此限制它们将是徒劳的。stdio 获得完全访问权限。
如果您需要通过HTTP进行写入操作,请使用带有 read-write 客户端的承载令牌,而不是开放端口。
配置来源
以下任意来源,适用于 --endpoints 和 --auth:
来源 | 示例 |
文件 |
|
内联环境变量 |
|
AWS Secrets Manager |
|
AWS SSM Parameter Store |
|
两个文档每 --config-refresh-seconds(默认300秒)重新读取一次。失败的刷新会记录日志并保留最后一个有效值,因此瞬时的AWS错误或写入一半的文件不会导致服务器宕机。添加实例无需重启;移除实例会停止其子进程。
启动前验证:
grafana-unified-mcp --endpoints … --auth … --check-config运行
# local, over stdio (no auth — the local caller already holds the config)
grafana-unified-mcp --endpoints ./examples/endpoints.json
# deployed, over streamable-HTTP behind a reverse proxy
grafana-unified-mcp \
--transport streamable-http \
--address 127.0.0.1:8900 \
--endpoints aws-secrets:prod/grafana/endpoints?region=us-west-2 \
--auth aws-secrets:prod/grafana/mcp-auth?region=us-west-2 \
--public-url https://grafana-mcp.example.com
--public-url很重要。 SDK 基于Host头部应用 DNS 重新绑定保护。在代理转发公共主机名的情况下,必须允许该主机,否则每个请求都会被拒绝。--public-url允许它(并用于 RFC 9728 资源元数据);--allowed-host添加更多。
GET /healthz 报告进程健康状态、活跃子进程和目录状态,无需接触 Grafana。
连接客户端
.mcp.json,用于本地 stdio 使用:
{
"mcpServers": {
"grafana": {
"command": "/opt/grafana-unified-mcp/.venv/bin/grafana-unified-mcp",
"args": ["--endpoints", "/etc/grafana-unified-mcp/endpoints.json"]
}
}
}对于部署的服务器——包括 Claude Code 例程和云会话,这正是承载令牌存在的原因:
{
"mcpServers": {
"grafana": {
"type": "http",
"url": "https://grafana-mcp.example.com/mcp",
"headers": {
"Authorization": "Bearer ${GRAFANA_UNIFIED_MCP_TOKEN}"
}
}
}
}在会话运行的环境中设置 GRAFANA_UNIFIED_MCP_TOKEN——对于 Web 上的 Claude Code,这是环境变量,因此计划例程和云会话可以获取它,而无需将机密信息存储在仓库中。为例程提供 read-only 客户端;为人类保留 read-write。
部署为 systemd 服务
参见 deploy/。简而言之:
sudo deploy/install.sh # user, dirs, venv, unit file
sudo systemctl edit grafana-unified-mcp # set the source URIs / region
sudo systemctl enable --now grafana-unified-mcp
curl -s localhost:8900/healthz | jq该单元以专用的非特权用户身份运行,并启用 ProtectSystem=strict、PrivateTmp 和 NoNewPrivileges。TLS 终止于前面的 nginx 或 ALB——参见 deploy/nginx.conf.example,它禁用了响应缓冲(SSE 流式传输所需)。
使用
首先让模型指向 list_grafana_instances:
list_grafana_instances()
→ { "instances": [ {"name": "tenant-a", "url": "…", "connection": "live"}, … ],
"routing_argument": "instance",
"access": { "client": "claude-routines", "scope": "read-only" } }然后每个其他工具都使用该名称:
search_dashboards(instance="tenant-a", query="latency")传递 check_health=true 还可以探测每个 Grafana——较慢,因为它会打开到每个实例的连接。
一个命名问题
上游的 grafana_api_request 已经有一个名为 endpoint 的必需参数(API 路径)。通过该名称注入路由参数会静默地遮蔽它,这就是为什么路由参数默认是 instance。如果您使用 --routing-param endpoint 重命名它,该工具自身的参数会自动重新发布为 api_path,并在传递过程中映射回来——无论您选择什么,都不会有工具因冲突而损坏。
开发
uv venv && uv pip install -e '.[dev,aws]'
uv run pytest # unit + integration集成测试驱动一个真实的 mcp-grafana 子进程,针对一个无法访问的 Grafana:足以证明目录发现、instance 注入和剥离、路由以及认证过滤,无需实时凭据。设置 MCP_GRAFANA_BINARY 指向二进制文件,否则它们会跳过。
路线图
OAuth 2.1 —— 认证层已经是一个接口,SDK 已经接受一个 OAuth 提供程序以及令牌验证器。填充
OAuth2Provider.verify_token就是全部工作;auth/oauth.py记录了三个步骤。将 IdP 组映射到现有的grafana:read/grafana:write/instance:<name>作用域,所有授权检查保持不变。扇出 ——
instance: "*"用于在所有实例上运行一个只读查询并合并结果。对于“哪个正在告警?”很有用;目前暂未实现,因为结果合并需要自己的设计。
This server cannot be deployed
Maintenance
Related MCP Connectors
An MCP server giving access to Grafana dashboards, data and more.
Stateless MCP gateway and OTel span-streaming bridge for hosted MCP servers.
Governed MCP gateway: one endpoint for your tools, with credential custody and audit log.
MCP-first control plane for ProAgentStore agents and private instances.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceProxy that aggregates multiple MCP servers and presents them as a unified interface, allowing clients to access resources from multiple servers transparently.4-
- AlicenseNot gradedqualityAmaintenanceA unified hub for centrally managing and dynamically orchestrating multiple MCP servers/APIs into separate endpoints with flexible routing strategies.846 npm2,488Apache 2.0
- AlicenseNot gradedqualityDmaintenanceEnables MCP-compatible agents to interact with Grafana instances for searching, creating, and updating dashboards, exploring logs via Loki, querying datasources, managing alerts, incidents, and on-call shifts, and accessing observability data.8Apache 2.0
- FlicenseNot gradedqualityDmaintenanceEnables interaction with multiple Jenkins instances from a single MCP server, using header-based authentication for multi-tenancy.1-