llmprobe
![]()
llmprobe
探测 LLM API 端点。测量 TTFT、延迟和吞吐量。单一二进制文件,零 SDK 依赖。
llmprobe 是一个 CLI 工具,用于探测 LLM API 端点并测量对生产可靠性至关重要的指标:首字延迟 (TTFT)、总延迟、生成吞吐量 (tokens/sec) 以及错误率。
你可以将其用作一次性健康检查、持续监控工具,或是在 LLM 提供商性能下降时阻止部署的 CI 门禁。

快速开始
从 最新发布版本 下载预构建的二进制文件 (支持 Linux、macOS、Windows;amd64 和 arm64 架构)。
或者从源码安装:
go install github.com/Jwrede/llmprobe@latest创建一个 probes.yml 文件(或复制提供的示例):
providers:
- name: openai
api_key: ${OPENAI_API_KEY}
models:
- name: gpt-4o
thresholds:
max_ttft: 2s
- name: gpt-4o-mini
thresholds:
max_ttft: 500ms
- name: anthropic
api_key: ${ANTHROPIC_API_KEY}
models:
- name: claude-sonnet-4-20250514
thresholds:
max_ttft: 1s运行探测:
$ llmprobe probe
Provider Model Status TTFT Latency Tok/s Tokens Error
-------- ----- ------ ---- ------- ----- ------ -----
openai gpt-4o healthy 312ms 2100ms 68.4 42
openai gpt-4o-mini healthy 98ms 814ms 112.3 56
anthropic claude-sonnet-4-20250514 healthy 420ms 2831ms 52.1 38
azure gpt-4o healthy 289ms 1950ms 71.2 44
bedrock anthropic.claude-3-5... degraded 1820ms 4510ms 28.1 38
4 healthy, 1 degraded, 0 errorsRelated MCP server: LLM API Benchmark MCP Server
测量指标
指标 | 含义 |
TTFT | 从发送请求到收到第一个内容 token 的时间。这是用户在响应开始流式传输前感受到的“延迟”。 |
Latency | 从请求发送到流关闭的总时间。 |
Tok/s | 生成吞吐量:首个 token 之后每秒生成的 token 数量。计算公式为 |
Tokens | 总输出 token 数。优先使用提供商的使用元数据(如果可用),否则回退到 SSE 事件计数。 |
Status | 如果所有阈值均通过则为 |
命令
llmprobe probe
一次性健康检查。探测所有配置的端点并打印结果。
llmprobe probe # table output
llmprobe probe -f json # JSON output
llmprobe probe --fail-on degraded # exit 1 if any endpoint is degraded
llmprobe probe -c custom-config.yml # custom config pathCI 的退出代码:
| 退出 0 | 退出 1 |
| 健康或降级 | 任何错误 |
| 仅健康 | 降级或错误 |
| 总是 | 从不 |
llmprobe watch
持续监控。按间隔探测所有端点,并为每次迭代打印一行摘要。
llmprobe watch # default 60s interval
llmprobe watch --interval 30s # custom interval
llmprobe watch --tui # live terminal dashboard with TTFT chart
llmprobe watch --tui --load data.jsonl # load historical data into the dashboard
llmprobe watch -f json # JSONL output (one line per result)--tui 标志会启动一个实时终端仪表板,包含 TTFT 图表、颜色图例和统计表格。使用 --load 导入历史 JSONL 数据(来自 llmprobe watch -f json > data.jsonl)。

$ llmprobe watch --interval 30s
Watching 4 endpoints every 30s (Ctrl+C to stop)
[14:01:02] All 4 endpoints healthy.
[14:01:32] All 4 endpoints healthy.
[14:02:02] 3 healthy, 1 degraded, 0 errors. DEGRADED: openai/gpt-4o (TTFT 1820ms)
[14:02:32] All 4 endpoints healthy.CI 集成
将 llmprobe probe 用作部署前门禁:
# .github/workflows/deploy.yml
- name: Check LLM providers
env:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
run: |
go install github.com/Jwrede/llmprobe@latest
llmprobe probe --fail-on degraded如果任何 LLM 提供商当前性能下降,这将阻止部署。
MCP 服务器
llmprobe 内置了一个 Model Context Protocol 服务器,允许 Claude Code 和其他 MCP 主机直接从代理工作流中检查 LLM API 的健康状况。
运行服务器
llmprobe mcp这将通过 stdio 启动 MCP 服务器。
在 Claude Code 中注册
claude mcp add --transport stdio llmprobe -- llmprobe mcp注册后,Claude Code 可以在任何对话中调用 llmprobe 工具。
可用工具
工具 | 描述 |
| 探测 |
| 在没有配置文件的情况下探测单个模型。需要 |
| 列出配置文件中所有提供商和模型及其阈值。在探测前使用此工具查看可用模型。 |
| 返回完整的解析配置,包括默认值、提供商、模型和阈值。 |
使用示例: 代理调用 list_providers 查看配置了哪些模型,然后调用 probe_all 在部署更改前验证它们是否健康。
配置
defaults:
prompt: "Hello" # probe prompt
max_tokens: 20 # max output tokens
timeout: 30s # per-probe timeout
concurrency: 5 # max parallel probes
providers:
- name: openai # openai, anthropic, google, azure, bedrock
api_key: ${OPENAI_API_KEY} # env var expansion
base_url: https://custom.api # optional, override endpoint
models:
- name: gpt-4o
prompt: "Say hello." # override default prompt
max_tokens: 10 # override default max_tokens
thresholds:
max_ttft: 2s # alert if TTFT exceeds this
max_latency: 10s # alert if total latency exceeds this
min_tokens_per_sec: 20 # alert if throughput drops below this
- name: azure
api_key: ${AZURE_OPENAI_API_KEY}
base_url: https://your-resource.openai.azure.com
api_version: "2024-10-21" # optional, defaults to 2024-10-21
models:
- name: gpt-4o # deployment name
- name: bedrock
access_key: ${AWS_ACCESS_KEY_ID}
secret_key: ${AWS_SECRET_ACCESS_KEY}
region: us-east-1
models:
- name: anthropic.claude-3-5-sonnet-20241022-v2:0API 密钥和 AWS 凭证支持 ${ENV_VAR} 语法。仅凭证字段会被展开,因此提示词或模型名称中的环境变量引用将保持原样。
OpenAI 兼容提供商
许多提供商(Groq、Together AI、Fireworks、DeepSeek、Mistral、OpenRouter、Ollama、vLLM)公开了与 OpenAI 兼容的 API。通过设置 base_url 即可直接使用:
providers:
# Groq
- name: openai
api_key: ${GROQ_API_KEY}
base_url: https://api.groq.com/openai
models:
- name: llama-3.3-70b-versatile
# DeepSeek
- name: openai
api_key: ${DEEPSEEK_API_KEY}
base_url: https://api.deepseek.com
models:
- name: deepseek-chat
# Together AI
- name: openai
api_key: ${TOGETHER_API_KEY}
base_url: https://api.together.xyz
models:
- name: meta-llama/Meta-Llama-3.1-70B-Instruct-Turbo
# Local Ollama
- name: openai
api_key: unused
base_url: http://localhost:11434/v1
models:
- name: llama3.2架构
probes.yml
-> Config loader (YAML + env var expansion)
-> Probe engine (concurrent goroutines per provider/model)
-> Provider clients (raw HTTP + SSE parsing, no SDKs)
-> Results (TTFT, latency, tokens/sec, status)
-> Output (table, JSON, JSONL)每个提供商客户端都是一个轻量级的 HTTP 包装器,用于发送流式请求并解析响应。不导入任何 LLM SDK。SSE 解析器同时处理仅数据事件(OpenAI、Google)和命名事件(Anthropic)。Bedrock 客户端从零实现了 SigV4 签名和 AWS 二进制事件流解析。
TTFT 是从发送 HTTP 请求的那一刻起,到包含实际内容文本(而非角色分配或元数据)的第一个事件为止进行测量的。
提供商
提供商 | 端点 | 认证 | 流式格式 |
OpenAI |
|
| SSE, |
Anthropic |
|
| 命名事件 SSE |
|
| SSE | |
Azure OpenAI |
|
| SSE, |
AWS Bedrock |
| SigV4 | AWS 二进制事件流 |
OpenAI-compat |
|
| SSE |
OpenAI 兼容涵盖:Groq、Together AI、Fireworks、DeepSeek、Mistral、OpenRouter、Ollama、vLLM 以及任何支持 OpenAI 聊天补全 API 的端点。
实时基准测试
llm-bench 使用 llmprobe 对主要 LLM API 进行持续的公开基准测试。结果以开放的 JSONL 数据集形式发布,并在 bench.jonathanwrede.de 提供实时终端仪表板。
路线图
基准跟踪:存储滚动百分位数,当当前探测超过 Nx 基准时发出警报
OpenTelemetry 指标导出,用于与 Grafana/Datadog 集成
Prometheus
/metrics端点结构化输出验证:验证 JSON 模式响应是否正确解析
许可证
MIT
Available Tools
4 toolsget_configA
Return the full parsed configuration including defaults, providers, models, and thresholds. Useful for understanding the current probe setup or debugging configuration issues.
| Name | Required | Description | Default |
|---|---|---|---|
| config | No | path to probes.yml config file |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Describes return contents but does not mention side effects, auth requirements, or rate limits. No annotations exist to supplement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two efficient sentences, no redundancy, front-loaded with purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Adequately covers purpose, return content, and common use cases for a simple tool with one optional parameter and no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers the single parameter with description. Description adds no extra meaning beyond 'full parsed configuration'; baseline 3 due to high coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it returns the full parsed configuration, listing included elements (defaults, providers, models, thresholds). Distinct from siblings like list_providers or probe_all.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Indicates usefulness for understanding setup or debugging, implying context. Lacks explicit when-not-to-use or comparison to siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_providersA
List all providers and models defined in the config file. Returns provider names, model identifiers, and any configured thresholds. Use this to discover what models are available before probing.
| Name | Required | Description | Default |
|---|---|---|---|
| config | No | path to probes.yml config file |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Implicitly a read operation, but with no annotations, the description should explicitly state it is read-only or disclose any side effects. It lacks explicit non-destructive guarantee.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences: first states what it does, second tells when to use it. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple single-parameter tool with no output schema, the description is sufficiently complete, covering purpose, returns, and usage context. Minor gap in behavioral transparency.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a single parameter described. The description adds no additional meaning beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it lists providers and models from a config file, specifying the exact information returned (names, identifiers, thresholds). This distinguishes it from sibling tools like 'probe_all' or 'get_config'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly advises using it 'before probing', providing clear usage context. However, it does not contrast with siblings or specify when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
probe_allA
Probe all configured LLM API endpoints. Returns TTFT (ms), total latency (ms), throughput (tokens/sec), and health status for every model in the config file.
| Name | Required | Description | Default |
|---|---|---|---|
| config | No | path to probes.yml config file |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. It discloses return values and that it uses a config file, but does not mention side effects, error handling, or whether it's read-only. Adequate but not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence that is concise, front-loaded, and contains no filler. Every word adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description adequately explains the return values. The single optional parameter is well-described. Sibling tool context implies complementarity with 'probe_model'. Complete for a probing tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a clear description for the 'config' parameter. The description adds context by explaining that the config file determines which endpoints are probed, enhancing the schema's meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Probe all configured LLM API endpoints') and specifies the exact metrics returned (TTFT, latency, throughput, health status). It distinguishes from sibling 'probe_model' which likely targets a single model.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives like 'probe_model' or 'get_config'. The description does not specify prerequisites or scenarios where this tool is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
probe_modelA
Probe a single LLM model by provider and model name. Use this for ad-hoc checks without a config file. Returns TTFT (ms), total latency (ms), throughput (tokens/sec), and health status.
| Name | Required | Description | Default |
|---|---|---|---|
| provider | Yes | provider name (openai, anthropic, google, azure, bedrock) | |
| model | Yes | model identifier (e.g. gpt-4o, claude-sonnet-4-20250514) | |
| api_key_env | Yes | environment variable name containing the API key | |
| base_url | No | optional base URL for OpenAI-compatible endpoints (e.g. http://localhost:8000) | |
| label | No | optional display name for the endpoint (e.g. vllm-local) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It discloses return values (TTFT, latency, throughput, health status) and states it probes a model, but fails to mention side effects (e.g., real API call) or safety properties (read-only vs. destructive). This is adequate but not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, each carrying essential information: first sentence states purpose and parameters, second sentence clarifies usage context and return values. No redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter tool with no output schema, the description adequately covers return values and usage context. It could mention that the tool makes a live API call, but overall completeness is high for a simple diagnostic tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description does not add parameter-level details beyond what the schema already provides; it only mentions 'provider and model name' generically.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Probe') and explicitly states the resource ('single LLM model by provider and model name'). It also distinguishes from siblings (probe_all, list_providers) by noting ad-hoc single-model use without a config file.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly recommends use for 'ad-hoc checks without a config file', implying when to use this tool. Sibling names (probe_all, list_providers, get_config) provide contrast, but no explicit exclusions or when-not-to-use guidance are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v1.4.0- Removed
check_model - Added
get_config - Added
list_providers - Removed
probe - Added
probe_all - Added
probe_model
2 tool updates
v0.1.0- First observed
check_model - First observed
probe
TDQS
Scored across 4 tools
Each tool has a distinct purpose: configuration retrieval, provider listing, bulk probing, and single model probing. No overlap.
All tool names follow a consistent verb_noun pattern (get_config, list_providers, probe_all, probe_model) with appropriate verbs.
4 tools is well-scoped for the LLM probing domain, covering necessary operations without bloat.
The tool set covers all essential probe operations: viewing config, listing providers, probing all endpoints, and probing a single endpoint, with no obvious gaps.
Maintenance
Related MCP Connectors
Measured latency, time to first token and uptime for ~45 AI inference APIs, by region.
Live LLM API price + status radar across 11 providers, with public per-model price HISTORY.
35-probe LLM/agent security red-team scan (injection, jailbreak, MCP abuse) with report.
Related MCP Servers
- FlicenseBqualityDmaintenanceProvides a universal bridge to interact with any OpenAI-compatible LLM API (local or cloud), enabling model testing, benchmarking, quality evaluation, and chat operations with performance metrics.8-
- AlicenseNot gradedqualityNot gradedmaintenanceEnables benchmarking of Large Language Model APIs by measuring performance metrics such as generation throughput, prompt throughput, and Time To First Token (TTFT) with configurable concurrency levels and parameters.1MIT
- AlicenseNot gradedqualityDmaintenanceExposes queryable GPU inference benchmark data (quantization, throughput, VRAM, concurrent users) as tools for LLM clients.MIT

qa-probeofficial
AlicenseAqualityCmaintenanceProbes your live API and classifies why each endpoint failed (root cause, evidence, and a calibrated confidence level), exposed over MCP so your AI assistant debugs from evidence instead of guessing. Works with FastAPI, Express, Next.js, tRPC, and GraphQL.83 npm2Apache 2.0