Skip to main content
Glama

llmprobe

llmprobe

LLM APIエンドポイントをプローブします。TTFT、レイテンシ、スループットを測定します。単一バイナリ、SDK不要。

CI Go License: MIT llmprobe MCP server

llmprobeは、LLM APIエンドポイントをプローブし、本番環境の信頼性において重要な指標である、最初のトークンまでの時間(TTFT)、合計レイテンシ、生成スループット(トークン/秒)、およびエラー率を測定するCLIツールです。

単発のヘルスチェック、継続的な監視、またはLLMプロバイダーのパフォーマンスが低下した際にデプロイをブロックするCIゲートとして使用できます。

demo

クイックスタート

最新リリースからビルド済みバイナリをダウンロードしてください (Linux、macOS、Windows、amd64およびarm64)。

または、ソースからインストールします:

go install github.com/Jwrede/llmprobe@latest

probes.ymlを作成します(または付属の例をコピーします):

providers:
  - name: openai
    api_key: ${OPENAI_API_KEY}
    models:
      - name: gpt-4o
        thresholds:
          max_ttft: 2s
      - name: gpt-4o-mini
        thresholds:
          max_ttft: 500ms

  - name: anthropic
    api_key: ${ANTHROPIC_API_KEY}
    models:
      - name: claude-sonnet-4-20250514
        thresholds:
          max_ttft: 1s

プローブを実行します:

$ llmprobe probe

Provider   Model                    Status    TTFT    Latency  Tok/s  Tokens  Error
--------   -----                    ------    ----    -------  -----  ------  -----
openai     gpt-4o                   healthy   312ms   2100ms   68.4   42
openai     gpt-4o-mini              healthy   98ms    814ms    112.3  56
anthropic  claude-sonnet-4-20250514 healthy   420ms   2831ms   52.1   38
azure      gpt-4o                   healthy   289ms   1950ms   71.2   44
bedrock    anthropic.claude-3-5...  degraded  1820ms  4510ms   28.1   38

4 healthy, 1 degraded, 0 errors

Related MCP server: LLM API Benchmark MCP Server

測定項目

指標

意味

TTFT

リクエスト送信から最初のコンテンツトークンまでの時間。ユーザーがレスポンスのストリーミング開始前に感じる「ラグ」です。

レイテンシ

リクエストからストリーム終了までの合計時間。

Tok/s

生成スループット:最初のトークン以降の1秒あたりの生成トークン数。token_count / (latency - ttft)として計算されます。

トークン

合計出力トークン数。利用可能な場合はプロバイダーの使用状況メタデータを優先し、利用できない場合はSSEイベントカウントにフォールバックします。

ステータス

すべてのしきい値をクリアした場合はhealthy、いずれかのしきい値を超えた場合はdegraded、リクエストが失敗した場合はerrorとなります。

コマンド

llmprobe probe

単発のヘルスチェック。設定されたすべてのエンドポイントをプローブし、結果を表示します。

llmprobe probe                        # table output
llmprobe probe -f json                # JSON output
llmprobe probe --fail-on degraded     # exit 1 if any endpoint is degraded
llmprobe probe -c custom-config.yml   # custom config path

CI用の終了コード:

--fail-on

終了コード 0

終了コード 1

error (デフォルト)

healthy または degraded

いずれかのエラー

degraded

healthy のみ

degraded または error

none

常に

なし

llmprobe watch

継続的な監視。一定間隔ですべてのエンドポイントをプローブし、反復ごとに要約行を表示します。

llmprobe watch                          # default 60s interval
llmprobe watch --interval 30s           # custom interval
llmprobe watch --tui                    # live terminal dashboard with TTFT chart
llmprobe watch --tui --load data.jsonl  # load historical data into the dashboard
llmprobe watch -f json                  # JSONL output (one line per result)

--tuiフラグは、TTFTチャート、カラー凡例、統計テーブルを備えたライブターミナルダッシュボードを起動します。過去のJSONLデータ(llmprobe watch -f json > data.jsonlから)をインポートするには--loadを使用します。

llmprobe

$ llmprobe watch --interval 30s

Watching 4 endpoints every 30s (Ctrl+C to stop)

[14:01:02] All 4 endpoints healthy.
[14:01:32] All 4 endpoints healthy.
[14:02:02] 3 healthy, 1 degraded, 0 errors. DEGRADED: openai/gpt-4o (TTFT 1820ms)
[14:02:32] All 4 endpoints healthy.

CI統合

llmprobe probeをデプロイ前のゲートとして使用します:

# .github/workflows/deploy.yml
- name: Check LLM providers
  env:
    OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
    ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
  run: |
    go install github.com/Jwrede/llmprobe@latest
    llmprobe probe --fail-on degraded

これにより、LLMプロバイダーのパフォーマンスが現在低下している場合にデプロイをブロックします。

MCPサーバー

llmprobe MCP server

llmprobeにはModel Context Protocolサーバーが組み込まれており、Claude Codeやその他のMCPホストがエージェントワークフローから直接LLM APIの健全性を確認できます。

サーバーの実行

llmprobe mcp

これにより、stdio経由でMCPサーバーが起動します。

Claude Codeへの登録

claude mcp add --transport stdio llmprobe -- llmprobe mcp

登録が完了すると、Claude Codeは会話中にllmprobeツールを呼び出せるようになります。

利用可能なツール

ツール

説明

probe_all

probes.ymlから設定されたすべてのエンドポイントをプローブします。すべてのモデルのTTFT、レイテンシ、スループット、ヘルスステータスを返します。カスタム設定パス用のオプションのconfigパラメータを受け入れます。

probe_model

設定ファイルなしで単一のモデルをプローブします。provider (openai, anthropic, google, azure, bedrock)、model (モデル識別子)、api_key_env (APIキーを保持する環境変数) が必要です。

list_providers

設定ファイル内のすべてのプロバイダーとモデルを、そのしきい値とともに一覧表示します。プローブ前に利用可能なモデルを確認するために使用します。

get_config

デフォルト値、プロバイダー、モデル、しきい値を含む、解析済みの完全な設定を返します。

使用例: エージェントがlist_providersを呼び出して設定されているモデルを確認し、次にprobe_allを呼び出して変更をデプロイする前にそれらが健全であることを確認します。

設定

defaults:
  prompt: "Hello"                                # probe prompt
  max_tokens: 20                                 # max output tokens
  timeout: 30s                                   # per-probe timeout
  concurrency: 5                                 # max parallel probes

providers:
  - name: openai                    # openai, anthropic, google, azure, bedrock
    api_key: ${OPENAI_API_KEY}      # env var expansion
    base_url: https://custom.api    # optional, override endpoint
    models:
      - name: gpt-4o
        prompt: "Say hello."        # override default prompt
        max_tokens: 10              # override default max_tokens
        thresholds:
          max_ttft: 2s              # alert if TTFT exceeds this
          max_latency: 10s          # alert if total latency exceeds this
          min_tokens_per_sec: 20    # alert if throughput drops below this

  - name: azure
    api_key: ${AZURE_OPENAI_API_KEY}
    base_url: https://your-resource.openai.azure.com
    api_version: "2024-10-21"       # optional, defaults to 2024-10-21
    models:
      - name: gpt-4o               # deployment name

  - name: bedrock
    access_key: ${AWS_ACCESS_KEY_ID}
    secret_key: ${AWS_SECRET_ACCESS_KEY}
    region: us-east-1
    models:
      - name: anthropic.claude-3-5-sonnet-20241022-v2:0

APIキーとAWS認証情報は${ENV_VAR}構文をサポートしています。認証情報フィールドのみが展開されるため、プロンプトやモデル名内の環境変数参照はそのまま残ります。

OpenAI互換プロバイダー

多くのプロバイダー(Groq、Together AI、Fireworks、DeepSeek、Mistral、OpenRouter、Ollama、vLLM)はOpenAI互換APIを公開しています。これらはbase_urlを設定することでそのまま動作します:

providers:
  # Groq
  - name: openai
    api_key: ${GROQ_API_KEY}
    base_url: https://api.groq.com/openai
    models:
      - name: llama-3.3-70b-versatile

  # DeepSeek
  - name: openai
    api_key: ${DEEPSEEK_API_KEY}
    base_url: https://api.deepseek.com
    models:
      - name: deepseek-chat

  # Together AI
  - name: openai
    api_key: ${TOGETHER_API_KEY}
    base_url: https://api.together.xyz
    models:
      - name: meta-llama/Meta-Llama-3.1-70B-Instruct-Turbo

  # Local Ollama
  - name: openai
    api_key: unused
    base_url: http://localhost:11434/v1
    models:
      - name: llama3.2

アーキテクチャ

probes.yml
  -> Config loader (YAML + env var expansion)
    -> Probe engine (concurrent goroutines per provider/model)
      -> Provider clients (raw HTTP + SSE parsing, no SDKs)
        -> Results (TTFT, latency, tokens/sec, status)
          -> Output (table, JSON, JSONL)

各プロバイダークライアントは、ストリーミングリクエストを送信してレスポンスを解析する軽量なHTTPラッパーです。LLM SDKはインポートされません。SSEパーサーは、データのみのイベント(OpenAI、Google)と名前付きイベント(Anthropic)の両方を処理します。Bedrockクライアントは、SigV4署名とAWSバイナリイベントストリーム解析をゼロから実装しています。

TTFTは、HTTPリクエストが送信された瞬間から、実際のコンテンツテキスト(ロール割り当てやメタデータではない)を含む最初のイベントまでの時間を測定します。

プロバイダー

プロバイダー

エンドポイント

認証

ストリーミング形式

OpenAI

/v1/chat/completions

Authorization: Bearer

SSE, [DONE] センチネル

Anthropic

/v1/messages

x-api-key ヘッダー

名前付きイベント SSE

Google

/v1beta/models/{model}:streamGenerateContent?alt=sse

key クエリパラメータ

SSE

Azure OpenAI

/openai/deployments/{model}/chat/completions

api-key ヘッダー

SSE, [DONE] センチネル

AWS Bedrock

/model/{model}/converse-stream

SigV4

AWS バイナリイベントストリーム

OpenAI-compat

/v1/chat/completions (カスタム base_url)

Authorization: Bearer

SSE

OpenAI互換には、Groq、Together AI、Fireworks、DeepSeek、Mistral、OpenRouter、Ollama、vLLM、およびOpenAIチャット補完APIを話すあらゆるエンドポイントが含まれます。

ライブベンチマーク

llm-benchはllmprobeを使用して、主要なLLM APIの継続的な公開ベンチマークを実行しています。結果はオープンなJSONLデータセットとして公開されており、bench.jonathanwrede.deでライブターミナルダッシュボードを確認できます。

ロードマップ

  • ベースライン追跡:ローリングパーセンタイルを保存し、現在のプローブ

Available Tools

4 tools
get_configA

Return the full parsed configuration including defaults, providers, models, and thresholds. Useful for understanding the current probe setup or debugging configuration issues.

ParametersJSON Schema
NameRequiredDescriptionDefault
configNopath to probes.yml config file

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Describes return contents but does not mention side effects, auth requirements, or rate limits. No annotations exist to supplement.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two efficient sentences, no redundancy, front-loaded with purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Adequately covers purpose, return content, and common use cases for a simple tool with one optional parameter and no output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema covers the single parameter with description. Description adds no extra meaning beyond 'full parsed configuration'; baseline 3 due to high coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states it returns the full parsed configuration, listing included elements (defaults, providers, models, thresholds). Distinct from siblings like list_providers or probe_all.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Indicates usefulness for understanding setup or debugging, implying context. Lacks explicit when-not-to-use or comparison to siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_providersA

List all providers and models defined in the config file. Returns provider names, model identifiers, and any configured thresholds. Use this to discover what models are available before probing.

ParametersJSON Schema
NameRequiredDescriptionDefault
configNopath to probes.yml config file

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Implicitly a read operation, but with no annotations, the description should explicitly state it is read-only or disclose any side effects. It lacks explicit non-destructive guarantee.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences: first states what it does, second tells when to use it. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple single-parameter tool with no output schema, the description is sufficiently complete, covering purpose, returns, and usage context. Minor gap in behavioral transparency.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with a single parameter described. The description adds no additional meaning beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it lists providers and models from a config file, specifying the exact information returned (names, identifiers, thresholds). This distinguishes it from sibling tools like 'probe_all' or 'get_config'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly advises using it 'before probing', providing clear usage context. However, it does not contrast with siblings or specify when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

probe_allA

Probe all configured LLM API endpoints. Returns TTFT (ms), total latency (ms), throughput (tokens/sec), and health status for every model in the config file.

ParametersJSON Schema
NameRequiredDescriptionDefault
configNopath to probes.yml config file

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description carries full burden. It discloses return values and that it uses a config file, but does not mention side effects, error handling, or whether it's read-only. Adequate but not comprehensive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence that is concise, front-loaded, and contains no filler. Every word adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description adequately explains the return values. The single optional parameter is well-described. Sibling tool context implies complementarity with 'probe_model'. Complete for a probing tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with a clear description for the 'config' parameter. The description adds context by explaining that the config file determines which endpoints are probed, enhancing the schema's meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Probe all configured LLM API endpoints') and specifies the exact metrics returned (TTFT, latency, throughput, health status). It distinguishes from sibling 'probe_model' which likely targets a single model.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives like 'probe_model' or 'get_config'. The description does not specify prerequisites or scenarios where this tool is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

probe_modelA

Probe a single LLM model by provider and model name. Use this for ad-hoc checks without a config file. Returns TTFT (ms), total latency (ms), throughput (tokens/sec), and health status.

ParametersJSON Schema
NameRequiredDescriptionDefault
providerYesprovider name (openai, anthropic, google, azure, bedrock)
modelYesmodel identifier (e.g. gpt-4o, claude-sonnet-4-20250514)
api_key_envYesenvironment variable name containing the API key
base_urlNooptional base URL for OpenAI-compatible endpoints (e.g. http://localhost:8000)
labelNooptional display name for the endpoint (e.g. vllm-local)

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It discloses return values (TTFT, latency, throughput, health status) and states it probes a model, but fails to mention side effects (e.g., real API call) or safety properties (read-only vs. destructive). This is adequate but not comprehensive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, each carrying essential information: first sentence states purpose and parameters, second sentence clarifies usage context and return values. No redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter tool with no output schema, the description adequately covers return values and usage context. It could mention that the tool makes a live API call, but overall completeness is high for a simple diagnostic tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description does not add parameter-level details beyond what the schema already provides; it only mentions 'provider and model name' generically.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Probe') and explicitly states the resource ('single LLM model by provider and model name'). It also distinguishes from siblings (probe_all, list_providers) by noting ad-hoc single-model use without a config file.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly recommends use for 'ad-hoc checks without a config file', implying when to use this tool. Sibling names (probe_all, list_providers, get_config) provide contrast, but no explicit exclusions or when-not-to-use guidance are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 6 tool updatesv1.4.0
    • Removedcheck_model
    • Addedget_config
    • Addedlist_providers
    • Removedprobe
    • Addedprobe_all
    • Addedprobe_model
  2. 2 tool updatesv0.1.0
    • First observedcheck_model
    • First observedprobe

TDQS

A4.3/5.0

Scored across 4 tools

Disambiguation5/5

Each tool has a distinct purpose: configuration retrieval, provider listing, bulk probing, and single model probing. No overlap.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern (get_config, list_providers, probe_all, probe_model) with appropriate verbs.

Tool Count5/5

4 tools is well-scoped for the LLM probing domain, covering necessary operations without bloat.

Completeness5/5

The tool set covers all essential probe operations: viewing config, listing providers, probing all endpoints, and probing a single endpoint, with no obvious gaps.

Maintenance

ActivityInactive
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Exposes queryable GPU inference benchmark data (quantization, throughput, VRAM, concurrent users) as tools for LLM clients.
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Probes your live API and classifies why each endpoint failed (root cause, evidence, and a calibrated confidence level), exposed over MCP so your AI assistant debugs from evidence instead of guessing. Works with FastAPI, Express, Next.js, tRPC, and GraphQL.
    8
    3 npm
    2
    Apache 2.0