Skip to main content
Glama

llmprobe

llmprobe

LLM API 엔드포인트를 프로빙합니다. TTFT, 지연 시간, 처리량을 측정합니다. 단일 바이너리, SDK 없음.

CI Go License: MIT llmprobe MCP server

llmprobe는 LLM API 엔드포인트를 프로빙하고 프로덕션 안정성에 중요한 지표인 첫 토큰까지의 시간(TTFT), 총 지연 시간, 생성 처리량(토큰/초) 및 오류율을 측정하는 CLI 도구입니다.

일회성 상태 점검, 지속적인 모니터링 또는 LLM 제공업체의 성능이 저하되었을 때 배포를 차단하는 CI 게이트로 사용하세요.

demo

빠른 시작

최신 릴리스에서 사전 빌드된 바이너리를 다운로드하세요 (Linux, macOS, Windows; amd64 및 arm64).

또는 소스에서 설치하세요:

go install github.com/Jwrede/llmprobe@latest

probes.yml을 생성하세요(또는 포함된 예제를 복사하세요):

providers:
  - name: openai
    api_key: ${OPENAI_API_KEY}
    models:
      - name: gpt-4o
        thresholds:
          max_ttft: 2s
      - name: gpt-4o-mini
        thresholds:
          max_ttft: 500ms

  - name: anthropic
    api_key: ${ANTHROPIC_API_KEY}
    models:
      - name: claude-sonnet-4-20250514
        thresholds:
          max_ttft: 1s

프로브를 실행하세요:

$ llmprobe probe

Provider   Model                    Status    TTFT    Latency  Tok/s  Tokens  Error
--------   -----                    ------    ----    -------  -----  ------  -----
openai     gpt-4o                   healthy   312ms   2100ms   68.4   42
openai     gpt-4o-mini              healthy   98ms    814ms    112.3  56
anthropic  claude-sonnet-4-20250514 healthy   420ms   2831ms   52.1   38
azure      gpt-4o                   healthy   289ms   1950ms   71.2   44
bedrock    anthropic.claude-3-5...  degraded  1820ms  4510ms   28.1   38

4 healthy, 1 degraded, 0 errors

Related MCP server: LLM API Benchmark MCP Server

측정 항목

지표

의미

TTFT

요청 전송부터 첫 번째 콘텐츠 토큰까지의 시간입니다. 사용자가 응답 스트리밍이 시작되기 전 느끼는 "지연"입니다.

Latency

요청부터 스트림 종료까지의 총 시간입니다.

Tok/s

생성 처리량: 첫 번째 토큰 이후 초당 생성된 토큰 수입니다. token_count / (latency - ttft)로 계산됩니다.

Tokens

총 출력 토큰 수입니다. 제공업체 사용량 메타데이터를 우선 사용하며, 없을 경우 SSE 이벤트 카운팅을 사용합니다.

Status

모든 임계값을 통과하면 healthy, 임계값을 초과하면 degraded, 요청 실패 시 error로 표시됩니다.

명령어

llmprobe probe

일회성 상태 점검입니다. 구성된 모든 엔드포인트를 프로빙하고 결과를 출력합니다.

llmprobe probe                        # table output
llmprobe probe -f json                # JSON output
llmprobe probe --fail-on degraded     # exit 1 if any endpoint is degraded
llmprobe probe -c custom-config.yml   # custom config path

CI를 위한 종료 코드:

--fail-on

종료 0

종료 1

error (기본값)

healthy 또는 degraded

모든 오류

degraded

healthy만

degraded 또는 error

none

항상

절대 없음

llmprobe watch

지속적인 모니터링입니다. 일정 간격으로 모든 엔드포인트를 프로빙하고 반복마다 요약 라인을 출력합니다.

llmprobe watch                          # default 60s interval
llmprobe watch --interval 30s           # custom interval
llmprobe watch --tui                    # live terminal dashboard with TTFT chart
llmprobe watch --tui --load data.jsonl  # load historical data into the dashboard
llmprobe watch -f json                  # JSONL output (one line per result)

--tui 플래그는 TTFT 차트, 색상 범례 및 통계 테이블이 포함된 실시간 터미널 대시보드를 실행합니다. --load를 사용하여 과거 JSONL 데이터(llmprobe watch -f json > data.jsonl에서 생성)를 가져올 수 있습니다.

llmprobe

$ llmprobe watch --interval 30s

Watching 4 endpoints every 30s (Ctrl+C to stop)

[14:01:02] All 4 endpoints healthy.
[14:01:32] All 4 endpoints healthy.
[14:02:02] 3 healthy, 1 degraded, 0 errors. DEGRADED: openai/gpt-4o (TTFT 1820ms)
[14:02:32] All 4 endpoints healthy.

CI 통합

llmprobe probe를 배포 전 게이트로 사용하세요:

# .github/workflows/deploy.yml
- name: Check LLM providers
  env:
    OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
    ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
  run: |
    go install github.com/Jwrede/llmprobe@latest
    llmprobe probe --fail-on degraded

이는 현재 LLM 제공업체의 성능이 저하된 경우 배포를 차단합니다.

MCP 서버

llmprobe MCP server

llmprobe에는 내장된 Model Context Protocol 서버가 포함되어 있어, Claude Code 및 기타 MCP 호스트가 에이전트 워크플로우에서 직접 LLM API 상태를 확인할 수 있습니다.

서버 실행

llmprobe mcp

stdio를 통해 MCP 서버를 시작합니다.

Claude Code에 등록

claude mcp add --transport stdio llmprobe -- llmprobe mcp

등록이 완료되면 Claude Code는 대화 중에 llmprobe 도구를 호출할 수 있습니다.

사용 가능한 도구

도구

설명

probe_all

probes.yml에 구성된 모든 엔드포인트를 프로빙합니다. 모든 모델에 대한 TTFT, 지연 시간, 처리량 및 상태를 반환합니다. 사용자 지정 구성 경로를 위한 선택적 config 매개변수를 허용합니다.

probe_model

구성 파일 없이 단일 모델을 프로빙합니다. provider(openai, anthropic, google, azure, bedrock), model(모델 식별자), api_key_env(API 키를 포함하는 환경 변수)가 필요합니다.

list_providers

구성 파일의 모든 제공업체와 모델을 임계값과 함께 나열합니다. 프로빙 전 사용 가능한 모델을 확인하는 데 사용하세요.

get_config

기본값, 제공업체, 모델 및 임계값을 포함하여 완전히 파싱된 구성을 반환합니다.

사용 예시: 에이전트가 list_providers를 호출하여 구성된 모델을 확인한 다음, probe_all을 호출하여 변경 사항을 배포하기 전에 상태가 정상인지 확인합니다.

구성

defaults:
  prompt: "Hello"                                # probe prompt
  max_tokens: 20                                 # max output tokens
  timeout: 30s                                   # per-probe timeout
  concurrency: 5                                 # max parallel probes

providers:
  - name: openai                    # openai, anthropic, google, azure, bedrock
    api_key: ${OPENAI_API_KEY}      # env var expansion
    base_url: https://custom.api    # optional, override endpoint
    models:
      - name: gpt-4o
        prompt: "Say hello."        # override default prompt
        max_tokens: 10              # override default max_tokens
        thresholds:
          max_ttft: 2s              # alert if TTFT exceeds this
          max_latency: 10s          # alert if total latency exceeds this
          min_tokens_per_sec: 20    # alert if throughput drops below this

  - name: azure
    api_key: ${AZURE_OPENAI_API_KEY}
    base_url: https://your-resource.openai.azure.com
    api_version: "2024-10-21"       # optional, defaults to 2024-10-21
    models:
      - name: gpt-4o               # deployment name

  - name: bedrock
    access_key: ${AWS_ACCESS_KEY_ID}
    secret_key: ${AWS_SECRET_ACCESS_KEY}
    region: us-east-1
    models:
      - name: anthropic.claude-3-5-sonnet-20241022-v2:0

API 키와 AWS 자격 증명은 ${ENV_VAR} 구문을 지원합니다. 자격 증명 필드만 확장되므로 프롬프트나 모델 이름의 환경 변수 참조는 그대로 유지됩니다.

OpenAI 호환 제공업체

많은 제공업체(Groq, Together AI, Fireworks, DeepSeek, Mistral, OpenRouter, Ollama, vLLM)는 OpenAI 호환 API를 노출합니다. base_url을 설정하면 즉시 작동합니다:

providers:
  # Groq
  - name: openai
    api_key: ${GROQ_API_KEY}
    base_url: https://api.groq.com/openai
    models:
      - name: llama-3.3-70b-versatile

  # DeepSeek
  - name: openai
    api_key: ${DEEPSEEK_API_KEY}
    base_url: https://api.deepseek.com
    models:
      - name: deepseek-chat

  # Together AI
  - name: openai
    api_key: ${TOGETHER_API_KEY}
    base_url: https://api.together.xyz
    models:
      - name: meta-llama/Meta-Llama-3.1-70B-Instruct-Turbo

  # Local Ollama
  - name: openai
    api_key: unused
    base_url: http://localhost:11434/v1
    models:
      - name: llama3.2

아키텍처

probes.yml
  -> Config loader (YAML + env var expansion)
    -> Probe engine (concurrent goroutines per provider/model)
      -> Provider clients (raw HTTP + SSE parsing, no SDKs)
        -> Results (TTFT, latency, tokens/sec, status)
          -> Output (table, JSON, JSONL)

각 제공업체 클라이언트는 스트리밍 요청을 보내고 응답을 파싱하는 가벼운 HTTP 래퍼입니다. LLM SDK는 가져오지 않습니다. SSE 파서는 데이터 전용 이벤트(OpenAI, Google)와 명명된 이벤트(Anthropic)를 모두 처리합니다. Bedrock 클라이언트는 SigV4 서명과 AWS 바이너리 이벤트 스트림 파싱을 처음부터 구현합니다.

TTFT는 HTTP 요청이 전송된 순간부터 실제 콘텐츠 텍스트(역할 할당이나 메타데이터가 아닌)를 포함하는 첫 번째 이벤트까지 측정됩니다.

제공업체

제공업체

엔드포인트

인증

스트리밍 형식

OpenAI

/v1/chat/completions

Authorization: Bearer

SSE, [DONE] 센티넬

Anthropic

/v1/messages

x-api-key 헤더

명명된 이벤트 SSE

Google

/v1beta/models/{model}:streamGenerateContent?alt=sse

key 쿼리 매개변수

SSE

Azure OpenAI

/openai/deployments/{model}/chat/completions

api-key 헤더

SSE, [DONE] 센티넬

AWS Bedrock

/model/{model}/converse-stream

SigV4

AWS 바이너리 이벤트 스트림

OpenAI-compat

/v1/chat/completions (사용자 지정 base_url)

Authorization: Bearer

SSE

OpenAI 호환에는 Groq, Together AI, Fireworks, DeepSeek, Mistral, OpenRouter, Ollama, vLLM 및 OpenAI 채팅 완료 API를 사용하는 모든 엔드포인트가 포함됩니다.

실시간 벤치마크

llm-bench는 llmprobe를 사용하여 주요 LLM API의 지속적인 공개 벤치마크를 실행합니다. 결과는 오픈 JSONL 데이터 세트로 게시되며 bench.jonathanwrede.de에서 실시간 터미널 대시보드로 확인할 수 있습니다.

로드맵

  • 기준 추적: 롤링 백분위수 저장, 현재 프로브가 Nx 기준을 초과할 때 알림

  • Grafana/Datadog 통합을 위한 OpenTelemetry 메트릭 내보내기

  • Prometheus /metrics 엔드포인트

  • 구조화된 출력 검증: JSON 모드 응답이 올바르게 파싱되는지 확인

라이선스

MIT

Available Tools

4 tools
get_configA

Return the full parsed configuration including defaults, providers, models, and thresholds. Useful for understanding the current probe setup or debugging configuration issues.

ParametersJSON Schema
NameRequiredDescriptionDefault
configNopath to probes.yml config file

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Describes return contents but does not mention side effects, auth requirements, or rate limits. No annotations exist to supplement.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two efficient sentences, no redundancy, front-loaded with purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Adequately covers purpose, return content, and common use cases for a simple tool with one optional parameter and no output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema covers the single parameter with description. Description adds no extra meaning beyond 'full parsed configuration'; baseline 3 due to high coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states it returns the full parsed configuration, listing included elements (defaults, providers, models, thresholds). Distinct from siblings like list_providers or probe_all.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Indicates usefulness for understanding setup or debugging, implying context. Lacks explicit when-not-to-use or comparison to siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_providersA

List all providers and models defined in the config file. Returns provider names, model identifiers, and any configured thresholds. Use this to discover what models are available before probing.

ParametersJSON Schema
NameRequiredDescriptionDefault
configNopath to probes.yml config file

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Implicitly a read operation, but with no annotations, the description should explicitly state it is read-only or disclose any side effects. It lacks explicit non-destructive guarantee.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences: first states what it does, second tells when to use it. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple single-parameter tool with no output schema, the description is sufficiently complete, covering purpose, returns, and usage context. Minor gap in behavioral transparency.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with a single parameter described. The description adds no additional meaning beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it lists providers and models from a config file, specifying the exact information returned (names, identifiers, thresholds). This distinguishes it from sibling tools like 'probe_all' or 'get_config'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly advises using it 'before probing', providing clear usage context. However, it does not contrast with siblings or specify when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

probe_allA

Probe all configured LLM API endpoints. Returns TTFT (ms), total latency (ms), throughput (tokens/sec), and health status for every model in the config file.

ParametersJSON Schema
NameRequiredDescriptionDefault
configNopath to probes.yml config file

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations provided, so description carries full burden. It discloses return values and that it uses a config file, but does not mention side effects, error handling, or whether it's read-only. Adequate but not comprehensive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single sentence that is concise, front-loaded, and contains no filler. Every word adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description adequately explains the return values. The single optional parameter is well-described. Sibling tool context implies complementarity with 'probe_model'. Complete for a probing tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with a clear description for the 'config' parameter. The description adds context by explaining that the config file determines which endpoints are probed, enhancing the schema's meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Probe all configured LLM API endpoints') and specifies the exact metrics returned (TTFT, latency, throughput, health status). It distinguishes from sibling 'probe_model' which likely targets a single model.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives like 'probe_model' or 'get_config'. The description does not specify prerequisites or scenarios where this tool is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

probe_modelA

Probe a single LLM model by provider and model name. Use this for ad-hoc checks without a config file. Returns TTFT (ms), total latency (ms), throughput (tokens/sec), and health status.

ParametersJSON Schema
NameRequiredDescriptionDefault
providerYesprovider name (openai, anthropic, google, azure, bedrock)
modelYesmodel identifier (e.g. gpt-4o, claude-sonnet-4-20250514)
api_key_envYesenvironment variable name containing the API key
base_urlNooptional base URL for OpenAI-compatible endpoints (e.g. http://localhost:8000)
labelNooptional display name for the endpoint (e.g. vllm-local)

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It discloses return values (TTFT, latency, throughput, health status) and states it probes a model, but fails to mention side effects (e.g., real API call) or safety properties (read-only vs. destructive). This is adequate but not comprehensive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, each carrying essential information: first sentence states purpose and parameters, second sentence clarifies usage context and return values. No redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter tool with no output schema, the description adequately covers return values and usage context. It could mention that the tool makes a live API call, but overall completeness is high for a simple diagnostic tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so baseline is 3. The description does not add parameter-level details beyond what the schema already provides; it only mentions 'provider and model name' generically.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Probe') and explicitly states the resource ('single LLM model by provider and model name'). It also distinguishes from siblings (probe_all, list_providers) by noting ad-hoc single-model use without a config file.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly recommends use for 'ad-hoc checks without a config file', implying when to use this tool. Sibling names (probe_all, list_providers, get_config) provide contrast, but no explicit exclusions or when-not-to-use guidance are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 6 tool updatesv1.4.0
    • Removedcheck_model
    • Addedget_config
    • Addedlist_providers
    • Removedprobe
    • Addedprobe_all
    • Addedprobe_model
  2. 2 tool updatesv0.1.0
    • First observedcheck_model
    • First observedprobe

TDQS

A4.3/5.0

Scored across 4 tools

Disambiguation5/5

Each tool has a distinct purpose: configuration retrieval, provider listing, bulk probing, and single model probing. No overlap.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern (get_config, list_providers, probe_all, probe_model) with appropriate verbs.

Tool Count5/5

4 tools is well-scoped for the LLM probing domain, covering necessary operations without bloat.

Completeness5/5

The tool set covers all essential probe operations: viewing config, listing providers, probing all endpoints, and probing a single endpoint, with no obvious gaps.

Maintenance

ActivityInactive
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Exposes queryable GPU inference benchmark data (quantization, throughput, VRAM, concurrent users) as tools for LLM clients.
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Probes your live API and classifies why each endpoint failed (root cause, evidence, and a calibrated confidence level), exposed over MCP so your AI assistant debugs from evidence instead of guessing. Works with FastAPI, Express, Next.js, tRPC, and GraphQL.
    8
    3 npm
    2
    Apache 2.0