prime-intellect-mcp
prime-intellect-mcp
Claude Code가 직접 Prime Intellect GPU 포드를 대여, 구동 및 종료하도록 하세요. 사용자가 제어하는 하드 지출 한도가 적용됩니다.
소개
Claude Code(또는 모든 MCP 클라이언트)를 Prime Intellect 계정에 연결하는 MCP 서버입니다. 이를 통해 에이전트는 다음 작업을 수행할 수 있습니다:
🔍 검색: 요구 사항에 맞는 가장 저렴한 GPU 포드 찾기
💸 견적: 비용을 지불하기 전에 가격 확인
🛒 프로비저닝: 포드 생성 (
confirm=True를 입력한 경우에만)🖥️ SSH: 포드에 접속 (연결 문자열은 에이전트의
Bash도구로 전달됨)🛑 종료: 작업 완료 후 포드 종료 — 잊어버린 경우 강력하게 경고
단일 워크플로우를 위해 구축되었습니다: Claude에게 *"가장 저렴한 H100을 대여해서 학습 스크립트를 실행하고, 끝나면 종료해"*라고 말하고 400달러 청구서를 받지 않도록 방지합니다.
Related MCP server: claude-colab
60초 만에 설치하기
Claude Code를 통해 GPU 대여를 시작하려면 다음만 있으면 됩니다:
1. Prime Intellect API 키 받기
여기를 클릭하여 생성 → 권한 설정:
범위 | 수준 |
Instances | 읽기 및 쓰기 |
Availability | 읽기 전용 |
Billing | 읽기 전용 |
SSH Keys | 읽기 전용 |
키를 복사하세요 (pit_…로 시작합니다).
2. Claude Code에 서버 추가
~/Library/Application Support/Claude/claude_desktop_config.json(macOS) 또는 프로젝트의 .mcp.json을 열고 다음을 붙여넣으세요:
{
"mcpServers": {
"prime-intellect": {
"command": "uvx",
"args": ["prime-intellect-mcp"],
"env": {
"PRIME_API_KEY": "pit_PASTE_YOURS_HERE",
"PRIME_MAX_HOURLY_USD": "5",
"PRIME_MAX_TOTAL_USD": "40"
}
}
}
}끝입니다. Claude Code를 재시작하고 다음과 같이 물어보세요: "지금 시간당 1달러 미만으로 사용 가능한 GPU가 뭐야?"
uvx가 없나요?curl -LsSf https://astral.sh/uv/install.sh | sh(또는brew install uv)로 설치하세요.uv패키지 관리자를 위한 한 줄 설치 프로그램이며, 다시는 가상 환경을 관리할 필요가 없습니다.
✨ SSH 추가 (선택 사항, +2분) — Claude가 포드에서 코드를 실행하려면 필요
위의 서버는 이미 포드를 프로비저닝/검사/종료할 수 있습니다. 하지만 Claude Code가 실행 중인 포드에 SSH로 접속하여 명령을 실행하려면, Prime Intellect가 사용자의 공개 SSH 키를 알고 있어야 합니다.
3. 머신에서 SSH 키 찾기 또는 생성
ls ~/.ssh/*.pub # if you have id_ed25519.pub or similar, you're set
# otherwise:
ssh-keygen -t ed25519 -C "you@example.com" # press Enter through the prompts4. Prime Intellect에 공개 키 등록
cat ~/.ssh/id_ed25519.pub # or whichever .pub file you have출력된 내용(ssh-ed25519 …로 시작하는 한 줄)을 복사하여 app.primeintellect.ai/dashboard/ssh-keys의 Add SSH key 양식에 붙여넣으세요.
끝입니다. 향후 생성되는 포드에는 authorized_keys에 사용자의 공개 키가 포함되며, Claude Code의 Bash 도구가 바로 SSH로 접속할 수 있습니다:
ssh ubuntu@<pod-ip-from-pod_status> "nvidia-smi"v0.2 예정: Claude 내부에서 4단계를 수행하는
register_ssh_keyMCP 도구 (브라우저 방문 불필요). 이슈 트래커에서 진행 상황을 확인하세요.
Claude가 수행할 수 있는 작업 (9가지 도구)
도구 | 사용 사례 |
| "Prime Intellect에서 제공하는 GPU 유형은 무엇인가요?" |
| "시간당 3달러 미만으로 사용 가능한 1×H100 포드를 보여줘." |
| "남은 크레딧이 얼마인가요?" |
| "200GB 디스크가 포함된 1×A100 견적을 내줘." (비용 발생 없음) |
| "해당 견적으로 포드를 프로비저닝해." ( |
| "실행 중인 포드를 보여줘." |
| "포드 X가 준비되었나요? SSH 정보가 나올 때까지 기다려." |
| "포드 X를 종료해." ( |
| "종료하지 않은 포드가 있나요?" |
안전: 자동 프로비저닝 방지
다음 3단계 안전장치가 있습니다:
견적 우선.
pod_quote는 가격과 60초 유효 토큰을 반환합니다. 부작용은 없습니다. 달러 금액은 이제 에이전트의 컨텍스트에 포함됩니다.명시적 확인.
pod_create(및pod_terminate)는confirm=True가 필요합니다. 없으면 드라이런(dry-run) 미리보기만 제공됩니다.환경 변수 하드 캡.
PRIME_MAX_HOURLY_USD는 해당 요금 이상의 포드를 차단합니다.PRIME_MAX_TOTAL_USD는 예산을 초과하는 (요금 × 최대_수명_시간) 포드를 차단합니다. 지갑 잔액도 강제 적용됩니다. 이러한 캡은 도구 인수로 재정의할 수 없으며, 호출될 때마다 읽어옵니다.
기본값: PRIME_MAX_HOURLY_USD=5, PRIME_MAX_TOTAL_USD=40. 설정의 env 블록에서 설정하세요.
모든 pod_create / pod_terminate는 ~/.prime-intellect-mcp/audit.log에 JSON으로 추가되므로, 에이전트가 사용자의 돈으로 무엇을 했는지 전체 기록을 확인할 수 있습니다.
예시 프롬프트 (Claude Code에 붙여넣기)
List the cheapest 1×H100 pods available right now. Show me the top 3 by hourly price.Quote a 1×A100 80GB with 100GB disk, 8 vCPU, 64GB RAM. Don't provision yet —
just show me what it would cost.I need to fine-tune a 7B model overnight. Find the cheapest 1×H100 with 200GB
disk, max $40 total budget, max 12 hours. Provision it, give me the SSH command,
and remind me to terminate when I'm done.Check if I have any running pods I forgot about and show me their hourly cost.Terminate pod abc123. Confirm before doing it.문제 해결
Claude Code 설정이 env 블록을 가져오지 못했거나, PRIME_API_KEY를 다른 변수로 입력했을 수 있습니다. 다음으로 확인하세요:
$ env | grep PRIMEClaude Code를 실행하는 동일한 셸 내부에서 확인하거나, (${PRIME_API_KEY}를 사용하는 대신) 키를 JSON env 블록에 직접 붙여넣으세요.
에이전트가 하드 캡을 초과하는 포드를 선택했습니다. 다음 중 하나를 수행하세요:
더 저렴한 GPU를 선택하세요 (
list_availability에 지역 필터를 사용하면 더 저렴한 커뮤니티 가격 행이 나타나는 경우가 많습니다).설정에서
PRIME_MAX_HOURLY_USD를 높이고 Claude Code를 재시작하세요.
견적은 60초 동안 유효합니다. 에이전트가 pod_quote와 pod_create 사이에서 너무 오래 기다렸습니다. pod_quote를 다시 호출하세요. 비용이 발생하지 않는 작업입니다.
프로비저닝이 완전히 완료되지 않았습니다. 포드는 살아있지만 여전히 설치 스크립트를 실행 중입니다. pod_status(pod_id, wait_for_ssh=True)를 호출하면 SSH가 활성화될 때까지 (5초마다 폴링하며) 대기합니다.
Prime Intellect에 공개 키를 알리지 않았거나(또는 등록하기 전에 포드가 프로비저닝됨) 발생합니다. 해결 방법:
app.primeintellect.ai/dashboard/ssh-keys에서 공개 키가 등록되었는지 확인하세요.
재프로비저닝 — 포드의
authorized_keys는 생성 시점에 설정되므로, 기존 포드는 나중에 등록한 키를 가져오지 않습니다.개인 키에 암호가 있는 경우, macOS에서
ssh-add --apple-use-keychain ~/.ssh/your_key를 한 번 실행하면 에이전트가 자동으로 잠금을 해제합니다.
app.primeintellect.ai/wallet에서 충전 후 다시 시도하세요.
왜 또 다른 서버인가요?
PyPI에 prime-mcp-server 0.1.2가 있습니다. 이는 간단한 개념 증명이며, 이 프로젝트는 포크가 아닙니다. 무인 야간 사용을 위한 차이점은 다음과 같습니다:
|
| |
2단계 견적 → 확인 | ✅ | ❌ |
환경 변수 하드 지출 캡 | ✅ | ❌ |
지갑 사전 확인 | ✅ | ❌ |
실행 중인 포드 감지 | ✅ | ❌ |
에이전트로 SSH 핸드오프 | ✅ | ❌ |
테스트 | 32개 단위 + 선택적 라이브 | 없음 |
로컬 개발
git clone https://github.com/kvrancic/prime-intellect-mcp
cd prime-intellect-mcp
uv sync
uv run pytest -m "not live" # 32 fast tests, no network, no spend
uv run ruff check .
uv run mypy src라이브 스모크 테스트 (가장 저렴한 GPU 프로비저닝, nvidia-smi 실행, 종료; 약 $0.05 소요):
PRIME_API_KEY=pit_... PRIME_LIVE_TEST=1 PRIME_LIVE_MAX_HOURLY=0.60 \
PRIME_MAX_HOURLY_USD=0.60 PRIME_MAX_TOTAL_USD=2.00 \
uv run pytest tests/test_smoke_live.py -v -s로드맵
v0.2 —
register_ssh_keyMCP 도구 (대시보드 단계 제거), 샌드박스(prime-sandboxesSDK), 환경 허브v0.3 — 선택적 자동 종료 데몬 (서버 측
max_lifetime_hours강제 적용); 비용 원격 측정v1.0+ — Prime Intellect가 OAuth를 출시할 때 호스팅/OAuth 배포; Anthropic 커넥터 디렉토리에 제출
감사의 말
Prime Intellect: 작업의 90%를 수행하는
primePython SDK 제공MIT 6.8610 (Advanced NLP): 테스트를 가능하게 한 Prime Intellect 크레딧 제공
FastMCP: 프레임워크 제공
라이선스
MIT — LICENSE 참조.
기여
이슈와 PR을 환영합니다. 제출하기 전에 uv run pytest -m "not live"와 uv run ruff check .를 실행해 주세요.
Available Tools
9 toolsget_wallet_balanceA
Return the current Prime Intellect wallet balance and recent billings.
Use this to estimate how long a quoted pod can run, or to check why pod_create returned an insufficient-funds error.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description adequately discloses a read-only behavior and the return of balance and billings. There is no mention of side effects, rate limits, or auth requirements, but the tool is simple with no parameters and an output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences long with no wasted words. It front-loads the purpose immediately and follows with practical usage guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has no parameters, an output schema, and no annotations, the description fully covers its functionality, including both the return value and practical use cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the input schema coverage is 100% (vacuously). The description does not need to add parameter information, so a baseline of 4 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Return the current Prime Intellect wallet balance and recent billings,' identifying a specific verb and resource. It distinguishes itself from sibling tools like pod_create and pod_quote by focusing on wallet balance.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit use cases: estimating pod runtime and debugging insufficient-funds errors. It lacks an explicit when-not-to-use section, but the context is clear enough for a simple getter tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_availabilityA
List currently-available GPU pods that match the filters.
Returns the SDK's GPUAvailability rows (cloud_id, gpu_type, gpu_count, prices, disk/vcpu/memory bounds, stock_status, ...). Use this to pick a target before pod_quote, or to show the user options.
| Name | Required | Description | Default |
|---|---|---|---|
| gpu_type | No | GPU type slug, e.g. 'H100_80GB'. Strongly recommended — the unfiltered response is large. | |
| gpu_count | No | Required GPU count per pod (1, 2, 4, 8). None means any. | |
| regions | No | Optional list of region slugs. None means any. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description adequately discloses that it returns GPUAvailability rows and lists fields. It implies a read-only operation and mentions the unfiltered response is large, but could add performance or reliability notes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise with three sentences: purpose, detail on returned data, and usage guidance. No superfluous words, well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has three optional parameters and an output schema, the description covers purpose, return type, and usage context. It could elaborate on pagination or filtering behavior, but overall complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and each parameter has a description. The tool's description adds value by noting that gpu_type is strongly recommended due to large unfiltered response, which goes beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists currently-available GPU pods matching filters, with a specific verb and resource. It distinguishes itself from siblings by mentioning its role before pod_quote or for showing options.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says 'Use this to pick a target before pod_quote, or to show the user options,' providing clear context. However, it does not explicitly state when not to use it or compare to alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_gpu_typesA
List every GPU type Prime Intellect currently offers (e.g. "H100_80GB", "A100_80GB").
Use this when the user is vague about what they want. Pass the result into list_availability or pod_quote.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description must cover behavioral traits. It accurately describes a safe, read-only list operation. While it doesn't discuss data freshness or rate limits, the simplicity of the tool (no parameters, no side effects) makes the implicit behavior clear. Slight deduction for not mentioning any potential delays or consistency guarantees.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, each earning its place. The first sentence states the action and gives examples. The second provides usage guidance. No wasted words, and critical information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (no parameters, clear output described), the description is complete. The existence of an output schema means return values are fully specified. The description directly addresses the agent's need to clarify vague user requests and chain to other tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has no parameters, so schema description coverage is 100%. With 0 parameters, the baseline is 4. The description adds value by providing examples of GPU types, which helps agents understand the output without needing to inspect the output schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the tool lists all GPU types offered by Prime Intellect, with specific examples like 'H100_80GB' and 'A100_80GB'. It distinguishes from siblings by specifying its role in clarifying vague user requests and directing results to list_availability or pod_quote.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use the tool: 'when the user is vague about what they want.' Also provides clear next steps: 'Pass the result into list_availability or pod_quote.' This leaves no ambiguity about context and downstream usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pod_check_runawayA
Return locally-tracked pods that have run past max_lifetime_hours OR whose accumulated cost is approaching PRIME_MAX_TOTAL_USD.
Call this at the start of long-running sessions to catch forgotten pods.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden for behavioral disclosure. It states the tool returns matching pods but does not mention whether it is read-only, side effects, rate limits, or refresh behavior. For a check tool, assuming read-only is reasonable but not explicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no wasted words. The first sentence states purpose, the second provides usage guidance. Front-loaded with key information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with no parameters and an output schema (not shown but exists), the description is fairly complete. It could note that the operation is read-only, but overall it covers what the tool does and when to use it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Tool has zero parameters, baseline is 4 per instructions. Description adds context about the filtering criteria (max_lifetime_hours and cost limit) which are not parameters but clarify the tool's logic. Schema coverage is 100% due to no parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states it returns 'locally-tracked pods that have run past max_lifetime_hours OR whose accumulated cost is approaching PRIME_MAX_TOTAL_USD'. This distinguishes it from sibling tools like 'pod_list' (list all) and 'pod_status' (status of specific pod), providing a specific verb+resource combination.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly advises 'Call this at the start of long-running sessions to catch forgotten pods', giving clear when-to-use context. While it doesn't exclude other scenarios, the guidance is sufficient for typical use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pod_createA
Provision a Prime Intellect GPU pod (or preview the provisioning).
With confirm=False: returns a dry-run preview describing what would happen. With confirm=True: validates spend caps + quote freshness, then provisions.
The server enforces:
quote_token must be fresh (TTL 60s)
hourly_usd ≤ PRIME_MAX_HOURLY_USD
hourly_usd × max_lifetime_hours ≤ PRIME_MAX_TOTAL_USD
estimated total ≤ wallet balance
On success, the pod is recorded in local state.json so pod_check_runaway can warn about overdue pods later.
| Name | Required | Description | Default |
|---|---|---|---|
| quote_token | Yes | Token returned by pod_quote. | |
| name | Yes | Human-readable pod name. | |
| max_lifetime_hours | No | Soft budget cap: hourly_usd × this must fit under PRIME_MAX_TOTAL_USD. | |
| confirm | No | Required True to actually provision. False returns a dry-run preview. | |
| env_vars | No | Optional env vars to inject into the pod. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, description discloses dry-run vs actual provisioning, server constraints, and side effects like recording in state.json for runaway detection.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
About 80 words, well-structured with bullet points, front-loaded with key action, no wasted sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers both modes, constraints, side effects. Output schema exists so return values not needed. Complete for provisioning tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
100% schema coverage but description adds context: explains confirm's dual role, constraints on max_lifetime_hours, and how quote_token is used.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states it provisions a GPU pod or previews provisioning, using specific verbs like 'provision' and 'preview'. It distinguishes from siblings like pod_quote and pod_terminate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains when to use confirm=False vs True and lists server-enforced constraints. No explicit 'when not to use' but implicit from context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pod_listA
List every pod the API key can see (active + provisioning + stopped).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description solely bears the burden. It mentions the statuses included but not any side effects, ordering, or pagination. Since an output schema exists, return value details may be covered there, but additional behavioral context (e.g., no mutations, read-only) is absent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. Every word contributes meaning, making it maximally concise for its purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given zero parameters and an existing output schema, the description is largely sufficient. However, it could be slightly more complete by clarifying that it lists all visible pods without filtering (vs. pod_status for a specific pod).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
No parameters exist, and schema coverage is 100% (trivially). The description adds no parameter information because none is needed. Baseline for zero parameters is 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly specifies the verb 'List', the resource 'pod', and the scope: 'every pod the API key can see' with explicit statuses (active, provisioning, stopped). It is distinctive from siblings like pod_status or pod_terminate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool vs. siblings such as pod_status for a specific pod. The description only states what it does, not when it is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pod_quoteA
Get a non-binding price quote + reserved provisioning payload.
Returns a quote_token (TTL=60s) that you pass to pod_create with confirm=True to actually provision. This tool has NO side effects.
The server picks the cheapest matching GPUAvailability row that satisfies the requested disk/vcpu/memory. If none matches, returns an error explaining what's available.
| Name | Required | Description | Default |
|---|---|---|---|
| gpu_type | Yes | GPU type slug, e.g. 'H100_80GB'. | |
| gpu_count | No | Number of GPUs per pod (1, 2, 4, 8). | |
| disk_size_gb | No | Disk size in GB. | |
| vcpus | No | vCPU count. | |
| memory_gb | No | Memory in GB. | |
| image | No | Container image slug. Use 'ubuntu_22_cuda_12' if unsure. | ubuntu_22_cuda_12 |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, description carries full behavioral burden. It discloses no side effects, TTL of 60s, server picks cheapest matching row, and returns error with available options if no match. Comprehensive and honest.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences with front-loaded purpose, then flow, behavior, and error case. No fluff, every sentence adds value. Very concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity and presence of output schema, the description covers essential aspects: return value, TTL, side-effect-free nature, selection logic, and error behavior. Complete for a quoting tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents each parameter well. Description adds overall logic (cheapest matching) but no extra per-parameter meaning beyond schema defaults and examples.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the tool gets a non-binding price quote and reserved provisioning payload, distinguishing it from sibling tools like pod_create. It specifies the verb 'Get' and the resource, making purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Description explains that the tool has no side effects and that the returned quote_token should be passed to pod_create with confirm=True to provision. It implicitly guides usage before creation, but lacks explicit when-not-to-use or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pod_statusA
Get the current status (provisioning / active / failed) for a pod.
With wait_for_ssh=True, blocks (polls every 5s) until ssh_connection is
available — that's when you can SSH in. Returns the SSH connection string
in ssh_connection (e.g. "root@1.2.3.4 -p 22000"). Use it from your Bash
tool: ssh -o StrictHostKeyChecking=no <ssh_connection> "<cmd>".
| Name | Required | Description | Default |
|---|---|---|---|
| pod_id | Yes | The id returned by pod_create. | |
| wait_for_ssh | No | If True, poll until ssh_connection is populated or timeout_s elapses. | |
| timeout_s | No | Max seconds to wait when wait_for_ssh=True. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. It discloses polling behavior every 5s, blocking until SSH available, and return format for SSH connection. It does not cover rate limits or permissions but is sufficient for safe use.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Description is brief (4 sentences), front-loaded with purpose, then explains the optional blocking behavior and SSH usage. Every sentence adds value with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given tool complexity (polling, SSH) and presence of output schema, description adequately explains the blocking behavior and SSH string usage. Lacks details on full return object but output schema covers that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema has 100% description coverage, so baseline is 3. Description adds marginal value by contextualizing SSH connection usage but essentially repeats parameter descriptions found in schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the tool gets pod status with specific statuses, and distinguishes from siblings like pod_create, pod_list, and pod_terminate by focusing on a single pod and offering SSH readiness detection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for checking status and waiting for SSH, but does not explicitly state when to use versus alternatives like pod_list or pod_create, nor provides when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pod_terminateA
Destroy (terminate) a pod. Idempotent on already-deleted pods.
Without confirm=True, returns a no-op preview so you can re-read your decision.
| Name | Required | Description | Default |
|---|---|---|---|
| pod_id | Yes | The pod to destroy. | |
| confirm | No | Required True to actually terminate. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry behavioral disclosure. It mentions idempotency and preview behavior, but it does not disclose potential side effects, required permissions, or data loss risks. While the preview feature adds transparency, the description lacks warnings about irreversibility, making it only partially transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences long, with the first sentence concisely stating purpose and idempotency, and the second explaining the preview feature. No unnecessary words or repetitions; every sentence earns its place, making it highly efficient and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple destructive tool with a clear output schema, the description covers purpose, idempotency, and preview behavior. However, it does not mention prerequisites (e.g., pod existence is handled by idempotency) or any contextual warnings about consequences. Slight gaps in completeness, but overall adequate given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with descriptions for both parameters. The description adds value by clarifying the confirm parameter's preview behavior beyond the schema's 'Required True to actually terminate.' This extra context improves understanding without redundancy, justifying a score above the baseline of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states 'Destroy (terminate) a pod' with a clear verb and resource. It also notes idempotency on already-deleted pods, adding clarity. The tool is uniquely positioned among siblings as the only destroy operation, making its purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the preview behavior with confirm=False, guiding when to preview vs execute. However, it does not provide explicit when-to-use or when-not-to-use guidance, nor does it compare to alternatives like pod_check_runaway. The usage guidelines are implied but not fully elaborated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.1.0- First observed
get_wallet_balance - First observed
list_availability - First observed
list_gpu_types - First observed
pod_check_runaway - First observed
pod_create - First observed
pod_list - First observed
pod_quote - First observed
pod_status - First observed
pod_terminate
TDQS
Scored across 9 tools
Each tool has a unique and clearly distinct purpose, from wallet balance and GPU availability listing to pod creation, quoting, and termination. There is no overlap that could cause an agent to select the wrong tool.
Tool names follow a consistent pattern: utility functions use verb_noun (e.g., get_wallet_balance, list_availability) and pod operations all start with pod_ (e.g., pod_create, pod_terminate). The naming is predictable and easily understood.
With 9 tools covering wallet, GPU types, availability, pod lifecycle (create, list, status, quote, terminate), and runaway monitoring, the count is well-scoped for the server's purpose. Each tool earns its place with no redundancy.
The tool surface covers the full lifecycle of GPU pod management: discovering availability, quoting, creating, monitoring status, listing, terminating, and checking for runaway pods. The inclusion of wallet balance and wait-for-SSH functionality addresses common operational needs.
Maintenance
Related MCP Connectors
On-demand GPU nodes for agents: create nodes, run commands, and submit jobs, billed by the minute.
Your AI Agent's Infrastructure Layer. Connect Claude, Copilot, Codex, or ChatGPT to 200+ managed open source services. Start databases, pipelines, and applications through natural language.
Hosted MCP server connecting claude.ai, ChatGPT and other AI apps to your own computer
Hosted Amazon Seller and Vendor MCP server for Claude, ChatGPT, Cursor, Codex, Gemini, Copilot.
Related MCP Servers
- AlicenseAqualityFmaintenanceA server that allows LLMs to run Claude Code with all permissions bypassed automatically, enabling code execution and file editing without permission interruptions.1588 npm1,314MIT
- AlicenseNot gradedqualityFmaintenanceEnables Claude Code to execute shell commands, Python code, and file transfers on a Google Colab T4 GPU via an MCP server, bridging the GPU gap for AI coding agents.6MIT
- AlicenseAqualityDmaintenanceThis MCP server enables remote control and management of Claude Code agents, allowing you to execute missions, configure agent personalities, and integrate with other MCP tools.720 npm1MIT
- AlicenseAqualityDmaintenanceMCP server that lets any agent or MCP host delegate tasks to Claude Code running headless, with tools for review, validation, analysis, and autonomous work.4151 npm2MIT