Ambient Code Platform MCP Server
The Ambient Code Platform (ACP) MCP Server lets you manage AI coding sessions (AgenticSessions) on an ACP cluster directly from your MCP client.
Session Management
List and filter sessions by status (running/stopped/creating/failed), age, with sorting and limits
Get detailed information about a specific session
Create sessions with a custom prompt, optional repos, model selection, and display name
Create sessions from predefined templates:
triage,bugfix,feature, orexplorationDelete, restart, clone, or update (rename/timeout) sessions
All mutating operations support
dry_runmode for safe previewing
Observability
Retrieve container logs for a session (configurable tail lines)
Get conversation transcripts in JSON or Markdown format
View usage metrics: tokens, duration, and tool calls
Labels
Add or remove key-value labels on sessions for organization
Filter/list sessions by label selectors
Bulk add or remove labels across up to 3 sessions at once
Bulk Operations
Delete, stop, or restart up to 3 sessions at once by name or label selector
All bulk destructive operations require
confirm=trueand supportdry_run
Cluster Management
List configured cluster aliases, check authentication/configuration status, switch between cluster contexts, and authenticate with Bearer tokens
Enables delegation of AI tasks to Kubernetes-hosted Claude agents running on OpenShift, with tools for managing agentic sessions, checking authentication status, listing projects/namespaces, and controlling session lifecycle.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Ambient Code Platform MCP Servercreate a session to analyze my codebase for security vulnerabilities"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
MCP ACP Server
A Model Context Protocol (MCP) server for managing Ambient Code Platform (ACP) sessions via the public-api gateway.
Table of Contents
Related MCP server: K8s MCP Server
Quick Start
# Install
git clone https://github.com/ambient-code/mcp
cd mcp
pip install . # installs the 'mcp-acp' command
# Configure
mkdir -p ~/.config/acp
cat > ~/.config/acp/clusters.yaml <<EOF
clusters:
my-cluster:
server: https://public-api-ambient.apps.your-cluster.example.com
token: your-bearer-token-here
default_project: my-workspace
default_cluster: my-cluster
EOF
chmod 600 ~/.config/acp/clusters.yamlThen add to your MCP client (Claude Desktop, Claude Code, or uvx) and try:
List my ACP sessionsFeatures
Session Management
Tool | Description |
| List/filter sessions by status, age, with sorting and limits |
| Get detailed session information by ID |
| Create sessions with custom prompts, repos, model selection, and timeout |
| Create sessions from predefined templates (triage/bugfix/feature/exploration) |
| Delete sessions with dry-run preview |
| Restart a stopped session |
| Clone an existing session's configuration into a new session |
| Update session metadata (display name, timeout) |
Observability
Tool | Description |
| Retrieve container logs for a session |
| Retrieve conversation history (JSON or Markdown) |
| Get usage statistics (tokens, duration, tool calls) |
Labels
Tool | Description |
| Add labels to a session for organizing and filtering |
| Remove labels from a session by key |
| List sessions matching label selectors |
| Add labels to multiple sessions (max 3) |
| Remove labels from multiple sessions (max 3) |
Bulk Operations
Tool | Description |
| Delete multiple sessions (max 3) with confirmation and dry-run |
| Stop multiple running sessions (max 3) |
| Restart multiple stopped sessions (max 3) |
| Delete sessions matching label selectors (max 3 matches) |
| Stop sessions matching label selectors (max 3 matches) |
| Restart sessions matching label selectors (max 3 matches) |
Cluster Management
Tool | Description |
| List configured cluster aliases |
| Check current configuration and authentication status |
| Switch between configured clusters |
| Authenticate to a cluster with a Bearer token |
Safety Features:
Dry-Run Mode — All mutating operations support
dry_runfor safe preview before executingBulk Operation Limits — Maximum 3 items per bulk operation with confirmation requirement
Label Validation — Labels must be 1-63 alphanumeric characters, dashes, dots, or underscores
Installation
From Source (end users)
git clone https://github.com/ambient-code/mcp
cd mcp
pip install . # installs the 'mcp-acp' commandFrom Wheel
Requires make to be installed.
# "make install" sets up the dev environment (.venv + dependencies),
# which "make build" then uses to produce the wheel in dist/
make install
make build
pip install dist/mcp_acp-*.whlDevelopment Install (contributors)
git clone https://github.com/ambient-code/mcp
cd mcp
uv venv
uv pip install -e ".[dev]"Virtual Environment Install
On some Linux distributions (Debian, Ubuntu, Fedora 38+), PEP 668
prevents pip install into the system Python. If you get an "externally-managed-environment" error,
install into a virtual environment instead:
git clone https://github.com/ambient-code/mcp
cd mcp
python3 -m venv .venv
source .venv/bin/activate
pip install .Note that the mcp-acp command will only be available inside the venv. To use it
from an MCP client like Claude Code, reference the full path:
claude mcp add mcp-acp -t stdio /full/path/to/mcp/.venv/bin/mcp-acpRequirements:
Python 3.12+
Bearer token for the ACP public-api gateway
Access to an ACP cluster
Configuration
Cluster Config
Create ~/.config/acp/clusters.yaml:
clusters:
vteam-stage:
server: https://public-api-ambient.apps.vteam-stage.example.com
token: your-bearer-token-here
description: "V-Team Staging Environment"
default_project: my-workspace
vteam-prod:
server: https://public-api-ambient.apps.vteam-prod.example.com
token: your-bearer-token-here
description: "V-Team Production"
default_project: my-workspace
default_cluster: vteam-stageThen secure the file:
chmod 600 ~/.config/acp/clusters.yamlAuthentication
Add your Bearer token to each cluster entry under the token field, or set the ACP_TOKEN environment variable:
export ACP_TOKEN=your-bearer-token-hereGet your token from the ACP platform administrator or the gateway's authentication endpoint.
Claude Desktop
Edit your configuration file:
macOS:
~/Library/Application Support/Claude/claude_desktop_config.jsonLinux:
~/.config/claude/claude_desktop_config.jsonWindows:
%APPDATA%\Claude\claude_desktop_config.json
{
"mcpServers": {
"acp": {
"command": "mcp-acp",
"args": [],
"env": {
"ACP_CLUSTER_CONFIG": "${HOME}/.config/acp/clusters.yaml"
}
}
}
}After editing, completely quit and restart Claude Desktop (not just close the window).
Claude Code (CLI)
claude mcp add mcp-acp -t stdio mcp-acpUsing uvx
uvx provides zero-install execution — no global Python pollution, auto-caching, and fast startup.
# Install uv (if needed)
curl -LsSf https://astral.sh/uv/install.sh | shClaude Desktop config for uvx:
{
"mcpServers": {
"acp": {
"command": "uvx",
"args": ["mcp-acp"]
}
}
}For a local wheel (before PyPI publish):
{
"mcpServers": {
"acp": {
"command": "uvx",
"args": ["--from", "/full/path/to/dist/mcp_acp-0.3.0-py3-none-any.whl", "mcp-acp"]
}
}
}Usage
Examples
# List sessions
List my ACP sessions
Show running sessions in my-workspace
List sessions older than 7 days in my-workspace
List sessions sorted by creation date, limit 20
# Session details
Get details for ACP session session-name
Show AgenticSession session-name in my-workspace
# Create a session
Create a new ACP session with prompt "Run all unit tests and report results"
# Create from template
Create an ACP session from the bugfix template called "fix-auth-issue"
# Restart / clone
Restart ACP session my-stopped-session
Clone ACP session my-session as "my-session-v2"
# Update session metadata
Update ACP session my-session display name to "Production Test Runner"
# Observability
Show logs for ACP session my-session
Get transcript for ACP session my-session in markdown format
Show metrics for ACP session my-session
# Labels
Label ACP session my-session with env=staging and team=platform
Remove label env from ACP session my-session
List ACP sessions with label team=platform
# Delete with dry-run (safe!)
Delete test-session from my-workspace in dry-run mode
# Actually delete
Delete test-session from my-workspace
# Bulk operations (dry-run first)
Delete these sessions: session-1, session-2, session-3 from my-workspace (dry-run first)
Stop all sessions with label env=test
Restart sessions with label team=platform
# Cluster operations
Check my ACP authentication
List my ACP clusters
Switch to ACP cluster vteam-prod
Login to ACP cluster vteam-stage with tokenTrigger Keywords
Include one of these keywords so your MCP client routes the request to ACP: ACP, ambient, AgenticSession, or use tool names directly (e.g., acp_list_sessions, acp_whoami). Without a keyword, generic phrases like "list sessions" may not trigger the server.
Quick Reference
Task | Command Pattern |
Check auth |
|
List all |
|
Filter status |
|
Filter age |
|
Get details |
|
Create |
|
Create from template |
|
Restart |
|
Clone |
|
Update |
|
View logs |
|
View transcript |
|
View metrics |
|
Add labels |
|
Remove labels |
|
Filter by label |
|
Delete (dry) |
|
Delete (real) |
|
Bulk delete |
|
Bulk by label |
|
List clusters |
|
Login |
|
Tool Reference
For complete API specifications including input schemas, output formats, and behavior details, see API_REFERENCE.md.
Category | Tool | Description |
Session |
| List/filter sessions |
| Get session details | |
| Create session with prompt | |
| Create from template | |
| Delete with dry-run support | |
| Restart stopped session | |
| Clone session configuration | |
| Update display name or timeout | |
Observability |
| Retrieve container logs |
| Get conversation history | |
| Get usage statistics | |
Labels |
| Add labels to session |
| Remove labels by key | |
| Filter sessions by labels | |
| Bulk add labels (max 3) | |
| Bulk remove labels (max 3) | |
Bulk |
| Delete multiple sessions (max 3) |
| Stop multiple sessions (max 3) | |
| Restart multiple sessions (max 3) | |
| Delete by label (max 3) | |
| Stop by label (max 3) | |
| Restart by label (max 3) | |
Cluster |
| List configured clusters |
| Check authentication status | |
| Switch cluster context | |
| Authenticate with Bearer token |
Troubleshooting
"No authentication token available"
Your token is not configured. Either:
Add
token: your-token-hereto your cluster in~/.config/acp/clusters.yamlSet the
ACP_TOKENenvironment variable
"HTTP 401: Unauthorized"
Your token is expired or invalid. Get a new token from the ACP platform administrator.
"HTTP 403: Forbidden"
You don't have permission for this operation. Contact your ACP platform administrator.
"Direct Kubernetes API URLs (port 6443) are not supported"
You're using a direct K8s API URL. Use the public-api gateway URL instead:
Wrong:
https://api.cluster.example.com:6443Correct:
https://public-api-ambient.apps.cluster.example.com
"mcp-acp: command not found"
Add Python user bin to PATH:
macOS:
export PATH="$HOME/Library/Python/3.*/bin:$PATH"Linux:
export PATH="$HOME/.local/bin:$PATH"
Then restart your shell.
MCP Tools Not Showing in Claude
Check Claude Desktop logs: Help → View Logs
Verify config file syntax is valid JSON
Make sure
mcp-acpis in PATHRestart Claude Desktop completely (quit, not just close)
"Permission denied" on clusters.yaml
chmod 600 ~/.config/acp/clusters.yaml
chmod 700 ~/.config/acpArchitecture
MCP SDK — Standard MCP protocol implementation (stdio transport)
httpx — Async HTTP REST client for the public-api gateway
Pydantic — Settings management and input validation
Three-layer design — Server (tool dispatch) → Client (HTTP + validation) → Formatters (output)
See CLAUDE.md for complete system design.
Security
Input Validation — DNS-1123 format validation for all resource names
Gateway URL Enforcement — Direct K8s API URLs (port 6443) rejected
Bearer Token Security — Tokens filtered from logs, sourced from config or environment
Resource Limits — Bulk operations limited to 3 items with confirmation
See SECURITY.md for complete security documentation including threat model and best practices.
Development
# One-time setup
uv venv && uv pip install -e ".[dev]"
# Pre-commit workflow
uv run ruff format . && uv run ruff check . && uv run pytest tests/
# Run with coverage
uv run pytest tests/ --cov=src/mcp_acp --cov-report=html
# Build wheel
uvx --from build pyproject-build --installer uvSee CLAUDE.md for contributing guidelines.
Roadmap
Current implementation provides 26 tools. 3 tools remain planned:
Contributing
Fork the repository
Create a feature branch
Add tests for new functionality
Ensure all tests pass (
uv run pytest tests/)Ensure code quality checks pass (
uv run ruff format . && uv run ruff check .)Submit a pull request
Status
Code: Production-Ready | Tests: All Passing | Security: Input validation, gateway enforcement, token security | Tools: 26 implemented (3 more planned)
Documentation
API_REFERENCE.md — Full API specifications for all 26 tools
SECURITY.md — Security features, threat model, and best practices
CLAUDE.md — System architecture and development guide
License
MIT License — See LICENSE file for details.
Support
For issues and feature requests, use the GitHub issue tracker.
Available Tools
41 toolsacp_add_repoB
Add a repository to a running session.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| session | Yes | Session ID | |
| repo_url | Yes | Repository URL to clone |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, and the description does not disclose any side effects, permissions, or behavior beyond the basic action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Extremely concise single sentence with no superfluous words, well front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations or output schema, the description lacks information on success/failure, idempotency, or what happens if the repository already exists.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so parameters are well-documented in the schema; the description adds no extra meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Add' and the resource 'repository to a running session,' distinguishing it from sibling tools like acp_remove_repo and acp_get_repos_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives, nor any prerequisites or context about what constitutes a 'running session.'
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_bulk_delete_sessionsA
Delete multiple sessions (max 3). DESTRUCTIVE: requires confirm=true. Use dry_run=true first!
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| sessions | Yes | List of session names (max 3) | |
| confirm | No | Required for destructive operations (default: false) | |
| dry_run | No | Preview without executing (default: false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It explicitly labels the operation as DESTRUCTIVE and warns that confirm=true is required, and recommends dry_run for preview. Additional constraints like max 3 sessions are disclosed, though side effects or error behavior are not detailed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no fluff. Key information is front-loaded: purpose, constraints, and usage guidance. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description provides core purpose and safety guidelines. It lacks details on success behavior or error responses, but for a simple bulk delete tool, the coverage is adequate and actionable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, baseline 3. The description adds value by stating 'max 3' for the sessions parameter and explaining the confirm and dry_run parameters' roles (required for destruction; preview without executing). This extends beyond the schema's default values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool deletes multiple sessions (max 3). The verb 'delete' distinguishes it from sibling tools like acp_bulk_stop_sessions or acp_bulk_restart_sessions, which perform different operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description advises using dry_run first and notes that confirm=true is required for destructive operations, providing clear operational context. However, it does not explicitly compare to alternatives like acp_delete_session for single deletions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_bulk_delete_sessions_by_labelA
Delete sessions matching label selectors (max 3 matches). DESTRUCTIVE: requires confirm=true.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| labels | Yes | Labels as key-value pairs (e.g., {"env": "test", "team": "qa"}) | |
| confirm | No | Required for destructive operations (default: false) | |
| dry_run | No | Preview without executing (default: false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses destructive nature and the need for confirm, plus a max of 3 matches. However, it does not detail irreversibility, error handling, or what happens when matches exceed 3.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences, front-loaded with the primary action and constraints, with no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations or output schema, the description is adequate but leaves gaps: no information on return values, behavior when matches exceed 3, or dry_run effects. It covers the key points but lacks completeness for a destructive tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters. The description doesn't add significant meaning beyond what the schema provides, but it contextualizes that 'labels' are used for matching.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Delete sessions matching label selectors', specifying the verb (Delete) and the resource (sessions by label), and distinguishes it from siblings like acp_bulk_delete_sessions which likely deletes all sessions without label filtering.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly notes 'DESTRUCTIVE: requires confirm=true', providing clear context for when to use the tool (i.e., with confirm=true for destructive operations). However, it does not mention alternatives or exclusion conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_bulk_label_resourcesA
Add labels to multiple sessions (max 3). DESTRUCTIVE: requires confirm=true.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| sessions | Yes | List of session names (max 3) | |
| labels | Yes | Labels as key-value pairs (e.g., {"env": "test", "team": "qa"}) | |
| confirm | No | Required for destructive operations (default: false) | |
| dry_run | No | Preview without executing (default: false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It discloses destructive nature and confirm requirement, but omits details like error handling, atomicity of batch operations, or potential side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Extremely concise: two sentences front-load the core action, constraint, and destructive requirement. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple labeling tool with 5 params, nested objects, and no output schema, the description covers purpose, limit, and destructive flag. It lacks detail on return behavior but is reasonably complete for its complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. Description adds 'max 3' (already in schema) and highlights 'DESTRUCTIVE' for confirm parameter, but overall adds limited new semantic value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the action ('Add labels'), resource ('multiple sessions'), and a key constraint ('max 3'). The term 'DESTRUCTIVE' distinguishes it from non-destructive siblings like acp_label_resource.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Mentions the need for confirm=true for destructive operations, providing some usage guidance. However, it does not explicitly tell when to prefer this tool over alternatives like acp_label_resource or acp_bulk_unlabel_resources.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_bulk_restart_sessionsA
Restart multiple stopped sessions (max 3). Requires confirm=true.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| sessions | Yes | List of session names (max 3) | |
| confirm | No | Required for destructive operations (default: false) | |
| dry_run | No | Preview without executing (default: false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It reveals the session limit (max 3) and the confirm requirement, which signals destructiveness. However, it does not disclose side effects, partial failure behavior, or prerequisites (e.g., sessions must be stopped).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences, front-loaded with the core action. Every word serves a purpose. No redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a bulk action with no output schema, the description lacks details on success/failure behavior, partial execution, and required session state. It meets minimal needs but leaves gaps that could mislead an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents each parameter. The description adds value only by restating the confirm requirement and implying the sessions limit, but does not clarify the 'dry_run' or 'project' parameters beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Restart'), the resource ('multiple stopped sessions'), and a constraint ('max 3'). This distinguishes it from singular restart or bulk stop operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies use for multiple stopped sessions by specifying 'multiple' and a limit of 3, but does not explicitly contrast with singular restart (acp_restart_session) or mention that sessions must be stopped. The mention of confirm=true hints at destructiveness but no explicit when-not-to-use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_bulk_restart_sessions_by_labelA
Restart sessions matching label selectors (max 3 matches). Requires confirm=true.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| labels | Yes | Labels as key-value pairs (e.g., {"env": "test", "team": "qa"}) | |
| confirm | No | Required for destructive operations (default: false) | |
| dry_run | No | Preview without executing (default: false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It discloses that the operation is destructive (requires confirm) and has a match limit. However, it does not detail what happens if more than 3 labels match (error or truncation) or any side effects of restarting sessions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise—a single sentence that conveys action, selection method, constraint, and requirement. Every word adds value with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that it's a bulk operation with no output schema, the description covers the basics but lacks information on error handling, behavior when no matches or too many matches, and overall workflow. More context would help the agent decide when to use this tool over alternatives.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already describes all 4 parameters comprehensively (100% coverage). The description adds the 'max 3 matches' context for labels but does not elaborate on the parameters beyond what's in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (restart sessions), the selection method (label selectors), and a key constraint (max 3 matches). It effectively distinguishes from siblings like acp_bulk_restart_sessions (which likely targets by session IDs) through the mention of labels and the limit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance on when the tool is appropriate: when you need to restart sessions by label selectors, and it warns that confirm=true is required. The max 3 matches hint implies that for larger sets, another tool might be better, though no alternative is named directly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_bulk_stop_sessionsA
Stop multiple running sessions (max 3). DESTRUCTIVE: requires confirm=true.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| sessions | Yes | List of session names (max 3) | |
| confirm | No | Required for destructive operations (default: false) | |
| dry_run | No | Preview without executing (default: false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so description is the sole source. It warns about destructiveness and confirm requirement, but does not detail side effects or state changes after stopping.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with no wasted words, front-loaded with the action and key constraint.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Coverage of input parameters is high, and description adds key constraints, but lacks return value information and fails to mention the label-based variant.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, yet description adds 'max 3' for sessions array and 'requires confirm=true' for safety, providing value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool stops multiple running sessions with a maximum of 3, distinguishing it from siblings like acp_bulk_delete_sessions (delete) and acp_bulk_restart_sessions (restart).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It specifies destructive behavior and the need for confirm=true, but does not differentiate from acp_bulk_stop_sessions_by_label or other siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_bulk_stop_sessions_by_labelA
Stop sessions matching label selectors (max 3 matches). DESTRUCTIVE: requires confirm=true.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| labels | Yes | Labels as key-value pairs (e.g., {"env": "test", "team": "qa"}) | |
| confirm | No | Required for destructive operations (default: false) | |
| dry_run | No | Preview without executing (default: false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses destructive behavior and confirm requirement, plus max 3 matches constraint. No annotations, so description carries burden. Missing details on error handling or no-match scenarios.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no fluff. Front-loads action and constraints effectively.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Adequate for a destructive tool with clear constraints. No output schema, so return value is unaddressed, but otherwise complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers all parameters with descriptions (100% coverage). Description does not add substantial meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb 'stop', resource 'sessions', selection method 'by label' with max 3 constraint. Clearly distinguishes from siblings like acp_bulk_stop_sessions and acp_stop_session.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States destructive nature and confirm requirement, implying when to use (with caution). Lacks explicit exclusions or alternatives, but context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_bulk_unlabel_resourcesA
Remove labels from multiple sessions (max 3). DESTRUCTIVE: requires confirm=true.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| sessions | Yes | List of session names (max 3) | |
| label_keys | Yes | List of label keys to remove | |
| confirm | No | Required for destructive operations (default: false) | |
| dry_run | No | Preview without executing (default: false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It discloses the destructive nature and confirm requirement, but lacks details on behavior when confirm is false, dry_run effects, authorization needs, or reversibility.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise (two sentences) and front-loaded with the core action and key condition. It wastes no words, though additional structure could improve readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description does not explain return values or error handling. It adequately covers the main purpose and a critical condition, but lacks completeness for a destructive batch operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema has 100% coverage with descriptions for all parameters. The description adds context about max 3 sessions and destructive nature, but does not significantly enhance understanding beyond the schema's existing parameter descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Remove labels from multiple sessions (max 3)', specifying the action and resource. It distinguishes from siblings like acp_unlabel_resource (single session) and acp_bulk_label_resources (adding labels).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions 'requires confirm=true' as a condition, indicating when the tool is safe to use. However, it does not explicitly state when to use this tool versus alternatives like acp_unlabel_resource or provide guidance on when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_clone_sessionA
Clone an existing session's configuration into a new session.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| source_session | Yes | Session ID to clone from | |
| new_display_name | Yes | Display name for the cloned session | |
| dry_run | No | Preview without executing (default: false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, and the description does not disclose behavioral traits such as whether the source session remains unmodified, the nature of the clone (independent vs linked), or any permission or rate limiting requirements. This lack of transparency is a gap for a mutation-like tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence that is concise and front-loaded with the action. No wasted words. Ideal for quick comprehension.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description lacks information about return values (new session ID?), side effects, and required context (e.g., source session must exist). With no output schema and no annotations, the description is minimal but not entirely inadequate for a straightforward clone operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds no additional meaning beyond the parameter descriptions in the schema. No extra clarification on parameter usage or interdependencies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool clones an existing session's configuration into a new session. The verb 'clone' is specific, and the resource 'session configuration' is distinct from sibling tools like acp_create_session (new from scratch) or acp_create_session_from_template (from template).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives such as acp_create_session or acp_create_session_from_template. The description only states what it does, without clarifying preferred contexts or prerequisites like requiring an existing session.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_create_scheduled_sessionA
Create a scheduled session backed by a Kubernetes CronJob. Requires a cron schedule and a session template.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| schedule | Yes | Cron expression (e.g., '0 2 * * *' for nightly at 2am) | |
| session_template | Yes | Session template (same as create_session: task, model, repos, displayName, etc.) | |
| display_name | No | Human-readable name for this schedule | |
| suspend | No | Create in suspended state (default: false) | |
| dry_run | No | Preview without executing (default: false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full behavioral burden. It discloses the CronJob backend, but lacks details on failure handling, resource creation, default suspend behavior, or permissions required. The `suspend` parameter (with default false) is not mentioned in the description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences: first states purpose, second states prerequisites. No extraneous words. Front-loaded and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema or annotations, the description should compensate with return value or error handling information, but it does not. The schema covers parameter details, but the description omits defaults (suspend, dry_run) and any post-creation behavior, making it only partially complete for a creation tool with nested objects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds marginal value by stating 'Requires a cron schedule and a session template' and referencing `create_session` for the template structure, which aids understanding. This slight addition justifies a 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool creates a scheduled session backed by a Kubernetes CronJob, requiring a cron schedule and session template. It differentiates from siblings like `acp_create_session` (non-scheduled) and `acp_trigger_scheduled_session` (immediate execution) by specifying scheduling behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions prerequisites (cron schedule and session template) but does not guide when to use this tool versus alternatives such as `acp_create_session` for one-off sessions or `acp_update_scheduled_session` for modifying existing schedules. No explicit context on when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_create_sessionB
Create an ACP AgenticSession with a custom prompt. Supports dry-run mode.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| initial_prompt | Yes | The prompt/instructions to send to the session | |
| display_name | No | Human-readable display name | |
| repos | No | Repository URLs to clone | |
| model | No | LLM model to use | claude-sonnet-4 |
| dry_run | No | Preview without executing (default: false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It adds the dry-run mode feature, but does not disclose other behaviors like whether the session starts immediately, authorization requirements, or what happens upon creation. Some transparency but incomplete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with action, no wasted words. Efficiently conveys core purpose and a key feature.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Missing details about return value (no output schema), prerequisites, error states, or side effects. For a creation tool with 6 params and no annotations, more context is needed for an agent to use it effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all parameters. The description mentions 'custom prompt' and 'dry-run' which align with initial_prompt and dry_run, but adds no new meaning beyond reinforcing the schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Create' and resource 'ACP AgenticSession', and adds key feature 'custom prompt' and 'dry-run mode'. Among siblings, it distinguishes from session creation variants but not explicitly enough to differentiate from similar tools like create_session_from_template.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool vs alternatives like create_session_from_template or create_scheduled_session. No context on when not to use it or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_create_session_from_templateB
Create a session from a predefined template (triage/bugfix/feature/exploration). Each template has optimized settings.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| template | Yes | Template name | |
| display_name | Yes | Display name for the session | |
| repos | No | Repository URLs to clone | |
| dry_run | No | Preview without executing (default: false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. Only states creation without disclosing side effects, authentication needs, rate limits, or what happens after creation (e.g., session starts immediately). Lacks detail for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with action and resource, second sentence adds value. No fluff, every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 5 parameters, no output schema, and no annotations, the description is minimal. Does not explain dry_run behavior, template differences, or whether session starts immediately. Leaves significant gaps for agent decision-making.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers 100% of parameters with descriptions. Description adds general context ('optimized settings') but no specific parameter-level value. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states verb (create), resource (session from template), and specific templates (triage/bugfix/feature/exploration). Differentiates from similar sibling tools like acp_create_session by emphasizing templates.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies usage for predefined templates but does not explicitly state when to use over alternatives like acp_create_session. No exclusions or when-not-to-use guidance provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_delete_scheduled_sessionA
Delete a scheduled session and its CronJob. Supports dry-run mode.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| name | Yes | Scheduled session name | |
| dry_run | No | Preview without executing (default: false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description adds value by specifying that the CronJob is also deleted and that dry-run is supported. However, it does not disclose irreversibility, required permissions, or side effects on associated resources.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence that efficiently conveys the core functionality and an important feature (dry-run). No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 3 parameters and no output schema or annotations, the description covers the basic action but lacks context on prerequisites, reversibility, and differentiation from related tools. Adequate but incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with clear parameter descriptions. The description reinforces dry-run mode and the deletion target but adds no new semantic meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the action (delete) and the resources affected (scheduled session and its CronJob). It also mentions dry-run mode, distinguishing it from sibling tools like suspend or resume.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for deletion scenarios and mentions dry-run, but does not explicitly state when to use this vs siblings like acp_suspend_scheduled_session or acp_delete_session. No exclusions or alternatives are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_delete_sessionB
Delete an AgenticSession. Supports dry-run mode.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| session | Yes | Session ID | |
| dry_run | No | Preview without executing (default: false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must fully disclose behavior. It states the tool is destructive ('Delete') and supports dry-run, but omits details like what happens to related resources, required permissions, or whether the action is reversible. The agent lacks sufficient warning for a destructive operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise at two sentences, with the primary action and key feature (dry-run) front-loaded. Every word is essential, and there is no redundant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of an output schema, the description should at least hint at the return value or confirmation of deletion. It does not. The tool is destructive, yet no mention of rollback, confirmation steps, or impact on other sessions is provided, leaving the agent underinformed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage for all three parameters. The description adds only the dry-run mode mention, which is already reflected in the schema's description and default value. Thus, it contributes minimal additional meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Delete' and the resource 'AgenticSession', making it unambiguous. It also mentions dry-run mode, which adds specificity. This distinguishes it from sibling tools like bulk delete variants.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives such as acp_bulk_delete_sessions or acp_bulk_delete_sessions_by_label. There is no mention of prerequisites, context, or when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_export_sessionB
Export session chat history as markdown.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| session | Yes | Session ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full burden. It states the action but does not disclose any behavioral traits such as what happens if the session is missing, whether it returns content or triggers a download, or if there are size limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence with no wasted words. Efficiently conveys the core purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple export tool with two parameters and no output schema, the description is adequate but lacks details about the output format (e.g., how the markdown is returned) and error behavior. Given the lack of annotations, more context would improve agent decision-making.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema has 100% description coverage for both parameters. The tool description adds no extra meaning beyond what the schema already provides, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the action (export), the resource (session chat history), and the output format (markdown). It distinguishes from siblings like 'acp_get_session' or 'acp_get_session_transcript' by specifying markdown export.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives like 'acp_get_session' or 'acp_get_session_transcript'. The description does not provide context about prerequisites or when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_get_repos_statusA
Check repository clone status for a session.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| session | Yes | Session ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden of disclosing behavior. It implies a read-only operation ('Check'), but does not explicitly state whether it has side effects, requires authentication, or any other behavioral traits. This is minimal disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no unnecessary words. It efficiently conveys the core purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no output schema and the description does not explain what the response looks like (e.g., status values). While the tool is simple, the lack of return format information makes it less complete for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the input schema already describes both parameters. The description adds no further semantic information beyond what the schema provides, so it meets the baseline but does not exceed it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Check'), resource ('repository clone status'), and scope ('for a session'), making it unambiguous. It distinguishes well from sibling tools like acp_add_repo and acp_remove_repo, which have different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like acp_get_session or acp_list_sessions. It merely states the function, leaving the agent to infer usage without explicit context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_get_scheduled_sessionB
Get details of a specific scheduled session by name.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| name | Yes | Scheduled session name |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description must convey behavioral traits. It implies a read operation ('Get details'), but does not explicitly state it is read-only, safe, or non-destructive. Missing disclosure of authentication or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no wasted words. However, it lacks any structural elements like bullet points or additional context that could improve scannability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with 2 parameters and no output schema. The description covers the basic purpose but does not describe the return value or provide any additional context that the schema doesn't already supply. It is minimally complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema covers both parameters with descriptions (100% coverage). The tool description adds no extra meaning beyond the schema, so baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('get details') and the resource ('scheduled session'), with the qualifier 'by name' making it distinct from listing or other operations. It effectively differentiates from sibling tools like acp_list_scheduled_sessions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives (e.g., acp_list_scheduled_sessions) or any preconditions. It only says 'by name', which is necessary but insufficient for optimal tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_get_sessionB
Get details of a specific session by ID.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| session | Yes | Session ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so the description carries full burden. It only states 'get details' without disclosing read-only nature, error handling, or access requirements. Minimal behavioral information beyond the verb.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, concise and to the point. No wasted words, but could benefit from a brief expansion on what 'details' includes without harming brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Adequate for a simple get-by-ID tool, but missing output details, error behavior, and usage context. Without output schema, the agent lacks information on expected return structure.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with clear parameter descriptions. The description adds 'by ID' but does not provide additional meaning beyond what the schema already indicates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states 'Get details of a specific session by ID,' specifying the verb (get), resource (session details), and method (by ID). It clearly distinguishes from siblings like acp_list_sessions and acp_create_session.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives (e.g., acp_list_sessions). Lacks context for prerequisites or decision criteria, leaving the agent to infer usage from just the operation name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_get_session_logsB
Retrieve container logs for a session. Useful for debugging.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| session | Yes | Session ID | |
| container | No | Container name (optional) | |
| tail_lines | No | Number of log lines (default: 1000, max: 10000) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description should disclose behavior. Merely states it retrieves logs, a read operation, but omits any details about side effects, rate limits, or output size. Bare minimum.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no redundant words. Efficiently conveys core purpose and typical use case.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema, yet description does not hint at return format (e.g., raw text or structured lines). Lacks details on important aspects like log line format or ability to filter by container. Incomplete for a 4-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema has 100% description coverage for all 4 parameters. Description adds no additional meaning beyond what schema provides. Baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description uses specific verb 'Retrieve' with precise object 'container logs for a session'. Clearly distinguishes from sibling tools like acp_get_session or acp_get_session_metrics, which do not retrieve logs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States 'Useful for debugging' implying when to use, but provides no guidance on when not to use or alternatives among sibling tools. Lacks explicit context for selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_get_session_metricsA
Get usage statistics for a session (tokens, duration, tool calls).
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| session | Yes | Session ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must convey behavioral traits. It implies a read operation and lists output fields, but does not disclose whether the tool is safe (idempotent), requires specific permissions, or has any limitations. The information is adequate but minimal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence with clear, front-loaded verb 'Get'. Every word serves a purpose, no waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description lists the key metrics (tokens, duration, tool calls) but lacks details on output format, pagination, or error states. Given the tool's simplicity, it is sufficient but not thorough.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema covers 100% of parameters, so the description need not add much. It adds no extra meaning beyond the schema, which is acceptable. The baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: to get usage statistics for a session, explicitly listing the statistics (tokens, duration, tool calls). This verb+resource+detail approach strongly distinguishes it from siblings like acp_get_session or acp_get_session_logs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. Given many sibling session tools, a note about when to query metrics vs. other session data would improve selection accuracy.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_get_session_transcriptA
Retrieve conversation history for a session in JSON or Markdown format.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| session | Yes | Session ID | |
| format | No | Output format | json |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description must cover behavioral traits. It only implies a read operation but does not disclose idempotency, error handling, or prerequisites (e.g., authentication). Lacks depth for a read tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, front-loaded with verb and resource, no redundant words. Efficient and clear.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description is minimal. It does not address pagination, return structure, or limits. Functional but could be more helpful for an agent to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with descriptions for each parameter. Description adds the concept of 'conversation history' which is not in schema, but otherwise does not enhance parameter understanding beyond what schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the action (Retrieve), resource (conversation history for a session), and output formats (JSON or Markdown). Distinguishes from siblings like acp_get_session_logs and acp_export_session.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool vs alternatives like acp_get_session_logs or acp_export_session. The name and description imply it's for transcripts, but no exclusions or context provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_get_workflow_metadataB
Get workflow metadata (steps, configuration) for a session.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| session | Yes | Session ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description bears full responsibility for behavioral disclosure. It only states the action is a read operation ('Get'), but does not mention permissions, side effects, rate limits, or other behavioral traits. This is insufficient for safe agent invocation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the core purpose. It is appropriately sized for a simple tool, though it could be slightly expanded to mention the return format without losing conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is adequate given the straightforward nature of the tool and full schema coverage. However, it lacks information about the output format or structure of the returned metadata, which would be helpful since there is no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds 'for a session', which mirrors the schema's session parameter description. It does not clarify the optional project parameter beyond the schema, nor does it explain how the parameters map to the metadata retrieval.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool retrieves workflow metadata (steps, configuration) for a session. The verb 'Get' and specific resource 'workflow metadata' differentiate it from siblings like acp_get_session (session info) and acp_set_workflow (sets workflow).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives. The description does not mention prerequisites, typical use cases, or when to prefer it over acp_get_session or acp_set_workflow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_label_resourceC
Add labels to a session. Labels are key-value pairs for organizing and filtering.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| name | Yes | Session name | |
| resource_type | No | Resource type | agenticsession |
| labels | Yes | Labels as key-value pairs (e.g., {"env": "test", "team": "qa"}) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description implies a mutation ('Add labels') but does not clarify whether labels are appended or replaced, nor does it disclose any side effects, permissions, or other behavioral traits. With no annotations provided, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise at two sentences, with the key action front-loaded. Every word contributes to the purpose, though it could be slightly more structured with headings or bullet points.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of the input (nested object with labels, multiple parameters) and lack of output schema, the description is insufficient. It does not explain return values, error conditions, or provide context about the labels lifecycle, making it hard for an agent to use correctly without additional knowledge.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema covers all parameters with descriptions (100% coverage). The description adds only that labels are 'key-value pairs', which is already implicit in the schema. Thus, it meets the baseline but provides no extra semantic value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Add labels') and resource ('session'), making the primary purpose unambiguous. However, it does not differentiate from sibling tools like acp_bulk_label_resources or acp_unlabel_resource, which could cause confusion in selection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool over alternatives (e.g., bulk labeling or unlableling). There is no mention of prerequisites (e.g., session must exist) or typical use cases, leaving the agent without decision-making context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_list_clustersA
List configured cluster aliases from clusters.yaml.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided. Description does not disclose behavioral traits like read-only nature, auth requirements, or effect on system. Simply states the action without safety or side-effect context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence with no filler. Directly communicates purpose and data source. Efficient and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, no-output-schema tool, the description fully specifies what the tool does and where the data comes from. No additional context needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
No parameters in schema, so description need not elaborate. However, it adds the source file (clusters.yaml), providing context beyond the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it lists configured cluster aliases from a specific source file (clusters.yaml). Uses a specific verb+resource combination, distinguishing it from sibling tools that manage sessions, repos, etc.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. Does not specify prerequisites, exclusions, or typical use cases. The agent must infer from the name and sibling context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_list_scheduled_session_runsB
List past runs (AgenticSessions) created by a scheduled session.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| name | Yes | Scheduled session name |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided; the description only indicates a read operation ('list past runs') but lacks details on pagination, ordering, error handling, or authentication needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no wasted words; highly concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description fails to hint at return structure or pagination, leaving agents underinformed about what the tool produces.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, but the description adds minimal context beyond the schema (e.g., 'created by a scheduled session'). Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'list' and the resource 'past runs (AgenticSessions) created by a scheduled session', differentiating from sibling tools like acp_list_sessions and acp_list_scheduled_sessions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives, nor any prerequisites or exclusions mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_list_scheduled_sessionsA
List all scheduled sessions (cron-based recurring sessions) in a project.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must fully disclose behavioral traits. It clarifies the resource type (cron-based recurring sessions) and scope (in a project), but does not mention whether the list is paginated, what fields are returned, or any potential side effects. For a read operation, this is minimal disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, concise and front-loaded. Every word adds value, stating the action, resource type, and scope. No wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one optional parameter, no output schema), the description is mostly complete. It identifies the resource and scope, but lacks details on the return format or default project behavior. However, the context of sibling tools and schema coverage compensates, making it adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% (the single 'project' parameter has a description). The description reinforces that the list is 'in a project', but adds no additional meaning beyond what the schema provides. With high coverage, a baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'List' and resource 'scheduled sessions', clarifying they are 'cron-based recurring sessions'. This clearly distinguishes it from sibling tools like acp_list_sessions (regular sessions) and acp_list_scheduled_session_runs (runs of a scheduled session).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: list all scheduled sessions in a project. However, it does not explicitly state when to use this tool versus alternatives (e.g., for a specific session use acp_get_scheduled_session) or provide any 'when not to use' guidance. The context is inferred but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_list_sessionsB
List and filter AgenticSessions in a project. Filter by status (running/stopped/failed), age. Sort and limit results.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| status | No | Filter by status | |
| older_than | No | Filter by age (e.g., '7d', '24h', '30m') | |
| sort_by | No | Sort field | |
| limit | No | Maximum number of results |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description must fully convey behavioral traits. It only states 'list and filter' without confirming read-only nature, authentication requirements, or any side effects. The description lacks critical behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence covering key capabilities. Efficient but could be improved by explicitly listing filter parameters. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema, so description should hint at return structure (e.g., list of session IDs or details). It does not specify what is returned, leaving the agent without full context. Score penalized for missing output information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Parameter schema coverage is 100% with descriptions for all 5 parameters. The description reiterates filtering by status and age, but adds minimal value beyond the schema. Slight mismatch: description mentions 'failed' but schema uses 'failed' as enum value (consistent). Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the tool lists and filters AgenticSessions in a project, specifying filter criteria (status, age) and sorting/limiting. This distinguishes it from sibling tools like acp_list_scheduled_sessions and acp_list_sessions_by_label.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives like acp_list_sessions_by_label or acp_get_session. The filtering capabilities are implied but not contrasted with other list tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_list_sessions_by_labelC
List sessions matching label selectors.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| labels | Yes | Labels as key-value pairs (e.g., {"env": "test", "team": "qa"}) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full responsibility for behavioral disclosure. It states only 'list' but does not confirm read-only behavior, required permissions, output format, or any side effects. This is insufficient for a tool with no annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, which is concise and front-loaded. It contains no filler. However, it could be slightly expanded to include matching logic without losing conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple (list with label filter) and the schema covers parameters, but the description lacks details on output, pagination, or label matching semantics (e.g., conjunction vs disjunction). With no output schema, more context about return values would be helpful.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description does not add any additional meaning beyond what the schema provides. The parameter descriptions in the schema are adequate, so no deduction.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (list) and resource (sessions) and distinguishes it from siblings like acp_list_sessions (which likely lists all sessions) by specifying the label-based filtering. However, it could be more explicit about the matching logic (e.g., all labels must match).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives such as acp_list_sessions or the bulk label-based tools. The description does not mention use cases or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_loginA
Authenticate to a cluster with a Bearer token. Sets the token in memory and verifies it works.
| Name | Required | Description | Default |
|---|---|---|---|
| cluster | Yes | Cluster alias name | |
| token | No | Bearer token for authentication |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must disclose behavior. It states 'Sets the token in memory and verifies it works,' indicating statefulness and verification, but lacks details on side effects, failure modes, or scope of the token.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences, no wasted words, immediate identification of purpose and action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema, yet description omits what the tool returns (e.g., success/failure, token info). Missing critical context for an authentication tool, such as error handling or next steps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and description echoes the schema's parameter meanings (Bearer token, cluster alias). No additional insight beyond schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'Authenticate' and identifies the resource 'cluster' with a Bearer token, clearly distinguishing it from sibling tools like session management or repo operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for initial authentication but does not explicitly state when to use it or alternatives like acp_switch_cluster. No guidance on prerequisites or when not to use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_remove_repoB
Remove a repository from a session.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| session | Yes | Session ID | |
| repo_name | Yes | Repository name to remove |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided and description is too brief, omitting behavioral details like permanence, permissions, or side effects of removal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, no fluff, appropriately concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Lacks details for a mutation tool with no output schema; does not explain return values, reversibility, or effects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description covers 100% of parameters; description does not add extra semantics beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action 'Remove' the resource 'repository' from a 'session', distinguishing it from sibling 'acp_add_repo'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool vs alternatives; no prerequisites or conditions mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_restart_sessionA
Restart a stopped session. Supports dry-run mode.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| session | Yes | Session ID | |
| dry_run | No | Preview without executing (default: false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It mentions dry-run mode but fails to disclose other behavioral traits like side effects, required permissions, or state changes upon restart.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Extremely concise: one sentence with essential information. No filler, front-loaded with the core action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity and lack of output schema, the description is adequate but not comprehensive. It doesn't address prerequisites or what constitutes a valid stopped session, leaving gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the description adds value by highlighting dry-run support beyond the schema's 'Preview without executing'. This enhances understanding for the agent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the tool restarts a stopped session, with a specific verb and resource. This distinguishes it from sibling tools like acp_clone_session or acp_delete_session.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives such as acp_bulk_restart_sessions or acp_create_session. The description only states what it does, not when to choose it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_resume_scheduled_sessionA
Resume a suspended scheduled session. The CronJob will start creating sessions again.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| name | Yes | Scheduled session name |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden but only states the basic effect. It lacks details on permissions, idempotency, or state assumptions, but is minimally adequate for a simple resume operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences, front-loaded with purpose. Every word contributes value with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with no output schema and no annotations, the description is minimally complete but could explain what 'suspended' means or that the session must be in suspended state. It covers the main function.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with clear parameter descriptions. The tool description adds no extra meaning beyond the schema, meeting the baseline expectation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action 'Resume' and the resource 'suspended scheduled session', and explains the effect on the CronJob. It distinguishes well from siblings like suspend.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no explicit guidance on when to use this tool versus alternatives, no prerequisites, and no exclusion conditions. It only implies usage for suspended sessions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_set_workflowB
Set the active workflow on a running session.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| session | Yes | Session ID | |
| workflow | Yes | Workflow name to activate |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, and the description only states the basic operation without disclosing side effects, state requirements, or potential failures.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, no extraneous information—efficient and to the point.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Adequate but minimal. Lacks explanation of what 'active workflow' means, constraints (session must be running), and no output schema provided.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so descriptions already explain each parameter. The description adds no extra meaning beyond schema definitions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (set) and resource (active workflow on a running session), distinguishing it from sibling tools like creation or deletion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool vs alternatives, nor any prerequisites or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_suspend_scheduled_sessionA
Suspend (pause) a scheduled session. The CronJob will stop creating new sessions.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| name | Yes | Scheduled session name |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must fully disclose behavior. It states that the CronJob stops creating new sessions, which is a key effect. However, it lacks details on reversibility (e.g., can it be resumed?), side effects on existing sessions, permission requirements, or idempotency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences with no extraneous information. It efficiently conveys the action and its effect.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple, schema coverage is complete, and no output schema exists. The description covers the primary effect. However, it could mention that the suspension is reversible via the sibling acp_resume_scheduled_session tool, which would enhance completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, providing clear definitions for 'project' (optional) and 'name' (required). The tool description does not add any additional meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Suspend (pause)' and the resource 'scheduled session', and explains the effect: 'The CronJob will stop creating new sessions.' This distinguishes it from sibling tools like acp_delete_scheduled_session (deletion) and acp_resume_scheduled_session (resume).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when wanting to temporarily stop new sessions from a schedule, but does not explicitly state when to use it vs alternatives, nor does it mention prerequisites or when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_switch_clusterB
Switch to a different cluster context.
| Name | Required | Description | Default |
|---|---|---|---|
| cluster | Yes | Cluster alias name |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided; the description only states the action without disclosing side effects, reversibility, authentication requirements, or impact on subsequent operations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, no extraneous words, front-loaded with the action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter tool, the description is minimally adequate but lacks context such as the need to first list clusters or that it affects session-related commands.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema has 100% coverage for the single parameter 'cluster', described as 'Cluster alias name'. The description adds no additional meaning beyond the schema, earning a baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Switch to a different cluster context' uses a specific verb (switch) and resource (cluster context), clearly distinguishing it from sibling tools like acp_list_clusters which only list clusters.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool vs alternatives (e.g., after listing clusters with acp_list_clusters), no prerequisites or exclusions provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_trigger_scheduled_sessionB
Manually trigger a scheduled session to run immediately, regardless of its cron schedule.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| name | Yes | Scheduled session name |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, and the description only states the action without disclosing side effects (e.g., effect on original schedule, permissions needed, whether it creates a new run).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, no redundant words, front-loaded with the core action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Minimally adequate given the simple action and lack of annotations, but missing behavioral context and return information would help an agent decide to use it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and descriptions are clear; the description adds no extra meaning beyond the schema for the two parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'trigger' and resource 'scheduled session', and the phrase 'regardless of its cron schedule' distinguishes it from scheduling operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives like resume or create, nor any prerequisites (e.g., session must exist and be active).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_unlabel_resourceB
Remove labels from a session by key.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| name | Yes | Session name | |
| resource_type | No | Resource type | agenticsession |
| label_keys | Yes | List of label keys to remove |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It does not disclose side effects (e.g., idempotency, behavior on missing labels, required permissions). Only the basic operation is described.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, no fluff, front-loaded. Could be slightly improved with additional context, but it is appropriately concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite low complexity, the description fails to explain behavioral details like error conditions, operation atomicity, or return values. An output schema is missing, so description should compensate, but does not.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema describes all parameters with 100% coverage. The description adds 'by key', but that is already implied by the 'label_keys' parameter. No meaningful extra semantics beyond schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (remove), resource (labels from a session), and method (by key). It distinguishes from siblings like 'acp_label_resource' and 'acp_bulk_unlabel_resources'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives, no prerequisites, and no contextual hints. The description only states what it does, not when to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_update_scheduled_sessionC
Update a scheduled session (schedule, template, display name, or suspend state).
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| name | Yes | Scheduled session name | |
| schedule | No | Cron expression (e.g., '0 2 * * *' for nightly at 2am) | |
| display_name | No | New display name | |
| session_template | No | Session template (same as create_session: task, model, repos, displayName, etc.) | |
| suspend | No | Suspend or unsuspend | |
| dry_run | No | Preview without executing (default: false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must disclose behavioral traits. It only states 'Update' without explaining side effects, idempotency, permissions, or the behavior of the dry_run parameter defined in the schema. This is insufficient for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single 11-word sentence with no fluff or repetition. It efficiently conveys the core purpose and the set of updatable attributes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of the input schema (nested session_template object) and many sibling tools for scheduled sessions, the description does not explain differences from suspend/resume tools or provide guidance on nested parameters. The tool feels under-described for its actual scope.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description summarizes the key parameter groups (schedule, template, display name, suspend) but does not add significant meaning beyond the schema descriptions. It provides a quick overview but no additional semantic detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states 'Update a scheduled session' and lists updatable fields (schedule, template, display name, suspend state), clearly specifying the verb and resource. It distinguishes from create, delete, and other sibling tools, but does not explicitly scope the operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives like acp_suspend_scheduled_session or acp_resume_scheduled_session. The description does not mention when not to use it or provide context for selecting this tool over siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_update_sessionC
Update session metadata (display name, timeout). Supports dry-run mode.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project/namespace name (uses default if not provided) | |
| session | Yes | Session ID | |
| display_name | No | New display name | |
| timeout | No | New timeout in seconds | |
| dry_run | No | Preview without executing (default: false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must disclose behavioral traits. It mentions dry-run mode, which is helpful, but omits important details like whether the operation is destructive, permission requirements, or side effects on other metadata.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise—two sentences covering purpose and the dry-run capability. Every sentence is necessary and front-loaded, with no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the 5 parameters, no output schema, and no annotations, the description is insufficient. It does not explain return values, error states, prerequisites (e.g., session must exist), or how optional fields behave when omitted.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds minimal value beyond listing 'display name' and 'timeout' from the schema, and it notes 'dry-run' but does not elaborate on parameter usage or constraints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (update) and resource (session metadata) with specifics (display name, timeout). It distinguishes from sibling tools like acp_delete_session or acp_create_session, though it does not explicitly contrast them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives (e.g., when to create vs update). The context implies an existing session, but lacks explicit usage context or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
acp_whoamiA
Get current configuration and authentication status.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full responsibility for behavioral disclosure. It indicates a read operation ('Get') but does not clarify what happens if the user is not authenticated or what specific information is returned. For a tool that potentially reveals sensitive state, more transparency (e.g., 'Returns user info if authenticated, otherwise error') is expected.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence of 6 words, conveying the essential purpose with zero redundancy. Every word contributes meaning, and the structure is front-loaded with the action verb.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given zero parameters and no output schema, the description is minimal. It covers the basic purpose but lacks details about the return value format or authentication requirements. For a simple status-check tool, this is adequate but not thorough, especially compared to siblings that may have richer descriptions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters with 100% coverage, so the schema already fully documents the absence of required inputs. The description adds nothing about parameters, which is acceptable since there are none. The high schema coverage justifies the baseline score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'Get' and clearly names the resources: 'current configuration and authentication status'. This distinguishes it from sibling tools like acp_login (which performs login) and acp_switch_cluster (which changes cluster). The purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context (retrieving current state) but provides no explicit guidance on when to use this tool versus alternatives, nor does it mention prerequisites such as being logged in. It leaves the agent to infer appropriate context from the tool name and behavior.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
15 tool updates
v0.4.0- Added
acp_add_repo - Added
acp_create_scheduled_session - Added
acp_delete_scheduled_session - Added
acp_export_session - Added
acp_get_repos_status - Added
acp_get_scheduled_session - Added
acp_get_workflow_metadata - Added
acp_list_scheduled_session_runs - Added
acp_list_scheduled_sessions - Added
acp_remove_repo - Added
acp_resume_scheduled_session - Added
acp_set_workflow - Added
acp_suspend_scheduled_session - Added
acp_trigger_scheduled_session - Added
acp_update_scheduled_session
26 tool updates
v0.3.0- First observed
acp_bulk_delete_sessions - First observed
acp_bulk_delete_sessions_by_label - First observed
acp_bulk_label_resources - First observed
acp_bulk_restart_sessions - First observed
acp_bulk_restart_sessions_by_label - First observed
acp_bulk_stop_sessions - First observed
acp_bulk_stop_sessions_by_label - First observed
acp_bulk_unlabel_resources - First observed
acp_clone_session - First observed
acp_create_session - First observed
acp_create_session_from_template - First observed
acp_delete_session - First observed
acp_get_session - First observed
acp_get_session_logs - First observed
acp_get_session_metrics - First observed
acp_get_session_transcript - First observed
acp_label_resource - First observed
acp_list_clusters - First observed
acp_list_sessions - First observed
acp_list_sessions_by_label - First observed
acp_login - First observed
acp_restart_session - First observed
acp_switch_cluster - First observed
acp_unlabel_resource - First observed
acp_update_session - First observed
acp_whoami
TDQS
Scored across 41 tools
Most tools have clearly distinct purposes due to consistent verb_noun naming. However, the large number of bulk operations (e.g., acp_bulk_* sessions) and similar tools like acp_clone_session vs acp_create_session_from_template could cause confusion if descriptions are not carefully read.
All tools follow the acp_verb_noun pattern in snake_case, which is highly consistent and predictable. The naming convention makes it easy to infer the action and target resource.
With 41 tools, the set is significantly larger than the recommended well-scoped range (3-15). Even for a comprehensive platform, this many tools risk overwhelming agents and users, and many operations could potentially be consolidated.
The tool set covers the full lifecycle of sessions (CRUD, clone, logs, metrics, transcript), scheduled sessions, repos, labels, clusters, authentication, and workflows. There are no obvious missing operations for the declared domain of managing an Ambient Code Platform.
Maintenance
Related MCP Connectors
Scoped agent execution. Server-side credentials, policy, budgets and verifiable receipts.
Fail-closed policy guardrails for AI agents running kubectl, terraform, helm, and argocd.
Your AI Agent's Infrastructure Layer. Connect Claude, Copilot, Codex, or ChatGPT to 200+ managed open source services. Start databases, pipelines, and applications through natural language.
On-demand GPU nodes for agents: create nodes, run commands, and submit jobs, billed by the minute.
Related MCP Servers
- FlicenseNot gradedqualityCmaintenanceEnables delegation of agentic sessions to Kubernetes-hosted Claude agents running on the Ambient Code Platform. Supports creating, managing, and communicating with remote AI agent sessions through OpenShift authentication.-
- AlicenseNot gradedqualityDmaintenanceEnables LLMs like Claude to securely execute Kubernetes CLI tools (kubectl, helm, istioctl, argocd) across multiple clusters through dynamic kubeconfig support, allowing natural language Kubernetes management and operations.5MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI assistants to create isolated Kubernetes Pod sessions for real-time shell command execution, file management, and Node.js code execution, with automatic TTL-based cleanup.MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI assistants to manage remote SSPCloud/Onyxia compute workspaces: push repositories, run stateful Python/bash via Jupyter kernels, switch GPU slots, expose services via HTTPS, and retrieve artifacts.MIT