kimi-swarm-bridge
This server provides an MCP bridge that lets you delegate tasks to Kimi Code, including native AgentSwarm parallel research/coding, manage sessions, recover long-running jobs, and retrieve handoffs with Git and swarm evidence.
Delegate tasks: start async (
kimi_delegate_task) or wait for completion (kimi_delegate_and_wait) with plan, acceptance criteria, depth, and optional swarm mode.Continue sessions: submit follow-up instructions to an existing Kimi session (
kimi_continue_task).Wait and monitor: poll a session until idle/failed/aborted with timeout (
kimi_wait_until_idle); check bridge health (kimi_bridge_status).Retrieve results: get handoff with changed files, committed/working-tree changes, and swarm evidence (
kimi_get_handoff); build review packages (kimi_review_package); read per-file diffs (kimi_get_diff).Abort: stop an in-flight session (
kimi_abort).Recover jobs: list durable jobs (
kimi_recent_jobs) or raw sessions (kimi_recent_sessions,kimi_find_recent_session) after timeouts/disconnects to avoid resubmitting.Configure swarm: set max workers and offer preference (
kimi_swarm_settings); choose coordinator/worker models (kimi_model_settings).
Reports Git repository state for delegation handoffs, including committed changes, working-tree changes, and pre-existing dirty paths.
Kimi AgentSwarm Bridge for ai&
An open-source MCP bridge for running Kimi Code — including native AgentSwarm — with inference provided through ai&.
This fork is built around one runtime policy:
inference always goes through
https://api.aiand.com/v1credentials are supplied with
AIAND_API_KEYthe model is configurable with
KIMI_MODEL_NAMEthe default model is
zai-org/glm-5.3(moonshotai/kimi-k3is also tested); users can switch models from chatKimi Code's REST API stays loopback-only
the externally exposed interface is MCP
The project began as a fork of ximenchuifeng/codex-kimi-bridge and remains available under the MIT License.
Organization deployment
To give every employee Kimi Swarm in Claude — per-employee isolated workspaces, sign-in through your identity provider, file upload/download, and persistence — deploy the Cloudflare edition: see docs/cloudflare-deploy.md. Share docs/using-kimi-swarm.md with employees. Also share the Kimi Swarm skill (kimi-swarm.zip on each release; Team and Enterprise owners can add it for everyone). With it, Claude offers to hand big, independent parts of a request to Kimi without being asked.
Related MCP server: firepass-mcp
Benchmarks
Web research briefs, with Kimi Swarm (GLM-5.3 on ai&) and Claude (Claude Code, Opus) given the same brief at the same time. Measured 2026-09-25 on the Cloudflare deployment. The number of agents is Kimi's own choice, up to the cap.
Brief | Kimi agents | Kimi time | Claude time | Empty cells ³ (Kimi / Claude) | Cost (Kimi / Claude) |
12 platforms, 7 fields | 4 | 3m49s | 1m00s | — | $0.78 / — |
12 platforms, 7 fields | 12 | 2m05s | 1m39s ¹ | — | $0.92 / — |
30 managed Postgres providers | 10 | 3m46s | 1m31s | 11 / 13 | — / — |
30 email APIs, with cost calculations | — | 4m46s | 1m59s | 7 / 45 | $1.42 / — |
30 vector databases, every cell required | 15 | 4m26s | 3m56s | 17 / 14 | $2.30 / ~$3–5 ² |
Same brief, after the v0.5.0 speed changes | 15 | 4m24s | 3m56s | 7 / 14 | $1.63 / ~$3–5 ² |
¹ Separate run of the same brief; Claude's simultaneous rerun reused its earlier work (28s), so it isn't a fair comparison. ² Measured from the account's usage before and after. That session carried a long context, which raises Claude's cost. ³ Table cells marked n/d, not documented, not published, not found or unknown, counted the same way in both reports. Accuracy was not graded against a reference.
What the numbers show:
Short briefs: Claude is faster. Kimi has a fixed overhead of about 2 minutes (starting the swarm and merging results) that small jobs can't hide.
Completeness: in normal runs Kimi left far fewer cells empty (7 vs 45 on the email brief, 11 vs 13 on Postgres), because each worker keeps searching its own items. When both were told to fill every cell, the first run finished about level (17 vs 14 empty cells out of 150); after the v0.5.0 changes Kimi left 7.
Large, complete briefs: the speed gap closes. Claude's extra checks ran one after another while Kimi's ran across 15 agents in parallel: 4m24s vs 3m56s, with Kimi costing about half ($1.63).
Background work: Kimi runs in its own sandbox, so Claude stays free for other work while a swarm runs.
Scaling beyond these tests: The agent cap goes up to 128 and can be raised live with no restart (POST /admin/sandboxes/<id>/limits); parallelism is set by SWARM_CONCURRENCY (20 here), which takes effect when the container restarts. The limit in practice is ai&'s per-organization limit on requests in flight at once, shared by every key in the org (100 on a new account, 1000 after the first payment at the time of writing). The Worker keeps the whole deployment under AIAND_CONCURRENCY_LIMIT so a busy organization queues briefly instead of hitting errors; a 20-worker run peaked at about 70 requests in flight. We expect Kimi to pull ahead on longer, wider jobs, but that is a projection, not yet measured.
Moonshot's own results for Agent Swarm (Kimi K2.5, not these tests) point the same way. In wide-search tasks, the swarm needed 3–4.5× fewer critical steps than a single Kimi agent, which Moonshot reports as up to 4.5× less wall-clock time. It also scored higher than Claude Opus 4.5 on BrowseComp and WideSearch, which measure accuracy rather than speed (Kimi K2.5 tech blog). Those runs used Moonshot's model and harness, with up to 100 sub-agents; this bridge defaults to GLM-5.3 and a cap of 20.
What it provides
The bridge exposes Kimi Code through MCP with support for:
task delegation
synchronous delegate-and-wait workflows
continuation of existing Kimi sessions
waiting and polling
handoff retrieval
cancellation
review packages
recent-session discovery
native Kimi Code
AgentSwarmauthenticated Streamable HTTP MCP
persistent Kimi and bridge state under
/data
The hosted container runs:
MCP client
|
v
Streamable HTTP MCP :3000
|
v
Kimi bridge
|
v
Kimi Code server 127.0.0.1:58627
|
v
ai& https://api.aiand.com/v1
|
v
selected ai& modelKimi's administrative REST API is intentionally not published outside the container.
Requirements
For development:
Git
Node.js 22.19 or newer
pnpm 10.x
The Docker image pins:
Node.js 22.19
@moonshot-ai/kimi-code@0.42.0pnpm 10.34.5
ai& configuration
AIAND_API_KEY is required by the official runtime.
The following inference settings are enforced:
Provider protocol: OpenAI-compatible
Base URL: https://api.aiand.com/v1
Credential source: AIAND_API_KEYThe model remains configurable:
export KIMI_MODEL_NAME="moonshotai/kimi-k3"If KIMI_MODEL_NAME is not set, the bridge defaults to:
zai-org/glm-5.3Other models exposed by ai& may work, but native AgentSwarm compatibility should be verified per model. zai-org/glm-5.3 is the default and moonshotai/kimi-k3 is also tested.
Native AgentSwarm
Swarm execution uses Kimi Code's native session profile and native AgentSwarm tool.
The bridge does not implement a custom swarm layer.
For a swarm task it:
creates or uses a Kimi session
updates the session profile with
swarm_mode=trueverifies swarm activation through Kimi status
submits the prompt
lets Kimi invoke its native
AgentSwarmtoolwaits for the coordinator to synthesize worker results
The Docker image defaults to 4 concurrent workers (KIMI_CODE_AGENT_SWARM_MAX_CONCURRENCY=4). The Cloudflare edition defaults to 20 (SWARM_CONCURRENCY) and lets Kimi choose how many agents to use, up to a ceiling of 20 that an admin can change live.
Docker
Build:
docker build -t kimi-swarm-bridge .Generate an MCP bearer token:
export KIMI_MCP_AUTH_TOKEN="$(openssl rand -hex 32)"Set your ai& API key in the environment:
export AIAND_API_KEY="..."Run:
docker run -d \
--name kimi-swarm-bridge \
-p 127.0.0.1:3000:3000 \
-e AIAND_API_KEY \
-e KIMI_MCP_AUTH_TOKEN \
-v kimi-swarm-data:/data \
kimi-swarm-bridgeOnly MCP port 3000 should be published. Kimi remains on loopback inside the container.
HTTP MCP
The Streamable HTTP endpoint is:
POST /mcp
GET /mcp
DELETE /mcpAuthentication:
Authorization: Bearer <KIMI_MCP_AUTH_TOKEN>Health endpoints:
GET /healthz
GET /ping/ping exists for managed-host health checks.
TLS is expected to be terminated by the hosting platform or reverse proxy.
Runtime environment variables
Required:
AIAND_API_KEY
KIMI_MCP_AUTH_TOKEN when HTTP transport is usedCommon optional settings:
KIMI_MODEL_NAME
KIMI_CODE_AGENT_SWARM_MAX_CONCURRENCY
KIMI_THINKING
KIMI_PERMISSION_MODE
KIMI_MCP_TRANSPORT
KIMI_MCP_HTTP_HOST
KIMI_MCP_HTTP_PORT
PORTThe official runtime intentionally overrides attempts to redirect KIMI_MODEL_BASE_URL, KIMI_MODEL_PROVIDER_TYPE, or KIMI_MODEL_API_KEY away from ai&.
Persistent data
The container uses:
/data
├── kimi-code/
├── jobs/
└── state/Mount /data on persistent storage for hosted deployments.
Handoff change metadata
Delegation handoffs report Git state using three separate fields:
committedChanges— changes committed during the delegated Kimi session.workingTreeChanges— uncommitted changes currently present in the worktree.initialDirtyPaths— paths that were already modified before delegation began.
Keeping these fields separate lets an MCP client distinguish work produced by the delegated task from pre-existing local changes.
Local development
Install dependencies:
pnpm install --frozen-lockfileType-check:
pnpm typecheckBuild:
pnpm buildRun tests:
pnpm testThe optional local Codex plugin validator test is skipped when the external Codex plugin-creator validator is not installed.
Claude Desktop with a self-hosted bridge
The Cloudflare edition is added as a connector in Claude (web and Desktop); no local setup is needed.
For a self-hosted Docker bridge, scripts/claude-desktop/ has a small stdio wrapper that forwards Claude Desktop to your bridge over Streamable HTTP. It reads the bearer token from the macOS Keychain, so the token stays out of Claude's config:
read -s KIMI_MCP_AUTH_TOKEN
/usr/bin/security add-generic-password -a "$(/usr/bin/id -un)" -s kimi-swarm-mcp -w "$KIMI_MCP_AUTH_TOKEN" -U
unset KIMI_MCP_AUTH_TOKEN
mkdir -p ~/.claude/bin
cp scripts/claude-desktop/kimi-mcp-bridge.py scripts/claude-desktop/kimi-mcp-desktop.sh ~/.claude/bin/
chmod 700 ~/.claude/bin/kimi-mcp-*Then merge this entry into mcpServers in ~/Library/Application Support/Claude/claude_desktop_config.json and restart Claude Desktop:
{
"mcpServers": {
"kimi-swarm": {
"command": "/Users/YOUR_USERNAME/.claude/bin/kimi-mcp-desktop.sh",
"env": { "KIMI_MCP_URL": "https://your-bridge.example.com/mcp" }
}
}
}Hosted delegations run in the container, so use cwd: /workspace; local macOS paths are not visible there.
Long jobs, recovery and AgentSwarm evidence
For long-running work, prefer the asynchronous sequence:
Call
kimi_delegate_task.Keep the returned
jobIdandsessionIdwhen available.Call
kimi_wait_until_idlewith that session ID.When the job is
idle, callkimi_get_handofforkimi_review_package.
kimi_delegate_and_wait remains convenient for work expected to finish while
the caller stays connected. Its wait may time out, and an MCP/client transport
may also disconnect before the tool response is delivered. Neither condition
means the underlying Kimi job should be submitted again.
When durable jobs are configured, recovery should start with
kimi_recent_jobs. The registry is persistent and connector-owned, so after a
client timeout, reconnect, or bridge restart it can recover the durable
jobId, bound Kimi sessionId, prompt ID, last known status, and cached
result/error data without relying on raw session-title guessing. After
recovering the session ID, continue with kimi_wait_until_idle and then
kimi_get_handoff.
The durable registry is an ownership and recovery index; Kimi remains the
source of truth for the live session and transcript. A direct
kimi_get_handoff refreshes the authoritative Kimi session status and
reconciles the durable job record.
Durable jobs are enabled only when both variables are configured:
KIMI_ORGANIZATION_ID=<customer-or-organization-id>
KIMI_CONNECTOR_INSTANCE_ID=<stable-connector-instance-id>The default SQLite database is:
/data/kimi-swarm-bridge/jobs.sqliteKIMI_JOB_DB_PATH can override that path. Hosted deployments should place the
database on persistent storage. Each Docker runtime serves one
organization; possession of a
job ID or Kimi session ID is not treated as authorization.
For native AgentSwarm acceptance, the final kimi_get_handoff includes a
fresh structured swarmEvidence snapshot. Verify values such as:
swarmEvidence.available: true
swarmEvidence.nativeAgentSwarmObserved: true
swarmEvidence.agentSwarmCallCount: 1
swarmEvidence.requestedWorkerCount >= 3
swarmEvidence.completedWorkerCount >= 3The evidence is derived from Kimi wire/session records rather than from the model's prose self-report.
stdio MCP
The original stdio transport remains available:
node dist/index.jsThe container defaults to Streamable HTTP transport for hosted use.
Security notes
Never commit
AIAND_API_KEY.Never commit
KIMI_MCP_AUTH_TOKEN.Kimi's REST API should remain bound to
127.0.0.1.Expose the MCP endpoint through HTTPS in hosted environments.
Durable job ownership requires both
KIMI_ORGANIZATION_IDandKIMI_CONNECTOR_INSTANCE_ID; job IDs and Kimi session IDs are not authorization credentials.Persistent MCP job/session state does not by itself provide hardened multi-tenant isolation. The Docker image is meant for one connector/runtime per organization; for many employees use the Cloudflare edition, which gives each signed-in employee their own container and keeps API keys out of containers.
Status
Validated so far:
ordinary Kimi inference through ai&
zai-org/glm-5.3(default) andmoonshotai/kimi-k3native Kimi AgentSwarm
up to 20 concurrent native workers in a single swarm (one per researched item)
coordinator and workers all using ai&
cancellation
authenticated Streamable HTTP MCP
MCP disconnect/reconnect with job recovery
Docker runtime (HTTP and stdio transports)
persistent Kimi state
managed-host-compatible
/pingandPORThandling
Validated on the Cloudflare edition (docs/cloudflare-deploy.md):
per-employee sign-in (Cloudflare Access OIDC) and isolated containers
files in and out of Claude chats through signed, single-use links
persistence of jobs, sessions and workspace files across container restarts
ai& and Brave Search keys held by the Worker, never inside containers
web research via Brave Search and a page reader that returns only relevant facts
per-employee daily ai& request budget and optional outbound logging/allowlist
per-employee agent ceiling set from chat
connector handshake and tool list served without waking a sleeping container
research, repository, download and coding tasks with web access
organization-wide ai& concurrency limit, changeable live
per-worker research time budget, so one slow worker doesn't hold up a swarm
the Kimi Swarm skill in the Claude desktop app: Claude offered Kimi for the independent half of a request, delegated on "yes" and combined both parts (the 20-item half took Kimi 2m44s)
See CHANGELOG.md for release notes.
Upstream and license
This repository is derived from:
ximenchuifeng/codex-kimi-bridge
The original copyright and MIT license are preserved in LICENSE.
Modifications in this fork are maintained by ryanameier.
MIT License.
Available Tools
14 toolskimi_abortAbort Kimi SessionADestructive
Abort an owned Kimi session that should no longer continue running. Use only after confirming the intended session because this interrupts in-flight delegated work and changes session state. When durable jobs are configured, a successful abort is persisted as aborted. Returns the sessionId and explicit abort confirmation. Session history may remain available after the abort.
| Name | Required | Description | Default |
|---|---|---|---|
| sessionId | Yes | Kimi session ID to stop. Use the ID returned by delegation or session-discovery tools and confirm it is the intended running job. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the annotations by disclosing that the abort interrupts in-flight delegated work, changes session state, persists as aborted for durable jobs, returns the sessionId and confirmation, and may keep history available. This is exactly the kind of behavioral context a destructiveHint alone cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded with the core purpose, and each subsequent sentence earns its place by addressing the destructive consequences, persistence behavior, return value, and post-abort history. No filler or repetition of structured metadata.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter destructive tool with rich schema coverage and clear annotations, the description covers when to act, what will happen, what is returned, and what may remain. No critical information is missing for an agent to select and invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already covers 100% of the single parameter with a rich description, so the baseline applies. The tool description reinforces the need to confirm the intended session but does not add new parameter meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise action ('Abort an owned Kimi session'), the target resource, and the condition for use ('that should no longer continue running'), which sets it apart from sibling tools like kimi_delegate_task or kimi_continue_task. The opening sentence makes the tool's purpose immediately clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear when-to-use context: abort only after confirming the intended session, and only for work that should no longer continue. It does not explicitly name sibling alternatives for when not to use it, so it falls short of fully explicit when-not/alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kimi_bridge_statusCheck Kimi Bridge StatusARead-onlyIdempotent
Check bridge and private Kimi-runtime readiness without submitting an LLM task. Use before delegation or when diagnosing connectivity, authentication, backend, or runtime problems. Returns health and auth status, safe Kimi backend/version metadata, diagnostics, and suggested next actions. It does not start a session, run AgentSwarm, expose credentials, or modify workspace content.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false. The description adds valuable operational transparency beyond those: no session creation, no AgentSwarm execution, no credential exposure, no workspace modification, and returns health/auth status plus safe backend metadata and diagnostics. This materially enriches the behavioral contract without contradicting annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, each earning its place: purpose, usage timing, return contents, and explicit negative side effects. The core function is front-loaded, and there is no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter diagnostic tool with no output schema, the description covers why and when to call it, what it returns, and what it will not do. It names the relevant diagnostic domains (connectivity, authentication, backend, runtime) and is sufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the empty schema is complete and the description does not need to document any parameter behavior. With no parameters, the baseline of 4 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description immediately states the specific verb and resource: 'Check bridge and private Kimi-runtime readiness' with a clear scope limiter 'without submitting an LLM task.' It clearly differentiates from the delegation and task-management siblings by framing this as a pre-delegation diagnostic.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says when to use the tool: 'Use before delegation or when diagnosing connectivity, authentication, backend, or runtime problems.' It also provides what-not conditions, such as not starting a session or running AgentSwarm, but it does not name a specific alternative sibling to use instead in those cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kimi_continue_taskContinue Kimi SessionADestructive
Submit follow-up instructions to an owned existing Kimi session while preserving prior context. Use for corrections, additional work, or deliberate continuation of a recovered job instead of creating a duplicate session. Returns the new promptId and running status; when durable jobs are configured, the registry is updated to the new prompt and running state. The continuation may execute commands and modify files, and optional model, thinking, and swarm settings apply to the new prompt.
| Name | Required | Description | Default |
|---|---|---|---|
| plan | No | Optional ordered steps for the follow-up work. | |
| task | Yes | Follow-up instruction, correction, or additional work to perform in the existing session. | |
| depth | No | Research depth: quick (one source per item), standard (default; one or two authoritative sources, gaps marked not documented), deep (cross-checked, thorough). | |
| model | No | Configured Kimi model alias. In the managed ai& deployment omit this field to keep the centrally configured model binding. | |
| thinking | No | Optional Kimi thinking setting. Omit to keep the bridge default. | |
| sessionId | Yes | Existing Kimi session ID whose context should be preserved for the follow-up task. | |
| swarmMode | No | Optional swarm-mode setting to verify before the continuation prompt. Set true only when the follow-up should allow native AgentSwarm. | |
| acceptanceCriteria | No | Optional verifiable conditions for the follow-up work. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds substantial behavioral detail beyond the annotations: it warns that the continuation may execute commands and modify files, describes return values (new promptId and running status), and notes registry updates when durable jobs are configured. This meaningfully informs an agent about side effects and state changes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with no filler, front-loading the purpose and usage guidance. Every sentence adds useful information, including return behavior and destructive side effects, without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers purpose, when to use, return values, and destructive behavior, which is strong for a tool with no output schema. It does not address failure cases or what happens when the session is not found or not owned, but the annotations and full schema coverage cover most operational needs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents every parameter. The description adds little beyond what the schema provides, only noting that optional model, thinking, and swarm settings apply to the new prompt, which is already reflected in the parameter descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Submit follow-up instructions') and names the resource ('owned existing Kimi session'), making the core action clear. It distinguishes the tool from creating a duplicate session, but it does not explicitly name or differentiate a specific sibling tool, so the distinction is slightly less sharp than it could be.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use the tool: for corrections, additional work, or deliberate continuation of a recovered job instead of creating a duplicate session. It gives clear use-case context but does not name alternative sibling tools or state when not to use it, leaving some routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kimi_delegate_and_waitDelegate Kimi Task and WaitADestructive
Start a Kimi task and wait until it becomes idle, blocked, failed, aborted, or the caller wait times out. Use this when the caller can remain connected for the expected task duration; prefer kimi_delegate_task for long-running work so the durable jobId is returned immediately. When durable jobs are configured, the result includes jobId. Idle results include handoff and review data; timeout does not abort the underlying Kimi session. After a client-side timeout or disconnect, use kimi_recent_jobs to recover the existing job rather than submitting the task again.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | Yes | Working directory visible to the Kimi runtime. In the managed hosted deployment use /workspace; local desktop paths are not automatically available to the remote runtime. | |
| plan | Yes | Ordered implementation or analysis steps Kimi should follow. For swarm work, use distinct non-conflicting scopes that can be delegated to workers. | |
| task | Yes | Concrete objective for Kimi to execute. Include the requested outcome and relevant constraints. | |
| depth | No | Research depth: quick (one source per item), standard (default; one or two authoritative sources, gaps marked not documented), deep (cross-checked, thorough). | |
| model | No | Configured Kimi model alias. In the managed ai& deployment omit this field to use the centrally configured model binding; do not pass a raw provider model ID unless Kimi exposes it as an alias. | |
| dedupe | No | Optional duplicate-session guard. When supplied, the bridge searches recent sessions before creating a new session and may reuse a compatible match. | |
| thinking | No | Optional Kimi thinking setting. Omit to use the bridge default; the managed pilot is configured for high thinking. | |
| sessionId | No | Existing Kimi session ID to submit into. Omit for a fresh session; fresh sessions are recommended for new swarm jobs. | |
| swarmMode | No | Set true to activate and verify Kimi native swarm mode before prompt submission. When true, the result includes structured swarmEvidence when Kimi wire records are available. | |
| timeoutMs | No | Maximum time in milliseconds to wait for this call. A timeout returns control without aborting the Kimi session, so the same session can be waited on later. | |
| acceptanceCriteria | Yes | Verifiable conditions that define successful completion. Pass an empty array only when there are genuinely no explicit acceptance checks. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the annotations by explaining timeout semantics ('timeout does not abort the underlying Kimi session'), durable jobId inclusion, and idle-result content such as handoff and review data. It also clarifies that a client-side disconnect does not destroy the underlying job, which is valuable behavioral context an agent would otherwise not know.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three dense sentences with every sentence earning its place: the first defines behavior, the second gives usage guidance and sibling differentiation, and the third covers timeout recovery and result contents. It is front-loaded with the core action and avoids repetition of schema details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter tool with no output schema, the description provides strong context by explaining when to use it, what happens on timeout, and how to recover a job. It does not enumerate result shapes for blocked/failed/aborted statuses, but enough is specified for an agent to invoke the tool and handle the common recovery path.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the input schema already provides detailed explanations for cwd, plan, task, dedupe, timeoutMs, and others. The description adds no extra parameter-level semantics beyond what the schema already documents, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb+resource statement: 'Start a Kimi task and wait until it becomes idle, blocked, failed, aborted, or the caller wait times out.' It also distinguishes itself from kimi_delegate_task by noting the waiting behavior and immediate jobId return differences, so an agent can clearly tell which tool to select.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage conditions are explicit: this tool is for cases where the caller can remain connected for the expected duration, while kimi_delegate_task is preferred for long-running work. It also gives a recovery instruction after timeout or disconnect, directing the caller to kimi_recent_jobs instead of resubmitting.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kimi_delegate_taskDelegate Kimi TaskADestructive
Start or continue a Kimi task and return immediately. Use this for long-running or asynchronous work where the caller should receive identifiers before completion. Returns a durable jobId when durable jobs are configured, plus sessionId, promptId, status, and webUrl. After a disconnect or client timeout, recover with kimi_recent_jobs, then use kimi_wait_until_idle and kimi_get_handoff instead of submitting a duplicate prompt. The task may execute commands and modify files in cwd. With swarmMode=true the bridge verifies Kimi swarm mode before submission; final AgentSwarm execution evidence is available from the handoff. Also consider it when the user did not mention Kimi: if part of a request is substantial (would take you several minutes, e.g. researching or comparing many items, reading many documents, building a report or file) and independent of the rest, offer to hand that part to Kimi so it runs while you do the rest; ask first unless kimi_swarm_settings shows offerKimi auto (off: only when asked). If the user agrees, send a self-contained brief here first, then do your part.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | Yes | Working directory visible to the Kimi runtime. In the managed hosted deployment use /workspace; local desktop paths are not automatically available to the remote runtime. | |
| plan | Yes | Ordered implementation or analysis steps Kimi should follow. For swarm work, use distinct non-conflicting scopes that can be delegated to workers. | |
| task | Yes | Concrete objective for Kimi to execute. Include the requested outcome and relevant constraints. | |
| depth | No | Research depth: quick (one source per item), standard (default; one or two authoritative sources, gaps marked not documented), deep (cross-checked, thorough). | |
| model | No | Configured Kimi model alias. In the managed ai& deployment omit this field to use the centrally configured model binding; do not pass a raw provider model ID unless Kimi exposes it as an alias. | |
| thinking | No | Optional Kimi thinking setting. Omit to use the bridge default; the managed pilot is configured for high thinking. | |
| sessionId | No | Existing Kimi session ID to submit into. Omit for a fresh session; fresh sessions are recommended for new swarm jobs. | |
| swarmMode | No | Set true to activate and verify Kimi native swarm mode before prompt submission. Activation does not by itself prove AgentSwarm executed; use kimi_delegate_and_wait for structured swarm evidence. | |
| acceptanceCriteria | Yes | Verifiable conditions that define successful completion. Pass an empty array only when there are genuinely no explicit acceptance checks. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description warns that the task may execute commands and modify files in cwd, going beyond the destructiveHint=true annotation with a concrete side-effect statement. It also discloses that swarmMode only verifies the mode before submission, that final evidence comes later via handoff, and that the call returns before completion.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core contract is front-loaded in the first sentence, followed by return values, recovery, side effects, swarm mode, and proactive-offer guidance. It is longer than strictly necessary, but each paragraph carries distinct information and avoids repeating schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex asynchronous tool with nine parameters and no output schema, the description covers what is returned, how to recover after interruptions, side-effect risk, swarm mode semantics, and when to offer Kimi proactively. An agent has enough context to call it correctly and select the right follow-up tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all nine parameters in detail. The description adds only mild extra context, such as swarmMode verification semantics, rather than materially enriching the parameters beyond the schemas.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a clear operation—start or continue a Kimi task and return immediately—so an agent understands the fire-and-forget contract. It identifies the resource, the immediate-return behavior, and the fact that the caller receives identifiers before completion, distinguishing it from the wait/recovery siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly scopes when to use it (long-running or asynchronous work) and gives a concrete recovery path after disconnect or timeout using kimi_recent_jobs, kimi_wait_until_idle, and kimi_get_handoff instead of submitting a duplicate prompt. It also covers the proactive case where Kimi was not mentioned, including ask-first and auto-off behavior.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kimi_find_recent_sessionFind Recent Kimi SessionARead-onlyIdempotent
Find raw Kimi sessions whose titles contain a requested substring, optionally filtered by status and working directory. Use this only when the exact session is unknown and durable recovery through kimi_recent_jobs is unavailable or insufficient. Once a session is identified, prefer sessionId-based tools. Discovery does not grant ownership; session-oriented tools still enforce durable ownership when configured. This tool does not create or modify a session.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Optional working directory used to constrain matches to the same workspace. Recommended for safe recovery/dedupe. | |
| status | No | Optional exact Kimi session-status filter, such as running, idle, awaiting_approval, awaiting_question, aborted, or failed. | |
| pageSize | No | Maximum number of recent sessions to inspect. Defaults to 20. | |
| matchAnyCwd | No | Set true only when intentionally allowing a title match from any working directory. Defaults to false when cwd is provided. | |
| excludeEmpty | No | Whether sessions with no messages should be excluded from the search. | |
| titleContains | Yes | Case-insensitive substring that must appear in the Kimi session title. Leading and trailing whitespace is ignored. | |
| includeArchive | No | Whether archived Kimi sessions should be included in the search. | |
| includeSummary | No | Fetch message count and latest meaningful user/assistant messages for candidates. Adds latency; leave false unless recovery context is needed. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations establish read-only/idempotent/non-destructive safety, and the description reinforces this with 'does not create or modify a session.' It also adds a non-obvious behavioral caveat: discovery does not grant ownership and session-oriented tools still enforce durable ownership.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four lean sentences arranged by function, usage, ownership caveat, and non-mutation. Every sentence adds a distinct piece of guidance without padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a rich schema, read-only annotations, and sibling context, the description covers purpose, precise invocation conditions, ownership behavior, and side-effect-free behavior. Nothing critical is missing for correct selection and use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema fully describes all 8 parameters and the description need not repeat them. The description adds only general context that results can be filtered by status and cwd, which is already evident from the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific action ('Find') and resource ('raw Kimi sessions') with an explicit matching criterion (title substring) and optional filters. Distinguishes itself by saying it is for when the exact session is unknown and contrasts with kimi_recent_jobs and sessionId-based tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use: only when exact session is unknown and durable recovery via kimi_recent_jobs is unavailable or insufficient. Also states what to do after identification (prefer sessionId-based tools), so an agent knows both when and when not to invoke it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kimi_get_diffGet Kimi File DiffARead-onlyIdempotent
Read the diff for one file in an owned Kimi session workspace. Use after kimi_get_handoff or kimi_review_package identifies a changed path and exact patch content is needed. Returns the requested path and diff content. This tool is read-only and does not modify the file, session, or durable job state.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Workspace-relative file path whose diff should be returned. Prefer a path reported by kimi_get_handoff or kimi_review_package. | |
| sessionId | Yes | Kimi session ID that owns the workspace/file change. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is known. The description adds value by explicitly stating what is not modified—file, session, or durable job state—and by clarifying the return behavior (requested path and diff content), which goes beyond the generic annotation hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with no filler. It front-loads the action, then provides the trigger context, the expected return, and the safety guarantee. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter read-only tool with no output schema, the description covers the purpose, when to call it, what it returns, and its side-effect behavior. There is no meaningful gap that would prevent an agent from invoking it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already provides 100% description coverage for both sessionId and path, including guidance to prefer paths from kimi_get_handoff or kimi_review_package. The tool description adds no additional parameter-level meaning beyond the schema, so the baseline score of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Read the diff') and a specific resource ('for one file in an owned Kimi session workspace'). It clearly distinguishes the tool as the diff-reading step after kimi_get_handoff or kimi_review_package, so an agent can identify what it does without needing to open the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use guidance: after kimi_get_handoff or kimi_review_package identifies a changed path, when exact patch content is needed. This tells the agent the preceding steps and the condition that should trigger this tool, which is strong contextual usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kimi_get_handoffGet Kimi HandoffARead-onlyIdempotent
Read the current or final handoff for an owned Kimi session. Use after kimi_wait_until_idle reports idle, or directly during recovery when the authoritative Kimi status and result need to be refreshed. Returns the assistant result, changed files, committed and working-tree changes, Git baseline evidence, and a fresh structured swarmEvidence snapshot. When durable jobs are configured, the bridge reconciles the durable status and caches the returned handoff. This tool does not submit another prompt or modify workspace content.
| Name | Required | Description | Default |
|---|---|---|---|
| sessionId | Yes | Kimi session ID whose current/final result and Git-change evidence should be retrieved. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already signal read-only/idempotent/non-destructive, and the description reinforces this by stating it does not submit or modify. It adds genuinely new behavior: durable-job reconciliation and caching of the returned handoff, which are side effects an agent might not expect from a read-only tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, each with a distinct job: scope, use timing, return payload, durable side-effect, and non-modification guarantee. Information is front-loaded and there is no filler, though the durable-job sentence adds nuance that could have been separate.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a single obvious parameter and no output schema, the description conveys the key return fields and the exact lifecycle timing, which is enough to call it correctly. The only minor gap is that 'current or final' and 'owned' are not formally defined, but siblings fill the surrounding workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter, sessionId, has a complete schema description. The tool description adds the 'owned' constraint and ties the ID to current/final result retrieval, giving a bit more selection meaning than the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence names a concrete read action ('Read') on a specific resource ('current or final handoff for an owned Kimi session'), which clearly distinguishes it from delegation, waiting, diff, abort, and status siblings. It also immediately scopes to 'current or final' and 'owned,' so an agent can classify the tool as a retrieval operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to use after kimi_wait_until_idle reports idle or during recovery when authoritative status/result must be refreshed. It does not name a sibling alternative for when not to use it, and the negative constraint ('does not submit another prompt') is more of a safety note than an exclusion, so this is clear context without full alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kimi_model_settingsKimi Swarm ModelsAIdempotent
Show or change which ai& models Kimi uses. The coordinator plans the task, delegates to AgentSwarm workers and writes the result; the workers do the parallel work, and most of a swarm's tokens are theirs, so a cheaper worker model cuts cost the most. Call with no arguments to list the available models with prices (USD per million tokens) and the current choice. Call with coordinatorModel and/or workerModel (model ids from the list) when the user asks to switch models; use "default" to return the coordinator to the deployment default and "same" to make workers use the coordinator model. Changes apply to tasks started afterwards; the setting is per user and persists.
| Name | Required | Description | Default |
|---|---|---|---|
| workerModel | No | ai& model id for AgentSwarm workers, or "same" (use the coordinator model). | |
| coordinatorModel | No | ai& model id for the coordinator, or "default". |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare idempotent and non-destructive hints, so the description carries the burden of explaining mutability. It does so thoroughly: changes are per-user, persistent, and only apply to tasks started afterwards. It also explains the cost implications of worker vs coordinator model choice, adding real behavioral context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and then builds logically: role explanation, cost rationale, read behavior, write behavior, and persistence semantics. Every sentence earns its place, and the length is justified by the need to explain a stateful setting with special values.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with two optional parameters and no output schema, the description covers all essential aspects: how to list, how to switch, what special values mean, when changes take effect, and scope. An agent has everything needed to invoke it correctly without additional inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, but the description goes further by explaining that parameter values come from the listed model ids and by defining the two special values 'default' and 'same.' It also ties each parameter to the coordinator/worker role, giving agents the conceptual model needed to pick the right parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource pairing: 'Show or change which ai& models Kimi uses.' It clearly distinguishes between the coordinator and worker roles and states that calling with no arguments lists models, while calling with parameters changes settings. This is unambiguous and easily separable from sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit call patterns: no arguments for listing, coordinatorModel/workerModel for switching, and special values 'default' and 'same' with their meanings. It does not name sibling tools or state when not to use this tool, but the context is clear enough for an agent to decide correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kimi_recent_jobsList Recent Durable JobsARead-onlyIdempotent
List recent connector-owned durable jobs from the persistent bridge registry. Use this first after a client timeout, disconnect, bridge restart, or lost response to recover the existing jobId and bound Kimi sessionId without guessing or submitting a duplicate task. Returns status, prompt/session identifiers, workspace, swarm mode, cached result or error data, and timestamps. It does not contact Kimi, so registry recovery remains available when the Kimi runtime is temporarily unavailable. This tool is read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | Optional durable job-status filter. | |
| pageSize | No | Maximum number of connector-owned durable jobs to return. Defaults to 10 and is capped by the registry at 100. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false Publishing. The description adds meaningful context beyond the annotations: it operates on a persistent registry and does not contact Kimi, making behavior transparent during runtime outages. No contradiction with the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, each earning its place: definition, usage timing, return payload, and a critical behavioral note. The primary statement is front-loaded)Skip and the guidance follows naturally without padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple two-parameter read-only tool and no output schema, the description fully covers what an agent needs: when to call it, what it returns (status, identifiers, workspace, swarm mode, cached results/errors, timestamps), and that it works even if Kimi is unavailable. The pagination default/cap is already present in the schema, so nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both 'status' and 'pageSize' fully documented including defaults, enum values, and caps. The description itself does not add parameter-level detail, but that is acceptable because the schema already does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List'), a precise object ('recent connector-owned durable jobs'), and the source ('persistent bridge registry'). The description also frames the tool as the recovery path for a lost jobId/sessionId, which clearly separates it from siblings like kimi_recent_sessions and kimi_get_handoff.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly directs the agent to 'use this first' after timeouts, disconnects, bridge restarts, or lost responses, and notes it remains usable when the Kimi runtime is down. It does not name alternative tools or state when not to use it, so it falls just short of the full 'when/when-not/alternatives' criterion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kimi_recent_sessionsList Recent Kimi SessionsARead-onlyIdempotent
List recent raw Kimi sessions with identifiers, statuses, titles, web links, and workspace metadata. Use this as a discovery fallback when durable job recovery is unavailable; when durable jobs are configured, prefer kimi_recent_jobs because it is connector-owned and persistent. Raw session discovery is not authorization: session-oriented tools still enforce durable ownership when configured. This tool only queries Kimi session metadata.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | Optional exact Kimi session-status filter, such as running, idle, awaiting_approval, awaiting_question, aborted, or failed. | |
| pageSize | No | Maximum number of recent sessions to return. Defaults to 10. | |
| excludeEmpty | No | Whether sessions with no messages should be excluded. | |
| includeArchive | No | Whether archived Kimi sessions should be included. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is known. The description adds context beyond those: it clarifies that this is a raw metadata query only, and it reveals a subtle behavioral rule that the listing itself does not grant authorization. This is useful, non-duplicative behavioral disclosure beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with no fluff. The first sentence states the action and scope, the second provides usage guidance with an explicit alternative, and the third adds a brief behavioral caveat. Every sentence earns its place, and the most important information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and four optional parameters, the description covers the essentials: what the tool returns, when to use it, why it differs from the durable-job alternative, and its authorization caveat. It also names the core return fields, which is sufficient for a read-only discovery tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already fully explains all four parameters (status, pageSize, excludeEmpty, includeArchive). The description mentions return fields such as statuses and titles but does not add parameter-specific details beyond the schema, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise action and resource: 'List recent raw Kimi sessions' and enumerates the returned metadata (identifiers, statuses, titles, web links, workspace metadata). It also draws an explicit contrast with kimi_recent_jobs, calling this tool a 'discovery fallback' and positioning it as the non-durable alternative, so an agent can distinguish it from siblings without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit routing guidance: 'Use this as a discovery fallback when durable job recovery is unavailable; when durable jobs are configured, prefer kimi_recent_jobs because it is connector-owned and persistent.' It also warns that raw session discovery is not authorization and that session-oriented tools still enforce durable ownership when configured, clarifying when and why this tool should be chosen.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kimi_review_packageBuild Kimi Review PackageARead-onlyIdempotent
Build a reviewer-oriented snapshot for an owned Kimi session from its handoff, changed files, Git statistics, and review checklist. Use after delegated implementation work when concise review evidence is needed, or when a fresh review snapshot is required after recovery. If kimi_delegate_and_wait already returned reviewPackage, reuse that unless newer evidence is needed. When durable jobs are configured, the package is cached in the job registry. This tool does not modify the session or workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| sessionId | Yes | Kimi session ID to package for review. Normally use a session that is idle or otherwise finished producing changes. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, and the description adds useful behavioral context: the package is cached when durable jobs are configured and the tool 'does not modify the session or workspace.' It could disclose failure behavior or freshness implications, but it goes beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, all information-bearing and front-loaded with the main action in the first sentence. The final sentence restates what annotations already cover, but it is brief and does not bloat the description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with read-only annotations and no output schema, the description communicates inputs, output components, when to reuse, caching behavior, and safety. Nothing an agent needs to select and invoke this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is only one parameter and schema coverage is 100%, so the schema already explains sessionId and its idle-session recommendation. The description adds no new parameter-level detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Build a reviewer-oriented snapshot for an owned Kimi session' and enumerates its inputs (handoff, changed files, Git statistics, review checklist). This makes its purpose distinct from siblings like kimi_get_handoff and kimi_get_diff, which cover only pieces of the snapshot.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit trigger conditions ('after delegated implementation work' or 'after recovery') and an explicit reuse rule: if kimi_delegate_and_wait already returned reviewPackage, reuse it unless newer evidence is needed. This is direct when-to-use guidance with an alternative named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kimi_swarm_settingsKimi Swarm SettingsAIdempotent
Show or change how many AgentSwarm workers Kimi may use per task. maxAgents is a ceiling: Kimi still decides how many workers each task needs and uses fewer for small tasks. Call with maxAgents when the user asks to raise or lower the agent limit (for example "increase the limit to 20"); call with no arguments to report the current settings. Values above the deployment cap are reduced to the cap. The response also reports concurrency, the number of workers that run at the same time (set by the deployment admin); extra workers queue, which keeps simultaneous ai& requests bounded, and Kimi backs off automatically if ai& rate-limits. offerKimi is the user's preference for Kimi taking independent parts of requests they did not send to Kimi: ask (default: offer and wait for a yes), auto (hand them over and say so) or off (only when the user asks for Kimi); read it before your first offer in a conversation, and save it when the user says to always do it or to stop asking.
| Name | Required | Description | Default |
|---|---|---|---|
| maxAgents | No | New ceiling on workers per task. Omit to only read the current settings. | |
| offerKimi | No | New preference for offering Kimi on parts of requests: ask, auto or off. Omit to keep it. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide idempotentHint and destructiveHint, but the description goes far beyond by explaining that maxAgents is a ceiling (not a hard number), that values above deployment cap are reduced, concurrency behavior, queueing, rate-limit backoff, and offerKimi semantics. No contradiction with annotations; description adds rich behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but dense; every sentence adds necessary detail. It front-loads the main purpose and then elaborates. It could be slightly tighter, but the complexity justifies the length. Not overly verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with two optional parameters and no output schema, the description explains what the response reports (concurrency, queueing, backoff) and how to use it. All needed information for correct invocation is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but description adds substantial meaning: explains maxAgents is a ceiling, how it gets reduced, offerKimi values (ask/auto/off) with default and behavior. This goes well beyond the schema's descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb ('Show or change') with a specific resource (AgentSwarm workers per task). It distinguishes itself from sibling tools like kimi_model_settings by specifying the exact resource. An agent can immediately understand what this tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly describes when to call with maxAgents (when user asks to raise/lower limit) and when to call with no arguments (to report current settings). Also gives guidance on offerKimi usage ('read it before your first offer', 'save it when the user says...'). This is excellent usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kimi_wait_until_idleWait for Kimi SessionARead-onlyIdempotent
Poll an owned Kimi session until it is idle, awaiting approval, awaiting a question, failed, aborted, or the caller wait times out. Use the sessionId returned by delegation or recovered through kimi_recent_jobs. Returns normalized status and pending approval/question data when applicable. A timeout only ends this wait attempt; it does not abort the Kimi session or mark the durable job timed out. This tool does not modify workspace content.
| Name | Required | Description | Default |
|---|---|---|---|
| sessionId | Yes | Kimi session ID returned by delegation or session-discovery tools. | |
| timeoutMs | No | Maximum time in milliseconds to poll before returning timeout. Timeout does not abort or otherwise mutate the session. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly/idempotent/destructive annotations, the description discloses meaningful behavior: a timeout only ends the wait attempt without aborting the session or marking the durable job timed out, and pending approval/question data is returned when applicable. These details align with the annotations, so there is no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core polling purpose and terminal states, and most sentences add unique value such as timeout semantics and return content. The final sentence about not modifying workspace content partially duplicates the readOnlyHint annotation, so it is not perfectly lean.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter poll tool with no output schema, the description adequately covers sessionId provenance, terminal states, timeout behavior, and return content. The exact shape or error signaling of 'normalized status' is underspecified, but an agent has enough information to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents both sessionId and timeoutMs, including the non-mutating timeout behavior. The description adds only marginal provenance guidance by naming kimi_recent_jobs as a recovery path, which the schema already implies via 'session-discovery tools'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Poll') and a specific resource (an owned Kimi session), and enumerates the exact terminal states. It distinguishes itself from sibling discovery/delegation/abort tools by describing the wait/poll role and by referencing session IDs from delegation or kimi_recent_jobs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear usage context: the sessionId must come from delegation or be recovered through kimi_recent_jobs, and it clarifies that a timeout is not an abort or a durable-job failure. It does not explicitly name alternatives like kimi_delegate_and_wait or state when not to use the tool, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v0.5.1- Changed
kimi_continue_task1 field changed- added
Input schema / properties / depthAdded value: +{ + "description": "Research depth: quick (one source per item), standard (default; one or two authoritative sources, gaps marked not documented), deep (cross-checked, thorough).", + "enum": [ + "quick", + "standard", + "deep" + ], + "type": "string" +}
- Changed
kimi_delegate_and_wait1 field changed- added
Input schema / properties / depthAdded value: +{ + "description": "Research depth: quick (one source per item), standard (default; one or two authoritative sources, gaps marked not documented), deep (cross-checked, thorough).", + "enum": [ + "quick", + "standard", + "deep" + ], + "type": "string" +}
- Changed
kimi_delegate_task1 field changed- added
Input schema / properties / depthAdded value: +{ + "description": "Research depth: quick (one source per item), standard (default; one or two authoritative sources, gaps marked not documented), deep (cross-checked, thorough).", + "enum": [ + "quick", + "standard", + "deep" + ], + "type": "string" +}
- Added
kimi_model_settings - Changed
kimi_swarm_settings1 field changed- added
Input schema / properties / offerKimiAdded value: +{ + "description": "New preference for offering Kimi on parts of requests: ask, auto or off. Omit to keep it.", + "enum": [ + "ask", + "auto", + "off" + ], + "type": "string" +}
1 tool update
v0.4.0- Added
kimi_swarm_settings
1 tool update
v0.3.4- Added
kimi_recent_jobs
11 tool updates
v0.3.2- Changed
kimi_abort1 field changed- added
Input schema / properties / sessionId / descriptionAdded value: +"Kimi session ID to stop. Use the ID returned by delegation or session-discovery tools and confirm it is the intended running job."
- Changed
kimi_bridge_status1 field changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#"
- Changed
kimi_continue_task7 fields changed- added
Input schema / properties / acceptanceCriteria / descriptionAdded value: +"Optional verifiable conditions for the follow-up work." - added
Input schema / properties / model / descriptionAdded value: +"Configured Kimi model alias. In the managed ai& deployment omit this field to keep the centrally configured model binding." - added
Input schema / properties / plan / descriptionAdded value: +"Optional ordered steps for the follow-up work." - added
Input schema / properties / sessionId / descriptionAdded value: +"Existing Kimi session ID whose context should be preserved for the follow-up task." - added
Input schema / properties / swarmMode / descriptionAdded value: +"Optional swarm-mode setting to verify before the continuation prompt. Set true only when the follow-up should allow native AgentSwarm." - added
Input schema / properties / task / descriptionAdded value: +"Follow-up instruction, correction, or additional work to perform in the existing session." - added
Input schema / properties / thinking / descriptionAdded value: +"Optional Kimi thinking setting. Omit to keep the bridge default."
- Changed
kimi_delegate_and_wait18 fields changed- added
Input schema / properties / acceptanceCriteria / descriptionAdded value: +"Verifiable conditions that define successful completion. Pass an empty array only when there are genuinely no explicit acceptance checks." - added
Input schema / properties / cwd / descriptionAdded value: +"Working directory visible to the Kimi runtime. In the managed hosted deployment use /workspace; local desktop paths are not automatically available to the remote runtime." - added
Input schema / properties / dedupe / descriptionAdded value: +"Optional duplicate-session guard. When supplied, the bridge searches recent sessions before creating a new session and may reuse a compatible match." - added
Input schema / properties / dedupe / properties / excludeEmpty / descriptionAdded value: +"Whether sessions with no messages should be excluded from the dedupe search." - added
Input schema / properties / dedupe / properties / includeArchive / descriptionAdded value: +"Whether archived sessions should be included in the dedupe search." - added
Input schema / properties / dedupe / properties / includeSummary / descriptionAdded value: +"Fetch recent user/assistant summary data for candidate sessions. Adds latency; leave false for normal dedupe checks." - added
Input schema / properties / dedupe / properties / matchAnyCwd / descriptionAdded value: +"Set true only when intentionally allowing reuse from a different working directory. Defaults to false for workspace safety." - added
Input schema / properties / dedupe / properties / pageSize / descriptionAdded value: +"Maximum number of recent sessions to inspect for a title match. Defaults to 20." - added
Input schema / properties / dedupe / properties / reuseIfStatus / descriptionAdded value: +"Statuses the caller permits for reuse. The bridge still only auto-reuses running, idle, awaiting_approval, and awaiting_question sessions." - added
Input schema / properties / dedupe / properties / status / descriptionAdded value: +"Optional exact Kimi session-status filter, such as running, idle, awaiting_approval, awaiting_question, aborted, or failed." - added
Input schema / properties / dedupe / properties / titleContains / descriptionAdded value: +"Case-insensitive substring used to find an existing recent session before creating a new one. Use a task-specific title fragment." - added
Input schema / properties / model / descriptionAdded value: +"Configured Kimi model alias. In the managed ai& deployment omit this field to use the centrally configured model binding; do not pass a raw provider model ID unless Kimi exposes it as an alias." - added
Input schema / properties / plan / descriptionAdded value: +"Ordered implementation or analysis steps Kimi should follow. For swarm work, use distinct non-conflicting scopes that can be delegated to workers." - added
Input schema / properties / sessionId / descriptionAdded value: +"Existing Kimi session ID to submit into. Omit for a fresh session; fresh sessions are recommended for new swarm jobs." - added
Input schema / properties / swarmMode / descriptionAdded value: +"Set true to activate and verify Kimi native swarm mode before prompt submission. When true, the result includes structured swarmEvidence when Kimi wire records are available." - added
Input schema / properties / task / descriptionAdded value: +"Concrete objective for Kimi to execute. Include the requested outcome and relevant constraints." - added
Input schema / properties / thinking / descriptionAdded value: +"Optional Kimi thinking setting. Omit to use the bridge default; the managed pilot is configured for high thinking." - added
Input schema / properties / timeoutMs / descriptionAdded value: +"Maximum time in milliseconds to wait for this call. A timeout returns control without aborting the Kimi session, so the same session can be waited on later."
- Changed
kimi_delegate_task8 fields changed- added
Input schema / properties / acceptanceCriteria / descriptionAdded value: +"Verifiable conditions that define successful completion. Pass an empty array only when there are genuinely no explicit acceptance checks." - added
Input schema / properties / cwd / descriptionAdded value: +"Working directory visible to the Kimi runtime. In the managed hosted deployment use /workspace; local desktop paths are not automatically available to the remote runtime." - added
Input schema / properties / model / descriptionAdded value: +"Configured Kimi model alias. In the managed ai& deployment omit this field to use the centrally configured model binding; do not pass a raw provider model ID unless Kimi exposes it as an alias." - added
Input schema / properties / plan / descriptionAdded value: +"Ordered implementation or analysis steps Kimi should follow. For swarm work, use distinct non-conflicting scopes that can be delegated to workers." - added
Input schema / properties / sessionId / descriptionAdded value: +"Existing Kimi session ID to submit into. Omit for a fresh session; fresh sessions are recommended for new swarm jobs." - added
Input schema / properties / swarmMode / descriptionAdded value: +"Set true to activate and verify Kimi native swarm mode before prompt submission. Activation does not by itself prove AgentSwarm executed; use kimi_delegate_and_wait for structured swarm evidence." - added
Input schema / properties / task / descriptionAdded value: +"Concrete objective for Kimi to execute. Include the requested outcome and relevant constraints." - added
Input schema / properties / thinking / descriptionAdded value: +"Optional Kimi thinking setting. Omit to use the bridge default; the managed pilot is configured for high thinking."
- Changed
kimi_find_recent_session8 fields changed- added
Input schema / properties / cwd / descriptionAdded value: +"Optional working directory used to constrain matches to the same workspace. Recommended for safe recovery/dedupe." - added
Input schema / properties / excludeEmpty / descriptionAdded value: +"Whether sessions with no messages should be excluded from the search." - added
Input schema / properties / includeArchive / descriptionAdded value: +"Whether archived Kimi sessions should be included in the search." - added
Input schema / properties / includeSummary / descriptionAdded value: +"Fetch message count and latest meaningful user/assistant messages for candidates. Adds latency; leave false unless recovery context is needed." - added
Input schema / properties / matchAnyCwd / descriptionAdded value: +"Set true only when intentionally allowing a title match from any working directory. Defaults to false when cwd is provided." - added
Input schema / properties / pageSize / descriptionAdded value: +"Maximum number of recent sessions to inspect. Defaults to 20." - added
Input schema / properties / status / descriptionAdded value: +"Optional exact Kimi session-status filter, such as running, idle, awaiting_approval, awaiting_question, aborted, or failed." - added
Input schema / properties / titleContains / descriptionAdded value: +"Case-insensitive substring that must appear in the Kimi session title. Leading and trailing whitespace is ignored."
- Changed
kimi_get_diff2 fields changed- added
Input schema / properties / path / descriptionAdded value: +"Workspace-relative file path whose diff should be returned. Prefer a path reported by kimi_get_handoff or kimi_review_package." - added
Input schema / properties / sessionId / descriptionAdded value: +"Kimi session ID that owns the workspace/file change."
- Changed
kimi_get_handoff1 field changed- added
Input schema / properties / sessionId / descriptionAdded value: +"Kimi session ID whose current/final result and Git-change evidence should be retrieved."
- Changed
kimi_recent_sessions4 fields changed- added
Input schema / properties / excludeEmpty / descriptionAdded value: +"Whether sessions with no messages should be excluded." - added
Input schema / properties / includeArchive / descriptionAdded value: +"Whether archived Kimi sessions should be included." - added
Input schema / properties / pageSize / descriptionAdded value: +"Maximum number of recent sessions to return. Defaults to 10." - added
Input schema / properties / status / descriptionAdded value: +"Optional exact Kimi session-status filter, such as running, idle, awaiting_approval, awaiting_question, aborted, or failed."
- Changed
kimi_review_package1 field changed- added
Input schema / properties / sessionId / descriptionAdded value: +"Kimi session ID to package for review. Normally use a session that is idle or otherwise finished producing changes."
- Changed
kimi_wait_until_idle2 fields changed- added
Input schema / properties / sessionId / descriptionAdded value: +"Kimi session ID returned by delegation or session-discovery tools." - added
Input schema / properties / timeoutMs / descriptionAdded value: +"Maximum time in milliseconds to poll before returning timeout. Timeout does not abort or otherwise mutate the session."
11 tool updates
v0.3.0- First observed
kimi_abort - First observed
kimi_bridge_status - First observed
kimi_continue_task - First observed
kimi_delegate_and_wait - First observed
kimi_delegate_task - First observed
kimi_find_recent_session - First observed
kimi_get_diff - First observed
kimi_get_handoff - First observed
kimi_recent_sessions - First observed
kimi_review_package - First observed
kimi_wait_until_idle
TDQS
Scored across 14 tools
Tools have largely distinct purposes, but there is some overlap between discovery tools (kimi_find_recent_session vs kimi_recent_sessions) and between delegation variants (kimi_delegate_task vs kimi_delegate_and_wait). The descriptions clarify the differences and intended usage, so selection should be reliable with careful reading.
All tool names follow the consistent 'kimi_' prefix + snake_case verb_noun pattern, such as kimi_delegate_task, kimi_get_handoff, and kimi_swarm_settings. No mixed conventions or inconsistent verb styles are present.
With 14 tools, the server is well-scoped for its bridge role, covering status, delegation, monitoring, recovery, settings, and finalization. Each tool serves a distinct operational need without redundancy or bloat.
The surface covers the full lifecycle of delegating to Kimi: check readiness, delegate (async/sync), wait, retrieve handoff, review, recover after disconnect, continue, and abort. Minor gaps like a separate 'list all sessions with filtering' beyond the raw find are mitigated by kimi_recent_jobs and kimi_recent_sessions.
Maintenance
Related MCP Connectors
Agent-native collaboration network: orchestrate a team of long-running agents from any MCP client.
MCP server for AI agents to plan, verify, and deploy Cloudflare-native apps.
Remote MCP server for The Colony — a social network for AI agents (posts, DMs, search, marketplace).
AI Reasoning Cache & Consensus Layer with 11 MCP tools via Streamable HTTP.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceMCP server for cross-platform agent onboarding. Registers external agents, translates intents from LangChain, CrewAI, AutoGen, and A2A formats, and proxies cross-ecosystem transactions.MIT
- AlicenseAqualityCmaintenanceMCP server that turns Kimi K2.6 Turbo into an agentic coding assistant with tools for file operations, shell commands, and code search.42MIT
- AlicenseAqualityDmaintenanceBridges any MCP client (like Claude Code, Zed, VS Code) to any ACP coding agent, enabling multi-agent orchestration from a single chat interface.24103 npm9Apache 2.0
- AlicenseDqualityBmaintenanceEnables Codex to delegate implementation tasks to Kimi Code while managing server lifecycle and authentication automatically.111MIT