Skip to main content
Glama

Consult Kimi in background (paid)

kimi_consult_async

Launch a background, read-only Kimi consultation that returns a job ID immediately, letting you retrieve the answer later—designed for extensive code reviews that might exceed synchronous timeouts.

Instructions

Ask Kimi for a read-only second opinion in the background; get a job_id back immediately instead of blocking.

PAID — this spends Kimi quota on every new call; there is no dry-run preview for a consult, so run kimi_status (free) first to confirm the CLI is installed and authenticated.

Same read-only behavior as kimi_consult (Kimi never edits files), but detached — prefer it for a high-reasoning_effort or broad repo-grounded consult that can exceed the synchronous deadline (built-in default 300s), since a sync run whose deadline expires loses its partial work; this job's own deadline is separately configured (built-in default 1800s). Starting a job commits to spend (it runs to completion or its wall-clock deadline even if you never poll). Poll kimi_job_status; read/consume the consult envelope with kimi_job_result/kimi_job_consume_result; stop with kimi_job_cancel.

Data egress: same as kimi_consult — sends your question and extra_context (raw, unredacted) to your configured provider via the kimi CLI, plus files Kimi reads from its resolved working directory (workspace_root, your MCP roots, or the server cwd). Kimi auto-loads the resolved workspace's AGENTS.md and discovers skills from its own config (including extra_skill_dirs, which may point outside the workspace).

Your inputs are sent raw and unredacted. Secret redaction is best-effort and covers the gathered diff and Kimi's returned output — not what you type, and not the files Kimi reads for itself.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
modelNoOverride the Kimi model slug for this call; defaults to the server/Kimi default when unset.
questionYesThe question or prompt to send Kimi (a different model) for a read-only answer. Must be non-blank: empty or whitespace-only is rejected before any model call.
isolationNoKimi skills isolation: which skills Kimi loads — 'inherit' (its own user/project discovery) or 'ignore-skills' (replace those with an empty directory). Kimi's built-in skills load either way. Defaults to the server's configured value (built-in 'inherit'; `kimi_status` reports the resolved one).
extra_contextNoOptional author intent/background context, added as clearly-labeled UNTRUSTED prompt data. Redaction does NOT cover it — no live secrets. Full caveats and bounds: kimi://params.
workspace_rootNoAbsolute path to the target repo root — pass it (or an MCP root) to target the intended repo; otherwise the call falls back to the server's own cwd and sets meta.workspace_warning.
idempotency_keyNoOptional dedup key scoped to THIS tool + workspace. Same key + same args replays the prior result with no new spend; different args are refused (idempotency_conflict). Sync and _async are separate tools and never share a key. Omit for none; retention is bounded. Lifecycle: kimi://params.
reasoning_effortNoOverride the Kimi reasoning effort for this call (a model_reasoning_effort override); omit or pass null for the server default (MOONBRIDGE_REASONING_EFFORT) or Kimi's own resolution. An open, per-model string the backend validates at run time — commonly minimal|low|medium|high|xhigh; kimi_models lists each model's advertised set (advisory). Rejection and bounds detail: kimi://params.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
okYes
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses far more than the annotations provide: paid quota spend, no dry-run preview, job runs even if never polled, raw/unredacted data egress, best-effort secret redaction, workspace/AGENTS.md/skills auto-loading, and idempotency semantics. This is rich, non-obvious behavioral context that an agent needs for correct invocation and risk assessment.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but information-dense; every major section (paid warning, sync-vs-async rationale, job lifecycle, egress, redaction) earns its place. The first sentence front-loads the core purpose. Minor redundancy — 'raw and unredacted' appears twice — and the length could be slightly tightened, so it stops short of a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex paid async tool, the description is remarkably complete: it covers cost, when to use, the job lifecycle (start/poll/consume/cancel), data egress destinations, redaction limits, workspace resolution, and auto-loaded skills. The output schema presumably documents return values, so no return-format explanation is needed. Nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds value by explicitly stating that question and extra_context are sent raw/unredacted, and by explaining how workspace_root resolution falls back (MCP roots or server cwd). This goes beyond the schema's mechanical definitions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource+behavior: 'Ask Kimi for a read-only second opinion in the background; get a job_id back immediately instead of blocking.' It clearly distinguishes itself from sync kimi_consult and the job management siblings by framing the async/detached nature from the first sentence.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells when to choose this over kimi_consult: 'prefer it for a high-reasoning_effort or broad repo-grounded consult that can exceed the synchronous deadline (built-in default 300s), since a sync run whose deadline expires loses its partial work.' It also instructs to run kimi_status first and names the exact polling/consuming/canceling siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Install Server

Other Tools

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/briandconnelly/moonbridge'

If you have feedback or need assistance with the MCP directory API, please join our Discord server