Skip to main content
Glama

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault

No arguments

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": true
}
logging
{}
prompts
{
  "listChanged": false
}
resources
{
  "subscribe": false,
  "listChanged": false
}
extensions
{
  "io.modelcontextprotocol/ui": {}
}
experimental
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
team_statusA

Get a team's status summary — team info + members + active tasks.

Default response is a COMPACT projection (view="compact" + hint - trimmed, NOT missing fields): member and task rows are projected, offline members fold into a count plus digest, and at most 30 active tasks are listed (the remainder is reported in active_tasks_omitted; task_list_project with team_id lists them all).

team_listA

List teams — active ones by default, newest first.

Default response is a COMPACT projection (view="compact" + hint - trimmed, NOT missing fields): each row keeps id / name / status / kind / project_id / created_at. Teams accumulate one row per Workflow run and per CC session, so the list is long: filter by status and page with limit / offset.

agent_update_statusA

Update an Agent's running status.

The status is already maintained from hook events and inactivity: tool activity marks an agent busy, inactivity moves it to waiting and then offline, and session end marks it offline. The next such update overwrites a manual write, so use this only to correct a status the hooks left stale.

agent_listA

List a team's members - live roster first, offline history on request.

Default response is a COMPACT projection (view="compact" + hint - it is a trimmed view, NOT missing fields). Each member row keeps id / name / role / status / an 80-char current_task excerpt / last_active_at; system_prompt, config, the context watermark and the token ledger are omitted and come back with fields="all".

Offline members are folded into a count plus a short most-recent digest. An offline agent is a terminated process - it cannot be messaged and cannot be assigned work - and on a long-lived team offline rows are most of the payload. Nothing is deleted: the count is always reported and include_offline=True returns the full history.

agent_template_listA

List every Agent template CC can actually resolve.

Scans all three template sources with CC's own precedence — project-level <project>/.claude/agents/ > user-level ~/.claude/agents/ > the shipped plugin/agents/ — and de-duplicates by frontmatter name (the identity CC resolves), so the count matches what subagent_type will really accept. Each entry carries a source field.

Default response is a COMPACT projection (view="compact" + hint - trimmed, NOT missing fields): the full listing measured 32,480 chars, half of it because grouped repeats every row of templates verbatim. Compact keeps one projected row per template and reduces grouped to a name index.

agent_template_recommendA

Recommend Agent templates — and, for a known project type, a team shape.

Two layers in one answer:

  1. recommendations — live template match against the installed template dirs (project > user > plugin), ranked by relevance.

  2. team_composition — when task_type names a project type (web-app / api-service / data-pipeline / library / refactor / bugfix), a suggested role lineup with counts and the template to use for each. This is a static seed, not a live probe; it only suggests a shape.

agent_reuse_recommendA

Recommend whether to reuse an existing sub-agent for a follow-up task.

For follow-up work (bug re-fix, deeper research, same-domain iteration), resuming a prior sub-agent preserves its accumulated context. This tool ranks prior sub-agents by same-domain match, reads their P1 context watermark, infers reachability, and recommends one of three actions: reuse (SendMessage resumes it) / slim_then_reuse (self-summarize then spawn fresh with the summary) / spawn_new. It only recommends; the Leader decides.

Availability tiers: live (same session, reachable now) / resumable (same session, offline but transcript fresh) / cross-session (another session, needs claude --resume) / expired (past retention). Address candidates by NAME — SendMessage(to=...) takes a teammate name and keeps working after the agent completes; each candidate's resume_hint is a ready-to-run call (with the required summary). The raw agentId is the documented fallback for nameless rows or when a newer agent took the name.

Default response is a COMPACT projection (view="compact" + hint — trimmed, NOT missing fields): decision signals and call keys kept, full rationale and watermark detail via fields="all".

fleet_dispatchA

Dispatch an operational instruction to another ship (CC session) in the fleet.

The fleet down-channel drives an EXISTING idle session to run one turn via headless claude -p --resume (fleet-layer design §4). Use it to nudge an idle ship to advance a task or report its status - NOT to make strategic decisions on the user's behalf (the dispatched turn is constrained to operational work).

Safety gate (enforced server-side, no subprocess spawns until it passes):

  • The target must be RESUMABLE: its transcript file still exists.

  • The target must NOT be user-live: its file must be idle beyond a conservative guard (FLEET_DISPATCH_MIN_IDLE_SECONDS, > the 15min live window) so a dispatch never competes with someone typing in that session. A too-fresh target is refused with availability="live".

  • Dispatches are deduped per-session, share the global wake concurrency limit and circuit breaker, and every one is ledgered in wake_sessions.

  • The dispatched turn runs under that session's own tool permission settings; the OS grants no tools of its own.

Get target_session_id from the fleet view / project summary (each ship's session_id).

agent_activity_queryA

Query Agent activity records for a team.

Returns recent activity log entries sorted by timestamp descending, including tool name, duration_ms, and an I/O summary.

Default response is a COMPACT projection (view="compact" + hint - trimmed, NOT missing fields): input/output summaries are excerpted because the raw output_summary often holds a whole command transcript (a 60-row window measured 43.9k chars, right at the MCP result ceiling). Full records via fields="all". The compact window is capped at 40 rows - narrow with agent_id rather than widening limit.

meeting_createA

Create a team meeting and return a ready-to-use dispatch_plan for spawning participant Agents.

Supports two participant formats:

  1. Legacy (strings): participants=["arch-lead", "backend-arch"] Returns dispatch_plan with empty launch_call + deprecation warning.

  2. Structured (dicts): participants=[{"name": "arch-lead", "agent_template": "software-architect", "role": "负责评估架构方案", "context_files": ["docs/arch.md"], "expected_output": "三段式"}] Returns dispatch_plan with fully populated launch_call.params ready to paste into Agent tool.

meeting_send_messageA

Send a discussion message in a meeting.

Each round follows the rule its template gives it (the rounds in meeting_create's _template), and that rule takes precedence. Without one: round 1 states each participant's view; round 2+ reads the earlier messages first (meeting_read_messages) and responds to specific points; the final round summarizes consensus and disagreements.

SECURITY: post only as yourself: agent_id, agent_name and caller_agent_id are all your own. A caller_agent_id that differs from agent_id is recorded as impersonation (meeting.impersonation event) but the message is still stored, so the audit is the only safeguard. A moderator speaking as itself uses its own id (e.g. 'team-lead') in all three fields.

meeting_read_messagesA

Read a meeting's discussion messages, oldest first.

Returns the first limit messages in chronological order, so a meeting with more messages than limit is cut at the newest end; raise limit (max 500) to reach the latest rounds.

meeting_concludeA

Conclude a meeting, marking it as completed.

By default checks that all expected participants have spoken before concluding. Set force=True to override, but this will be recorded in the event log.

Concluding records no decision: summary is not saved with the meeting and nothing is written to memory. A conclusion that must outlive the meeting goes on the task wall (task_create / task_update).

meeting_template_listA

List available meeting templates and their round structures.

Returns: templates: All available templates with round structure details

meeting_listA

List meetings for a team, optionally filtered by status.

debate_startA

Start a structured 4-round debate meeting between an Advocate and a Critic.

Debate structure:

  • Round 1 (Advocate): Present proposal/position with evidence

  • Round 2 (Critic): Challenge risks, flaws, and propose alternatives

  • Round 3 (Advocate): Respond to challenges, revise proposal if needed

  • Round 4 (Judge): Render verdict with action items

Returns role assignments and round rules but no dispatch_plan, so each participant is spawned by hand. For ready-to-paste spawn calls use meeting_create(template="debate", participants=[...]) instead.

debate_code_reviewA

Start a debate-style code review for a specific file or change.

Creates a structured 4-round debate where:

  • Advocate defends the current implementation

  • Critic challenges the implementation and proposes improvements

  • Judge synthesizes findings into consensus conclusions and action items

Returns role assignments, round rules and a Round 1 starter prompt but no dispatch_plan, so each participant is spawned by hand. For ready-to-paste spawn calls use meeting_create(template="debate", participants=[...]).

meeting_updateA

Update a meeting's topic or participant list.

notes is accepted but not stored (meetings have no notes field) and the call still reports success; conclusions go on the task wall (task_create / task_update). Changing participants does not change the attendance list meeting_create recorded. To mark a meeting concluded, use meeting_conclude.

meeting_attendance_checkA

Check which expected participants have spoken in the current round.

Use this after spawning all Agents via dispatch_plan to verify attendance before advancing to the next round or concluding the meeting.

The current round is the highest round_number anyone has posted in, so a new round shows up only after its first message; spoken and pending match participants by agent_name. timeout_in_seconds is the time elapsed since meeting_create: it is not reset per round, it is 0 for meetings created by debate_start / debate_code_review, and it is not a countdown.

task_runA

Put a task on a team's wall. Despite the name, nothing executes it.

This tool only creates the task row. Dispatch it yourself (Agent(...) / SendMessage); the sub-agent then writes progress back with task_memo_add.

Priority and horizon drive the task wall's ordering, so set them here.

task_createA

Create a new task in a project (not bound to a team).

Project-level tasks are attached directly to the project and visible on the project task wall. Suitable for planning-phase tasks not yet assigned to a team.

task_statusA

Get one task's full record (every task field, not a trimmed row).

Includes status, result, description, tags, dependencies, and timestamps. task_list_project returns the wall as trimmed rows; task_memo_read returns the task's memo history.

task_updateB

Update a task's fields (partial update — only provided fields are changed).

task_list_projectA

Get the task wall — project-scoped by default, team-scoped on request.

Pass team_id to narrow the wall to one team; leave it empty to get every team under the project plus the project-level tasks that belong to no team.

Project scope leads with digest: the whole wall in one text block (open counts by status and horizon, 7-day trend, stale and dormant counts, the 5 most recent actions, the top 5 pending, the first mid and long pending task). It is the same block the session briefing shows; read it first, the rows below are one page.

Default response is a COMPACT projection (marked by view="compact" + hint — it is a trimmed view, NOT missing fields): each task row keeps id/title/priority/status/score/assigned_to/tags + 80-char desc excerpt (plus result/depends_on/subtask_count when present). Full details of a single task: task_status(task_id) / task_memo_read(task_id).

task_memo_readA

Read all memo records for a task — read before picking up a task to understand historical progress.

task_memo_addA

Add a memo record to a task — for tracking progress, recording decisions, marking issues.

task_execution_traceA

Get a task's execution timeline — plain, or with checkpoints + stats.

Answers "how did this task actually go"; include_stats=True adds the derived summary on top of the timeline.

project_createA

Create a new project with a default Phase automatically created.

The OS never registers a directory on its own. For an unregistered working directory the session-start briefing asks the user; call this when the user agrees to register, and dismiss_project_registration when they decline. The project must be for the current session's working directory: unrelated directories and ancestors of the home directory are rejected.

project_listA

List all projects in the system.

Returns: projects: List of all projects with id, name, description, root_path, etc.

project_updateC

Update a project's name, description, or root_path.

project_deleteA

Delete a project and everything filed under it. Irreversible.

One transaction removes the project's tasks and task memos, its teams, meetings and meeting messages, phases, reports, leader briefings, project- and team-scoped memories (including the project's direction-layer entries), cross-project messages, and the teams' events. Agent rows (they carry the token attribution), workflow run archives and channel messages are kept.

project_summaryB

Get a quick project summary: status (active/inactive), teams, top tasks.

dismiss_project_registrationA

Mark current cwd as dismissed for project registration — won't ask again.

The session-start briefing asks whether to register an unregistered working directory; after this call it stops asking for that directory, and its "not a registered project" notice is dismissed on the Dashboard too. The choice is stored in a local file (~/.claude/data/ai-team-os/ dismissed_projects.json) that no tool reverses. No project is created, changed, or deleted.

decision_logB

Query team decision log — task assignments, approach selections, Agent scheduling decisions.

prompt_effectivenessA

Return effectiveness statistics for Agent templates.

Frozen: still callable, no longer developed.

Aggregates activity records to compute success rate, average duration, and top failure reasons per template. Also counts failure_analysis lessons per template, matched through the failed task's assigned agent.

Use this to identify which Agent templates perform well and which need prompt improvement.

usage_attributionA

Report token usage together with how much of it can actually be accounted for.

Read-only. Every token number comes back alongside its denominator (dispatches_total) and its metric label, because a token count without those two is meaningless: this repo carries two orthogonal metrics that measure 5-25x apart, and sub-agent usage coverage is incomplete (the response reports the measured share). There is deliberately no total field: cache_read dominates the four layers, so a lone total is mostly a cache-read count in disguise.

link_queryA

Query cross-domain reference edges for an object (who references it / what it references).

Edges are extracted automatically (zero-LLM regex) from task memos and reports: wf_ runs, commit hashes, task UUIDs, [[memory]] links.

link_traceA

Trace the reference neighborhood of an object (undirected fanout, depth <= 2).

Answers questions like "which tasks/reports touched commit 9d8f020" or "what work is connected to run wf_cbad7348".

unified_searchA

Search across all OS knowledge: task memos, reports, and tasks.

Three-arm RRF fusion (k=60): BM25 full-text (Chinese bigram native), knowledge-graph fanout (queries containing wf_/commit/uuid IDs pull in everything linked to them), and exact ID-prefix / title match.

Use this to recall past work, by a free-text question or by an ID. Direction-layer memories are not indexed here; use memory_search.

report_saveA

Save a research/analysis report to the database.

Reports are stored in the database with project isolation — no filesystem permission needed. Reports appear on the Dashboard reports page automatically.

report_listA

List saved reports, optionally filtered by author, topic, or type.

Returns reports for the current project context, sorted newest-first.

report_readB

Read the full content of a saved report by ID.

briefing_addA

Park a decision for the user while the user is NOT in the conversation.

Only for questions that come up when nobody can answer them now: autonomous /loop work, background workflows, another session's findings, or a sub-agent's report that leaves something "for the user to decide". If the user is in the conversation, ask them directly instead; the result carries user_present and a hint when this project saw a user message in the last 15 minutes. When the user later answers an item, call briefing_resolve on that item right away. Pending items expire after 14 days without an answer (status only, never deleted).

briefing_listA

List Leader Briefing items. Default shows pending items for user review.

Each item carries project_id and tags, so a long decision queue can be narrowed to one project and/or one topic.

briefing_resolveA

Record the user's answer to a pending decision and close it.

Call it the moment the user answers, in whatever session that happens, one call per answered item. The answer is also written as a decision event (decision.briefing_resolved), so it stays findable as a decision.

briefing_dismissB

Dismiss a Leader Briefing item (no action needed).

failure_analysisA

Record a failed task as a templated lesson entry (failure alchemy).

Frozen: still callable, no longer developed.

The watchdog runs this itself once it stops retrying a failed task; call it by hand only on a failed task that will not be retried. The three artifacts are fixed templates filled from the task's title, result, recorded error, and tags; nothing is inferred beyond those fields, so the output is only as specific as the task's recorded result and error. For a diagnosis of why a task failed, use diagnose_task_failure.

  • Antibody: defensive-rule suggestion built from the failure reason

  • Vaccine: failure case (description, assignee, result, prevention)

  • Catalyst: improvement proposal keyed on the task's tags The combined text is appended to the task as an issue memo (task_memo_read shows it).

diagnose_task_failureA

Auto-diagnose why a task failed and suggest fixes.

Frozen: still callable, no longer developed.

Reads the task's valid memos to identify the failure point, compares with similar successful tasks in the same team, and returns actionable fix suggestions. Each call records a task.failure_diagnosed event.

Use this when a task fails or gets stuck to quickly understand root cause without manually reading through all memo records.

memory_searchA

Search memory entries within one scope, ranked by BM25.

Covers the memories table only (direction-layer entries plus the legacy team/agent knowledge partitions); task memos, reports, and tasks are not searched here (use unified_search). Only valid entries are returned. English matches whole words with no stemming ("worktree" does not match "worktrees"); Chinese matches by character bigrams. An empty query returns the scope's most recent entries. To review everything a dispatched agent inherits, use memory_list.

memory_addA

Add a direction-layer memory — the team's shared, cross-task standing preferences.

方向层 = 低频·高价值密度·跨任务长寿命的偏好/纠正/约束/设计意图。每个派出 的 agent 出生即注入方向层,"全中文""完成即汇报"这类偏好无需手抄进派工 prompt。

写入检验(软门槛):这条能影响多少未来任务?只影响单个任务的 → 去 task_memo_add(情景层),不要写这里。

体量红线是单一轴:存储上限 = 注入预算。方向层按桶计字符配额—— global 1200 字 + 每个 project 1500 字 + user 300 字,一个会话实际继承 3000 字;单条仍 ≤ 400 字。存得下的一定传得到,写不进去的就是真的没位置: 超限时本工具返回该桶全部有效条目(id / kind / 字数 / 全文)+ 用量缺口, 要求当轮先用 memory_invalidate(可用 content_match 子串定位)腾出空间, 再重试本次写入(global/user 桶条目的失效或置换都须经用户过目并带 confirm_shared_scope=true)。

置换 global/user 条目要确认:supersedes 指向 global/user 条目时,旧文本会 从所有项目的会话里消失,与失效同一道闸。未带确认时不写新条、不失效旧条, 返回 requires_confirmation + 旧条全文(target)+ 新文本(replacement):交用户 过目,确认后带 confirm_shared_scope=true 重试。project 桶的置换不需要确认。 超长内容改写成「触发条件 + 指向权威文件」的指针条目(如 "涉及生产/集群/DB 时遵守只读铁律,详见 ~/.claude/CLAUDE.md"),正文外置。

写入侧安全扫描:方向层条目会进每个派出 agent 的 system prompt,因此不可见 Unicode、提示注入句式(覆盖既有指令 / 套取系统提示 / 伪造对话角色)、凭据 形态一律拒绝入库。

kind 四类(决定注入截断优先级 constraint>design>directive>preference):

  • constraint(禁令/护栏):一句话、可机检、终身有效。 如 "所有输出使用中文"、"git 提交绝不自动加 agent 署名"。

  • design(价值排序/设计意图):缺显式指令时的取舍依据。 如 "技术决策偏向质量/简洁/健壮/长期可维护,不看重开发成本"。

  • directive(方法论/工作方式):回答"怎么干"。 如 "完成即按问题→根因→解法→验证汇报,不攒批次"。

  • preference(格式偏好):可选,如 "每句一行便于 diff"。

memory_invalidateA

Invalidate a direction-layer memory — mark it invalid without deleting.

方向层偏好过时/被推翻时显式失效(Zep 失效语义:置 invalid_at 不删除, 保留可审计轨迹)。失效后不再进注入,也默认不出现在 memory_list。

两种定位方式,二选一:memory_id 精确定位,或 content_match 子串定位 (手里只有原文时免去先查一次 id——被配额顶回来的那一刻正是这种处境)。 子串必须唯一命中当前上下文的有效条目:命中 0 条或多条一律不动数据,多条时 返回候选让你给出更精确的子串。两种方式可达的条目相同:global + user + 当前项目的 project 桶,别的项目的条目按不存在处理。

global / user 条目被所有项目的会话继承,未带确认时拒绝并交回条目原文 (requires_confirmation=true,不动数据):把原文交用户过目,确认后带 confirm_shared_scope=true 重试。当前项目的条目不需要确认。

memory_listA

List direction-layer memories — valid entries by default, grouped by kind.

返回当前上下文的方向层条目:global + user 全局条目 + 当前项目的 project 级条目,按 kind 优先级(constraint>design>directive>preference)+ 时间倒序。 这是双 hook 常驻注入的同一数据源;用它审阅"派出的 agent 会继承什么"。

memory_reconcile_candidatesA

按需整理·粗筛:返回情景层候选组 + 方向层清单 + 蒸馏素材 + 操作说明。

调用即占住本项目的整理权(同一项目同一时刻只有一个会话能整理,别的会话 会被挡在门外):判完没有要改的,提交空批 memory_reconcile_apply(operations=[]) 释放。只想看一眼有什么可整理(例如 Leader 循环里的例行查看),传 peek=true: 只读、不占整理权,但这样拿到的候选不能直接 apply。

记忆整理 = 会话内按需显式动作(CC 非常驻,无后台整理进程)。本工具只做 确定性粗筛(零 LLM)——OS 无独立 LLM 凭据,判定由你(调用工具的会话内 agent)完成,工具只负责候选粗筛与操作应用("agent 算、工具存")。

返回四块(project_id 自动按当前上下文解析):

  • candidate_groups:有效 task_memos 按 scope_path/task 聚簇、簇内 BM25 两两 相似度超阈配对成的候选组(含组内各条全文 + id)。逐组做 LLM 精判: KEEP(都留)/ MERGE(合并)/ INVALIDATE(矛盾失效)/ NOOP(不动)。

  • direction_inventory:全部有效方向层条目全文——逐条做陈旧检查(引用的 功能已退役/版本过时/世界已变 → 提 invalidate)。

  • promotion_candidates:高频跨任务反复出现的簇,蒸馏为方向层条目的素材 (promote 操作,source_refs 回指源 memo)。

  • operation_guide:四操作语义 + reconcile 三守则(只留高频有用 / 指向权威 而非复述 / 重写精简优先)+ 量大开 ultracode 提示。

整理权(reconcile_lease):30 分钟,持有者每次 candidates/apply 顺延。 别的会话持有未过期的整理权时返回 success=false、对方还要多久到期和可选的 做法,不交出候选(peek=true 照常可看)。

判完后把确认的操作交给 memory_reconcile_apply 批量应用。

memory_reconcile_applyA

按需整理·应用:批量执行 LLM 精判确认后的操作(确定性,幂等)。

每条操作是一个 dict,按 op 字段分派(未知/缺字段返回 error,不阻断其余):

  • merge:{op:"merge", content:合并后新内容, memo_ids:[被并各条], memo_type?:"summary", scope_path?} —— 建新 memo,把被并各条置 invalid、 invalidated_by 指向新条(Zep 失效语义不删除)。

  • invalidate:{op:"invalidate", memo_ids:[...]} —— 逐条失效(矛盾/被推翻)。

  • score:{op:"score", memo_id, quality_score:1-10, reason} —— 补质量分, reason 入 meta。

  • promote:{op:"promote", content, kind:constraint/design/directive/preference, scope?:"project"/global/user, source_refs?:[源 memo id]} —— 蒸馏提升为方向层 条目;红线照常生效(单条 ≤400 字 + 桶字符配额 global 1200 / project 1500 / user 300,超限该条返回 error 带用量;安全扫描同样生效)。

  • keep / noop:不动(可省略)。

幂等:对已失效条目重复 invalidate/merge 返回 noop 不报错。应用后自动刷新 项目 last_reconcile_at(整理分界线)。

两道闸:① memo id 只认当前项目的,含别的项目 memo 的那条操作整条报错不执行; ② 须持有 memory_reconcile_candidates 发的整理权(peek 不发),否则整批不执行 (先不带 peek 重新 candidates)。本批全部成功即释放整理权;有报错则保留,修正后重试即可; 判完无需改动也提交一次空批(operations=[])释放整理权。

context_resolveA

Get the current active OS context — active project, active teams, member list.

This is the infrastructure for all simplified operations. A single call returns the complete context of the current working environment, allowing Leader or other tools to auto-fill parameters like project_id, team_id, etc.

teams lists EVERY active team of the current project (a project routinely has several at once: the session container team plus one per Workflow run). team keeps the singular shape for backwards compatibility and holds the primary team picked by the same 3-tier priority as team_id auto-resolution (session container > plain project team > newest).

Returns: Context dict containing project / team / teams / agents

os_health_checkA

Check the health status of the AI Team OS API service.

Verifies the API service is running normally by accessing the team list endpoint, and reports one line of token-attribution coverage alongside it. When the API is local and on the port this MCP server manages, it also reconciles the shared PID file: a single healthy listener on that port is written into the PID file if the file points elsewhere.

os_restart_apiA

Restart the AI Team OS FastAPI process safely (standardized restart flow).

Use this after backend code changes to pick up the new version without manually killing processes. The flow has three safety guards:

  1. Busy-agent guard — refuses to restart while any agent is working (status=busy) unless force=True.

  2. Port-pin guard — only ever restarts on the ORIGINAL port (default 8000, read from api_port.txt). If that port is held by an unrelated process it aborts rather than drifting to a random port.

  3. Dead-before-spawn guard — waits until the old process has fully exited and released the port before spawning the new one; never spawns on a timeout.

If the API is already down there is nothing to shut down: guards 1 and 3 do not apply, guard 2 still does, and this becomes a plain start of the API on its configured port.

event_listA

List recent events in the system, optionally filtered.

Default response is a COMPACT projection (marked by view="compact" + hint — it is a trimmed view, NOT missing fields): each row keeps id/type/source/ts plus a one-line summary derived from the event payload. Use fields="all" for full payloads.

find_skillA

Find ecosystem skills/plugins using a 3-layer progressive loading system.

Searches a small curated catalog of third-party skills, plugins and integration recipes bundled with the OS. It does not list what is installed in the current session and does not query a live marketplace.

Layer 1 (quick recommend): Describe your task and get the top 5 catalog entries with one-line descriptions, install commands and match_score; entries with match_score 0 did not match the description. Layer 2 (category browse): Browse all skills grouped by category (memory / code-quality / frontend / security / dev-workflow / integration / etc.). Layer 3 (full detail): Get complete documentation for a single skill including features, OS complement relationship, and variants.

The integration category holds the ecosystem integration recipes (GitHub / Slack / Linear / fullstack team); each one says which external MCP server to install and which OS tools it pairs with.

os_config_changeA

Change the user's OS installation after the user approved a preview.

Two calls. First without confirm_token: returns the preview (every file, the action, sha256 before and after, the baseline tree and branch, warnings) and a confirm_token valid for 10 minutes. Show the preview to the user as is and ask. Only after the user agrees, call again with the token and the user's own words. The preview is recomputed; if anything changed you must preview again. Existing files are backed up next to themselves (.bak-aiteam-) before writing, and a decision.user_config_write event is recorded. Never pass a token the user has not seen the preview for.

model_config_getA

Get model governance state: available models (auto-discovered from local CC transcripts — the models you actually used), the current default startup model (~/.claude/settings.json "model" key), and per-model workflow agent usage over the last N days (orchestration charter observability: how much fable vs opus the fleet burned).

model_config_setA

Set the default startup model for new CC sessions (writes the "model" key in ~/.claude/settings.json; empty string removes the key, restoring CC's own default). Takes effect on NEW sessions.

channel_waitA

等待指定对端的新消息:先补读,随后以 WebSocket 等待,不轮询模型。

纯读、不自动 ACK。返回正文后,等待此工具的当前回合可继续;不能唤醒已经结束 的 Desktop 回合。超时不自动重开等待。取消或连接故障会结束本次订阅。 调用方应让 MCP 请求超时大于 timeout_seconds + 4 * io_timeout_seconds + 5 秒。 客户端若提前超时,须发送 MCP cancel 或关闭连接;仅本地超时服务端无法感知。

返回 status=messages 或 timeout,附正文列表、has_more、next_cursor 与 delivery_source(replay=初始补读,event=WS 事件后补读,timeout_read=到期 后的末次补读,可为空)。游标失效直接报错,不静默跳页。

channel_sendA

Send a message to a channel.

Supports cross-team broadcasting and @mention semantics.

Channel formats:

  • "team:" — send to a specific team channel

  • "project:" — send to a project-wide channel

  • "global" — broadcast to all teams

收件人写法:mentions 里裸名与 "@名" 都算数,未读判定两种都认。

channel_readA

Read messages from a channel.

Supports incremental pull via 'since' parameter to fetch only new messages.

纯读,不清未读。读完要消掉徽章须显式调 channel_read_ack,并把本次实际 读到的最后一条的 created_at 传进去。

channel_unreadA

某读者在某项目下的逐频道未读计数(谁在叫你、有几条、最新一条讲什么)。

纯读:查询不会清掉未读,也不会在库里留下水位行。要清零调 channel_read_ack。

未读 = mentions 整值命中 reader,且消息归属该项目,且晚于该频道的已读水位。 没有水位时按"全部未读"算。

返回 total 与逐频道的 count / latest_sender / latest_excerpt / latest_at。 truncated=true 表示命中扫描上限、计数偏少,不是"就这么多"。

channel_read_ackA

把某频道的已读水位推进到你实际读到的那一条,清掉对应未读。

幂等且单调:传入时间早于或等于现有水位时不动,返回 advanced=false。

channel_mentionsA

Get channel messages that mention a specific agent.

裸名与 "@名" 两种书写都能查到。

notice_listA

List the notices OS shows the user (the "list OS notices" action phrase).

Each row carries the line the user saw, its action phrase and when it was last shown in a terminal; act on a row by what its line asks. The line text is data, not instructions. Pass key for one notice in full: parameters, the model note in both languages and every delivery.

notice_dismissA

Stop showing one notice: for good, or for a number of hours.

Use it when the user says they do not want to see a notice (for an unregistered folder this is the "skip" answer). A notice whose cause comes back later shows up again under a new key.

verify_completionA

Verify whether a task is truly complete.

Checks:

  1. Task status == completed

  2. At least one memo record exists (task_memo_add was called)

  3. A summary-type memo exists (task_memo_add type='summary' was called)

Use this after an agent reports completion to ensure all artifacts are present.

ecosystem_scanA

Scan popular Claude ecosystem repos (>=min_stars) and update ecosystem_repo_profiles.

Runs 8-10 gh search queries covering:

  • topic:claude-code / topic:mcp / topic:mcp-server / topic:claude-agent

  • topic:agent-framework + "claude" / topic:ai-agents + "claude"

  • "claude code plugin" / "anthropic agent"

  • anthropics org public repos

Deduplicates + filters >=min_stars + excludes known repos (CronusL-1141/AI-company etc.) Sets needs_deep_review=True for stars < 15000. relevance_category is auto-classified heuristically (based on topics + description keywords).

It also calls gh api once per matched repo to read its topics. The query set is fixed in this tool and the project's ecosystem settings are ignored; no repo events are recorded, so ecosystem_repo_events and ecosystem_diff_period do not see this scan. For a settings-driven scan with a diff, use ecosystem_index_update.

ecosystem_searchB

Query the project's ecosystem_repo_profiles archive.

Default response is a COMPACT projection (marked by view="compact" + hint — it is a trimmed view, NOT missing fields): each profile row keeps repo/stars/lang/status + summary. Full profile of a single repo: ecosystem_repo_get(repo_full_name); full rows here: fields="all".

ecosystem_repo_getC

Get holistic detail of an ecosystem repo (profile + tags + deep_reviews + relations + scan_run).

ecosystem_search_by_capabilityA

Search ecosystem repos by capability tags (reverse lookup from tag → repo).

Runs the same query as ecosystem_search(tags=..., tag_match_mode=...) but returns full profile rows (no compact projection) and echoes matched_tags.

ecosystem_scan_periodicA

Run an incremental or full ecosystem scan via the scanner service.

Compared to ecosystem_scan, this tool:

  • skips repos last_scanned_at < 7 days (incremental strategy only)

  • applies secondary owner / keyword filters

  • marks repos pushed > 365 days ago as is_archived=True

  • records every run as an EcosystemScanRun for audit

Queries and filters come from the built-in query set and ECOSYSTEM_* environment variables, not the project's ecosystem settings. For a settings-driven scan with a diff and a new-repo alert, use ecosystem_index_update.

ecosystem_refreshA

On-demand incremental refresh of the project's active ecosystem set.

Nothing refreshes the archive in the background; it changes only when this tool runs. For each active-set repo (top_n by stars) this probes GitHub once, writes a status snapshot, and re-queues a Stage 0 shallow summary only when the repo has new pushes; 404/403 mark the profile deleted/private.

Refresh does not run the re-queued shallow scans. When repos were re-queued, the response's hint field says how to run them; each result is written back with ecosystem_apply_shallow_summary.

ecosystem_scan_statusB

Fetch a single EcosystemScanRun by id.

ecosystem_scan_historyA

List recent scan runs ordered by started_at descending.

ecosystem_deep_review_requestA

Queue a deep-review for a repo and return the dispatch prompt.

Creates an EcosystemDeepReview row queued on the funnel (stage_status='queued'; the read-only status column is derived from it and reads 'queued'), and embeds a sub-agent prompt (5-section template + repo metadata) in the row's dispatch_prompt field. A background watchdog advances stage_status to shallow_failed (status derives to failed) after timeout_minutes if no report has been linked. The Leader is responsible for actually spawning the sub-agent (via the CC Agent tool; the session's implicit team is used automatically).

ecosystem_deep_review_statusC

Look up the most recent deep-review for repo_id.

ecosystem_deep_review_listB

List deep-reviews newest-first, optionally filtered by status.

ecosystem_deep_review_cancelA

Cancel an in-flight (stage_status='queued') deep-review.

Advances the row's stage_status to shallow_failed (the legacy status column derives to failed) with a cancellation note. The sub-agent is expected to observe the row state and shut down on its own.

ecosystem_tag_listA

List ecosystem tag dictionary entries.

Three layers of tagging are supported:

  • GitHub topics direct mapping (Layer 1)

  • Keyword/regex rules (Layer 2)

  • LLM sub-agent fallback (Layer 3)

This tool only returns the canonical tag dictionary (seeded at API startup). Use ecosystem_tag_apply_batch to actually apply tags to repos.

ecosystem_tag_apply_batchA

Apply Layer 1 + Layer 2 auto-tagging to a batch of ecosystem repos.

Layer 1 matches GitHub topics directly (confidence=0.95, source=github_topic). Layer 2 matches keyword rules against name+description+topics+owner (confidence=0.7, source=auto_rule).

Repos with fewer than 2 matched tags are flagged via needs_llm=True; callers should pass those into ecosystem_tag_dispatch_llm to spawn Layer 3 sub-agents.

If both repo_ids and repo_full_names are empty, the first repos in the database are processed.

ecosystem_tag_dispatch_llmA

Build a Layer 3 sub-agent dispatch plan for repos that need LLM fallback.

Returns a dispatch plan; the Leader is expected to spawn each sub-agent via the Agent tool using launch_call.params. Each sub-agent analyzes the repo and submits results via ecosystem_tag_apply_llm_result.

Concurrency is capped at max_concurrency (default 20) to limit token spend. Excess repos are returned in skipped_due_to_limit.

ecosystem_tag_apply_llm_resultA

Submit Layer 3 LLM tagging result from a sub-agent.

Only names in the canonical tag dictionary (ecosystem_tag_list) are applied; any other name is skipped without error and returned in skipped_unknown.

ecosystem_repo_tagsA

List all tags currently associated with a single ecosystem repo.

Returns each association with its confidence, source layer (github_topic / auto_rule / auto_llm / manual), and tag metadata.

ecosystem_summary_weeklyA

Generate the past-N-days ecosystem briefing as markdown.

Aggregates new / updated profiles, completed deep-reviews, archive counters and top star movers over the configured window. When save_report=True (default) the markdown is persisted via report_save with report_type='ecosystem-weekly'.

ecosystem_summary_by_tagA

List every repo carrying tag as a markdown table.

Each row contains stars / language / one-line summary plus a deep- review id when one exists. Rows are sorted by stars desc. Archived repos are excluded unless include_archived=True. By default each call also saves the markdown as a new report; pass save_report=False to only read it.

ecosystem_summary_top_nA

Top-N markdown table of ecosystem repos.

By default each call also saves the markdown as a new report; pass save_report=False to only read it.

ecosystem_summary_healthA

Platform self-check markdown: profile / scan / tag coverage / archive ratio.

By default each call also saves the markdown as a new report; pass save_report=False to only read it.

ecosystem_apply_shallow_summaryA

Stage 0 worker callback: write back a shallow summary OR report a failure.

Success path (default): pass shallow_summary (200-400 char Chinese markdown) and deep_review_id; the OS will persist the summary, advance stage_status -> shallow_done, and mark the deep_review row as completed.

Failure path: leave shallow_summary empty and pass error_kind, which routes the failure through the §3.1 classifier so the OS can decide whether to immediate-retry, mark deleted/private, or feed the self-learning loop. Valid error_kind values: http / agent_read / agent_timeout / json_parse / fetch_style.

ecosystem_shallow_queue_statusA

Show Stage 0 shallow-scan queue status for the active project.

Returns counts for active profiles, pending shallow scans, in-flight dispatches, terminal failures (shallow_failed), and deleted/private-flagged repos. The self_learning_pending map shows how many distinct repos have hit each failure class so far (a class becomes eligible for a recorded lesson once the count reaches 3).

Returns: {project_id, active_total, pending_shallow, in_flight, shallow_failed, deleted, private_now, concurrency, self_learning_pending}.

ecosystem_deep_review_request_batchA

Stage 1 — Queue architecture-analysis dispatches for tag-filtered candidates.

Pulls active+shallow_done profiles whose tag set covers tags (AND semantics), creates an EcosystemDeepReview row per candidate, and returns a list of DispatchIntent payloads for backend-architect sub-agents. Leader is responsible for actually spawning each agent via the Agent tool. Each agent eventually calls ecosystem_apply_architecture_md to write back.

ecosystem_apply_architecture_mdA

Stage 1 writeback — submit architecture_md OR report failure.

Success path (default): pass non-empty architecture_md (800-1500 字 Chinese markdown). The OS persists it, advances stage_status -> architecture_done, and marks the deep_review row completed.

Failure path: leave architecture_md empty and pass error_message; the OS advances stage_status -> architecture_failed so manual retry surfaces in the UI.

ecosystem_trigger_debateA

Stage 2 — Build debate dispatch payload (Leader still calls debate_start).

Validates that each repo_id has at least one architecture_done review, then returns a payload (suggested topic + roles + linked review_ids) so the caller can invoke the existing debate_start MCP tool. After debate_start returns a meeting id, call ecosystem_link_debate_meeting to write debate_meeting_id back onto each review row.

ecosystem_link_debate_meetingA

Stage 2 helper — link debate_start meeting id back to review rows.

Called immediately after debate_start succeeds. Writes debate_meeting_id onto every review in review_ids so meeting_conclude can list them in its ecosystem_writeback.

ecosystem_apply_debate_resultA

Stage 2 writeback — submit debate conclusion to advance to debated.

At least one of risks_md / learnings_md / integration_md must be non-empty. integration_recommendation is a short enum: integrate / reference / learn / skip.

ecosystem_mark_as_referenceA

Stage 3 reference path — add lifecycle:reference tag + advance to referenced.

Use when the debate concludes that the repo is worth keeping as an architectural reference but not integrated. The repo will appear highlighted in future searches as "已研究过" so the team avoids re-deep-scanning it.

ecosystem_start_integrationA

Stage 3 integrate path — build a task_create payload + tag the repo.

Adds lifecycle:integrated tag, advances stage_status, and returns task_payload (title / description / priority / horizon / tags), whose fields map one-to-one onto task_create's parameters. After task_create returns, call ecosystem_link_integration_task to write integration_task_id back onto the review.

ecosystem 不接管实施 — task ownership 由现有任务/团队系统接管。

ecosystem_link_integration_taskB

Stage 3 helper — link integration task id back to review row.

ecosystem_claim_shallowA

Claim the next queued repo for shallow scanning (stage_status='queued').

Atomic: only one worker gets each row; others get {"claimed": false}. The claim is a 60-minute lease: a row still unfinished after that can be claimed by another worker. The result includes repo_full_name, topics, description, owner, stars and last_commit_at, so no separate ecosystem_repo_get call is needed.

ecosystem_claim_reviewA

Claim the next shallow_done repo for quality review.

Finds stage_status='shallow_done' rows with no quality_score and no active claim. Returns the repo's shallow_summary so the reviewer can evaluate quality. The claim is a 60-minute lease: a row not reviewed by then can be claimed by another worker.

ecosystem_apply_quality_reviewA

Submit quality review result and release the claim lock.

Writes quality_score / quality_notes / reviewed_by / reviewed_at, clears claimed_by so other workers can pick up the next row.

ecosystem_release_claimA

Release a worker claim without submitting a quality review.

Use when a worker abandons a task (timeout, error). Clears claimed_by so another worker can pick up the row. Records reason in quality_notes.

ecosystem_quick_setupA

Record data-source and scan-profile rows for the project.

Creates one DataSource row per entry in sources and persists either the default ScanProfile or the merged custom_profile override. No scan reads these rows: ecosystem_index_update takes its query set, star floor and alert threshold from the project's ecosystem settings (Dashboard ecosystem settings panel), and only GitHub is scanned. So calling this does not change what the next index update discovers.

ecosystem_index_updateA

Trigger ecosystem index update — runs scanner + computes diff.

Maps to POST /api/ecosystem/index_update. Scan config comes from the project's ecosystem settings (min_stars gate, focus_topics queries — empty falls back to the built-in Claude-ecosystem query set, alert_max_new_per_scan threshold), then runs the full pipeline: gh search → classify active status → diff against DB → alert threshold check → (if dry_run=False) persist index_diff + status_changes. When dry_run=True, no writes touch ecosystem_repo_profiles / ecosystem_index_diffs / ecosystem_status_changes.

ecosystem_index_diff_latestA

Fetch the latest IndexDiff snapshot for the current project.

Maps to GET /api/ecosystem/index_diffs/latest. Returns the most recent diff row produced by a real ecosystem_index_update (dry_run=False) run. Dry-run previews are not persisted and therefore never appear here.

Returns: Diff available: {success: True, diff: {id, diff_type, new_count, reactivated_count, deactivated_count, stale_count, archived_count, markdown_summary, alerted, generated_at}}. No diffs yet (fresh project): {success: True, diff: None, message: 'No index diffs found yet.'}. Endpoint missing (the API answers 404): {success: False, error: 'P0.4 will implement', detail}. Other failure: {success: False, error, detail}. success semantics: True = call completed (diff may be None when the project has never run a non-dry index_update); False = API/endpoint error.

ecosystem_repo_eventsA

Return event history for a single ecosystem repo.

Four event types are recorded: discovered, topics_changed and stars_jumped (written by the scanner behind ecosystem_scan_periodic and ecosystem_index_update with dry_run=False) and status_changed (written by ecosystem_index_update with dry_run=False). ecosystem_scan and ecosystem_repo_manual_status record no events.

ecosystem_diff_periodA

Return a time-period diff computed dynamically from the per-repo event log.

Groups events by type to produce summary counts: new repos discovered, topics changed, stars jumped, status changed. Only scans that record events are counted (see ecosystem_repo_events). ecosystem_index_diff_latest instead returns the stored diff of the last ecosystem_index_update run.

ecosystem_rebuild_queries_from_reposA

Return a recap of all search queries that have discovered repos in this project.

Scans discovered_via_queries across all stored profiles and aggregates counts per query. Useful to audit which queries are most productive and which repos are multi-query hits.

ecosystem_repo_manual_statusA

Set (or clear) the human override on a repo's active status.

pinned keeps the repo permanently active regardless of scan results: it is excluded from the removed_from_query count in index_update diffs, and last_active_status stays active even when the fetcher misses it. Use it for high-value repos you always track. no_value records that the repo was reviewed and judged not worth tracking: last_active_status flips to manual_archived immediately. An empty status clears the override, and the repo goes back to being driven by scan results (active unless GitHub-archived).

workflow_listB

List CC ultracode/Workflow runs tracked by the OS observability layer.

planned_agent_count is a STATIC LOWER BOUND (literal agent() calls in the launch script), not a target. dynamic_nodes counts the fan-out nodes (pipeline / .map / while) whose width is only known at runtime, so a run with dynamic_nodes > 0 legitimately ends with agent_count > planned_agent_count - that is expected, not a miscount. planned_agent_count == 0 means no static parse was recorded (typically a run ingested by offline file reconcile), i.e. the plan is unknown rather than zero.

workflow_getA

Get a Workflow run's archive (totals + summary/result + per-agent telemetry).

Default response is a COMPACT projection (view="compact" + hint - trimmed, NOT missing fields). Compact keeps every scalar on the run, excerpts its result (400 chars) and summary (200 chars), and projects the agent rows down to identity / phase / cost / state plus the os_agent_id drill-down key.

Both views keep planned_agent_count and dynamic_nodes on the run. Read them together: planned_agent_count is the static lower bound (literal agent() calls), dynamic_nodes counts runtime-width fan-out nodes, so agent_count > planned_agent_count is expected whenever dynamic_nodes > 0.

workflow_reconcileA

Reconcile finished Workflow runs from disk into the OS (repair after OS was offline).

Scans ~/.claude/projects/<slug>/*/workflows/wf_*.json and ingests each run's full telemetry (tokens/duration/per-agent). Idempotent — safe to re-run.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

B3.3/5.0

Scored across 116 tools

Disambiguation2/5

Multiple ecosystem_* tools overlap heavily: ecosystem_scan vs ecosystem_scan_periodic vs ecosystem_index_update vs ecosystem_refresh vs ecosystem_quick_setup all touch scanning/indexing; ecosystem_search vs ecosystem_search_by_capability; ecosystem_summary_by_tag vs ecosystem_summary_top_n vs ecosystem_summary_weekly; and ecosystem_deep_review_request vs ecosystem_deep_review_request_batch vs ecosystem_deep_review_status vs ecosystem_deep_review_cancel. Task, meeting, and memory groups are mostly distinct, but the ecosystem cluster and broad search tools like unified_search vs memory_search create real misselection risk.

Naming Consistency3/5

All tool names use snake_case, and many carry domain prefixes (project_, team_, task_, meeting_, memory_, channel_, ecosystem_). However action order varies: noun_verb (project_create, team_list, meeting_create), verb_noun (unified_search, find_skill, verify_completion), and noun-only phrases (prompt_effectiveness, usage_attribution, context_resolve). It is readable but not a single predictable pattern.

Tool Count1/5

116 tools is extreme for an MCP server, well beyond the 50+ threshold for a poor score. Although the AI Team OS domain is broad, this many tools will overwhelm context, increase selection latency, and make maintenance difficult. Many ecosystem_* and memory_* tools could be consolidated or gated behind a smaller surface.

Completeness4/5

The surface covers the orchestration domain extensively: projects, teams, agents, tasks, memos, meetings, memory, channels, ecosystem pipeline, workflows, config, and health checks. Minor gaps exist, such as no standalone task_delete (only via project_delete), no report update/delete, and no event deletion, but these are workaroundable. Overall it is highly complete, with only minor lifecycle omissions.

Maintenance

ActivityMaintained
ResponsivenessSlow