steward
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@stewardWhat's the current situation across workspaces?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Steward(管家账本)
The ledger behind the floating orb: every workspace reports what it did, the orb reads the report. An append-only event log plus four MCP tools — no chat transcripts, no models, no dependencies beyond the standard library.
悬浮球(管家)背后的账本:每个工作区干完活交一行回执,管家随时汇报。只增不改的事件表 + 四个 MCP 工具,纯标准库。
Platform support: macOS — supported (developed & verified on this machine) · Windows — unverified (paths are cross-platform in the codebase, but no real-machine testing yet).
中文说明 · License: AGPL-3.0-only
Why a ledger instead of a transcript
The orb has to answer "what is going on right now?" without holding every conversation in context. Events do that: who, what, outcome, file, at most three numbers. A transcript grows until the context window dies; a ledger stays bounded and stays answerable.
The injectable column is the gate: progress chatter is written down but never reaches the
brief. Only events that answer a question get injected.
Related MCP server: factlog
Tools (MCP)
Tool | Input | What it does |
|
| A workspace's one-line receipt. Returns |
|
| Pull event detail on demand — the orb asks for specifics. |
|
| What is going on right now: recent key events, unresolved failures, which actors are moving, plus |
|
| Page / project / selection changes are events too — that is how the orb knows where you are after you switch. |
The contract
Workspaces never write the ledger directly. They call
report_doneand hand back a receipt. One writer, one schema, no two code paths disagreeing about what an event means.Append-only. No updates, no deletes in the tool surface; history is the point.
Bounded injection.
metricskeeps the first three keys;injectable=falsemarks progress events that belong in the log but not in the brief.
Wire it up
# stdio MCP server, no third-party packages
python3 mcp_server.pyStorage: ~/Documents/ShadowRoom/_steward/ledger.db (SHADOWROOM_STEWARD_DB overrides it).
Delete that file to reset the ledger; nothing else references it.
Development
python3 -c "import ast,sys; ast.parse(open('mcp_server.py').read())"
printf '%s\n' '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' | python3 mcp_server.pyAvailable Tools
4 toolsfocus_changeC
记录焦点变化(切页面/切工程/选中片段)。
| Name | Required | Description | Default |
|---|---|---|---|
| actor | No | ||
| object_id | No | ||
| workspace | No | ||
| object_type | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden, and it only says a focus change gets recorded. It does not say whether recording is idempotent, whether duplicate events are deduplicated, whether it requires an active session/workspace, or what happens to the record afterward. For a mutation/logging tool with zero annotation coverage this is a meaningful gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler, and the examples earn their place by narrowing scope. It is efficient, though the brevity contributes to the under-specification penalized elsewhere rather than being a virtue by itself.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 4 undocumented parameters, no annotations, and no output schema, the description would need to do much more work. As written it leaves field meanings, required-vs-optional expectations, and behavioral guarantees entirely to the agent's guesswork.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Four parameters (actor, object_id, workspace, object_type) have 0% schema description coverage, and the description supplies no semantics for any of them. The examples hint at 'workspace' and 'object' concepts but never map them to named fields or formats, so the schema gap is left entirely unmitigated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('记录焦点变化') and disambiguates with concrete examples (page switch / project switch / selected fragment), which is enough to separate it from siblings like recent_events or situation_report. It stops short of naming those siblings explicitly, so it is clear but not maximally differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical examples imply the kind of events to log, but there is no statement of when to call this versus report_done or situation_report, nor any prerequisite or timing guidance. The agent must infer the trigger condition entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recent_eventsC
按 id 拉事件明细(管家要细节时才用)。
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| since_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full behavioral burden. It only implies a read operation ('拉事件明细') and says nothing about permissions, return format, pagination, or what happens with the limit parameter. Key behavioral traits are undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with a parenthetical hint, so it is very concise. However, the parenthetical is cryptic and the overall terseness may leave the reader unsure of the exact scope, slightly undermining the structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations, no output schema, and two parameters at 0% schema coverage, the description is incomplete. It does not explain what 'events' are, what the return contains, or how the parameters work, leaving significant gaps for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so both parameters (limit and since_id) are undocumented in the schema. The description only vaguely references 'id', which does not clarify the meaning of either parameter, especially limit, and provides no syntax or interaction guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('拉' fetch) and resource ('事件明细' event details) with a scoping condition ('按 id'). It is clear what the tool does, but it does not distinguish this tool from its siblings (report_done, situation_report, focus_change), so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical '管家要细节时才用' provides a usage condition (use only when the steward needs details), which implies a context. However, it names no alternatives or exclusions, leaving the agent to infer when this tool should be chosen over siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
report_doneC
一个工作区干完活后的回执(谁、什么、结果、文件、≤3 个数字)。
| Name | Required | Description | Default |
|---|---|---|---|
| file | No | ||
| type | No | ||
| actor | Yes | ||
| job_id | No | ||
| metrics | No | ||
| outcome | No | ||
| object_id | No | ||
| injectable | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full behavioral burden. It says the tool creates a receipt and limits numeric metrics to three, but it does not disclose side effects, authentication needs, idempotency, or what the tool returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence with no filler, and the field list is front-loaded in parentheses. It is under-specified for an 8-parameter tool, but it is structurally efficient rather than verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 8 parameters, a nested metrics object, no annotations, and no output schema, the description is far too sparse. It gives only a high-level receipt concept and does not explain parameter behavior, usage, or output expectations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 8 parameters and 0% schema description coverage, the description needs to compensate. It loosely maps several fields (who, what, result, file, numbers) but omits job_id, injectable, object_id, and type distinctions, leaving many parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states that the tool produces a receipt after work in a workspace and lists the kinds of content it captures (who, what, result, file, up to three numbers). It does not use a clear action verb, and it never distinguishes this tool from siblings like situation_report or recent_events.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'after finishing work in a workspace' implies a usage condition, so some context is present. However, there is no explicit guidance on when to choose this tool over the sibling reporting tools, and no exclusions or prerequisites are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
situation_reportC
随时汇报:最近关键事件、未解决的问题、活跃的工作区。
| Name | Required | Description | Default |
|---|---|---|---|
| workspace_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does not state whether this is a read-only aggregation, whether it requires authentication, how the report is assembled, or any cost/latency characteristics — a notable gap for a summarizing tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact line that front-loads the purpose and enumerates the three content categories. Nothing is wasted, though the brevity comes at the cost of the missing information noted elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with a parameter, no output schema, and no annotations, the description should explain the parameter's effect and the shape of the returned report. It does neither, leaving the agent unable to predict what a call returns or how workspace_id changes results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one parameter, workspace_id, with 0% schema description coverage, and the description never mentions it directly; '活跃的工作区' refers to report content rather than the parameter's role. It also doesn't clarify the default behavior when no workspace_id is supplied (0 required parameters).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names concrete content the report covers (recent key events, unresolved issues, active workspaces), which goes slightly beyond restating the name 'situation_report'. However it never states the verb/resource clearly, and it does not distinguish this from the sibling recent_events, which appears to cover a subset.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no indication of when to invoke this tool versus report_done, recent_events, or focus_change. The phrase '随时汇报' (report at any time) implies a general-purpose status query but gives no trigger conditions or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
focus_change - First observed
recent_events - First observed
report_done - First observed
situation_report
TDQS
Scored across 4 tools
report_done creates a completion receipt, situation_report summarizes current status, recent_events fetches event details by id, and focus_change records a focus shift. The purposes are mostly distinct, though report_done and situation_report both involve reporting and could be confused in edge cases.
All names use snake_case and are readable, but they follow mixed patterns: report_done is verb-participle, recent_events is adjective-noun, while situation_report and focus_change are noun-noun. There is no single predictable verb_noun convention.
Four tools is a reasonable, well-scoped set for a lightweight steward assistant. Each tool appears to earn its place, though the surface is minimal enough that it may feel slightly thin for full workspace tracking.
The set covers reporting completion, summarizing status, retrieving event details, and recording focus changes. However, it lacks explicit workspace lifecycle operations and event discovery beyond lookup by known id, which may cause dead ends for agents.
Maintenance
Related MCP Connectors
Hosted MCP memory and agent control plane for durable conversations, jobs, and operations.
- OctopadOAuthapp.octopad
The back-office workspace for your team's AIs: tasks, knowledge and context shared over MCP.
Hosted MCP messaging across owners, tools, and machines, with readable transcripts.
Control plane for autonomous software labor. Agents claim objectives over MCP with audit trail.
Related MCP Servers
- AlicenseBqualityAmaintenanceLocal MCP event log for multi-agent workspaces, enabling structured event publishing, cursor-based polling, and status checks via SQLite-backed sidecar.10MIT
- FlicenseNot gradedqualityBmaintenanceA shared, append-only fact log that enables AI agents to coordinate on a workspace via MCP, CLI, REST, and web interfaces.-
- FlicenseNot gradedqualityBmaintenanceCaptures AI-assisted work into a searchable ledger and exposes it via MCP for querying past conversations and decisions.-

roxabi-senseofficial
AlicenseBqualityBmaintenanceLocal workstation attention journal that tracks focus, idle, and agent sessions, exposing timeline data via MCP for AI agents to query current or past activity.5AGPL 3.0