harness-homies
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@harness-homieswhat are my currently running coding-agent sessions?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
harness-homies
A read-only MCP server that lists coding-agent sessions on your machine — Claude Code, OpenCode, Cursor, and Codex — in one place: what's running, its topic, its todos, its transcript.
It never writes to another agent's session and never sends messages.
Install
Claude Code (plugin)
/plugin marketplace add nikp29/harness-homies
/plugin install harness-homies@harness-homiesThis installs the MCP server (run via npx -y harness-homies) and the inspect-agents skill.
Agent Plugins clients
plugin/ is an Agent Plugins 1.0.0 package: plugin.json, mcp.json, and skills/inspect-agents/SKILL.md. Point any conformant client at that directory. The same directory also carries .claude-plugin/plugin.json and .mcp.json for Claude Code.
OpenCode
OpenCode plugins are JS modules and can't register MCP servers or skills, so set it up with config instead:
Merge
opencode/opencode.jsoninto~/.config/opencode/opencode.json(or a project'sopencode.json).Copy the skill:
cp -r plugin/skills/inspect-agents ~/.config/opencode/skills/
Codex
codex mcp add harness-homies -- npx -y harness-homiesCursor
Add an entry to ~/.cursor/mcp.json:
{ "mcpServers": { "harness-homies": { "type": "stdio", "command": "npx", "args": ["-y", "harness-homies"] } } }From source
pnpm install
pnpm build
claude mcp add --scope user harness-homies -- node /path/to/harness-homies/dist/cli.jsRelated MCP server: Hua PlanRelay
Tools
Tool | Args | Does |
|
| List sessions across one or all agents, live and historical, most recent first |
|
| Title, cwd, status, todos, message count |
|
| Paginated transcript, secrets redacted by default |
|
| Substring search over titles and recent transcript content |
agent is one of claude-code | opencode | cursor | codex. All tools are readOnlyHint: true.
Develop
pnpm test # builds, then runs the suite against dist/
pnpm typecheck
node dist/cli.js doctor # per-adapter availability + session countsNotes
Redaction is on by default for transcript text (API keys, tokens, PEM blocks, etc.). Pass
raw: trueto skip it.Claude Code: reads
~/.claude/sessions/*.json(live) and~/.claude/projects/**/*.jsonl(history).OpenCode: reads
~/.local/share/opencode/opencode.dbdirectly.Cursor: reads
~/Library/Application Support/Cursor/User/globalStorage/{conversation-search,state}.db. macOS only. Status is alwaysunknown— no running signal is available.Codex: reads
~/.codex/state_5.sqlite, falling back to scanning~/.codex/sessions/**/*.jsonldirectly if the DB schema doesn't match.All formats are undocumented by their vendors. Adapters degrade (fewer fields) rather than crash on an unrecognized record.
setupsubcommand and Cursor auto-registration are not implemented — register manually as shown above.The plugin configs pin
harness-homies@<version>. When releasing, bump the version inpackage.json,plugin/plugin.json,plugin/.claude-plugin/plugin.json, both MCP configs inplugin/,opencode/opencode.json, andsrc/server.ts.
Available Tools
4 toolsget_sessionGet session detailARead-only
Get a single session's title, cwd, status, todos, and message count. Use the agent and id fields from a list_sessions/search_sessions result. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| agent | Yes | Which agent this session belongs to | |
| session_id | Yes | The session's id, as returned by list_sessions/search_sessions |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description explicitly lists returned fields (title, cwd, status, todos, message count) which goes beyond the readOnlyHint annotation. It also states 'Read-only,' consistent with annotations. No contradiction or missing behavioral detail like error cases, but for a simple read operation this is sufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no fluff. The main purpose is front-loaded, followed immediately by the usage prerequisite. Every word earns its place; it is concise while packed with necessary information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with 2 parameters, no output schema, and no nested objects, the description covers the returned data, the source of parameters, and the read-only nature. Nothing critical is missing for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with descriptions for both agent and session_id. The description adds value by clarifying that these come from list_sessions/search_sessions results, giving origin context that helps the agent know exactly what values to pass.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action: 'Get a single session's title, cwd, status, todos, and message count.' This differentiates it from sibling tools like list_sessions (listing) and get_transcript (transcript). It also specifies the exact data returned.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides a direct prerequisite: 'Use the agent and id fields from a list_sessions/search_sessions result.' This implies a workflow but does not explicitly mention when not to use alternatives like get_transcript. Still, it gives clear context on how to obtain the required parameters.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_transcriptGet session transcriptARead-only
Fetch a session's transcript, paginated (most recent messages by default). Secrets are redacted unless raw is set. Use the agent and id fields from a list_sessions/search_sessions result. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| raw | No | Skip secret redaction. Off by default — use only when you trust the destination. | |
| role | No | Restrict to one role | |
| agent | Yes | Which agent this session belongs to | |
| limit | No | Max messages to return (default 50) | |
| cursor | No | Opaque pagination cursor from a previous call | |
| session_id | Yes | The session's id, as returned by list_sessions/search_sessions |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds substantial behavioral context beyond the readOnlyHint annotation: secrets are redacted by default and only exposed when raw is set; results are paginated with most recent messages first; and it repeats 'Read-only' for emphasis. The raw parameter's trust caveat is also surfaced. These are meaningful disclosures that help the agent predict side effects and data handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three short sentences with zero filler. The core action is front-loaded, followed by the critical redaction detail and then the input-source note. Every sentence earns its place and the structure is ideal for quick scanning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has six parameters, pagination, and role filtering, but no output schema. The description covers the essentials—purpose, default ordering, redaction behavior, and input provenance. It does not explicitly describe the response structure (e.g., list of messages with roles/content), but for a transcript fetch this is inferable. A minor gap regarding error conditions or rate limits exists, but overall it is sufficiently complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters (e.g., raw, limit, cursor, session_id). The description adds only a slight reminder to source agent and id from list/search, but that is largely redundant with the schema's session_id description. No new parameter meaning is introduced, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific verb (fetch) and resource (a session's transcript), and immediately adds pagination and default ordering. It distinguishes itself from siblings by focusing on transcript content, not session metadata or searching. The phrase 'Use the agent and id fields from a list_sessions/search_sessions result' also implies this tool is downstream of those, further clarifying its role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly instructs the agent to obtain agent and session_id from list_sessions/search_sessions, which is a clear prerequisite. However, it does not explicitly state when to prefer this tool over get_session or search_sessions, nor does it give an exclusion condition. The guidance is implied by the resource type, but not as explicit as the best examples.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_sessionsList agent sessionsARead-only
List coding-agent sessions on this machine across Claude Code, OpenCode, Cursor, and Codex — what's running, its topic, and its status. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | No | Restrict to sessions in this exact working directory | |
| agent | No | Restrict to one agent kind | |
| limit | No | Max sessions to return, per adapter | |
| running_only | No | Only include sessions with a known-live status |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotation readOnlyHint=true already discloses the read-only nature, and the description redundantly states 'Read-only.' The description adds a bit of behavioral context by describing the output fields (running, topic, status), but it does not disclose potential limitations like pagination, error behavior, or rate limits. Since the annotation covers the safety profile, the description adds minimal extra value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that front-loads the verb and resource, then provides the scope and return info. There is no redundancy or filler—every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by stating what fields are returned (running, topic, status), which is helpful. It does not explicitly mention pagination or limit behavior, but the schema already defines 'limit' as 'Max sessions to return, per adapter.' The description is adequately complete for a listing operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, so all 4 parameters (cwd, agent, limit, running_only) are already well-documented in the schema. The tool description does not add any semantic meaning beyond what the schema provides, so it meets the baseline of 3 without further contribution.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List'), a precise resource ('coding-agent sessions on this machine'), and enumerates the agent types (Claude Code, OpenCode, Cursor, Codex). It also clarifies what is returned ('what's running, its topic, and its status'), which distinguishes it from sibling tools like get_session that fetch a single session. This is clear and specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage (e.g., use this to get a broad overview of sessions on this machine) but does not explicitly state when to prefer this over search_sessions or get_session. There is no mention of alternatives or when not to use it, leaving some inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_sessionsSearch agent sessionsARead-only
Search session titles (and recent transcript content) for a substring across one or all agents. Read-only, case-insensitive.
| Name | Required | Description | Default |
|---|---|---|---|
| agent | No | ||
| limit | No | ||
| query | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already set readOnlyHint=true, and the description reinforces this by stating 'Read-only'. It adds value by mentioning 'case-insensitive' and the dual search targets (titles and recent transcript content), which are non-obvious behaviors not captured in the schema or annotations. No contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one sentences with no fluff, front-loading the main purpose. It includes 'Read-only' and 'case-insensitive' efficiently. The only minor issue is the parenthetical 'and recent transcript content' could be clearer, but it does not waste words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple enough that the description mostly suffices, but it lacks parameter semantic details that are critical due to 0% schema coverage. An agent might misuse the 'agent' filtering or misunderstand 'limit' behavior. Without output schema, the return format is also undocumented, but that is less critical for a search tool. On balance, it falls short of full completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must explain parameters, but it does not. It omits meaning for 'agent' (which has an enum), 'limit', and 'query'. The description only hints at search functionality but provides no detail on how these parameters affect behavior, such as the default agent scope or limit interpretation. This is a significant gap given the low coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states 'Search session titles (and recent transcript content) for a substring', combining a specific verb with clear resources. It also distinguishes itself by mentioning read-only and case-insensitive behavior, setting it apart from siblings like list_sessions or get_session.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage (searching across sessions) but does not explicitly compare to siblings. It lacks guidance on when not to use this tool, such as when full transcript retrieval is needed (get_transcript) or when listing all sessions without search is desired (list_sessions). However, the search scope is clear enough for an agent to infer common use cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
get_session - First observed
get_transcript - First observed
list_sessions - First observed
search_sessions
TDQS
Scored across 4 tools
Each tool targets a distinct concern: listing all sessions, retrieving one session's metadata, fetching a transcript, and searching across sessions. The overlap between list_sessions and search_sessions is minimal and clearly differentiated by behavior.
All tool names follow a consistent verb_noun snake_case pattern: list_sessions, get_session, get_transcript, search_sessions. The naming is predictable and instantly readable.
Four tools is well-scoped for a read-only session inspection server. Each tool covers a distinct aspect of the domain without redundancy or missing essentials.
For a read-only monitoring tool, the surface is complete: discover sessions, inspect one session, read its transcript, and search across sessions. No obvious lifecycle operations are expected here, so there are no meaningful gaps.
Related MCP Connectors
Read-only access to your CodeMouse accounts, repositories, and AI pull-request reviews.
Local-first memory and continuity for AI coding agents. No cloud backend; optional hosted lane.
Read-only AI coding tools for change verification, release readiness, capacity, and guidance.
Read-only MCP access to authorized Vocci sessions, notes, files, and memory search.
Related MCP Servers
- AlicenseAqualityBmaintenanceProvides a read-only interface to audit and continue coding agent sessions by extracting plans, intents, and edit authorship from history across multiple agents (Claude, Codex, OpenCode, Antigravity, Pi) via MCP, CLI, and Python SDK.1827 PyPI3MIT
- AlicenseAqualityAmaintenanceProvides a workspace-safe, read-only bridge between browser-based AI planning/review and local coding agents, enabling structured plan, execution summary, and review handoffs without granting shell, file write, or Git push access.11MIT
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to search and retrieve past coding-tool conversations across Claude Code, Claude Desktop, Codex, and Cursor through a local, read-only index.1MIT
- FlicenseNot gradedqualityCmaintenanceEnables read-only access to Kilo Code and Opencode conversation history stored in a local SQLite database, allowing users to list projects and sessions, read messages, and perform full-text keyword searches across all messages.-