mcp-ashigaru
This server runs background dev workflows on GitHub repos: start a coding run for an issue, check its progress, and promote an approved PR.
Start a dev run (
work_ticket): clone a crunchtools repo, work a GitHub issue with Claude, run repo gates, and open a PR; returns arun_idimmediately.Check run status (
status): get a digest of the run's phase, recent agent actions, gate result, and PR URL.Promote a PR (
promote): squash-merge and ship a reviewed PR, but only with the human-approval token from a confirmed Signal approval.Model escalation is built in: runs start on Sonnet and escalate to Opus only if gates fail.
Notifications and heartbeat channels (webhook, Matrix, shell command) fire on run milestones and activity.
Allows creating pull requests from GitHub issues, running repository gates, and promoting to production on approval.
Reports live CI/build check status for pull requests.
Handles human approval prompts for promoting changes to production.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@mcp-ashigaruwork crunchtools/ashigaru #12"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
ashigaru
Issues in, fixed bugs out. Ashigaru is a set of reusable GitHub Actions workflows that watch a repository's issues, sort bug-class work from feature work, and fix the bugs: it opens the pull request, answers the code review, and lets GitHub merge when the required checks pass. Feature requests are labeled and left for a maintainer. It runs on GitHub-hosted runners with Claude Code, on a Claude subscription token; there is no server to operate.
Named for the ashigaru (足軽), the foot soldiers a daimyo sent into the field.
Capabilities
Triage: every new issue is read against the code and labeled. Bug, documentation and performance issues go to the code agent; features wait for a person. An issue from someone without write access is labeled too, but is fixed only when a maintainer says so.
Code: a
ready-to-codeissue becomes a pull request fromashigaru/issue-Nwith auto-merge on. The repository's own pre-commit hooks run before the pull request opens, and failures go back to the same agent session.Fix: after each review of a pull request the code agent opened, every finding gets an answer in its thread and failed checks get fixed, for at most five rounds.
Sweep: every hour the oldest untriaged issues across enrolled repositories are queued for triage, and up to five waiting fixes are started, so a backlog drains without spending the model budget.
Release: a repository that asks for it with a topic gets its merged work released: a release pull request from the changelog, then the tag and the GitHub Release, so a fix does not sit merged and unshipped.
Review: a pull request from an outside contributor gets one first-pass comment for the maintainers, and everything waiting on a person is listed in a single inbox issue.
Authority and limits: the job that runs the model holds no credential that can write; the job that writes runs no model. One organization variable stops everything.
Related MCP server: devflow-mcp
Quick Start
Organization, once:
Create a GitHub App with repository permissions Contents, Issues and Pull requests set to read and write, install it on the organization, and store its id and private key as the secrets
ASHIGARU_APP_IDandASHIGARU_APP_KEY.Run
claude setup-tokenand store the result as the secretCLAUDE_CODE_OAUTH_TOKEN.Set the variable
ASHIGARU_ENABLEDtotrue.Require your CI and review checks in a branch ruleset the app cannot bypass, and allow auto-merge.
Each repository:
mkdir -p .github/workflows
curl -fsSL https://raw.githubusercontent.com/crunchtools/ashigaru/v2/examples/ashigaru.yml \
-o .github/workflows/ashigaru.ymlDetails and the full list of inputs are in docs/enrollment.md.
Documentation
Page | What it covers |
Categories, the confidence threshold, which labels mean what | |
The two-job split, the pre-commit inner loop, what is refused | |
What triggers a round, how findings are answered, the round limit | |
Backlog selection and the hourly budget | |
Turning releases on per repository, how the version is chosen and rewritten, the releases record | |
First-pass review of outside pull requests, and the inbox issue | |
Who holds which credential, untrusted input, the switch and the caps | |
Organization setup, enrolling a repository, inputs, releasing |
Development
A developer machine needs git, podman, pre-commit and gh.
pre-commit install
podman run --rm -v .:/src:Z -w /src docker.io/library/python:3.12-slim \
bash -c 'pip -q install pytest==9.1.1 pyyaml==6.0.3 ruff==0.16.9 && ruff check . && pytest -q'
podman run --rm -v .:/repo:Z -w /repo docker.io/rhysd/actionlint:1.7.12Credits
The design follows Fullsend: labels as the state, a triage gate that routes bugs to a code agent and parks features, and deterministic scripts that hold push and merge authority. The label names are theirs, so a repository can move between the two.
License
AGPL-3.0-or-later
Available Tools
3 toolspromoteA
GATED. Promote a reviewed PR to production via the repo's deploy path.
Refuses unless approval_token matches the human-approval marker recorded
for this PR (set only through a confirmed Signal approval). This is the hard
gate on the one irreversible action — the coding agent never reaches it.
| Name | Required | Description | Default |
|---|---|---|---|
| pr | Yes | PR number to promote. | |
| approval_token | Yes | the approval marker from the confirmed human gate. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, but description fully discloses gating, token requirement (set only via Signal), and irreversible nature. Adds critical behavioral context beyond schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four efficient sentences, front-loaded purpose, no filler. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Simple tool (2 required params, output schema exists). Description explains purpose, gating, and token source comprehensively, leaving no ambiguity for use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers 100% of parameters with clear descriptions. Description does not add new info about each parameter beyond what schema already provides, so baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states verb 'promote' and resource 'PR to production'. Distinguishes from siblings 'status' and 'work_ticket' which are non-promotion tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly notes gating (requires approval_token) and refusal condition. Implies when to use: after review and human approval. Does not mention alternatives but siblings are unrelated, so no confusion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
statusA
Return an on-demand digest of a run: current phase, recent agent actions, gate result, and PR URL. This is what Kagetora answers from when you ask "what's the runner doing?".
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided. Description implies a read operation and lists output contents, but does not disclose behavioral traits like side effects, auth requirements, or performance impact.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with purpose and contents, followed by a helpful analogy. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a single simple parameter and an output schema, the description adequately explains the return value contents. Could mention constraints (e.g., run must exist) but not necessary for this simple tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema has one required string parameter 'run_id' with 0% schema description coverage. The description does not describe the parameter, leaving the agent to infer from context. While obvious, it does not compensate for the low coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it returns a digest of a run with specific components (phase, actions, gate result, PR URL). Differentiates from siblings 'promote' and 'work_ticket' by focusing on status retrieval.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'on-demand digest' and gives a user-friendly context ('what's the runner doing?'). No explicit when-not or alternatives, but sibling tools are clearly different.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
work_ticketA
Start a dev run: clone , fix GitHub issue # with Claude (Sonnet), run the repo's quality gates, and open a PR. Returns a run_id immediately; the run continues in the background. Poll status(run_id) for progress.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | Yes | crunchtools repo name (e.g. "rotv"). | |
| issue | Yes | GitHub issue number to work. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description fully carries the burden. It details the workflow (clone, fix, quality gates, PR), immediate return of run_id, background execution, and how to poll for progress. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, each carrying essential information: first sentence explains the action, second describes return behavior, third instructs on monitoring. No unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema (implied by context signals) and the description mentions the return value (run_id). It covers the complete workflow and expected behavior for a background dev run tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with both parameters described. The description adds minimal extra meaning beyond the schema, restating that repo is a 'crunchtools repo name' and issue is a 'GitHub issue number'. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description specifies the action (start a dev run), the resources (repo and issue), and the steps (clone, fix, quality gates, PR). It clearly distinguishes from sibling tools 'promote' and 'status' by mentioning polling status separately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use: to start a new dev run for a given repo and issue. It also suggests using status(run_id) to monitor progress, but does not explicitly state when not to use this tool or mention alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.0- First observed
promote - First observed
status - First observed
work_ticket
TDQS
Scored across 3 tools
Each tool serves a unique purpose: starting a run, monitoring status, and promoting to production. There is no overlap in functionality.
All names are simple and readable, but there is a slight inconsistency: 'promote' and 'status' are single words while 'work_ticket' is a compound. However, the style is uniform overall.
With only 3 tools, the server feels minimal for a deployment pipeline. While it covers the essential steps, additional tools like cancellation or listing runs would be expected for broader utility.
The tool set covers the core workflow (start, check, promote), but lacks common operations such as canceling a run, retrying, or viewing history. The surface is functional but not comprehensive.
Maintenance
Related MCP Connectors
An MCP server that gives your AI access to the source code and docs of all public github repos
A MCP server built for developers enabling Git based project management with project and personal…
Create, deploy, and operate MCP servers directly from your GitHub repositories.
Devopness MCP server for DevOps happiness! Empower AI Agents to deploy apps and infra, to any cloud.
Related MCP Servers
- FlicenseAqualityDmaintenanceAn MCP server that automates the full software development lifecycle through an AI-driven TDD state machine. It handles everything from task decomposition and test-driven development to integration testing and automated pull request creation.4-
- AlicenseNot gradedqualityDmaintenanceA production-ready MCP server that provides AI assistants with comprehensive GitHub developer tooling including PR analysis, code review, changelog generation, dependency auditing, commit summarization, and refactoring suggestions.8 npmISC
- AlicenseNot gradedqualityDmaintenanceAn MCP server that enables AI agents to directly manage GitHub repositories, including PRs, issues, and code search, using natural language.MIT
- AlicenseNot gradedqualityAmaintenanceAn MCP server that lets Claude Code manage GitHub issues, branches, and pull requests through natural language, automating the full development workflow from planning to closing.55 npm2MIT