legwork-mcp
Provides tools for searching GitHub repositories, evaluating repository metadata such as stars, license, recency, and MCP availability, and cloning/installing selected open-source repositories as MCP tools.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@legwork-mcpbuild an MCP server for owner/repo"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Legwork
Your AI finds the open-source tool it needs on GitHub. Legwork installs it safely.
Add Legwork to Claude, Cursor or any MCP client once. When you ask for something your AI has no tool for, it searches GitHub, picks a repo, and asks you to approve installing it. Legwork then checks the code for malware, builds a working tool for it in a sandbox, tests it, and hands it over. It can only touch the folders you allow.
claude mcp add legwork -- uvx legwork-mcp hub --allow-read ~/DownloadsThen just ask: "Pull the risk table out of ~/Downloads/launch-plan.pdf."
Claude finds pdfplumber, you
approve the install, and it reads the table. Or build one tool yourself:
uvx legwork-mcp owner/repo.
Why not just ask your AI to install it?
Most repos have no MCP server to install. Your AI would have to write one on the spot, untested, every time. Legwork writes it once, proves it works with a self-test, and caches it so the next person doesn't pay for it.
Whatever your AI installs runs with full access to your machine: your SSH keys, your files, your environment variables. That's how the two malware repos below would have got in. Legwork scans the code first, and every tool it installs runs sandboxed, seeing only the folders you grant.
It works across your tools: the same install works in Claude Code, Claude Desktop and Cursor.
Related MCP server: github-repo-intel-mcp
Proof, not a pitch
On 2026-09-26 we ran Legwork against the 10 most-starred AI repos created on GitHub in the previous week, taken exactly as search ranked them, with nothing skipped for being hard:
Outcome | Repos |
✅ Built a working MCP server | 3 — an npm CLI, a Python library, a Go binary |
↩️ Refused, with the right reason | 5 — two Android/macOS apps, a desktop app, a repo with no usage docs, a tool that needs its own API keys |
🛑 Blocked as malware | 2 |
Two of those top-10 "AI tools", about 700 stars each and three days old at the time, carried the same byte-identical obfuscated dropper under different file names. Legwork's pre-install scan refused both in about a second. Nothing from them was installed or run. Stars are not a trust signal.
Full write-up, including what failed along the way and what we fixed: docs/designs/legwork-trending-trial-2026-09-26.md.
A second, larger run on the next 22 new trending repos, with two models: gpt-5 built 12 and refused 10 with reasons, with no crashes; OpenRouter's free Nemotron built 4. That run also caught two false malware alarms and a gap where only published packages could be installed, both fixed. Write-up.
Let your AI find its own tools (hub)
legwork hub is one MCP server that gives your AI five tools: find_tools,
install_tool, install_status, list_installed_tools and use_tool.
claude mcp add legwork -- uvx legwork-mcp hub --allow-read ~/DownloadsFor Claude Desktop or Cursor, add this to the MCP config (for Claude Desktop, while the app is quit):
{ "mcpServers": { "legwork": { "command": "uvx",
"args": ["legwork-mcp", "hub", "--allow-read", "/Users/you/Downloads"] } } }You approve every install: your client shows the repo and the folders it asks for.
The hub's flags are the ceiling. An install can ask for the folders you allowed or less, never more, and no network unless you started the hub with
--allow-net. Secret folders (~/.ssh,~/.aws, ...) are refused outright.Installed tools appear by name (
pdfplumber__extract_tables), anduse_toolworks in clients that don't refresh their tool list. Installs are remembered across restarts.Search results tell your AI what matters: stars, license, how recently the repo was updated, whether it's in the Legwork cache, whether it already has an MCP server, and a warning on very new repos (the malware we caught was three days old). Descriptions are marked as untrusted text.
Repos in the public cache install in about a minute with no API key. Others need a model key (see below), and building takes 1–5 minutes.
legwork find "extract tables pdf" runs the same search from your terminal.
Quick start
You need macOS or Linux, uv, and any OpenAI-compatible chat-completions endpoint. Legwork itself is free; the only cost is your own model usage, which is 1–3 calls per build.
The Linux sandbox is bubblewrap:
sudo apt install bubblewrap python3-venv # Debian/Ubuntu
sudo dnf install bubblewrap # FedoraUbuntu 24.04 and later block the user namespaces bubblewrap needs, unless a
program's AppArmor profile allows them. Allow them for bwrap only (this
is what Ubuntu does for Chrome and Flatpak), rather than switching the
restriction off system-wide:
sudo tee /etc/apparmor.d/bwrap <<'EOF'
abi <abi/4.0>,
include <tunables/global>
profile bwrap /usr/bin/bwrap flags=(unconfined) {
userns,
include if exists <local/bwrap>
}
EOF
sudo apparmor_parser -r /etc/apparmor.d/bwrapIf anything's missing, Legwork stops and prints these instructions.
Tested on every commit: macOS on Apple silicon and Intel; Ubuntu on x86-64 and ARM64; Fedora; Debian. Windows: not supported natively. WSL2 should behave like Ubuntu (install bubblewrap) but hasn't been tested.
export LEGWORK_LLM_ENDPOINT=https://api.openai.com/v1
export LEGWORK_LLM_API_KEY=... # your own key; never a CLI flag
export LEGWORK_LLM_MODEL=gpt-5
uvx legwork-mcp 2akouwu/reverifyA build takes 1–3 minutes and ends with the line to connect it:
Built an MCP wrapper for 2akouwu/reverify on attempt 2 of 3.
What it wraps: Verify claims about binaries using Reverify's deterministic tools (wraps `reverify verify --json`)
Connect it to Claude Code:
claude mcp add reverify -- ~/.local/bin/uvx legwork-mcp serve 2akouwu/reverifyIt also prints an mcpServers block for Claude Desktop, Cursor and other
clients. For Claude Desktop, add it to claude_desktop_config.json while
the app is quit: the running app writes its settings back on exit and
drops edits it didn't make. Tested with Claude Code, Claude Desktop and Cursor.
For a permanent legwork command: uv tool install legwork-mcp.
Repos already in the public cache need no model call and no API
key. uvx legwork-mcp 2akouwu/reverify works as is.
Model providers
Any OpenAI-compatible chat-completions endpoint works. Put three lines in a
private file (chmod 600) and source it before building:
Provider |
|
| Cost |
OpenAI |
|
| Your API credits |
Anthropic |
|
| Your API credits |
OpenRouter |
|
| Free models, rate-limited |
Ollama (local) |
|
| Free; 14 GB download, ~16 GB RAM |
LEGWORK_LLM_API_KEY is your key (for Ollama, any non-empty value).
Ollama: start it with a larger context window, or it silently cuts
Legwork's prompt to ~2,000 tokens and the model never sees the
instructions: OLLAMA_CONTEXT_LENGTH=32768 ollama serve. Legwork prints
this tip when it detects Ollama.
OpenRouter free models come and go and are often busy (HTTP 429). If one keeps failing, pick another from their list.
Measured 2026-09-27, same four repos on each (reverify, golive-skill, AnyJev, phone-harness), counting wrappers built:
Model | Built | Notes |
gpt-5 | 4 / 4 | all on the first attempt |
Nemotron 3 Super (OpenRouter, free) | 3 / 4 | |
gpt-oss:20b (Ollama, local) | 3 / 4 | slowest: 1–3 min per reply on an M5 |
Claude Sonnet 5 | 1 / 4 | the most cautious: refused three, citing third-party accounts, a paired iPhone, and "a research pipeline" |
A refusal isn't a crash: Legwork reports the model's reason and stops without installing anything.
How it works
Cache. If the repo is in the public cache, Legwork uses that wrapper instead of steps 3–4. It still clones and scans the repo, scans the cached wrapper too, and installs and self-tests it locally; if the self-test fails, it writes a fresh wrapper.
--no-cacheskips the cache.Fetch. Checks the repo is public and reachable, then shallow-clones it.
Scan. Statically scans every Python and JavaScript/TypeScript file for obfuscated payloads: XOR-decoded byte arrays, computed imports, javascript-obfuscator output,
evalof decoded strings, code hidden off-screen behind whitespace. A hit stops everything before install.Read. The README, plus setup docs it links to (
install.md,docs/getting-started.md, …), capped at 50KB with install and usage sections kept first.Write. Your model writes the install command and a Python MCP wrapper with a self-test, or refuses with a specific reason: no programmatic entrypoint, needs hardware, needs its own credentials, needs a toolchain the sandbox doesn't have, and so on.
Install, sandboxed, network on.
Self-test, sandboxed, network off. A failure goes back to the model and it tries again, up to 3 attempts and 30 minutes. Attempts share one download cache, so a retry doesn't re-download multi-GB dependencies like PyTorch.
Serve.
legwork serve owner/reporuns the wrapper as an MCP server over stdio, sandboxed, network off.
Permissions: locked down unless you say so
A served tool sees none of your files and has no network. Grant exactly what a tool needs when you add it to your client:
# Let markitdown read one folder (read-only):
claude mcp add markitdown -- uvx legwork-mcp serve microsoft/markitdown --allow-read ~/Downloads
# A tool that fetches web pages:
uvx legwork-mcp serve owner/repo --allow-net--allow-read is repeatable and always read-only. Some folders can't be
granted even on request: /, your whole home folder, and places
credentials live (~/.ssh, ~/.aws, ~/.config, ~/.gnupg, keychains,
browser cookies). With --allow-net, the tool still can't reach local
sockets such as your SSH agent or Docker.
On macOS, the first time a tool reads a protected folder (Downloads,
Documents, Desktop), macOS itself asks whether uvx may access it. That's
the operating system's own privacy check, on top of Legwork's grant; allow
it once.
What "built" means
The self-test calls the simplest documented command and checks the shape of its result. So "built" means the wrapper starts and its simplest tool works. It does not mean every tool works. In the trial:
golive-skill's wrapper exposes its read-only commands, not its deploy flow, which needs accounts.
AnyJev's decision tool needs a running vLLM server, which the self-test didn't have.
Tools that must reach the network at run time won't work under serve,
because the network is off.
Risks, stated plainly
Legwork runs code you didn't write: the target repo's install steps and a wrapper an LLM wrote from a README a stranger wrote. What contains it:
Install time is the biggest risk. Package installs and
install.shscripts run with network on, because they have to. The sandbox limits them to their own working folder: no access to your home directory (SSH keys, cloud credentials, dotfiles), no writes outside the folder, and an environment with none of your variables (including your model key). A malicious package can still misbehave inside that folder and reach the network during install.The scan covers the repo's source, not what installers download. magpie, for example, installs with the vendor's
curl … | sh. For downloaded code, the sandbox is the protection.Prompt injection is not mitigated. A README can steer the model into writing a wrapper that does something other than what you asked for. The wrapper still runs sandboxed with network off, which bounds the damage; it doesn't prevent a wrong or misleading tool.
Linux on architectures other than x86-64 and ARM64: install code can reach "abstract" Unix sockets (such as an X11 display), because the seccomp filter that blocks Unix sockets during install only covers those two. The run phase has no network and isn't affected.
The scanner is heuristic. It catches the obfuscation patterns seen in real payloads so far, and a determined author can get past it.
The sandbox is sandbox-exec on macOS and bubblewrap on Linux, with the
same rules on both: the system read-only, your home directory hidden,
writes only in the build's own folder, no local sockets (SSH agent, Docker,
display servers), only the system services builds need, and network only
during install. If the sandbox isn't available, Legwork
refuses to run. It never falls back to running unsandboxed.
Commands
Command | What it does |
| One MCP server through which your AI finds, installs and uses tools, within the limits you set |
| Search GitHub for tools that do something, with facts to choose by |
| Build a wrapper (also |
| Run the built wrapper as an MCP server over stdio. Needs no model key. |
| Write the build as a public-cache entry, ready for a PR (below) |
| Free disk space: remove old and failed builds, keep the ones you serve. |
Builds live in ~/.legwork (override with LEGWORK_HOME). A failed build
saves its full error output to attempts.log in its build folder. Each
build keeps its own environment, which is several GB for repos that use
PyTorch: legwork clean removes old and failed builds and keeps the ones
you serve (--dry-run to preview, --all to remove everything).
Contributing a wrapper
legwork contribute owner/repo writes cache/<owner>__<repo>/ containing
wrapper.py and manifest.json, and prints a PR description. Before
writing anything, it:
blocks the entry if the wrapper copies 50+ words in a row from the source repo
re-runs the self-test in the sandbox
blocks the write if any output contains something shaped like an API key, including your own configured key
A GPL, AGPL, missing or unrecognized source license is flagged in the
manifest and the PR description, but not blocked. If the repo has commits
newer than the build or the existing entry, you get a warning. Open the PR
yourself; a bad entry is removed with a plain git revert.
Status
Early. What's next:
Better multi-tool verification than a single simplest-command self-test.
Development
See CONTRIBUTING.md. Security reports: SECURITY.md.
uv sync
uv run pytest # sandbox tests need macOS, or Linux with bubblewraptests/fixtures/ai_data_extractor_payload.py is a real malicious file with
its payload destroyed (every encoded byte randomized, code shape kept), so
the scanner is tested against a real technique. Tests only parse it with
ast; see tests/fixtures/README.md.
License
MIT
This server cannot be deployed
Maintenance
Related MCP Connectors
Create, deploy, and operate MCP servers directly from your GitHub repositories.
Find, vet, and run MCP tools through a secure audited gateway with prompt-injection risk scoring
An MCP server that gives your AI access to the source code and docs of all public github repos
AI-native git hosting — repos, PRs, issues, CI gates, and AI code review over MCP (60 tools).
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables analysis of any GitHub repository to get architecture, file roles, execution flows, system design Q\&A, and structured agent context. Works with MCP-compatible clients like Claude Desktop, Cursor, and Windsurf.628 npm1MIT
- AlicenseNot gradedqualityDmaintenanceProvides AI coding agents with structured intelligence about any GitHub repository including overview, PRs, contributors, hot files, CI status, and dependencies via a hosted MCP endpoint.36 npmMIT
- AlicenseNot gradedqualityBmaintenanceA production-grade MCP server that provides LLMs with safe, structured, tool-based access to GitHub repositories, including issue management, semantic search, and guarded write operations.MIT
- AlicenseBqualityBmaintenanceA security-first MCP gateway that enables AI assistants to safely inspect and interact with GitHub repositories through a controlled, auditable tool layer with policy enforcement and human approval for mutations.27MIT