Fieldkit
Provides pipelines and triage for Debian kernel and packaging failures, enabling staged, resumable builds and diagnosis of build issues.
Provides a pipeline for the Firefox Windows build harness and triage for Firefox on Windows build failures.
Allows gathering tools from GitHub repositories into a local toolbox, recording commit and file hashes, and optionally running their tests.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Fieldkitis this Word file safe to send?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Fieldkit
Tested tools that let a small AI model on your own laptop do real work: find code, check and clean documents, run builds, and refuse to publish anything it cannot prove.
Built for the people every other tool prices out: kids with no credit card, 15-year-old laptops, data sold by the megabyte. Free forever, by design, not as a trial. Why, with the numbers: PHILOSOPHY.md
The model is a small part of the job. The tools around it decide whether the job gets done. Fieldkit is those tools: deterministic Python that does the same thing every time, checks its own result, and answers in plain JSON. The model only has to choose the right tool.
Measured on one ordinary laptop (Intel i7-1255U, LM Studio): the same small model, Gemma, got one of five tasks right with basic tools and five of five with Fieldkit, in about half the time and with fewer tokens. Five tasks is a small test; run it on your own model with
fieldkit exam.
What it does (Layman Track)
You run an AI helper on your own computer, and it is usually a small model. Left alone, a small model guesses: it opens file after file, loses track, and sometimes makes things up.
Fieldkit gives it tools that do the hard part and hand back a plain answer:
"Where is this defined?" It answers
DEFINED at file:lineorNOT DEFINED, so there is nothing to guess."Is this Word file safe to send?" It checks the file is not broken, removes people's names from its hidden properties, checks again and scans it for private data. It answers
SAFE TO SENDorNOT SAFE, with the reason."Why did the build fail?" It names the known cause from the log and says the fix.
"Is this release true?" It refuses to publish until every claim in the release notes has a passing proof.
When the AI wants to change your files, Fieldkit shows you a preview first, keeps a backup, checks the result and can undo it. Anything that changes your system, or cannot be undone, waits for your approval. The AI cannot approve on your behalf.
It runs on Windows and Linux from the same code. It needs no account, sends nothing anywhere and costs nothing.
Related MCP server: mini-mcp-demo
Technical Definition (Developer Track)
Python 3.11+, one codebase for Windows 11 and Debian, AGPL-3.0-or-later. Every command
takes --json; exit codes are fixed: 0 fine, 1 error, 2 bad usage, 3 findings.
Agent interface.
discover,describe,run,undo, served on the command line and over MCP (fieldkit mcp, JSON-RPC 2.0 on stdio).runenforces one sequence: refuse draft cards, validate inputs, preview, require approval where the card says so, back up the scope, apply, verify, restore on failure. The MCP surface has no approval argument.Tool cards. Each tool declares inputs (read from argparse by AST, never by running it), effects, safety (
read-only,reversible,irreversible), modes (preview, apply, undo, verify) and tests. Trust ladder: gathered, carded, tested (a pass recorded on the file's current sha256), verified.Pipelines. YAML stages with fingerprinted, resumable state and verify checks (
files_exist,file_contains,output_contains,output_lacks,python). A stage is done only when its checks pass, never on exit code alone.Release gate.
release checkexports the tag withgit archive, runs its tests there, compares published files by sha256, privacy-scans, and demands a passing proof for every claim in the notes.release proverecords evidence from another machine (platform, git tree id, result). No--force.Tests.
python -m pytest. A clean copy on the author's Windows 11 laptop: 187 passed, 41 skipped, 1 known fault kept as a strict xfail. GitHub Actions on every push: Windows 183 passed, 46 skipped; Ubuntu 181 passed, 48 skipped. A skipped test names the tool it needs, such as a repositoryfieldkit gatherhas not copied in yet.
Release notes: GitHub releases.
Table of contents
1. Install
Step-by-step guide, for people and for developers: INSTALL.md. It shows how to ask your AI to install Fieldkit or do it yourself, where to put it, how to connect it to Gorilla OpenCode and LM Studio, and how to make a small model aware of it.
The short version (Python 3.11 or newer and Git):
git clone https://github.com/gorillanobakaa-dot/Gorilla.Fieldkit
cd Gorilla.Fieldkit
python -m pip install -e ".[test]"
python -m pytest -q
fieldkit hostPass: the tests end with zero failed. Some say
skippedand name a tool: those needfieldkit gatherfirst (section 3).Pass:
fieldkit hostprints your system and Python version.If
fieldkitis not recognised: usepython -m fieldkitinstead. It does the same.
To remove it: python -m pip uninstall fieldkit, then delete the folder. Fieldkit changes
nothing else on your computer.
2. Connect it to your AI helper
Full steps, backups and checks: INSTALL.md, steps 4 to 6.
Fieldkit speaks MCP, the standard way AI helpers use outside tools. Your helper then sees
six tools: discover, describe, run, undo, next and readiness.
Gorilla OpenCode - add to your config.json (on Windows,
%USERPROFILE%\.config\gorilla-opencode\config.json), then restart it:
"mcpServers": {
"fieldkit": { "type": "stdio", "command": "fieldkit", "args": ["mcp"] }
}Claude Code:
claude mcp add fieldkit -- fieldkit mcpLM Studio (versions with MCP support) - add the same fieldkit entry to its
mcp.json.
If your helper cannot find the fieldkit command, give the full path to it instead, for
example the fieldkit.exe in your Python Scripts folder.
Pass: the first time the AI uses Fieldkit, your helper asks you to allow it.
3. What is in the box
Job | Command |
Read a Word, Excel, PowerPoint or PDF file as text |
|
Create one from a JSON or YAML description |
|
Check it is not broken |
|
Remove people's names from its hidden properties |
|
All of the above, safe to send? |
|
Find secrets, home paths, emails, your private words |
|
Name the cause of a failed build |
|
Run a staged, resumable build |
|
The one next thing to do in a pipeline |
|
Every named file, link or module must exist |
|
Record state, then prove what changed |
|
Install, check, uninstall; list leftovers |
|
Prove a release before publishing it |
|
Find a tool by describing the job |
|
Index what every script in a folder does |
|
Bring in tools from GitHub, with their tests |
|
Measure a model with and without the kit |
|
Pipelines that ship: office-deliver, firefox-windows (drives the Gorilla Firefox build
harness) and debian-kernel. Triage knows Firefox on Windows, the Debian kernel and Debian
packaging failures. New failures are added as data, in fieldkit/build/signatures/.
fieldkit gather copies these repositories into toolbox/ and records the commit and
sha256 of every file: pfind, searchfox-tools, model-eval, code-review-toolkit,
dual-track-doc-generator, debian-kernel, sensors-gorilla, speaker-loudness-fix,
black-gorilla-theme, apple-superdrive-enabler, respect-your-llms and the HDMI harness.
fieldkit gather --test runs each one's own tests.
4. How it keeps changes safe
Every tool has a card that says what it changes. When the AI asks to run one:
Read-only tools run straight away.
Tools that change files show a preview first. When applied, Fieldkit backs up the files, runs the tool, checks the result, and puts the backup back if the check fails.
undorestores them later.Tools that change the system (registry, services, power, hardware) or cannot be undone wait for your approval. Over MCP there is no way for the AI to give it.
The process runner records the programs it starts and stops only those, never a program of the same name that you are running yourself.
5. Your own tools stay private
Put your own tools in a local/ folder next to the code. Git ignores it, and Fieldkit
merges it in when it loads:
File | What goes in it |
| Cards for your own scripts (same format as |
| Your own folders for |
| Tests for your own scripts; |
| Machine paths and your private words for the privacy scan |
Copy fieldkit.local.example.json to fieldkit.local.json to start. Your name and email go
in its privacy.terms list, so the scan finds them without them ever being written in code.
6. Measure it on your own model
fieldkit exam builds the same small test project every time and gives your model five
tasks, once with basic tools and once with Fieldkit. It talks to an OpenAI-compatible
server on your own computer (LM Studio's address by default).
fieldkit exam run --model <model id as your server lists it>
fieldkit exam reportThe result on the author's laptop:
Model | Basic tools | With Fieldkit |
Gemma | 1 of 5, 762 s, 8,349 tokens | 5 of 5, 410 s, 6,630 tokens |
Qwen3 30B-A3B | 3 of 5, 1,415 s, 60,379 tokens | 5 of 5, 1,030 s, 34,455 tokens |
Five tasks on one test project, written by the same author as the tools. It shows the effect; it does not prove it for every job.
7. The release gate
A release spec in releases/ lists the tag, its tests, the files that must match what was
tested, and the claims the release notes may make, each with a proof.
fieldkit release check releases/<name>.yaml CLEAR, or DO NOT PUBLISH with every reason
fieldkit release prove releases/<name>.yaml run the tag's tests on THIS machine; record evidenceA claim such as "runs on Windows and Linux" passes only with a passing prove from each
platform for the same code. There is no --force: a failing gate means you fix the
release or fix the check.
8. Honest limits
The Debian kernel pipeline has run as a dry run only; its full build has not run yet.
The
firefox-windowspipeline needs the Gorilla Firefox build harness.The privacy scan finds patterns and the words you list. It cannot know a secret it has no pattern for.
Fieldkit reads commands and files. It is not a sandbox.
Office text extraction handles ordinary documents; for complex layouts, install the optional
[readers]extra (markitdown, docling).Known faults are kept as tests marked "expected to fail", so a fix is noticed.
9. Layout
fieldkit/core host, settings, process runner, privacy, pipeline engine, next, snapshot
fieldkit/office read, create, check, scrub, deliver
fieldkit/build triage + signatures, kernel, refcheck, lifecycle, pipelines
fieldkit/desk registry (tools.yaml), cards, discover, readiness
fieldkit/exam fixture, tasks and graders, toolsets, runner
fieldkit/agent.py, mcp.py, release.py, gather.py, harvest.py
skills/ short SKILL.md pointers; install_skills.py puts them where agents look
tests/ python -m pytest
AGENTS.md instructions for AI agents working on this codeLicence: AGPL-3.0-or-later. If you run a changed version as a service, its source must stay open too.
This server cannot be deployed
Maintenance
Related MCP Connectors
Persistent memory, hybrid search and a goal graph for AI agents, over stdio or remote HTTP.
Runtime permission, approval, and audit layer for AI agent tool execution.
Deterministic reasoning stack for AI agents: simulate, decide & compute, plus cross-domain tools.
Shared control plane for AI coding agents — tasks, memory, decisions, file locks. 12 tools.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceLocal-first AI agent for approval-gated automation and verifiable LLM workflows.1MIT
- AlicenseNot gradedqualityCmaintenanceEnables local tool calling over Model Context Protocol via stdio, providing deterministic tools such as calc.add, text.word_count, and text.summarize_naive after JSON-RPC handshake and discovery.MIT
- AlicenseNot gradedqualityCmaintenanceEnables registering custom agent tools via a minimal JSON-RPC over stdio implementation, with zero external dependencies, to expose them to AI agents like Claude, Cursor, and Gemini.MIT
- AlicenseAqualityBmaintenanceEnables an AI agent to run local tools on the user's own machine via stdio, including command execution, workspace file read/write, and system status checks.52MIT