Skip to main content
Glama

Fieldkit

Tested tools that let a small AI model on your own laptop do real work: find code, check and clean documents, run builds, and refuse to publish anything it cannot prove.

Built for the people every other tool prices out: kids with no credit card, 15-year-old laptops, data sold by the megabyte. Free forever, by design, not as a trial. Why, with the numbers: PHILOSOPHY.md

tests

The model is a small part of the job. The tools around it decide whether the job gets done. Fieldkit is those tools: deterministic Python that does the same thing every time, checks its own result, and answers in plain JSON. The model only has to choose the right tool.

Measured on one ordinary laptop (Intel i7-1255U, LM Studio): the same small model, Gemma, got one of five tasks right with basic tools and five of five with Fieldkit, in about half the time and with fewer tokens. Five tasks is a small test; run it on your own model with fieldkit exam.


What it does (Layman Track)

You run an AI helper on your own computer, and it is usually a small model. Left alone, a small model guesses: it opens file after file, loses track, and sometimes makes things up.

Fieldkit gives it tools that do the hard part and hand back a plain answer:

  • "Where is this defined?" It answers DEFINED at file:line or NOT DEFINED, so there is nothing to guess.

  • "Is this Word file safe to send?" It checks the file is not broken, removes people's names from its hidden properties, checks again and scans it for private data. It answers SAFE TO SEND or NOT SAFE, with the reason.

  • "Why did the build fail?" It names the known cause from the log and says the fix.

  • "Is this release true?" It refuses to publish until every claim in the release notes has a passing proof.

When the AI wants to change your files, Fieldkit shows you a preview first, keeps a backup, checks the result and can undo it. Anything that changes your system, or cannot be undone, waits for your approval. The AI cannot approve on your behalf.

It runs on Windows and Linux from the same code. It needs no account, sends nothing anywhere and costs nothing.

Related MCP server: mini-mcp-demo

Technical Definition (Developer Track)

Python 3.11+, one codebase for Windows 11 and Debian, AGPL-3.0-or-later. Every command takes --json; exit codes are fixed: 0 fine, 1 error, 2 bad usage, 3 findings.

  • Agent interface. discover, describe, run, undo, served on the command line and over MCP (fieldkit mcp, JSON-RPC 2.0 on stdio). run enforces one sequence: refuse draft cards, validate inputs, preview, require approval where the card says so, back up the scope, apply, verify, restore on failure. The MCP surface has no approval argument.

  • Tool cards. Each tool declares inputs (read from argparse by AST, never by running it), effects, safety (read-only, reversible, irreversible), modes (preview, apply, undo, verify) and tests. Trust ladder: gathered, carded, tested (a pass recorded on the file's current sha256), verified.

  • Pipelines. YAML stages with fingerprinted, resumable state and verify checks (files_exist, file_contains, output_contains, output_lacks, python). A stage is done only when its checks pass, never on exit code alone.

  • Release gate. release check exports the tag with git archive, runs its tests there, compares published files by sha256, privacy-scans, and demands a passing proof for every claim in the notes. release prove records evidence from another machine (platform, git tree id, result). No --force.

  • Tests. python -m pytest. A clean copy on the author's Windows 11 laptop: 187 passed, 41 skipped, 1 known fault kept as a strict xfail. GitHub Actions on every push: Windows 183 passed, 46 skipped; Ubuntu 181 passed, 48 skipped. A skipped test names the tool it needs, such as a repository fieldkit gather has not copied in yet.

Release notes: GitHub releases.


Table of contents

  1. Install

  2. Connect it to your AI helper

  3. What is in the box

  4. How it keeps changes safe

  5. Your own tools stay private

  6. Measure it on your own model

  7. The release gate

  8. Honest limits

  9. Layout


1. Install

Step-by-step guide, for people and for developers: INSTALL.md. It shows how to ask your AI to install Fieldkit or do it yourself, where to put it, how to connect it to Gorilla OpenCode and LM Studio, and how to make a small model aware of it.

The short version (Python 3.11 or newer and Git):

git clone https://github.com/gorillanobakaa-dot/Gorilla.Fieldkit
cd Gorilla.Fieldkit
python -m pip install -e ".[test]"
python -m pytest -q
fieldkit host
  • Pass: the tests end with zero failed. Some say skipped and name a tool: those need fieldkit gather first (section 3).

  • Pass: fieldkit host prints your system and Python version.

  • If fieldkit is not recognised: use python -m fieldkit instead. It does the same.

To remove it: python -m pip uninstall fieldkit, then delete the folder. Fieldkit changes nothing else on your computer.

2. Connect it to your AI helper

Full steps, backups and checks: INSTALL.md, steps 4 to 6.

Fieldkit speaks MCP, the standard way AI helpers use outside tools. Your helper then sees six tools: discover, describe, run, undo, next and readiness.

Gorilla OpenCode - add to your config.json (on Windows, %USERPROFILE%\.config\gorilla-opencode\config.json), then restart it:

"mcpServers": {
  "fieldkit": { "type": "stdio", "command": "fieldkit", "args": ["mcp"] }
}

Claude Code:

claude mcp add fieldkit -- fieldkit mcp

LM Studio (versions with MCP support) - add the same fieldkit entry to its mcp.json.

If your helper cannot find the fieldkit command, give the full path to it instead, for example the fieldkit.exe in your Python Scripts folder.

  • Pass: the first time the AI uses Fieldkit, your helper asks you to allow it.

3. What is in the box

Job

Command

Read a Word, Excel, PowerPoint or PDF file as text

fieldkit office read FILE

Create one from a JSON or YAML description

fieldkit office create SPEC OUT

Check it is not broken

fieldkit office check FILE

Remove people's names from its hidden properties

fieldkit office scrub FILE

All of the above, safe to send?

fieldkit office deliver FILE

Find secrets, home paths, emails, your private words

fieldkit privacy scan PATH [--git]

Name the cause of a failed build

fieldkit triage LOG

Run a staged, resumable build

fieldkit pipeline run NAME

The one next thing to do in a pipeline

fieldkit next NAME

Every named file, link or module must exist

fieldkit refcheck ...

Record state, then prove what changed

fieldkit snapshot take / diff

Install, check, uninstall; list leftovers

fieldkit lifecycle SPEC --approve

Prove a release before publishing it

fieldkit release check / prove

Find a tool by describing the job

fieldkit agent discover WORDS

Index what every script in a folder does

fieldkit harvest --find WORDS

Bring in tools from GitHub, with their tests

fieldkit gather [--test]

Measure a model with and without the kit

fieldkit exam run --model ID

Pipelines that ship: office-deliver, firefox-windows (drives the Gorilla Firefox build harness) and debian-kernel. Triage knows Firefox on Windows, the Debian kernel and Debian packaging failures. New failures are added as data, in fieldkit/build/signatures/.

fieldkit gather copies these repositories into toolbox/ and records the commit and sha256 of every file: pfind, searchfox-tools, model-eval, code-review-toolkit, dual-track-doc-generator, debian-kernel, sensors-gorilla, speaker-loudness-fix, black-gorilla-theme, apple-superdrive-enabler, respect-your-llms and the HDMI harness. fieldkit gather --test runs each one's own tests.

4. How it keeps changes safe

Every tool has a card that says what it changes. When the AI asks to run one:

  1. Read-only tools run straight away.

  2. Tools that change files show a preview first. When applied, Fieldkit backs up the files, runs the tool, checks the result, and puts the backup back if the check fails. undo restores them later.

  3. Tools that change the system (registry, services, power, hardware) or cannot be undone wait for your approval. Over MCP there is no way for the AI to give it.

The process runner records the programs it starts and stops only those, never a program of the same name that you are running yourself.

5. Your own tools stay private

Put your own tools in a local/ folder next to the code. Git ignores it, and Fieldkit merges it in when it loads:

File

What goes in it

local/tools.yaml

Cards for your own scripts (same format as fieldkit/desk/tools.yaml)

local/imports.yaml

Your own folders for fieldkit gather

local/tests/

Tests for your own scripts; pytest runs them with the rest

fieldkit.local.json

Machine paths and your private words for the privacy scan

Copy fieldkit.local.example.json to fieldkit.local.json to start. Your name and email go in its privacy.terms list, so the scan finds them without them ever being written in code.

6. Measure it on your own model

fieldkit exam builds the same small test project every time and gives your model five tasks, once with basic tools and once with Fieldkit. It talks to an OpenAI-compatible server on your own computer (LM Studio's address by default).

fieldkit exam run --model <model id as your server lists it>
fieldkit exam report

The result on the author's laptop:

Model

Basic tools

With Fieldkit

Gemma

1 of 5, 762 s, 8,349 tokens

5 of 5, 410 s, 6,630 tokens

Qwen3 30B-A3B

3 of 5, 1,415 s, 60,379 tokens

5 of 5, 1,030 s, 34,455 tokens

Five tasks on one test project, written by the same author as the tools. It shows the effect; it does not prove it for every job.

7. The release gate

A release spec in releases/ lists the tag, its tests, the files that must match what was tested, and the claims the release notes may make, each with a proof.

fieldkit release check releases/<name>.yaml    CLEAR, or DO NOT PUBLISH with every reason
fieldkit release prove releases/<name>.yaml    run the tag's tests on THIS machine; record evidence

A claim such as "runs on Windows and Linux" passes only with a passing prove from each platform for the same code. There is no --force: a failing gate means you fix the release or fix the check.

8. Honest limits

  • The Debian kernel pipeline has run as a dry run only; its full build has not run yet.

  • The firefox-windows pipeline needs the Gorilla Firefox build harness.

  • The privacy scan finds patterns and the words you list. It cannot know a secret it has no pattern for.

  • Fieldkit reads commands and files. It is not a sandbox.

  • Office text extraction handles ordinary documents; for complex layouts, install the optional [readers] extra (markitdown, docling).

  • Known faults are kept as tests marked "expected to fail", so a fix is noticed.

9. Layout

fieldkit/core     host, settings, process runner, privacy, pipeline engine, next, snapshot
fieldkit/office   read, create, check, scrub, deliver
fieldkit/build    triage + signatures, kernel, refcheck, lifecycle, pipelines
fieldkit/desk     registry (tools.yaml), cards, discover, readiness
fieldkit/exam     fixture, tasks and graders, toolsets, runner
fieldkit/agent.py, mcp.py, release.py, gather.py, harvest.py
skills/           short SKILL.md pointers; install_skills.py puts them where agents look
tests/            python -m pytest
AGENTS.md         instructions for AI agents working on this code

Licence: AGPL-3.0-or-later. If you run a changed version as a service, its source must stay open too.

Related MCP Connectors

Related MCP Servers