galton
Integrates with Ollama to discover and run local Ollama models, including using the Ollama API or the same GGUF file via llama-server, detecting model capabilities such as thinking/reasoning, sending think:false when needed, and inferring server activity from /api/ps for benchmarking and routing.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@galtonrun the reasoning and code suites on my local models and show the ranking"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Galton's Hoard
A local test bench for the language and vision models on your computer. It gives each model tasks whose answers can be checked automatically (reasoning, maths, Python code with hidden tests, JSON extraction, tool calling, instruction following, retrieval in long text, answers with citations, images, translation, summaries and Spanish writing, plus cases you add yourself), stores every answer with its timing and memory use, ranks and compares the models with confidence intervals and paired tests, notices when a model is new or has changed, and publishes a routing table that says which measured model is best for each kind of task.
Part of the Hoard family of local apps: it runs on your PC, keeps its data in data/, works on its own in the browser and can be driven by an assistant over MCP.
What it does
Finds your models. Models already served by Ollama, by llama-server (ports 8080 to 8090) or listed in the Faustus registry, the GGUF files in the folders you configure (default
D:\LocalAI\modelswhen it exists) and the Ollama models whose file is on disk. For an Ollama model both ways of running are offered: through the Ollama API, or on Galton's own llama-server using the same file (the default when the file is found). On Windows the Ollama store is found throughOLLAMA_MODELSfrom the environment or, when the app was not started from a shell that has it, from the user and machine variables in the registry; the folder whosemanifestsreally hold models wins. Aliases are only names a server or an app can report (Ollama tags, the llama-server alias, the id from/v1/models, a GGUF file stem); file paths andsha256-…blob names stay in their own fields and are removed from aliases stored by earlier versions on start. A model whose file changed is detected by its digest and its old results are marked stale, never mixed in. A llama-server that loads an Ollama blob runs the same weights as the Ollama tag: both entries answer to each other's names and the model page lists them as the same weights. A server that answers/v1/modelsbut cannot chat (no chat template, or an error to one test request) is marked as not a chat model once, switched off and never announced as a new model. The identity of a contestant on a shared server is its address plus what the server says it serves (ids of/v1/models, alias and model file of/props, or the Ollama digest). A run checks it before the contestant starts and again before every case (the look is cached about 20 s) and compares themodelfield of each reply; if somebody restarted the server with another model the contestant stops with an error and nothing measured after the change is stored.models_refreshmarks a contestant whose address now serves another model as not served ("no se sirve ahora"): it keeps its history, run forms exclude it with the reason, and it comes back when the server serves it again; the new model gets its own contestant.Measures with checkable tasks. 225 cases in 13 built-in suites, in Spanish of Spain unless the task is about another language. Sixteen kinds of checker: exact text, contains, regular expression, multiple choice, number (Spanish and English formats, tolerance; after an answer marker such as «Respuesta:» it reads the result of a calculation on that line, so
Respuesta: 2^10 = 1024is 1024; when that line has no digits it reads numbers written with words and ordinals in Spanish and English, soRespuesta: El octavo díais 8), mathematical equivalence (LaTeX such as\(3x^2\ln(x)+x^2\)is normalised before the symbolic comparison, and only the part after the last=of the answer line counts), JSON with a schema, tool call, Python with hidden tests, instruction constraints, needle in a long text, citations, a judge model with a rubric, a call to another Hoard app, instruction-following translation (ifmt, see below) and none (kept for the arena). A word or constraint check can be limited to the final answer (scope: "answer": only the text after the last «Respuesta:» or «Answer:» line), and a field of an extraction case can be compared ascontainsornorm(ignoring a leading generic word such as sala, calle or avda.), so a model that sayssala Magallanesis not failed forMagallanes. Every deterministic built-in case carries a reference answer and a test checks that it passes its own checker.Gives reasoning models room to think. Whether a model reasons is read from the chat template of llama-server (
enable_thinkingor<think>), from the capabilities Ollama reports (thinking), or learned when an answer arrives with reasoning content. For such a model, unless the run's effort isoff, the output limit is the answer budget of the case plusrunner.reasoning_tokens(8192 by default, 0 switches it off; a bigger effort asks for more), and the timeout grows with it. Effortoffsendsenable_thinking: falseto llama-server andthink: falseto Ollama. A result whose visible answer is empty because the limit was reached is stored astruncated; a model not known to reason that shows it is asked once more with the room and remembered. The run page and the leaderboard show how many results were truncated per model, and the leaderboard, the routing table andrecommendwarn when more than 5 % of a model's results in a category are (the budget was too small, so the score understates the model).Runs models on the GPUs you allow. A GGUF file runs on a llama-server started by Galton on a free port in 8091 to 8099, under a lease for the estimated memory (file size plus KV cache plus headroom), on one allowed GPU or split across several, and the process tree is killed afterwards. GPUs 0 and 1 are reserved by default and allowed GPUs are 2 and 3; reserved GPUs cannot be allowed without an explicit confirmation. Running servers are called as they are, but never while somebody else is using them: a busy server puts that model in the state
waiting_server(the run page andrun_statusshow how long it has waited), and the run starts once the server has been quiet forrunner.idle_grace_s(20 s by default); it gives up only afterrunner.wait_idle_max_s(an hour by default, 0 waits for ever; a run'swait_soverrides it). Between two questions it looks again and pauses the same way if someone else started a chat (llama-server shows its slots; for Ollama the activity is inferred from theexpires_atof/api/ps, so a request still in flight cannot be seen); an Ollama model that is not loaded is only loaded if you enable it.Falls back to the CPU for small files. When a GGUF fits no allowed GPU right now (every GPU taken by another workload, or none of the allowed ones present) and its file is at most
runner.cpu_max_gb(4 GB by default), Galton starts its own llama-server with-ngl 0,CUDA_VISIBLE_DEVICES=""(it never touches a GPU),-t= physical cores minus 2 (at least 2) and no lease, instead of waiting for a GPU: a model that could take a GPU is still given one when it is free. The run settingdeviceisauto(this behaviour),gpu(never the CPU: exactly the old behaviour, including the wait and the error) orcpu(always; the size limit does not apply).runner.cpu_fallback(on by default) switches the automatic choice off. The CPU needs the memory too: with too little free RAM the old GPU error stays. The plan says where each model would run ("CPU (no hay GPU libre)"), the run card and the results are flagged as CPU, and a CPU speed is never mixed with GPU speeds: for one model the leaderboard,compareand the routes use its GPU measurements when it has any (the CPU figure is shown apart), otherwise its CPU ones with a "CPU" badge; a CPU-only speed does not count in the speed weight of the ranking. The watch never puts a background run on the CPU.Says how sure it is. Pass rate with a Wilson 95 % interval; mean score with a seeded bootstrap 95 % interval, weighted by case weight; two models are compared on the cases they share with an exact McNemar test and a paired bootstrap, and the verdict is one of better, worse, no clear difference or no data. Speed (tokens per second, time to first token, load time) and memory come from the runs, not from guesses.
Publishes the routing table. For each task category the candidates are the enabled models with recent results for their current file; remote endpoints are left out unless allowed; too slow or too large models are excluded; the rest are ranked by the lower bound of the 95 % interval, with decode speed as tie-break. Each choice has a sentence that explains it. The table is written atomically to
routes.json(HOARD_ROUTES_FILEor~/.hoard/routes.json) and Hoard Link reads it, so the other apps of the family use the measured best model. A category with fewer checked cases than the policy asks for is never published. See what would change before publishing.Uses the models your other machines serve. The language models that another machine of your own network serves (a GPU cluster serving vLLM, for one) become contestants of kind
server, one per model. The sources are the Faustus registry (items filed as local language models on a private address or a.localname; Hoard Link'sis_lan_item) and, when Faustus lists none or is down, Prometheus's Hoard (prometheus.url,/api/endpoints); the model must also be listed by the server's own/v1/models(a bounded 3 s per server, at most 16 servers per refresh, so a refresh never hangs on them). They are the person's own, so they are notremote(noremote_okneeded) and carrymeta.network = "lan"and the badge "Red local" / "Local network"; the key isserver:lan:<host>:<port>:<model>, so a server that stops and comes back is the same contestant: while it does not answer it isup: false(never deleted, its measurements stay), and a server restarted with another model marks the old entry as not served as for any shared server. A server without/slots(vLLM) is judged busy by the request gauges of its/metrics. Their results flow into the leaderboard and, through the normal publication, into the routing table.Watches. New or changed models are measured with the quick suite when that disturbs nobody: a model that a server is already serving (a llama-server or Ollama on this PC, or a server of your network) once that server has been idle for 10 minutes; never in quiet hours, 01:00 to 08:00 by default. By default the watch never loads anything on this PC: with
watch.load_localoff (the default) it does not start a llama-server or load a model into Ollama, and it only reports the GGUF files it left alone (held_backin the watch status); switchwatch.load_localon (Settings > Regression watch, orsettings_set) to let it also measure a GGUF that fits on a free allowed GPU. Manual runs (run_start,measure_new, the buttons) can still load models; a run's own settingload_local: falsemakes any run refuse to load. A model the watch (or a run) could not measure is not tried again by the watch until its file (digest, path) or its server changes, a server that was down answers again, or somebody runs it by hand: the reason stays on the model (watch_gave_upinmodels_list, a chip on the model page) and the failure is announced once per model, file and reason, not on every cycle. After each run it compares a model with the results it had before its file changed and raises a notice and the family eventgalton.regressionwhen it got significantly worse.Reports its work to the family. Every run emits the canonical job events (queued, started, progress with the GPU and an estimate, done, failed, cancelled), so the hub shows it next to the jobs of the other apps.
galton_run {models}queues the quick suite on the named models (a name it has not seen yet triggers one rediscovery first; names it still cannot find are listed, never ignored); it is what the hub calls after a model is published.Arena. Two anonymous answers to the same prompt, taken from stored runs (no model time spent), you vote, and Bradley-Terry ratings are computed per category.
Your own suites. Create a suite, add cases by hand, or import JSONL or CSV (columns may be in Spanish or English; a one-line checker such as
contains:uno|dosornumber:42is enough). Built-in suites are read-only and can be duplicated.case_tryruns one case on one model without saving anything.External benchmark: IFMTBench.
benchmark_import(or Pruebas > Importar benchmark) downloads the translation instruction-following benchmark from a pinned commit into the data folder, checks its SHA-256 and builds a suite of your own from a seeded stratified sample (default 210 cases: 30 for each of six constraints, glossary, style, background context, layout, structured data and code or tags, plus 30 that combine several). Rule checks are exact; style, context and a glossary the rule rejects go to the judge and wait, uncounted, while there is none. An item with several constraints scores the product of the all-or-nothing ones times the mean of the graded ones, and the leaderboard shows the pass rate per constraint so you can see where a model fails. Data CC BY 4.0, scoring ported from the benchmark's Apache-2.0 code (NOTICE); details, limits and differences from the original in docs/IFMTBENCH.md.Judge. A model you choose grades open answers against a rubric on a 0 to 10 scale. It never silently grades itself (such results are marked
self_judged), grades are cached by what the judge saw, a judge that is the model being measured grades only after that model's last case (so grading never delays a timed question), and answers wait as pending, not failed, while the judge is not available.
Related MCP server: Coval MCP Server
Screens
Spanish by default, English with one click, dark.
Panel: routing table with score, interval, speed and memory; the run in progress; notices; the GPUs with their leases; buttons to measure what is new and to publish the routes.
Modelos: every model with kind, family, size, quantisation, context, whether it fits in 16 GB, last measurement and stale marks; enable, rename, add aliases, add a file or a server.
Pruebas: suites and cases with their checkers; create, duplicate, import and try a case; import the IFMTBench benchmark.
Ejecutar: choose suites and models, see the plan (cases, memory, where each model would run, problems), start, follow the progress case by case, cancel (the message says what is stopped: only a llama-server that Galton started itself; a shared server is never touched), continue a run that failed (Galton was restarted) or was cancelled, asking only the cases it did not measure, discard a finished run with a reason (it stays in the history with a badge but no statistic, ranking, route, comparison or regression check uses it; restore undoes it), read every answer with the checker detail.
Clasificación: ranking by category or suite with interval bars, speed and memory columns, a pass rate per constraint for models measured on IFMTBench, and the paired comparison of two models with the list of cases where they differ.
Arena: blind votes and ratings.
Rutas: the routing table, the difference with what is published, the history of publications.
Ajustes: allowed GPUs, llama-server path, judge, runner limits, regression watch, routing policy, language.
What it cannot do
It measures what its cases check. A high score on a suite says little about tasks that are not in the suite; add cases of your own for the work you care about.
Galton's own llama-server runs on the CPU only for files up to
runner.cpu_max_gb(or when a run saysdevice: cpu), and its speed is not comparable with GPU speeds. A GGUF over that size that does not fit in the allowed GPUs together is reported as impossible with the numbers, not attempted.It needs
llama-server(llama.cpp) installed to run GGUF files itself; without it only servers that are already running can be measured.Judge grades are as good as the judge model you choose, and an answer that is still waiting for the judge does not count until it is graded. A request that fails because the server broke is stored with its error and left out of the score; it is never counted as a wrong answer.
Python cases run model-written code in a separate interpreter with a timeout and a cleaned environment. That is not a security sandbox; switch it off with the setting
checks.allow_code_executionif you do not want it.Vision cases need a model with its projector file; models without one are skipped with the reason stored.
It does not train, quantise or download models, and it does not call remote endpoints unless you enabled that model one by one.
The room given to thinking is a limit, not a measurement of how much a model thinks; a model that needs more than
runner.reasoning_tokensis still cut and shown as truncated.IFMTBench is a sample of a larger benchmark, scored with the benchmark's own rules, which have known limits (docs/IFMTBENCH.md): Markdown that is not a table is refused even for the reference answer, a glossary term must appear as a substring, and a sample of 30 items per type gives wide intervals. Its score is not the published score of the whole benchmark.
Small samples give wide intervals. The app shows them instead of hiding them and never publishes a category with too few cases.
Install
Requirements: Python 3.11 or newer (tested on 3.11 and 3.13), Node 22 only to rebuild the UI, a llama-server binary for GGUF files and Ollama if you use it.
python -m venv venv
venv/Scripts/python -m pip install -r requirements.txt # Windows; venv/bin/python elsewhereThe UI is prebuilt in galton_hoard/static. To rebuild it: npm install && npx vite build.
Run
venv/Scripts/python -m galton_hoard # http://127.0.0.1:5201
python scripts/launch.py # free port, opens the browserDemo mode shows the whole app without touching hardware: GALTON_FAKE=1 python -m galton_hoard uses invented GPUs and three models that answer from a table, and publishes to data/routes-demo.json.
The UI is Spanish or English; the messages of the backend (errors with their hints, plan lines, run warnings, notices) reach it as a stable key plus parameters and are formatted in the chosen language, while assistants always get the English sentence. Without a console (pythonw) the start-up messages go to the log only. Starting is the shared launcher of Hoard Link: a second copy exits at once without touching the data folder of the first, and PORT_STRICT=1 refuses a taken port.
Environment: GALTON_PORT (5201), GALTON_DATA_DIR, PORT_STRICT=1, GALTON_ALLOWED_HOSTS, GALTON_HTTP_TIMEOUT_S, GALTON_SCHEDULER=0 (no background jobs), GALTON_OFFLINE=1 (no discovery), GALTON_FAKE=1 (demo), HOARD_ROUTES_FILE, and GALTON_FAUSTUS_TOKEN in .env (see .env.example). Everything else is a setting stored in data/galton.db and changed in the app.
Assistants (MCP)
mcp_server.py is a stdio MCP bridge named galton-hoard (the shared bridge of Hoard Link). It never opens the database: it proxies every call to the running app with the token in data/mcp-token, refreshes the tool list while it runs, and starts the app when it is not answering. Calls may take up to 660 seconds; the tools that wait for a run (wait_s, up to 600 seconds) or for the judge say so in the catalogue, and a run that outlasts wait_s answers still_running: true and goes on. faustus-plugin.json describes the app, its health check and the bridge for Faustus and the Hoard Hub.
Tools (46): galton_overview, galton_status, models_list, models_refresh, model_get, model_add, model_update, model_remove, suites_list, suite_get, case_get, suite_create, suite_update, suite_duplicate, suite_remove, case_add, case_update, case_remove, cases_import, benchmark_import, case_try, run_plan, run_start, run_status, run_cancel, run_discard, run_restore, run_resume, runs_list, run_results, measure_new, galton_run, judge_run, leaderboard, compare, recommend, routes_get, routes_publish, arena_next, arena_vote, arena_ratings, settings_get, settings_set, gpu_status, notices_list, housekeeping_run. Arguments in docs/API.md.
Events on the family bus: galton.routes.updated (the published table changed), galton.regression (a model got worse), galton.models.updated (new, changed or missing models) and the canonical job events of every run, galton.job.queued|started|progress|done|failed|cancelled with {job_id, title, kind: "run", progress, gpu, eta_s, url, error} (progress events at most every 5 seconds). The hub's Work tab shows them, and a rule of the hub calls galton_run {models} after a model is published, so a freshly published model is measured without asking.
Data and privacy
Everything lives in data/ (or GALTON_DATA_DIR): galton.db (models, suites, runs, results, settings), images/ (images attached to your cases), logs/ (one llama-server log per model, and galton-hoard.log, a rotating log of the app itself, which is what remains when it is started without a console), cache/, external/ (benchmark data, downloaded from raw.githubusercontent.com only when you ask for it and verified against a pinned SHA-256), servers.json (the llama-server processes Galton started, so a crash cannot leave a model loaded: they are stopped at the next start), mcp-token and url. Nothing leaves the computer: the app listens on 127.0.0.1 only, rejects cross-site requests, and calls only local servers unless you enable a remote endpoint for a model. The prompts and answers of your own cases are stored as they are; do not put secrets in them. Back up the folder to back up the app.
Development
python -m pytest # no network, no GPU, no llama-server needed
python scripts/gen_api_doc.py # regenerate docs/API.md after changing a tool (a test fails when it is stale)
npx vite build # rebuild the UI into galton_hoard/staticSee AGENTS.md and docs/ARCHITECTURE.md. The tests use a fake world (fake GPUs and leases, models that answer from the reference answers, a fake launcher) and mocked HTTP; nothing in them touches real hardware.
License
MIT
This server cannot be deployed
Maintenance
Related MCP Connectors
Run data annotation and evaluation tasks with vetted domain experts: create, publish, get results.
Human judgment for AI agents: discover capabilities, get quotes, and track paid human tasks.
Pay-per-call AI evaluation MCP server. Score LLM outputs against benchmark rubrics via Workers AI.
Test the voice agents you run: scored transcripts, pass/fail verdicts, latency and WER metrics.
Related MCP Servers
AlicenseNot gradedqualityCmaintenanceProvides advanced evaluation tools for assessing AI safety, alignment, and performance of LLM outputs. Enables programmatic evaluation of quality, safety metrics like toxicity and PII detection, and operational metrics including carbon footprint and cost estimation.4Apache 2.0
Coval MCP Serverofficial
AlicenseAqualityBmaintenanceEnables AI assistants to interact with Coval's evaluation platform for launching and monitoring evaluation runs, managing agents and test sets, and retrieving evaluation metrics.1814 npm3MIT- AlicenseNot gradedqualityBmaintenanceEnables LLM evaluation and observability by uploading documents, building test sets, running RAG pipelines, and automatically scoring answers for groundedness, hallucination risk, retrieval quality, latency, and cost, with tools exposed to MCP-compatible clients.1MIT
- AlicenseNot gradedqualityBmaintenanceEnables agents to run evals mid-task, providing tools to validate suites, execute model comparisons, and inspect calibration and win-rate reports.MIT