Verinoda
Exports the analyzed code graph and named maps as Mermaid diagrams, alongside other graph export formats.
Imports OpenTelemetry traces and links their spans to the corresponding code, letting runtime observations be analyzed against the repository's code graph.
Imports Sentry error traces into the code graph so a stack trace or error report can be mapped onto the code that produced it.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Verinodaindex this repo and show what depends on UserService"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Verinoda Symbiosis
Verinoda + Claude Code, together. This repository is Verinoda with verinoda-live, a Claude Code mod that keeps the index fresh while Claude edits, can nudge the agent to use it (measured: the same facts; fewer tokens on small corpora, 14 % more input tokens on real repositories), reviews commits, and shows it all in a pane with an animated mascot. Start with SYMBIOSIS.md; the mod is claude-mods/verinoda-live. Everything below is Verinoda's own README.
Verinoda
Durum / Status: BETA. Verinoda beta aşamasındadır: çekirdek komutlar (
scan,update,query,analyze,trace,check,review) ve MCP çekirdek araçları kullanıma hazırdır;docs/DESIGN.md'de "partial" olarak işaretli özellikler deneyseldir ve değişebilir. Cevaplar kanıt satırlarıyla verilir, ama yanlış olabilir: kritik kararlarda kanıtı kendiniz okuyun. Hiçbir doğruluk, güvenlik ya da uygunluk garantisi verilmez. Verinoda is in beta: the core commands (scan,update,query,analyze,trace,check,review) and the core MCP tools are ready for use; features marked "partial" indocs/DESIGN.mdare experimental and may change. Answers come with their evidence lines but can be wrong: read the evidence yourself before a critical decision. No warranty or guarantee of correctness, security or fitness for any purpose is given (Apache-2.0 "AS IS" terms inLICENSE). Performance claims are published only with their measurement (model, version, date, raw logs). Published: 0.4.0 (beta) on PyPI and npm (2026-10-01; 0.3.0 and 0.3.2 on 2026-09-27, 0.1.0 and 0.2.0 on 2026-09-26). Formerly developed under the working name "RepoAtlas".
What is in this snapshot (2026-10-01)
Part | State |
Release | 0.4.0 (beta), 2026-10-01: |
Documents and images ( | PDF, Word, Excel and PowerPoint files in the repository are read as text (a heading per page, sheet, slide or Word heading), become graph nodes and search passages, and are quoted as evidence that can be checked again; on Windows the text in screenshots and diagrams is read with the OCR engine built into Windows (no model, nothing downloaded; textures, icons and small images skipped; |
Java name check ( | Java files are checked against the project's sources, its classpath and the JDK: imports, types, methods with their number of arguments, fields, constructors and Fabric Mixin targets ( |
Kotlin name check ( | Kotlin files are checked in the same world as Java (the project's Java and Kotlin sources, the classpath, the JDK; Java code now sees the project's Kotlin classes, objects and companions): imports, types, and members and properties on receivers of known type (a Java getter counts for a property). Kotlin keeps more names open, and they stay |
TypeScript/JavaScript imports ( | The invented name an agent writes most often on the web is an import. Checked: a relative path that names no file, a package not in |
Minecraft mods and the JVM (D47-D54, 2026-09-26) |
|
English questions over Turkish-named code (D55, 2026-09-26) | a symbol's doc comment is its own search text; the Turkish-English seed dictionary is read backwards for code tokens (on a 30-question mixed set of a Turkish-named mod: top-3 8/30 -> 18/30) |
Answers read as a person would (D56-D59, 2026-09-26/27) | an analysis says the changed files once, without repeated uncertainties (-1.6 % characters, same facts); Java overloads are separate symbols and a call binds to the overload its argument count fits; "how does X work" with one subject is answered by what X calls (an inference), not by unrelated entry-to-storage paths; a storage question gets the paths through its own code; "what breaks if I change X" names X's callers at their call sites; the JSON answer keeps room for the next candidates (JSON facts 253 -> 262 over the nine benchmark sets); context the critique refuted is counted, not printed; a why-answer quotes the section that gives the reason; a config question also finds settings read by a string key and the file line that sets them |
Faster update (D42, 2026-09-26) | Same graph, faster: Leiden communities in native code by default, graph.json written through json's C encoder, memoised path work; about 12% on Verinoda's own repository with Leiden on both sides (31-38 s -> 27-33 s), more without the old Leiden extra. An update proportional to the change is not done yet |
| The changed files are taken in at once (search index, lexicon, syntax facts, stale claims) and the graph is rebuilt by a background |
Graphify port ( | done. Upstream suite at port time: 5436 passed / 50 failed, and every failure also fails on unmodified upstream on the same Windows machine; not re-run since ( |
Core: claims, evidence, critique, experiments, research/compare, feedback, memory, installers, MCP server (44 tools; several projects from one process, stdio or HTTP) | implemented |
Repository groups ( | Calls from one repository into another as links with both locators, computed from the members' own indexes. On |
Language servers ( | Opt-in, CLI only, trusted projects only (a server runs third-party code over the project). A small stdio LSP client asks the installed server of the file's language ( |
Pinned regression tests ( | Python, opt-in, trusted projects only. Runs the tests that reach one function in a throw-away copy, records each call's inputs and its result or exception when every value rebuilds from Python source, writes |
Round 3: search engine, question plans with Turkish support, reference resolver, trust engine (anchors, entailment, facet-level staleness), runtime observation, precise call resolution | implemented and wired into the CLI, MCP and |
Product test suite | 3,881 passed, 27 failed, 6 skipped, 2 deselected at 8868179 (before D134-D136), Windows 11 / Python 3.13, 2026-10-01: 26 failures are installer and setup tests that depend on this machine's environment (they fail the same way on every recent run), one was a test of the MCP review's key order, fixed since; D134-D136 ran their own tests |
Agent integration | Claude Code ( |
Benchmarks (measured, | Graphify's own code (226 files, 37 facts, in-sample): Verinoda text retrieval 36/37 at 1,424 tokens/question, ~0.14 s; Graphify 7/37. Set on Verinoda's own earlier code (33 facts): 25/33 vs Graphify 8/33 and raw reading 9/33; it was held out until the 2026-09-23 ranking change, which was chosen with it in view (22/33 before). Turkish paraphrases of the example app: 32/32. Regressions: |
Game mods and data files (2026-09-24) | data packs, JSON/yml configs and other data files indexed; resource-id links between code and data; reference trees; Java calls the extractor drops; translation pairs from locale files. New example set |
Notes and graph view ( | a note per symbol, file, section and data file with its code and links, the line each link is written on and editor links; a 3D graph with a panel that says what is in view, follow, a walk through a file's links, regions, a tour and a watch list; a command bar ( |
Answer quality (2026-09-25) |
|
Truth rules ( | Four ways a false sentence reached a verified or "likely" status are closed. Word overlap with the cited lines never verifies (only a verbatim |
Name check ( | Python here; Java since 2026-09-26 (see Java name check). Since 2026-09-26 a file in another language (named, under a named directory, changed in the diff, or a snippet's |
Decisions stay human ( | A "should we / which one / how will this scale" question is |
Debug ledger ( | Every attempt at fixing one symptom is recorded against one repro command: the tree it ran on, the patch against the base, the failure as exception at |
Change review ( | For a diff, a staged change or a planned one ( |
Behaviour probe ( | Python only. For a changed function: inputs from its signature, annotations, call sites and boundary constants, run on the old and the new version in throw-away copies, compared; differences are reported as behaviour changes (not bugs), with the example and an optional pinning test. A side-effect gate refuses functions that write files, open sockets, start processes or keep global state it cannot isolate, and says why. Measured on hand-written fixtures and mutants (in-sample): 22/22 planted regressions and 25/25 mutants found; 5 of 60 behaviour-preserving edits reported, all float rounding marked as such; the gate refused 9 of 9 unsafe functions and none of the safe ones. It says "no difference found in N inputs", never "verified". Since 2026-09-26 an environment variable set while a library module loads (numpy sets |
Exact names and a fresh index ( |
|
Token cost ( | What a model reads got shorter with no gold fact lost, found or shown, on the eight public sets (86 questions, 319 facts) or on 10 out-of-sample questions (chars/4 tokens). Per question: |
Honest verdicts (D39, 2026-09-26) |
|
Cross-platform | tested on Windows 11; CI (Linux, Windows, macOS x Python 3.10/3.12/3.13) runs on every push since 2026-09-25 - its first run found and fixed macOS isolation, macOS path forms in the export and the Python 3.10 call tracer |
Known limitations (see also docs/DESIGN.md for per-decision gaps):
A Turkish question about a repository whose code is English but whose UI strings and docs are Turkish (Verinoda itself) finds the Turkish text first: the
verinoda_user_trset scores 13/30 (query text and analyze; 11/30 before an English code word in a Turkish question stopped taking the names it merely appears next to, 2026-09-25), against 36/37 on the Englishgraphify_core. Each dictionary gloss of a Turkish word still counts as its own search term, which is why adding correct dictionary entries has not helped yet. Two fixes were measured and not kept (docs/BENCHMARKS.md, Update 2026-09-25): weighing such words below their translation cost a mod set, where those words are how the data files are found.On Windows with Python 3.10 or 3.11, a source file nested thousands of levels deep can crash the extractor: the upstream pipeline raises Python's recursion limit to 10,000, and before Python 3.12 that can exhaust the C stack before a
RecursionError(found by CI, 2026-09-25; Python 3.12+ and other systems are not affected).
Open findings of the second acceptance audit (2026-09-23 05:00), not yet fixed:
A project's tests are its own code: they run under process isolation only in a project the user trusts (
verinoda trust); an untrusted project's tests run in a container (docker/podman) or are refused. Pytest@filearguments,-pplugins,-o addopts=, paths outside the copy or naming an environment variable, and the same in the project's pytest config files are refused. OS confinement without a container is not done yet.Running
scan/init/observewith the home directory itself as the project is not refused.Questions that are only partly about the code can come back
metwith verified but irrelevant claims.The reference resolver can still merge or drop references in some multi-reference sentences. Fixed on 2026-09-24: "PR #123 and issue #456 in psf/requests" bound both numbers to the local project's origin remote (an
owner/repoafter a preposition, or before "reposundaki"/"'teki", is now a repository reference when its sentence is about git objects; fractions, protocols, word pairs and folders such as1/3,HTTP/2,read/write,services/billingare not); a bare number keeps the origin remote over a repository named in another sentence, and with two repositories in a sentence each number goes to the one written after it. A name right before a version ("fancylib 2.31") is looked up in the project's ecosystems, then PyPI, npm and crates.io, only with--network on; by default it stays unbound, common words and units ("took 2.5 seconds", "macOS 14.2") are never package names, and only a clearly written name ("the X package", a name beforev0.3) makes the resultpartialwith a question.Fixed after that audit: user
claim add --kindwith unrelated text no longer verifies;.verinoda/and.git/files are no longer accepted as evidence;packagingis now a declared dependency.
Found by running Verinoda on its own repository (2026-09-23), not yet fixed:
updaterebuilds the code graph over the whole corpus when a file of the graph or a new code file changed (unchanged files come from the AST cache): about 33 s on a repository of about 2,100 files, about 27 s on Verinoda's own repository (about 1,200 files), 21 s when the edit leaves the graph as it was (34 s before the per-file caches and the kept graph of 2026-09-25, on the same quiet machine; 44 s before each path was resolved once per build, measured under load) and 94 s on Python's standard library copied as a project (2,305 files, 79,526 nodes; measured before that change), soverinoda ui --watchtrails an edit by that much; edits to other files (data files, documents outside the graph) only refresh the search index. The upstream incremental pass was faster (about 20 s) but lost the cross-file edges of every file it re-extracted, so after an edit a function's imports and calls into other files were missing until the next full scan (fixed 2026-09-24).analyzerefreshes first without charging it to its 60 s budget; on a project of 300 or more files where the last graph build took over 15 s (or over 200 files changed, or another build is running) it answers from the previous index and names the changed files instead, and the MCP server startsverinoda updatein the background. Runverinoda update .when you want the answer on the new code. Two builds of one project never run at once: the second waits (the CLI up to 10 minutes) or, where it may not wait, does nothing and says who is building.A frozen copy of the code inside the repository (here
benchmarks/corpora/heldout_repoatlas_7371990/) answered self-queries unless it was marked by hand (verinoda setup . --reference benchmarks/corpora=heldout,snapshot). Since 2026-09-24 scan and update find such copies themselves and rank them the same way (see Reference trees); a copy they cannot tell apart from the original still needs--reference. (Fixed on 2026-09-24: command names now map to their handlers,scankomutu ->cmd_scan.)
Game mods and data files (2026-09-24):
The kind of resource an id names is taken from a fixed table of command words and JSON keys; an id in plain Java is "kind not stated" (an inference) unless one file carries it. Bare names count only as an argument of an id constructor or of a helper whose name says what it loads.
The extra call pass covers Java and Kotlin only. Turkish stems that folding merges (
öldie /olbe) are matched as written only in retrieval; the question plan still reads folded words.Windows only so far. The POSIX resource limits and the container isolation path (docker/podman) are coded but have never been run.
experiment,observeandanalyze --run-tests/--observerun the project's own tests. Under process isolation (the default) the command's path arguments are confined to a throw-away copy, but the tests themselves can read and write outside it and use the network. Only run them on code you trust (experiment run --isolation containerneeds docker/podman and has not been tried against a real one).Test runs and observation are Python/pytest only. They use the project's own
.venv/venvwhen it has one, otherwise Verinoda's interpreter; pytest must be installed there, or the run isinconclusive. The fast tracer needs Python 3.12+ (sys.monitoring). The fallback is much slower, and child processes are not traced.Precise call resolution is optional (
verinoda[precise], jedi) and covers Python only. Without it, method calls on typed parameters staystrong_inference. Other languages need a SCIP index that you produce yourself (scan --scip FILE).Turkish questions work but trail English. The Turkish question sets and the intent gold table were written by the rule author, so those scores are in-sample.
Staleness of relation claims is tracked at the level of the whole calling function, so edits elsewhere in that function can mark a still-true claim stale.
Some reference mismatch codes fire only in narrow cases, PEP 740 provenance is not used for pinning, and there is no authenticated GitHub access.
No model-in-the-loop measurement exists. Token counts are chars/4 estimates, and the harness scores whether gold facts are present in the delivered context, not answer accuracy.
The sections below describe the product; where they are ahead of the code, the lists above say so.
Evidence-first codebase analysis for people and coding agents.
Verinoda looks at a repository as a running system: code, call and data
flow, tests, configuration, git history, decision records and, when you let
it, targeted test runs. It answers questions as claims with evidence.
Every claim carries a status (statically_verified, experiment_verified,
strong_inference, unknown, stale, …), the exact file:line or commit it
rests on, and what is still uncertain. When a claim can't be supported, it says
unknown and names the next check to run. Questions may be in English or
Turkish. References in them (repositories, versions, packages, PRs, papers)
are pinned to the exact version the user meant before anything is compared.
Verinoda is derived from Graphify (commit
20a20d30, Apache-2.0). It is an independent project, not an official Graphify release. See docs/UPSTREAM.md. The name "Verinoda" was checked as free on PyPI, npm and GitHub on 2026-09-23 (see docs/NAMING.md); 0.1.0 and 0.2.0 were published there on 2026-09-26 from the tagsv0.1.0andv0.2.0, 0.3.0 and 0.3.2 on 2026-09-27 fromv0.3.0andv0.3.2, 0.4.0 on 2026-10-01 fromv0.4.0(docs/RELEASING.md).
Related MCP server: Cordyceps Search
Install
Verinoda is a Python package (verinoda on PyPI, with a small verinoda wrapper on npm). Pick one:
uv tool install --link-mode copy "verinoda[precise]" # recommended: isolated, uv fetches a Python if needed
pipx install "verinoda[precise]" # the same with pipx
pip install "verinoda[precise]" # into the current environment (Python 3.10+)
npx -y verinoda --version # Node users: runs the PyPI release through uvx / pipx / a private venv
uvx --from "verinoda[precise]" verinoda --version # run once without installingOn Windows, get uv first with winget install --id astral-sh.uv -e (then open a new terminal) and run
uv tool update-shell once so verinoda is on PATH. [precise] adds the optional precise call-site resolver
(jedi); leave it out for a smaller install. Upgrade with uv tool upgrade verinoda, pipx upgrade verinoda or
pip install -U verinoda. Releases are published from tags by .github/workflows/release.yml
(docs/RELEASING.md). MCP clients that start servers with npx can use
npx -y verinoda mcp serve; verinoda setup registers the installed command instead.
The development version (main), straight from GitHub:
# Windows (PowerShell or cmd)
uv tool install --force --reinstall-package verinoda --link-mode copy "verinoda[precise] @ https://github.com/ozcinax-star/verinoda/archive/main.zip"# macOS / Linux (or Git Bash on Windows)
curl -LsSf https://raw.githubusercontent.com/ozcinax-star/verinoda/main/install.sh | shBoth development-version commands install from the GitHub archive with the optional precise resolver and copy files
instead of hardlinking them (so sandboxed agents such as Codex can import the package).
Run the same uv tool install ... line, or the script, again to upgrade. The script
(install.sh, read it first) installs uv if it is missing, runs that
command and uv tool update-shell; options are environment variables: VERINODA_REF
(branch, tag or commit; default main), VERINODA_EXTRAS (none to skip the precise
extra), VERINODA_NO_MODIFY_PATH=1.
Why there is no irm ... | iex one-liner for Windows: Microsoft Defender blocked
powershell -ExecutionPolicy ByPass -c "irm <script url> | iex" for this project's
script as Trojan:Win32/Commando.A!ml, a machine-learning verdict on that
download-and-run command line (the script file itself was not flagged). The plain uv
commands above avoid the pattern.
Then, once per project:
cd my-project
verinoda setup # index the code + connect Claude Code / Codex if they are installedverinoda setup is safe to re-run (it updates the index and leaves unchanged skills
alone). --agents claude,codex|all|none, --scope user for all projects, --no-mcp.
It ends with a first question about the project, built from its most connected
function or class, and how to open the graph.
Or skip it and just ask: in a git work tree without an index, the first
verinoda query, analyze, trace, map or ui indexes the project once and
says so on stderr (never outside a git work tree or in the home folder;
VERINODA_NO_AUTO_INDEX=1 turns it off):
cd my-project
verinoda query "where is the discount threshold configured?" # indexes first, then answersOther ways to install
Python 3.10+ (3.12+ recommended: the runtime tracer uses sys.monitoring).
Graphify (graphifyy) is not required; the extractor is part of this
package.
# from a checkout or a wheel you built (uv build)
uv tool install --link-mode copy . # or a wheel path
pipx install ./dist/verinoda-*-py3-none-any.whl
pip install ./dist/verinoda-*-py3-none-any.whl # into an existing venv
# optional: precise call-site resolution (jedi)
uv tool install --link-mode copy --with "jedi>=0.19.2,<0.21" .
pip install ".[precise]"
verinoda --version
verinoda doctorBuild the wheel with uv build --wheel (or python -m build --wheel).
Languages
The index (the graph behind query, trace, map, tq and the rest) reads these with what a plain install
brings: Python, JavaScript and TypeScript (also in Vue, Svelte and Astro files), Go, Rust, Java, Kotlin, Scala,
Groovy and Gradle, C, C++, CUDA and Metal, C#, Razor, XAML and .NET project and solution files, Ruby, PHP and
Blade, Swift, Objective-C, Lua and Luau, Zig, PowerShell, Elixir, Julia, Verilog and SystemVerilog, Fortran,
Bash, Dart, Pascal and Delphi forms, Apex, COBOL (programs, paragraphs, copybooks, PERFORM and CALL), JSON,
Markdown, PDF and Office documents, and Minecraft datapack functions.
These need an optional grammar, one extra each:
Language | Extra |
VB.NET |
|
R |
|
Erlang |
|
Solidity |
|
SQL |
|
Terraform / HCL |
|
OCaml |
|
Common Lisp |
|
DreamMaker |
|
Robot Framework |
|
pip install "verinoda[languages]" installs every optional grammar that ships prebuilt wheels (all of the
above but dm and robot, plus tree-sitter-pascal, which reads Pascal more closely than the built-in
fallback); R and Erlang come from tree-sitter-language-pack, a prebuilt wheel used offline.
Without its grammar a file is not extracted: scan and update name such files with the reason and the
install line (not_extracted), and the next update after installing the grammar reads them.
Codex on Windows: use --link-mode copy. With uv's default link mode
the installed package files are hardlinks into uv's cache. Codex's
workspace-write sandbox on Windows could not read them (PermissionError),
so the verinoda CLI failed inside Codex while the MCP tools still worked
(docs/AGENT-VERIFICATION.md). uv tool install --link-mode copy … (or
UV_LINK_MODE=copy) avoids this. verinoda doctor warns when the install
is hardlinked or editable (sandbox_readable). On such an install the Codex
MCP entry is registered with --profile full: the Codex sandbox may not
import the package, MCP is then the way in, and it must serve every tool the
skill uses. Reinstall in copy mode and run install again to get the default
(core) entry.
Quick start
cd my-project
verinoda setup # once: index + agent skills (or `verinoda scan .` for the index only)
verinoda map . --view dataflow # entry points -> persistence, with limits stated
verinoda map . --view dead # code nothing the project starts from reaches, as claims
verinoda map . --view hotspots # files and functions that change often and are complex
verinoda map . --view sides # Minecraft: client-only code reachable from server code, with the path
verinoda map . --view dsm # dependency structure matrix between folders (--depth N, --group-by tag)
verinoda map . --view model --model docs/workspace.dsl # a C4 model against the code's dependencies
verinoda map . --view repo --max-tokens 2000 # the files to read first and their signatures
verinoda ui # notes + graph in the browser (local)
verinoda ui --graph # ... opened straight on the graph view
verinoda ui --watch # ... and kept up to date while you edit
verinoda ui --export --open # the graph + file notes as one HTML file, no server
verinoda notes --changed # your own notes whose code changed since you wrote them
verinoda query "where is the discount threshold configured?" # plain-text context
verinoda query "retry path:src/** lang:kt NOT is:vendored" # filters narrow the ranked results
verinoda trace create_order_handler OrderRepository.save
verinoda when RepairScheduler.tick # when it runs: the event or caller, the conditions on the way
verinoda plan draft "Sipariş API'den veritabanına nasıl ulaşıyor?" # -> .verinoda/plans/plan-001.json
verinoda plan check plan-001.json # grounds every mention; 0 ready, 2 invalid, 3 needs clarification
verinoda analyze --plan plan-001.json # or: verinoda analyze "How does an order reach the database?"
verinoda resolve "compare with requests 2.31 sessions.py" # pin the references first
verinoda observe --for apply_discount # which tests reach it at runtime (isolated copy)
verinoda resolve-call orders/service.py:22 save # precise extra: which definition?
verinoda claim show clm_… # evidence (with grades), uncertainties, full history
verinoda challenge clm_… # adversarial re-check; can only lower confidence
verinoda update . # after edits: re-index changed files, mark affected claims stale
verinoda verify clm_… # re-check evidence (moved lines are relocated)Try it on the bundled example: copy examples/orders_app somewhere, git init
and commit it, then run the commands above inside the copy. For observe,
give the copy a .venv with pytest installed.
Commands
Command | What it does |
| Python, package layout, upstream base, graph/snapshot freshness, schema, claim counts, search index, lexicon, precise/SCIP availability, |
| One step per project, safe to re-run: |
| Create |
| Trust a project (recorded outside it): only then do its tests run with process isolation, and only then does its own |
| Full / incremental index + snapshot, then the derived search index, lexicon and symbol facts; |
| Between sessions: re-verifies stale claims with |
| Your own notes on the code (written in |
| notes and graph of the project in the browser: a note per symbol, file and data file, local and global graphs, search (see Notes and graph view); |
| The graph for other tools: GraphML (Gephi, yEd, Cytoscape), Neo4j Cypher ( |
| hierarchy, dependencies, dataflow, config, tests, history, impact ( |
| A trace (as |
| Bounded retrieval from the passage index; plain text for a model by default (skeleton first, each item with why it was chosen; nothing printed twice: no question echo, a signature once, windows dedented), |
| Directed paths, each hop with relation, confidence and call-site location; hints when an endpoint does not resolve; an endpoint that names several symbols is listed ( |
| Database tables read from the code: ORM models (SQLAlchemy |
| The cross-service links: the route table (Flask, Quart, Sanic, FastAPI with router and mount prefixes, a prefix given as a constant or a pydantic settings default resolved and cited, one nothing spells kept as |
| When a method runs: paths back through its callers to the event or scheduler that starts it (JVM registrations and lambdas: "at the end of every server tick", "80 ticks later"), each call with the conditions around it as written and its file:line; exit 2 when the name does not resolve (candidates listed) |
| One symbol with its callers and callees as two trees, or, for a type (the default for a class), what it extends and implements and what extends or implements it; each link with the |
| The whole function or class around a location, with its lines (a top-level statement or a Markdown section when there is no definition): |
| Git history, read only. |
| Code references in the repository's own Markdown, reStructuredText and AsciiDoc files checked against the working tree: Markdown link targets and inline code spans that read as paths, with |
| Requirement criteria traced to their evidence. A spec is a Markdown file in the specs folder ( |
| Git hooks that keep the index current: post-commit (not during a rebase), post-checkout (a branch switch only), post-merge and post-rewrite (the end of a rebase or an amend) start |
| Hooks in the agents' own hook files that put what the graph knows next to the agent's searches and reads, with no Verinoda call: after a Grep, or |
| Infrastructure as code linked to the code it runs: every Dockerfile (its last stage's |
| What the project says about one file: the enforced decision records whose accepted guards name it in a glob (only_in, no_edge, layers, allow_edges, public; |
| A CodeTour file ( |
| Path-scoped review rules. AGENTS.md, CLAUDE.md, BUGBOT.md and REVIEW.md in any folder (and |
| Saved searches that must not gain matches, kept in the committed |
| Who knows this code, read only. TARGET is a file, a folder, |
| A Mermaid diagram with its evidence as |
| The wiki outline: an overview page with the architecture diagram and a page per folder, or the pages a committed |
| Exact or regular-expression search (Python syntax, multiline: |
| Paths from untrusted sources to dangerous calls in the project's Python code (taint analysis), each with every hop at |
| Where a Python line's values come from: a backward slice inside its function - every statement whose definitions reach the line (reaching definitions over the function's control-flow graph, loops and |
| An inventory computed, not estimated: each named search (the trigram |
| A declarative query over the graph, answered with evidence: |
| Typed questions, batched: up to 20 closed questions in one call, each answered with one value - |
| Structural search: a code-shaped pattern such as |
| Every line a rename of one symbol would touch, without editing: the definition, the calls, imports and references the index ties to it (a call graded as |
| Simulate moving or renaming files and folders without editing: the graph's file paths are relabelled in memory, then the edge guards of the decision records ( |
| A numbered backlog item ( |
| The comments that say why: |
| AGENTS.md, CLAUDE.md, GEMINI.md, Copilot, Cursor, Windsurf and Cline rules and Claude Code's memory for the project, checked against the tree line by line: paths exist (with the case as written; a |
| The project brief, not the decision brief ( |
| Minecraft datapacks: entity tags checked but never added, tags added but never checked, objectives written but never read, objectives, teams and boss bars used but never declared ( |
| The stack traces and GameTest results of a log: the project's frames mapped to methods with their callers, the game's folded, a trace through a test's |
| Secrets (token formats, key blocks, |
| Where a GLSL uniform block field ( |
| Minecraft translation keys: keys missing from a locale or only in it, written twice, placeholders that differ from the default locale ( |
| Access wideners and access transformers against the class files of the build's classpath ( |
| Mixins against the bytecode of their |
| Question plans: draft from the message (TR/EN rules), check and ground a plan file, print the schema, re-judge an analysis' sub-questions later |
| Budgeted loop per sub-question → claims + evidence + critique + unknowns, each sub-question judged against its |
| Inspect claims, or record one with source evidence ( |
| Claims in time. |
| Derived facts: a result kept under a name with what it rests on. |
| Re-check evidence against the current tree (anchored relocation; optionally re-run its test) |
| Critique and counter-hypothesis probes; lowers status/confidence when support is weak |
| Pin every reference in a message to the exact version meant; reports mismatches and questions for the user |
| Reference repo pinned to an exact commit (or the pin from |
| Assumption diff: data structures, errors, concurrency, environment, dependencies |
| User critique handled as a hypothesis → confirmed / qualified / corrected / unresolved (references resolved first) |
| Isolated targeted experiment on a copy of the working tree, or with |
| Decisions stay human (D33). |
| Debug ledger (D34): every attempt at fixing one symptom with the tree it ran on, the patch vs the base, the failure at |
| The ledger; the working tree against the base or an attempt; close (resolved needs a pass of the repro command on the current tree, no Verinoda run of that tree failing, and the user's decision on any test changed since the first attempt; when the baseline passed, a failure seen only after an edit is the edit's own and does not count as the symptom; Verinoda never says "fixed") |
| Strategies in throw-away copies, each recorded: the repro at the base with the symptom's test held fixed and the diff's hunks ranked; a bisect that runs both ends first (over commits, or with |
| Each test's pass rate over the runs the debug ledger made (every run Verinoda makes of a repro, |
| Change review (D35): the changed definitions, their dependents with via-chains, findings by concern (persistence, security, performance, public API, config, entry points), a verdict per public definition added, removed or re-signed ( |
| SARIF 2.1 in and out. In: a linter's or CodeQL's SARIF files read as evidence, each result a claim at its |
| JVM checkers' findings read as evidence; nothing is run, the files are logs and reports the build already wrote. Error Prone and NullAway diagnostics as javac prints them ( |
| The workspace packages of a monorepo a change affects (the same block as |
| Coverage reports read into lines and symbols, in any language: lcov (istanbul, c8, Jest, Vitest, cargo-llvm-cov, ...), Cobertura XML (coverage.py, gcovr, ...), JaCoCo XML (Gradle, Maven) and coverage.py JSON, found at the usual paths (also in sub-project folders) or given with |
| Code health per function (default: every code file that is not a test): cyclomatic and cognitive complexity, deepest nesting with its line, length and parameters, counted on the syntax tree (Python |
| Mutation testing scoped to the diff, Python: on the lines changed against the base (default HEAD) in files that are not tests, one mutant per operator site (a comparison, |
| Behaviour probe (D36, Python): generated inputs on the old and the new version of a changed function in throw-away copies; behaviour differences, new exceptions and non-determinism with examples; refuses functions with side effects it cannot isolate. With |
| Pinned regression tests (Python, opt-in; runs the project's tests, so trusted projects only, like |
|
|
| Precise resolution of one call site (needs the |
| The installed language server (opt-in, CLI only; never started in a project you have not trusted: tsserver plugins, Gradle and Maven imports, build scripts and macros run in it). Server: |
| Python, Java, Kotlin, and TypeScript/JavaScript imports. Do the names code uses exist? Python: imports, from-imports, attributes, keyword arguments and constant dict keys, checked in the project's own environment; Java: imports, types, methods with their number of arguments, fields, constructors and Mixin targets; Kotlin: imports, types, members and properties (D45); TypeScript/JavaScript: imports (D46); against the project, its classpath (Loom or |
| Opt-in package existence and squatting check over |
| Declared against used dependencies: pyproject/requirements (with PEP 735 and Poetry groups), package.json (with workspaces) and Gradle/Maven builds against the imports of the project's own files. Four findings, each a claim with |
| The real members of a Python module, class or function in the project's environment, or of a Java class as the build sees it (project, classpath with the Minecraft jars, JDK: signatures, access, where inherited members come from), with signatures, file:line and the installed version; |
| Versioned learnings, invalidated (never deleted) with their source claim, when their time-to-live ( |
| Skill (+ MCP) for Claude Code ( |
| MCP server (stdio) over the same core, 44 tools; |
| The HTTP server in the background: |
| The HTTP transport's bearer token (generated on first use, kept in the user config folder readable by you alone) and the file it is in; |
| Names for project folders one MCP server serves together ( |
| Several repositories, each indexed on its own (registered names, |
| The ready workflows the MCP server also offers as prompts ( |
| The Graphify-derived CLI (advanced, unsupported); installer, hook and |
| Raw search vs Graphify baseline vs Verinoda on question sets with gold facts; |
| Staleness harness (history replay, mutation suite), critique precision/recall, the wrong- |
Exit codes: 0 done, 1 error, 2 usage error / invalid plan / blocked command /
a trace with no path or an endpoint that does not name one symbol / an impact
--target that does not name one symbol / a health path that matches no code file,
3 "needs more" (clarification, partial resolution, refused experiment,
incomplete observation, no precise answer, an absent name or a lock mismatch in check, a name
api did not find, a debug attempt that says stop, a debug strategy that did not settle it;
decide check exits 1 on VIOLATED, 3 when something could not be checked and 2 on an error);
4 (check, api): nothing absent, but something asked for was not checked (another language,
a file that does not parse). Every command except memory, mcp serve and index accepts
--json (plan schema always prints JSON); JSON is compact when stdout is not a terminal.
In CI: commit the decisions folder and name it in verinoda.toml ([decisions] /
dir = "docs/decisions"); run verinoda decide check --base origin/main (exit 1 violated,
3 not checked) and verinoda check --diff origin/main (exit 3 an absent name, 4 a changed file
it does not read; a clean checkout has nothing changed against HEAD, so check then says
nothing_to_check).
Notes and graph view
verinoda ui opens the project as linked notes in the browser, in the manner of
Obsidian, built from Verinoda's own index rather than from hand-written notes:
A note per symbol, source file, document section and data file: qualified name (
Wisp.spawn(),search_index.rank()), signature, doc text, the code (highlighted, with line numbers), and its links in sections: defined in, members, calls, called by, extends / implemented by, imports, imported by, references, the data files it names by resource id and the lines that name it, and the claims recorded about it with their status. A link the tool inferred rather than read in the code is marked?.Local graph beside every note (depth 1 to 3; tests, data files and external types can be hidden) and a graph view of the whole project at file level, coloured by folder (or by community), with a filter that highlights matching notes; files with no links ring the linked ones, and a note's links list the project's own code before tests. Both are force-directed: drag, zoom, hover to see a note's neighbours, click to open it.
The graph in 3D (3D in the graph view, or
V): the same files and links laid out in three dimensions and drawn with perspective, no WebGL or library; the flat graph inflates into depth when you switch. Click a file and the camera flies to it; a panel says what it is in words ("extract.py is used by 136 files and uses 42"; a document mentions files, a data file is named by code) and numbers its linked files:1-9fly to one,Nwalks round all of them,Backspacegoes back along your trail,Ffollows the file (the camera circles it),Ilights up what a change to it may affect,Enteropens its note. A region (a folder, or a community when coloured by community) is framed with its busiest files and the regions it works with most;[and]step through them, and Tour (T) visits the product's own regions first, then tests, examples and docs, one sentence each.Pputs the selected file or region on a watch list: a chip at the bottom brings you back to it, a small diamond marks a watched file, and with the local server the page says when a watched file has changed since the index (it looks every 15 seconds; the list stays in your browser). Measured on Python's standard library as a project: about 60 frames a second in Chrome with 1,745 files and 8,259 links.Command bar (
Ctrl+K), in English or Turkish:focus rank/odak rankflies to a note in 3D,region ui/bölge uiframes a region,impact store/etki storeandpath parse to rank/yol parse ile ranklight up the files concerned,tour,changed,open …, or a question, which is answered as below; with an empty bar it lists what it can do.?shows every shortcut;Escundoes one step at a time (the tour, what is lit, the view).Search by name (exact, prefix, part of the name, path) or with a question, which runs the same ranking as
verinoda query; a file tree; back and forward; Turkish and English; light and dark. A question (three words or more, a question word, or a?) also offers Answer the question: the passagesverinoda queryanswers with, each with its lines (highlighted, opening the editor), why it was chosen and its note, then the other places found. The search index is never written for it (not in an exported file: it has no code).The line a link is written on under each call, import, reference and resource-id link (not for a file edited since the last index: its line numbers would point elsewhere), and open in your editor: every
file:lineand a button on each note open VS Code, Cursor or VSCodium at that line (chosen at the top of the page).Preview on hover: resting the mouse on a link to a note shows its kind, file and line, signature, first doc lines, your note on it, its link counts and the first lines of its code, without leaving the page.
What changed: Changed in the graph view rings the files edited, added or deleted since the index (what
verinoda updatewould take in) and, in another colour, the files that use them; a note whose file changed since the index says so, since its links and lines may be off.Impact and path: Impact on a note lists what may be affected when it changes: what calls, imports, extends or names it, then what uses those, up to three links back (the project's own code first, tests on or off), and turns the local graph into that set; Path… finds the shortest chain of calls, imports and references from the note to another one, or the other way round. In an exported file both work at file level.
Butterfly: Butterfly on a function, method or class puts the note in the middle, what calls it on the left and what it calls on the right, each a tree one to four links out; for a class it opens on the inheritance tree (what it extends or implements, and what extends or implements it, library types as leaves). Every link shows the line it is written on and the ceiling of an unchecked graph edge (
strong_inferenceorweak_inference), and the local graph turns into the same set.verinoda butterfly NAMEprints it (not in an exported file: it has no symbol graph).Notes of your own on any symbol, file, section or data unit: plain Markdown (
**bold**,`code`, lists,[[Name]]links another note), kept as.mdfiles in.verinoda/notes/(notes.dirin.verinoda/config.jsonputs them in a folder you commit). Each note is anchored to the code it was written about (a symbol or section by its fingerprint, found again wherever it moved; a whole file or a data unit by a hash of its lines) and shows its status: up to date, code changed (read it again, then Read it: still right, which anchors it to the code as it is now) or code gone (the symbol was renamed or deleted: edit it onto something else or delete it). The start page lists your notes, the changed ones first;verinoda notes --changeddoes the same on the command line and exits 1 when any note needs reading again, for CI.
It is local: a standard-library server on 127.0.0.1 (a free port unless
--port is given) that answers only requests addressed to that host and port;
the page loads nothing from outside (no CDN, fonts or telemetry;
Content-Security-Policy: default-src 'none', scripts and styles only from the
server). The one thing it writes is your notes: POST /api/usernote needs the
random token of that server run, which only the page it serves carries, JSON, and
this origin, so another site cannot write through it; --read-only turns writing
off. It follows the index: a few seconds after verinoda update (or any
rebuild) the open page redraws the note or graph it shows, keeping its scroll
position, and says so (an unseen tab looks when it is shown again).
verinoda ui --watch also runs verinoda update itself when the project's files
change (once the edits stop; one update at a time; while another process
builds the index, the watcher skips that round, the page says why, and it
tries again soon). The operating system's file events wake it
(ReadDirectoryChangesW on Windows, inotify on Linux, watchdog elsewhere when
installed); an event only wakes it, the files' sizes and times still decide
whether anything changed, and without events it polls as before.
verinoda mcp serve --watch does the same for the MCP server, with the fast
update index_update runs. verinoda ui --graph opens
straight on the graph view.
One file, no server. verinoda ui --export [FILE] writes the graph view and
a note per source file, document and data file into one HTML file (default
.verinoda/index/verinoda-graph.html; about 4.4 MB for Verinoda's own 1,150
files) that opens with a double click, or with --open right away; the command
also prints its file:/// address. Checked in Chrome and Edge from file://.
It is the same page with its data inside:
the graph with its filters and colours, the file tree, each file's links, outline
and claims, a file-level local graph and a name search (a symbol opens the note
of its file). It holds no code (verinoda ui shows it) and no path of the
machine it was made on (the project root and the home folder are taken out of
every name); its Content-Security-Policy allows only its own script and style
(by hash) and no connections, so it fetches nothing. It is a snapshot: export
again after verinoda update. (The index step no longer writes the upstream
Graphify graph.html, which loaded vis-network from a CDN when opened, and
removes an old one; GRAPHIFY_VIZ_NODE_LIMIT set to a positive number keeps it.)
Limits: the global graph shows at most 2,500 files (the best connected ones, and it says how many it left out); code is read, not edited; the exported file has file notes only (no symbol notes, no code, no question search) and shows your notes read-only.
Measured on Python's standard library copied as a project (2,305 files, 79,526
notes, 140,342 links; Windows 11, headless Chrome): the server starts in 1.6 s,
the start page shows in 1.5 s, a search in 0.8 s, a note in 0.5 s; the graph
view (1,745 files, 8,249 links) draws in 0.4 s and runs at about 60 frames a
second while it settles; impact and path on the most connected notes take
under 20 ms. --export writes 9.3 MB in about 6 s. With --watch each look at
the tree takes 0.2 s, and an update takes as long as verinoda update (above,
Known issues). On Verinoda's own checkout (2,735 files listed, one look 0.92 s)
an edit wakes the update 1.2 s after the save with file events, against 12-22 s
when polling (a slow listing spaces the looks out), and an idle watcher uses no
CPU (polling: 1.5 s of CPU every 20 s).
Game mods, data packs and other data files
Code often names its data only through strings: a Minecraft mod runs the
data-pack function mymod:wisp_death, loads config/mymod.yml, registers
the item whose model is assets/mymod/models/item/x.json. Verinoda indexes
those files too and follows the strings between them.
Data files are searchable. Text files the code graph has no node for (
.mcfunction, JSON, YAML/TOML/INI/properties configs, SQL, shaders, CSV, Gradle scripts, skipped sources) becomedataunits; configs are split by top-level section. Files it leaves out are listed with the reason (binary,may hold secrets- the graph's own secret rule -,.graphifyignore, dependency or build-output folder, generated output such asresults/orlogs/, large generated JSON, minified);doctorcounts them and a query names a left-out file whose name matches the question.Resource ids link code and data in repositories that are packs or mods (a
pack.mcmetaor a Fabric/Quilt/Forge/NeoForge manifest):ns:pathids,#ns:tags, worldgen ids,function ns:x, translation keys"item.ns.x",Identifier.of("ns", "x"), full asset paths, and bare names passed to an id constructor or a helper whose name says what it loads (runFunction(server, "wisp_death")). The context picks the registry (advancement revoke ... only ns:xnames an advancement,"parent"a model). Query output showsnames:/named by:lines; a link whose namespace is assumed or whose kind the line does not state is marked inferred, and ananalyzeclaim built from it isstrong_inference.Identical copies of a data file (a data pack shipped twice) rank once, as the copy in the source set; the others are listed as
same content:.Java and Kotlin calls the extractor drops (a class name that exists twice in the repository, calls through typed variables, Kotlin calling Java) are added when the file's imports or package bind the class, and graded like any call site.
Callbacks (
docs/DESIGN.mdD38): a method reference passed on (END_SERVER_TICK.register(RepairScheduler::tick),createTickerHelper(..., Block::serverTick)) is aregistersedge, never a call and never weighed by the ranking.tracefollows it when no call path exists and labels the hopcallback; impact, the UI andreviewlist the method that registers a changed one; a claim "A calls B" that only a method reference supports staysweak_inference.map --view dataflowstarts at the mod's entry points (fabric.mod.json, Fabric initializers,@Mod,@SubscribeEvent, mixin handlers, registered callbacks) and knows JVM file writes,NbtIoand dirty flags; all of these are text heuristics, each with its reason.Reference trees:
verinoda setup --reference original-plugin/=original,pluginkeeps an original implementation searchable but ranks it at 0.6x unless the question says "original", "plugin" or the folder name. A folder that holds a copy of the project's own code (a benchmark corpus with an older version, a vendored snapshot) is found at every scan and update and ranked the same way: most of its files have a twin elsewhere defining the same names, nothing outside it uses it, and the twins are in code the project does use (copies.json;index.not_copiesorindex.detect_copies: falsein.verinoda/config.jsonundo it). A port next to its original with nothing else using either is left to--reference.Turkish names from the repository: parallel locale files (
lang/en_us.json+lang/tr_tr.json,locales/en.json+locales/tr.json) teach the lexicon that "Fener Asası" islantern_staff.
examples/glow_mod/ is a small fictional Fabric mod with a data pack, a
config file, a reference tree and a copied data pack; its question set
(glow_mod, 14 questions, 7 Turkish) was written without running Verinoda
on it. Results: docs/BENCHMARKS.md.
Coding agents
verinoda install --agent claude --scope project # .claude/skills/verinoda/SKILL.md + .mcp.json entry
verinoda install --agent codex --scope project # .agents/skills/verinoda/SKILL.md (+ MCP config)
verinoda install --agent cursor --scope project # .cursor/rules/verinoda.mdc + .cursor/mcp.json entry
verinoda uninstall --agent claude --scope project # removes only what install recorded
verinoda setup . --agents all # every supported agent found hereAgent | MCP server entry | Instructions |
Cursor |
|
|
Gemini CLI |
| a marked block in |
GitHub Copilot (VS Code) |
| a marked block in |
Kiro |
|
|
Continue |
|
|
Aider | none (Aider has no MCP client) |
|
Claude Code:
/verinoda how does checkout reach the database?Codex: mention
$verinodain the prompt. (Codex has no/verinodacommand.)
The skills describe the working method; all logic lives in the CLI/MCP core.
They read the CLI's plain text and add --json only for a field the text
leaves out (JSON cost 2-5x the tokens for the same content).
Two protocols come first:
References the user gives. When the message has links, repository or package names, versions, commits, PR/issue numbers, papers or docs, the agent runs
verinoda resolve "<message>"(MCPreference_resolve) before researching or answering. It reports each reference as<name> @ <pin> (basis: …)with each mismatch on its own line. It never substitutes the default branch for a version the user named, asks only the returnedquestions_for_user, and states every unresolved part with its next step.Understand the question first.
verinoda plan draft "<message>"(MCPquestion_plan_draft). Then the agent edits the plan: it splits compound questions, glosses domain words, copies versions exactly as written, and never invents candidates. Thenverinoda plan check: exit 0 ready, 2 invalid, 3 needs clarification. The agent asks only the returned clarifications (AskUserQuestionin Claude Code;request_user_inputor plain text in Codex) and records the answers. Thenverinoda analyze --plan <file>. The answer starts with "Understood as / Anladığım: …", followed by one block per sub-question with its verdict, claims and unknowns.Check the names code uses (Python, Java). After every edit, and before proposing code, the agent runs
verinoda check --diff(MCPcode_check; code not written yet:--stdin --as <path>). It never keeps anabsentsite: it fixes it fromnearest/elsewhereor fromverinoda api <module.or.Class>(MCPapi_members).unknownis unverified, not fine. The report names the environment it checked. It reads Python and Java: a Kotlin or TypeScript file, or one that does not parse, comes back undernot_checked(exit 4), and the agent says so instead of reporting a pass.Confirm your own sentences. A sentence the agent writes about the code is recorded with a typed kind (
claim add --kind relation|config|order|location --symbol X) or a verbatim quote; plain prose is at mostweak_inference, and anot_foundname is reported, not replaced.Decisions are the user's. For a should-we / which-one question the agent runs
verinoda decide brief, asks thequestions_for_human, records the user's explicit choice (decide record) and never picks for them; before finishing a code change it runsverinoda decide check --changed.Keep a debug ledger.
debug startbefore the first edit of a bug fix,debug try --hypothesisafter every edit; onstopit stops editing and follows the first strategy, and it asks before changing a test's expectation.
Then the evidence discipline: report claims with their status, never upgrade
a status by wording, challenge what you rely on, report unknown with its
next step, and treat user critique as a hypothesis (feedback add --process).
How claims stay honest
Relevant evidence, checked in one place. A
*_verifiedstatus needs one evidence group that is verifying, fresh and mechanically entails the claim (for example: an AST call to the target at the cited line inside the claimed caller; a definition spanning exactly the cited lines). Every stored status change passes through this check, so unrelated evidence cannot verify a claim on any path (API, verify, feedback, experiments, runtime runs, MCP).Word overlap never verifies. Every word of "apply_discount returns the subtotal above the threshold" is in the lines that return
subtotal * 0.9there. Term coverage makes evidence relevant (partial), never a verification; only a verbatim quote (path:12 contains: <text>, which verifies the quoted text and nothing around it) or a kind's typed check (call site, definition span, environment read, call order) verifies. A written claim that states more than its check binds (another callee, a condition or bound, a negation, the arguments of a call, another file than the cited one) stays unverified.Every role is bound. A written relation must name the caller and the callee in the right direction ("OrderRepository.save calls place_order" is checked against
save's body); a written config claim ("the discount threshold is read from ORDERS_MAX_ITEMS") must be about the name the read is bound to.claim addanswers a definitive miss at once, with its scope: "no direct call to save in create_order_handler (orders/api.py:16-21); calls through other names are not followed".A name written as code is never replaced by a similar one.
analyze,plan checkandtracereportnot_foundwithdid_you_mean("no symbol namedplace_ordersin this repository; nearest: place_order (orders/service.py:19)"); the sub-question isunmet. A name spelled only in a file changed since the index is "not in the index yet"; one the index spells but has no symbol for (a constant, an attribute) isnot_a_symbol, with where it occurs.trace,map --view impactandnode_inspectshare one resolver: a detected copy of the project gives way to the original, test, example, fixture and vendored code to the product's own, and a name that several symbols still carry is listed, not picked (passpath/file.py::Name).A stale index is never silent. Reading commands list the files changed since the index;
analyzesays when it answered from the previous index, and a changed file that spells the question's subject caps that sub-question atmet_with_inference.A graph edge (
EXTRACTED/INFERRED) is never enough on its own. Search results, model summaries and user feedback are not even support for an inference. A claim with no evidence isunknown.Definitive vs heuristic refutation. Only an exhaustive check within a stated scope (no call to the target on the cited line, a precise resolver's definitive different target, …) makes a claim
contradicted. A heuristic doubt lowers it one step and adds an uncertainty.Facet-level staleness. Claims depend on symbol facets (signature, body, name bindings, doc sections, the test set). An edit makes a claim
staleon the nextupdate/analyzeonly if something it depends on changed. Code that only moved is relocated through anchors, and a duplicated line is reportedambiguousrather than guessed.Critique and re-verification never raise a claim above its assessed ceiling. Critique is idempotent and never restores a stale or contradicted claim.
Runtime observations are run-scoped ("observed in run R at commit C"). They never support an "always" claim, and calls seen through test doubles never support production edges.
Nothing is deleted: user corrections supersede (the old claim is kept as
contradictedwithsuperseded_by), claim text is immutable, and history, plans, reference resolutions and runtime runs are append-only.Heuristics state their method and limits (
coverage.limits,uncertainties,derived_by). Budget exhaustion or irrelevant retrieval yieldsunknownwith the next verification step.
Name check (Python)
verinoda check answers one question for code an agent (or you) just wrote: do
the modules, functions, methods, keyword arguments and dict keys it uses exist,
in this project's environment?
$ verinoda check orders/ai/export.py
orders/ai/export.py:7:28 ABSENT import orders.service.place_orders
not found in module orders.service in this project (orders/service.py)
nearest: place_order (orders/service.py:19)
orders/ai/export.py:15:43 ABSENT kwarg compute_total(currency=)
keyword currency= not found in the signature compute_total(items: list[dict]) (orders/pricing.py:6)
orders/ai/export.py:19:10 unknown attribute repo.save_order
`repo` is a parameter: its runtime type is not knownWhich environment:
--env PATH, else the project's.venv,venvorenvwhen its base interpreter is a known Python installation outside the project (otherwise the note names the program--env .venvwould start), else Verinoda's own interpreter for the standard library only: third-party names are thennot_installed, neverabsent. The MCP tools never start a program from the project.When it says absent: only when the container's names are all known (a module without
__getattr__or dynamic writes, a class without descriptors or code that sets attributes from outside, an instance made right there, one known signature without**kwargs, the dict literals a function returns), and jedi also found nothing. The wording is "not found in as installed in ()", never "does not exist".Unknown is not fine: parameters, annotations, inferred return values,
**kwargs, module__getattr__, names assigned elsewhere,sys.pathchanges inconftest.py, and standard-library names of another platform or Python version (collections.Mapping) stayunknownwith the reason.What it does not read is never a pass: a file in another language, a notebook, a Cython file or a Python file that does not parse is listed under
not_checkedwith the reason; with nothing absent the exit is 4 (3 means an absent name or a version that differs from the lock).The project's own type checker (opt-in):
--checker tsc|pyright|mypy|autoruns the checker the project already has (node_modules/.bin/tsc, the virtual environment's pyright or mypy, else PATH) with the project's configuration, and lists its errors on the files and lines being checked as sites: a name, member or import it did not find isabsent, a call or type that does not fit ismismatch(exit 3), eachobservedwith the tool, its version, the configuration and the diagnostic code. TypeScript files the compiler read are then type-checked, not imports only. Nothing is installed or downloaded (nonpx, nopip; pyright's PyPI wrapper is not started); a checker that is not found, runs past--checker-timeout(300 s) or prints output that cannot be read is exit 4 with the next step. Trust is the boundary: in a project you have not trusted (verinoda trust), no program inside the repository is started, pyright is not run, and mypy is not run when the repository's configuration namespluginsor apython_executable. CLI only: the MCP server never starts a program of the project.verinoda api packaging.specifiers.SpecifierSetlists the real members before a call is written. Existence and signature shape only: a real name used wrongly is not detected.
Documentation
docs/ARCHITECTURE.md — modules, state on disk, invariants
docs/DESIGN.md — design decisions D1-D40 and their implementation status
docs/BENCHMARKS.md — measured comparison (no unmeasured savings claims)
docs/UPSTREAM.md — Graphify base commit, feature inventory, port method, runtime patch
docs/UPGRADING.md — versioning, schema migrations, calibration changes, derived files
docs/AGENT-VERIFICATION.md — what was verified with the real agents
docs/GENEL-BAKIS.md — Türkçe genel bakış (ürün sahibi için)
docs/NAMING.md — name availability
docs/RELEASING.md — how a release reaches PyPI, npm and GitHub
License
Apache-2.0 (see LICENSE); portions originally under MIT (LICENSE-MIT).
NOTICE records the Graphify origin and the modifications.
Available Tools
5 toolsanalyzeA
Answer a question as claims with evidence: one verdict per sub-question, claims with file:line, unknowns with next steps, and the passages. needs_clarification is a normal result (ask the user). run_tests / observe also run or trace the tests that reach the answer (isolated copy). Re-indexes first if the tree changed.
| Name | Required | Description | Default |
|---|---|---|---|
| question | Yes | The question. | |
| budget_seconds | No | Wall-time budget in seconds (1-600). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, and the description explains why: it re-indexes first if the tree changed. It also discloses that test runs/traces happen on an isolated copy and that needs_clarification can be a legitimate result. That is meaningful behavioral context beyond the annotation flags, though the side effects of re-indexing are not fully spelled out.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded in the first clause, which is good, but the remaining text is a dense run-on mixing result format, clarification behavior, test execution, and re-indexing. Phrases like 'run_tests / observe also run or trace the tests' are cryptic and cost the reader clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must carry the return contract, and it does: verdicts, claims with file:line, unknowns with next steps, passages, plus clarification and test behavior. For a complex analysis tool this is fairly complete, though side-effect and failure-mode detail remains thin.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with two parameters (question, budget_seconds), so the schema already documents both. The description adds no syntax, format, or budget-interaction detail beyond what the schema provides, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource: it answers a question and returns verdicts, claims with file:line, unknowns with next steps, and passages. That is far more concrete than the bare name 'analyze' suggests. However, it never names or contrasts with the siblings (project_query, code_check), so an agent cannot cleanly route between them from this text alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied (pose a question, expect claims/evidence) and it usefully notes that needs_clarification is a normal outcome to escalate to the user, plus the budget parameter exists. But there is no explicit when-to-use vs. when-not, no prerequisites, and no routing guidance against project_query or code_check, leaving selection to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
code_checkARead-onlyIdempotent
Python, Java, Kotlin, TS/JS imports; other languages: not_checked (exit 4). For code you wrote or edited: do the modules, names, methods, arguments and keys it uses exist in the project and its environment? Input: paths, diff (a revision; nothing: changes against HEAD), or snippet + as_path. Each site: exists | absent (nearest names) | unknown | not_installed | guarded; exit 3 = absent.
| Name | Required | Description | Default |
|---|---|---|---|
| deps | No | Instead: the manifests' dependencies vs the imports. | |
| diff | No | A revision (e.g. 'HEAD'): only lines changed against it. | |
| paths | No | Repository-relative files or directories. | |
| as_path | No | With snippet: the file it is for. | |
| snippet | No | Code not written yet, checked as if it were in as_path. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only/idempotent/non-destructive, so the bar is lower, and the description still adds real behavioral detail: unsupported languages return not_checked (exit 4), exit 3 signals an absent site, and the per-site vocabulary (exists/absent/unknown/not_installed/guarded). This is well beyond what the annotations convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the language coverage, which is useful, but the rest is compressed into telegraphic fragments ('Input: paths, diff (a revision; nothing: changes against HEAD), or snippet + as_path') that are hard to parse. Every element earns its place, but the structure is cramped and reads like notes rather than guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly carries the full burden of explaining results via exit codes and site statuses, which it does. It stops short of explaining some statuses precisely (guarded, not_installed) and the exact output shape, leaving minor gaps for a tool with five optional params.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description meaningfully supplements it: diff means 'a revision; nothing: changes against HEAD', snippet pairs with as_path, and deps switches to manifest-vs-import comparison. That default-behavior clarification is genuine added value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: checking whether modules, names, methods, arguments and keys referenced by code 'exist in the project and its environment,' and names the exact languages covered. It is not tautological and scopes the tool well, though it never explicitly distinguishes itself from the vague siblings (analyze, project_query).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an implied use case ('For code you wrote or edited') and enumerates the three input modes (paths, diff, snippet + as_path), plus a one-line alternative-mode hint for deps. However it never states when to prefer this over siblings like analyze or project_query, and offers no exclusions beyond the language restriction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
index_updateA
Re-index after editing files (on a folder never scanned: the first scan); claims whose files changed become stale. mode: noop | incremental | full | first_scan.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare a non-read-only, non-idempotent mutation, and the description usefully adds that claims whose files changed 'become stale' — a real behavioral consequence. However, it never explains what re-indexing actually rebuilds or whether concurrent runs are safe, and the listed modes are unexplained, so it is adequate rather than rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences with the core action front-loaded and no filler. The parenthetical and mode tail are slightly compressed/awkward but not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with annotations covering safety and no output schema, the description covers purpose, trigger, effect on claims, and mode options — enough for correct invocation. The unexplained mode list and schema mismatch are the only gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema declares zero properties, which would baseline at 4, but the description introduces a 'mode' parameter with enum values (noop | incremental | full | first_scan) that do not exist in the schema — a confusing mismatch rather than genuine parameter elucidation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Re-index after editing files') and adds scope context (first scan on a never-scanned folder). It does not differentiate itself from siblings like analyze or code_check, which it plausibly overlaps with, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The trigger is implied ('after editing files') and the parenthetical hints at the first-scan case, but no alternative tool is named and there is no explicit when-not-to-use guidance. Usage is inferable but not spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
project_queryARead-onlyIdempotent
Where is X / what handles Y: ranked code locations as plain text (path:lines headers, call outlines, numbered source lines; up to 6000 chars, truncation stated). format='json' for programs. Hits are leads, not verified claims.
| Name | Required | Description | Default |
|---|---|---|---|
| format | No | 'text' (default): plain text for reading, skeleton first; 'json': structured items. | text |
| question | Yes | Question, symbol or file names to look up; filters narrow it: path:GLOB lang: symbol: is:vendored /regex/ AND OR NOT. | |
| max_items | No | Maximum code locations to return (1-25). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, and closed-world behavior, so the safety profile is covered. On top of that the description discloses output shape (path:lines headers, call outlines, numbered source lines), a 6000-char cap with stated truncation, and the important caveat that hits are unverified leads rather than confirmed claims.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Extremely dense and front-loaded: purpose first, then output shape, then format switch, then the reliability caveat. The long parenthetical is information-rich rather than padded, though the telegraphic style borders on terse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so: it describes the plain-text layout, the char cap, and truncation reporting, plus the 'leads not verified claims' caveat. This is nearly complete for a read-only query tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents question, format, and max_items with enums and defaults. The description's only parameter-level addition is the 'format=json for programs' guidance, which largely restates the enum meaning, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: answering location/handler questions by returning ranked code locations. The 'Where is X / what handles Y' framing makes the retrieval purpose immediately concrete. It does not differentiate itself from siblings analyze, code_check, or run_tool, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: it is clearly a lookup tool, and 'format=json for programs' hints at a caller profile. However, no when-to-use/when-not conditions and no alternatives (analyze, code_check) are named, leaving the agent to infer routing between siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_toolB
Run one more Verinoda tool (name, arguments). node_inspect {name}: a symbol's definition and edges with file:line; relation_trace {source, target, mode?: flow|any}: call paths between symbols; map_view {view: hierarchy|dependencies|dataflow|config|tests|history|impact|cycles|outline|dead|hotspots|sides|repo|saved, targets?}; claim_list {status?}, claim_inspect {claim_id}, evidence_inspect {evidence_id}: earlier claims, evidence re-checked; change_review {targets?, change?: body|signature|remove} before editing, {since_last?} after: what it touches; history_search {text, regex?, path?}: when text came/went; {symbol}: its commits; {message?, author?, since?, until?, diff?, path?}: commits; {base, head?}: compare.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | The tool. | |
| arguments | No | Its arguments, e.g. {"name": "Cls.method"}. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, so the safety profile is partly covered. The description adds semantic context about what each sub-tool inspects (definitions, call paths, claims, history), but never says which sub-tools mutate state or what side effects a dispatch can cause, which matters given readOnlyHint=false.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The opening sentence is front-loaded, but the rest is a dense run-on catalog using inconsistent notation (bare names, {braces} for some optional params, '?' markers) that is hard to scan. The information is dense but poorly structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-sub-tool dispatcher with an open-ended 'arguments' object and no output schema, the catalog is the necessary core and is mostly present, but two enum values are undocumented and per-sub-tool argument semantics are terse or missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema documents 'name' (enum) and leaves 'arguments' as a free-form object with only a generic example, so the description carries the real payload documentation by spelling out each sub-tool's arguments. It falls short of 5 because grep_context and read_context appear in the enum but are never described.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The lead sentence states a specific verb and resource ('Run ... Verinoda tool') and the enumeration makes clear this is a dispatcher over named sub-tools. However, the phrasing 'one more' is odd and no sibling tool (project_query, analyze, code_check) is referenced, so the agent gets no differentiation from alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The catalog implies routing — e.g. 'change_review before editing, {since_last?} after' and 'history_search ... when text came/went' — which helps pick a sub-tool. But there is no guidance on when to call run_tool at all versus the sibling tools, and no explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v0.4.0- First observed
analyze - First observed
code_check - First observed
index_update - First observed
project_query - First observed
run_tool
TDQS
Scored across 5 tools
project_query, analyze, and code_check all answer questions about code and overlap meaningfully (analyze subsumes existence-checking that code_check specializes). run_tool further muddies boundaries by bundling node_inspect, relation_trace, and change_review, which duplicate project_query and index_update concerns. Descriptions help, but an agent must reason carefully to pick the right entry point.
Names are all snake_case and readable, but the convention is mixed: project_query, index_update, and code_check use noun_verb, analyze is a bare verb, and run_tool uses verb_noun. There is no single predictable pattern, though it is far from chaotic.
Five top-level tools is a reasonable, well-scoped surface for a code-intelligence server. However, run_tool smuggles roughly twenty distinct operations (node_inspect, map_view, history_search, etc.) into a single tool, artificially compressing what is effectively a much larger toolset.
The surface covers querying, indexing, evidence-backed analysis, structure/impact/history inspection, and test tracing — solid lifecycle coverage for a read/analysis server. Gaps are minor, e.g. no direct workaround for edit application or non-listed language import checking beyond the not_checked fallback.
Maintenance
Related MCP Connectors
Search indexed code, trace dependencies, assess change impact, and recall repository memory.
Code intelligence platform for AI agents. 20 tools for architecture, security & impact analysis.
Code intelligence for coding agents: semantic, AST, graph, and full-text search. 279+ languages.
Codebase intelligence for AI agents — dead code, blast radius, ownership.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceProvides semantic code search and code insights via a knowledge graph, enabling AI to understand, navigate, and modify complex projects with deep dependency and architecture analysis.MIT
- FlicenseNot gradedqualityAmaintenanceEnables developer agents to perform semantic codebase search, dependency and impact analysis, cross-file refactoring, and full-stack API tracing through a unified query DSL over a high-performance graph engine.-
- AlicenseNot gradedqualityBmaintenanceEnables local code intelligence for AI agents and editors by indexing source code, dependencies, Git history, tests, and documentation, and provides change-impact analysis and relevant test selection through MCP.1AGPL 3.0
- AlicenseNot gradedqualityAmaintenanceProvides local-first repository architecture analysis with file-level proof, enabling agents to map codebases, locate implementations, trace call paths, and assess change impact with deterministic evidence.4 npm2MIT