Skip to main content
Glama

v3.0.5 · Tests in TypeScript and JavaScript count: calls inside describe/it/test callbacks now belong to the test, so arbor callers, diff and check see Jest, Vitest and Mocha tests. MCP tool output uses project-relative paths. Release notes


Why Arbor

Most AI coding tools treat code as text. Arbor builds a semantic dependency graph — functions, classes, and modules as nodes; calls, imports, and inheritance as edges — then answers execution-aware questions with deterministic precision:

Question

Arbor answer

If I change this symbol, what breaks?

Blast radius with depth, confidence, and risk level

Who calls this — directly and transitively?

Caller/callee traversal on the call graph

What's the shortest path between A and B?

A* path through real dependencies

Is this PR too risky to merge?

CI gate on blast-radius thresholds

No keyword guessing. No embedding hallucinations. One graph, every interface.

Where the graph is unsure, it says so — edges carry a confidence, and ambiguous resolutions are labelled rather than hidden. An honest unknown beats a confident wrong answer.


Related MCP server: CodeGraph CLI MCP Server

What's new in v3.0.5

A fix release: tests written in TypeScript and JavaScript now count.

Fix

What was wrong

Jest, Vitest and Mocha tests are callers

Calls inside describe/it/test callbacks belonged to no symbol, so arbor callers missed every test and diff/check reported that nothing tested a change. Each test and hook is now a function named after it, such as it: accepts two, that owns the calls in its callback. On the repro in #235, arbor callers decide went from 2 callers to 4.

Project-relative MCP paths

Tool output could include absolute paths from your machine. Every file path is now relative to the project, and the namespaced io.modelcontextprotocol/protocolVersion request key is read first.

Cached graphs rebuild once on first use (extract-4). See the 3.0.4 release notes for the previous release.

A fix release. Every item below was a wrong answer, not a missing feature, and each ships with a regression test.

Fix

What was wrong

Rust calls resolve

crate::jobs::enqueue(), Type::new(), self.helper() and calls inside assert!/format! produced no edge, so heavily used Rust functions reported no callers. Paths now resolve through the module tree, and macro arguments are read as code.

Per-symbol arbor diff

Every symbol in a touched file counted as changed, so adding a function to a busy file reported that file's whole blast radius. New symbols now carry none; only modified ones do. Tests that call the change are listed as worth running, not counted as impact.

--base and --staged

arbor diff, check and summary saw only uncommitted edits unless ARBOR_DIFF_BASE was set. --base origin/main compares against the merge base, which is what a pull request shows; --staged checks just what is staged.

Languages stay apart

A TypeScript enqueue could be reported as a caller of a Rust enqueue. Resolution stays within a language family, and callers/callees list each same-named definition separately.

No stale graphs

After a branch switch, answers came from the old branch's graph when no remaining file was newer than it. The graph now records the commit it was built from and refreshes when HEAD moves.

Inheritance edges

class Middle(Base) produced no edge, so changing a base class showed zero blast radius. extends/implements edges are emitted and inherited methods stay reachable.

Call cycles rank as one

A closed ring of functions filled the top of every centrality ranking. Cycles are condensed before PageRank.

HTTP bridge hardening

arbor bridge --http accepted cross-origin and DNS-rebinding requests. It now checks Origin/Host and requires JSON bodies.

Quieter bridge

The bridge re-indexed ignored build directories in a busy loop.

Also new: arbor receipt. After each coding-agent turn it explains, in plain English, what changed and what was touched that you didn't ask for, with arbor receipt undo to put a turn back.

Cached graphs from 3.0.0 are rebuilt automatically on first use.

One fix, measured.

Symbol resolution consults the importing file. When a bare name matched definitions in several modules, resolve_ref fell through to SameDir and attached the edge to whichever definition sat in the caller's own directory — not a dropped edge, a confidently misrouted one, stamped at 0.55 confidence.

GraphBuilder already kept a per-file import map, but only apply_import_validation read it, and that scores an edge after one has been chosen. It never saw the references going to the wrong node. Consulting it between the same-file and same-directory checks keeps a local definition shadowing an import, while letting a written import beat mere adjacency. Resolution::ViaImport scores 0.93, above SameDir's 0.55.

Measured

A fixture of 260 modules across 10 layers, each layer defining the same 26 function names. Ground truth is derived from the generator's own edge list, so the expected answer is exact rather than estimated.

True downstream

v2.6.0

v3.0.0

179

0

163

178

0

161

161

0

133

143

22

133

122

22

119

36

22

61

16

22

46

Previously flat at about 22 regardless of the real answer. Now it tracks. Risk on the largest hub moves from LOW to CRITICAL.

Total edge count barely moves (1335 → 1334). That is the signature of misrouting rather than loss: the edges were always there, pointing at the wrong nodes.

Breaking

  • Resolution gains a ViaImport variant — an exhaustive match will not compile

  • Edges land on different nodes, so cached graphs, stored node ids, and centrality baselines from 2.6.0 will differ

Call cycles are condensed before PageRank. Each strongly connected component is ranked once and that mass is shared across its members, so a closed ring does not fill the top of the ranking and a cycle that calls out keeps its members together. The number CentralityScores reports is still the v2.6.0 percentile, i / (n - 1).

Known and still open

Written down rather than left to be discovered:

  • Small targets now over-report (36 → 61, 16 → 46). Safer direction than silence, but not yet correct.

  • Dynamic and reflective imports (importlib, __import__, import(), eval(require(...))) are unresolvable by construction and are documented as expected misses in the fixture rather than counted as defects.

Correctness, not speed. Each of these was silently wrong before.

Fix

Why it mattered

Colliding symbols are kept

SymbolTable used HashMap::insert, so a second handler, new, or process replaced the first. The loser had zero callers and was invisible to blast radius.

Resolution is deterministic

Same-directory locality was decided by iterating a HashMap. Rust seeds RandomState per process, so the same binary on the same input could build different edges between runs. Now asserted across eight fresh processes.

Edges carry confidence

A proven same-file call and a same-directory guess were identical evidence. Each edge now scores [0,1] by how it resolved.

Exported TS symbols indexed once

export_statement recursed into its children, then the generic loop recursed again — every exported symbol became two vertices sharing one node id. 133 phantom nodes on a 149-file app, 25% of the graph.

Method calls on untyped receivers resolve

obj.method() was dropped outright, leaving the graph nearly edgeless on TS/JS — and an empty graph reports a blast radius of zero, which reads as "safe" rather than "unknown".

Centrality is a percentile rank

Scores were divided by the graph maximum, so the top node was 1.0 by construction and a 0.6 threshold meant nothing consistent between repos. Adding one hub rescaled every other node.

Resolution is O(1), not O(refs × nodes × files)

Unresolvable references — stdlib and third-party calls, most call sites in real code — paid the worst case. Suffixes are now indexed.

New capability — concept search. Substring matching cannot find get_authenticated from login; they share no substring. Identifiers are now tokenized and expanded through curated concept clusters, and docstrings, signatures, and paths are indexed alongside names. Deterministic, offline, no model. Available on the library as ArborGraph::search_ranked (arbor query remains literal-substring for now).

New capability — hunk-level impact. changed_node_ids_for_ranges keeps only symbols whose lines actually changed, instead of every symbol in a touched file.

Measured on identical node sets, after the duplicate-extraction fix:

Codebase

Before

After

TypeScript (149 files)

172 edges

196 (+14%)

Rust (arbor-graph)

116 edges

167 (+44%)

Graph caches from earlier versions are invalidated — centrality now means something different, so a stale cache would be read wrong.

Change

Measured

PageRank rewrite — flat call-graph adjacency replaces per-iteration traversal

149.8ms → 6.6ms on a 10k-node graph (23x), verified side-by-side vs the old implementation

Parallel indexing — parse fans out across all cores, deterministic assembly

Arbor: 253ms → 95ms · tokio (178k LOC): 2.7s → 1.6s

Warm-start centrality — watcher recomputes seed from previous scores

Converges in ~2 rounds after a one-file patch instead of the full 20-iteration budget

Convergence early-exit

Iteration stops at 1e-9 max delta — the budget is a ceiling, not a sentence

Think a number is wrong? cargo bench -p arbor-graph and prove it: BENCHMARKS.md.

Feature

What it does

MCP 2026-07-28

Stateless server/discover, response caching (ttlMs/cacheScope), dual-version fallback for 2025-03-26 clients

Tasks extension

tasks/get · tasks/update · tasks/cancel — cold-start indexing returns task handles, not errors

MCP Apps

Interactive blast-radius graph (ui://arbor/blast-radius) and architecture map (ui://arbor/architecture-map) inside agent hosts

HTTP transport

arbor bridge --http --port 3333 — stateless MCP behind load balancers

Real get_blast_radius

Git-diff-aware impact analysis via shared arbor-graph::compute_blast_radius

Pagination

offset / limit / hasMore on search_symbols and get_map

Benchmarks

Criterion suite + CI regression gate — see BENCHMARKS.md


Quickstart

# Install (crates.io can lag the latest release: see docs/INSTALL.md)
cargo install arbor-graph-cli

# Index your project (one command)
cd your-project && arbor setup

# Explore before you edit
arbor map . --exclude-test          # ranked project skeleton (~1k tokens)
arbor refactor parse_file           # blast radius of changing a symbol
arbor diff                          # impact of uncommitted git changes
arbor diff --base origin/main       # impact of this branch, as its PR shows it

# Wire up your AI agent
claude mcp add --transport stdio --scope project arbor -- arbor bridge

After every turn: arbor hook claude makes Claude Code show a receipt of what it changed, what it touched that you didn't ask for, and what to test. Receipts →

Agent workflow: call get_map first → search_symbols / get_file_graph to locate code → Read only the target file. Full MCP guide →


For AI agents (MCP)

Arbor ships a production MCP server via arbor bridge. Stdio is the default; HTTP is opt-in for remote/enterprise.

# Stdio (Claude, Cursor, VS Code)
arbor bridge

# HTTP (MCP 2026-07-28)
arbor bridge --http --port 3333

Cursor / VS Code

{
  "mcpServers": {
    "arbor": {
      "type": "stdio",
      "command": "arbor",
      "args": ["bridge"]
    }
  }
}

Templates: templates/mcp/ · Setup scripts: scripts/setup-mcp.sh · scripts/setup-mcp.ps1

16 MCP tools

Tier

Tools

Use when

Orientation

get_map

First call — token-budgeted project skeleton ranked by PageRank

Surgical

list_entry_points · get_callers · get_callees · search_symbols · get_file_graph · get_node_detail

Navigate to a specific symbol or file

Broad

get_logic_path · analyze_impact · find_path · get_knowledge_path

Trace dependencies, blast radius, paths

Agent-native

get_blast_radius · explain_symbol · audit_security · get_architecture_overview · batch_query

PR impact, onboarding, security audit, bulk lookup

Every tool returns { ok, tool, data, meta: { suggested_next_tool, suggested_next_args } } so agents chain calls without re-prompting.

Registry: io.github.Anandb71/arbor · Official API lookup · Glama listing


CLI reference

Command

Description

arbor setup

One-shot init + index

arbor map

Ranked, token-budgeted project skeleton

arbor query <term>

Fuzzy symbol search (supports | OR)

arbor callers / callees <sym>

One-hop graph traversal. Same-named symbols are listed per definition; narrow with jobs::enqueue, Type.method or src/jobs.rs:enqueue

arbor entry-points

HTTP handlers, main, jobs, webhooks

arbor file-graph <path>

Symbols + edges in one file

arbor inspect <sym>

Full symbol detail

arbor path <a> <b>

Shortest call-graph path

arbor refactor <sym>

Blast radius before refactoring

arbor diff

Git-change impact report, per symbol: new symbols carry no blast radius. --base <ref> compares against the merge base (what a PR shows), --staged only staged changes

arbor check

CI safety gate (--max-blast-radius N, --base <ref>)

arbor summary

Auto-generate PR description (--base <ref>)

arbor agent review

Autonomous PR architecture review

arbor agent onboard

Codebase onboarding guide

arbor agent guard

Real-time architectural safety gate

arbor bridge

MCP server (add --http for HTTP transport)

arbor watch

Live re-index on file changes

arbor receipt list / show / undo

Plain-English receipts of what each agent turn changed, and undo for a turn (guide)

arbor hook claude

Wire Arbor into Claude Code: directives, receipts after every turn

arbor gui

Native desktop UI

All query commands support --json. map additionally supports --tokens N, --focus "pattern", --focus-changed.


Visual tour


Installation

# macOS / Linux: prebuilt binary from the latest GitHub release
curl -fsSL https://raw.githubusercontent.com/Anandb71/arbor/main/scripts/install.sh | bash

# Windows (PowerShell)
irm https://raw.githubusercontent.com/Anandb71/arbor/main/scripts/install.ps1 | iex

# Scoop (Windows)
scoop install https://raw.githubusercontent.com/Anandb71/arbor/main/packaging/scoop/arbor.json

# Rust / Cargo (crates.io can lag the latest release)
cargo install arbor-graph-cli

# npm wrapper (cross-platform)
npx @anandb71/arbor-cli

# Docker
docker pull ghcr.io/anandb71/arbor:latest

Pinned installs: docs/INSTALL.md


Language support

Production parsers: Rust · TypeScript / JavaScript · Python · Go · Java · C / C++ · C# · Dart

Fallback parsers: Kotlin · Swift · Ruby · PHP · Shell

Adding languages →


CI & pull requests

arbor diff --markdown
arbor check --max-blast-radius 30 --markdown
arbor summary

GitHub Action (pre-built binary, ~5s vs ~3–5min compile):

name: Arbor Check
on: [pull_request]

jobs:
  arbor:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
        with:
          fetch-depth: 0

      - uses: Anandb71/arbor@v3.0.5
        with:
          command: check . --max-blast-radius 30 --markdown
          comment-on-pr: true
          github-token: ${{ secrets.GITHUB_TOKEN }}

Architecture

arbor-core (Tree-sitter parsing)
    └── arbor-graph (petgraph + PageRank + impact analysis)
            ├── arbor-cli      — CLI + MCP bridge
            ├── arbor-mcp      — MCP protocol server
            ├── arbor-server   — WebSocket JSON-RPC
            ├── arbor-watcher  — incremental file watcher
            └── arbor-gui      — desktop UI

Docs: Quickstart · Architecture · Graph schema · MCP integration · Receipts · Benchmarks · Roadmap · Philosophy

Release channels: GitHub Releases · crates.io · GHCR · npm · VS Code / Open VSX · Scoop — Releasing guide


Philosophy

  1. Consumer first — beautiful, intuitive, instantly useful

  2. Accessibility second — works across ecosystems, runs anywhere

  3. Affordability next — minimal overhead, from laptops to monoliths

Arbor is local-first: no mandatory data exfiltration, offline-capable, open source. Security policy →


Contributing

cargo build --workspace
cargo test --workspace
cargo clippy --workspace --all-targets --all-features

CONTRIBUTING.md · Good first issues · Code of conduct


Contributors


License

MIT — see LICENSE.

Available Tools

2 tools
analyze_impactC

Analyzes the impact (blast radius) of changing a specific node.

ParametersJSON Schema
NameRequiredDescriptionDefault
node_idYesID or name of the node to analyze

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions analyzing impact but doesn't specify what the analysis entails (e.g., computational cost, side effects, permissions required, or output format). This leaves significant gaps in understanding the tool's behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that directly states the tool's purpose without any fluff or redundancy. It is appropriately sized and front-loaded, making it easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of impact analysis, no annotations, and no output schema, the description is incomplete. It doesn't explain what 'blast radius' entails, the nature of the analysis, or what results to expect, leaving the agent with insufficient context for effective use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, with 'node_id' documented as 'ID or name of the node to analyze'. The description adds no additional parameter semantics beyond this, so it meets the baseline of 3 for high schema coverage without extra value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Analyzes the impact (blast radius) of changing a specific node.' It specifies the verb ('analyzes') and resource ('impact of changing a specific node'), though it doesn't explicitly differentiate from the sibling tool 'get_logic_path' (which might retrieve paths rather than analyze impact).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives, such as the sibling tool 'get_logic_path'. It lacks context on prerequisites, scenarios where this analysis is needed, or any exclusions, leaving the agent without usage direction.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_logic_pathC

Traces the call graph to find dependencies and usage of a function or class.

ParametersJSON Schema
NameRequiredDescriptionDefault
start_nodeYesName of the function or class to trace

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions tracing and finding dependencies/usage, which suggests a read-only analysis operation, but doesn't clarify if it's safe, has side effects, requires permissions, or details output format (e.g., graph structure, depth limits). This leaves significant gaps for a tool with potential complexity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that directly states the tool's purpose without unnecessary words. It's appropriately sized and front-loaded, making it easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and no output schema, the description is incomplete for a tool that traces call graphs. It lacks details on behavioral traits (e.g., safety, performance), output format, or how it differs from siblings, making it inadequate for an agent to fully understand usage without additional context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, clearly documenting the single parameter 'start_node'. The description adds context by specifying it traces 'dependencies and usage of a function or class', which aligns with the schema but doesn't provide additional syntax or format details beyond what's already covered.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with specific verbs ('traces', 'find') and resources ('call graph', 'dependencies and usage', 'function or class'), making it easy to understand what it does. However, it doesn't explicitly differentiate from its sibling tool 'analyze_impact', which might have overlapping or related functionality.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus its sibling 'analyze_impact' or any alternatives. It implies usage for tracing dependencies and usage, but lacks explicit context, prerequisites, or exclusions, leaving the agent to infer appropriate scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv1.0.0
    • First observedanalyze_impact
    • First observedget_logic_path

TDQS

B3.1/5.0

Scored across 2 tools

Disambiguation5/5

The two tools have clearly distinct purposes: analyze_impact focuses on assessing the blast radius of a node change, while get_logic_path traces dependencies and usage in a call graph. There is no overlap or ambiguity between these functions.

Naming Consistency5/5

Both tools follow a consistent verb_noun pattern (analyze_impact, get_logic_path) with clear, descriptive names. The naming style is uniform and predictable across the set.

Tool Count2/5

With only two tools, the server feels thin for a domain like code or system analysis, where more operations (e.g., for managing nodes, viewing graphs, or updating logic) might be expected. This limited set could hinder agent workflows.

Completeness2/5

Inferring a domain of code or system dependency analysis, the toolset is severely incomplete. It lacks basic CRUD operations (e.g., create, update, delete nodes) and essential functions like listing or searching dependencies, leaving significant gaps for agent tasks.

Maintenance

ActivityActive
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    An intelligent server that provides semantic code search, domain-driven analysis, and advanced code understanding for large codebases using LLMs and vector embeddings.
    10
    MIT
  • F
    license
    Not graded
    quality
    C
    maintenance
    A high-performance CLI tool that provides semantic code search, advanced architectural analysis, and codebase indexing with vector embeddings across multiple programming languages. Enables AI assistants to understand and navigate large codebases through graph-based relationships and intelligent code pattern detection.
    888
    -
  • A
    license
    A
    quality
    A
    maintenance
    A local-first codebase intelligence tool that enables AI assistants to research codebases using semantic search, multi-hop relationship discovery, and structural parsing. It allows users to extract architectural patterns and institutional knowledge across 30+ programming languages through an MCP-compatible interface.
    2
    1,468 PyPI
    1,446
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    A minimalist indexing tool that provides AI agents with semantic search and structural AST parsing for deep codebase understanding. It enables autonomous agents to navigate large codebases predictably using vector embeddings and native language server capabilities like definition and reference tracking.
    -