Skip to main content
Glama

OntoPrune

Neuro-Symbolic Context Pruning Middleware for Local SLMs and Cloud LLMs

Tests Python License: MIT MCP Server Languages

English | Español | 📄 Technical Whitepaper (PDF) | 📄 Whitepaper en Español (PDF)


OntoPrune is an ultra-lightweight (<12ms CPU) neuro-symbolic middleware that transforms multi-file source code into minimal typed dependency contracts. By isolating closed-world functional boundaries before attention computation, OntoPrune slashes input tokens by ~60% in modular enterprise architectures, collapses Time-to-First-Token ($TTFT$) latency by 62% on CPU-bound local models (Ollama), guarantees 100% functional test success (Pass@1), and prevents 100% of proprietary code leakage.


📊 Rigorous Empirical Benchmark: 24 Independent Sandbox Runs

Tested across 3 real-world software archetypes in isolated execution sandboxes with live evaluation via pytest:

  • Cloud Frontier: Google Gemini (gemini-3.8-flash via native SSE streaming)

  • Local SLM: Ollama (qwen2.5-coder:7b running on AMD Ryzen CPU)

Software Archetype

Backend

Treatment

Input Tokens (Median)

Token Reduction

TTFT (Local/Cloud)

Pass@1 (pytest)

Leaked IP (Private Lines)

Archetype 1: Isolated Algorithm

gemini

OntoPrune

818

-24.3%

0.1 ms

100.0% (5/5)

0 lines

Archetype 1: Isolated Algorithm

gemini

Naive Full

1,081

Baseline

0.1 ms

100.0% (5/5)

0 lines

Archetype 1: Isolated Algorithm

ollama

OntoPrune

731

-23.5%

155 ms

100.0% (5/5)

0 lines

Archetype 1: Isolated Algorithm

ollama

Naive Full

956

Baseline

12,985 ms

100.0% (5/5)

0 lines

Archetype 2: Multi-Module Service

gemini

OntoPrune

931

-59.7%

0.1 ms

100.0% (8/8)

0 lines (100% shielded)

Archetype 2: Multi-Module Service

gemini

Naive Full

2,313

Baseline

0.1 ms

100.0% (8/8)

74 private lines exposed

Archetype 2: Multi-Module Service

ollama

OntoPrune

797

-59.7%

10,569 ms

100.0% (8/8)

0 lines (100% shielded)

Archetype 2: Multi-Module Service

ollama

Naive Full

1,979

Baseline

28,032 ms

100.0% (8/8)

74 private lines exposed

Archetype 3: Clean Architecture

gemini

OntoPrune

1,147

-60.1%

0.1 ms

100.0% (10/10)

0 lines (100% shielded)

Archetype 3: Clean Architecture

gemini

Naive Full

2,873

Baseline

0.1 ms

100.0% (10/10)

91 private lines exposed

Archetype 3: Clean Architecture

ollama

OntoPrune

952

-60.5%

13,092 ms

100.0% (10/10)

0 lines (100% shielded)

Archetype 3: Clean Architecture

ollama

Naive Full

2,412

Baseline

34,376 ms

100.0% (10/10)

91 private lines exposed

Key Scientific Findings:

  1. Zero Semantic Degradation (100.0% Pass@1): In all 24 sandbox runs, code generated with OntoPrune stubs passed 100% of unit tests, proving that contract interfaces and ontology metadata provide all the context an LLM needs.

  2. 62% TTFT Speedup on Local CPU: In multi-module and clean architecture projects, time-to-first-token dropped from ~34s to ~13s on consumer CPU.

  3. 100% Intellectual Property Shield: Naive tools (Cursor/Copilot) send entire method bodies (up to 91 private algorithmic lines). OntoPrune sends exactly 0 lines of internal dependencies.

  4. Architectural Dynamics: In monolithic single scripts, compression is moderate (~24%). In structured multi-module systems, compression is consistent at ~60% net token reduction.


Related MCP server: mcp-sophon

💡 Why OntoPrune Makes Local AI-Assisted Programming Truly Feasible

Running local code intelligence models (Qwen 2.5 Coder 7B, Llama 3 8B, DeepSeek Coder via Ollama or llama.cpp) on consumer laptops and developer workstations has historically faced an insurmountable barrier. OntoPrune dismantles this barrier through three architectural pillars:

1. Collapsing the "Pre-fill Latency Wall" on CPU ($TTFT$)

  • The Bottleneck: On CPUs and consumer laptops without high-end 24GB GPUs, token decoding speed is acceptable (15–25 t/s). However, the Pre-fill phase (Prompt Evaluation) is heavily memory-bandwidth bound and saturates RAM.

  • Without OntoPrune: Standard AI assistants inject entire raw files and bloated multi-module contexts (2,500 – 3,000 tokens), causing a 35-second freeze before the first token appears. A half-minute freeze on every code turn shatters the developer's flow state.

  • With OntoPrune: By pruning raw context into concise typed contracts (~800 tokens), pre-fill latency collapses from 34.3s down to 13.0s (a 62% to 74% drop in TTFT) on standard CPUs (AMD Ryzen 5600G). Local inference shifts from unusable to completely interactive.

2. Eliminating Attention Dilution in Small Language Models (SLMs)

  • Small models (3B to 7B parameters) lack the massive context comprehension of hundred-billion-parameter cloud models. Dumping hundreds of lines of irrelevant database or third-party implementations dilutes attention, leading to hallucinated method calls.

  • OntoPrune constructs a mathematically sound closed-world contract. The SLM only sees allowable methods and domain types. In our 24 empirical sandbox runs, Qwen 2.5 Coder 7B achieved 100.0% Pass@1 on pytest suites.

3. Absolute Data Sovereignty and Air-Gapped Privacy

  • OntoPrune operates in $<12\text{ ms}$ on CPU using local AST visitors and in-memory RDF graphs with zero network calls. It provides an airtight guarantee: 0 lines of internal proprietary code are ever leaked to external APIs.


📦 Installation

# Standard installation (native Python support):
pip install ontoprune

# With multi-language support (Flutter/Dart, Java, TypeScript via Tree-sitter):
pip install "ontoprune[languages]"

# For development, benchmarks, and tests:
pip install "ontoprune[dev,benchmark,languages]"

🌐 Universal Multi-Language Support

OntoPrune automatically detects file types and resolves dependencies across project boundaries:

Language

Extension

AST Engine

Output Format

Python

.py

Native Python ast

def name(args) -> Ret: ...

Flutter / Dart

.dart

tree-sitter-dart

abstract class ... { Ret method(); }

Java / Spring Boot

.java

tree-sitter-java

public interface ... { Ret method(); }

TypeScript / React

.ts, .tsx, .js

tree-sitter-typescript

export interface ... { method(): Ret; }


🛠️ Usage Modes

1. Native Model Context Protocol (MCP) Server

OntoPrune runs out of the box as an MCP server (ontoprune-mcp) compatible with Claude Desktop, Cursor, Gemini CLI, or Antigravity IDE:

ontoprune-mcp

Configuration in claude_desktop_config.json:

{
  "mcpServers": {
    "ontoprune": {
      "command": "ontoprune-mcp"
    }
  }
}

Exposed MCP Tools:

  • prune_context(file_path, target_symbol, format='stubs'): Extracts the minimal dependency contract resolving cross-file imports.

  • verify_response(response_code, contract_or_file): Deterministically validates generated code against authorized contracts.


2. Command Line Interface (CLI)

# Prune a target method across multi-file projects:
ontoprune translate src/services/OrderService.java processOrder --format stubs

# Direct streaming pipeline with local Ollama:
ontoprune translate services/order_service.py procesar_orden | ollama run qwen2.5-coder:3b

# Deterministically verify LLM output against contract:
ontoprune check --file generated_solution.py --contract contract.py

3. Python API

import ontoprune

# 1. Prune a multi-module project to a minimal typed contract
context = ontoprune.translate(
    "services/order_service.py",
    target="procesar_orden",
    fmt="stubs",
    multi_module=True,
)
print(context)

# 2. Verify model output against the contract
violations = ontoprune.check(llm_code_response, against=context)
if not violations:
    print("Code is 100% compliant and free of hallucinations!")

🧪 Test Suite

pytest tests/
# 46 passed in 1.40s

🔬 Reproduce Empirical Benchmarks

You can independently replicate the benchmark evaluation in isolated sandboxes on your machine:

# 1. Quick deterministic validation in sandbox (no API key needed):
ontoprune benchmark --backend mock --analyze
# or via Make:
make benchmark-mock

# 2. Live evaluation against Google Gemini (requires GEMINI_API_KEY):
ontoprune benchmark --backend gemini --analyze
# or via Make:
make benchmark-gemini

# 3. Live local evaluation against Ollama CPU (requires running Ollama):
ontoprune benchmark --backend ollama --analyze
# or via Make:
make benchmark-ollama

📄 Whitepapers & Publications


👤 Author

Vigmar Carlo

Available Tools

2 tools
prune_contextA

Extracts a minimal, hallucination-resistant context contract for a given function or method. Reduces input tokens by >80% while preserving exact type signatures and dependencies across project modules.

Args: file_path: Absolute or relative path to the Python source file. target_symbol: Identifier of the function or method (e.g. 'procesar_orden' or 'OrderService.procesar_orden'). format: Desired output format: 'stubs' (default, optimal for code LLMs), 'turtle' (RDF), 'json', or 'nl'. include_body: Whether to include the raw source code body of the target function. project_root: Optional root directory of the project (auto-detected if omitted).

Returns: The compact pruned contract string.

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNostubs
file_pathYes
include_bodyNo
project_rootNo
target_symbolYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description bears the full burden. It discloses useful behavioral traits (token reduction, signature/dependency preservation, available output formats) but never states that this is a safe read-only extraction with no side effects, nor any performance or filesystem-access limits implied by project_root auto-detection.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded purpose followed by cleanly scoped Args/Returns sections; each line earns its place. Slightly padded by marketing-flavored wording ('hallucination-resistant', '>80%') that adds tone more than callable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return-value detail is unnecessary, and the parameter documentation is thorough for a 5-param, 0%-coverage schema. What's missing is the operational framing — when to use it and confirmation of read-only/no-side-effect behavior — which is meaningful for a tool with no annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate — and it does: every one of the five parameters is explained, including path semantics ('absolute or relative'), symbol syntax with concrete examples ('OrderService.procesar_orden'), the accepted format values ('stubs'/'turtle'/'json'/'nl' — not present as an enum in the schema), and project_root's auto-detection fallback.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Extracts a minimal... context contract for a given function or method,' and quantifies the effect ('>80%' token reduction). The term 'context contract' is jargon, but the following clauses (type signatures, dependencies across modules) make the intent concrete. It does not distinguish itself from the sibling verify_response, though that sibling is clearly unrelated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied — the tool is framed as an LLM-oriented context extractor, but there is no statement of when to call it versus reading the file directly, nor any preconditions or exclusions. The format/default guidance partially informs selection but is about output shape, not when-to-use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_responseA

Verifies LLM-generated code against an ontological software contract to detect hallucinated or unauthorized method/function calls.

Args: response_code: The Python code snippet or markdown text produced by the model. contract_or_file: Either the rendered contract string OR the path to the original Python file. target_symbol: If contract_or_file is a file path, specify the target symbol to extract its contract.

Returns: Dictionary with validation status, count of violations, and list of invalid calls.

ParametersJSON Schema
NameRequiredDescriptionDefault
response_codeYes
target_symbolNo
contract_or_fileYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses the return shape (status, violation count, invalid call list), which is useful, but says nothing about how violations are reported, whether the call can fail on malformed contracts, or any cost/limits. Adequate but shallow for a no-annotation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded one-sentence purpose followed by clean Args/Returns sections; every line carries information. Minor redundancy: the Returns block restates what the output schema already provides.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values needn't be explained, yet the description adds them anyway. Combined with thorough parameter explanation, an agent has enough to invoke the tool correctly; only failure modes and violation-reporting detail are absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it largely does: it explains response_code as a code snippet or markdown, contract_or_file as either a rendered contract string OR a file path, and target_symbol as conditionally required only when a file path is given. That conditional relationship is genuine added meaning the schema does not convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb+resource: verifies LLM-generated code against a software contract to detect hallucinated/unauthorized calls. The scope ('hallucinated or unauthorized method/function calls') is specific enough that an agent knows exactly what output to expect, and the unrelated sibling prune_context creates no ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the description (run it on model-generated code against a contract), but there is no explicit when-to-use, when-not-to-use, or prerequisite guidance. Nothing tells the agent whether this should run before or after execution, or what to do when a contract is unavailable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.1.0
    • First observedprune_context
    • First observedverify_response

TDQS

A4.1/5.0

Scored across 2 tools

Disambiguation5/5

The two tools have clearly distinct purposes: prune_context extracts a compact contract from source code, while verify_response validates LLM-generated code against such a contract. Their descriptions and argument sets make the boundary unambiguous.

Naming Consistency5/5

Both names follow a consistent verb_noun pattern in snake_case: prune_context and verify_response. There are no mixed conventions or vague verbs.

Tool Count4/5

Two tools is slightly below the typical 3-15 sweet spot, but each tool is substantial, non-redundant, and directly supports a focused extract-and-verify workflow. The set does not feel bloated or padded.

Completeness4/5

The pair covers the core lifecycle of extracting a contract and verifying code against it, with flexible input options for files or rendered contracts. Minor gaps exist for symbol discovery, batch pruning, or contract management, but agents can work around them.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Model-agnostic context management middleware for LLM-powered workflows, providing persistent selective memory across sub-conversations using RAG over embedded summaries.
    MIT
  • F
    license
    A
    quality
    A
    maintenance
    Deterministic repository context packing for AI coding agents: selects, compresses, and budgets only the files a task needs. Measured 83% fewer input tokens at the same task coverage, fully local, no LLM in the loop.
    9
    9
    -