Skip to main content
Glama

openground

PyPI version

tldr: openground lets you give controlled access to documentation to AI agents. Everything happens on-device.

openground is an on-device RAG system that extracts documentation from git repos and sitemaps, embeds it for semantic search, and exposes it to AI agents via MCP. It uses a local embedding model, and local lancedb for storing embeddings and for hybrid vector similarity and BM25 full-text search.

Architecture

      ┌─────────────────────────────────────────────────────────────────────┐
      │                           OPENGROUND                                │
      ├─────────────────────────────────────────────────────────────────────┤
      │                                                                     │
      │       SOURCE                  PROCESS              STORAGE/CLIENT   │
      │                                                                     │
      │    ┌──────────┐      ┌───────────┐   ┌──────────┐   ┌──────────┐    │
      │    │ git repo ├─────>│  Extract  ├──>│  Chunk   ├──>│ LanceDB  │    │
      │    |   -or-   |      │ (raw_data)│   │   Text   │   │ (vector  │    │
      │    │ sitemap  │      └───────────┘   └──────────┘   │  +BM25)  │    │
      │    │   -or-   │                           │         └────┬─────┘    │
      │    │ local dir│                           │              │          │
      │    └──────────┘                           │              │          │
      │                                           ▼              │          │
      │                                    ┌───────────┐         │          │
      │                                    │   Local   |<────────┘          │
      │                                    │ Embedding │         │          │
      │                                    │   Model   │         ▼          │
      │                                    └───────────┘  ┌─────────────┐   │
      │                                                   │ CLI / MCP   │   │
      │                                                   │  (hybrid    │   │
      |                                                   |   search)   |   |
      │                                                   └─────────────┘   │
      │                                                                     │
      └─────────────────────────────────────────────────────────────────────┘

Related MCP server: tech-doc-mcp

Quick Start

Installation

Recommended to install with uv:

uv tool install openground # Larger package size, automatic GPU/MPS/CPU support
uv tool install 'openground[fastembed]' # Lightweight CPU support
uv tool install 'openground[fastembed-gpu]' # Experimental CUDA/GPU support through fastembed

or

pip install openground

Add Documentation

Openground can source documentation from git repos, sitemaps, or local directories.

To add documentation from a git repo to openground, run:

openground add library-name \
  --source https://github.com/example/example.git \
  --docs-path docs/ \
  --version v1.0.0 \ # gets v1.0.0 docs using git tags
  -y

The --version flag specifies a git tag to checkout (defaults to latest).

To add documentation from a sitemap to openground:

openground add library-name \
  --source https://docs.example.com/sitemap.xml \
  --filter-keyword docs \ 
  --filter-keyword blog \
  -y

To add documentation from a local path to openground:

# Absolute path
openground add library-name --source /path/to/docs -y

# Home directory
openground add library-name --source ~/path/to/docs -y

# Relative path (from current directory)
openground add library-name --source ./docs -y
openground add library-name --source ../docs -y
openground add library-name --source docs -y

Git and local directory additions support .md, .rst, .txt, .mdx, .ipynb, .html, and .htm files.

This will download the docs, embed them, and store them into lancedb. All locally.

Multiple versions of the same library can be stored and queried independently.

Sources Files

Openground uses sources.json files to store library source configurations. When you add documentation with --source, openground remembers the source URL so you can add/update the same library later by just specifying its name.

How Sources Files Work

There are two types of sources files:

  1. User Sources File (~/.openground/sources.json)

    • Shared across all your projects

    • Created automatically when you first use --source flag

    • This is where new sources are saved by default

  2. Project Sources File (.openground/sources.json)

    • Project-specific overrides

    • Created automatically in each project when you add a source

    • Takes priority over user sources when both exist

Priority Order

When looking up a library by name, openground checks:

  1. Custom path via --sources-file flag

  2. Project-local .openground/sources.json (if exists)

  3. User ~/.openground/sources.json (if exists)

  4. Package-level bundled sources

Example Workflow

# In project1: Add library with source
cd project1/
openground add fastapi --source https://github.com/tiangolo/fastapi.git --docs-path docs/

# In project2: Same library is now available by name!
cd ../project2/
openground add fastapi  # Finds source from ~/.openground/sources.json

# Project-specific override: Add a different version for this project
echo '{"fastapi": {"type": "git_repo", "repo_url": "https://github.com/tiangolo/fastapi.git", "docs_paths": ["docs"], "languages": ["python"]}}' > .openground/sources.json

Managing Sources

Sources are stored as JSON with this structure:

{
  "fastapi": {
    "type": "git_repo",
    "repo_url": "https://github.com/tiangolo/fastapi",
    "docs_paths": ["docs"],
  },
  "numpy": {
    "type": "sitemap",
    "sitemap_url": "https://numpy.org/doc/sitemap.xml",
    "filter_keywords": ["docs/"]
  }
}

To disable automatic source saving:

openground config set sources.auto_add_local false

Use with AI Agents

To install the MCP server:

# For Cursor
openground install-mcp --cursor

# For Claude Code
openground install-mcp --claude-code

# For OpenCode
openground install-mcp --opencode

# For any other agent
openground install-mcp

Now your AI assistant can search your stored documentation automatically!

Example Workflow

Here's how to add the fastembed documentation and make it available to Claude Code:

# 1. Install openground
uv tool install openground

# 2. Add fastembed to openground
openground add fastembed --source https://github.com/qdrant/fastembed.git --docs-path docs/ --version v0.7.4 -y

# 3. Configure Claude Code to use openground MCP
openground install-mcp --claude-code

# 4. Restart Claude Code
# Now you can ask: "What models are available in fastembed?"
# Claude will search the fastembed docs automatically!

Claude Code Agent

Openground includes a custom Claude Code agent that searches official documentation without polluting your main conversation context. See docs/claude-code-agent.md for installation and usage instructions.

MCP Usage Statistics

To see how many times each tool in the MCP server has been called:

openground stats show # show stats
openground stats clear # reset stats

Development

To contribute or work on openground locally:

git clone https://github.com/poweroutlet2/openground.git
cd openground
uv sync .

License

MIT

Available Tools

3 tools
get_full_content_toolA

Retrieve the full content of a document by its URL and version.

Use this tool when you need to see the complete content of a page that was returned in search results. The URL and version are provided in the search result's tool hint.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
versionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full behavioral burden. It implies a read-only retrieval but does not disclose authentication needs, error behavior, rate limits, content limitations, or any side effects. The only added behavioral context is that arguments come from the search result hint, which is more parameter guidance than behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The primary action is front-loaded, and the usage guidance and parameter sourcing are each given their own concise sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity retrieval tool with an output schema and two self-descriptive parameters, the description covers what the tool does, when to use it, and where to get the inputs. The main missing pieces are behavioral caveats, but the presence of an output schema covers return-value structure.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description names both parameters and explains where to obtain their values ('provided in the search result's tool hint'). The parameter names are self-explanatory, but the description does not clarify version format, constraints, or edge cases beyond the names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Retrieve the full content of a document by its URL and version.' It clearly distinguishes this from sibling tools by stating it is for pages returned in search results, not for searching or listing libraries.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit trigger condition: 'Use this tool when you need to see the complete content of a page that was returned in search results.' It also points the agent to the search result's tool hint for the URL and version. It does not explicitly name alternatives, but the context makes the intended placement clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_libraries_toolA

Retrieve a dictionary of available documentation libraries/frameworks with their versions.

Returns a dictionary mapping library names to lists of available versions. Use this tool to see what documentation is available before performing a search. If the desired library is not in the list, you may prompt the user to add it.

Args: search_term: Optional search term to filter library names (case-insensitive). If provided, only libraries whose names contain the search term will be returned.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the return shape (dictionary mapping library names to version lists), the case-insensitive filter behavior, and the fallback action of prompting the user if a library is missing. It could mention side effects or limits, but this appears to be a safe read/list operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured and front-loaded with purpose, then return format, usage, and fallback. The Args section is concise and clear, though it would be stronger if it matched the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple listing tool, the description covers purpose, behavior, and fallback well. However, the mismatch between the documented search_term argument and the empty input schema creates a real ambiguity that prevents the agent from confidently invoking the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description documents an optional search_term parameter with meaningful semantics, but the input schema declares zero properties. This is a direct contradiction: an agent cannot tell whether passing search_term is valid. The 0-param baseline is 4, but the misleading documentation lowers the score substantially.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool's purpose: retrieving available documentation libraries/frameworks and their versions. It uses a specific verb and resource, and it distinguishes itself from sibling tools by framing this as a pre-search discovery step.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says to use this tool before performing a search, which gives clear usage context. It does not name sibling alternatives or state when not to use it, but the 'before a search' positioning is enough to guide selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_documents_toolA

Search the official documentation knowledge base to answer user questions.

Always used this tool when a question might be answered or confirmed from the documentation.

First call list_libraries_tool to see what libraries and versions are available, then filter by library_name and version.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes
versionYes
library_nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It states the core behavior—searching documentation and filtering by library and version—which implies a read-only operation. However, it does not disclose details such as matching behavior, result limits, error conditions, or whether only snippets are returned. Adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise at three sentences and front-loads the core purpose. Every sentence adds information about when or how to use the tool. The minor grammar issue in 'Always used this tool' slightly detracts from polish, but the structure is effective.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that an output schema exists, return details need not be in the description. The description covers the tool's purpose, usage trigger, and required pre-step via list_libraries_tool. It could be more complete by mentioning get_full_content_tool as the logical follow-up for retrieving full documents, but nothing essential to invoking search_documents_tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does clarify library_name and version by directing the agent to list_libraries_tool to see valid values and to filter by them. However, it gives no guidance on what constitutes a valid query or how the query is matched, leaving one of the three parameters semantically under-specified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Search the official documentation knowledge base.' This clearly distinguishes it from the sibling tools, which list libraries and retrieve full content. It also frames the tool's purpose as answering user questions, making its role unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says to use this tool when a question might be answered or confirmed from documentation, and it instructs the agent to call list_libraries_tool first to discover available libraries and versions. It does not mention when to prefer get_full_content_tool or when not to use the tool, so it falls just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.14.0
    • First observedget_full_content_tool
    • First observedlist_libraries_tool
    • First observedsearch_documents_tool

TDQS

A4.2/5.0

Scored across 3 tools

Disambiguation5/5

Each tool serves a distinct, non-overlapping role: listing available documentation, searching within it, and retrieving full content. There is no ambiguity about which tool to use for a given step in the workflow.

Naming Consistency5/5

All tools follow a consistent snake_case verb_noun_tool naming pattern (search_documents_tool, list_libraries_tool, get_full_content_tool). This is highly predictable and aids agent selection.

Tool Count5/5

With only 3 tools, the server is tightly scoped to its documentation-searching purpose. Each tool is essential and the count is well within the ideal range for a focused server.

Completeness5/5

The tool set covers the full lifecycle of documentation interaction: discover available libraries, search for relevant documents, and fetch full content when needed. No obvious gaps or dead ends exist for the intended use case.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    Provides up-to-date documentation for AI agents by locally querying a community-driven registry of pre-built docs packages.
    Apache 2.0
  • A
    license
    A
    quality
    A
    maintenance
    Provides fast, token-efficient search over coding agent documentation (e.g., Claude Code, Cursor) using local SQLite FTS5 indexing, with tools for searching snippets, reading pages, and grepping markdown.
    5
    MIT
  • A
    license
    B
    quality
    B
    maintenance
    Enables local agents to search and retrieve cited evidence from PDFs and Markdown notes, including page-specific passages and rendered page images.
    6
    GPL 3.0