Skip to main content
Glama

kb-assistant

A permission-aware RAG assistant for your GitHub docs and code. It works as an MCP server and a Slack bot, and every claim about it is measured.

Python MCP Slack LLM Tests License

🇹🇷 Özet: Şirket içi bilgi asistanı. GitHub reposundaki dokümanları ve kodu indeksler. Hibrit arama (BM25 + vektör) ve reranking ile ilgili parçaları bulur. Kaynak gösteren, uydurmayan cevaplar üretir. Bunu MCP sunucusu ve Slack botu olarak sunar. Yetki kontrolü, PII redaksiyonu ve audit log içerir. 49 soruluk bir eval seti ile retrieval doğruluğunu, halüsinasyon oranını, gecikmeyi ve maliyeti ölçer. Türkçe soruları da İngilizce dokümanlarda bulabilir.


At a glance

🎯 Retrieval

93% Hit@5, 0.87 MRR on 45 hand-written questions; Hit@1 rose from 0.56 to 0.82 with reranking

🇹🇷 Turkish → English

Turkish questions over English docs: Hit@5 rose from 0.17 to 0.83 after an embedding + reranker ablation

🛡️ Abstention

4/4 questions the corpus can't answer were declined, with nothing invented

🔒 Security

ACLs enforced inside retrieval, secrets and PII redacted before indexing, a JSONL audit log of who saw what

💸 Cost

$0 per query with a local LLM (Ollama). Claude is one setting away

🔌 Interfaces

MCP server · Slack bot · REST API · CLI. All of them run through one code path

✅ CI

Every PR rebuilds the index and fails if retrieval quality drops

The demo corpus is encode/httpx pinned at b5addb64: 52 files of Markdown docs and Python source, standing in for an internal repo. To index your own repo, edit config/settings.yaml.

Related MCP server: doc-mcp

How it works

flowchart LR
    subgraph Ingest["Ingest (nightly job)"]
        GH[GitHub repo<br/>pinned commit] --> RED[Redact secrets & PII]
        RED --> CH[Structure-aware chunking<br/>md headings · Python AST]
        CH --> IDX[(BM25 index<br/>+ embeddings)]
    end

    subgraph Serve["Answer a question"]
        U1[Slack] --> SVC
        U2[MCP client] --> SVC
        U3[REST API] --> SVC
        SVC[KnowledgeService] --> ACL{ACL mask<br/>by user groups}
        ACL --> B[BM25]
        ACL --> V[Vector]
        B --> RRF[RRF fusion]
        V --> RRF
        RRF --> RR[Cross-encoder<br/>rerank]
        RR --> LLM[LLM<br/>Ollama or Claude]
        LLM --> ANS[Answer + citations<br/>+ GitHub permalinks]
    end

    IDX -.-> B
    IDX -.-> V
    SVC --> AUD[(Audit log)]

Example: a Turkish question answered from English docs, by the local 3B model, at $0 (from the eval run):

Q: HTTP/2 desteğini nasıl açarım?

A: HTTP/2 desteğini açmak için, httpx client'inizde HTTP/2 desteği etkinleştirmeniz gerekmektedir. Bu, pip install httpx[http2] komutunu kullanarak httpx paketini güncellemeniz ve ardından http2=True parametresiyle bir AsyncClient veya Client nesnesi oluşturmanız gerekmektedir. […]

Quickstart

python -m venv .venv && source .venv/bin/activate     # Windows: .venv\Scripts\activate
pip install -e ".[dev]"
python -m kbassist.cli index                          # clone at pinned commit, chunk, embed (~3 min, CPU)
python -m kbassist.cli search "how do I make httpx ignore HTTP_PROXY"

ollama pull qwen2.5:3b                                # free local LLM (default provider)
python -m kbassist.cli ask "httpx'te varsayılan timeout nedir?"

To use Claude instead, set ANTHROPIC_API_KEY and KB_LLM__PROVIDER=anthropic.


Results

Every number below comes from eval/run_eval.py and is written into this README by python eval/run_eval.py report. Nothing is typed in by hand.

The eval set

eval/questions.yaml holds 49 questions, written by hand against the pinned commit and phrased the way people actually ask in Slack:

Slice

n

What it tests

Docs

27

How-to questions answered by docs/**

Code

12

Answers that only exist in source, e.g. redirect method rewriting or the digest-auth cnonce

Turkish

6

Turkish questions over an English corpus

Unanswerable

4

Company questions (on-call policy, SLOs) that the bot must decline

Each question has gold file paths, used for retrieval scoring, and reference facts, used for answer grading. One question has a false premise: it asks how to set the TTL of a cache that httpx doesn't have.

1 · Retrieval: does the right file reach the LLM?

Embedding sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2 · reranker jinaai/jina-reranker-v2-base-multilingual · 45 answerable questions · 776 chunks

Retrieval mode

Hit@1

Hit@3

Hit@5

Recall@5

MRR@10

p50 latency

p95 latency

BM25 only

0.556

0.733

0.822

0.752

0.668

0.1 ms

0.2 ms

Vector only

0.511

0.644

0.689

0.630

0.589

5.8 ms

7.3 ms

Hybrid (BM25 + vector, RRF)

0.556

0.800

0.844

0.778

0.672

5.8 ms

8.1 ms

Hybrid + cross-encoder rerank

0.822

0.911

0.933

0.852

0.869

2,791 ms

3,109 ms

Hit@k: at least one gold file is in the top-k chunks. Recall@5: the fraction of gold files found in the top 5. Latency is measured warm on a laptop CPU with no GPU.

Slice

n

BM25 only Hit@5

Vector only Hit@5

Hybrid (BM25 + vector, RRF) Hit@5

Hybrid + cross-encoder rerank Hit@5

category: code

13

0.923

0.615

1.000

1.000

category: doc

32

0.781

0.719

0.781

0.906

lang: en

39

0.872

0.718

0.872

0.949

lang: tr

6

0.500

0.500

0.667

0.833

2 · How the models were chosen (ablation)

Embedding model

Reranker

Vector Hit@5

Hybrid Hit@5

Hybrid+rerank Hit@5

Hybrid+rerank MRR@10

Turkish Hit@5 (hybrid / +rerank)

Hybrid+rerank p50

BAAI/bge-small-en-v1.5

Xenova/ms-marco-MiniLM-L-12-v2 (pool 20)

0.733

0.889

0.844

0.718

0.500 / 0.167

2,774 ms

BAAI/bge-small-en-v1.5

jinaai/jina-reranker-v2-base-multilingual (pool 20)

–

0.889

0.911

0.798

0.500 / 0.500

9,911 ms

sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2

jinaai/jina-reranker-v2-base-multilingual (pool 20)

0.689

0.844

0.978

0.885

0.667 / 1.000

10,186 ms

sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2

jinaai/jina-reranker-v2-base-multilingual (pool 20, 600 chars)

–

–

0.956

0.847

– / 1.000

3,639 ms

chosen sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2

jinaai/jina-reranker-v2-base-multilingual (pool 10, 1000 chars)

0.689

0.844

0.933

0.869

0.667 / 0.833

2,791 ms

What I learned, in the order I ran the experiments:

  1. The popular reranker made results worse. ms-marco-MiniLM dropped Hit@5 from 0.889 to 0.844 and Turkish from 0.50 to 0.17. It was trained on English web passages, not on code or Turkish. A reranker has to be measured on your own corpus before you trust it.

  2. A code-aware multilingual reranker fixed the ordering (MRR 0.72 → 0.80). Turkish stayed at 0.50, because the right chunks never reached the candidate pool.

  3. A multilingual embedding model is worse on its own but better in the pipeline. Vector Hit@5 fell from 0.73 to 0.69, but Turkish candidates now reached the reranker: overall Hit@5 rose to 0.978, and Turkish to 1.00. Picking each component on its own would have chosen the wrong embedding model.

  4. Latency tuning. That setup needed about 10 s p50 on CPU. Reranking 10 candidates instead of 20, truncated to 1,000 chars, brought it to 2.8 s p50 (p95 14 s → 3.1 s) for a loss of 0.016 MRR. That is the shipped default.

Caveats: there are 45 answerable questions and only 6 are Turkish, so one question moves that column by 0.17. Gold labels are file-level and strict. Some "misses" retrieved the implementing source file where the label expected the docs page, and I didn't relabel them after seeing the results.

3 · Generation: accuracy, hallucination, latency, cost

Generator qwen2.5:3b (local via Ollama, CPU only), judge qwen2.5:3b, n = 49; second-pass labels: eval/results/audit.yaml (claude-opus-5-5 (dev session, not human))

Quality metric

Local 3B judge (automatic)

Second-pass review

Answer accuracy (correct = 1, partial = 0.5)

80.0%

63.3%

Strictly correct answers

77.8%

55.6%

Hallucination rate (answers with a claim not supported by the sources)

16.2%

16.2%

Declines unanswerable questions (higher is better)

100.0%

100.0%

False "not in the sources" on answerable questions (lower is better)

17.8%

11.1%

Judge agrees with review: correctness / hallucination flag

69.4% / 83.8%

Operational metric

Value

Answers citing a gold file

89.2%

End-to-end latency p50 / p95

37,443 ms / 51,306 ms

of which retrieval p50

2,755 ms

of which LLM p50 / p95

34,679 ms / 48,342 ms

Mean tokens in / out

1,666 / 85

Cost per query / per 1k queries

$0.0000 / $0.00

What I learned:

  1. The judge needs evaluating too. A 3B judge agreed with a careful second review on only 69% of verdicts, and it was lenient. It accepted a claim that the Authorization header survives cross-domain redirects (the code strips it). Accuracy fell from 80% to 63% once those cases were caught. In production I'd use a stronger judge and keep a human-labelled sample.

  2. The bottleneck is the generator, not retrieval. In 17 of 20 imperfect answers the right file was in the context. The 3B model misread it, or said "not in the sources". A stronger model is one setting away (KB_LLM__PROVIDER=anthropic).

  3. Abstention works: all 4 unanswerable company questions were declined, and nothing was invented for them.

  4. The eval caught a product bug. In JSON mode the small model sometimes writes "here's an example:" and then ends without the code block. The fix is to generate prose first and extract citations separately.

  5. 37 s p50 is fine for an offline demo, not for Slack. Interactive use needs a GPU or a hosted model.

  • Hallucination means a claim that the retrieved sources don't support. A judge splits each answer into atomic claims and checks every one against the exact sources the generator saw (eval/judge.py).

  • Accuracy is graded against the hand-written reference facts.

  • Abstention is measured both ways: the bot should decline unanswerable questions, and should not decline answerable ones.

  • Cost comes from the model's reported token usage × the price table in config/settings.yaml. Retrieval runs locally, so the LLM is the only per-query cost.

  • Second-pass review: every answer was also labelled against the references and the source code (eval/results/audit.yaml) to measure how far the automatic judge can be trusted. A stronger model wrote those labels during development, not a human, and the file says so.


Design decisions

Decision

Why

Trade-off

Hybrid BM25 + dense, fused with RRF

BM25 nails exact identifiers (trust_env, DEFAULT_MAX_REDIRECTS). Dense retrieval catches paraphrases. RRF needs no score calibration

Two retrievers to maintain

Cross-encoder rerank (top 10, 1,000 chars)

The biggest single gain: Hit@1 0.56 → 0.82

About 2.8 s on CPU. Tuned in the ablation

Code-aware tokenizer

Indexes max_keepalive_connections whole and split into words

Slightly larger index

Structure-aware chunking

Markdown is split by heading, with a breadcrumb (Timeouts > Fine tuning). Python is split per function/class via ast

Language-specific code

Local ONNX models (fastembed)

Source code never leaves the network to be embedded, and indexing is free

Reranking on CPU is slow

numpy, not a vector DB

About 800 chunks: exact search takes about 1 ms with perfect recall

Swap to Qdrant/pgvector at scale; it's one function

ACLs inside retrieval

Restricted chunks are masked before ranking, so they can't reach the prompt, citations or MCP tools

A mask per group set

Redact before indexing

Tokens, keys, cards, TCKN, IBAN, emails and phone numbers are never embedded or stored. Checksums keep false positives low

Redacted text isn't searchable, which is intended

Structured output

answerable / answer / citations make abstention and citations machine-checkable

–

Untrusted-source framing

Retrieved text goes inside <source> tags and is treated as data. This defends against prompt injection hidden in repo content

–

Pluggable LLM

Ollama runs offline at $0. Claude gives higher quality. The same schema works on both

A local 3B model is weaker, and the eval measures how much

Security & governance

  • Permission-aware retrieval. config/acl.yaml maps path globs and users to groups. In production, groups come from the IdP and are keyed on the Slack user's verified email. A test checks that a restricted chunk can't be retrieved even when it's the best lexical match.

  • Audit log. Every search, answer, document read, denied access and 👍/👎 becomes one JSONL record: the user, their groups, the chunks shown and cited, latency per stage, tokens and cost. Questions are stored redacted. Answers are stored as a hash.

  • Secrets. Tokens come only from env vars or k8s Secrets. GITHUB_TOKEN is sent as an HTTP header and never stored in git config. Containers run as non-root, and egress is limited by a NetworkPolicy.

Interfaces

Mention @kb in a channel or DM it. It replies in a thread with the answer, links to the cited sources (GitHub permalinks to the exact commit and lines), and 👍/👎 buttons that feed an online quality signal into the audit log. It uses Socket Mode, so no public ingress is needed. The app manifest is in deploy/slack-manifest.yaml.

python -m kbassist.slack_app      # needs SLACK_BOT_TOKEN and SLACK_APP_TOKEN

Tool

Purpose

search_knowledge_base(query, k)

Hybrid + rerank search; returns chunks with permalinks

read_file(path, start_line, end_line)

Reads more context around a hit, with ACL checks

ask_knowledge_base(question)

One-shot grounded answer with citations

list_sources()

Indexed repos, pinned commits, index stats

claude mcp add kb -- python -m kbassist.mcp_server     # use it from Claude Code
python -m kbassist.mcp_server --http 8765              # or serve it over streamable HTTP
python scripts/mcp_smoke_test.py                       # end-to-end check with a real MCP client
uvicorn kbassist.api:app --port 8080       # /search  /ask  /documents/{path}  /healthz  /readyz
docker compose --profile jobs run --rm indexer && docker compose up api slack
kubectl apply -f deploy/k8s/               # nightly index CronJob + api/slack Deployments

Running the evals

pytest -q                                  # 17 unit tests, no network or API key
python eval/run_eval.py retrieval          # retrieval metrics, free
python eval/run_eval.py generate           # answers + LLM judge (Ollama by default)
python eval/run_eval.py report             # refresh the tables in this README
python eval/run_eval.py gate --min-hit5 0.85 --min-mrr 0.65    # what CI runs

CI (.github/workflows/ci.yml) runs unit tests, then an index build at the pinned commit, then the retrieval eval, then a quality gate that fails the PR if Hit@5 or MRR regress. The LLM eval runs only on manual dispatch.

Project layout

src/kbassist/
├── ingest/         github.py (pinned clone) · chunking.py (md headings / Python AST)
├── index/          bm25.py · embeddings.py (fastembed) · store.py
├── security/       acl.py · pii.py · audit.py
├── retrieval.py    BM25 + vector + RRF + rerank, ACL-masked
├── generation.py   grounded answers, citations, refusal handling
├── llm.py          OllamaLLM (local) · AnthropicLLM (Claude)
├── service.py      the one code path behind CLI / MCP / Slack / API
└── mcp_server.py · slack_app.py · api.py · cli.py · observability.py
eval/               questions.yaml · run_eval.py · judge.py · results/
deploy/             k8s manifests · Slack app manifest

Roadmap

  • Stronger generator and judge. Run the same eval with Claude and put the local and hosted results side by side

  • Query translation for BM25. Turkish BM25 Hit@5 is still 0.50, and an English rewrite of non-English queries should lift it

  • Reranker on a GPU to bring back the 20-candidate pool (Hit@5 0.978)

  • Incremental indexing from GitHub push webhooks (chunk ids are already content hashes)

  • Online eval loop: feed 👎 answers from the audit log back into the eval set

  • OpenTelemetry export of the existing trace spans

License

MIT

Available Tools

4 tools
ask_knowledge_baseA
Read-onlyIdempotent

Answer a question end-to-end with citations, grounded only in indexed sources.

Returns answerable: false instead of guessing when the sources don't cover it.

ParametersJSON Schema
NameRequiredDescriptionDefault
questionYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnlyHint, idempotentHint, destructiveHint=false), so the bar is lower. The description adds genuine behavioral context beyond them: it discloses the refusal semantics ('returns answerable: false instead of guessing') and that output is citation-grounded, which an agent cannot learn from the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core capability and followed immediately by the notable edge-case behavior. No filler or restated information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description partially covers the return shape by naming citations and the answerable flag, which is the most important thing an agent needs to branch on. It does not describe citation format or answer structure, a minor gap for a single-parameter QA tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single 'question' parameter has no description, so the description carries the burden but adds nothing about it. The parameter is self-evident by name and type, which keeps this at baseline rather than below it, but no format or scope guidance (e.g. natural-language length) is provided.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Answer a question end-to-end with citations, grounded only in indexed sources'), which clearly distinguishes it from the snippet-returning search_knowledge_base sibling. It stops short of naming or explicitly contrasting the alternative, so the differentiation is implied rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'end-to-end with citations' framing implies you use this when you want a synthesized answer rather than raw search results, and the answerable:false behavior implies a fallback path. However, no explicit when-to-use/when-not guidance or named alternative is given, so usage must be inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_sourcesA
Read-onlyIdempotent

List indexed repositories, pinned commits and index statistics.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description contributes by enumerating what the listing contains, which matters given there is no output schema, but it says nothing about pagination, auth, or result size limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with the verb first and the three returned categories following. Nothing is redundant and nothing needs trimming.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully enumerates the returned entities, and annotations cover the safety profile. It is adequate for a zero-parameter list tool, with only minor gaps around result volume or ordering.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is no parameter semantics to document and the baseline of 4 applies. The description correctly does not invent filtering options that the schema does not support.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and three concrete resources (indexed repositories, pinned commits, index statistics). The purpose is unmistakable, though it does not explicitly differentiate itself from the sibling search/read/ask tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: an agent can infer this is the discovery/inventory call to run before search_knowledge_base or read_file. There is no explicit when-to-use statement, no prerequisite, and no named alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_fileA
Read-onlyIdempotent

Read an indexed file (or a line range of it) by repository-relative path.

Use after search_knowledge_base to see surrounding context. Only files in the index that you have permission to read are available.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
end_lineNo
start_lineNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the safety profile is covered. The description adds non-obvious behavioral context: reads are restricted to files present in the index and to those the caller may read, which is an authorization/availability constraint not expressible in the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler; the core action and its scope are front-loaded, and the usage hint follows. Every clause contributes new information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need no explanation, and the description covers purpose, sequencing, and access constraints. It still leaves edge-case behavior (path not in index, permission denied, out-of-range lines) unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the burden: it clarifies that 'path' is repository-relative and that start_line/end_line select a range. However, it never specifies whether the range is inclusive, how end_line interacts with the default start_line=1, or what happens if the range exceeds the file.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (read an indexed file) plus the scope options (whole file or line range) and the path convention (repository-relative). This clearly separates it from search_knowledge_base and ask_knowledge_base, which retrieve knowledge rather than file content.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: 'Use after search_knowledge_base to see surrounding context,' naming the alternative and the sequencing condition. It adds a real availability constraint (only indexed and permission-granted files), though it gives no explicit when-not guidance or fallback behavior when a path is unavailable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_knowledge_baseA
Read-onlyIdempotent

Hybrid search (BM25 + embeddings + reranker) over indexed repository docs and code.

Use for any question about how our code or documentation works. Returns the top-k chunks with path, line range, a GitHub permalink and the chunk text. Write the query the way the answer would be phrased, and include exact identifiers (function, class, config key names) when you know them.

ParametersJSON Schema
NameRequiredDescriptionDefault
kNo
queryYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/idempotentHint/destructiveHint=false, so safety is covered. The description adds genuine value beyond that: the retrieval pipeline (BM25 + embeddings + reranker) and the exact shape of results (path, line range, GitHub permalink, chunk text). Auth/rate-limit behavior is not mentioned, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the retrieval mechanism, then usage, then parameter guidance. No filler; each sentence adds distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be spelled out, yet the description usefully summarizes them. The one gap is the unresolved relationship with ask_knowledge_base, leaving the agent to infer which retrieval tool fits a given question.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must carry the load, and it does: "top-k chunks" explains the purpose of k, and "Write the query the way the answer would be phrased, and include exact identifiers" is concrete guidance for the required query parameter. It stops short of documenting k's default/range explicitly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (hybrid BM25/embeddings/reranker search over indexed repository docs and code) and names the mechanism, so an agent can tell it apart from read_file and list_sources without opening a schema. The scope (repository docs and code) is explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Use for any question about how our code or documentation works" gives clear context, and the query-phrasing tip tells the agent how to drive it. It does not, however, resolve when to prefer this over the sibling ask_knowledge_base, and gives no exclusions. Clear context without alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedask_knowledge_base
    • First observedlist_sources
    • First observedread_file
    • First observedsearch_knowledge_base

TDQS

A4.1/5.0

Scored across 4 tools

Disambiguation4/5

read_file and list_sources are clearly distinct from the query tools. search_knowledge_base and ask_knowledge_base both accept question-style queries about the same corpus, which introduces mild overlap, but their descriptions clearly separate raw-chunk retrieval from a synthesized cited answer.

Naming Consistency5/5

All four tools use snake_case and follow a consistent verb_noun pattern: list_sources, search_knowledge_base, read_file, ask_knowledge_base. No stylistic deviations.

Tool Count5/5

Four tools is a tight, well-scoped set for a read-only knowledge base assistant: enumerate sources, search, read context, and ask. Each tool earns its place without redundancy or bloat.

Completeness4/5

The read-only RAG lifecycle is well covered (list sources, search, fetch file context, answer with citations), and ask_knowledge_base handles the 'no answer' case gracefully. Minor gaps like listing files within a repo or fetch-by-permalink are workable around via search/read_file.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables semantic code search across indexed codebases using natural language queries, with support for CLI and MCP interfaces.
    1
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables semantic search and AI-powered Q&A over ingested GitHub documentation repositories via MCP tools.
    -
  • F
    license
    A
    quality
    C
    maintenance
    Enables natural-language Q&A over codebases via MCP, using AST-aware chunking, hybrid retrieval, reranking, and call-graph expansion to answer with file:line citations and impact analysis.
    5
    -