kb-assistant
Indexes a GitHub repository's Markdown documentation and source code at a pinned commit, enabling permission-aware retrieval with citations and GitHub permalinks.
Provides a Slack bot interface for asking questions against the indexed GitHub repository, returning cited answers while enforcing access controls.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@kb-assistanthow do I make httpx ignore HTTP_PROXY?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
kb-assistant
A permission-aware RAG assistant for your GitHub docs and code. It works as an MCP server and a Slack bot, and every claim about it is measured.
🇹🇷 Özet: Şirket içi bilgi asistanı. GitHub reposundaki dokümanları ve kodu indeksler. Hibrit arama (BM25 + vektör) ve reranking ile ilgili parçaları bulur. Kaynak gösteren, uydurmayan cevaplar üretir. Bunu MCP sunucusu ve Slack botu olarak sunar. Yetki kontrolü, PII redaksiyonu ve audit log içerir. 49 soruluk bir eval seti ile retrieval doğruluğunu, halüsinasyon oranını, gecikmeyi ve maliyeti ölçer. Türkçe soruları da İngilizce dokümanlarda bulabilir.
At a glance
🎯 Retrieval | 93% Hit@5, 0.87 MRR on 45 hand-written questions; Hit@1 rose from 0.56 to 0.82 with reranking |
🇹🇷 Turkish → English | Turkish questions over English docs: Hit@5 rose from 0.17 to 0.83 after an embedding + reranker ablation |
🛡️ Abstention | 4/4 questions the corpus can't answer were declined, with nothing invented |
🔒 Security | ACLs enforced inside retrieval, secrets and PII redacted before indexing, a JSONL audit log of who saw what |
💸 Cost | $0 per query with a local LLM (Ollama). Claude is one setting away |
🔌 Interfaces | MCP server · Slack bot · REST API · CLI. All of them run through one code path |
✅ CI | Every PR rebuilds the index and fails if retrieval quality drops |
The demo corpus is encode/httpx pinned at b5addb64:
52 files of Markdown docs and Python source, standing in for an internal repo. To index
your own repo, edit config/settings.yaml.
Related MCP server: doc-mcp
How it works
flowchart LR
subgraph Ingest["Ingest (nightly job)"]
GH[GitHub repo<br/>pinned commit] --> RED[Redact secrets & PII]
RED --> CH[Structure-aware chunking<br/>md headings · Python AST]
CH --> IDX[(BM25 index<br/>+ embeddings)]
end
subgraph Serve["Answer a question"]
U1[Slack] --> SVC
U2[MCP client] --> SVC
U3[REST API] --> SVC
SVC[KnowledgeService] --> ACL{ACL mask<br/>by user groups}
ACL --> B[BM25]
ACL --> V[Vector]
B --> RRF[RRF fusion]
V --> RRF
RRF --> RR[Cross-encoder<br/>rerank]
RR --> LLM[LLM<br/>Ollama or Claude]
LLM --> ANS[Answer + citations<br/>+ GitHub permalinks]
end
IDX -.-> B
IDX -.-> V
SVC --> AUD[(Audit log)]Example: a Turkish question answered from English docs, by the local 3B model, at $0 (from the eval run):
Q: HTTP/2 desteğini nasıl açarım?
A: HTTP/2 desteğini açmak için,
httpxclient'inizde HTTP/2 desteği etkinleştirmeniz gerekmektedir. Bu,pip install httpx[http2]komutunu kullanarakhttpxpaketini güncellemeniz ve ardındanhttp2=Trueparametresiyle birAsyncClientveyaClientnesnesi oluşturmanız gerekmektedir. […]
Quickstart
python -m venv .venv && source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -e ".[dev]"
python -m kbassist.cli index # clone at pinned commit, chunk, embed (~3 min, CPU)
python -m kbassist.cli search "how do I make httpx ignore HTTP_PROXY"
ollama pull qwen2.5:3b # free local LLM (default provider)
python -m kbassist.cli ask "httpx'te varsayılan timeout nedir?"To use Claude instead, set ANTHROPIC_API_KEY and KB_LLM__PROVIDER=anthropic.
Results
Every number below comes from eval/run_eval.py and is written into this README by
python eval/run_eval.py report. Nothing is typed in by hand.
The eval set
eval/questions.yaml holds 49 questions, written by hand
against the pinned commit and phrased the way people actually ask in Slack:
Slice | n | What it tests |
Docs | 27 | How-to questions answered by |
Code | 12 | Answers that only exist in source, e.g. redirect method rewriting or the digest-auth |
Turkish | 6 | Turkish questions over an English corpus |
Unanswerable | 4 | Company questions (on-call policy, SLOs) that the bot must decline |
Each question has gold file paths, used for retrieval scoring, and reference facts, used for answer grading. One question has a false premise: it asks how to set the TTL of a cache that httpx doesn't have.
1 · Retrieval: does the right file reach the LLM?
Embedding sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2 · reranker jinaai/jina-reranker-v2-base-multilingual · 45 answerable questions · 776 chunks
Retrieval mode | Hit@1 | Hit@3 | Hit@5 | Recall@5 | MRR@10 | p50 latency | p95 latency |
BM25 only | 0.556 | 0.733 | 0.822 | 0.752 | 0.668 | 0.1 ms | 0.2 ms |
Vector only | 0.511 | 0.644 | 0.689 | 0.630 | 0.589 | 5.8 ms | 7.3 ms |
Hybrid (BM25 + vector, RRF) | 0.556 | 0.800 | 0.844 | 0.778 | 0.672 | 5.8 ms | 8.1 ms |
Hybrid + cross-encoder rerank | 0.822 | 0.911 | 0.933 | 0.852 | 0.869 | 2,791 ms | 3,109 ms |
Hit@k: at least one gold file is in the top-k chunks. Recall@5: the fraction of gold files found in the top 5. Latency is measured warm on a laptop CPU with no GPU.
Slice | n | BM25 only Hit@5 | Vector only Hit@5 | Hybrid (BM25 + vector, RRF) Hit@5 | Hybrid + cross-encoder rerank Hit@5 |
category: code | 13 | 0.923 | 0.615 | 1.000 | 1.000 |
category: doc | 32 | 0.781 | 0.719 | 0.781 | 0.906 |
lang: en | 39 | 0.872 | 0.718 | 0.872 | 0.949 |
lang: tr | 6 | 0.500 | 0.500 | 0.667 | 0.833 |
2 · How the models were chosen (ablation)
Embedding model | Reranker | Vector Hit@5 | Hybrid Hit@5 | Hybrid+rerank Hit@5 | Hybrid+rerank MRR@10 | Turkish Hit@5 (hybrid / +rerank) | Hybrid+rerank p50 |
|
| 0.733 | 0.889 | 0.844 | 0.718 | 0.500 / 0.167 | 2,774 ms |
|
| – | 0.889 | 0.911 | 0.798 | 0.500 / 0.500 | 9,911 ms |
|
| 0.689 | 0.844 | 0.978 | 0.885 | 0.667 / 1.000 | 10,186 ms |
|
| – | – | 0.956 | 0.847 | – / 1.000 | 3,639 ms |
chosen |
| 0.689 | 0.844 | 0.933 | 0.869 | 0.667 / 0.833 | 2,791 ms |
What I learned, in the order I ran the experiments:
The popular reranker made results worse.
ms-marco-MiniLMdropped Hit@5 from 0.889 to 0.844 and Turkish from 0.50 to 0.17. It was trained on English web passages, not on code or Turkish. A reranker has to be measured on your own corpus before you trust it.A code-aware multilingual reranker fixed the ordering (MRR 0.72 → 0.80). Turkish stayed at 0.50, because the right chunks never reached the candidate pool.
A multilingual embedding model is worse on its own but better in the pipeline. Vector Hit@5 fell from 0.73 to 0.69, but Turkish candidates now reached the reranker: overall Hit@5 rose to 0.978, and Turkish to 1.00. Picking each component on its own would have chosen the wrong embedding model.
Latency tuning. That setup needed about 10 s p50 on CPU. Reranking 10 candidates instead of 20, truncated to 1,000 chars, brought it to 2.8 s p50 (p95 14 s → 3.1 s) for a loss of 0.016 MRR. That is the shipped default.
Caveats: there are 45 answerable questions and only 6 are Turkish, so one question moves that column by 0.17. Gold labels are file-level and strict. Some "misses" retrieved the implementing source file where the label expected the docs page, and I didn't relabel them after seeing the results.
3 · Generation: accuracy, hallucination, latency, cost
Generator qwen2.5:3b (local via Ollama, CPU only), judge qwen2.5:3b, n = 49; second-pass labels: eval/results/audit.yaml (claude-opus-5-5 (dev session, not human))
Quality metric | Local 3B judge (automatic) | Second-pass review |
Answer accuracy (correct = 1, partial = 0.5) | 80.0% | 63.3% |
Strictly correct answers | 77.8% | 55.6% |
Hallucination rate (answers with a claim not supported by the sources) | 16.2% | 16.2% |
Declines unanswerable questions (higher is better) | 100.0% | 100.0% |
False "not in the sources" on answerable questions (lower is better) | 17.8% | 11.1% |
Judge agrees with review: correctness / hallucination flag | 69.4% / 83.8% |
Operational metric | Value |
Answers citing a gold file | 89.2% |
End-to-end latency p50 / p95 | 37,443 ms / 51,306 ms |
of which retrieval p50 | 2,755 ms |
of which LLM p50 / p95 | 34,679 ms / 48,342 ms |
Mean tokens in / out | 1,666 / 85 |
Cost per query / per 1k queries | $0.0000 / $0.00 |
What I learned:
The judge needs evaluating too. A 3B judge agreed with a careful second review on only 69% of verdicts, and it was lenient. It accepted a claim that the Authorization header survives cross-domain redirects (the code strips it). Accuracy fell from 80% to 63% once those cases were caught. In production I'd use a stronger judge and keep a human-labelled sample.
The bottleneck is the generator, not retrieval. In 17 of 20 imperfect answers the right file was in the context. The 3B model misread it, or said "not in the sources". A stronger model is one setting away (
KB_LLM__PROVIDER=anthropic).Abstention works: all 4 unanswerable company questions were declined, and nothing was invented for them.
The eval caught a product bug. In JSON mode the small model sometimes writes "here's an example:" and then ends without the code block. The fix is to generate prose first and extract citations separately.
37 s p50 is fine for an offline demo, not for Slack. Interactive use needs a GPU or a hosted model.
Hallucination means a claim that the retrieved sources don't support. A judge splits each answer into atomic claims and checks every one against the exact sources the generator saw (
eval/judge.py).Accuracy is graded against the hand-written reference facts.
Abstention is measured both ways: the bot should decline unanswerable questions, and should not decline answerable ones.
Cost comes from the model's reported token usage × the price table in
config/settings.yaml. Retrieval runs locally, so the LLM is the only per-query cost.Second-pass review: every answer was also labelled against the references and the source code (
eval/results/audit.yaml) to measure how far the automatic judge can be trusted. A stronger model wrote those labels during development, not a human, and the file says so.
Design decisions
Decision | Why | Trade-off |
Hybrid BM25 + dense, fused with RRF | BM25 nails exact identifiers ( | Two retrievers to maintain |
Cross-encoder rerank (top 10, 1,000 chars) | The biggest single gain: Hit@1 0.56 → 0.82 | About 2.8 s on CPU. Tuned in the ablation |
Code-aware tokenizer | Indexes | Slightly larger index |
Structure-aware chunking | Markdown is split by heading, with a breadcrumb ( | Language-specific code |
Local ONNX models (fastembed) | Source code never leaves the network to be embedded, and indexing is free | Reranking on CPU is slow |
numpy, not a vector DB | About 800 chunks: exact search takes about 1 ms with perfect recall | Swap to Qdrant/pgvector at scale; it's one function |
ACLs inside retrieval | Restricted chunks are masked before ranking, so they can't reach the prompt, citations or MCP tools | A mask per group set |
Redact before indexing | Tokens, keys, cards, TCKN, IBAN, emails and phone numbers are never embedded or stored. Checksums keep false positives low | Redacted text isn't searchable, which is intended |
Structured output |
| – |
Untrusted-source framing | Retrieved text goes inside | – |
Pluggable LLM | Ollama runs offline at $0. Claude gives higher quality. The same schema works on both | A local 3B model is weaker, and the eval measures how much |
Security & governance
Permission-aware retrieval.
config/acl.yamlmaps path globs and users to groups. In production, groups come from the IdP and are keyed on the Slack user's verified email. A test checks that a restricted chunk can't be retrieved even when it's the best lexical match.Audit log. Every search, answer, document read, denied access and 👍/👎 becomes one JSONL record: the user, their groups, the chunks shown and cited, latency per stage, tokens and cost. Questions are stored redacted. Answers are stored as a hash.
Secrets. Tokens come only from env vars or k8s Secrets.
GITHUB_TOKENis sent as an HTTP header and never stored in git config. Containers run as non-root, and egress is limited by a NetworkPolicy.
Interfaces
Mention @kb in a channel or DM it. It replies in a thread with the answer, links to the
cited sources (GitHub permalinks to the exact commit and lines), and 👍/👎 buttons that
feed an online quality signal into the audit log. It uses Socket Mode, so no public
ingress is needed. The app manifest is in
deploy/slack-manifest.yaml.
python -m kbassist.slack_app # needs SLACK_BOT_TOKEN and SLACK_APP_TOKENTool | Purpose |
| Hybrid + rerank search; returns chunks with permalinks |
| Reads more context around a hit, with ACL checks |
| One-shot grounded answer with citations |
| Indexed repos, pinned commits, index stats |
claude mcp add kb -- python -m kbassist.mcp_server # use it from Claude Code
python -m kbassist.mcp_server --http 8765 # or serve it over streamable HTTP
python scripts/mcp_smoke_test.py # end-to-end check with a real MCP clientuvicorn kbassist.api:app --port 8080 # /search /ask /documents/{path} /healthz /readyz
docker compose --profile jobs run --rm indexer && docker compose up api slack
kubectl apply -f deploy/k8s/ # nightly index CronJob + api/slack DeploymentsRunning the evals
pytest -q # 17 unit tests, no network or API key
python eval/run_eval.py retrieval # retrieval metrics, free
python eval/run_eval.py generate # answers + LLM judge (Ollama by default)
python eval/run_eval.py report # refresh the tables in this README
python eval/run_eval.py gate --min-hit5 0.85 --min-mrr 0.65 # what CI runsCI (.github/workflows/ci.yml) runs unit tests, then an
index build at the pinned commit, then the retrieval eval, then a quality gate that
fails the PR if Hit@5 or MRR regress. The LLM eval runs only on manual dispatch.
Project layout
src/kbassist/
├── ingest/ github.py (pinned clone) · chunking.py (md headings / Python AST)
├── index/ bm25.py · embeddings.py (fastembed) · store.py
├── security/ acl.py · pii.py · audit.py
├── retrieval.py BM25 + vector + RRF + rerank, ACL-masked
├── generation.py grounded answers, citations, refusal handling
├── llm.py OllamaLLM (local) · AnthropicLLM (Claude)
├── service.py the one code path behind CLI / MCP / Slack / API
└── mcp_server.py · slack_app.py · api.py · cli.py · observability.py
eval/ questions.yaml · run_eval.py · judge.py · results/
deploy/ k8s manifests · Slack app manifestRoadmap
Stronger generator and judge. Run the same eval with Claude and put the local and hosted results side by side
Query translation for BM25. Turkish BM25 Hit@5 is still 0.50, and an English rewrite of non-English queries should lift it
Reranker on a GPU to bring back the 20-candidate pool (Hit@5 0.978)
Incremental indexing from GitHub push webhooks (chunk ids are already content hashes)
Online eval loop: feed 👎 answers from the audit log back into the eval set
OpenTelemetry export of the existing trace spans
License
MIT
Available Tools
4 toolsask_knowledge_baseARead-onlyIdempotent
Answer a question end-to-end with citations, grounded only in indexed sources.
Returns answerable: false instead of guessing when the sources don't cover it.
| Name | Required | Description | Default |
|---|---|---|---|
| question | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint, idempotentHint, destructiveHint=false), so the bar is lower. The description adds genuine behavioral context beyond them: it discloses the refusal semantics ('returns answerable: false instead of guessing') and that output is citation-grounded, which an agent cannot learn from the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core capability and followed immediately by the notable edge-case behavior. No filler or restated information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description partially covers the return shape by naming citations and the answerable flag, which is the most important thing an agent needs to branch on. It does not describe citation format or answer structure, a minor gap for a single-parameter QA tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single 'question' parameter has no description, so the description carries the burden but adds nothing about it. The parameter is self-evident by name and type, which keeps this at baseline rather than below it, but no format or scope guidance (e.g. natural-language length) is provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Answer a question end-to-end with citations, grounded only in indexed sources'), which clearly distinguishes it from the snippet-returning search_knowledge_base sibling. It stops short of naming or explicitly contrasting the alternative, so the differentiation is implied rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'end-to-end with citations' framing implies you use this when you want a synthesized answer rather than raw search results, and the answerable:false behavior implies a fallback path. However, no explicit when-to-use/when-not guidance or named alternative is given, so usage must be inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_sourcesARead-onlyIdempotent
List indexed repositories, pinned commits and index statistics.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description contributes by enumerating what the listing contains, which matters given there is no output schema, but it says nothing about pagination, auth, or result size limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the verb first and the three returned categories following. Nothing is redundant and nothing needs trimming.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully enumerates the returned entities, and annotations cover the safety profile. It is adequate for a zero-parameter list tool, with only minor gaps around result volume or ordering.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is no parameter semantics to document and the baseline of 4 applies. The description correctly does not invent filtering options that the schema does not support.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and three concrete resources (indexed repositories, pinned commits, index statistics). The purpose is unmistakable, though it does not explicitly differentiate itself from the sibling search/read/ask tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: an agent can infer this is the discovery/inventory call to run before search_knowledge_base or read_file. There is no explicit when-to-use statement, no prerequisite, and no named alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_fileARead-onlyIdempotent
Read an indexed file (or a line range of it) by repository-relative path.
Use after search_knowledge_base to see surrounding context. Only files in the index that you have permission to read are available.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| end_line | No | ||
| start_line | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the safety profile is covered. The description adds non-obvious behavioral context: reads are restricted to files present in the index and to those the caller may read, which is an authorization/availability constraint not expressible in the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler; the core action and its scope are front-loaded, and the usage hint follows. Every clause contributes new information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need no explanation, and the description covers purpose, sequencing, and access constraints. It still leaves edge-case behavior (path not in index, permission denied, out-of-range lines) unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden: it clarifies that 'path' is repository-relative and that start_line/end_line select a range. However, it never specifies whether the range is inclusive, how end_line interacts with the default start_line=1, or what happens if the range exceeds the file.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (read an indexed file) plus the scope options (whole file or line range) and the path convention (repository-relative). This clearly separates it from search_knowledge_base and ask_knowledge_base, which retrieve knowledge rather than file content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent: 'Use after search_knowledge_base to see surrounding context,' naming the alternative and the sequencing condition. It adds a real availability constraint (only indexed and permission-granted files), though it gives no explicit when-not guidance or fallback behavior when a path is unavailable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_knowledge_baseARead-onlyIdempotent
Hybrid search (BM25 + embeddings + reranker) over indexed repository docs and code.
Use for any question about how our code or documentation works. Returns the top-k chunks with path, line range, a GitHub permalink and the chunk text. Write the query the way the answer would be phrased, and include exact identifiers (function, class, config key names) when you know them.
| Name | Required | Description | Default |
|---|---|---|---|
| k | No | ||
| query | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/idempotentHint/destructiveHint=false, so safety is covered. The description adds genuine value beyond that: the retrieval pipeline (BM25 + embeddings + reranker) and the exact shape of results (path, line range, GitHub permalink, chunk text). Auth/rate-limit behavior is not mentioned, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the retrieval mechanism, then usage, then parameter guidance. No filler; each sentence adds distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be spelled out, yet the description usefully summarizes them. The one gap is the unresolved relationship with ask_knowledge_base, leaving the agent to infer which retrieval tool fits a given question.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry the load, and it does: "top-k chunks" explains the purpose of k, and "Write the query the way the answer would be phrased, and include exact identifiers" is concrete guidance for the required query parameter. It stops short of documenting k's default/range explicitly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (hybrid BM25/embeddings/reranker search over indexed repository docs and code) and names the mechanism, so an agent can tell it apart from read_file and list_sources without opening a schema. The scope (repository docs and code) is explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use for any question about how our code or documentation works" gives clear context, and the query-phrasing tip tells the agent how to drive it. It does not, however, resolve when to prefer this over the sibling ask_knowledge_base, and gives no exclusions. Clear context without alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
ask_knowledge_base - First observed
list_sources - First observed
read_file - First observed
search_knowledge_base
TDQS
Scored across 4 tools
read_file and list_sources are clearly distinct from the query tools. search_knowledge_base and ask_knowledge_base both accept question-style queries about the same corpus, which introduces mild overlap, but their descriptions clearly separate raw-chunk retrieval from a synthesized cited answer.
All four tools use snake_case and follow a consistent verb_noun pattern: list_sources, search_knowledge_base, read_file, ask_knowledge_base. No stylistic deviations.
Four tools is a tight, well-scoped set for a read-only knowledge base assistant: enumerate sources, search, read context, and ask. Each tool earns its place without redundancy or bloat.
The read-only RAG lifecycle is well covered (list sources, search, fetch file context, answer with citations), and ask_knowledge_base handles the 'no answer' case gracefully. Minor gaps like listing files within a repo or fetch-by-permalink are workable around via search/read_file.
Maintenance
Related MCP Connectors
Ask any GitHub repository a question. Get source-backed answers.
The Cortex MCP server provides read-only access to real-time engineering context from the Cortex developer portal, allowing AI coding assistants to answer natural language questions about your organization's catalog (microservices, libraries, domains, teams, infrastructure), scorecards (engineering standards and best practices), initiatives (goals and deadlines), and Engineering Intelligence metrics. It includes tools for querying documentation, tracking personal entities, and accessing AI-assisted insights across the entire Cortex ecosystem.
An MCP server that gives your AI access to the source code and docs of all public github repos
Team docs served to AI agents over MCP - search, Markdown reads, version pinning, read audit.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceEnables semantic code search across indexed codebases using natural language queries, with support for CLI and MCP interfaces.1MIT
- FlicenseNot gradedqualityDmaintenanceEnables semantic search and AI-powered Q&A over ingested GitHub documentation repositories via MCP tools.-
- AlicenseAqualityAmaintenanceEnables AI assistants to search, analyze, and understand multi-language codebases by providing indexed code intelligence via MCP.161,409 npm8MIT
- FlicenseAqualityCmaintenanceEnables natural-language Q&A over codebases via MCP, using AST-aware chunking, hybrid retrieval, reranking, and call-graph expansion to answer with file:line citations and impact analysis.5-