sufler
Connects GitHub to the knowledge base and chat by ingesting issue, pull request, comment, review, CI, and issue-closure events, with two-way thread replies and the ability to open issues or answer in GitHub threads from Teams.
Provides read-only Jira access for listing the user's open and finished tasks, viewing an issue with comments, searching, and looking up a team member's tasks through a trusted identity map.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@suflersearch notes for the API design decision from last week"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Sufler
An MCP server and agent runtime built for one team: a shared, searchable knowledge base of project notes and status, reachable from Claude Code, Microsoft Teams, a command line and GitHub. Meeting notes and project status go into one store instead of scattered files, GitHub and Teams exchange events, and Jira answers "what are my open tasks" without leaving the conversation.
About this copy. I built Sufler during my internship at BIAP (Intelligent Technologies division, July to September 2026) and publish it with the company's consent. Personal data and company infrastructure (hosts, tenant, team and channel identifiers) were removed from every commit, and
#Nin commit messages refers to pull requests in the private original. The documentation underdocs/is in Polish.
Status
Deployed as a pilot in one team. The code of phases 1 to 4 is complete: the MCP server, the agent
runtime, the Teams, CLI and GitHub entry points, the GitHub ↔ event store ↔ Teams bridge with gated
writes, read-only Jira, the Teams Shifts schedule and lexical retrieval over notes. Open: the HTTP
deployment (docs/roadmap.md).
Related MCP server: Personal RAG MCP Server
Architecture
One core, several entry points. All logic lives in a core with no I/O (core/); Claude Code over
MCP, Teams, the CLI and GitHub are thin adapters over the same tool catalogue, so a new tool is
written once. The core never imports an adapter, and import-linter checks that in CI.
Quick start
Requires Python 3.11 and uv.
uv sync # core; the MCP server runs on local files with no secrets
uv run sufler # MCP server (stdio)
uv run mcp dev src/sufler/server.py # inspect the toolsDetails
Area | What it can do |
Knowledge base (MCP) | 4 read tools (search, read note, list projects, project status) and one gated write tool ( |
Agent on Teams, CLI and GitHub | The same tool catalogue drives an agent runtime (Claude in a loop) with conversation memory and history compaction. |
| Commands from the model run in a separate executor container with no network, created per conversation and able to see only that conversation's scratch space, behind a membership gate (ADR 0057, ADR 0063). |
|
|
GitHub ↔ event store ↔ Teams bridge | Ingests issue, pull request, comment, review, CI and issue-closure events; posts to a channel and to 1:1 chats; two-way threads; from Teams, opens an issue or answers in the GitHub thread. |
Jira (read-only) | My open and finished tasks, one issue with comments, search, and a team member's tasks through a trusted identity map. There is no write path (ADR 0054). |
Teams Shifts schedule (read) | Who works today or this week, on site or remote, and who is off (ADR 0059). |
Meeting notes from transcripts | The agent reads a Teams meeting transcript and writes a structured note; participants are counted from the transcript, not by the model. |
Lexical retrieval (BM25) | Search over Polish lemmas ( |
Both MCP tool counts, five frozen and three additive, are held by the golden test
tests/adapters/test_mcp_tool_surface.py. The agent has eight tools with the shell enabled and
fourteen without it; the full set, each tool's actions and a byte ceiling on tool descriptions are
held by tests/core/test_tool_descriptions.py.
Jira is read-only. It does not create issues, comment or change status; there is no write path in the code, and the Jira account never comes from the model.
Writes are off by default everywhere. Code for a mutating capability can be in the tree and still be unreachable until an operator opens its gate.
tests/test_gates_closed_by_default.pyfinds the gates by reflection and checks the configuration templates too.The HTTP entry point does not write. The
streamable-httptransport forces writes off regardless of configuration and does not expose Jira tools (ADR 0007).The MCP tool surface is frozen. A new tool there breaks the golden test and needs an ADR first.
Meeting and thread notes are immutable; immutability is how idempotency works here.
Dense retrieval is built but switched off until it passes the evaluation gate (ADR 0039).
src/sufler/
├── core/ # no I/O, no SDKs
│ ├── domain/ # models and pure logic
│ ├── ports/ # interfaces (repositories, LLM, GitHub, Jira, notifications)
│ ├── application/ # use cases and the single tool catalogue
│ └── agent/ # agent runtime
├── adapters/
│ ├── inbound/ # entry points: mcp, teams, teams_graph, cli, github
│ └── outbound/ # clients of external APIs
├── server.py # MCP server wiring
└── config/ # SUFLER_* settings, one module per domainThe bridge connects GitHub, a shared append-only event store with deduplication, and Teams; no
component calls another directly. Jira is not part of the bridge: "my tasks" is a direct, read-only
query scoped to the asking user's account. Diagrams: docs/explanation/architecture.md.
The repository ships a project-scoped .mcp.json, so Claude Code opened in this
directory picks up the sufler server. save_note is off by default; copy
.env.example to .env to enable it locally. Other entry points need extras
(uv sync --extra <name>): agent, teams, teams-graph, github, jira, retrieval,
retrieval-dense, file-reply, seed. The full list of console scripts and settings is in
docs/reference/ and docs/how-to/gate-matrix.md.
uv run --no-sync pytest
uv run --no-sync ruff check src tests eval deploy scripts
uv run --no-sync ruff format --check src tests eval deploy scripts
uv run --no-sync mypy
uv run --no-sync lint-imports # core must not import adaptersThis is the same gate CI runs (.github/workflows/ci.yml). The
sub-projects have their own environments and are tested separately. A security suite covers
injection, path traversal and secret leakage (tests/security/).
Reading is the default; every mutating capability has its own gate, off by default, opened per entry point.
Note and event content is data, not instructions. Only control characters are rejected or stripped by code; trust classes T0 to T3 label content rather than block it (ADR 0066). The defence against prompt injection is the capability gates, read-only mounts, the network-less executor and reversibility.
Secrets are read from the environment or files outside
data/, never from the repository.
sufler(src/,tests/,docs/,deploy/,eval/,scripts/): the core described above.Powiadomienia_teams/: weekly Microsoft Shifts reminders in Teams, writing shifts back only after an explicit "yes".claude_summary/: a CLI that summarises a person's day from Claude Code prompt history (with explicit consent) and commit history, redacting sensitive content at parse time.ceidg-tool/: exports data on sole proprietorships from the CEIDG register API to Excel, with a wizard and an offline--demomode.krs-tool/: reads a court-register (KRS) extract saved by the operator and reports registry signals about a company; it has no network client by design.
Each sub-project has its own environment, CI entry and README.
docs/ follows Diátaxis: tutorial, how-to, reference,
explanation, decision records in docs/adr/ and research notes. Changes between versions:
CHANGELOG.md.
License
MIT, see LICENSE.
Available Tools
4 toolsget_noteA
Pobierz pełną treść jednej notatki po jej identyfikatorze.
Identyfikator ma postać ``<firma>/<projekt>/<plik-bez-rozszerzenia>``,
np. 'mpwik/scada-integration/2025-06-12-przeglad-api-scada' (z wyników search_notes).
| Name | Required | Description | Default |
|---|---|---|---|
| note_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden of disclosing behavior. It clearly implies a read-only retrieval operation and specifies the required ID format, but it does not mention behavior on missing/incorrect IDs, authorization, or rate limits. It is minimally adequate for a simple fetch tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences with no filler. The action is stated first, and the critical parameter format and example are front-loaded in the second sentence. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read tool with an output schema present, the description covers the essential invocation details: how to construct the ID and where to obtain it. Minor omissions like error behaviors or permissions are unlikely to block correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema only defines note_id as a string with no description. The tool description compensates fully by specifying the exact structure ('<company>/<project>/<file-without-extension>') and giving a realistic example, plus the source of the value (search_notes results).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Pobierz' / retrieve) and names the exact resource: the full content of a single note by identifier. It is clearly distinguished from siblings like search_notes, which searches, and list_projects, which lists projects.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says the identifier comes from search_notes results and provides a concrete format with an example. It implies the correct workflow: first search_notes, then get_note. It does not explicitly name alternatives or exclusions, but the usage context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_project_statusA
Zwróć status projektu: stan zadeklarowany + syntezę z notatek.
``project`` to klucz projektu (np. 'workmate'). W odpowiedzi m.in. firma,
zdrowie, faza, podsumowanie oraz liczba notatek i otwartych action items.
| Name | Required | Description | Default |
|---|---|---|---|
| project | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. It discloses output fields but not side effects, auth requirements, rate limits, or behavior on missing projects. Being a read operation is implied but not stated. Minimal behavioral disclosure beyond purpose.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action, no redundant wording. Efficient and well-structured, with the essential parameter explanation in the second sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with an output schema (not shown), the description covers the main purpose, parameter semantics, and expected output fields. Lacks error-handling context and explicit usage guidance, but adequate for basic invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema has no description, so description compensates by explaining `project` as a project key with example 'workmate'. Gives clear semantics for the single parameter, exceeding what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States verb 'return' and resource 'project status', lists output fields (company, health, phase, summary, note count, open action items). Clearly distinguishes from siblings like list_projects (lists projects) and get_note (fetches a note).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implied usage is for retrieving status of a specific project, but no explicit when/when-not or alternatives are mentioned. Does not clarify when to use list_projects instead or exclude other tools. Lacks explicit routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_projectsA
Wypisz projekty pionu dostępne w bazie wiedzy (klucz, nazwa, opis).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It clearly indicates a read-only operation ('Wypisz' = list) and specifies the data source ('dostępne w bazie wiedzy'). It does not mention side effects, but none are expected for a list operation. It adds the context that the projects are from the knowledge base, which is useful. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, compact sentence that efficiently conveys the purpose and return fields. No filler or redundancy. The critical information (action, resource, returned attributes) is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that there are no parameters and an output schema exists, the description is largely complete. It specifies what the tool does and what fields are returned. The only minor gap is that 'pionu' (vertical) is not defined, but in context of the knowledge base and sibling tools this is likely understood. Overall, an agent can call this tool correctly without ambiguity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so the baseline per the rubric is 4. The description does not need to add parameter-level detail. It provides a concise summary of what the tool returns (key, name, description), which aligns with the expected output.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'Wypisz' (list) with a clear resource 'projekty pionu' (projects of the vertical) and states the returned fields (key, name, description). This clearly distinguishes it from siblings like search_notes and get_note, which deal with notes, and get_project_status, which presumably reports status rather than listing projects.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not explicitly state when to use this tool versus the sibling tools. However, the sibling names and the tool's purpose make it implicit: use it to list projects, not to search notes or get a single note/status. There are no usage exclusions or prerequisites mentioned, so it leaves some inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_notesA
Przeszukaj notatki ze spotkań po słowach kluczowych i metadanych.
Zwraca listę dopasowań (metadane + fragment), posortowaną po trafności.
Opcjonalne filtry: ``project`` (klucz projektu, np. 'scada-integration') oraz
``participant`` (fragment nazwiska uczestnika).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes | ||
| project | No | ||
| participant | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses that the tool returns a list of matches (metadata + snippet) sorted by relevance, which is useful behavioral context. It does not mention read-only behavior, error conditions, authentication needs, or how limit affects results, so transparency is partial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the main action appears in the first sentence, followed by return behavior and optional filters. Every sentence adds value, though the formatting with blank lines and code spans is slightly loose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the core search purpose, return shape, sorting, and both optional filters. Since an output schema exists, return-value documentation is not required. Missing context includes query syntax, limit semantics, and when to use get_note instead, so it is adequate but not complete for a tool with no annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains project with a concrete example ('scada-integration') and participant as a surname fragment, adding real meaning. However, it does not clarify the query parameter format beyond 'keywords' or explain the limit parameter's behavior, leaving gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Przeszukaj notatki ze spotkań' (search meeting notes) by keywords and metadata. It clearly distinguishes itself from siblings like get_note, list_projects, and get_project_status by stating it returns a ranked list of matches rather than a single note or project data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: it is for searching meeting notes and optionally filtering by project or participant. However, it never explicitly says when to prefer this over get_note or when to use get_note instead, leaving the choice between search and retrieval tools to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v1.16.0- First observed
get_note - First observed
get_project_status - First observed
list_projects - First observed
search_notes
TDQS
Scored across 4 tools
Each tool has a clearly distinct purpose: search notes, retrieve a single note, list projects, and get project status. The related search_notes/get_note pair is well-separated by search versus full retrieval.
All tool names follow the same snake_case verb_noun pattern: search_notes, get_note, list_projects, get_project_status. Verbs and object nouns are used consistently and predictably.
Four tools is well-scoped for a focused read-only knowledge base server. Each tool earns its place and covers a distinct part of the workflow without unnecessary bloat.
The core read workflows are covered: searching notes, retrieving note content, listing projects, and summarizing project status. Minor gaps include no direct way to list all notes for a project or retrieve raw action items, but these are workable through search and status.
Maintenance
Related MCP Connectors
- SeturosOAuthcom.seturos
Shared work memory for Claude Code, Codex, Cursor and chat, scoped to each repository.
Persistent, governed institutional memory for Claude Code — specs, decisions, learnings.
- OneLoreOAuthai.onelore
Shared project context for AI agents and teams: docs, tasks, and messages that stay current.
Shared project memory that keeps teammates and AI agents aligned across sessions.
Related MCP Servers
- AlicenseAqualityNot gradedmaintenanceProvides centralized knowledge management for projects, allowing users to store, search, and maintain project-specific knowledge that persists across sessions.2714 npm1-
- FlicenseNot gradedqualityBmaintenanceEnables storing and searching personal notes, documents, and snippets using semantic search and RAG capabilities across Claude Desktop, VS Code, and Open WebUI.2-
- AlicenseNot gradedqualityCmaintenanceCross-project memory for Claude Code, enabling local semantic recall and secure, git-versioned markdown storage of reusable knowledge across repositories.MIT
- AlicenseNot gradedqualityBmaintenanceEnables recall of promoted team and org knowledge from GitHub repositories via Claude.ai, supporting search and status checks.MIT