project-guides
This MCP server delivers a project's AI instruction kit — bundles, modules, overlays and profile metadata — to connected clients, and can bootstrap the kit's files into an application repo.
Read instruction bundles for the current profile (
list_bundles,get_bundle):backend,frontend,payments,architecture,infra,devops,full, each returned as concatenated Markdown plus overlay.Browse individual modules (
list_modules,get_module) by ID, e.g.stack:django-drf:backend-standard.Read the project's own context: overlay files like
.ai/project.md(get_overlay) and the resolved profile index of enabled modules/bundles (get_index).Inspect configuration metadata: prose-language rules (
get_language), API codegen choice — orval/none/graphql (get_codegen), configured AI clients (get_clients), available presets (list_presets).Check kit freshness (
check_kit_status): compares the repo's bootstrap stamp commit against the kit HEAD and lists files a re-bootstrap would overwrite.Install or refresh kit files (
bootstrap_workspace): agents, commands, hooks,mcp.jsonand stamp; the only write tool, defaulting todry_run=True(plan only) withdry_run=Falserequired to actually write.Everything besides
bootstrap_workspaceis read-only — no code execution, no repo mutation.Caveat: the schema still exposes
preset/codegenonbootstrap_workspace(documented as removed) and lacks the README'sreload_workspaceandlist_questionstools, so README and schema are out of sync.
Provides infrastructure module for task queue with Celery.
Provides stack module for Django REST Framework.
Provides stack module for Expo Router (React Native).
Provides code review integration with GitHub (Bugbot).
Provides infrastructure module for message queue with RabbitMQ.
Provides infrastructure module for cache and queue with Redis.
Provides capabilities module for payments with Stripe.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@project-guidesshow me the Django guides for this project"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Instruction Kit — MCP z instrukcjami projektów
Centralne repo MD + serwer MCP. Projekty wybierają Stack per Tier (backend/web/mobile w .ai/project.profile.yaml) + overlay.
Szybki start
Trzy kroki cyklu życia kita w projekcie. Szczegóły każdego kroku są niżej.
Krok 1 — Instalacja
uv tool install git+https://github.com/radthenone/ai-instruction-kit-mcp
cd /m/projects/moja-appka # repo aplikacji
kit-ai installBez @ instalujesz gałąź domyślną (master). Nowsze, jeszcze nie wydane zmiany są na
dev — wtedy dopisz ref, np. …/ai-instruction-kit-mcp@dev.
Krok 2 — Konfiguracja
Zrestartuj IDE / CLI, potem w kliencie AI:
/kit-project-begin # pierwszy raz: pytania o Stack, karta Profilu i .ai/project.md
/kit-project-edit "…" # później: jedna zmiana, np. "zmień web na angular"/kit-project-begin proponuje odpowiedzi wykryte w repo, a /kit-project-edit zmienia jedną
odpowiedź bez całego wywiadu (pełny opis). Po zapisie
zrestartuj klienta.
Krok 3 — Update
uv tool upgrade guides-mcp # nowa wersja kita
kit-ai reload # odśwież pliki kita w projekcie
kit-ai status # w agencie: check_kit_statusRelated MCP server: CodeGuard MCP Server
Szczegóły instalacji
Jedna komenda robi instalację i update. Bootstrap jest idempotentny: pliki generowane (agenci, komendy,
hooki, mcp.json) nadpisuje świeżą kopią, a pliki z Twoją treścią (AGENTS.md,
.ai/project.md, BUGBOT.md) zostawia w spokoju.
Wymagania: uv w PATH, bash (Windows: Git for Windows), node (dla hooków),
opcjonalnie npx (skille zewnętrzne).
1. Instalacja / update w projekcie — kit-ai
Bez klona kita — kit-ai jako narzędzie uv (ref po @: branch, tag albo commit):
uv tool install git+https://github.com/radthenone/ai-instruction-kit-mcp # master
uv tool install git+https://github.com/radthenone/ai-instruction-kit-mcp@dev # albo gałąź dev
cd /m/projects/moja-appka # repo aplikacji
kit-ai install # bez ścieżki = bieżący katalog
kit-ai reload | kit-ai status | kit-ai remove [--dry-run]
# jednorazowo, bez instalowania narzędzia:
uvx --from git+https://github.com/radthenone/ai-instruction-kit-mcp kit-ai installWybrany ref trafia do mcp.json projektu (np. uvx --from git+…@dev), a commit kita do
stampu — kit-ai status porównuje go z commitem zainstalowanego narzędzia. Update:
uv tool upgrade guides-mcp (albo uv tool install --force git+…@<inny-ref>) +
kit-ai reload. Instalacja z lokalnej ścieżki (uv tool install .) nie zna commitu —
status mówi wtedy „nieznany commit”.
install pyta o dwie rzeczy (Język [pl/en] (pl), Klienci (…) (all)) — flagi --language
i --clients pomijają pytania, bez TTY pytań nie ma wcale. Zakłada .ai/project.profile.yaml
(backend/web/mobile: none = sam core) i .ai/project.md, robi Bootstrap, a na końcu
wypisuje JSON serwera MCP i gdzie leży per klient. Repo z kitem install odrzuca — wtedy
reload.
Jedyna konfiguracja to Profil (ADR-0007): język, klienci, Stacki per Tier, codegen:.
Zmiana czegokolwiek = edycja .ai/project.profile.yaml + kit-ai reload. reload nie
rusza .ai/project.md; repo ze starą konfiguracją (stamp z --preset, brak Profilu) dostaje
Profil core + none i nowy mcp.json.
Klucz Profilu | Kiedy zmienić |
|
|
| Stack per Tier (puste Tiery = sam core) |
|
|
|
|
Niskopoziomowo to samo robi scripts/bootstrap-project.sh "$APP" --from "$KIT" --clients … --language …
(dodatkowo --with-overlay, --with-plugins, --keep-unselected-clients). --preset,
--profile, --with-profile i --codegen zostały usunięte — skrypt odmawia i odsyła do kit-ai reload.
2. Po instalacji (kroki, których skrypt nie zrobi za Ciebie)
kit-ai install kończy się blokiem „Dalej” (Superpowers + /kit-project-begin). Pełna
lista poniżej — rób tylko to, czego jeszcze nie masz, i tylko dla klientów z clients:.
Wszystko poza 2.5 robisz raz na maszynę, nie per projekt.
Komendy są te same w Git Bash (Windows) i na Linuksie, chyba że wiersz mówi inaczej. Różnice Windows zebrane są w 2.4.
2.1. rtk — sprawdź, czy jest
rtk --version && rtk gain # oba muszą zadziałać; "command not found" = brak rtkrtk gain nie działa, a rtk --version tak → masz inne narzędzie o tej nazwie
(reachingforthejack/rtk), nie Rust Token Killer. Brak rtk niczego nie psuje — agenci
wykonują wtedy komendy bez prefiksu (core:tooling-rtk).
Instalacja:
# Linux / macOS
curl -fsSL https://raw.githubusercontent.com/rtk-ai/rtk/refs/heads/master/install.sh | sh
# Windows (winget działa też z Git Bash)
winget install rtk-ai.rtkHook, który przepisuje komendy na rtk … poza modelem — po jednym na klienta:
Klient | Komenda | Uwagi |
Claude Code |
| Hook kita |
Codex |
| Potem |
OpenCode |
| |
VS Code (Copilot) | — | Hook przychodzi z kita: |
Antigravity |
|
Po rtk init zrestartuj klienta. Weryfikacja: rtk init --show.
2.2. Pluginy i skille zewnętrzne — per klient
Kolejność bez znaczenia; wszystkie są opcjonalne poza Superpowers (warstwa 3 w
AGENTS.md). npx skills przyjmuje -a claude-code|codex|opencode|github-copilot|antigravity
i dla klientów poza Claude instaluje do .agents/skills/ (dodaj -g, żeby globalnie).
Wyjątek: Antigravity CLI (agy) nie czyta globalnego ~/.agents/skills/, gdzie -g
kładzie skille — globalnie widzi tylko ~/.gemini/config/skills/
(antigravity-cli#103).
Superpowers (obra/superpowers)
Klient | Komenda |
Claude Code | Proponuje się sam przy otwarciu repo (bootstrap wpisuje go do |
Codex | W TUI: |
OpenCode | Napisz agentowi: |
VS Code (Copilot) | Tylko Copilot CLI: |
Antigravity |
|
Skille Matta Pococka (mattpocock/skills, aihero.dev/skills) — /grill-me, /tdd
Klient | Komenda |
Claude Code |
|
Codex, OpenCode, VS Code |
|
Antigravity — per projekt |
|
Antigravity — globalnie |
|
Caveman (JuliusBrussee/caveman)
Klient | Komenda |
Claude Code |
|
Codex |
|
OpenCode |
|
VS Code (Copilot) |
|
Antigravity |
|
Ponytail (DietrichGebert/ponytail)
Klient | Komenda |
Claude Code |
|
Codex |
|
OpenCode | Klon repo + w |
VS Code (Copilot) | Tylko Copilot CLI: |
Antigravity |
|
Context7 (upstash/context7) — docs bibliotek (MCP)
Najprościej npx ctx7 setup --claude / --opencode (OAuth + klucz API + konfiguracja).
Reszta ręcznie — serwer zdalny https://mcp.context7.com/mcp, nagłówek
Authorization: Bearer <KLUCZ> (wszystkie klienty):
Klient | Gdzie | Wpis |
Claude Code | CLI |
|
Codex |
|
|
OpenCode |
|
|
VS Code |
|
|
Antigravity |
|
|
CodeGraph (@colbymchenry/codegraph) — graf kodu (MCP + CLI, nie skill)
npm i -g @colbymchenry/codegraph
codegraph install --target claude,codex,opencode,copilot-vscode,antigravity --location global
codegraph init -i # w każdym repo — buduje indeks .codegraph/Zostaw w --target tylko swoich klientów. codegraph install --print-config <id> pokazuje
wpis bez zapisu.
GitHub MCP (github/github-mcp-server) — issues/PR z poziomu agenta
Serwer zdalny https://api.githubcopilot.com/mcp/. Poza VS Code (OAuth) potrzebny PAT
(github.com/settings/tokens, scope repo, read:org, read:user)
w zmiennej środowiskowej, nie w pliku w repo:
Klient | Komenda / wpis |
Claude Code |
|
Codex |
|
OpenCode |
|
VS Code |
|
Antigravity |
|
2.3. Gdzie co ląduje
.agents/skills/ jest wspólny dla Codexa, OpenCode, Copilota i Antigravity — skill dodany
przez npx skills -a codex zobaczą też pozostali. Kit ignoruje ten katalog w .gitignore,
więc na nowej maszynie instalację powtarzasz.
Globalnie Antigravity CLI czyta tylko ~/.gemini/config/skills/ — npx skills update
odświeża ~/.agents/skills/, więc po aktualizacji skopiuj skille ponownie.
2.4. Windows (Git Bash) vs Linux
npx skillsdomyślnie robi symlinki, a te na Windowsie wymagają Developer Mode albo admina. Bez tego dodaj--copy.Lokalny (stdio) serwer MCP odpalany przez
npxuruchamiaj na Windowsie jakocmd /c npx …. Serwery zdalne (Context7, GitHub wyżej) tego nie potrzebują.Pełny instalator Caveman: na Windowsie
install.ps1, nieinstall.sh— hooki i tak wołają wersje PowerShell.Hooki Ponytail i Caveman (Claude, Codex) potrzebują
nodew PATH powłoki nieinteraktywnej — przy nvm to częsta pułapka.Zmienne z tokenami: Git Bash / Linux
export GITHUB_PAT_TOKEN=…, PowerShell$env:GITHUB_PAT_TOKEN = "…". Klient musi wystartować już z ustawioną zmienną.
2.5. Konfiguracja projektu
/kit-project-begin — patrz Krok 2.
Hooka pre-push kit już nie dostarcza (guardraile siedzą w hookach klientów);
istniejący git-hooks/pre-push w Twoim repo zostaje nietknięty.
Zrestartuj IDE / CLI. MCP i komendy ładują się przy starcie — bez restartu zobaczysz stan sprzed bootstrapu.
3. Weryfikacja
# MCP odpowiada i widzi właściwy kit
# w Claude Code: poproś o wywołanie narzędzia check_kit_status
# oczekiwane: "Kit status: aktualny"
# hooki działają (powinno wypisać "deny")
printf '%s' '{"tool_input":{"command":"git reset --hard HEAD"}}' \
| node .claude/hooks/git-guard.mjs
# konfiguracja AI wchodzi do repo, lokalny stan nie
git status --short -uall .claude .codex .github/prompts4. Kiedy aktualizować
Narzędzie MCP check_kit_status porównuje commit kita zapisany przy bootstrapie
(.ai/.kit-bootstrap.json) z aktualnym HEAD i mówi, co się zmieniło. Rozdziela dwie
rzeczy: pliki, które re-bootstrap wciągnie sam, i te wymagające ręcznego
przeniesienia (AGENTS.md, BUGBOT.md, .ai/project.md — kopiowane
tylko gdy brak, żeby nie zdeptać Twojej treści). Gdy pokaże zmiany: kit-ai reload
(z terminala to samo pokazuje kit-ai status).
5. Zanim odpalisz update na repo z pracą w toku
Bootstrap nadpisuje .claude/{agents,commands,hooks}/, .codex/, .github/prompts/,
copilot-instructions.md i pliki MCP. Jeśli edytowałeś je ręcznie — git diff najpierw.
Nie chcesz oglądać planu na sucho? Z poziomu agenta:
reload_workspace() # dry run — lista plików nowych/nadpisanych/usuniętych
reload_workspace(dry_run=False) # odświeżenie z Profilu (= kit-ai reload)Gdzie żyje konfiguracja AI po instalacji: wszystko poza .claude/settings.local.json
i .agents/skills/ idzie do repo — bootstrap wstawia do .gitignore sekcję między
markerami # >>> instruction-kit >>>. Szczegóły: sekcja „.gitignore" niżej.
Gdzie czytać / zmieniać konfigurację:
Co | Gdzie pisać |
Argumenty MCP ( | ten README (sekcja niżej) + szablony |
Stack per Tier i fork | |
Kanon agentów / reguł (niezależny od IDE) | |
Multi-client design | |
Szczegóły jednego produktu |
|
Inny zestaw modułów niż Tiery |
|
Docelowy kontrakt | |
Cursor |
|
Skille kita (wszyscy klienci) |
|
Struktura docs/
docs/
├── adr/ — decyzje architektoniczne (format Nygarda)
├── agents/ — kontrakt issue trackera, etykiety triage, docs domenowe
├── specs/ — projekty przed implementacją
└── plans/ — plany implementacyjne
docs/specs/idocs/plans/nazywały się wcześniejdocs/superpowers/{specs,plans}. Zmiana jest celowa: Superpowers imattpocock/skillsto zewnętrzne biblioteki, z których kit korzysta — ich nazwa nie powinna strukturyzować drzewa docs tego repo.
Konfiguracja projektu — argumenty guides-mcp
Wszystkie flagi serwera MCP wpisujesz w args klienta (Cursor: .cursor/mcp.json). Kolejność: najpierw --from / nazwa pakietu (guides-mcp), potem flagi poniżej.
Warstwy (co gdzie należy)
Warstwa | Mechanizm | Przykład |
Stack backendu | Tier |
|
Stack webu | Tier |
|
Stack mobile | Tier |
|
Powtarzalny wariant |
|
|
Fakty jednego repo |
| porty, Taskfile |
Inny zestaw modułów niż Tiery |
|
|
Nie mieszaj: nazwa produktu ≠ Stack; porty ≠ tag.
Flagi (aktualne)
{
"mcpServers": {
"project-guides": {
"command": "uvx",
"args": [
"--from", "git+https://github.com/radthenone/ai-instruction-kit-mcp.git",
"guides-mcp",
"--language", "pl",
"--clients", "all",
"--workspace", "${workspaceFolder}"
]
}
}
}Flaga | Wymagana? | Rola | Gdzie / jak zmieniać |
| przy | Źródło zdalne kita: |
|
`--language pl | en` | nie | Język prozy (odpowiedzi, docstringi, body issue/PR, commity). Tytuły issue/PR/branch zawsze EN. Domyślnie: |
| nie | Metadane IDE: | mcp.json / bootstrap |
| nie | Klon kita, z którego serwer czyta | mcp.json — bootstrap dodaje sam przy źródle lokalnym; env |
| zalecane | Root aplikacji — stąd profil | mcp.json; Cursor/VS: |
| nie | Extra MD (można wielokrotnie) | mcp.json — rzadko; zwykle wystarczy workspace |
Stare konfiguracje klienta z --preset / --profile / --codegen nadal startują serwer, ale
te flagi są ignorowane, a bundle i indeks niosą ostrzeżenie o migracji — kit-ai reload
przepisze mcp.json. Bootstrap zapisuje tylko --language i --clients (z Profilu).
Język: MCP tool get_language. Priorytet: --language / GUIDES_LANGUAGE → language: w YAML profilu → pl. Moduł w bundle: core:language-pl albo core:language-en.
Klienci AI: MCP tool get_clients — tylko metadane instalacji; treść get_bundle jest identyczna dla każdego klienta.
Codegen (Orval): wyłącznie codegen: w .ai/project.profile.yaml (orval | none | graphql, domyślnie orval; ADR-0007), odczyt przez MCP tool get_codegen. Bez pary backend + klient (web/mobile) efektywny codegen to zawsze none. Reviewery FE/BE to honorują (przy orval wymagają regeneracji klienta po zmianie API; graphql → moduł arch:api-contract:graphql zamiast REST).
Profil z Tierami (backend/web/mobile)
# .ai/project.profile.yaml w repo aplikacji
name: moja-appka
language: pl
clients: claude,codex
backend: django # none | django | django-html | fastapi | flask
web: react # none | react | react@legacy | angular | angular@rxjs | expo
mobile: none # none | expo | react-native
codegen: orval # orval | none | graphql (bez pary backend + klient: zawsze none)
capabilities:
- auth
- payments
decisions:
database: postgres
auth: jwtBundle liczone są z Tierów: get_bundle backend zawiera Stack z Tieru backend, pusty profil daje sam core. Nierozpoznana wartość Tieru nie wywraca serwera — ląduje w „Nierozpoznanych decyzjach" w get_index (ADR-0004). Stare klucze stacks: / patterns: czytane są nadal.
Katalog pytań o projekt (list_questions)
manifest.yaml → questions: trzyma pytania o projekt (Tiery, warianty, codegen, Docker,
Taskfile, CI/CD, monorepo, capability-provider, webhooki, ścieżki per Tier) z opcjami,
domyślnymi (defaults: zależne od Profilu, np. django-html → web/mobile none),
warunkiem when: i sygnałami detect: (glob + opcjonalny regex pattern w treści pliku).
Narzędzie MCP list_questions zwraca katalog z warunkami ocenionymi na bieżącym Profilu —
sygnały sprawdza agent w plikach repo, Python niczego nie skanuje. Odpowiedź ląduje tam, gdzie
wskazuje sets: (klucz Profilu albo paths.<tier> → ## Ścieżki w .ai/project.md), a pytania
tak/nie dopisują include: / patterns: z on_yes:.
Moduły układu katalogów (stack:frontend:*) nie wchodzą do Bundli — layouts: w manifeście
wybiera podpowiedź drzewka dla kombinacji web/mobile (np. web: react + mobile: expo →
react-expo-split), web: expo + mobile: react-native daje ostrzeżenie, brak drzewka →
domyślne ścieżki (backend/, frontend/web/, frontend/mobile/; Expo unified: frontend/).
Tagi / facety (planowane — jeszcze nie w CLI)
Gdy wiele projektów dzieli ten sam powtarzalny wariant instrukcji (np. sklep fizyczny vs cyfrowy), zamiast mnożyć presety shop-jewelry / shop-tokens:
W
manifest.yaml→mappings.tiers.<tier>dopisać dozwolone Stacki (np.fulfillmentnie — Tiery to backend/web/mobile; nowy wymiar trafia dodecisionsalbocapabilities).W mcp.json dodać np.
"--tag", "physical"albo"--facet", "fulfillment=physical"(docelowa składnia przy implementacji).Resolver dołoży wtedy dodatkowe MD z
modules/— bez lokalnego forka, jeśli zestawy capabilities są te same.
Teraz: różnice jubiler vs tokeny → .ai/project.md. Tagi włączaj dopiero gdy wariant wraca w ≥2–3 projektach.
Szkic (nie działa jeszcze):
"args": [
"--from", "…",
"guides-mcp",
"--tag", "physical",
"--tag", "b2c",
"--workspace", "${workspaceFolder}"
]Bootstrap
# Generyczny — profil z Tierami + --language pl
./scripts/bootstrap-project.sh /sciezka/do/projektu \
--from /absolutna/sciezka/do/ai-instruction-kit-mcp \
--with-overlay
# Tylko Cursor
./scripts/bootstrap-project.sh /sciezka/do/projektu \
--clients cursor \
--from /absolutna/sciezka/do/ai-instruction-kit-mcp
# Proza EN, wszyscy klienci AI (Stacki potem w .ai/project.profile.yaml)
./scripts/bootstrap-project.sh /sciezka/do/moj-sklep \
--language en \
--clients all \
--from /absolutna/sciezka/do/ai-instruction-kit-mcpAgenci per Tier: agenci z tier: we frontmatterze (templates/shared/agents/) trafiają do klienta tylko przy wybranym Tierze — tier: backend (review-backend, teacher-backend, subagent-backend) gdy backend ≠ none, tier: client (review-frontend, teacher-frontend, subagent-frontend, review-ui) gdy web lub mobile ≠ none. Tier zmieniony na none + kit-ai reload = ich pliki znikają u wszystkich klientów. Agenci nie zakładają Stacka — biorą go z get_bundle. BUGBOT.md dostaje sekcje (<!-- tier:backend -->, <!-- tier:client -->) tylko wybranych Tierów.
Zapisuje m.in. MCP per klient (--language, --clients, --workspace), agents z templates/shared/agents, BUGBOT.md w root (wszyscy klienci) + .cursor/BUGBOT.md (natywny Cursor BugBot), skill Cursor /compact, hooki gate-* (Cursor), stamp .ai/.kit-bootstrap.json (patrz "Update kita w projekcie"). Wymaga Python 3 (python3 albo python z major==3).
Declarative sync klientów: domyślnie bootstrap usuwa kitowe pliki klientów spoza --clients (np. przełączenie z --clients all na --clients claude sprząta .cursor/, .codex/ itd. wygenerowane przy poprzednim bootstrapie). Flaga --keep-unselected-clients wyłącza to sprzątanie — zostają pliki wszystkich klientów kiedykolwiek bootstrapowanych.
Sprzątanie kasuje wyłącznie pliki kita, po nazwie — listę bierze z przebiegu tych samych funkcji instalacji w pustym katalogu. Własne agenty, komendy, hooki i skille w .claude/, .opencode/, .codex/, .github/prompts/ itd. zostają; katalog znika tylko, gdy po kicie jest pusty. Plik użytkownika o nazwie identycznej z plikiem kita (np. własny .claude/agents/git-start.md) zostanie usunięty razem z kitowymi.
kit-ai remove [ścieżka] [--dry-run] — odinstalowanie: pliki kita wszystkich klientów, konfiguracje MCP, wpisy kita w .claude/settings.json, sekcja # >>> instruction-kit >>> w .gitignore, Profil i stamp (także stara konfiguracja z presetem). Pliki tworzone raz (AGENTS.md, BUGBOT.md, .gitattributes) znikają tylko, gdy są identyczne z bieżącym szablonem kita — zmienione przez Ciebie zostają. Zawsze zostają .ai/project.md, CONTEXT.md, docs/adr/ i Twoje pliki. Kit nie robi kopii, więc nic nie przywraca. --dry-run pokazuje listę bez usuwania.
.gitignore — co z tego wersjonować
Bootstrap wstawia do .gitignore repo aplikacji sekcję między markerami
# >>> instruction-kit >>> i # <<< instruction-kit <<<. Przy kolejnych przebiegach
podmienia ją w całości, więc wpisy się nie duplikują, a reguły spoza markerów zostają
nietknięte. Źródło: templates/gitignore-kit.txt.
Zasada: konfiguracja AI jest częścią repo. Hooki bezpieczeństwa, agenci i komendy mają działać u każdego, kto sklonuje projekt — nie tylko na maszynie, gdzie odpalono bootstrap. Poza gitem zostaje lokalny stan klienta, to, co i tak żyje globalnie, oraz pliki, które bootstrap renderuje ze ścieżką tej maszyny:
Wersjonowane | Ignorowane |
|
|
|
|
|
|
|
|
— |
|
Konfigi MCP i stamp są per maszyna, nie per repo. Przy --from <lokalny klon> bootstrap
wpisuje do nich absolutną ścieżkę klona (uv run --project, --kit-root), a dla Codex
i opencode absolutny --workspace. Zacommitowane z Windowsa (M:/projects/…) na Linuksie
dają CONNECTION_CLOSED bez czytelnego powodu. Każdy odbiornik — PC, laptop, serwer —
odpala kit-ai reload u siebie; język i klienci są w zacommitowanym Profilu, więc nic
więcej nie trzeba pamiętać.
Repo zbootstrapowane wcześniej mają te pliki w indeksie — sam wpis w .gitignore ich nie
odśledzi. Bootstrap wykrywa to i wypisuje gotową komendę (pliki zostają na dysku):
git -C "$APP" rm --cached .mcp.json .vscode/mcp.json .codex/config.toml .ai/.kit-bootstrap.jsonTypowy .gitignore ma .claude/ wpisane hurtem — wtedy hooki i komendy nigdy nie trafiają
do repo, a bootstrap trzeba powtarzać na każdej maszynie. Reguły kita są w formie „ignoruj
katalog, odwróć dla plików kita", bo git nie wchodzi do zignorowanego katalogu i sam wyjątek
na plik by nie wystarczył.
Bootstrap bez klona kita — narzędzie MCP bootstrap_workspace
Jeśli projekt ma już podłączony serwer MCP project-guides, kita nie trzeba klonować ani ręcznie odpalać skryptu — serwer ma szablony pod ręką i uruchamia ten sam bootstrap-project.sh u siebie. Poproś agenta o wywołanie narzędzia:
bootstrap_workspace() # dry run — tylko lista plików
bootstrap_workspace(dry_run=False) # instalacja
bootstrap_workspace(clients="claude", with_overlay=True, dry_run=False)Odświeżenie z Profilu (= kit-ai reload, łącznie z migracją starej konfiguracji) robi reload_workspace() / reload_workspace(dry_run=False).
Argumenty bootstrap_workspace (clients, language, with_overlay, keep_unselected_clients) odpowiadają flagom skryptu; pominięte biorą wartość z parametrów startowych serwera MCP. Cel zapisu to --workspace / GUIDES_WORKSPACE — bez niego narzędzie odmawia, zamiast zapisywać do katalogu, z którego przypadkiem wystartował proces serwera.
dry_run=True jest domyślne i nic nie zapisuje: skrypt leci na kopii kitowej powierzchni repo w katalogu tymczasowym, a raport pokazuje pliki nowe, nadpisane i usunięte przez sprzątanie klientów spoza --clients. Plan pochodzi więc z faktycznego przebiegu skryptu, nie z drugiej listy ścieżek w Pythonie.
Uwaga na bramki. Hooki kita (
PreToolUse) łapiąBash,PowerShelliEdit|Write|MultiEdit|NotebookEdit— nie nazwy narzędzi MCP. To jedyne zapisujące narzędzie tego serwera i hooki go nie zatrzymają;dry_run=Truejako domyślka plus wymóg jawnegodry_run=Falsesą tu całą ochroną. Reszta narzędzi serwera pozostaje tylko do odczytu.
MCP w innych klientach (multi-client)
Kanon treści: templates/shared/{agents,rules}. Adaptery IDE trzymają tylko format MCP / ścieżki natywne. Bootstrap --clients instaluje wybrane pakiety (default all).
Klient | Id | Plik MCP w aplikacji | Klucz top-level | Szablon |
Cursor |
|
|
|
|
Claude Code |
|
|
|
|
Codex CLI |
|
|
|
|
GitHub Copilot (VS Code) |
|
|
|
|
Kiro |
|
|
|
|
Kilo |
|
|
|
|
Antigravity |
|
|
|
|
opencode |
|
|
|
|
Zmienna dla --workspace:
Klient | Zmienna |
Cursor, VS Code, Kiro, Kilo, Antigravity |
|
Claude Code |
|
Codex CLI, opencode | ścieżka absolutna (brak stabilnej zmiennej) |
Instalacja per klient (krok po kroku)
Wspólne dla wszystkich: git clone / masz kita lokalnie → uruchom bootstrap-project.sh w repo aplikacji (nie w repo kita) z --from wskazującym na kita → zrestartuj IDE.
./scripts/bootstrap-project.sh /sciezka/do/mojej-appki \
--from /m/projects/ai-instruction-kit-mcp \
--clients cursor \
--with-overlayKlient |
| Wymaga poza kitem | Extra config po bootstrapie |
Cursor |
| Cursor IDE | Ustaw |
Claude Code |
|
|
|
Codex CLI |
|
|
|
GitHub Copilot (VS Code) |
| VS Code + rozszerzenie GitHub Copilot Chat |
|
Kiro |
| Kiro IDE |
|
Kilo Code |
| rozszerzenie Kilo Code |
|
Google Antigravity |
| Antigravity IDE |
|
opencode |
|
|
|
Wiele klientów naraz: --clients cursor,claude albo --clients all. Każdy klient dostaje ten sam --language/--workspace — różni się tylko format pliku MCP i ścieżka komend.
Po bootstrapie zawsze: zrestartuj IDE/CLI (MCP i komendy ładują się przy starcie), potem sprawdź że MCP wstał (np. get_bundle / lista narzędzi w kliencie).
Czego kit nie robi / brakujące komendy
Świadome braki — nie zgłaszaj jako bug, tylko sprawdź czy potrzebujesz obejścia niżej:
Brak | Status | Obejście |
| Zaprojektowane, nie w CLI | Różnice trzymaj w |
| Usunięte (serwer je ignoruje, bootstrap odmawia) | Profil zawsze |
| Nie istnieje w | To skill user/global (Cursor) — dodaj we własnym środowisku, kit go nie dostarcza |
| Nie istnieje dla Claude/Codex/inne | To alias Cursor UI Summarize; Claude Code ma wbudowane |
Natywna weryfikacja formatu VS Code/Kilo/Antigravity/opencode | Oparta o dokumentację (sierpień 2026), nie testowana na żywych klientach | Jeśli |
Auto-instalacja Superpowers/Autopilot | Niemożliwa ze skryptu (marketplace pluginów Claude/Cursor, wymaga interaktywnego |
|
Pluginy zewnętrzne — schemat użycia (4 warstwy)
Kit nie bundluje tych pluginów w guides-mcp (różna dystrybucja: MCP vs Claude/Cursor plugin marketplace vs npx skill). Pełna tabela warstw i priorytet źródeł: AGENTS.md.
1. Fundament — ten kit (MCP + /git-* + /review-*) → instaluje bootstrap
2. Proces — mattpocock/skills (/grill-me, /tdd) → npx skills@latest add mattpocock/skills
3. Meta/izolacja — Superpowers (worktree, finishing…) → Claude Code: /plugin marketplace add obra/superpowers-marketplace
/plugin install superpowers@superpowers-marketplace
4. PR → green — Autopilot (Cursor) → Cursor: Settings → Extensions/Skills → AutopilotNie mieszaj warstw: kit = prawda o stacku i nazwach branchy, Matt = proces feature, Superpowers = sesja/worktree/finisz, Autopilot = dociąganie PR.
Auto-instalacja przy bootstrapie: --with-plugins (best-effort, opt-in — nic nie instaluje się bez tej flagi):
./scripts/bootstrap-project.sh ../moj-projekt \
--clients claude \
--from /m/projects/ai-instruction-kit-mcp \
--with-pluginsCo robi: odpala npx skills@latest add mattpocock/skills w TARGET (wymaga npx/Node.js w PATH; best-effort — błąd nie przerywa bootstrapu), i wypisuje gotowe komendy do Superpowers/Autopilot (te dwa wymagają interaktywnego kroku w kliencie, nie da się ich odpalić z bash). TDD: jeden path na feature — domyślnie Matt /tdd, nie mieszaj z Superpowers TDD.
Katalog modułów
modules/
core/ repo-first, workflow, typing, code-review, language-*, tooling-rtk
architecture/ platforms, CI/CD, API (REST/GraphQL), security, testing, i18n,
taskfile, docker-structure, …
stacks/
django-drf/ (+ django/, fastapi/, flask/ layouts)
expo-router/
frontend/ warianty Expo/React (macierz web/mobile — design)
capabilities/ auth (+ allauth/jwt/custom warianty), files, payments (+ expo-stripe gdy Tier expo), …
patterns/ capability-provider, providers-and-settings, gateway, webhooks, …
infra/ database, cache, queue, storage, tasks, search, vps-lightweight (auto na VPS)
templates/
shared/ kanon agents + rules (źródło prawdy)
cursor|claude|… adaptery MCP / format IDESloty infrastruktury (decisions)
decisions:
database: postgres # → infra:database:postgres
cache: redis # → infra:cache:redis
queue: redis # → infra:queue:redis (lub rabbitmq)
storage: s3 # → infra:storage:s3
tasks: celery # → infra:tasks:celery
search: postgres # → infra:search:postgres (lub meilisearch)Moduły infra trafiają automatycznie do bundle infra i devops.
Dodanie nowej technologii nie wymaga Pythona (ADR-0001). Trzy kroki:
Napisz
modules/infra/queue/kafka.md.Zarejestruj go w
manifest.yaml→modules:.Dopisz wartość w
manifest.yaml→mappings.slots.queue.kafka.
Nierozpoznana Decyzja (literówka postgress, technologia bez modułu) nie wywraca
serwera — ląduje w sekcji „Nierozpoznane decyzje" w get_index (ADR-0004).
Lekki tryb VPS (infra:vps-lightweight)
Agenci na słabym VPS-ie (np. 1 vCPU, 4,5 GB RAM, bez swapu, LXC) nie mogą zachowywać
się jak na maszynie dewelopera — pełny docker compose up, next dev czy e2e
potrafią zabić produkcję obok przez OOM.
Wykrywanie bez nazw hostów: Linux, brak CI/WSL, małe zasoby (CPU ≤ 2, RAM ≤ 8 GB, brak swapu) i sygnał serwerowy (wirtualizacja z
systemd-detect-virtalbo brak desktopu). Nadpisanie:KIT_HOST_PROFILE=vps|local(aliasGUIDES_HOST_PROFILE). CI zawsze dostajelocal— pełna weryfikacja należy do CI.get_bundlena takim hoście doklejainfra:vps-lightweightdo każdego bundle'a (teżbackend/frontend/architecture);get_indexpokazuje- Host: vps|local. Bez hooka blokującego — tylko reguła tekstowa, żeby nie zatrzymać celowego deployu.Moduł operuje kategoriami (kontenery, dev serwery, buildy, e2e vs lint/typecheck/ unit in-memory), nie zakłada Django/React/Next ani innego stacku.
Podział Twojego repo dopisz w overlay (
.ai/project.md), np.task lintlekkie vstask test:e2eciężkie — jedna zmiana przez/kit-project-edit.
Wariant auth (decisions.auth)
decisions:
auth: custom # default — brak enforced pakietu, opisz w .ai/project.md
# auth: allauth # → capability:auth:allauth (django-allauth headless)
# auth: jwt # → capability:auth:jwt (djangorestframework-simplejwt)Inny mechanizm niż infra: nie tworzy osobnego bundle'a — dokleja się zaraz po
capability:auth wszędzie tam, gdzie ten moduł już jest wypisany w bundle
(capabilities: [auth] albo include: z ID modułu — routing wg tagów).
W manifeście to mappings.variants.auth (Wariant = wstaw po module bazowym),
w odróżnieniu od mappings.substitutions.codegen (Substytucja = podmień moduł bazowy).
Słownik i decyzje
Plik | Rola |
Ubiquitous language kita — Bundle, Preset, Slot, Wariant, Alias, Overlay… | |
Decyzje architektoniczne z uzasadnieniem (dlaczego tak, a nie inaczej) |
Nazwy z CONTEXT.md obowiązują w kodzie, docstringach i review. Zanim zaproponujesz
zmianę architektury, sprawdź docs/adr/ — część rzeczy już rozstrzygnięto.
Bundle'e MCP
Bundle | Zastosowanie |
| Stack z Tieru backend, capabilities BE |
| Stacki z Tierów web i mobile, UI/UX |
| Stripe, webhooks |
| monorepo, kontrakt API, capability-provider |
| postgres, redis, queue, s3, celery |
| CI/CD + infra |
| wszystko + infra |
Bootstrap w projekcie docelowym
W repo aplikacji uruchom scripts/bootstrap-project.sh albo skopiuj z templates/:
Plik | Rola | Wymagany? |
| uvx → | tak (Cursor) |
| MCP per klient z | wg wybranego klienta |
| Overlay — Taskfile, Docker, porty | zalecany |
| .ai/project.profile.yaml | Tiery (backend/web/mobile) + codegen: — jedyna konfiguracja kita | tak |
| .cursor/rules/use-guides.mdc | Bootstrap MCP | tak |
| .cursor/rules/code-review.mdc | Review przed pushem | tak |
| .cursor/rules/git-branch-pr.mdc | /git-start+/git-check+/git-commit+/git-end, issue#, chronione main/master/dev | tak |
| .cursor/BUGBOT.md | Reguły Bugbota | tak |
| .cursor/hooks.json + hooks/invoke-hook.js + hooks/*.mjs | Guardy: git-guard + sensitive-files (adapter → node) | tak |
| AGENTS.md | Cienki — odsyła do MCP | tak |
| .cursor/agents/*.md | Subagenty /review-*, /subagent-*, /git-* | zalecany |
| .cursor/skills/compact/ | Tylko Cursor: /compact = alias UI Summarize (nie Claude/Codex) | zalecany (Cursor) |
W projekcie docelowym nie duplikuj modules/ — wystarczy profil z Tierami + opcjonalny overlay.
Update kita w projekcie
Komendy update: Szybki start, krok 3. Poniżej to, co update robi z plikami, i lokalny klon kita.
Bootstrap to jednorazowy stempel, nie sync. Trzy różne zachowania:
Co | Przy ponownym |
| Zawsze nadpisane świeżą kopią z kita — traktuj jak wygenerowany kod, nie edytuj ręcznie. |
| Kopiowane tylko jeśli brak — bootstrap nigdy więcej ich nie tyka, update ręczny. |
| W ogóle nie kopiowane — MCP czyta je z |
Lokalny klon: uv run --project, nie uvx --from
uvx --from <katalog> nie czyta kita z tego katalogu w czasie działania. uv buduje koło,
w którym manifest.yaml i modules/ lądują jako guides/_data
(force-include w pyproject.toml), i cache'uje je pod wersję pakietu. Wersja nie rośnie
przy zwykłej edycji modułu ani kodu serwera, więc klient dostaje kopię sprzed builda.
Do tego find_kit_root() woli _data od repo, więc check_kit_status traci historię gita.
Objaw: poprawiasz modules/…, restartujesz klienta, a get_bundle wciąż zwraca starą treść.
Bez komunikatu błędu. To samo dotyczy poprawek w src/guides/ — serwer nadal biegnie na
starym kodzie.
Dlatego przy źródle lokalnym bootstrap generuje:
"command": "uv",
"args": ["run", "--project", "/sciezka/do/klona", "guides-mcp", …,
"--kit-root", "/sciezka/do/klona", …]Pakiet ma układ src/, więc uv run instaluje go jako editable — _data w ogóle nie
powstaje, a kod i moduły czytane są wprost z klonu. --kit-root nie jest wtedy konieczny,
ale zostaje: nazywa klon wprost, zamiast pozwalać serwerowi go wnioskować.
Przy źródle zdalnym (git+https://…) nic się nie zmienia — zostaje uvx --from, bo klonu
nie ma, a _data z koła jest jedyną i aktualną kopią.
Projekty zbootstrapowane przed tą zmianą mają w mcp.json stare uvx --from albo
uv run --directory — wystarczy kit-ai reload.
Skąd wiedzieć kiedy re-bootstrapować (bez ciągłego czytania plików kita — tanie, jedno porównanie commitów):
MCP tool: check_kit_statusBootstrap zapisuje .ai/.kit-bootstrap.json (commit kita w momencie bootstrapu). check_kit_status
porównuje go z aktualnym HEAD kita (git rev-parse + git diff --name-only tylko na ścieżkach
które bootstrap faktycznie kopiuje) i zwraca: aktualny / zmienił się (+ lista plików) / brak stampu
(stary bootstrap sprzed tej funkcji) / brak lokalnej historii git (gdy --from to zdalny URL, nie
lokalny klon). Zero kosztu tokenów na nawigację plików — jedno wywołanie tool, agent woła je kiedy
chce sprawdzić stan (np. na początku sesji), nie w pętli.
Gdy pokaże zmiany: bootstrap-project.sh ponownie z tymi samymi flagami co poprzednio.
Slash commands — konwencja nazw
Prefiks | Rola | Przykłady |
| Cursor only — alias UI Summarize w tym projekcie |
|
| Start / sync issue / commit / PR |
|
| Pomysł → ocena na tle repo → issue (bez brancha) |
|
| Pomysł na skill → skill czy agent → issue (bez brancha) |
|
| Konfiguracja projektu po |
|
| Jedna zmiana konfiguracji / odstępstwo od modułu |
|
| Review tylko do odczytu, raport |
|
| Praca w dwóch oknach (wymiana raportów) |
|
| Nocna praca na liście issue pod |
|
/compact (wyłącznie Cursor)
Skill: templates/cursor/skills/compact/SKILL.md → tylko .cursor/skills/compact/.
Bootstrap nie kopiuje tego do Claude / Codex. Nie nadpisuje ani nie „tłumaczy” ich wbudowanego /compact.
Po co: w Cursorze jedna komenda
/compactzamiast szukania UI Summarize.Nie jest wspólną konwencją kita cross-tool.
Nie mylić z
/handoff(plik + nowy chat).
/compact/git-start, /git-check, /git-commit, /git-end + Superpowers + Autopilot
Wymaga gh + git. Konwencja: feat/42-add-cart-coupon. Pełne zasady: .cursor/rules/git-branch-pr.mdc.
Podział ról (czytelnie)
Krok | Narzędzie | Uwagi |
Scope / TDD | Matt |
|
Issue + branch |
| Numeracja issue, Conventional name |
Sync issue ↔ diff |
| Gdy tytuł/body rozjechały się z plikami |
Commit(y) |
| Conventional; |
Izolacja (opc.) | Superpowers | Na branchu z |
Review przed pushem |
| Nie wszystkie |
Push + PR |
| Jedno z dwóch. |
CI / komentarze aż green | Autopilot | Po istniejącym PR; bez auto-merge |
Merge | Ty / | Gdy green → GitHub zamyka issue ( |
Krótko: /git-start → kod → [/git-check] → /git-commit → /review-bugbot → /git-end → [Autopilot]
Długo: [/grill-me] → /git-start → worktree → kod → [/git-check] → /git-commit → finishing| /git-end → Autopilot → mergeKomenda kit | Co robi |
| Ocena pomysłu na tle repo → karta issue → utworzenie po akceptacji; |
| Rozstrzyga skill vs agent, potem karta issue z nazwą, |
| Pytania z MCP |
| (A) jedna odpowiedź: Stack per Tier, codegen, klienci, język, sekcja |
|
|
| Dopasuj tytuł (EN) i body (język MCP) issue do realnego diffa; |
| Conventional Commit(s) z diffa; |
| Push + PR z |
| Cursor only — skrót czatu (alias UI Summarize); nie Claude/Codex |
/git-start --help
/git-start feat add cart coupon
/git-start fix #108 login returns 500
/git-start # auto z lokalnego diffa
# … praca zmieniła scope …
/git-check
/git-commit # lub --one / --split / --dry-run
# … review …
/git-end --help
/git-endRęczny odpowiednik:
gh issue create --title "Add cart coupon" --body "…"
gh issue develop 42 --name feat/42-add-cart-coupon --base dev --checkout
# … praca …
git push -u origin HEAD
gh pr create --base dev --title "feat: add cart coupon" --body "Closes #42"UI: GitHub Issue → Development → Create a branch (potem nazwij spójnie typ/N-slug).
Slash | Plik szablonu |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| skille Cursor (user/global), nie ten kit |
/teacher-* — tryb nauki (przed kodem, nie po)
/review-* sprawdza gotowy diff i zwraca tabelę findingów. /teacher-* działa zanim napiszesz kod: bierze Twoją koncepcję (albo bieżący diff, gdy nie podasz argumentu), tłumaczy o co w problemie naprawdę chodzi, pokazuje max 3 opcje z kosztami, wskazuje jedną rekomendację i zostawia Ci zadanie do zrobienia samodzielnie.
Komenda | Zakres |
| Django/DRF (opc. FastAPI, Flask+Pydantic): warstwy, modele, migracje, transakcje, Celery, ACL, pytest, uv/ruff |
| React, React Native/Expo Router (opc. Angular): stan serwera vs klienta, granice komponentów, re-rendery, web/native, typy TS, RTL/Playwright |
| granice FE/BE, kontrakt API, kiedy nie dzielić, infra (Postgres/Redis/Celery/S3), Docker+Taskfile, odwracalność decyzji, ADR |
| praca z samymi agentami: skille, worktree i izolacja zadań, delegacja do subagentów, autonomia ( |
Kontrakt tych agentów: readonly — nie edytują plików, nie dają gotowca do wklejenia (szkic ≤ 20 linii), nazywają wzorce po imieniu i mówią wprost, gdy koncepcja jest zła. Czytają get_bundle + get_overlay, więc uczą na Twoim stacku i Twoim kodzie, nie na Foo/Bar.
/teacher-backend czy walidację ceny dać do serializera czy do serwisu
/teacher-frontend # bez argumentu → uczy o tym, co masz w git diff
/teacher-architecture czy dodać Redisa pod cache koszyka/night-run — lista issue przez noc
Za dnia grillujesz issue (kryteria akceptacji, relacje blocked-by). W nocy /goal pilnuje pętli, a /night-run daje procedurę: na każdy ticket /git-start → test-first → szybkie bramki → /git-commit → review na diffie → /git-end → CI → merge → jeden raport na PR i zamknięcie issue. Problem zamiast pytania kończy się komentarzem needs-human z pytaniami Q1/Q2 na issue i agent idzie dalej. Pełna procedura: templates/shared/agents/night-run.md.
Agent jest orkiestratorem w głównej sesji (w Claude przez Skill, nie jako subagent). Sam nie czyta kodu: na ticket odpala świeżego subagenta ticketu, a review robi osobny świeży subagent na samym git diff. Dzięki temu żaden kontekst nie puchnie do 200k, a każda tura nie czyta go od nowa. Niczego nie dopisujesz do overlay:
Gałąź bazowa:
dev, jeśliorigin/devistnieje i nie jest w tyle za gałęzią domyślną; inaczej gałąź domyślna repo. Porzuconydevnie przejmie nocy.Plik kontekstu nocy (
/tmp/night-run-<repo>-<data>/context.md): mapa aplikacji, konwencje, pułapki toolchainu z pamięci projektu i overlay. Po każdym tickecie orkiestrator dopisuje, co doszło (modele, serwisy, endpointy). Subagent czyta ten plik zamiast AGENTS.md, BUGBOT.md i wszystkich ADR-ów.Bramki jakości: kroki
run:z.github/workflows/*.yml+ sekcja kontroli z.ai/project.md(icodegen:). Lokalnie tylko szybkie (lint, typecheck,makemigrations --check, testy dotknięte ticketem); pełny zestaw testów tylko w CI.Model subagenta ticketu:
model: <nazwa>w tekście celu; brak → model sesji.Koszt: po każdym tickecie snippet
python3liczy z transkryptów subagentów tury, tokeny (input / cache_creation / cache_read / output), maks. kontekst i czas. Wynik trafia doNIGHT-RUN REPORT.Wybrana baza, bramki, model i każde założenie trafiają do
NIGHT-RUN REPORT.
Sędzia /goal widzi tylko transkrypt, więc warunek żąda dowodów w rozmowie:
/goal Wykonaj issue #150–#157 wg /night-run, model: sonnet. Koniec, gdy w transkrypcie jest
NIGHT-RUN REPORT, w którym każdy ticket ma: MERGED (wynik gh pr view --json state)
albo needs-human (link do komentarza), albo jest wpis "night-run halted".Uwagi dopisujesz za warunkiem („#155 bez PDF”, „bez merge, same PR-y”, „model: sonnet”) — polecenia z celu mają pierwszeństwo przed procedurą. W Claude model: przyjmuje tylko aliasy (sonnet, opus, haiku, fable); konkretną wersję modelu wybierasz dla całej sesji: claude --model <id>.
Pomiar kosztu ticketu na różnych modelach
Tańszy token nie znaczy tańszy ticket: mocniejszy model może zrobić mniej tur i mniej poprawek, a koszt nocy to głównie ponowne czytanie kontekstu (cache_read). Porównanie robisz tak:
Wybierz jeden ticket średniej wielkości (kilka kryteriów akceptacji, jedna aplikacja backendu), bez decyzji o pieniądzach i zgodach, żeby needs-human nie zepsuł porównania. Zapisz commit bazowy:
git rev-parse origin/<BASE>.Na każdy model osobny worktree z tego commitu i osobna sesja z dokładnym id modelu:
git worktree add ../measure-<model> <commit> cd ../measure-<model> && claude -p --model <id> "/night-run #<N>, bez merge"PR służy tylko do pomiaru: po zebraniu wyników
gh pr close <PR> --delete-branchigit worktree remove ../measure-<model>.Metryki: snippet z sekcji „Pomiar kosztu ticketu” w agencie, uruchomiony na transkryptach sesji i jej subagentów (
~/.claude/projects/<projekt>/<sesja>.jsonli…/<sesja>/subagents/*.jsonl). Koszt liczysz osobno dla input, cache_creation, cache_read i output według aktualnego cennika, nie jedną stawką.Jakość: CI zielone za pierwszym razem (t/n), liczba rund poprawek, potwierdzone findingi review, needs-human (t/n).
Limit: ile ticketów mieści się w jednym oknie limitu sesji. Okno nie jest publiczne, więc szacujesz: zużycie na ticket w stosunku do zużycia skumulowanego w chwili HTTP 429 we wcześniejszym przebiegu.
Wynik (tabela + rekomendacja modelu domyślnego) trafia do issue pomiaru; zmiana domyślnego modelu to jedna linijka w agencie.
Bootstrap (--clients) kopiuje/renderuje shared agents do natywnych ścieżek każdego klienta. Format i mechanizm różnią się per klient:
Cursor:
.cursor/agents/— natywne slash commands, działa 1:1.Claude Code:
.claude/agents/(subagenty, wywołanie przez Task/Agent tool) oraz.claude/commands/(prawdziwe slash commands/git-startitd. —$ARGUMENTSwstrzyknięty automatycznie przy kopiowaniu).Codex: agenty instalowane jako natywne skille w
.codex/skills/<nazwa>/(renderowane ztemplates/shared/agents/*.mdprzezscripts/install_shared_skills.py). Custom prompts (.codex/agents/*.toml) zostały wycofane w Codex CLI — Codex sam ładuje SKILL.md, gdydescriptionpasuje do sytuacji.Kiro:
.kiro/agents/— kopiowane 1:1, format niezweryfikowany na żywym Kiro.VS Code/Copilot:
scripts/render_agent_commands.py vscode→.github/prompts/*.prompt.md(wywołanie/nazwaw Copilot Chat).Kilo:
scripts/render_agent_commands.py kilo→.kilocode/workflows/*.md(wywołanie/nazwa,$ARGUMENTSwspierane).Antigravity:
scripts/render_agent_commands.py antigravity→.agents/workflows/*.md(wywołanie/nazwa; limit 12 000 znaków/plik, kit przycina jeśli trzeba).opencode:
scripts/render_agent_commands.py opencode→.opencode/command/*.md(wywołanie/nazwa,$ARGUMENTSwspierane). Do tego/goali/loopjak w Claude Code:templates/opencode/command/{goal,loop}.md+ plugin.opencode/plugins/kit-loop.js, który posession.execution.succeededwysyła kolejną turę. Stop:<promise>DONE</promise>w odpowiedzi, Esc,/goal clear//loop stop, limit tur (goal 25, loop 10,max=N);/loop 5m <zadanie>powtarza co interwał.
Formaty VS Code/Kilo/Antigravity/opencode oparte o publiczną dokumentację tych klientów (sierpień 2026) — nie testowane na żywych instalacjach; jeśli coś nie zadziała, zgłoś różnicę i popraw scripts/render_agent_commands.py.
Po skopiowaniu/wyrenderowaniu zrestartuj okno IDE — agenty/komendy ładują się przy starcie.
Wywołanie
/git-start feat #42 cart coupon # lub bez # — utworzy issue
/git-check # gdy diff rozjechał się z opisem issue
/git-commit # Conventional Commit(s)
/review-backend przejrzyj zmiany w backend/apps/products/
/git-end/subagent-backend przejrzyj zmiany… # potem wklej raport do /subagent-frontend w drugim oknieSkille kita — wspólne źródło
Skill to wiedza, którą model ładuje sam, gdy description pasuje do sytuacji —
w odróżnieniu od agenta (/nazwa), którego ktoś musi wywołać. Jedno źródło:
templates/shared/skills/<nazwa>/SKILL.md (+ opcjonalne references/, scripts/,
assets/). Rozkłada je scripts/install_shared_skills.py.
Klient | Gdzie ląduje | Jak działa |
claude |
| natywnie, z zasobami |
cursor |
| natywnie, z zasobami (obok Cursor-only |
antigravity |
| natywnie, z zasobami |
codex |
| natywnie, z zasobami |
vscode |
| degradacja: komenda |
kiro |
| degradacja: komenda |
kilo |
| degradacja: komenda |
opencode |
| degradacja: komenda |
Degradacja kosztuje dwie rzeczy: skill przestaje odpalać się sam (trzeba wpisać
/nazwa) i gubi wszystko poza SKILL.md, bo komenda to jeden plik. Instalator mówi
o gubionych katalogach na stderr. Skill, którego sens leży w scripts/, będzie
w pięciu na osiem klientów wydmuszką — wtedy to prawdopodobnie powinien być agent.
.claude/skills/ i .agents/skills/ dzielisz ze skillami spoza kita (npx skills add),
więc odznaczenie klienta kasuje tam tylko katalogi o nazwach ze wspólnego źródła,
nigdy całego katalogu skilli.
Nowy skill zakładasz przez /create-skill (issue), a piszesz według skilla
skill-authoring — to on trzyma zasady frontmatter, sufity długości i kryteria odpalania.
Guardrails — bezpieczeństwo
Jedno źródło polityki: templates/shared/guards/. Bootstrap kopiuje je do katalogu
hooków wybranego klienta (--clients), więc Cursor i Claude Code egzekwują dokładnie
te same reguły.
Guard | Klient | Zachowanie |
| Claude, Cursor | deny: |
| Claude, Cursor | deny odczyt i zapis sekretów ( |
| Claude | Tylko Windows: deny |
| Claude | PostToolUse po Edit/Write: format → lint edytowanego pliku (ruff, prettier, eslint, shellcheck, hadolint, yamllint), tylko gdy repo ma config danego narzędzia; wynik wraca do modelu jako |
| Claude | SessionStart: brak |
| Copilot |
|
Zero ask (ADR 0006): Guard odpowiada allow albo deny. W auto mode ask z hooka
blokuje tak samo jak prompt, więc bramka, która pyta, nie jest automatyczna. Model dostaje
permissionDecisionReason i sam dobiera bezpieczną alternatywę. Git odzyska wszystko
w repo; poza repo pilnujemy tylko katalogów systemowych i sekretów — resztę gate'uje
natywna permission klienta (cwd + additionalDirectories).
Jeden dialekt, adapter na brzegu. Skrypty polityki mówią wyłącznie kontraktem
Claude Code (hookSpecificOutput.permissionDecision). Cursor ma własny kształt
(permission), więc invoke-hook.js tłumaczy — i to jedyne miejsce w kicie, które
wie o różnicy między klientami.
Klient | Wywołanie | Kontrakt |
Claude Code |
| natywny, bez adaptera |
Cursor |
| tłumaczony przez adapter |
Cursor: beforeShellExecution → git-guard, beforeReadFile → sensitive-files-guard
(--tool Read), preToolUse z matcherem Write → sensitive-files-guard (--tool Write).
--tool dopisuje tool_name, którego payload Cursora nie niesie. Wszystkie wpisy mają
failClosed: true — padnięty Guard (brak JSON) blokuje akcję, a nieczytelny payload
daje deny. invoke-hook.js po wypisaniu JSON zawsze kończy exit 0 (niezerowy
exit ukrywa payload przy failClosed).
Guardy są w .mjs i idą przez node — bez basha, więc bez wykrywania Git Basha na
Windows i bez otwartych okien konsoli.
Regresja: uv run python -m unittest tests.test_guards (tabela allow/deny każdego Guarda)
i bash tests/test_guard_adapter.sh (tłumaczenie kontraktu) — odpalane też przez CI
(tests/test_shell_suites.py wciąga suity powłoki do unittest discover).
Code review (Bugbot + GitHub)
Moduł MCP: core:code-review (bundle devops lub architecture).
Minimalny zestaw przed pushem (nie odpalaj całego wachlarza):
Zmiana | Minimum |
Drobna |
|
Backend / Frontend | Bugbot + |
API + UI | Bugbot + BE+FE lub para |
Auth / płatności |
|
Dowód „działa” |
|
Bugbot = blocking/security. Stack /review-* = konwencje z MCP (Severity | Location | Finding | Fix).
Przy codegen: orval w overlay — po zmianie API regeneruj klienta.
Warstwa | Plik / akcja |
Lokalnie |
|
Przed push |
|
Na PR | Bugbot (GitHub integration) |
Reguły |
|
CI (ten kit) |
|
Hook regresja |
|
Suity powłoki w CI |
|
Zależności Python (pin majora)
mcp>=1.0.0,<2 # FastMCP (1.x); mcp 2.0 usuwa mcp.server.fastmcp
pyyaml>=6.0,<7uvx resolvuje zależności od zera (nie bierze lokalnego uv.lock) — upper bound chroni konsumentów przed breaking major.
Skills / pluginy zewnętrzne (poza tym kitem)
Trzy warstwy — nie bundluj Matt/Superpowers w guides-mcp:
Warstwa | Przykłady | Gdzie | Rola |
Fundament | Context7, | MCP + agents/skills z bootstrap | stack, git, skrót czatu (Cursor) |
Proces |
|
| |
Meta | superpowers, caveman, Autopilot | user / plugin Cursor | worktree, finishing, CI loop |
Priorytet w AGENTS.md: użytkownik → overlay+MCP → review kita → Matt → Superpowers.
TDD: jeden path na feature (preferuj Matt). Setup Matt: po instalacji uruchom /setup-matt-pocock-skills.
Context7 (docs Django/Expo): globalnie npx ctx7 setup --cursor.
Subagenty — szczegóły
Każdy plik agentów jest cienkim wrapperem: przy starcie woła get_bundle / get_overlay z MCP project-guides. Wiedza merytoryczna żyje w modules/.
Praca w dwóch oknach: /subagent-backend ↔ /subagent-frontend — sekcja „Raport do przekazania” na końcu odpowiedzi.
Available Tools
12 toolsbootstrap_workspaceA
Zainstaluj pliki kita (hooki, agenci, komendy, mcp.json, stamp) w repo aplikacji.
Jedyne narzędzie MCP, które zapisuje na dysk — reszta serwera jest tylko do
odczytu. Bez tego trzeba sklonować kit lokalnie i ręcznie odpalić
scripts/bootstrap-project.sh; tu ten sam skrypt uruchamia serwer, który już
ma wszystkie szablony pod ręką.
dry_run=True jest domyślne i nic nie zapisuje: skrypt leci na kopii kitowej
powierzchni repo w katalogu tymczasowym, a wynikiem jest lista plików, które
powstałyby, zostałyby nadpisane albo usunięte. Zapis wymaga jawnego
dry_run=False — hooki PreToolUse łapią Bash/Edit/Write,
a nie nazwy narzędzi MCP, więc ta domyślka jest tu jedyną bramką.
Args:
clients: --clients: all | cursor | claude | codex | vscode | kiro | kilo |
antigravity | opencode (po przecinku). Domyślnie: wartość startowa serwera.
preset: Kategoria presetu (_base, shop). Domyślnie: preset serwera.
language: Język prozy pl/en. Domyślnie: język bieżącego profilu.
codegen: orval/none/graphql. Domyślnie: codegen bieżącego profilu.
with_overlay: Skopiuj szablon .ai/project.md, jeśli repo go nie ma.
keep_unselected_clients: Nie usuwaj kitowych plików klientów spoza clients.
dry_run: True (domyślnie) — tylko plan. False — faktyczny zapis.
Returns: str: Markdown — plan zmian (dry run) albo raport z instalacji i stampu.
| Name | Required | Description | Default |
|---|---|---|---|
| preset | No | ||
| clients | No | ||
| codegen | No | ||
| dry_run | No | ||
| language | No | ||
| with_overlay | No | ||
| keep_unselected_clients | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and fully delivers: it discloses that this is the only mutating tool, that dry runs execute on a temporary copy of the repo surface, that the output lists files that would be created, overwritten, or deleted, and that real writes are destructive and update the kit stamp. The risk profile is surfaced, not hidden.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but appropriately so for a high-risk mutation tool with 7 parameters; the critical write-vs-read-only distinction is front-loaded and the Args/Returns structure aids scanning. Minor redundancy: dry_run default semantics are fully explained in the third paragraph and then restated in the Args block.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex mutating tool with 7 optional parameters, no annotations, and only a string output schema, the description covers everything needed to call it correctly: purpose, side effects, the safety gate, all parameter semantics, and the Markdown return format. An agent could invoke it safely and correctly with no additional information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description fully compensates: all 7 parameters receive value domains (clients list, preset categories, language choices, codegen options), default sources (server start value, profile setting, or fixed default), and behavioral meaning (with_overlay copies, keep_unselected_clients prevents deletion). Nothing is left to inference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific verb and resource: install kit files (hooks, agents, commands, mcp.json, stamp) into the app repo. It also explicitly differentiates from all siblings by declaring this is the only MCP tool that writes to disk while the rest of the server is read-only, making selection unambiguous without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the manual alternative (cloning the kit and running scripts/bootstrap-project.sh) and explains why this tool supersedes it. It also prescribes the safe workflow: dry_run=True is the default planning mode, actual writes require explicit dry_run=False, and it warns that PreToolUse hooks catch Bash/Edit/Write rather than MCP tool names — clear, actionable guidance on when and how to invoke.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_kit_statusA
Sprawdź czy kit zmienił się od ostatniego bootstrap-project.sh w tym repo.
Tanie: porównuje commit zapisany w .ai/.kit-bootstrap.json z aktualnym HEAD
kita (git rev-parse/diff --name-only) — nie czyta treści modułów instrukcji.
Moduły (modules/*.md) są i tak czytane live przez get_bundle/get_overlay,
więc nigdy nie "gniją" — ten tool dotyczy tylko plików które bootstrap
kopiuje (agents/commands/mcp.json), bo te są statyczną migawką.
Returns: str: Markdown — aktualny / zmienił się (+ lista plików które re-bootstrap by nadpisał) / brak stampu / brak lokalnej historii git do porównania.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the mechanism (git rev-parse/diff --name-only), the fact that it does not read module contents, and that it only concerns static snapshot files. It also lists the return cases: current/changed (+list of files), no stamp, no local git history. This is thorough behavioral disclosure, including limitations and edge cases.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Although the description is somewhat long, every sentence serves a purpose: it states the core check, the cheap method, the scope exclusion (modules), and the return format. It is front-loaded with the main purpose and then details. No redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no parameters, an output schema exists, and the description covers the behavior, the return format (Markdown with status), and edge cases (no stamp, no history). It is complete for an agent to understand what it does and what to expect. The description also clarifies the distinction from sibling tools, so nothing is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline is 4. The description does not need to explain parameters, and the schema coverage is 100% (vacuously). No additional parameter semantics are required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Sprawdź czy kit zmienił się od ostatniego bootstrap-project.sh' (Check if the kit changed since the last bootstrap-project.sh). It specifies the resource (kit) and the action (check status), and distinguishes itself from siblings by explicitly noting that modules are handled live by get_bundle/get_overlay, so this tool only concerns files that bootstrap copies. This is a specific verb+resource with clear differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the context of use: it's a cheap check comparing a stored commit to HEAD, and it clarifies that module files are not covered because they are read live elsewhere. This implies when to use this tool (to decide if re-bootstrap is needed) and when not to (for modules). It doesn't explicitly name alternatives but does reference get_bundle/get_overlay, providing enough guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_bundleA
Pobierz pełną treść bundle'a instrukcji.
Args: name: Nazwa bundle'a, np. backend, frontend, architecture, full.
Returns: str: Połączona treść Markdown wszystkich modułów w bundle'u + overlay (jeśli dotyczy).
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does disclose the key behavior: it concatenates all module Markdown and appends an overlay if present. It also signals a read operation via 'Pobierz'. It does not discuss failure modes or permissions, but for a simple fetch tool that is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loads the core purpose, and uses a clear Args/Returns structure. Every sentence adds information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-required-parameter read tool with an output schema, the description covers purpose, parameter meaning, and return format. It omits exact allowed bundle names and error behavior, but these are not critical for a straightforward fetch by name.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the schema only labels the property 'Name', so the description compensates by defining name as the bundle identifier and giving concrete examples ('backend', 'frontend', 'architecture', 'full'). It could be stronger with an exhaustive enum or validation rules, but the compensation is sufficient for correct invocation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Pobierz' – fetch) and a precise resource ('pełną treść bundle'a instrukcji'), and clarifies that the result is the combined Markdown of all modules plus optional overlay. This distinguishes it from siblings like get_module (single module) and list_bundles (listing bundles).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this tool is for retrieving the full assembled content of a bundle and gives name examples, but it does not explicitly state when to prefer it over get_module or get_overlay, nor does it mention exclusions. Usage context is left largely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_clientsA
Lista klientów AI skonfigurowanych przy starcie MCP (--clients / GUIDES_CLIENTS).
Metadane instalacji szablonów — nie zmieniają treści bundle.
Returns: str: Markdown z wartością flagi i rozwiniętą listą id klientów.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses the return type (Markdown string with flag value and client IDs) and explicitly states it does not modify bundle content, implying a read-only operation. It does not cover edge cases like errors or auth, but for a simple getter this is adequate and adds meaningful behavior context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise: one sentence states the purpose, a second clarifies scope, and a return line specifies output. Every sentence earns its place, and the primary purpose is front-loaded. No fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with an output schema (the return type is described), the description is fully adequate. It tells the agent what the tool does, what it returns, and that it is side-effect-free. Nothing essential is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema already covers everything trivially. The description adds no parameter details, but that is not a gap since there are none. The baseline of 4 applies because the description could have mentioned that no parameters are needed, but it is not necessary given the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear, specific purpose: listing AI clients configured at MCP startup, with the flag and environment variable named. It also clarifies it deals with template installation metadata and explicitly says it does not alter bundle content, distinguishing it from content-related tools. The verb 'Lista' (list) and resource 'klientów AI' make the action unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context by noting it covers metadata and does not change bundle content, which hints at when to use it (for client configuration info). However, it does not explicitly mention when not to use it or name alternative tools, and no comparison to siblings is provided. The guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_codegenA
Aktualny wybór generatora klienta API (--codegen / codegen: w profilu).
Returns:
str: Markdown — orval (schema → frontend/src/api/generated + mutatory),
none (tool-agnostyczny klient, konkret w overlay projektu) albo
graphql (GraphQL zamiast REST, patrz arch:api-contract:graphql).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must disclose behavior. It only states the return value and its possible values, but does not mention whether the operation is read-only, if it can error, or any prerequisites. It does not explicitly state it is a safe getter, leaving behavioral aspects undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, front-loaded with the purpose, and uses a structured 'Returns' block to enumerate possible values. It contains no fluff and is appropriately sized for a simple getter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no parameters and an output schema is present. The description explains the semantic meaning of each return value, which is valuable context beyond the schema. It could mention default behavior or error handling, but for a simple getter, it is reasonably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, and schema coverage is 100% (since there are no parameters to document). The description adds value by explaining the meaning of the return values, which is beyond schema requirements. Baseline for no parameters is 4, and the description meets that.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool returns the current API client generator selection ('Aktualny wybór generatora klienta API') and lists the possible return values (orval, none, graphql) with their meanings. It distinguishes itself from sibling getters by focusing on this specific configuration.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is used to check the current codegen setting but does not explicitly state when to use it vs. alternatives. It lacks exclusions or context about typical invocation (e.g., before generating clients). The purpose is clear enough for basic use, but no explicit usage guidance is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_indexB
Indeks włączonych modułów i bundle'i dla bieżącego profilu.
Returns: str: Markdown z podsumowaniem konfiguracji profilu.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description must carry the behavioral disclosure burden. It does disclose that the return value is a Markdown string, but it says nothing about side effects, read-only behavior, prerequisite profile state, or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short lines with no filler. The main purpose is front-loaded and the return type is stated directly, which is ideal for a tool of this simplicity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that this is a zero-parameter getter with an output schema, the description covers the essential scope and return format. It would benefit from explicit usage guidance or a read-only statement, but the low complexity makes it sufficiently complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the no-parameter baseline is 4. The description appropriately avoids inventing parameter details and focuses on what the tool returns.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource ('enabled modules and bundles') and a clear scope ('current profile'), so an agent understands what the tool returns. It does not explicitly contrast itself with siblings like list_modules or list_bundles, but the 'enabled ... for current profile' phrasing narrows the purpose enough to avoid major ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as list_modules, list_bundles, get_bundle, or get_overlay. 'For current profile' implies a profile-specific use case, but no conditions, exclusions, or alternative recommendations are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_languageA
Aktualny język instrukcji i polityka tytułów vs prozy.
Returns:
str: Markdown z kodem języka, modułem core:language-* i regułami EN/PL.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the burden. It discloses the return type and content, and the 'get' verb implies a read-only operation, but it does not explicitly state whether the tool has side effects, requires context, or depends on any session state. For a zero-parameter getter this is minimally acceptable but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact: one purpose sentence followed by a structured 'Returns:' line with the type and content. Every piece contributes useful information, and the purpose is front-loaded before the return detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given zero parameters and an existing output schema, the description covers what an agent needs to know to invoke the tool and interpret the result. It could be slightly more explicit about the intended use case or what 'title vs prose policy' means, but for a simple getter it is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there are no parameter semantics to document. Schema description coverage is effectively complete with an empty properties object, matching the baseline for no-parameter tools.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource ('get_language' means current instruction language) and specifies what is returned: Markdown with the language code, the `core:language-*` module, and EN/PL rules. It is specific enough to distinguish the tool from siblings like get_overlay or get_codegen, though the phrase 'tytułów vs prozy' remains somewhat ambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to call this tool versus alternatives such as get_overlay or get_codegen. It does not state contexts, exclusions, or relationships to sibling tools, so an agent must infer usage from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_moduleA
Pobierz pojedynczy moduł instrukcji po identyfikatorze.
Args: module_id: Id modułu z manifestu, np. stack:django-drf:backend-standard.
Returns: str: Treść Markdown modułu.
| Name | Required | Description | Default |
|---|---|---|---|
| module_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does add useful behavioral context by stating the return format (Markdown string) and that the ID comes from a manifest with an example. However, it does not disclose failure behavior, side effects, or availability constraints, though these are relatively minor for a simple read-style getter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured: a one-sentence purpose, then Args and Returns sections. There is no filler, and each line adds necessary information for invoking the tool correctly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity tool with one required parameter and a string return, the description is largely complete: it names the resource, explains the parameter, and specifies the output format. Missing usage context and failure semantics are gaps, but they do not prevent a correct call in the common case.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only a type and title, with 0% description coverage. The description compensates fully by explaining that module_id is an ID from the manifest and giving a concrete example (stack:django-drf:backend-standard). This makes the parameter meaning and expected format clear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('Pobierz' / fetch) and resource ('pojedynczy moduł instrukcji') by identifier, so an agent can understand the core action. It does not explicitly distinguish this tool from siblings such as get_index or get_bundle, but the 'module' resource is distinct enough to make the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to use get_module versus alternatives like list_modules, get_bundle, or get_overlay. The description only explains the mechanism (fetch by module_id), not the selection context or any exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_overlayA
Pobierz overlay projektu — unikalne instrukcje tylko dla tego repo.
Returns:
str: Treść .ai/project.md i innych overlay wskazanych w profilu.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It discloses that the tool returns the content of `.ai/project.md` and other overlays indicated in the profile, which is useful behavioral context. However, it does not mention whether the tool can fail (e.g., missing overlay), whether it reads from the local filesystem, or any side effects. It is a read operation by nature, but that is not explicitly stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded with the main purpose in the first sentence. The Returns section adds useful detail about the specific files involved. It is efficient, though the Polish language may be slightly less universally accessible than English for an AI agent, and the Returns block could be considered redundant with the first sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has no parameters and an output schema exists, the description is mostly complete for a simple fetch operation. However, it lacks context about failure modes (e.g., what happens if `.ai/project.md` does not exist), whether the overlay is loaded from the current workspace, and how the returned string is formatted. These gaps are minor for a zero-parameter read tool but still leave some ambiguity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema provides no parameter semantics to rely on. The description compensates by explaining what the tool returns and where the content comes from (`.ai/project.md` and profile-indicated overlays). With 0 params, a baseline of 4 is appropriate because the description clarifies the tool's input-free nature and its output source.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Pobierz' = download/fetch) and resource ('overlay projektu' = project overlay), and clarifies it contains unique instructions for this repo. It distinguishes itself from siblings by specifying it retrieves `.ai/project.md` and other overlays indicated in the profile, which is a clear purpose even though it doesn't explicitly name a sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it: when you need the project-specific overlay instructions. It does not explicitly state when not to use it or name alternatives among siblings like get_language or get_codegen. The context is clear enough for an agent to infer usage, but there is no explicit routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_bundlesA
Lista dostępnych bundle'i instrukcji dla bieżącego profilu projektu.
Returns: str: Markdown z listą bundle'i i liczbą modułów.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral burden. It discloses that the return value is a Markdown string listing bundles and module counts, which adds useful information. It does not discuss side effects or errors, but 'list' implies a read-only operation and the tool has no parameters.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short, purposeful lines: an action sentence followed by a return-value line. There is no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless list tool with an output schema, this description is complete: it names the resource, the scope, and the return format. Nothing essential is missing for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema is empty with zero parameters, so the baseline is 4. There are no parameter semantics to explain, and the description still adds the useful implicit context of 'current project profile'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Lista' / list), a clear resource ('bundle'i instrukcji'), and a scope ('dla bieżącego profilu projektu'). It distinguishes itself from siblings like get_bundle (which implies retrieving one specific bundle) and list_modules (which targets a different resource).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it: when you need the list of bundles for the current project profile. However, it does not explicitly mention alternatives or when not to use it, leaving sibling differentiation to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_modulesB
Lista wszystkich modułów dostępnych w instruction-kit (manifest).
Returns: str: Markdown z identyfikatorami i tytułami modułów.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the return type (Markdown string with identifiers and titles) but does not state read-only nature, side effects, auth requirements, or any failure behavior. Minimal behavioral detail for a tool with no annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with no fluff. The purpose is front-loaded, and the return type is stated clearly. Perfectly sized for a simple list tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 0-parameter list tool with an output schema present, the description covers what it returns. It doesn't explain the concept of 'instruction-kit' or the manifest, but that context may be understood from the tool family. Slightly lacking in contextual depth but adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the description need not explain parameter meanings. The schema covers 100% with no parameters, and the description adds nothing about parameters, but per rubric 0 params baseline is 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Lista wszystkich modułów') and resource ('instruction-kit manifest'), clearly indicating it retrieves all modules. It distinguishes from siblings like get_module by specifying 'all modules', though it doesn't explicitly name alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives like get_module, list_presets, or list_bundles. There are no prerequisites or conditions provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_presetsA
Lista dostępnych presetów w instruction-kit (profiles/*.yaml).
Returns:
str: Markdown z kategoriami (_base, shop, …).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations to carry the behavioral burden, so the description must disclose traits itself. It partially does so by stating the return format (Markdown with categories such as '_base', 'shop'), and listing is naturally read-only. However, it does not mention whether it scans nested directories, errors on missing profiles, or reflects current workspace state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two purpose-built lines: one stating scope and source, one stating return type and content. It is front-loaded with the main action and contains no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless listing tool this is nearly complete: it names the resource, the source path, and the returned Markdown format. It lacks explicit mention of failure behavior or environment context, but the scope is simple enough that the gap is minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and no required parameters, so there is nothing for the description to clarify about arguments. The 100% schema coverage on an empty properties object supports assigning the baseline of 4 for no-parameter tools.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Lista dostępnych presetów' = list available presets), identifies the exact resource location ('profiles/*.yaml' in instruction-kit), and differentiates from sibling listing tools like list_modules and list_bundles by naming the preset domain. There is no ambiguity about what the tool does or returns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The purpose implicitly tells an agent to use this tool when it needs available presets in profiles/*.yaml, but the description gives no explicit guidance about when not to use it or how it compares to sibling list tools. With zero parameters and a narrow scope this is minimally adequate but leaves routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
12 tool updates
v0.1.0- First observed
bootstrap_workspace - First observed
check_kit_status - First observed
get_bundle - First observed
get_clients - First observed
get_codegen - First observed
get_index - First observed
get_language - First observed
get_module - First observed
get_overlay - First observed
list_bundles - First observed
list_modules - First observed
list_presets
TDQS
Scored across 12 tools
Each tool has a clearly distinct target: get_* returns a specific configuration piece, list_* enumerates available options, check_kit_status compares state, and bootstrap_workspace is the only write operation. Even the overlapping get_overlay and get_bundle are disambiguated by descriptions: overlay-only vs combined bundle.
All tool names follow a consistent snake_case verb_noun pattern with predictable prefixes: get_, list_, check_, bootstrap_. The naming makes the tool surface easy to scan and select from.
12 tools is well within the ideal scope for a configuration/instruction server. Every tool covers a distinct need—reading project guides, listing modules/presets/bundles/clients, checking kit status, and bootstrapping the workspace—with no obvious filler.
The tool surface covers the full lifecycle for this domain: read current project context, discover available content, inspect kit state, and perform the only intended mutation (bootstrap install with dry-run). No critical dead-end operations are missing.
Maintenance
Related MCP Connectors
Read-only AI project discovery, verification, comparison, shortlisting, and stack planning.
Project memory for coding agents: requirements, decisions, code graph and delivery telemetry.
Durable, shareable and governed project memory with smart triage and explicit project composition.
Provide your AI coding tools with token-efficient access to up-to-date technical documentation for…
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceProvides AI agents with professional coding standards, development best practices, and context-aware guidance through static documentation and AI-powered custom recommendations. Enables agents to access comprehensive development guidelines including coding rules, debugging techniques, and AI steering instructions.-
- AlicenseBqualityCmaintenanceProvides centralized security instructions for AI-assisted code generation by matching context-aware rules to the user's programming language and file patterns. It ensures generated code adheres to security best practices without requiring manual maintenance of instruction files across individual repositories.28 npm1MIT
- AlicenseAqualityDmaintenanceManages project standards, configurations, and API debugging for AI-assisted development, ensuring unified development practices across teams and machines.1320 npm5MIT
- FlicenseNot gradedqualityDmaintenanceProvides AI agents with structured access to project conventions, technology stacks, and architectural patterns to ensure consistency across development teams.1-