Gemmein MCP Server
OfficialThe Gemmein MCP server gives a coding agent the full Gemmein build contract as tools — docs, rule/error/relay explainers, a collection-name validator, a CI harness, and a live integration checker.
guide — the full builder's guide (llms.txt): fit assessment and verdicts (FITS / FITS EXCEPT / DOESN'T FIT), auth flow, seven collection safety rules, record shapes, links/expand, uploads, contention, payments, drafts, keys, sync, go-live, relays, AI tools, credits. Call first, before any install or code.
reference — every
@gemmein/sdkmethod, exact signature, return shape, and stable error-code table; use while writing code.search_docs — targeted search over the guide and reference, returning matching passages with 3 lines of context (max 6 blocks per doc) for mid-build questions.
explain_rule — one of the seven rules (private, shared, admin_write, public_read, community, addressed, direct) explained: access contract, what it's right for, mistakes that leak data; no argument returns the all-seven cheat-sheet plus what no rule supports.
explain_error — what a
GemmeinErrorcode means and the exact next step, including whether the refusal is final; no code lists all stable codes.validate_collection_name — checks a planned collection name against the naming law (lowercase letters/numbers/underscores, starts with a letter, 2–63 chars) and suggests a fix if invalid.
reaffirm_template — fetches the ready-to-edit
reaffirm.mjsCI harness that re-proves app boundaries against live Gemmein on every deploy.explain_relay — validates a
gemmein/relays/<name>.jsondefinition offline, returning the English sentence the dashboard would show or the single refusal naming the field and fix.check_integration — runs live boundary checks against an app and returns structured pass/fail: Tier A (public
pk_key) validates the collection name and anonymous read/write refusals; Tier B (addsk_dev) proves cross-user isolation with throwaway dev test sessions; the only writes are Tier B probe records, deleted afterwards (sk_liveis refused by design).
gemmein
The distribution home of the Gemmein CLI — the
gemmein package on npm, and the release record of the local engine it runs.
Gemmein is the trustworthy backend for AI-built apps: accounts, data,
payments — with the boundaries enforced by the platform, not by your app's
code. The CLI brings that to your machine: npx gemmein init shapes your
product's boundaries in plain language, and gemmein dev runs a real local
backend that enforces them while you build.
MCP server
@gemmein/mcp gives a coding agent the whole Gemmein contract as tools — the
builder guide, the SDK reference, rule and error explainers, and a live check
of an app's access boundaries. Eight tools are read-only. check_integration
runs live checks against your app; with a development secret key it also
writes in the development environment: it creates two test people and a
probe record, deletes the record, and signs the test people out of earlier
sessions.
Install for Claude Code:
claude mcp add gemmein -- npx -y @gemmein/mcpInstall for Claude Desktop, Cursor, or any client that takes an mcpServers
block:
{
"mcpServers": {
"gemmein": {
"command": "npx",
"args": ["-y", "@gemmein/mcp"]
}
}
}No environment variables are required to run it.
Tools it serves:
guide— the full builder's guide, covering auth flow, the seven safety rules, record shapes, links, uploads, contention, payments.reference— every SDK method, exact signature, return shape, and error code.search_docs— targeted search over both, when it needs one fact.explain_rule— any rule's contract, what it's right for, and the mistakes to avoid, or a cheat-sheet of all seven at planning time.explain_error— what aGemmeinErrorcode means and exactly what to do about it.validate_collection_name— catches a misnamed collection before every read starts returning empty results.explain_relay— agemmein/relays/<name>.jsondefinition in, the sentence the dashboard would show out, or the one refusal naming the field; offline, over the eleven verbswrite_record,grant_access,revoke_access,grant_credits,email_person,call_url,fulfil_product,refund_product,grant_plan,revoke_planandstart_run.reaffirm_template— the CI harness, ready to copy.check_integration— runs an app's isolation and access checks live against itself and hands back structured pass/fail.
The server's full source is in mcp/ (MIT). It is a standalone
Node package: cd mcp && npm install && npm run build, then run
node dist/index.js. It does not use the Gemmein engine.
Docs: docs.gemmein.com/mcp.
Related MCP server: Convex MCP server
Claude plugin
plugins/gemmein/ is a Claude plugin that bundles the
MCP server above, pinned to an exact version, with one skill,
build-on-gemmein: it has Claude read the Gemmein guide, ask what the app
is for, say whether Gemmein fits, and then follow the first-hour path
(npx -y gemmein dev, collections, npx gemmein sync, going live). This
repository is also a plugin marketplace. In Claude Code:
/plugin marketplace add gemmeinhq/gemmein-release
/plugin install gemmein@gemmeinWhat the plugin runs, sends and stores is in
plugins/gemmein/README.md.
How the launcher works
The npm package you install is a small launcher, published as readable
source — what you see in shim/ is exactly what runs on your machine.
On first run it downloads the Gemmein engine, verifies it against SHA-256 hashes pinned inside this package, caches it locally, and runs it. After that one download, everything works fully offline. If a downloaded file ever fails verification, the launcher refuses to run it.
shim/— thegemmeinnpm package (plain JS, no build step)manifest/releases.json— every engine release and its file hashes
Licensing
Two licenses, on purpose:
The MCP server (
mcp/, published as@gemmein/mcp) is MIT.The npm launcher (
shim/) is MIT — read it, audit it, it's yours.The downloaded engine is proprietary, under the Gemmein Engine License — you can run and cache it freely for building against Gemmein (offline once cached, even if a version is retired), but not redistribute or extract it. It's also served at downloads.gemmein.com/engine/LICENSE.
Using it
npx gemmein initDocs: docs.gemmein.com · Questions: hello@gemmein.com
Available Tools
9 toolscheck_integrationAIdempotent
Call after wiring the app to Gemmein and before telling your human it is done — and again before go-live. Runs the reaffirm boundary checks live against the caller's own app; returns structured pass/fail (structuredContent: checks, notes, failedCount, passed). Tier A (public pk_ key only): the collection name is valid, anonymous reads and writes of a private collection are refused, an optional public collection reads as its rule intends — safe against any environment, live included. Tier B (add the sk_dev secret key): proves one user cannot read another's private records, using two throwaway test sessions in the DEV environment. sk_live is refused by design — never pass a live secret to any tool; dev and live enforce the same rules, so isolation proven in dev holds in live. The only writes anywhere are Tier B's own probe records in the caller's dev environment, deleted at the end of the check. A failed check means the app's assumptions drifted from its rules — fix before shipping.
| Name | Required | Description | Default |
|---|---|---|---|
| apiUrl | No | optional: API base URL override (local/dev API); omit for production Gemmein | |
| publicKey | Yes | the app's public pk_ key | |
| secretKey | No | optional: the sk_dev secret key — enables Tier B isolation proof (sk_live is refused) | |
| testUsers | No | optional: the two Tier-B test emails (default reaffirm-a/b@test.dev) | |
| timeoutMs | No | overall time budget, default 30000 | |
| publicCollection | No | optional: a community/public_read collection to confirm anonymous readability | |
| privateCollection | Yes | a collection with the `private` rule |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnly=false, destructive=false, idempotent=true, openWorld=true; the description adds far more: the two-tier model (pk_ vs sk_dev), that sk_live is refused by design, that the only writes are Tier B probe records in dev which are deleted at the end, and that dev/live enforce identical rules. It never states an explicit required permission scope or the exact runtime/rate profile beyond the timeout param, but the disclosure is well above the annotation baseline.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the actionable timing guidance, and each sentence carries real content (tiers, safety guarantees, failure meaning). It is dense and slightly repetitive in reasserting the sk_live prohibition, but not padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter tool with no output schema, the description covers what matters: which keys trigger which tier, what writes occur and that they are cleaned up, the dev/live equivalence argument, and the return shape. An agent has enough to call it correctly and interpret results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the description goes further by tying parameters to behavior — public pk_ enables Tier A, adding sk_dev enables Tier B, publicCollection confirms anonymous readability, and testUsers are the two Tier-B sessions. It does not add syntax/format detail beyond the schema, but the tier mapping is genuinely useful semantic context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific action (runs reaffirm boundary checks live against the caller's own app) and a concrete return shape (structured pass/fail with checks, notes, failedCount, passed). It is clearly distinguishable from the doc-oriented siblings (guide, reference, explain_rule, validate_collection_name) since it executes live checks rather than explaining concepts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit timing windows: after wiring the app and before declaring it done, and again before go-live. It also states the remediation condition ('a failed check means the app's assumptions drifted — fix before shipping'), so the agent knows exactly when to invoke and what to do with the result.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
explain_errorARead-only
Call the moment a GemmeinError reaches you (err.code: conflict, forbidden, unknown_collection, invalid_shape, html_not_allowed, …): what the code means and the exact next step — including whether the refusal is final (a forbidden repeats on retry; fix the approach, not the request). Parsed from the installed API reference, so codes match the SDK version the app runs. Call with no code to list every stable code.
| Name | Required | Description | Default |
|---|---|---|---|
| code | No | the err.code to explain; omit to list all |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only supply readOnlyHint=true, and the description adds real behavior the annotation cannot: retry finality semantics ('a forbidden repeats on retry; fix the approach, not the request') and provenance ('parsed from the installed API reference, so codes match the SDK version the app runs'). Those are the two things an agent actually needs to decide its next action after a failure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the trigger in the first clause and every sentence carries information (trigger, returned content, retry semantics, provenance, no-arg mode). The em-dash and parenthetical nesting makes it slightly dense to scan, but there is no wasted sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description does tell the agent what comes back (meaning of the code plus the exact next step) and how to enumerate all codes. Minor gaps remain: behavior for an unrecognized code string and the shape/verbosity of the explanation are not stated, and annotations carry only the read-only hint.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter and schema description coverage is 100%, so the baseline is 3. The description goes beyond the schema by naming concrete code values (conflict, forbidden, unknown_collection, invalid_shape, html_not_allowed) and by clarifying that omitting the code lists every *stable* code, adding a nuance the schema's 'list all' does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (explain) and resource (GemmeinError code from a GemmeinError object), with the exact trigger condition 'the moment a GemmeinError reaches you'. That scope cleanly separates it from the sibling explainers (explain_rule, explain_relay) and the doc-lookup tools (guide, reference, search_docs).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger ('the moment a GemmeinError reaches you') plus a second invocation mode ('call with no code to list every stable code'), which is exactly the alternative an agent needs. It stops short of a when-not-to-use or explicit routing away from sibling explain tools, so it is clear context without full exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
explain_relayARead-only
Call while WRITING or FIXING gemmein/relays/.json — before gemmein sync carries it to the cloud. A relay is one trigger (receiver: a provider's webhook; schedule: a clock; data_change: a record changing) and up to ten actions in Gemmein's own verbs (write_record, grant_access, revoke_access, grant_credits, email_person, call_url, fulfil_product, refund_product, grant_plan, revoke_plan) — Gemmein runs it: receives the event, maps the fields, authorises, carries out the action. Pass the definition JSON; the answer is the English sentence the dashboard shows ("When gocardless-paid receives an event where event_type is confirmed → grant Pro, email the person, call https://…") or the ONE refusal sentence the cloud would answer, naming the field and the fix. Offline and read-only: nothing is created. Two checks run only in the cloud and are stated in the answer (the API's own hosts; the address's resolved network at call time).
| Name | Required | Description | Default |
|---|---|---|---|
| definition | Yes | the relay definition — the contents of gemmein/relays/<name>.json ({ name, trigger, actions }) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description reinforces it ('Offline and read-only: nothing is created') while adding genuinely new behavioral context: the answer is the dashboard sentence or a single refusal, and two checks (the API's own hosts, the address's resolved network at call time) run only in the cloud. This discloses what cannot be validated locally, which is exactly what an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense paragraph, front-loaded with the call trigger and progressively filling in relay anatomy, return shape, and limitations. Every clause carries information, though the parenthetical verb list and dash-heavy construction make it heavier than strictly necessary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly explains the return value (sentence or refusal) and the offline/cloud split. Purpose, timing, inputs, outputs, and safety are all covered, so an agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so the 'definition' parameter is already documented, setting a baseline of 3. The description adds meaning by unpacking what the definition contains — trigger kinds (receiver/schedule/data_change) and the action verbs — which helps an agent construct a valid payload.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (explain) on a specific resource (gemmein/relays/<name>.json) and defines a 'relay' concretely (one trigger plus up to ten actions in named verbs). It also states exactly what the tool returns — an English sentence or the ONE refusal sentence naming the field and fix. An agent can distinguish it from siblings like explain_rule and explain_error without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear when-to-use context: 'Call while WRITING or FIXING ... before `gemmein sync` carries it to the cloud.' The condition (pre-sync validation) is explicit. It does not name alternative tools or exclusions, but no other sibling covers relays, so routing ambiguity is low.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
explain_ruleARead-only
Call while DESIGNING a collection — which rule fits this data? — or when a rule refuses something at runtime. One of the seven rules (private, shared, admin_write, public_read, community, addressed, direct) returns its exact access contract, what it is right for, and the mistakes that leak data. Call with no rule for the all-seven cheat-sheet plus the cross-cutting law, including what NO rule supports (team/group/workspace scope, per-user visibility inside a rule) — if the app needs those shapes, that is a fit gap to report to your human, never something to approximate with client-side filtering.
| Name | Required | Description | Default |
|---|---|---|---|
| rule | No | the rule to explain; omit for the all-rules cheat-sheet |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=true already establishes this as a safe read, and the description adds meaningful behavioral context: it discloses the return payload (access contract, fit, leak mistakes), enumerates the seven rules and the cross-cutting law, and describes what NO rule supports (team/group/workspace scope, per-user visibility). It stops short of stating whether the explanations are static or context-dependent, but the disclosure goes well beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is front-loaded with the call trigger and packs the payload description into two dense sentences with little wasted wording. The final clause about fit gaps is long and parenthetical-heavy, slightly blurring the structure, but each part carries actionable content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With one optional enum parameter, full schema coverage, and no output schema, the description carries the return-value burden itself and does: it spells out what each invocation returns, the all-rules fallback, and the boundary of what the rules can express. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the enum values are documented, so the baseline is 3. The description adds real meaning beyond the schema by explaining the omission behavior ('call with no rule for the all-seven cheat-sheet plus the cross-cutting law') and by enumerating the seven rule names inline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource — it explains one of seven named rules, returning that rule's 'exact access contract, what it is right for, and the mistakes that leak data', plus the no-arg cheat-sheet variant. An agent can tell exactly what it gets back and how it differs from a generic guide/reference sibling without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names both trigger conditions explicitly: 'Call while DESIGNING a collection' and 'when a rule refuses something at runtime'. It also supplies a when-not rule ('never something to approximate with client-side filtering') and routes unsupported shapes to a human instead of a workaround, so the agent knows what NOT to use this for.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
guideARead-only
Call this FIRST — before any install, account, or code — when your human asks to build an app on Gemmein, to move an existing app onto it, or whether their app can use it at all. The guide (llms.txt) opens with two doors — starting from an idea with nothing built yet, or already holding an app — and both lead to the same fit assessment: the in-scope map, the out-of-scope list (each item downgrades the verdict; none may be approximated), and the three verdicts you deliver to your human before installing anything — FITS, FITS EXCEPT , or DOESN'T FIT. After the verdict it is the full build contract: auth flow, the seven collection safety rules, record shapes, links/expand, uploads, contention patterns, payments (g.subscriptions.checkout / g.payments.buy), drafts, error philosophy, pricing. It also teaches the keys (server · CLI · sync), gemmein sync and sync --live, go-live and promotion, relays, AI tools defined on the server and run with g.ai.run, and credits.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already declares this as a safe read operation. The description adds that the guide is an llms.txt document containing a fit assessment, verdicts, and a build contract, which helps the agent know what kind of static guidance it returns. However, it does not describe return format details beyond naming llms.txt.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with a strong imperative ('Call this FIRST'), but it is a very long, dense single paragraph that enumerates many guide topics. Several clauses, such as specific payment commands and sync flags, are useful for understanding the guide's contents but could be trimmed for tool-selection purposes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only guide tool, the description is complete: it explains when to call it, what the guide covers, the fit-assessment verdicts, and the build contract topics. No output schema exists, so explaining return values is unnecessary.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so the baseline is 4. The description does not need to compensate for any parameter documentation, and the schema correctly shows an empty object with additionalProperties false.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource (the guide/llms.txt) and a clear trigger: call first before install, account, or code when building, moving, or assessing an app on Gemmein. It does not explicitly name or differentiate itself from siblings like search_docs or reference, so sibling differentiation is missing despite strong resource clarity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit usage context: call this FIRST when the human asks to build an app on Gemmein, move an existing app, or check whether an app can use it. It does not state when not to use it or name alternative tools, so it stops short of full when/when-not/alternatives guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reaffirm_templateARead-only
Fetch this when you wire up the app's CI, or when you hand the finished app to your human: reaffirm.mjs, the ready-to-edit harness that re-proves the app's boundaries against live Gemmein on every deploy (also shipped inside the @gemmein/sdk package). Copy it next to the app, set the CONFIG block, run it in CI. For a one-off check right now, call check_integration — the same checks with no file to install.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint=true, so the safety profile is covered; the description goes further by disclosing that the artifact is also shipped inside the @gemmein/sdk package and the required installation steps (copy next to the app, set the CONFIG block, run in CI). It does not mention auth requirements or any limits on the live checks it performs, so it stops short of full behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the use case and the artifact identity, followed by installation steps and the sibling alternative. A few phrases ('the ready-to-edit harness that re-proves the app's boundaries') lean promotional, but no sentence is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of explaining what the agent receives and what to do with it, and it does so: a named file, its source, and the steps to adopt it. Auth and any runtime prerequisites for the harness itself remain unspecified, which is a minor gap for a zero-parameter fetch.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters and the description correctly presents it as a no-argument artifact fetch with no inputs to configure. There is nothing to document beyond what the schema already conveys, so the baseline of 4 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the concrete artifact ('reaffirm.mjs, the ready-to-edit harness') and states its purpose: re-proving the app's boundaries against live Gemmein on every deploy. It also distinguishes itself from the sibling check_integration by contrasting an installable file with a one-off check, so an agent can pick between them without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit trigger conditions ('when you wire up the app's CI, or when you hand the finished app to your human') and explicitly routes the alternative case ('For a one-off check right now, call check_integration'). This is a clear when/when-not/alternative statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
referenceARead-only
Reach for this while WRITING code against @gemmein/sdk: every method, exact signature, return shape, and the stable error-code table (REFERENCE.md). Use guide for how the model works and whether the app fits at all; use search_docs when you need one fact from either document.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so safety is covered. The description contributes additional content context by naming the artifact (REFERENCE.md) and its payload (signatures, return shapes, stable error codes), which tells the agent what it will get back. It stops short of describing format or size, keeping it just under full credit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero waste, and the highest-value information (use this while writing SDK code) is front-loaded. The sibling-routing clause is compact and earns its place by preventing wrong tool selection.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, no-output-schema document retrieval tool, the definition covers purpose, contents, and sibling routing completely. There is no return-value contract to explain and no arguments to document, so nothing an agent needs is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so the schema imposes no semantic burden and the baseline is 4. Nothing about arguments needs explaining, and the description correctly spends no words on parameter syntax.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (the @gemmein/sdk API reference / REFERENCE.md) and enumerates exactly what it delivers: methods, exact signatures, return shapes, and the error-code table. It also explicitly contrasts itself with `guide` and `search_docs`, so an agent can disambiguate without opening other schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear trigger ('while WRITING code against @gemmein/sdk') plus explicit alternatives with selecting conditions: `guide` for conceptual/how-it-works questions and app fit, `search_docs` for pulling a single fact. This is textbook when/when-not/alternative guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_docsARead-only
Use when one question comes up mid-build ('keyed create', 'ifVersion', 'addressed', 'expand') and reading a full document would waste context. Searches the guide and the API reference; returns matching passages with 3 lines of context either side, at most 6 match blocks per document. Not the tool for the fit verdict — search finds what the docs say, not what Gemmein refuses to support; call guide for that.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | term or phrase to find |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, yet the description still adds substantive behavior: scoped search over guide + API reference, passage returns with 3 lines of context either side, and a cap of 6 match blocks per document. That return-shape and truncation detail goes well beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the usage trigger and alternative routing, and the illustrative query terms make the intent immediately legible. Nothing is redundant with the structured fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read-only lookup with no output schema, the definition covers purpose, trigger, disambiguation, and the shape/limits of results. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter exists and schema coverage is 100%, so the schema already fully documents `query` ('term or phrase to find'). The description adds contextual examples ('keyed create', 'ifVersion') but no syntax or format detail beyond the schema, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (searches) plus two concrete resources (the guide and the API reference), and explicitly distinguishes itself from the sibling `guide` by refusing the 'fit verdict' job. An agent can route between search_docs and guide without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger ('when one question comes up mid-build... reading a full document would waste context') and an explicit exclusion ('Not the tool for the fit verdict... call `guide` for that'). Both when-to-use and when-not/alternative are present.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_collection_nameARead-only
Run at PLANNING time on every collection name you intend to use, before any g.collection(name) call is written. The naming law: lowercase letters, numbers, underscores; starts with a letter; 2-63 characters. A bad name throws from g.collection(name) before any network call — at module load that blanks the whole app with no console error. An invalid name comes back with a suggested fix.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond readOnlyHint=true by disclosing the consequence of skipping it: a bad name throws from g.collection(name) *before* any network call, and at module load that blanks the whole app with no console error. It also notes that invalid input returns a suggested fix, which is real behavioral context about the failure path.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences: the when, the rule, the failure consequence, and the return behavior — each earns its place, and the strongest constraint (run before writing g.collection) is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-param validator with no output schema and only a readOnlyHint annotation, the description covers timing, the validity rules, the error consequence, and the shape of the failure response. It leaves the exact success return (boolean? echo of the name?) unstated, a minor gap given no output schema exists.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single 'name' parameter is undocumented in the schema, so the description carries the burden — and it does, spelling out the naming law (lowercase letters, numbers, underscores, starts with a letter, 2-63 chars). That fully compensates for the gap, though it never restates the parameter by name or its type.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('validate collection name') and immediately frames it as a pre-flight check for g.collection(name) calls, which no sibling tool covers. An agent can distinguish this from guide/reference/check_integration without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit trigger: 'Run at PLANNING time on every collection name you intend to use, before any g.collection(name) call is written.' This names both the timing (planning, not runtime) and the exact point of use, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.1.0- First observed
check_integration - First observed
explain_error - First observed
explain_relay - First observed
explain_rule - First observed
guide - First observed
reaffirm_template - First observed
reference - First observed
search_docs - First observed
validate_collection_name
TDQS
Scored across 9 tools
Most tools have distinct purposes, and descriptions explicitly route usage (e.g., guide vs reference vs search_docs; reaffirm_template vs check_integration). However, guide/reference/search_docs all provide documentation access, so some overlap remains that could cause hesitation.
Seven tools follow verb_noun snake_case (search_docs, explain_rule, validate_collection_name, etc.), but guide and reference are bare nouns. All names are in snake_case and readable, making this a minor deviation.
Nine tools is well within the 3-15 sweet spot and each maps to a distinct stage of app-building guidance, validation, or checking. No tool feels redundant or missing in count.
The surface covers fit assessment, SDK reference, rule/error explanations, name validation, relay semantics, CI harness setup, and live integration checks. No obvious lifecycle or documentation gap exists for the stated guidance/validation purpose.
Maintenance
Related MCP Connectors
Zero-setup MCP gateway securely connecting AI to your tools with authentication and workflows
Paid remote MCP for persistent AI agent memory, analytics, checkout, and search-readiness.
MCP server connecting AI agents to 100+ apps (Gmail, Slack, Notion, GitHub) via one-click OAuth.
A paid remote MCP for AI SDK MCP gateway registry, built to return verdicts, receipts, usage logs, a
Related MCP Servers
AlicenseNot gradedqualityBmaintenanceLets you use Claude Desktop, or any MCP Client, to use natural language to accomplish things with Neon.3,094 npm650MIT
Convex MCP serverofficial
FlicenseNot gradedqualityAmaintenanceConvex’s MCP server lets you introspect tables, call functions, and read/write data seamlessly. Agents can generate one-off queries safely—thanks to Convex’s sandboxed queries, ensuring data integrity. Perfect for AI automation, real-time apps, and dynamic data access.12,615-
Appwrite MCP Serverofficial
AlicenseCqualityBmaintenanceA Model Context Protocol server that allows AI assistants to interact with Appwrite's API, providing tools to manage databases, users, functions, teams, and other resources within Appwrite projects.4973MIT- AlicenseAqualityDmaintenanceOne MCP server for the SaaS back office. Stripe, HubSpot, and Google Sheets exposed as typed, read-only-by-default tools for Claude and any MCP client.11MIT