Skip to main content
Glama
Pappa

mcp-oidc-proxy

by Pappa

mcp-oidc-proxy

Local prototype for exercising FastMCP OIDCProxy against a demo OIDC provider.

The repository packages two cooperating apps:

  • Auth server — a NanoIDP demo OIDC provider on http://127.0.0.1:9000

  • MCP server — a FastMCP HTTP server on http://127.0.0.1:8000 with a single protected hello_world tool

Together they prove that an MCP client can authenticate through OIDCProxy, call a protected tool, and be rejected when unauthenticated.

Prerequisites

  • uv

  • Python 3.14+

Related MCP server: mcp-keycloak

Setup

uv sync
cp .env.example .env

Demo client credentials are committed in config/ with defaults matching .env.example. Both servers read the same environment variable names so credentials stay aligned.

Running locally

Run each server in its own terminal from the repository root:

uv run python -m nanoidp
uv run launch-mcp

The auth server listens on http://127.0.0.1:9000 and exposes standard OIDC discovery, JWKS, authorize, and token endpoints. NanoIDP reads committed configuration from ./config.

Demo-only warning: the auth server exists solely to validate the OIDC proxy flow during local development and testing. It is not a production identity provider.

Demo credentials:

  • User: admin / admin

  • OAuth client: values from .env.example (mcp-proxy-client / dev-secret)

Tests

uv run pytest          # unit/integration tests (smoke excluded)
uv run pytest -m smoke # end-to-end smoke tests (spawns both servers)

Smoke tests launch real subprocesses for NanoIDP and the MCP server, complete the OAuth flow with FastMCP HeadlessOAuth, and verify both authenticated and unauthenticated tool calls.

Available Tools

26 tools
clear_audit_logC

Clear the audit log

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. 'Clear' implies a destructive action, but it does not state that the operation is irreversible, that it affects all audit log entries, or whether special permissions are required. This is a significant gap for a potentially destructive tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise and front-loaded, but it is arguably under-specified rather than well-structured. It conveys the core action in a single sentence, yet omits critical context that would make it appropriately sized for an agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive tool with no annotations, no output schema, and no parameters, the description must provide sufficient context for safe invocation. It fails to mention that the log is permanently cleared, that all entries are removed, or any operational implications. This is incomplete for a tool of this nature.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the baseline is 4. The description does not need to explain parameters, and the schema already confirms no inputs are required. The description adds no parameter-specific information, but none is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb (clear) and resource (audit log), which is specific enough to distinguish it from siblings like get_audit_log or get_audit_stats. However, it does not elaborate on the scope (e.g., all entries, entire log) or the consequence of clearing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool, when not to, or any prerequisites. The description does not differentiate from alternatives or mention any cautionary context, leaving the agent to infer usage entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_clientB

Create a new OAuth client

ParametersJSON Schema
NameRequiredDescriptionDefault
client_idYesUnique client identifier
descriptionNoHuman-readable description (optional)
footer_colorNoHex color (e.g. '#ffffff') for the /authorize login card footer band (optional)
header_colorNoHex color (e.g. '#0d6efd') for the /authorize login card header band (optional)
client_secretYesClient secret for authentication
redirect_urisNoRegistered redirect URIs; when non-empty, /authorize enforces exact matching, except a registered loopback URI (http://127.0.0.1:{port}/..., http://[::1]:{port}/...) matches any port per RFC 8252 section 7.3; reverse-domain private-use schemes like com.example.app:/cb are accepted, schemes without a period such as myapp:// are rejected per section 7.1 (optional)
allowed_scopesNoPer-client scope allow-list (#186); when non-empty, /authorize and /token reject a requested scope outside this set with invalid_scope (RFC 6749 4.1.2.1/5.2). Empty = any scope in the global oauth.scopes_supported vocabulary is allowed (optional)
show_client_idNoShow client_id on the /authorize login page (optional, default true)
background_colorNoHex color (e.g. '#1a1a2e') behind the /authorize login card (optional)
show_descriptionNoShow description on the /authorize login page (optional, default false)
additional_audiencesNoExtra audiences added to the ID Token 'aud' alongside the client_id (optional)

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description carries the full burden of behavioral disclosure. It states the mutation but does not mention effects such as whether creation conflicts with existing client IDs, whether defaults apply, idempotency, or authorization requirements. The schema documents parameter-level behavior, but the description itself adds no behavioral context beyond the verb.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence with no filler or repetition. It front-loads the core purpose and does not duplicate information already present in the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The schema extensively covers parameter semanticshol, and the description correctly states the operation. However, given the absence of annotations and a missing output schema, the agent is left without explicit usage context, behavioral effects, or return-value expectations. This is adequate for selection but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema already documents each parameter thoroughly, including defaults and validation rules. The one-line description adds no parameter-level meaning, which fits the baseline of 3 when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Create a new OAuth client' pairs a specific verb with a specific resource and includes the word 'new', clearly marking this as an insertion operation. This distinguishes it from sibling tools like update_client, get_client, delete_client, and list_clients without needing additional context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided about when to choose create_client over other client-related siblings, nor are alternatives named or excluded. The intended use is only implied by the verb 'Create' and the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_persona_userA

Create a password-less user for persona login mode (local dev/testing convenience, 'login.mode: persona' in settings). The user can only authenticate by identity selection in the interactive login UI - never via password-mode login or the OAuth password grant. To keep 'create_user' unambiguous (always creates a normal, password-protected user), this is a separate tool rather than an optional password on create_user.

ParametersJSON Schema
NameRequiredDescriptionDefault
emailNoEmail address (optional)
rolesNoList of roles (optional, default: ['USER'])
groupsNoList of groups (optional)
tenantNoTenant identifier (optional, default: 'default')
usernameYesUsername for the new persona-mode-only user
attributesNoCustom key-value attributes (optional)
source_aclNoSource ACL entries for document-level security
descriptionNoDisplay-only note shown in the persona login picker (optional, max 200 chars)
entitlementsNoList of entitlements
identity_classNoIdentity class (e.g., INTERNAL, EXTERNAL)

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses that the user is password-less, can only authenticate via identity selection, and never via password-mode or OAuth password grant. It also states the purpose is a dev/testing convenience. However, it does not mention permissions, reversibility, or side effects beyond creation, so it is not perfect but strong.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, front-loaded with the core purpose, then explains the auth constraint and the reason for being a separate tool. Every sentence earns its place with no fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 10 parameters and no output schema, the description covers the essential usage context (when, why, and auth behavior). It could mention the response format or any prerequisites (e.g., admin role), but the schema covers parameter details, so the description is reasonably complete for an agent to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so each parameter already has a clear description. The tool description does not add additional meaning to parameters beyond the schema, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb (create), the resource (password-less user for persona login mode), and explicitly distinguishes it from the sibling create_user, which creates a normal password-protected user. The purpose is unambiguous and specific.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly explains when to use this tool (persona login mode, local dev/testing) and contrasts it with create_user, stating that create_user is for normal password-protected users. This gives clear usage direction without ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_userC

Create a new user in NanoIDP

ParametersJSON Schema
NameRequiredDescriptionDefault
emailNoEmail address (optional)
rolesNoList of roles (optional, default: ['USER'])
groupsNoList of groups (optional)
tenantNoTenant identifier (optional, default: 'default')
passwordYesPassword for the new user
usernameYesUsername for the new user
attributesNoCustom key-value attributes (optional)
source_aclNoSource ACL entries for document-level security
descriptionNoDisplay-only note shown in the persona login picker (optional, max 200 chars)
entitlementsNoList of entitlements
identity_classNoIdentity class (e.g., INTERNAL, EXTERNAL)

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It only restates the operation ('Create a new user') and adds nothing about side effects, validation, default role assignment, idempotency, or error behavior. For a mutation tool, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler or redundant phrasing. It is appropriately brief, though minimal; it could arguably use a second sentence about behavior or sibling distinction, but structurally it is clean.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 11 parameters, nested objects, no output schema, and no annotations, this tool is complex, yet the description offers only the bare operation statement. It omits behavioral defaults (e.g., default roles/tenant from the schema), possible auth requirements, and the critical distinction from create_persona_user, so the agent lacks enough context to call it correctly in all cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema documents all 11 parameters with descriptions, so schema coverage is 100% and the baseline is 3. The description itself contributes no parameter-level semantics, but it doesn't need to because the schema already carries that load. No penalty beyond baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb ('Create') and resource ('a new user in NanoIDP'), so an agent can tell it creates a user. However, it does not distinguish this from the sibling tool 'create_persona_user' – both appear to create users, so the agent cannot tell which one fits a persona-user scenario.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives like create_persona_user, nor any prerequisites or preconditions for creating a user. The description simply announces the operation without context, so the agent is left to guess based on parameter names.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

decode_tokenA

Decode and display the claims in a JWT token (without signature verification)

ParametersJSON Schema
NameRequiredDescriptionDefault
tokenYesJWT token to decode

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it discloses the most important behavioral trait: this tool does not verify signatures, meaning it will decode tokens regardless of validity. It also implies a read-only, non-mutating operation by saying 'decode and display', which is sufficient for this simple tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that conveys the action, the target resource, and the key limitation with zero wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with no options and no output schema, the description adequately covers what the tool does and its key caveat. It could mention what happens with malformed tokens or the exact return shape, but the minimal complexity makes the current description sufficient for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already covers the only parameter with a description ('JWT token to decode'), and the tool description adds no new semantic detail such as expected format, error behavior, or edge cases. With 100% schema coverage, a baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Decode and display') with a clear resource ('claims in a JWT token'), and immediately distinguishes itself from signature verification via the parenthetical. This makes the tool's purpose unambiguous even among siblings like verify_token and generate_token.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'without signature verification' clearly tells an agent that this tool is for inspection, not validation, which is a useful usage boundary. However, it stops short of explicitly naming verify_token as the alternative or stating when verification should be used instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_clientB

Delete an OAuth client

ParametersJSON Schema
NameRequiredDescriptionDefault
client_idYesClient ID to delete

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description bears full responsibility for disclosing behavioral traits. It states the action ('delete') but omits essential details such as irreversibility, potential invalidation of existing tokens, required permissions, or other side effects. This is a significant gap for a destructive operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence with no filler or repetition. It front-loads the action and resource, making it immediately scannable and direct.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple delete tool, the description is minimal but leaves out important operational context: permanence, side effects on tokens, and any required privileges. Given the lack of annotations or output schema, the description should carry more of this burden to ensure an agent can call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides a description for client_id, achieving 100% schema description coverage. The tool description adds no additional meaning or usage hints beyond what the schema already documents, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Delete') and resource ('OAuth client'), clearly distinguishing it from sibling tools like create_client, update_client, get_client, and list_clients. The meaning is unambiguous and immediately actionable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The use case is implicitly clear from the verb, and the sibling set offers no alternative delete-like tool for OAuth clients. However, there is no explicit when-to-use guidance, prerequisites, or exclusions, leaving the agent to infer context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_userC

Delete a user from NanoIDP

ParametersJSON Schema
NameRequiredDescriptionDefault
usernameYesUsername to delete

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. 'Delete' implies destructive mutation, but the description discloses nothing about irreversibility, cascading effects on tokens or sessions, or success/error behavior. For a destructive operation this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with zero waste and the operation front-loaded. Not a 5 because it is so minimal it borders on under-specification rather than deliberate conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive operation with no annotations, no output schema, and no behavioral disclosure, a one-line description is inadequate. It doesn't cover error cases, side effects, or restrictions that an agent needs to call it safely and correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the schema already describes username as 'Username to delete', so the description adds nothing beyond the structured field. Baseline 3 is appropriate since the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb+resource ('Delete a user from NanoIDP'). The verb 'Delete' distinguishes it from siblings like create_user, update_user, get_user, and list_users without needing to open any schema. Not a 5 because it doesn't explicitly scope what kind of user or any special cases.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this vs alternatives, no prerequisites, and no exclusions (e.g., cannot delete the last admin, cannot delete a user with active sessions). The name makes the intent obvious, but there is zero contextual routing help for the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_tokenB

Generate an OAuth2 access token for a user

ParametersJSON Schema
NameRequiredDescriptionDefault
scopeNoSpace-separated OAuth scopes (optional). Include 'openid' to also receive an ID Token; the scope is persisted in the refresh token so refreshing re-issues an ID Token (OIDC Core §12.2)
usernameYesUsername to generate token for
extra_claimsNoAdditional claims to include in the token
id_token_claimsNoClaim names to embed in the ID Token, mirroring the OIDC `claims` request parameter (§5.5). Requires an 'openid' scope. Resolved from the user (e.g. 'email', 'preferred_username', or a custom attribute); names nanoidp cannot supply are skipped.
userinfo_claimsNoClaim names /userinfo should return for this access token, mirroring the `userinfo` member of the OIDC `claims` request parameter (§5.5). Stamped on the access token as `req_userinfo_claims` and honoured by /userinfo even under a stricter profile that would scope-gate them out.
expires_in_minutesNoToken expiration in minutes (optional, default: 60)

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full burden, but it only states that a token is generated. It does not disclose side effects, such as refresh-token persistence, auditability, or authorization requirements for this sensitive operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler or repetition. It states the core action efficiently, even though additional context would be welcome elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a credential-issuing operation among admin/audit tools, with no output schema and no annotations, yet the description provides no context about expected output, permissions, or when to choose it. The schema is rich, but the description leaves important operational context missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and every parameter already has a detailed explanation, including scope persistence and OIDC claims behavior. The description adds no parameter-level meaning, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb (Generate), a clear resource (OAuth2 access token), and a beneficiary (for a user). This clearly differentiates it from siblings like decode_token and verify_token, which inspect rather than create tokens.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as verify_token or decode_token. Prerequisites, like admin permissions or conditions for generating a token, are also left entirely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_audit_logB

Get audit log entries (what the IdP recorded: token requests, logins, SAML flows)

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum entries to return (default: 100)
usernameNoFilter by username
event_typeNoFilter by event type (e.g. token_request, authorization_request)

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behaviors. It only states 'Get audit log entries' and lists content types, but doesn't mention any behavioral traits such as pagination, ordering, default limits, or that it is read-only. The parenthetical adds context about content but not behavior, leaving the agent with limited understanding of side effects or response characteristics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, succinct sentence that leads with the core action and resource, then adds a clarifying parenthetical. No fluff, no redundancy. It is front-loaded and efficient for a simple read tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a straightforward read operation with fully documented parameters and no output schema, the description is mostly adequate. It explains what the entries contain, but does not specify the return structure (e.g., list of objects) or any ordering/pagination details. The lack of annotations increases the need for such context, so it falls short of excellent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema provides 100% coverage with descriptions for all three parameters (limit, username, event_type). The description does not add any parameter-specific meaning, but the baseline of 3 applies because the schema handles parameter documentation adequately.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'Get' with the resource 'audit log entries' and clarifies the content with examples (token requests, logins, SAML flows). It implicitly differentiates from sibling tools like get_audit_stats (statistics) and clear_audit_log (deletion) by focusing on entries.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives. It doesn't mention get_audit_stats for aggregated statistics or clear_audit_log for deletion, nor any exclusions or preferred contexts. The agent must infer usage from the name and siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_audit_statsA

Get audit log statistics (event counts by type/status)

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the burden. The verb 'Get' implies a read-only operation and the parenthetical describes the result as event counts, but the description does not explicitly state that the audit log is unchanged, nor does it mention response format, limits, or scoping behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One tight sentence with the core action first and a parenthetical adding the key output detail. There is no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read-only tool, the description is sufficient for an agent to invoke it and understand it returns aggregated counts by type/status. It does not specify time range, whether counts are global, or an output schema, but the tool is simple enough that those gaps are minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters, so no parameter documentation is needed; baseline 4 applies. The description correctly avoids inventing parameter details that do not exist.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description uses a specific verb and resource, 'Get audit log statistics', and the parenthetical 'event counts by type/status' clarifies the output. It does not explicitly contrast with the sibling get_audit_log, but the statistics framing makes the distinction mostly clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use this tool versus get_audit_log, clear_audit_log, or any other sibling. There are no stated exclusions, prerequisites, or alternative conditions, so an agent must infer usage from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_clientB

Get details of a specific OAuth client

ParametersJSON Schema
NameRequiredDescriptionDefault
client_idYesClient ID to look up

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must disclose behavior itself. It only says 'Get details', implying a read-only operation but does not explicitly state that it has no side effects, does not describe permissions or rate limits, and gives no indication of error behavior or return format.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that immediately conveys the core action. There is no redundancy or unnecessary detail, and it is appropriately front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple with one parameter, but there is no output schema and the description does not explain what details are returned, whether a missing client produces an error, or how errors are presented. This leaves some ambiguity for an agent, though the core operation is clear.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already fully describes the single parameter client_id as 'Client ID to look up', covering 100% of parameters. The description adds no additional meaning or nuance beyond the schema, so it meets the baseline without adding value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the verb 'Get' and the resource 'details of a specific OAuth client', which clearly differentiates it from sibling tools like list_clients that enumerate all clients. It is specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives such as list_clients, nor does it mention any prerequisites like the client_id needing to be known or valid. There is no context for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_jwksA

Get JSON Web Key Set for token verification

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden. 'Get' implies a read-only retrieval, and 'for token verification' adds context, but the description does not explicitly disclose side-effect-freeness, authentication requirements, or any operational behaviors such as caching or failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, tightly worded sentence that conveys the object and purpose with no filler. It is front-loaded and appropriate for a tool with no parameters and no complex configuration.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter retrieval tool, the description tells the agent what will be fetched (JWKS) and why (token verification). It lacks explicit detail about the exact return shape or whether any auth is needed, but given the tool's simplicity, it is reasonably complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and the schema is empty, so there are no parameter semantics to document. The baseline for a zero-parameter tool is 4, and the description provides no irrelevant parameter information.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get') and resource ('JSON Web Key Set') and clarifies its purpose ('for token verification'). It is distinct from common siblings like get_keys_info or get_oidc_discovery by naming the exact artifact, though it does not explicitly contrast itself with those tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to prefer this tool over alternatives such as get_oidc_discovery, get_keys_info, or verify_token. There is no mention of typical use cases, prerequisites, or situations where a sibling would be more appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_keys_infoB

Get information about the signing keys (active kid, previous keys)

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. 'Get information' implies a read-only operation with no destructive side effects, but it does not disclose auth requirements, response shape, or behaviors around key rotation. This is minimal but not misleading.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence with no filler, front-loading the verb and object and parenthetically listing the key data returned. It is as compact as the tool's simplicity allows.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-parameter info tool, the description covers the core payload: active and previous signing keys. However, it omits guidance on when to prefer get_keys_info over get_jwks, and with no annotations or output schema the agent still lacks indication of auth or output format, leaving moderate gaps for a key-related tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and an empty input schema, so there are no parameter semantics to document. Under the zero-parameter baseline, the description has no obligation to explain parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'Get' and clearly identifies the resource 'signing keys' with concrete content '(active kid, previous keys)'. It conveys a distinct read purpose and is distinguishable from write-oriented siblings like rotate_keys, though it does not explicitly contrast with the similar get_jwks.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to call this tool versus alternatives such as get_jwks or rotate_keys. It does not mention preconditions, permissions, or conditions that would favor another tool, so the agent must infer usage from sibling names alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_oidc_discoveryA

Get OIDC discovery document (/.well-known/openid-configuration)

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It conveys a read-only retrieval action via 'Get' and identifies the resource, but it does not disclose response format, authentication needs, or potential errors. Those omissions are relatively minor for a well-known static discovery endpoint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with the action and resource front-loaded)Skip followed by the precise well-known path in parentheses. Every word earns its place and there is no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless GET to a standard OIDC discovery endpoint, the description supplies the exact URL and the resource name, which is nearly complete. It does not explicitly say the response is a JSON document, but the well-known path and 'discovery document' phrasing make that obvious.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters with 100% coverage, so there is nothing the description must add about parameters. The zero-parameter baseline applies here, and the description correctly implies that no arguments are needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Get') and resource ('OIDC discovery document'), and includes the exact well-known path. This clearly distinguishes it from sibling tools like get_jwks or get_keys_info, which serve different discovery and key endpoints.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the usage case: whenever the OIDC discovery document is needed. However, it does not explicitly state when to use this tool versus alternatives, nor does it mention any exclusions or prerequisites. For a zero-parameter standard endpoint, the implication is acceptable but not explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_settingsB

Get current NanoIDP settings

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of behavioral disclosure. The verb 'Get' implies a read-only, side-effect-free operation, which is adequate but minimal. It does not mention authentication requirements, response shape, or whether settings are returned in full, leaving some ambiguity for a tool with zero annotation support.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with zero filler. It states the action and resource directly, making it easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple no-parameter getter, the description is minimally sufficient, but it lacks detail about what is included in the returned settings or whether there are any operational caveats. Since there is no output schema and no annotations, a bit more context about the response content would make it fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the baseline is 4. The description adds no parameter-specific meaning, but none is needed since the input schema is empty and no configuration choices are required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Get') and resource ('current NanoIDP settings'), making it clear this is a read operation for the settings state. It is easily distinguishable from siblings like update_settings or save_config, though it does not explicitly contrast with get_oidc_discovery or other config-related getters.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No usage guidance is provided. The description does not state when to call this tool versus siblings like update_settings or get_oidc_discovery, nor does it mention any prerequisites or context in which it should be used.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_userC

Get details of a specific user

ParametersJSON Schema
NameRequiredDescriptionDefault
usernameYesUsername to look up

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It does not state whether this is a read-only operation (likely, given 'Get'), whether it requires authentication, what happens if the user does not exist, or any throttling/error behavior. It mentions nothing beyond the purpose, leaving the agent to guess at side effects and failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence with no fluff, and it is front-loaded with the main purpose. It is efficient, though it could have added more detail in the same length. The conciseness itself is good, but it sacrifices necessary context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a single parameter and no output schema, the description is minimally sufficient for a simple lookup, but it lacks context about return format, error handling, and relationship to sibling tools. For an API with many similar tools, the absence of routing or context makes it incomplete for reliable agent selection.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema coverage is 100% and the sole parameter 'username' is described as 'Username to look up', which aligns with the description. The description adds no additional meaning beyond the schema; it just reiterates the intent. A baseline 3 is appropriate since the schema is sufficient for the single parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get details of a specific user' names the verb and resource clearly, but it is generic and does little to differentiate from sibling tools like list_users, create_user, update_user, or delete_user. It states what it does but not the specific scope (e.g., vs. a user list). A 3 is appropriate as it is clear but not differentiating.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like list_users or get_client. The agent must infer usage from the name alone, which is insufficient given there are multiple user-related tools and other entity lookup tools. This is a critical gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_clientsA

List all configured OAuth clients

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the read-only scope and that all configured clients are returned, but it does not mention output shape, pagination, or whether sensitive fields like secrets are included. Acceptable for a minimal zero-parameter list tool, but still thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with no filler, front-loading the verb and object. Everything included is relevant and nothing needs to be trimmed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a simple zero-input tool, and the description is enough for an agent to invoke it correctly. It could add output shape or pagination notes, but those are secondary for such a low-complexity operation, especially with no output schema to complement it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are no parameters, so there is nothing extra to explain. The base level for a zero-parameter tool is appropriate, and the 'all' language correctly signals an unfiltered global operation rather than a target-specific lookup.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description uses a specific imperative verb ('List') with a clear object ('all configured OAuth clients') and scope ('all configured'). This distinguishes it from get_client for single-client lookups and from create/update/delete_client mutations among siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: an agent would use this when it needs the full set of configured OAuth clients. However, there is no explicit guidance about when to prefer this over get_client or list_users, and no exclusions or alternatives are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_usersA

List all configured users in NanoIDP

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full behavioral burden. It only states the basic action and doesn't disclose permissions, pagination, ordering, response format, or any side effects. 'List' implies read-only, but nothing more is added.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that is front-loaded with the verb and resource. There is no fluff or redundant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple zero-parameter list tool, the description is mostly complete, but it doesn't describe the return value structure or any pagination behavior, which could matter to an agent. The core purpose, however, is unambiguous.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the baseline is 4. The description adds useful scope context ('all configured users') that clarifies the tool's breadth, adding value beyond the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') with a clear resource ('all configured users') and scope ('in NanoIDP'). Distinguishes from siblings like get_user, which likely targets a single user.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrasing 'all configured users' implies this is for full listing, but it lacks explicit guidance on when to use it versus get_user for individual lookups. No exclusions or conditions are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reload_configB

Reload configuration from files

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It only states the action and source; it doesn't disclose side effects (e.g., overwriting unsaved changes), failure modes, validation behavior, or what it returns. This is a significant gap for a mutating operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise—five words with no waste, leading with the verb. It is front-loaded and easy to parse, though it sacrifices behavioral detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, no-output-schema action, the core action is captured, but side effects and preconditions are missing. The agent does not know if this is a safe read or a mutating reload, nor what happens on failure. Given the lack of annotations and output schema, more context is needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has zero properties, so there is nothing to document. Baseline 4 applies because the description cannot add meaning to parameters that don't exist, and the schema fully covers all parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('reload') on a specific resource ('configuration') with a source ('from files'). This is clear and distinguishes the tool from siblings like save_config or validate_config, though it doesn't explicitly name alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided about when to use this tool versus alternatives. The agent gets no context for selecting reload_config over validate_config, save_config, or get_settings, nor any indication of prerequisites or conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rotate_keysA

Rotate the signing keys: the active key moves to 'previous' (still valid for verification) and a new active key is generated - useful to test clients' JWKS refresh handling

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses the behavioral side effect: the active key is rotated, but the previous key remains valid for verification. It also notes the generated new key, which implies a mutation. Since there are no annotations, the description carries the full burden, and it does an excellent job of stating the security-relevant side effect. The only minor gap is that it doesn't explicitly mention that this changes the JWKS endpoint results, but that is implied by the context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one sentence that packs all essential information: the operation, the outcome, and the use case. It is efficient and front-loaded with the action. Nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has no parameters, no output schema, and no annotations, the description is nearly complete for a mutation tool. It explains the side effect and the typical use case. The only missing piece is what the response looks like (e.g., does it return the new key info?), but since there is no output schema, the description could have briefly mentioned the return value. That is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so parameter semantics are inherently not an issue. The description doesn't need to explain parameters. A score of 4 is justified because the description effectively eliminates any need for parameter documentation, and the baseline for zero parameters is 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (rotate), the resource (signing keys), and the specific behavior (active key moves to 'previous', new key generated). It also includes a purpose ('test clients' JWKS refresh handling'), which distinguishes it from other key-related operations like get_keys_info or get_jwks.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool: when testing clients' JWKS refresh handling. It doesn't explicitly state when not to use it or name alternative tools, but given the tool has no parameters and is specialized, the usage context is reasonably clear. However, it could mention that it's not needed for simple key inspection (get_keys_info) or general key retrieval (get_jwks).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_configA

Save current configuration to YAML files (persists changes made via create/update tools without persist=True support). Writes users.yaml and settings.yaml as one coordinated, conflict-checked save (#229) and then refreshes the running configuration from what was just written. To refuse the save if another writer (the web UI, another agent, a second nanoidp process on the same directory) changed a file since you read it, pass the expected_users_revision / expected_settings_revision a read tool handed back (list_users and get_user carry users_revision; list_clients, get_client and get_settings carry settings_revision; reload_config and a successful save_config carry both). save_config always writes both files, so there are exactly two modes: omitting both revisions keeps today's unconditional last-write-wins, and supplying either makes the WHOLE save conflict-checked - the omitted revision defaults to the one this runtime was loaded from, so a save guarded on users.yaml cannot silently overwrite a settings.yaml another writer changed, or vice versa. A failure response's 'kind' distinguishes four outcomes: 'conflict' (nothing was written - a supplied revision was stale; call reload_config, reapply your change on the fresh state and save with the revisions from its response), 'lock_timeout' or 'lock_unsupported' (nothing was written either - the write never started; lock_timeout is worth retrying, lock_unsupported means this config directory's filesystem does not support advisory locks and will not succeed on retry), a hook's own 'kind' under hooks.strict (both files ARE written; only the mirror push failed), or 'reload_after_save' (both files ARE written but the runtime could not adopt them - do not retry expecting a different result, the file on disk is authoritative).

ParametersJSON Schema
NameRequiredDescriptionDefault
expected_users_revisionNousers.yaml revision from a read tool; the save is refused with kind 'conflict' if the file no longer matches it. Supplying either revision makes the whole two-file save conflict-checked (the omitted one defaults to this runtime's loaded revision); omit both for unconditional last-write-wins.
expected_settings_revisionNosettings.yaml revision from a read tool; same contract as expected_users_revision.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description fully carries the behavioral burden. It discloses that the tool writes both files, coordinates the save, refreshes the runtime config, refuses the save on stale revisions, and defines four distinct failure kinds with their consequences (nothing written vs. files written but mirror/reload failed). This is exceptionally transparent about side effects and edge cases.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense and long, but every sentence carries necessary information for a complex tool. It is front-loaded with the primary purpose and then logically progresses through modes and failure outcomes. It loses one point because a structured list or clearer paragraph breaks for the four failure kinds would improve scannability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having no output schema and no annotations, the description is complete enough for an agent to invoke the tool correctly and handle all outcomes. It explains return semantics via the failure 'kind' field, gives recovery steps for conflicts, warns about non-retryable lock_unsupported conditions, and clarifies when files are actually written. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds substantial semantics beyond the schema. It explains the interaction between the two revisions, the defaulting of an omitted revision to the runtime's loaded revision, and the fact that supplying either revision makes the entire two-file save conflict-checked. This deeper meaning is critical for correct invocation and is not fully derivable from the schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'Save current configuration to YAML files', and immediately names the exact outputs (users.yaml, settings.yaml). It distinguishes itself from siblings by stating it persists changes made via create/update tools that lack persist=True, while related tools like reload_config focus on runtime state. This leaves no ambiguity about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use the tool (to persist changes from create/update tools without persist=True support) and gives conditional guidance for the two modes: omit revisions for last-write-wins, or supply them to enable conflict checking. It also names alternatives in context, such as calling reload_config after a conflict and using read tools to obtain revisions. This is clear, actionable usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_clientC

Update an existing OAuth client

ParametersJSON Schema
NameRequiredDescriptionDefault
client_idYesClient ID to update
descriptionNoNew description (optional)
footer_colorNoNew hex color (e.g. '#ffffff') for the /authorize login card footer band; empty string clears it (optional)
header_colorNoNew hex color (e.g. '#0d6efd') for the /authorize login card header band; empty string clears it (optional)
client_secretNoNew client secret (optional)
redirect_urisNoReplace the client's registered redirect URIs (loopback URIs match any port per RFC 8252 section 7.3, reverse-domain private-use schemes accepted, myapp:// rejected per section 7.1); empty list removes the restriction (optional)
allowed_scopesNoReplace the client's scope allow-list (#186); empty list removes the restriction (optional)
show_client_idNoShow client_id on the /authorize login page (optional)
background_colorNoNew hex color (e.g. '#1a1a2e') behind the /authorize login card; empty string clears it (optional)
show_descriptionNoShow description on the /authorize login page (optional)
additional_audiencesNoReplace the client's extra ID Token audiences (optional)

TDQS

C2.7/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure, but it only restates the tool's name. It does not disclose side effects (e.g., whether updating client_secret invalidates existing tokens), whether the update is partial or a full replacement, permission requirements, or any consequences for the login page. This is a serious gap for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no wasted words, which is appropriately concise. However, it is so minimal that it fails to convey any useful information beyond the tool name, so it does not earn its place as a meaningful description. It lacks the field enumeration seen in the calibration example for update_drive.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 11 parameters, no output schema, and no annotations, the description is severely incomplete. It omits essential context such as whether the update is partial or full, the effect on existing tokens, authorization requirements, and how the login page is affected. An agent cannot reliably invoke this tool correctly based solely on the provided description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents every parameter's meaning and constraints. The tool description itself adds no parameter-level information, which aligns with the baseline of 3 for high coverage. It does not introduce any additional context beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb-resource pair ('Update an existing OAuth client') that distinguishes this from create_client and delete_client by specifying 'existing'. It unambiguously identifies the operation and target, leaving no doubt about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives such as create_client, delete_client, or get_client. It does not mention prerequisites (e.g., the client must already exist), nor does it explain scenarios where partial updates are preferred over full replacement. The agent is left to infer usage from the name and schema alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_settingsA

Update NanoIDP settings (issuer, audience, token expiry, SAML options, etc.). hooks: and plugins: (#185) are YAML-only, like secret_key and require_ui_login: they are reported by get_settings but cannot be changed here, since a command editable through the surface it observes would be a remote-execution primitive.

ParametersJSON Schema
NameRequiredDescriptionDefault
issuerNoOAuth2/OIDC issuer URL
audienceNoDefault token audience
login_modeNoInteractive login mode: 'password' (default) requires the configured password on /login, /authorize, /saml/sso and the device flow; 'persona' lists the configured users and logs in by selecting one, no password prompt. Opt-in, off by default - a local development/testing convenience, not an authentication mode for deployed environments. Orthogonal to 'security_profile' and to the OAuth password grant, which is unaffected either way.
require_pkceNoReject /authorize requests without a PKCE code_challenge (#47)
saml_sso_urlNoSAML SingleSignOnService location. Empty string clears it so it is derived again as <issuer>/saml/sso (#181)
saml_entity_idNoSAML IdP entityID. Empty string clears it so it is derived again from the effective issuer as <issuer>/saml (#181)
verbose_loggingNoInclude usernames/client_ids in log messages (dev convenience)
issuer_allowlistNoOrigins (e.g. 'http://localhost:8000') allowed to be reflected back by 'issuer_from_request'. Empty (default) allows any Host header. A non-matching Host falls back to the fixed 'issuer'.
saml_export_rolesNoEmit the user's roles as a SAML attribute (off by default)
saml_export_groupsNoEmit the user's groups as a SAML attribute (off by default)
issuer_from_requestNoDerive the issuer from each request's own Host header instead of the fixed 'issuer' (dev convenience for setups reachable under more than one hostname). MCP tools have no request of their own, so this only affects HTTP discovery/token/device-flow responses, never MCP ones.
saml_c14n_algorithmNoXML canonicalization algorithm: 'c14n' (1.0), 'c14n11' (1.1), or 'exc_c14n' (Exclusive 1.0)
saml_sign_responsesNoEnable/disable SAML response signing
strict_saml_bindingNoEnforce strict SAML binding compliance (reject GET with uncompressed data)
saml_roles_attr_nameNoSAML attribute name for the roles (default: 'roles')
saml_sp_certificatesNoPEM certificate files of SPs whose AuthnRequest signatures are accepted
token_expiry_minutesNoToken expiration in minutes
saml_groups_attr_nameNoSAML attribute name for the groups (default: 'groups')
refresh_token_rotationNoRotate refresh tokens: each refresh invalidates the consumed refresh token (#46)
issuer_from_proxy_headersNoTrust 'X-Forwarded-Proto'/'X-Forwarded-Host'/'X-Forwarded-For' from a single reverse-proxy hop in front of NanoIDP (applies werkzeug's ProxyFix). Only affects the 'issuer_from_request' derivation - and only when that toggle is also on; it always affects rate-limit client IP attribution regardless. Only enable this when NanoIDP is deployed directly behind exactly one trusted proxy - these headers are otherwise spoofable by any client. Takes effect on the next app restart, not the running process.
device_verification_base_urlNoFixed base URL for the device flow's verification_uri (e.g. 'https://idp.example.com'), used instead of the request-derived issuer so a backend/container caller's Host doesn't leak into a URL a human's browser can't reach. Only consulted when 'issuer_from_request' is on; empty string clears it back to following the request Host.
saml_want_authn_requests_signedNoRequire and verify AuthnRequest signatures, both bindings (#69)

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses the YAML-only limitation and the remote-execution security reasoning, which is valuable behavioral context. However, it does not mention whether changes are persisted, require a restart, or return the updated settings—though some of that is covered per-parameter in the schema. It is transparent about what it cannot do, which is a strong point.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with the main purpose front-loaded, followed by a focused clarification about exclusions. No wasted words; every part earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an update tool with 22 optional parameters and no output schema, the description is reasonably complete in scope and limitation. It does not describe return values or error handling, but given the schema's thorough parameter explanations and the clear purpose, it is adequate for an agent to invoke correctly. The YAML-only caveat is a critical piece that is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides 100% coverage with detailed descriptions for all 22 parameters, including enums and examples. The tool description adds no additional parameter-specific meaning beyond the overall scope. Baseline 3 is appropriate since the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Update' and the resource 'NanoIDP settings', listing example categories (issuer, audience, token expiry, SAML options). It also differentiates from the sibling get_settings by explicitly noting which settings are YAML-only and cannot be changed here, so an agent can distinguish the tool's scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains what the tool can and cannot do, specifically excluding hooks/plugins/secret_key/require_ui_login and giving a security rationale. However, it does not explicitly contrast with siblings like save_config or reload_config, nor does it state prerequisites (e.g., 'use get_settings first to see current values'). The guidance is implicit rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_userC

Update an existing user's attributes

ParametersJSON Schema
NameRequiredDescriptionDefault
emailNoNew email (optional)
rolesNoNew roles list (optional)
groupsNoNew groups list (optional)
tenantNoNew tenant (optional)
passwordNoNew password (optional)
usernameYesUsername to update
source_aclNoNew source ACL entries (optional)
descriptionNoNew display-only persona picker note (optional, max 200 chars)
entitlementsNoNew entitlements list (optional)
identity_classNoNew identity class (optional)

TDQS

C2.1/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of disclosing behavioral traits. It only says 'update' without indicating whether it performs a partial or full replacement of attributes, what happens to unmentioned fields, required permissions, or side effects. This is a serious gap for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence, but it is under-specified rather than appropriately concise. For a tool with 10 parameters, it lacks critical context that would justify its brevity. It is minimal but not efficiently informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (10 parameters), absence of annotations, and lack of an output schema, the description is severely incomplete. An agent cannot determine expected behavior, required fields beyond username, or what constitutes a successful update. The description fails to provide necessary operational context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents each parameter clearly. The description adds no additional meaning beyond the generic term 'attributes', which does not map to specific parameters. It meets the baseline for high schema coverage but does not exceed it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the action (update) and the resource (existing user) but does not specify which attributes are updatable, relying entirely on the schema. It distinguishes from create/delete but is too generic to guide an agent on scope, making it a minimal but not tautological purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus siblings like create_user or update_settings. There is no mention of prerequisites, exclusion criteria, or when this is the appropriate choice, leaving the agent to infer based on the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_configA

Validate the running configuration directory (settings.yaml, users.yaml, bootstrap.yaml): unknown keys as warnings, wrong types and refused values as errors. settings.yaml and users.yaml findings are what a startup or the next reload would hit; bootstrap.yaml findings are what would stop the NEXT startup (the bootstrap surface loads at startup only). Read-only and inert: it re-reads the files through the same loaders, runs no hook and loads no plugin. 'valid' is false on any error, and on a warning too under strict mode, which is when a start would refuse. 'strict' defaults to this server's effective validation mode; pass it explicitly to override.

ParametersJSON Schema
NameRequiredDescriptionDefault
strictNoTreat warnings as failures, like the server's --strict-config. A directory declaring config_validation: strict is strict regardless.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full behavioral burden. It explicitly states read-only, inert, runs no hook, loads no plugin, and explains the strict mode behavior and the impact of bootstrap findings on next startup. This fully discloses the tool's effects and semantics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is moderately long but every sentence contributes: purpose, file-specific behaviors, read-only/inert nature, and strict mode semantics. It is front-loaded with the core purpose and then explains nuances, making it efficient without being verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema, the description explains what 'valid' means and how strict mode affects the result. It covers the single optional parameter and its default behavior. The tool is fully specified for an agent to call correctly, including edge cases like bootstrap.yaml affecting next startup.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the parameter description in the schema already explains strict behavior. The tool description adds value by stating that strict defaults to the server's effective mode and can be overridden, providing context beyond the schema. This is a meaningful addition.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool validates configuration files (settings.yaml, users.yaml, bootstrap.yaml) and specifies the action (validate). It distinguishes itself from siblings like get_settings and reload_config by focusing on validation rather than retrieval or mutation. The phrase 'running configuration directory' specifies the resource scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for pre-flight checks before reload or startup, but doesn't explicitly say when to use it versus alternatives. It does provide clear context that it's read-only and inert, which tells the agent it's safe and non-mutating, but lacks an explicit 'use this instead of X' statement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_tokenA

Verify a JWT token's signature and expiration

ParametersJSON Schema
NameRequiredDescriptionDefault
tokenYesJWT token to verify

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without annotations, the description carries the burden of behavioral disclosure and does state the verification scope (signature and expiration). However, it omits outcome behavior—what happens for invalid, expired, malformed, or valid tokens—and any side effects, so the disclosure is incomplete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler or redundancy. Every word adds information, making it easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's low complexity and one documented parameter, the description is broadly adequate. Still, with no output schema and no mention of return/error behavior, an agent cannot be fully certain how to interpret the verification result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already describes the single token parameter with 100% coverage. The description adds semantic framing by mentioning signature and expiration, but it does not meaningfully extend the parameter-level meaning beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('verify') with a concrete resource ('JWT token') and states exactly what is checked: signature and expiration. This clearly differentiates it from sibling tools like decode_token, which parses rather than validates.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool should be used when a token's authenticity/validity needs to be confirmed, but it provides no explicit when-to-use or when-not-to-use guidance. It does not name alternatives or exclusion conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 26 tool updatesv0.1.0
    • First observedclear_audit_log
    • First observedcreate_client
    • First observedcreate_persona_user
    • First observedcreate_user
    • First observeddecode_token
    • First observeddelete_client
    • First observeddelete_user
    • First observedgenerate_token
    • First observedget_audit_log
    • First observedget_audit_stats
    • First observedget_client
    • First observedget_jwks
    • First observedget_keys_info
    • First observedget_oidc_discovery
    • First observedget_settings
    • First observedget_user
    • First observedlist_clients
    • First observedlist_users
    • First observedreload_config
    • First observedrotate_keys
    • First observedsave_config
    • First observedupdate_client
    • First observedupdate_settings
    • First observedupdate_user
    • First observedvalidate_config
    • First observedverify_token

TDQS

B3.2/5.0

Scored across 26 tools

Disambiguation5/5

Every tool targets a distinct resource and action with clear descriptions that explicitly disambiguate similar pairs (e.g., decode_token vs verify_token, create_user vs create_persona_user). No two tools have overlapping purposes.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern in snake_case, using standard verbs like get, list, create, update, delete, rotate, and save. Minor variations (list_users vs get_user) reflect conventional list-vs-fetch semantics.

Tool Count2/5

At 26 tools, the server exceeds the 25-tool threshold for 'too many' defined in the rubric. While the broad scope of an OIDC proxy justifies many operations, the count is above the well-scoped range and may overwhelm an agent with options.

Completeness4/5

The tool set provides comprehensive coverage: CRUD for users and clients, token generation and verification, key management, audit logging, and config lifecycle (validate, reload, save). Minor gaps exist (e.g., no granular audit log deletion or token introspection), but they do not create dead ends.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables secure MCP tool calls (add and multiply numbers) by validating OAuth2 tokens via Keycloak token introspection.
    -
  • F
    license
    Not graded
    quality
    B
    maintenance
    Enables an MCP server with OAuth authentication, protecting tools like user CRUD operations behind session tokens obtained through a browser-based authentication flow.
    -
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables deploying a remote MCP server with OAuth 2.1 resource-server authorization, protecting tools behind validated bearer tokens and RFC 9728 discovery.
    MIT