Skip to main content
Glama
shigechika

keycloak-mcp

by shigechika

daily_brief

Run a morning Keycloak health check to detect anomalies: failed logins, active sessions, password updates, admin events, and flag brute-force or credential-spraying attacks.

Instructions

Run a morning Keycloak health check.

Checks (all scoped to the last since_hours hours):

  • Login statistics (success / failure totals, top failing IPs)

  • Active sessions by client

  • Password update events

  • Admin events (CREATE/UPDATE/DELETE on USER/CLIENT resources)

A single IP with login failures >= ip_failure_threshold is flagged as WARNING (possible brute-force). Independently, the same login events are run through the spray_check rule (external IP, >= 10 distinct users, success rate < 20%); a match is a [SPRAY] WARNING and the "Spray check" section lists the breached accounts with their evidence tuples (time / ip / username / client). Only accounts in that list may be called breached — see spray_check for the full row shape and to widen the window or tune the thresholds.

since_hours defaults to 18 (≈ previous 15:00 for a 09:00 morning run).

Output tiers:

  • CRITICAL — API connection failure

  • WARNING — anomalies detected

  • OK — clean

Args: since_hours: Look-back window in hours (default 18). ip_failure_threshold: Login failures from a single IP that triggers a WARNING (default 50).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
since_hoursNo
ip_failure_thresholdNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv0.14.1
  2. Removedv0.13.1
  3. First observedv0.1.0

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden, and it delivers: it discloses output tiers (CRITICAL/WARNING/OK), default thresholds, scoping to since_hours, the independent brute-force and spray-check rules, and the caveat that only accounts in the Spray check list may be called breached. It even explains the evidence tuple shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Although longer than average, the description is well structured with bullets, sections, and an Args block. It front-loads the purpose, groups checks logically, and every sentence adds meaningful information about thresholds, defaults, output tiers, or routing to spray_check.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with two parameters, an existing output schema, and non-trivial detection logic, the description is complete: it covers all checks, thresholds, defaults, output tiers, and the exact semantics of the two arguments. It also points to spray_check for further tuning and row details, so an agent has everything needed to invoke and interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description fully compensates by explaining both parameters semantically: since_hours is the look-back window with a default and morning-run rationale, while ip_failure_threshold is the single-IP failure count that triggers a WARNING. This goes well beyond the bare integer schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description opens with a specific verb and resource: 'Run a morning Keycloak health check.' It then enumerates the exact checks, making clear this is an aggregate tool that sits above granular siblings like get_login_stats, get_session_stats, and get_admin_events, and it explicitly references spray_check for the spray-detection portion.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear usage context as a morning health check, with since_hours defaulting to 18 to approximate a previous-day window for a 09:00 run. It also directs users to spray_check when they need the full row shape or want to widen the window/tune thresholds, but it does not explicitly enumerate when to prefer the granular sibling tools over this aggregate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.