Skip to main content
Glama

save_alert_rule

Create or update an alert rule for the calling tenant in one call. WRITE: available to any authenticated user. Omit ruleId to create a new rule; supply ruleId to REPLACE an existing one.

UPDATE IS A WHOLE-OBJECT REPLACE, NOT A MERGE. Every field you leave out is cleared — omitting description sets it to null. To change one thing, fetch the rule with get_alert_rules and re-send its full spec with that one field altered. Two things are carved out and survive omission: status — ACTIVE/DISABLED is preserved; change it with set_alert_rule_status delivery — notifyOnResolve is preserved. Where a rule's alerts go is not held on the rule at all: routing lives in the notification gateway, so set it with set_notification_destinations, passing source "alerting" and the rule id as subject.

Re-validates exactly like preview_alert_rule: if the spec is invalid, nothing is persisted and problems[] is populated instead of rule — preview_alert_rule first to calibrate the threshold, then save once problems[] is empty there.

The rule still saves even when warnings[] is non-empty — warnings are advisory, never a reason to withhold saving, unlike problems[]. warnings[] currently carries one code, FIELD_NEVER_OBSERVED: a filter/groupBy field querysql couldn't resolve to a known column (so it silently falls back to reading it from the JSON catch-all) and that has never appeared in this customer's recent telemetry — almost always a typo'd field name, especially when preview_alert_rule also reported dataCoverage.status = NO_MATCHING_DATA. Fix the spelling and re-preview rather than treat it as a calibration problem.

Authors a single metric or anomaly rule (one measure over a rolling window). Compound multi-condition rules can't be created here — build those in the web editor.

METRIC RULE (structured) — watches one measure over a rolling time window: source: telemetry source (required): LOGS, SPANS, METRICS filter: optional QuerySQL boolean filter, e.g. service = 'my-svc' fn: catalog measure function, e.g. count, error_rate, p95, error_burn_rate arg: optional field the measure operates on, e.g. duration_ms for p95 params: optional named measure params, e.g. {"budget":"0.001"} (error_burn_rate) expression: optional free-form aggregate (used instead of fn) — a ratio/calculation, e.g. countIf(status_code = 'ERROR') * 100.0 / count() (this is exactly fn: error_rate; use fn instead unless you need a custom ratio — both already return 0-100, don't divide by 100 again) metricName/metricType: required only when source is METRICS unit: optional explicit display unit for the measure, e.g. BYTES or DURATION_MS — set it when the metric name doesn't self-describe its unit (OTel names like jvm.memory.used or http.server.request.duration carry no unit suffix); omit it to let the server infer the unit from the metric name or measure function windowMinutes: rolling window length in minutes (required) groupBy: optional list of plain field names to group the series by, e.g. service. Field names only: a string that looks like an expression (it contains a parenthesis or a space) is rejected in problems[], because the rule page and the alert title would then show the raw expression as the group name. derivedGroupBy: optional list of computed group keys, each an object {"expression": "...", "label": "..."}, e.g. {"expression": "regexp_extract(message, 'customerId=([0-9a-f-]+)', 1)", "label": "customer"}. Both fields are required and non-blank. The label is what rule pages and alert titles show, so give it a short readable name. groupBy and derivedGroupBy can be used together. comparator: threshold comparator: GT, GTE, LT, LTE (required for static) warningThreshold: the warning-tier threshold the measure is compared against (required for static) warningConsecutiveWindows: consecutive breaching windows for the warning tier (default 1) criticalThreshold / criticalConsecutiveWindows: optional escalation tier

ANOMALY METRIC (structured) — flags a measure that deviates from its own historical baseline instead of a fixed threshold. Supply zScoreThreshold + direction instead of comparator/warningThreshold; groupBy must be empty. zScoreThreshold: robust z-score magnitude that counts as anomalous (> 0) direction: HIGH (spikes above baseline) or LOW (drops below baseline) anomalyConsecutiveWindows: consecutive anomalous windows required (>= 1)

Common fields: name: human-readable rule name (required, non-blank) description: optional free text notifyOnResolve: whether to notify when the alert resolves (default true) active: create-only — whether the rule starts ACTIVE (default true) or DISABLED

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
fnNoCatalog measure function: count, error_rate, p95, error_burn_rate, ...
argNoOptional field the measure operates on, e.g. duration_ms
nameYesHuman-readable rule name (non-blank)
unitNoExplicit display unit for the measure, e.g. BYTES or DURATION_MS — set it when the metric name doesn't self-describe its unit (OTel names like jvm.memory.used or http.server.request.duration carry no unit suffix); omit to let the server infer the unit from the metric name or measure function
checkNoCompound AND/OR check tree as JSON, in the same shape get_alert_rules returns in checkDetails.check. Supply instead of the structured measurement and condition fields, which are ignored when this is set.
activeNoCreate-only: start the rule ACTIVE (default true) or DISABLED
filterNoOptional QuerySQL boolean filter, e.g. service = 'my-svc'
paramsNoOptional named measure params, e.g. {"budget":"0.001"}
ruleIdNoRule id (UUID) to update; omit to create a new rule
sourceNoTelemetry source: LOGS, SPANS, METRICS
groupByNoOptional plain field names to group the series by, e.g. service. Field names only: a string that looks like an expression is rejected, pass it in derivedGroupBy instead
directionNoAnomaly direction: HIGH or LOW
comparatorNoThreshold comparator: GT, GTE, LT, LTE (static rules)
expressionNoOptional free-form aggregate expression measuring the source, used instead of fn (takes precedence when set). QuerySQL over the source's fields, e.g. a ratio 'countIf(status_code = ''ERROR'') * 100.0 / count()' (this is exactly fn: error_rate, which already returns 0-100 — don't divide by 100 again) or a metric ratio 'avg(if(metric_name = ''a'', value, null)) / avg(if(metric_name = ''b'', value, null))'.
metricNameNoMetric name (required only when source is METRICS)
metricTypeNoMetric type: GAUGE, SUM, HISTOGRAM, ... (only when source is METRICS)
descriptionNoOptional free-text description
windowMinutesNoRolling window length in minutes
derivedGroupByNoOptional computed group keys, each an object with an expression and a label, e.g. {"expression": "regexp_extract(message, 'customerId=([0-9a-f-]+)', 1)", "label": "customer"}. The label is what rule pages and alert titles show
notifyOnResolveNoNotify when the alert resolves (default true)
zScoreThresholdNoAnomaly z-score threshold (> 0) — supply instead of comparator/warningThreshold
warningThresholdNoWarning-tier threshold (static rules)
criticalThresholdNoOptional critical-tier threshold (escalation)
anomalyConsecutiveWindowsNoConsecutive anomalous windows required (>= 1)
warningConsecutiveWindowsNoConsecutive breaching windows for the warning tier (default 1)
criticalConsecutiveWindowsNoConsecutive breaching windows for the critical tier (default 1)

Schema Changelog

Changes observed during successful MCP inspections.

  1. Added
  2. Removed
  3. Changed2 schema fields changed
    • addedInput schema / properties / derivedGroupBy
      Added value: +{
      +  "description": "Optional computed group keys, each an object with an expression and a label, e.g. {\"expression\": \"regexp_extract(message, 'customerId=([0-9a-f-]+)', 1)\", \"label\": \"customer\"}. The label is what rule pages and alert titles show",
      +  "items": {
      +    "properties": {
      +      "expression": {
      +        "description": "QuerySQL expression whose value the series are grouped by, e.g. regexp_extract(message, 'customerId=([0-9a-f-]+)', 1)",
      +        "type": "string"
      +      },
      +      "label": {
      +        "description": "Short readable name for the expression, e.g. customer. This is what rule pages and alert titles show, never the expression itself.",
      +        "type": "string"
      +      }
      +    },
      +    "required": [
      +      "expression",
      +      "label"
      +    ],
      +    "type": "object"
      +  },
      +  "type": "array"
      +}
    • changedInput schema / properties / groupBy / description
      Previous value: -"Optional fields to group the series by"New value: +"Optional plain field names to group the series by, e.g. service. Field names only: a string that looks like an expression is rejected, pass it in derivedGroupBy instead"
  4. Changed1 schema field changed
    • removedInput schema / properties / channelIds
      Removed value: -{
      -  "description": "Create-only: registry channel UUIDs (the id field from list_alert_channels, not a Slack/provider channel id) to route this rule's alerts to; omit to inherit the tenant's default channel. Use set_alert_rule_delivery to reroute an existing rule.",
      -  "items": {
      -    "type": "string"
      -  },
      -  "type": "array"
      -}
  5. Changed1 schema field changed
    • addedInput schema / properties / check
      Added value: +{
      +  "description": "Compound AND/OR check tree as JSON, in the same shape get_alert_rules returns in checkDetails.check. Supply instead of the structured measurement and condition fields, which are ignored when this is set.",
      +  "type": "string"
      +}
  6. Changed1 schema field changed
    • addedInput schema / properties / unit
      Added value: +{
      +  "description": "Explicit display unit for the measure, e.g. BYTES or DURATION_MS — set it when the metric name doesn't self-describe its unit (OTel names like jvm.memory.used or http.server.request.duration carry no unit suffix); omit to let the server infer the unit from the metric name or measure function",
      +  "type": "string"
      +}
  7. First observed

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description bears full responsibility for behavioral disclosure, and it excels. It warns that update is a whole-object replace, not a merge, and lists the specific carve-outs for status and delivery. It details re-validation behavior, the difference between warnings and problems (with a concrete warning code FIELD_NEVER_OBSERVED), and how to interpret dataCoverage. It even explains edge cases like groupBy expression rejection and derivedGroupBy label semantics. This transparency is exceptional.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but appropriately sized for 26 parameters. It front-loads the most critical behavioral warning (whole-object replace) in the first paragraph. The use of section headers (METRIC RULE, ANOMALY METRIC, Common fields) and bullet-like indentation keeps it scannable. Every sentence provides actionable information; there's no fluff or redundancy. It's structured to maximize comprehension for a complex tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description covers the response semantics well: it mentions problems[] and warnings[], states that problems[] populates instead of rule on validation failure, and clarifies that warnings are advisory and never block saving. It also gives guidance on actionable next steps (fix spelling, re-preview). The description fully equips an agent to call, interpret, and troubleshoot. No critical gaps are apparent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Even though schema coverage is 100%, the description adds substantial semantic value beyond the schema. It explains relationships (expression vs fn), conditional requirements (metricName/metricType only for METRICS), format constraints (groupBy field names only), and the meaning of labels in derivedGroupBy. It clarifies defaults (notifyOnResolve, warningConsecutiveWindows), and warns about common pitfalls (don't divide by 100 again). This far exceeds the baseline for fully-covered schemas.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's core function: 'Create or update an alert rule for the calling tenant in one call.' It also distinguishes between metric and anomaly rules, and explicitly notes limitations ('Compound multi-condition rules can't be created here'). It differentiates itself from siblings like set_alert_rule_status and set_notification_destinations by carving out those concerns, which makes the purpose unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use and when-not-to-use guidance: it names preview_alert_rule for validation and calibration, set_alert_rule_status for status changes, set_notification_destinations for delivery routing, and the web editor for compound rules. It also provides concrete directives like 'use fn instead unless you need a custom ratio' and 'omit it to let the server infer the unit.' This is exemplary usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources