silentwatch-mcp
silentwatch-mcp
捕获那些监控系统无法察觉的 cron 失败。 一个 MCP 服务器,用于向任何 Claude 或支持 MCP 的代理展示定时任务状态——包括运行情况、逾期任务以及退出代码为 0 但未产生任何有效输出的静默失败。开箱即用,支持 OpenClaw 调度程序、系统 cron 和 systemd 定时器。
功能概述
每个运行定时任务的团队都至少遇到过以下情况之一:
静默失败 — 任务运行了,返回退出代码 0,但没有产生任何有效输出(例如:返回空结果的网页搜索 cron、写入 0 字节文件的备份、正文中包含
<no rows>的摘要邮件)。传统监控显示为绿色勾选,但数据实际上已经损坏。逾期无警报 — 任务已经 3 天没有运行;因为没人关注,所以没人发现。
最后成功时间漂移 — 任务每小时运行一次,但在过去 12 次尝试中仅成功了一次;每个人都认为它很健康,因为最近一次运行是绿色的。
审计追踪缺失 — 你需要知道某个特定任务最后一次完成的时间以进行合规性检查,而唯一的“日志”是上周已经轮转的
journalctl输出。
silentwatch-mcp 将这些可见性作为 MCP 工具暴露出来,供你的 AI 代理直接查询。无需指标管道,无需单独的仪表板,无需 SaaS 订阅。
> claude: which of my cron jobs have silent failures in the last 24 hours?
[MCP tool: find_silent_failures]
3 jobs flagged:
• web-search-refresh — ran 12× successfully but output empty in 8 (66% silent fail rate)
• daily-summary — ran 1× successfully (24× expected); output normal
• audit-snapshot — last success 5 days ago, all subsequent runs returned exit 0 with empty bodyRelated MCP server: task-orchestrator
为什么选择 silentwatch-mcp
现有工具(Cronitor、Healthchecks.io、Datadog、Prometheus)无法做到以下三点:
检测静默失败,而不仅仅是退出代码。 传统的 cron 监控假设
exit 0 = 成功。我们根据可配置的规则检查输出:空输出、长度相对于历史中位数的异常、尽管退出代码为 0 但 stdout 中仍包含错误关键字、持续时间异常。那些“运行成功”但没有返回任何有效内容的任务——这正是隐藏数周的失败模式。我们能捕捉到它。MCP 原生,无需集成层。 Claude Desktop、Cline、Continue、OpenClaw 代理——任何支持 MCP 的客户端都可以直接查询。无需 Grafana 插件,无需 API 包装器,无需手动解析 JSON。
开箱即用的多源支持。 OpenClaw 原生 JSONL 日志、系统 crontab (
/etc/crontab+/etc/cron.d/*+ 用户级crontab -l) 以及 systemd 定时器 (systemctl list-timers+journalctl)——所有四个后端都在 v0.3 中提供,因此你可以针对你拥有的任何调度程序运行silentwatch-mcp。没有供应商锁定。
专为运行 40 美元 VPS 的 SMB 自托管用户打造,对于他们来说 Datadog 太过昂贵,而“0 美元/月的开源 MCP”是合适的切入点——但静默失败检测在企业基础设施中同样具有价值。
工具界面
服务器注册了以下 MCP 工具(完整规范见 SPEC.md):
工具 | 功能 |
| 枚举所有已知的 cron 任务及最后运行摘要 |
| 单个任务的详细状态:最后运行、最后成功、窗口期内的成功率 |
| 近期运行历史,包含时间、状态和输出片段 |
| 按照计划应该运行但尚未运行的任务 |
| “成功”运行但输出看起来可疑的任务 |
| 单个任务的近期日志输出 |
资源:
cron://jobs— 所有任务列表(清单)cron://job/{id}— 单个任务清单 + 近期运行记录cron://run/{id}— 带有完整输出的单个运行实例
提示词:
diagnose-overdue— 用于逾期任务的诊断提示词模板summarize-cron-health— cron 活动 + 异常的每日摘要
快速入门
v0.3 beta — 已发布全部 4 个后端 + 通过 cron 计划解析 (croniter) 实现真正的逾期检测。 Mock、OpenClaw JSONL、crontab 和 systemd 后端均已达到生产就绪状态。74 个测试通过。v1.0 现已进入完善阶段:PyPI 发布 + GitHub Actions CI + MCP 注册表提交。
安装
pip install silentwatch-mcp # not yet on PyPI; install from source for now:
pip install -e .配置 Claude Desktop
添加到 ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) 或 %APPDATA%\Claude\claude_desktop_config.json (Windows):
{
"mcpServers": {
"silentwatch": {
"command": "python",
"args": ["-m", "silentwatch_mcp"],
"env": {
"SILENTWATCH_BACKEND": "mock"
}
}
}
}后端(v0.3 已发布全部四个):
SILENTWATCH_BACKEND=mock— 返回示例数据(开发默认值)SILENTWATCH_BACKEND=openclaw-jsonl— 解析 OpenClaw 的原生 cron 运行 JSONL 文件(设置SILENTWATCH_OPENCLAW_LOGS为目录,默认为~/.openclaw/cron-runs/);数据最丰富——包含完整运行历史 + 静默失败检测SILENTWATCH_BACKEND=crontab— 解析/etc/crontab+/etc/cron.d/*+ 用户 crontabs (crontab -l);最后运行时间从/var/log/syslog或/var/log/cron推断(设置SILENTWATCH_SYSLOG可覆盖)SILENTWATCH_BACKEND=systemd— 解析systemctl list-timers --all --output=json+journalctl -u <unit>以获取运行历史;将OnCalendar=提取到计划字段中
所有非 mock 后端在底层工具不存在的平台/主机上会优雅地返回空结果,因此在不同环境中保留配置是安全的。
重启 Claude Desktop
服务器注册为 silentwatch。测试:
显示我所有的 cron 任务及其最后运行状态。
路线图
版本 | 范围 | 状态 |
v0.1 | 协议连接,mock 后端,所有 6 个工具注册存根数据,测试通过 | ✅ 完成 |
v0.2 | 实现 OpenClaw JSONL 后端(真正的 cron 运行解析,格式错误行处理,静默失败增强) | ✅ 完成 (2026-05-02) |
v0.3 | Crontab + systemd 后端;用于真正逾期检测的 cron 计划解析 (croniter);35 个新测试 | ✅ 完成 (2026-05-02) |
v1.0 | 完善:PyPI 发布,GitHub Actions CI,MCP 注册表提交 (Glama + PulseMCP),细化静默失败规则配置 | ⏳ 第一阶段发布目标 (5月18日当周) |
v1.x | 附加后端(Cowork 调度程序,Claude Code 后台任务,通用 JSON 配置),用于警报的 webhook 发射器 | ⏳ 第二阶段及以后 |
需要适配你的技术栈?
silentwatch-mcp 附带 4 个后端(mock、OpenClaw JSONL、crontab、systemd)。如果你的调度程序是其他工具——AWS EventBridge、GCP Cloud Scheduler、Hangfire、Sidekiq、Temporal、Apache Airflow、Prefect、Dagster 或自定义任务运行器——并且你希望为其获得相同的静默失败检测 MCP 可见性,那属于 自定义 MCP 构建 服务。
层级 | 范围 | 投资 | 时间线 |
简单 | 为具有记录 API 的现有调度程序(例如 GCP Cloud Scheduler)提供单个后端适配器 | $8,000–$10,000 | 1–2 周 |
标准 | 自定义后端 + 自定义静默失败规则 + 与现有警报系统(PagerDuty、Slack 等)集成 | $15,000–$20,000 | 2–4 周 |
复杂 | 多后端(跨区域/集群/租户的联合 cron)+ RBAC + 审计日志集成 + 值班工作流 | $25,000–$35,000 | 4–8 周 |
参与方式:
发送邮件至 admin@pixelette.tech,主题为
Custom MCP Build inquiry包括:一段关于你的调度程序技术栈的描述 + 你正在考虑的层级
在 2 个工作日内回复,预约 30 分钟的发现通话
此服务器也是 AI 生产纪律框架 的一部分——这是我进行生产 AI 审计所依据的方法论。
生产 AI 审计
如果你正在运行生产级 AI,并希望外部从业者评估就绪情况、发现已存在的失败模式并编写纠正行动计划——这正是此 MCP 所支持的内容。独立审计服务:
层级 | 范围 | 投资 | 时间线 |
审计 Lite | 一个系统,前 5 大发现,书面报告 | $1,500 | 1 周 |
审计标准 | 全面审计,所有 14 种模式,5C 发现,90 天跟进 | $3,000 | 2–3 周 |
审计 + 工作坊 | 标准审计 + 2 天团队工作坊 + 包含首次月度审计 | $7,500 | 3–4 周 |
相同邮件渠道:admin@pixelette.tech,主题为 AI audit inquiry。
贡献
欢迎提交 PR。结构特意保持扁平,以便轻松添加自定义后端——请参阅 src/silentwatch_mcp/backends/ 获取现有示例。
添加新后端:
在
backends/<your_backend>.py中继承CronBackend实现
list_jobs、get_job_runs、tail_logs在
backends/__init__.py中注册在
tests/test_backend_<your_backend>.py中添加测试
错误报告 + 功能请求:请开启 GitHub Issue。
许可证
MIT — 请参阅 LICENSE。
相关
AI 生产纪律框架 — Notion 模板,$29
SPEC.md — 完整服务器设计
Model Context Protocol — 协议概述
由 Temur Khan 构建 — 生产 AI 系统独立从业者。 联系方式:admin@pixelette.tech
Available Tools
6 toolsfind_overdue_jobsA
Returns jobs whose schedule indicates they should have run but haven't, beyond a grace window.
| Name | Required | Description | Default |
|---|---|---|---|
| grace_minutes | No | Tolerance to avoid flagging jobs about to run (default 5) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the core behavior (returns overdue jobs) but does not mention whether the operation is read-only, performance implications, or pagination. Adequate but not thorough.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that is efficient and front-loaded with the core purpose, no redundant words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with one optional parameter and no output schema, the description is largely complete. It could clarify what 'schedule indicates' means or the output format, but overall it provides sufficient context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema coverage, the description adds value by explaining 'beyond a grace window' which directly connects to the grace_minutes parameter, providing context beyond the schema's technical description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool returns overdue jobs with a grace window, distinguishing it from siblings like find_silent_failures (different failure mode) and list_jobs (all jobs).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for checking missed jobs but lacks explicit when-to-use or when-not-to-use guidance, nor does it mention alternatives among siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_silent_failuresC
Jobs that returned exit code 0 but output was flagged by silent-fail rules (empty output, length anomaly, error keywords, duration anomaly).
| Name | Required | Description | Default |
|---|---|---|---|
| window_hours | No | Lookback window in hours (default 24) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided; description only lists detection criteria. It does not disclose behavioral traits like read-only nature, prerequisites, or potential side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence efficiently conveys purpose but lists multiple anomaly types in a somewhat dense manner. No wasted words, but readability could improve.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given one optional parameter and no output schema, the description lacks details on return format, pagination, or usage context. Incomplete for a search/filter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with clear description for window_hours. The tool description adds no extra meaning beyond the schema, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it finds jobs with exit code 0 flagged by silent-fail rules, using specific verb and resource. However, it does not differentiate from siblings like find_overdue_jobs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool vs alternatives. The description implies usage for detecting silent failures but lacks when-not or alternative tool references.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_job_runsA
Recent run history for a job (newest first) with timing, exit code, status, silent-fail indicators, output snippet.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | Job identifier | |
| limit | No | Max runs to return (default 20, max 500) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so the description carries the full burden. It discloses ordering, data fields, and indicators (e.g., silent-fail), but does not mention side effects, rate limits, access requirements, or return format specifics. This is adequate but not thorough for a tool with no annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence with all key information front-loaded (recent, newest first, data fields). No wasted words; every part adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so description should explain return structure. It lists fields but does not specify if results are an array, pagination behavior (beyond limit), or error handling. Adequate for a simple list but incomplete for comprehensive understanding.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description does not add meaningful information beyond the schema: 'job_id' and 'limit' are already described in the schema with default and max values. No additional context for parameter usage is provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly specifies the resource ('run history for a job'), action ('get'), ordering ('newest first'), and included data fields ('timing, exit code, status, silent-fail indicators, output snippet'). It effectively distinguishes from sibling tools like get_job_status or tail_job_logs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives like get_job_status or find_silent_failures. The description implies usage for recent runs but lacks when-not-to-use conditions or comparisons.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_job_statusA
Detailed status for one job: last run, last success, success rates over 24h + 7d, overdue state, silent-fail indicators on the last run.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | Job identifier from list_jobs |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided; the description describes return fields but does not disclose any behavioral traits such as read-only nature, authentication needs, or cost. For a read-like tool, this is a notable gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The entire description is a single concise sentence that front-loads the core purpose and lists specific details without any wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's low complexity (one parameter, no output schema), the description adequately covers return values. It could mention error handling or prerequisites, but is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema has 100% description coverage with 'job_id' documented. The description adds no further parameter information, so a baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Detailed status for one job' and enumerates specific fields (last run, success rates, overdue state, silent-fail indicators), distinguishing it from siblings like find_overdue_jobs or get_job_runs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for obtaining a comprehensive snapshot of a single job's health, but does not explicitly state when to use this tool over alternatives or provide exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_jobsA
Enumerate all known cron jobs with last-run summary. Returns id, name, schedule, last run time + status, runs/successes in last 24h, silent-fail count, overdue flag.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Despite lacking annotations, the description transparently lists all returned fields, including last-run summary and overdue flags. This adequately discloses the read-only behavior and output structure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that efficiently conveys the tool's purpose and output. No extraneous information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's low complexity (no parameters, no output schema), the description covers its functionality and output comprehensively. It could mention it as a read-only operation, but the field list suffices.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With zero parameters, the description adds value by detailing the output fields, compensating for the absence of an output schema. It provides richer semantics than the empty input schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool enumerates all known cron jobs with a last-run summary, specifying the returned fields (id, name, schedule, last run time + status, etc.). It distinguishes itself from sibling tools like find_overdue_jobs and find_silent_failures, which target specific subsets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
While the description implies general-purpose enumeration, it does not explicitly state when to use this tool versus siblings like get_job_status or get_job_runs. No exclusion criteria or alternative recommendations are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tail_job_logsB
Most recent N log lines for a job (newest last).
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | ||
| lines | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. Only states the result order and count, but omits read-only hint, error handling, or limitations like max lines.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, no wasted words. Efficient but could include more contextual info without becoming verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Adequate for a simple tool with 2 parameters and no output schema. Lacks details on behavior for edge cases and result format, but core purpose is clear.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%. Description hints at 'lines' parameter ('N log lines') but does not explain 'job_id' or provide format/constraints for either parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the action ('Most recent N log lines'), resource ('a job'), and ordering ('newest last'). Distinguishes from siblings like get_job_status or list_jobs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool vs siblings. Does not mention exclusions or alternatives despite related tools (e.g., find_silent_failures, get_job_runs).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v0.3.0- First observed
find_overdue_jobs - First observed
find_silent_failures - First observed
get_job_runs - First observed
get_job_status - First observed
list_jobs - First observed
tail_job_logs
TDQS
Scored across 6 tools
Each tool addresses a distinct aspect of job monitoring: overdue detection, silent failure detection, run history, job status, job listing, and log tailing. No overlap in purpose.
All tool names follow a consistent verb_noun pattern with underscores, using clear verbs like find, get, list, and tail. No mixing of conventions.
Six tools cover the core functionality of a job monitoring server: listing, anomaly detection, status, logs. Neither too few nor too many for the scope.
The tool set provides complete coverage for monitoring cron jobs: discovery, anomaly detection (overdue and silent failures), status, history, and logs. No obvious gaps for a read-only monitoring use case.
Maintenance
Related MCP Connectors
Monitoring that agents set up for themselves: cron jobs, CI/CD pipelines and AI agent runs.
Control plane for autonomous software labor. Agents claim objectives over MCP with audit trail.
Personal assistant MCP server with search, execute, packages, jobs, secrets, and integrations.
MCP server for AI agents to plan, verify, and deploy Cloudflare-native apps.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceAn MCP server for scheduling and executing Claude Code CLI tasks via cron expressions, featuring a web dashboard and webhook support. It enables users to dynamically create custom MCP servers, manage recurring AI jobs, and track execution history with token and cost analytics.24MIT
- AlicenseNot gradedqualityAmaintenanceServer-enforced workflow discipline for AI agents. An MCP server providing persistent work items, dependency graphs, quality gates, and actor attribution. Schemas define what agents must produce — the server blocks the call if they don't. Works with any MCP-compatible client.207MIT
- AlicenseNot gradedqualityDmaintenanceA hosted remote MCP server that lets your AI agent schedule tasks for later — reminders, delayed webhook callbacks, and recurring jobs. Read-only by design.1 npmMIT

elvatis-mcpofficial
AlicenseBqualityAmaintenanceMCP server for OpenClaw that enables AI clients to control smart home devices, manage memory and cron jobs, and orchestrate multiple AI sub-agents.3752 npmApache 2.0