Telegram OSINT
Integrates with a user's Telegram account to list dialogs, read and search message history, summarize unread messages, and monitor incoming messages via local rules and judging criteria. It can generate alerts for matching messages, with no Telegram write operations.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Telegram OSINTsearch my telegram chats for anything about the project deadline"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Telegram OSINT
An MCP server for your own Telegram account. Read your chats through an AI assistant, and run a background monitor that watches every incoming message for things you care about.
The monitor works in two stages: cheap rules (substring / whole-word / regex) decide what's worth looking at, then Claude judges each hit against your plain-language criteria and decides whether it's worth surfacing. On a firehose of busy channels, stage one keeps the cost down and stage two keeps the noise down. The judge is optional — without it, every rule hit becomes an alert.
Everything runs locally. Nothing leaves your machine except calls to Telegram's API and, if you enable judging, Anthropic's.
What it runs and sends
Local server. The plugin starts a Python MCP server on your computer with
uv run --locked. On first start, uv downloads Python and the exact package versions pinned inuv.lockfrom PyPI.Telegram. The server and the watcher log in to your Telegram account (a user session, not a bot) and talk to Telegram's servers to list chats, read and search messages, and receive new ones. Nothing is ever sent, deleted, or left on Telegram — there are no write tools.
Anthropic. Only when judging is on and
ANTHROPIC_API_KEYis set, the watcher sends the text of each rule-matched message, plus your criteria, to Anthropic's API for a verdict. Messages that match no rule are never sent anywhere.Local files. Session files, your API credentials (
.env, mode0600), the alerts database, and the watcher log live in~/.telegram-mcp/. Nothing is uploaded elsewhere.Background watcher. A separate process you start yourself; the plugin never starts it for you. On macOS it can show desktop notifications via
osascript.
Related MCP server: tdl-mcp
Where it works
Claude app | Works |
Claude Code (terminal, IDE, desktop app Code tab) | Yes |
Cowork in the desktop app, on your computer | Yes |
Chat on claude.ai web, desktop, or mobile | Skills only — chat can't start a local server, so the Telegram tools aren't available |
Install
Requires uv. It provisions the right Python
version itself, so you don't need to manage one. On macOS, brew install uv;
for other systems, see uv's
installation guide.
As a Claude Code plugin — this is the easy path, and gives you the setup skill:
claude plugin marketplace add alamri-intel/telegram-mcp
claude plugin install telegram-osint@telegram-osintOr from inside Claude Code: /plugin marketplace add alamri-intel/telegram-mcp,
then /plugin install telegram-osint@telegram-osint.
Start a new session, then run /telegram-osint:setup and answer the
questions — it writes your rules and criteria for you.
Or as a plain Python package, if you'd rather not use the plugin:
uv tool install "telegram-monitor-mcp[judge] @ git+https://github.com/alamri-intel/telegram-mcp"
claude mcp add telegram -- telegram-mcpThe judge extra pulls in the Anthropic SDK. Omit it if you only want rule
matching.
You need a Telegram API_ID / API_HASH pair from
https://my.telegram.org (API development tools). These identify the app,
not your account.
export TELEGRAM_API_ID=...
export TELEGRAM_API_HASH=...
export ANTHROPIC_API_KEY=... # only if you want judgingStart the watcher
telegram-mcp-watcherThe first run prompts for your phone number and the login code Telegram sends. After that the saved session is reused and startup is silent. Leave it running — alerts are only recorded while it's up.
To run it in the background:
nohup telegram-mcp-watcher > ~/.telegram-mcp/watcher.log 2>&1 &Registering the server by hand
The plugin does this for you. If you installed the plain package instead:
claude mcp add telegram -- telegram-mcpClaude Desktop — add to claude_desktop_config.json:
{
"mcpServers": {
"telegram": {
"command": "telegram-mcp",
"env": {
"TELEGRAM_API_ID": "...",
"TELEGRAM_API_HASH": "..."
}
}
}
}The server and the watcher use separate Telegram sessions on purpose: one session file can only be used by one process at a time. A side effect worth knowing is that the monitor tools work even if the server itself is never logged in — they only read the local database.
Tools
Monitoring — local database only, no Telegram login needed
Tool | |
| Create a prefilter rule. Live within ~5s. |
| Rules with hit counts. |
| Pause or resume a rule. |
| Delete a rule and its alerts. |
| Add a plain-language judging criterion. |
| Criteria with hit counts. |
| Pause or resume a criterion. |
| Delete a criterion. |
| What matched and what the judge decided. |
| Clear alerts from the default view. |
| Watcher liveness, counts, judge throughput. |
Skills (plugin install only)
/telegram-osint:setup interviews you about what you watch for and writes the
rules and criteria. /telegram-osint:watcher covers starting the daemon and
working out why alerts aren't arriving.
Reading — talks to Telegram
list_dialogs, read_messages, search_messages, unread_summary,
preview_rule. list_dialogs is how you find the numeric chat ids for rule
scopes; search_messages searches history, which rules never see;
preview_rule tests a candidate rule against real message history before you
create it.
Login — sets up your Telegram session
Tool | |
| Whether credentials are saved and each session is logged in. |
| Save your Telegram app credentials to |
| Ask Telegram to send a login code. |
| Finish login; the session file is saved locally. |
Rules and criteria
They answer different questions. A rule asks "does this string appear?" A criterion asks "does this matter?"
Rules:
match_type—substring(default),word(whole words only, sodeploywon't fire onredeployment), orregexcase_sensitive— off by defaultchat_ids/exclude_chat_ids— scope it; omitchat_idsto watch everything. Exclusions apply first.include_outgoing— off by default, so your own messages don't matchnotify— desktop notification on a surfaced alert (macOS)
Criteria are prose. Write them the way you'd brief a person doing the triage for you — say what qualifies and what doesn't, because near-misses are where a judge earns its keep:
remote-backend-roles — A specific open role for a backend engineer that is remote or remote-friendly in Europe, posted by someone hiring for it. Must name the company and the role. Excludes: recruiter cold-calls with no named company, roles requiring relocation, and anyone advertising their own availability.
Nothing about the subject matter is built in — the criteria supply it. The same machinery works for job leads, release announcements, a research field, or a resale market.
Each alert keeps the judge's verdict, severity, one-line summary, and its
reasoning — including why it rejected something. list_alerts hides
rejected alerts by default; pass include_irrelevant=True to audit them.
Gotchas
Rules are forward-looking. The watcher only sees messages that arrive while it's running. Use
search_messagesfor history.Check
monitor_status()first. If the watcher died, the alert tools keep answering from a stale database. Status is what tells you. It reports the watcher down when its heartbeat is over 90 seconds old.Judging is best-effort. If the API call fails, the alert is marked
errorand surfaced anyway — dropping a possible hit is worse than a false positive. Repeated auth failures disable judging for that run rather than retrying on every message.A broken rule is skipped, not fatal. An invalid regex logs a warning; the other rules keep running.
Configuration
Variable | Default | |
| — | required |
| — | required for judging |
|
| database and session files |
|
| database path |
|
| server's session name |
|
| watcher's session name |
|
|
|
|
| judging model |
|
|
|
|
|
|
|
| watcher log level |
Security
Session files under
~/.telegram-mcp/(mode0700) are equivalent to login credentials for your Telegram account. Never commit or share one.The database holds the full text of every message that matched a rule.
These are user sessions, not bots — the watcher sees everything you see. Scope rules with
chat_idsif you'd rather not store message text from every chat you're in.The server exposes no Telegram write operations: nothing here can send, delete, or leave anything. Add write tools deliberately and narrowly if you need them.
License
MIT
Available Tools
20 toolsack_alertsA
Mark alerts as acknowledged so they drop out of the default list_alerts view. Pass exactly one of the three selectors.
Args: alert_ids: specific alert ids to acknowledge. rule_id: acknowledge every unacked alert from this rule. all_alerts: acknowledge every unacked alert.
| Name | Required | Description | Default |
|---|---|---|---|
| rule_id | No | ||
| alert_ids | No | ||
| all_alerts | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose the bulk blast radius (rule_id and all_alerts affect 'every unacked alert') and the mutual-exclusivity constraint, but says nothing about reversibility/unack, permissions, idempotency, or error behavior when the selector rule is violated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The effect and the selector constraint are front-loaded in the first two sentences, then the Args list adds only what is needed. No filler sentences and nothing duplicated from the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-param mutation with no annotations and no output schema, the description covers purpose, effect, and all selector semantics, which is most of what an agent needs. It omits confirmation/return behavior and reversibility, the remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does: each of the three parameters is given meaning beyond its title, including the critical mutual-exclusivity contract. It stops short of stating value formats or what happens with an empty list.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('mark alerts as acknowledged') and names the observable effect on a named sibling view ('drop out of the default list_alerts view'). An agent can distinguish this from list_alerts without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit selection rule ('Pass exactly one of the three selectors') plus the scope of each selector, which is real when-to-use guidance for this family of tools. It never names a sibling alternative or an explicit when-not case, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_alert_ruleA
Create an alert rule. The watcher picks it up within a few seconds — no restart needed. Rules only match messages that arrive after they are created; use search_messages to look backwards.
Args: name: unique label for the rule, used in alerts and notifications. pattern: what to look for. match_type: 'substring' (default), 'word' (whole-word only), or 'regex'. case_sensitive: match case exactly (default: case-insensitive). chat_ids: only watch these chats; omit to watch every chat. exclude_chat_ids: never match in these chats; applied before chat_ids. include_outgoing: also match messages you send (default: incoming only). notify: fire a macOS desktop notification on match.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| notify | No | ||
| pattern | Yes | ||
| chat_ids | No | ||
| match_type | No | substring | |
| case_sensitive | No | ||
| exclude_chat_ids | No | ||
| include_outgoing | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses pickup latency ('within a few seconds — no restart needed') and the critical temporal constraint that rules only match future messages. It still leaves permissions/auth requirements and behavior on duplicate names unstated, so it stops short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads purpose and the two highest-value behavioral facts (instant pickup, forward-only matching) before the structured Args list. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description covers behavior, timing, alternatives, and every parameter. It does not state what the call returns (e.g., a rule id) or any auth prerequisite, which keeps it just under complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, yet the Args section documents all 8 parameters with meaningful semantics: match_type options, case-sensitivity default, chat-scoping behavior, exclude-before-include ordering, and include_outgoing default. This fully compensates for the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Create an alert rule') that is immediately distinguishable from siblings like list_alert_rules, delete_alert_rule, and set_alert_rule_enabled. The follow-on behavioral sentence clarifies it further, leaving no ambiguity about the operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent to search_messages for backwards lookups, which is the key alternative-use case, and frames the rule's forward-only scope. However, it omits the sibling preview_rule, which an agent would plausibly want before committing a pattern, so guidance is clear but not comprehensive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
add_criterionA
Add a natural-language criterion the judge uses to decide whether a prefiltered post actually matters. The watcher picks it up on the next judged alert — no restart needed.
Rules decide what gets looked at; criteria decide what gets surfaced. Write a criterion the way you'd brief someone doing the triage for you: say what qualifies and what doesn't. "A specific open remote backend role that names the company" beats "jobs".
Args: name: short unique label, shown on matching alerts. description: what qualifies, in plain language.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| description | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses activation semantics ('the watcher picks it up on the next judged alert — no restart needed'), which is genuine value. However it says nothing about whether a new criterion starts enabled (a set_criterion_enabled sibling implies that matters), whether duplicates are rejected, persistence, or permissions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then the rules-vs-criteria distinction, then authoring guidance, then args. Every sentence adds information rather than restating the name, though the prose is slightly longer than strictly necessary for a two-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with no annotations and no output schema, the definition supplies purpose, activation timing, vocabulary disambiguation, and both parameter semantics — enough to call it correctly. The notable omission is the default enabled/disabled state of a newly added criterion.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does: 'name' is explained as a 'short unique label, shown on matching alerts' (adding uniqueness and display semantics absent from the bare string schema) and 'description' as 'what qualifies, in plain language.' Both parameters gain meaning, though the effect of a non-unique name is not specified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (add a criterion) and immediately clarifies the mechanism: 'the judge uses to decide whether a prefiltered post actually matters.' It explicitly distinguishes criteria from the sibling concept of rules ('Rules decide what gets looked at; criteria decide what gets surfaced'), so an agent can place it against add_alert_rule and list_criteria without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly frames when this tool applies (adding a triage criterion) and how it differs from rules, plus concrete guidance on how to author the criterion text with a good/bad example. It stops short of naming alternatives such as preview_rule for testing a criterion before adding it, or list_criteria for reviewing existing ones.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
auth_statusA
Check whether Telegram is logged in, and as whom.
Reports both sessions: the watcher's (the one that must be logged in for monitoring to work) and this server's (needed only for the read tools). Call this before anything else if alerts or reads are failing.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden, and it delivers meaningfully: it explains that two sessions are inspected, that the watcher session is required for monitoring, and that this server's session is only for read tools. For a zero-parameter read-only check the safety profile is essentially self-evident via 'Check', so nothing critical is missing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose in the first clause, then adds only the two-session context and the call condition. Every sentence earns its place with no repetition of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description is the only source of return-value semantics, and it supplies them by naming the two sessions and their roles. It does not spell out the exact status values or failure shapes, but for a trivial status check this is nearly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4 and there is no parameter meaning to add or omit. The description appropriately spends its words on what is reported rather than on inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (check Telegram login status) and goes further by specifying the two distinct sessions it reports and what each one governs. An agent can distinguish this diagnostic from the sibling login_* and monitor_status tools without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit triggering condition: 'Call this before anything else if alerts or reads are failing.' It does not name alternative tools or describe when-not-to-call, but the failure-diagnosis context is clear and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_alert_ruleA
Delete an alert rule and every alert it produced. Irreversible — prefer set_alert_rule_enabled(rule_id, False) to just pause it.
Args: rule_id: id from list_alert_rules.
| Name | Required | Description | Default |
|---|---|---|---|
| rule_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full behavioral burden and does so well for the critical trait: it declares the destructive cascade (all produced alerts) and that the operation is irreversible. It does not cover permission/auth requirements or failure behaviour, so it falls short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, irreversibility and the safer alternative front-loaded ahead of the Args note. Every clause earns its place and nothing repeats structured data.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter destructive tool with no annotations and no output schema, the description supplies the two things an agent most needs: the blast radius and the reversible alternative. It omits what the call returns and any auth/permission requirement, leaving a small but real gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% — the schema only offers a bare 'Rule Id' integer — so the description must compensate, and it does by tying rule_id to list_alert_rules. It adds provenance meaning beyond the schema, though no format/validity detail (e.g. what an invalid id does).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Delete an alert rule') and immediately widens the scope to the cascade ('and every alert it produced'), which is the distinguishing detail versus sibling delete_criterion or set_alert_rule_enabled. An agent can tell this apart from its siblings without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the preferred alternative — set_alert_rule_enabled(rule_id, False) — and the condition that selects it ('to just pause it'). It also tells the agent where to obtain the argument (list_alert_rules), closing the loop on invocation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_criterionA
Delete a criterion. Alerts already judged against it keep their verdicts. Prefer set_criterion_enabled to pause one.
Args: criterion_id: id from list_criteria.
| Name | Required | Description | Default |
|---|---|---|---|
| criterion_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It does disclose one genuinely non-obvious effect — existing alert verdicts are preserved rather than retroactively invalidated — but says nothing about irreversibility, confirmation requirements, or required permissions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences plus the arg note; the destructive consequence is front-loaded and the alternative is stated immediately after. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter delete with no output schema, the description covers purpose, the alternative, and the main side effect. The remaining gap — whether deletion is permanent and whether confirmation is needed — is minor but real given zero annotation coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It names criterion_id and its provenance ('id from list_criteria'), which is exactly the lookup an agent needs; it stops short of format or validation details the schema also omits.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Delete a criterion') and explicitly names the sibling alternative (set_criterion_enabled) for the adjacent pause case, so an agent can distinguish intent without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear when-not-to-use rule: 'Prefer set_criterion_enabled to pause one.' It does not describe a positive trigger condition for deleting (e.g., obsolete or superseded criteria), but the alternative routing is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_alert_rulesB
List alert rules with their hit counts.
Args: include_disabled: include rules that are currently turned off.
| Name | Required | Description | Default |
|---|---|---|---|
| include_disabled | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. The verb 'List' implies a read-only operation and it adds that hit counts are returned, but it says nothing about permissions, result size, or ordering. Annotations covering the safety profile would have lowered the bar; here the coverage is only partial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with the core purpose front-loaded and the argument explanation following. No filler, though the 'Args:' formatting is slightly verbose for a single optional flag.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. For a simple read-only list tool with one optional boolean, the description is nearly sufficient, missing only a note on default behavior (default is true) and usage context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does by explaining that include_disabled brings in rules that are 'currently turned off' — clarifying what 'disabled' means beyond the bare boolean. Only one parameter exists and it is adequately explained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (list alert rules) plus a distinguishing detail (their hit counts), which separates it from the similarly named list_alerts and list_criteria. It does not explicitly name which sibling to use instead, so it stops short of 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no mention of prerequisites, and no reference to alternative tools such as list_alerts or preview_rule. The agent must infer the context entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_alertsA
List recorded alerts, newest first.
By default this hides alerts the judge ruled irrelevant — that filtering is the whole point of the judge. Pass include_irrelevant=True to audit what it rejected and why (each one keeps its reasoning).
Args: limit: max number of alerts to return. unacked_only: only alerts you haven't acknowledged yet (default). rule_id: restrict to one prefilter rule. chat_id: restrict to one chat. since_hours: only alerts matched within this many hours. verdict: exact verdict — relevant, irrelevant, pending, skipped, error. min_severity: lowest severity to include (low, medium, high, critical). include_irrelevant: include alerts the judge rejected.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| chat_id | No | ||
| rule_id | No | ||
| verdict | No | ||
| since_hours | No | ||
| min_severity | No | ||
| unacked_only | No | ||
| include_irrelevant | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does disclose meaningful non-obvious behavior: default filtering of judge-rejected alerts, default unacked_only, and that rejected alerts retain their judge reasoning. It omits auth/permission requirements and pagination behavior, so it is strong but not complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and the critical default-filtering nuance are front-loaded, then parameters are listed compactly. The Args block repeats parameter names already visible in the schema, but each line adds real meaning, so waste is minimal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists, so return-shape explanation is unnecessary. All 8 optional parameters are covered and the default-filtering semantics are explained, leaving only peripheral gaps like pagination and permission requirements.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the Arg list fully compensates by documenting all 8 parameters, including enumerated values for verdict (relevant, irrelevant, pending, skipped, error) and min_severity (low, medium, high, critical) that the schema does not encode. Definitions like 'limit: max number to return' are terse but adequate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List recorded alerts') plus ordering ('newest first'), and immediately clarifies the non-obvious scope of what is returned by default. An agent can distinguish it from siblings like list_alert_rules or ack_alerts without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the key use-case split: default hides judge-rejected alerts, and include_irrelevant=True is for auditing what was rejected and why. It does not explicitly name alternative sibling tools (e.g., ack_alerts) for adjacent tasks, so it stops short of full when/when-not routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_criteriaA
List the judging criteria and how often each has matched.
Args: include_disabled: include criteria that are currently turned off.
| Name | Required | Description | Default |
|---|---|---|---|
| include_disabled | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose a real behavioral trait beyond the name — the response includes match counts per criterion — but says nothing about permissions, ordering, result size, or that the listed criteria feed into rules/alert evaluation. Behavior is partially, not fully, disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence states purpose and scope, followed by a short param note. The only waste is the 'Args:' formatting convention, which is slightly redundant with the schema but harmless.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the single parameter is covered in prose. The remaining gap is contextual: without annotations or prose, the agent learns nothing about ordering, limits, or how the disabled-criterion state ties to set_criterion_enabled.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does by explaining that include_disabled pulls in criteria that are currently turned off. It omits the schema's default of true, which is a minor gap but leaves the practical 'what happens if I omit it' question partly unanswered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('List the judging criteria') plus a scope detail ('how often each has matched'), which distinguishes it from the mutating siblings add_criterion, delete_criterion and set_criterion_enabled. It never names a sibling explicitly, so an agent must infer the routing from verb semantics alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by the verb 'List'. There is no statement of when to call this versus browsing criteria another way, nor any prerequisite or context such as needing auth first. For a simple enumeration tool this is adequate but thin.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_dialogsA
List your Telegram chats/channels, most recently active first.
Use this to find the numeric chat ids you need for alert rule scopes.
Args: limit: max number of dialogs to return. unread_only: only return dialogs with unread messages. include_archived: include archived chats (excluded by default).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| unread_only | No | ||
| include_archived | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does disclose a real behavioral trait — archived chats are excluded by default and can be re-included — but says nothing about read-only safety, authentication/prerequisite state, rate limits, or result ordering/pagination limits beyond the ordering note.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose and routing hint in the first two sentences, then a short Args block. It is appropriately sized; the only minor redundancy is that the Args list repeats information inherent in the parameter names.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the description covers the tool's behavior and all inputs adequately for a read-only listing call. It stops short of noting pagination/truncation behavior or prerequisite auth state, which an agent might want when a long dialog list is expected.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it documents all three parameters (limit as max dialogs, unread_only as unread filter, include_archived with its default of excluded). The explanations are brief and largely restate the names, but they cover every parameter and supply the non-obvious default behavior of include_archived.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource ('List your Telegram chats/channels') and adds the ordering ('most recently active first'), which no sibling tool provides. An agent can immediately distinguish it from search_messages, read_messages, or unread_summary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit purpose in context: 'Use this to find the numeric chat ids you need for alert rule scopes,' which ties it to the alert-rule workflow. It does not, however, name alternatives or state when-not-to-use (e.g., when you already know the chat id or want message content, use read_messages/search_messages).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
login_request_codeA
Start logging in: ask Telegram to send a login code.
Telegram delivers the code in the Telegram app itself if the user is signed in elsewhere, otherwise by SMS. Follow with login_submit_code.
Ask the user for their phone number rather than guessing it.
Args: phone: phone number in international format, e.g. +14155550123. session: which session to log in — 'watcher' (default, the one monitoring needs) or 'mcp' (for this server's read tools).
| Name | Required | Description | Default |
|---|---|---|---|
| phone | Yes | ||
| session | No | watcher |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses how the code is delivered (in-app if signed in elsewhere, otherwise SMS) and what each session value means. It omits failure modes such as invalid numbers or whether a new request invalidates a previous code, so it falls short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and sequencing, then a compact Args block. Every sentence carries information, though the two-session explanation is somewhat verbose relative to its size.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description covers the call's purpose, follow-up step, and both parameters, which is sufficient for an agent to invoke it correctly. Error and retry behavior is not covered, keeping it just under a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate fully. It specifies the phone format with a concrete example ('+14155550123') and enumerates the session values with their semantics and default ('watcher' for monitoring, 'mcp' for read tools) — all information absent from the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Start logging in: ask Telegram to send a login code') and explicitly names the next sibling tool (login_submit_code). An agent can distinguish this from login_submit_code and auth_status without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear sequencing ('Follow with login_submit_code') and a strong usage instruction ('Ask the user for their phone number rather than guessing it'). It does not state a when-not condition (e.g., what to do if already signed in via auth_status), so it stops short of full exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
login_submit_codeA
Finish logging in with the code Telegram sent.
If the account has two-factor authentication, this returns an error asking for the password; call it again with both.
Args: code: the login code the user received. password: two-factor password, only if the account has one. session: the session being logged in; must match login_request_code.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | ||
| session | No | watcher | |
| password | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the behavioral burden. It usefully discloses the 2FA error-and-retry flow, which an agent could not otherwise infer. However it omits what a successful call returns, whether the session persists, and any rate-limit or failure behavior beyond the 2FA case.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and the retry rule, then a compact Args block. Every line carries meaning, though the Args formatting is slightly heavier than needed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-param login tool with no annotations or output schema, the definition covers the action, the error path, and all parameters including the session linkage. The main remaining gap is the nature of the successful response and the session default.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must document the three params, and it does: 'code' is the received login code, 'password' is the 2FA password, and 'session' is the session being logged in and must match login_request_code. The session cross-reference is a genuine constraint the schema does not express, though the 'watcher' default is left unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Finish logging in with the code Telegram sent') and the second sentence clarifies the two-factor path. The phrasing 'Finish logging in' implicitly separates it from the sibling login_request_code, which starts the flow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the retry condition clearly: if 2FA is enabled, the first call errors asking for the password and the tool must be called again with both. It also ties session to login_request_code, implying ordering. It stops short of an explicit 'call this after login_request_code' instruction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
monitor_statusA
Check whether the watcher daemon is running and what it has seen.
Reports liveness from the daemon's heartbeat, so a 'stale' result means watcher.py has stopped and no new alerts are being recorded.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full burden. It usefully discloses that liveness comes from a daemon heartbeat and that a 'stale' result indicates watcher.py has stopped, but it omits auth requirements, polling/caching behavior, and the shape of the returned status.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core action and followed by the one piece of interpretive context. Little waste, though the second sentence could be tightened slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description should carry more of the return-value burden for a status tool. It explains the meaning of a 'stale' result but does not describe the other possible states or the fields the agent should expect to read.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline of 4 applies. No parameter insight is missing because none exist.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Check whether the watcher daemon is running', which is clearly distinct from sibling tools like auth_status or list_alert_rules. The only softness is 'what it has seen', which is vague about the return payload, and no sibling is named for contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the note that a 'stale' result means 'no new alerts are being recorded' hints this is a diagnostic tool for missing alerts, but there is no explicit when-to-use, when-not-to-use, or named alternative among the many sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
preview_ruleA
Test a candidate rule against real message history before creating it.
Shows what the rule would have matched, so the user can judge whether it is too broad or too narrow while they are still writing it. Creates nothing. Use this during setup, before add_alert_rule, and show the user the samples — a rule that looks sensible in the abstract often turns out to match mostly noise.
The matched samples are also the right input for testing draft criteria: read them and decide which ones a criterion should keep, then tell the user which would have been surfaced and which filtered out.
Args: pattern: the candidate pattern. match_type: 'substring', 'word', or 'regex'. case_sensitive: match case exactly. chat_ids: restrict to these chats; omit to search everywhere. search_hint: a plain word to search Telegram for when match_type is 'regex' — Telegram cannot search by regex, so this narrows what gets scanned locally. Required for an unscoped regex preview. limit: max samples to return. scan: how many messages to pull and test.
| Name | Required | Description | Default |
|---|---|---|---|
| scan | No | ||
| limit | No | ||
| pattern | Yes | ||
| chat_ids | No | ||
| match_type | No | substring | |
| search_hint | No | ||
| case_sensitive | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the burden and does well: it declares 'Creates nothing' (read-only semantics) and discloses a non-obvious platform constraint — Telegram cannot search by regex, hence search_hint narrows what is scanned locally. It stops short of stating cost/latency implications of a large scan.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then caveats, then the args block. Slightly wordy in the middle paragraph about draft criteria, but each sentence carries actionable content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-param, annotation-free, no-output-schema preview tool, the description covers purpose, side-effect profile, workflow routing, and every parameter. Only return-shape details and scan-cost behavior are left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage the description must compensate, and it documents all seven parameters. The search_hint explanation (required for an unscoped regex preview, and why) and the match_type enum values are genuinely useful; the remaining entries are terse but adequate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: test a candidate rule against real message history before creating it. It also explicitly distinguishes itself from the mutation sibling add_alert_rule and from list/search tools via 'Creates nothing.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use: 'Use this during setup, before add_alert_rule,' plus a follow-on workflow ('show the user the samples'). It also names a second use case — testing draft criteria — and says which samples to feed it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_messagesA
Read the most recent messages from a chat/channel.
Args: chat_id: numeric dialog id, as returned by list_dialogs. limit: max number of messages to return (most recent first).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| chat_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden. It discloses the ordering behavior ('most recent first') and the limit semantics, which is genuinely useful, but omits auth requirements, behavior on empty chats, and whether older messages are reachable. An output schema exists, so return format is covered elsewhere.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences plus a lean Args block, with the purpose front-loaded and zero filler. Every line contributes information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read tool with an output schema, the definition covers purpose, both parameters, and ordering. It stops short of describing pagination to older messages or behavior when the chat is empty, which would fully round it out.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does: it defines chat_id as a 'numeric dialog id, as returned by list_dialogs' and limit as the max number returned, most recent first. That covers both parameters meaningfully, though no default value is stated for limit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read the most recent messages from a chat/channel'), making the scope (recent, not filtered) clear. It implicitly distinguishes itself from search_messages by emphasizing recency, but never names the sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a useful prerequisite by noting chat_id is 'as returned by list_dialogs,' which implies the correct workflow. However, it gives no guidance on when to use this versus search_messages or unread_summary, and no exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_messagesA
Search your Telegram messages for a keyword.
Searches history, unlike alert rules which only see new messages.
Args: query: text to search for. chat_id: restrict search to one chat; omit to search all chats. limit: max number of results.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes | ||
| chat_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full disclosure burden. It usefully reveals that the search is retroactive over history, but says nothing about authentication, rate limits, ordering, or how many results are actually returned; with an output schema present, return-shape details are excusable but operational behavior is thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence with the sibling comparison immediately after, followed by a compact Args block. Nothing is wasteful, though the Args formatting and blank-line spacing are slightly heavier than necessary for three short parameters.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description need not describe return values, and it covers purpose, usage context, and all three parameters. The remaining gaps (auth expectations, result ordering/pagination) are minor for a simple keyword-search tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it meaningfully documents all three parameters: query as the search text, chat_id as a single-chat restriction with 'omit for all' semantics, and limit as a max-results cap. It lacks format details (e.g., chat_id must be numeric) and does not restate the default of 30, but covers the intent of every parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Search your Telegram messages for a keyword.' It further differentiates the tool by contrasting it with the alert-rule family ('searches history, unlike alert rules which only see new messages'), so an agent can position it against at least one sibling without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states a clear use condition (retroactive search of history vs. the real-time alert rules) and gives scoping guidance ('omit to search all chats'). It does not explicitly distinguish this from the closer sibling read_messages, so it stops short of full when/when-not coverage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_alert_rule_enabledA
Turn an alert rule on or off without deleting it or its alerts.
Args: rule_id: id from list_alert_rules. enabled: True to resume matching, False to pause.
| Name | Required | Description | Default |
|---|---|---|---|
| enabled | Yes | ||
| rule_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose the key behavioral trait that alerts and the rule are preserved. However, it says nothing about permissions required, whether paused rules re-evaluate existing alerts on resume, or what the call returns, leaving gaps for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded single-sentence purpose followed by a compact Args block; every line earns its place given the empty schema descriptions. The Args formatting is slightly verbose but justified by 0% schema coverage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter mutation with no output schema and no annotations, the description covers both parameters and the non-destructive behavior. It omits permission/auth requirements and return behavior, minor gaps but the essential calling information is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does: rule_id is tied to list_alert_rules (telling the agent how to obtain a valid id) and enabled is annotated with the semantic mapping True=resume matching, False=pause. This goes meaningfully beyond the bare boolean/integer types in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('turn on or off') and resource ('an alert rule') and immediately distinguishes the operation's non-destructive nature from sibling delete_alert_rule. An agent can tell it apart from add_alert_rule/delete_alert_rule without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly frames the use case as toggling state without deleting the rule or its alerts, which contrasts with delete_alert_rule. It stops short of explicitly naming the alternative tool or stating when a pause is preferable to deletion beyond the implied 'keep alerts intact' condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_api_credentialsA
Save the Telegram app credentials needed before any login.
The user gets these from https://my.telegram.org under "API development tools" — they identify the application, not the account, and are not secret in the way a login code is. Stored owner-only in /.env.
Args: api_id: numeric App api_id from my.telegram.org. api_hash: 32-character App api_hash from my.telegram.org.
| Name | Required | Description | Default |
|---|---|---|---|
| api_id | Yes | ||
| api_hash | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and largely meets it: it discloses persistence ('Stored owner-only in <home>/.env') and the security semantics ('not secret in the way a login code is', 'identify the application, not the account'). It omits idempotency (overwrite behavior on repeat calls) and return/confirmation behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose in the first sentence, then separates rationale, source, storage, and an explicit Args list. Slightly verbose but every sentence adds usable context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers purpose, prerequisite, credential origin, storage, and both parameters. No output schema exists, so return values need not be described; only repeat-call behavior and failure modes are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% (only bare titles), so the description must compensate and mostly does: api_id documented as 'numeric', api_hash as '32-character', both sourced from my.telegram.org. Minor tension: schema types both as string while the description calls api_id 'numeric', which could confuse the agent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Save the Telegram app credentials.' The phrase 'needed before any login' distinguishes it from the login_* siblings without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Establishes the prerequisite ordering ('before any login'), which tells the agent when this step belongs in the flow, and the source for obtaining values. It stops short of naming the specific alternative tools (login_request_code/login_submit_code) or stating when NOT to call it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_criterion_enabledA
Turn a criterion on or off without deleting it.
Args: criterion_id: id from list_criteria. enabled: True to resume judging against it, False to pause.
| Name | Required | Description | Default |
|---|---|---|---|
| enabled | Yes | ||
| criterion_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden. It usefully clarifies the non-destructive nature of the toggle (the criterion is preserved), but says nothing about permissions, whether the change takes effect immediately, or what is returned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Very short, front-loaded single sentence for purpose followed by compact arg notes. Nothing is padded, though the 'Args:' formatting is slightly more verbose than needed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description covers the two parameters and the non-destructive behavior adequately for a simple toggle. It omits side effects on judging state, permissions, and failure behavior, leaving some gaps for an unannotated mutation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does: it sources criterion_id ('id from list_criteria') and gives the operational meaning of each boolean value. This is meaningful added context beyond the bare schema types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (turn on/off) and resource (criterion), and the phrase 'without deleting it' distinguishes it from sibling delete_criterion. It does not explicitly name add_criterion or list_criteria as alternatives, but the action is unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parameter note 'True to resume judging against it, False to pause' implies when to use each mode, which is a reasonable usage hint. However, there is no explicit when-to-use/when-not-to-use guidance or named alternative for removing vs. pausing a criterion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
unread_summaryA
Summarize unread chats: for each, the unread count and the last few unread message previews.
Args: limit_per_chat: how many recent messages to preview per chat. max_chats: max number of unread chats to include.
| Name | Required | Description | Default |
|---|---|---|---|
| max_chats | No | ||
| limit_per_chat | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full burden. It does convey useful scoping behavior (only unread chats, per-chat aggregation, bounded preview window), which implies a non-destructive read, but it never states read-only semantics, auth requirements, or rate limits explicitly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose in the first sentence, then a compact Args block for the two parameters. Sized appropriately for a simple two-parameter tool with no wasted filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the description need not spell out return values, and both input parameters are covered. For a simple read-only summary tool this is largely complete; only the absence of any sibling routing or auth context keeps it from being fully sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does: both parameters are explained in plain language ('how many recent messages to preview per chat', 'max number of unread chats to include'). Defaults live in the schema, which is fine, though no guidance is given on sensible values or truncation behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb (summarize) and a specific resource scope (unread chats), and even specifies the shape of the result: unread count plus previews. It does not, however, position itself against close siblings like list_dialogs, read_messages, or search_messages, so the differentiation is by scope only.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The word 'unread' implies the retrieval context (triage of pending conversations), but there is no explicit when-to-use guidance, no statement of when to prefer read_messages or search_messages, and no prerequisites such as requiring an authenticated session.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
20 tool updates
v0.1.0- First observed
ack_alerts - First observed
add_alert_rule - First observed
add_criterion - First observed
auth_status - First observed
delete_alert_rule - First observed
delete_criterion - First observed
list_alert_rules - First observed
list_alerts - First observed
list_criteria - First observed
list_dialogs - First observed
login_request_code - First observed
login_submit_code - First observed
monitor_status - First observed
preview_rule - First observed
read_messages - First observed
search_messages - First observed
set_alert_rule_enabled - First observed
set_api_credentials - First observed
set_criterion_enabled - First observed
unread_summary
TDQS
Scored across 20 tools
Each tool targets a distinct action/resource: auth flow (status, credentials, request/submit code), dialog reading, search, alert-rule CRUD, criterion CRUD, alert listing/acking, and a monitor status check. The two-tier design (rules prefilter, criteria judge, alerts result) is explicitly documented, so even similar-sounding pairs like list_alert_rules/list_alerts and set_alert_rule_enabled/set_criterion_enabled are clearly separated.
The vast majority follow a clean verb_noun pattern (list_alert_rules, add_alert_rule, delete_criterion, search_messages, ack_alerts). A few deviate to noun-phrases with no verb (auth_status, unread_summary, monitor_status), which is a minor inconsistency but still readable and predictable.
At 20 tools this is on the heavier side, but the surface spans several genuinely distinct concerns (auth, reading, rules, criteria, alerts, monitoring), and each tool maps to a real operation without obvious redundancy. It sits just above the ideal band rather than being bloated.
Coverage is strong: full login/auth flow, dialog listing, reading, searching, preview-before-create, and CRUD for both rules and criteria, plus alert listing/acking and daemon status. Minor gaps remain (no delete/prune or un-ack for alerts, no fetch-by-id), but these are workaroundable.
Maintenance
Related MCP Connectors
Search, read and reply to your Telegram Business chats, transcribed voice included.
1Multi-tenant Telegram gateway for AI agents — HTTP+stdio, 8 tools, MTProto User API
Read Google News across 80 editions and any public Telegram channel, without an API key
Search your AI chat history (ChatGPT, Claude, Codex) from any MCP client. Remote, private, read-only
Related MCP Servers
- AlicenseNot gradedqualityDmaintenancePrivacy-first Telegram MCP server enabling maintainers to triage chats, inspect context, search messages, draft replies, and send authorized messages locally without a cloud relay.557 npm1MIT
- AlicenseAqualityDmaintenanceRead-only Telegram access for Claude and other MCP hosts. Provides tools to list chats, read recent messages, and download media from your own Telegram account without needing an api_id/api_hash.5MIT
- AlicenseNot gradedqualityDmaintenanceA read-only Telegram MCP server that retrieves messages from your DMs, groups, and channels, enabling Claude to generate executive briefings from Telegram conversations.MIT
- AlicenseNot gradedqualityCmaintenanceAn MCP server that gives AI assistants read-only access to your personal Telegram account via MTProto, enforcing a whitelist of allowed chats and never marking messages as read.MIT