txtreel
Server Details
Free iMessage, WhatsApp, Instagram DM and Reddit chat videos (1080x1920 MP4) and screenshots.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 5 tools
Each tool targets a distinct format+action combination: chat vs reddit, and render_video vs screenshot vs validate. An agent can unambiguously pick the right tool for the intent.
All names follow a strict {format}_{action} snake_case pattern (chat_render_video, chat_screenshot, chat_validate, reddit_render_video, reddit_screenshot). The only asymmetry is the missing reddit_validate, which is a coverage issue rather than a naming inconsistency.
Five tools is well-scoped for a renderer serving two content formats. Each tool earns its place with no redundancy.
Both formats support video rendering and screenshots, and chat additionally has a fast, free validate tool. Reddit lacks a dedicated validate tool, though validation is folded into the render/screenshot calls as documented, so it's a minor asymmetry rather than a blocking gap.
Available Tools
5 toolschat_render_videoRender a txtreel videoAInspect
Render a txtreel conversation (script or messages) to an MP4 and wait for it to finish. This queues a render, waits for it, and returns once it is done (or after failing/timing out). A short render (a handful of messages) typically takes 20-60 seconds.
Returns a public link to the MP4; the file is deleted after 24 hours. If txtreel is busy or you hit the rate limit, the result is an error with the message.
Tip: call chat_validate first to catch script errors before paying for a render.
Conversation script format (one event per line):
them: hey, you up? me: yeah why [pause 1.5] wait 1.5 s --- Today 9:41 PM date/time separator me: [photo URL] caption a photo (URL, local file path, or an uploaded photo's number) [them reacts ❤️] reaction on my last message [read] they read my messages === everything above is already on screen at the start === scroll 3 same, but open on the oldest message and skim down in 3 s
comment
"Name: text" also means "them" when Name matches contact.name (e.g. contact.name "Sam" lets you write "Sam: omg" instead of "them: omg"). A literal "\n" inside a message becomes a line break. Lines starting with # are comments.
Guidance: one short message per line, like real texting — split up what a real person would send as separate texts rather than one long paragraph. A typical reel is 8-16 messages and 15-30 seconds. Use "[pause N]" to hold a beat before a reply lands (dramatic timing). Use "===" to start the video with everything above it already on screen, e.g. for a "catch up on this conversation" reel. Photos: "me: [photo https://example.com/image.jpg] optional caption" (https URLs only; local file paths are not available). Use "=== scroll 3" instead to open on the oldest message of a longer history and skim down to the live part in 3 s — too fast to read, so viewers pause and rewind (good for comments: hide a detail in the history). Leave keyboard at its default (true) so the newest messages stay above where the Reels caption/UI usually sits.
| Name | Required | Description | Default |
|---|---|---|---|
| speed | No | Pacing multiplier; 2 = twice as fast, 0.5 = half speed. Default: 1. | |
| theme | No | Color theme, "light" or "dark". Default: light. | |
| script | No | Conversation as a plain-text script. Use this OR messages, not both. Conversation script format (one event per line): them: hey, you up? me: yeah why [pause 1.5] wait 1.5 s --- Today 9:41 PM date/time separator me: [photo URL] caption a photo (URL, local file path, or an uploaded photo's number) [them reacts ❤️] reaction on my last message [read] they read my messages === everything above is already on screen at the start === scroll 3 same, but open on the oldest message and skim down in 3 s # comment "Name: text" also means "them" when Name matches contact.name (e.g. contact.name "Sam" lets you write "Sam: omg" instead of "them: omg"). A literal "\n" inside a message becomes a line break. Lines starting with # are comments. Guidance: one short message per line, like real texting — split up what a real person would send as separate texts rather than one long paragraph. A typical reel is 8-16 messages and 15-30 seconds. Use "[pause N]" to hold a beat before a reply lands (dramatic timing). Use "===" to start the video with everything above it already on screen, e.g. for a "catch up on this conversation" reel. Photos: "me: [photo https://example.com/image.jpg] optional caption" (https URLs only; local file paths are not available). Use "=== scroll 3" instead to open on the oldest message of a longer history and skim down to the live part in 3 s — too fast to read, so viewers pause and rewind (good for comments: hide a detail in the history). Leave keyboard at its default (true) so the newest messages stay above where the Reels caption/UI usually sits. | |
| sounds | No | Real UI sounds: iOS key clicks while typing, the app's send and receive sounds. Default: true. | |
| contact | No | Contact header info. | |
| endHold | No | Seconds to hold on the final frame before the video ends. Default: 2. | |
| autoRead | No | Mark "me" messages as read as soon as the other person starts typing. Default: true. | |
| keyboard | No | Show the iOS keyboard, which keeps the latest messages above the Reels caption area. Default: true. | |
| messages | No | Conversation as an array of event objects instead of a script string (use this OR script, not both). Passed through to the txtreel API as-is; each item is one of: {from:"me"|"them", text, delay?, typing?, hold?, time?, instant?} (type "message" is the default and can be omitted), {type:"pause", seconds}, {type:"timestamp", text, instant?}, {type:"read", time?}, {type:"react", from:"me"|"them", emoji}. | |
| platform | No | Chat app to render: "imessage", "whatsapp", or "instagram". Default: imessage. | |
| startTime | No | Clock used for messages, read receipts, and WhatsApp bubble times, e.g. "9:41 PM". Default: 9:41 PM. | |
| statusBar | No | Phone status bar shown at the top of the frame. | |
| composerTyping | No | Type "me" messages into the input bar before sending them. Default: true. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well past the annotations: states the call blocks until the render completes (20-60s typical), returns a public MP4 link, the file is deleted after 24 hours, and busy/rate-limit conditions surface as an error message. Those latency, lifecycle, and failure traits are exactly the behavior an agent needs and are not in readOnlyHint/openWorldHint/idempotentHint. No conflict with the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loading is good: purpose, blocking behavior, return value, and the validate tip come first. But the large conversation-script format section duplicates the script parameter's schema description almost word-for-word, inflating the text without adding new information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 13-parameter, nested-object tool with no output schema, the description covers the essentials an agent needs: synchronous wait, the returned public link and its 24-hour expiry, and error behavior. Remaining parameter detail is fully carried by the schema, so the omission there is acceptable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 13 parameters (including nested contact/statusBar) are already documented, making 3 the baseline. The description's script-format block is essentially a verbatim copy of the script property's schema description, so it adds no meaning beyond what the schema already carries.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and resource: render a txtreel conversation (script or messages) to an MP4, and clarifies it is synchronous (queues, waits, returns when done). The video-vs-message framing is clear. However, it never distinguishes this from the sibling chat_screenshot, so an agent choosing between image and video output gets no explicit routing help from the text.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives strong in-context guidance: call chat_validate first to catch script errors before paying for a render, plus practical advice on message length, reel duration, [pause] timing, and === usage. It stops short of explicit when-not-to-use or contrast with chat_screenshot / reddit_render_video, so it is clear context without exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chat_screenshotScreenshot a txtreel conversationAInspect
Render a single frame of a txtreel conversation (script or messages) to a PNG and return it as an image. Defaults to the last frame — pass "frame" to capture an earlier moment (30 fps).
Returns the PNG as image content, plus a text line with the PNG's URL (expires after 24 hours).
Conversation script format (one event per line):
them: hey, you up? me: yeah why [pause 1.5] wait 1.5 s --- Today 9:41 PM date/time separator me: [photo URL] caption a photo (URL, local file path, or an uploaded photo's number) [them reacts ❤️] reaction on my last message [read] they read my messages === everything above is already on screen at the start === scroll 3 same, but open on the oldest message and skim down in 3 s
comment
"Name: text" also means "them" when Name matches contact.name (e.g. contact.name "Sam" lets you write "Sam: omg" instead of "them: omg"). A literal "\n" inside a message becomes a line break. Lines starting with # are comments.
Guidance: one short message per line, like real texting — split up what a real person would send as separate texts rather than one long paragraph. A typical reel is 8-16 messages and 15-30 seconds. Use "[pause N]" to hold a beat before a reply lands (dramatic timing). Use "===" to start the video with everything above it already on screen, e.g. for a "catch up on this conversation" reel. Photos: "me: [photo https://example.com/image.jpg] optional caption" (https URLs only; local file paths are not available). Use "=== scroll 3" instead to open on the oldest message of a longer history and skim down to the live part in 3 s — too fast to read, so viewers pause and rewind (good for comments: hide a detail in the history). Leave keyboard at its default (true) so the newest messages stay above where the Reels caption/UI usually sits.
| Name | Required | Description | Default |
|---|---|---|---|
| frame | No | Frame index to capture, at 30 fps (e.g. 30 = one second in). Omit to capture the final frame of the conversation. | |
| speed | No | Pacing multiplier; 2 = twice as fast, 0.5 = half speed. Default: 1. | |
| theme | No | Color theme, "light" or "dark". Default: light. | |
| script | No | Conversation as a plain-text script. Use this OR messages, not both. Conversation script format (one event per line): them: hey, you up? me: yeah why [pause 1.5] wait 1.5 s --- Today 9:41 PM date/time separator me: [photo URL] caption a photo (URL, local file path, or an uploaded photo's number) [them reacts ❤️] reaction on my last message [read] they read my messages === everything above is already on screen at the start === scroll 3 same, but open on the oldest message and skim down in 3 s # comment "Name: text" also means "them" when Name matches contact.name (e.g. contact.name "Sam" lets you write "Sam: omg" instead of "them: omg"). A literal "\n" inside a message becomes a line break. Lines starting with # are comments. Guidance: one short message per line, like real texting — split up what a real person would send as separate texts rather than one long paragraph. A typical reel is 8-16 messages and 15-30 seconds. Use "[pause N]" to hold a beat before a reply lands (dramatic timing). Use "===" to start the video with everything above it already on screen, e.g. for a "catch up on this conversation" reel. Photos: "me: [photo https://example.com/image.jpg] optional caption" (https URLs only; local file paths are not available). Use "=== scroll 3" instead to open on the oldest message of a longer history and skim down to the live part in 3 s — too fast to read, so viewers pause and rewind (good for comments: hide a detail in the history). Leave keyboard at its default (true) so the newest messages stay above where the Reels caption/UI usually sits. | |
| sounds | No | Real UI sounds: iOS key clicks while typing, the app's send and receive sounds. Default: true. | |
| contact | No | Contact header info. | |
| endHold | No | Seconds to hold on the final frame before the video ends. Default: 2. | |
| autoRead | No | Mark "me" messages as read as soon as the other person starts typing. Default: true. | |
| keyboard | No | Show the iOS keyboard, which keeps the latest messages above the Reels caption area. Default: true. | |
| messages | No | Conversation as an array of event objects instead of a script string (use this OR script, not both). Passed through to the txtreel API as-is; each item is one of: {from:"me"|"them", text, delay?, typing?, hold?, time?, instant?} (type "message" is the default and can be omitted), {type:"pause", seconds}, {type:"timestamp", text, instant?}, {type:"read", time?}, {type:"react", from:"me"|"them", emoji}. | |
| platform | No | Chat app to render: "imessage", "whatsapp", or "instagram". Default: imessage. | |
| startTime | No | Clock used for messages, read receipts, and WhatsApp bubble times, e.g. "9:41 PM". Default: 9:41 PM. | |
| statusBar | No | Phone status bar shown at the top of the frame. | |
| composerTyping | No | Type "me" messages into the input bar before sending them. Default: true. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare generic hints (readOnlyHint=false, idempotent=false, openWorld=false), so the description carries real weight and does: it discloses the return shape (PNG image content plus a text line with a URL that expires after 24 hours) and the default-last-frame behavior. It does not mention rate limits or persistence beyond the URL TTL, but this is solid added context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and return behavior are front-loaded, but roughly two-thirds of the description is a verbatim duplicate of the "script" parameter's schema description, including the same code block and authoring guidance. That is a large block of unearned text that an agent already receives in the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 14-parameter, nested-object tool with no output schema, the description does supply the return format and the key authoring conventions, and the schema covers all parameters. The remaining gap is tool-selection context against the video/validate/reddit siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, including a full restatement of the script format, so the baseline is 3. The description's parameter-level content (defaults to last frame, frame at 30 fps) duplicates the schema rather than adding syntax or edge-case meaning beyond it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence gives a precise verb+resource+output: render a single frame of a txtreel conversation to PNG and return it as image content. The "single frame" scope implicitly separates it from the video sibling, but the description never names chat_render_video, so a sibling distinction is left to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is substantial "Guidance:" text, but it is about how to author the script (message pacing, [pause N], ===, === scroll), not about when to pick this tool over chat_render_video, chat_validate, or reddit_screenshot. No exclusions or prerequisites for tool selection are given, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chat_validateValidate a txtreel conversationARead-onlyIdempotentInspect
Check a txtreel conversation (script or messages) without rendering anything. Fast and free — call this before chat_render_video, and again after fixing any reported errors.
On success returns { ok: true, durationSeconds, messageCount, conversation } where "conversation" is the fully-resolved object (all defaults applied). On failure returns { ok: false, errors: string[] } — one message per problem; when the input was a script, errors are prefixed with the offending line number. Videos longer than the server limit are rejected here too.
Conversation script format (one event per line):
them: hey, you up? me: yeah why [pause 1.5] wait 1.5 s --- Today 9:41 PM date/time separator me: [photo URL] caption a photo (URL, local file path, or an uploaded photo's number) [them reacts ❤️] reaction on my last message [read] they read my messages === everything above is already on screen at the start === scroll 3 same, but open on the oldest message and skim down in 3 s
comment
"Name: text" also means "them" when Name matches contact.name (e.g. contact.name "Sam" lets you write "Sam: omg" instead of "them: omg"). A literal "\n" inside a message becomes a line break. Lines starting with # are comments.
Guidance: one short message per line, like real texting — split up what a real person would send as separate texts rather than one long paragraph. A typical reel is 8-16 messages and 15-30 seconds. Use "[pause N]" to hold a beat before a reply lands (dramatic timing). Use "===" to start the video with everything above it already on screen, e.g. for a "catch up on this conversation" reel. Photos: "me: [photo https://example.com/image.jpg] optional caption" (https URLs only; local file paths are not available). Use "=== scroll 3" instead to open on the oldest message of a longer history and skim down to the live part in 3 s — too fast to read, so viewers pause and rewind (good for comments: hide a detail in the history). Leave keyboard at its default (true) so the newest messages stay above where the Reels caption/UI usually sits.
| Name | Required | Description | Default |
|---|---|---|---|
| speed | No | Pacing multiplier; 2 = twice as fast, 0.5 = half speed. Default: 1. | |
| theme | No | Color theme, "light" or "dark". Default: light. | |
| script | No | Conversation as a plain-text script. Use this OR messages, not both. Conversation script format (one event per line): them: hey, you up? me: yeah why [pause 1.5] wait 1.5 s --- Today 9:41 PM date/time separator me: [photo URL] caption a photo (URL, local file path, or an uploaded photo's number) [them reacts ❤️] reaction on my last message [read] they read my messages === everything above is already on screen at the start === scroll 3 same, but open on the oldest message and skim down in 3 s # comment "Name: text" also means "them" when Name matches contact.name (e.g. contact.name "Sam" lets you write "Sam: omg" instead of "them: omg"). A literal "\n" inside a message becomes a line break. Lines starting with # are comments. Guidance: one short message per line, like real texting — split up what a real person would send as separate texts rather than one long paragraph. A typical reel is 8-16 messages and 15-30 seconds. Use "[pause N]" to hold a beat before a reply lands (dramatic timing). Use "===" to start the video with everything above it already on screen, e.g. for a "catch up on this conversation" reel. Photos: "me: [photo https://example.com/image.jpg] optional caption" (https URLs only; local file paths are not available). Use "=== scroll 3" instead to open on the oldest message of a longer history and skim down to the live part in 3 s — too fast to read, so viewers pause and rewind (good for comments: hide a detail in the history). Leave keyboard at its default (true) so the newest messages stay above where the Reels caption/UI usually sits. | |
| sounds | No | Real UI sounds: iOS key clicks while typing, the app's send and receive sounds. Default: true. | |
| contact | No | Contact header info. | |
| endHold | No | Seconds to hold on the final frame before the video ends. Default: 2. | |
| autoRead | No | Mark "me" messages as read as soon as the other person starts typing. Default: true. | |
| keyboard | No | Show the iOS keyboard, which keeps the latest messages above the Reels caption area. Default: true. | |
| messages | No | Conversation as an array of event objects instead of a script string (use this OR script, not both). Passed through to the txtreel API as-is; each item is one of: {from:"me"|"them", text, delay?, typing?, hold?, time?, instant?} (type "message" is the default and can be omitted), {type:"pause", seconds}, {type:"timestamp", text, instant?}, {type:"read", time?}, {type:"react", from:"me"|"them", emoji}. | |
| platform | No | Chat app to render: "imessage", "whatsapp", or "instagram". Default: imessage. | |
| startTime | No | Clock used for messages, read receipts, and WhatsApp bubble times, e.g. "9:41 PM". Default: 9:41 PM. | |
| statusBar | No | Phone status bar shown at the top of the frame. | |
| composerTyping | No | Type "me" messages into the input bar before sending them. Default: true. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safe-read profile (readOnlyHint, idempotentHint, non-destructive, non-open-world), so the bar is lower, and the description still adds real value: the exact success/failure return shapes, that errors are line-number-prefixed for scripts, that defaults are fully resolved, and that oversized videos are rejected here. The only gap is that the referenced 'server limit' is left unquantified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is well front-loaded — purpose, placement, and return values come before any format detail. However, the entire script-format block and usage guidance is duplicated almost word-for-word from the `script` parameter's schema description, making the definition substantially longer than it needs to be for a tool whose schema is already fully documented.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly carries the return-value burden and does so thoroughly, plus it explains the error-reporting model. It is slightly incomplete on scope — it never clarifies whether validation covers only the conversation payload or the rendering settings too (platform, contact, theme).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 13 parameters in detail; per the rubric this establishes a baseline of 3. The description's format section largely restates the `script` parameter's schema description verbatim rather than adding new semantics, and it says nothing about how the other parameters (platform, theme, contact, etc.) participate in validation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a specific verb and resource — 'Check a txtreel conversation (script or messages) without rendering anything' — and immediately distinguishes it from the sibling chat_render_video by saying it renders nothing. An agent can tell exactly what this tool does versus the render tools without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit workflow placement: 'call this before chat_render_video, and again after fixing any reported errors', which names the alternative and the condition that selects it. There is no explicit 'when not to use' or statement about what a valid result guarantees downstream, so it stops just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reddit_render_videoRender a Reddit thread videoAInspect
Render a Reddit thread (script, or post+comments) to an MP4 and wait for it to finish. This queues a render, waits for it, and returns once it is done (or after failing/timing out). A short render typically takes 20-60 seconds.
Returns a public link to the MP4; the file is deleted after 24 hours. If txtreel is busy or you hit the rate limit, the result is an error with the message.
If the thread is invalid, this returns an error result with the validation messages (script errors are prefixed with the offending line number) instead of rendering — fix and call again; there is no separate reddit_validate tool.
Reddit thread script format (post + nested comments):
title: AITA for not going to my sister's wedding? body: First paragraph of the post (optional, one line per paragraph) body: Second paragraph [photo URL] a photo in the post (URL, file path or uploaded photo's number) sleepy_joe: A top-level comment
OP: A reply by the post's author (blue OP label)
sleepy_joe [2.1K 3h]: A reply to that reply, with its votes and age [91 more replies] collapsed replies under the comment above [pause 2] wait 2 s
comment
This is a screen recording of someone reading a real-looking thread on mobile reddit.com in iOS Safari: the page is laid out once and the camera scrolls from block to block, holding each one long enough to read. The title is the hook and is on screen from frame 0 — make it count. A typical reel is a post with 1-3 short body paragraphs plus 4-10 comments, about 30-60 seconds. Use nested replies (">", ">>", …) and a reply from the post's author ("> OP: …") for drama — someone getting called out and the poster jumping in reads as real. A "[N more replies]" line under a comment adds realism: it's what a real collapsed thread looks like. Leave scores and ages out unless they matter to the story — they're made up realistically (decreasing down the thread, never older than the post) when omitted. Use "speed" 1.3-1.5 for a faster-paced reel.
| Name | Required | Description | Default |
|---|---|---|---|
| post | No | The post. Use this OR script, not both. | |
| speed | No | Pacing multiplier; 2 = twice as fast, 0.5 = half speed. Default: 1. Use 1.3-1.5 for a faster reel. | |
| theme | No | Color theme, "light" or "dark". Default: light. | |
| script | No | Thread as a plain-text script. Use this OR post/comments, not both. Reddit thread script format (post + nested comments): title: AITA for not going to my sister's wedding? body: First paragraph of the post (optional, one line per paragraph) body: Second paragraph [photo URL] a photo in the post (URL, file path or uploaded photo's number) sleepy_joe: A top-level comment > OP: A reply by the post's author (blue OP label) >> sleepy_joe [2.1K 3h]: A reply to that reply, with its votes and age [91 more replies] collapsed replies under the comment above [pause 2] wait 2 s # comment This is a screen recording of someone reading a real-looking thread on mobile reddit.com in iOS Safari: the page is laid out once and the camera scrolls from block to block, holding each one long enough to read. The title is the hook and is on screen from frame 0 — make it count. A typical reel is a post with 1-3 short body paragraphs plus 4-10 comments, about 30-60 seconds. Use nested replies (">", ">>", …) and a reply from the post's author ("> OP: …") for drama — someone getting called out and the poster jumping in reads as real. A "[N more replies]" line under a comment adds realism: it's what a real collapsed thread looks like. Leave scores and ages out unless they matter to the story — they're made up realistically (decreasing down the thread, never older than the post) when omitted. Use "speed" 1.3-1.5 for a faster-paced reel. | |
| endHold | No | Seconds to hold on the final frame before the video ends. Default: 2. | |
| comments | No | Comments as an array of event objects instead of a script string (use this OR script, not both). Each item is one of: {author, text, depth (0 top-level, 1 reply to the nearest depth-0 comment above, 2 reply to that, …), score?, age?, op?, avatar?, moreReplies?, hold?} or {type:"pause", seconds}. "op" or an author matching the post author gets the blue OP label automatically. | |
| statusBar | No | Phone status bar shown above the Safari toolbar. | |
| subreddit | No | The subreddit the thread is posted in. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, it discloses the synchronous queue-and-wait behavior, typical latency (20-60s), the artifact's lifetime (public link, deleted after 24h), the failure modes (busy/rate-limit errors, validation errors with line-number prefixes), and that no separate validate tool exists. Annotations only say it is a non-read-only, non-idempotent, non-destructive local operation; the description supplies all the operational detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well front-loaded: the core behavior, return value, and error contract come first, with creative guidance last. It is long, and the script-format block is largely a verbatim duplicate of the same text embedded in the script parameter's schema description, which is redundancy rather than new information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so ('returns a public link to the MP4; the file is deleted after 24 hours') along with duration expectations and both failure shapes. For an 8-parameter, nested-object tool with a creative DSL, an agent has everything needed to call it correctly on the first attempt.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds prescriptive guidance the schema does not: recommending speed 1.3-1.5, advising to omit scores/ages unless they matter (and noting they are invented plausibly), and explaining that nested replies and '> OP:' authorship drive the drama. It does restate the script grammar that already lives in the schema, which caps the score below 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence gives a specific verb and resource: 'Render a Reddit thread (script, or post+comments) to an MP4 and wait for it to finish.' It also distinguishes itself from the sibling reddit_screenshot (video vs. still) and from chat_render_video by the thread/script domain. An agent can identify the tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the two mutually exclusive input paths (script vs. post+comments) and identifies both alternatives explicitly ('use this OR script, not both'), and it tells the agent there is 'no separate reddit_validate tool' so validation must be done by calling this and reading errors. What it lacks is a direct routing statement against reddit_screenshot for the still-image case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reddit_screenshotScreenshot a Reddit threadAInspect
Render a single frame of a Reddit thread (script, or post+comments) to a PNG and return it as an image. Defaults to the last frame — pass "frame" to capture an earlier moment (30 fps).
Returns the PNG as image content, plus a text line with the PNG's URL (expires after 24 hours).
If the thread is invalid, this returns an error result with the validation messages instead of an image — fix and call again; there is no separate reddit_validate tool.
Reddit thread script format (post + nested comments):
title: AITA for not going to my sister's wedding? body: First paragraph of the post (optional, one line per paragraph) body: Second paragraph [photo URL] a photo in the post (URL, file path or uploaded photo's number) sleepy_joe: A top-level comment
OP: A reply by the post's author (blue OP label)
sleepy_joe [2.1K 3h]: A reply to that reply, with its votes and age [91 more replies] collapsed replies under the comment above [pause 2] wait 2 s
comment
This is a screen recording of someone reading a real-looking thread on mobile reddit.com in iOS Safari: the page is laid out once and the camera scrolls from block to block, holding each one long enough to read. The title is the hook and is on screen from frame 0 — make it count. A typical reel is a post with 1-3 short body paragraphs plus 4-10 comments, about 30-60 seconds. Use nested replies (">", ">>", …) and a reply from the post's author ("> OP: …") for drama — someone getting called out and the poster jumping in reads as real. A "[N more replies]" line under a comment adds realism: it's what a real collapsed thread looks like. Leave scores and ages out unless they matter to the story — they're made up realistically (decreasing down the thread, never older than the post) when omitted. Use "speed" 1.3-1.5 for a faster-paced reel.
| Name | Required | Description | Default |
|---|---|---|---|
| post | No | The post. Use this OR script, not both. | |
| frame | No | Frame index to capture, at 30 fps (e.g. 30 = one second in). Omit to capture the final frame of the thread. | |
| speed | No | Pacing multiplier; 2 = twice as fast, 0.5 = half speed. Default: 1. Use 1.3-1.5 for a faster reel. | |
| theme | No | Color theme, "light" or "dark". Default: light. | |
| script | No | Thread as a plain-text script. Use this OR post/comments, not both. Reddit thread script format (post + nested comments): title: AITA for not going to my sister's wedding? body: First paragraph of the post (optional, one line per paragraph) body: Second paragraph [photo URL] a photo in the post (URL, file path or uploaded photo's number) sleepy_joe: A top-level comment > OP: A reply by the post's author (blue OP label) >> sleepy_joe [2.1K 3h]: A reply to that reply, with its votes and age [91 more replies] collapsed replies under the comment above [pause 2] wait 2 s # comment This is a screen recording of someone reading a real-looking thread on mobile reddit.com in iOS Safari: the page is laid out once and the camera scrolls from block to block, holding each one long enough to read. The title is the hook and is on screen from frame 0 — make it count. A typical reel is a post with 1-3 short body paragraphs plus 4-10 comments, about 30-60 seconds. Use nested replies (">", ">>", …) and a reply from the post's author ("> OP: …") for drama — someone getting called out and the poster jumping in reads as real. A "[N more replies]" line under a comment adds realism: it's what a real collapsed thread looks like. Leave scores and ages out unless they matter to the story — they're made up realistically (decreasing down the thread, never older than the post) when omitted. Use "speed" 1.3-1.5 for a faster-paced reel. | |
| endHold | No | Seconds to hold on the final frame before the video ends. Default: 2. | |
| comments | No | Comments as an array of event objects instead of a script string (use this OR script, not both). Each item is one of: {author, text, depth (0 top-level, 1 reply to the nearest depth-0 comment above, 2 reply to that, …), score?, age?, op?, avatar?, moreReplies?, hold?} or {type:"pause", seconds}. "op" or an author matching the post author gets the blue OP label automatically. | |
| statusBar | No | Phone status bar shown above the Safari toolbar. | |
| subreddit | No | The subreddit the thread is posted in. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations only stating readOnly/idempotent/destructive/openWorld hints, the description carries real behavioral weight: it discloses the return shape (PNG as image content plus a text line with a URL that expires in 24 hours) and the failure path (error result containing validation messages instead of an image). That goes well beyond what the structured hints convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose, return behavior, and error handling are front-loaded, and the thread-authoring guidance is genuinely useful for generating a realistic result. However, the script-format block is reproduced almost verbatim from the schema's 'script' property, which is redundant padding within the description itself.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description compensates by explaining exactly what is returned and how long the URL lives. For a 9-parameter tool with nested post/subreddit/statusBar objects and two alternative input modes, the description plus schema together leave nothing essential unexplained for a correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description ties parameters together into a workflow: frame at 30 fps, speed as a pacing multiplier with a recommended 1.3-1.5 range, and the post-vs-script mutual exclusivity. Much of the script-format text is duplicated verbatim from the schema's 'script' property, which limits the added meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence names a specific verb and resource ('Render a single frame of a Reddit thread ... to a PNG') and immediately scopes it to single-frame stills as opposed to the sibling reddit_render_video. It also clarifies the two mutually exclusive input modes (script vs post+comments), so an agent can identify the tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It tells the agent when to pass 'frame' (capture an earlier moment) versus omitting it (default last frame), how to pace a reel with 'speed', and explicitly notes that validation is folded in because there is no separate reddit_validate tool. What is missing is an explicit contrast with reddit_render_video for when a still is preferable to a video.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
- First observed
chat_render_video - First observed
chat_screenshot - First observed
chat_validate - First observed
reddit_render_video - First observed
reddit_screenshot
Related MCP Connectors
Public Instagram reels, posts, profiles and comments, plus reel-to-text transcripts. No login.
Instagram, WhatsApp and Messenger DMs through official Meta Business APIs.
Instagram audits, carousel decks, AI images and video with one consistent face — from chat.
Public Instagram and YouTube data for AI agents: reels, videos, channels, comments. Pay per result.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceLocal-first WhatsApp and iMessage memory for AI coding agents. It provides MCP tools to search messages, get recent context, summarize relationships, draft replies, and view unreplied threads.MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI assistants to read, search, and send iMessages with features like contact name resolution, session grouping, and attachment listing. It provides intent-aligned tools to efficiently navigate conversation history and manage messages through natural language queries.6MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents on macOS to securely read and search the local Messages database, catch up on missed messages via a persistent inbox, and send texts or files to allowlisted chats, with optional voice note transcription and text-to-speech.MIT
- MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.