ttagent
Harvests public TikTok videos and soundtracks from video URLs, short links, or bare IDs. Provides verified downloads of video, music, and covers, plus metadata and a JSON manifest, with metadata-only lookup and schema tools for AI agents.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@ttagentdownload this TikTok video and music: https://vt.tiktok.com/ZMabc123/"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
ttagent
An agentic TikTok video & soundtrack harvester for AI agents — no login, no API keys, no browser. MCP wrapper included.
ttagent is the TikTok sibling of xthread-agent (X/Twitter threads), ytagent (YouTube), and igagent (Instagram). Same doctrine: a deterministic, slot-based, stdlib-only state machine that a cloud agent can call with one URL and read back one JSON contract.
python ttagent.py "https://www.tiktok.com/@user/video/<ID>"
# → verified video + soundtrack + covers on disk + video_manifest.jsonWhy this exists
TikTok is the friendliest of the big platforms to sessionless machines — a whole ecosystem of public mirror workers and open CDN paths grew around it — but it is also the most volatile: mirror domains rotate, regions differ, and the official surface offers metadata only. ttagent is built for exactly that landscape: compose many legitimate access paths, verify every byte, and tell the truth about what worked.
What it gives you:
One input, one artifact. A video URL (or short link, or bare id) in; the video, its soundtrack, its covers, and
video_manifest.jsonout.The soundtrack is media. A TikTok post is a video plus its music — the track is harvested alongside the video (opt-out with
--no-music), with its identity (title, author, original?, duration, id) in the envelope.Honest negatives. A deleted video is not an exception — it is
status: "empty"with structured errors and a slot-by-slot provenance trace.Slots, not brands. The four decode surfaces are interchangeable implementations of one contract. When a mirror dies, you replace the slot — the pipeline never restructures.
Verified delivery. Files exist only after passing the CDN allowlist, a
Content-Lengthcheck, and a magic-byte identity check (MP4ftyp, MP3ID3/frame-sync, JPEG, PNG). A tool that reports a file is vouching for its bytes.Politeness as a hard constraint. Bounded retries, 0.6 s decode sleeps, response caps, one post per invocation. The public surfaces this tool depends on are free; restraint is the rent.
Related MCP server: TubePull MCP Server
Architecture
┌──────────────────────────────────────────────────────┐
│ DISCOVERY │
│ normalize @user/video|photo URLs · vm./vt./t/ │
│ short links (one hop, dead-link verdict) · bare │
│ 15–20 digit video ids │
└───────────────────────────┬──────────────────────────┘
▼
┌──────────────────────────────────────────────────────┐
│ DECODE (slots) │
│ 1. tikwm — mirror worker API (richest: HD, │
│ watermark variants, music, stats, slideshows) │
│ 2. embed_v2 — TikTok's own embed hydration blob │
│ 3. tiklydown — second mirror (documented slot) │
│ 4. oembed — official, metadata-only last resort│
│ provenance trace [{"slot","outcome"}] on every run │
└───────────────────────────┬──────────────────────────┘
▼
┌──────────────────────────────────────────────────────┐
│ DELIVER │
│ https + tiktok CDN families / tikwm mirror / │
│ byteoversea / akamaized / muscdn only │
│ stream → .part → Content-Length ✓ → magic bytes ✓ │
│ → os.replace (atomic) → files + video_manifest.json│
└──────────────────────────────────────────────────────┘Each tier fails closed: a slot that finds nothing hands control to the next slot and the attempt is recorded; only "everything failed" becomes a run-level error.
Quickstart
# full harvest (decode + verified downloads of video, music, covers)
python ttagent.py "https://www.tiktok.com/@user/video/<ID>" --out ./tiktok_media
# metadata only
python ttagent.py "<ID>" --no-download
# video only, skip the soundtrack
python ttagent.py "<url>" --no-music
# the machine contract: exactly one JSON object on stdout
python ttagent.py "<input>" --json --quietFor AI agents
The three-command contract:
command | returns | exit |
| summary JSON on stdout, full envelope at | 0 ok/partial · 1 empty · 2 invalid |
| same, no files on disk | same |
|
| 0 |
CLI reference
ttagent.py <input> [--out DIR] [--no-download] [--no-music] [--json] [--quiet] [--version]
<input> tiktok.com/@user/video/<id> · /photo/<id> · vm./vt. short
links · tiktok.com/t/<code> · bare 15–20 digit id
--out DIR output directory (default: tiktok_media)
--no-download decode only; manifest still written
--no-music skip the soundtrack download
--json one JSON summary object on stdout (stderr silenced)
--quiet silence human logs on stderrJSON output (for AI agents)
{
"ok": true,
"status": "ok",
"video_id": "6718335390845095173",
"canonical_url": "https://www.tiktok.com/@scout2015/video/6718335390845095173",
"extraction_source": "tikwm",
"author": "scout2015",
"video_file": "6718335390845095173_video.mp4",
"music_file": "6718335390845095173_music.mp3",
"images": 0,
"downloaded": 3,
"failed_downloads": 0,
"out_dir": "tiktok_media",
"manifest_path": "tiktok_media/video_manifest.json",
"errors": [],
"duration_sec": 3.3
}The full envelope (video_manifest.json) adds the video object: title,
region, created_at, author, the soundtrack identity, stats
(plays/likes/comments/shares/collects/downloads), media.video with
url / hd_url / watermarked_url and per-variant download state,
media.images[] for slideshows, errors[] with stable codes, and
metadata.decode_slots_tried — the provenance trace. The draft-07 schema
is bundled at schema/video-result.schema.json.
For MCP hosts
python mcp_server.py speaks newline-delimited JSON-RPC 2.0 on stdio —
stdlib only, no mcp package. Register it in Claude Desktop / Zed:
tool | what it does |
| full harvest → verified video/music/covers + envelope |
| metadata-only decode (no downloads) |
| return an existing |
| the draft-07 envelope schema |
Error policy: an honest empty result is not an MCP error; a bad tool argument, timeout, or crash is.
Documentation
file | role |
the operating manual for AI agents — read this file and nothing else | |
perfection-based role prompts for the five pipeline roles | |
maintainer manifesto: why, load-bearing walls, fragile parts, debts | |
living endpoint status table + maintenance protocol | |
the output contract, machine-checkable | |
version history |
Constraints (non-negotiable)
No login, no cookies, no OAuth, no browser. Public content only; everything fails closed.
No GUI, no interactive prompts. 100% non-interactive CLI.
No LLM at runtime. Deterministic state machine — the "agent" is designed for AI agents, not made of one.
stdlib only. Single file, zero pip dependencies, Python 3.9+.
Logs on stderr, data on stdout. Always pipe-safe.
Files stay under the output directory. CDN allowlist, ID validation, atomic writes, magic-byte verification.
Requirements
Python 3.9+ and outbound HTTPS. Nothing else — no pip, no ffmpeg, no browser, no env vars, no config files.
Installation
# as a tool (after PyPI publish)
pip install ttagent
ttagent "<url>" --json
# from source
git clone https://github.com/Bilal140202/ttagent.git
python ttagent/ttagent.py "<url>"
# or
python -m ttagent "<url>"Testing
python -m unittest discover -s tests -p "test_*.py"
# 66 tests, ~1s, zero network — synthetic fixtures only, never real IDsVerified behavior (as shipped)
claim | evidence |
| live run: |
mirror-relative media paths are absolutized before download | offline tests |
dead videos produce honest empties | live run: |
magic-byte gate rejects HTML masquerading as media | offline tests |
lookalike CDN hosts ( | offline tests |
MCP handshake + 4-tool registry | offline stdio probe |
Re-verify against your own vantage point and update
docs/endpoint-matrix.md — that is the protocol.
Limitations (the honest section)
Region walls are real. A mirror worker in region A may not see a video visible in region B; the envelope reports what the doors showed, and nothing else.
Mirrors are borrowed ground. If a mirror rate-limits or dies, the trace says so and the next slot takes over — but if all mirrors and the embed page are down,
oembedmetadata is the ceiling.Slideshow posts depend on the mirror exposing
images[]; the official surface does not, and the envelope says which slot produced what.Livestreams and private accounts are out of scope — both require authentication, which violates constraint 1.
Legal / ethics
ttagent accesses only publicly served documents over unauthenticated HTTP(S), with politeness sleeps and hard caps, and it never circumvents a paywall, a login, or a private post. Respect creators: videos, music and captions remain the property of their authors; downstream use is your responsibility. Do not use this tool at volumes that constitute abuse.
FAQ
Why is there no login option? Because the calling agent is a cloud VM by definition — no session, no cookies. The whole design is "what can a sessionless machine legitimately get?" — and the answer is documented, versioned, and fail-closed.
Why is the music download an MP3 even when TikTok serves M4A? The music URL decides: the file keeps the extension the CDN serves, and the magic-byte check accepts both ID3 and raw AAC frame syncs. The manifest records the real file name either way.
A decode stopped working — is ttagent broken? Check
docs/endpoint-matrix.md first. Surfaces flip; the matrix is the
impersonal record of flips, and slots are how they get absorbed.
Does it work for livestreams or private videos? No, by design — both require authentication, which violates constraint 1.
License
MIT — see LICENSE.
Acknowledgments
xthread-agent— the architecture, the docs-as-contract system, the MCP wrapper, and the phrase "restraint is the rent".ytagent— "never trust a method's self-report", the verification doctrine this project merged into its delivery tier.igagent— the Instagram sibling; the slot-failover provenance design was born there.The mirror-worker ecosystem — the doors that were already open.
Links
Repository: https://github.com/Bilal140202/ttagent
Siblings:
xthread-agent·ytagent·igagent
Available Tools
4 toolsextract_videoA
Harvest a public TikTok video: given a video/photo URL, short link or bare video id, decodes the public surface and downloads the video (HD when exposed), the soundtrack and covers (all magic-byte verified) to out_dir. Returns the summary plus the full video_manifest envelope (author, title, stats, music identity, media URLs, local files, errors, decode-slot provenance). No login, no API keys.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | TikTok video/photo URL, short link, or bare video id | |
| out_dir | No | output directory; a fresh temp dir when omitted | |
| download_media | No | false = manifest only, no files on disk |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does it well: it discloses that files are written to disk, HD selection when exposed, magic-byte verification of downloads, and the no-auth requirement. It does not cover overwrite behavior for an existing out_dir, rate limits, or failure modes beyond noting an errors field in the returned envelope.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action and resource, followed by the return payload. Parenthetical qualifiers ('HD when exposed', 'all magic-byte verified') earn their space, though the second sentence packs a long field list that borders on dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description compensates by enumerating the manifest envelope contents (author, title, stats, music identity, media URLs, local files, errors, decode-slot provenance), and it covers the auth precondition. Only minor gaps remain, such as out_dir overwrite behavior and any rate-limit expectations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents url, out_dir, and download_media, making 3 the baseline. The description adds mild value by restating accepted URL forms and clarifying that the video, soundtrack and covers all land in out_dir, but it introduces no syntax or format detail the schema lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb (Harvest) and resource (public TikTok video), enumerates what is fetched (video in HD, soundtrack, covers) and what is returned. It implicitly contrasts with lookup_video by emphasizing downloading rather than just metadata, but never names a sibling explicitly, so differentiation is left to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a precondition ('given a video/photo URL, short link or bare video id') and a note that no login or API keys are needed, but offers no when-to-use/when-not guidance and never mentions lookup_video or read_manifest as alternatives. An agent choosing between harvesting, looking up, and reading a manifest gets no routing help.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_schemaA
Return the JSON Schema (draft-07) describing the video_manifest envelope (schema_version 1.0) — use it to validate or explore the output contract.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It discloses the return artifact (a draft-07 JSON Schema) and the contract version, and 'Return' plus zero parameters makes the read-only, side-effect-free nature evident. It doesn't describe error cases, but for a static schema fetch that is a minor omission.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the artifact and its standard, with the usage hint trailing. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only schema-retrieval tool with no output schema, the definition covers what is returned and why. A note on sibling boundaries (schema vs. actual manifest data) would make it complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, so per the baseline a 4 is appropriate. The description correctly implies a parameterless call and spends its wording on what is returned instead.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Return) and resource (JSON Schema draft-07 for the video_manifest envelope), and pins the contract version (schema_version 1.0). It is distinguishable from read_manifest by being the schema rather than the data, but it never explicitly names that sibling boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear use context ('use it to validate or explore the output contract'), which tells the agent when the tool is relevant. It stops short of stating when not to use it or pointing to alternatives like read_manifest for actual content.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lookup_videoA
Fast metadata-only lookup of a public TikTok video: same decode pipeline as extract_video but with --no-download — author, title, stats, music identity, media URLs, without writing media files. Use this when you only need to read the video.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | TikTok video/photo URL, short link, or bare video id | |
| out_dir | No | output directory for the manifest; temp dir when omitted |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden and does disclose the key traits: metadata-only, no media files written, fast, and restricted to public videos (implying private/auth-gated videos are unsupported). It also enumerates the returned fields. It does not mention failure modes, rate limits, or auth requirements, keeping it below a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the distinguishing claim, and every clause (pipeline equivalence, return fields, no-write guarantee) adds information. The internal CLI-flag reference ('--no-download') is slightly jargon-heavy but earns its place as the fastest contrast with extract_video.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by listing what comes back (author, title, stats, music identity, media URLs), which is the right call for a metadata lookup. A note on private-video behavior or error cases would close the remaining gap; for a simple two-parameter read tool this is nearly sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with only two parameters, so the schema already documents url and out_dir fully; baseline is 3. The description adds one small clarification — that out_dir holds the manifest — but no format or syntax detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('metadata-only lookup of a public TikTok video') and immediately differentiates from the sibling by noting it is 'the same decode pipeline as extract_video but with --no-download'. An agent can tell exactly which of the two extraction tools to pick without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear selection rule: 'Use this when you only need to read the video,' framed against extract_video's download behavior. The exclusion (use extract_video if you need media files) is implied by the --no-download contrast rather than stated outright, so it falls just short of an explicit when-not.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_manifestA
Read an existing video_manifest.json produced by a previous extract/lookup and return the full envelope. Refuses any file not named video_manifest.json.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | absolute path to a video_manifest.json |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It usefully discloses an unexpected guardrail ('Refuses any file not named video_manifest.json') and the return shape ('full envelope'), but says nothing about error behavior, permissions, or what the envelope contains beyond the vague term.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences with no filler, and the core purpose is front-loaded before the constraint clause. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple single-parameter read with no output schema and no annotations, the description covers purpose, provenance, and the naming guardrail adequately. It could note failure behavior or envelope contents, but nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a single documented 'path' parameter, so the schema already does the work. The description adds only an implicit naming constraint on the target file, not syntax or format detail beyond what the schema provides, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (Read) and resource (video_manifest.json) and clarifies the file's provenance ('produced by a previous extract/lookup'), which distinguishes it from the sibling extract_video and lookup_video tools. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly situates usage after an extract/lookup and states what it returns, implying when to reach for it. It stops short of naming an explicit alternative or saying when NOT to use it, so it is clear context but not full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v1.0.0- First observed
extract_video - First observed
get_schema - First observed
lookup_video - First observed
read_manifest
TDQS
Scored across 4 tools
Each tool has a clearly distinct purpose: extract_video downloads media, lookup_video fetches metadata only, read_manifest reads a stored manifest, and get_schema returns the contract schema. The extract/lookup pair is well-differentiated by the download flag and use-case wording, leaving no realistic ambiguity.
All four tools follow a clean verb_noun pattern: extract_video, lookup_video, read_manifest, get_schema. No mixed conventions or vague verbs.
Four tools is tight and well-scoped for a TikTok public video harvester: one for full extraction, one for metadata, one for reading stored output, and one for schema introspection. Every tool earns its place.
The core lifecycle for a single public video — fetch metadata, extract media, persist and re-read a manifest, and inspect the output contract — is covered. Missing only peripheral operations like listing or discovery, which are outside the stated per-URL harvesting scope.
Maintenance
Related MCP Connectors
Download YouTube, TikTok, Vimeo, SoundCloud and 6 more platforms from any MCP AI chatbot.
TikTok data for AI agents: videos, creators, sounds, hashtags, trends. Content + creator research.
TikTok data for agents: videos, creators, comments, search, transcripts. 25 tools, pay per call.
OCR, transcription, file extraction, and image generation for AI agents via MCP.
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables AI agents to download, transcribe, and inspect video or audio URLs from YouTube, TikTok, X, and 1000+ other sites using server-side yt-dlp, residential proxies, and speech-to-text.929 npmMIT
- AlicenseNot gradedqualityDmaintenanceEnables AI assistants to download video or audio from YouTube, TikTok, Twitter/X, SoundCloud, Vimeo, Twitch, Streamable, Bandcamp, Mixcloud, and Dailymotion via a hosted MCP server.MIT
- AlicenseAqualityBmaintenanceEnables downloading TikTok videos and photos without watermarks, extracting metadata and analytics, and performing bulk downloads of user profiles via CLI or as an MCP server for AI assistants.516 npm1MIT
- AlicenseNot gradedqualityCmaintenanceEnables downloading and probing images/videos from public Instagram, TikTok, and Facebook accounts through MCP tools, returning structured results for AI agents.MIT