Skip to main content
Glama

ttagent

An agentic TikTok video & soundtrack harvester for AI agents — no login, no API keys, no browser. MCP wrapper included.

Python Tests License LLM

ttagent is the TikTok sibling of xthread-agent (X/Twitter threads), ytagent (YouTube), and igagent (Instagram). Same doctrine: a deterministic, slot-based, stdlib-only state machine that a cloud agent can call with one URL and read back one JSON contract.

python ttagent.py "https://www.tiktok.com/@user/video/<ID>"
# → verified video + soundtrack + covers on disk + video_manifest.json

Why this exists

TikTok is the friendliest of the big platforms to sessionless machines — a whole ecosystem of public mirror workers and open CDN paths grew around it — but it is also the most volatile: mirror domains rotate, regions differ, and the official surface offers metadata only. ttagent is built for exactly that landscape: compose many legitimate access paths, verify every byte, and tell the truth about what worked.

What it gives you:

  1. One input, one artifact. A video URL (or short link, or bare id) in; the video, its soundtrack, its covers, and video_manifest.json out.

  2. The soundtrack is media. A TikTok post is a video plus its music — the track is harvested alongside the video (opt-out with --no-music), with its identity (title, author, original?, duration, id) in the envelope.

  3. Honest negatives. A deleted video is not an exception — it is status: "empty" with structured errors and a slot-by-slot provenance trace.

  4. Slots, not brands. The four decode surfaces are interchangeable implementations of one contract. When a mirror dies, you replace the slot — the pipeline never restructures.

  5. Verified delivery. Files exist only after passing the CDN allowlist, a Content-Length check, and a magic-byte identity check (MP4 ftyp, MP3 ID3/frame-sync, JPEG, PNG). A tool that reports a file is vouching for its bytes.

  6. Politeness as a hard constraint. Bounded retries, 0.6 s decode sleeps, response caps, one post per invocation. The public surfaces this tool depends on are free; restraint is the rent.

Related MCP server: TubePull MCP Server

Architecture

            ┌──────────────────────────────────────────────────────┐
            │                    DISCOVERY                        │
            │  normalize @user/video|photo URLs · vm./vt./t/      │
            │  short links (one hop, dead-link verdict) · bare    │
            │  15–20 digit video ids                              │
            └───────────────────────────┬──────────────────────────┘
                                        ▼
            ┌──────────────────────────────────────────────────────┐
            │                     DECODE  (slots)                 │
            │  1. tikwm      — mirror worker API (richest: HD,    │
            │     watermark variants, music, stats, slideshows)   │
            │  2. embed_v2   — TikTok's own embed hydration blob  │
            │  3. tiklydown  — second mirror (documented slot)    │
            │  4. oembed     — official, metadata-only last resort│
            │  provenance trace [{"slot","outcome"}] on every run │
            └───────────────────────────┬──────────────────────────┘
                                        ▼
            ┌──────────────────────────────────────────────────────┐
            │                     DELIVER                        │
            │  https + tiktok CDN families / tikwm mirror /       │
            │  byteoversea / akamaized / muscdn only              │
            │  stream → .part → Content-Length ✓ → magic bytes ✓  │
            │  → os.replace (atomic) → files + video_manifest.json│
            └──────────────────────────────────────────────────────┘

Each tier fails closed: a slot that finds nothing hands control to the next slot and the attempt is recorded; only "everything failed" becomes a run-level error.

Quickstart

# full harvest (decode + verified downloads of video, music, covers)
python ttagent.py "https://www.tiktok.com/@user/video/<ID>" --out ./tiktok_media

# metadata only
python ttagent.py "<ID>" --no-download

# video only, skip the soundtrack
python ttagent.py "<url>" --no-music

# the machine contract: exactly one JSON object on stdout
python ttagent.py "<input>" --json --quiet

For AI agents

The three-command contract:

command

returns

exit

ttagent.py "<input>" --json --quiet

summary JSON on stdout, full envelope at manifest_path

0 ok/partial · 1 empty · 2 invalid

ttagent.py "<input>" --no-download --json --quiet

same, no files on disk

same

ttagent.py --version

ttagent 1.0.0

0

CLI reference

ttagent.py <input> [--out DIR] [--no-download] [--no-music] [--json] [--quiet] [--version]

<input>         tiktok.com/@user/video/<id> · /photo/<id> · vm./vt. short
                links · tiktok.com/t/<code> · bare 15–20 digit id
--out DIR       output directory (default: tiktok_media)
--no-download   decode only; manifest still written
--no-music      skip the soundtrack download
--json          one JSON summary object on stdout (stderr silenced)
--quiet         silence human logs on stderr

JSON output (for AI agents)

{
  "ok": true,
  "status": "ok",
  "video_id": "6718335390845095173",
  "canonical_url": "https://www.tiktok.com/@scout2015/video/6718335390845095173",
  "extraction_source": "tikwm",
  "author": "scout2015",
  "video_file": "6718335390845095173_video.mp4",
  "music_file": "6718335390845095173_music.mp3",
  "images": 0,
  "downloaded": 3,
  "failed_downloads": 0,
  "out_dir": "tiktok_media",
  "manifest_path": "tiktok_media/video_manifest.json",
  "errors": [],
  "duration_sec": 3.3
}

The full envelope (video_manifest.json) adds the video object: title, region, created_at, author, the soundtrack identity, stats (plays/likes/comments/shares/collects/downloads), media.video with url / hd_url / watermarked_url and per-variant download state, media.images[] for slideshows, errors[] with stable codes, and metadata.decode_slots_tried — the provenance trace. The draft-07 schema is bundled at schema/video-result.schema.json.

For MCP hosts

python mcp_server.py speaks newline-delimited JSON-RPC 2.0 on stdio — stdlib only, no mcp package. Register it in Claude Desktop / Zed:

tool

what it does

extract_video

full harvest → verified video/music/covers + envelope

lookup_video

metadata-only decode (no downloads)

read_manifest

return an existing video_manifest.json (refuses anything else)

get_schema

the draft-07 envelope schema

Error policy: an honest empty result is not an MCP error; a bad tool argument, timeout, or crash is.

Documentation

file

role

agent.md

the operating manual for AI agents — read this file and nothing else

agents.md

perfection-based role prompts for the five pipeline roles

PROJECT_CONTEXT.md

maintainer manifesto: why, load-bearing walls, fragile parts, debts

docs/endpoint-matrix.md

living endpoint status table + maintenance protocol

schema/video-result.schema.json

the output contract, machine-checkable

RELEASE_NOTES.md

version history

Constraints (non-negotiable)

  1. No login, no cookies, no OAuth, no browser. Public content only; everything fails closed.

  2. No GUI, no interactive prompts. 100% non-interactive CLI.

  3. No LLM at runtime. Deterministic state machine — the "agent" is designed for AI agents, not made of one.

  4. stdlib only. Single file, zero pip dependencies, Python 3.9+.

  5. Logs on stderr, data on stdout. Always pipe-safe.

  6. Files stay under the output directory. CDN allowlist, ID validation, atomic writes, magic-byte verification.

Requirements

Python 3.9+ and outbound HTTPS. Nothing else — no pip, no ffmpeg, no browser, no env vars, no config files.

Installation

# as a tool (after PyPI publish)
pip install ttagent
ttagent "<url>" --json

# from source
git clone https://github.com/Bilal140202/ttagent.git
python ttagent/ttagent.py "<url>"
# or
python -m ttagent "<url>"

Testing

python -m unittest discover -s tests -p "test_*.py"
# 66 tests, ~1s, zero network — synthetic fixtures only, never real IDs

Verified behavior (as shipped)

claim

evidence

tikwm slot decodes a live public video with HD + music + covers

live run: status=ok, video (2.0 MB, ftyp-verified), music (ID3-verified), cover (JPEG-verified), downloaded: 3

mirror-relative media paths are absolutized before download

offline tests

dead videos produce honest empties

live run: status=empty, E_VIDEO_UNAVAILABLE, exit 1

magic-byte gate rejects HTML masquerading as media

offline tests

lookalike CDN hosts (tiktokcdn.com.evil.io) refused

offline tests

MCP handshake + 4-tool registry

offline stdio probe

Re-verify against your own vantage point and update docs/endpoint-matrix.md — that is the protocol.

Limitations (the honest section)

  • Region walls are real. A mirror worker in region A may not see a video visible in region B; the envelope reports what the doors showed, and nothing else.

  • Mirrors are borrowed ground. If a mirror rate-limits or dies, the trace says so and the next slot takes over — but if all mirrors and the embed page are down, oembed metadata is the ceiling.

  • Slideshow posts depend on the mirror exposing images[]; the official surface does not, and the envelope says which slot produced what.

  • Livestreams and private accounts are out of scope — both require authentication, which violates constraint 1.

ttagent accesses only publicly served documents over unauthenticated HTTP(S), with politeness sleeps and hard caps, and it never circumvents a paywall, a login, or a private post. Respect creators: videos, music and captions remain the property of their authors; downstream use is your responsibility. Do not use this tool at volumes that constitute abuse.

FAQ

Why is there no login option? Because the calling agent is a cloud VM by definition — no session, no cookies. The whole design is "what can a sessionless machine legitimately get?" — and the answer is documented, versioned, and fail-closed.

Why is the music download an MP3 even when TikTok serves M4A? The music URL decides: the file keeps the extension the CDN serves, and the magic-byte check accepts both ID3 and raw AAC frame syncs. The manifest records the real file name either way.

A decode stopped working — is ttagent broken? Check docs/endpoint-matrix.md first. Surfaces flip; the matrix is the impersonal record of flips, and slots are how they get absorbed.

Does it work for livestreams or private videos? No, by design — both require authentication, which violates constraint 1.

License

MIT — see LICENSE.

Acknowledgments

  • xthread-agent — the architecture, the docs-as-contract system, the MCP wrapper, and the phrase "restraint is the rent".

  • ytagent — "never trust a method's self-report", the verification doctrine this project merged into its delivery tier.

  • igagent — the Instagram sibling; the slot-failover provenance design was born there.

  • The mirror-worker ecosystem — the doors that were already open.

Available Tools

4 tools
extract_videoA

Harvest a public TikTok video: given a video/photo URL, short link or bare video id, decodes the public surface and downloads the video (HD when exposed), the soundtrack and covers (all magic-byte verified) to out_dir. Returns the summary plus the full video_manifest envelope (author, title, stats, music identity, media URLs, local files, errors, decode-slot provenance). No login, no API keys.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesTikTok video/photo URL, short link, or bare video id
out_dirNooutput directory; a fresh temp dir when omitted
download_mediaNofalse = manifest only, no files on disk

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does it well: it discloses that files are written to disk, HD selection when exposed, magic-byte verification of downloads, and the no-auth requirement. It does not cover overwrite behavior for an existing out_dir, rate limits, or failure modes beyond noting an errors field in the returned envelope.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the action and resource, followed by the return payload. Parenthetical qualifiers ('HD when exposed', 'all magic-byte verified') earn their space, though the second sentence packs a long field list that borders on dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description compensates by enumerating the manifest envelope contents (author, title, stats, music identity, media URLs, local files, errors, decode-slot provenance), and it covers the auth precondition. Only minor gaps remain, such as out_dir overwrite behavior and any rate-limit expectations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents url, out_dir, and download_media, making 3 the baseline. The description adds mild value by restating accepted URL forms and clarifying that the video, soundtrack and covers all land in out_dir, but it introduces no syntax or format detail the schema lacks.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb (Harvest) and resource (public TikTok video), enumerates what is fetched (video in HD, soundtrack, covers) and what is returned. It implicitly contrasts with lookup_video by emphasizing downloading rather than just metadata, but never names a sibling explicitly, so differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a precondition ('given a video/photo URL, short link or bare video id') and a note that no login or API keys are needed, but offers no when-to-use/when-not guidance and never mentions lookup_video or read_manifest as alternatives. An agent choosing between harvesting, looking up, and reading a manifest gets no routing help.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_schemaA

Return the JSON Schema (draft-07) describing the video_manifest envelope (schema_version 1.0) — use it to validate or explore the output contract.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It discloses the return artifact (a draft-07 JSON Schema) and the contract version, and 'Return' plus zero parameters makes the read-only, side-effect-free nature evident. It doesn't describe error cases, but for a static schema fetch that is a minor omission.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the artifact and its standard, with the usage hint trailing. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only schema-retrieval tool with no output schema, the definition covers what is returned and why. A note on sibling boundaries (schema vs. actual manifest data) would make it complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes no parameters, so per the baseline a 4 is appropriate. The description correctly implies a parameterless call and spends its wording on what is returned instead.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Return) and resource (JSON Schema draft-07 for the video_manifest envelope), and pins the contract version (schema_version 1.0). It is distinguishable from read_manifest by being the schema rather than the data, but it never explicitly names that sibling boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear use context ('use it to validate or explore the output contract'), which tells the agent when the tool is relevant. It stops short of stating when not to use it or pointing to alternatives like read_manifest for actual content.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

lookup_videoA

Fast metadata-only lookup of a public TikTok video: same decode pipeline as extract_video but with --no-download — author, title, stats, music identity, media URLs, without writing media files. Use this when you only need to read the video.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesTikTok video/photo URL, short link, or bare video id
out_dirNooutput directory for the manifest; temp dir when omitted

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden and does disclose the key traits: metadata-only, no media files written, fast, and restricted to public videos (implying private/auth-gated videos are unsupported). It also enumerates the returned fields. It does not mention failure modes, rate limits, or auth requirements, keeping it below a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the distinguishing claim, and every clause (pipeline equivalence, return fields, no-write guarantee) adds information. The internal CLI-flag reference ('--no-download') is slightly jargon-heavy but earns its place as the fastest contrast with extract_video.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by listing what comes back (author, title, stats, music identity, media URLs), which is the right call for a metadata lookup. A note on private-video behavior or error cases would close the remaining gap; for a simple two-parameter read tool this is nearly sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with only two parameters, so the schema already documents url and out_dir fully; baseline is 3. The description adds one small clarification — that out_dir holds the manifest — but no format or syntax detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('metadata-only lookup of a public TikTok video') and immediately differentiates from the sibling by noting it is 'the same decode pipeline as extract_video but with --no-download'. An agent can tell exactly which of the two extraction tools to pick without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear selection rule: 'Use this when you only need to read the video,' framed against extract_video's download behavior. The exclusion (use extract_video if you need media files) is implied by the --no-download contrast rather than stated outright, so it falls just short of an explicit when-not.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_manifestA

Read an existing video_manifest.json produced by a previous extract/lookup and return the full envelope. Refuses any file not named video_manifest.json.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesabsolute path to a video_manifest.json

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It usefully discloses an unexpected guardrail ('Refuses any file not named video_manifest.json') and the return shape ('full envelope'), but says nothing about error behavior, permissions, or what the envelope contains beyond the vague term.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences with no filler, and the core purpose is front-loaded before the constraint clause. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple single-parameter read with no output schema and no annotations, the description covers purpose, provenance, and the naming guardrail adequately. It could note failure behavior or envelope contents, but nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with a single documented 'path' parameter, so the schema already does the work. The description adds only an implicit naming constraint on the target file, not syntax or format detail beyond what the schema provides, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (Read) and resource (video_manifest.json) and clarifies the file's provenance ('produced by a previous extract/lookup'), which distinguishes it from the sibling extract_video and lookup_video tools. An agent can identify the operation without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly situates usage after an extract/lookup and states what it returns, implying when to reach for it. It stops short of naming an explicit alternative or saying when NOT to use it, so it is clear context but not full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv1.0.0
    • First observedextract_video
    • First observedget_schema
    • First observedlookup_video
    • First observedread_manifest

TDQS

A4/5.0

Scored across 4 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: extract_video downloads media, lookup_video fetches metadata only, read_manifest reads a stored manifest, and get_schema returns the contract schema. The extract/lookup pair is well-differentiated by the download flag and use-case wording, leaving no realistic ambiguity.

Naming Consistency5/5

All four tools follow a clean verb_noun pattern: extract_video, lookup_video, read_manifest, get_schema. No mixed conventions or vague verbs.

Tool Count5/5

Four tools is tight and well-scoped for a TikTok public video harvester: one for full extraction, one for metadata, one for reading stored output, and one for schema introspection. Every tool earns its place.

Completeness4/5

The core lifecycle for a single public video — fetch metadata, extract media, persist and re-read a manifest, and inspect the output contract — is covered. Missing only peripheral operations like listing or discovery, which are outside the stated per-URL harvesting scope.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers