igagent
Harvests public Instagram posts without login, cookies, or API keys. Given a post URL, reel/tv link, share link, or bare shortcode, it decodes post metadata (caption, created_at, author username/followers/verified, carousel membership) via public embed/OpenGraph surfaces, downloads the associated images and videos from allowlisted CDN hosts with atomic writes and magic-byte verification, and emits a verified media directory plus a JSON manifest. Available as both a CLI and MCP tools (extract_post, lookup_post, read_manifest, get_schema) for metadata-only lookups or full media harvests, with honest empty results for deleted posts.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@igagentharvest this post: https://www.instagram.com/p/ABC123/ — media plus manifest"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
igagent
An agentic Instagram content & media harvester for AI agents — no login, no API keys, no browser. MCP wrapper included.
igagent is the Instagram sibling of xthread-agent (X/Twitter threads) and ytagent (YouTube). Same doctrine: a deterministic, slot-based, stdlib-only state machine that a cloud agent can call with one URL and read back one JSON contract.
python igagent.py "https://www.instagram.com/p/<CODE>/"
# → verified media on disk + post_manifest.jsonWhy this exists
Cloud-based AI agents (GitHub Actions runners, sandboxed VMs, MCP hosts) cannot log into Instagram, and Instagram keeps proving that a login is not even the interesting boundary — the interesting boundary is which door a machine is allowed through with no session at all. Public embed pages are exactly such a door: Instagram serves them to any referrer, and for link unfurlers it renders the actual media straight into the HTML. igagent is built entirely on that observation.
What it gives you:
One input, one artifact. A post URL (or share link, or bare shortcode) in; a verified media directory plus
post_manifest.jsonout.Honest negatives. A deleted post is not an exception — it is
status: "empty"with structured errors and a slot-by-slot provenance trace.Slots, not brands. The three decode surfaces are interchangeable implementations of one contract. When Instagram kills one, you replace the slot — the pipeline never restructures.
Verified delivery. Files exist only after passing the CDN allowlist, a
Content-Lengthcheck, and a magic-byte identity check. A tool that reports a file is vouching for its bytes.Truth layers over heuristics. Carousel membership comes only from the payload (
edge_sidecar_to_children) and is flaggedcarousel.known— never inferred from page layout. The shortcode itself decodes to the numeric media id by pure math, cross-checked against Instagram'sig_cache_key.Politeness as a hard constraint. Bounded retries, 0.6 s decode sleeps, response caps, one post per invocation. The public surfaces this tool depends on are free; restraint is the rent.
Related MCP server: mcp-instagram
Architecture
┌──────────────────────────────────────────────────────┐
│ DISCOVERY │
│ normalize p/reel/reels/tv URLs · instagr.am & │
│ ddinstagram mirrors · share links (1 hop) · bare │
│ shortcodes · shortcode ↔ media-id base64 math │
└───────────────────────────┬──────────────────────────┘
▼
┌──────────────────────────────────────────────────────┐
│ DECODE (slots) │
│ 1. embed_json — legacy rich payload │
│ (caption, metrics, sidecar children, author) │
│ 2. embed_html — unfurler SSR page │
│ (media img/video, author, followers, cache key) │
│ 3. og_meta — OpenGraph tags │
│ (title/caption, og:image, og:video:secure_url) │
│ provenance trace [{"slot","outcome"}] on every run │
└───────────────────────────┬──────────────────────────┘
▼
┌──────────────────────────────────────────────────────┐
│ DELIVER │
│ https + *.cdninstagram.com / *.fbcdn.net only │
│ stream → .part → Content-Length ✓ → magic bytes ✓ │
│ → os.replace (atomic) → files + post_manifest.json │
└──────────────────────────────────────────────────────┘Each tier fails closed: a slot that finds nothing hands control to the next slot and the attempt is recorded; only "everything failed" becomes a run-level error.
Quickstart
# full harvest (decode + verified downloads)
python igagent.py "https://www.instagram.com/p/<CODE>/" --out ./ig_media
# metadata only
python igagent.py "<CODE>" --no-download
# the machine contract: exactly one JSON object on stdout
python igagent.py "<input>" --json --quietFor AI agents
python igagent.py agent-instructions 2>/dev/null || true # not needed —
# the operating manual IS agent.md: read that file and nothing else.The three-command contract:
command | returns | exit |
| summary JSON on stdout, full envelope at | 0 ok/partial · 1 empty · 2 invalid |
| same, no files on disk | same |
|
| 0 |
CLI reference
igagent.py <input> [--out DIR] [--no-download] [--json] [--quiet] [--version]
<input> instagram.com /p|reel|reels|tv/<code>/ · mirror hosts ·
share links (instagram.com/share/…) · bare shortcodes
--out DIR output directory (default: ig_media)
--no-download decode only; manifest still written
--json one JSON summary object on stdout (stderr silenced)
--quiet silence human logs on stderrAccepted input shapes: https://www.instagram.com/p/<CODE>/,
instagram.com/reel/<CODE>?utm_source=…, m./l./ddinstagram/instagr.am
mirrors, https://www.instagram.com/share/<token>/, and bare shortcodes
(5–32 chars of A-Za-z0-9_- that decode to a positive media id).
JSON output (for AI agents)
{
"ok": true,
"status": "ok",
"shortcode": "<CODE>",
"media_id": "1949525278281554174",
"canonical_url": "https://www.instagram.com/p/<CODE>/",
"post_type": "image",
"extraction_source": "embed_html",
"images": 1, "videos": 0,
"downloaded": 1, "failed_downloads": 0,
"out_dir": "ig_media",
"manifest_path": "ig_media/post_manifest.json",
"errors": [],
"duration_sec": 2.2
}The full envelope (post_manifest.json) adds the post object: caption,
created_at, author (username, followers, verified, … — explicit
nulls when a slot does not expose them), media.images[] /
media.videos[] with url / file / downloaded / downloadable / reason,
carousel{known, child_count}, errors[] with stable codes, and
metadata.decode_slots_tried — the provenance trace. The draft-07 schema
is bundled at schema/post-result.schema.json.
For MCP hosts
python mcp_server.py speaks newline-delimited JSON-RPC 2.0 on stdio —
stdlib only, no mcp package. Register it in Claude Desktop / Zed:
tool | what it does |
| full harvest → verified files + envelope |
| metadata-only decode (no downloads) |
| return an existing |
| the draft-07 envelope schema |
Error policy: an honest empty result is not an MCP error; a bad tool argument, timeout, or crash is.
Documentation
file | role |
the operating manual for AI agents — read this file and nothing else | |
perfection-based role prompts for the five pipeline roles | |
maintainer manifesto: why, load-bearing walls, fragile parts, debts | |
living endpoint status table + maintenance protocol | |
the output contract, machine-checkable | |
version history |
Constraints (non-negotiable)
No login, no cookies, no OAuth, no browser. Public content only; everything fails closed.
No GUI, no interactive prompts. 100% non-interactive CLI.
No LLM at runtime. Deterministic state machine — the "agent" is designed for AI agents, not made of one.
stdlib only. Single file, zero pip dependencies, Python 3.9+.
Logs on stderr, data on stdout. Always pipe-safe.
Files stay under the output directory. CDN allowlist, ID validation, atomic writes, magic-byte verification.
Requirements
Python 3.9+ and outbound HTTPS. Nothing else — no pip, no ffmpeg, no browser, no env vars, no config files.
Installation
# as a tool (after PyPI publish)
pip install igagent
igagent "<url>" --json
# from source
git clone https://github.com/Bilal140202/igagent.git
python igagent/igagent.py "<url>"
# or
python -m igagent "<url>"Testing
python -m unittest discover -s tests -p "test_*.py"
# 69 tests, ~1s, zero network — synthetic fixtures only, never real IDsVerified behavior (as shipped)
claim | evidence |
| live run: |
shortcode ↔ media-id math is exact | round-trip property tests (50 random ids) |
dead posts produce honest empties | live run: |
magic-byte gate rejects HTML masquerading as media | offline tests |
lookalike CDN hosts ( | offline tests |
MCP handshake + 4-tool registry | offline stdio probe |
Re-verify against your own vantage point and update
docs/endpoint-matrix.md — that is the protocol.
Limitations (the honest section)
Reel/video URLs are frequently not exposed by the embed surface. The video entry then reports
downloadable: false, reason: "no_video_url_exposed"— this is the truth, not a bug. Stories, live, and private accounts are out of scope entirely.Carousel depth is slot-dependent. The commonly-served unfurler view shows only the primary item;
carousel.known=falsesays so honestly. Full sidecars require the legacyembed_jsonpayload.Datacenter IPs get walls. Instagram's structured APIs (graphql, media info) are walled from cloud ranges and deliberately not slots;
og_metaoften needs a residential IP.Instagram changes surfaces without notice. The slot architecture and the endpoint matrix exist precisely because everything upstream of this tool is borrowed ground.
Legal / ethics
igagent accesses only publicly served documents over unauthenticated HTTP(S), with politeness sleeps and hard caps, and it never circumvents a paywall, a login, or a private post. Respect creators: media and captions remain the property of their authors; downstream use is your responsibility. Do not use this tool at volumes that constitute abuse.
FAQ
Why is there no login option? Because the calling agent is a cloud VM by definition — no session, no cookies. The whole design is "what can a sessionless machine legitimately get?" — and the answer is documented, versioned, and fail-closed.
Why is status often partial instead of ok? partial means the
post was decoded but something failed (e.g. a video URL was never
exposed so nothing could be downloaded). The media entries carry the
per-item reason. errors[] is the full story.
A decode stopped working — is igagent broken? Check
docs/endpoint-matrix.md first. Surfaces flip; the matrix is the
impersonal record of flips, and slots are how they get absorbed.
Does it work for stories or private accounts? No, by design — both require authentication, which violates constraint 1.
License
MIT — see LICENSE.
Acknowledgments
xthread-agent— the architecture, the docs-as-contract system, the MCP wrapper, and the phrase "restraint is the rent".ytagent— "never trust a method's self-report", the verification doctrine this project merged into its delivery tier.The mirror/unfurler ecosystem — the doors that were already open.
Links
Repository: https://github.com/Bilal140202/igagent
Siblings:
xthread-agent·ytagent
This server cannot be deployed
Maintenance
Related MCP Connectors
Instagram data for AI agents: profiles, posts, reels, followers. Influencer + brand research.
Instagram for AI agents: publish, read comments and DMs, insights, and engage from your account.
Instagram search and posts for AI agents, with images and videos AI-parsed to text.
Your agent needs public Instagram data — a creator's posts and reels, what a hashtag is producing, what a video actually says. The official Graph API only sees accounts you already own, and needs app review to see those. **What you can ask for** • "Pull this creator's last 50 posts and reels with engagement counts." • "What is trending under #skincare this week, and which profiles keep appearing?" • "Transcribe this reel and tell me what the hook in the first three seconds is." • "Read the comments on this post and group the objections." • "Which reels use this song right now?" **How to use it** Point any MCP client at https://mcp.aisa.one/instagram/mcp and sign in with OAuth — there is no key to create or paste. 17 read tools: profiles (basic and full), a user's posts, reels and highlights, post and profile digests, post comments, reels search, trending reels, reels by song, hashtag and profile search, and media transcripts. **Why this rather than the source** Public profiles without owning the account, and no app review to sit through. **It is also a door to the rest** The same login reaches 26 sources and 580+ operations. Size a creator's audience here, then ask the same agent what their brand's site traffic looks like or who to contact there — without adding a second server. **What it costs** Finding and inspecting an operation is free. Running one is billed per call at API prices, with no seat and no monthly minimum, and every call takes max_price_usd so an agent cannot overspend by accident. **Where else it reaches** https://mcp.aisa.one/social/mcp for X plus Instagram, Reddit, Pinterest and YouTube; https://mcp.aisa.one/gtm/mcp for those plus Similarweb and Apollo.
Related MCP Servers
- AlicenseAqualityCmaintenanceEnables AI assistants to download Instagram content including posts, videos, reels, stories, highlights, and profile pictures using Instaloader, with optional metadata and caption extraction.5MIT
- AlicenseNot gradedqualityBmaintenanceMCP server for downloading Instagram content (videos, reels, audio, carousels) using yt-dlp, with tools for downloading media and fetching metadata. Supports stdio and HTTP transports.MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to scrape Instagram Reels and public profile data (metadata, engagement, owner info) without the official API, via tools for scraping, status checks, cookie import, stopping, and exporting results.MIT
- AlicenseNot gradedqualityCmaintenanceEnables local research of public Instagram Reels through a dashboard, CLI, and 22 MCP tools, with optional Codex or Claude Code analysis.MIT