moreel
Most video-to-transcript tools stop at the audio. Moreel treats the video itself as the source of truth: it transcribes what was said, reads what was shown (on-screen text, slides, charts, products), and resolves what a vague "this one" or a pointing gesture actually meant — then makes all of it searchable and jumpable, down to the exact second.
It's built to be used three ways:
As a web app — paste a URL, get video + transcript side by side, with in-transcript search and a "what did I miss?" panel for anything shown but never said out loud.
As an HTTP API — the same pipeline, for your own backend/integration.
As an MCP server — a set of tools (
understand_video,search_video,find_moment,get_video_map, ...) so an AI agent can query a video without ever "watching" it or scrubbing through a transcript by hand.
1. What it does
Transcription — accurate, timestamped speech-to-text (OpenAI Whisper), adaptively re-chunked from word-level timestamps so a segment reflects how the person actually spoke, not an arbitrary decode boundary.
Visual understanding (opt-in) — bounded frame sampling + a vision model surfaces on-screen text, slides, charts, products, and scene context, kept as a separate layer from the transcript, never merged into it.
The Video Map (opt-in, layered on visual understanding) — resolves "this"/"that"/"this one" and pointing/showing/holding gestures to a specific visual entity, with a confidence level and never a fabricated target when the evidence is weak.
Unified search — one query surfaces matches across speech, on-screen text, and resolved references, each labeled by source and one click from seeking the player. Lexical by default; opt-in embeddings add semantic matching (a query for "cost" also finds on-screen text that says "pricing").
"What did I miss?" — surfaces exactly what was visible but never said, classified (on-screen text, visual context, an unspoken visual reference, ...), each anchored to a timestamp — never a generic summary.
Supported sources: public Instagram Reels, TikTok videos, and YouTube videos/Shorts. Private, login-gated, or otherwise access-controlled content is never attempted — it fails with a typed error by design.
2. Installation
Requirements:
Node.js >= 22
yt-dlponPATH(or setYTDLP_PATH)ffmpegonPATH(or setFFMPEG_PATH)An OpenAI API key (transcription; also vision/Video Map/embeddings if you enable those)
git clone https://github.com/medaharrat/moreel.git
cd moreel
npm install
cp .env.example .env
# edit .env and set OPENAI_API_KEY
cd web && npm install && cd ..3. Running it
Quickest path — Docker Compose (API + web UI + Postgres + Redis):
export OPENAI_API_KEY=sk-...
# optional: export VISION_ENABLED=true VIDEO_MAP_ENABLED=true
docker compose -f docker-compose.prod.yml up --buildWeb UI at http://localhost:3001, API at http://localhost:8080.
From source, without Docker:
npm run build && npm run migrate:up # Postgres schema (optional — see below)
npm run dev:http # HTTP API + web-facing endpoints, from source
cd web && npm run dev # web UI (separate terminal)
npm run dev # or: the MCP server over stdio, for an agent clientPostgres/Redis are optional in development — without DATABASE_URL, video
records fall back to an in-process cache; without REDIS_URL, rate
limiting/dedup fall back to in-memory, single-replica-only behavior.
4. Configuration
All configuration is via environment variables — see .env.example for the
full, documented list. The important ones:
Variable | Default | Purpose |
| required | Transcription (and vision/Video Map/embeddings, if enabled) |
|
| On-screen text/scene understanding alongside the transcript |
|
| Resolve "this"/"that"/pointing references (requires |
|
| Semantic matching on top of lexical search |
|
| Reject media larger than this |
|
| Reject media longer than this |
|
| Bounded concurrency; excess fails fast with |
|
| Cache processed videos (in-memory, or Postgres/Redis if configured) |
| (none) | Enables durable, cross-replica video storage + accounts/API keys |
| (none) | Enables distributed rate limiting/dedup |
|
| Structured log verbosity |
Never hardcode secrets — always via environment/.env.
5. MCP tools
Point any MCP client at node dist/mcp/server.js (stdio transport):
{
"mcpServers": {
"moreel": {
"command": "node",
"args": ["/absolute/path/to/moreel/dist/mcp/server.js"],
"env": { "OPENAI_API_KEY": "sk-..." }
}
}
}Tool | Does |
| Speech only — timestamped transcript, opt-in visual observations |
| Speech + visual understanding, always on — entry point for most agent use |
| Query across speech, on-screen text, and resolved references |
| The single best timestamped answer to a specific question |
| The full chronological merge of every modality |
| Entities, interactions, and references the Video Map resolved |
| Everything tied to one specific entity (every mention, every moment) |
| The underlying evidence behind one specific map fact, by id |
search_video/find_moment/get_video_map operate on a video_id
returned by a prior transcribe_video/understand_video call — an agent
watches a video once, then queries it repeatedly without re-processing.
Every response is structured JSON — timestamps, confidence, and observed/inferred/uncertain evidence levels — never prose an agent has to parse.
6. Architecture
INGEST src/providers/ VideoProvider interface; Instagram/TikTok/YouTube
(yt-dlp), each behind a circuit breaker + rate limit
↓
MEDIA src/media/ SSRF-guarded download, audio extraction, bounded
frame sampling (periodic + scene-change)
↓
TRANSCRIPTION src/transcription/ Transcriber interface; OpenAI Whisper; word-timestamp
re-chunking; hallucination/confidence normalization
↓
VISION src/vision/ VisionProvider (on-screen text/scene) +
VideoInteractionAnalyzer (the Video Map), opt-in
↓
APPLICATION src/app/ Timeline merge, lexical+semantic search, "what did I
miss", video-map resolution — orchestration only,
never provider-specific
↓
TRANSPORTS src/mcp/ MCP tools (stdio)
src/http/ Web UI's API + REST-ish JSON routessrc/domain/ holds the shared types (Transcript, VisualObservation,
VideoRecord, the Video Map's Interaction/LinguisticReference) and a
typed MoreelError taxonomy — nothing above it knows which platform or
which model vendor produced the data. Adding a platform means one more
VideoProvider; adding a model vendor means one more Transcriber/
VisionProvider — nothing else changes.
Cross-cutting concerns live alongside, not inside, the pipeline:
src/cache/, src/observability/, src/usage/, src/config/, src/util/.
Security posture
URL validation & SSRF protection — every input URL is parsed and rejected if malformed, non-http(s), or pointing at a private/loopback/ link-local IP (including the cloud-metadata address); redirects are re-validated the same way (
src/media/downloader/ssrf.ts).No shell interpolation —
yt-dlp/ffmpegrun viaexecFilewith argument arrays, never a shell.No path traversal — downloaded files live under a server-generated, randomly named per-request temp directory, always cleaned up in a
finallyblock, regardless of success or failure.Bounded everything — size/duration caps enforced while streaming, hard timeouts on every network/subprocess call, bounded concurrency that fails fast (
RATE_LIMITED) instead of queuing unboundedly.No secret leakage — typed errors carry a client-safe message; stack traces, file paths, and provider internals are logged server-side only.
Bounded-retention storage, not an archive — processed videos (when persisted to Postgres/Redis) carry an explicit TTL and are treated as a cache, matching the same minimal-retention principle as the in-process cache. See
docs/privacy.md.
7. Development
npm run dev # MCP server from source (tsx), no build step
npm run dev:http # HTTP API from source
npm run typecheck # tsc --noEmit
npm run lint # eslint
npm run format # prettier --write8. Testing
npm test # everything
npm run test:unit # provider selection, media limits, normalization,
# timeline/search/video-map resolution, config, ...
npm run test:integration # full pipeline, real Postgres video-store round-trips
npm run test:protocol # real MCP Client <-> Server over an in-memory transport
npm run test:regression # transcription quality regression corpusAll network access and subprocess execution (yt-dlp, ffmpeg, the OpenAI
API) are mocked at their interface boundaries in unit/protocol tests — the
integration suite's Postgres-backed tests are the exception, and skip
automatically when DATABASE_URL isn't reachable.
9. Limitations
Public content only, by design — Moreel never bypasses login, CAPTCHAs, or other access controls.
Automatic transcription/visual analysis is best-effort, not human-verified — surfaced via
low_confidence/evidenceLevelrather than silently guessed.No speaker diarization — multi-speaker audio transcribes in order without speaker labels (Whisper doesn't expose this).
The Video Map doesn't track entity identity via visual similarity across a whole video — it relies on the model reusing a consistent label for a recurring entity within one analysis pass.
Video Map events don't yet carry their own evidence frame thumbnail (timestamp + confidence still make them fully traceable via the player).
10. Contributing
Issues and PRs are welcome. npm run typecheck && npm run lint && npm test
should pass before opening one — CI runs the same checks (plus
integration/protocol/regression suites against real Postgres+Redis service
containers) on every PR.
11. License
MIT — see LICENSE.