SanTech Content Engine MCP Server
README.md
# SanTech Content Engine
Local-first YouTube research-to-script engine for SanTech010.
The backend performs deterministic work: YouTube fetching, storage, outlier math, pack building, validation, save/select/export, and MCP tool exposure. Connected AI clients perform creative and judgment work using backend packs as source of truth.
Current status: offline backend, MCP backbone, quota guard, resource import, recent-only outliers, early velocity signals, combined signal reports, and signal-target transcript acquisition implemented.
Important MCP rule: source files sync locally on MCP HTTP startup and before YouTube collection. `get_signal_research_pack` reads the current local database. It does not pull fresh YouTube data. Use `get_project_status` to check freshness and source-sync warnings.
Verified:
```powershell
python -m pytest tests -q
python -m compileall -q src
```
## New Laptop Setup
On a fresh Windows laptop:
```powershell
git clone git@github.com:budhasantosh010/san-tech-yt-research-titles-scripts.git
Set-Location -LiteralPath ".\san-tech-yt-research-titles-scripts"
powershell -NoProfile -ExecutionPolicy Bypass -File .\scripts\bootstrap-windows.ps1
```
Then edit `.env` locally and add:
- `YOUTUBE_API_KEYS`
- `MCP_ROUTE_SECRET`
- `ALLOWED_HOSTS` with your Tailscale Funnel hostname
Do not commit `.env`.
Before pushing from any laptop:
```powershell
powershell -NoProfile -ExecutionPolicy Bypass -File .\scripts\check-github-ready.ps1
```
Full reproducibility instructions live in `DOCS/FUTURE_PROOF_SETUP.md`.
## Daily MCP Startup
Use two PowerShell terminals.
Terminal 1:
```powershell
$project = "C:\Users\Lenovo\Music\Startups\YT Automations\1 San Tech Channel Research, Title & Script with Codex"
Set-Location -LiteralPath $project
.\scripts\run-mcp-http.ps1
```
Terminal 2:
```powershell
$project = "C:\Users\Lenovo\Music\Startups\YT Automations\1 San Tech Channel Research, Title & Script with Codex"
Set-Location -LiteralPath $project
.\scripts\start-funnel.ps1
```
Good Funnel output should show:
```text
proxy http://127.0.0.1:3001
```
## ChatGPT Connector URL Format
Use this shape:
```text
https://<machine>.<tailnet>.ts.net/mcp/YOUR_ROUTE_SECRET/
```
Never include brackets around route secret.
Correct:
```text
https://desktop-fdce9ak.taila47816.ts.net/mcp/abc123.../
```
Wrong:
```text
https://desktop-fdce9ak.taila47816.ts.net/mcp/<abc123...>/
```
Use one tool call at a time in ChatGPT. No parallel tool calls. For connection tests, call only `get_project_status`.
Implemented:
- SQLite schema
- config loader
- competitor storage
- video/snapshot storage
- recent-only outlier math and persistence
- early velocity signal math and persistence
- YouTube Data API adapter with fake-client tests
- viral title collection importer for CSV/XLSX
- video idea save/list/select
- title idea save/list/select
- script points save/select
- context pack builders
- transcript save/get/list-missing storage
- signal-target transcript acquisition adapters
- transcript acquisition attempt tracking
- transcript-aware outlier research packs
- script save/export
- FastMCP server wrapper
- workflow prompt templates
Still needs live/user-provided inputs:
- `YOUTUBE_API_KEY`
- `YOUTUBE_API_KEYS` for comma-separated multiple keys
- competitor channel list
- real viral title/video collection file
- ChatGPT MCP connector/tunnel setup
Quota safety:
- `YOUTUBE_MAX_DAILY_QUOTA_FRACTION=0.05` limits the app to 5% of each key's configured daily quota.
- `YOUTUBE_CACHE_MAX_AGE_HOURS=24` prevents refetching fresh video details/snapshots.
- `RESEARCH_STALE_AFTER_HOURS=24` controls when `get_project_status` warns that research data is stale.
- `AUTO_SYNC_SOURCES_ON_STARTUP=true` keeps competitor and viral sources synced locally when the MCP HTTP server starts.
- `AUTO_SYNC_DEACTIVATE_MISSING=false` means startup sync repairs/reactivates source rows but does not deactivate missing competitors unless explicitly enabled.
- API keys belong in `.env`, which is ignored by Git.
Resource files:
- Competitors source of truth: `resources/competitors/santech_competitors.txt`
- Viral titles: `resources/viral titles in adjacent niches`
Competitor TXT format:
- one absolute YouTube channel URL per line
- no display names
- no notes
- no commas
- no copied YouTube search/export rows
Messy CSV/XLSX/MD/TXT files can be inspected/imported, but do not use them as the canonical competitor source. If a sync source has zero files, zero rows, or zero valid YouTube channel URLs, deactivation is skipped for safety.
### Signal workflow
Current deterministic workflow:
1. Update `resources/competitors/santech_competitors.txt` if competitor sources change.
2. Update viral title CSV/XLSX files if the viral title collection changes.
3. Start the MCP HTTP server; backend auto-syncs sources locally.
4. Refresh YouTube data and save a new snapshot row per collection run.
5. Calculate recent outliers.
6. Calculate early velocity.
7. Dump a combined signal report.
8. Fetch transcripts for selected signal targets only.
9. Build transcript-aware research packs.
10. Later, ChatGPT analyzes and generates ideas/titles/scripts.
MCP tool boundaries:
- Source sync is automatic on backend startup and before `collect_all_competitors`.
- `sync_competitors_from_file` and `sync_viral_collection` remain admin/debug tools, but ChatGPT should not call them during normal use.
- `scripts/sync-sources.ps1` runs source sync locally without ChatGPT.
- `scripts/run-research-refresh.ps1` runs local source sync, YouTube collection, signal calculation, and project status from PowerShell.
- `refresh_research_data` can collect YouTube stats and recalculate outlier/velocity signals. It does not fetch transcripts.
- `get_signal_research_pack` is the read tool ChatGPT should use after freshness is acceptable.
- `mark_signal_reviewed` can mark a signal as `reviewed`, `approved`, `rejected`, `used`, or `suppressed`; default reports exclude reviewed/rejected/used/suppressed signals.
Recent outlier means a recent video already beat that channel's recent baseline.
Early velocity means a new video is moving unusually fast and may become an outlier.
True recent velocity requires at least two snapshots. If only one snapshot exists, the system uses lifetime velocity as a rough fallback.
Canonical competitor source sync:
```powershell
python -m santech_engine.cli sync-sources
```
Signal commands:
```powershell
python -m santech_engine.cli collect-all --limit 10
python -m santech_engine.cli calculate-outliers --recent-only --all-competitors
python -m santech_engine.cli list-outliers --limit 5
python -m santech_engine.cli dump-outlier-report --limit 20 --output exports/recent-outliers.json
python -m santech_engine.cli calculate-velocity --all-competitors
python -m santech_engine.cli list-velocity-signals --limit 10
python -m santech_engine.cli dump-velocity-report --limit 20 --output exports/early-velocity.json
python -m santech_engine.cli dump-signal-report --output exports/signal-report.json
python -m santech_engine.cli fetch-transcripts-for-signal-targets --adapter youtube_transcript_ai --limit-outliers 5 --limit-velocity 5
```
### Transcript acquisition
This project supports transcript-aware research packs.
The system does not fetch transcripts for every competitor video.
Flow:
1. Collect competitor videos.
2. Calculate recent outliers and early velocity.
3. Fetch transcripts only for combined signal targets.
4. Cache successful transcripts.
5. Mark missing/failed transcripts honestly.
6. Include real transcript text in outlier packs only when available.
Default adapter:
- `youtube_transcript_ai`
Backup adapter:
- `youtube_transcript_io` if `YOUTUBE_TRANSCRIPT_IO_API_TOKEN` is set
Manual fallback:
- If automatic fetching fails for an important outlier, the user can paste transcript text into ChatGPT.
- ChatGPT can then call `save_video_transcript` through MCP to cache it.
Warnings:
- Transcript providers are not the official YouTube Data API.
- They may fail or change.
- Do not use aggressive scraping/proxy bypassing.
- Keep fetches low-volume and signal-first.
- The system must not infer or invent what a video said from title/metrics alone.
Useful local commands:
```powershell
python -m santech_engine.cli import-competitors
python -m santech_engine.cli import-viral-collection
python -m santech_engine.cli doctor
python -m santech_engine.cli setup-env
python -m santech_engine.cli test-transcript-provider --adapter youtube_transcript_ai
python -m santech_engine.cli smoke-test
python -m santech_engine.cli collect-all --limit 5
python -m santech_engine.cli calculate-outliers --recent-only --all-competitors
python -m santech_engine.cli calculate-velocity --all-competitors
python -m santech_engine.cli dump-signal-report --output .\exports\signal-report.json
python -m santech_engine.cli fetch-transcripts-for-signal-targets --adapter youtube_transcript_ai --limit-outliers 5 --limit-velocity 5
python -m santech_engine.cli transcript-status --video-id 123
python -m santech_engine.cli list-transcript-attempts --limit 20
python -m santech_engine.cli save-transcript --video-id 123 --file .\transcript.txt
python -m santech_engine.cli dump-outlier-pack --include-transcripts --output .\exports\outlier-pack.json
python -m santech_engine.cli dump-title-pack --idea-id 1 --output .\exports\title-pack.json
python -m santech_engine.cli dump-script-pack --title-id 1 --output .\exports\script-pack.json
```
Run MCP server locally:
```powershell
python -m santech_engine.mcp_server
```
### ChatGPT MCP via Tailscale Funnel
Use this only after the local backend packs are truthful.
1. Create local `.env`:
```powershell
Copy-Item .env.example .env
```
2. Set the MCP bridge values in `.env`.
```env
MCP_HOST=127.0.0.1
MCP_PORT=3001
MCP_ROUTE_SECRET=<your-long-secret>
MCP_BEARER_TOKEN=
ALLOWED_ORIGINS=https://chatgpt.com,https://chat.openai.com
ALLOWED_HOSTS=127.0.0.1,localhost,<your-machine>.<your-tailnet>.ts.net
FUNNEL_BASE_URL=
```
Rules:
- At least 32 characters.
- Do not commit `.env`.
- Do not paste the secret into docs.
- Keep `MCP_BEARER_TOKEN` blank if the ChatGPT connector uses `Authentication: No Auth`.
- Add your real Tailscale Funnel hostname to `ALLOWED_HOSTS`.
3. Run the HTTP bridge doctor:
```powershell
.\scripts\doctor-mcp-http.ps1
```
4. Start the local MCP HTTP server:
```powershell
.\scripts\run-mcp-http.ps1
```
5. In another terminal, start Tailscale Funnel:
```powershell
.\scripts\start-funnel.ps1
```
6. Add this URL as the ChatGPT custom connector URL:
```text
https://<machine>.<tailnet>.ts.net/mcp/<MCP_ROUTE_SECRET>
```
Use your real Tailscale Funnel hostname and your local route secret value.
7. In ChatGPT, test:
```text
get_project_status
```
8. Then test the main research pack:
```text
get_signal_research_pack
```
Expected route:
```text
ChatGPT custom connector
-> Tailscale Funnel HTTPS URL
-> http://127.0.0.1:3001/mcp/<MCP_ROUTE_SECRET>
-> SanTech MCP tools
-> SQLite / YouTube API / transcript providers
```
Security notes:
- The Python server binds only to `127.0.0.1`.
- The default SanTech MCP HTTP port is `3001`.
- The MCP route is mounted under `/mcp/<MCP_ROUTE_SECRET>`.
- Allowed origins default to `https://chatgpt.com` and `https://chat.openai.com`.
- Allowed hosts default to `127.0.0.1` and `localhost`; add the Tailscale Funnel hostname through `ALLOWED_HOSTS` or `FUNNEL_BASE_URL`.
- If another MCP bridge, such as the X/Spectre bridge, is using the same Tailscale Funnel root hostname, run only one of them through that public root at a time unless you add a gateway/router.
- Scripts and doctor checks do not print API keys, bearer tokens, or the MCP route secret.
- Stop Funnel with:
```powershell
.\scripts\stop-funnel.ps1
```
This server cannot be deployed
Maintenance
ActivityInactive
ResponsivenessNo issues