movie-rec
# movie-rec
`movie-rec` is a movie and television recommendation service implemented as a FastMCP server. It stores audience-specific evidence in SQLite, retrieves candidates from MovieLens ALS factors and TMDB, and records recommendation runs and reactions.
## Architecture
The repository contains the following parts:
- `movie_rec/server.py`: FastMCP tools and optional GitHub OAuth authorization.
- `movie_rec/service.py`: application-level operations shared by the MCP surface.
- `movie_rec/store.py`: SQLite schema, immutable evidence events, derived profiles, recommendation runs, and candidate exposures.
- `movie_rec/retrieval.py`: candidate assembly, filtering, ranking, and evidence attachment.
- `movie_rec/als.py` and `movie_rec/movielens.py`: MovieLens ALS training artifacts and per-audience fold-in retrieval.
- `movie_rec/tmdb.py`, `movie_rec/tmdb_recommendations.py`, and `movie_rec/discover.py`: TMDB title resolution, recommendation expansion, and property-based discovery.
- `scripts/`: offline MovieLens training, fixture generation, and evidence-shape inspection.
- `tests/`: unit and integration tests with synthetic data.
The service recognizes two neutral audience namespaces, `primary` and `shared`. Evidence and recommendation history remain isolated by audience.
## MCP tools
The server exposes these tools:
- `resolve_title`: resolve a movie or television title through TMDB and store its canonical record.
- `add_evidence`: append an immutable evidence event.
- `get_evidence_context`: return current evidence and a derived profile for an audience.
- `get_candidates`: retrieve and filter candidates from the available retrieval lanes.
- `record_recommendations`: mark shortlisted and shown titles for a recommendation run.
- `get_recommendation_history`: return recent recommendation runs and shown titles.
- `record_reaction`: store a reaction linked to an audience and optional run.
- `discover_titles`: query TMDB by media type, genre, year, language, keywords, votes, and sorting.
## Local setup
Python 3.11 or later and [uv](https://docs.astral.sh/uv/) are required.
```sh
uv sync --all-extras
cp .env.example .env
```
Replace the placeholder values in `.env`. For local shell use, load the file before starting the service:
```sh
set -a
. ./.env
set +a
uv run movie-rec serve --transport stdio
```
The default database path is under the current user's local data directory. `MOVIE_REC_DB_PATH` and `MOVIE_REC_ARTIFACTS_PATH` can select project-local or external runtime storage.
For Streamable HTTP transport, provide the bind port and configure every GitHub OAuth variable listed in `.env.example`:
```sh
uv run movie-rec serve --transport http --host 127.0.0.1 --port 8000
```
The HTTP transport restricts access to the configured GitHub login. Deployment-specific TLS termination, public routing, and secret storage are outside this repository.
## External APIs
TMDB access is required for title resolution, metadata, discovery, and TMDB recommendation expansion. Set either `TMDB_READ_ACCESS_TOKEN` or `TMDB_API_KEY` in the file selected by `MOVIE_REC_TMDB_ENV_FILE`.
GitHub OAuth is required only for HTTP transport. Configure a GitHub OAuth application and set `MOVIE_REC_PUBLIC_URL`, `MOVIE_REC_GITHUB_CLIENT_ID`, `MOVIE_REC_GITHUB_CLIENT_SECRET`, `MOVIE_REC_OAUTH_SIGNING_KEY`, and `MOVIE_REC_ALLOWED_GITHUB_LOGIN`. Stdio transport does not require GitHub OAuth.
MovieLens training downloads the checksum-pinned MovieLens 32M archive when the configured local archive is absent. Train and map artifacts with:
```sh
uv run --extra offline python scripts/train_movielens.py train
uv run --extra offline python scripts/train_movielens.py map
```
## Tests
Run the full test suite with:
```sh
uv run pytest
```
Regenerate the deterministic ALS fold-in fixture with:
```sh
uv run --extra offline python scripts/generate_als_fixture.py
```
The generator uses seeded synthetic factors and the `implicit` library's user-factor recalculation as the independent reference implementation.
## Privacy
This public repository contains no personal recommendation evidence, preference history, taste-derived model rows, family context, runtime database, deployment secrets, or private infrastructure configuration. Tests and committed fixtures use synthetic or fictional data. Runtime evidence, databases, generated model artifacts, exports, backups, and secret files are excluded by `.gitignore`.
## License
This project is licensed under the MIT License. See `LICENSE`.
TDQS
Scored across 8 tools
Each tool has a clearly distinct purpose: resolving titles, adding evidence, retrieving context, generating candidates, recording recommendations, fetching history, recording reactions, and discovering titles. No two tools appear to overlap in functionality, and the descriptions provide strong differentiation.
All tool names follow a consistent verb_noun pattern: resolve_title, add_evidence, get_evidence_context, get_candidates, record_recommendations, get_recommendation_history, record_reaction, discover_titles. The pattern is systematic and predictable, making tool selection straightforward.
With 8 tools, the server is well-scoped for its purpose of movie recommendation and user feedback. Each tool addresses a distinct part of the workflow without redundancy, and the count is within the typical well-scoped range (3-15).
The tool set covers the core lifecycle: resolving user inputs, adding evidence, generating candidates, recording recommendations, retrieving history, and capturing reactions. Minor gaps exist (e.g., no explicit delete/update for evidence or reactions), but these are not critical for the main recommendation flow and can be worked around.