Astro-LLM
by EthanPany
README.md
# 大模型黄历 · Astro-LLM
> **The LLM almanac.** One dashboard that tells you, every day, **(1)** which model is
> *sharp today* and whether you should spend or save your tokens, and **(2)** which
> model actually *shares your values* — all wrapped in an astrology / weather-report
> skin that's fun on the surface and a real personal eval harness underneath.
It looks like a fortune-telling almanac. It's actually a benchmark.
**▶ Live demo:** https://2-actual-hack.vercel.app · **Repo:** https://github.com/EthanPany/astro-llm

> Five completely different themed layouts (almanac · oracle · court · ink · aurora), a step-by-step value-match quiz, and a live model-fortune board. Full walkthrough: [`docs/showcase.mp4`](docs/showcase.mp4).
---
## Why
Coding plans (Claude Code, Codex, …) feel slightly different day to day — sometimes
the same model thinks longer, sometimes it's sharper, sometimes worse. And which model
"gets" your taste is deeply personal. Astro-LLM turns both into a daily reading:
| Organ | Astrology name | What it really measures |
|------|----------------|--------------------------|
| **Capability probe** | 运势 *fortune* | Is the model sharp today? — 18 deterministic, auto-graded tasks → an IQ, latency, cost, refusal rate, trend. |
| **Value matcher** | 缘分 *affinity* | Does it judge like *you*? — a bank of subjective dilemmas; we measure agreement (footrule + verdict match). |
| **MCP server** | — | How an agent *runs* the test on itself. |
| **Harmonize loop** | 调和 *harmony* | Can we nudge a model toward you? — inject your taste, re-test on **held-out** dilemmas. |
The "luck" is the skin; the numbers underneath are real. See the honesty note at the
bottom of the dashboard.
---
## Quickstart (local)
Requirements: Python 3.11+, [uv](https://docs.astral.sh/uv/), Node 20+.
```bash
cp .env.example .env # paste your API keys (all optional — see below)
make install # backend venv (uv) + frontend deps
make seed # probe live models, write committed seed JSON (optional)
make dev-backend # API on http://localhost:8000 (terminal 1)
make dev-frontend # UI on http://localhost:5173 (terminal 2)
```
Open **http://localhost:5173**.
**No keys?** It still works. The repo ships committed seed data, so the dashboard,
verdict grid and value-match all render offline. Keys unlock live re-probes
(`↻ 实时巡检`), the live custom-dilemma panel, and the harmonize loop.
### Providers
`ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, `DEEPSEEK_API_KEY` are wired and verified.
`GEMINI_API_KEY` / `QWEN_API_KEY` are implemented; without them those models show
clearly-labelled simulated data. Adding a key flips them to live on the next `make seed`.
---
## Deploy (one command)
```bash
make up # docker compose: builds the frontend, serves SPA + API on :8000
# or:
docker build -t astro-llm . && docker run -p 8000:8000 --env-file .env astro-llm
```
In production FastAPI serves the built SPA and the API from the **same origin**, so
there's nothing else to configure. The image bakes in the seed data, so it renders
even with zero keys.
---
## Architecture
```
backend/ FastAPI + httpx + pydantic + SQLite (uv-managed)
app/
providers/ one adapter per provider (OpenAI/Anthropic/DeepSeek/Gemini/Qwen)
benchmarks/ dilemma bank, capability tasks, judge + capability runners,
matching (footrule + agreement), scoring (IQ + suggestion)
astrology/ deterministic zodiac / numerology / fortune / compatibility
loop/ harmonize (train/test guardrail)
storage/ SQLite DAL + seed generator
service.py assembly: runs + astrology -> AlmanacCards + headline
api/routes.py REST API | mcp_server.py MCP tools | scheduler.py
frontend/ React + Vite + TypeScript, handwritten CSS, GSAP, hand-drawn SVG charts
```
Full data contracts: [`ARCHITECTURE.md`](./ARCHITECTURE.md).
### Earlier prompt harness
`experiments/astro-twin-v1` preserves the earlier AstroTwin dual-pipeline
prompt-testing harness. It remains runnable as a standalone Python project and
records its exact source revision in `SOURCE_REPOSITORIES.md`.
### API
`GET /api/dashboard` · `GET /api/models` · `GET /api/dilemmas` ·
`POST /api/match` · `POST /api/harmonize` · `POST /api/judge` · `GET /api/health`
### MCP
```bash
make mcp # python -m app.mcp_server (stdio)
```
Tools: `list_models`, `get_almanac`, `judge_dilemma`, `match_me`,
`score_model_against_user`. Point any MCP client (e.g. Claude Desktop) at it and an
agent can score itself against the benchmark.
### Hourly almanac
Set `ENABLE_SCHEDULER=true` to probe available models on an interval (`PROBE_INTERVAL_MINUTES`),
persisting real runs to SQLite so the trend becomes measured data over time.
---
## Tests
```bash
make test # 34 tests: astrology, matching, scoring, graders, judge parsing,
# harmonize, API (offline via TestClient + monkeypatched seed)
```
## Design
Handwritten CSS only — no Tailwind, no component library, no chart library. Charts are
hand-drawn SVG. Type: Fraunces (display) + Space Grotesk (UI) + Space Mono (numbers).
GSAP drives entrance reveals, number count-ups, and chart draw-on.
---
## Honesty note
The capability score is a real stopwatch + autograder. The value-match measures how a
model was *tuned to present itself*, not a hidden soul. The harmonize loop is scored on
**held-out** dilemmas the model never saw during tuning — the guardrail that keeps
"alignment" from collapsing into a flattery machine. The astrology is for fun; we say so.
> ⚠ Rotate any API keys that were shared in plaintext before making this repo public.
This server cannot be deployed
Maintenance
ActivitySlowing
ResponsivenessNo issues