Nunchi Bridge
by zadong12
README.md
# Nunchi Bridge
**Alexa+ relays messages between family members who don't share a language — in the register their relationship requires — and turns replies into plans both sides confirm.**
Built for the Amazon Developer Hackathon 2026 (Alexa+ track) as a **self-hosted MCP server** (MCP spec 2025-11-25, Streamable HTTP) plus a **simulated Alexa+ household web app**. It does not run on a real Echo: Alexa+ add-on developer tools are in partner-only preview.
**Demo video (1:37):** https://youtu.be/MXJBo56_VDk — playback of a recorded live run; MCP calls, register checks and saves executed for real; AI text-to-speech narration.
## What it does
1. Minjun (English) says: *"Tell grandma I can't make it Sunday. Something came up."*
2. The host agent looks up the household (`get_household`): grandson → grandmother requires polite Korean (해요체).
3. For comparison, a generic translation with no family context (a recorded Gemini output: `할머니한테 일요일에 못 간다고 전해줘. 일이 생겼어.`) is **refused** by `check_register`: plain endings are curt from a grandchild to a grandmother. `send_message` would refuse it too, whatever the LLM wants. It is never sent.
4. The agent's own relationship-aware draft passes the gate and is delivered. Minjun hears a back-translation to confirm the meaning.
5. Halmeoni replies in Korean: `그럼 토요일에 와. 갈비찜 해놓을게. 김치통 꼭 가져오고.`
6. Minjun gets it in English. The agent calls `propose_plan`, and **each item must quote its evidence verbatim** from her message. Nothing is scheduled yet.
7. Minjun presses **Confirm**, and `commit_plan` stores the Saturday visit and "bring the kimchi container" in shared state.
## What the register gate is, and is not
**It is:** a deterministic safety gate, separate from the LLM, that catches *selected high-risk* Korean register violations before a message is sent. It checks three things:
- plain (반말) sentence endings;
- plain or confrontational "you" (너/너희/당신);
- a few honorific-vocabulary slips. These are soft hints and do not block.
**It is not:** a complete judge of Korean politeness. Known limits (clause-level plain speech, subject honorifics, context-dependent 당신, family norms) are listed in [`docs/LINTER_RED_TEAM.md`](docs/LINTER_RED_TEAM.md). Native-speaker review results are in [`docs/register_eval_set_reviewed.csv`](docs/register_eval_set_reviewed.csv) and summarized in the red-team doc; no general accuracy figure is claimed.
## Family setting: the family's norm, not a textbook rule
By default, grandson → grandmother requires polite Korean. Some close families text their grandparents in casual speech (반말), and for them a refusal would be wrong. So a person can switch a pair to **casual allowed** (the toggle in the simulator header; `POST /household/settings` on the server).
- This is **not an MCP tool**. An agent must never be able to relax the gate it is checked by.
- Even in a casual family, plain/confrontational "you" (너, 당신) toward an elder is still refused.
- Formal (하십시오체) or archaic endings are never blocked. They get a non-blocking warning with a suggested 해요체 replacement, e.g. `연락드리겠습니다` → `연락드릴게요`, `잡수셨사옵니까?` → `드셨어요?`. The suggestion is only offered when a safe rule applies; irregular stems get no guess.
## Implemented vs. simulated
| Part | Status |
|---|---|
| MCP server, 7 tools, Streamable HTTP, protocol 2025-11-25 | Implemented, tested |
| Register gate (Korean) | Implemented, tested; scope as above |
| Propose → human confirm → commit, with verbatim-evidence check | Implemented, tested |
| Host agent loop (LLM chooses tools, real MCP calls) | Two live flows passed on Windows on 2026-09-30 with Gemini 3.5 Flash-Lite and prompt revision `original-evidence-language-v3-no-demo-answers`: the unchanged demo and a new library/borrowed-books reply. Each quoted the Korean original, required human confirmation, and stored one event and one task. These two checks do not establish general reliability. See [VALIDATION_V3_KO.md](VALIDATION_V3_KO.md) |
| REPLAY MODE | Default script is converted from the new unchanged-demo live run (`live-recordings/validation-demo-v3.json`, SHA-256 in metadata). A full browser replay passed. LLM decisions and wording are recorded; MCP calls, register checks and state changes run again. Previous default and hand-authored scripts are retained |
| Literal translation baseline | Comparison only, never sent. The first requested live flow produced a plain generic translation (REFUSED); the second produced a polite generic translation (PASS). Both relationship-aware agent drafts passed immediately. The default replay reproduces the first flow's recorded REFUSED baseline and labels its source; this is not an agent refusal-and-rewrite sequence |
| Alexa+ device, voice | Simulated. Browser speech recognition/synthesis (Chrome) or typed input |
| Family casual-speech setting | Implemented, tested (in replay the scripted agent draft stays polite; a live LLM is told the pair allows casual speech) |
| Other languages | Not built. The policy table and gate interface are language-keyed |
## Run it
```bash
pip install -r requirements.txt -c constraints-windows.txt # Python 3.11+; pins MCP SDK v1 (2.x removed FastMCP)
python -m nunchi.server # MCP server -> http://127.0.0.1:8000/mcp
python -m nunchi.sim_app # simulator -> http://127.0.0.1:8080
# optional live mode:
GEMINI_API_KEY=... python -m nunchi.sim_app
# optional fixed demo date (weekday resolution): NUNCHI_TODAY=2026-10-07 python -m nunchi.server
```
On Windows PowerShell, see the [Windows first-run guide](WINDOWS_FIRST_RUN_KO.md). Do not enable billing for this demo; check current Gemini free-tier model availability before choosing `GEMINI_MODEL`.
## Recorded live demo
In live mode the simulator records responses (tool arguments, translations, final replies) to `live-recordings/latest-live-run.json`, and archives a copy after confirmation. API keys and HTTP headers are not recorded. After one clean grandchild → grandmother → confirm flow:
```bash
python -m nunchi.recording --input live-recordings/successful-live-run.json --output nunchi/replay/demo_script.json
```
The converter rejects incomplete, failed or unconfirmed runs. Only the proposal's message ID and next-Saturday date become placeholders. Replay accepts only the recorded inputs (use **Use demo line**); start the server with `NUNCHI_TODAY=2026-09-30` to reproduce the recorded dates exactly. See [VALIDATION_V3_KO.md](VALIDATION_V3_KO.md). [LIVE_RUN_REPORT_KO.md](LIVE_RUN_REPORT_KO.md) is the earlier run's historical report.
Tests: `python -m unittest discover -s tests` (unit, MCP-over-HTTP integration, end-to-end demo scenario).
Any MCP client works against the server, for example MCP Inspector with transport "Streamable HTTP" and URL `http://127.0.0.1:8000/mcp`.
## Tools
| Tool | Purpose |
|---|---|
| `get_household` | Members, languages, today's date, required register per sender→reader pair |
| `check_register` | Per-sentence register status with English explanations |
| `send_message` | Deliver in the reader's language; **refuses** Korean text that fails the gate |
| `get_inbox` | Messages for a member |
| `propose_plan` | Events/to-dos from a message; evidence must be a verbatim quote; nothing scheduled |
| `commit_plan` | Store only items the named person confirmed; once per proposal |
| `get_board` | Read-only shared state |
The demo reset is an HTTP route (`POST /demo/reset`), deliberately not an MCP tool. The host also blocks any LLM attempt to call `commit_plan`.
## Data and privacy
The household is fictional. State is in memory only.
## License
MIT (see `LICENSE`).
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues