Skip to main content
Glama
epoko77-ai

Korean Assembly Speech MCP

by epoko77-ai
README.md
# Korean Assembly Speech MCP

<!-- mcp-name: io.github.epoko77-ai/korean-assembly-speech-mcp -->

> Follow a Korean policy from bill status to the people, committees, and actual words behind it.

Korean & English queries · local-first · no paid API · FTS5 + E5 + FAISS · MCP + CLI

[한국어 문서](README.ko.md) · [Architecture](docs/architecture.md) ·
[Deployment](docs/deployment.md) · [MCP clients](docs/mcp-clients.md) ·
[Data sources](docs/data-sources.md)

![KASM demo](assets/demo.gif)

Ask “Who raised concerns about sovereign AI?”, “What happened to bill 2200001?”, or “Show the
subcommittee debate and the government's answers.” KASM traverses **bill/agenda → status →
committee/subcommittee → meeting → member → speech → surrounding Q&A**. It combines official
structured records with speech-level retrieval, acting as a small, evidence-first GraphRAG for
legislative research.

| | Bill lookup MCP | Korean Assembly Speech MCP |
|---|---|---|
| Result | an isolated API row | a connected bill-and-debate evidence graph |
| Retrieval | structured API lookup | lexical + multilingual semantic + RRF |
| English query | client chooses an API | directly retrieves Korean passages |
| Context | API metadata | status, committee/subcommittee, previous/next turns, Q&A |
| Verification | API record | original text, locator, meeting, official PDF |
| End-user key | often required | none for demo or public prepared-index MCP |

## Try it

The bundled demo is deliberately synthetic and clearly labeled. After package installation it
proves the CLI and MCP contract without a key, model download, or upstream data call.

```bash
uvx korean-assembly-speech-mcp demo
uvx --from 'korean-assembly-speech-mcp[mcp]' kasm mcp
```

For a public deployment, an MCP client mounts one endpoint without credentials:

```json
{
  "mcpServers": {
    "korean-assembly": {"url": "https://YOUR_HOST/mcp"}
  }
}
```

The operator's Open Assembly key belongs only in a separate refresh job. It is never required
by clients and is not present in the public search container.

## Search and synchronize

```bash
# Search a configured local/prepared index
kasm search "AI 기본법에 대한 정부 측 답변" --committee 과학기술정보방송통신위원회 \
  --database kasm.sqlite3 --vector-index kasm-vectors.faiss

# Operator-only official synchronization
kasm sync --source committee --assembly-term 22 --month 2025-01 --ingest \
  --all-pages --max-meetings 100 --database kasm.sqlite3

# Bills/agendas and their current processing results
kasm sync-bills --assembly-term 22 --all-pages --database kasm.sqlite3

# Local multilingual E5 + FAISS index
kasm index --database kasm.sqlite3 --output kasm-vectors.faiss --backend faiss

# stdio or stateless Streamable HTTP
kasm mcp --database kasm.sqlite3 --vector-index kasm-vectors.faiss
kasm mcp --transport streamable-http --host 0.0.0.0 --port 8000
```

Official synchronization uses only `open.assembly.go.kr` metadata and
`record.assembly.go.kr` minutes. `data.go.kr` and third-party parliamentary datasets are out of
scope. Set `ASSEMBLY_OPEN_API_KEY` only for `kasm sync`; raw caches, PDFs, `.env`, databases, and
indexes are ignored by Git.

## MCP tools

- `explore_issue` — one-query GraphRAG traversal across bills, committees, people and speeches
- `search_bills` — natural-language bill/agenda discovery with term, committee and status filters
- `get_bill_status` — current outcome plus connected debate evidence
- `search_speeches` — Korean/English policy-opinion retrieval with eight filters
- `get_speech` — full speech record and stable provenance
- `get_speech_context` — ordered surrounding speech turns
- `list_committees` — indexed committees and covered dates
- `list_meetings` — indexed meetings by committee, date, and type

Public clients need no Assembly key. The server operator uses a key only when producing refreshed
SQLite/FAISS artifacts; clients query those prepared artifacts exactly like a public search index.

## Measured gates

Checked-in scripts and artifacts make the results reproducible; synthetic evaluation data is
explicitly labeled and never represented as Assembly speech text.

| Gate | Corpus | Result |
|---|---:|---:|
| Official parser review | 20 distinct official PDFs | 20/20 reviewed boundaries pass |
| SQLite FTS5 latency | 50,000 synthetic speeches | p95 4.46 ms |
| E5 English → Korean | 25 qrels | Recall@10 1.00 |
| Bilingual hybrid | 50 queries | Recall@10 1.00, MRR@10 0.99 |

```bash
uv sync --extra dev --extra mcp --extra semantic
HF_HOME=.hf-cache uv run python scripts/evaluate_e5.py
HF_HOME=.hf-cache uv run python scripts/evaluate_hybrid.py
uv run python scripts/benchmark_fts.py
uv run ruff check . && uv run mypy && uv run pytest
```

## Architecture and data integrity

Refresh and search are separate trust domains:

```text
Open Assembly key → scheduled refresh → validated SQLite + FAISS artifacts
                                             ↓ atomic release
keyless MCP client → HTTPS /mcp → read-only prepared search artifacts
```

The fetcher allowlists the official minutes host, verifies PDF signatures, records SHA-256 and
retrieval metadata, reports parser failures, and refuses mismatched vector metadata. The public
ASGI service exposes `/mcp` and `/healthz`; mount `/data` read-only and rate-limit at ingress.

## Development

Python 3.12 and 3.13 are tested in GitHub Actions.

```bash
uv sync --extra dev --extra mcp
uv run ruff check .
uv run mypy
uv run pytest --cov=kasm
```

See [CONTRIBUTING.md](CONTRIBUTING.md), [SECURITY.md](SECURITY.md), and the roadmap in
[SPEC.md](SPEC.md).

## License and records

Code is licensed under Apache-2.0. Parliamentary records and excerpts remain subject to their
official source terms; source URLs and hashes are retained. The repository contains only small
review fixtures, not the full parliamentary corpus. See [DATA_LICENSE.md](DATA_LICENSE.md).

TDQS

C2/5.0

Scored across 5 tools

Disambiguation4/5

The tool names are distinct: get_speech, get_speech_context, list_committees, list_meetings, and search_speeches each target different resources or actions. However, without descriptions, an agent might not fully understand the nuance between get_speech and get_speech_context, causing slight ambiguity.

Naming Consistency5/5

All tools follow a consistent verb_noun pattern using underscores (get_speech, list_committees, search_speeches), which is predictable and easy to parse.

Tool Count4/5

With 5 tools, the set is focused and appropriate for a read-only API covering speeches, committees, and meetings. It is slightly minimal but fits the domain well.

Completeness4/5

The tool surface covers core operations: retrieving and searching speeches, listing committees and meetings. Minor gaps exist (e.g., no individual committee/meeting detail), but the main workflows are supported.