Skip to main content
Glama
PrattyT85

theosis-oracc-mcp

by PrattyT85
README.md
# theosis-oracc-mcp

Read-only MCP server wrapper for the [ORACC](https://oracc.museum.upenn.edu/) (Open Richly Annotated Cuneiform Corpus) JSON archive API, built for integration with the [Theosis](https://github.com/PrattyT85/theosis-mcp) theological research stack.

## Features

- **`list_projects`** — List public ORACC projects with optional filtering
- **`get_project_metadata`** — Project config, formats, and witnesses
- **`get_project_manifest`** — Available JSON files for a project
- **`list_project_texts`** — Catalogue entries with designation, language, period, genre
- **`get_text`** — Bounded CDL excerpt with transliteration and translation
- **`search_project`** — Search catalogue by designation, author, or title

## Setup

```bash
uv sync
uv run python -m oracc_mcp.server  # stdio transport
```

Or add to your Hermes config:

```yaml
mcp:
  servers:
    oracc-mcp:
      transport: stdio
      command: uv
      args: ["--directory", "/path/to/theosis-oracc-mcp", "run", "python", "-m", "oracc_mcp.server"]
```

## Testing

```bash
uv run pytest -v                    # offline tests only
uv run pytest -m live -v           # live tests (ORACC_LIVE=1 required)
ORACC_LIVE=1 uv run pytest -m live -v
```

## ORACC Attribution

This server is a read-only client for the [ORACC JSON archive API](https://oracc.museum.upenn.edu/json/). `projects.json` is used for discovery; project metadata, catalogues, corpora, and text editions are read from the current `/json/<project-archive>.zip` downloads. Archives are bounded and cached only in memory for the lifetime of one client; they are not persisted or redistributed. Results include the archive URL and member path as provenance. ORACC data is provided by the University of Pennsylvania and its project contributors; follow the licence and attribution for the specific source.

**Citation**: Steve Tinney & Eleanor Robson, 'Oracc JSON Data: A brief introduction for programmers', *Oracc: The Open Richly Annotated Cuneiform Corpus*, Oracc, 2019 [http://oracc.museum.upenn.edu/doc/opendata/json/]

## License

MIT — see [LICENSE](LICENSE).

TDQS

A3.9/5.0

Scored across 6 tools

Disambiguation4/5

Each tool targets a distinct resource or action, but list_project_texts and search_project both query the project catalogue and could be confused. The descriptions make the browse-versus-search distinction clear, so this is only a minor overlap.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern using snake_case: list, get, and search verbs are applied predictably to projects, metadata, manifests, texts, and project text lists. There are no abrupt style or casing changes.

Tool Count5/5

Six tools is a well-scoped set for a read-only ORACC corpus access server. Each tool covers a meaningful step in the workflow from discovering projects to retrieving text content, with no obvious redundancy.

Completeness4/5

The tool surface covers project discovery, metadata, file manifests, text listing, text retrieval, and catalogue search, which forms a complete read-only workflow. The main limitation is that get_text only returns a bounded excerpt rather than full text editions, but this appears intentional for output management.

Maintenance

ActivityMaintained
ResponsivenessNo issues