io.github.mischuh/canonic
Connects to local DuckDB .duckdb files as well as CSV and Parquet files, enabling the server to query and build context from these data sources.
Integrates with Keycloak as an OAuth 2.1 identity provider for remote/enterprise deployments, enabling role and tenant enforcement, masking, and run_sql gating for differently-scoped users.
Connects to local SQLite .db files as a data source, allowing the server to query and generate semantics from the database schema.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@io.github.mischuh/canonicWhat was last month's revenue using the canonical definition?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
canonic
The context layer that lets AI agents query your data correctly.
Point canonic at your database and it builds the context an agent needs to answer data questions accurately: definitions, relationships, business meaning, and the guardrails that stop confidently-wrong answers. It keeps that context up to date as your data changes, and it never touches your warehouse beyond reading it.
๐ Full documentation: https://docs.getcanonic.app
Package and image names below show the shape of each install channel; exact names are confirmed per release.
The problem
An AI agent connected straight to your warehouse sees tables and columns, not meaning. It doesn't know that revenue lives in orders.amount but excludes refunds, or that "active customer" has a specific definition your finance team agreed on. So it guesses. Schema access makes an agent fluent. It doesn't make it correct.
Real output, captured from a live run against the ecommerce example:
$ canonic sql "SELECT SUM(amount) FROM fct_orders"
โโโโโโโโโโโ
โ sum โ
โกโโโโโโโโโโฉ
โ 4050.50 โ
โโโโโโโโโโโThis total includes two refunded orders ($260), a confident, well-formatted number that's off by 6.4%.
$ canonic --json query --metrics revenue
{
"result": { "rows": [["3790.50"]] },
"compiled": {
"sql": "SELECT SUM(\"orders\".\"amount\") AS \"total_revenue\" FROM \"fct_orders\" AS \"orders\" WHERE \"orders\".\"status\" <> 'refunded'"
},
"metadata": {
"guardrails_fired": [{ "id": "revenue-excludes-refunds", "kind": "mandatory_filter" }]
}
}canonic resolves "revenue" to its canonical definition, compiles the guardrail into the SQL whether or not anyone asked for it, and returns the right number with the reasoning attached.
canonic is not a BI tool and not a chat interface: it's the layer that feeds the tools you already have (a BI dashboard, an agent, a notebook) correct, governed answers.
Related MCP server: autokg
The three layers
canonic's context lives in three committed surfaces: plain files in your git repo, reviewed like code.
Layer | File | Answers | Owned by |
Semantics |
| "How do I query this safely?" | auto-maintained |
Knowledge |
| "What does this mean to the business?" | auto-maintained |
Contracts |
| "Which definition is canonical, and what must the answer obey?" | human-owned |
Changes how the SQL runs โ semantics. A human needs it to trust the answer โ knowledge. Governs which definition is authoritative โ contracts. See Concepts: the three layers.
Install
uv (dev machines, primary):
uvx canonic --version # ephemeral, no install step
uv tool install canonic # persistent, global commandpip (fallback for environments without uv):
pip install canonicDocker (CI, headless, air-gapped):
docker pull ghcr.io/mischuh/canonic:latestVerify with canonic --version. Air-gapped install and offline wheels: see Installation.
Quickstart
The fastest path uses local connectors, no server, no network. Point at a SQLite .db or DuckDB .duckdb/CSV/Parquet file:
canonic setup
The wizard names your project, connects a source, optionally configures an LLM, drafts your semantics from the live schema, then runs a real query and shows the answer with its freshness and definition. Postgres or an LLM provider need a credential in an environment variable before you run canonic setup (canonic never stores secrets in canonic.yaml directly).
Don't have a database handy? examples/ ships 5 ready-to-run sample projects (dbt Jaffle Shop, e-commerce, vehicle rental, SaaS analytics, Dutch railway), see the guides.
You now have a working context layer committed to your repo:
canonic overview # what's askable
canonic query --metrics revenue --dimensions order_date # ask it
canonic review && canonic status # review what it draftedConnect your agent (MCP)
canonic exposes its capabilities over a local, on-demand MCP server, verified with Claude Code, Cursor, and Codex:
canonic mcp start{
"mcpServers": {
"canonic": {
"command": "uvx",
"args": [
"canonic",
"mcp",
"start",
"--project",
"/path/to/canonic/examples/rental",
"--suggestions"
]
}
}
}GUI-launched clients (Claude Desktop, Cursor) don't source your shell profile, so pass connection credentials via the config's env field, not export. Every answer-producing tool of the 11 registered (query, run_sql, search_knowledge, ...) returns a metadata band: resolved definition, guardrails fired, freshness, trust_score. On ambiguity, the agent gets a structured reason, not a guess.
See Connecting your agent for remote/enterprise deployment (--transport http, per-client bearer tokens) and the tools reference.
Want a full example with a real identity provider, including role/tenant enforcement end-to-end? scripts/local_idp spins up a local Keycloak plus a dockerized canonic serving the marketplace example via OAuth 2.1, so you can log in as differently-scoped test users and see masking, run_sql gating, and tenancy scoping applied live. Full walkthrough: Marketplace with Keycloak.
What you can rely on
Read-only. canonic never mutates your warehouse.
Propose-only, refuse-and-ask. Every change is a reviewable diff; ambiguous or unsafe answers get a structured reason, not a guess.
No LLM in the answer path. Queries compile deterministically. An LLM is optional and only drafts context during setup, four providers supported (Anthropic, OpenAI, any OpenAI-compatible endpoint, GitHub Copilot), see Configuring an LLM.
Local-first & air-gapped-capable. Run entirely on your machine; nothing has to leave your network.
Documentation
Quickstart: first answer in minutes.
Concepts: the three layers and the split rule.
CLI reference: every command, flag by flag.
MCP / agent integration: wiring canonic into Claude Code, Cursor, Codex, or any MCP client.
Guides: 5 ready-to-run example projects.
Reference: error codes and the full
canonic.yamlconfig schema.
License
This server cannot be deployed
Maintenance
Related MCP Connectors
Let AI agents query data and act across all your business apps via MCP.
Give AI agents identity, scoped access, trusted context, and verifiable actions through MCP.
Query your org's data in natural language โ read-only MCP access to SQL, NoSQL, files & warehouses.
Governed data discovery, exact queries, decisions, simulations, and runtime utilities over MCP.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to query live schema, lineage, and query-context across data warehouses, dbt projects, orchestration systems, and BI tools via MCP tools.Apache 2.0
- AlicenseNot gradedqualityDmaintenanceTurns warehouse/lakehouse tables into a governed entity-relationship knowledge graph exposed through MCP, enabling AI agents to answer multi-table business questions without hard-coded SQL or large schema prompts.Apache 2.0
- AlicenseNot gradedqualityBmaintenanceServes an automatically inferred semantic layer from your warehouse over MCP, enabling AI agents to query with correct business context, joins, and filters.2Apache 2.0
- AlicenseNot gradedqualityBmaintenanceEnables enterprise AI agents to query governed data lineage, PII-aware schema documentation, and semantic metadata from SQL logs via MCP, with role-based access and vector search.MIT