OMNIA MCP
OfficialClick on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@OMNIA MCPEvaluate this news article: classify event type, relevance score, and source support."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
OMNIA MCP brings JEV and local LAYA into agent workflows through the Model Context Protocol. Give an agent the context, define the questions that matter, and receive decisions it can use: classifications, scored assessments, and probabilities for specific claims.
A single request can assess several dimensions of a document, conversation, software update, or market observation. OMNIA validates the answers, applies your review thresholds, and records the result with its model identity, policy, and evidence references.
Intelligence in context
The intelligence is in evaluating meaning against context and criteria. The answer format makes that assessment usable by software.
For a news item, an agent can assess the event type, its relevance, who made a claim, and whether the supplied source supports it. For a token observation, it can assess whether fields are complete, whether reported conditions match a policy, and which review category applies. Each question addresses a distinct part of the decision.
Assessment | Question you can define | Native result |
Event classification | Is this a product launch, a security update, routine maintenance, or something else? |
|
Relevance and priority | How significant is the update under our editorial rubric? |
|
Evidence support | Does the supplied source support, contradict, or leave a claim unaddressed? |
|
Attribution | Is the claim made by the main author, a quoted author, or is attribution unclear? |
|
Completeness | Does this observation contain the fields required by our policy? |
|
Workflow routing | Which review queue or next step best fits this context? |
|
These are questions you configure through the same evaluation interface. Your application supplies the records and defines the available answers. Exact arithmetic, transaction limits, and mandatory requirements belong in deterministic application checks.
Several assessments can inform one action. An application can combine relevance, source support, and attribution before routing an item to an editor. It can combine field completeness, reported risk, and freshness checks before requesting a deeper market review. OMNIA returns the assessments; the application owns their combination and the action that follows.
Related MCP server: jev
JEV/LAYA, natively
OMNIA uses the providers' native decision types. It preserves selected labels, ordinal score semantics, probabilities, and available confidence fields in a common MCP response.
JEV | LAYA | |
Inference | TypeSafe's hosted evaluation API | Local inference through the public |
Model identity | Explicit version; default | Explicit checkpoint family and full revision SHA |
Input | Supplied text, JSON objects, or JSON arrays | The same request contract, within the selected checkpoint's input budget |
Answers |
|
|
Operation | API credential, bounded HTTP requests, retry controls | Optional runtime, isolated worker, configured device and threads |
Data path | Context and questions go to TypeSafe | Inference stays local; offline operation uses prepared dependencies and weights |
Choice compares the alternatives you define. Score evaluates an ordered rubric and can return a value between levels. Noul expresses the probability of a specified condition, retaining uncertainty instead of reducing the result to a bare boolean.
Questions in one request share the supplied state and are evaluated independently. For a dependent sequence, the calling agent can pass an earlier result into the next request. Provider selection is explicit; OMNIA does not silently move local context to a hosted service.
The decision path
flowchart LR
Agent[MCP client or agent] --> Context[Context and questions]
Context --> Model{JEV / LAYA}
Model --> Answers[Native typed answers]
Answers --> Policy[Validation and review policy]
Policy --> Record[Decision record]
Record --> AgentThe record includes a decision ID, input fingerprint, provider/model identity, answers, applied policy, review reasons, evidence references, timing, and available usage metadata. Repeated identical requests can reuse the configured cache. Earlier decisions can be retrieved by ID.
accepted means the response passed the configured checks. needs_review preserves answers that fall below a probability or margin threshold. Neither status grants permission to publish, trade, or execute an external action.
Evaluate several dimensions
This example asks three independent questions about one software update. It keeps the maintainer's statement separate from the quoted comment.
Call omnia_capabilities first, then send these arguments to omnia_evaluate:
{
"request": {
"state": {
"source": "GitHub",
"project": "Atlas SDK",
"main_author": "maintainer",
"main_update": "Version 2.4 is available. It adds streaming responses and fixes connection recovery.",
"quoted_comment": {
"author": "community_member",
"text": "This will double every customer's revenue."
}
},
"questions": {
"event_type": {
"type": "choice",
"instructions": "Classify main_update. Keep quoted_comment separate from the maintainer's statement.",
"criteria": {
"release": "An available software version with described changes",
"maintenance": "Routine repository upkeep without a release announcement",
"other": "Neither category is supported by the main update"
}
},
"relevance": {
"type": "score",
"instructions": "Rate the user-facing significance described in main_update, excluding quoted_comment.",
"criteria": [
"Routine bookkeeping with no described user-facing change",
"A minor improvement to existing behavior",
"A new capability or a described reliability fix"
]
},
"author_claim": {
"type": "noul",
"instructions": "Does main_update itself claim that customer revenue will double? Do not attribute quoted_comment to the maintainer."
}
},
"policy": {
"min_probability": 0.8,
"min_margin": 0.1
}
}
}The response contains a separate answer for each question, plus the decision record and review status. The request above is an illustrative input; recorded provider results are in the validation report.
Quickstart
Requires Python 3.11–3.14. Install the released source with uv:
uv tool install "git+https://github.com/Omniaeye/omnia-mcp.git@v0.1.0"
omnia-mcp version
omnia-mcp doctorTo use JEV, set these variables in the environment that launches your MCP server:
OMNIA_DEFAULT_PROVIDER=jev
OMNIA_ALLOWED_PROVIDERS=jev
TYPESAFE_API_KEY=<your TypeSafe API key>Register the installed executable in your client:
{
"mcpServers": {
"omnia": {
"command": "omnia-mcp",
"args": ["serve"]
}
}
}Use an absolute executable path when the client cannot find omnia-mcp. The server reads its process environment; it does not automatically load .env files. .env.example lists every setting.
Client | Configuration example |
Claude Code | |
Claude Desktop | |
Cursor | |
VS Code |
These recipes follow each client's documented configuration format. See compatibility for client checks and supported transport.
Run LAYA locally
Install from a checkout when configuring the local runtime:
git clone https://github.com/Omniaeye/omnia-mcp.git
cd omnia-mcp
git checkout v0.1.0
python -m venv .venvPlatform | Install the local runtime |
macOS / Linux |
|
Windows PowerShell |
|
Configure the provider and checkpoint:
OMNIA_DEFAULT_PROVIDER=laya
OMNIA_ALLOWED_PROVIDERS=laya
OMNIA_LAYA_MODEL=english
OMNIA_LAYA_REVISION=55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851
OMNIA_LAYA_DEVICE=cpu
OMNIA_LAYA_THREADS=2Point your MCP client at .venv/bin/omnia-mcp or .venv\Scripts\omnia-mcp.exe. This configuration selects the checkpoint used in the recorded local smoke tests. typed-decisions and multilingual are also supported checkpoint-family selections; choose and validate the corresponding revision and input limits for your task.
For offline operation, prepare the exact snapshot and dependencies before setting OMNIA_LAYA_OFFLINE=true. To enable both providers, use OMNIA_ALLOWED_PROVIDERS=jev,laya, configure each, and select one through request.provider or OMNIA_DEFAULT_PROVIDER. See configuration.
MCP interface
Tool | Arguments | Purpose |
|
| Evaluate the questions against one supplied state. |
|
| Evaluate a bounded collection of requests with per-item results or errors. |
| None | Inspect enabled providers, question types, and operating limits. |
|
| Retrieve a persisted decision without another inference call. |
Resources: omnia://guide for usage and omnia://schemas for complete contracts. The same schemas are available locally:
omnia-mcp schemasdoctor checks configuration and package presence without loading a model or making a provider call. serve starts the stdio server; running omnia-mcp with no subcommand does the same.
Operating controls
Configure input size, question count, batch size, concurrency, queue capacity, deadlines, retries, and per-minute/per-day evaluation allowances. Responses are validated against the submitted questions, including rubric identity and rounded probability consistency. Failed requests return explicit errors.
The SQLite ledger retains answers, labels, rubric descriptions, evidence references, fingerprints, and model metadata. It does not retain the full input state by default. Cache lifetime and automatic retention are configurable; records can still contain sensitive application data.
Evidence references are carried with a decision, not fetched. Supply the actual source text when a question needs it. Applications own collection, exact rule enforcement, publication, and execution. OMNIA MCP provides typed assessments rather than free-form text generation or autonomous trading.
Validation
The release has recorded real JEV and local LAYA inference checks, stdio integration tests, Docker checks, and Windows/Linux CI across Python 3.11–3.13.
The multilingual JEV acceptance suite covers software updates, quoted authorship, missing fields, and explicit snapshot rules in English, Portuguese, Chinese, and Japanese. Its initial 100 evaluations exposed a rounding-validation defect. After the correction, a fresh run passed 20/20 cases and 48/48 answer checks. The complete controlled suite passed 175 tests.
Read the results and original evidence. These results describe the recorded cases; choose thresholds against reviewed examples from your own workload.
OMNIA software
Project | Focus |
The agent-facing interface to hosted JEV and local LAYA. | |
OMNIA's independent Laya integration fork for local decision workflows. | |
Source-content assessment and configurable feed decisions. | |
Market, holder, liquidity, and risk observation assessments. |
These are separate repositories. This MCP package connects directly to TypeSafe and the public LAYA runtime; it does not automatically install or run the News and Trading pipelines.
Documentation and development
Architecture · Configuration · Compatibility · Releasing · Security · Contributing
python -m pip install --upgrade pip
python -m pip install -e . --group dev
python -m ruff check .
python -m ruff format --check .
python -m pytest
python -m build
python -m twine check dist/*The default suite requires no provider credentials. Live inference remains explicit opt-in.
Provider references: JEV state and context · Composing assessments · Evidence assessment · LAYA upstream.
License and attribution
Apache License 2.0. Original OMNIA MCP code is maintained by OMNIA EYE Corporation. JEV is provided through TypeSafe; LAYA is developed by Convai Innovations. Upstream authorship, licenses, service terms, and model terms remain with their respective projects. See NOTICE.
Available Tools
4 toolsomnia_capabilitiesARead-only
Report configured capabilities; this does not load or probe models.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is known. The description adds meaningful behavioral context by explicitly stating the tool does not load or probe models, which goes beyond the annotations and clarifies side-effect expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. Every clause earns its place by stating the tool's purpose and its key non-behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless introspection tool with an output schema and readOnly annotations, the description fully covers what an agent needs to know: what it reports and what it does not do. No critical context is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With zero parameters, there is nothing for the description to explain. Baseline for zero-parameter tools is 4, and no parameter semantics are needed beyond the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Report') and resource ('configured capabilities'), making the purpose immediately clear. It also distinguishes itself from the evaluation-focused siblings by clarifying it does not load or probe models.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: this tool reports configured capabilities, not model behavior. It does not explicitly name sibling alternatives or state when-not-to-use, but the negative clarification ('does not load or probe models') helps an agent select it for capability introspection rather than evaluation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
omnia_evaluateBDestructive
Evaluate a request, persist its audit record, and expire records under retention.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| model | Yes | |
| usage | No | |
| cached | No | |
| policy | Yes | |
| status | Yes | |
| answers | Yes | |
| evidence | Yes | |
| identity | Yes | |
| provider | Yes | |
| created_at | Yes | |
| input_hash | Yes | |
| decision_id | Yes | |
| duration_ms | Yes | |
| review_reasons | Yes | |
| schema_version | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false, and idempotentHint=false. The description adds context beyond those: it explicitly says the tool persists an audit record and expires records under retention, helping an agent anticipate side effects before calling. It does not contradict the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that names the main action and side effects without excess. 'Expire records under retention' is vague jargon, but the overall length and ordering are efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the highly complex input schema and presence of an output schema, the description is minimal but covers the key side-effect concerns. It omits usage context and any clarification of what 'evaluate' produces or how the retention expiration behaves, so it is adequate but not complete for a tool of this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description must compensate, but it only says 'a request' — essentially repeating the parameter name without explaining state, questions, policy, evidence, or provider. The schema itself is rich, but the description adds no meaning beyond what the parameter name already conveys.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action (evaluate) on a specific resource (a request) and names two side effects (persist audit record, expire records under retention). The singular 'a request' distinguishes it from the batch sibling omnia_evaluate_batch, but it does not explicitly contrast with omnia_get_decision or omnia_capabilities.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus omnia_evaluate_batch, omnia_get_decision, or omnia_capabilities. The singular wording implies single evaluation, but the description never tells an agent 'use this for one request and omnia_evaluate_batch for multiple' or 'use omnia_get_decision to retrieve decisions'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
omnia_evaluate_batchBDestructive
Evaluate a batch, persist audit records, and expire records under retention.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| items | Yes | |
| schema_version | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds behavioral context beyond annotations by specifying that it persists audit records and expires records under retention. This complements the destructiveHint=true and readOnlyHint=false annotations without contradiction, providing useful side-effect information for an agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence that front-loads the core action 'Evaluate a batch' and appends two side-effect clauses. It contains no wasted words and is structurally efficient, though it sacrifices depth for brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of the tool (batch evaluation with multiple question types, policies, and evidence), the description is far too sparse. It does not explain when to use it versus single evaluation, expected outcomes, or how the batch relates to other tools. While an output schema exists, the description itself is inadequate for a tool of this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, meaning the description provides no explanation of the 'request' parameter or its nested structure. The schema itself is self-documenting with types and constraints, but the description fails to add any semantic meaning, leaving the agent to rely entirely on the raw schema without guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb 'Evaluate', a resource 'a batch', and adds concrete side effects (persist audit records, expire under retention). The word 'batch' implicitly distinguishes it from the sibling omnia_evaluate, which likely handles a single evaluation, making the tool's purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for multiple evaluations via the word 'batch' but does not explicitly compare against omnia_evaluate or omnia_get_decision. There is no guidance on when not to use this tool, nor any mention of prerequisites or alternatives, leaving the agent to infer context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
omnia_get_decisionARead-only
Read a persisted local decision by its opaque ID, without re-evaluating it.
| Name | Required | Description | Default |
|---|---|---|---|
| decision_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| model | Yes | |
| usage | No | |
| cached | No | |
| policy | Yes | |
| status | Yes | |
| answers | Yes | |
| evidence | Yes | |
| identity | Yes | |
| provider | Yes | |
| created_at | Yes | |
| input_hash | Yes | |
| decision_id | Yes | |
| duration_ms | Yes | |
| review_reasons | Yes | |
| schema_version | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the read-only nature is covered. The description adds that it operates on a 'persisted local decision' and explicitly says 'without re-evaluating it', which discloses that it does not trigger recomputation. This adds context beyond the annotation, though it could mention behavior like error handling for nonexistent IDs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. It states the action, resource, and a key distinguishing characteristic, earning its place with zero waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values are already documented. The tool is simple (one parameter, read-only), and the description covers its purpose, the nature of the ID, and the key distinction from siblings. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. The description labels decision_id as an 'opaque ID', clarifying that it is a meaningless token rather than a human-readable key. This adds semantic meaning beyond the schema's type and pattern constraints, making the parameter's role clear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'Read' and identifies the resource as 'a persisted local decision', clearly distinguishing it from evaluation tools. The phrase 'without re-evaluating it' further separates it from siblings like omnia_evaluate, making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool (retrieve a stored decision without recomputation) but does not explicitly name alternatives or exclusion conditions. The phrase 'without re-evaluating it' hints at the contrast with evaluation tools, but it would be stronger to explicitly state 'use this when you have a decision ID and need the stored result, not a fresh evaluation.'
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
omnia_capabilities - First observed
omnia_evaluate - First observed
omnia_evaluate_batch - First observed
omnia_get_decision
TDQS
Scored across 4 tools
The tools are mostly distinct: evaluate handles single requests, evaluate_batch handles groups, get_decision reads persisted results, and capabilities reports configuration. There is minor potential confusion between evaluate and evaluate_batch, but the descriptions clarify the difference.
Most tools follow an omnia_<verb>_<object> pattern, but omnia_capabilities is a bare noun and omnia_evaluate lacks an object, breaking the otherwise consistent style. All names share the omnia_ prefix and snake_case, so the inconsistency is moderate rather than chaotic.
Four tools is a well-scoped size for this server's purpose. Each tool covers a distinct operation—single evaluation, batch evaluation, reading decisions, and capability reporting—without unnecessary bloat.
The surface covers the core evaluation lifecycle: evaluate single, evaluate batch, and read persisted decisions, plus capability discovery. A listing or search over persisted decisions would round it out, but the existing tools support the primary workflow without dead ends.
Maintenance
Related MCP Connectors
- DatagoatOAuthio.datagoat
Governed decision engine: yes/no, score, choice and rank answers about cases, from past outcomes.
Deterministic contextual decision arbitration and action routing for autonomous software. Takes current state, context, or intent plus caller-supplied candidate actions, state transitions, routes, refusals, escalations, tools, or models and returns a deterministic ordered candidate field. Also provides persistent machine representations for memory, retrieval, indexing, and downstream coherence measurement.
Deterministic allow/require_approval/deny verdicts for agent actions, before they happen.
Append-only decisions with provenance, supersession, retrieval, and audited MCP actions.
Related MCP Servers
- AlicenseAqualityBmaintenanceProvides coding agents and CI with a typed decision layer that sends bounded state and questions to Jev, then returns deterministic actions for review, risk assessment, requirement checks, and verification.91,141 npmMIT
- AlicenseNot gradedqualityCmaintenanceEnables agents to evaluate single records or batches with dynamically authored typed questions, returning structured decisions and probabilities.MIT
- AlicenseAqualityCmaintenanceEnables coding or reasoning agents to request structured judgments from TypeSafe's Jev model at decision points, including choices, scores, claim verification, and code reviews, with probabilities and confidence returned as data.5173 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables typed decisions (yes/no, choice, score) with transparent preflight checks, honest confidence reporting, and health monitoring.510 npmApache 2.0