Skip to main content
Glama

omniseek_graph

Map connections between papers, authors, and entities using typed evidence edges; run bounded neighborhood, path, similarity, and accretion queries from any resolved anchor.

Instructions

Use WHEN you want HOW two entities connect, or what OmniSeek already knows AROUND a paper / author / entity; read-only, budgeted projections of its accumulated evidence graph with typed edges (ONE graph).

Fully-qualified MCP name: mcp__omniseek__omniseek_graph (server name is omniseek; there is no omniseek-eye server).

Everything OmniSeek perceives is a statement with provenance ("X relates to Y, per Z"); the evidence graph accumulates those typed edges in ONE store surfaced through N indexes. It stores FACTS + labeled CANDIDATES, never verdicts: mechanical world edges (tier M: cites, authored, coauthored, affiliated, published_in, about, observed, exact-id same_as) and alignment CANDIDATES (tier A: title-fingerprint / fuzzy-name same_as, name-match authored, string mentions, signal conflicts). Judgment (claims, gaps, identity rulings) is tier J and is STRUCTURALLY excluded from OmniSeek's store — the views project structure, YOU judge it.

ONE STABLE VERB: omniseek_graph(view, args). view names the projection; args is that view's OWN parameter dict (the views are an open family, their params disjoint per view, so the ABI is (view, args), not a flat union). THE SCHEMA IS FROZEN: future views and future per-view parameters change NOTHING in this signature; a no-view call returns the live view catalog (the surface is self-describing), and content is NEVER inlined (every view returns node ids + labels

  • edge tuples, so you zoom with the other eye tools).

• view="find", args={"label_query": ..., "kind"?: ...} -> the ENTRY POINT. A node id is minted by the backend that knows it, so a NAME ("Siva Reddy") is not a node until you resolve it: find does the mechanical token/substring match over node labels and returns candidate ids + kinds. Every other view takes an anchor id; find is how you get one. • view="stats", args={} -> counts by kind / type / tier. The cheap orientation call (also the cold-start check: see below). • view="neighborhood", args={"anchor": ..., "depth"?<=2, "types"?, "policy"?, "max_nodes"?} -> the bounded subgraph around a node. • view="between", args={"a": ..., "b": ..., "types"?, "policy"?, "max_nodes"?} -> bounded connection paths between two anchors, the "how do these relate" question. Bidirectional BFS, <=2 hops per side, up to 8 shortest paths; capped when more existed. No path -> paths:[]. • view="voices", args={"doc_ids": [...], "policy"?: ...} -> collapse a doc set to distinct upstream VOICES via same_as + authored; the independence counter (mirror collapse, shared-speaker docs merge, docs with zero evidence land in unresolved and are NEVER counted as a voice). Input capped at 64 doc ids by explicit error; non-doc: ids come back in skipped. • view="since", args={"anchor": ..., "date": ..., "types"?, "max_nodes"?} -> the accretion log: what accreted around an anchor after a date (YYYY-MM-DD or full ISO), STORED edges only, tier + method shown on every row, NO collapsing (accretion is a fact stream, not an identity question). Derived edges carry no timestamps and are structurally absent. The sensor consumer. • view="similar", args={"anchor": , "k"?: ...} -> vector-nearest doc CANDIDATES for an anchor doc, method align:embed, by RANK (k is a budget, never a score threshold). PROPOSALS only, never collapsed by any policy; verify, then ratify with omniseek_ruling. Coverage: any doc with an embedded title, ranked across the UNION of the indexed vec matrix AND the thin-title vec_thin matrix (P7), so a thin arXiv original and an indexed post rank in ONE space; candidates may be thin docs. A doc with no vector in either store (un-embedded yet) -> an error naming that. NON-GOAL: vec_thin does NOT feed search's recall arm (similar + future P5 consumers only). A no-view call (view="") returns the live view catalog: each view's params + one-line blurb, DERIVED from the registry, so new views appear here without a client restart. Identity rulings are WRITTEN via omniseek_ruling (this tool stays read-only, hence gather-safe).

policy (an arg on the collapsing views) = conservative | working | exploratory: NAMED METHOD-SETS for how far to trust identity (same_as) edges when collapsing, NOT numeric thresholds (a hand-picked constant is pseudo-precision; the METHOD is the honest epistemic unit, as recall fuses by rank only):

  • conservative: collapse on exact-id equality only (DOI / OpenAlex / ORCID / arXiv; default)

  • working: conservative + agent identity rulings from graph_rulings.json

  • exploratory: working + title-fingerprint / fuzzy-name alignment CANDIDATES Identity is an EVIDENCE-CARRYING EDGE, never a destructive merge: same_as edges carry tier + method, collapse is reversible, and a not_same_as ruling beats a same_as. OmniSeek never MAKES an identity ruling; it only applies the ones you already recorded.

COLD START (set the expectation or the first stats reads as failure): documents and same-work edges are LIVE FROM DAY ONE (derived over recall's docs — the wall is born pre-populated by construction). Document THIN rows (title + url only, from NON-indexed sources) now accumulate from EVERY search (stats.node_kinds.document_thin), so the perception history is complete, not just the ~40 enumerable sources. Entity kinds (work / person / institution / venue / topic) still fill in as the P2/P3 write taps ship and calls happen; emptiness of those kinds early is CORRECT, not broken.

BUDGETS (the no-silent-caps discipline): depth is clamped to <=2, max_nodes caps the node count, and any capped result stamps capped: true so a bounded view never reads as complete. Schema + the view registry live in omniseek.core.recall.graph.

FAIL-OPEN: a graph failure returns an error dict, never an exception — the graph is memory, it must NEVER break search or recall.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
argsNo
viewNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.2.0

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full burden, and it thoroughly discloses behavior: read-only and gather-safe, budget clamping with capped:true, fail-open returning error dicts instead of exceptions, structural exclusion of tier J verdicts, non-destructive identity policy, input caps, and vector-missing errors. It also states that content is never inlined and that derived edges carry no timestamps. This is exemplary transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long, but the tool is a polymorphic eight-view API with no schema or annotations to lean on, so most sentences earn their place. It is front-loaded with the core query intent, uses clear bullets and capitalization for structure, and avoids filler. A few concepts are restated for emphasis, but the density is justified by the complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, no annotations, and only a two-property input schema, the description is remarkably complete: it covers all current views, step-by-step usage for find-to-anchor flows, budget semantics, identity policy behavior, cold-start expectations, failure mode, and the frozen ABI. An agent has everything needed to invoke the tool correctly and interpret its constraints.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must fully compensate for the bare view/args schema. It does: every view name is documented with its args, types, defaults, and semantic meaning, and the self-describing no-view catalog behavior is specified. The policy values are defined as named method-sets rather than numeric thresholds, and budget params are explained. This is exactly the compensation needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise usage statement: "Use WHEN you want HOW two entities connect, or what OmniSeek already knows AROUND a paper / author / entity." It names the tool's resource (accumulated evidence graph with typed edges), identifies it as read-only, and differentiates it from sibling tools like omniseek_ruling and the other eye tools. The one-stable-verb (view, args) design is clearly explained, leaving no ambiguity about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance for each view, names the entry point (view="find"), and states when alternatives apply: writes go to omniseek_ruling, zooming uses other eye tools, and similar does not feed search recall. It even warns about cold-start expectations so an agent does not misinterpret early stats. This is far beyond minimum viable usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.