Skip to main content
Glama
J-X0

Kingfisher Streaming Inference

by J-X0

Kingfisher Streaming Inference

Incremental model scoring over an append-only event log, with a hash-chained audit trail so every decision can be reviewed and independently re-verified. Built for Quarryman Labs (Project Kingfisher).

Why it is shaped this way

The governing constraint is that every decision must carry a reviewable audit trail. That drives the architecture:

  • Events land in an append-only log (kingfisherstreaming/log.py). There is no update or delete path; duplicate ids are rejected.

  • Each event is scored incrementally (kingfisherstreaming/scoring.py) using Welford's online mean/variance. Scoring one event is O(1) and never rescans the log, which is what makes the stream tractable over time.

  • Every decision is written to a hash-chained audit log (kingfisherstreaming/audit.py). Each record's SHA-256 covers its contents and the previous record's hash, so any later edit, reorder, or deletion is detectable with verify().

Model behaviour sits behind a provider interface (kingfisherstreaming/providers/base.py) with a deterministic, offline stub (stub.py) used by the tests and an HTTP-backed implementation (real.py). The full test suite runs offline with no API key.

Related MCP server: DCL Evaluator

Layout

kingfisherstreaming/
  types.py          domain types (Event, Score, Decision, ...)
  log.py            append-only event log (in-memory + optional JSONL)
  scoring.py        Welford running stats + incremental scorer (core algorithm)
  audit.py          hash-chained, tamper-evident audit log
  engine.py         wires log + scorer + audit into one ingest path
  providers/
    base.py         ScoringProvider interface
    stub.py         deterministic offline provider
    real.py         HTTP-backed provider (optional 'real' extra)
tests/

Install and test

make install      # creates .venv and installs with the dev extra
make test         # runs pytest

Override the interpreter if you manage your own venv:

make test PY=python

Usage

from kingfisherstreaming import StreamingInferenceEngine, Event
from kingfisherstreaming.providers.stub import StubProvider

engine = StreamingInferenceEngine(StubProvider(), threshold=3.0, min_samples=30)

score, record = engine.ingest(
    Event(event_id="lease-001", stream="office-cbd", timestamp=0.0, value=100.0)
)
print(score.decision, score.normalized_score)
assert engine.verify_audit()   # re-checks the entire chain

During cold start (fewer than min_samples observations on a stream) the decision is REVIEW, not a guessed pass/flag: with too little history a z-score is not trustworthy, and saying so is part of the audit trail.

To score against a real endpoint:

pip install '.[real]'
export KINGFISHER_SCORING_ENDPOINT=https://scoring.internal/score
export KINGFISHER_API_KEY=...

Running the MCP server

The server speaks JSON-RPC 2.0 over stdio (one JSON object per line), implemented with the standard library only. Logs are JSON on stderr; stdout carries only protocol frames.

python -m kingfisherstreaming serve

Tools exposed: ingest_event, verify_audit, stream_stats, export_audit. verify_audit and export_audit exist so a reviewer can pull and re-check the hash chain through the same interface that produced the decisions.

Example exchange (request on stdin, response on stdout):

{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"ingest_event","arguments":{"event_id":"lease-001","stream":"office-cbd","timestamp":0.0,"value":100.0}}}

Batch scoring a file

python -m kingfisherstreaming score events.jsonl

One event object per line. Malformed lines are logged and skipped, not fatal; the exit code is non-zero only if the resulting audit chain fails verification.

Configuration

All configuration is read from the environment and validated at startup:

Variable

Default

Meaning

KINGFISHER_PROVIDER

stub

stub (offline) or real (HTTP)

KINGFISHER_THRESHOLD

3.0

z-score magnitude that flags an event

KINGFISHER_MIN_SAMPLES

30

per-stream samples before scoring leaves REVIEW

KINGFISHER_LOG_PATH

unset

JSONL path to persist events

KINGFISHER_AUDIT_PATH

unset

JSONL path to persist the audit chain

KINGFISHER_MAX_EVENT_BYTES

65536

reject events larger than this

KINGFISHER_MAX_FEATURES

256

reject events with more features than this

KINGFISHER_LOG_LEVEL

INFO

logging level

Invalid configuration exits non-zero with a message on stderr rather than a traceback.

Available Tools

4 tools
export_auditA

Export audit records as a list for offline review.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states that the tool exports a list, but it does not clarify whether this is a read-only operation, whether it has side effects, whether authentication or permissions are needed, or what the exact output format is. For a simple export tool this is a notable gap, but the description is not misleading.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence that contains the action, resource, output format, and purpose. There is no filler or repetition, and every word adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool with no output schema, the description is reasonably complete: it names the resource, the action, the output shape ('list'), and the purpose. It could have elaborated on the content or format of the audit records, but this is a minor omission given the low complexity of the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there are no parameter semantics to document. The schema coverage is 100%, and the description correctly focuses on behavior rather than parameter details. A baseline of 4 is appropriate for a parameterless tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Export'), the resource ('audit records'), and the output form ('as a list'). It also conveys the intended purpose ('for offline review'). This distinguishes it from siblings like ingest_event, verify_audit, and stream_stats, which involve different operations on related resources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for offline review' gives some implied context for when to use the tool, but the description does not explicitly mention alternatives or exclusion criteria. Since sibling tools exist, some guidance on when to choose export_audit over them would have been helpful, though the core use case is inferable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ingest_eventB

Score one event incrementally and append it to the audit chain.

ParametersJSON Schema
NameRequiredDescriptionDefault
valueYes
streamYes
payloadNo
event_idYes
featuresNo
timestampYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It does reveal a key behavior: appending to the audit chain, which implies a persistent, mutating effect. Still, it omits crucial behavior such as idempotency, validation, duplicate handling, and whether the operation is reversible.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler. It front-loads the core action and effect, and every word carries meaning. Despite being terse, it is appropriately concise for what it conveys.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has six parameters, nested objects, no annotations, and no output schema, yet the description provides only a one-line behavioral summary. An agent would lack guidance on how to populate payload and features, what response to expect, and what side effects or prerequisites matter. This is inadequate for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it does not explain any of the six parameters. Terms like 'Score' hint at the 'value' parameter, and 'event' hints at event_id, but payload, features, stream, and timestamp are entirely unexplained. This leaves the agent guessing about how to construct a valid request.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states an action ('Score'/'append'), a resource ('event'), and the target ('audit chain'), which broadly distinguishes it from the sibling tools verify_audit, stream_stats, and export_audit. However, 'Score' is somewhat ambiguous and could be read as evaluation rather than ingestion, so it is not perfectly crisp.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'one event incrementally' implies this tool is for single-event ingestion rather than bulk operations or verification, but there is no explicit statement of when to use this tool versus siblings. No alternatives or exclusions are mentioned, leaving usage mostly to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

stream_statsC

Return running statistics for a stream.

ParametersJSON Schema
NameRequiredDescriptionDefault
streamYes

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the behavioral disclosure burden. The verb 'Return' suggests a read-only operation, but nothing is said about side effects, live versus historical data, the meaning of 'running', error conditions, or whether the stream must exist. This is minimal transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The sentence is short and free of filler, so it is concise. But the extreme brevity under-specifies the tool's meaning and context, making it less useful than a slightly longer, structured description would be.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is too thin for a tool with no output schema, no annotation, and one undocumented parameter. There is no return-value shape, no parameter guidance, and no relationship to sibling tools, so an agent cannot confidently invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the description does not define the single 'stream' string parameter beyond repeating the noun. It does not specify whether 'stream' is a name, ID, ARN, or what format is expected, so an agent must guess the input semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Return') and names a resource ('running statistics for a stream'), which distinguishes it from the audit/ingest/export siblings. However, 'running statistics' and 'stream' are not defined, so it falls short of a fully specific, unambiguous purpose statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance is given about when to use this tool versus the sibling tools. The phrasing implies a stats-retrieval use case, but there are no exclusions, alternative references, or contextual cues to help the agent decide among the available tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_auditA

Recompute and verify the entire audit hash chain.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It reveals only the action ('recompute and verify') but not whether the operation mutates state, what side effects might occur, how expensive it might be, or what failure behavior looks like. This is a significant gap for a tool that likely performs a substantial computation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One precise sentence with no filler. The verb and target are front-loaded, and every word contributes to understanding the tool's purpose. This is appropriately concise for a no-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description lacks information about the outcome of the verification: what does it return, does it throw on failure, does it produce a report, and is it a read-only check or a recomputing mutation? With no output schema and no annotations, the agent is left without enough context to confidently use the tool and interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there is no schema information to supplement. The baseline for 0 parameters is a 4, and the description correctly avoids inventing parameters. It does not need to explain parameter semantics because none exist.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Recompute and verify') and a specific resource ('the entire audit hash chain'), making the tool's function unambiguous. It is clearly distinct from the sibling tools ingest_event, stream_stats, and export_audit, which handle ingestion, statistics, and export respectively.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The use case is implied: if you want to verify the audit hash chain, this is the tool. However, there is no explicit guidance about when to prefer it over alternatives, nor any exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 4 tool updatesv0.1.0
    • First observedexport_audit
    • First observedingest_event
    • First observedstream_stats
    • First observedverify_audit

TDQS

A3.5/5.0
Disambiguation5/5

Each tool has a clear, distinct role: ingest_event adds data, verify_audit validates the chain, stream_stats provides metrics, and export_audit retrieves records. There is no functional overlap between any of the four tools.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern in snake_case (ingest_event, verify_audit, stream_stats, export_audit). The naming is predictable and immediately communicates each tool's action and target.

Tool Count5/5

With 4 tools, the set is tightly scoped to the server's apparent purpose of event ingestion and audit verification. Each tool earns its place, and the number is well within the typical range for a focused toolkit.

Completeness4/5

The set covers the core lifecycle of an audit chain: ingest, verify, export, and stream-level statistics. Minor gaps exist, such as no per-record retrieval or stream configuration, but the essential workflows are complete.

Maintenance

ActivityInactive
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    Enables AI agents to sign decisions with post-quantum cryptographic proofs and maintain secure audit trails for compliance. It provides tools for stamping events, verifying chain integrity, and exporting audit data across industries like finance and healthcare.
    4
    87
    MIT

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/J-X0/quarryman-labs-streaming-inference-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server