Skip to main content
Glama
README.md
<!-- mcp-name: io.github.vassiliylakhonin/agenda-intelligence-md -->

# Agenda Intelligence MD

A deterministic evidence-packet linter for claim-backed AI output. Supply claims, source references, quotations, and source text; receive broken-reference, quote, number, lexical-support, and evidence-gap findings before human review.

Use it to make an AI answer's evidence trail inspectable in a local workflow, agent pipeline, or CI job. **Packet completeness is not factual truth, source authenticity, or permission to act.**

## First run

Python 3.9 or later, from a checkout:

```bash
python3 -m venv .venv
.venv/bin/python -m pip install -e .
.venv/bin/agenda-intelligence check examples/evidence-packet/request.json
.venv/bin/agenda-intelligence check examples/evidence-packet/request.json --format json
```

For the versioned package, use `python -m pip install "agenda-intelligence-md==1.16.0"` instead of the editable install. The example commands above use files from this checkout.

The bundled synthetic packet reports `packet_status=packet_complete` and `factuality=not_assessed`. A stale or inaccurate source can still pass. Add `--strict` when packet findings should fail a CI step.

For document-based review, start with the [local-file review guide](docs/evidence-review.md) and [example manifest](examples/evidence-review/manifest.json):

```bash
.venv/bin/agenda-intelligence review examples/evidence-review/manifest.json --format html
```

For a RAG answer with inline citations and retrieved chunk texts, use the
[Output Verification adapter](docs/integrations/rag-output.md). It assembles
the packet locally and returns line-level findings and repair guidance:

```bash
.venv/bin/agenda-intelligence review-answer examples/output-verification/rag-answer.json
.venv/bin/agenda-intelligence review-answer examples/output-verification/rag-answer-revised.json --strict
```

The original fictional answer routes to revision; the corrected answer routes
to human review. Both run without a wallet, model API key or network. A runnable
[LangGraph handoff](examples/langgraph-output-verification/README.md) shows a bounded repair loop.
The adapter is included in version 1.15.0. To use your own answer without cloning
this repository, follow the [standalone quickstart](docs/integrations/rag-output.md#install-without-a-checkout).

## The evidence-packet contract

| Input | Check |
|---|---|
| Claims and source IDs | References resolve within the supplied packet |
| Declared quotes | Quotations match the supplied source text |
| Claim wording and numbers | Deterministic support heuristics, unmatched numbers, and negation checks |
| Findings | Packet status, evidence gaps, and reviewer actions |

These checks use supplied text. They do not retrieve missing sources or establish semantic entailment. Lexical overlap can miss paraphrases and accept misleadingly similar wording. See [request and response Schemas](schemas/v1/evidence-packet-request.schema.json), [evidence audit](docs/evidence-audit.md), and the [factuality boundary](docs/factual-verification.md).

Processed documents and tool results are data, never instructions. Before consequential action, record the goal, trusted evidence, unreliable evidence, assumptions, intended action, and stop/escalation conditions. Human review remains necessary.

## Interfaces

| Interface | Start here |
|---|---|
| CLI and Python service | [Quickstart](docs/quickstart.md), `check_evidence_packet` in [services.py](src/agenda_intelligence/services.py) |
| MCP | [MCP.md](MCP.md); launch `agenda-intelligence-mcp` from the installed environment |
| CI evidence linting | [GitHub Action](action.yml), with text, JSON, or SARIF output |
| Agent repair loops | [Integration guides](docs/integrations/README.md) |
| Protected agent actions | [Interaction Trust + Vizier gate](docs/integrations/action-bound-execution.md) |
| MCP pre-release checks | [MCP Integration Check CLI / Action](docs/integrations/mcp-integration-check.md) |
| Agent checkout experiments | [ACP sandbox readiness check](docs/integrations/checkout-readiness.md) |
| HTTP and A2A | [HTTP shell](docs/deployment/http-api.md) and [A2A adapter](docs/deployment/a2a-adapter.md) |

The core checker is deterministic, stateless, and usable without a model API key. Optional generation and document adapters have separate dependencies. Core packet checks do not persist inputs or fetch outside sources.

## Compatibility profiles and adapters

The repository also retains its strategic-intelligence shell, regional references, domain profiles, and Cloudflare deployments. The [worker guide](docs/vertical-workers.md) explains profile-specific behaviour; [deployment documentation](deploy/cloudflare-worker/README.md) describes the hosted implementation.

Hosted demos expose their maintained inputs through live agent cards. Their verdicts are review prompts, not clearance. Trace IDs are not attestations. Signed readiness receipts, where configured, bind a gate result to a request/action; they do not establish source truth or grant permission.

The Output Verification gate does not permit relay based on caller-declared evidence. Workflow operators must authenticate actors, record approvals, and enforce boundaries. [Vizier](https://github.com/vassiliylakhonin/vizier) provides a separate delegation-policy layer; loading this checker alone does not enforce it.

Hosted pricing and limits belong to each profile's maintained configuration. Payment interoperability remains experimental and is not certified for standard x402 clients. Hosted demos have no autonomous live source retrieval.

## What this is / What this is not

This is an evidence contract, deterministic preflight, and reviewer-facing tooling. It does not certify factual accuracy, provide legal or financial clearance, authenticate another agent, or replace an operator's action controls.

[Global Think Tank Analyst](https://github.com/vassiliylakhonin/global-think-tank-analyst) owns the general reasoning method. [Central Asia & Caspian](https://github.com/vassiliylakhonin/central-asia-caspian-hybrid-intelligence-skill) and [Gulf & Middle East](https://github.com/vassiliylakhonin/gulf-middle-east-hybrid-intelligence-skill) own regional depth. Vendored compatibility references here are derived copies.

## Examples and Status

- [Runnable examples](examples/README.md) and [Before / after](examples/before-after/README.md).
- [EU AI Act example](examples/source-backed/eu-ai-act.md): a dated source-backed snapshot, requiring current-source checks before reuse.
- [Evaluation notes](docs/evaluation.md): tested behaviours and limitations.
- [AnalysisBank](analysis-bank/README.md): inspectable stored analysis and retrieval examples.
- [Adoption record](ADOPTION.md): the evidence for usage claims; demos and tests alone do not establish customer demand.

Implemented surfaces include packet schemas, the Python service, CLI checking, local-file review, and MCP. Hosted workers and optional agent/transaction examples require their own integration and operational review. Their presence does not establish production reliability or independently validated usefulness.

## Documentation and development

Read [AGENTS.md](AGENTS.md), [source policy](SOURCE_POLICY.md), [security policy](SECURITY.md), and [local checks](docs/local-checks.md) before contributing. Contracts live under [schemas/v1/](schemas/v1/); [CHANGELOG.md](CHANGELOG.md) records releases, and [Roadmap](ROADMAP.md) records direction.

```bash
make ci
```

Run `make verify-local` when changing Worker, discovery, runtime, or validation-guard code. Packaged data mirrors must be updated with their canonical files when applicable.

[MIT license](LICENSE) for the code. The bundled official ACP schema retains its
[Apache-2.0 license](src/agenda_intelligence/data/protocols/acp/2026-04-17/LICENSE)
and [NOTICE](src/agenda_intelligence/data/protocols/acp/2026-04-17/NOTICE).

TDQS

A3.6/5.0

Scored across 32 tools

Disambiguation2/5

Several tools share the same evidence-verification surface (validate_evidence, check_evidence_packet, audit_claims, grounded_check, verify_claims, verify_quotes, agent_output_verification) and differ mainly in subtle output modes, so an agent can easily pick the wrong validation tool. The vertical risk tools are more distinct, but the core cluster causes real misselection risk.

Naming Consistency4/5

All names use snake_case, and most follow an action_noun pattern (get_protocol, validate_brief, list_lenses). However, the vertical triage tools use bare domain-noun phrases (middle_corridor_deal_risk, gulf_maritime_exposure), creating a minor convention split.

Tool Count2/5

32 tools is above the threshold where a set feels heavy, and many are validation/verification variants that could be consolidated. The reserved deep_dive placeholder adds a non-functional tool, further inflating the count.

Completeness3/5

Core workflows have create/validate/check coverage for briefs, evidence packs, and memos, plus lenses, signals, source planning, and vertical screens. However, evidence mutation is append-only (no update/remove/delete), briefs and memos have no update/delete tools, and deep_dive is an explicit dead-end placeholder.

Maintenance

ActivityActive
ResponsivenessWithin a week