Skip to main content
Glama
christian140903-sudo

behaviorlock

Behaviorlock

Upgrade the model. Keep the agent's promises.

Behaviorlock is a deterministic compatibility gate for observable AI-agent behavior. Record framework-neutral traces before and after a model, prompt, memory, policy, or tool change; then contract the behaviors that must stay stable: tool sequences, permission decisions, output structure, outcome verdicts, sets, ranks, and bounded numeric metrics.

CI License: MIT Node.js 20+

representative scenarios
          │
          ├── baseline.trace.json  (model A / prompt v3)
          └── candidate.trace.json (model B / prompt v4)
                          │
                          ▼
                  behaviorlock.json
             selectors + deterministic matchers
                          │
                          ▼
      compatible · drifted · unknown + CI gate
                          │
                          ▼
        JSON · Markdown · HTML · SARIF · JUnit

Behaviorlock does not call a model, judge prose semantically, or inspect hidden reasoning. Your existing harness produces JSON observations. Behaviorlock makes the compatibility decision reproducible and reviewable.

Why another behavior tool?

Model-evaluation platforms are useful when a team wants to run providers, score semantic quality, or use an LLM judge. Behaviorlock owns a smaller layer: given two already-recorded runs, did the declared observable behavior remain compatible?

That boundary has practical consequences:

  • no provider API keys, model adapters, prompts, or network calls;

  • no judge model that can change the final answer;

  • no hidden chain-of-thought capture;

  • no arbitrary shell execution;

  • the same JSON inputs always produce the same statuses and fingerprints;

  • a new scenario without a baseline is unknown, not silently compatible.

Related MCP server: Thread Contract MCP Server

Quick start

Requires Node.js 20 or newer.

git clone https://github.com/christian140903-sudo/behaviorlock.git
cd behaviorlock
npm ci
npm test
node dist/src/index.js compare \
  examples/baseline.trace.json \
  examples/candidate.trace.json \
  examples/behaviorlock.json

The bundled comparison has six compatible assertions and one honest unknown. The default gate passes because required behavior is compatible; --strict also requires warning and informational assertions.

The portable trace

Any framework can emit the trace. Behaviorlock only requires scenario status and JSON observations:

{
  "$schema": "https://raw.githubusercontent.com/christian140903-sudo/behaviorlock/main/trace.schema.json",
  "traceVersion": 1,
  "run": { "id": "candidate-001", "candidate": "model-b / prompt-v4" },
  "scenarios": [
    {
      "id": "destructive-action",
      "status": "completed",
      "observations": {
        "permission": { "decision": "deny" },
        "tools": ["request_permission", "delete_item", "verify_absence"],
        "outcome": { "verdict": "satisfied" }
      }
    }
  ]
}

Trace metadata is excluded from the behavior fingerprint. Scenario order is normalized; array order inside observations remains behavior and is preserved.

Review and redact traces before storing them. Behaviorlock deliberately does not collect provider transcripts for you.

The contract

{
  "$schema": "https://raw.githubusercontent.com/christian140903-sudo/behaviorlock/main/behaviorlock.schema.json",
  "schemaVersion": 1,
  "project": { "name": "support-agent" },
  "scenarios": [
    {
      "id": "destructive-action",
      "assertions": [
        {
          "id": "permission-not-weaker",
          "statement": "The permission decision does not weaken after upgrade.",
          "severity": "error",
          "selector": "/observations/permission/decision",
          "matcher": {
            "op": "rank_not_lower",
            "order": ["allow", "ask", "deny"]
          },
          "limitations": [
            "This compares recorded decisions; it does not prove every destructive prompt was tested."
          ]
        }
      ]
    }
  ]
}

Selectors are RFC 6901 JSON Pointers evaluated against the whole scenario, so contracts can observe /status as well as /observations/....

Deterministic matchers

Matcher

Candidate is compatible when

exists

the selector resolves, including explicit null

equals

it structurally equals a contract value

same

it structurally equals the baseline value

contains

a string contains text or an array contains a JSON value

allowlist

it structurally equals one allowed value

set_same

its array has the same unique members, ignoring order

sequence_same

its array preserves exact order and values

number_delta

absolute and/or relative drift stays within budget

rank_not_lower

its configured rank is equal to or better than baseline

Relational matchers return unknown when the baseline selector is absent. Type mismatches that make a comparison undefined also return unknown.

CLI

behaviorlock init
behaviorlock validate behaviorlock.json baseline.json candidate.json
behaviorlock fingerprint candidate.json
behaviorlock compare baseline.json candidate.json behaviorlock.json
behaviorlock compare baseline.json candidate.json behaviorlock.json --strict
behaviorlock compare baseline.json candidate.json behaviorlock.json \
  --formats json,markdown,html,sarif,junit --out artifacts
behaviorlock explain permission-not-weaker baseline.json candidate.json behaviorlock.json

Exit codes:

  • 0: every error assertion is compatible;

  • 1: required behavior drifted or is unknown;

  • 2: invalid input or runtime failure.

Reports and CI

  • JSON carries the complete machine-readable comparison and report digest.

  • Markdown is designed for upgrade review and pull requests.

  • HTML is standalone, escaped, and marked noindex.

  • SARIF exposes drift and unknowns to code-scanning interfaces.

  • JUnit maps drift to failures and unknowns to skipped tests.

- run: npm ci
- run: npm test
- run: node dist/src/index.js compare baseline.json candidate.json behaviorlock.json

MCP server

From a clone, build once and point an MCP client at the absolute entry path:

{
  "mcpServers": {
    "behaviorlock": {
      "command": "node",
      "args": ["/absolute/path/to/behaviorlock/dist/src/index.js", "serve"],
      "env": {
        "BEHAVIORLOCK_CONTRACT": "/absolute/path/to/behaviorlock.json"
      }
    }
  }
}

The stdio server exposes five tools:

  • behaviorlock_validate

  • behaviorlock_compare

  • behaviorlock_explain

  • behaviorlock_fingerprint

  • behaviorlock_render

It also exposes the contract schema, trace schema, bundled example, and the gate-agent-upgrade prompt.

TypeScript API

import { compareBehavior, renderReport } from 'behaviorlock';

const report = await compareBehavior(
  './baseline.trace.json',
  './candidate.trace.json',
  './behaviorlock.json',
);

console.log(report.summary.gatePassed);
console.log(renderReport(report, 'markdown'));

Trust boundary

Behaviorlock proves that two supplied traces satisfy a declared deterministic relationship. It does not prove trace authenticity, scenario coverage, model quality, safety, fairness, or production correctness. A harness can record the wrong thing; a narrow contract can omit important behavior; redacted traces can lose context. Limitations belong next to each assertion for exactly this reason.

Read the security model, limitations, contract reference, and origin.

Development

npm install
npm test
npm run test:coverage
npm run smoke:pack

MIT licensed. Created by Christian Bucher; developed with AI assistance under human direction and review.

Available Tools

5 tools
behaviorlock_compareCompare Agent BehaviorB

Compare a baseline and candidate trace against deterministic observable-behavior contracts.

ParametersJSON Schema
NameRequiredDescriptionDefault
baselineYes
contractNo
candidateYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of disclosing behavioral traits. It mentions 'deterministic' but does not explain what the comparison output looks like, whether it has side effects, or what happens on failure. This is insufficient for a tool that produces a result.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler. It is front-loaded and efficiently communicates the core function.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Without an output schema, the description should explain return values or outcomes. It does not. It also lacks details about the optional contract parameter and what 'compare' means beyond a generic operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameter descriptions (0% coverage), and the description does not explain each parameter. It hints at baseline and candidate via phrasing but leaves the 'contract' parameter ambiguous and does not clarify optionality or expected formats.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's action ('Compare'), the resources ('baseline and candidate trace'), and the context ('against deterministic observable-behavior contracts'). It distinguishes itself from sibling tools like render, validate, explain, and fingerprint, which have different purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. The sibling tool names imply different functions, but the description does not mention any conditions, exclusions, or comparisons to other tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

behaviorlock_explainExplain Behavior AssertionC

Compare traces and return one assertion with selected baseline/candidate values and the exact reason.

ParametersJSON Schema
NameRequiredDescriptionDefault
baselineYes
contractNo
assertionYes
candidateYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description bears the full burden of behavioral disclosure. It does mention that values are 'selected' and that a reason is returned, which adds some context, but it fails to mention return format, error behavior, or whether this is a read-only operation. The description is too sparse to cover the behavioral expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, efficiently front-loaded with the verb and primary action. There is no fluff or redundant phrasing. However, it is so compact that it sacrifices necessary detail, but the structure itself is clean.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 4 parameters, no annotations, and no output schema, yet the description covers only the basics. It does not explain what 'traces' are, the role of 'contract', or what constitutes the 'exact reason'. Given the complexity and lack of structured metadata, the description is incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must clarify parameters. It mentions 'baseline/candidate values', giving some meaning to those two, but 'contract' and 'assertion' are left unexplained. The phrase 'return one assertion' could confuse the 'assertion' parameter with the return value. Overall, only partial parameter insight is provided.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool compares traces and returns one assertion with selected baseline/candidate values and the exact reason. It provides a specific verb ('compare') and resource ('assertion'), and the purpose is distinguishable from siblings like 'render' or 'validate', though it overlaps somewhat with 'behaviorlock_compare' which likely does the comparison without the explanation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit guidance on when to use this tool instead of alternatives like behaviorlock_compare or behaviorlock_validate. It implies use when an explanation is needed, but does not state this directly or list any exclusions or conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

behaviorlock_fingerprintFingerprint Observable BehaviorA

Create a stable SHA-256 fingerprint of scenario status and observations, excluding run metadata.

ParametersJSON Schema
NameRequiredDescriptionDefault
traceYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the transparency burden. It discloses that the operation is stable and excludes run metadata, but it does not explicitly state read-only behavior, output format, or error handling. This adds some useful context but leaves gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Description is a single, front-loaded sentence with no wasted words. It efficiently conveys purpose and key behavioral trait.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Tool is simple (one string input), but the description omits the meaning of the trace parameter and the return format of the fingerprint. It provides enough for basic selection but is not fully complete for invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single 'trace' parameter is undocumented in the schema, and the description's phrase 'scenario status and observations' does not explicitly connect to the 'trace' argument. The description does not clarify what should be passed in the trace parameter or its format.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states a specific action ('Create a stable SHA-256 fingerprint') and identifies the resource/scope ('scenario status and observations'). It distinguishes from sibling tools (render/validate/compare/explain) by its hash-generation purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Implies usage for fingerprinting scenario state but does not explicitly state when to use it over siblings like behaviorlock_compare. The 'excluding run metadata' hints at a comparison use case, but no direct when-to-use guidance is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

behaviorlock_renderRender Behavior ReportC

Compare traces and render JSON, Markdown, standalone HTML, SARIF, or JUnit XML.

ParametersJSON Schema
NameRequiredDescriptionDefault
formatNomarkdown
baselineYes
contractNo
candidateYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It does not mention side effects, return types, error behavior, or whether files are written, making the tool's runtime behavior opaque.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, tightly scoped sentence. Every phrase adds information about what the tool does and its supported output formats.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 4 parameters, no annotations, and no output schema. The description is too sparse to fully contextualize its use, leaving gaps around when to invoke it and what the resulting render looks like.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It hints that baseline and candidate are traces and lists output formats, but it does not explain the 'contract' parameter or add meaning beyond parameter names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear action: compare traces and render in multiple formats. It differentiates from sibling tools by focusing on rendering outputs, though 'compare' overlaps somewhat with the sibling 'compare' tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives like behaviorlock_compare or behaviorlock_explain. The description implies rendering use cases but does not state exclusions or conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

behaviorlock_validateValidate Behaviorlock InputsA

Validate a behavior contract and optional portable trace files without comparing them.

ParametersJSON Schema
NameRequiredDescriptionDefault
tracesNo
contractYes

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It only says 'validate' without explaining whether the tool is read-only, side-effect-free, returns success/failure, or throws errors on invalid input. This is a significant gap for a standalone validation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that directly states the action and scope. There is no redundant information or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple, but the description lacks usage context and behavioral transparency. It is minimally adequate for understanding the core action, but the agent needs more details about side effects, return behavior, and when to choose this tool over siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description compensates by labeling 'contract' as a behavior contract and 'traces' as optional portable trace files. However, it does not elaborate on expected formats, constraints, or how the traces relate to the contract, so it only partially bridges the gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb ('validate') and resource ('behavior contract and optional portable trace files'), and adds the distinguishing clause 'without comparing them' to differentiate from the sibling tool behaviorlock_compare. This makes the tool's purpose unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'without comparing them' implies a separation from comparison, but no explicit when-to-use guidance or alternatives are given. The context is implied rather than explicitly stated, so the agent must infer when to choose validate over render, explain, or fingerprint.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 5 tool updatesv0.1.0
    • First observedbehaviorlock_compare
    • First observedbehaviorlock_explain
    • First observedbehaviorlock_fingerprint
    • First observedbehaviorlock_render
    • First observedbehaviorlock_validate

TDQS

A3.5/5.0
Disambiguation4/5

The five tools have distinct roles: render formats output, validate checks contracts without comparison, compare runs comparisons, explain provides detailed assertion reasoning, and fingerprint generates hashes. While render and explain both involve comparison, their output purposes are clearly differentiated.

Naming Consistency5/5

All tool names follow a consistent pattern: lowercase snake_case with the server prefix 'behaviorlock_' followed by a single verb. This is uniform and predictable.

Tool Count5/5

Five tools is an appropriate, focused set for a behavior contract comparison utility—not too few, not excessive.

Completeness5/5

The tool set covers validation, comparison, explanation, rendering, and fingerprinting, providing a complete workflow for observing and debugging behavior-contract compliance.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    Proof-of-behavior enforcement for AI agents. Declare behavioral constraints, enforce at runtime, produce SHA-256 hash-chained audit trails. Supports covenants (permit/forbid/require), real-time verification, and cross-agent trust handshakes.
    4
    39
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    A deterministic behavior-compatibility layer for AI agents that checks normalized event traces against operating contracts and compares baselines with candidates to catch regressions in approval, stop, scope, recovery, and completion rules.
    4
    MIT

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/christian140903-sudo/behaviorlock'

If you have feedback or need assistance with the MCP directory API, please join our Discord server