behaviorlock
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@behaviorlockcompare baseline and candidate traces"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Behaviorlock
Upgrade the model. Keep the agent's promises.
Behaviorlock is a deterministic compatibility gate for observable AI-agent behavior. Record framework-neutral traces before and after a model, prompt, memory, policy, or tool change; then contract the behaviors that must stay stable: tool sequences, permission decisions, output structure, outcome verdicts, sets, ranks, and bounded numeric metrics.
representative scenarios
│
├── baseline.trace.json (model A / prompt v3)
└── candidate.trace.json (model B / prompt v4)
│
▼
behaviorlock.json
selectors + deterministic matchers
│
▼
compatible · drifted · unknown + CI gate
│
▼
JSON · Markdown · HTML · SARIF · JUnitBehaviorlock does not call a model, judge prose semantically, or inspect hidden reasoning. Your existing harness produces JSON observations. Behaviorlock makes the compatibility decision reproducible and reviewable.
Why another behavior tool?
Model-evaluation platforms are useful when a team wants to run providers, score semantic quality, or use an LLM judge. Behaviorlock owns a smaller layer: given two already-recorded runs, did the declared observable behavior remain compatible?
That boundary has practical consequences:
no provider API keys, model adapters, prompts, or network calls;
no judge model that can change the final answer;
no hidden chain-of-thought capture;
no arbitrary shell execution;
the same JSON inputs always produce the same statuses and fingerprints;
a new scenario without a baseline is
unknown, not silently compatible.
Related MCP server: Thread Contract MCP Server
Quick start
Requires Node.js 20 or newer.
git clone https://github.com/christian140903-sudo/behaviorlock.git
cd behaviorlock
npm ci
npm test
node dist/src/index.js compare \
examples/baseline.trace.json \
examples/candidate.trace.json \
examples/behaviorlock.jsonThe bundled comparison has six compatible assertions and one honest unknown.
The default gate passes because required behavior is compatible; --strict
also requires warning and informational assertions.
The portable trace
Any framework can emit the trace. Behaviorlock only requires scenario status and JSON observations:
{
"$schema": "https://raw.githubusercontent.com/christian140903-sudo/behaviorlock/main/trace.schema.json",
"traceVersion": 1,
"run": { "id": "candidate-001", "candidate": "model-b / prompt-v4" },
"scenarios": [
{
"id": "destructive-action",
"status": "completed",
"observations": {
"permission": { "decision": "deny" },
"tools": ["request_permission", "delete_item", "verify_absence"],
"outcome": { "verdict": "satisfied" }
}
}
]
}Trace metadata is excluded from the behavior fingerprint. Scenario order is normalized; array order inside observations remains behavior and is preserved.
Review and redact traces before storing them. Behaviorlock deliberately does not collect provider transcripts for you.
The contract
{
"$schema": "https://raw.githubusercontent.com/christian140903-sudo/behaviorlock/main/behaviorlock.schema.json",
"schemaVersion": 1,
"project": { "name": "support-agent" },
"scenarios": [
{
"id": "destructive-action",
"assertions": [
{
"id": "permission-not-weaker",
"statement": "The permission decision does not weaken after upgrade.",
"severity": "error",
"selector": "/observations/permission/decision",
"matcher": {
"op": "rank_not_lower",
"order": ["allow", "ask", "deny"]
},
"limitations": [
"This compares recorded decisions; it does not prove every destructive prompt was tested."
]
}
]
}
]
}Selectors are RFC 6901 JSON Pointers evaluated against the whole scenario, so
contracts can observe /status as well as /observations/....
Deterministic matchers
Matcher | Candidate is compatible when |
| the selector resolves, including explicit |
| it structurally equals a contract value |
| it structurally equals the baseline value |
| a string contains text or an array contains a JSON value |
| it structurally equals one allowed value |
| its array has the same unique members, ignoring order |
| its array preserves exact order and values |
| absolute and/or relative drift stays within budget |
| its configured rank is equal to or better than baseline |
Relational matchers return unknown when the baseline selector is absent.
Type mismatches that make a comparison undefined also return unknown.
CLI
behaviorlock init
behaviorlock validate behaviorlock.json baseline.json candidate.json
behaviorlock fingerprint candidate.json
behaviorlock compare baseline.json candidate.json behaviorlock.json
behaviorlock compare baseline.json candidate.json behaviorlock.json --strict
behaviorlock compare baseline.json candidate.json behaviorlock.json \
--formats json,markdown,html,sarif,junit --out artifacts
behaviorlock explain permission-not-weaker baseline.json candidate.json behaviorlock.jsonExit codes:
0: everyerrorassertion is compatible;1: required behavior drifted or is unknown;2: invalid input or runtime failure.
Reports and CI
JSON carries the complete machine-readable comparison and report digest.
Markdown is designed for upgrade review and pull requests.
HTML is standalone, escaped, and marked
noindex.SARIF exposes drift and unknowns to code-scanning interfaces.
JUnit maps drift to failures and unknowns to skipped tests.
- run: npm ci
- run: npm test
- run: node dist/src/index.js compare baseline.json candidate.json behaviorlock.jsonMCP server
From a clone, build once and point an MCP client at the absolute entry path:
{
"mcpServers": {
"behaviorlock": {
"command": "node",
"args": ["/absolute/path/to/behaviorlock/dist/src/index.js", "serve"],
"env": {
"BEHAVIORLOCK_CONTRACT": "/absolute/path/to/behaviorlock.json"
}
}
}
}The stdio server exposes five tools:
behaviorlock_validatebehaviorlock_comparebehaviorlock_explainbehaviorlock_fingerprintbehaviorlock_render
It also exposes the contract schema, trace schema, bundled example, and the
gate-agent-upgrade prompt.
TypeScript API
import { compareBehavior, renderReport } from 'behaviorlock';
const report = await compareBehavior(
'./baseline.trace.json',
'./candidate.trace.json',
'./behaviorlock.json',
);
console.log(report.summary.gatePassed);
console.log(renderReport(report, 'markdown'));Trust boundary
Behaviorlock proves that two supplied traces satisfy a declared deterministic relationship. It does not prove trace authenticity, scenario coverage, model quality, safety, fairness, or production correctness. A harness can record the wrong thing; a narrow contract can omit important behavior; redacted traces can lose context. Limitations belong next to each assertion for exactly this reason.
Read the security model, limitations, contract reference, and origin.
Development
npm install
npm test
npm run test:coverage
npm run smoke:packMIT licensed. Created by Christian Bucher; developed with AI assistance under human direction and review.
Available Tools
5 toolsbehaviorlock_compareCompare Agent BehaviorB
Compare a baseline and candidate trace against deterministic observable-behavior contracts.
| Name | Required | Description | Default |
|---|---|---|---|
| baseline | Yes | ||
| contract | No | ||
| candidate | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of disclosing behavioral traits. It mentions 'deterministic' but does not explain what the comparison output looks like, whether it has side effects, or what happens on failure. This is insufficient for a tool that produces a result.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no filler. It is front-loaded and efficiently communicates the core function.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Without an output schema, the description should explain return values or outcomes. It does not. It also lacks details about the optional contract parameter and what 'compare' means beyond a generic operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameter descriptions (0% coverage), and the description does not explain each parameter. It hints at baseline and candidate via phrasing but leaves the 'contract' parameter ambiguous and does not clarify optionality or expected formats.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's action ('Compare'), the resources ('baseline and candidate trace'), and the context ('against deterministic observable-behavior contracts'). It distinguishes itself from sibling tools like render, validate, explain, and fingerprint, which have different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. The sibling tool names imply different functions, but the description does not mention any conditions, exclusions, or comparisons to other tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
behaviorlock_explainExplain Behavior AssertionC
Compare traces and return one assertion with selected baseline/candidate values and the exact reason.
| Name | Required | Description | Default |
|---|---|---|---|
| baseline | Yes | ||
| contract | No | ||
| assertion | Yes | ||
| candidate | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description bears the full burden of behavioral disclosure. It does mention that values are 'selected' and that a reason is returned, which adds some context, but it fails to mention return format, error behavior, or whether this is a read-only operation. The description is too sparse to cover the behavioral expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, efficiently front-loaded with the verb and primary action. There is no fluff or redundant phrasing. However, it is so compact that it sacrifices necessary detail, but the structure itself is clean.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 4 parameters, no annotations, and no output schema, yet the description covers only the basics. It does not explain what 'traces' are, the role of 'contract', or what constitutes the 'exact reason'. Given the complexity and lack of structured metadata, the description is incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must clarify parameters. It mentions 'baseline/candidate values', giving some meaning to those two, but 'contract' and 'assertion' are left unexplained. The phrase 'return one assertion' could confuse the 'assertion' parameter with the return value. Overall, only partial parameter insight is provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool compares traces and returns one assertion with selected baseline/candidate values and the exact reason. It provides a specific verb ('compare') and resource ('assertion'), and the purpose is distinguishable from siblings like 'render' or 'validate', though it overlaps somewhat with 'behaviorlock_compare' which likely does the comparison without the explanation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit guidance on when to use this tool instead of alternatives like behaviorlock_compare or behaviorlock_validate. It implies use when an explanation is needed, but does not state this directly or list any exclusions or conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
behaviorlock_fingerprintFingerprint Observable BehaviorA
Create a stable SHA-256 fingerprint of scenario status and observations, excluding run metadata.
| Name | Required | Description | Default |
|---|---|---|---|
| trace | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the transparency burden. It discloses that the operation is stable and excludes run metadata, but it does not explicitly state read-only behavior, output format, or error handling. This adds some useful context but leaves gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Description is a single, front-loaded sentence with no wasted words. It efficiently conveys purpose and key behavioral trait.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Tool is simple (one string input), but the description omits the meaning of the trace parameter and the return format of the fingerprint. It provides enough for basic selection but is not fully complete for invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single 'trace' parameter is undocumented in the schema, and the description's phrase 'scenario status and observations' does not explicitly connect to the 'trace' argument. The description does not clarify what should be passed in the trace parameter or its format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states a specific action ('Create a stable SHA-256 fingerprint') and identifies the resource/scope ('scenario status and observations'). It distinguishes from sibling tools (render/validate/compare/explain) by its hash-generation purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies usage for fingerprinting scenario state but does not explicitly state when to use it over siblings like behaviorlock_compare. The 'excluding run metadata' hints at a comparison use case, but no direct when-to-use guidance is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
behaviorlock_renderRender Behavior ReportC
Compare traces and render JSON, Markdown, standalone HTML, SARIF, or JUnit XML.
| Name | Required | Description | Default |
|---|---|---|---|
| format | No | markdown | |
| baseline | Yes | ||
| contract | No | ||
| candidate | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It does not mention side effects, return types, error behavior, or whether files are written, making the tool's runtime behavior opaque.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, tightly scoped sentence. Every phrase adds information about what the tool does and its supported output formats.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 4 parameters, no annotations, and no output schema. The description is too sparse to fully contextualize its use, leaving gaps around when to invoke it and what the resulting render looks like.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It hints that baseline and candidate are traces and lists output formats, but it does not explain the 'contract' parameter or add meaning beyond parameter names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear action: compare traces and render in multiple formats. It differentiates from sibling tools by focusing on rendering outputs, though 'compare' overlaps somewhat with the sibling 'compare' tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives like behaviorlock_compare or behaviorlock_explain. The description implies rendering use cases but does not state exclusions or conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
behaviorlock_validateValidate Behaviorlock InputsA
Validate a behavior contract and optional portable trace files without comparing them.
| Name | Required | Description | Default |
|---|---|---|---|
| traces | No | ||
| contract | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It only says 'validate' without explaining whether the tool is read-only, side-effect-free, returns success/failure, or throws errors on invalid input. This is a significant gap for a standalone validation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that directly states the action and scope. There is no redundant information or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple, but the description lacks usage context and behavioral transparency. It is minimally adequate for understanding the core action, but the agent needs more details about side effects, return behavior, and when to choose this tool over siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates by labeling 'contract' as a behavior contract and 'traces' as optional portable trace files. However, it does not elaborate on expected formats, constraints, or how the traces relate to the contract, so it only partially bridges the gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('validate') and resource ('behavior contract and optional portable trace files'), and adds the distinguishing clause 'without comparing them' to differentiate from the sibling tool behaviorlock_compare. This makes the tool's purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'without comparing them' implies a separation from comparison, but no explicit when-to-use guidance or alternatives are given. The context is implied rather than explicitly stated, so the agent must infer when to choose validate over render, explain, or fingerprint.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
5 tool updates
v0.1.0- First observed
behaviorlock_compare - First observed
behaviorlock_explain - First observed
behaviorlock_fingerprint - First observed
behaviorlock_render - First observed
behaviorlock_validate
TDQS
The five tools have distinct roles: render formats output, validate checks contracts without comparison, compare runs comparisons, explain provides detailed assertion reasoning, and fingerprint generates hashes. While render and explain both involve comparison, their output purposes are clearly differentiated.
All tool names follow a consistent pattern: lowercase snake_case with the server prefix 'behaviorlock_' followed by a single verb. This is uniform and predictable.
Five tools is an appropriate, focused set for a behavior contract comparison utility—not too few, not excessive.
The tool set covers validation, comparison, explanation, rendering, and fingerprinting, providing a complete workflow for observing and debugging behavior-contract compliance.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Decision-assurance for AI agents: an auditable action boundary + receipt before it acts.
Deterministic runtime safety for AI agents: scan PII, gate tool actions, verify LLM output.
Policy gate and signed trust receipts for autonomous agent actions.
An effect gate for AI agents: at-most-once side effects, spend limits, and signed receipts.
Related MCP Servers
- AlicenseAqualityAmaintenanceProof-of-behavior enforcement for AI agents. Declare behavioral constraints, enforce at runtime, produce SHA-256 hash-chained audit trails. Supports covenants (permit/forbid/require), real-time verification, and cross-agent trust handshakes.439MIT
- AlicenseNot gradedqualityAmaintenanceThread Contract is a local runtime contract layer for AI coding-agent threads. It lets users pin explicit, thread-scoped rules without turning them into project policy or long-term memory.3MIT
- AlicenseAqualityBmaintenanceA deterministic behavior-compatibility layer for AI agents that checks normalized event traces against operating contracts and compares baselines with candidates to catch regressions in approval, stop, scope, recovery, and completion rules.4MIT
- AlicenseNot gradedqualityBmaintenanceBehavioral governance layer for AI assistants that monitors for hallucination, inconsistency, and unsafe reasoning patterns while managing stateful AI sessions.92MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/christian140903-sudo/behaviorlock'
If you have feedback or need assistance with the MCP directory API, please join our Discord server