Skip to main content
Glama

Nobulex

El protocolo de prueba de comportamiento para agentes de IA autónomos.

Cada agente de IA hace promesas: "No transferiré más de $500", "Solo accederé a APIs aprobadas", "No tocaré datos de producción". Pero hoy en día, no hay forma de probar que un agente cumplió esas promesas. Los registros son escritos por el mismo software que está siendo auditado. El cumplimiento se afirma, nunca se prueba.

Nobulex cambia eso. Define reglas de comportamiento. Aplícalas antes de la ejecución. Prueba el cumplimiento con criptografía, no con confianza.

CI Tests License TypeScript

¿Qué es la Prueba de Comportamiento?

No puedes auditar una red neuronal. Pero sí puedes auditar acciones frente a compromisos establecidos.

verify(covenant, actionLog) → { compliant: boolean, violations: Violation[] }

Esto es siempre decidible, siempre determinista, siempre eficiente. Sin ML, sin heurísticas: prueba matemática.

Prueba de comportamiento significa que cada acción de un agente autónomo es:

  • Declarada — reglas de comportamiento definidas antes del despliegue en un lenguaje formal

  • Aplicada — violaciones bloqueadas en tiempo de ejecución, antes de la ejecución

  • Probada — cada acción encadenada mediante hash en un registro de auditoría a prueba de manipulaciones que terceros pueden verificar de forma independiente

Related MCP server: Agent Receipts

Inicio rápido

npm install @nobulex/sdk
import { createDID } from '@nobulex/identity';
import { parseSource } from '@nobulex/covenant-lang';
import { EnforcementMiddleware } from '@nobulex/middleware';
import { verify } from '@nobulex/verification';

// 1. Create an agent identity
const agent = await createDID();

// 2. Write behavioral rules
const spec = parseSource(`
  covenant SafeTrader {
    permit read;
    permit transfer (amount <= 500);
    forbid transfer (amount > 500);
    forbid delete;
  }
`);

// 3. Enforce at runtime
const mw = new EnforcementMiddleware({ agentDid: agent.did, spec });

// $300 transfer — allowed
await mw.execute(
  { action: 'transfer', params: { amount: 300 } },
  async () => ({ success: true }),
);

// $600 transfer — BLOCKED before execution
await mw.execute(
  { action: 'transfer', params: { amount: 600 } },
  async () => ({ success: true }),  // never runs
);

// 4. Prove compliance
const result = verify(spec, mw.getLog());
console.log(result.compliant);    // true
console.log(result.violations);   // []

Protocolo de verificación entre agentes

Antes de que dos agentes realicen una transacción, verifican la prueba de comportamiento del otro. Sin prueba, no hay transacción.

import { generateProof, verifyCounterparty } from '@nobulex/sdk';

// Agent A generates its proof-of-behavior
const proof = await generateProof({
  identity: agentA,
  covenant: spec,
  actionLog: middleware.getLog(),
});

// Agent B verifies Agent A before transacting
const result = await verifyCounterparty(proof);

if (!result.trusted) {
  console.log('Refusing transaction:', result.reason);
  return; // No proof, no transaction
}

// Safe to transact — Agent A is verified
await executeTransaction(proof.agentDid, amount);

El protocolo verifica seis cosas en orden: firma del pacto, firma de la prueba, integridad del registro, cumplimiento, historial mínimo y pacto requerido. Si alguna verificación falla, la transacción es rechazada.

Por qué es importante la Prueba de Comportamiento

Lo que existe hoy

Lo que falta

Guardrails filtran prompts y salidas

No hay prueba de que el agente siguió las reglas en la capa de acción

Monitoreo observa lo que hacen los agentes después del hecho

No hay aplicación antes de la ejecución

Identidad verifica quién es el agente

No hay verificación de lo que hizo el agente

Plataformas de gobernanza proporcionan paneles y políticas

No hay evidencia criptográfica que un tercero pueda verificar independientemente

La prueba de comportamiento llena el vacío: declarar → aplicar → probar.

El DSL de Pactos (Covenant DSL)

covenant SafeTrader {
  permit read;
  permit transfer (amount <= 500);
  forbid transfer (amount > 500);
  forbid delete;
  require counterparty.compliance_score >= 0.8;
}

Prohibir gana. Si algún forbid coincide, la acción se bloquea inmediatamente independientemente de los permisos. Denegación predeterminada para acciones no coincidentes. Las condiciones soportan >, <, >=, <=, ==, != en campos numéricos, de cadena y booleanos.

Tres palabras clave. Sin archivos de configuración. Sin YAML. Sin esquemas JSON. Solo reglas.

Arquitectura

┌─────────────────────────────────────────────────────────────┐
│                        Platform                             │
│              cli  ·  sdk  ·  mcp-server                     │
├─────────────────────────────────────────────────────────────┤
│                  Proof-of-Behavior Stack                    │
│                                                             │
│  ┌──────────┐  ┌──────────────┐  ┌────────────┐            │
│  │ identity │  │ covenant-lang│  │ action-log │            │
│  │  (DID)   │  │    (DSL)     │  │(hash-chain)│            │
│  └──────────┘  └──────────────┘  └────────────┘            │
│                                                             │
│  ┌────────────┐  ┌──────────────┐  ┌───────────────┐       │
│  │ middleware  │  │ verification │  │ composability │       │
│  │(pre-exec)  │  │ (post-hoc)   │  │(trust graph)  │       │
│  └────────────┘  └──────────────┘  └───────────────┘       │
├─────────────────────────────────────────────────────────────┤
│                      Foundation                             │
│            core-types  ·  crypto  ·  types                  │
└─────────────────────────────────────────────────────────────┘

Paquetes principales

Paquete

Qué hace

@nobulex/identity

Creación de W3C DID con claves Ed25519

@nobulex/covenant-lang

DSL inspirado en Cedar: analizador léxico, sintáctico y compilador

@nobulex/action-log

Registro a prueba de manipulaciones encadenado con hash SHA-256 con pruebas de Merkle

@nobulex/middleware

Aplicación previa a la ejecución: bloquea violaciones antes de que se ejecuten

@nobulex/verification

Verificación determinista del cumplimiento

@nobulex/sdk

API unificada que combina todas las primitivas

@nobulex/mcp-server

Servidor de cumplimiento MCP para cualquier agente compatible con MCP

@nobulex/cli

Línea de comandos: nobulex init, verify, inspect

@nobulex/langchain

Integración de middleware de LangChain (PyPI)

Integraciones

  • npmnpm install @nobulex/sdk

  • PyPIpip install langchain-nobulex

  • MCPnpx @nobulex/mcp-server (funciona con Claude Desktop, Cursor, VS Code)

  • LangChain — middleware de cumplimiento de inserción directa

  • ElizaOS — plugin para acciones, evaluadores, proveedores

Comparación conceptual

Bitcoin

Ethereum

Nobulex

Qué verifica

Transferencias monetarias

Ejecución de contratos

Comportamiento del agente

Mecanismo

Prueba de trabajo

Prueba de participación

Prueba de comportamiento

Qué se prueba

Validez de la transacción

Transiciones de estado

Cumplimiento del comportamiento

Garantía

Dinero sin confianza

Contratos sin confianza

Agentes sin confianza

Demo en vivo

npx tsx demo/covenant-demo.ts

Crea dos agentes, define reglas de comportamiento, aplica en tiempo de ejecución, bloquea una transferencia prohibida y verifica criptográficamente el cumplimiento, todo en un solo script.

Desarrollo

git clone https://github.com/arian-gogani/nobulex.git
cd nobulex
npm install
npx vitest run    # 4,237 tests, 80 files, 0 failures

Documentación

Enlaces

Licencia

MIT: úsalo para lo que quieras.

Available Tools

4 tools
check_actionA

Check whether an action is allowed or blocked by the current covenant rules.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesThe action name to check, e.g. 'delete_user'
paramsNoOptional parameters for the action

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so description carries full burden. It implies a read-only check with no side effects, but does not disclose auth needs, rate limits, or behavior for missing actions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, concise sentence with no wasted words. The purpose is front-loaded with the verb 'Check'.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with 2 parameters and no output schema, the description is mostly complete. It could benefit from mentioning the return format (e.g., boolean or status), but the core behavior is clear.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, with clear descriptions for both parameters. The description adds only an example ('e.g. delete_user'), which is marginally helpful but not necessary.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool checks whether an action is allowed or blocked by covenant rules, using specific verb 'Check' and resource 'action'. It distinguishes from siblings like set_rules and verify_log.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use or when-not-to-use guidance is provided. The purpose implies use for permission checking, but no alternatives or exclusions are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_audit_logB

Returns the full hash-chained audit trail of all compliance checks.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations and minimal description. 'Returns' implies read-only, but doesn't disclose potential size limits, authentication needs, or what 'full' means. Important behavioral traits unaddressed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single, front-loaded sentence. Efficient but could be slightly more detailed without losing conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given zero parameters and no output schema, description adequately states scope ('full...all compliance checks'). However, lacks output structure hints. Sibling tools provide context but description doesn't leverage them.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

No parameters exist (schema coverage 100%), so baseline is 4. Description adds no parameter info, but not needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states it returns the full hash-chained audit trail of compliance checks, clearly identifying the verb and resource. It distinguishes from siblings like 'check_action' and 'set_rules' which are action-oriented.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives like 'verify_log'. Lacks context on prerequisites or scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_rulesA

Set covenant rules using permit/forbid/require syntax. Each rule is a string like 'forbid delete_user' or 'permit read_data safe to read'.

ParametersJSON Schema
NameRequiredDescriptionDefault
rulesYesArray of rule strings, e.g. ['forbid delete_user', 'permit read_data']

TDQS

A3.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description bears full responsibility for behavioral disclosure. It states 'Set covenant rules' but does not indicate whether this is a destructive or reversible operation, what permissions are needed, or any side effects. This is insufficient for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the core purpose, and contains no unnecessary words. Each sentence contributes meaning: the first states what it does, the second gives format examples.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one parameter, no nested objects, no enums, no output schema), the description covers the key aspects: purpose and parameter format. It could mention whether rules are appended or replaced, but overall it is adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the parameter is described in the schema. The description adds value by providing concrete syntax examples ('forbid delete_user', 'permit read_data safe to read'), which clarify the expected format beyond the generic schema description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: 'Set covenant rules using permit/forbid/require syntax.' It specifies the verb 'Set' and the resource 'covenant rules', and provides example syntax. This distinguishes it from siblings like check_action, get_audit_log, and verify_log, which have different purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides syntax examples but lacks explicit guidance on when to use this tool versus alternatives (e.g., check_action, get_audit_log). It does not mention prerequisites or when not to use it. The usage context is implied but not clarified.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_logA

Independently verify the integrity of the hash-chained audit log. Detects any tampering.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must fully convey behavioral traits. It only says 'Detects any tampering' but does not disclose the tool's return value (e.g., boolean), side effects (if any), or whether it checks against a remote source. This lack of detail forces the agent to guess the behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that directly states the purpose. It is front-loaded and contains no fluff. Every word contributes to understanding the tool's function.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no parameters and no output schema, the description is adequate but not fully complete. It lacks details about the return value (e.g., does it return a boolean, raise an exception, or log results?). The agent needs more context to know how to handle the tool's output.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are no parameters, so the input schema is fully covered. Baseline for 0 parameters is 4, and the description adds no param info (none needed). The agent can invoke the tool without any parameter confusion.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it verifies the integrity of the hash-chained audit log and detects tampering. The verb 'verify' and resource 'audit log' are specific, and the tool is easily distinguished from siblings like 'check_action' or 'get_audit_log'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives. It does not mention prerequisites, exclusions, or typical use scenarios. Without explicit instructions, an agent may not know when verification is needed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 4 tool updatesv1.0.0
    • First observedcheck_action
    • First observedget_audit_log
    • First observedset_rules
    • First observedverify_log

TDQS

A3.8/5.0
Disambiguation5/5

Each tool has a clearly distinct purpose: checking actions, retrieving audit logs, setting rules, and verifying log integrity. No overlap or ambiguity.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern in snake_case (check_action, get_audit_log, set_rules, verify_log).

Tool Count5/5

Four tools is an appropriate scope for a compliance/auditing server; each tool serves a necessary function without redundancy.

Completeness4/5

Covers core operations: rule setting, action checking, audit log retrieval, and log integrity verification. Minor gap: no explicit rule deletion or modification beyond full replacement, but this is acceptable for the domain.

Maintenance

ActivitySlowing
ResponsivenessWithin a week

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    Provides covenant rule enforcement, hash-chained audit logs, and integrity verification for MCP-compatible agents. It enables users to define granular permission rules and maintain a tamper-evident audit trail of all actions.
    4
    14
    1
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    AI agent provenance, trust, and auditability layer. VERITAS multi-gate scoring, Cortex approval gates, S.E.A.L. hash-chain audit ledger, and semantic RAG with cryptographic provenance tracking for every decision an agent makes.
    27
    5
    MIT

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/arian-gogani/nobulex'

If you have feedback or need assistance with the MCP directory API, please join our Discord server