Skip to main content
Glama

wasmagent-js

npm version License: Apache-2.0 CI Quick Start Docs

WasmAgent adds a verifiable evidence layer to agent tool use: protect tool calls, record what happened, audit the result, and admit trusted traces into downstream systems.

Protect → Record → Audit → Admit · Sync — agent↔UI shared state


Start in 30 seconds

Pick your entry point:

Goal

Install

Protect tools — runtime firewall, policy enforcement, taint tracking

npm add @wasmagent/mcp-firewall

Record evidence — signed AEP records after every agent run

npm add @wasmagent/aep

Admit from traces — compliance scoring produces ComplianceEvalRecords for downstream training

npm add @wasmagent/aep @wasmagent/compliance

Sync state — reducer-backed agent↔UI shared state, agent reads projections + writes intent

npm add @wasmagent/core (/shared-state subpath)

Trust Pack — 30-minute end-to-end: docs/quickstarts/trust-pack-30min.md


Related MCP server: Code Executor MCP Server

Quickstart

Three paths — pick the one that fits your use case:

Path 1 — Protect: MCP runtime firewall

Wrap any MCP server: vet tools before execution, enforce policy per call, track taint across results.

npm install @wasmagent/mcp-firewall
import { evaluatePolicy, snapshotTool, taintObservation, vetTool } from "@wasmagent/mcp-firewall";

const entry = {
  name: "read_file",
  description: "Read a file from disk",
  inputSchema: { type: "object", properties: { path: { type: "string" } } },
};
const args = { path: "/tmp/report.txt" };
const consentRecords = [];

// Before calling a tool
const snap     = snapshotTool(entry, "my-server");   // hash descriptor at registration
const vetting  = vetTool(entry);                     // static scan: injection / exfil / rug-pull
const decision = evaluatePolicy(entry.name, args, vetting, consentRecords);

if (decision.decision === "deny")   throw new Error(`Blocked: ${decision.reasons.join("; ")}`);
if (decision.decision === "ask_user") {
  // surface consent UI, then call recordConsent(...)
}

// After receiving result
const rawResult = "example report contents";
const obs = taintObservation(entry.name, rawResult);  // boundary-tagged, safe to assemble into prompt

Security pack · OWASP Agentic Top 10 · Attack demos

Path 2 — Record: AEP evidence export

Emit a signed evidence record after every agent run — consumable by trace-pipeline for audit and training.

npm install @wasmagent/aep
import { AEPEmitter } from "@wasmagent/aep";

const emitter = new AEPEmitter({ run_id: "run-001", model_id: "claude-sonnet-4-6" });

// During the run — add tool call evidence
emitter.addAction({ tool_name: "bash", outcome: "pass", exit_code: 0 });

// At the end — emit the record
const record = emitter.build();
// record satisfies aep/v0.1 JSON Schema — ready for evomerge validate-aep

AEP schema · trace-pipeline 10-min tutorial

Path 3 — Execute: Sandboxed code execution

Run agent-generated code in an isolated WASM kernel — no host-process access.

npm install @wasmagent/aisdk @wasmagent/kernel-quickjs
import { sandboxedJsTool } from "@wasmagent/aisdk";
import { QuickJSKernel } from "@wasmagent/kernel-quickjs";

// Drop into any AI SDK / LangChain / OpenAI Agents setup
const codeTool = sandboxedJsTool({ kernel: new QuickJSKernel() });

Kernel comparison · Getting started

Path 4 — Sync: Human-agent shared state

Reducer-backed collaborative state where the LLM reads projections, dispatches semantic actions, and respects affordances — all through standard tools.

npm install @wasmagent/core
import { defineStateModel, SharedStateStore, stateTools } from "@wasmagent/core/shared-state";

// 1. One reducer, shared by both UI and agent.
const model = defineStateModel({
  initial: () => ({ page: "list", selectedId: null as string | null }),
  reduce: (s, a) => {
    if (a.type === "SELECT") return { ...s, page: "detail", selectedId: a.id };
    if (a.type === "BACK")   return { ...s, page: "list", selectedId: null };
    return s;
  },
  project: (s) => ({ page: s.page, selectedId: s.selectedId }),
  affordances: (s) => s.page === "list" ? ["SELECT"] : ["BACK"],
});

// 2. Server-side store keyed by session.
const store = new SharedStateStore(model);

// 3. Give the agent read_state + dispatch_action tools.
const tools = stateTools(store, "session-001");
// Pass `tools` to any ToolCallingAgent — the LLM reads state and dispatches intent.

The semantic action stream doubles as AEP evidence — every dispatch is a provenance-ready record (see #141 for the full confluence design).


📚 Docs · Getting started · Kernels · OWASP governance · Security pack · Changelog


What is shipped vs alpha

WasmAgent uses a five-tier maturity scale to prevent "shipped" from becoming a vague claim:

Tier

Meaning

Semver guarantee

Production use

stable

Public API locked; breaking changes require major-version bump

Yes

Yes

beta

Functional and used in production, but a specific limitation is documented (e.g. first-line filter only, contract still evolving)

Minor/patch only

Yes, with caveats documented

alpha

Schema versioned; fields may be added without a breaking-change bump

No

Informed use

demo

Demonstration or example code; not hardened for production

No

No

research

Research-grade prototype; interfaces may change without notice

No

No

Packages not listed here (model adapters, UI cards, etc.) follow the same scale — see each package's README or package.json wasmagent.stability field.


Package maturity

Package

Maturity

Notes

@wasmagent/core

stable

Public API; semver guaranteed

@wasmagent/kernel-quickjs

stable

@wasmagent/kernel-remote

stable

@wasmagent/mcp-gateway

stable

Published 0.1.0; gateway composes all firewall layers

@wasmagent/mcp-firewall

beta

First-line filter, not adversarial-grade — keyword bag + lightweight n-gram classifier; use defence-in-depth

@wasmagent/aep

beta

v0.2 signature contract (Ed25519) shipped; schema versioned

@wasmagent/otel-exporter

alpha

GENAI_SEMCONV, AEP↔OTel bridge

@wasmagent/aisdk / @wasmagent/mastra-sandbox

alpha

API stable, may add fields

@wasmagent/compliance

alpha

Schema versioned; may add fields without breaking

@wasmagent/mcp-policy

alpha — private

Not yet published to npm

@wasmagent/mcp-attestation

alpha — private

Not yet published to npm

@wasmagent/evals-runner

alpha

@wasmagent/devtools

alpha


WasmAgent Ecosystem

WasmAgent is a portable, governable agent runtime for safe code execution, verifiable rollouts, and post-training data loops.

Repo

Role

wasmagent-js (this repo)

Embedded Agent Runtime / WASM Kernel / policy / verifier / adapters

bscode

Cloudflare flagship demo and deploy template for safe coding agents

trace-pipeline

Public datafactory and eval-trust backend for rollout data

Task → Safe Runtime → Verifiable Rollout → Trajectory Export → DPO/PPO Data → Better Models

What makes wasmagent different

Three wedges where wasmagent stands apart from generic agent frameworks:

Wedge

What it means

Sandboxed execution

Three isolation tiers — VmKernel / WASM (QuickJS·Pyodide·Wasmtime) / microVM — with a single CapabilityManifest and MCP runtime firewall across all

Runtime compliance

TaskSpecConstraintIRComplianceEvalRecord — every run produces an auditable, cross-repo training contract, not just a log

Trace-to-training contract

Verifiable rollout branching, objective scoring, DPO/PPO export — the loop from runtime evidence to training data is first-class, not an afterthought

#

Axis

Status

1

Multi-provider adapters — one Model interface across Anthropic, OpenAI, Doubao, DeepSeek, Kimi, Qwen, GLM, MiniMax, local llama.cpp

shipped

2

Three isolation tiersVmKernel (in-process) / QuickJS·Pyodide·Wasmtime (WASM) / RemoteSandboxKernel (microVM) — same CapabilityManifest across all

shipped

3

Cross-runtime + offline — Node / edge / browser / air-gapped laptop; @wasmagent/model-local + WASM kernel = zero outbound traffic

shipped

4

Memory layersMemoryBlockSet (prompt-cache stable) + observational memory + Checkpointer + 4 KV backends

shipped

5

Durable workflowsLocalWorkflowEngine + CloudflareWorkflowEngine — observable, terminable, resumable

shipped

6

Code-mode MCP — N tools → 2 tools (docs_search + execute_code); 13.6% token cost at N=30

shipped

7

Devtools + OTel — local Studio, gen_ai.* semantic conventions (Datadog / Honeycomb / Grafana)

shipped

8

Goal-directed loop — agent synthesises success criteria, verifies, retries with hints

shipped 2026-06-18

9

Adaptive execution — registered fallbacks (L1) → synthesised tool (L2) → relaxed goal (L3)

shipped 2026-06-18

10

MCP runtime firewall@wasmagent/mcp-firewall: descriptor snapshot, static vetting (injection / exfiltration / rug-pull / taint), per-call policy, consent ledger

shipped 2026-06-25

Full comparison with Vercel AI SDK, LangGraph.js, OpenAI Agents JS, Mastra, CF Agents SDK: docs/compare.md


Quick Start

Tool-Calling Agent

import { ToolCallingAgent, AnthropicModel } from "@wasmagent/core";
import { z } from "zod";

const agent = new ToolCallingAgent({
  model: new AnthropicModel("claude-haiku-4-5-20251001"),
  tools: [{
    name: "search", description: "Search the web",
    inputSchema: z.object({ query: z.string() }),
    readOnly: true, idempotent: true,
    forward: async ({ query }) => `Results for: ${query}`,
  }],
  stopPolicies: ["steps:10", "cost:0.5"],
});

for await (const ev of agent.run("Search for recent AI news")) {
  if (ev.event === "final_answer") console.log(ev.data.answer);
}

Sandboxed Code Agent

import { CodeAgent, AnthropicModel } from "@wasmagent/core";

const agent = new CodeAgent({
  model: new AnthropicModel("claude-sonnet-4-6"),
  tools: [],  // kernel executes code; no extra tools needed
  maxSteps: 10,
});

for await (const ev of agent.run("What is 42 * 1337?")) {
  if (ev.event === "final_answer") console.log(ev.data.answer);
}

CLI

npm install -g @wasmagent/cli

# Agent runs
wasmagent run "What is the square root of 144?"
wasmagent run "Summarise AI news" --stream | jq .

# Rollout / training data
wasmagent rank-rollout rollouts.jsonl --out ranked.jsonl
wasmagent validate-rollouts ranked.jsonl
wasmagent export-rollouts --in ranked.jsonl --format dpo --out dpo.jsonl

# MCP security (scan → guard → evidence)
wasmagent init --guard               # generate wasmagent.policy.yaml
wasmagent scan-mcp tools.json        # static risk scan, exits 1 on critical findings
wasmagent guard --config wasmagent.policy.yaml --upstream tools.json
wasmagent evidence export --input aep-records.jsonl --format json

GitHub Action — enforce policy in CI:

- uses: WasmAgent/wasmagent-js/.github/actions/agent-evidence-gate@main
  with:
    policy: wasmagent.policy.yaml
    tools-file: mcp-tools.json
    fail-on-policy-violation: "true"

MCP Guard guide · Attack demos


Key Capabilities

Capability

Guide

Shared state — reducer-backed agent↔UI sync, projections, affordances

packages/core/src/shared-state/

MCP firewall — vetTool, ScopeLease, ApprovalReceipt

docs/guides/mcp-guard.md

AEP v0.2 evidence — causal chain, scope lease, taint, memory refs

packages/aep/src/types.ts

OWASP MCP Top 10 crosswalk

docs/security/standards-crosswalk.yaml

OWASP security demo (10 scenarios)

examples/owasp-demo/

Security benchmark runner

examples/security-benchmark/

AEP ↔ OTel bidirectional mapping

packages/otel-exporter/src/aep-otel-bridge.ts

AgentTeam delegation chain

packages/core/src/agents/AgentTeam.ts

Claim dashboard

node scripts/verify-claims.mjs --htmldocs/claims/claims.html

Quality runners (self-consistency, reflect-refine, parallel fork-join)

docs/guides/quality-runners.md

Durable runtime (checkpoints, SSE resume, HITL)

docs/guides/durable-runtime.md

Observational memory — ~22% tokens on 50-turn traces

docs/guides/observational-memory.md

Goal-directed agent with verifiers

docs/guides/goal-directed.md

Production APIs (retry, evals, OTel, React hook)

docs/api/production-apis.md

API stability policy

docs/api/stability-policy.md


Model Providers

First-class adapters: Anthropic · OpenAI · Doubao · DeepSeek · Kimi · Qwen · GLM · MiniMax · local llama.cpp

// Chinese providers with thinking support
import { DoubaoModel, DoubaoModels } from "@wasmagent/model-doubao";
import { DeepSeekModel, DeepSeekModels } from "@wasmagent/model-deepseek";

// Local / offline
import { LocalModel } from "@wasmagent/model-local";  // node-llama-cpp, multi-mirror download

Full provider reference and proxy/custom endpoint setup: docs/guides/openai-compat-recipes.md


Ecosystem

Project

Role

bscode

Flagship Cloudflare deploy template — wires every wasmagent-js capability into a real edge product

trace-pipeline

Training data factory — converts ranked rollouts into DPO/PPO datasets

Upstream integration status (PRs filed to Vercel AI SDK, Mastra, LangChain.js, ElizaOS, MCP registry, …) is tracked in docs/distribution/upstream-prs.md.


Development

bun install && bun run build
bun test packages/
bun run typecheck
bun run bench          # reproduce all README benchmarks
bun run check:branding # CI guard: no old brand references
bun run verify:claims  # CI guard: all benchmark claims have evidence scripts

See CONTRIBUTING.md · Changelog · License: Apache-2.0

Available Tools

2 tools
execute_codeA

Run a JavaScript snippet inside a sandboxed kernel. The snippet may call callTool(name, args) to invoke any downstream tool. Only the snippet's final return value (or the value assigned to __finalAnswer__) is returned to you — intermediate tool outputs stay in the sandbox. Use this to chain many tool calls in one round.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYesJavaScript source. May use top-level `await`. Must end with a return value or set `__finalAnswer__ = ...`.

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must fully cover behavioral traits. It mentions sandboxing, the ability to call downstream tools via callTool, and that only the final return value is returned. However, it lacks details on error handling, timeouts, or limitations on callTool, which are important for a code execution tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded with the main action. Every sentence adds necessary information without waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of code execution and chaining, the description covers the core functionality: sandboxed execution, chaining via callTool, and final return value. Minor gaps remain in error behavior and environment specifics, but for an agent it is mostly complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 100% coverage with a description for the 'code' parameter. The tool's description adds value by explaining that the snippet may call `callTool(name, args)`, which is not in the schema description. It also reinforces the return value requirement.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool runs a JavaScript snippet in a sandboxed kernel, distinguishing it from the only sibling tool 'docs_search' which is for documentation search. It uses specific verbs and resources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Use this to chain many tool calls in one round,' providing clear context for when to use. It does not exclude alternatives, but with only one sibling, the distinction is implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

A4.4/5.0
Disambiguation5/5

The two tools have completely distinct purposes: docs_search for discovering downstream tools, and execute_code for running code that calls them. There is no overlap or confusion.

Naming Consistency5/5

Both tool names follow a consistent verb_noun pattern with snake_case: docs_search and execute_code.

Tool Count5/5

Two tools is exactly right for this server's purpose as a meta-agent: search to discover, execute to run. It's minimal but complete.

Completeness5/5

The tool surface fully covers the intended workflow: discovering available downstream tools and executing code that chains them. There are no obvious gaps.

Maintenance

ActivityActive
ResponsivenessWithin a week

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    Not graded
    maintenance
    A lightweight and fast MCP server that enables AI agents to efficiently discover and execute tools through progressive disclosure, minimizing context consumption while supporting safe code execution in external environments.
    12
  • A
    license
    Not graded
    quality
    D
    maintenance
    Universal MCP server for executing TypeScript and Python code with progressive disclosure, reducing token usage by 98% by enabling on-demand access to all other MCP tools through code execution rather than loading tool definitions directly.
    22
    130
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    A self-hosted MCP server that provides a single execute_code tool, enabling agents to write TypeScript to call multiple REST APIs via fetch() with transparent credential injection, reducing token usage by keeping intermediate results in the sandbox.
    11
    BSD 3-Clause
  • A
    license
    B
    quality
    A
    maintenance
    Agent-optimized MCP server that replaces built-in file, search, exec, and git tools with compact, structured JSON equivalents. Benchmarked 20–45% token savings for AI coding agents.
    20
    2
    MIT

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/WasmAgent/wasmagent-js'

If you have feedback or need assistance with the MCP directory API, please join our Discord server