Retrieve detailed metadata for any AI model, including family, parameter size, quantization, and context length. Optionally specify a provider to filter results.
Creative Commons Attribution Non Commercial No Derivatives 4.0 International
Execute multi-step AI workflows with reduced context usage by keeping intermediate results in the workflow engine, supporting multiple model calls and tool integrations.
Generate a signed token proving human authorization for a specific AI agent action, embedding action details, context, and expiry for verification before execution.
Exposes a get_context_usage tool that reports raw token usage from session transcripts for Claude Code and OpenAI Codex CLI, enabling agents to check context size and branch behavior.
Enables access to Usage and Billing APIs for managing accounts, products, meters, plans, and usage reporting. Supports operations like creating products/plans, reporting usage, and retrieving billing information.
Query x402 consumption records for any EVM wallet address. Get per-call details including model, token usage, credits, USD paid, transaction hash, and timestamp — no login needed.
Retrieve your account's AI model limits and tier configuration, including subscription tokens, context limits, and account constraints like repository and file caps.
Record compact model token usage for value proof. Capture input, output, cached, and total tokens, cost, latency, and provider details while excluding sensitive content.
Query any AI model with a prompt and receive its response with metadata including latency and token usage. Optionally limit response tokens with automatic distillation.
Generate single-turn or multi-turn responses with DeepSeek V4. Choose flash or pro models, control thinking, and persist conversation context. Provide a message or full chat history to get AI-generated replies.
Compare AI model performance by testing 1-5 models simultaneously with identical prompts. Get output text, latency, token usage, and cost estimates for informed model selection.
Retrieve managed AI gateway token counts, response count, and exact USD cost per model for a project. Use to report spend or identify the dominant model.
Check remaining context capacity by viewing message count, token usage, and bloat indicators. Helps decide if pruning old messages is needed before continuing.