Skip to main content
Glama
509,341 tools. Updated 2026-09-02 20:30

"Optimizing AI Model Thinking, Token Usage, and Context Size" matching MCP tools:

  • Find keyword mentions in AI model outputs from ChatGPT and Google AI. Returns mention context and sources.
    MIT
  • Exports the AI Context file for a given master, combining network state and CLI commands for language model planning.
    Apache 2.0
  • Retrieve detailed metadata for any AI model, including family, parameter size, quantization, and context length. Optionally specify a provider to filter results.
    Creative Commons Attribution Non Commercial No Derivatives 4.0 International
  • Execute a single AI model call to test prompts before building full workflows. Returns output, token usage, estimated provider cost, and trace URL.
    MIT
  • Execute multi-step AI workflows with reduced context usage by keeping intermediate results in the workflow engine, supporting multiple model calls and tool integrations.
    MIT
  • Generate a signed token proving human authorization for a specific AI agent action, embedding action details, context, and expiry for verification before execution.
    MIT

Matching MCP Servers

  • A
    license
    A
    quality
    B
    maintenance
    Exposes a get_context_usage tool that reports raw token usage from session transcripts for Claude Code and OpenAI Codex CLI, enabling agents to check context size and branch behavior.
    1
    20
    MIT
  • A
    license
    C
    quality
    D
    maintenance
    Enables access to Usage and Billing APIs for managing accounts, products, meters, plans, and usage reporting. Supports operations like creating products/plans, reporting usage, and retrieving billing information.
    18
    MIT

Matching MCP Connectors

  • Query x402 consumption records for any EVM wallet address. Get per-call details including model, token usage, credits, USD paid, transaction hash, and timestamp — no login needed.
    MIT
  • Execute LLM requests by burning Shells to get AI-generated responses with calculated costs based on model and token usage.
  • Record compact model token usage for value proof. Capture input, output, cached, and total tokens, cost, latency, and provider details while excluding sensitive content.
    MIT
  • Track AI token usage and estimated cost for the current MCP session, covering calls from design analysis, documentation, and composition tools.
    MIT
  • Query any AI model with a prompt and receive its response with metadata including latency and token usage. Optionally limit response tokens with automatic distillation.
    MIT
  • Generate single-turn or multi-turn responses with DeepSeek V4. Choose flash or pro models, control thinking, and persist conversation context. Provide a message or full chat history to get AI-generated replies.
    MIT
  • Compare AI model performance by testing 1-5 models simultaneously with identical prompts. Get output text, latency, token usage, and cost estimates for informed model selection.
    MIT
  • Retrieve managed AI gateway token counts, response count, and exact USD cost per model for a project. Use to report spend or identify the dominant model.
    MIT
  • Check remaining context capacity by viewing message count, token usage, and bloat indicators. Helps decide if pruning old messages is needed before continuing.
    MIT
  • Record token usage and cost per task after each AI interaction to track spending and enable budget monitoring.
    MIT