Skip to main content
Glama
m2rads
by m2rads

Limetest

Limetest is the most light weight end to end testing framework with AI capabilities that can run in your CI workflows. Define your test cases in natural language and let AI handle the execution.

Key Features

  • Optimized for AI: Define your test cases in plain language and let AI execute them end to end.

  • Lightweight & Efficient: Leverages Playwright snapshot instead of pixel analysis for faster, more reliable execution.

  • Vision Capabilities: Falls back to vision mode when snapshot mode fails during more sophisticated test scenarios.

Installation

npm install @limetest/limetest

npx playwright install

User data directory

limtest will launch Chrome browser with the new profile, located at

- `%USERPROFILE%\AppData\Local\ms-limetest\mcp-chrome-profile` on Windows
- `~/Library/Caches/ms-limetest/mcp-chrome-profile` on macOS
- `~/.cache/ms-limetest/mcp-chrome-profile` on Linux

Related MCP server: Playwright MCP

Usage

Run Tests

Use --headless for running tests headlessly in CI workflows

npx limetest example

limetest MCP Server

https://github.com/user-attachments/assets/b801f239-dc66-4b3b-bcf2-42e2a9a68721

A Model Context Protocol (MCP) server powered by Playwright that streamlinse end to end testing for your MCP client.

Use Cases

  • Automated testing planned and executed by LLMs

Example config

After cloning this repo, build and add the E2E MCP server to your MCP Client as such: Notice that you need OpenAI API key to run this MCP server in end to end mode.

npm install @limetest/mcp

npx playwright install

Then:

{
    "mcpServers": {
        "limetest": {
            "command": "npx",
            "args": [
                "npx @limetest/mcp",
                "--api-key=<your openai api key>"
            ]
        }
    }
}

All the logged in information will be stored in that profile, you can delete it between sessions if you'd like to clear the offline state.

Acknowledgements

Limetest is based on Microsoft's Playwright MCP and optimized for automated end-to-end testing as a standalone framework. This project is distributed under the Apache 2.0 License.

Available Tools

1 tool
browser_endtoendC

Run an end to end test suit in the browser

ParametersJSON Schema
NameRequiredDescriptionDefault
testDefinitionYesThe test case definition

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions 'run' which implies execution, but doesn't describe what happens during the test (e.g., browser behavior, timeouts, outputs, or side effects). This leaves significant gaps in understanding the tool's behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, straightforward sentence that efficiently conveys the core purpose. However, the typo 'suit' (likely meant 'suite') slightly detracts from clarity, preventing a perfect score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations and output schema, the description is insufficient for a tool that executes tests. It doesn't explain what 'run' entails (e.g., whether it returns results, logs, or errors), leaving the agent without key operational context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, with the single parameter 'testDefinition' documented as 'The test case definition'. The description doesn't add any meaning beyond this, such as format examples or constraints. Since the schema does the heavy lifting, the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Run an end to end test suit in the browser' clearly states the action (run) and target (end to end test suit in the browser). It's specific enough to understand the tool's function, though 'test suit' appears to be a typo (likely 'test suite'). Since there are no sibling tools, differentiation isn't needed, so it doesn't reach the highest score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives, prerequisites, or context for its application. It merely states what the tool does without indicating scenarios or constraints for its use.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

B3/5.0
Disambiguation5/5

With only one tool, there is no possibility of ambiguity or overlap between tools. The tool's purpose is clearly defined and distinct by default.

Naming Consistency5/5

A single tool inherently has perfect naming consistency, as there are no other tools to compare it against. The name 'browser_endtoend' follows a clear snake_case pattern.

Tool Count2/5

One tool is too few for a server's purpose, as it severely limits functionality and suggests an incomplete or trivial implementation. This is a borderline case leaning toward inadequacy.

Completeness1/5

The server is severely incomplete; with only a single end-to-end test tool, there are obvious gaps in functionality for any meaningful domain, such as setup, teardown, or other testing operations.

Maintenance

ActivityInactive
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    D
    maintenance
    A Model Context Protocol server that provides browser automation capabilities using Playwright, enabling LLMs to interact with web pages through structured accessibility snapshots without requiring screenshots or vision models.
    22
    5,881,527
    1
    Apache 2.0
  • A
    license
    B
    quality
    D
    maintenance
    A Model Context Protocol server that provides browser automation capabilities using Playwright, enabling LLMs to interact with web pages through structured accessibility snapshots without needing screenshots or visually-tuned models.
    22
    5,881,527
    Apache 2.0
  • A
    license
    B
    quality
    D
    maintenance
    A Model Context Protocol server that provides browser automation capabilities using Playwright, enabling LLMs to interact with web pages through structured accessibility snapshots without requiring screenshots or visually-tuned models.
    24
    5,881,527
    Apache 2.0
  • A
    license
    A
    quality
    C
    maintenance
    A Model Context Protocol server that provides browser automation capabilities using Playwright, enabling LLMs to interact with web pages through structured accessibility snapshots without requiring screenshots or visually-tuned models.
    22
    5,881,527
    Apache 2.0

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/m2rads/limetest'

If you have feedback or need assistance with the MCP directory API, please join our Discord server