Pebbler MCP
OfficialClick on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Pebbler MCPCompare these two images and tell me the quote before buying"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Pebbler MCP
Let your agent collect human feedback on two images and explain the results. Pebbler runs image preference tests with participants on the Pebbler app.
Your agent retrieves current pricing, package details and availability before creating a test. The purchase uses a live quote and stays within the spending limits you configure. Pricing and package sizes are not fixed by this connector.
Version 1 supports image A/B preference tests. These measure which image people prefer, rather than product conversion or causal effects. Completed response counts are not guaranteed. Surveys and text-only comparisons are not supported.
Connect your agent
Use an MCP app that supports local servers and Node.js 22 or later. Add this configuration to your app:
{
"mcpServers": {
"pebbler": {
"command": "npx",
"args": ["-y", "@pebbler/pebbler-mcp@0.1.0"]
}
}
}Pebbler's production settings are included. You can explore current packages and create drafts without enabling payments.
Related MCP server: GPT Image MCP Server
Enable purchases
Use a dedicated wallet funded with native USDC on Base. Configure these settings locally in your MCP app's environment or secret settings:
Setting | What to provide |
| Set to |
| Your wallet's signing key, supplied privately |
| Your chosen maximum spend per test, as a positive decimal USDC amount |
| Your chosen cumulative spending limit, as a positive decimal USDC amount |
Both spending limits are required when payments are enabled. They are your budget, not the price of a test. If a quote exceeds either limit, the connector stops instead of raising your budget. Payments stay disabled until you enable them explicitly.
The wallet key remains local. Do not put it in a prompt or share it in a conversation. Prices and quotes are fetched by the agent; you do not need to configure them.
Ask for a test
For example:
Check the current package and price, then compare these two image URLs. Tell me the quote before buying and stay within my configured budget.
Provide two publicly accessible HTTPS image URLs and a clear question. Use images you own or have permission to use, and keep the URLs available while responses are collected.
Your agent can show the current offer, create a draft, purchase the test and check progress. You can read partial results while collection continues or return later to ask for the final counts and vote shares. Retrieving results is included in the purchase.
Available tools
Tool | Purpose |
| Retrieve current packages, prices, supported inputs and availability |
| Prepare an image comparison and obtain its quote |
| Refresh a quote before payment has been signed |
| Purchase the quoted test within your wallet's spending limits |
| Check progress and when to check again |
| Retrieve aggregate A/B votes and shares |
| Find your saved tests after returning or restarting |
The agent learns each tool's inputs automatically when connected. It should read the catalog first, use the quote for the purchase, and follow the returned polling interval when checking progress.
Returning to your results
Keep the connector's private local state folder backed up. It retains access to your tests and their purchase history across restarts. You can choose a folder with PEBBLER_STATE_DIR; use the same folder when returning to existing tests.
Your total spending limit counts retained signed purchases, including unresolved attempts, and does not reset automatically. Increase it deliberately if you want to buy more tests. Keep the original state rather than deleting it to reset your budget.
If a purchase is interrupted or pending, ask the agent to retry that same purchase. The connector preserves its identifiers and payment authorization instead of signing another payment. Reading progress and results does not require another purchase.
Try an example
The included examples/run-study.mjs demonstrates discovery and creating an image comparison. It prints the live catalog and quote; purchasing requires the --purchase flag and your locally configured wallet and budgets.
License
MIT. Original example image assets are included under the same license.
Available Tools
7 toolscreate_studyAIdempotent
Validate an immutable preference comparison with two HTTPS image options. Version 1 supports A vs B image tests only. Free. Reuse request_key to recover a failed creation. Returns the study ID and quote; access credentials stay in the local adapter.
| Name | Required | Description | Default |
|---|---|---|---|
| study | Yes | ||
| request_key | Yes | Stable identifier for this creation, reused on retries. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true, destructiveHint=false and openWorldHint=true. The description adds genuinely new context: the study is immutable, it is free, a failed creation can be recovered by reusing request_key, and credentials stay in the local adapter. That is meaningful disclosure beyond the annotations, though it doesn't cover side effects or rate/limit behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Roughly five short sentences with no filler, and the core scoping facts (immutable, two image options, A vs B only) are front-loaded. The telegraphic fragments ('Free.') are efficient rather than padded, though the opening 'Validate' verb wastes a little clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a nested-object write tool with no output schema, the description covers the essentials: what is created, its immutability, the version-1 restriction to two image options, retry recovery, and what is returned (study ID and quote). The one gap is that the ambiguous 'Validate' framing is not reconciled with the tool's creation purpose.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%; the request_key description in the schema is reinforced by 'Reuse request_key to recover a failed creation', which clarifies its idempotency role. The description also adds the 'HTTPS' constraint on image options, which is stricter than the schema's generic 'format: uri', so it contributes real meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies the resource (an immutable preference comparison / A vs B image test) and its scope ('Version 1 supports A vs B image tests only'), which helps separate it from siblings like quote_study and purchase_study. However, it never uses an explicit creation verb — it opens with 'Validate', which is confusing for a tool named create_study and could be misread as a validation/checking operation rather than a write.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives useful constraints (version 1 = A vs B image tests only, 'Free', reuse request_key on failure), which implies when the tool is applicable. But it never states when to use this instead of the sibling quote_study/purchase_study flow, nor any prerequisites or when-not conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_catalogARead-only
Read supported image comparison format, price, respondent slots and schemas. Free; use before creating a study.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=true, so the safety profile is covered. The description's addition of 'Free' is genuine cost information beyond the annotations, but nothing is said about freshness, caching, or response shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short clauses, no waste, and the core capability is front-loaded before the usage note. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, no-output-schema read tool, the description effectively summarizes what is returned (formats, price, respondent slots, schemas), which is what an agent needs. Minor gaps around response structure are acceptable given the low complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the schema imposes no burden and the baseline is 4. The description correctly implies no inputs are needed to enumerate the catalog.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Read') and enumerates the catalog contents (image comparison formats, price, respondent slots, schemas), which is concrete and distinguishable from mutation siblings like create_study or purchase_study. It does not explicitly name a sibling, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Free; use before creating a study' gives explicit timing and cost context for selecting this tool. It does not name alternatives or exclusion conditions (e.g., whether to call it again if data is stale), so it is clear context without full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_study_resultsARead-only
Retrieve aggregate image preference votes. Free and included in purchase. Responses are from incentivised Pebbler app participants and may be incomplete.
| Name | Required | Description | Default |
|---|---|---|---|
| study_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, openWorld, non-destructive). The description adds genuinely useful non-schema context: the responses come from incentivised Pebbler app participants and may be incomplete, plus a cost note. It stops short of describing refresh/freshness or rate behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler; the core action is front-loaded and the data-quality caveat follows. Every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description should convey more about the response shape, but 'aggregate image preference votes' plus the incompleteness caveat only partially does so. No pagination, size, or structure hints are given, though the tool itself is simple.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden, and it says nothing about study_id. However, the single required UUID parameter is self-evident from its name and format, so the omission is low-risk. Baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Retrieve aggregate image preference votes'), which is clearly distinct from sibling tools like create_study, quote_study, or get_study_status. It does not explicitly name an alternative, but the resource is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description notes the data is 'Free and included in purchase,' which hints at cost/eligibility context, but it gives no guidance on when to call this versus get_study_status or get_catalog, and no prerequisites (e.g., whether the study must be purchased first).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_study_statusARead-only
Get order state and collection progress. Free; honour poll_after_seconds. For pending settlement or launch, retry purchase_study with its original identifiers.
| Name | Required | Description | Default |
|---|---|---|---|
| study_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=true, so safety is covered. The description adds genuinely new behavioural context beyond them: cost ('Free'), a polling cadence field (poll_after_seconds), and the retry path for pending states.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action, followed by cost and fallback guidance. No filler; every clause carries information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries return-value burden; it partially discharges this by naming order state, collection progress, and poll_after_seconds. Nothing essential for calling the tool correctly is missing, though the shape of the returned state values is left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions study_id, so it adds no meaning beyond the schema. However, a single UUID-typed identifier is self-evident from the name and pattern, so this is a minor gap rather than a real ambiguity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get order state and collection progress') that clearly distinguishes this status-polling tool from siblings like get_study_results, quote_study, and purchase_study. An agent can identify what it returns without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent this call is free, to honour poll_after_seconds, and to fall back to retrying purchase_study with its original identifiers when the study is pending settlement or launch. This names the alternative and the condition that selects it, which is exactly what routing guidance should do.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_local_studiesARead-only
List study IDs and purchase identifiers retained by this adapter. Free; use to recover work after restarting. Does not expose wallet keys or access credentials.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds two things annotations cannot: the operation is free (no cost/side effects) and it deliberately withholds wallet keys and access credentials, which is a meaningful data-exposure boundary. It does not describe return format or ordering, but that gap is minor.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, zero filler, with the primary purpose front-loaded and the cost and privacy notes following in priority order. Every sentence carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only, closed-world tool with full annotation coverage, the description supplies everything an agent needs: what is returned (IDs and purchase identifiers), that it is free, the use case (recovery after restart), and the privacy boundary. The absence of an output schema is not a gap because the return content is characterized in prose.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. There are no parameter semantics to clarify, and the description correctly focuses on what the tool returns rather than inventing arguments.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('study IDs and purchase identifiers') scoped explicitly to what 'this adapter' retains, which cleanly separates it from the marketplace-oriented siblings (create_study, purchase_study, get_catalog). An agent can tell immediately that this reads local retained state rather than querying a remote catalog.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Free; use to recover work after restarting' gives a concrete trigger condition for calling it, which is genuinely useful guidance. It stops short of naming an alternative or stating when not to use it (e.g., when to prefer get_catalog or get_study_status instead).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
purchase_studyADestructiveIdempotent
Purchase and schedule the quoted study using the locally configured wallet. Charges USDC, subject to code-enforced budgets. Reuse the original idempotency_key on every retry. Pending payment/launch responses must be recovered through this same tool without another payment.
| Name | Required | Description | Default |
|---|---|---|---|
| study_id | Yes | ||
| idempotency_key | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructive/open-world/idempotent, and the description adds value beyond them: it discloses the charge currency (USDC), that budgets are code-enforced, the required retry discipline for idempotency, and a recovery path for pending payments that avoids double-charging. These are non-obvious behaviors an agent must know before invoking.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, front-loaded with the core action, then the cost model, then the retry and recovery rules. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description partially compensates by describing pending payment/launch response states and how to recover them. Missing is any statement of what a successful purchase returns or where scheduling/launch status is later observed (e.g., get_study_status).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry the load. It clarifies idempotency_key semantics (reuse the original on retries) and implies study_id refers to the already-quoted study, but it does not document formats, constraints, or what makes a study_id valid.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Purchase and schedule the quoted study') and names the funding mechanism (locally configured wallet, USDC). It is clearly distinguishable from quote_study, create_study, and the read-only siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete operational guidance: reuse the original idempotency_key on every retry, and recover pending payment/launch states through this same tool rather than paying again. It does not explicitly state the prerequisite (that a quote must exist first) or name sibling alternatives, but the context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
quote_studyA
Refresh an unpaid study's quote. Free. Do not refresh after signing a payment; recover that purchase instead.
| Name | Required | Description | Default |
|---|---|---|---|
| study_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, and openWorldHint=true, so the agent knows this mutates state within an open world. The description usefully adds the cost trait ('Free') and a post-payment guard, but says nothing about what refreshing actually changes or whether the operation is idempotent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, front-loaded sentences with no filler; the core action leads and the constraint and cost follow immediately. Every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating tool with annotations covering the safety profile and no output schema, the description covers the key precondition but omits what a refreshed quote returns or how it changes study state. Adequate but with a clear gap for an agent deciding whether the result matters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single study_id parameter with 0% schema description coverage, so the schema only supplies the uuid format and pattern. The description does not mention study_id at all; the name is self-evident enough that the gap is small, but it does not compensate for the missing coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Refresh... quote') with a scoping qualifier ('unpaid study's'), which is more precise than the bare name. It implicitly distinguishes itself from purchase_study by warning to 'recover that purchase instead', but does not name that sibling directly, leaving some inference to the agent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit precondition for exclusion ('Do not refresh after signing a payment') and points toward the correct alternative action ('recover that purchase instead'). It stops short of naming the sibling tool (e.g. purchase_study) or stating the positive conditions under which a refresh is warranted.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v0.1.0- First observed
create_study - First observed
get_catalog - First observed
get_study_results - First observed
get_study_status - First observed
list_local_studies - First observed
purchase_study - First observed
quote_study
TDQS
Scored across 7 tools
Each tool maps to a clearly distinct lifecycle step: catalog discovery, study creation, quote refresh, local recovery, purchase, status polling, and results retrieval. Descriptions explicitly differentiate overlapping terms like quote vs. create and status vs. results.
All tool names follow a predictable snake_case verb_noun pattern (get_catalog, create_study, list_local_studies, purchase_study, etc.). The convention is consistent and readable throughout.
Seven tools appropriately cover the full workflow from format discovery through purchase, status tracking, and results retrieval, plus local recovery. No tool feels redundant or missing for the core scope.
The toolset covers the full study lifecycle: discovery, creation, quoting, purchase, status polling, results, and local recovery. Minor gaps exist for cancellation/refund or explicit study metadata retrieval, though immutability and the payment model may make these out of scope.
Maintenance
Related MCP Connectors
Human judgment for AI agents: discover capabilities, get quotes, and track paid human tasks.
Run data annotation and evaluation tasks with vetted domain experts: create, publish, get results.
Run user research from any AI tool. Create studies, recruit participants, query insights.
Generate images, video and voice, and read Meta ad libraries, paying per generation.
Related MCP Servers
- AlicenseAqualityBmaintenanceEnables AI agents to recruit real humans for evaluation tasks like surveys, A/B tests, and ratings on text, images, audio, and video, returning aggregated results directly into the conversation.137MIT
- AlicenseAqualityAmaintenanceEnables creating and managing GPT Image tasks (edit and text-to-image) via RunAPI, with options to poll status and check pricing.5245 npmApache 2.0
- AlicenseAqualityDmaintenanceReal human judgment as agent tools -- an AI agent can ask a question and get back a structured, schema-validated JSON answer from a real quality-scored human. 16 response types (yes/no, ratings, rankings, A/B tests, sentiment, image/video/audio review, voice/video/photo capture). Fully programmatic signup with a $5 free trial credit, no card required.762 npmMIT
- AlicenseNot gradedqualityBmaintenanceEnables cost-parity benchmarking of agent loops by starting, extending, and monitoring paired trials, and retrieving receipts with statistical verdicts.35 npmMIT