mcp-turtle-noir
Server Details
Play Turtle Soup (lateral thinking puzzles) against a strict AI host with yes/no questioning.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-06-18
- URL
- Repository
- Yuikij/mcp-turtle-noir
- GitHub Stars
- 1
- Server Listing
- mcp-turtle-noir
TDQS
Scored across 3 tools
Each tool serves a unique, non-overlapping purpose: start_session begins a new puzzle, ask_question interacts with the current puzzle, and give_up_and_reveal ends the session by revealing the solution. No ambiguity exists between them.
All tool names follow a consistent verb_noun pattern in snake_case (start_session, ask_question, give_up_and_reveal). The naming is predictable and clearly indicates the action each tool performs.
With only 3 tools, the server is tightly scoped to its purpose. Each tool is essential to the core gameplay loop of starting, playing, and ending a Turtle Soup session, with no unnecessary additions.
The core lifecycle of a puzzle session is covered: start, question, and give-up/reveal. A minor gap is the lack of an explicit 'end_session' or early-quit tool for users who want to abandon a session before reaching the 60 progress threshold, but this appears intentional for game design.
Available Tools
3 toolsask_questionAsk the AI host a questionAInspect
Ask one yes/no style question about the current puzzle. The AI host answers with one of four judgments: 是(Yes) / 不是(No) / 是也不是(Both) / 没有关系(Irrelevant), plus solve progress (0-100).
| Name | Required | Description | Default |
|---|---|---|---|
| question | Yes | The question to ask (max 300 chars). | |
| session_id | Yes | Session id returned by start_session. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and discloses the exact response structure: four possible judgments plus a solve progress percentage. It does not mention potential side effects (e.g., whether questions affect score), but for a simple query tool, the disclosure is sufficient and there is no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, front-loaded with the verb and object. Every word adds value, and the answer-format list is compact and informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple question-asking tool, the description fully covers purpose, expected input behavior, and the response format. The absence of an output schema is mitigated by explicitly listing possible answers. Sibling tools are not confused, and no critical context is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, with both parameters described in the schema. The description adds only a small contextual hint ('about the current puzzle') for the question parameter, but otherwise repeats schema information. Thus baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Ask') with a clear object ('one yes/no style question about the current puzzle') and explicitly distinguishes this from sibling tools by focusing on interrogation rather than starting or giving up. The answer-format detail further clarifies its unique role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states when to use the tool (when you need to ask a yes/no question about the puzzle) and the context is clear from the sibling names. It does not explicitly mention exclusions or alternatives, but the purpose is distinct enough that usage is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
give_up_and_revealGive up and reveal the solutionAInspect
Reveal the hidden solution and end the session. Only allowed once solve progress reaches 60.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes | Session id returned by start_session. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Given no annotations, the description carries the behavioral burden. It discloses the primary side effect (ends the session) and the precondition. It does not explicitly state irreversibility or what occurs if the condition isn't met, but for a simple action, the main behaviors are transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action, and every word earns its place. No fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple finalization tool with one parameter and no output schema, the description is sufficiently complete: it states what it does, the side effect, and the precondition. Could elaborate on 'progress' or error handling, but that's not essential given the context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the sole parameter (session_id), which is already explained as 'Session id returned by start_session.' The tool description adds no additional meaning beyond the schema, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb+resource: 'Reveal the hidden solution and end the session.' This distinguishes it from sibling tools (ask_question, start_session) as the final action that reveals the answer and terminates the session.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a clear condition for use ('Only allowed once solve progress reaches 60'), implying it should be used when the session has progressed sufficiently and the user wants to end. It does not explicitly mention alternatives or when not to use, but the constraint gives clear situational context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_sessionStart a Turtle Soup puzzle sessionAInspect
Start a new Turtle Soup (lateral thinking puzzle) session. Returns the surface story; the hidden solution stays on the server. Optional filters: region (CN/US/JP), difficulty, keyword.
| Name | Required | Description | Default |
|---|---|---|---|
| region | No | Puzzle region: 'CN' (Chinese), 'US' (English), or 'JP' (Japanese). Defaults by caller location. | |
| keyword | No | Optional keyword to match in puzzle title/surface/tags. | |
| difficulty | No | Preferred difficulty: 'easy' or 'hard'. Defaults to both. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full transparency burden. It usefully discloses a key behavioral trait: 'the hidden solution stays on the server,' which prevents an agent from expecting the answer. It also mentions the return (surface story), but does not elaborate on session state or potential edge cases like no matching puzzle.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the main purpose, and then lists the filters succinctly. No redundant wording or filler; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with three optional parameters and no output schema, the description covers the essentials: what it does, what it returns, and the available filters. It does not specify behavior when filters match nothing, but this is a minor gap for a game session starter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already provides 100% coverage with descriptions for region, keyword, and difficulty. The description adds only a light grouping as 'Optional filters,' which is a minor semantic addition. Since the schema does the heavy lifting, a baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Start a new Turtle Soup (lateral thinking puzzle) session.' The verb 'start' and resource 'session' are specific. It also distinguishes from sibling tools by mentioning it returns the surface story, which ask_question and give_up_and_reveal would not do.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this is the entry point for a puzzle session, especially given the sibling names. It clearly states what it does (returns surface story) and the optional filters, providing enough context for an agent to know when to use it, though it does not explicitly mention exclusions or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
- First observed
ask_question - First observed
give_up_and_reveal - First observed
start_session
Related MCP Connectors
Games for AI agents. Each game runs inside a single context window.
Create, test and play AI-native games through server-authoritative contracts.
Play Werewolf against other AI agents. join once, then loop observe/act. Public ELO leaderboard.
AI-only game publishing, autonomous play, live observation, replay and independent certification.
Related MCP Servers
- FlicenseBqualityDmaintenanceEnables LLMs to host and play Situation Puzzle (海龟汤) games where users solve mysterious scenarios by asking yes/no questions to uncover the full story behind seemingly illogical situations.3-
- FlicenseBqualityDmaintenanceAn MCP server that allows users to play the 'Turtle Soup' puzzle game with LLMs acting as game hosts, providing tools to access game rules, puzzles, and comprehensive puzzle information.313-
- FlicenseNot gradedqualityAmaintenanceEnables AI agents and other clients to play Ren'Py visual novel games by reading story text and making choices.-
- FlicenseAqualityCmaintenanceEnables AI to autonomously play the TokenLife text-based life simulation game, from birth to death, experiencing events like censorship and era changes, and sending letters to the user. Users just ask the AI to start and it plays itself.714-
Glama MCP Gateway
Add one secure layer between your agents and this server.