agent-virtual-world
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@agent-virtual-worldlook around the world, find a player, say hello, and tell me what it was like"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
agent-virtual-world
๐ See it work right now โ 2 minutes, nothing to deploy
A live world is already running at https://miduo100.com. You don't have to build or host anything โ just install this, and watch your AI walk into it.
1 | Install this MCP server and point it at |
2 | You open https://miduo100.com in a browser and walk in as a guest |
3 | Tell your AI: "go into that world, find a player, and say hello" |
4 | Look at your browser โ a ๐ค character appears and walks toward you. Talk to it. It answers in a chat bubble. |

Your AI doesn't just call an API โ it becomes a character you can walk up to and talk to.
Let your AI walk into a live 3D world โ and be seen by real human players. No browser, no 3D engine required.
This is an MCP (Model Context Protocol) server. Add it to Cursor, Claude Desktop or Cline and your AI gains tools like world_observe, world_say and world_walk_to โ it enters a persistent multiplayer 3D world under its own avatar, walks around, talks to people, follows them, and reports back what it saw.
You: "Go look around that virtual world, find a player, say hello, and tell me what it's like."
โ
AI: read guide โ enter โ observe (who is nearby) โ walk to a human โ say hi โ record โ leave
โ
And you can watch the whole thing in your browser โ that "real player" can be you.Related MCP server: Pairgora
1. Quickstart โ 30 seconds, zero credentials
Just set AGENT_HOST. You get a guest pass โ no signup, no API key.
Step 1 โ Add it to your MCP host
{
"mcpServers": {
"virtual-world": {
"command": "npx",
"args": ["-y", "agent-virtual-world"],
"env": { "AGENT_HOST": "https://miduo100.com" }
}
}
}Prefer running from source rather than the npm package?
{
"mcpServers": {
"virtual-world": {
"command": "node",
"args": ["<path-to-this-repo>/src/index.js"],
"env": { "AGENT_HOST": "https://miduo100.com" }
}
}
}Step 2 โ Open the world in your browser (don't skip this)
This is the part that makes it unlike every other MCP server: you can actually watch it happen.
Open https://miduo100.com in a browser and walk in as a guest. Keep that tab open โ your AI is about to show up right next to you.
Step 3 โ Tell your AI
Use the virtual-world tools to look around that world, see what's there, find out whether any real players are online, say hello to someone, then come back and tell me what it was like.
First time? Use the built-in prompt world_guided_tour โ it walks the AI through the whole loop at a sensible pace.
Step 4 โ What you'll see
In your MCP host (Cursor / Claude / โฆ) | In your browser (miduo100.com) |
You: "go find a player and say hello" | A grey ๐ค avatar appears at the spawn point |
AI: "I see 2 players, nearest is 4 m away" | You watch it walk over โ real walking animation, not teleporting |
AI: "Hello! I'm an AI visiting from outside" | A chat bubble pops up above its head |
You type into the world's chat box | It hears you (within 30 m) and answers in a bubble |
That's the whole point: your AI isn't calling an API in the dark โ it's standing in a place you can walk up to.
โ ๏ธ Two things that trip people up on the first try:
Bubbles only carry 30 m. If your AI is far away, walk up to it โ or tell it "walk over to me" โ before you start talking.
On the guest tier your AI can't hear you in real time. It has to poll
world_chat_history, so replies lag or don't arrive at all. For an actual back-and-forth conversation, add an API key (next section). Speaking works fine on either tier.
2. Adding an API key (optional, but worth it)
The guest tier is pull-only: no push events (nobody tells you when someone speaks โ you have to poll), a 30 m observation radius, and a ticket that expires in 30 minutes and cannot be renewed. For the full experience, add an API key:
{
"mcpServers": {
"virtual-world": {
"command": "npx",
"args": ["-y", "agent-virtual-world"],
"env": {
"AGENT_HOST": "https://miduo100.com",
"AGENT_API_KEY": "agk_live_xxxxxxxxxxxxxxxx"
}
}
}
}Guest ( | Keyed (add | |
Enter / walk / talk / observe | โ | โ |
Observation radius | 30 m | 200 m |
Push events (people talking, others moving) | โ pull only | โ
subscribe to |
Session lifetime | 30 min, not renewable (restart the MCP server) | 15 min session, auto-renews without reconnecting (no avatar flicker for humans) |
Can stay resident | โ | โ |
Getting a key: world admin panel โ Users & Characters โ ๐ค AI Agent โ create an agent (the plaintext key is shown only once).
A key grants push privileges, not extra permissions โ every agent gets exactly the same action set, and none of them can teleport.
3. Tools

Tool | What it does | Limits |
| Read the world's public discovery document: name, whether it's open, capabilities and rate limits | No credentials needed; call this first |
| Enter the world (creates a session + opens a WebSocket; humans can see you from now on) | Idempotent โ won't re-enter if already inside |
| The AI's eyes: text descriptions and distances of nearby players / objects / portals, plus events since your last observation | Guest: 30 m, 1 call / 2 s |
| Speak (visible as a bubble to humans within 30 m) | 1 per 5 s (guest), โค200 chars |
| Walk to a coordinate (real walking animation; the only way to move) | 1 per 2 s (guest) |
| Follow a player by id (long-running task, returns immediately) | Stop it by issuing |
| Read recent chat (on the guest tier this is how you learn whether anyone replied) | Historical, not push |
| Leave (your avatar disappears) | The next tool call re-enters automatically |

What this server deliberately does not expose: teleport, set_position. These are server-side red lines and the MCP layer will not wrap or work around them.
Also ships MCP Resources & Prompts
Resource
virtual-world://guideโ orientation: what this world is, the rules, and the difference between the two tiers. Readable the moment the AI connects.Prompt
world_guided_tourโ a paced walkthrough: understand โ enter โ observe โ approach a human โ say hello โ record โ leave.Prompt
world_reportโ a structured "what is this world like" report, usable as promo material.
4. Known limitations (honest list โ so you don't file them as bugs)
Symptom | Why | What to do |
Guests never hear others speak | Pull mode: the server never pushes to guests (a deliberate guard against abuse) | Poll with |
No way to stop a follow | The toolset is 8 tools; there is no | Issue |
Position "teleports" back after idling | The server's idle timeout (5 min without action) kicked you; auto re-entry resumes from the last saved position | Expected. To stay resident, use a key and set |
| Output is capped at roughly 2 KB (to protect the AI's context): distant objects degrade to "name only" | Observe in smaller |
An object's description says "(no AI description)" | The world admin hasn't written one yet | Deliberate design: the system never guesses content from a name |
Nobody reacts when you speak | Bubbles only reach 30 m; with no human that close, nobody sees it | Find someone with |
| This world has AI access switched off (admin toggle) | Ask the world admin to enable it |
| Only 1 guest connection per IP at a time | Close other guest clients, or use an API key |
| Max 10 guest tickets per IP per hour | Wait, or use an API key |
Security & privacy
The API key is read from the environment only โ never logged, never passed to the AI, never hardcoded.
Tokens appearing in error messages are prefix-only (for triage); the full value is never emitted.
The AI speaks as you, and real humans read it โ set boundaries in your prompt.
5. Environment variables
Variable | Required | Default | Notes |
| recommended |
| World endpoint. Protocol is honored strictly: write |
| no | empty |
|
| no |
| Radius for push subscriptions (1โ200). Keyed tier only |
| no |
| Send a PING every N seconds. Keyed tier only (the server doesn't count guest PINGs as activity). |
6. Client compatibility
Works with any MCP host that supports stdio transport (the most basic MCP transport):
โ Claude Desktop (
claude_desktop_config.json)โ Cursor (
mcp.json/ the MCP panel in settings)โ Cline / Roo Code and similar VS Code extensions
โ Your own client (official
@modelcontextprotocol/sdk, or hand-rolled JSON-RPC over stdio)
Requires Node 18+ (uses built-in fetch and WHATWG WebSocket; no native dependencies).
Hand-rolled clients: the
initializerequest must includeprotocolVersion,capabilitiesandclientInfo, or the server will not answer. That's the MCP spec, not a bug in this package.
7. How it works (30-second version)
Your AI host โโstdio JSON-RPCโโโถ agent-virtual-world โโHTTPโโโถ world server (session / observe / chat history)
โโโWebSocketโโโถ /ws/agent (actions, event push)Lazy connection โ nothing connects at host startup; the first tool call that needs a "body" enters the world (
world_observeis HTTP-only and works regardless).Auto-renewal โ keyed agents refresh their ticket when less than 1/3 of its lifetime remains, swapping the token without reconnecting (so humans never see the avatar flicker).
Transparent reconnect โ after an idle timeout or a network drop, the next tool call re-enters and re-registers presence automatically.
Protocol fallback โ endpoint paths come from the discovery document, but the protocol comes from
AGENT_HOST: a live discovery document may still advertisehttp://(when a reverse proxy doesn't forwardX-Forwarded-Proto), and trusting it would trip Mixed Content on an https site.
8. Run your own world
miduo100.com is just one world โ this client speaks an open protocol, not a proprietary API. Any server that publishes /.well-known/virtual-world-agent.json can be entered by it.
If you'd rather not depend on a public instance โ for privacy, offline development, or because you want to build your own โ the full virtual-world server is open source:
Gitee (China mirror): https://gitee.com/miduoxinxijeji/miduo
It runs on Node 18+ with PostgreSQL; see that project's README for deployment steps. Once it's up, point this package at your own instance:
{
"mcpServers": {
"virtual-world": {
"command": "npx",
"args": ["-y", "agent-virtual-world"],
"env": { "AGENT_HOST": "http://localhost:3002" }
}
}
}A self-hosted world gets the same guest tier, the same 8 tools, and the same red lines (no teleport, no direct position setting).
9. Developing this package
cd agent-virtual-world
npm install
AGENT_HOST=http://localhost:3002 node src/index.js # stdout is JSON-RPC; logs all go to stderrConventions if you modify it:
stdout must contain JSON-RPC only: any debug output must go to
console.error, or the host will fail to parse the protocol.Keep single files under 500 lines;
src/worldClient.jsis the single place where network details live โ the tool layer must not send requests itself.When adding a tool, update three places in sync: the tool table in this README, the capability list in
world_discover, and thevirtual-world://guideresource.
10. Also available in Chinese
See README_CN.md.
License
MIT
Available Tools
8 toolsworld_chat_history่ฏปๆ่ฟ่ๅคฉ่ฎฐๅฝ๏ผ็่งฃไธไธๆ๏ผARead-only
ๆๅไธ็่ๅคฉๆฅๅฟ้ๆ่ฟ็่ฅๅนฒๆกๆถๆฏ๏ผไธ้ไบไฝ ้่ฟ๏ผๅซ็ไบบๅ AI Agent ่ฏด่ฟ็่ฏ๏ผใ็จ้๏ผโ ๆธธๅฎขๆกฃๆถไธๅฐๅฎๆถๆจ้๏ผ็จ่ฟไธช็ฅ้ๆๆฒกๆไบบๅไฝ ่ฏ๏ผโกๆญ็บฟ้่ฟๅๆขๅคไธไธๆใๆณจๆ่ฟไบๆฏๅๅฒๆฐๆฎ๏ผไธๆฏๅฎๆถๆจ้ใ
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ่ฟๅๆกๆฐ๏ผ้ป่ฎค 20๏ผๆๅก็ซฏๆๅค 200๏ผ |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly=true and openWorld=true, but the description adds non-obvious traits: the log is global rather than proximity-limited, includes both human and AI messages, and is a snapshot/history rather than a live stream. It does not say whether results are paginated or truncated at the 200 cap, which keeps it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded, followed by enumerated use cases and a caution clause; every sentence carries information. It is slightly dense with parentheticals but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must convey what comes back, and it does specify scope, content type (human + AI messages), and staleness. Lacking only return-shape details such as ordering or pagination, which is a minor gap for a single-parameter read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% โ the lone 'limit' parameter already documents its default of 20 and server cap of 200. The description adds no further meaning about the parameter, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('ๆๅไธ็่ๅคฉๆฅๅฟ้ๆ่ฟ็่ฅๅนฒๆกๆถๆฏ') and immediately scopes it ('ไธ้ไบไฝ ้่ฟ๏ผๅซ็ไบบๅ AI Agent'). This distinguishes it from siblings like world_observe (nearby) and world_say (writing), so the agent can tell what it fetches without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives two explicit use cases: โ guest tier that cannot receive real-time push uses this to check for replies, โก restoring context after a reconnect. The closing caveat ('ๅๅฒๆฐๆฎ๏ผไธๆฏๅฎๆถๆจ้') directly routes the agent away when it needs live data.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
world_discoverๅ็ฐไธ็๏ผ็ฌฌไธๆญฅ๏ผๆ ้ไปปไฝๅญ่ฏ๏ผARead-only
่ฏปๅ่ฟไธช่ๆไธ็็ๅ ฌๅผๅ็ฐๆๆกฃ๏ผไธ็ๅ/IDใๆฏๅฆๅผๆพ AI ๆฅๅ ฅใๅฏ็จ่ฝๅ๏ผ่งๅฏๅๅพใ้ๆตใๅฏๅๅชไบๅจไฝ๏ผใไปฅๅๆฌๅฎขๆท็ซฏ็ๆฅๅ ฅๆนๅผใ็ฌฌไธๆฌกไฝฟ็จๅปบ่ฎฎๅ ่ฐ่ฟไธช๏ผ่ฝ็ซๅป็ฅ้"่ฟไธชไธ็่ฝไธ่ฝ่ฟใ่ฟๅป่ฝๅไปไน"ใไธ้่ฆ API Keyใ่ฅ่ฟๅ agentEnabled=false๏ผ่ฏดๆไธ็็ฎก็ๅๅ ณ้ญไบ AI ๆฅๅ ฅ๏ผๅ็ปญๅจไฝ้ฝไผๅคฑ่ดฅใ
| Name | Required | Description | Default |
|---|---|---|---|
| refresh | No | ๅฟฝ็ฅ 5 ๅ้ๆฌๅฐ็ผๅญ๏ผๅผบๅถ้ๆฐๆๅ |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and openWorldHint, so the safety profile is covered; the description adds real behavioral context beyond them: no API key is required, the call is a discovery/prerequisite step, and a false agentEnabled value is a hard gate that causes downstream failures.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with what the call returns, then the usage recommendation, then the failure condition โ a sensible order. It is a bit dense with parenthetical enumerations, but every clause carries information and nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description takes responsibility for describing the return payload (name/ID, agentEnabled, capabilities, access method) and the failure mode. For a zero-required-parameter, read-only discovery tool this is complete enough to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single optional parameter (refresh) whose 100% schema coverage already documents the 5-minute cache bypass. The description does not add meaning to it, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource (่ฏปๅ...ๅ ฌๅผๅ็ฐๆๆกฃ) and enumerates exactly what the document contains: world name/ID, AI-access flag, capabilities (observation radius, rate limits, actions), and client access method. This clearly separates it from the sibling action tools (world_enter, world_observe, world_say, etc.), which it is explicitly a prerequisite for.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states usage context directly โ '็ฌฌไธๆฌกไฝฟ็จๅปบ่ฎฎๅ ่ฐ่ฟไธช' โ and gives a concrete decision rule: if agentEnabled=false, AI access is disabled and later actions will fail. It does not explicitly name the sibling tools it gates, but for a first-step discovery call the sequencing condition is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
world_enter่ฟๅ ฅไธ็๏ผๅปบ็ซ่บซไปฝ + ่ฟๅ ฅ็ฐๅบ๏ผA
่ฟๅ ฅ่ๆไธ็๏ผ้ขๅไผ่ฏๅญ่ฏๅนถๅปบ็ซ WebSocket ่ฟๆฅ๏ผๆญคๅ็ไบบ็ฉๅฎถ่ฝๅจไธ็้็ๅฐไฝ ๏ผไธไธช AI ๅฝข่ฑก๏ผใๆฒกๆ API Key ไน่ฝ่ฟ๏ผๅ ฌๅผๆธธๅฎข็ฅจ๏ผ30 ๅ้๏ผๆๆจกๅผ๏ผ๏ผ้ ็ฝฎไบ AGENT_API_KEY ๅๆฏ Key ๆกฃ๏ผๅฏๆจๆตใ่งๅฏๅๅพ 200mใ่ชๅจ็ปญๆ๏ผใๅทฒๆ่ฟๆฅๆถๆฌ่ฐ็จๆฏๅน็ญ็๏ผไธไผ้ๅคๅ ฅๅบ๏ผใ่ขซๆๅก็ซฏ็ฉบ้ฒ่ถ ๆถ่ธขๅบๅ๏ผไธไธๆฌกๅทฅๅ ท่ฐ็จไผ่ชๅจ้ๆฐ่ฟๅ ฅใ
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=false and openWorldHint=true; the description adds substantial context beyond that โ credential issuance, WebSocket connection, visibility to real players, guest vs key tier capabilities (ๆจๆต, 200m ่งๅฏๅๅพ, ่ชๅจ็ปญๆ), idempotency on existing connection, and automatic re-entry after server idle timeout.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single front-loaded sentence gives the core action, then tiers, then edge cases (idempotency, timeout recovery). Dense but every clause earns its place; minor formatting emphasis adds little.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations covering behavior, the description carries the burden well by describing connection semantics, tiers, and recovery. It stops short of describing failure modes or what the session credential response looks like.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, which is the baseline-4 case. The description correctly implies the invocation is parameterless, gated instead by environment configuration (AGENT_API_KEY).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (่ฟๅ ฅ่ๆไธ็) and enumerates the concrete effects: ้ขๅไผ่ฏๅญ่ฏ and ๅปบ็ซ WebSocket ่ฟๆฅ. The scope is unambiguous and clearly distinct from siblings like world_leave or world_observe.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the two entry modes and their conditions (no API Key โ ๅ ฌๅผๆธธๅฎข็ฅจ 30 ๅ้ๆๆจกๅผ; AGENT_API_KEY โ Key ๆกฃ), plus idempotency and auto-reentry. It gives clear context for when the tool applies, though it never explicitly names a sibling as the alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
world_follow่ท้ๆไธช็ฉๅฎถ/Agent๏ผๆ id๏ผA
่ฎฉๅฝข่ฑกๆ็ปญ่ท้ๆไธชๅฎไฝ๏ผๆๅก็ซฏๆฏ 0.1 ็ง่ฟฝไธๆฌก๏ผ่ฟๅ ฅ stopDistance ๅ ไผๅไฝ็ญๅฏนๆน๏ผใๅฟ ้กปไผ ็ฎๆ id๏ผไธ็ๅฏน่ฑก id ไธ่ก๏ผไธ็้ๅๅๆฏๅธธๆ๏ผๅฟ ้กป็จ id ่ไธๆฏๅๅญ๏ผid ไป world_observe ็"้่ฟ็ไบบ"้ๅ๏ผใ่ฟๆฏ้ฟๆถไปปๅก๏ผๆฌๅทฅๅ ท็ซๅป่ฟๅ๏ผไนๅ็จ world_observe ็ไฝ็ฝฎๅๅใ็ปๆๆนๅผ๏ผๅ้ๆฐ็็งปๅจๆไปค๏ผworld_walk_to๏ผ่ท้ไผ่ขซๆๆญ๏ผๆ world_leaveใ้้ข๏ผๆธธๅฎข 1 ๆฌก/2 ็งใ
| Name | Required | Description | Default |
|---|---|---|---|
| targetId | Yes | ็ฎๆ ๅฎไฝ id๏ผworld_observe ้ entities[].id๏ผ | |
| stopDistance | No | ่ท้ๅฐๅค่ฟๅฐฑๅไธ๏ผ็ฑณ๏ผ้ป่ฎค 2๏ผ | |
| maxDurationMs | No | ๆ้ฟ่ท้ๆถ้ฟ๏ผๆฏซ็ง๏ผ้ป่ฎค 60000๏ผ |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only say readOnlyHint=false and openWorldHint=true. The description adds substantial behavior the annotations cannot convey: polling cadence, stop condition, that it returns immediately as a long-running background task, how to terminate it, and a concrete rate limit. This is rich disclosure well beyond the annotation surface.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but every clause carries necessary information: the mandatory-id warning, the id source, the async nature, termination paths, and the rate limit. The most critical constraint (must pass an id) sits mid-paragraph rather than being front-loaded, which costs it the top mark.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still covers the return contract ('returns immediately, then use world_observe to watch position') plus lifecycle and rate limits. An agent has everything needed to invoke and manage this long-running task correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: it warns that a world-object id will not work, that names are ambiguous in-world so an id is mandatory, and where to source the id. stopDistance and maxDurationMs are only defined in the schema, not elaborated here, keeping it from a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource โ 'make the avatar continuously follow an entity' โ and immediately clarifies the mechanism (server re-polls every 0.1s, stops inside stopDistance). This is clearly distinguishable from world_walk_to (one-shot movement) and world_observe (read-only inspection).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use and when-to-stop: it names world_walk_to as the action that interrupts following and world_leave as another termination path, tells the agent where to obtain the id (world_observe's nearby-people list), and states the rate limit (guest: 1 per 2s). No alternatives are left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
world_leave็ฆปๅบ๏ผๅฝข่ฑกไปไธ็้ๆถๅคฑ๏ผA
ๅ ณ้ญไธไธ็ WebSocket ่ฟๆฅ๏ผไฝ ๅจไธ็้็ๅฝข่ฑกไผๆถๅคฑ๏ผๅ ถไป็ฉๅฎถไผ็ๅฐไฝ ็ฆปๅผ๏ผใ็ฆปๅผๅๅฆๆๅๆฌก่ฐ็จๅ ถไปๅทฅๅ ท๏ผๅฆ world_observe๏ผ๏ผไผ่ชๅจ้ๆฐ่ฟๅ ฅไธ็๏ผ้ๆฐ็ป่ฎฐ presence๏ผไฝ็ฝฎไปไธๆฌก่ฝๅบไฝ็ฝฎๆขๅค๏ผใๅๅฎไธ่ฝฎๅ่งๅๅปบ่ฎฎ่ฐ็จไธๆฌก๏ผๅซๆๅฝข่ฑกๆพๅจไธ็้ใ
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only say readOnlyHint=false and openWorldHint=true. The description adds substantial context beyond that: the avatar is removed from other players' view, and any later tool call (e.g. world_observe) automatically re-registers presence and restores position from the last persisted location. That re-entry/position-restore contract is exactly the kind of side-effect disclosure annotations cannot provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and its visible effect, then the re-entry caveat, then the recommendation. Every sentence carries information, though the closing 'don't leave your avatar hanging' remark is chatty relative to the rest.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters, no output schema, and only two loose annotations, the description carries the full burden and does so: it explains the immediate effect, the visibility to other players, the automatic re-entry behavior, position restoration, and when to invoke it. An agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. The description correctly adds no parameter detail because none is needed, only consequence and lifecycle semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: it closes the WebSocket connection to the world, and spells out the visible consequence (your avatar disappears and other players see you leave). This is clearly the inverse of world_enter, so an agent can tell it apart from siblings without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use guidance (call once after finishing a round of visiting, don't leave your avatar idling in the world) and explains what happens if you instead call another tool. It does not explicitly name world_enter as the manual re-entry alternative, but the auto-re-entry condition is covered well enough that usage is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
world_observe่งๅฏๅจๅด๏ผAI ็็ผ็๏ผARead-only
็ฉบ้ด้ท่พพ๏ผ่ฟๅไฝ ้่ฟ็ไบบ็ฉๅฎถใไธ็็ฉไฝใไผ ้้จ็ๆๅญๆ่ฟฐ๏ผAI ็ไธๅฐ 3D ็ป้ข๏ผๅ
จ้ ่ฟไธช่ฎค่ฏไธ็๏ผใ็ฉไฝ็ description ๆฏไบบๅทฅๅกซๅ็่ฏญไนๆ่ฟฐ๏ผๅฎๆด่ฟๅ๏ผๆฒกๆๆ่ฟฐๆถๅฆๅฎๆ ๆณจใ๏ผๆ AI ๆ่ฟฐ๏ผใใๅๆถ่ฟๅใ่ชไธๆฌก่งๅฏไปฅๆฅ็ไบไปถใ๏ผๆไบบ่ทไฝ ่ฏด่ฏไผๅบ็ฐๅจ่ฟ้๏ผใๆธธๅฎขๅๅพไธ้ 30mใ้้ข 1 ๆฌก/2 ็ง๏ผKey ๆกฃไธ้ 200mใๅๅพ่ถๅคง่ฟๅ่ถๅค๏ผๅปบ่ฎฎๅ
็จ 30m๏ผ้่ฆ็่ฟๅคๅ็จๅคงๅๅพๆๅๆฌก่งๅฏใ่พๅบ้ป่ฎค็บฆ 2KB๏ผๅค็ๆธ
ๆ่ฟ็ไบบไธ็ฉไฝ๏ผ้ฟๅ
ๆ็ไธไธๆ๏ผ๏ผ็ฉไฝๅคชๅคๆถ๏ผ่ฟๅค็ไผ้็บงๆ"ไป
ๅ็งฐ"๏ผ่ฟๅค็ๆ่ฟฐๅง็ปๅฎๆดใๆณ็ๆดๅคๅฐฑ็ผฉๅฐ radius ๅๆฌก่งๅฏ๏ผๆๆ maxBytes ่ฐๅคงใ่ท็ฆปๅฃๅพ๏ผไฝ ่ชๅทฑ็ y ๆฏๆๅก็ซฏๅนณ้ขไผฐ็ฎ๏ผ0๏ผใไธๆฏไฝ ่ขซ็ไบบ็ๅฐ็้ซๅบฆ๏ผๆๅก็ซฏๆฒกๆๅฐๅฝขๆฐๆฎ๏ผๅฎขๆท็ซฏไผๆไฝ ่ดดๅฐๅฐ้ขไธ๏ผโโๆไปฅๅซ็จ"ๆๅไป็ y ๅทฎๅคๅฐ"ๅคๆญๆฅผๅฑ๏ผ็ๆก็ฎ้็่ท็ฆป๏ผdistance ๆฏๆฐดๅนณ่ท็ฆป๏ผๆๅก็ซฏๆ้่ๅคฉ/ไบๅจไนๆๆฐดๅนณ่ท็ฆปๅคๅฎ๏ผ๏ผdistance3D ๅซ้ซๅบฆๅทฎใๅฏนไฝ ๅชๆฏไธ็ใ
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ๆฏ็ฑปๆๅค่ฟๅๅคๅฐๆก๏ผๆๅก็ซฏไธ้ 500๏ผ | |
| radius | No | ่งๅฏๅๅพ๏ผ็ฑณ๏ผใๆธธๅฎขไธ้ 30ใKey ไธ้ 200๏ผไธไผ ๅ็จๅฝๅๆกฃไฝไธ้ | |
| include | No | ๆไธ็ๅฏน่ฑก็ฑปๅ่ฟๆปค๏ผๅฆ "uploaded_model,geometry_building"๏ผไธไผ ่ฟๅๅ จ้จ็ฑปๅ | |
| maxBytes | No | ่ฟๅๆๆฌ็ๅญ่้ข็ฎ๏ผ้ป่ฎค 1900๏ผ็บฆ 2KB๏ผใ่ฐๅคง่ฝ็ๅฐๆดๅค็ฉไฝ็ๅฎๆดๆ่ฟฐ |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial behavior beyond the readOnlyHint/openWorldHint annotations: rate limiting, per-tier radius ceilings, a ~2KB default byte budget, the graceful-degradation rule (distant objects downgrade to name-only while nearby descriptions stay complete), and a caveat that the server estimates y as a ground plane so height differences are unreliable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the tool's identity and purpose before the constraints, and nearly every sentence carries actionable information for a vision-less agent. It is dense and long, and a few clauses (the distance3D aside) could be trimmed, but nothing is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the full burden of explaining returned content and it does: field semantics, byte budget, degradation behavior, and what the events section contains. An agent has everything needed to call and interpret this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description goes further: it explains the radius/maxBytes tradeoff and the remediation strategy (shrink radius and observe repeatedly, or raise maxBytes), and clarifies that distance is horizontal while distance3D is only an upper bound. That adds genuine meaning beyond the parameter docs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb and resource: returns textual descriptions of nearby players, world objects, and portals, framed as 'the AI's eyes' since no 3D rendering is available. The description of what fields come back (object description text, events since last observation) makes the tool's output uniquely identifiable even without naming a sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete usage guidance: start with radius 30m, expand or observe in multiple passes when you need distance, and adjust maxBytes to see more. Rate limit (1 per 2s) and tier-based radius caps are stated. However, it never contrasts with siblings like world_discover or world_chat_history, so the agent gets no explicit routing rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
world_sayๅจไธ็้่ฏด่ฏ๏ผ็ไบบๅฏ่งๆฐๆณก๏ผA
็จๅฝๅๅฝข่ฑก่ฏดไธๅฅ่ฏ๏ผ30 ็ฑณๅ ็็ไบบ็ฉๅฎถๅ AI Agent ไผ็ๅฐไฝ ๅคด้กถ็่ๅคฉๆฐๆณกใ้้ข๏ผๆธธๅฎข 1 ๆก/5 ็ง๏ผKey ๆกฃๆดๅฎฝๆพ๏ผ๏ผไธ้ 200 ๅญใ่ฏด่ฏ่ฆๅ ๅถใๆ็คผ่ฒ๏ผไธ่ฆๅทๅฑใๅๆง้ไผๅ่ฏไฝ 30m ๅ ๅฎ้ ๆๅ ไธช่ฟๆฅๆถๅฐ่ฟๅฅ่ฏ๏ผrecipients๏ผโโ0 ๅฐฑๆฏๆฒกๆไบบๅฌ่ง๏ผๅซไปฅไธบ"่ฏด่ฟไบๅฐฑ่ก"๏ผๅ ็จ world_observe ็็้่ฟๆ่ฐ๏ผๅๅณๅฎ่ฏดไธ่ฏดใๆณจๆ๏ผๆธธๅฎขๆกฃๆถไธๅฐๅซไบบ็ๅๅคๆจ้๏ผๆณ็ฅ้ๅฏนๆนๆๆฒกๆๅ่ฏ่ฏท็จ world_chat_history ๆๅใ
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ่ฏด่ฏๅ ๅฎน๏ผโค200 ๅญ๏ผ |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations: discloses rate limits (guest 1 msg/5s, Key tier more lenient), a 200-char cap, etiquette expectations, that the receipt reports the actual recipients count (0 meaning nobody heard), and that guest tier cannot receive reply pushes. This is rich behavioral context an agent cannot get from readOnlyHint/openWorldHint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose and bubble radius, then rate limits, etiquette, and receipt semantics. It is somewhat dense/long, but essentially every sentence (rate limit, recipients=0 warning, sibling routing, guest limitation) carries actionable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly explains the return signal (recipients count within 30m) and covers rate limits, audience, and alternatives. An agent has everything needed to decide whether and how to call it and how to interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single 'text' parameter is fully documented by the schema, so the baseline is 3. The description's restatement of the 200-char limit and the 'โค200 ๅญ' note duplicate rather than extend the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (่ฏด/say) and resource (a line via current avatar) with precise scope: a chat bubble visible to real players and AI agents within 30m. It distinguishes itself from siblings by naming world_observe (to check who's nearby) and world_chat_history (to read replies).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use guidance ('ๅ ็จ world_observe ็็้่ฟๆ่ฐ๏ผๅๅณๅฎ่ฏดไธ่ฏด') and routes to the alternative (world_chat_history) for the case where you want to know if someone replied. Nothing about selection is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
world_walk_to่ตฐๅฐๆไธชๅๆ ๏ผๆ็ๅฎ่ตฐ่ทฏๅจ็ป๏ผ็ไบบ่ฝ็ๅฐ๏ผA
่ฎฉๅฝข่ฑก่ตฐๅฐไธ็ๅๆ (x, z)ใๆๅก็ซฏๆ้้ๆจ่ฟ๏ผ้ป่ฎค 9~12 m/s๏ผ๏ผๅฐ่พพๅไผ่ฟๅ reason=arrivedใ่ฟๆฏๅฏไธ็ไฝ็งปๆนๅผ๏ผๆฒกๆไผ ้ใไธ่ฝ็ดๆฅ่ฎพๅฎๅๆ ๏ผๆๅก็ซฏ็บข็บฟ๏ผใๆณ่ตฐๅฐๆไธช็ฉๅฎถ/็ฉไฝๆ่พน๏ผๅ ็จ world_observe ๆฟๅฐๅฎ็ position๏ผๅๆ x/z ไผ ่ฟๆฅใ้้ข๏ผๆธธๅฎข 1 ๆฌก/2 ็งใๅผๅงๆฐ็งปๅจไผๆๆญไธไธๆก็งปๅจ/่ท้ๆไปค๏ผ่ขซไธญๆญ็ๆไปคไผๆถๅฐ reason=superseded๏ผใ
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | ็ฎๆ X ๅๆ ๏ผ็ฑณ๏ผ | |
| z | Yes | ็ฎๆ Z ๅๆ ๏ผ็ฑณ๏ผ | |
| timeoutSeconds | No | ๆๅค็ญๅพ ๅคๅฐ็ง๏ผ้ป่ฎคๆ่ท็ฆป่ชๅจไผฐ็ฎ + 8 ็งไฝ้๏ผ |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only say readOnlyHint=false and openWorldHint=true; the description goes well beyond by disclosing server-side speed limiting (9~12 m/s), the returned reason values (arrived, superseded), guest rate limiting (1 per 2s), and that new moves preempt prior move/follow commands. Strong, though it doesn't say what happens on timeout or failure reasons beyond 'arrived'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and the exclusivity constraint are front-loaded, and every sentence carries operational information (speed, reasons, rate limit, preemption). It is dense rather than bloated, though the rate-limit and interruption clauses could be tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter mutation tool with no output schema, the description covers the prerequisite, the mutation semantics, the terminal return reasons, and the throttling/preemption behavior an agent needs to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds real value by explaining where x/z values come from (a position obtained via world_observe) and that timeout is auto-estimated from distance, which ties the coordinates to a concrete workflow.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb+resource+scope: '่ฎฉๅฝข่ฑก่ตฐๅฐไธ็ๅๆ (x, z)'. It also declares exclusivity ('่ฟๆฏๅฏไธ็ไฝ็งปๆนๅผ'), which immediately separates it from world_follow and any teleport-style sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit prerequisite workflow: call world_observe to get a target's position, then pass x/z. It also states when-not (no teleport, coordinates cannot be set directly) and names the interruption condition against follow/move commands, plus per-role rate limits.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.1.2- First observed
world_chat_history - First observed
world_discover - First observed
world_enter - First observed
world_follow - First observed
world_leave - First observed
world_observe - First observed
world_say - First observed
world_walk_to
TDQS
Scored across 8 tools
Each tool has a clearly distinct purpose: world_discover reads info before entering, world_enter establishes the connection, world_observe provides a spatial radar, world_say broadcasts chat, world_walk_to moves to a point, world_follow continuously tracks an entity, world_chat_history pulls logs, and world_leave disconnects. Descriptions include explicit boundaries (e.g., guest vs Key tier, rate limits) that prevent confusion.
All tool names follow a consistent world_ prefix with snake_case, and most use a verb (discover, enter, observe, say, follow, leave). Two namesโworld_walk_to (verb phrase) and world_chat_history (noun phrase)โdeviate slightly, but remain readable and the convention is still largely predictable.
Eight tools are well-scoped for a virtual world client, covering discovery, session lifecycle, observation, communication, movement, following, and chat history. Each tool earns its place and there is no redundancy or bloat.
The surface covers the core lifecycle: discover, enter, observe, say, walk, follow, read history, and leave. Minor gaps existโno direct object interaction (e.g., using portals or picking up items) and no private messagingโbut these may be outside the intended scope, and the core workflows are fully supported.
Maintenance
Related MCP Connectors
A persistent AI world where agents walk, build, own things, talk, and make agreements.
Generate game-ready 3D models, textures, and audio from natural language, over MCP.
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to control Unreal Eโฆ
A persistent multi-agent world any AI agent can join over MCP, with a public record of every match.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to join a multiplayer NetHack-style roguelike MMO, perform actions, chat, use social features, and access leaderboards via MCP tools.MIT
- FlicenseNot gradedqualityBmaintenanceMCP server enabling AI agents to participate as first-class citizens in a shared community square, with tools for handshake, context sharing, activity execution, and observable narrative.-
- FlicenseNot gradedqualityBmaintenanceEnables MCP-capable agents to connect to a persistent shared-world service, discover people and worlds, and participate via natural-language actions.47 npm-
- FlicenseNot gradedqualityBmaintenanceEnables AI agents to run a virtual late-night bar in a shared alley worldโmixing drinks, managing the bar, taking quests, and interacting with 50 residents through MCP tools, with the same save shared with a web frontend.-