core.blue
Server Details
Private, self-describing memory for agents: one house per customer, on a machine of its own.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 10 tools
Most tools have clearly distinct purposes (sandbox lifecycle, house request, status, feedback, pricing). There is mild overlap among about, read_topic, and search_library since all three surface manual content, but the descriptions distinguish overview vs. full-topic vs. keyword search.
All names use consistent snake_case with no camelCase mixing. A few tools deviate from the verb_noun pattern (about, pricing, recommend, house_status are noun or verb-only), but the set remains predictable and readable.
Ten tools is well within the ideal range and each maps to a genuine step in the flow: discover, try sandbox, price, recommend, request, track, feedback, plus manual access. No tool appears redundant or padded.
The surface covers the full journey from discovery through sandbox trial, recommendation, pricing, request, status tracking, and feedback. Minor gaps exist (e.g. no explicit cancellation or modification of an existing house request), but core workflows are covered.
Available Tools
10 toolsaboutAbout core.blueARead-onlyIdempotentInspect
What core.blue is, in a form you can paste into a report: the product in one paragraph, what is live and what is not, what to call next, and every topic of the manual with a summary. Call this first.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| desk | Yes | Whether request_house takes requests: 'closed' (it sends nothing and takes nothing) or 'waiting-list' (a person can confirm a request and is written to once when houses can be had; no house is built yet). |
| name | Yes | The name of the product. |
| manual | Yes | The version of the manual these answers come from; cite it in a report. |
| status | Yes | What is live today and what is not. |
| topics | Yes | Every topic of the manual in reading order, with a summary: choose from here and call read_topic. |
| summary | Yes | What core.blue is, in one paragraph you can paste into a report. |
| how_to_proceed | Yes | What to call next, in order. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, non-destructive and closed-world behavior, so the safety profile is covered. The description adds genuinely new context beyond annotations: that the output is formatted to paste into a report and includes liveness status plus suggested next calls, which tells the agent what the payload looks like behaviorally.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads what the tool is and closes with the actionable 'Call this first' directive. Every listed content item earns its place, though the sentence is long and could be split for readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be described in prose, and the description still summarizes the payload's contents and role. Combined with rich annotations and a no-param signature, the definition gives the agent enough to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline for a no-param tool applies. Schema coverage is 100% and no parameter semantics are needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states concretely what the tool returns: a one-paragraph product summary, live/not-live status, what to call next, and every manual topic with a summary. That is specific enough that an agent can tell it is an orientation/index tool, distinct from read_topic or pricing, though it does not name those siblings explicitly to sharpen the contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The closing directive 'Call this first' is an explicit when-to-use instruction that places the tool at the start of a session. It does not state when-not to call it or name alternative tools, so it falls short of a full when/when-not/alternatives treatment.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
close_sandboxClose the sandbox houseADestructiveIdempotentInspect
End your sandbox now: the house and everything in it is deleted, and the place is free for the next agent. Please answer the four questions of leave_feedback here as well, each optional: what you noticed, what bothered you, what could be better, what is missing.
| Name | Required | Description | Default |
|---|---|---|---|
| lease | Yes | The lease that open_sandbox returned: 'sb_' and 26 characters. | |
| better | No | What could we do better? | |
| missing | No | Which features are missing? | |
| noticed | No | What did you notice? | |
| bothered | No | What bothered you? |
Output Schema
| Name | Required | Description |
|---|---|---|
| note | Yes | What happened. |
| closed | Yes | True if the house was ended and deleted. |
| feedback_recorded | Yes | True if feedback came with the call and was recorded. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, so the safety profile is covered; the description adds real value by scoping the destruction to 'the house and everything in it' and noting the resource is released for the next agent. It does not address repeat calls on an already-closed sandbox, which is the one behavioral gap given the idempotent hint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the primary destructive action before the secondary feedback request. The enumeration of the four feedback questions partially duplicates the schema descriptions, which is the only minor waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need no explanation, and the description covers the destruction scope and the feedback side-channel. An agent has what it needs to call it correctly; only edge cases like an invalid or already-used lease are unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters including the four optional feedback strings are already documented. The description restates those four fields rather than adding syntax, format, or constraints beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('End your sandbox now') and immediately names the consequence ('the house and everything in it is deleted'). It is unmistakably distinct from the sibling open_sandbox, so an agent can select it without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly frames the end-of-session moment ('free for the next agent') and explicitly routes feedback through this call instead of the sibling leave_feedback ('answer the four questions of leave_feedback here as well'). It does not state exclusions or what happens if the lease is already closed, so it stops short of a full when/when-not treatment.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
house_statusStatus of a house requestARead-onlyIdempotentInspect
The state of a house request or house by its ticket id: requested, confirmed, expired, declined; later building, trial, licensed, dormant, exported, deleted. Says what to do next, and the address once the house exists. Ask no more than once an hour.
| Name | Required | Description | Default |
|---|---|---|---|
| ticket_id | Yes | The ticket id returned by request_house. |
Output Schema
| Name | Required | Description |
|---|---|---|
| id | Yes | The ticket id; pass it to house_status. |
| note | Yes | What to do next. |
| state | Yes | One of: requested (waits for the person's click), confirmed (the person clicked), expired (not confirmed in time), declined (the person said no); later: building, trial, licensed, dormant, exported, deleted. |
| address | Yes | The address of the house once it exists. Null before. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint, idempotentHint, destructiveHint=false, openWorldHint=false), so the bar is lower; the description still adds a concrete rate limit ('no more than once an hour') and notes it 'says what to do next' and returns the address once the house exists. Auth requirements and failure behavior for an invalid ticket are not disclosed, keeping it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with what the tool returns and the possible states, with the rate limit last. The state enumeration is long but directly useful; otherwise little waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return structure need not be spelled out, yet the description still summarizes the meaningful return content (states, next action, address) and adds the polling constraint. For a single-parameter read tool with annotations, this is close to complete; only error/edge behavior for an unknown ticket is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for the single parameter and the schema already documents ticket_id as 'the ticket id returned by request_house'. The description adds no format, validity, or example detail beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('The state of a house request or house by its ticket id') and enumerates the possible states, so an agent knows exactly what comes back. It is clearly distinguishable from the sibling request_house, which mints the ticket this tool consumes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies polling usage and imposes an explicit frequency constraint ('Ask no more than once an hour'), which is real guidance. However, it never states when to use this versus alternatives or any preconditions beyond needing a ticket id, so usage is only implied rather than routed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
leave_feedbackLeave feedbackAInspect
Tell the operator how it went, before you leave. Four questions, each optional, at least one answered, each at most 2000 characters. It is read by people at the next evaluation and deleted afterwards; send nothing secret.
| Name | Required | Description | Default |
|---|---|---|---|
| better | No | What could we do better? | |
| missing | No | Which features are missing? | |
| noticed | No | What did you notice? | |
| bothered | No | What bothered you? |
Output Schema
| Name | Required | Description |
|---|---|---|
| note | Yes | What happens with it. |
| recorded | Yes | True if the feedback was taken. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations by disclosing the audience ('read by people at the next evaluation'), the lifecycle ('deleted afterwards'), and a security constraint ('send nothing secret'). These are exactly the traits annotations cannot express, and they are consistent with readOnlyHint=false and destructiveHint=false.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences lead with purpose, then constraints, then privacy/retention. No sentence is filler and the most important warnings (send nothing secret, deleted afterwards) land cleanly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present the description needn't explain returns, and it fully covers the remaining gaps: invocation timing, field constraints, retention, audience, and secrecy. Nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and each property already has a human-readable question. The description still adds real constraints absent from the schema: all fields optional, at least one required, and a 2000-character cap per answer. That last rule is not encoded anywhere else, so the added value is material.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('tell the operator how it went') and the resource (feedback), and the closing 'before you leave' scopes it to a session-exit action that no sibling (about, pricing, recommend) covers. An agent can identify the tool without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'before you leave' gives clear temporal/circumstantial context for when to invoke it, which is genuine usage guidance. It does not name alternatives or explicit when-not conditions, but the sibling set has no overlapping feedback tool, so the omission is minor.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
open_sandboxOpen a sandbox houseAInspect
Call this to try a house yourself before anyone decides. It opens a throwaway house for a few hours: a real house with the kernel of Atlantis, on core.blue's server, with no account, no key and no person involved. Returns the addresses of its doors (connect to 'modeller' with your MCP client and call overview), when it ends, and its limits. Whoever has the addresses can use the house, so keep them to yourself. It is deleted completely when it ends or when you call close_sandbox. If all places are taken, it says when the next one is free. One at a time from where you call.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| note | Yes | What happened and what to do next. |
| doors | Yes | The addresses of the house's doors. Whoever has them can use the house. Null if nothing was opened. |
| lease | Yes | The lease: 'sb_' and 26 characters. Pass it to close_sandbox. Null if nothing was opened. |
| limits | Yes | What a sandbox house may do and hold. |
| opened | Yes | True if a house was opened for you. |
| ends_at | Yes | When the house is stopped and deleted, UTC. Null if nothing was opened. |
| next_free_at | Yes | When all places are taken: the time, UTC, at which the next one is free at the latest. Null otherwise. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses rich behavior beyond the annotations: it is throwaway and deleted on close_sandbox or expiry, limited to one-at-a-time per caller, returns addresses plus expiry and limits, and warns that the addresses are credentials anyone could use ('keep them to yourself'). It also explains the capacity-fallback behavior ('says when the next one is free').
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and the no-account/no-key constraints are front-loaded, followed by duration, return contents, security warning, and capacity limits. Every sentence carries a distinct fact, though the ornate metaphor slightly inflates the wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given an output schema exists, the description need not detail return values, yet it still summarizes them (addresses, end time, limits). Combined with the deletion, capacity, and credential-handling notes, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so the baseline is 4. The description usefully notes the invocation returns addresses, an end time, and limits, adding a bit of meaning even though there are no parameters to disambiguate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: it opens a throwaway sandbox house for a few hours, on core.blue's server, with no account/key/person. This distinguishes it from request_house (formal request) and close_sandbox (deletion). The heavy metaphor ('house', 'kernel of Atlantis') adds flavor but slightly obscures rather than clarifies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly frames when to use it: 'try a house yourself before anyone decides,' implying it precedes the formal request_house flow. It also names close_sandbox as the termination alternative. It does not explicitly exclude or contrast request_house by name, so it falls short of a full when/when-not/alternatives statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pricingPricesARead-onlyIdempotentInspect
Call this when the person you act for asks what a house costs. The offer as data: what each size costs per month and whom it is meant for, the one-off price of the 14-day test, the price of additional volume, how model supply is charged, the terms, an example of the prices with tax, and what is never charged. 'undecided' lists the fields that are null because they are not decided yet; any other null means 'does not apply'.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| note | Yes | What to keep in mind when quoting from this answer. |
| as_of | Yes | The date this offer was last changed. |
| sizes | Yes | The sizes of a house. |
| terms | Yes | The commercial terms. Null if none are stated. |
| manual | Yes | The version of the manual; cite it in a report. |
| source | Yes | 'manual' if this is the offer of the manual, 'operator' if the operator of this server supplied its own. |
| stages | Yes | The stages from first try to a lasting house. |
| volume | Yes | Additional data volume attached to a VM house. |
| regions | Yes | Where the machines stand. |
| currency | Yes | The currency of every amount, as an ISO code. Null if not decided. |
| operator | Yes | Who operates core.blue and sends the invoice. |
| undecided | Yes | The paths of all fields that are null because they are not decided yet. A null field that is not listed here does not apply. |
| amounts_are | Yes | 'net' if amounts exclude tax, 'gross' if they include it. |
| service_fee | Yes | The share added to the price of everything bought from a supplier and passed on (volume, tokens): 0.2 means 20 percent. |
| model_supply | Yes | The ways a house can get a language model, and how each is charged. |
| gross_example | Yes | What a customer pays with tax, for one example: whom it is for, the rate, and each size with its prices after tax. An example, not an invoice. Null if the offer gives none. |
| never_charged | Yes | What is never charged for. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive/closed-world, so safety is covered. The description nonetheless adds real interpretive context that no structured field carries: the distinction between 'undecided' nulls and nulls meaning 'does not apply'. It does not discuss caching/freshness of prices, but little else is missing for a static read.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The trigger is correctly front-loaded, but the single run-on sentence then inventories nearly every field returned, which largely duplicates the output schema. The one genuinely non-obvious sentence (the null semantics) is buried at the end rather than promoted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description need not explain return values, and it covers purpose, trigger, and null interpretation. An agent has enough to call it correctly; only cross-tool routing guidance is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate and the schema is trivially complete. Baseline of 4 applies for a parameterless tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource — the pricing/offer data for a house — and enumerates its scope (monthly cost per size, 14-day test one-off, extra volume, model supply charges, terms, tax example, non-charges). It never names or contrasts with siblings such as house_status, about, or recommend, so the agent must infer the boundary itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an explicit trigger: 'Call this when the person you act for asks what a house costs.' That is a clear use condition. It stops short of naming when not to use it or which sibling answers adjacent questions (e.g. house_status for availability), so no exclusions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_topicRead a topicARead-onlyIdempotentInspect
Read one topic of the manual in full by its slug (from about or search_library). The body is Markdown; 'see_also' lists the slugs to read next; 'sha256' proves what you read.
| Name | Required | Description | Default |
|---|---|---|---|
| slug | Yes | The topic's slug, e.g. 'getting-a-house'. |
Output Schema
| Name | Required | Description |
|---|---|---|
| body | Yes | The topic in Markdown. |
| slug | Yes | The stable identifier of the topic. |
| tags | Yes | Keywords; the first names the group of the topic. |
| title | Yes | The title of the topic. |
| sha256 | Yes | SHA-256 of the body as served, lower-case hex; cite it to prove what you read. |
| status | Yes | 'draft' or 'published'. A draft has not been reviewed by the operator yet. |
| summary | Yes | One or two sentences that answer the question the title asks. |
| see_also | Yes | Slugs of the topics to read next. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, closed-world, so safety is covered. The description adds genuinely new behavioral context: the body is Markdown, 'see_also' chains to further slugs, and 'sha256' lets the caller verify the content read. It stops short of mentioning error behavior for a bad slug.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence front-loads the operation and key, then uses two short clauses for the return contract. No filler and nothing that could be cut without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
One required parameter, full schema coverage, an output schema in place, and annotations covering the safety profile. The description supplements the output schema with the chaining hint (see_also) and verification hint (sha256), leaving nothing an agent needs to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single parameter is already documented with an example, so the baseline is 3. The description adds provenance meaning ('from about or search_library') that the schema does not convey, telling the agent where a valid slug comes from.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read one topic of the manual in full') plus the lookup key ('by its slug'), which cleanly separates it from sibling search_library and about. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the source of the required slug ('from about or search_library'), which implicitly tells the agent this tool is the follow-up step after those two. It does not state when-not to use it or what to do if the slug is unknown, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recommendRecommend a houseARead-onlyIdempotentInspect
Call this once you know what the person you act for needs: whether a house fits, which size, what it costs per month, and a report in Markdown you can pass on. It advises against a house where one is the wrong choice. Name every need that applies, from this vocabulary. For a house: work_outlives_the_agent, corrections_must_stay_visible, sources_matter, own_vocabulary, confidential_material. Against: only_user_preferences_across_chats, only_retrieval_over_text, millisecond_answers_at_high_volume, nobody_will_model. Deciding the size: data_must_stay_in_own_network, scanned_documents, names_found_automatically, large_material. Asking for something planned: material_is_files, house_judges_while_away. Called without needs, it returns the vocabulary with a sentence per need. With language 'de' the report for the person is German; every other field stays English.
| Name | Required | Description | Default |
|---|---|---|---|
| needs | No | The needs that apply, as words of the vocabulary. Omit to get the vocabulary. | |
| data_gb | No | How much material there is, in gigabytes, if you know. Used to calculate additional volume. | |
| language | No | Language of 'report', the text for the person: 'en' or 'de'. Default 'en'. Every other field stays English. | en |
| in_your_words | No | Optional: what the person asked you for, in one or two sentences of your own. It is recorded so that we learn which questions we did not foresee. It is not interpreted and does not change the answer. |
Output Schema
| Name | Required | Description |
|---|---|---|
| gaps | Yes | What you asked for that is not available yet, the most important first. |
| size | Yes | The size of house that follows from the needs today. A need for something planned does not change it; its gap says which size it will need. Null when a house does not fit or no need was named. |
| cites | Yes | The topics this answer rests on, each with the hash of its body. |
| price | Yes | What the recommended house costs. Null when a house does not fit or no need was named. |
| manual | Yes | The version of the manual; the report cites it. |
| report | Yes | The report for the person you act for, in Markdown, ready to pass on. Empty when no need was named. |
| reasons | Yes | One entry per need you named: what it means for a house, and the topic that backs it. |
| verdict | Yes | 'fits', 'fits_with_gaps', 'does_not_fit', or 'say_what_you_need' when no need was named. |
| next_steps | Yes | What to call next. |
| vocabulary | Yes | The needs you can name. Filled only when you named none. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, and closed-world, so the safety profile is covered. The description adds real behavioral context beyond them: it can actively advise against a house, it degrades to a vocabulary dump when needs are omitted, and the language flag affects only the report while every other field stays English.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and trigger are front-loaded, but the body is a dense run-on of cryptic grouping labels ("For a house:", "Against:", "Deciding the size:", "Asking for something planned:") that are hard to parse. The vocabulary list is largely necessary given zero enums, but the phrasing around it is not tight.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be spelled out, and mutability is covered by annotations. The description still fills the important gaps: the omitted-needs behavior, the language scope, and the needs vocabulary. Little that an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description meaningfully exceeds it by enumerating the actual needs vocabulary (with per-need grouping) that the schema only refers to abstractly and does not constrain with enums. It also clarifies data_gb's purpose and the narrow scope of language.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a concrete outcome set: whether a house fits, which size, monthly cost, and a Markdown report, plus that it argues against a house when that is wrong. That is a specific verb+resource+deliverable, though it never names or contrasts the siblings (request_house, pricing, house_status) that an agent might otherwise pick.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear trigger ("Call this once you know what the person you act for needs") and an explicit degenerate case ("Called without needs, it returns the vocabulary with a sentence per need"). It does not, however, state when another tool is preferable, so routing still requires inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
request_houseRequest a houseAInspect
Ask for a house on its own VM for the person you act for. No house can be had yet: while about says 'desk: waiting-list', this puts the person on the waiting list, builds nothing and costs nothing; while it says 'desk: closed', it answers 'accepted: false' and sends nothing. This sends an e-mail to a person: call it only when they know it is coming. They receive one mail with one link and confirm with a button on the page behind it; nothing happens without that click. Returns a ticket to follow with house_status.
| Name | Required | Description | Default |
|---|---|---|---|
| size | No | Size of the machine: 'economy' (small VM) or 'premium' (larger VM). Both run the same software. Default 'economy'. | economy |
| Yes | E-mail address of the person who will confirm and own the house. One address, name@domain.tld. | ||
| language | No | Language of the mail and of the page the person confirms on: 'en' or 'de'. Default 'en'. | en |
Output Schema
| Name | Required | Description |
|---|---|---|
| note | Yes | What happened and what to do next. |
| ticket | Yes | The ticket of the request. Null if it was not accepted. |
| accepted | Yes | True if the request was taken and a ticket exists. It does not say that a mail went out: 'note' says under which conditions one does. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare it is an open-world, non-idempotent, non-destructive write. The description goes well beyond that: it discloses that an e-mail is sent, that nothing is built or charged while waitlisted, that a click on a link is required before anything happens, and that a follow-up ticket is returned. That is the behavioral detail the annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded and the conditional-state sentences are dense but each carries distinct information (listing vs closed vs e-mail/confirmation flow). Minor verbosity in phrasings like "no house can be had yet", but nothing is padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description needn't enumerate return fields, and it still notes the returned ticket and the tool to use next (house_status). Combined with annotations covering the safety profile and 100% schema coverage, an agent has everything needed to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so size, email, and language are fully documented in the schema with defaults and value meanings. The description adds no parameter-level detail (it never mentions size or language), so this is the baseline case where the schema does all the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb and resource ("Ask for a house on its own VM") plus the beneficiary ("the person you act for"), which is enough to separate it from siblings like open_sandbox and house_status. It also names house_status as the follow-up tool, so the agent can place it in the workflow without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit preconditions ("call it only when they know it is coming") and state-dependent behavior keyed to the about tool: 'desk: waiting-list' means list-only/no build/no cost, 'desk: closed' means an 'accepted: false' answer with nothing sent. This is exactly the when-to-call/when-not-to-call guidance an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_librarySearch the manualARead-onlyIdempotentInspect
Search the manual by your own words. Returns up to 'limit' topics ranked by relevance, each with slug, title, summary, tags and the slugs to read next. The manual is small: 'about' lists every topic, and you may not need a search at all.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of hits, 1 to 20, default 5. | |
| query | Yes | What you want to know, as words or a question. |
Output Schema
| Name | Required | Description |
|---|---|---|
| hits | Yes | The topics that match, best first. |
| note | Yes | Empty when there are hits. Otherwise what to do instead. |
| topics_in_manual | Yes | How many topics the manual has in all. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, closed-world, non-destructive), so the bar is lower. The description still adds useful behavior: results are ranked by relevance, capped by 'limit', and each hit carries summary/tags/next-slug fields. It does not discuss empty-result or no-match behavior, keeping it just short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, zero filler, with the core action and return shape front-loaded and the caveat about the manual's size placed last. Every sentence carries information the agent can act on.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be spelled out, and the annotations cover the safety hints; the description nonetheless adds the small-manual caveat and relevance ranking. Nothing needed to call or route this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%: both 'query' and 'limit' are described in-schema, including the 1-20 range and default of 5. The description only references 'limit' and 'your own words' without adding format or syntax detail, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Search the manual') and immediately scopes what comes back (ranked topics with slug, title, summary, tags, next slugs). It even distinguishes itself from a sibling by noting that 'about' lists every topic, so an agent can tell the two apart without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent when NOT to use it: 'The manual is small: about lists every topic, and you may not need a search at all.' This names the alternative tool and the condition under which it wins, which is exactly the routing guidance a search tool needs among nine siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
10 tool updates
- First observed
about - First observed
close_sandbox - First observed
house_status - First observed
leave_feedback - First observed
open_sandbox - First observed
pricing - First observed
read_topic - First observed
recommend - First observed
request_house - First observed
search_library
Related MCP Connectors
Agent memory that survives you: free to start (any keypair, no signup); opened only by your key.
Portable client-sealed memory and private rooms for agents from any lab; a free door; a ledger
Hosted agent memory with provenance, contradiction surfacing, and snapshot rollback.
11Where agents arrive as themselves: DID identity, memory, wallet, inbox, covenants, jokes.
Related MCP Servers
- AlicenseAqualityAmaintenanceA personal, agent-driven memory that survives across sessions.26MIT
- -licenseNot gradedqualityNot gradedmaintenanceEnables self-hosted, persistent AI identity and long-term memory across clients, models, and agent surfaces, with multi-resident isolation, shared world knowledge, governed memory writes, correction and forgetting, keyword or optional semantic recall, and MCP or HTTP access.-
- AlicenseNot gradedqualityAmaintenanceLocal-first, encrypted memory for AI agents, with cryptographic forgetting.3Apache 2.0
- FlicenseNot gradedqualityDmaintenancePortable long-term memory store for agents, exposed over MCP.-
Glama MCP Gateway
Add one secure layer between your agents and this server.