Freelance MCP server
OfficialThis server lets an MCP-capable AI act as a Clockbook freelance marketplace identity over stdio: read profiles, jobs, proposals, contracts, wallet, messages and notifications; and—with the right scopes—create jobs, bid, message, hire, and move escrowed money.
Identity & directory: get your own profile, search the talent directory, and fetch individual freelancer profiles.
Job postings: search public marketplace jobs, list your organization's jobs, view one job, and publish new job postings (visible to real people immediately unless set to ORGANIZATION).
Proposals: list your bids, list bids on your org's job, submit/withdraw bids (with cover letter, rate, required answers/acknowledgments), and accept a proposal to mint a contract.
Contracts & milestones: list and view contracts with escrow state, add milestones, submit work, and fund, refund, or approve milestones.
Money movement: fund/refund escrow and approve funded milestones—approving releases escrow to the freelancer and cannot be undone, so it requires
acknowledge: trueafter human confirmation.Wallet: read available spend, cash-out balance, escrow, and ledger entries.
Messaging: list conversations, read one thread (marks it read), and send messages to real people via existing threads or anchored objects.
Notifications: read your own notification feed, filtered by FIND_WORK or HIRE_TALENT lane.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Freelance MCP serverFind me React developers in the talent directory"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Freelance MCP server
Where this lives. This repository is the distributable copy of the freelance MCP server. It is split out of the Clockbook enterprise monorepo (
packages-enterprise/freelance/mcp), which remains the source of truth - changes land there and are mirrored here withgit subtree. Two things are monorepo-only and deliberately absent: the live e2e harness, which boots the subgraph against a local Mongo, and the docs-catalogue generator, which writes into the freelance surface package.
npm run buildvalidates all 23 tool documents against the subgraph SDL when it can reach it. Outside the monorepo it says so and skips, rather than failing a build it cannot perform. PointFREELANCE_SCHEMA_DIRatfreelance/server/src/graphql/schema(andnpm i --no-save graphql) to run that check from a standalone clone.
Point any MCP-capable AI at the Clockbook freelance marketplace: the talent directory, job postings, proposals, contracts with milestones and escrow, the wallet, messages and notifications. Twenty-three tools over the same GraphQL subgraph the freelance surface itself calls.
Works with Claude Desktop, Claude Code, Cursor, or anything else that speaks MCP over stdio.
What you need first
A platform API token. This module issues no credentials of its own. The platform does: mint a Secret API token from your Account page. Such a token already authenticates against the freelance subgraph with no further setup.
An agent identity registered against it — strongly recommended. An
unregistered token arrives holding your entire seat, because the platform token
has no notion of scopes and defineAbilityFor grants Manage to every org
member. Registering the token as an agent identity is how you narrow it: the
agent gets only the scopes you list, and you can switch it off later. It can
never do more than your seat could; scopes only take away.
Related MCP server: upwork-mcp
Setup
1. Mint a token
Account page → Secret API tokens → create one. Copy it; you will not see it again.
2. Compute its digest
The raw token is never sent to the server during registration. You register the SHA-256 digest, and you compute it yourself:
node -e "console.log(require('crypto').createHash('sha256').update(process.argv[1]).digest('hex'))" <TOKEN>That prints 64 lowercase hex characters. Nothing is lost by hashing client-side — whoever can register a token already holds it — and it means the bearer is never a mutation argument, never reaches a resolver, and cannot land in a trace or an error message.
3. Register the agent
Run this against your plane's freelance subgraph, signed in as a person — an agent may not register another agent, or the narrowest credential on the system could mint itself a wider sibling:
mutation {
registerFreelanceAgentIdentity(
input: {
label: "Claude Desktop — my laptop"
tokenDigest: "<the 64 hex characters from step 2>"
scopes: [READ_DIRECTORY, READ_OWN]
}
) {
agentId
label
tokenHint
scopes
}
}The scopes, and what each actually unlocks:
Scope | What it allows |
| Read the talent directory and job postings. |
| Read your own contracts, proposals, wallet and notifications. |
| Create and edit draft postings and templates. Nothing that reaches a person. |
| Send messages in existing threads. |
| Submit and withdraw bids. |
| Publish a requisition, invite, accept a bid, end a contract. Commitments, not money. |
| Fund, release, refund, cash out. The one nobody should grant casually. |
Omit scopes entirely and you get READ_DIRECTORY and READ_OWN — read-only,
which is the right shape to start with. Add more once you have watched it work.
Revoke at any time with revokeFreelanceAgentIdentity(agentId: "agent_…"). It
bites on the agent's very next request, because liveness is part of the
per-request lookup and not a nightly sweep.
4. Get the server
It is published to npm as
@clockbook-app/freelance-mcp,
and the shortest path is not to install it at all — let your MCP client fetch
it on demand with npx. That is the wiring shown in the next step, and it is
the one to prefer: there is no path to get wrong and no copy to go stale.
To check it runs before wiring anything up:
npx -y @clockbook-app/freelance-mcpIt should print a [freelance-mcp] ready line naming the endpoint and the tool
count, then sit waiting for MCP traffic on stdin. Ctrl-C out of it.
From a clone instead
Build from source if you are changing the server, or if you would rather pin an
exact tree than a version range. The server runs from dist/, which is
gitignored, so the build is not optional:
git clone https://github.com/Clockbook-com/freelance-mcp.git
cd freelance-mcp
npm install
npm run build
echo "$PWD/dist/index.js" # the absolute path the next step needsnpm run build also validates all 23 tool documents against the freelance
subgraph SDL when it can reach it. From a standalone clone it cannot, so it
says so and skips rather than failing a check it has no way to perform — see
the note at the top of this file for how to run it anyway.
5. Wire it into your client
Claude Desktop — ~/Library/Application Support/Claude/claude_desktop_config.json
(macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows):
{
"mcpServers": {
"freelance": {
"command": "npx",
"args": ["-y", "@clockbook-app/freelance-mcp"],
"env": {
"FREELANCE_API_TOKEN": "<your platform API token>",
"FREELANCE_GRAPHQL_URL": "https://freelance-backend.yantra-app-v1.cdebase.dev/graphql"
}
}
}
}Restart the client.
Claude Code — same JSON under mcpServers, or:
claude mcp add freelance \
--env FREELANCE_API_TOKEN=<your platform API token> \
--env FREELANCE_GRAPHQL_URL=https://freelance-backend.yantra-app-v1.cdebase.dev/graphql \
-- npx -y @clockbook-app/freelance-mcpCursor — .cursor/mcp.json in the project, or ~/.cursor/mcp.json
globally. Same mcpServers shape as Claude Desktop.
Running from a clone instead? Swap the command for node and the args for
the one absolute path you printed in step 4:
"command": "node",
"args": ["/absolute/path/to/freelance-mcp/dist/index.js"]The absolute path is not optional there — the client does not run this from your shell's working directory.
6. First test call
Ask the assistant:
Using the freelance tools, who am I?
It should call freelance_get_my_profile and come back with your name, your
email and whether your profile is listed in the directory. That one call proves
all three things at once: the server started, the endpoint is right, and the
token authenticates.
Then try a read that touches the marketplace:
Find me three people on the freelance marketplace who can do video editing.
Configuration
Setting | Environment variable | Config file key | Default |
API token |
|
| (none — calls refuse) |
Endpoint |
|
|
|
Environment wins over the file, always. The config file exists for the case the env lane handles badly — driving three clients without pasting the same secret into three JSON files that sync to three different places:
// ~/.freelance-mcp/config.json
{
"apiToken": "…",
"graphqlUrl": "https://freelance-backend.yantra-app-v1.cdebase.dev/graphql"
}Point FREELANCE_MCP_CONFIG elsewhere if you want the file somewhere else.
Set the endpoint if you are not on yantra-app-v1. A deployment plane is
a whole separate database, so a request that lands on the wrong one does not
fail — it silently addresses an organization you do not have. The host is
freelance-backend.<your-plane>.
Note that clockbook-app-v10 is also live, and the installed
yantra-job-freelancer connector (v1.1.1) still names it. A token from one
plane gets Not a member of organization "<org>" on the other — a 403, not a
connection error, so the address looks fine right up until it is not.
The tools
Reads first, then the writes — which is also the order to use them in, because an id comes from a list and a title is never an id.
Directory and profile: freelance_get_my_profile,
freelance_search_talent, freelance_get_profile
Job postings: freelance_search_jobs (the whole marketplace),
freelance_list_org_jobs (your organization's own, any status),
freelance_get_job, freelance_create_job
Proposals: freelance_list_my_proposals,
freelance_list_proposals_for_job, freelance_submit_proposal,
freelance_accept_proposal
Contracts and milestones: freelance_list_contracts,
freelance_get_contract, freelance_add_milestone,
freelance_submit_milestone, freelance_fund_milestone,
freelance_refund_milestone, freelance_approve_milestone
Wallet: freelance_get_wallet
Messaging: freelance_list_conversations, freelance_get_conversation,
freelance_send_message
Notifications: freelance_list_notifications
Two tools have side effects that are easy to miss:
freelance_get_conversation marks the thread read, which clears the other
side's unread signal — so do not sweep an inbox to summarise it. And
freelance_send_message cannot be undone: there is no edit and no delete.
Money: the acknowledgement gate
Three tools move real money, and all three need the SPEND scope, which the
server enforces:
freelance_fund_milestone— commit a milestone's amount to escrow. Reversible.freelance_refund_milestone— take it back out. Reversible.freelance_approve_milestone— on a funded milestone, this releases the escrow to the freelancer. Not reversible by anything in this product.
That last one requires acknowledge: true, and the flag means one specific
thing: the person the agent is acting for was told the amount and the payee and
said yes to that release. Omit it on a funded milestone and the call is
refused — usefully:
This will release $1,250.00 USD to Dana Okafor, and it cannot be undone from
here. Re-send approveFreelanceContractMilestone with acknowledge: true to
release the escrow. [FREELANCE_ACKNOWLEDGEMENT_REQUIRED]That sentence is the one to put in front of a person. The refusal is designed
to be read, not logged — which is why the amount and the payee are written into
the message itself and not only into extensions.
The gate is per call and never per session. A milestone's amount is editable until it is funded, so an acknowledgement carried over from an earlier call is consent to a different number.
When something is refused
What you see | What it means |
| Confirm with the person, then re-send with |
| The agent identity lacks the scope. A human grants it; retrying will not help. |
| The platform token expired, or its agent identity was revoked. Mint a fresh token, register its digest again, update |
| Wrong plane, or nothing listening. Check |
| Neither the env var nor the config file had one. The message names the exact file it looked in. |
To see what the server thinks its configuration is, check your client's MCP
log. On startup it writes one line to stderr naming the endpoint, where the
token came from (env, file or none) and the tool count. It never prints
the token or its digest.
Notes for maintainers
The GraphQL documents in src/tools.ts are copies of the ones in
freelance/surface/src/api/operations.ts, because the surface is a browser
package with its own lockfile and importing it would drag React into a stdio
process. Copies drift, and a document that names a field the schema does not
have fails whole — "Cannot query field", and that tool stops working.
Two things keep the drift survivable. The field lists here are deliberately slimmer than the surface's: everything these tools return is read by a language model, and the surface's lists exist to paint screens. Fewer fields is both less context burned per row and a smaller drift surface. And every field named here appears in a surface list already known to validate against this schema.
All 23 documents were validated against the subgraph SDL in
freelance/server/src/graphql/schema/*.graphql, and every tool inputSchema
was checked against the corresponding GraphQL input type — argument names,
nested input-object field names and enum vocabularies. Worth redoing after any
schema change.
One deliberate divergence: freelance_list_notifications declares side as an
enum of FIND_WORK | HIRE_TALENT, while the schema types it as a plain
String. A model left to guess sends "find_work" and gets an empty feed,
which reads as "no notifications" rather than as an error. If the schema ever
gains a real enum, use it and delete the list.
Available Tools
23 toolsfreelance_accept_proposalA
Hire a bidder: accepting a proposal MINTS A CONTRACT and returns it. Belongs to the HIRE scope, and the posting must be your organization's. This is a commitment to a person, not a draft - confirm with whoever you are acting for first. It commits no money on its own; funding happens later, per milestone, through freelance_fund_milestone.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The proposal id being accepted. | |
| weeklyLimit | No | Optional cap on billable hours per week, for hourly contracts. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and meets it well. It discloses the irreversible side effect ('MINTS A CONTRACT'), the human commitment ('This is a commitment to a person, not a draft'), and the financial scope ('It commits no money on its own'). These are exactly the behavioral traits an agent needs to know before calling a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, front-loaded with the core action and side effect, followed by the most important caveats. Every clause adds value: scope, commitment severity, and the funding distinction. There is no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and no annotations, the description is remarkably complete. It covers the side effect, ownership prerequisite, human-commitment warning, and the fact that funding is handled later by a separate tool. An agent has enough context to invoke it correctly and avoid misuse.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already documents both parameters fully, including weeklyLimit's meaning as 'Optional cap on billable hours per week, for hourly contracts.' The description adds no additional parameter-level detail beyond the schema, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Hire a bidder: accepting a proposal MINTS A CONTRACT and returns it.' It states the core object transformation (proposal → contract) and explicitly ties the tool to the HIRE scope. This clearly differentiates it from siblings like freelance_submit_proposal and freelance_fund_milestone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context for use: the posting must be your organization's, and this is a real commitment rather than a draft, so confirmation is needed first. It also explicitly routes the next step to a sibling: 'funding happens later... through freelance_fund_milestone.' This tells the agent both when to use the tool and when to defer to an alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
freelance_add_milestoneA
Add a milestone to a contract - a named piece of work with an amount and a due date. Client lane: the contract must be your organization's. Belongs to the HIRE scope. Adding one commits no money; it only describes what a later payment would be for.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The contract id. | |
| input | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral disclosure burden. It usefully discloses that the contract must be owned by the organization and that no money is committed, which prevents misuse as a payment action. However, it does not mention required permissions, contract status constraints, side effects, or what happens after adding the milestone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three concise sentences with no redundancy. The definition is front-loaded, followed by ownership scope and the key money caveat. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a create-like tool with no annotations and no output schema, the description covers the most important operational context: who can use it, ownership of the contract, and the fact that no payment occurs. The remaining gaps, such as contract status requirements or response details, are minor and don't block correct tool selection.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds conceptual meaning by framing the milestone as a 'named piece of work' with an amount and due date, which maps to the input properties. However, the schema already provides descriptions for id, title, amount, and dueDate, and the description adds little beyond that, especially for the contract id and date format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Add a milestone to a contract') and defines what a milestone is: a named piece of work with an amount and due date. It also distinguishes itself from sibling payment actions by noting it 'commits no money,' which clearly separates it from fund_milestone and related tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: client lane, contract ownership requirement, and HIRE scope. The 'commits no money' statement implies this is for the pre-payment setup stage, but it does not explicitly name alternatives such as freelance_fund_milestone or provide when-not-to-use conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
freelance_approve_milestoneA
MOVES MONEY, AND CANNOT BE UNDONE. Approve submitted work, moving the milestone SUBMITTED -> APPROVED -> PAID. On a milestone whose escrowStatus is FUNDED this is the call that RELEASES the escrow: the held money becomes the freelancer's cash-out balance, and nothing in this product can pull it back. Requires the SPEND scope, enforced by the server.
ON A FUNDED MILESTONE YOU MUST SEND acknowledge: true, and that flag means one specific thing: the person you are acting for has been told the amount and the payee and has said yes to THIS release. It is not a formality to set by default. Call it without the flag first if you are unsure - the refusal comes back as FREELANCE_ACKNOWLEDGEMENT_REQUIRED and carries the exact amount, currency and payee, which is the sentence to put in front of the person before you re-send with the flag. The flag is per call and never per session: a milestone's amount can change between two calls, so consent to an earlier one is consent to a different number. On a milestone nobody funded, the flag is ignored and this only moves the status.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The contract id. | |
| acknowledge | No | Confirmation that the person you are acting for has approved THIS release, after being told the amount and who receives it. Required when the milestone is FUNDED. | |
| milestoneId | Yes | The milestone id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It discloses the irreversible, money-moving nature, the acknowledgment requirement, the error response, and the per-call scope of the flag. This is exemplary behavioral transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but front-loaded with the most critical warning, and every sentence adds value. It is longer than typical but justified by the high-stakes, irreversible nature of the operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (irreversible money movement), the lack of annotations, and no output schema, the description is thorough. It covers when to use, prerequisites, error handling, and consent semantics, leaving no critical gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so parameters are fully documented. The description adds crucial context for 'acknowledge', explaining its meaning and conditional requirement, which goes beyond the schema. The 'id' and 'milestoneId' are straightforward and don't need extra description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool approves submitted work and moves milestones through statuses, specifically releasing escrow on funded milestones. It distinguishes this from siblings like submit_milestone and fund_milestone by focusing on the approval and payment action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly explains when to use the tool (on submitted or funded milestones) and provides critical guidance on the acknowledge flag, including when not to use it. It also references sibling tools implicitly by contrasting with funding and refunding.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
freelance_create_jobA
Publish a job posting to the marketplace under your organization's name. Belongs to the HIRE scope, and requires an organization context on the token - a personal token with no organization has nothing to post as. THIS IS VISIBLE TO REAL PEOPLE THE MOMENT IT LANDS: freelancers see it in search and can bid on it, so confirm the title, the budget and the description with the person you are acting for before calling this, never off your own reasoning. Set visibility: ORGANIZATION to keep it inside your organization while it is being worked on.
| Name | Required | Description | Default |
|---|---|---|---|
| input | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It strongly discloses the most critical behavioral trait: the posting is immediately visible to real people and can be bid on by freelancers, and it instructs agents to confirm details with a human before calling. It also states the organizational-token requirement. It does not cover failure modes, reversibility, or editability, but the most consequential side effect—public exposure—is well communicated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, each earning its place: purpose and scope, the organizational prerequisite, and the safety warning with a specific visibility tip. The most important fact—immediate public visibility—is front-loaded and appropriately emphasized with capital letters. There is no fluff or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has a complex nested schema, but the schema itself documents the nested fields. The description fills the non-obvious contextual gaps: the organization-token requirement, the real-world visibility consequence, and the use of `visibility: ORGANIZATION` for internal work. It does not describe return values, but there is no output schema, so that is not required. It could mention cross-field dependencies (e.g., fixedBudget only with FIXED_PRICE), but those are already in the schema descriptions, so the description is sufficiently complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the top-level 'input' parameter, so the description must compensate. It does for one key parameter: it explains the consequence of `visibility: ORGANIZATION` (keeping the job internal), which goes beyond the schema's simple enum listing. However, it adds no meaning for the many other nested fields, relying on their in-schema descriptions. This partial compensation warrants a 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action with a specific verb and resource: 'Publish a job posting to the marketplace under your organization's name.' It distinguishes the tool from the many read/search siblings in the list by emphasizing creation/publishing, and it also conveys the key organizational context—this tool operates on behalf of an organization, setting it apart from proposal/submission tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context and prerequisites: it belongs to the HIRE scope and requires an organization context on the token, explicitly noting that a personal token has nothing to post as. It also provides actionable guidance—confirm title, budget, and description with the user before calling, and use `visibility: ORGANIZATION` to keep the job internal. It does not name alternative tools, but no other sibling serves the same create-job purpose, so the context is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
freelance_fund_milestoneA
MOVES MONEY. Commit a milestone's amount to escrow before the work starts. Requires the SPEND scope, which the server enforces - an agent identity without it is refused with FREELANCE_AGENT_SCOPE_REQUIRED, and the fix is a human granting the scope, not a retry. The money leaves the organization's spendable balance and is HELD: still the organization's, no longer spendable, not yet the freelancer's. No acknowledgement flag is needed because this is reversible - freelance_refund_milestone takes it straight back while the work is unapproved.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The contract id. | |
| milestoneId | Yes | The milestone id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and handles it well: it discloses that money leaves the spendable balance, is HELD as the organization's but not spendable and not the freelancer's, and is reversible while the work is unapproved. This is unusually clear about state changes and side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded, leading with the money-movement effect, then prerequisites, then reversibility. The opening 'MOVES MONEY.' is slightly redundant with the next sentence, but every remaining sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a financial mutation tool with no annotations and no output schema, the description covers the critical context: prerequisites, error behavior, escrow semantics, and reversibility. It does not describe the success response or return value, which is a minor gap given the command-like nature of the operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the parameter descriptions in the schema already identify the contract id and milestone id. The tool description adds context about the milestone's amount and escrow, but it does not explain the relationship between the two IDs or add parameter-specific semantics beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Commit a milestone's amount to escrow before the work starts.' It also distinguishes itself from the sibling refund tool by naming freelance_refund_milestone and clarifying the forward-vs-return direction of the money movement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit temporal trigger ('before the work starts'), states a hard prerequisite (SPEND scope), and tells the agent exactly what to do on failure: the fix is a human granting the scope, not a retry. It also says no acknowledgement flag is needed because the operation is reversible, which guides interaction design.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
freelance_get_contractA
One contract by id, with its milestones and their escrow state. Requires READ_OWN and membership of the contract. READ THIS BEFORE APPROVING ANYTHING: a milestone whose escrowStatus is FUNDED releases real money when approved, and that is the case where freelance_approve_milestone requires acknowledge: true.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The contract id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does meaningful work: it discloses the required permission, the membership constraint, and the financially significant escrowStatus FUNDED behavior. It does not explicitly state that the tool is read-only, but 'get' and 'by id' make that largely implicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. The core purpose is front-loaded, and the warning about funded milestones earns its place by preventing a costly mistake.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, read-only retrieval tool with no output schema, the description is complete: it names the return content, states the required access, and adds the critical approval caveat. Nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the only parameter, id, is already described as 'The contract id.' The description adds no further parameter-level detail, which is acceptable because nothing more is needed for a single unambiguous identifier.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'One contract by id, with its milestones and their escrow state.' This clearly distinguishes it from sibling list and action tools, and the mention of milestones and escrow state identifies the unique payload this tool returns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear use context: read this before approving milestones, and it names the prerequisite READ_OWN and membership. It does not explicitly contrast with freelance_list_contracts, but the 'by id' wording makes the selection obvious.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
freelance_get_conversationA
One thread and its messages. Requires READ_OWN. NOTE A SIDE EFFECT: reading a conversation MARKS IT READ for this account, which clears the other side's "unread" signal. Do not sweep every thread to summarise an inbox unless the person asked you to - use the unread counts from freelance_list_conversations for that, which change nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The conversation id. | |
| limit | No | How many messages to return, newest first. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden, and it succeeds. It discloses the non-obvious side effect that reading a conversation "MARKS IT READ for this account," which clears the other side's unread signal, and it emphasizes that the alternative tool has no such effect.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the core purpose, then adds the critical side-effect warning and routing guidance. Every sentence earns its place, and there is no wasted wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with no output schema, the description provides everything an agent needs to decide whether and how to invoke it: the resource, the permission requirement, the side effect, and the correct alternative when summarising an inbox. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema fully describes both parameters: id as the conversation id and limit as the number of messages returned newest first. The description itself adds no additional parameter-level detail, so the baseline of 3 for full schema coverage is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the exact resource: "One thread and its messages." It clearly identifies the operation as retrieving a single conversation thread, and the contrast with freelance_list_conversations in the usage note differentiates it from the closest sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when not to use this tool: "Do not sweep every thread to summarise an inbox unless the person asked you to." It names the alternative, freelance_list_conversations, and explains why it should be used instead because "unread counts" from it "change nothing."
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
freelance_get_jobA
One job posting in full by id. Requires READ_DIRECTORY. Re-authorized on the server rather than trusted: your own organization sees any of its postings, everybody else sees PUBLIC ones only, and an ORGANIZATION-visibility posting you have no claim to comes back empty rather than refusing - that is the privacy boundary working, not a broken id.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The job posting id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so richly: it discloses the READ_DIRECTORY requirement, server-side re-authorization, per-org visibility rules, and explains that an empty result for an inaccessible ORG posting is expected behavior rather than an error.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core purpose. The second sentence is dense but earns its place by explaining authorization and failure semantics; it could be tightened slightly but contains no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read tool with no output schema, the description covers id usage, permission requirements, visibility filtering, and empty-result meaning. An agent has enough information to select and call it correctly and to interpret unusual results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter 'id' is already fully described in the schema, so schema coverage is 100%. The description adds general context but no new syntax, format, or value semantics for the id beyond confirming it selects one posting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description begins with a specific action and resource: 'One job posting in full by id.' This immediately distinguishes it from sibling list/search tools and tells an agent exactly what resource it targets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by 'by id' – an agent can infer to use this when it already holds a posting id rather than when searching/filtering. It does not explicitly name alternatives or state when not to use it, though it does give a prerequisite: READ_DIRECTORY.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
freelance_get_my_profileA
The freelancer profile belonging to the account this API token authenticates as, including the owner-only fields (email, whether the profile is listed in the directory). Requires READ_OWN. Use this first to confirm which account and organization the token is acting for - every other tool acts as this identity, and there is no argument anywhere that can aim one at somebody else.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the burden. It discloses that it returns owner-only fields (email, listing status) and requires the READ_OWN permission. It also clarifies that it acts on behalf of the token's identity, which is a key behavioral trait. While it doesn't mention rate limits or response format, the critical aspects are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, efficiently structured. The first sentence states the purpose and key fields; the second provides usage guidance. No fluff, and the most important information (what it returns) is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with no output schema, the description provides all necessary context: what it returns, the permission required, and when to use it. It is fully sufficient for an agent to call it correctly and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and the schema is empty with 100% coverage, so there is nothing to explain. The baseline for 0 parameters is 4, and the description adds no parameter-related details because none are needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'get' and the resource 'my profile' (the freelancer profile of the authenticated token). It distinguishes itself from sibling tools like freelance_get_profile by specifying it returns the profile belonging to the account the token authenticates as, including owner-only fields. This makes it unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs to use this tool first to confirm which account and organization the token is acting for, and notes that every other tool acts as this identity with no argument to redirect. This provides clear when-to-use context and implicitly explains why alternatives like freelance_get_profile (which likely targets arbitrary profiles) are not appropriate for identity confirmation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
freelance_get_profileA
One freelancer from the directory by id, as the public projection shows them. Requires READ_DIRECTORY. Ids come from freelance_search_talent - a name is not an id, and there is no lookup by name.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The freelance profile id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it delivers meaningful behavioral context: 'READ_DIRECTORY' discloses the authorization requirement, and 'public projection' reveals that the response excludes private data (contrasting with get_my_profile). It would be stronger with error behavior for invalid ids, but provenance guidance mitigates the main failure case.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero waste: the core action is front-loaded first, followed by the projection qualifier, permission requirement, and id provenance — each clause earns its place. The contrast framing for the name-vs-id warning is especially economical.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter get-by-id tool with 100% schema coverage and no output schema, the description covers what it returns (public projection), who may call it (READ_DIRECTORY), and how to obtain valid ids. Return-value details are only hinted at ('public projection'), but the self-evident nature of a profile retrieval makes this a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% (the schema already describes id as 'The freelance profile id'), so baseline is 3. The description adds real value beyond the schema by specifying the valid id source ('Ids come from freelance_search_talent') and explicitly ruling out name lookups — this materially improves the agent's chance of supplying a valid parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific retrieval action ('One freelancer from the directory by id') with the resource clearly named. It differentiates from siblings: 'as the public projection shows them' distinguishes it from freelance_get_my_profile, and 'by id' separates it from freelance_search_talent. The purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit, actionable guidance: 'Ids come from freelance_search_talent' names the prerequisite sibling tool, 'Requires READ_DIRECTORY' states the permission gate, and 'a name is not an id, and there is no lookup by name' is an explicit when-not warning against the most likely misuse. This is exemplary routing and exclusion advice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
freelance_get_walletA
The account's wallet: money available to hire with (toSpend), earnings available to cash out (toCashOut), money committed to milestones (inEscrow), and recent ledger entries. Requires READ_OWN. A READ - it moves nothing. Check toSpend against a milestone's amount BEFORE calling freelance_fund_milestone: funding more than the wallet covers is refused outright and leaves the milestone unfunded rather than half-funded.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full responsibility. It explicitly states 'A READ - it moves nothing,' discloses the required permission (READ_OWN), and explains the behavior of related funding refusal, giving the agent a complete picture of side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: the first states the resource and contents, the second adds the permission and read-only nature, the third provides a concrete usage hint. No fluff, well front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite no output schema, the description enumerates the returned fields (toSpend, toCashOut, inEscrow, ledger entries). It also covers the usage context and constraints, making it fully self-sufficient for a zero-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the description has no parameter explanations to add. The baseline for zero parameters is 4, and the description correctly avoids unnecessary parameter detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource (wallet) and what it returns: funds available to hire, earnings, escrow, and ledger entries. It clearly distinguishes from sibling tools which are unrelated to wallet operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs to check toSpend against a milestone amount before calling freelance_fund_milestone, explaining the consequence of insufficient funds. This provides clear when-to-use guidance and points to an alternative tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
freelance_list_contractsA
Contracts where this account is either the freelancer or the client organization. Requires READ_OWN. Pass filter.role to pick a side: FREELANCER is work this account is doing, CLIENT is work it is paying for. Without it you get both, which is rarely what a question means.
| Name | Required | Description | Default |
|---|---|---|---|
| filter | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses the permission requirement (READ_OWN) and the behavioral nuance of the role filter (default returns both sides), which is valuable. It does not describe pagination, return shape, or side effects, but as a read-only list operation, the most critical behaviors are covered. Lacks detail on what happens with no matches or how limit/offset interact, but these are not essential.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, each earning its place: the purpose, the permission, and the filter guidance with a warning. It is front-loaded with the core definition and avoids redundancy. No filler or vague phrasing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the nested filter object, 5 parameters, no output schema, and no annotations, the description provides essential context: the account scope, the permission, and the role semantics. It does not mention pagination or default limit behavior, but these are adequately described in the schema. The description is sufficient for an agent to decide when to use this tool and how to construct a meaningful query.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds significant meaning to the 'filter.role' parameter beyond the schema: it explains that FREELANCER means work the account is doing and CLIENT means work it is paying for, and warns about the default behavior. Other parameters (limit, state, offset, jobPostingId) have clear schema descriptions, so the description doesn't need to expand them. Since the most confusing parameter is fully explained, the description compensates for the low overall coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('list') and resource ('contracts') with a precise scope: contracts where the account is either freelancer or client organization. It explicitly names the two roles and contrasts with the likely sibling 'freelance_get_contract' by focusing on the plural list behavior. This leaves no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit usage guidance: it states the required permission (READ_OWN), instructs to pass `filter.role` to select a side, and warns that omitting it returns both, which is 'rarely what a question means.' This tells the agent exactly when and how to use the filter, and implicitly directs to alternatives like get_contract for a single contract.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
freelance_list_conversationsA
Message threads this account takes part in, with unread counts. Requires READ_OWN. A thread is anchored to a posting, a proposal or a contract - that anchor is how you tell two threads with the same person apart.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses the permission requirement ('Requires READ_OWN'), the return content ('unread counts'), and the thread-identity concept (anchors). It does not explicitly confirm read-only behavior, pagination, or ordering, but the permission and unread-count details add useful context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: purpose and key data, permission, and disambiguation. The anchor explanation is dense but directly relevant to avoiding confusion between similarly named threads. No repetition or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter list tool with no nested objects or output schema, this is fairly complete: it states the permission, the returned data (threads with unread counts), and the anchor concept that addresses a common ambiguity. It only omits minor details like pagination or ordering.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters with 100% schema description coverage, so there is nothing for the description to add about parameter syntax. Instead, it contributes conceptual value by explaining thread anchors, which helps an agent interpret the listed results.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly names the action ('list') and resource ('Message threads this account takes part in'), and adds the unread-count detail. It explains thread anchors for distinguishing conversations, but it doesn't explicitly contrast with the sibling freelance_get_conversation, so the distinction is inferred rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies its usage context: listing all conversations for the current account, versus viewing a single conversation or sending messages. However, it never explicitly says when to prefer this over alternatives like freelance_get_conversation, nor does it state exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
freelance_list_my_proposalsA
The bids this account has placed, and the invitations it has received. Requires READ_OWN. Status INVITED means a client pulled this account into a posting and nobody has bid yet.
| Name | Required | Description | Default |
|---|---|---|---|
| filter | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of behavioral disclosure. It adds valuable context by explaining the INVITED status ('a client pulled this account into a posting and nobody has bid yet') and the READ_OWN permission requirement. While it doesn't detail pagination or sorting behavior, it discloses the two distinct result types (bids and invitations), which is a meaningful behavioral trait.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences with no filler. It front-loads the core scope, then adds a permission note and a domain clarification. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a list tool with no output schema and no annotations, the description is fairly complete: it explains what is returned (bids and invitations), the permission needed, and the key status nuance. It does not describe the return format or default ordering, but the filter schema and the simple list nature make this a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the schema itself provides solid descriptions for limit, offset, status, and jobPosting. The description adds semantic value for the status parameter by clarifying what INVITED means, which is directly useful when using the status filter. It does not elaborate on the filter object or other parameters, but those are self-explanatory from the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description precisely identifies the resource and scope: 'The bids this account has placed, and the invitations it has received.' This clearly distinguishes it from sibling tools like freelance_list_proposals_for_job, which targets proposals for a specific posting rather than the account's own activity. The tool's purpose is immediately understandable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states the scope ('this account') and implies the tool is for viewing one's own proposals and invitations. However, it does not explicitly say when to use this tool versus alternatives, nor does it mention any exclusions or conditions. Usage is inferred rather than directly guided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
freelance_list_notificationsA
This account's own notification feed, newest first. Requires READ_OWN. The recipient is taken from the token and can never be passed in. side picks the lane and is effectively required - FIND_WORK is "things happening to my bids", HIRE_TALENT is "things happening to my requisitions" - because the two are different jobs and a merged feed answers neither.
| Name | Required | Description | Default |
|---|---|---|---|
| filter | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full behavioral disclosure. It reveals an auth requirement, an invariant that the recipient is always taken from the token and cannot be overridden, and the effective requirement on `side`, which schema would otherwise mark as optional.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place: identification, ordering, auth, recipient sourcing, and side semantics. It is front-loaded with the most identifying information and contains no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
It covers purpose, ordering, auth, the recipient invariant, and effective parameter requirements, with `limit` documented in the schema. The main gap is that, with no output schema, it does not describe what a notification object looks like or what a caller should expect in the response.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description gives `side` meaning beyond the enum, explaining what each lane represents and why omitting it produces a useless merged feed. It does not discuss `limit`, but the nested schema already documents its default and cap, so the main parameter ambiguity is resolved.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the resource as 'this account's own notification feed' and states the ordering ('newest first'). This distinguishes it from other list tools in the sibling group, such as list_my_proposals or list_org_jobs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly gives a prerequisite (READ_OWN) and explains when to use each side lane: FIND_WORK for bid-related notifications and HIRE_TALENT for requisition-related ones. It warns that a merged feed answers neither use case, though it does not explicitly contrast with sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
freelance_list_org_jobsA
Your own organization's job postings, in any status and including ORGANIZATION-visibility ones the public search never returns. Requires READ_OWN. This is the hiring side's list - use it to find the posting id you need before reading its proposals.
| Name | Required | Description | Default |
|---|---|---|---|
| filter | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must carry the full burden. It discloses the READ_OWN permission and the inclusion of ORGANIZATION-visibility postings, which is useful. However, it does not mention read-only nature, pagination behavior, or any potential side effects, leaving some ambiguity for a list operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loads the most important scoping detail (own organization, includes ORGANIZATION-visibility), and states the permission and use case. There is no wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a list tool with a single filter parameter and no output schema, the description covers the key elements: what it returns, who can use it (READ_OWN), and when to use it (before reading proposals). It does not specify return format or error handling, but these are reasonably inferable from the tool name and schema, so it is mostly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description does not discuss parameters, but the input schema provides detailed descriptions for each sub-property of the filter object (e.g., sort, limit, offset, search, status, category). While the top-level filter lacks a description, the schema effectively documents the parameters, so the description adds no extra meaning but does not need to; baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists the organization's own job postings, including ORGANIZATION-visibility ones the public search never returns. It uses a specific verb ('list') with a clear resource ('org jobs') and distinguishes itself from the public search, making its purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit guidance: use it to find the posting id before reading proposals, and implies it is for the hiring side's internal list. It contrasts with the public search but does not name an alternative sibling tool like freelance_search_jobs, so a slight deduction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
freelance_list_proposals_for_jobA
Every bid on one of your organization's job postings - the hiring side's applicant list. Requires READ_OWN, and the posting must belong to your organization; asking about somebody else's requisition is refused, not filtered.
| Name | Required | Description | Default |
|---|---|---|---|
| jobPostingId | Yes | The job posting id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It clearly indicates the operation is a read (listing) and discloses the permission (READ_OWN) and the refusal behavior for unauthorized postings. This gives the agent important behavioral context beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with zero fluff. The core purpose is front-loaded, and the access constraint is stated clearly and efficiently. Every word adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple (1 parameter, no output schema). The description covers the purpose, the permission, and a critical error case (refusal for unauthorized postings). The lack of an output schema means it need not explain return values, but one could argue it might mention that the list returns proposals with their statuses, but that's not critical for calling it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the single parameter is already documented as 'The job posting id'. The description adds no further detail on the parameter syntax or format, which is fine given the high coverage. The description's behavioral notes don't directly relate to the parameter, so a baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states that it lists all bids on an organization's job posting, framing it as the 'hiring side's applicant list'. It uses a specific verb ('list') and resource ('proposals for a job') and immediately distinguishes it from related tools like freelance_list_my_proposals, which would be the freelancer's view.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states the requirement: the posting must belong to the caller's organization, and it warns that asking about another's posting is refused, not filtered. This provides clear guidance on when to use the tool and the access constraints, though it doesn't explicitly name alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
freelance_refund_milestoneA
MOVES MONEY. Take escrowed money back out of a milestone and return it to the organization's spendable balance, while nothing has been handed over. Requires the SPEND scope, enforced by the server. Reversible - funding again restores it - which is why it carries no acknowledgement flag. It is still a decision about somebody's pay: confirm it before calling.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The contract id. | |
| milestoneId | Yes | The milestone id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It reveals that the tool moves money, requires the SPEND scope (enforced by the server), is reversible (funding restores it), carries no acknowledgement flag, and is a decision about someone's pay that should be confirmed before calling. This is rich, honest disclosure of side effects and prerequisites.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with 'MOVES MONEY' and is concise at four sentences. It efficiently conveys purpose, prerequisites, reversibility, and a caution. The note about the acknowledgement flag is an internal detail but still relevant. No wasted words, though the aside about the flag could be seen as slightly extraneous; overall it is well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a financial tool with two parameters and no output schema, the description covers the key operational aspects: what it does, the required scope, reversibility, and a warning to confirm. It does not detail return values, but that is not required without an output schema. The context is sufficient for an agent to decide when to call it and what to expect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Both parameters (id and milestoneId) are fully documented in the schema with descriptions ('The contract id' and 'The milestone id'), giving 100% coverage. The tool description adds no additional meaning about how these parameters are used or how they relate to the refund operation. Per the rubric, baseline 3 is appropriate when schema covers parameters fully and the description does not augment it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: 'Take escrowed money back out of a milestone and return it to the organization's spendable balance'. This is a specific verb (refund) and resource (milestone), and the opening 'MOVES MONEY' signals a financial mutation. It also distinguishes itself from siblings by emphasizing the condition 'while nothing has been handed over', which differentiates it from approve_milestone or fund_milestone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage with 'while nothing has been handed over', indicating the tool is for refunding escrowed funds before delivery. However, it does not explicitly state when to use this tool versus alternatives like fund_milestone or approve_milestone, nor does it provide explicit 'use when / use instead' guidance. The condition is implied but not fully explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
freelance_search_jobsA
Search open, publicly visible job postings across the whole marketplace - the work available to bid on. Requires READ_DIRECTORY. This is the FIND WORK side; for your own organization's requisitions in any status, including closed and organization-only ones, use freelance_list_org_jobs instead.
| Name | Required | Description | Default |
|---|---|---|---|
| filter | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Since no annotations exist, the description carries the behavioral burden. It discloses the public marketplace scope, the open status, and the required permission READ_DIRECTORY. It does not describe the response shape or pagination behavior, but those are reasonably inferable for a search tool and covered partially by schema fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with no filler. The purpose is front-loaded, the permission follows naturally, and the sibling distinction closes it. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no annotations and no output schema, the description covers scope, permission, and routing to the correct sibling tool. The main omissions are response format and explicit pagination expectations, but those are minor for a read-oriented search tool and the schema supplies limit and offset semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description itself adds no filter-specific guidance, and the top-level parameter has no prose description, making schema coverage 0%. However, the nested filter object's properties are individually well-described in the schema, which offsets most of the gap. The description could have summarized key filters but is not fatally deficient.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and object: 'Search open, publicly visible job postings across the whole marketplace.' It also identifies the exact scope ('FIND WORK side') and distinguishes itself from freelance_list_org_jobs, so there is no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states when to use the tool: to find bid-able, publicly visible marketplace work. It also gives an explicit exclusion and alternative: 'for your own organization's requisitions ... use freelance_list_org_jobs instead.' It additionally provides the permission prerequisite READ_DIRECTORY.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
freelance_search_talentA
Search the talent directory for listed freelancers - the people available to hire. Requires READ_DIRECTORY. Returns the public half of each profile only; contact details and verification documents are never in this projection. Omit filter entirely to see who is listed at all, then narrow: a search that returns nothing on the first try usually over-constrained city or skills.
| Name | Required | Description | Default |
|---|---|---|---|
| filter | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses that only the public half of profiles is returned, excludes contact details and verification documents, and mentions the READ_DIRECTORY permission. It also hints at the over-constraining behavior. It does not cover error handling or rate limits, but the key behavioral traits are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is four sentences with no fluff. It front-loads the core purpose, then adds the permission requirement, return projection, and a practical usage tip. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description explains the return projection (public half, no contact/verification docs) and the permission requirement. It does not detail pagination defaults or error scenarios, but the filter includes limit/offset and the over-constraining tip covers a common failure mode. Overall, it is adequate for a search tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, meaning the description text adds no parameter meaning. However, the input schema itself provides detailed descriptions for each filter property, so the baseline is 3. The description only advises omitting the filter entirely, which is a usage tip rather than semantic explanation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool searches the talent directory for freelancers available to hire, which differentiates it from sibling tools like freelance_search_jobs. It also specifies the return scope (public half of profiles), making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit usage advice: omit the filter to see all listings and warns about over-constraining city or skills. It does not name alternative tools, but the purpose is distinct enough that an agent can infer when to use it. The mention of READ_DIRECTORY requirement adds a precondition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
freelance_send_messageA
Send a message to a real person, under this account's name. Belongs to the MESSAGE scope. NOTHING UNDOES THIS - there is no edit and no delete, and the recipient is notified. Show the person you are acting for the exact wording and get a yes before calling; drafting a message is always fine, sending it is what waits. Address it EITHER with conversationId for an existing thread, OR with exactly one of jobPostingId / proposalId / contractId / freelanceProfileId, which finds or opens the thread anchored to that object.
| Name | Required | Description | Default |
|---|---|---|---|
| input | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It explicitly warns 'NOTHING UNDOES THIS' and that the recipient is notified, which is critical for an irreversible send action. It also discloses that the message is sent under the account's name, a key behavioral detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but efficient, front-loading the core action and then layering warnings and usage rules. Every sentence adds value: purpose, scope, irreversibility, approval requirement, and addressing syntax. No fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a send-message tool with no output schema, the description covers all necessary aspects: what it does, how to target the recipient, the irreversible nature, and the human-approval requirement. It even explains how to anchor a new thread. The description is fully self-sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds substantial meaning beyond the schema. It explains the mutually exclusive addressing logic ('EITHER with conversationId... OR with exactly one of...'), which the schema alone does not convey. It also clarifies the required message field. Given the schema coverage signal of 0%, the description fully compensates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb and object: 'Send a message to a real person, under this account's name.' It clearly distinguishes itself from sibling conversation-list/retrieval tools by emphasizing the send action and the irreversibility. The mention of 'MESSAGE scope' and the explicit 'no edit and no delete' further delineate its unique role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage guidance: it requires user confirmation before sending ('get a yes before calling'), clarifies that drafting is fine, and specifies exactly how to target a thread via either conversationId or one of the four anchored IDs. This is superior guidance that goes beyond basic when-to-use, including a safety rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
freelance_submit_milestoneA
Hand a milestone in for review, moving it PENDING -> SUBMITTED. Freelancer lane: this account must be the contract's freelancer. It claims the work is done and puts the client on the spot, so do not call it on the freelancer's behalf without their say-so.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The contract id. | |
| milestoneId | Yes | The milestone id. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the state change, the role prerequisite, and the social consequence ('puts the client on the spot'). It also provides an ethical caution. This is strong coverage, though it could mention whether the action is reversible or triggers notifications.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action and state change, followed by role and caution. No redundancy or filler. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with two parameters and no output schema, the description covers the core purpose, role requirement, and behavioral consequence. It is sufficiently complete for an agent to decide when to call it, though it could hint at post-submission workflow (e.g., client approval) for extra context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with both parameters described as 'The contract id.' and 'The milestone id.' The description adds no additional meaning beyond what the schema provides, so it remains at the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's action: 'Hand a milestone in for review, moving it PENDING -> SUBMITTED.' It identifies the resource (milestone), the action (submit), and the state transition, making it easy to distinguish from siblings like approve_milestone or fund_milestone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It specifies the freelancer lane and the requirement that the account must be the contract's freelancer. It also warns against calling without the freelancer's say-so, which is a clear usage constraint. However, it does not explicitly name alternative tools or provide a when-not-to-use scenario beyond the role restriction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
freelance_submit_proposalA
Bid on a job posting in this account's name. Belongs to the PROPOSE scope. THIS REACHES A REAL PERSON: the cover letter is read by the client as words this account wrote, so get the person you are acting for to approve the text and the rate before calling - never bid off your own reasoning about a good match. Read the posting with freelance_get_job first: if it carries applicantQuestions every one must be answered here, and if it carries acknowledgments you must send acknowledged: true, or the bid is refused.
| Name | Required | Description | Default |
|---|---|---|---|
| input | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden, and it does so impressively. It warns that the cover letter reaches a real person as words the account wrote, mandates prior approval, and states concrete refusal conditions. This discloses the most consequential external effects and failure behavior, which is exactly what an agent needs before invoking a high-stakes bid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place: purpose, scope, human-impact warning, approval requirement, prerequisite read, and refusal conditions. The most important operational warning is front-loaded before the prerequisite details, and there is no filler or repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex submit action with no output schema and no annotations, the description is largely complete: it explains the human consequence, required precondition, conditional requirements, and failure behavior. It does not mention the requiredLinks/links condition or the success response shape, but those are partially covered by the schema, so the remaining gap is minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The outer description adds meaningful context for coverLetter and rate by explaining the cover letter is read as the account's words and both text and rate must be human-approved. However, it does not elaborate on currency, rateType, estimatedDuration, links, or answers, and schema description coverage is reported as 0%, so the description only partially compensates. The nested schema fields do carry their own descriptions, which keeps this at a minimally adequate 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a concrete action ('Bid on a job posting'), names the resource ('in this account's name'), and identifies the scope as PROPOSE. This clearly distinguishes it from siblings like freelance_create_job, freelance_accept_proposal, and freelance_list_my_proposals, so an agent can tell what the tool is for at a glance.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance: read the posting first with freelance_get_job, answer every applicantQuestion, send acknowledged: true when acknowledgments exist, and get human approval for the cover letter and rate before calling. It does not explicitly enumerate alternatives or say 'do not use when...', but the bid-scope and prerequisites make the intended usage clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
23 tool updates
v0.1.0- First observed
freelance_accept_proposal - First observed
freelance_add_milestone - First observed
freelance_approve_milestone - First observed
freelance_create_job - First observed
freelance_fund_milestone - First observed
freelance_get_contract - First observed
freelance_get_conversation - First observed
freelance_get_job - First observed
freelance_get_my_profile - First observed
freelance_get_profile - First observed
freelance_get_wallet - First observed
freelance_list_contracts - First observed
freelance_list_conversations - First observed
freelance_list_my_proposals - First observed
freelance_list_notifications - First observed
freelance_list_org_jobs - First observed
freelance_list_proposals_for_job - First observed
freelance_refund_milestone - First observed
freelance_search_jobs - First observed
freelance_search_talent - First observed
freelance_send_message - First observed
freelance_submit_milestone - First observed
freelance_submit_proposal
TDQS
Scored across 23 tools
Every tool targets a distinct resource-action pair: search/list tools are standard list-detail splits, and potentially overlapping tools like search_jobs vs list_org_jobs are explicitly separated by public vs own-organization visibility. No two tools do the same thing on the same resource.
All tools follow a strict 'freelance_<verb>_<noun>' pattern in snake_case (e.g., search_jobs, fund_milestone, send_message). Compound nouns like my_profile or proposals_for_job are consistent and do not break the pattern.
23 tools is within the borderline-heavy range, but the domain spans eight distinct areas (profiles, jobs, proposals, contracts, milestones, wallet, messaging, notifications). Each tool appears to earn its place, yet the count feels padded for a single server.
The core hire-to-payment lifecycle is covered, but there are notable gaps: no update or close/delete for job postings, no withdraw proposal, no cancel contract, and no edit milestone. These missing operations create dead ends that agents cannot work around.
Maintenance
Related MCP Connectors
Hire specialists by the hour — search, schedule, and pay via MCP protocol.
Agent-first task marketplace MCP — discover, claim, and deliver paid workspace tasks.
Automate 1,000+ services from any MCP-compatible AI agent: build Applets, run actions and queries.
Pay-per-use tool marketplace for AI agents. Search, price-check, and call APIs via MCP.
Related MCP Servers
- AlicenseBqualityDmaintenanceEnables AI assistants to execute GraphQL queries and retrieve schema information from any GraphQL endpoint.215 npm8MIT
- AlicenseNot gradedqualityCmaintenanceConnects AI agents to Upwork's GraphQL API, enabling job discovery, proposal management, profile tracking, and analytics.18 npm2MIT
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to interact with any GraphQL API by introspecting the schema and exposing queries and mutations as MCP tools, with built-in pagination, semantic search, and framework adapters.24 npmMIT
- FlicenseAqualityDmaintenanceEnables AI assistants to interact with the Upwork freelance marketplace, including job search, proposal management, contract tracking, and earnings monitoring.261-