colleag-mcp-ups
Server Quality Checklist
Latest release: v0.1.0
- Disambiguation5/5
Each tool targets a completely distinct UPS workflow: tracking an existing shipment, getting carrier rates, and estimating import costs. There is no overlap or realistic chance of selecting the wrong tool for a task.
Naming Consistency4/5track_shipment and get_rates follow a clear verb_noun pattern in snake_case. landed_cost is still readable and consistent in style, but it breaks the imperative verb pattern, making it a minor deviation.
Tool Count5/5Three tools is a well-scoped count for this focused UPS visibility and cost-estimation server. Each tool covers a meaningful capability with no redundant entries.
Completeness4/5The set covers the apparent purpose of tracking and cost estimation well: track an existing shipment, compare rates, and estimate landed cost. Shipping execution operations like creating or canceling a label are absent, but that seems outside the server's deliberate scope.
Average 4.3/5 across 3 of 3 tools scored.
See the Tool Scores section below for per-tool breakdowns.
- No community issues in the last 6 months
- 3 commits in the last 12 weeks
- Last stable release on
- No critical vulnerability alerts
- No high-severity vulnerability alerts
- No code scanning findings
- CI is passing
This repository is licensed under MIT License.
This repository includes a README.md file.
No tool usage detected in the last 30 days. Usage tracking helps demonstrate server value.
Tip: use the "Try in Browser" feature on the server page to seed initial usage.
Add a glama.json file to provide metadata about your server.
If you are the author, simply .
If the server belongs to an organization, first add
glama.jsonto the root of your repository:{ "$schema": "https://glama.ai/mcp/schemas/server.json", "maintainers": [ "your-github-username" ] }Then . Browse examples.
Add related servers to improve discoverability.
How to sync the server with GitHub?
Servers are automatically synced at least once per day, but you can also sync manually at any time to instantly update the server profile.
To manually sync the server, click the "Sync Server" button in the MCP server admin interface.
How is the quality score calculated?
The overall quality score combines two components: Tool Definition Quality (70%) and Server Coherence (30%).
Tool Definition Quality measures how well each tool describes itself to AI agents. Every tool is scored 1–5 across six dimensions: Purpose Clarity (25%), Usage Guidelines (20%), Behavioral Transparency (20%), Parameter Semantics (15%), Conciseness & Structure (10%), and Contextual Completeness (10%). The server-level definition quality score is calculated as 60% mean TDQS + 40% minimum TDQS, so a single poorly described tool pulls the score down.
Server Coherence evaluates how well the tools work together as a set, scoring four dimensions equally: Disambiguation (can agents tell tools apart?), Naming Consistency, Tool Count Appropriateness, and Completeness (are there gaps in the tool surface?).
Tiers are derived from the overall score: A (≥3.5), B (≥3.0), C (≥2.0), D (≥1.0), F (<1.0). B and above is considered passing.
Tool Scores
- Behavior4/5
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It explains that the tool produces an estimate, defines who pays at import, notes the DDP nuance, and exposes accuracy sensitivity to HS codes. This is substantial behavioral context, though it does not cover data sources or handling of missing export country.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Conciseness5/5Is the description appropriately sized, front-loaded, and free of redundancy?
Three efficient sentences, front-loaded with the main purpose, and every sentence adds useful information. No filler or redundant restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Completeness3/5Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description provides a solid high-level picture and output schema exists, but with five parameters and zero schema descriptions it leaves important gaps. Import country, currency, and export country semantics are missing, and commodity item structure is only vaguely implied. It is adequate but not fully complete for confident invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Parameters2/5Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It touches on commodities ('one commodity line per product', HS code) and indirectly on incoterm via DDP, but it never mentions import_country, export_country, or currency. Required parameters and accepted incoterm values are left undocumented, leaving an agent to guess critical input semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Purpose5/5Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Estimate import duties, VAT and brokerage fees for an international shipment.' It also clarifies the payment responsibility (receiver vs. shipper under DDP), which gives precise scope and differentiates it from sibling tools like track_shipment and get_rates.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Usage Guidelines4/5Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states when the tool is not applicable ('Not applicable to intra-EU shipments') and provides practical usage guidance ('One commodity line per product; an HS code improves accuracy'). However, it does not explicitly name alternatives among the sibling tools or explain when to prefer landed_cost over get_rates.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
- Behavior4/5
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It explains that the tool performs live quoting, uses negotiated account prices, can filter by service code, does not book anything, and defaults origin to the configured ship-from address. It omits auth/rate-limit details, but the key behavioral traits are well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Conciseness5/5Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no wasted words. The main purpose is front-loaded, and the service_code clarification, no-booking warning, and origin default each add distinct value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Completeness4/5Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists and annotations are absent, the description covers the essential invocation semantics: quote-only behavior, service selection, and default origin. It could be more explicit about how optional dimensions interact with rating and how this tool relates to landed_cost, but it is sufficient for correct common use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Parameters3/5Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It meaningfully explains the otherwise cryptic service_code parameter with concrete values (07, 65, 08, 11) and clarifies the origin default. It does not expand on required destination or weight parameters, but those names are fairly self-explanatory from the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Purpose5/5Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool gets live UPS shipping quotes for a lane, which is a specific verb plus resource. It also distinguishes itself from siblings by focusing on rate comparison rather than tracking or landed cost calculation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Usage Guidelines4/5Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context: use this tool for quote-only UPS rates, optionally narrowed to one service code. It does not explicitly say when to prefer landed_cost or track_shipment, but the quote-centric phrasing and the 'nothing is booked' note make the intended use case unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
- Behavior4/5
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral disclosure burden. It discloses what data will be returned, implies a read-only tracking operation, and usefully notes the 120-day UPS data retention limit. It does not mention error behavior or authorization, but these are minor for a simple read-style tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Conciseness5/5Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact: one sentence defines the action, resource, and expected returned data, and one short note adds a relevant behavioral limitation. Every part earns its place with no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Completeness5/5Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read-only tracking tool, the description is complete: it explains the input format, the outputs, and the retention limitation. The output schema covers the exact response shape, so no additional return-value detail is necessary.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Parameters5/5Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the sole parameter. It explicitly names the tracking number as the lookup key and gives a concrete format example ('1Z...'), which helps an agent construct a valid call. This exceeds what the input schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Purpose5/5Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the verb 'Track', the resource 'a UPS shipment', and the specific outputs: current status, scheduled/actual delivery date, and recent scan events. The tracking-number example and UPS scope distinguish it from sibling tools like get_rates and landed_cost.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Usage Guidelines4/5Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for use: when a UPS tracking number is available and shipment status information is needed. It does not explicitly name alternatives or exclusion criteria, but the purpose is specific enough that an agent can infer when to choose it over the rate/cost-focused siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
GitHub Badge
Glama performs regular codebase and documentation scans to:
- Confirm that the MCP server is working as expected.
- Confirm that there are no obvious security issues.
- Evaluate tool definition quality.
Our badge communicates server capabilities, safety, and installation instructions.
Card Badge
Copy to your README.md:
Score Badge
Copy to your README.md:
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/colleag-ai/mcp-ups'
If you have feedback or need assistance with the MCP directory API, please join our Discord server