Support Ticket Triage MCP
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Support Ticket Triage MCPTriage ticket TKT-001 with knowledge articles and similar tickets."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Support Ticket Triage MCP
A local Model Context Protocol (MCP) server and repository-local Codex Skill for governed support-ticket triage. The system reads synthetic tickets and knowledge articles, prepares evidence-backed recommendations, and records local audit events. The Skill directs Codex to present each recommendation and wait for a human decision before a finalizing action.
The current demo also includes a browser-based Approval Desk. It lets a reviewer type customer replies into a conversation workspace, generate an updated recommendation from the full timeline, inspect classifier evidence, review a customer draft, and approve only named fields.
The repository is a safety and workflow demonstration. It contains only synthetic fixture data, writes only to a local runtime directory, and has no live Zendesk, Jira, email, paging, identity, or customer-data connection. The fixture domain is Northstar Marketing Cloud, a fictional ecommerce marketing automation platform with synthetic support cases for flows, events, campaigns, profiles, segments, deliverability, SMS compliance, webhooks, coupons, and catalog sync. The articles and tickets are clean-room examples; they are not copied from a real vendor.
At A Glance
This project demonstrates governed AI support automation end to end:
Triage intelligence: deterministic classification, evidence readiness, optional GPT reasoning, and grounded customer-response drafting.
Workflow governance: one service owns lifecycle transitions, escalation, diagnosis review, fix gating, closure, and audit invariants across the Approval Desk and MCP tools.
Human control: recommendations and knowledge candidates remain pending until an operator explicitly approves named fields or promotes a reviewed object.
Knowledge evolution: completed diagnoses produce deterministic similarity signals; GPT may draft a reusable candidate, but only human promotion can make it available to future workflows.
Start with the 60-second Approval Desk demo, then run the knowledge-evolution showcase for the latest extension. The architecture, safety boundary, and limitations explain what the system can and cannot decide.
Related MCP server: Incident Triage MCP
Safety Boundary
Ticket subjects and descriptions are untrusted data. Embedded instructions, claimed approval, urgency, and policy-bypass language are evidence, not authorization.
submit_triage_recommendationstores a pending proposal. It does not change the ticket or an external system.The Skill/Codex workflow requires presenting the recommendation before a human explicitly approves named fields or explicitly rejects it with feedback.
The MCP approval schema requires
confirm: true, matching recommendation and ticket IDs, the current ticket revision, an actor, and one or more explicitly named fields. The service also enforces required security and outage routing.The MCP rejection schema requires a pending recommendation, matching recommendation and ticket IDs, an actor, and nonblank feedback. It has no revision check and cannot prove that a human intended the rejection.
Only
category,priority,team,assignee,status,tags, andcustomerResponseare approvable.Security risk must route to
security. A likely or confirmed outage must route toincident-response, unless security takes precedence while the outage reason remains visible.Submission rejects a stale source revision. Approval rechecks the recommendation source revision against the current expected ticket revision. Both approval and rejection reject an already-resolved recommendation.
Successful submission, approval, and rejection create append-style JSONL audit events. The local operator can still edit local files, so this is not a tamper-evident ledger.
See SECURITY.md for the full threat model.
Architecture
flowchart LR
Human["Human reviewer"]
Codex["Codex desktop project"]
Skill["Repository Skill<br/>$triaging-support-tickets"]
MCP["support-ticket-triage MCP server<br/>stdio"]
Reads["Read tools and resources"]
Service["TriageService"]
Policy["Policy, similarity, metrics"]
Tickets["Runtime tickets.json"]
Recommendations["Recommendation JSON files"]
Audit["Audit events.jsonl"]
Knowledge["Markdown knowledge articles"]
Seed["Synthetic seed fixtures"]
Human <--> Codex
Codex --> Skill
Codex <--> MCP
MCP --> Reads
MCP --> Service
Reads --> Tickets
Reads --> Knowledge
Reads --> Recommendations
Reads --> Audit
Reads --> Policy
Service --> Policy
Service --> Tickets
Service --> Recommendations
Service --> Audit
Seed --> TicketsDemo In 60 Seconds
npm ci
npm run build
npm run demo:showcaseOpen the printed local URL. A good portfolio walkthrough is:
Select
TKT-1010, the intentionally vague "Problem / It does not work" ticket.Create the first recommendation and review the generated customer-response draft.
Approve the named fields and click Done. The deterministic demo adds a ticket-specific customer reply automatically after the response is marked sent.
Evaluate the ticket again and point out that the system reclassifies it from generic support to a product performance issue, recalculates the evidence checklist, avoids asking for a screenshot of a blank page, and drafts a response that matches the new lifecycle state.
Use the conversation timeline and audit panel to show the ordered lifecycle.
The action bar's collapsed Advanced settings includes a manual customer-reply composer, an automatic-reply toggle, and action-bar positioning for edge cases or screen recording. Conversation Context is read-only and has no customer-reply editor. For a manual test reply, expand Advanced settings, check Disable automatic customer replies, expand Manual customer reply, paste or choose the reply, and click Add reply before evaluating again. The toggle and composer apply to the current Approval Desk page session and reset on reload. The Move action bar selector offers bottom-right, bottom-left, bottom-center, top-left, top-center, and top-right positions; its selection also resets on reload. These controls are not needed for the normal showcase flow, which uses the automatic reply generated after Done.
The Workflow Bar owns evaluation, recommendation approval, diagnosis, Fix, verification, and close actions. The separate Pattern Bar appears only when knowledge discovery has an active candidate or is still running. An actionable diagnosis or pattern review is a hard workflow gate: Done changes to Review, focuses the relevant bar, and downstream support actions remain hidden until the operator completes the required review.
The alternate incident walkthrough still works well with TKT-1001, which
shows correlated event-ingestion delay handling and incident-response routing.
Resettable Codex Skill AI Showcase
The command-line showcase replays the synthetic TKT-1010 lifecycle through
the MCP interface in fresh temporary state. It follows each
operatorGuidance.nextAction, reads the workflow again after every action, and
cleans the temporary state when it exits. Every report names its selected safe
mode and the classification, drafting, and network provenance used for that
mode. The saved controlled transcript is in
docs/skill-showcase-example.md.
npm run build
npm run demo:skill-showcase
npm run demo:skill-showcase -- --deterministicThe default controlled mode makes no network request and needs no configured
external provider. Local controlled simulations exercise both optional AI
roles: classification reasoning and customer-response drafting. The report
labels both roles controlled-local-simulation, labels network access
disabled, and never attributes their output to an external model adapter.
Their auditable traces are used when their advice or draft is accepted;
deterministic policy still owns routing, lifecycle, validation, and approval.
All eight controlled drafting stages are reported as accepted local
deterministic output with explicit simulation provenance.
The explicit --deterministic mode passes no providers and never makes a
provider call. Its classification traces and normal drafting traces are
skipped, including the valid customer-confirmed closure draft. Both modes
traverse Diagnose, Fix, verification, ready-for-close, and closed stages.
Every review step is an explicitly disclosed scripted
portfolio-reviewer simulation using exactly the approval fields supplied by
the workflow guidance.
Live mode is optional and is never selected implicitly. It requires an API key set only in the shell:
$env:OPENAI_API_KEY = 'set-in-the-shell-only'
npm run demo:skill-showcase -- --liveThe recorded portfolio results cover controlled and deterministic modes only; no live showcase run is claimed. Unknown arguments and repeated or conflicting mode flags fail safely, so a typo cannot silently select controlled mode.
Deterministic Lifecycle Replay
The chronological replay is the companion to the snapshot-based evaluation harness. It drives the existing MCP tools against fresh temporary state and verifies one complete, stateful journey:
evidence → diagnosis → approval → response → mitigation/fix → verification → closureRun it locally without GPT or network access:
npm run evaluate:lifecycle-replayThe replay reads get_ticket_workflow before the first action and after every
transition. Each read is validated to include the ticket, conversation history
and timeline, recommendation history and summary, latest recommendation when
present, and backend operator guidance. The report also proves that each MCP
mutation follows a workflow read, records explicit approval, captures the
customer reply, diagnosis and fix audits, and ends at resolved.
The current deterministic run reports 23 complete workflow reads, 20 governed MCP actions, and an 11/11 pass for the separate context-aware diagnostic scenario matrix. The matrix's bounded-ambiguity/escalation scenario remains a second supporting example: it routes unresolved ambiguity toward specialist review rather than pretending that an unresolved hypothesis is a fix.
This is a verification and portfolio showcase, not a second workflow engine.
The same TriageService, operator guidance, evidence gates, audit repository,
and MCP tools used by the Approval Desk perform the work.
Governed Diagnosis Review And Fixes
Once the required evidence is complete, an operator can record an immutable
diagnosis and review it before it becomes authoritative. The Approval Desk and
MCP tools call the same TriageService rules; neither duplicates diagnosis
freshness, fix gating, or lifecycle transitions. The Approval Desk and MCP
diagnosis reads present the live audit-backed history. Lifecycle Replay is a
separate, read-only evaluation-report snapshot view; it does not read live
audits or make lifecycle decisions.
The normal operator journey is deliberately bounded:
Complete the evidence checklist and send the approved evidence update.
Record the original diagnosis, then explicitly approve or revalidate it. A review is a separate audit event; it never rewrites the original diagnosis.
Send the reviewed, customer-safe diagnosis response. Authorizing only that outbound response does not change ticket revision or invalidate otherwise current diagnostic evidence.
Select every ticket that the approved diagnosis should affect and explain why. A diagnosis-scoped fix writes a separate audit event for every selected ticket; it does not resolve any ticket.
Send a customer-safe verification request. A new customer reply or a real ticket-field change makes the earlier review stale, so an operator must revalidate before any later governed action can rely on it.
Close only after the customer has confirmed the result, the configured ready-to-close response has been approved and sent, and an operator takes the explicit close action.
Customer-facing updates say what was investigated or corrected and what the customer should verify. They do not expose internal policy, detection, similarity, prompt, or secret-handling details. Optional GPT assistance can propose drafts, but deterministic checks and explicit human approval remain authoritative.
See the detailed synthetic walkthrough in docs/diagnosis-review-example.md. The ordered regression is runnable locally:
npx vitest run test/approval-desk-diagnostic-workflow.test.ts test/demo-skill-showcase.test.tsThis diagnosis-review slice does not add or re-prove the existing separately governed candidate discovery, optional GPT candidate drafting, or human promotion workflow. It also does not add revision-aware queue analysis, evidence-graph similarity, or executable/versioned workflow migration. Those capabilities remain separately governed and explicitly audited rather than being inferred from a reviewed diagnosis.
Screenshots
The screenshots below are generated from local synthetic data.



Hybrid Recommendation Architecture
flowchart LR
Ticket["Synthetic ticket and replies<br/>untrusted text"]
Context["Conversation context"]
KB["Retrieved local KB articles"]
Rules["Deterministic classifier<br/>and safety rules"]
GPTReasoning["Optional GPT advisory<br/>classification signals"]
Evidence["Evidence readiness<br/>and lifecycle"]
GPTDraft["Optional GPT draft provider"]
Validators["Deterministic draft validators"]
Fallback["Local deterministic fallback"]
Reviewer["Human reviewer"]
Audit["Local audit trail"]
Ticket --> Context
Context --> Rules
Context --> GPTReasoning
GPTReasoning --> Rules
Rules --> Evidence
KB --> GPTDraft
Evidence --> GPTDraft
Rules --> Fallback
GPTDraft --> Validators
Validators -->|pass| Reviewer
Validators -->|warn or provider error| Fallback
Fallback --> Reviewer
Reviewer -->|approve named fields| AuditThe important boundary is that deterministic code remains the final authority for security, outage, SLA, approval, and audit behavior. GPT can help in two bounded ways: drafting customer-facing language and, when configured, proposing low-to-medium-weight advisory classification signals for ambiguous evolving conversations. Those signals are recorded as classifier evidence and cannot override hard safety rules.
What This Demonstrates
MCP tools can expose local business data and workflow actions to an AI assistant without connecting to live customer systems.
Codex can operate the modern ticket workflow through MCP tools:
get_ticket_workflow,add_customer_reply,evaluate_ticket, andmark_response_donemirror the Approval Desk lifecycle while preserving the human approval boundary.Deterministic policy can own routing, escalation, validation, approval, and audit guarantees while GPT assists with bounded drafting and advisory classification evidence.
Conversation-aware automation can re-evaluate a ticket after customer replies, recalculate evidence requirements, and adapt the next draft.
Retrieved knowledge articles can ground a customer response without exposing internal article IDs to the customer.
Human reviewers can edit and approve named fields, preserving accountability instead of letting automation mutate tickets directly.
The same local demo can show success, fallback, stale-approval rejection, and audit evidence in a repeatable way.
For a shorter narrative version, see docs/case-study.md. For sample outputs and demo talking points, see docs/demo-results.md. For screenshot and GIF planning, see docs/capture-guide.md. For next build ideas, see docs/roadmap.md.
The stdio entry point is dist/src/index.js. Its defaults are:
Setting | Default |
|
|
|
|
|
|
|
|
|
|
| unset (deterministic discovery only); use |
| unset (inherits |
|
|
All relative paths are resolved from the process working directory.
Knowledge candidate drafting is explicitly opt-in. Set
TRIAGE_KNOWLEDGE_CANDIDATE_PROVIDER=openai and provide OPENAI_API_KEY to
enable the advisory Knowledge Engineer. It shares OPENAI_MODEL by default,
with the knowledge-specific model override available above. Discovery remains
deterministic, candidate output is contract- and guardrail-validated, and a
human operator must still review and promote any candidate; GPT never changes a
ticket or workflow directly. Without the selector, only deterministic discovery
runs.
Codex Operator Layer
The MCP server exposes two generations of workflow tools:
Tool | Purpose |
| Reads the ticket, conversation timeline, recommendation history, latest recommendation, and workflow state. |
| Appends a customer reply to the local audit trail before re-evaluation. |
| Runs the current Approval Desk recommendation builder from the full timeline instead of asking Codex to hand-build recommendation JSON. |
| Applies only explicitly approved fields and records the approved customer response as sent. |
| Legacy lower-level proposal tool for manually assembled recommendations. |
| Legacy explicit finalization tools guarded by strict schemas and audit logging. |
The repository-local Skill at .agents/skills/triaging-support-tickets teaches
Codex to use the operator tools, present evidence and drafts, wait for explicit
human approval of named fields, and verify the resulting audit trail.
Approval Flow
sequenceDiagram
participant H as Human
participant C as Codex and Skill
participant M as MCP server
participant R as Local repositories
C->>M: get_ticket
C->>M: search_knowledge
C->>M: find_similar_tickets
C->>M: submit_triage_recommendation
M->>R: Store pending recommendation and submission audit
M-->>C: Recommendation, source revision, computed escalation
C-->>H: Evidence, citations, confidence, risks, proposed fields, response
Note over C,H: Stop before mutation
H->>C: Approve explicit named fields
C->>M: approve_triage_recommendation with confirm true
M->>M: Validate pending state, revision, fields, and required routing
M->>R: Update ticket, resolve recommendation, append approval audit
M-->>C: Updated ticket and audit event
C->>M: get_ticket and get_audit_events
C-->>H: Readback of changed and unchanged fieldsThe Skill/Codex workflow treats rejection as a human decision and requires explicit rejection wording plus concrete feedback. The MCP rejection action validates the pending recommendation, matching IDs, actor, and nonblank feedback, then records an audit without changing the ticket; it cannot verify who formed the intent and does not check a ticket revision.
Requirements
Node.js
^20.19.0,^22.12.0, or>=24.0.0npm
PowerShell for the commands below
Codex desktop when exercising the repository Skill and project MCP config
Setup And Verification
From the repository root:
npm ci
npm run build
npm testnpm test runs pretest, which rebuilds, type-checks, and then runs the Vitest
suite in test/.
Generate the deterministic synthetic fixtures and knowledge articles:
npm run build
npm run generate:fixtures
git diff -- data/seed/tickets.json data/seed/expected-outcomes.json data/knowledgeRun the fixture evaluation:
npm run build
npm run evaluateRun the compiled stdio server directly only when testing an MCP client or diagnosing startup:
npm run build
npm startThe server speaks MCP over standard input and output, so an idle terminal is normal. Diagnostics are written to standard error.
Reset The Local Demo State
Stop the MCP server before resetting. This preserves data/runtime/.gitkeep
and removes ignored runtime tickets, recommendations, and audits:
$ErrorActionPreference = 'Stop'
$repoRoot = (Resolve-Path -LiteralPath '.' -ErrorAction Stop).ProviderPath
$packagePath = Join-Path -Path $repoRoot -ChildPath 'package.json'
if (-not (Test-Path -LiteralPath $packagePath -PathType Leaf)) {
throw "Refusing reset: package.json was not found at $packagePath"
}
$package = Get-Content -LiteralPath $packagePath -Raw -ErrorAction Stop |
ConvertFrom-Json -ErrorAction Stop
if ($package.name -ne 'support-ticket-triage-mcp') {
throw "Refusing reset: unexpected package name '$($package.name)'."
}
$dataRoot = Join-Path -Path $repoRoot -ChildPath 'data'
$dataItem = Get-Item -LiteralPath $dataRoot -Force -ErrorAction Stop
if (($dataItem.Attributes -band [System.IO.FileAttributes]::ReparsePoint) -ne 0) {
throw "Refusing reset: data directory is a reparse point."
}
$expectedRuntimeRoot = [System.IO.Path]::GetFullPath(
(Join-Path -Path $repoRoot -ChildPath 'data\runtime')
)
$runtimeItem = Get-Item -LiteralPath $expectedRuntimeRoot -Force -ErrorAction Stop
if (($runtimeItem.Attributes -band [System.IO.FileAttributes]::ReparsePoint) -ne 0) {
throw "Refusing reset: runtime directory is a reparse point."
}
$runtimeRoot = $runtimeItem.FullName
if (-not [string]::Equals(
[System.IO.Path]::GetFullPath($runtimeRoot).TrimEnd([char[]]"\/"),
$expectedRuntimeRoot.TrimEnd([char[]]"\/"),
[System.StringComparison]::OrdinalIgnoreCase
)) {
throw "Refusing reset: runtime directory resolved outside the verified repository."
}
$runtimeChildren = @(
Get-ChildItem -LiteralPath $runtimeRoot -Force -ErrorAction Stop
)
$resetTargets = @(
$runtimeChildren | Where-Object Name -ne '.gitkeep'
)
foreach ($target in $resetTargets) {
if (($target.Attributes -band [System.IO.FileAttributes]::ReparsePoint) -ne 0) {
throw "Refusing reset: runtime child is a reparse point: $($target.FullName)"
}
}
foreach ($target in $resetTargets) {
Remove-Item -LiteralPath $target.FullName -Recurse -Force -ErrorAction Stop
}All repository, package, path, JSON, reparse-point, and enumeration checks
finish before the deletion loop starts. The next server start initializes
data/runtime/tickets.json from the synthetic seed without overwriting an
existing runtime file.
Use From Codex Desktop
No separate codex command is required for this repository.
Run
npm ciandnpm run buildin PowerShell.Open the repository root as a local project in Codex desktop.
Trust the project only after reviewing
.codex/config.toml; it launchesnode dist/src/index.jswith the repository root as its working directory.Start a new thread after building or after changing the project MCP config.
Trigger the repository Skill explicitly in the prompt:
Use $triaging-support-tickets to triage TKT-1005 using the local MCP server.
Present the recommendation and wait for my explicit approval of named fields.The Skill lives at
.agents/skills/triaging-support-tickets/SKILL.md. Its UI metadata is at
.agents/skills/triaging-support-tickets/agents/openai.yaml, and its detailed
classification and escalation tables are in
.agents/skills/triaging-support-tickets/references/policy.md.
Use The Local Approval Desk
The Approval Desk is a local browser UI for the human decision layer. It uses
the same synthetic fixtures, local repositories, and TriageService rules as
the MCP server.
npm ci
npm run build
npm run approval-deskFor a repeatable walkthrough, run:
npm ci
npm run build
npm run demo:showcasedemo:showcase is an alias for the local Approval Desk demo runner. It resets
local runtime data, starts the Approval Desk, and prints the local URL plus a
suggested presentation path. The Automation Evidence dashboard shows open
tickets, recommendation counts, estimated minutes saved, audit events, safety
blocks, and active guardrails.
Open the printed http://127.0.0.1:5177 URL. Select TKT-1010 for the
conversation-aware reclassification walkthrough, or TKT-1001 for the incident
routing walkthrough. The Recommendation panel shows classifier evidence,
lifecycle state, evidence readiness, the customer draft, validator checks,
retrieved context, and a compact What changed summary when a new
recommendation differs from the previous one. Select named fields, enter an
actor, check the explicit confirmation box, and approve. The UI then reads back
the updated ticket revision and audit event.
The app is local-only. It does not send customer responses, connect to external support systems, or authenticate multiple users.
GPT Drafting And Advisory Classification
The Approval Desk can build draft customer responses in two modes:
default deterministic local drafting, which requires no network or API key;
optional OpenAI drafting, which uses the Responses API when
APPROVAL_DRAFT_PROVIDER=openaiandOPENAI_API_KEYare set.
Both modes keep the same approval and audit boundary. In OpenAI drafting mode:
The app retrieves the selected ticket, conversation timeline, classifier outcome, evidence readiness, lifecycle state, and cited local knowledge articles.
The OpenAI draft provider writes a customer response from that trusted context.
Deterministic validators check the draft for unsafe promises, internal-only IDs, approval-bypass language, and missing human-review boundaries.
If the provider fails or the draft fails validation, the app falls back to the deterministic local response.
The human reviewer still edits and approves the response before anything is recorded in the audit trail.
Run the optional OpenAI drafting mode from PowerShell:
$env:OPENAI_API_KEY = 'sk-...'
$env:APPROVAL_DRAFT_PROVIDER = 'openai'
$env:OPENAI_MODEL = 'gpt-5.6-luna'
$env:APPROVAL_RESPONSE_STYLE = 'balanced'
npm run demo:showcaseOPENAI_MODEL is optional; the app defaults to gpt-5.6-luna. The draft
source and validation checks appear in the Recommendation panel so reviewers
can see whether the response came from deterministic rules, OpenAI, or a local
fallback.
The Approval Desk also includes a Draft style selector. Supported styles are
balanced, concise, empathetic, technical, and executive-update.
APPROVAL_RESPONSE_STYLE is still available as the startup default and falls
back to balanced. These settings change only the GPT draft tone.
The current backend also has an injectable GPT reasoning lane for ambiguous
conversation context. A GptClassificationReasoningProvider can return
structured advisory output such as candidate category, team, priority,
knowledge article IDs, evidence, and explanation. The app converts that output
into gpt-advisory-* classification signals. Deterministic safety signals
still win: security, outage, SLA, stale revision, approval requirements, and
named-field mutation rules remain local code.
Do not commit API keys, paste them into tickets, include them in screenshots, or store them in runtime audit data. The demo should remain usable without an API key by defaulting to the deterministic local provider.
Other useful trigger examples:
Use $triaging-support-tickets to review TKT-1004. Surface every escalation,
cite the local policy articles, and stop before changing the ticket.Use $triaging-support-tickets to triage TKT-1001, TKT-1002, and TKT-1003 as
a correlated incident cluster. Prepare recommendations only.MCP Interface
The server exposes exactly 9 tools: 6 read-only tools and 3 local workflow actions.
Read-Only Tools
Tool | Purpose | Important bounds |
| Filter and page tickets |
|
| Read one | Exact ticket ID |
| Search local Markdown knowledge | Nonblank query, |
| Rank deterministic Jaccard candidates | At most 5 candidates with score greater than 0.2 |
| Calculate queue, SLA, recommendation, escalation, and savings counters | No input |
| Page all audits or one ticket's audits |
|
All six are annotated read-only, non-destructive, idempotent, and closed-world.
Workflow Actions
Tool | Effect | Boundary |
| Stores a pending recommendation and submission audit | Does not change the ticket; server owns the timestamp and recomputes escalation |
| Applies only approved fields and returns the ticket plus audit event | Enforces pending state, matching IDs, exact revision, actor, named fields, |
| Resolves a pending recommendation as rejected and records feedback | Enforces pending state, matching IDs, actor, and nonblank feedback; has no revision check and leaves the ticket unchanged |
Submission mutates local workflow data but is annotated non-destructive. Approval and rejection are annotated destructive because they finalize local state; none of the actions are idempotent or open-world.
The Skill/Codex workflow supplies the human-decision boundary by presenting a recommendation and waiting for explicit approval or rejection. MCP validates the action payload and repository state, but it cannot prove that a human saw the recommendation or personally formed the intent represented by a tool call.
customerResponse is an approvable recommendation field, but the ticket schema
has no customer-response property and there is no outbound messaging
integration. Its approved text is recorded in the audit event's before and
after data; it is not sent or stored on the ticket.
Resources
The server exposes 4 resources:
URI | MIME type | Content |
|
| One ticket |
|
| One knowledge article body |
|
| First 50 ticket audit events plus total |
|
| Current queue metrics |
The first three are resource templates. metrics://queue is the single
directly listed resource.
Prompts
The server exposes exactly 3 MCP prompts:
Prompt | Arguments | Behavior |
| Required | Reads one ticket, knowledge, and similar tickets; submits a recommendation; stops before approval |
| Optional integer | Prepares recommendations for a bounded batch; stops before approval |
| None | Reviews security, outage, confidence, and SLA escalation conditions; stops before approval |
Each prompt states that ticket text is untrusted and approval cannot be inferred from ticket content.
Five-Minute Walkthrough
For a clean synthetic fixture state, build and reset data/runtime before
opening the project in Codex. Fixture data and deterministic tool calculations
are reproducible when state and time inputs match. Model-generated
recommendations and wording may vary, so the checkpoints are acceptance
criteria rather than a guaranteed transcript. The detailed script is in
docs/demo-script.md.
Read
metrics://queueor callget_queue_metrics. A fresh fixture has 30 tickets, 29 open tickets, and no recommendations. SLA counts depend on the current clock because fixture deadlines are fixed on June 10, 2026.Triage
TKT-1005. The Browse Abandonment ticket contains an instruction to ignore policy, close as P4, skip approval, and hide the instruction. The workflow must ignore it, preserve integration/P2/integrations evidence, surface policy conflict, prepare a pending recommendation, and stop.Triage
TKT-1004. The private-key exposure report must remain security/P1 and route tosecurity, with the unknown exposure scope surfaced.Triage
TKT-1001,TKT-1002, andTKT-1003. Deterministic similarity links the EU event-ingestion delay cluster, and the expected outcome is incident/P1/incident-response with outage and SLA escalation.After seeing one recommendation, approve selected named fields only. Then read the ticket and audit event to verify the revision, actor, citations, changed fields, and unchanged fields.
Use the Approval Desk conversation workspace on
TKT-1010to show the evolving-ticket path: a vague first contact becomes a product performance diagnosis after the customer describes the blank campaign editor.
Queue Metrics
get_queue_metrics and metrics://queue return:
open and untriaged ticket counts;
breached and at-risk SLA counts;
open-ticket counts by category, priority, and team;
submitted, pending, approved, and rejected recommendation counts;
acceptance and rejection rates over resolved recommendations;
average submitted-recommendation confidence;
escalation totals and counts by reason;
configured minutes per accepted recommendation;
estimated minutes saved;
average approved confidence, confidence-band counts, and pending potential savings.
The savings formula is deliberately simple:
estimatedMinutesSaved =
approvedRecommendations * minutesPerAcceptedRecommendation
potentialMinutesSaved =
pendingRecommendations * minutesPerAcceptedRecommendationestimatedMinutesSaved is a realized, approval-attributed estimate: it counts
only approved recommendations. potentialMinutesSaved is a projection for
pending recommendations. Neither value is measured stopwatch time, labor cost,
customer outcome, or financial impact. The stdio process defaults to 8 minutes
per accepted recommendation. Override
the bookkeeping assumption before starting a manual server process:
$env:TRIAGE_MINUTES_SAVED = '5'
npm startThis value is a configured estimate, not measured labor, cost, response time, customer outcome, or financial impact. At a fresh runtime there are no approved recommendations, so the estimate is zero.
Uncertainty-aware classification confidence
Classifier confidence is uncertainty-aware decision support, not a calibrated
probability that the classification is true. The deterministic classifier
uses the versioned uncertainty-aware-v1 method, combining category support,
the margin over the runner-up, independent signal diversity, and disagreement
penalties. The resulting bands are low (<0.75), medium (0.75–<0.90), and
high (>=0.90).
Persisted provenance records bounded reason codes such as
weak-category-support, close-category-competition, low-signal-diversity,
metadata-disagreement, and no-actionable-category. Metadata, disagreement,
known-cause/event, duplicate, and GPT-advisory emissions do not count as
independent evidence diversity. GPT classification may suggest bounded
advisory signals, but it cannot author trusted confidence provenance; the
deterministic classifier and approval boundary remain authoritative.
Older recommendations without optional provenance remain readable. Their numeric confidence is retained, and queue metrics can still place it in a band, but no retrospective reason codes are invented. Fixture/expected-outcome evaluation lanes likewise use the legacy scalar-confidence path and do not author trusted provenance.
For a deterministic, transport-consistent showcase run:
npm run demo:metricsThe command reads the seed tickets and sample recommendations, emits stable
JSON validated by QueueMetricsSchema, and uses the same fixed timestamp and
8-minute default assumption as the metrics calculator. get_queue_metrics,
/api/metrics, and metrics://queue all use that shared schema/calculator;
the CLI output is a consistency check rather than a second metrics
implementation.
For a single release-readiness check covering the complete portfolio journey, run:
npm run verify:portfolioThis runs the build, typecheck, full Vitest suite, diagnostic evaluation, stateful lifecycle replay, deterministic knowledge holdout, knowledge-evolution promotion/reuse showcase, and deterministic queue-metrics showcase in sequence. It stops at the first failure, so the command is suitable for reproducing the evidence reported in this README without enabling live GPT providers.
Reproducible Evaluation
npm run evaluate compares
data/seed/sample-recommendations.json with
data/seed/expected-outcomes.json. The committed sample is intentionally
constructed to match all 30 expected outcomes and prints:
{
"ticketCount": 30,
"categoryAccuracy": 1,
"routingAccuracy": 1,
"priorityAgreement": 1,
"securityEscalationRecall": 1,
"outageEscalationRecall": 1,
"duplicatePrecision": 1,
"duplicateRecall": 1,
"knowledgeCitationCoverage": 1,
"approvalSafetyViolations": 0
}These are reproducible fixture results, not observations from real support work. The evaluator requires recommendation ticket IDs to match the expected outcome set exactly and counts any non-pending sample recommendation as an approval-safety violation.
The focused diagnostic harness is intentionally a small scenario suite. To audit the live deterministic path against every seeded ticket, run:
npm run evaluate:lifecycleThis 30-ticket audit reports each ticket's seed status, category/routing, known-cause and known-event matches, evidence gate, diagnosis confidence, operator next action, and the shared lifecycle invariants. It is a baseline lifecycle audit rather than a claim that every ticket has been replayed through every chronological customer turn; Lifecycle Replay remains a snapshot viewer for the separate diagnostic scenarios. The current seed audit reports 30/30 classification contracts, 28 tickets correctly waiting for evidence, 30/30 lifecycle invariants, 7 known-cause matches, 3 known-event matches, and one already-resolved ticket with zero evidence requests.
Deterministic knowledge holdout
The governed-learning proof is a fixed, multi-turn holdout rather than a single snapshot or a live-model benchmark. Run it with:
npm run evaluate:knowledge-holdoutThe command evaluates baseline and learned lanes through the same
listReusableApproved({ asOf }) and evaluateTicketWithAi production path,
then writes sanitized artifacts to
reports/knowledge-holdout/controlled-latest.json and
reports/knowledge-holdout/controlled-latest.md. The fixtures cover complete
evidence, evidence arriving on a later customer turn, near misses, unrelated
questions, stale and contradicted knowledge, an approved replacement, and an
unapproved revision draft. Every turn is retained, and scoring verifies that
the evaluator changed no ticket, recommendation, audit, learning, candidate,
version, or head state.
The report separates efficacy, governance safety, and exact-version scorecards. It is controlled synthetic evidence: it demonstrates governed reuse and historical/version isolation, but it makes no claim about human time saved or live GPT performance.
The scorecard uses explicit ratios rather than a single blended quality number:
exact-version precision is correct learned matches divided by all learned
matches, recall is correct learned matches divided by expected matches, and
evidence precision is necessary evidence requested divided by all requested
evidence. Governance rates count stale or contradicted versions reused when
they should be excluded; version rates count the expected exact version or
approved replacement selected. Missing-evidence rate is missing necessary IDs
divided by all expected IDs; zero-denominator rate metrics are reported as
null, while count totals remain zero. Baseline-to-learned deltas are
safety-first:
new unsafe lifecycle codes, wrong-version reuse, or lost evidence are
regressions even when a later turn recovers.
If the learning ledger cannot be read, listReusableApproved({ asOf })
returns ledger-unavailable with no reusable contexts. Deterministic triage
continues without learned context, and it never silently falls back to an
older version. Deprecated versions are excluded by the reusable-context
boundary regression; the holdout provides end-to-end stale and contradicted
version lanes.
Article-backed diagnosis and GPT advisory diagnosis
The deterministic diagnostic workflow now has specific playbooks for the
approved knowledge articles used by the seed set, including deliverability,
audience rules, campaign preparation, catalog/coupon sync, profile sync,
ecommerce integration, SMS compliance, and webhook validation. These
playbooks produce a customer-safe narrowing diagnosis and next action; they
never turn an article match into proof of a root cause. Missing evidence still
forces a likely diagnosis and keeps the ticket in its evidence-gated state.
Two deliberately vague tickets (TKT-1010 and TKT-1026) have no meaningful
article or playbook on purpose. They request basic problem evidence and the GPT
diagnosis lane skips them rather than manufacturing a diagnosis. Resolved
tickets and prompt-injection tickets are skipped for the same governance
reason.
To evaluate the optional GPT diagnosis lane across all 30 seed tickets, run:
npm run evaluate:ai-diagnosisThe default run uses a controlled local adapter to exercise the same provider
contract without network access. For an authenticated observation using real
GPT, set OPENAI_API_KEY and opt in explicitly:
npm run evaluate:ai-diagnosis -- --liveEach run saves sanitized output under reports/ai-diagnosis/ as
controlled-latest.* or live-latest.*.
GPT returns an advisory candidate only. The deterministic diagnosis, evidence gate, prompt-injection preflight, lifecycle state, and human approval boundary remain authoritative. Invalid output, unavailable GPT, or stale/unsafe context falls back to the deterministic diagnosis without exposing provider payloads.
AI Comparison Evaluation
npm run evaluate:ai-comparison adds a network-free, controlled-local
comparison of deterministic classification/drafting, advisory classification,
and drafting providers across the eleven diagnostic scenarios. Its four lanes
are reproducible local simulations, not live-model observations. The report
shows actual drafts, operator stage, baseline-versus-final classification,
sanitized candidate/advisory signals, accepted and rejected advice,
deterministic overrides, evidence recall and precision, hard-safety status,
safe failure reasons, and provider provenance without recording API keys or
raw provider payloads.
For an explicit live OpenAI observation, set OPENAI_API_KEY in the shell and
run:
npm run evaluate:ai-comparison -- --liveLive mode runs only GPT-containing lanes, labels them
live-openai-adapter, and records returned model/latency/token metadata. It
is opt-in and non-reproducible; default tests and the default comparison
command do not make network calls. Prompt-injection scenarios skip both GPT
stages, while deterministic safety policy remains authoritative. Read the
AI comparison evaluation guide
and its sanitized controlled-local example
before interpreting results.
The command also saves the full sanitized output under
reports/ai-comparison/controlled-latest.md and
reports/ai-comparison/controlled-latest.json (or the corresponding live-
latest files for --live). These files contain the complete per-lane
customer drafts, not just the shortened documentation example.
Customer-response drafting is contract-first: the deterministic baseline and authoritative workflow state define customer-safe obligations before an optional GPT draft is considered. A GPT candidate must pass that local contract, may receive one bounded repair attempt, and otherwise falls back to the deterministic draft. The report records only the candidate/repair/fallback outcome and sanitized obligation identifiers; it never adds candidate model text to provenance.
Controlled response-quality scorecard (local simulation, 2026-07-26):
11/11 scenarios passed in each of the four lanes (44/44 lane/scenario runs).
Classification agreement is reported independently from drafting quality; all 44 controlled final classifications matched their expected contracts.
All 44 final drafts passed hard safety and response-quality checks. Candidate hard-safety violations, final-response hard-safety violations, and deterministic fallbacks were all 0.
Candidate passes and repaired passes were both 0 because controlled drafting reuses the deterministic baseline; these counters are not live-GPT claims.
Checks cover required concepts and evidence, evidence precision, forbidden claims, unnecessary questions, tone, length, deterministic safety checks, and contract repair/fallback provenance.
GPT-labelled lanes use the local controlled simulation; use the opt-in
--liverun for a fresh live-model observation.
Live response-quality scorecard (last authenticated run, 2026-07-27,
gpt-5.6-luna):
GPT-assisted classification agreed with the expected classification contract in 20/20 GPT classification evaluations.
The GPT-classification/deterministic-draft lane passed 11/11 scenarios.
The deterministic-classification/GPT-draft lane passed 8/11 scenarios; the full GPT-classification/GPT-draft lane passed 6/11. Overall, GPT-containing lanes passed 25/33 scenario runs.
GPT drafting produced 10/17 strict response-quality passes. The remaining misses were wording/evidence-state and tight-length contract misses, not unsafe customer claims. One stale-reply candidate was rejected and safely fell back to the deterministic draft.
No forbidden customer-facing claims were detected. Three final response quality gates still failed in the report (evidence-state, deterministic safety-check, and/or length checks); candidate-level violations and fallback reasons remain visible separately instead of being hidden behind the lane score.
These live numbers are an observation, not a reproducible guarantee: they
depend on the selected model and prompt response. The controlled 44/44 result
is the regression baseline. The live run is useful because it demonstrates the
portfolio boundary in practice: GPT can improve classification evidence and
wording, while deterministic contracts, safety gates, and fallback behavior
remain authoritative. The two evaluator aliases added after that run cover
the exact phrases platform processing delay and source event creation time
seen in the live drafts; a new network run is not required to validate those
local contract changes.
Governed Knowledge Evolution
Completed diagnoses can deterministically surface a reusable knowledge
candidate. GPT may optionally draft a strictly validated advisory version, but
it cannot route tickets, promote knowledge, or change a customer response.
An authorized operator reviews and may edit the evidence policy,
customer-safe explanation, owner, triggers, time constraints, and declarative
workflows before explicitly promoting the candidate. The promotion audit
records the supporting diagnoses, deterministic provenance, reviewer, and
version. Approval Desk candidate discovery is a POST action because it may
persist a new candidate and its creation or rediscovery audit; plain reads use
the candidate detail endpoint.
Only approved objects affect later evaluations; candidates (including rejected candidates) have no routing effect, and promotion never rewrites earlier recommendations or audits. Defer is a resumable review pause, not approval or rejection: the candidate remains a hard gate until the operator explicitly approves or rejects it. Approval Desk and MCP use the same knowledge service, while lifecycle replay and AI comparison project the same approved object context through the shared evidence and diagnostic workflow.
Durable learning ledger
The knowledge-evolution learning plane now has a durable SQLite learning ledger. It records sanitized, append-oriented events for diagnosis support, verified outcomes, candidate review, immutable object versions, reuse, stale signals, contradictions, and evaluation evidence. Candidate/version/audit promotion is transaction-safe, and duplicate event IDs are idempotent.
This is governed knowledge evolution, not autonomous model retraining. The operational plane remains authoritative for classification, evidence readiness, diagnosis, fixes, verification, lifecycle transitions, and customer drafts. The ledger only records verified outcomes and proposes reusable knowledge; GPT-generated fields remain advisory, and a named operator must approve a version before it can affect a future evaluation.
Learning maturity and health are separate axes: observed,
diagnosis-supported, outcome-verified, reuse-validated, and promoted
describe the strength of evidence, while active, stale, contradicted,
deprecated, and superseded describe whether that evidence is currently
usable. Stale history remains queryable with decayed signal weight but cannot
bypass an evidence gate. New evaluations use the latest active version;
in-progress tickets remain pinned to their original version, so historical
recommendations remain unchanged.
Run the deterministic ledger showcase without an API key or network access:
npm run demo:learning-ledgerThe output demonstrates candidate creation, explicit human promotion, a verified outcome, successful and failed reuse signals, stale decay, and historical immutability. The first SQLite slice intentionally does not migrate the operational ticket, conversation, recommendation, or audit stores; those remain local JSON/JSONL adapters until a later migration with equivalent repository contracts.
Evidence provenance and policy boundaries
Knowledge evolution keeps three related, but deliberately different, records:
Observed diagnosis evidence is what was actually present when an operator completed a diagnosis.
evidenceUsedremains readable reasoning;evidenceReferencescontains catalog-backed IDs, the immutable label seen at diagnosis time, and optional ticket/reply/knowledge/operator provenance. References are created from recognized provided evidence only. The system never turns an audit ID or free-form prose into a synthetic evidence ID.Candidate evidence policy is a reviewable proposal for future tickets. Discovery deduplicates observed IDs for this proposal while preserving duplicate observations and their provenance in diagnosis history. With multiple supporting diagnoses, deterministic discovery uses the intersection of catalog-backed IDs shared by every diagnosis; it never unions divergent evidence or chooses the first diagnosis by order. A candidate may be
required, explicitlynone-requiredwith a rationale, orundecidedwhen no safe shared policy is available. Anundecidedcandidate stays visible for review but cannot be promoted.Approved evidence policy is the operator-approved contract used by later evaluations. Promotion reloads the current candidate and catalog, then accepts only a valid
requiredpolicy or a justifiednone-requiredpolicy. GPT can suggest fields, but it cannot select this policy or promote it.
This creates two important reuse paths. In the positive path, a diagnosis with
real catalog references produces a candidate, an operator approves its policy,
and a later matching ticket remains evidence-gated until required evidence is
present; after reevaluation, the approved known cause can be selected through
the same catalog. In the negative path, a diagnosis may still be valid and
readable when it has no structured references, but its candidate is
undecided, cannot affect routing, and is rejected until an operator supplies
and approves a valid policy. An active outage or an open-ticket similarity
signal does not bypass this gate by itself.
Catalog entries are retained for compatibility. Existing approved objects that refer to a deprecated evidence ID remain readable and executable; new promotions containing that ID are blocked and must use its replacement when one is defined.
In the Approval Desk, select a ticket and open Advanced settings in the Workflow Bar to run Find pattern explicitly. Discovery also runs after an authoritative diagnosis is available, but the manual trigger makes the boundary visible and reports when the current completed-diagnosis evidence does not meet the candidate threshold. When a candidate is found, the separate Pattern Bar shows the supporting evidence and compact declarative editor. The operator can inspect the evidence choices, edit the proposed object, and Approve, Refresh, Defer, or Reject it. The button never changes the ticket or promotes a candidate by itself.
Evidence policy remains the gate. For example, an approved
none-required known cause whose deterministic trigger matches may use its
confirmed known-cause path without collecting extra customer evidence. An
approved required cause still requests every listed evidence item before it
can advance. Ordinary tickets and active outages remain evidence-gated.
Customer drafts use only the approved customer-safe explanation and safe next
step; candidate rationale, GPT advisory details, and internal detection stay
in the operator review and audit surfaces. If a candidate is rejected, its
reason remains audit provenance and future routing stays unchanged.
Run the bounded portfolio showcase locally:
npm run demo:knowledge-evolution
npm run demo:knowledge-evolution -- --verboseIt uses the controlled provider, so it requires no API key or network access.
The output shows deterministic discovery, an advisory candidate draft, the
explicit human approval boundary, the promoted v1 object, and the
corresponding audit actions. It then evaluates a future matching ticket before
promotion, after promotion while required evidence is still missing, and after
that evidence arrives. The final line proves that the historical
pre-promotion recommendation is byte-for-byte unchanged. Add --verbose to
show the sanitized supporting diagnoses, tickets, evidence IDs, scores,
provenance, similarity reasons, and per-action audit detail. The fixture
includes two matching completed diagnoses, an unrelated completed diagnosis,
and an open-ticket corroboration so the output shows the difference between
confirmed support and an early signal. This is suitable for a repeatable
recording or screenshot; a live OpenAI run remains an optional separate
evaluation.
Lifecycle Replay Viewer
The read-only Lifecycle Replay page makes evaluation output inspectable in the same customer context as the Approval Desk. Run an evaluation first, then start the local browser server:
npm run evaluate:ai-comparison -- --live # optional; controlled output also works
npm run demo:approval-deskOpen /lifecycle-replay on the printed local URL. The page groups snapshots by
ticket, shows customer replies and previous support responses, and lets you
compare deterministic and GPT-labelled lanes. Operator view includes
classification agreement, quality breakdown, failure reasons, and sanitized
provider provenance. Customer view shows only the draft that a customer would
see and the explicit approval pause.
Replay reads reports/ai-comparison/live-latest.json (and the controlled report
when present); it makes no OpenAI calls, sends no responses, and never mutates
ticket or audit state. Snapshots are labeled evaluation states rather than an
invented chronological journey, so the viewer does not imply that unrelated
scenario runs happened in a particular order. See
the replay viewer guide for a portfolio walkthrough.
Extension To Zendesk Or Jira
No live connector is included. A future adapter can preserve the current governance model by:
Mapping external ticket fields into the validated
Ticketcontract while retaining the external ID separately.Implementing read adapters for tickets and knowledge without exposing credentials or raw provider errors through MCP.
Keeping recommendations in a local or durable pending store separate from provider mutation.
Translating only explicitly approved named fields into provider updates with revision or version checks and idempotency keys.
Writing an audit event that records the external request identifier and outcome without secrets or full customer content.
Adding provider-specific authorization, rate limiting, retry, webhook verification, and reconciliation tests.
The approval gate should remain above the provider adapter. Ticket text, webhook payloads, provider comments, and imported macros remain untrusted.
Limitations And Residual Risks
Fixtures and knowledge are synthetic and local. The server has no network integration or identity boundary.
Similarity is token-based Jaccard scoring, not semantic retrieval. It can miss paraphrases and produce lexical false positives.
Policy is deterministic and intentionally narrow. Human review remains necessary for ambiguous facts, conflicting policy, and customer messaging.
Operational tickets, conversations, recommendations, and their audit data still use local JSON/JSONL repositories. Knowledge candidates, immutable versions, knowledge audits, and learning events use the SQLite learning ledger, which is designed for one local process rather than distributed writers.
Locks are in-process. Ticket update, recommendation resolution, and audit append include compensation paths, but they are not cross-process ACID transactions.
Local users with filesystem access can edit or delete tickets, recommendations, knowledge, and audits.
Linked-path checks reject symbolic links and multi-link files. Node pathname APIs cannot fully prevent a hostile concurrent Windows parent-junction swap.
Directory
fsyncis best effort because it is not supported consistently on Windows. Rename, hard-link publication, antivirus scanning, sync clients, and filesystem behavior can affect durability and startup.Unexpected tool errors are generic to the MCP client, while diagnostic details are written to local standard error. Do not forward those logs to an untrusted destination.
Fixture SLA deadlines are fixed on June 10, 2026. Runs after that date classify due open tickets as breached unless an explicit historical
asOfvalue is used withlist_tickets.The official Python Skill validator was run in the current Skill evaluation and reported
Skill is valid!.test/skill.test.tsadds narrower structural regression checks.
Repository Guide
src/server.ts: MCP tools, resources, prompts, annotations, and safe errorssrc/triage-service.ts: submission, approval, rejection, and compensationsrc/policy.ts: escalation and approved-field rulessrc/metrics.ts: queue metrics and savings formulasrc/evaluation.ts: deterministic evaluation metricsdata/seed/: tickets, expected outcomes, and sample recommendationsdata/knowledge/: local policy and troubleshooting articles.codex/config.toml: project MCP launch configuration.agents/skills/triaging-support-tickets/: Codex Skill and policy reference
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- Alicense-qualityCmaintenanceA deterministic MCP server for legal intake triage that provides practice-area lookup, conflict screening, matter validation, follow-up drafting, and triage logging with a hard conflicts gate.Apache 2.0
- AlicenseCqualityCmaintenanceIncident Triage MCP is a Model Context Protocol (MCP) server for incident triage. It provides safe, auditable tools for evidence retrieval, deterministic summaries, ticket workflows, and notifications.28Apache 2.0
- Flicense-qualityCmaintenanceA governed MCP server for integrating AI agents with customer data, featuring role-based access control, field redaction, and human-in-the-loop approval for secure support operations.1
- AlicenseAqualityBmaintenanceA local-first MCP server for retrieving a small evidence set and recording reviewed conclusions, policy-gated and redacted without giving an agent general filesystem access.5MIT
Related MCP Connectors
A paid remote MCP for CLI tool MCP, built to return verdicts, receipts, usage logs, and audit-ready
A paid remote MCP for HyperFrames, built to return verdicts, receipts, usage logs, and audit-ready J
A paid remote MCP for hosted MCP server, built to return verdicts, receipts, usage logs, and audit-r
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/MatiasLaukka/Support-Ticket-Triage-archive'
If you have feedback or need assistance with the MCP directory API, please join our Discord server