When people ask about Smart Desk, they usually want to talk about the AI model. That's the least interesting part of the build. The architecture decisions around it — how tickets flow in, what happens when the model is wrong, how escalation actually gets triggered — are what determined whether this became a production tool or another abandoned pilot.
The Starting Point
The brief was simple: reduce manual ticket triage effort without disrupting the existing ServiceNow workflow or asking analysts to learn a new tool. That constraint shaped almost every decision that followed.
Architecture: Where the AI Actually Sits
Smart Desk uses n8n as the orchestration layer between ServiceNow and Claude API. Tickets land in ServiceNow exactly as they always did. An n8n workflow trigger picks up new tickets, sends the relevant content to Claude for classification, and routes the result back into ServiceNow's existing assignment and priority fields — so from an analyst's point of view, tickets just arrive better-classified, with no new tool to open.
If your AI layer requires the end user to change their workflow, you've added a tool, not an improvement. Smart Desk's biggest design win was being invisible to the people it helped most.
The First Failure Mode: Overconfidence
Early versions auto-resolved too aggressively. A handful of edge cases got incorrectly closed before a human could catch the misclassification — not a disaster, but exactly the kind of trust-eroding mistake that gets a tool shut off. The fix wasn't a better prompt; it was a stricter confidence threshold and a much narrower initial scope, limited to ticket categories with very clear, low-ambiguity patterns.
The Second Failure Mode: Silent Drift
Classification accuracy that looks great in week one can quietly degrade as ticket patterns shift — new product launches, seasonal spikes, org changes. Smart Desk now includes a lightweight accuracy-tracking loop that flags when auto-resolution rates or override frequency move outside expected ranges, so drift gets caught before it becomes a trend.
What I'd Do Differently
I'd build the accuracy-tracking and drift-detection loop from day one rather than retrofitting it after the fact. Everything else — the narrow initial scope, the confidence thresholds, the invisible integration — I'd keep exactly as it was.
If you want the deeper operational case study with full before/after metrics, it's written up in detail here.