The service desk is one of the highest-volume, highest-repetition environments in IT. Thousands of tickets per month. Dozens of agents fielding the same questions. Manual triage. Copy-pasted RCA templates. This is exactly the environment where well-designed prompts deliver outsized value.

I've spent the past year embedding GenAI into ITSM workflows — using Claude, ChatGPT, and Gemini in different configurations. What follows is the practical prompt architecture that actually works in production, not in demos.

The Three Prompt Layers

Effective ITSM prompt engineering operates across three layers, each building on the last:

Layer 1: Classification Prompts

The first job of any GenAI ITSM integration is accurate ticket classification. A well-structured classification prompt should output a structured JSON object, not a narrative response.

Example Classification Prompt

You are an ITSM triage specialist. Analyse the following ticket and return ONLY a JSON object with: category (one of: hardware, software, network, access, other), priority (P1–P4 using ITIL criteria), suggested_team, confidence_score (0–1), and reasoning (max 30 words). Ticket: [TICKET_TEXT]

The key design decisions here: force structured output, constrain the category set to your actual service catalogue, and request a confidence score so your workflow can route low-confidence tickets to human review.

Layer 2: Enrichment Prompts

Once a ticket is classified, enrichment prompts add operational context — pulling from your knowledge base, CMDB, and recent incident history. This is where RAG (Retrieval-Augmented Generation) comes in, which I'll cover in a separate article.

A basic enrichment prompt might ask the model to:

Layer 3: Communication Prompts

The third layer handles user-facing and stakeholder communications. This is where tone and brand voice matter. Your service desk has a communication standard — the LLM should follow it.

"The best ITSM prompt is the one your agents never have to rewrite."

Common Mistakes I've Seen

Vague Instructions

Prompts like "summarise this ticket" produce inconsistent output. Always specify format, length, audience, and constraints. The more specific your prompt, the more reliable your output.

Ignoring Edge Cases

P1 incidents, VIP users, and security-related tickets need special handling. Build explicit conditional logic into your prompts: "If the ticket contains keywords related to security or data breach, set priority to P1 regardless of other factors and add flag: SECURITY."

No Human Override Path

Every AI-generated action needs a visible override mechanism. Agents must be able to see what the AI recommended, why, and change it with one click. This isn't just good UX — it's how you build the feedback data that improves the model over time.

42%
Tickets Auto-Resolved
90s
AI Triage vs 8min Manual
2.5h
Agent Time Saved/Day

Starting Small

You don't need a platform overhaul to start. The fastest path to production value is a single, well-defined use case: auto-generating the first response to password reset tickets, or summarising incident timelines for post-incident reviews.

Pick one high-volume, low-risk ticket type. Design a prompt. Test it on 100 historical tickets. Measure accuracy against human-labelled outcomes. Iterate. That's the prompt engineering loop that scales.