After twenty years in IT operations — managing enterprise infrastructure across APAC, EMEA, and the Americas — I've developed a finely tuned radar for industry hype. I've watched ITIL get over-complicated, cloud get over-promised, and DevOps get over-simplified. So when AIOps started appearing in every vendor deck around 2022, my instinct was healthy scepticism.
That scepticism has since been replaced by something more useful: a framework for separating genuine AIOps maturity from expensive dashboards with a machine learning veneer.
What AIOps Actually Is
The term was coined by Gartner in 2017 to describe platforms that combine big data and machine learning to enhance and partially replace IT operations processes. The operative word is enhance. AIOps is not a replacement for ITIL, not a substitute for skilled engineers, and definitely not a magic box that eliminates on-call rotations.
At its core, genuine AIOps does three things well:
- Event correlation — connecting thousands of alerts across siloed tools into meaningful incident narratives
- Anomaly detection — identifying deviation from baseline before humans notice a symptom
- Predictive routing — directing the right incident to the right team with pre-populated context
"The measure of AIOps maturity isn't how many alerts it suppresses. It's how many incidents it prevents."
The Three Failure Modes I've Seen
1. Telemetry Poverty
AIOps is only as intelligent as the data it ingests. Organisations that deploy AIOps on top of fragmented, inconsistent monitoring data end up with highly confident predictions about the wrong things. Before any AI layer, you need unified observability: metrics, logs, traces, and events in a single data fabric. Dynatrace and Splunk are excellent — but only when properly instrumented.
2. Workflow Isolation
AIOps tools that don't integrate with your ticketing, change, and configuration management systems create a parallel universe. Insights stay inside the AIOps platform. Remediation still happens manually. The AI becomes a very expensive alert dashboard.
3. No Feedback Loop
Machine learning models degrade without reinforcement. If your AIOps platform doesn't learn from human override decisions — when engineers dismiss predictions, escalate incidents, or route differently — it stops improving. Operationalising the feedback loop is the most overlooked engineering challenge in AIOps deployments.
True AIOps maturity requires three prerequisites: high-quality telemetry across all infrastructure layers, bidirectional integration with ITSM workflows, and an active feedback loop that continuously trains the model on real operational decisions.
What Good Looks Like
In my work at AVAYA, we achieved 99.9%+ availability across enterprise managed services accounts by combining Dynatrace observability with structured incident workflows. The shift from reactive monitoring to predictive operations reduced mean-time-to-detect by over 60% on our highest-complexity accounts.
The difference wasn't the tool — it was the architecture. Telemetry was standardised across all accounts. Alerts fed directly into ServiceNow with auto-populated CMDB context. Every incident closure included a structured outcome tag that the system learned from.
Where GenAI Changes the Equation
The arrival of large language models has added a new layer to AIOps that deserves separate attention. LLMs don't replace the pattern-recognition and anomaly-detection functions of traditional AIOps platforms. But they dramatically improve the human interface.
Specifically, GenAI excels at:
- Generating natural-language incident summaries from structured telemetry data
- Synthesising root cause analysis narratives from multiple correlated alerts
- Drafting post-incident reports that comply with ITIL Problem Management standards
- Coaching L1 agents through resolution steps in real time
The combination of traditional AIOps (pattern detection, correlation, prediction) with GenAI (language, synthesis, communication) is where the real productivity multiplier lives. Neither alone is sufficient.
The Honest Assessment
AIOps is not a buzzword — but it is frequently sold as something it isn't. The technology is real. The value is provable. The failure rates in enterprise deployments are high because organisations skip the foundational work: clean telemetry, integrated workflows, and disciplined model governance.
If you're evaluating AIOps investments, my recommendation is to spend 60% of your budget on data infrastructure and integration — and 40% on the AI platform. The inverse ratio is what most organisations get wrong.
In the next article, I'll cover how to structure an AIOps proof-of-concept that generates board-ready ROI data within 90 days — without disrupting your production environment.