After twenty years in IT operations — managing enterprise infrastructure across APAC, EMEA, and the Americas — I've developed a finely tuned radar for industry hype. I've watched ITIL get over-complicated, cloud get over-promised, and DevOps get over-simplified. So when AIOps started appearing in every vendor deck around 2022, my instinct was healthy scepticism.

That scepticism has since been replaced by something more useful: a framework for separating genuine AIOps maturity from expensive dashboards with a machine learning veneer.

What AIOps Actually Is

The term was coined by Gartner in 2017 to describe platforms that combine big data and machine learning to enhance and partially replace IT operations processes. The operative word is enhance. AIOps is not a replacement for ITIL, not a substitute for skilled engineers, and definitely not a magic box that eliminates on-call rotations.

At its core, genuine AIOps does three things well:

"The measure of AIOps maturity isn't how many alerts it suppresses. It's how many incidents it prevents."

The Three Failure Modes I've Seen

1. Telemetry Poverty

AIOps is only as intelligent as the data it ingests. Organisations that deploy AIOps on top of fragmented, inconsistent monitoring data end up with highly confident predictions about the wrong things. Before any AI layer, you need unified observability: metrics, logs, traces, and events in a single data fabric. Dynatrace and Splunk are excellent — but only when properly instrumented.

2. Workflow Isolation

AIOps tools that don't integrate with your ticketing, change, and configuration management systems create a parallel universe. Insights stay inside the AIOps platform. Remediation still happens manually. The AI becomes a very expensive alert dashboard.

3. No Feedback Loop

Machine learning models degrade without reinforcement. If your AIOps platform doesn't learn from human override decisions — when engineers dismiss predictions, escalate incidents, or route differently — it stops improving. Operationalising the feedback loop is the most overlooked engineering challenge in AIOps deployments.

Key Takeaway

True AIOps maturity requires three prerequisites: high-quality telemetry across all infrastructure layers, bidirectional integration with ITSM workflows, and an active feedback loop that continuously trains the model on real operational decisions.

What Good Looks Like

In my work at AVAYA, we achieved 99.9%+ availability across enterprise managed services accounts by combining Dynatrace observability with structured incident workflows. The shift from reactive monitoring to predictive operations reduced mean-time-to-detect by over 60% on our highest-complexity accounts.

99.9%
SLA Achieved
60%
MTTD Reduction
95%+
P1 Resolution Rate

The difference wasn't the tool — it was the architecture. Telemetry was standardised across all accounts. Alerts fed directly into ServiceNow with auto-populated CMDB context. Every incident closure included a structured outcome tag that the system learned from.

Where GenAI Changes the Equation

The arrival of large language models has added a new layer to AIOps that deserves separate attention. LLMs don't replace the pattern-recognition and anomaly-detection functions of traditional AIOps platforms. But they dramatically improve the human interface.

Specifically, GenAI excels at:

The combination of traditional AIOps (pattern detection, correlation, prediction) with GenAI (language, synthesis, communication) is where the real productivity multiplier lives. Neither alone is sufficient.

The Honest Assessment

AIOps is not a buzzword — but it is frequently sold as something it isn't. The technology is real. The value is provable. The failure rates in enterprise deployments are high because organisations skip the foundational work: clean telemetry, integrated workflows, and disciplined model governance.

If you're evaluating AIOps investments, my recommendation is to spend 60% of your budget on data infrastructure and integration — and 40% on the AI platform. The inverse ratio is what most organisations get wrong.


In the next article, I'll cover how to structure an AIOps proof-of-concept that generates board-ready ROI data within 90 days — without disrupting your production environment.