// Reference Architecture · Interactive
AI Resilience by Design
Designing enterprises that can evolve with AI — and operate without it. An interactive reference architecture for AI-driven IT service operations, built for CIOs, enterprise architects, IT operations, risk and AI governance teams.
// The Idea
Resilience Is a Property of the Business, Not the Model
AI resilience is not about having a backup AI model. It is about designing the business so that AI can evolve, fail, disappear or be replaced without bringing the business to a halt.
Most AI resilience planning asks how to keep the model available: a second provider, a fallback endpoint. That helps with outages. It does nothing for a model that is up, fast and confidently wrong, a vendor that changes pricing or behaviour, a regulator that restricts a model, or an agent that is given more autonomy than the design assumed.
This reference architecture starts from a different question: what must the business keep doing when the AI underneath it changes? Business processes, decisions, knowledge and automation are separated from AI capability, so the capability can be swapped, degraded or switched off while service outcomes continue.
// Try It
The Interactive Model
Switch between AI-on, AI-degraded and AI-off, inject failures, replace a model through evaluation gates, and score your own environment. The preview below is the full model; open it full screen for the best experience.
Open Full Screen →
Scenario figures, thresholds and evaluation data are illustrative examples, not benchmarks of any product.
// The Architecture
Ten Layers, With the Business on Top
Every layer has a defined behaviour when AI fails. The business layers never reference a model or a vendor; the AI capability layers are designed to be replaced.
Business Services
→
Business Process Layer
→
Decision & Policy Engine
→
AI Gateway
→
Model Router
→
Models (A, B, Local)
→
Knowledge / RAG
→
Agent & Tool Orchestration
→
Human Control
→
Automation / ITSM / Infrastructure
Observability, security, governance, audit and resilience run across every layer. Governance is a runtime capability here: policy is evaluated on each call rather than described in a document.
// What the Model Demonstrates
Six Ideas, Each Interactive
1. The AI dependency trap
Thirteen realistic failures (provider outage, latency, hallucination, vector database loss, prompt regression, vendor pricing change, regulatory restriction, incorrect remediation and more) shown side by side against a tightly coupled design and a resilient one.
2. AI on, AI degraded, AI off
Every business-critical capability (receive incidents, monitor, run runbooks, escalate, access knowledge, approve changes, communicate, restore services) has a defined way to be performed in every mode. AI off is not business off.
3. Wrong AI, not just down AI
Availability, reliability and trustworthiness are separated. Confidence thresholds, grounding checks, risk-tiered human approval and runtime policy contain an AI that is available but wrong.
4. Model and vendor replacement
A candidate model advances through evaluation gates, shadow testing and canary rollout, with the previous model kept as the rollback target. The vendor view lists honestly what stays portable and what does not, rather than claiming full vendor neutrality.
5. Technology evolution and scale
New model architectures, multimodal input, reasoning models, agents and edge inference enter as replaceable capabilities under stable business contracts. Controls such as agent registries and policies become required as users and agents grow.
6. A scorecard and an executive view
Fourteen dimensions produce an AI Resilience Maturity Score and seven CIO-level answers: can we operate without AI, replace our provider or model, detect wrong AI, stop automation safely, hand over to people, and absorb the next generation.
// Connection to My Work
Same Domain as Smart Desk and AI Incident Co-Pilot
The reference scenario, AI-powered IT service operations, is the domain of Smart Desk, AI Incident Co-Pilot and the Enterprise AI Knowledge Assistant. This model describes the design principles such tools need to survive provider changes, model drift and outages: AI recommends, policy decides, automation executes with checks, and people stay accountable for high-risk decisions.
AI GatewayModel RoutingRAGPolicy as CodeHuman-in-the-LoopAI ObservabilityRunbook AutomationITSM
// Common Questions
Frequently Asked Questions
Is this a product?
No. It is a reference architecture and interactive model. The numbers in it (thresholds, evaluation scores, costs) are illustrative examples; each enterprise sets its own.
Is the architecture future-proof?
No architecture can honestly claim that. This one is designed for technology evolution and change: it does not try to predict which model, vendor or technique will win, and it keeps those choices replaceable.
Does it depend on a particular AI vendor?
No. Applications talk to a gateway, not to a vendor. The model also states plainly what is not portable between vendors, such as embeddings, fine-tunes, contracts and exact model behaviour.
Can this be applied to our environment?
Yes. A good starting point is to score your current state in the model and review the weakest dimensions together. The free assessment and a 30-minute consultation are both available.
// Who This Is For
Built For Technology and Risk Leaders
CIOs and CTOs deciding how far to let AI into operations
Enterprise architects defining AI platform boundaries
IT operations leaders who must keep services running through AI changes
Risk and AI governance teams that need controls enforced at runtime