AI SRE is having a moment. The category pulled in massive funding rounds over the last two years, Gartner published its first market guide, and vendors are promising everything from 90% faster resolution to fully autonomous incident response. If you run an engineering organization, someone has probably pitched you an AI SRE in the last quarter.

And let’s be honest: faster triage, less alert fatigue, and automated frontline response are wins for understaffed teams. But that doesn’t mean they’re a one-stop solution to all your reliability problems. While AI SREs can respond incredibly quickly, there are still five gaps in their capabilities, and if you don't understand those gaps before you deploy, you'll find out during an outage.

1. AI SREs only act after failure has started

No matter how sophisticated, AI SREs will only engage when something is going wrong. That's not a criticism—it's the job description. These tools watch your telemetry, detect anomalies, correlate signals, and respond. The whole point is for them to respond as quickly as possible.

But that means by the time your AI SRE is up to speed, your customers are already feeling the impact. Yes, a 10-minute outage is better than a 20-minute outage, but an outage is still an outage. Let’s quickly do the availability target math: Four nines gives you less than an hour of downtime for the entire year, or about four minutes a month. Even at ten minutes, that outage completely destroys two months of budget. 

Faster incident response is important, but it’s still just a better version of the strategy engineering teams have always had: wait for things to break, then fix them quickly. It compresses the timeline without changing the strategy.

2. AI SREs can’t predict sudden failures

AI SREs sometimes claim to prevent incidents by spotting early warning signs before problems become full-blown outages. This can, in fact, serve as an early-detection system for some kinds of outage. But it assumes that all failures start small and escalate, giving you a window of detectable pre-signal to act on.

There are failures that behave this way, such as when a memory leak grows, error rates creep up, or latency degrades over hours. But there are plenty that don’t. What about if a certificate expires? Or a config change takes out a load balancer? Or dependency suddenly hard-fails during a deployment? These failures drop out instantly without any pre-signal or warning.

Sadly, these failures are behind some of the biggest and most painful outages. Once they happen, AI SREs can react quickly, but they’re simply unable to prevent these failures.

3. AI SRE’s don’t tests for failure

To be fair, the better AI SREs do more than wait for pages. They triage alert floods, rank emerging signals by likely impact, and catch some problems while they're still small. There is value there, but it’s still detection and reaction rather than actively uncovering risks.

What they don't do is test anything. An AI SRE won't validate that your failover actually fires before you need it, confirm your service survives a zone outage, or check that last month's dependency change didn’t break your redundancy. The industry has spent decades cataloging how systems fail—dependencies time out, zones go down, failovers don't fire—and the tests that expose those failure modes are well understood. A detection-based tool simply has no way to run those tests.

4. AI SRE root cause analysis is based on extrapolation

The leading AI SREs are investing heavily in root cause analysis—world models of your production environment, causal reasoning over telemetry and deploy history. But even the best RCA drawn from telemetry is an inference about what probably happened, reconstructed from the signals the incident left behind. Vendors themselves cite accuracy rates in the low 80s. While genuinely impressive, it also means roughly one in five diagnoses is wrong, with no reliable way to know which one you just got.

Potential problems really emerge if you take it one step further. An explanation for this incident doesn't tell you whether the same weakness exists and is waiting to fail in the eleven other services that share the pattern. A model built from passive signals knows your topology, but it doesn't know how your systems actually behave under failures they haven't had. Without that, every incident stays an isolated event instead of becoming a lesson, and your reliability posture improves one outage at a time.

5. AI SREs can't prove the fix worked

When an AI SRE remediates an incident, how do you know the fix actually worked? In most cases, the evidence is that things came back, but that’s a signal for successful triage, not effective repair The AI restarted the instance, rerouted the traffic, or rolled back the deploy, then the graphs went green. But did it fix the underlying problem, or did it just reset the conditions until the same failure happens again?

Think of it as duct tape on the hull of a ship. Each patch works, in the moment. But nobody goes back to check whether the hull was actually repaired, and unverified fixes accumulate. Eventually you don't have a small leak—you have a hull made of duct tape, and you find out how strong it is at the worst possible time. You need a way to verify that the same failure won’t just happen again, and that requires testing.

What Foresight AI does instead

Gremlin’s Foresight AI uncovers risks before the failure, when things are calm and you can do prudent engineering instead of emergency surgery.

Foresight AI is built on 10 years of cause and effect resilience testing data, so it knows the failure modes that take down real systems and recommends the tests that safely expose them in yours. The known failure conditions get found and fixed on your schedule, not during an incident.

From there, it explains what went wrong, why it went wrong, and how to approach a fix (often even providing the exact code or config to change), all grounded in observed failure behavior, not reconstructed from telemetry after the fact. Then it runs tests to verify the fix. And because Gremlin can reproduce the failure condition safely—with controlled blast radius and automatic halt conditions—you don't have to hope the remediation worked. You run the failure again and prove it. 

Better together: give your AI SRE more to work with

None of this means you shouldn't run an AI SRE. It means the two approaches solve different problems, and each one makes the other better.

There’s a clean division of labor:

  • Foresight AI handles the known: it systematically finds the failure conditions you can anticipate—dependency failures, zone outages, broken failovers, resource exhaustion—and helps your team fix them before they become incidents. Along the way, it builds a real understanding of how reliable your systems are and how they behave under failure, shared across teams instead of siloed in a few heads. 
  • Your AI SRE handles the unknown: the novel failures and edge cases that no amount of testing will fully eliminate. For that slice, fast AI-assisted response is exactly the right tool—and it's a much smaller slice once the known failure modes are handled.

The architecture should include prevention as the primary strategy, response as the backup for what slips through. Most organizations today have it inverted—reactive response is the default and proactive work is the aspiration.

Get ahead of the failure

Every dollar spent making incident response faster is a bet that the incident will happen. Sometimes that's the right bet, and when it is, an AI SRE is a strong way to make it.

But the failures you can find in advance—the ones you already suspect are lurking in your failover configs, your dependency chains, your Kubernetes clusters—those shouldn't be incidents at all. Prudent engineering earlier in the cycle, in dev and in production while things are calm, shrinks the surface area your AI SRE has to cover and turns your reactive layer into what it should have been all along: the edge case, not the everyday.

Your AI SRE gets better the less it has to do. Foresight AI is how you get there.

Ready to see how Foresight AI can help you leverage AI to get ahead of incidents and outages? Join the waitlist for early access.

No items found.
Start your free trial

Gremlin's automated reliability platform empowers you to find and fix availability risks before they impact your users. Start finding hidden risks in your systems with a free 30 day trial.

sTART YOUR TRIAL
Book a demo

Schedule a time with a reliability expert to see how reliability management and Chaos Engineering can help improve the reliability, resilience, and availability of your systems.

Schedule now
Ryan Detwiller
Ryan Detwiller
VP, Marketing