
Why AI development creates a reliability blind spot for humans, and what to do about it
Part 1 of The Intellyx Building Agentic Resilience Series by Jason English, for Gremlin
Application development and operations teams are adopting AI coding tools at an exponentially increasing rate, from a rounding error of 6% of code output by AI in 2023, to as much as 51%-75% for a majority of enterprises.
Agents are making pull requests (PRs) faster than human developers could have ever dreamed. As features are pushed to market faster, there’s a sharp increase in production incidents, with 80% of development shops specifically tracing production outages to AI.
If AI coding agents were expected to deliver software 10 times faster, those results haven’t shown up in the industry surveys, as another useful benchmark notes only a 20% boost in successful PR velocity, with change failure rates up 30%.
Looking across these reports from leading software quality, observability, and security testing firms, we can see an industry-wide drag on the expected AI productivity boost for software development, as the rate of failures increases linearly with the increased volume of AI code delivered, and the time spent responding to production outages increases.
But despite the unknown risks of an adolescent market, companies have already opened the garage doors for AI coding agents and handed them the keys to the Ferrari.
Clearly we are missing something: the ability to reliably keep AI-driven software releases on track and avoid costly incidents that will slow our delivery progress in the future.
Looking for root cause through observability and SRE lenses
The software industry has made great strides in advancing observability, which looks at streams of system telemetry such as logs and metrics to understand the inner workings of software as it runs, and comparing real-time data to historical data to spot anomalies and alert human teams about concerning trends.
For years, the holy grail of observability was using the golden signals of latency, traffic, errors, and saturation to discover the root cause of any perceived failure. A new role of SRE (site reliability engineering) emerged, consisting of highly skilled individuals who were very well-versed in the architectural details of a system, so they can drop into an on-call incident and hopefully use an esoteric suite of analytics and tools to find that root cause and implement a fix.
One early shortcoming of observability was being able to “cause the cause” and reproduce the conditions under which a failure occurs. Observability and testing vendors built real user monitoring (RUM) and synthetic data to help reproduce these failure conditions in pre-production environments.
A new class of “AI SRE” solutions has now entered the market focused on the detection, triage, and AI-assisted remediation of issues, so teams can respond to them as quickly as possible. Every team wants to be responsive and reduce all the MTTs such as MTTI (mean time to identification), MTTA (acceptance), MTTR (resolution) and so on.
While very useful today, many of these solutions are still reactive instead of proactive. They are looking in the rear view mirror at things that have already gone wrong. But what should we do about the huge blind spot of unexpected problems that we haven’t yet observed?
An alternative approach to proactive validation
Taking another tack, an open source project called Chaos Monkey was contributed by Netflix in 2012, which popularized the concept of chaos engineering: automated failure injection at different levels of a software stack in order to test its resilience, induce unexpected results and see what happens.
Gremlin was an early commercial vendor in this new space with their fault injection solution. Over several years, they gathered extensive data on exactly how any component within an application environment can misbehave or fall down in production when pushed by the traffic of a chaotic world, which enabled them to build an automated fault detection solution.
These millions of failure data points Gremlin collected would come in especially handy as a corpus of data for training AI, a “Failure Atlas” if you will, of observed causes and effects of failures in the software world.
Here’s where agentic AI usefully enters the story, with semi-autonomous fault injection and fault detection agents that can use the context of the Failure Atlas to continuously look for potential scalability and failure conditions within the software delivery pipeline and recommend safe fixes to avoid failures, before they have even occurred in production.
Building an agentic resilience loop with Foresight AI
What if AI agents could actually see our application stack with the vision of a software architect and the practical debugging knowledge of an expert SRE?
Today, Gremlin introduced their new Foresight AI solution, now in general release, which provides agentic AI guardrails to keep development projects and engineering changes from introducing failures that can escape into production.
They built an AI harness for agentic development, populated with the contextual data, tools and skills necessary to conduct actionable reviews and provide demonstrable advice about how to ensure reliability.

Foresight AI accomplishes this by instantiating an agentic resilience loop, where specialized agents for data gathering, analysis, testing, operations, and program management continuously work together in preproduction to find potential failures, automate or recommend fixes, and generate the conditions that prove the case for making significant code and configuration changes.
Rather than putting a “human in the loop” of reviewing the details of each potential change, which can cause reviewer fatigue and basically erase the velocity benefits of agentic software development, this system puts humans in the driver’s seat, focused on navigation – keeping their development agents delivering features with less prospective risk on the road ahead.
While most of the action behind the scenes here is agents talking to each other, human users can interact with Foresight AI through a familiar LLM-style chat interface, which is ideal for displaying and documenting the output of several other specialized agents and deterministic processes.
The Intellyx Take
Agentic loops aren’t novel to software development. In the cybersecurity world, we are already seeing parallel proofs for such loops. Specialized agents working together to conduct attack simulations in emulated environments and recommend remediations, in order to relieve security operators from deluge of alerts so they can focus on resolving the most critical vulnerabilities.
While the power of a solution like Foresight AI may seem like a panacea, reliability engineering isn’t just a software package a company can just turn on, nor is it just an extension of platform engineering or continuous delivery practices.
Sustained and repeatable success on this new course still depends upon a shared program management vision and accountability between engineering and executive stakeholders. We need to change our mindset about risk, in order to justify the continuation of an agentic resilience initiative and safely reach our destination.
©2026 Intellyx B.V. Intellyx is editorially responsible for this document. At the time of writing, Gremlin is an Intellyx customer. AI was not used to write a single word of this content. Image source: Gremlin (feature), Gremlin (diagram).
Start your free trial
Gremlin's automated reliability platform empowers you to find and fix availability risks before they impact your users. Start finding hidden risks in your systems with a free 30 day trial.
START YOUR TRIAL
Why agentic AI development needs reliability guardrails
Companies are moving faster than ever with agentic AI, but that means more risks. Without reliability guardrails, they risk costly outages.


Companies are moving faster than ever with agentic AI, but that means more risks. Without reliability guardrails, they risk costly outages.
Read more