How a Leading Customer Experience Software Company Made Reliability an Org-Wide Standard with Gremlin

A global leader in customer experience software used Gremlin to migrate critical systems with zero incidents — then scaled resilience testing from a single project into an organization-wide standard embedded in their SDLC and internal developer platform.

Proved critical business data is safe

Simulated failures to verify zero data loss during outages

110
+

services onboarded so far

0

failures during migration to AWS Lambda.

60

business-critical integrations running Failure Flags

Executive Summary

A multi-billion dollar global leader in customer experience software with over 8,000 employees embarked on a migration of essential business systems that had zero margin for error. The systems had to launch on a specific date, couldn't have any downtime during the migration, and needed to retain all critical business data.

Using Gremlin Failure Flags, the IT engineering team launched on day one of the new fiscal year without a single incident. Additionally, by testing all services using a standardized group of tests, the team proved to business leadership that that critical business data was safe. That success became a foundation. Today, reliability isn't a per-project conversation at the company. It's the default for every service the organization builds.

This is some text inside of a div block.
The Challenge

How do you meet high availability needs for complex transactions?

Migrations are always complex, and the migration of the company's key financial and reporting systems was no exception. Made the highest priority by leadership, it came with stiff requirements: it had to happen on the first day of the new fiscal year, and it had to be reliable from day one.

The system dealt with critical financial data, which raised the stakes further. The team needed to prove to leadership that data wouldn't be lost in case of errors or outages, even if the failure originated with a third-party system. To add a layer of complication, the new system incorporated key serverless components built on AWS Lambda.

Faced with an inflexible deadline and the highest reliability standards, the IT team needed a way to detect and prevent outage-causing issues before launch so they could deploy with complete confidence.

The core problem we're solving is confidence with the business. The business needs to know that we have their data and that we're taking care of it."  
IT Director of Engineering Practices, Leading Customer Experience Software Company

‍

The Solution

Create a culture that addresses failures before they cause outages

The IT Director of Engineering Practices turned to Gremlin Failure Flags. Experienced with Chaos Engineering in previous roles, he knew the advantage of proactive reliability testing. Failure Flags—Gremlin's application-level testing for serverless and containers—let the team run resilience tests directly against their AWS Lambda functions.

The team developed a standardized list of failures based on experience and the causes of past outages, then refined it with further testing. Before launch, every service and Lambda call ran through those tests, and any issues were addressed. Shifting left this way meant errors surfaced in dev instead of production.

The result? They deployed the AWS Lambda functions in under five minutes without a single issue.

Of all the parts that went fine, the Gremlin stuff went the most fine. That was without incident. Flip the switch and spend 99 percent of the rest of the time doing everything else. And you can verify that."
IT Director of Engineering Practices, Leading Customer Experience Software Company

‍

How they built reliable and resilient critical business systems

1. Prove critical business data is safe from failure

When it comes to financial and invoice data, there is zero margin for error. Every system will experience failure at some point, especially when third-party dependencies are involved. The IT team built their new system to retain all data when a failure occurs—but designing software to work a certain way and knowing it will perform as intended are two different things.

Using Failure Flags, the team simulated failures and faults directly in their Lambda functions, including latency that leads to API call overload, network outages, and more. They even worked with third-party technical teams to verify that Gremlin's simulated faults accurately recreated past failures.

They resolved the reliability issues they found, confirmed no data would be lost during outages, timeouts, or latency, and—just as importantly—had the test results to show leadership.

‍

Here are 15 points where it can break in this process, and here's what happens in every one of them. We've proactively done it, and we will continue to proactively redo that weekly as part of our revalidation process."
IT Director of Engineering Practices, Leading Customer Experience Software Company

‍

2. Migrate with zero failures

Migrations are notoriously complex, often bringing transition periods marked by outages and degraded service. At this company, those outcomes were unacceptable.

The team built their test list from experience with AWS Lambda and past outages they and other companies had encountered. As services began passing those tests, they used their observability integration to run exploratory testing and uncover unknown failure conditions. Before long they had a comprehensive standardized suite that could run against every service—and when launch day arrived, the new Lambda services went live without incident.

‍

Of all the parts that went fine, the Gremlin stuff went the most fine. That was without incident. Flip the switch and spend 99 percent of the rest of the time doing everything else. And you can verify that."
IT Director of Engineering Practices, Leading Customer Experience Software Company

‍

3. Make reliability the default, not a project

A strong reliability posture takes more than a single successful launch. In the year since the migration, the team has scaled proactive resilience testing into the way the entire organization builds software.

Protecting the integrations the business runs on. The team expanded Failure Flags across roughly 60 business-critical integrations—the systems powering order management, quoting, revenue recognition, and billing. Every Lambda function behind them now runs through failure testing.

Hardening a new observability platform before launch. When the organization set out to build a centralized collector standardizing telemetry across every service, reliability was non-negotiable. The team ran failure experiments throughout the build, using Gremlin's out-of-the-box scenarios and failure points for their Kubernetes-hosted platform. Gremlin's automatic dependency detection proved especially valuable: when the team went to build custom scenarios for third-party dependencies, they found Gremlin had already identified them and run experiments against them—including scenarios they hadn't thought of. The platform has since been rolled out to more than 110 services without a single reliability issue.

‍

A lot of software recommends stuff. The fact that you just did it — you found it, ran the experiments, included it as part of the run, and showed us it was already taken care of — saved us a lot of time. Literally everything we'd thought about building as a custom scenario was already covered, and more."
Engineering Leader, Leading Customer Experience Software Company

‍

Building reliability into the SDLC and the platform. The team injected resilience testing directly into the organization's SDLC standards, approved for org-wide rollout, then embedded Gremlin into the internal developer platform itself. Every service hosted there is automatically annotated, appears in Gremlin, and is added to standardized test suites that run weekly against availability and latency targets. Teams without deep infrastructure expertise get enterprise-grade reliability automatically — application owners focus on business logic while the platform handles infrastructure resilience. The organization plans to onboard 75 services, with all new services built on the platform by default.

‍

The question becomes: are you on the platform, yes or no? If you're part of the platform, there are a lot of reliability standards baked in with Gremlin as part of the ecosystem."
Engineering Leader, Leading Customer Experience Software Company

‍

Getting there required more than tooling. To win executive buy-in, the team analyzed their business-impacting RCAs, identified how many stemmed from reliability gaps that proactive testing would have caught, and connected those gaps to metrics leadership already tracked—making the cost of inaction feel tangible.

"It has become easy to build software and very difficult to build reliable software," said the engineering leader. "Reliability in the agentic and AI world is very, very important, so that we don't deploy things we don't really understand. That's why putting reliability into the standard charter of how things get done is key, so there are gates that check how things go wrong before things go wrong."

By embedding Gremlin into their SDLC and developer platform, the organization has built exactly those gates—ensuring that as they move faster, the foundations stay in place.

Avoid downtime. Use Gremlin to turn failure into resilience.

Gremlin empowers you to proactively root out failure before it causes downtime. See how you can harness chaos to build resilient systems by requesting a demo of Gremlin.

Product Hero ImageShape