Gremlin and AWS EC2

Test EC2 reliability with Gremlin. Run CPU, Memory, Disk, and Shutdown experiments plus built-in reliability tests to validate autoscaling and failover.

Why EC2 reliability is important

Amazon EC2 runs a significant portion of workloads on AWS, which means EC2 failures rarely stay contained to EC2. A single lost instance can cascade into slow response times, failed transactions, and lost revenue long before your monitoring tools show anything unusual.

Track down those failures and prevent them from causing outages with the Gremlin Reliability Platform. Install the Gremlin Agent on your instances, then use Chaos Engineering experiments and standardized Reliability Test Suites to prove your services survive the failures AWS will eventually hand you. With Gremlin, you can:

  • Prove your services survive instance loss. AWS retires hosts, degrades hardware, and reclaims Spot capacity on its own schedule. Shutdown experiments let you take an instance offline on your terms and confirm that traffic reroutes, replacements launch, and your customers never notice.
  • Validate the Multi-AZ architecture you're paying for. Redundancy on an architecture diagram is not the same as redundancy under load. Blackhole experiments isolate an availability zone so you can watch the remaining zones absorb the traffic, turning an assumption into evidence you can show your leadership team.
  • Find resource limits before your customers do. CPU, memory, and disk experiments apply pressure to an instance while you watch how your service responds. You learn where autoscaling triggers, where health checks fire, and where your thresholds were set years ago and never revisited.

Building resilience on EC2 with Gremlin

EC2 is an infrastructure service, which means it supports the Gremlin agent*. This lets you run the full suite of Gremlin experiments. Using the Gremlin agent, you can find deep, infrastructure-level reliability risks in your EC2 workloads, such as:

  • Scaling problems due to missing or misconfigured autoscaling rules
  • Processes that don't auto-recover after a crash
  • Limited redundancy on the host and network level
You can also run these expert-built workflows designed to replicate real-world failure modes on EC2:

* Gremlin is designed to work on any cloud platform that provides Linux or Windows hosts. We haven't individually tested every service we cover, and not all are officially supported. Check our compatibility documentation for tested operating systems and known caveats, or get in touch if you don't see yours.

resources

Learn more about Gremlin and EC2

All product names, logos, and brands are property of their respective owners. AWS is a trademark of Amazon.com, Inc.; Azure is a trademark of Microsoft Corporation; Google Cloud is a trademark of Google LLC. Use of these names is for identification purposes only and does not imply endorsement or affiliation unless otherwise stated.

Avoid downtime. Use Gremlin to turn failure into resilience.

Gremlin empowers you to proactively root out failure before it causes downtime. See how you can harness chaos to build resilient systems by requesting a demo of Gremlin.

Product Hero ImageShape