Gremlin and AWS EMR

Test AWS EMR reliability with Gremlin. Simulate node failures and resource pressure to validate big-data job resilience on EC2-backed clusters.

Why EMR reliability is important

EMR failures are expensive in a way that's easy to measure. A node lost partway through a six-hour job costs you the compute you already spent plus the compute to run it again, and it moves every downstream deadline that depended on that output.

Most teams assume their job framework handles node loss for them. Spark and Hadoop both have fault tolerance built in, but how much it protects you depends on checkpointing, replication settings, and how the job was written. Those details vary job to job, and few teams have verified them under actual node loss rather than in theory.

Testing gives you a concrete answer for the jobs that matter most: this one recovers, that one restarts from the beginning, this other one produces output that looks complete but isn't. Knowing which is which changes how you prioritize both engineering effort and cluster spend.

Building resilience on EMR with Gremlin

EMR is an infrastructure service, which means it supports the Gremlin agent*. This lets you run the full suite of Gremlin experiments. Using the Gremlin agent, you can find deep, infrastructure-level reliability risks in your EMR workloads, such as:

  • Scaling problems due to missing or misconfigured autoscaling rules
  • Processes that don't auto-recover after a crash
  • Limited redundancy on the host and network level
You can also run these expert-built workflows designed to replicate real-world failure modes on EMR:

* Gremlin is designed to work on any cloud platform that provides Linux or Windows hosts. We haven't individually tested every service we cover, and not all are officially supported. Check our compatibility documentation for tested operating systems and known caveats, or get in touch if you don't see yours.

resources

Learn more about Gremlin and EMR

All product names, logos, and brands are property of their respective owners. AWS is a trademark of Amazon.com, Inc.; Azure is a trademark of Microsoft Corporation; Google Cloud is a trademark of Google LLC. Use of these names is for identification purposes only and does not imply endorsement or affiliation unless otherwise stated.

Avoid downtime. Use Gremlin to turn failure into resilience.

Gremlin empowers you to proactively root out failure before it causes downtime. See how you can harness chaos to build resilient systems by requesting a demo of Gremlin.

Product Hero ImageShape