Gremlin and AWS S3

Test how your application handles an AWS S3 outage. Run Gremlin network experiments and reliability tests to validate timeouts and fallbacks.

Why S3 reliability is important

Amazon S3 has a durability record so strong that teams design around it without designing for it. Calls get written without timeouts, without fallbacks, and often without error handling beyond a log line, because the service is assumed to always answer.

Durability and availability are separate guarantees. Your objects can be perfectly safe and still temporarily unreachable, and an application that never considered that possibility handles it badly. Requests hang instead of failing. Threads tie up waiting. A page that needs one image waits on a call that will never timeout, leaving your customers to look at a blank page.

Gremlin finds these behaviors before your customers do. Run a blackhole experiment to validate that your services can tolerate an S3 outage. Run a latency experiment to ensure you can render the rest of your application while waiting for the object in the background. The fix is usually small—a timeout and a placeholder—but it can make a huge difference.

Building resilience on S3 with Gremlin

S3 is a dependency: a managed service that your service connects to. You can run tests by deploying the Gremlin agent or Failure Flags sidecar to a service that consumes S3*. Using the experiments shown to the right, you can prepare your service for S3 failure modes, including:

  • Network outages making S3 unavailable
  • Slow performance due to network latency
  • Expiring TLS certificates
You can also run these expert-built workflows designed to replicate real-world failure modes on S3:

* Testing a dependency doesn't require installing anything on it. Gremlin runs the experiment from the service that calls it, so compatibility depends on that host rather than on the dependency itself. Check our compatibility documentation for supported operating systems and platforms, or get in touch if you don't see yours.

resources

Learn more about Gremlin and S3

All product names, logos, and brands are property of their respective owners. AWS is a trademark of Amazon.com, Inc.; Azure is a trademark of Microsoft Corporation; Google Cloud is a trademark of Google LLC. Use of these names is for identification purposes only and does not imply endorsement or affiliation unless otherwise stated.

Avoid downtime. Use Gremlin to turn failure into resilience.

Gremlin empowers you to proactively root out failure before it causes downtime. See how you can harness chaos to build resilient systems by requesting a demo of Gremlin.

Product Hero ImageShape