Test how your application handles a Cosmos DB outage. Run Gremlin network experiments and reliability tests to validate multi-region failover.
Azure Cosmos DB gives you controls most databases don't: consistency level, partition strategy, multi-region write configuration, and more. But these are often configured early in the deployment process, yet they determine how your application recovers from failures. You might not know how your database will handle a region outage until it happens.
Gremlin helps you find out whether your configuration matches your expectations. With Gremlin, you can:
- Validate your failover policies. Use a blackhole experiment to make a regional endpoint unreachable. You'll learn whether your service reconnects to a healthy region, or fails.
- Check what your consistency level actually costs. Teams that choose session consistency for its performance might not know what to expect after a region failover. Latency experiments surface replication lag so you can see whether stale reads impact your users.
- Hold every service to the same standard. Gremlin detects Cosmos DB as a dependency wherever it appears. Run our suite of recommended tests on each service that relies on Cosmos DB, then use the resulting reliability scores to prioritize your riskiest services.
Dependencies: Failure Test
Simulate a failed dependency by dropping all network traffic to the dependency.
Dependencies: Latency Test
Recreate poor network conditions by delaying all network traffic to a dependency by 100ms.
Unreliable Networks - Dependencies
Simulate unreliable network conditions when communicating with dependencies by adding latency to API calls. Test whether your users are affected when dependency response times degrade.
* Testing a dependency doesn't require installing anything on it. Gremlin runs the experiment from the service that calls it, so compatibility depends on that host rather than on the dependency itself. Check our compatibility documentation for supported operating systems and platforms, or get in touch if you don't see yours.
Intelligent Health Checks - Azure
Automatically monitor your Azure services during testing in one-click with Intelligent Health Checks.
Avoid downtime. Use Gremlin to turn failure into resilience.
Gremlin empowers you to proactively root out failure before it causes downtime. See how you can harness chaos to build resilient systems by requesting a demo of Gremlin.
