Test how your application handles a Bigtable outage. Run Gremlin network experiments and reliability tests to validate pipeline retry and backoff logic.
Google Cloud Bigtable performance depends heavily on how your row keys distribute load. A hot tablet, where one server handles a disproportionate share of traffic, produces latency spikes that affect only some requests. That makes them hard to spot and easy to dismiss as noise.
For a high-throughput pipeline, a short period of elevated latency compounds:
- An uneven key distribution concentrates traffic on a single tablet.
- Reads and writes to that tablet slow down while the rest of the cluster looks healthy.
- Your pipeline keeps writing at full rate, because nothing has errored and nothing applies backpressure.
- The backlog grows faster than the slowdown lasts, and draining it takes considerably longer than the original disruption.
Gremlin reproduces this from your side of the connection. A latency experiment against the Bigtable endpoint slows reads and writes so you can watch whether your pipeline degrades gracefully or accumulates work it can't process.
For sustained high volume, the ability to apply backpressure matters more than the ability to retry, and it's the property least likely to have been built in.
Dependencies: Failure Test
Simulate a failed dependency by dropping all network traffic to the dependency.
Dependencies: Latency Test
Recreate poor network conditions by delaying all network traffic to a dependency by 100ms.
Unreliable Networks - Dependencies
Simulate unreliable network conditions when communicating with dependencies by adding latency to API calls. Test whether your users are affected when dependency response times degrade.
* Testing a dependency doesn't require installing anything on it. Gremlin runs the experiment from the service that calls it, so compatibility depends on that host rather than on the dependency itself. Check our compatibility documentation for supported operating systems and platforms, or get in touch if you don't see yours.
Intelligent Health Checks - GCP
Automatically monitor your Google Cloud services during testing in one-click with Intelligent Health Checks.
Avoid downtime. Use Gremlin to turn failure into resilience.
Gremlin empowers you to proactively root out failure before it causes downtime. See how you can harness chaos to build resilient systems by requesting a demo of Gremlin.
