Test Bare Metal Solution reliability with Gremlin. Simulate server failures and resource pressure to validate dedicated hardware failover.
Google Cloud Bare Metal Solution exists for workloads that can't run virtualized, which in practice means licensing-constrained databases and legacy enterprise systems. Those workloads are usually critical, and the tier they run on has none of the conveniences your team has grown used to.
There's no live migration, no instant reprovisioning, and no autohealing. A hardware problem means real downtime measured against hardware replacement timelines. The applications running here often predate the idea of designing for failure, and whatever redundancy exists was configured by people who have since left.
Install the Gremlin Agent on your bare metal servers and you can measure recovery before you have to perform it. Shutdown experiments establish how long failover to redundant hardware genuinely takes. CPU and memory experiments show whether these systems degrade gracefully under pressure or fall over, which matters more here than anywhere else in your estate.
When the answer to "how long until we're back" might be hours rather than seconds, knowing the number changes what you promise the business.
Host redundancy - Linux
Test resilience to host failures by shutting down a randomly selected Linux host. Verify that your platform automatically restarts or replaces it.
CPU scalability - Linux
Test that your Linux-hosted service scales as expected when CPU capacity is limited. Gremlin consumes CPU in three stages—50%, 75%, and 90%—to validate scaling thresholds.
Scalability: Memory
Verify that your Linux-hosted service scales as expected when memory is limited. Gremlin increases memory utilization in three stages—50%, 75%, and 90%—to validate memory management.
* Gremlin is designed to work on any cloud platform that provides Linux or Windows hosts. We haven't individually tested every service we cover, and not all are officially supported. Check our compatibility documentation for tested operating systems and known caveats, or get in touch if you don't see yours.
Intelligent Health Checks - GCP
Automatically monitor your Google Cloud services during testing in one-click with Intelligent Health Checks.
Avoid downtime. Use Gremlin to turn failure into resilience.
Gremlin empowers you to proactively root out failure before it causes downtime. See how you can harness chaos to build resilient systems by requesting a demo of Gremlin.
