Chaos Engineering

How to make your services resilient to slow dependencies

Apr 24, 2024 By Andre Newman In Gremlin

When discussing reliability, we tend to focus on the things that we have control over: applications, virtual machine instances, deployment patterns, etc. But this ignores a significant and ever-growing part of nearly all modern software: dependencies. Dependencies are services that provide extra functionality for other services and applications. For instance, many websites depend on databases, caches, payment processors, and similar services in order to function.

Read Post

Gremlin

Read more about How to make your services resilient to slow dependencies

Hitting reliability goals in the face of layoffs

Apr 23, 2024 By Jeff Nickoloff In Gremlin

It’s never easy when layoffs hit your organization. In addition to the personal impact of losing friends and coworkers from your team, those who remain are left trying to achieve the same business goals with less people and resources. Unfortunately, layoffs and restructuring have become a common part of business. But you’re not alone. Your partners (including Gremlin) are here to help you navigate your new reality.

Read Post

Gremlin

Read more about Hitting reliability goals in the face of layoffs

How to ensure your Kubernetes Pods and containers can restart automatically

Apr 16, 2024 By Andre Newman In Gremlin

As complex as Kubernetes is, much of it can be distilled to one simple question: how do we keep containers available for as long as possible? All of the various utilities, features, platform integrations, and observability tools surrounding Kubernetes tend to serve this one goal. Unfortunately, this also means there’s a lot of complexity and confusion surrounding this topic. After all, most people would agree that availability is important, but how exactly do you go about achieving it?

Read Post

Gremlin

Read more about How to ensure your Kubernetes Pods and containers can restart automatically

How to ensure your Kubernetes cluster can tolerate lost nodes

Apr 12, 2024 By Andre Newman In Gremlin

Redundancy is a core strength of Kubernetes. Whenever a component fails, such as a Pod or deployment, Kubernetes can usually automatically detect and replace it without any human intervention. This saves DevOps teams a ton of time and lets them focus on developing and deploying applications, rather than managing infrastructure.

Read Post

Gremlin

Read more about How to ensure your Kubernetes cluster can tolerate lost nodes

How to test your systems for scalability and redundancy with Fault Injection

Apr 11, 2024 By Gremlin In Gremlin

Part of the Gremlin Office Hours series: A monthly deep dive with Gremlin experts. Do you know if your services can tolerate losing a node? What about an entire availability zone? Or a region?‍ Large-scale outages aren’t unheard of. When you’re running critical services, it’s vital that those services can keep running even if an AZ or region fails. In addition to failing over, these services also need to scale quickly so traffic shifts don’t overwhelm your systems. How do you prove that a service is both scalable and redundant? The answer is with Fault Injection.

View Video

Gremlin

Read more about How to test your systems for scalability and redundancy with Fault Injection

Top 7 Kubernetes Chaos Engineering Tools

Apr 10, 2024 By Daniel Olaogun In Speedscale

Enhance kubernetes resilience with AWS Fault Injection Simulator, LitmusChaos, Gremlin, Chaos Monkey, ChaosBlade, Azure Chaos Studio, and Speedscale.

Read Post

Speedscale

Read more about Top 7 Kubernetes Chaos Engineering Tools

How to standardize resiliency on Kubernetes

Apr 10, 2024 By Gavin Cahill In Gremlin

There’s more pressure than ever to deliver high-availability Kubernetes systems, but there’s a combination of organizational and technological hurdles that make this ‌easier said than done. Technologically, Kubernetes is complex and ephemeral, with deployments that span infrastructure, cluster, node, and pod layers. And like with any complex and ephemeral system, the large amount of constantly-changing parts opens the possibility for sudden, unexpected failures.

Read Post

Gremlin

Read more about How to standardize resiliency on Kubernetes

Where to automate resilience testing in your SDLC

Apr 9, 2024 By Ryan Detwiller In Gremlin

When organizations begin to deploy resilience testing or Chaos Engineering, there’s a natural question: can we integrate this with our CI/CD pipeline or release automation tools? After all, you’re likely running unit, performance, and integration tests already—is resiliency different? The short answer is yes—to both. Integration is possible, but resiliency is different, so automation is a nuanced conversation.

Read Post

Gremlin

Read more about Where to automate resilience testing in your SDLC

Resiliency is different on AWS: Here's how to manage it

Apr 2, 2024 By Andre Newman In Gremlin

There’s a common misconception about running workloads in the cloud: the cloud provider is responsible for reliability. After all, they’re hosting the infrastructure, services, and APIs. That leaves little else for their customers to manage, other than the workloads themselves…right?

Read Post

Gremlin

Read more about Resiliency is different on AWS: Here's how to manage it

Fault Injection in your release automation

Mar 18, 2024 By Sam Rossoff In Gremlin

One of the real successes of the Agile Software development movement has been the push to have regular, frequent deployments. This has manifested as build and deployment automation and the general adoption of CI/CD. As engineers automate more processes of their software release lifecycle, an important question is how to automate Quality Assurance, which includes resilience testing and, more specifically, Fault Injection.

Read Post