Chaos Engineering

Chaos Engineering: The Path to Reliability - Kolton Andrus

Oct 15, 2020 By Gremlin In Gremlin

We’re all here for the same purpose: to ensure the systems we build operate reliably. This is a difficult task, one that must balance people, process and technology during difficult conditions. We operate with incomplete information, assessing risks and dealing with emerging issues. We’ve found Chaos Engineering to be a valuable tool in addressing these concerns. Learn from real world examples what works, what doesn’t, and what the future holds.

View Video

Gremlin

Read more about Chaos Engineering: The Path to Reliability - Kolton Andrus

Identifying Hidden Dependencies - Liz Fong Jones

Oct 15, 2020 By Gremlin In Gremlin

You don't need to write automation or deploy on Kubernetes to gain benefits from resilience engineering! Learn how Honeycomb improved the reliability of our Zookeeper, Kafka, and stateful storage systems through terminating nodes on purpose. We'll discuss the initial manual experiments we ran, the bugs in our automatic replacement tools we uncovered, and what steps we needed to progress towards continuously running the experiments. Today, no node at Honeycomb lives longer than 12 months, and we automatically recycle nodes every week.

View Video

Gremlin

Read more about Identifying Hidden Dependencies - Liz Fong Jones

Lessons from Incident Management and Postmortems at Atlassian - Jim Severino

Oct 15, 2020 By Gremlin In Gremlin

How do you run incidents and postmortems at a company with thousands of engineers spread across the globe? Jim Severino shares what worked (and didn't worked) for Atlassian.

View Video

Gremlin

Read more about Lessons from Incident Management and Postmortems at Atlassian - Jim Severino

Looking back on Chaos Conf 2020

Oct 15, 2020 By Andre Newman In Gremlin

It’s already been a week since we closed our third annual Chaos Conf! While we were forced to take the conference online, this meant that more of you could join us. Over 3,500 people signed up to help make this the world’s largest Chaos Engineering conference. That’s 5x more than 2019, and nearly 10x more than 2018! This is a testament to the growth of Chaos Engineering as a practice across many different industries and around the world.

Read Post

Gremlin

Read more about Looking back on Chaos Conf 2020

Incident Ready: How to Chaos Engineer Your Incident Response Process | FireHydrant

Oct 15, 2020 By FireHydrant In FireHydrant

We’re pretty sure using a real incident to test a new response process is not the best idea. So, how do you test your process ahead of time? In this video, FireHydrant CEO, Robert Ross, will share how FireHydrant customers leverage best practices to break, mitigate, resolve, and fireproof incident processes. We’ll show you how to use chaos engineering philosophies to stress test 3 critical parts of a great process.

View Video

FireHydrant

Read more about Incident Ready: How to Chaos Engineer Your Incident Response Process | FireHydrant

Lead Times and Psychological Safety within the Five Ideals - Gene Kim

Oct 15, 2020 By Gremlin In Gremlin

The biggest challenges engineering organizations face are not technical. They’re fundamental problems with how we think and go about doing work, and the environments that we work in. In this talk, Gene Kim will share the Five Ideals and how they relate to Chaos Engineering. He’ll also show how the Five Ideals help build stronger, better performing, and ultimately more reliable companies.

View Video

Gremlin

Read more about Lead Times and Psychological Safety within the Five Ideals - Gene Kim

Chaos Engineering Processes

Oct 13, 2020 By FireHydrant In FireHydrant

You can use chaos engineering to test processes as much as you can test how systems fail.

View Video

FireHydrant

Read more about Chaos Engineering Processes

Is your microservice a distributed monolith?

Sep 30, 2020 By Andre Newman In Gremlin

Your team has decided to migrate your monolithic application to a microservices architecture. You’ve modularized your business logic, containerized your codebase, allowed your developers to do polyglot programming, replaced function calls with API calls, built a Kubernetes environment, and fine-tuned your deployment strategy. But soon after hitting deploy, you start noticing problems.

Read Post

Gremlin

Read more about Is your microservice a distributed monolith?

Technology Business Management and Chaos Engineering

Sep 18, 2020 By Matthew Helmke In Gremlin

Get started with Gremlin’s Chaos Engineering tools to safely, securely, and simply inject failure into your systems to find weaknesses before they cause customer-facing issues. Technology Business Management (TBM) is a decision-making tool that helps organizations maximize the business value of information technology (IT) spending by adjusting management practices. With TBM, IT is transformed to run like a business instead of merely a cost center.

Read Post

Gremlin

Read more about Technology Business Management and Chaos Engineering

Understanding your application's critical path

Sep 14, 2020 By Andre Newman In Gremlin

Don’t wait for an incident to focus on reliability. Learn concrete steps for preventing incidents in the first place in our two-part series, Planning and Architecting for Reliability. It’s 3 a.m. You’re lying comfortably in bed when suddenly your phone starts screeching. It’s an automated high-severity alert telling you that your company’s web application is down. Exhausted, you open the website on your phone and do some basic tests.

Read Post