Incident Management

The latest News and Information on Incident Management, On-Call, Incident Response and related technologies.

What is Mean Time Between Failures - and why does it matter for service availability

Oct 5, 2023 By Amy Brennen In BigPanda

Mean Time Between Failures (MTBF) measures the average duration between repairable failures of a system or product. MTBF helps us anticipate how likely a system, application or service will fail within a specific period or how often a particular type of failure may occur. In short, MTBF is a vital incident metric that indicates product or service availability (i.e. uptime) and reliability.

Read Post

BigPanda

Read more about What is Mean Time Between Failures - and why does it matter for service availability

Enhance Your Customer Service with PagerDuty for ServiceNow CSM

Oct 5, 2023 By Hadijah Creary In PagerDuty

In today’s fast-paced, digital-first landscape, delivering exceptional customer experience is paramount to business success. For customer service teams, that means maintaining service level agreements (SLAs) and ensuring swift responses to customer issues that can make or break your company’s reputation. Fortunately, PagerDuty has improved the way companies handle customer service teams and has built applications into ServiceNow’s CSM platform.

Read Post

PagerDuty

Read more about Enhance Your Customer Service with PagerDuty for ServiceNow CSM

The Rise of Generative AI

Oct 5, 2023 By Blameless In Blameless

Revolutionizing Business: The Rise of Generative AI - Actionable Strategies to Integrate Advanced AI Seamlessly into Your Engineering Operations.

View Video

Blameless

Read more about The Rise of Generative AI

Alerting, Incident Management and the SDLC | Better Incidents Podcast Ep. 8

Oct 5, 2023 By FireHydrant In FireHydrant

In this episode we chat with veteran cloud architect Masaru Hoshi about the challenges of alert fatigue, the importance of effective alerting systems, and fostering ownership in software teams. Masaru shares insights from his 30-year career, emphasizing the need for balance, trust, and collaboration in incident response.

View Video

FireHydrant

Read more about Alerting, Incident Management and the SDLC | Better Incidents Podcast Ep. 8

Global Event Rulesets: Streamlining Alert Routing Across Services

Oct 4, 2023 By Vishal Padghan In Squadcast

In the fast-paced world of organizations handling numerous microservices and projects, tackling the challenges that arise can be a daunting task. As many of our customers come with infrastructures that included a large number of microservices we set out to make it easier for them to streamline alert source management. Enter Global Event Rulesets (GER). This feature is designed to redefine the way you manage alerts.

Read Post

Squadcast

Read more about Global Event Rulesets: Streamlining Alert Routing Across Services

The Link Between Early Detection and Internet Resilience: A Lesson from Salesforce's Outage

Oct 4, 2023 By Madan Gopal N In Catchpoint

Almost every study examining the hourly cost of outages invariably leads to a clear and undeniable conclusion: outages are expensive. According to a 2016 study, the average cost of downtime was estimated at approximately $9,000 per minute. In a more recent study, 61% of respondents stated that outages cost them at least $100,000, with 32% indicating costs of at least $500,000 and 21% reporting expenses of at least $1 million per hour of downtime.

Read Post

Catchpoint

Read more about The Link Between Early Detection and Internet Resilience: A Lesson from Salesforce's Outage

Practicing SDLC the right way #shorts #incidentresponse #sre #softwareengineer

Oct 4, 2023 By FireHydrant In FireHydrant

View Video

FireHydrant

Read more about Practicing SDLC the right way #shorts #incidentresponse #sre #softwareengineer

The problem with noise in Alerting #shorts #incidentresponse #sre #softwareengineer

Oct 4, 2023 By FireHydrant In FireHydrant

View Video

FireHydrant

Read more about The problem with noise in Alerting #shorts #incidentresponse #sre #softwareengineer

Whose fault was it anyway? On blameless post-mortems

Oct 4, 2023 By incident.io In Incident.io

No one wants to be on the receiving end of the blame game—especially in the wake of a major incident. Sure, you know you were the one who made the final change that caused the incident. And hopefully, it was a small one that didn’t cause any SEV-1s. Still, the weight of knowing you caused something bad should be enough, right? Unfortunately, sometimes fingers get pointed, your name gets called, and suddenly, everyone knows that you’re the person who created more work for everyone.

Read Post

Incident.io

Read more about Whose fault was it anyway? On blameless post-mortems

Choosing the Right Metrics for Noiseless K8s Alerting

Oct 4, 2023 By Zenduty In Zenduty

Watch Ankur Rawal and Dheeraj Reddy talk about how to choose the right metrics for noise K8s alerting, with insights and suggestions based on the mistakes made by hundreds of companies while implementing Prometheus Alertmanager in their production systems, and learn how much bad monitoring could be costing you. This talk was delivered at PromCon'2023 in Berlin.

View Video