Operations | Monitoring | ITSM | DevOps | Cloud

Assign Bugs and Tickets to the Current On-Call, Automatically

There is a particular kind of ticket that costs more than it should. It is filed correctly, it has a good description, it is in the right project, and it has nobody's name on it. It sits in the queue for two days because everyone who looks at the board assumes someone else has it. Then a customer follows up, someone notices, and the fix takes twenty minutes. The gap was never engineering time. It was ownership.

Identity Provider Outage: The On-Call Blast Radius

An identity provider outage is the one failure mode where your dashboards stay green and your entire company stops working anyway. On August 31, 2026, Microsoft acknowledged a widespread Exchange Online incident at 5:30 PM UTC, tracked as EX1464935, and described it in the admin center as "a common failure pattern across affected Exchange Online requests that is associated with authentication and protocol connectivity." Tens of thousands of users were affected according to Downdetector.

Run an Incident Response Game Day for Your On-Call Team

An incident response game day is the cheapest reliability investment most engineering teams still refuse to make. The 2026 Catchpoint SRE Report, based on 418 responses from reliability professionals worldwide, names it as one of five defining trends: resilience has to be practiced. Teams that deliberately test failure report more confidence and better preparedness. And yet the same report finds production chaos engineering is still far from standard practice.

Cloud Outage Resilience: On-Call Lessons for 2026

Cloud outage resilience has quietly become the most important reliability topic of the year. Analysts now treat large scale cloud downtime as a matter of when, not if. Forrester has predicted at least two major multi day hyperscaler outages in 2026, and the reasoning is hard to argue with. AWS, Azure, and Google Cloud together account for well over half of enterprise cloud spending, so when any one of them stumbles, a huge slice of the digital economy stumbles with it.

Incident Communication Lessons From Spotify Outages

Good incident communication is the difference between an outage your users forgive and an outage that quietly pushes them toward a competitor. That lesson landed hard in late July 2026, when Gergely Orosz of The Pragmatic Engineer publicly walked away from publishing video podcasts on Spotify after a run of reliability failures. The bug that broke publishing was almost beside the point.

Alert Fatigue Is Now a Reliability Risk in 2026

Two big reliability surveys landed in 2026, and together they deliver an uncomfortable verdict: alert fatigue has stopped being a morale complaint and turned into a measurable production risk. Engineers are drowning in signals, most of which mean nothing, and the noise is now directly causing outages.

AI-Related Outages Are Reshaping On-Call in 2026

AI-related outages just moved from a fringe worry to a mainline reliability problem, and the on-call rotation is where that shift lands first. A new StackGen analysis of nearly 178,000 public status-page records found that incidents disclosed by AI model and AI application companies now account for more than one in ten reported outages, a sixfold jump from 1.7 percent in 2023 to 10.7 percent so far in 2026.

AI Provider Outages: An On Call Playbook

On the morning of August 5, 2026, a major AI provider went dark for roughly seven and a half hours, and thousands of engineering teams learned in real time what an AI provider outage actually costs them. Anthropic's Claude models returned elevated error rates and failed API requests starting around 3:00 AM Eastern, and applications that quietly route user traffic through a large language model suddenly had no model to route to. Chatbots stopped answering. Summarization pipelines stalled.

How to Fix On-Call Burnout Before It Breaks Your Team

On-call burnout is no longer a fringe complaint. It is one of the loudest signals in the 2026 reliability data. A wave of fresh industry research this year points to the same uncomfortable conclusion: the people who keep systems running are running on empty. In the DuploCloud 2026 AI and DevOps Report, 47 percent of engineers said DevOps overload contributes to burnout, with on-call rotations and repetitive maintenance singled out as primary culprits.