Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on Incident Management, On-Call, Incident Response and related technologies.

Alert Fatigue Is Now a Reliability Risk in 2026

Two big reliability surveys landed in 2026, and together they deliver an uncomfortable verdict: alert fatigue has stopped being a morale complaint and turned into a measurable production risk. Engineers are drowning in signals, most of which mean nothing, and the noise is now directly causing outages.

AI-Related Outages Are Reshaping On-Call in 2026

AI-related outages just moved from a fringe worry to a mainline reliability problem, and the on-call rotation is where that shift lands first. A new StackGen analysis of nearly 178,000 public status-page records found that incidents disclosed by AI model and AI application companies now account for more than one in ten reported outages, a sixfold jump from 1.7 percent in 2023 to 10.7 percent so far in 2026.

AI Provider Outages: An On Call Playbook

On the morning of August 5, 2026, a major AI provider went dark for roughly seven and a half hours, and thousands of engineering teams learned in real time what an AI provider outage actually costs them. Anthropic's Claude models returned elevated error rates and failed API requests starting around 3:00 AM Eastern, and applications that quietly route user traffic through a large language model suddenly had no model to route to. Chatbots stopped answering. Summarization pipelines stalled.

SRE Agent Enhancements: Faster Triage, Greater Access Controls, Deeper System Connectivity

This blog post is part of PagerDuty’s ongoing series on how we’re helping customers navigate their journey towards autonomous operations. Read on to learn about how recent SRE Agent Enhancements build towards this vision. During an incident, everything is competing for attention at once. Responders lose time swiveling between tools, insights gathered by AI stay siloed instead of feeding into the next decision, and the pressure to move fast means learnings rarely stick.

Bring Your Backstage Context Into Every PagerDuty Incident

This blog post is part of PagerDuty’s ongoing series on how we’re helping customers navigate their journey towards autonomous operations. Read on to learn about how Custom Field Mapping for PagerDuty’s plugin for both Spotify for Backstage and Spotify Portal for Backstage now generally available builds towards this vision. It’s 2am. A Sev-1 fires, and your on-call responder opens the incident in PagerDuty. What’s waiting for them? A service name, and not much else. No tier.