Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on Incident Management, On-Call, Incident Response and related technologies.

IT problem management VS. IT incident management, and how agentic ITOps improves both

Picture a familiar scene: a critical application goes down during peak business hours, and your on-call engineers scramble to restore service. Two weeks later, the same application fails again, frustrating your teams with the same symptoms, the same scramble, and the same customer frustration. If this pattern feels familiar, your organization may be strong at IT incident management, but underinvested in IT problem management.

Why Regular IT Health Checks Help Prevent Downtime and Improve Business Resilience

Most IT problems do not announce themselves. A backup job quietly fails for three weeks before anyone notices. A firewall rule left open "temporarily" during a project stays open for a year. A former employee's account still has admin rights nobody remembered to remove. None of these cause trouble on the day they happen. They cause trouble later, usually at the worst possible time. This very difference between when these problems start and when they finally become a source of trouble is precisely what a routine IT health check is supposed to bridge.

Your AI agents are lost: give them a graph

The biggest limitation facing enterprise AI agents may not be the model. It may be the context surrounding it. Anthony Alcaraz, Senior AI/ML Portfolio Growth Manager at AWS and co-author of O'Reilly's *Agentic GraphRAG*, joins Humans of Reliability to explain why reliable agents need more than a vector database and a large context window. They need structured knowledge they can navigate, memory they can prune, constraints they can follow, and feedback loops that help them improve.

Don't add a read replica until you've read this

As the size and complexity of their relational database workload grows, every company eventually goes through the process of off-loading work on a read replica. It comes with lots of benefits, but at a cost of increased complexity. This article is about how we dealt with that, a lot of learnings, and some useful techniques. incident.io is an incident management product relied on by thousands of customers to be the thing that supports them through anything from a minor blip to a full outage.

Custom shifts for one-off requirements or complex schedules

While most on-call schedules are built to represent regular rotations, often on a weekly basis, not all of your on-call needs require the same coverage every week. We’ve added Custom Shifts to the Shift-Based Schedules for maximum flexibility. Custom shifts are a feature of our new Shift-Based Schedules. With Custom Shifts, your team can cover ad hoc needs for special events, major deploys, Failure Fridays, gamedays, or whatever comes up that needs some extra coverage.

From AIOps to agentic ITOps: Why AI for IT operations has entered a new era

Enterprise IT has reached an inflection point. Your teams are responsible for hybrid cloud infrastructure, microservices, third-party dependencies, and shipping AI-generated code at unprecedented velocity. IT environments are becoming more complex faster than traditional tools and processes can keep pace. Alert volumes keep climbing. Institutional knowledge keeps walking out the door. And the pressure to do more with flat or shrinking budgets isn’t letting up.

H1 2026 Cloud and SaaS Reliability Report

The first half of 2026 reinforced a key idea about Cloud and SaaS reliability - dependency risk. IncidentHub tracked 30,246 outages across 1,082 providers between January and June 2026. May was the busiest month, with 6,070 incidents. Cloud providers led in the total number of outages (4,723), followed closely by developer tools (4,589).

The July 2026 AWS CloudFront Outage: VPC Origins, Cascade Impact, and What Broke

On July 16, 2026, AWS experienced a disruption in its CloudFront service, which affected a large number of websites and applications. The outage was caused by a configuration loading failure in CloudFront's VPC Origins feature. This was AWS's most widely-felt outage after last year's outage on October 20th, which caused widespread damage.

Trust, Resilience & AI: A Customer Panel with TD Bank & New York Life

What does it really take to be "the calm in the storm" during a major incident? In this candid panel from PagerDuty on Tour, Chris Conklin (Technology Executive AIOPs, TD Bank) and Sam Brinley (CVP Enterprise Cloud Solution Architect & Engineer at New York Life) sit down with PagerDuty to talk through two decades of evolution in IT operations – from the "Wild West" of early network management to today's push into AI and agentic operations.