Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on Incident Management, On-Call, Incident Response and related technologies.

Introducing Spike's new look: designed to scale.

Today, we are introducing Spike’s new logo, a new website, and honestly, a new identity from the ground up. This is very exciting day for all of us at Spike. More companies are being built today than at any other point in history. Small teams are making big products. And no matter the size of the team, every single one of them needs reliability. Reliability should not be a second-class citizen for any company, no matter where they are in their journey.

Incident Chat and Virtual War Rooms: How to Improve Incident Response

How do you keep incident communication organized during a critical incident? Effective incident response requires more than getting an alert to the right person. Once responders are engaged, they need a shared place to exchange information, coordinate actions, and track decisions. A dedicated incident chat – or virtual incident war room – keeps that collaboration tied directly to the incident instead of scattering it across email, Microsoft Teams, Slack, text messages, and phone calls.

Assign Bugs and Tickets to the Current On-Call, Automatically

There is a particular kind of ticket that costs more than it should. It is filed correctly, it has a good description, it is in the right project, and it has nobody's name on it. It sits in the queue for two days because everyone who looks at the board assumes someone else has it. Then a customer follows up, someone notices, and the fix takes twenty minutes. The gap was never engineering time. It was ownership.

PD Automation Runner: Automation That Finally Reaches Your On-Prem Stack

This blog post is part of PagerDuty’s ongoing series on how we’re helping customers navigate their journey towards autonomous operations. Read on to learn about how PD Automation Runner, now in Early Access, builds towards this vision. Your on-call engineer gets paged for a critical alert on a self-hosted Kubernetes cluster. If this were a cloud-hosted service, Workflow Actions and the SRE Agent would already be working the problem: pulling logs, checking health, surfacing a fix.

See It, Approve It, Revoke It: Scoped OAuth for Public Apps

This blog post is part of PagerDuty’s ongoing series on how we’re helping customers navigate their journey towards autonomous operations. Read on to learn about how Scoped OAuth for Public Apps, now in Early Access, builds towards this vision. Your security team asks a simple question during a routine review: which third-party apps can reach our PagerDuty data right now, and what exactly can they do with it?