Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on Incident Management, On-Call, Incident Response and related technologies.

Managing Dedicated Lines in OnPage | How to Set up Dedicated Lines

Learn how to set up and manage Dedicated Lines in OnPage. In this step-by-step tutorial, we’ll walk through how to configure your Dedicated Lines to control how incoming calls are received, processed, and routed. You’ll learn how to:✓ Configure basic line details✓ Customize call processing behavior✓ Set up interactive routing menus✓ Add caller instructions and prompts✓ Review and save your Dedicated Line settings.

Identity Provider Outage: The On-Call Blast Radius

An identity provider outage is the one failure mode where your dashboards stay green and your entire company stops working anyway. On August 31, 2026, Microsoft acknowledged a widespread Exchange Online incident at 5:30 PM UTC, tracked as EX1464935, and described it in the admin center as "a common failure pattern across affected Exchange Online requests that is associated with authentication and protocol connectivity." Tens of thousands of users were affected according to Downdetector.

Incident Management: 20 Years of Change

Incident management fundamentals still apply, even as hybrid cloud and Kubernetes reshape incidents. Garrett Douglas joins LogicMonitor's Coffee and Context on why one failure floods ITOps with alerts and fuels alert fatigue. The bottleneck is incident context: knowing which alerts matter for incident response. Forrester's report Incident Management Has Outgrown Its Playbook frames context as the new competitive advantage.

OnPage Enterprise (Dispatcher) Console Dashboard Overview

Welcome to the OnPage Dispatcher Console, your central hub for messaging activity and performance insights. In this video, we’ll walk through the main dashboard and show how it gives dispatchers a consolidated view of messaging activity, performance metrics, on-call status, and upcoming schedules, all in one place. See how the OnPage Dispatcher Dashboard helps teams stay informed and maintain visibility into critical communications.

Your next internal developer platform is a library of agent skills

What happens when AI agents become direct users of your infrastructure? Michael Kutsch, Staff SRE and Team Lead for Cloud Foundations at PostHog, argues that the next internal developer platform may be a library of agent skills. Instead of forcing every task through a portal, his team is giving agents structured context, reusable workflows, and deterministic scripts they can call when reliability and governance matter.

Run an Incident Response Game Day for Your On-Call Team

An incident response game day is the cheapest reliability investment most engineering teams still refuse to make. The 2026 Catchpoint SRE Report, based on 418 responses from reliability professionals worldwide, names it as one of five defining trends: resilience has to be practiced. Teams that deliberately test failure report more confidence and better preparedness. And yet the same report finds production chaos engineering is still far from standard practice.

Building an Operational Runbook for Exporting and Managing WhatsApp Business Chats

WhatsApp is widely used for customer support, supplier coordination and project communication. However, important information can become difficult to manage when it remains inside individual conversations or employee accounts. An operational runbook creates a consistent process for exporting relevant chats, reviewing the files and storing them securely. It also reduces the risk of losing business context during employee offboarding, account changes or project handovers.

Shift Scheduling: 10 Signs You've Outgrown Spreadsheets (And What to Look for Next)

It’s Friday afternoon. Two employees have requested time off. Someone calls in sick. Another wants to swap shifts. Then you realize the only certified technician scheduled for the night shift is also marked as being on vacation. What looked like a perfectly organized spreadsheet this morning can quickly turn into a puzzle. As organizations grow, scheduling gets more complex.

Extending autonomous L1 ops with new suppression and runbook capabilities

Earlier this year, I had the chance to meet with one of our airline customers. During the meeting, we discussed how to use agentic technology to automate L1 workflows. As one of the largest global airlines, they have many applications and service teams focused on flight-critical, tier 1 environments. Any downtime can cause costly delays and unhappy customers.