Operations | Monitoring | ITSM | DevOps | Cloud

Creating Message Templates in OnPage

Learn how to create and configure a message template in OnPage to make critical communication faster, more consistent, and easier for your team. OnPage message templates are predefined message formats that help teams standardize frequently used communications without having to compose every alert or message from scratch, and can be pulled up on both web message dispatcher and OnPage's phone app. In this step-by-step tutorial, we show you how to navigate to Settings then Template, create a new message template, and configure the information users will see when that template is selected.

Suppressing OnPage Notifications During Maintenance

Learn how to suppress OnPage notifications during a scheduled maintenance window so your on-call team isn’t unnecessarily alerted while planned work is underway. OnPage’s Suppress Notifications option temporarily prevents notifications from being sent during a defined period. This is especially useful during scheduled maintenance, planned downtime, testing, or other known activities that could otherwise generate unnecessary alert noise.

How Facilities Management Supports Business Continuity and IT Resilience

When organizations build business continuity plans, they usually focus on software backups and cyber threats, while ignoring the physical building itself. But a single power fluctuation, HVAC failure, or roof leak can shut down critical IT assets just as fast as a digital attack. Facilities management bridges the gap between physical infrastructure and digital uptime. This guide breaks down how proactive building systems protect your technology, minimize expensive operational disruptions, and turn theoretical recovery plans into daily operational defense.

How to Build a Self-Improving Operations System in 5 Steps

With AI agents and AI-generated code becoming the norm in modern enterprise software, backend systems are evolving faster than ever. And it’s leaving most operations teams with an impossible choice: burn out senior talent on repetitive firefighting, or hand production over to untrained AI agents. With disruptions costing enterprises an average of $300,000 per hour, manual firefighting isn’t an option.

Sending HighPriority Messages from OnPage's Dispatcher

See how OnPage Dispatcher helps teams centralize and streamline time-sensitive communication. In this demo, a request is created in Dispatcher, routed to the appropriate on-call specialist, and tracked through acknowledgment and response. We also show how the communication is bi-directional, and responses can be seen within the dispatcher, with a complete audit trail.

Proofpoint outage on August 14, 2026: DNS failure disrupts email worldwide

A DNS failure at Proofpoint broke email delivery for organizations around the world on August 14, 2026. Records for pphosted.com stopped resolving, so inbound and outbound mail routed through Proofpoint bounced or stalled for nearly four hours. StatusGator flagged the incident with an Early Warning Signal at 12:48 UTC, about an hour before Proofpoint acknowledged it publicly on its status page at 13:50 UTC. Here is what happened, who it hit, and how some teams kept mail moving.

The August 13, 2026 Namecheap Outage

Namecheap took more than 5,000 servers offline on August 13, 2026 after cooling systems failed at RadiusDC's Phoenix datacenter, and brought services back in stages over roughly 28 and a half hours. The shutdown was deliberate, intended to protect hardware from overheating. It reached most of the product line - hosting, EasyWP, Private Email, DNS management, URL redirect management and the support helpdesk - while DNS zone resolution was unaffected.

Creating an AfterHours OnCall Schedule

Learn how to create and configure an **on-call schedule in OnPage** with this step-by-step how-to guide. In this video, we walk through how to select a group, create a new schedule, assign on-call team members, set escalation priority, and configure coverage for specific days and times. Using OnPage’s on-call scheduling capabilities, teams can ensure the right responders are automatically available for critical alerts, incidents, messages or calls during designated coverage periods, including after-hours, weekends, and other shifts.

How to Create and Import Contacts in the NEW OnPage Web Console

A step-by-step guide for OnPage’s new web management console, including how to create a single contact and how to create multiple contacts at once by importing an Excel spreadsheet. Feel free to comment below with any questions! Whether you’re in IT or healthcare, OnPage helps teams manage critical alerting and communication to ensure urgent messages reach the right people at the right time. If you’re not yet using OnPage or want to see how it works, request a demo or speak with a member of our team to learn more.

Creating Escalation and Regular Groups in OnPage

Learn how to create and configure an Escalation Group in OnPage with this step-by-step how-to guide. This video walks through how to create an escalation group, a regular group, configure escalation intervals and factors, enable Round Robin, set failover OPIDs, and add a Fail Report email address. With escalation groups, OnPage can route critical alerts to team members in a predefined order and automatically move to the next responder when needed—helping ensure time-sensitive notifications don’t go unanswered.

Introducing the next generation of the BigPanda AI Incident Assistant

Effective incident response depends on having all of the context surrounding what’s happening. You have to understand your systems, services, architecture, and teams deeply enough to correctly interpret whatever alert just fired. Too often, that context doesn’t arrive packaged neatly in one place. Gathering and interpreting context correctly under time pressure is one of the most difficult parts of the job.

How to Investigate a Production Incident Using an AI Agent (AppSignal MCP)

An incident has hit your product. I've been there: you're context-switching between hosting, CI/CD, codebase, AppSignal for monitoring, and whatever else your product depends on to minimize downtime and potential losses. You're trying to piece everything together, but it takes a lot of time, and that's something you don't have. AI agents connected to your tooling and your monitoring data via MCP free up that time for you.

Reduce duplicate alert noise with Alert Deduplication

A single incident can generate the same alert several times in quick succession. These duplicate alerts create unnecessary noise and alert fatigue, and make it harder for on-call teams to focus on the issue that needs their attention. OnPage Alert Deduplication reduces repeat notifications while preserving visibility into every incoming alert.

What data sources does agentic ITOps use

Agentic IT operations have arrived. It’s no longer a question of if enterprise IT departments will adopt agentic ITOps, but how quickly. The question we hear most often at BigPanda isn’t “what are agentic ITOps,” it’s “what data do we actually need to get started?” That’s the right question to ask. Agentic AI is only as good as the data and context that feeds it. Real-time observability and telemetry data from machines. Structured ITSM and workflow records.

AI Is Outpacing Code Review. Here's How to Catch Up (Without Slowing Down)

In a 2025 analysis spanning over 100 large language models, Veracode found that nearly half (45%) of AI-generated code causes known security issues and vulnerabilities. Novel risks are being introduced into your operations systems faster than humans can manage or review. At the same time, studies suggest that human review isn’t all that effective, especially beyond 400 lines of code. But AI-generated code isn’t inherently bad. It just doesn’t always work across your whole system.

Cloud Incident Management: Process, Tools, and Practices

How do you resolve an outage your organization has no authority to fix? A managed database drops into read-only mode and stops accepting writes. There's no host to reach, no configuration file to edit, and no restart command available to your engineers. Cloud incident management begins at that boundary, where the response depends on a support channel and a provider status page. Plenty of what you already know still applies here.

How to build a resilient incident management workflow using ilert

Your payment API suddenly returns 503 errors. Within seconds, your infrastructure monitors, application checks, and dependency monitors begin generating their own alerts. And while the dashboards keep flashing, the clock is still running. Your customers are waiting, internal teams are asking for updates, and engineers are trying to separate the real problem from the noise before the situation gets worse.

5 Ways IT Leaders Are Using AI to Improve Operations in 2026

As the world is racing to plug AI into nearly every part of business, especially software engineering, the stakes to maintain operational integrity have never been higher. AI-generated code and AI-agents ship faster than human SREs can prepare for, which can create costly issues down the line: incidents get harder to predict and more expensive to recover from.

We turned off Pub/Sub and nobody noticed

Like many modern software stacks, the incident.io platform is predominantly event-driven. For example, whenever you send us an alert, post a message to our agent on Slack, or update an entry in your Catalog - these are all events that then get enqueued on a message topic, meaning any of our downstream components that are interested in that event can subscribe and react asynchronously, such as sending a push notification or posting a reply to you in Slack.

Cloud Outage Resilience: On-Call Lessons for 2026

Cloud outage resilience has quietly become the most important reliability topic of the year. Analysts now treat large scale cloud downtime as a matter of when, not if. Forrester has predicted at least two major multi day hyperscaler outages in 2026, and the reasoning is hard to argue with. AWS, Azure, and Google Cloud together account for well over half of enterprise cloud spending, so when any one of them stumbles, a huge slice of the digital economy stumbles with it.

Stop Chasing Field Technicians - Track Work Progress with One-Tap Status Updates

Field service managers need visibility into more than just whether a technician has arrived. They need to know when work begins, when it’s completed, and when the technician is heading to the next job. Without a simple, consistent way to communicate these milestones, dispatchers and supervisors are left making phone calls and sending text messages just to find out what’s happening. Most of the time, nothing is wrong.

Incident Communication Lessons From Spotify Outages

Good incident communication is the difference between an outage your users forgive and an outage that quietly pushes them toward a competitor. That lesson landed hard in late July 2026, when Gergely Orosz of The Pragmatic Engineer publicly walked away from publishing video podcasts on Spotify after a run of reliability failures. The bug that broke publishing was almost beside the point.

From Incident Data to Operational Knowledge: A Safer Role for Generative AI in IT Ops

IT operations teams produce an enormous amount of information. Alerts, logs, incident messages, deployment records, support tickets, runbooks and post-incident reviews all contain operational knowledge. The problem is that much of this knowledge remains fragmented and difficult to reuse. Generative artificial intelligence can help organise and transform this information, but its safest role is not unrestricted control over production infrastructure. Its strongest initial use cases involve reading, summarising, classifying and drafting information for an engineer to review.

Alert Fatigue Is Now a Reliability Risk in 2026

Two big reliability surveys landed in 2026, and together they deliver an uncomfortable verdict: alert fatigue has stopped being a morale complaint and turned into a measurable production risk. Engineers are drowning in signals, most of which mean nothing, and the noise is now directly causing outages.

The August 6, 2026 GitHub Actions Outage: Queued Jobs, Throttled Webhooks, Impact Lasting 10 Hours

On August 6, 2026, GitHub opened an incident for degraded Actions performance at 15:22 UTC. Within about twenty minutes, Actions availability was listed as degraded, workflow runs were failing to start or failing partway through, and the Actions REST API was returning errors. Pages was pulled into the same incident shortly afterwards. The status page marked Actions and Pages as mitigated at 00:05 UTC on August 7, and closed the incident at 02:04 UTC.

AI-Related Outages Are Reshaping On-Call in 2026

AI-related outages just moved from a fringe worry to a mainline reliability problem, and the on-call rotation is where that shift lands first. A new StackGen analysis of nearly 178,000 public status-page records found that incidents disclosed by AI model and AI application companies now account for more than one in ten reported outages, a sixfold jump from 1.7 percent in 2023 to 10.7 percent so far in 2026.

AI Provider Outages: An On Call Playbook

On the morning of August 5, 2026, a major AI provider went dark for roughly seven and a half hours, and thousands of engineering teams learned in real time what an AI provider outage actually costs them. Anthropic's Claude models returned elevated error rates and failed API requests starting around 3:00 AM Eastern, and applications that quietly route user traffic through a large language model suddenly had no model to route to. Chatbots stopped answering. Summarization pipelines stalled.

How to Fix On-Call Burnout Before It Breaks Your Team

On-call burnout is no longer a fringe complaint. It is one of the loudest signals in the 2026 reliability data. A wave of fresh industry research this year points to the same uncomfortable conclusion: the people who keep systems running are running on empty. In the DuploCloud 2026 AI and DevOps Report, 47 percent of engineers said DevOps overload contributes to burnout, with on-call rotations and repetitive maintenance singled out as primary culprits.

Get the Context Your Alerts Are Missing with Event Enrichment

This blog post is part of PagerDuty’s ongoing series on how we’re helping customers navigate their journey towards autonomous operations. Read on to learn about how Event Enrichment builds towards this vision. Every on-call engineer knows the drill. An alert fires. It tells you something is wrong, but not what it means. Is this asset in maintenance? Which team owns it? Is it customer-facing?

SRE Agent Enhancements: Faster Triage, Greater Access Controls, Deeper System Connectivity

This blog post is part of PagerDuty’s ongoing series on how we’re helping customers navigate their journey towards autonomous operations. Read on to learn about how recent SRE Agent Enhancements build towards this vision. During an incident, everything is competing for attention at once. Responders lose time swiveling between tools, insights gathered by AI stay siloed instead of feeding into the next decision, and the pressure to move fast means learnings rarely stick.

Bring Your Backstage Context Into Every PagerDuty Incident

This blog post is part of PagerDuty’s ongoing series on how we’re helping customers navigate their journey towards autonomous operations. Read on to learn about how Custom Field Mapping for PagerDuty’s plugin for both Spotify for Backstage and Spotify Portal for Backstage now generally available builds towards this vision. It’s 2am. A Sev-1 fires, and your on-call responder opens the incident in PagerDuty. What’s waiting for them? A service name, and not much else. No tier.

On-Call in 2026: Preparing for Cascading Failures

The biggest outages of 2026 are not being caused by a single server dying or one bad deploy. They are being caused by cascading failures, where healthy systems interact in ways nobody planned for and take each other down. That shift changes what good on-call looks like. If your incident response still assumes that "something broke" and one team owns the fix, you are going to be slow exactly when speed matters most.

Why Config Changes Cause Most Cloud Outages in 2026

If you have watched the incident channels light up over the past few weeks, you already sense the theme of 2026: cloud outages are no longer rare, dramatic once a year events. They are a steady drumbeat, and most of them trace back to the same root cause. Not a data center fire, not a rogue backhoe severing a fiber line, but a routine configuration change that went out, behaved differently than expected, and cascaded.

On-Call Incident Response When Outages Are the New Normal

If your engineering team feels like the outage alerts have gotten louder in 2026, the data agrees with you. Strong on-call incident response has quietly become the difference between a five minute blip and a headline. In the week of July 20 to 26, 2026, ThousandEyes tracked 610 global network outage events, up 4 percent from the 587 the week before, with United States outages rising 10 percent to 457 (Network World).

Institutional knowledge doesn't scale: Building an agentic data analyst

We’ve previously written about how deeply embedded data is in people’s day-to-day work at incident.io, and I’d have it no other way — demand for data is undoubtedly a good thing. What risks breaking at scale, however, is everything downstream of that demand: data-team capacity gets stretched thin, dashboard sprawl outpaces anyone's ability to maintain it, and stakeholders can't reach an answer without going through the data team.

Third-Party Outages: On-Call Lessons From Q2 2026

Cloudflare just published its Q2 2026 Internet Disruption Summary, and the biggest lesson for on-call engineers is uncomfortable: most of the outages that ruin your week are not caused by your own code. Third party outages, upstream provider failures, cable cuts, DNS misconfigurations, and government shutdowns all produce the exact same symptom your users care about, which is that your product stops working.

Cloud Outage Response: Lessons From July 2026

Cloud outage response got a brutal stress test in July 2026. In the span of nine days, three separate cloud infrastructure failures took large chunks of the internet offline: AWS CloudFront on July 16, Microsoft Azure West US on July 23, and AWS us-west-2 on July 24. None of them were caused by a dramatic data center fire or a nation state attack. They were routing faults, configuration translation bugs, and a piece of networking hardware on the path between a region and a metro area.

Incident Response Lessons From a 3 GW Grid Drop

When a transmission line faulted in Ashburn, Virginia on July 22, 2026, more than 3 GW of data center load vanished from the PJM grid in seconds. That is roughly three percent of total grid demand at the moment it happened, and the grid took about ten minutes to stabilize instead of the milliseconds a routine disturbance normally requires. For anyone who owns a pager, this is more than an energy story.

Incident Response When the Outage Isn't Yours

Most of the outages that will page your team this quarter did not start in your code. They started in the physical world: a storm, a severed fiber cable, a data center losing power, or a government flipping a national switch. That is the uncomfortable takeaway from Cloudflare's Q2 2026 Internet Disruption Summary, published on July 29, and it has real consequences for how on-call teams practice incident response.