Operations | Monitoring | ITSM | DevOps | Cloud

Cloud Outage Preparedness: On-Call Lessons for 2026

Cloud outage preparedness stopped being a nice-to-have this month. In a span of roughly 48 hours, Microsoft Azure lost a big chunk of its West US footprint and Amazon Web Services dropped connectivity between its us-west-2 region in Oregon and the Seattle metro. The AWS event alone rippled outward and knocked DoorDash, Reddit, Hulu, Apple Pay, Snapchat, Fortnite, and the PlayStation Network offline for millions of users, according to incident trackers. Neither outage was caused by anything exotic.

Say You're Struggling - People Want to Help

Being honest about what you're struggling with isn't weakness — it's the fastest way to get real help. This IT leader says it's okay to tell your team "I'm struggling with X, can I just vent for 15 minutes" and skip the technical fix for a minute. Most people want to help — you just have to let them. Watch the full IT Leadership Lab session.

The Upsides to Keeping it Boring: How Fiber-to-the-Home Providers Protect Their Investments

What does "good" actually look like for a fiber-to-the-home network? According to two of the people who build them, it looks like nothing at all. No complaints. No support calls. No coffee-talk the next morning about the stream that buffered. Just a light switch that turns all the way on, every time.

Tech Talk | From Insight to Action: Reducing Metrics Costs in Splunk Observability

Explore how to leverage observability insights to reduce metric costs with Span Observability. Led by Tomasz Romaniuk and Martyna Karbownik from Cisco, this tech talk covers the significance of metric time series in impacting cardinality and consequently driving metric costs. After a theoretical introduction, the presentation dives into practical real-life use cases demonstrating how to effectively utilize available tools for cost reduction. The session concludes with an overview of future developments and a Q&A segment for audience inquiries.

4 Cloud-Native Challenges AI SRE Is Solving in 2026 and the 3 New Ones to Look Out For

AI SRE is making real strides in resolving some of the greatest pains related to incident response, troubleshooting, and complex root cause analysis. The on-call rotation, the war room, the week-long RCA, and the ticket queue that ate a third of every platform engineer’s week all look different now than they did two years ago.