Operations | Monitoring | ITSM | DevOps | Cloud

Session Replay: Reproducible User Sessions

See what your users see. Don't try and guess what happened, watch it. Session Replay is a feature for Mobile and Web apps on Sentry to record user sessions as reproducible sessions. These are not videos. On web, a session replay contains the entire DOM reconstruction and DevTools. It's exactly like inspecting in Chrome DevTools locally on your machine, but in a real production session.

AI Amplifies Your Existing Practices: Lessons from Our Shift to an AI-First Strategy

In this two-part blog series, I give a detailed report-out on how our Honeycomb engineering team 2.5x-ed our throughput using AI without breaking everything or lowering our standards for quality. Part 1 explains how we did it and shows data about how that ramp-up happened. In this blog, I share what we learned. The “platform engineering” frame and the “autonomy, ownership, feedback loops” frame are the same frame, spoken in two different vocabularies.

30 to 70 PRs a Day: How We Managed to Not Wreck Our Systems

In this two-part blog series, I give a detailed report-out on how our Honeycomb engineering team 2.5x-ed our throughput using AI without breaking everything or lowering our standards for quality. Part 1 explains how we did it and shows data about how that ramp-up happened. Part 2 shares what we learned.

The Investigator That Remembers: Inside Klaudia Memory

There is a particular kind of incident every SRE team is familiar with. A common component of your stack, say your Redis database, starts misbehaving. Someone spends two hours tracing it back to a connection pool exhausted by a misconfigured client, the fix goes in, and everyone moves on, for today. The following Tuesday it happens again, and whoever is on call investigates it from scratch, because the person who solved it last week is asleep, on vacation, or working somewhere else now.

Tealium's Dr. Martin Nettling on reviewing AI-generated work

Cortex co-founder and CTO Ganesh Datta sits down with Dr. Martin Nettling, Senior Director of Engineering and Head of QA at Tealium, to explore the distinction between trusting people and having confidence in tools, and why that difference matters as AI becomes part of every engineering workflow.

What is MTTR, and how can agentic ITOps reduce it?

Mean time to resolution (MTTR) measures the average duration to restore regular operation for an application, service, or infrastructure component. It’s a key performance indicator (KPI) for IT incident management. To tie MTTR directly to customer satisfaction, you first need to understand how it affects service and application reliability and availability. From there, you can make informed decisions, operate efficiently, and provide a seamless customer experience.

Why multicloud has become a governance decision

For most IT leaders, multicloud didn't arrive as a decision. It arrived as a fait accompli. A team chose AWS for one workload. Azure came in through a Microsoft enterprise agreement. A SaaS acquisition brought its own cloud dependencies. A DR requirement pointed to a second region with a different provider. Nobody declared a multicloud strategy; the organization just became one. Today, 87% of organizations run a multicloud strategy, balancing an average of 2.6 public cloud providers simultaneously.