Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on Monitoring for Websites, Applications, APIs, Infrastructure, and other technologies.

What Is the MITRE ATT&CK Framework? A Guide for IT Ops Teams

Most IT operations teams cannot say how much of the MITRE ATT&CK framework they already cover. The framework gets explained in the language of threat hunting and red teams. The parts that belong to infrastructure work are easy to miss. And then, coverage questions get answered with a guess. The mismatch costs time on both sides. Security asks for a coverage answer that ops has no clean way to produce. Yet the controls that stop a large share of those techniques already sit with your team.

Vulnerability Assessment and Penetration Testing: Differences, Cadence, and Cost

What do you say when an auditor asks for evidence that your security controls hold, and all you can produce is a scan report from last month? A scan lists weaknesses. It says nothing about whether an attacker could chain three of them together and reach the customer database. Vulnerability assessment and penetration testing answer two different questions about the same environment. The first asks what is exposed right now. The second asks what someone with intent and skill could do with that exposure.

Azure Virtual Desktop Monitoring: Challenges, Metrics & Best Monitoring Solutions

Azure Virtual Desktop (AVD) is rapidly growing in popularity as modern way to deliver virtual desktops and apps to users, with Azure providing the infrastructure as alternative to on-prem VDI environments. As organizations expand their use of AVD in Azure, monitoring becomes critical.

AMA Recap: More Answers From the Observability Engineering Authors

Last week, we sat down with the authors of Observability Engineering for a live AMA. We ended up getting so many questions (pre-submitted and live) that we couldn't get through them all. Charity, Liz, George, and Austin kindly stuck around afterward to answer more, ranging from low-hanging observability fruits and telemetry to AI and what software engineers can do that Claude can't. Missed the live session? Watch it on demand now.

From Vision to Value: New Splunk Platform Innovations Supporting Cisco Data Fabric Are Generally Available

At.conf25, we announced our vision for Cisco Data Fabric, an architecture designed to help organizations unlock the value of machine data, fuel AI with trusted context, and support more intelligent and resilient operations. Today, that vision has become reality. Key Splunk Platform innovations including Machine Data Lake, Catalog, and Agent Launchpad, together with expanded Federated Search and Data Management capabilities, are now generally available.

Building an AI Observability Agent: Lessons from the Trenches - Stripe at O11yCon 2026

Stripe shares lessons from building an incident investigation agent, from context-window blowups to why the final 5% still needs a human. In this O11yCon 2026 talk, they dig into what it takes to go from 'it works' to 'it works reliably,' including how pointing agents at like Honeycomb's speeds up on-call investigations.

Signal vs. Spend: Building Cost-Aware Observability at Slack - O11yCon 2026

It started with a single log line taking up a massive amount of volume: 500 million emissions per hour. Pulling that thread led Emma and Steven into Slack's broader logging pipeline: 311 billion logs per day at 4.4M/sec peak, with no volume limits, no per-service attribution, and no feedback to the teams generating the noise.

Where Historians Fall Short for Physical AI

Summary Physical AI—machines and industrial systems that sense conditions, reason, and act in the real world—needs two things from operational data: detailed history for training, and real-time telemetry for inference. Traditional data historians weren’t built for either at the speed Physical AI requires. Four gaps result: limited real-time access, compression that strips model-relevant signal, IT/OT fragmentation, and site-by-site architectures.

Introducing MCP Connections: Netdata AI Now Reads From the Tools You Already Run

Netdata AI can now connect outward to the tools your team already runs, like GitHub, PagerDuty, Atlassian, or any custom MCP server, and read from them during an investigation. We call this MCP Connections. It’s the missing piece in the middle of every root-cause investigation: the alert tells you what changed, but the why is usually somewhere else entirely.