Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on Monitoring for Websites, Applications, APIs, Infrastructure, and other technologies.

Introducing Infrastructure Knowledge: Teach Netdata AI What Your Metrics Can't Show

Netdata AI sees everything your infrastructure does: every metric, every anomaly, every alert. It does not see what your infrastructure is: which services matter, which host is supposed to run hot, who owns what, what your team considers normal. Without that context, “CPU at 91%” is just a finding. With it, it might be a machine doing exactly its job.

Assisted, Augmented or Agentic? Choose Your Splunk Starting Point

Episode two of Beyond the Thread explores how organizations can leverage a solid data foundation for AI-driven actions. Hosted by Courtney Wright and featuring experts Greg Ainsley-Malik and Sonal Pardeshi, the discussion delves into the Cisco Data Fabric, powered by the Splunk platform, and its role in transforming machine data into actionable insights. The episode highlights the journey towards agentic operations, addressing the challenges faced in moving from AI-ready data to effective implementations, and examines different adoption strategies that organizations may pursue.

Agentic Operations Start with Context: Build the Right Data Foundation

Episode 1, "Beyond the Thread: Deconstructing the Cisco Data Fabric Powered by the Splunk Platform," explores the intersection of data strategy and operational efficiency. Hosted by Splunk's Courtney Wright, the session features insights from experts Keith McClellan and Michael Sondag on the complexities organizations face in data management and operational models.

From Audit Readiness to Continuous Control: Making Compliance Part of IT Operations

Compliance is often treated as a governance responsibility. But many of the conditions that determine whether controls continue to hold are created inside day-to-day IT operations. Operations teams manage the devices, configurations, changes, dependencies, and remediation activities where compliance can either remain aligned or begin to drift. Governance defines the requirements. Operations manages much of the environment where those requirements must remain true.

Azure Virtual Desktop Monitoring: A Complete Guide

Azure Virtual Desktop (AVD) puts the user’s desktop at the end of a long delivery chain: the Azure control plane, host pools, session hosts, profile storage, the network, and the endpoint on the user’s desk. Any one of them can make a session feel slow, and none of them looks broken from inside the others. That is why performance work on AVD starts with continuous monitoring across the whole chain rather than at either end of it. Azure Virtual Desktop Monitoring is what closes that gap.

AppSignal vs the tools it replaces (PagerDuty, Cronitor, Rollbar etc.)

There’s no scenario in which you should be required to run six monitoring tools at once. OK, I may have been a bit dramatic there, you might actually be at a scale where you need it. But for the rest of us, it’s certainly overkill. Using UptimeRobot for, “Is the site up?”, Papertrail for logs, PagerDuty so someone actually gets notified… Tons of logins, tons of invoices, tons of separate configs, and the worst thing is, they are all unaware of each other.

How Observability and Real-Time Data Can Improve Warehouse Operations

Warehouse operations generate a constant stream of information. Goods are received, inventory moves between locations, orders enter picking workflows, stock levels change, and shipments leave the facility. When these activities are managed through disconnected systems or delayed manual updates, managers can struggle to understand what is actually happening on the warehouse floor.

The rise of autonomous digital operations

Monitoring has come a long way. Your team has dashboards, alerts, and automation that would've looked like magic a decade ago. Most days, things just work. But underneath all that tooling, a lot of the actual work still happens manually. An alert fires, and you pull the page-load metric from one tool, the user session logs from another, the backend trace from a third, and line them up until the story makes sense. Ten minutes, maybe fifteen pass, then you are able to fix it and move on.