Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on Monitoring for Websites, Applications, APIs, Infrastructure, and other technologies.

Datadog achieves ISO 42001 certification for responsible AI

As AI-powered products and services become central to how organizations operate, the need for responsible AI governance has never been greater. Customers, partners, and regulators are seeking assurance that AI systems are built, managed, and monitored responsibly and effectively. Datadog is committed to the responsible use of AI, both in how we build our products and in how we help customers observe their AI workloads.

Monitor Nutanix clusters, hosts, and VMs with Datadog

Nutanix is a hyperconverged infrastructure (HCI) platform that combines compute, storage, and virtualization into a single software-defined stack. By collapsing traditional infrastructure tiers into one platform, Nutanix simplifies provisioning and operations for virtualized workloads. Clusters are managed through Prism Central, which provides visibility into health, performance, capacity, and operational activity across hosts and VMs.

What Are Containers? (And Why "It Works on My Machine" Finally Dies)

What are containers in DevOps—and why do they solve the classic “it works on my machine” problem? In this episode of Cloud Security in a Minute, Sysdig breaks down containers in simple terms: what they are, how they work, and why they’ve become the backbone of modern cloud applications. You’ll learn: Containers package everything an application needs—code, dependencies, and system tools—so it runs consistently anywhere: your laptop, the cloud, or at massive scale.

Olivier Pomel and Alexis Lê-Quôc on Datadog's origin, AI, and more | This Month in Datadog

Get an insider’s view of Datadog from the people who built it. On a special episode of This Month in Datadog, co-founders Olivier Pomel and Alexis Lê-Quôc sit down for a rare, in-depth look at the challenge that inspired them to build the Datadog platform, what the company is working on today, AI, and more. This Month in Datadog brings you the latest updates on our newest product features, announcements, resources, and events.

Beyond the Queue: Modernizing Legacy Middleware with Apache Kafka 4.x

Apache Kafka 4.x eliminates the final barriers to legacy middleware modernization. With KRaft mode removing ZooKeeper dependency and native queue semantics bridging the gap, enterprises can finally transition from point-to-point messaging to event-driven architectures.

Monitoring Your App Without Running Your Own Prometheus Stack

Prometheus and Grafana are the default monitoring recommendations across DevOps blogs, Reddit, and Hacker News, and for good reason. Prometheus is open-source and backed by the CNCF, but it’s not actually a complete monitoring system. It’s more of a metric collection engine.

Debunking the Myth of the Homogeneous Network

If you have been in network operations for more than a week, you know the dream of the single vendor shop is exactly that, just a dream. In the practical reality of your daily job, the network is a diverse, chaotic ecosystem. It is a complex stack in which layers of technology from different times and vendors coexist, often uneasily.

Observability Is Now a Boardroom Priority Even If Nobody Wants to Say It Out Loud

Executives rarely state the full truth publicly, but inside boardrooms the conversation has changed. Observability, once viewed as a technical capability deep within operations, has become a strategic requirement for understanding business performance. Leaders may not always use the term itself, yet they focus intensely on the outcomes it promises. Their environments have grown too fast, too fragmented, and too interdependent for traditional visibility approaches to keep pace.