Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on Monitoring for Websites, Applications, APIs, Infrastructure, and other technologies.

When Vendor Support Ends, Your IT Monitoring Doesn't Have To

IT environments change. Technologies evolve, infrastructure vendors change their strategies, and sometimes support for a monitoring integration ends. For IT teams, that can create an immediate challenge: a previously monitored part of the infrastructure suddenly becomes a monitoring gap.

10 Top Network Traffic Analysis Tools for Faster Troubleshooting and Capacity Planning

A request to upgrade a saturated circuit is easy to raise and hard to defend. The interface graph proves the link is full. It says nothing about which application, host or conversation filled it, so the spend gets approved on assumption instead of evidence. The same missing detail turns up everywhere else. Incidents run long because the cause is guessed at, capacity planning rests on estimates, and security questions arrive weeks after the traffic record expired.

Analyze your experiments in ChatGPT with the Datadog Experiments plugin

ChatGPT Work has become a common starting point for data and product teams. Analysts open it to compare launch adoption across segments, diagnose a metric that moved overnight, or turn a week of scattered numbers into a readout that a leader can act on. But the moment teams ask whether their experiment actually caused an effect they’ve observed, the conversation stalls.

Understanding NetFlow duplication: Why it happens, and how to deduplicate

NetFlow is a popular network protocol for collecting metadata about traffic flows across your environment so that it can be exported for analysis and monitoring. One of the most common issues that users encounter is NetFlow duplication, which occurs when identical flow records from the same conversation are recorded from different sources. Flow duplication inflates traffic data, undermining capacity planning and making top-talker rankings unreliable.

The most expensive half-hour of an incident.

It’s not the outage, it’s the stretch before you know what actually broke In short: VictoriaMetrics Enterprise support is expertise, not a ticket queue. It’s reactive by design (you reach engineers who know the stack when something breaks), with one proactive service, Monitoring of Monitoring, that watches the health of your VictoriaMetrics observability stack (metrics, logs, and traces).

Icinga vs Checkmk: Setup, Cost, Flexibility and Support

Icinga and Checkmk are the two open-source monitoring tools that turn up most often on the same shortlist. Both monitor IT infrastructure, and on a feature list they look close to interchangeable. In practice they are built on different assumptions, and those assumptions decide which one fits. This page works through the differences section by section: setup, customization, integrations, Windows, distributed monitoring, multi-tenancy, licensing, and cost.

Moving Your Business Website: How to Avoid Email and Hosting Disruption

Website migration from one host to another is more than copying a few files to another server. A website might depend on databases, e-mail accounts, DNS records, SSL certificates, sub-domains, and other external applications that should continue to function after migration. Proper planning ensures the safety of the information and eliminates any risk of losing access to emails or having people visit a partially transferred website. This can be achieved by preparing the new hosting in advance.

RISE with SAP: Successfully Managing SAP Operations During the Transition

RISE with SAP is SAP’s methodology and commercial program for implementation and migration to SAP Cloud ERP Private. The product is SAP Cloud ERP Private, renamed in July 2025. Like RISE, GROW also transitioned from product to program, and the product parallel is SAP Cloud ERP Public.

Troubleshoot Kafka issues across every layer of your stack with Kafka Console

Kafka is a crucial and widely used technology: 80% of the Fortune 100 rely on the event streaming platform as part of their stack, according to Apache. But Kafka issues can be complex to manage and even more difficult to troubleshoot, as the same symptom can point to very different problems. Suppose consumer lag on your checkout-events topic suddenly exceeds its SLA.

How we built Datadog Experiments

When Datadog acquires a company, we usually rebuild the product rather than plugging it in as is. That’s exactly what we did with Eppo, an experimentation and feature-management platform. Eppo’s feature-management capabilities became Datadog Feature Flags, while experimentation became Datadog Experiments. This post focuses on the experimentation platform and four changes we made to help you get to a decision faster.