Operations | Monitoring | ITSM | DevOps | Cloud

When Vendor Support Ends, Your IT Monitoring Doesn't Have To

IT environments change. Technologies evolve, infrastructure vendors change their strategies, and sometimes support for a monitoring integration ends. For IT teams, that can create an immediate challenge: a previously monitored part of the infrastructure suddenly becomes a monitoring gap.

The most expensive half-hour of an incident.

It’s not the outage, it’s the stretch before you know what actually broke In short: VictoriaMetrics Enterprise support is expertise, not a ticket queue. It’s reactive by design (you reach engineers who know the stack when something breaks), with one proactive service, Monitoring of Monitoring, that watches the health of your VictoriaMetrics observability stack (metrics, logs, and traces).

Icinga vs Checkmk: Setup, Cost, Flexibility and Support

Icinga and Checkmk are the two open-source monitoring tools that turn up most often on the same shortlist. Both monitor IT infrastructure, and on a feature list they look close to interchangeable. In practice they are built on different assumptions, and those assumptions decide which one fits. This page works through the differences section by section: setup, customization, integrations, Windows, distributed monitoring, multi-tenancy, licensing, and cost.

Analyze your experiments in ChatGPT with the Datadog Experiments plugin

ChatGPT Work has become a common starting point for data and product teams. Analysts open it to compare launch adoption across segments, diagnose a metric that moved overnight, or turn a week of scattered numbers into a readout that a leader can act on. But the moment teams ask whether their experiment actually caused an effect they’ve observed, the conversation stalls.

10 Top Network Traffic Analysis Tools for Faster Troubleshooting and Capacity Planning

A request to upgrade a saturated circuit is easy to raise and hard to defend. The interface graph proves the link is full. It says nothing about which application, host or conversation filled it, so the spend gets approved on assumption instead of evidence. The same missing detail turns up everywhere else. Incidents run long because the cause is guessed at, capacity planning rests on estimates, and security questions arrive weeks after the traffic record expired.

Understanding NetFlow duplication: Why it happens, and how to deduplicate

NetFlow is a popular network protocol for collecting metadata about traffic flows across your environment so that it can be exported for analysis and monitoring. One of the most common issues that users encounter is NetFlow duplication, which occurs when identical flow records from the same conversation are recorded from different sources. Flow duplication inflates traffic data, undermining capacity planning and making top-talker rankings unreliable.

Moving Your Business Website: How to Avoid Email and Hosting Disruption

Website migration from one host to another is more than copying a few files to another server. A website might depend on databases, e-mail accounts, DNS records, SSL certificates, sub-domains, and other external applications that should continue to function after migration. Proper planning ensures the safety of the information and eliminates any risk of losing access to emails or having people visit a partially transferred website. This can be achieved by preparing the new hosting in advance.

RISE with SAP: Successfully Managing SAP Operations During the Transition

RISE with SAP is SAP’s methodology and commercial program for implementation and migration to SAP Cloud ERP Private. The product is SAP Cloud ERP Private, renamed in July 2025. Like RISE, GROW also transitioned from product to program, and the product parallel is SAP Cloud ERP Public.

How to troubleshoot JMX metric collection issues | Datadog Tips & Tricks

Missing JMX metrics make it hard to know what’s happening in a Java application, especially when vague errors or configuration mismatches make the cause difficult to diagnose. In this video, you’ll see how to troubleshoot common JMX metric collection issues and isolate the cause in less time.

ISO 20000 in ITSM: What the Standard Actually Requires From Your Service Desk

Certification against ISO 20000 puts your service desk under audit. That audit runs on what your team wrote down at the time. The standard does not care how your team describes its process. It does not care which ITIL 4 practices you adopted. It cares what your records show, so auditors spend their time in your tickets, approvals, and review minutes. In this blog, you will: You will finish knowing which of your records would survive an audit.

Grafana Tempo + Pyroscope: Profiles Traces (Sept 2026 Community Call )

Profiles + Traces and span redaction Can't comment in the chat? You may need to create a channel. Join us live for an introduction to flame graphs. We’ll cover what they are, how to read them, and how to use them to find performance bottlenecks in your applications. Bring your questions! Grafana Cloud is the easiest way to get started with Grafana dashboards, metrics, logs, traces, and profiles. Our forever-free tier includes access to 10k metrics, 50GB logs, 50GB traces and more.

Redact PII at the edge - and still be able to search for it

Ask a platform team why their application logs aren't in their observability backend and you'll often get a one-sentence answer: And, that's where the conversation ends. The logs stay in a silo. Or, they don't get collected at all. The team loses the troubleshooting signal, and nobody revisits the decision because the alternative looks like a compliance violation. Application logs in healthcare, aviation, insurance, and retail are full of personal information that should not be stored in plain text.

Custom labels in Grafana Cloud Synthetic Monitoring: New updates for consistency and ease-of-use

Labels are a powerful way to organize telemetry and define policies across Grafana Cloud, helping to streamline alerting, attribution, access control, and more. But traditionally, custom labels in Synthetic Monitoring have worked a little differently: they only lived on a single sm_check_info metric, and Grafana Cloud prefixed each one with label_.

9 Best Log Management Tools and What They Cost

Most log management tools bill you on log ingestion, the volume of data you send them. That works until your log volume doubles, and the invoice doubles with it. The best log management tools let you control what gets indexed and kept, so growth stops being a budget problem. In this blog, you will see: By the end you will know which fits your volume. Log management is the full lifecycle of your log data, from the moment it is collected to the moment it is deleted.

Best API Monitoring Tools in 2026 [31 Analyzed]

The best API monitoring tools are Hyperping for HTTP and API checks with on-call and status pages, Checkly for API monitoring as code, Postman Monitors for teams that already keep collections in Postman, Datadog for connecting failed checks to traces and logs, Grafana Cloud for teams using k6, Better Stack for checks inside a broader incident workflow, and UptimeRobot for inexpensive availability checks.

Troubleshoot Kafka issues across every layer of your stack with Kafka Console

Kafka is a crucial and widely used technology: 80% of the Fortune 100 rely on the event streaming platform as part of their stack, according to Apache. But Kafka issues can be complex to manage and even more difficult to troubleshoot, as the same symptom can point to very different problems. Suppose consumer lag on your checkout-events topic suddenly exceeds its SLA.

How we built Datadog Experiments

When Datadog acquires a company, we usually rebuild the product rather than plugging it in as is. That’s exactly what we did with Eppo, an experimentation and feature-management platform. Eppo’s feature-management capabilities became Datadog Feature Flags, while experimentation became Datadog Experiments. This post focuses on the experimentation platform and four changes we made to help you get to a decision faster.

Monitor smarter with Applications Manager's GenAI capabilities

GenAI has moved well past the pilot stage. According to a Gartner finding, by 2026, more than 80% of enterprises will have used GenAI APIs or deployed GenAI-enabled applications in production. Today, GenAI is becoming an integral part of how infrastructure and application teams work every day. Organizations are depending on LLMs from a diverse range of vendors—OpenAI, Anthropic, Google AI, and DeepSeek—based on the strengths each offer for different use cases.

MSP Ticketing System: How to Evaluate and Choose the Right Platform

Most managed service providers do not replace their ticketing tool because ticket volume got too high. They replace it because a client asked for a report the tool could not produce. That moment usually arrives between the third and tenth client. The shared inbox and the general help desk that carried you this far stop holding the accounts apart. Per-client reporting is usually the first thing to give way. An MSP ticketing system takes support requests from every client you serve.

Using writable bind mounts with Icinga 2 in rootless Podman

Running Icinga 2 in a rootless Podman container is pretty straightforward, it works just the same as on Docker, so all the examples on our Docker Hub page work as expected. For example this one to generate certificates and initialize the master configuration: Same as with mounting some existing configuration into the container: But once you want to combine the two, for example to store the certificates on the host and mount them into the container, generating or renewing certificates will fail.

The Essential Eight: Patching Applications and Operating Systems at Maturity Level Two

Why do so many patching programs pass every internal check and still come back from an Essential Eight assessment rated at Maturity Level One? The answer is rarely speed. Teams that miss the mark are usually patching their servers, browsers, and office suites on schedule, then losing the rating on the fifty applications nobody put on a list. Maturity Level Two is where the Essential Eight stops asking how fast you patch and starts asking how much you can see.

On a Network, an Agent Acts Where the Blast Radius Is Largest

Every network engineer carries an instinct that outsiders mistake for caution: a change in one place can travel. Reroute a path, push a policy, drop an interface, and the effect can ripple across campus, data center, WAN, and cloud before the first alert is read. The blast radius of a network change is the reason operators move deliberately, and it is the single most important thing an AI agent takes on the moment it is allowed to act on the network instead of merely describe it.

AI Norms & Values, Part 3 of 3: Things We Hold True

Welcome to the third and final part of our series on AI norms and values. Parts of this doc were extracted and published separately on substack; as a whole, they describe the principles we hold pertaining to technology and AI, and the ethical commitments we make to each other and our customers. We set out to write about AI, and ended up writing about ourselves. These documents are not meant to be aspirational ones; they are derived from how we do our work every day in honeycomb.

How we built data-driven AI Golden Paths at Datadog

As teams rush to adopt AI, they often find themselves with conflicting workflows unique to each individual developer. To manage costs and promote good development practices, organizations need to establish Golden Paths around AI usage. AI Golden Paths are standardized flows that help developers work with agents more reliably and effectively. But how do you sift through all the possible workflows to decide what these Golden Paths should be?

How to monitor Cypress tests with Grafana Cloud

If your Cypress suite has tests that fail more often or run slower, you know it can be hard to figure out the pattern from a single job. It could be one spec that slowed down, or a single test that fails, or maybe the entire suite is trending slower. The root cause could be a bug in the app, or a flaky test, or something else.

Challenges and Limitations of Open Source Software

Open source software (OSS) offers organizations access to flexible, customizable, and often free software. It can reduce licensing costs and provide access to a large community of developers and users. However, open source does not automatically mean free, simple, secure, or risk-free. Organizations using OSS can face challenges around technical support, security, licensing, maintenance, functionality, and internal expertise.

McKinsey Says Agentic Enterprises Need "Automated Guardrails." Here's What That Means

TLDR/: McKinsey’s new research on AI transformation, published August 28, 2026, studied 20 companies that have created real economic value from AI and found that only a small number have reached “Stage 3: Agentic AI enterprise.” The capability that separates Stage 3 from Stage 2, per McKinsey’s own maturity framework, is orchestration layers and automated guardrails: the ability to govern agent actions automatically, in real time, rather than reviewing them after the fact.

Log ingestion: you are probably paying to store logs you will never read

The default way to adopt log management is to ship everything and search it later. It is the path every vendor’s quickstart puts you on, and it is the reason log bills surprise people: ingestion is priced by volume, so“ship everything” is a spending decision disguised as a configuration default. The uncomfortable part is that most of that volume is never read. Nobody greps last Tuesday’s 200 OK access lines.

Raygun APM Agent 3.1: async traces that stay with the right request

Raygun APM Agent 3.1 introduces more accurate asynchronous request tracing for Windows, Linux, and Azure App Service. Version 3.0 rebuilt the foundation of the Agent, profiler, installers, and release pipeline. Version 3.1 builds on that work with a focused improvement for ASP.NET Core: automatic request correlation that follows asynchronous execution without requiring developers to instrument their application.

AppSignal vs the tools it replaces (PagerDuty, Cronitor, Rollbar etc.)

There’s no scenario in which you should be required to run six monitoring tools at once. OK, I may have been a bit dramatic there, you might actually be at a scale where you need it. But for the rest of us, it’s certainly overkill. Using UptimeRobot for, “Is the site up?”, Papertrail for logs, PagerDuty so someone actually gets notified… Tons of logins, tons of invoices, tons of separate configs, and the worst thing is, they are all unaware of each other.

From Audit Readiness to Continuous Control: Making Compliance Part of IT Operations

Compliance is often treated as a governance responsibility. But many of the conditions that determine whether controls continue to hold are created inside day-to-day IT operations. Operations teams manage the devices, configurations, changes, dependencies, and remediation activities where compliance can either remain aligned or begin to drift. Governance defines the requirements. Operations manages much of the environment where those requirements must remain true.

Azure Virtual Desktop Monitoring: A Complete Guide

Azure Virtual Desktop (AVD) puts the user’s desktop at the end of a long delivery chain: the Azure control plane, host pools, session hosts, profile storage, the network, and the endpoint on the user’s desk. Any one of them can make a session feel slow, and none of them looks broken from inside the others. That is why performance work on AVD starts with continuous monitoring across the whole chain rather than at either end of it. Azure Virtual Desktop Monitoring is what closes that gap.

Introducing Infrastructure Knowledge: Teach Netdata AI What Your Metrics Can't Show

Netdata AI sees everything your infrastructure does: every metric, every anomaly, every alert. It does not see what your infrastructure is: which services matter, which host is supposed to run hot, who owns what, what your team considers normal. Without that context, “CPU at 91%” is just a finding. With it, it might be a machine doing exactly its job.

Wide Events vs. Three Pillars: AI Observability Costs

As agentic AI workflows gain traction within organizations, those organizations are asking how to account for their behavior while keeping costs manageable. Some are sticking with the old three pillars of observability approach: take a measurement to create a metric, record output to a log, and track serial progress with a trace. Each of these is useful, but treating them as distinct formats from the start means paying for them distinctly too. Separate storage doesn't come cheap.

PCI DSS Requirement 10: Logging and Monitoring in v4.0.1

Version 4.0 renumbered PCI DSS Requirement 10 from end to end, and the Council retired v3.2.1 on 31 March 2024. Sub-requirement numbers written before then mostly point somewhere else now. Four more Requirement 10 rules changed status on 31 March 2025, automated log review among them. Checking your numbering against v4.0.1 costs an afternoon and saves a finding. In this blog, you will: PCI DSS Requirement 10 covers audit logging and monitoring across the cardholder data environment.

IT Service Management for Government and Public Sector Organizations

What happens when a citizen-facing portal fails on the last day of a filing deadline, and the only record of the outage lives in an email thread between two engineers? In a commercial organization, that is an operational embarrassment. In a government department, it is a hole in a statutory record that an auditor will eventually ask about.

Agentic Operations Start with Context: Build the Right Data Foundation

Episode 1, "Beyond the Thread: Deconstructing the Cisco Data Fabric Powered by the Splunk Platform," explores the intersection of data strategy and operational efficiency. Hosted by Splunk's Courtney Wright, the session features insights from experts Keith McClellan and Michael Sondag on the complexities organizations face in data management and operational models.

Assisted, Augmented or Agentic? Choose Your Splunk Starting Point

Episode two of Beyond the Thread explores how organizations can leverage a solid data foundation for AI-driven actions. Hosted by Courtney Wright and featuring experts Greg Ainsley-Malik and Sonal Pardeshi, the discussion delves into the Cisco Data Fabric, powered by the Splunk platform, and its role in transforming machine data into actionable insights. The episode highlights the journey towards agentic operations, addressing the challenges faced in moving from AI-ready data to effective implementations, and examines different adoption strategies that organizations may pursue.

How Observability and Real-Time Data Can Improve Warehouse Operations

Warehouse operations generate a constant stream of information. Goods are received, inventory moves between locations, orders enter picking workflows, stock levels change, and shipments leave the facility. When these activities are managed through disconnected systems or delayed manual updates, managers can struggle to understand what is actually happening on the warehouse floor.

The rise of autonomous digital operations

Monitoring has come a long way. Your team has dashboards, alerts, and automation that would've looked like magic a decade ago. Most days, things just work. But underneath all that tooling, a lot of the actual work still happens manually. An alert fires, and you pull the page-load metric from one tool, the user session logs from another, the backend trace from a third, and line them up until the story makes sense. Ten minutes, maybe fifteen pass, then you are able to fix it and move on.

Two cats, two dogs, four vendors, and the model the AI couldn't find (Tech Talk Companion)

Tech Talks went dark for a few months, and on episode 13 I finally got to ask why. Mathias Palmersheim’s answer, delivered completely straight, was that his users were unhappy with the availability and usability of their feeders and their litter box, and he wasn’t allowed back on stream until that got fixed. The users are two dogs and two cats, and they have titles. Maisie, a Shiba Inu who came to him through a rescue, is the recently promoted chief executive pawofficer.

Relational Query Superpowers

I'm investigating repeated errors in my e-commerce application, and I need to get enough context in a single Honeycomb query to piece the entire picture together. Each query returns events based on the event's WHERE clauses, but I want to know several things from outside of the event that recorded an error. Things like: Those attributes are all over the trace. That's going to make a single query tough, right? Wrong!

ITSM for Healthcare: IT Service Management in Hospitals and Health Systems

How long does a nurse stand at a workstation waiting for a record to load before the ward gives up and reaches for paper? In most hospitals nobody measures that number, and the ticket that reaches IT describes a symptom instead of a cause. Hospital IT support runs on a different clock from corporate IT. There is no quiet Sunday night and no safe window for maintenance.

Log Filtering: How to Cut Log Ingest Volume Without Losing Evidence

Every log estate reaches a point where volume grows faster than the value inside it. The usual response is to find the biggest source and drop it. Cutting volume is the easy part. Cutting the right half takes judgment. Log filtering is only one of four options for an expensive source, and the other three matter just as much. In this blog, you will: By the end you can defend every rule you write, including the ones that keep data. A volume cut fails in two directions.

From Monitoring to Prediction: How Fleet Data Is Changing Maritime Operations

Most maritime operators already collect more fleet data than their shore teams can meaningfully use. Positions appear on screens, real-time data arrives from onboard systems, and reports document vessel performance throughout a voyage. Tracking where a ship is has become the easy part. The harder question is what happens when that information starts indicating what the vessel will do next. The change in maritime operations comes down to shifting from reviewing events to anticipating them early enough to alter an operational decision.

Monitor HTTPS and SVCB Records with DNS Check

DNS Check now supports monitoring HTTPS records and SVCB records, DNS record types 65 and 64, both standardized in RFC 9460. They tell a client how to connect to a service rather than only where it is: which HTTP versions the endpoint speaks, which port it listens on, which addresses it can start connecting to, and which keys it needs for Encrypted ClientHello, all before it opens a connection.

Hybrid cloud management: 6 challenges IT teams need to solve in 2026

In 2026, a hybrid cloud is no longer something organizations are working toward; it's already where they are. According to Forrester's The State Of Cloud Series 2026, the vast majority of enterprises across major markets, including the United States, India, Australia and New Zealand, Canada, and the Asia-Pacific region, are running some form of a hybrid cloud, combining public cloud platforms with private infrastructure, colocation data centers, and sovereign cloud providers.

Build incident response workflows with Datadog Bits Chat

See how Bits Chat turns a natural-language request into an automated incident response workflow. In this demo, Bits Chat builds a workflow that investigates a monitor alert, identifies whether a recent deployment caused the issue, rolls it back when appropriate, and sends a summary to Slack.

Can we live dangerously? Sandboxing Claude, and the Claude foreman that runs the rest

While logging into one’s LinkedIn will spew out endless talk of AI possibilities from “thought leaders” and the semi-disconnected alike, another pocket of the world spent the last few weeks watching the Shai-Hulud worm chew through npm. A self-propagating credential stealer that hit 400-plus packages and, delightfully, planted Claude Code and VS Code hooks so just opening the repo could run its payload.

Live Debugging for Critical Systems: MTBF, MTTR & MTTA

A critical system has to stay reliable without new failures or added downtime, and live debugging, confirming the root cause without stopping the system, is often the only way to do that. In practice, this means having runtime context: on-demand evidence generated at the point of failure rather than logging configured months earlier, which is what keeps MTBF up, MTTR, and MTTA down.

Prompts, skills, and the AGENTS.md nobody wants to write (and how Anthropic writes theirs)

You’ve watched Claude Code compact a conversation. The context bar fills, it pauses, a summary appears, and it carries on like nothing happened. You probably assumed a housekeeping script trimmed the transcript in the background. It didn’t. The model compacted itself. When the window fills, Claude Code sends a long, specific prompt telling the model how to summarize its own conversation. Then it does, same model, same turn. The thing managing your context window is just another instruction.

Reliability Is the Test Agentic NetOps Has to Pass

It is 2:14 a.m. An agent has correlated a latency spike to an asymmetric routing condition and is ready to reroute traffic away from the affected path. The plan looks right. The only question that matters to the on-call SRE is whether to let it run, and that question is not really about the agent. It is about whether the picture the agent reasoned from is complete enough to trust at 2 a.m. with production on the line.

Agent vs Agentless Monitoring and How to Decide What Goes Where

Why does half the infrastructure end up returning no monitoring data? The standard plan is to install collection software on everything, which moves quickly across servers and stops dead at the first device running closed firmware. Storage arrays, firewalls, and switches will never accept an install, and the rollout stalls there. That plan usually gets set once for the whole environment, with a single collection model applied to hardware it was never suited for.

Help Desk Software for Schools: Managing IT Support Across Campuses

How many support requests reach school IT staff each week without ever becoming a ticket? A teacher stops a technician in the corridor about a projector that will not connect. An office administrator sends a direct email about a locked account, and a student tells the librarian their laptop stopped charging during second period. Help desk software for schools collects those requests into one queue, routes them by site and category, and keeps a record of what was done.

How to Cut SIEM Ingest by 90% Without Losing Detection Coverage

Every SOC team knows the trade-off. Send everything to the SIEM platform and pay for it. Or filter aggressively and risk missing something. Filter lists are written once, during onboarding. Detection content keeps moving after that. Smart Engine, the new core of the VirtualMetric DataStream pipeline, takes the guesswork out of that decision. It reduces SIEM ingest using your registered detection rules. An event that no registered detection could match is dropped.

Top Tips: How to be a tech-savvy traveler

Top tips is a weekly column where we highlight what’s trending in the tech world and share ways to stay ahead. This week, let's look at a few ways you can be a tech-savvy traveler. Being a traveler is not easy, but with today's modern technology, it has become much easier. When we travel to places with no network, we sometimes forget about the ways we can use technology. Excluding the more familiar, I'm going to list some lesser-known tips. 1.
Sponsored Post

Raygun APM Agent 3.0.14: faster, simpler, and ready for ARM64

Today we are releasing Raygun APM Agent 3.0.14 for Windows, Linux, and Azure App Service. This release is the result of a substantial modernization of the Agent, profiler, installers, and release pipeline. It makes Raygun APM easier to deploy, reduces overhead in several critical paths, adds native Linux ARM64 support, and lets developers investigate APM data through Raygun API v3 and the Raygun MCP server. If you are upgrading from version 2.3.0, there is much more here than a version-number change.

10 Top Website Monitoring Tools for Uptime, Page Speed and Real User Data

Your uptime tool reports the site as available all month. Support reports something else, because customers in one region spent the morning unable to complete a checkout. Both records are accurate, and that is the problem. An external check confirms the page answered. It cannot tell you that the answer took nine seconds for everyone routed through one CDN edge, or what caused the delay.

Automate Product Analytics reports with your agent and the CX CLI

Every page view, click, and session your RUM SDK captures lands in Coralogix as a log event under the cx_rum subsystem — the raw data behind how people actually use your product. You can turn it into a shareable report without writing a single query. Just ask your coding agent. Your agent queries that data through the CX CLI and writes the report for you: describe what you want in plain English, get a formatted report back — without leaving the terminal.

The pager shouldn't be what starts the investigation

Authored by Greg Janco, Engineering Manager at Mezmo I've always thought there was something backwards about incident response. An alert fires at 3 a.m. PagerDuty does its job. Somebody wakes up, grabs a laptop, connects to the VPN, opens the alert, and then starts answering the same basic questions we ask at the beginning of almost every incident. What changed? What else is broken? Have we seen this before? Some of those need a human eventually. A lot of the first pass doesn't.

How to use Grafana auto grid dashboard option for flexible layout across devices

Learn how to use the auto-grid option, a flexible panel layout that adapts to varying screen sizes and dynamic content. Creators can now define the max number of columns or max height of panels, making dashboard layouts more responsive and maintainable.

Build and run Datadog workflows from Bits Chat or AI agents

Teams use AI coding agents and Bits Chat to troubleshoot systems and handle complex tasks, often uncovering repetitive work worth automating. But turning those routines into workflows can still require switching tools and recreating context manually. Through the Datadog MCP Server, Workflow Automation now lets you build workflows from Bits Chat or AI coding agents like Claude Code, Cursor, and Codex.

Cribl On Your Coffee Break Episode 4 - Gathering REST data

In the 4th installment of our series, Leon looks at Cribl’s ability to collect REST API data. By the time the month (and the series) is over, you will have a pretty good idea of what Cribl can do, and how to do it. You’ll also have consumed your body weight in caffeinated beverages...

August 2026 Early Warning Signals

August brought notable outages across developer platforms, SaaS tools, communications services, and cloud applications. StatusGator detected 789 Early Warning Signals during the month. Of those, 153 incidents (19.39%) were acknowledged by providers, while 636 (80.61%) were not officially acknowledged. StatusGator’s Early Warning Signals often surface service disruptions before providers post an official update.

Getting Started with InfluxDB 3 and Grafana Tutorial

Summary This guide walks through an end-to-end Grafana and InfluxDB 3 integration using a realistic dataset you generate yourself. The tutorial covers getting data in, transforming it, connecting Grafana, and building real dashboards. Table of Contents InfluxDB and Grafana are the most common pairing in time series monitoring, and division of labor between them is simple.

Icinga Web SSO walkthrough

The ability to log into all corporate applications with one username and password is pretty convenient, even compared to a password manager. As a benefit, the IT department can centrally enforce one desired two-factor auth mechanism. Now we, Icinga, also provide a so-called OpenID Connect integration for single sign-on. By the end of this text you’ll know how to connect your Icinga Web instance to the ID provider of your choice.

SAP HANA Monitoring Tools 2026: How to Compare Options and Simplify Monitoring

If your team works with SAP HANA, the main challenge usually isn’t finding another dashboard. The real difficulty is identifying which tool can quickly help you trace vague complaints about slow transactions to their root cause. In many setups, SAP HANA monitoring is divided among native SAP interfaces, cloud monitoring, infrastructure dashboards, and the broader monitoring systems used by the rest of IT. This fragmented approach can slow down root cause analysis.

Cloud Cost Management for Observability: A Practical Guide

Observability spend is outgrowing infrastructure budgets. What drives the cost up, how pricing models work, and a practical framework to manage it. Sejal Pandey works on content and growth at Last9, writing about observability, reliability, and SRE practices.

6 Signs a Dedicated Log Tool Fits Better Than a Full Observability Platform

Most growing teams eventually consolidate onto a full observability platform, and for teams correlating logs, metrics, and traces across a complex system, that’s often the right call. But a dedicated log tool still wins for a specific set of teams: ones that need to move fast, keep costs simple, and get real answers from logs without carrying the weight of a platform they don’t fully need yet. Here’s when that’s you.

Cribl On Your Coffee Break Episode 3 - Configuring Prometheus Remote-Write

In day 3 of our coffee break series, Leon continues to explore common observability data types and how to get them into Cribl. Today, we’ll look at setting up a simple Prometheus ingestion. By the time the month (and the series) is over, you will have a pretty good idea of what Cribl can do, and how to do it. You’ll also have consumed your body weight in caffeinated beverages...

Monitor prompt caching to optimize your token usage

Datadog’s 2026 State of AI Engineering report showed organizations’ LLM inputs swelling rapidly as context engineering expands. In March 2026, 69% of all input tokens in Datadog customer traces were for system prompts: internal instructions, policy definitions, and tool guidance providing context and guardrails around the user input. This suggests that most context engineering spend among Datadog customers is going toward optimizing repeating system prompts in heavily scaffolded agent systems.

Why you should (not) build your own observability stack

If you are able to build it better than your vendor, then change your vendor. Not build it. Rishi builds large-scale observability systems at Last9, focusing on reliable and cost-efficient telemetry infrastructure, and writes about the practical lessons learned while operating ClickHouse, VictoriaMetrics, and OpenTelemetry in production.

PII Redaction in Logs: Mask, Redact, Hash, or Drop?

Sensitive values reach your logs without anyone deciding they should. A debug line prints a whole request object. An error message carries the query string. A customer email address is suddenly stored in three systems. PII redaction in logs then gets treated as one setting to switch on. In practice it covers four separate treatments. The value is already inside the message before log ingestion finishes. In this blog, you will: By the end you can write a rule for each field and defend it.

Log Processing: What Happens to a Log Line Before You Can Search It

A log line arrives as plain text and leaves as a record you can query. Six steps sit between those two states. Each one adds something useful, and each one costs you time, CPU, or storage. Most teams never look at that chain until a search comes back empty. Here is what log processing does to an event, step by step: By the end you can look at your own chain. You will know what each step buys you. Six steps turn a raw log line into a searchable record.

Azure integration now supports service principal authentication

We’ve released some improvements to our Azure status integration. StatusGator can now read your Azure Resource Health events via a service principal. Previously the only supported authentication mechanism was OAuth. Both pull the same data and produce the same alerts – the difference is who the connection belongs to, and what happens to it over time.

Bringing the Most Advanced Sampling to the OpenTelemetry Collector

Sampling is a core skill that everyone who runs an observability pipeline at scale will learn. There are lots of tradeoffs within the various decisions you'll make from reducing bandwidth, CPU, and memory, to reducing costs and making the observability backend's performance better for users. Historically, there have only been three mechanisms, each with their own tradeoffs: However, there is a secret fourth option: adaptive tail sampling—which changes those tradeoffs.

From traces to experiments: A loop for improving AI agents

Let’s say your team shipped a support agent last quarter. The launch demo went well, stakeholders were pleased, and everyone moved on. A few months later, things start to look off. Summaries of long conversations are truncated, and monitors show latency spikes on tool calls to the billing API. Your team’s first instinct is to ship fixes such as tweaking prompts or upgrading the model.

Visualize how CUPED adjusts experiment results with Datadog

CUPED (Controlled-experiment Using Pre-Experiment Data) is a powerful tool that can reduce metric variance and help teams obtain precise experiment results with less data. However, the difference between an experiment’s CUPED-adjusted lift and raw lift can be difficult to explain, especially when an experiment uses many pre-exposure metrics and subject properties. The CUPED adjustments visualization in Datadog Experiments breaks the difference into a sequence of specific adjustments.

We Let AI Agents Rewrite a 92M-Message-a-Day Service in Go. Zero Incidents.

Our Results Daemon processes about 92 million messages a day. We recently rewrote it from Node.js to Go, and we let Claude Code write it. We wanted to know whether we could trust an agentic rewrite for a critical, high-throughput production service rather than a prototype. It shipped with zero incidents, a 70% reduction in running pods, and a lighter database load.

Best Storage Monitoring Software: 10 Tools Compared

Storage rarely fails loudly. A pool fills. Latency climbs on one LUN. The first to notice is a user whose application timed out. The best storage monitoring software catches it earlier. It watches capacity, IOPS, latency and drive health across your arrays, which is what storage resource monitoring is for. In this blog, you will see: By the end you will know which one fits. Storage monitoring software tracks the health, capacity and performance of your IT storage.

How NIST Compliance Turns Observability Data Into Audit Evidence

Can you prove, on demand, which production systems were under continuous monitoring last quarter? Buyers, auditors, and insurers all ask a version of that question, and the answer decides contracts as often as audit findings. NIST compliance means aligning security controls and operations with standards from the National Institute of Standards and Technology, then holding evidence that the alignment stayed continuous. The frameworks are precise about outcomes and quiet about mechanics.

Compliance Doesn't Fail on Audit Day. It Drifts Every Day in Between.

For CIOs, compliance is no longer simply a box to check at audit time. It has become part of the operating standard for resilient, accountable, and well-managed enterprise IT. The reason is straightforward: enterprise technology environments change continuously. Infrastructure scales. Configurations change. Cloud resources move. Exceptions accumulate. Dependencies evolve across hybrid and distributed architectures.

Application Metrics caught my broken size estimator

There’s a very specific kind of frustration that comes from waiting several minutes for a video to encode, dragging it into a message, and getting hit with a “file too large” error. Then you’re blindly trying to shave off a few more megabytes by re-encoding, maybe at a lower resolution or a smaller bitrate, hoping you won’t have to do it more than one or two more times. Here’s how I used Sentry’s Application Metrics to make a more accurate video size estimator.

Cribl On Your Coffee Break Episode 2 - Setting up Syslog

In our second video Leon picks on Syslog (because honestly, it deserves it). Cribl is the perfect tool to whip that disorganized, loud, unruly mess of a data stream into shape. By the time the month is over, you will have a pretty good idea of what Cribl can do, and how to do it. You’ll also have consumed more caffeinated beverages than is strictly appropriate...

Data-Driven Decisions Accelerate IT Results

Modern IT teams, having moved beyond the traditional reliance on hunches and personal experience that once shaped their day-to-day choices, no longer operate on intuition, since every meaningful decision now rests upon measurable, verifiable evidence gathered from their systems and workflows. Every deployment, capacity change, and incident response now depends on measurable evidence, not guesswork. Companies that base their operations on concrete numbers ship faster, recover quicker, and allocate budgets with far greater accuracy.