Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on Monitoring for Websites, Applications, APIs, Infrastructure, and other technologies.

Monitoring Fundamentals in Six Lessons: From Golden Signals to Error Budgets

Summer is the perfect time to step back from the alert stream and revisit the ideas your whole monitoring practice quietly rests on. Not the newest tool or the shiniest dashboard — the fundamentals: what to measure, how to read it, when to wake someone up, and what to do when something breaks at 2 a.m. We packed six of those fundamentals into a free, no-signup PDF — The DevOps & Monitoring Summer Workbook, complete with hands-on exercises and an answer key.

Best Anomaly Detection Software: 9 Tools Compared on Cost and Coverage

Does the anomaly you need to catch show up in infrastructure, in security logs, in a data pipeline, or in a revenue figure? If you already know the answer, you probably learned it from an incident. A service degraded quietly, nobody got paged, and the post-mortem showed the signal had been in the data for hours. The thresholds were set correctly, and they still could not separate a busy Tuesday from a failure.

Triage Production Incidents with a Single Prompt Using the AppSignal CLI

At AppSignal, we love talking to our customers to learn how they're using the product. Recently, one of them showed us something worth sharing: with a single chat prompt, his AI agent searches production logs, closes incidents, and checks whether a pull request has shipped. Automating that kind of work turned out to be a matter of building the right agent skill.

CLIs are more token-efficient than MCP. Or are they?

MCP servers have a reputation: they eat your context window. CLIs paired with skills, on the other hand, are more token efficient. But is this still true? I dropped all my MCP servers five months ago. Five months is a long time in AI land. When Anthropic came up with the concept of skills, many people stopped using MCP servers in favor of CLI tooling and skills.

GPU monitoring in OpManager: Full visibility for every AI workload

AI has moved to be a core part of enterprise infrastructure. GPUs are the engines behind that shift. Every training run, every inference request, and every fine-tuning job depends on GPU chipsets that are expensive and delicate. A GPU that overheats, runs out of memory, or sits idle for hours doesn't just slow a project down, it quietly drains the IT budget. Most monitoring tools weren't built with this hardware in mind. This leaves AI and DevOps teams blindsided when a job fails or a chipset degrades.

LAS Migration Aftermath: What Happened to Your License Usage Metrics?

When Citrix introduced License Activation Service (LAS), customers transitioned from the traditional file-based licensing model to a new cloud-connected licensing architecture. For most administrators, the migration itself was straightforward. The biggest surprise came afterwards, when familiar licensing metrics such as licenses in use, licenses available, and peak license usage disappeared from the Citrix License Server. The answer is no.

Choose a group when adding monitors

We’re making it easier to organize your monitors from the moment you add them. Previously, clicking Add would immediately add a monitor to your board. If your board already contained one or more groups, new monitors were automatically placed in the Ungrouped section, requiring an extra step to move them into the correct group. Now, clicking Add opens a list of your monitor groups, allowing you to choose exactly where the monitor should be added.

AI-Powered Spacecraft Operations with InfluxDB 3

Summary The InfluxDB satellite telemetry demo is a live mission-control application that monitors a simulated fleet of 12 satellites in real-time. It shows how the InfluxDB 3 Processing Engine can detect anomalies as data is written, enrich time series data with third-party data, and power a grounded AI agent using the InfluxDB 3 MCP server to investigate and explain fleet health.

Our 3-month AI roadmap - the future of smart dashboards

AI is set to transform our technology landscape. For many of us working in software, it already has — developers are now writing more code, building more features, and deploying more applications, faster. For the teams supporting IT and software services that means more applications to support, across a greater breadth of technologies, and with more complexity (that is probably less well understood by the developers who created it). Your operational tooling needs to keep pace.

See, Tell, Do: The Auvik Story

See. Tell. Do. For nearly 15 years, these three words have shaped how we build Auvik. First, help IT teams see what's happening across their networks. Then, help them understand why it's happening and what needs attention. Finally, empower them to take action with confidence. As we prepare to celebrate our 15th anniversary, we're excited to share the story behind that philosophy - and why we believe the best is still to come.