Operations | Monitoring | ITSM | DevOps | Cloud

That 2am incident cost you more than downtime

It's 2:14am. A pager alert wakes an on-call engineer for a checkout failure hitting a slice of customers. By the time they've pulled the right logs, cross-checked the deploy history, and finally gotten the bug to happen again in front of them, the sun's coming up. The incident report will list two hours of downtime. It won't say anything about the day that engineer just lost, or the fact that nobody on your team could have told you in advance how long that reproduction step was going to take.

Governing AI Agents From the Inside: What We Learned Building AgentIQ

When every employee is building AI agents, seeing what they did afterward isn't enough. AgentIQ's in-flow governance runs inside the agent's execution flow—pausing for human approval, enforcing policies by value, masking PII, and stopping runaway agents before costs spiral.

Icinga Web Filter Syntax: Operators, Wildcards and Custom Variables

An Icinga Web filter is a chain of conditions, each one a column, an operator, and a value, combined with & and |. You type it into the search bar of Icinga DB Web, but the same syntax shows up in URLs, dashboards, role restrictions, and the event rules of Icinga Notifications Web. This post covers the whole thing: operators, wildcards, grouping, and filtering all the way down into arrays and dictionaries.

8 Best Network Diagnostic Tools for 2026

If your team manages a network across several offices and cloud systems, finding the cause of a slow or failed connection can take time. You may run into problems such as: As your network grows, a quick test from one computer may not show where the problem started or when it first appeared. Network diagnostic tools help you check devices, ports, routes, packets, and link performance.

9 Best Help Desk Software for Small Business in 2026

Why is your network slow or failing even when your devices seem to be working? Finding the cause of a network problem can take time. You may need to check packet loss, trace the network path, find a busy link, or see why an application is not connecting. A simple ping test can tell you whether a device responds, but it cannot always show what is causing the problem. Network diagnostic tools help you narrow down the cause by checking devices, ports, routes, packets, traffic, and link performance.

Air-Gapped vs Private Network Backup: Which Is More Secure?

Compare air-gapped and private network backup approaches for ransomware defense, data protection, and secure recovery. Ransomware doesn’t stop at production data. Attackers are increasingly targeting the backups organizations depend on for recovery, turning a manageable incident into a much longer outage. These attacks make the choice between air-gapped and private network backup important — but it isn’t an either-or decision.

Why AI development creates a reliability blind spot for humans, and what to do about it

Application development and operations teams are adopting AI coding tools at an exponentially increasing rate, from a rounding error of 6% of code output by AI in 2023, to as much as 51%-75% for a majority of enterprises. Agents are making pull requests (PRs) faster than human developers could have ever dreamed. As features are pushed to market faster, there’s a sharp increase in production incidents, with 80% of development shops specifically tracing production outages to AI.

Announcing Gremlin Foresight AI

Today we're launching Foresight AI, Gremlin’s agentic resilience product that analyzes and tests your systems for potential failures, fixes them, and verifies reliability at the speed of AI. I've spent most of my career on call. At Amazon and Netflix, I served as a Call Leader, the person running the bridge when something big broke. Those years taught me the same lesson we founded Gremlin on: the best incident is the one that never happens.

Context Engineering for AI Agents: What to Feed an Agent and What to Leave Out

Picture a Monday at 03:10 UTC. An agent investigating an out-of-memory alert on payment-service works out the pattern: it runs out of memory every Monday between 03:00 and 04:00, and the spike lines up with the batch reconciliation job. The next Monday the alert fires again. A different agent picks it up and starts from zero, because nothing it can see holds what the first one learned.