Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on AIOps, alerting in complex systems and related technologies.

The enterprise changed. ITOps didn't.

The modern enterprise runs on a technology stack that changes faster than the operating model responsible for keeping it available. Applications that once moved through scheduled releases now change continuously. Infrastructure is distributed and dynamic. Services depend on other services, teams depend on other teams, and operational data arrives from more places than any individual can reasonably inspect. The business asked for agility, flexibility, and velocity, and the tech delivered.

On a Network, an Agent Acts Where the Blast Radius Is Largest

Every network engineer carries an instinct that outsiders mistake for caution: a change in one place can travel. Reroute a path, push a policy, drop an interface, and the effect can ripple across campus, data center, WAN, and cloud before the first alert is read. The blast radius of a network change is the reason operators move deliberately, and it is the single most important thing an AI agent takes on the moment it is allowed to act on the network instead of merely describe it.

How to Build an HR PTO AI Agent with Resolve Agent Lab

See how to build an HR PTO agent with Resolve Agent Lab. In this Resolve Reels demo, we create a purpose-built AI agent by adding automation skills, instructions, conversation starters, and guardrails. The agent can answer PTO questions, check balances, account for calendar conflicts, and submit requests through systems like Workday or ADP. See how Resolve helps teams build AI agents that take action across enterprise systems.

From Audit Readiness to Continuous Control: Making Compliance Part of IT Operations

Compliance is often treated as a governance responsibility. But many of the conditions that determine whether controls continue to hold are created inside day-to-day IT operations. Operations teams manage the devices, configurations, changes, dependencies, and remediation activities where compliance can either remain aligned or begin to drift. Governance defines the requirements. Operations manages much of the environment where those requirements must remain true.

Introducing Swarm Investigation from BigPanda: Autonomous, multi-agent IT incident investigation

When a major incident opens, the opening minutes often become a race across disconnected tools and competing theories. One engineer checks a monitoring tool. Another scrolls through change records, looking for the one line that explains everything. A third pings Slack, asking if anyone has seen this before. While these are reasonable steps, taken one at a time, in sequence, they are far too slow.

Reliability Is the Test Agentic NetOps Has to Pass

It is 2:14 a.m. An agent has correlated a latency spike to an asymmetric routing condition and is ready to reroute traffic away from the affected path. The plan looks right. The only question that matters to the on-call SRE is whether to let it run, and that question is not really about the agent. It is about whether the picture the agent reasoned from is complete enough to trust at 2 a.m. with production on the line.

How Is AI Changing IT Operations? Building Production-Ready AI Agents with Alex Zinovy

How is AI changing IT operations, and what does it take to move AI agents from impressive demos to production-ready systems? In this episode of Agents of IT, Resolve’s Zack Austin sits down with Alex Cinovoj, Founder and CTO of TechTide AI, to explore what enterprise AI looks like when it has to work in the real world. Alex brings years of hands-on IT, infrastructure, DevOps, and AI engineering experience to a conversation about the shift from experimenting with AI to building trustworthy systems that deliver measurable outcomes.