Operations | Monitoring | ITSM | DevOps | Cloud

Building AI SRE Agents, Part 3: Autonomous in the Cloud

Your agent has earned trust in shadow mode. Now it runs on its own: an alert fires, the agent starts, investigates and proposes a fix before anyone opens a laptop. Here is what it takes to make that safe, scalable and better every week. This is the third article in a series on taking an AI SRE agent from a weekend experiment to production. Part 1 built a local, read-only agent on a throwaway cluster and refined it against a synthetic eval set.

Context Engineering for AI Agents: What to Feed an Agent and What to Leave Out

Picture a Monday at 03:10 UTC. An agent investigating an out-of-memory alert on payment-service works out the pattern: it runs out of memory every Monday between 03:00 and 04:00, and the spike lines up with the batch reconciliation job. The next Monday the alert fires again. A different agent picks it up and starts from zero, because nothing it can see holds what the first one learned.

Why We Built the Komodor Agentic Operations Platform: Q&A with CEO Ben Ofiri

Komodor spent years building an AI SRE platform before the category had a name. With the launch of the Komodor Agentic Operations Platform, it’s opening that engine up so enterprises can build, run, govern and optimize their own agents in production. Following the launch, co-founder and CEO Ben Ofiri sat down to talk about why now is the right time for agentic operations, what breaks between prototype and production, and where operations will head next.