Operations | Monitoring | ITSM | DevOps | Cloud

Building a Self-Service Knowledge Base Employees Want to Use

Most organizations already have policy documents, troubleshooting guides, knowledge articles, resolved tickets, runbooks, and internal wikis. They're not starting at zero, but despite that investment, employees continue to open tickets for questions the organization has already answered. The problem is not always a lack of knowledge. More often, employees cannot find the right information quickly enough to trust self-service as their first option.

How savepoints quietly throttled our Postgres queue

At incident.io we are huge fans of Postgres; we've written about it a lot over the years, including how to choose the right indexes and how we're proud of being boring (The Pet Shop Boys). We use Postgres as our primary transactional database, which as of today has ~900 tables, and counting! The vast majority of our codebase does something along the following lines: read some data from Postgres, execute some business logic, then write that data back to Postgres. It is not, however, always that simple.

PostgreSQL MCP: Manage Postgres From Your AI Assistant

TL;DR Most of the time we understand the tasks we're working on, but inevitably something comes up that we don't know much about, and we need just enough skill to cope. That used to mean reading paper documentation, then searching vendor sites and the web. Now it tends to mean a dialog with an LLM that has already ingested the documentation we don't have time to find and read. That's only half the problem, though.

How to Guarantee a Website or Service Never Goes Down (And What You Can Actually Promise)

No one can guarantee that a website or service never goes down. What you can promise is a measured availability target, and with a multi-location, active-active design you can reach 99.999% (five nines), about 5 minutes 15 seconds of downtime a year. That takes redundancy at every layer, automatic health-based failover across regions and ideally providers, safe deployments, failure testing, and outside-in monitoring. Control Plane is built for that tier.

Beyond Traditional Observability: Turning Technical Insight into Operational Intelligence

Observability has become a central part of modern IT operations and for good reason. Metrics, logs and traces give technical teams detailed evidence about how applications, infrastructure and services are behaving. Such evidence helps them investigate performance degradation, identify abnormal behavior and understand what changed around the time an issue occurred.

Azure in Bleemeo: your subscription next to your servers, with one read-only role

Most teams that run on Azure do not run only on Azure. There is a database on a VM nobody wants to move, a Kubernetes cluster somewhere else, a few servers in a rack, and a monitoring setup that grew around all of it. Azure Monitor sees the Azure part very well and nothing else, so the picture of an incident ends up split across two consoles, two alerting configurations and two sets of dashboards. Bleemeo now connects to Azure the same way it already connects to AWS.

Why We Built the Komodor Agentic Operations Platform: Q&A with CEO Ben Ofiri

Komodor spent years building an AI SRE platform before the category had a name. With the launch of the Komodor Agentic Operations Platform, it’s opening that engine up so enterprises can build, run, govern and optimize their own agents in production. Following the launch, co-founder and CEO Ben Ofiri sat down to talk about why now is the right time for agentic operations, what breaks between prototype and production, and where operations will head next.

Observability vs Monitoring: Why Does IT Still Find Out After the Business Does?

✓ operational truth IT finds out late because traditional monitoring is built to detect what goes wrong, not what has quietly stopped happening. Closing that gap requires observability that validates business journeys end to end, detects missing activity, checks its own coverage, and predicts degradation before a threshold is ever crossed.

The Data Race That Wasn't a Bug (and the One That Was)

Imagine this: you are testing the performance of some part of your application. Everything is going smoothly, the numbers look good, and as a last check you turn on Go’s race detector. Then, out of nowhere, it prints a warning you didn’t expect: So you look at it. You look at it again, and again, and you think: “What the…?” The race is between your code and a goroutine you never started, somewhere deep inside net/http. You have no idea how that is possible, or why.

AI Agent Context Explained: What Agents Can't See in Your Infrastructure

The "C" word is a controversial subject in the US, but we have to talk about "context", and what it means to an AI Agent. For starters, agents can only act on what's in their context window. Everything outside of it is a guess. In application code that limit is usually an annoyance.