Operations | Monitoring | ITSM | DevOps | Cloud

Building Investigations: what it takes to build an AI SRE

At incident.io, we've spent the last two years building Investigations, our AI SRE. When you get paged, it starts investigating straight away, looking across your telemetry, recent deploys, past incidents, docs and code, and posts what it's found in your incident channel (or on your phone, if it's 2am and you're still deciding whether you need to get out of bed). By the time you open your laptop, you're starting at step six of triage rather than step one.

Monitor warehouse data quality beyond pipeline health

You get a Slack message from the VP of Sales: They have asked an AI agent connected to Snowflake for the past quarter’s revenue and the numbers look wrong. First, you verify the agent’s query and, when that looks fine, check the pipelines that populate the underlying table. All jobs completed, the data is recently refreshed. Then it’s time to check the logs for errors. Nothing.

Android belongs in your CI/CD pipeline

How on-demand Android environments turn validation into a repeatable pipeline stage In the first article, we looked at automation: how Android environments can be created and managed programmatically. In the second, we looked at scaling: how shared infrastructure can make those environments available to more developers, tests, and workloads. This third article looks at the next step: integrating those environments directly into CI/CD.

Dr. Cat Hicks on the Psychology of Software Teams

What happens when a software engineer who has put their entire identity into being the resident expert of an obscure technology or language with ten years of experience and who knows the codebase like the back of their hand now has to compete with AI? Psychologically speaking, according to Dr. Cat Hicks, author of The Psychology of Software Teams and founder of Catharsis, that's called identity threat, and it's a surefire way to feel unsafe. We were honored Dr.

OnlineOrNot updates from June through August 2026

It's been a while since I last updated you on what's new in OnlineOrNot. For the most part, I've continued making it useful for software teams: there's now an MCP server, I've improved the terraform provider significantly, and there's a new TypeScript SDK. You can also (finally?) monitor DNS/TCP with it.

Megaport Advanced Services: From Network Design to Deployment

Network changes often stall after the design is done. See how Megaport Advanced Services helps teams move from planning to implementation and support. Ask a network team where their last big network change got stuck and you’ll rarely hear “we couldn’t work out the design.” The completed design is usually sitting in a document somewhere, reviewed and signed off. What stalls is everything needed after that. Somebody has to write the runbook. Somebody has to sit in the 2 a.m.

Best Monitoring Tools With MCP Servers for AI Coding Agents (2026)

Your AI coding agent can read your code, run your tests, and open a pull request. Until recently it could not see what that code does in production. MCP servers from monitoring vendors close that gap. Connect one, and Claude Code, Cursor, or Copilot can pull the error, the trace, and the slow query behind a bug report without you copying anything out of a dashboard. Most major monitoring vendors now offer an official MCP server. But they aren’t interchangeable. Some only expose errors.