Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on DevOps, CI/CD, Automation and related technologies.

Learn these 4 Chaos Engineering Principles Before You Break Anything | Resilience Testing | Harness

Want to start chaos engineering? Don't randomly break stuff and hope for the best. Real chaos engineering starts with defining your system's steady state metrics like latency, throughput, and error rates. Then you form a clear hypothesis about what should happen when failures occur. Next, you inject controlled failures, starting small with single pod kills or network drops, not production meltdowns. Finally, you limit the blast radius by running experiments in safe environments first.

Harness Lives Inside Cursor Now - Plus Everything Else That Shipped in April

April was a big month at Harness. AI is changing how code gets written — and the rest of the SDLC is catching up. In this update, Dewan Ahmed walks through Harness product releases across three themes: AI in the developer workflow, security and governance for AI assets, and self-service maturity for developers and platform teams. What's covered (with timestamps): Found this useful? Subscribe for monthly product updates, and drop a comment telling us which release you want a deep dive on next.

What is an ASN? Understanding the backbone of the Internet

Using the internet often feels effortless when clicking a link or joining a call, but behind that simplicity lies a highly structured system that ensures data moves efficiently across the globe. One of the key building blocks of this system is the Autonomous System Number (ASN).

A Guide to 400G Connectivity

Ready to scale beyond 100G? Learn why 400G is on the rise, when to use it, and how to deploy it. Network traffic is growing exponentially. Cloud adoption, AI, large-scale data replication, video streaming, and generative applications are all drivers, and enterprises with traditional connectivity setups may find themselves struggling to keep up. Enter 400-gigabit Ethernet (400G): a high-capacity, scalable networking standard that enables you to build faster and more cost-efficient networks at scale.

What is alert fatigue? (And how does it happen)

Alert fatigue doesn’t announce itself. It builds quietly over weeks and months until one day a critical incident triggers and nobody responds with the urgency it deserves. By that point, the damage is already done. This guide walks through what alert fatigue actually is, how it happens, and what you can do about it.

AI in Software Delivery: Engineering Excellence or Just Market Hype? | Harness Blog

AWS re:Invent 2025 made one thing very clear: enterprise interest in AI is no longer theoretical. The conversation has moved beyond curiosity. Teams are actively experimenting, leaders are looking for production-ready use cases, and engineering organizations are trying to figure out where AI can create real leverage across software delivery, security, platform engineering, and operations.

The most debated DORA metric (even Google debates this)

What's the most debated DORA metric? Nathen H from Google's DORA team breaks down the change lead time debate — and why even the experts can't fully agree on when a change is "committed." Is it at commit? After merge? The answer matters more than you think. Subscribe for more DevEx and DORA insights from our Web Summit series.

AI Supply Chain Attacks Are Here. And Most Organizations Aren't Ready

When I read about the Vercel breach tied to a Context AI compromise, I wasn’t surprised. I’ve been talking with customers for a while now about how AI was going to introduce a new kind of supply chain risk. This is exactly what that looks like. What stands out to me is how familiar the pattern is. We saw it with open source, then again with SaaS, and again with cloud.

AI Enablement for Dev Teams: The 6-Pillar Flywheel

AI adoption is already happening on your team, whether you have a strategy or not. Tracy Lee (CEO of This Dot Labs, Microsoft MVP, Google Developer Expert) breaks down the AI Enablement Flywheel — a 6-pillar framework used by successful engineering organizations to move from scattered experimentation to scalable, ROI-positive AI workflows.