Operations | Monitoring | ITSM | DevOps | Cloud

Reliability Engineering in the AI Era

Engineering leaders have been claiming to “shift quality left” for years but production remains stubbornly stuck out of reach of software engineers. The realm of production remains mysterious with tools no one has access to and UIs that wouldn’t make sense to engineers anyway. I’ve noticed a small but growing trend of large enterprises hiring Reliability Engineers instead of Site Reliability Engineers. Dropping one word looks cosmetic but I think it points to a much bigger change.

From Telemetry to Traffic

A metric says latency increased. A log says a request failed. A trace identifies the slow dependency. An APM agent points to the method. Manual instrumentation explains the business operation. Traffic capture shows the exact request and response that triggered it. Each layer answers a question the previous layer could not. Each also introduces a new cost, blind spot, and failure mode.

Why AI agents don't need infrastructure running 24/7

AI agents don't need a server sitting there while they wait for the next prompt. They need infrastructure that shows up, does the job, and disappears. In this Product Highlights conversation, Nicolas Gommenginger, director of product at Upsun with more than four years leading the platform's core functionality, breaks down Upsun's newest primitive: task containers. His take: "What really triggered us to go ahead and implement it is the rise of AI agents, because tasks are something ideal for running them." We get into.

Stop paying for compute that does nothing six days a week

Most teams are paying for a container that stays awake all week to run one job on Friday. Task containers are the fix: they spin up on an API call, run exactly one command, then remove themselves. In this Product Highlights conversation, Corey Dockendorf, a senior solutions architect at Upsun with most of 25 years in tech behind him and five at the company, breaks down how ephemeral task containers cut idle spend and give AI agents production context without production access. His take: "They spin up when they're triggered directly by an API call.

Your preview environment is lying to you without production data

A preview link with an empty database tells you almost nothing about a content-managed site. Most teams find that out the expensive way, in front of a client. In this Product Highlights conversation, Andrew Kester, a senior cloud support engineer at Upsun whose team handles everything from product questions to sites under active attack, twenty four hours a day, every day of the year, breaks down why cloning services and data into every environment is the feature he would have paid for at his old agency job. His take: "You click a button. We clone your services. We clone the data." We get into.

On Release Days We Wear Teal for release 4.19.1

In this episode, Leon explores some of the new features, functions, updates, and improvements in release 4.19.1, which includes private links, setting a region per workspace, choosing a release channel per workspace, and a look at Output Routers, an incredibly useful (but unfortunately under-loved) feature that’s been around for a hot minute.

How to Integrate Multiple Help Desks Into One ITSM Solution

Managing separate help desks for different departments can make support harder to keep consistent. Each team may have its own processes, tools, and data, while employees have to figure out where to go for help. InvGate Service Management brings those teams together in one platform. IT, HR, Finance, Maintenance, and other departments can manage their services from the same place, with shared processes, centralized data, and a single point of contact for employees.

Shipped: A changelog that keeps up with how fast we ship

When the changelog doesn’t keep pace with the product, two things can happen. One, you keep working around something that was already fixed weeks ago. Or two, a behavior changes, you assume it’s a bug, and you spend an afternoon on triage and a support ticket before learning it was an intentional improvement. CloudZero now ships around 30 improvements a week, a pace driven by the Next Gen Platform and the AI-first approach we’re building for our customers.