Operations | Monitoring | ITSM | DevOps | Cloud

How to Cut Cloud Compute Costs Without Rewriting Your Apps

The fastest way to cut cloud compute costs is to stop paying for capacity your workloads do not use. Right-size CPU and memory to real usage, scale idle workloads to zero, and make cost policy a platform default instead of a quarterly review. Control Plane does all three at the platform level: Capacity AI right-sizes running workloads, autoscaling scales idle ones to zero, and customers typically spend 30 to 50 percent less on compute than running directly on AWS, GCP, or Azure.

Tempo 3.1 release: new features for Kafka, TraceQL metrics updates, trace redaction, and more

Building on the major release of Tempo 3.0, Tempo 3.1 is here, delivering community-contributed Kafka client improvements, query-based trace redaction, sampling-aware TraceQL metrics, and more. Together, the updates in 3.1 make it easier to operate Tempo, get accurate insights from your trace data, and investigate issues more efficiently. You can continue reading and check out the video below to learn more about the latest features.

Telemetry Talks ep 8 - Fireside chat with OpenTelemetry maintainers

Telemetry Talks episode 8 is here We sat down with OTel maintainers to talk about the future of the community, GenAI semantic conventions, contributing beyond code, OTel in Practice, and what they’re currently building, writing, organizing, and experimenting with across the CloudNative and OpenSource ecosystem. A conversation about where OTel is today and what comes next. Playlist Resources for Further Learning.

MTTR Is Not a Time Problem. It Is a Context Problem

Your Mean Time to Resolution (MTTR) has likely stayed flat for three or four quarters. The investment was real: scheduling tools, dispatch optimization, new training modules, and more technicians. Operations reviews still dissect response time, travel time, and wrench time. The metric still refuses to move. Most field service leaders measure MTTR from the start of the repair to the moment the asset returns to service.

AI cost allocation: how to attribute AI spend by team, product, and customer

AI cost allocation is the practice of attributing every dollar of AI spend to the team, product, feature, or customer that generated it. That spend includes API tokens, GPU compute, per-seat tools, and shared infrastructure. It's harder than cloud allocation because AI spend arrives untagged, spans vendors, and pools in shared resources. Four methods cover most cases: tag-based, key-based attribution, proportional split, and usage-telemetry.

Shipped: One CloudZero for everyone, starting October 1

On June 3, we made the new CloudZero experience the default for every customer. Since then, we’ve shipped around 30 improvements a week: side-by-side period comparisons in Explorer, budgets you can create and edit right in the app, threshold alerts on dashboard tiles, and Monitors, which flags AI and cloud spend that moves outside its normal pattern and shows you what changed. Pages load 28 to 61% faster. JavaScript execution is 85% faster.

Auto-Generate Richer Azure Architecture Diagrams

Azure architecture diagrams go out of date fast, and management-plane data alone misses the runtime connections that matter most. In v5.4, diagrams move to their own Diagrams tab in Azure Documenter. Alongside the enhanced network and workload diagrams, there is a new Resource Visualizer diagram. Scope it by subscription and resource group, or write your own custom Azure Resource Graph query to define exactly which resources to include.

Change the Cloud Cost Conversation from Spend to Margin

"Our Azure bill went up 20% last month." Without context, finance only sees a rising cost. Unit economics gives them the full picture. Turbo360 lets you overlay business KPIs on your Azure spend. Track units like orders, active users, document views, or monthly recurring revenue alongside cost, and see your cost per unit month over month. Now the conversation becomes: "Orders went up 150% and our cost per order came down." That is a story about efficiency, not overspend.

IT Service Continuity Management: How to Build an ITSCM Plan

Most IT teams have a recovery plan somewhere. It was written for a disruption that has not happened yet, and tested less often than anyone admits. The gap rarely sits in the technology. Nobody agreed which services come back first, or how fast. There was time to settle that calmly, and it went unused. IT service continuity management is the ITIL practice that settles those questions in advance. In this blog, you will: By the end you will know what belongs in an ITSCM plan and who has to agree to it.