Operations | Monitoring | ITSM | DevOps | Cloud

AI in IT Operations: How to Build Trust, Automate Smarter & Prepare for Autonomous IT

What does it take to make AI and automation actually work in enterprise IT? In this episode of Agents of IT, Resolve’s Zach Austin sits down with Nick Dimmock, Co-Founder and CEO of TechWorks, to discuss how IT automation is evolving, why trusted data matters, and what organizations need to build before AI can deliver meaningful business outcomes.

How to Use Megaport Storage as a Veeam Backup Target

Learn how to use Megaport Storage as an S3-compatible Veeam backup target for scalable, private, offsite backup storage. Table of Contents A backup is only useful if it can be retrieved when needed. Where backups are stored determines how well they’re protected from incidents at the primary site, how long recovery transfers take, and what it costs to bring the data back.

Every 404 in Your Rails App Might Be Allocating 13 MB

A bot requests /wp-login.php on your Rails app. Rails can’t route it, raises ActionController::RoutingError, and returns a 404. That should cost almost nothing. On a Rails 8.1 app with a few thousand compiled templates, it can cost 13 MB of allocations and 27 ms of CPU. A detailed report on rails/rails#58887 traces the cost to one method, ActionDispatch::ExceptionWrapper#build_backtrace.

How to Cut Cloud Compute Costs Without Rewriting Your Apps

The fastest way to cut cloud compute costs is to stop paying for capacity your workloads do not use. Right-size CPU and memory to real usage, scale idle workloads to zero, and make cost policy a platform default instead of a quarterly review. Control Plane does all three at the platform level: Capacity AI right-sizes running workloads, autoscaling scales idle ones to zero, and customers typically spend 30 to 50 percent less on compute than running directly on AWS, GCP, or Azure.

Tempo 3.1 release: new features for Kafka, TraceQL metrics updates, trace redaction, and more

Building on the major release of Tempo 3.0, Tempo 3.1 is here, delivering community-contributed Kafka client improvements, query-based trace redaction, sampling-aware TraceQL metrics, and more. Together, the updates in 3.1 make it easier to operate Tempo, get accurate insights from your trace data, and investigate issues more efficiently. You can continue reading and check out the video below to learn more about the latest features.

Telemetry Talks ep 8 - Fireside chat with OpenTelemetry maintainers

Telemetry Talks episode 8 is here We sat down with OTel maintainers to talk about the future of the community, GenAI semantic conventions, contributing beyond code, OTel in Practice, and what they’re currently building, writing, organizing, and experimenting with across the CloudNative and OpenSource ecosystem. A conversation about where OTel is today and what comes next. Playlist Resources for Further Learning.

MTTR Is Not a Time Problem. It Is a Context Problem

Your Mean Time to Resolution (MTTR) has likely stayed flat for three or four quarters. The investment was real: scheduling tools, dispatch optimization, new training modules, and more technicians. Operations reviews still dissect response time, travel time, and wrench time. The metric still refuses to move. Most field service leaders measure MTTR from the start of the repair to the moment the asset returns to service.

AI cost allocation: how to attribute AI spend by team, product, and customer

AI cost allocation is the practice of attributing every dollar of AI spend to the team, product, feature, or customer that generated it. That spend includes API tokens, GPU compute, per-seat tools, and shared infrastructure. It's harder than cloud allocation because AI spend arrives untagged, spans vendors, and pools in shared resources. Four methods cover most cases: tag-based, key-based attribution, proportional split, and usage-telemetry.