Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on Cloud monitoring, security and related technologies.

Kubernetes GPU Scheduling for MLOps and GPU Sharing

The default Kubernetes scheduler was built for stateless services: web servers, APIs, databases. It schedules a pod, checks that a node has enough of whatever resources were requested, and binds it. For CPU and memory, that model works fine. For GPUs, it falls apart in three specific ways. First, GPUs are treated as an opaque integer resource.

Monitor your Amazon Bedrock workloads with Applications Manager

Organizations are increasingly integrating GenAI capabilities into their applications to deliver richer, more contextual user experiences—from AI-powered customer support and enterprise search to content generation, virtual assistants, and automated workflows. To build and scale these GenAI-powered experiences, they are turning to platforms such as Amazon Bedrock, which provides access to foundation models that developers can integrate into their applications.

Shipped: Catch a cost spike before it hits your bill

You’re probably already tracking the metrics that matter most in your Analytics dashboards like unit economics, AI ROI, and spend by team. Now you can put a target on any of them. Pick the metric, set the threshold, and CloudZero emails you when it’s crossed, with no ticket to us, no custom build.

How to measure AI ROI: metrics and a framework finance can actually run

To measure AI ROI, compare attributable value (revenue lift, cost savings, engineering time recovered, risk reduction) against fully loaded AI spend (API usage, subscriptions, infrastructure, people time) at the unit level: per initiative, per team, per task. The formula is simple. The instrumentation is the hard part, and it's where most organizations are failing: in CloudZero's 2026 survey, 34% of finance leaders couldn't produce a credible ROI number at all.
Sponsored Post

Building a Modern Cloud Outage Response Workflow in Slack and Microsoft Teams

On May 7 and 8, 2026, a thermal event in a single AWS data center hall knocked out power to EC2 instances and EBS volumes in a single Availability Zone in us-east-1. Within hours, more than 150 cloud services went down, including Coinbase, Reddit, HubSpot, and Atlassian's suite of tools, Jira, Confluence, and Trello among them. For teams without a structured cloud outage response workflow, the next several hours looked familiar: Slack DMs asking "is it down for you too?", tab-switching between status pages, and incident commanders repeating the same update in three different channels.