Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on DevOps, CI/CD, Automation and related technologies.

LLM cost optimization: 7 strategies to cut inference spend

LLM cost optimization is the practice of cutting what you spend on large language models, mostly inference, without losing the quality that makes the AI worth running. The biggest levers are routing requests to cheaper models, caching repeated tokens, batching anything that can wait, trimming prompts, right-sizing models, cutting calls you do not need, and putting one gateway and cost view in front of all of it.

Kepler Is in Public Preview: One Task, Every Repo, Every Agent

A faster car doesn’t get you home faster if the freeway is still jammed. That is the problem most teams run into once they add a second, third, or fourth AI coding agent to the mix. More agents generate more code. They do not automatically generate more finished work, because someone still has to track which agent is waiting on input, which one just opened a pull request, and which one has been quietly stuck for twenty minutes. Kepler is GitKraken’s answer to that traffic jam.

Ubuntu 26.04 LTS Resolute Raccoon is Now Supported!

We are excited to announce support for Ubuntu 26.04 LTS Resolute Raccoon on all Cloud 66 products, including registered servers. From this point onward, new applications are provisioned on Ubuntu 26.04 by default, on both x86_64 and ARM64. Don't forget, you can control your target Ubuntu version via the selection dropdown when scaling up via the UI, or through your manifest!

The 2026 pocket guide to engineering metrics

Most engineering leaders are drowning in data but starved for insight. We have dashboards full of metrics, but they often create more questions than answers and rarely tell us what to do next. In the age of AI, where development velocity is accelerating at an unprecedented rate, this problem is only getting worse. Shipping code faster than you can fix it is an existential risk, and a dashboard that doesn't lead to action is just a distraction.

Netdata AI & MCP: Root cause in one investigation

An alert tells you what changed. It rarely tells you why. That answer usually lives somewhere else: the pull request that shipped minutes earlier, the incident already open in PagerDuty, the runbook sitting in Confluence. Netdata AI can now reach those systems directly. Through the Model Context Protocol, it connects to the tools your team already runs and reads from them while it investigates. It pulls that context alongside the metrics, logs and anomalies Netdata detects on every node, then reasons across both.

Your platform is your business, encoded onto your infra, with Syntasso's Abby Bangser

Cortex co-founder and CTO Ganesh Datta sits down with Abby Bangser, a platform engineering leader at Syntasso and former lead of the CNCF Platforms Working Group, to talk about why AI agents need real platform APIs, not raw cloud credentials.

4 Cloud-Native Challenges AI SRE Is Solving in 2026 and the 3 New Ones to Look Out For

AI SRE is making real strides in resolving some of the greatest pains related to incident response, troubleshooting, and complex root cause analysis. The on-call rotation, the war room, the week-long RCA, and the ticket queue that ate a third of every platform engineer’s week all look different now than they did two years ago.

Shipped: Compare Periods, side-by-side cost comparison in Explorer

When spend moves, the first question is always “compared to what?” Answering it used to mean pulling up two tabs with two different date ranges and toggling between them, or squinting at a single delta number without the ability to customize what you’re comparing against. Compare Periods puts both periods in front of you at once, so a spike, a regression, or a shift shows up on the chart and in the cost table, without you doing the math yourself.