Operations | Monitoring | ITSM | DevOps | Cloud

Platform engineering in the age of AI

94% of engineering leaders say their AI metrics are missing. Here's how platform engineering is changing to close that gap. Based on the InfoQ webinar "Platform Engineering in the Age of AI," featuring panelists from Harness, DKB, and Shine, August 18, 2026. 94% of engineering leaders say the AI metrics that matter most to them are missing.

How to build a Language Server Protocol (LSP) plugin for Claude Code

Language servers give editors structured, real-time feedback: diagnostics, hover docs, autocomplete, and other guidance that would otherwise surface later. Language servers already exist for many of the languages and tools developers use every day, but Claude Code doesn’t automatically receive their feedback.

Log Analysis with Machine Learning: An Automated Approach to Analyzing Logs Using ML/AI

AI log analysis helps IT teams turn massive volumes of operational data into actionable insight. By applying statistical methods, machine learning (ML), semantic analysis, and generative AI, organizations can identify unusual behavior, connect related signals, and investigate probable root causes faster. But AI-generated answers should not be mistaken for proof.

Best AI Infrastructure Providers for Power, Cooling, and Compute

AI infrastructure is becoming a facilities problem as much as a compute problem. Adding accelerators is only useful when the surrounding environment can support them. Power has to reach the rack reliably. Cooling has to remove the heat produced under sustained load. The network fabric has to keep accelerators communicating. Storage has to feed the workload. Orchestration and monitoring then determine whether expensive capacity spends its time doing useful work.

HITL for autonomous agents: Where does the human go?

Human approval is easy when you are sitting in front of the agent. For an agent running by itself in a cluster, almost none of that holds. You’re in a meeting and your agent is running in a cluster. It has a service account, it has been asked to keep a service healthy, and it has just worked out that the right fix is to roll back a database migration. Nobody is watching it. That was rather the point of deploying it. You want to get notified to approve such an important action.

Day 2 Operations for AI-Generated Code: What Changes When You Didn't Write It

We have all watched AI speed up the way we write software. With tools like Copilot and ChatGPT, developers can spin up boilerplate, write complex functions, or draft entire micro services in minutes instead of days. It feels like magic. But there is a silent catch that we do not talk about enough: writing the code is only Day 1. The real challenge is Day 2 operations, which is everything that happens after that code is deployed.

Why AI Creative Workflows Are Moving Beyond Single-Purpose Generators

AI generation is no longer the difficult part of creating digital content. Generating an image from a prompt can take seconds. Creating a short AI video is also becoming increasingly accessible. Editing a background, modifying an object, or producing another visual variation can often be handled with a few instructions. The harder problem appears after the first generation.

How to operate shared platforms safely at agent scale

A platform engineering team can design robust Golden Paths for agent use yet still be unprepared for what happens after adoption. An agent may authenticate properly, call the correct tools, adhere to approval gates, and complete tasks without incident, but new operational risks arise once multiple teams begin running agents continuously and in parallel. We’ve encountered these risks firsthand at Datadog.

Monitor TAS and gang scheduling for AI training in Kubernetes

Distributed AI training workloads impose complex scheduling requirements that Kubernetes’s built-in scheduler can’t meet. Kubernetes schedules pods individually and independently, but distributed training introduces two requirements that break this model: Pods must land on hardware with the right inter-GPU bandwidth, and all pods must be scheduled simultaneously. If either requirement goes unmet, training stalls or runs far below the hardware’s potential.

Manage Cursor costs with Datadog Cloud Cost Management

AI coding tools such as Cursor are becoming a significant source of engineering spend. But Cursor costs can be difficult for FinOps teams to manage. Cursor’s usage data alone doesn’t tell you how costs break down across users and models, and fixed-threshold alerts may not catch an unusual cost spike if spend remains below the threshold. Datadog Cloud Cost Management (CCM) brings Cursor costs into the same place where you monitor cloud, SaaS, and other AI spend.

Bindplane Agent Is Here: Build, Edit, and Understand Pipelines in Plain Language

Pipeline Intelligence already recommends processors, reads live telemetry, detects log types, and generates processor bundles from natural language. But most of it lives inside a single processor node. You still have to know which one to open and what to ask for. This changes today. Bindplane Agent is your AI assistant inside Bindplane, ready to act on what you describe in plain language.

150+ AI statistics for 2026: spend, cost, and AI ROI

Worldwide AI spending will reach $2.59 trillion in 2026, up 47% from 2025, according to Gartner. Yet only 37% of organizations report any earnings impact from AI, McKinsey finds. These AI statistics cover what companies spend, what AI costs to run, and the ROI they're actually getting. That gap between the two headline numbers is the story of AI in 2026.

From a $60K invoice to a $200B earnings call, few can explain the AI bill

CloudZero’s own AI Economics Pulse for September found the 75th percentile of its 430-company customer panel crossed 10% of its cloud bill on AI for the first time in August. Gartner’s latest survey found only 22% of organizations have scaled AI successfully and 11% don’t know what their own function spent on it last year. CJ Gustafson showed what that gap looks like on an actual invoice this week.

The 12 Point Checklist Before Your AI Voice Agent Takes Live Calls

You have tested your AI voicebot solutions for weeks. The calls sound natural, the answers are accurate, and the demo works exactly as planned. Then you put it on a live call. A caller interrupts while the agent is speaking. Speech recognition misses an account number. An API takes three seconds to respond. The LLM slows down under load. One service fails, and suddenly your polished voice agent has no idea what to do next.

Are AI agents about to break cloud computing?

AI agents can write code and run tools on their own now. The problem is they still need somewhere to actually do it. exe.dev's David Crawshaw joins Michael Reid to break down why persistent virtual machines might become the backbone of agentic AI, and what happens to cloud economics when one person is running dozens of agents at once.

The AI Economy Has a Senior Engineer Problem. Here's How to Solve It

According to a 2025 report from Ravio, entry-level hiring (especially in engineering roles) has collapsed by more than 73% due to increasing AI capabilities. That means junior developer jobs are disappearing. At the same time, demand for senior engineers keeps climbing because organizations need more people to manage and optimize their complex AI agents.

How Canvas Powers the AI Agent Development Feedback Loop

For teams building AI agents, the feedback loop should already be a familiar idea: watch how the agent behaves, find what needs improvement, ship a change, and measure the result. In theory, each turn builds on the last until the loop becomes a flywheel and your agent is getting more effective with each turn. In practice, many of us are still in reaction mode. A user reports something strange, costs spike, or an eval score drops.

First Look: Build Grafana Dashboards with AI using the MetricFire MCP Server

Get a first look at what’s coming next to the MetricFire MCP Server: AI-powered dashboard creation and management. We’re connecting the Hosted Graphite HTTP Dashboard API to our MCP Server, letting compatible AI clients work with your monitoring data and Grafana dashboards directly through an AI-assisted workflow. Soon, you’ll be able to use natural language prompts to reference metrics stored in Hosted Graphite and create, update, and manage dashboards.

ilert AI SRE is generally available

When you get paged at 3am, it takes about 30 seconds for the notification to reach you and maybe two minutes until you're in front of a laptop, awake enough to read. What you see then is usually a raw alert. A metric name, a threshold, a link to a dashboard. Then the ritual starts: open the dashboard, check what deployed in the last few hours, grep the logs for the first error, ask in Slack whether anyone touched the database. ‍ Most of that time is search.

Datadog named the Company to Beat for observability platforms in 2026 Gartner AI Vendor Race report

Datadog has been named the Company to Beat for observability platforms in the August 2026 Gartner AI Vendor Race research. Datadog has also been named a Leader in the 2026 Gartner Magic Quadrant for Observability Platforms for the sixth consecutive year. We believe that these recognitions reflect what we have been building toward for more than a decade: a single platform where teams can observe, secure, and act on everything that matters across their technology stack.

Shipped: Know what you actually pay per token on OpenAI

Picture two teams running the same million input tokens through the same model. One team’s tokens are cache hits, queued through the batch API. The other team’s are fresh, sent live. On a current-generation OpenAI model, cached input runs about a tenth the price of a fresh token, and batch processing cuts whatever’s left in half. Stack the two: at a list rate of $2 per million tokens, one team’s bill comes to 10 cents, the other’s to two dollars.

How is the SaaS Market Performing in 2026?

Software as a Service (SaaS) dominates corporate technology purchases in 2026. The cloud powers communications, finance, customer management, analytics, cybersecurity, HR, and more. Software spending is soaring, but purchasers choose value-based subscriptions. Business conditions affect SaaS enterprises. Higher costs, shifting technology spending, mergers, and corporate shutdowns limit demand but also create opportunities. A company shutdown 2026 report might cover software spend, customer retention, and investment trends. Gartner anticipated $1.47 trillion in software spending in July 2026, up 15.5% from July 2025.

Top 5 AI Gateways for Enterprise (2026 Guide)

Enterprise AI infrastructure has become considerably more complicated than connecting an application to a single large language model. Production systems increasingly use several model providers, while AI agents may also communicate with tools, MCP servers and other agents. Every additional connection introduces questions around security, reliability, cost, access control and observability.

Can Technology Become Smarter Without Becoming More Complicated?

Our phone runs a neural network to pick the sharpest frame every time we take a photo, yet the whole interaction is one tap. That gap is the story of modern technology: the machinery keeps getting denser while the thing in your hand keeps getting quieter. We tend to assume intelligence and complication rise together, but the most important products of the last decade prove the opposite can happen. The real question is not whether tools can get smarter without getting harder to use, but where all that hidden difficulty actually goes.

Anthropic's Fable 5 Falls Short on Enterprise Adoption

Two months ago, Claude Fable 5 launched with the kind of fanfare usually reserved for a flagship phone. Anthropic called it the most capable model on the market. The U.S. Department of Commerce briefly suspended access over export-control concerns, which only added to the buzz when it returned three weeks later. Then the spending data came in, and it told a much quieter story. According to Ramp's August AI Index, which tracked spending across 70,000 businesses, Fable 5 accounts for just 11.4% of what companies spend on Anthropic's models and only 6% of the tokens they actually use.

The Synthetic Claims Crisis: How Generative AI Is Reshaping Insurance Fraud and Visual Verification

The insurance industry is currently facing a fundamental shift driven by rapid advancements in generative technology. Automated intake pipelines have significantly sped up processing, yet they have simultaneously exposed insurance companies to an entirely new category of financial risk. Today's generative networks and image editing tools allow anyone to craft hyper-realistic visual evidence in moments. Bad actors no longer rely solely on physical staging; they can fabricate fake vehicular accidents or digitally amplify minor home damage with incredible accuracy.

Why engineers ignore cloud costs, and how AI Cost Management Agents fix it

Engineers ignore cloud costs because of broken feedback loops, not apathy. Learn what AI cost management is, why AEO matters more than ever, and how a cost management agent embeds accountability directly into engineering workflows. Engineers ignore cloud costs because cost data arrives too late and too disconnected from their workflow to act on.

Best LLM gateways in 2026: 30+ AI gateways compared on cost control

An LLM gateway is a proxy that sits between your applications and model providers, handling routing, failover, caching, and cost controls through one API. The strongest picks in 2026: LiteLLM for self-hosted control, OpenRouter for instant multi-model access, Portkey for managed governance, and Bifrost for production-scale throughput. Enterprises spent $37 billion on generative AI in 2025, a 3.2x jump in one year, per Menlo Ventures.

Analyze your experiments in ChatGPT with the Datadog Experiments plugin

ChatGPT Work has become a common starting point for data and product teams. Analysts open it to compare launch adoption across segments, diagnose a metric that moved overnight, or turn a week of scattered numbers into a readout that a leader can act on. But the moment teams ask whether their experiment actually caused an effect they’ve observed, the conversation stalls.

Debugging our AI search assistant with agent tracing

In order for users to get the most out of the data being sent to Sentry, it’s important that we make it easy to find that data. Our team works on features to help users browse their data to find a particular event using search queries and filters. The search bar enables users to find their data by specifying search terms. Searching uses the Sentry Search Syntax, which can be barrier for users.

How AI and Digital Transformation Are Changing the Way Consumers Shop for Eyewear Online

The way people buy eyewear has changed significantly in recent years. A process that once required visiting multiple optical stores can now happen from a smartphone or computer. Digital platforms have made it easier for consumers to browse styles, compare options, and make informed purchasing decisions without leaving home.

A simpler way to run AI agents in Bitbucket Pipelines

AI agents can help investigate failed builds, fix flaky tests and automate other development tasks. But setting up those agents has required more Pipelines configuration than it should. Agent-powered steps often need different compute, permissions and runtime settings from ordinary build and test steps. Until now, teams have either repeated those settings across every agent-powered step or tried to make one set of global defaults work for everything.

Every Deployment Platform Is Pivoting to AI. Day 2 Operations Aren't Going Anywhere

Over the summer, Fly.io founder Kurt Mackey announced a complete pivot for the company toward "Computers for Agents", which are ephemeral virtual machines (called Sprites) optimized for AI coding workflows. He was refreshingly explicit about what this means: they are not trying to do both traditional application hosting and AI agent compute. They are choosing one over the other. This is a completely rational bet on the future of developer tooling.

Megaport Collaborates With NVIDIA to Boost AI in Australia

Australia’s home-grown global automated infrastructure platform is part of a cohort of companies with Australian operations that will provide regional businesses and institutions access to NVIDIA accelerated computing and NVIDIA Nemotron open models. It’s a point of pride for all of us at Megaport that we’ve built a global automated infrastructure platform while maintaining our deep Australian roots.

How to right-size the handoff between two agents

model-right-sizer-schema is a Claude Code skill that designs the typed contract between one agent and the controller that dispatches it. Point it at an agent plus its controller and it returns a JSON prescription with typed in/out fields, an exclusion list that keeps raw logs out of the reply, a before/after size delta, then writes the contract into the agent's own file. It picks from nine portable output-shape families, or your repo's own.

How AI-Based Crop Counting and Health Analysis Boost Agricultural Yield

Modern farms need faster, more reliable ways to understand plant population, detect stress early, and act before losses spread, because manual scouting is time-consuming, resource-intensive, and highly dependent on human expertise. That challenge matters at the yield level, since pests and diseases can significantly reduce crop productivity, and early recognition is critical for protecting both output and quality. AI-based crop counting and health analysis address this problem by turning images, sensor data, and field observations into structured decisions that support more precise crop management.

What an "Agent Harness" Actually Is - and Why Raw Model Calls Don't Survive Production

There's a demo that convinces every engineering team that agents are ready: someone gives a model a goal, it calls a couple of tools, and it produces a result that would have taken a person an hour. The gap between that demo and a system real users depend on is enormous, and most of that gap is not the model. It's everything around the model - the layer that decides what to do next, calls tools safely, remembers what happened, asks for help when it should, and records the whole run so you can debug it. That layer has a name: the agent harness.

Nvidia Forecasts 70% Revenue Growth as AI Spending Heads Toward $1.3 Trillion

Nvidia has decided to break with its usual practice of providing guidance only for the upcoming quarter and, for the first time, has given investors an outlook for the next fiscal year. The company expects to increase revenue by approximately 70% to $673 billion, which would further cement its position in the Dow Jones index and make it the secondlargest U.S. technology company by revenue, after Amazon. This projection substantially surpasses market expectations based on growth of about 44% and is effectively an attempt to convince the market that the AI boom is far from over. Against this backdrop, Nvidia stock moved higher.

Monitor smarter with Applications Manager's GenAI capabilities

GenAI has moved well past the pilot stage. According to a Gartner finding, by 2026, more than 80% of enterprises will have used GenAI APIs or deployed GenAI-enabled applications in production. Today, GenAI is becoming an integral part of how infrastructure and application teams work every day. Organizations are depending on LLMs from a diverse range of vendors—OpenAI, Anthropic, Google AI, and DeepSeek—based on the strengths each offer for different use cases.

n8n pricing in 2026: every plan, the execution math, and what AI agents change

n8n pricing runs €24 per month for 2,500 workflow executions (Starter), €60 for 10,000 (Pro), and €800 for 40,000 (Business), with 17 percent off on annual billing and custom Enterprise pricing above that. Every plan includes unlimited users and unlimited workflows. The self-hosted Community Edition is free with unlimited executions; you pay only for your server.

CoreWeave pricing in 2026: every GPU rate and what a node really costs

CoreWeave, a GPU cloud provider, prices start at $6.16 per GPU hour for an Nvidia H100 and reaches $8.60 for a B200, sold as fixed multi-GPU nodes: an 8x H100 node lists at $49.24 per hour on demand. Spot rates run up to 60 percent below on demand, reserved contracts discount up to 60 percent, and egress is free.

How we built data-driven AI Golden Paths at Datadog

As teams rush to adopt AI, they often find themselves with conflicting workflows unique to each individual developer. To manage costs and promote good development practices, organizations need to establish Golden Paths around AI usage. AI Golden Paths are standardized flows that help developers work with agents more reliably and effectively. But how do you sift through all the possible workflows to decide what these Golden Paths should be?

AI Norms & Values, Part 3 of 3: Things We Hold True

Welcome to the third and final part of our series on AI norms and values. Parts of this doc were extracted and published separately on substack; as a whole, they describe the principles we hold pertaining to technology and AI, and the ethical commitments we make to each other and our customers. We set out to write about AI, and ended up writing about ourselves. These documents are not meant to be aspirational ones; they are derived from how we do our work every day in honeycomb.

On a Network, an Agent Acts Where the Blast Radius Is Largest

Every network engineer carries an instinct that outsiders mistake for caution: a change in one place can travel. Reroute a path, push a policy, drop an interface, and the effect can ripple across campus, data center, WAN, and cloud before the first alert is read. The blast radius of a network change is the reason operators move deliberately, and it is the single most important thing an AI agent takes on the moment it is allowed to act on the network instead of merely describe it.

AI Code Review Loop in the Terminal: Introducing Harness CLI for Harness Code

Every developer knows the fatigue of the "12-tab code review dance": Agents have become first class citizens in SDLC and AI coding agents author code alongside human engineers, thus the above context switching destroys flow state. GitHub's gh CLI proved developers love the terminal, but modern delivery is tied to AI reviews, pipeline executions, risk scoring, and autonomous agents, not just git hosting.

Questions to Ask About AI Agent Orchestration

Running AI coding agents in parallel across repositories is no longer experimental. It’s how high-performing engineering teams ship faster. But the tools you pick to orchestrate those agents can either multiply your output or introduce new bottlenecks. GitKraken gives your team a purpose-built surface for AI coding agent orchestration through Kepler, its agent-agnostic development environment. Before you commit to any orchestration tool, though, you need to ask the right questions.

McKinsey Says Agentic Enterprises Need "Automated Guardrails." Here's What That Means

TLDR/: McKinsey’s new research on AI transformation, published August 28, 2026, studied 20 companies that have created real economic value from AI and found that only a small number have reached “Stage 3: Agentic AI enterprise.” The capability that separates Stage 3 from Stage 2, per McKinsey’s own maturity framework, is orchestration layers and automated guardrails: the ability to govern agent actions automatically, in real time, rather than reviewing them after the fact.

Top Legal AI Tools for Reducing Manual Work Across the Personal Injury Case Lifecycle in 2026

Personal injury cases create a lot of work that has little to do with making legal decisions. Someone still has to review medical records, find details buried in case files, build chronologies, prepare demands, draft documents, organize evidence, and keep case information up to date. Legal AI can take some of that work off the team's plate. The most useful tools are not necessarily the ones with the most features. They are the ones that address the parts of a case where attorneys, paralegals, and case managers are spending hours on repetitive work.

The VM Boom For AI Agents | David Crawshaw Co-Founder & CEO, exe.dev

What happens when AI agents stop simply answering questions and start using computers of their own? It could create an entirely new boom in virtual machines. In this episode of Uplink, David Crawshaw, Co-Founder and CEO of exe.dev, joins host Michael Reid to explore the infrastructure behind the rapidly emerging world of AI agents. As agents become capable of writing code, running applications, operating tools, maintaining state, and working independently, they need more than access to an AI model. They need computing environments where they can actually get work done.

From idea to working software: what the full development lifecycle needs to look like

GitHub's research found that developers using Copilot completed tasks 55% faster than those who didn't. Tools like GitHub Copilot and Cursor, powered by large language models such as Claude or GPT, are designed to automate the tedious parts of programming so engineers can focus on harder, more creative problems. With this. new repos spin up every week. The promise is being kept. But where are the products?

Your AI Economics Pulse for September 2026

Across a same-store panel of 430 CloudZero customer organizations, AI reached 2.66% of the median company's cloud bill in August 2026, up from 2.61% in July and roughly four times its level a year ago. The 75th percentile crossed 11%. The share of organizations with at least 10% of cloud spend attributed to AI jumped to 28.2% from 23.9%, the largest one-month move that tier has posted. Two-thirds of the panel now spends at least $1,000 a month on AI. The typical bill barely shifted.

AI Agents on Kubernetes 101: From Laptop Script to Production Pod

In short, this is a beginner’s guide to deploying an AI agent on Kubernetes. You will containerize an agent, store its API key as a Kubernetes secret, write a deployment with health probes and resource limits, expose it with a service, and lock down its network egress, in that order, with a working manifest at every step. On a local kind cluster the whole walkthrough takes about an hour.

How to Build an HR PTO AI Agent with Resolve Agent Lab

See how to build an HR PTO agent with Resolve Agent Lab. In this Resolve Reels demo, we create a purpose-built AI agent by adding automation skills, instructions, conversation starters, and guardrails. The agent can answer PTO questions, check balances, account for calendar conflicts, and submit requests through systems like Workday or ADP. See how Resolve helps teams build AI agents that take action across enterprise systems.

Agentic Operations Start with Context: Build the Right Data Foundation

Episode 1, "Beyond the Thread: Deconstructing the Cisco Data Fabric Powered by the Splunk Platform," explores the intersection of data strategy and operational efficiency. Hosted by Splunk's Courtney Wright, the session features insights from experts Keith McClellan and Michael Sondag on the complexities organizations face in data management and operational models.

Assisted, Augmented or Agentic? Choose Your Splunk Starting Point

Episode two of Beyond the Thread explores how organizations can leverage a solid data foundation for AI-driven actions. Hosted by Courtney Wright and featuring experts Greg Ainsley-Malik and Sonal Pardeshi, the discussion delves into the Cisco Data Fabric, powered by the Splunk platform, and its role in transforming machine data into actionable insights. The episode highlights the journey towards agentic operations, addressing the challenges faced in moving from AI-ready data to effective implementations, and examines different adoption strategies that organizations may pursue.

Introducing Infrastructure Knowledge: Teach Netdata AI What Your Metrics Can't Show

Netdata AI sees everything your infrastructure does: every metric, every anomaly, every alert. It does not see what your infrastructure is: which services matter, which host is supposed to run hot, who owns what, what your team considers normal. Without that context, “CPU at 91%” is just a finding. With it, it might be a machine doing exactly its job.

Wide Events vs. Three Pillars: AI Observability Costs

As agentic AI workflows gain traction within organizations, those organizations are asking how to account for their behavior while keeping costs manageable. Some are sticking with the old three pillars of observability approach: take a measurement to create a metric, record output to a log, and track serial progress with a trace. Each of these is useful, but treating them as distinct formats from the start means paying for them distinctly too. Separate storage doesn't come cheap.

Two cats, two dogs, four vendors, and the model the AI couldn't find (Tech Talk Companion)

Tech Talks went dark for a few months, and on episode 13 I finally got to ask why. Mathias Palmersheim’s answer, delivered completely straight, was that his users were unhappy with the availability and usability of their feeders and their litter box, and he wasn’t allowed back on stream until that got fixed. The users are two dogs and two cats, and they have titles. Maisie, a Shiba Inu who came to him through a rescue, is the recently promoted chief executive pawofficer.

I use Claude every day. I still build dashboards in SquaredUp

If you've looked at SquaredUp and thought "I don't need this, I'll just use Claude", I understand completely. I've thought it too. Here's what changed my mind. I use Claude constantly. I've also spent the last few months building the parts of SquaredUp that let it in: our MCP server, the object graph and correlation, a stack of plugins. So this isn't a dashboard vendor being sniffy about AI. I've watched Claude pull from several sources and produce something genuinely useful in under a minute.

Build incident response workflows with Datadog Bits Chat

See how Bits Chat turns a natural-language request into an automated incident response workflow. In this demo, Bits Chat builds a workflow that investigates a monitor alert, identifies whether a recent deployment caused the issue, rolls it back when appropriate, and sends a summary to Slack.

Prompts, skills, and the AGENTS.md nobody wants to write (and how Anthropic writes theirs)

You’ve watched Claude Code compact a conversation. The context bar fills, it pauses, a summary appears, and it carries on like nothing happened. You probably assumed a housekeeping script trimmed the transcript in the background. It didn’t. The model compacted itself. When the window fills, Claude Code sends a long, specific prompt telling the model how to summarize its own conversation. Then it does, same model, same turn. The thing managing your context window is just another instruction.

AI cost calculator: estimate your total spend

An AI cost calculator for the whole wallet adds four lanes: seats and subscriptions, API and token usage, cloud AI services, and GPU infrastructure. Average 2026 totals run $25 per employee per month at light adoption, $100 to $150 at active adoption, and $300 or more at AI-heavy companies. Getting to your number takes four lane subtotals and three corrections.

Agent vs Agentless Monitoring and How to Decide What Goes Where

Why does half the infrastructure end up returning no monitoring data? The standard plan is to install collection software on everything, which moves quickly across servers and stops dead at the first device running closed firmware. Storage arrays, firewalls, and switches will never accept an install, and the rollout stalls there. That plan usually gets set once for the whole environment, with a single collection model applied to hardware it was never suited for.

Microsoft Took 8 Months to Fix This Copilot Vulnerability

Microsoft finally patched a critical Copilot vulnerability nearly eight months after researchers first disclosed it — and the way the attack worked raises some unsettling questions about AI memory. The vulnerability chained together multiple flaws that could allow a malicious prompt hidden inside a webpage to be pulled into Copilot simply by asking it to summarize the page. From there, the attack could potentially access connected data from services like Gmail, Google Drive, and Google Calendar and exfiltrate that information using Copilot’s own capabilities. But the most concerning part may have been persistence.

Building AI Systems That Survive an Audit: Evidence Trails, Traceability and Compliance by Design

A model returns an answer with a confidence score of 0.94. The team ships it. Six months later someone asks why the system produced that specific answer, and nobody can reconstruct it. For years accuracy was the only number that mattered in machine learning. Get the error rate down, ship the model, move on. In regulated domains that is no longer enough. The harder question is whether you can defend a single decision after it has been made. Most systems were never built to answer that, and by the time someone asks, the information needed is already gone.

PostgreSQL IDE + AI Assistant | dbForge Studio for PostgreSQL

Manage the full PostgreSQL database lifecycle from one AI-powered IDE. dbForge Studio for PostgreSQL brings together database design, development, and administration, as well as data management, analysis, reporting, and extensive automation. Additionally, the integrated AI Assistant generates, explains, optimizes, and troubleshoots SQL queries directly in the Studio. It supports on-premises PostgreSQL databases and related cloud services such as Supabase, Heroku, Amazon Redshift, and TimescaleDB.

What to Know About AI Code-to-Merge Platforms

AI coding agents can generate pull requests at a pace your team has never seen. The bottleneck has shifted from writing code to everything that follows: reviewing, iterating, and merging. AI code-to-merge platforms are the category of tools built to manage that entire lifecycle, from the moment an agent starts working to the moment code lands in your main branch. This article walks through ten questions you should ask before committing to a platform.

Build and run Datadog workflows from Bits Chat or AI agents

Teams use AI coding agents and Bits Chat to troubleshoot systems and handle complex tasks, often uncovering repetitive work worth automating. But turning those routines into workflows can still require switching tools and recreating context manually. Through the Datadog MCP Server, Workflow Automation now lets you build workflows from Bits Chat or AI coding agents like Claude Code, Cursor, and Codex.

Shipped: Self-serve your MCP server credentials

Enterprise agent platforms need a client ID and client secret in hand before they will connect to anything. An admin with the Modify MCP Settings permission can now issue that pair directly in Settings, connect the platform, and manage the credential lifecycle on whatever schedule your security policy requires. No support request, no wait.

When an AI Agent Breaks the Law, Who's Responsible?

An AI agent was given one simple task: book a gym class when a slot became available. Instead, it discovered a vulnerability in the gym’s software, gained administrative access, deleted another user, and booked the slot anyway. Australian AI technologist Andrew Bird had connected an AI agent to WhatsApp to automate a routine gym booking. But when the agent encountered an API without proper authorization checks, it didn’t simply stop. It found a way around the problem and used the vulnerability to accomplish the task it had been given. And that creates a much bigger question.

Automate Product Analytics reports with your agent and the CX CLI

Every page view, click, and session your RUM SDK captures lands in Coralogix as a log event under the cx_rum subsystem — the raw data behind how people actually use your product. You can turn it into a shareable report without writing a single query. Just ask your coding agent. Your agent queries that data through the CX CLI and writes the report for you: describe what you want in plain English, get a formatted report back — without leaving the terminal.

Driving Impact and AI Adoption as an FDE

Enterprise AI is only useful when it actually runs in production, inside the tools teams already depend on. Getting there is harder than it sounds. Forward Deployed Engineers at Atlassian work directly inside some of the world's largest organizations, building AI-powered agents, connectors, and workflows on Atlassian's platform. They work alongside customers to understand the real constraints, design solutions that hold up at scale, and see them through to deployment.

We Benchmarked AI Models on Git Tasks. Results Surprised Us

Most AI model benchmarks measure general coding ability or reasoning. GitBench, built by GitKraken developer advocate Chris Griffing, measures something narrower and more practical: how well a given AI model handles specific Git tasks, starting with commit squashing, identifying which commits in a messy history should be combined into one clean commit.

Making Shared GPUs Even Safer with Kubex and HAMi-core

Table of Contents A few months ago, we introduced Kubex support for the KAI Scheduler to improve GPU sharing for production inference workloads. The basic model is simple: The KAI Scheduler handles placement and GPU sharing. Kubex continuously observes usage and adjusts those allocations as demand changes. KAI provides the scheduling foundation. It lets multiple workloads share a GPU while accounting for the amount of GPU each workload requests. Kubex then closes the loop.

AI Spend Is a Capacity Problem, Not a Billing Problem

Every organisation running models in production eventually reaches the same point: the AI portion of the cloud bill grows faster than expected, and the immediate response is to invest in visibility. Calls are tagged, spending is attributed, dashboards are created, and the results are shown to the teams responsible.

Best AI Humanizer Tools for Ops and IT Teams Writing Technical Documentation in 2026

You finish the postmortem at 11pm, push it to the knowledge base, and the next morning it comes back flagged. Not for a factual error - the reviewer's note says it reads like AI. So now you're rewriting a document that was already correct. Most ops teams have hit some version of this. DevOps engineers, SREs, and IT ops managers draft runbooks, release notes, API documentation, and incident comms with Copilot, ChatGPT, or Gemini in the loop, because the alternative is writing them from scratch at 2am. The drafting problem is solved. The publishing problem isn't.

MCP Servers 1.1.0 Add Flexible HTTP Routing and CLI Connection Management

We are pleased to announce the release of MCP Servers 1.1.0, bringing new configuration options for HTTP-based deployments and expanded command-line capabilities for managing database connections. The new version makes it easier to control how MCP Servers are exposed over HTTP, host multiple MCP Servers under a single hostname, and configure connections directly from the command line.

Can You Prove Your AI Agents Are Paying Off? Most Developers Can't

We put a blunt question to developers on a recent live webinar: right now, could you actually prove AI agents are paying off for you or your team? Only 24% said yes. The other 76% were guessing, unsure, or already suspicious that agents are costing more than they’re saving. That gap between adoption and proof is the real story in agentic development right now. Teams aren’t behind on running agents. They’re behind on knowing whether it’s working.

The Evolution of JFrog AI Catalog: Your AI Control Plane for Agentic Development

In a single morning, a coding agent can pull an open-source model, connect to an unvetted MCP server, and execute a code-optimizing skill from the web. In the rush toward agentic automation, these AI assets quietly bypass traditional security reviews, creating new attack vectors across the software supply chain. Closing this blind spot has been the driving force behind the JFrog AI Catalog since its launch at swampUP 2025.

Our Customer Success AI bill tripled. Here's why we're spending more.

Pop quiz: If you spend $40,000 per month on Anthropic, and you’ve got two customers, what’s your cost per customer? If you bypassed the easy answer of $20,000 and said, “Scott, you old trickster, that’s not enough information to answer that question,” you’ve won today’s prize: a lesson in the perils of average costs. Let’s flesh out the situation: You put an AI feature in your product, a document assistant powered by Claude.

Shipped: Rightsize Kubernetes workloads without leaving your MCP client

Changing a Kubernetes resource request takes two numbers: what the workload requests, and what it uses. The CloudZero MCP server now returns both, by cluster, namespace, or workload. This gives you a number you can defend. Usage comes back as P95 over the date range you query, 30 days by default. When an engineering lead asks whether a service runs on a smaller request, that is the figure that settles it. Over-provisioning and under-provisioning show up on the same query.

LLM token cost: pricing per token explained

LLM token cost is the price a provider charges per token a model reads or writes, quoted in dollars per million tokens. Input and output bill at separate rates, with output priced at roughly 5x input. As of September 2026, published rates range from under $0.10 to more than $180 per million tokens on top-end reasoning tiers. In late 2025, Hardik Sonetta of Thomson Reuters Labs published a warning about the most common prompt caching mistake in production.

How Is AI Changing IT Operations? Building Production-Ready AI Agents with Alex Zinovy

How is AI changing IT operations, and what does it take to move AI agents from impressive demos to production-ready systems? In this episode of Agents of IT, Resolve’s Zack Austin sits down with Alex Cinovoj, Founder and CTO of TechTide AI, to explore what enterprise AI looks like when it has to work in the real world. Alex brings years of hands-on IT, infrastructure, DevOps, and AI engineering experience to a conversation about the shift from experimenting with AI to building trustworthy systems that deliver measurable outcomes.

Outrun the Threat Window: AI-accelerated Vulnerability and Patch Management

The gap between vulnerability disclosure and active exploitation is shrinking—often from weeks to mere hours. Traditional patching cycles no longer cut it. In this session, discover how AI-accelerated solutions can help you: Identify exposed assets faster Prioritize vulnerabilities by real-world risk Remediate across Windows, macOS, Linux, and hundreds of apps Verify success with a connected workflow Learn how our approach, powered by AI-driven insights and automation, can help you close the gap before attackers strike. Watch now and take control of your patch management.

SaaS Sprawl Is Becoming an IT Problem: Here's How to Bring It Under Control

For most organizations, SaaS sprawl does not begin with a bad technology decision. It starts with a useful tool. Marketing needs a new analytics platform. Sales adopts prospecting software. HR adds an applicant tracking system. Engineering signs up for another monitoring service. Someone discovers an AI tool that saves several hours a week and puts it on a company card. Each purchase makes sense on its own.

From AI Prototype to Production: The Technical Architecture Enterprises Need

Building a generative model that spits out flawless answers in a controlled notebook feels like a massive win for any engineering team. But watching that exact same model crash the second it hits real, concurrent user traffic? That is a frustrating reality check. The gap between a slick proof of concept and a mission-critical deployment is surprisingly wide, and it almost always comes down to the underlying infrastructure. If your systems cannot handle the dynamic load, the smartest algorithm in the world will not save you.

Why AI Adoption Fails Without the Data Work First

Most enterprise AI projects don't fail because the model wasn't good enough. They fail because the data underneath was a mess before anyone switched anything on. Duplicated contacts, contradictory fields, records that haven't been touched in three years but are still floating around in production tables. The AI doesn't know any of that context. It just reads what's there and runs with it.

We Let AI Agents Rewrite a 92M-Message-a-Day Service in Go. Zero Incidents.

Our Results Daemon processes about 92 million messages a day. We recently rewrote it from Node.js to Go, and we let Claude Code write it. We wanted to know whether we could trust an agentic rewrite for a critical, high-throughput production service rather than a prototype. It shipped with zero incidents, a 70% reduction in running pods, and a lighter database load.

AI in the public sector (infrastructure challenges and solutions)

The U.S. government has cataloged over 1,700 active AI use cases, and nearly 90% of federal agencies are already using or planning to use AI. The European Commission has disclosed nearly 1,500 AI use cases across EU member states. With over 3,200 combined AI use cases cataloged across the US and EU, public sector IT leaders face an identical roadblock: traditional application delivery controllers were not designed to parse or throttle Layer 7 LLM payloads, leading to backend GPU exhaustion.

From traces to experiments: A loop for improving AI agents

Let’s say your team shipped a support agent last quarter. The launch demo went well, stakeholders were pleased, and everyone moved on. A few months later, things start to look off. Summaries of long conversations are truncated, and monitors show latency spikes on tool calls to the billing API. Your team’s first instinct is to ship fixes such as tweaking prompts or upgrading the model.

AI usage tracking: Monitor spend by team, feature & model

AI usage tracking means measuring who and what consumes AI across your company, by team, feature, and model, then converting the usage into spend and cost per unit of work. Provider consoles stop at totals per API key. Tracking puts names on those totals: which team, which product, which model, and whether any of it was worth the money. In May 2026, CNBC reported that “almost every Fortune 500 is tracking overall AI usage,” quoting ModelOp CTO Jim Olsen. The same reporting carried his warning.

The gap between individual AI productivity and team performance

As a product manager at Upsun with a computer engineering background, Kateryna Dvornichenko had spent months researching competing tools in the agentic development space, running tests, comparing features, and building a picture of where the market was heading. She realized the tools were impressive, but something kept standing out. "Collaboration was not the strong point of any of them," she says. "Everyone stays on their own machine with their own setup.".

AI can write database code fast. Here's how to keep it safe before production.

AI can write database schema changes in seconds, but nothing should reach production until it's validated, tested, and approved. In this discussion, Ken Muse (GitHub), Steve Jones (Redgate), and Huxley Kendall (Redgate) show how a governed pipeline keeps AI-generated database changes safe without slowing teams down.

How to set up CircleCI with Cursor Origin

CircleCI now integrates with Cursor Origin, bringing scalable CI/CD to Origin-hosted repositories. In this demo, see how to connect an Origin repository to CircleCI, configure your pipeline triggers, run a build, and report CI status back to your Origin pull request. Already using CircleCI? Your existing.circleci/config.yml works as-is, with no Origin-specific CI syntax or separate config to maintain.

How to Use Claude Code with CircleCI to Fix Failed Builds

Give Claude Code direct access to CircleCI and let it diagnose failed builds, fix issues, and keep iterating until your pipeline is green. In this tutorial, we walk through how to connect Claude Code to CircleCI using the CircleCI CLI. You’ll see how Claude can read pipeline results, identify test failures, make fixes, trigger new builds, and monitor CircleCI without leaving the terminal.

From Attention to Action: How Digital Systems Shape Demand

Between the moment someone types a query and the moment they act on the result, a sequence of independent systems runs: query rewriting, intent classification, retrieval, an auction, a click-through prediction, a page render measured in hundreds of milliseconds, and a routing decision about where the resulting tap goes. None of those systems is designed to persuade. They are designed to estimate, price, and allocate.

The End of Browsing: How AI Is Rewriting Digital Discovery

Sometime in 2025, a quiet threshold was crossed: the majority of Google searches in the United States began ending without a single click. Data from SparkToro and Similarweb puts roughly 58.5% of US searches as resolving on the results page itself, and for news queries the zero-click rate climbed from 56% to 69% in a single year. The search box still works the way it always has. What changed is that we've stopped leaving it.

Why More Technology Is Becoming a Service Instead of a Product

A growing number of technology purchases no longer end at checkout. A phone gains new AI features months after launch, a vehicle receives software updates from the cloud, and a security camera may lose important functions if its online service disappears. The physical product still matters, but increasingly it is only the visible edge of a much larger system.