Operations | Monitoring | ITSM | DevOps | Cloud

OpenTelemetry eBPF Instrumentation (Grafana OTel Community Call #9)

In this episode of the Grafana OTel Community Call, we're exploring OpenTelemetry eBPF Instrumentation (OBI). OpenTelemetry eBPF Instrumentation (OBI) offers a powerful way to instrument applications at the system kernel level, capturing essential “RED metrics” — request rate, error rate, and duration — and network flows without requiring code changes, rebuilds, or redeployments. We will cover the project's architecture, discuss its origins as Grafana Beyla, and look ahead to the roadmap for language and runtime coverage.

Can a T-Shirt Fool AI? Why AI Guardrails Matter for IT | Zero Ticket Minute

Can AI be influenced by something as simple as a T-shirt? New research suggests irrelevant context can affect how some AI models respond. In this Zero Ticket Minute, Ian explains why AI guardrails matter and what IT leaders should consider as they adopt agentic AI and autonomous operations.

GCP Monitoring: A Complete Guide to Monitoring Google Cloud Applications and Infrastructure

Most production incidents in Google Cloud don't announce themselves as infrastructure problems. A checkout service on GKE starts timing out, a Cloud Function cold-starts under load, a Cloud SQL replica falls behind, and a Pub/Sub subscription quietly backs up until messages start expiring. None of that shows up as a red node in a compute dashboard. It shows up as slow requests, failed webhooks, and a support queue filling up faster than anyone can triage it.

Switching Between AI Agents Like This Is a Game Changer #ai #productivity

AI didn't just change how fast code gets written. It exposed a new bottleneck: everything around the code. Reviews slow down. Context gets lost. Planning drifts from implementation. Teams move fast and still feel stuck. That's the problem GitKraken is built to solve, and this Friday we're going live to walk through what's changed. We'll cover the latest Code Flow Company features we've shipped, how they connect developers, AI agents, and production into one system, and what it actually looks like to go from plan to main without the chaos.

Agentic Pipelines | Bitbucket Blitz | Atlassian

Most CI/CD pipelines are fragile bash scripts that break when things change. What if your pipeline could think? Agentic Pipelines lets you add AI agents as steps in Bitbucket Pipelines. In this video, I show an agent that reads a design spec from Confluence, generates frontend code, runs tests, and opens a PR, all inside a pipeline. With Agentic Pipelines, Bitbucket goes from a CI/CD platform to a full workflow and automation engine you can use far beyond builds and deploys.

Building with AI: Our Approach to Responsible Agentic Development in Open Source

The tech world has been building up towards the shift to a fully agentic development life cycle for a few years now. AI is changing how software gets built. Across the Puppet ecosystem, we’re seeing a shift toward more agentic engineering workflows. AI helps generate code, shape documentation, and accelerate how Puppet modules evolve.

How to Choose the Right Infrastructure Monitoring Tool

A production service degrades, and one question decides the next hour: is it the server, the network, or a cloud dependency? Each layer usually reports into a separate console, so pinning down the answer can absorb an hour the business would rather not lose. The right infrastructure monitoring tool is what turns that hour into minutes. On paper, most monitoring platforms look identical. Each one promises full-stack visibility and shows a polished dashboard.

Ensuring Business Continuity in Adverse Conditions

Businesses will always face disruptions. Whether it's a big storm, a broken supply chain, or a power outage, unexpected problems can bring operations to a halt, hurting your income, your reputation, and how much customers trust you. The companies that make it through these tough times, and those that don't, often come down to one thing: resilience. Being a resilient organization isn't about building an unshakeable fortress. It's about being flexible, thinking ahead, and having the right systems to bounce back when disruptions occur.

The Failure Mode Your Runbook Probably Does Not Cover

Operations teams rehearse plenty of scenarios. Failed deployments, database corruption, certificate expiry, a region going dark, the on-call engineer who cannot be reached. What gets rehearsed far less often is the building losing power for eleven hours, because that feels like somebody else's problem, filed under facilities alongside the air conditioning and the parking barrier. It stops being somebody else's problem at the moment the UPS batteries drain and everything still running on premises goes down at once.

Incident Response Communication: Why Ops Teams Own the Narrative

Your monitoring stack flagged the outage in 90 seconds. A customer posted about it in 40. That gap is now the defining challenge of incident response communication. Ops teams have spent years driving down recovery times, yet very few track how quickly a public explanation takes shape. This article looks at how teams can monitor both timelines - and respond before speculation hardens into accepted fact.