Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on Monitoring for Websites, Applications, APIs, Infrastructure, and other technologies.

PostgreSQL Connection Pool Exhausted in Kubernetes: Causes, Diagnosis and Fixes

"Connection pool exhausted" is one of PostgreSQL's most misleading production errors. The database is often healthy, while your application waits behind full connection pools. In Kubernetes, every pod and replica maintains its own pool, amplifying issues like connection leaks and slow queries across the cluster.

Scheduled Sharing in SquaredUp

Getting your dashboards in the right hands is vital for many teams. This typically means that someone has to remember to check them, and that rarely happens consistently. Stakeholders miss weekly updates, managers ask for reports that already exist, and your carefully-built dashboards sit unread while people wait for someone to send them a screenshot. Scheduled sharing is the frictionless feature that solves this issue.

Install AURA to Debug Incidents Using an Open Source SRE Agent

AURA is a fully open-source agentic harness built for SRE and production operations work. In this walkthrough, Mezmo forward deployed engineer Jeff iinstalls AURA on a local desktop, runs `aura init` to generate the config and connect it to an Anthropic Sonnet model, then wires in a Grafana MCP server pointed at his homelab. He hands AURA a live incident: a set of addressable LED lights that stopped responding to Home Assistant.

How GPS Ankle Monitors Work in Ohio: A Complete Guide to Electronic Monitoring and Bail Bonds

When someone you care about is arrested, the experience can be overwhelming. Along with concerns about bail, court dates, and legal representation, many families are surprised to learn that a judge may require a GPS ankle monitor as a condition of release. If you've never dealt with the Ohio criminal justice system before, you probably have questions. What is a GPS ankle monitor? How does it work? Can someone still go to work? Do you still need a bail bond? What happens if the monitor stops working?

What SREs Can Learn from Revenue Operations (and Vice Versa)

Site reliability engineers and revenue operations teams rarely sit at the same desk. Software engineers look after cloud infrastructure while operations professionals look after data pipelines and sales funnels. Yet both teams spend their days managing complex systems that can't afford to crash. When you look past the different tools they use, the underlying principles of both roles are almost identical. Let's examine how these two technical worlds can share practical insights to build better business systems.

Are You Audit-Ready? SecOps for SAP in the Age of Constant Change

Drift Happens. Six months ago, your SAP landscape passed audit cleanly. Today? You couldn’t say for certain without a manual scramble across Basis, security, and infrastructure teams to reconstruct what’s true right now. You did exactly that for the audit, after all. That gap between “was compliant” and “is compliant” is where most SAP security programs quietly fail.
Sponsored Post

No SAP Expertise? No Problem. Automation Just Got Easier

Avantra 21 introduced the concept of Automation workflows. By Avantra 23, compatibility with Ansible was added. Along the way, Avantra became the tool SAP teams reached for when they wanted system copies, refreshes, and backup orchestration to just run - predictably, on a schedule, without a senior engineer babysitting them. So, automation isn't new to Avantra users. Avantra made automation for SAP practical to deploy, predictable in operation, and easier to maintain.

Real-Time Satellite Monitoring with InfluxDB 3 & Claude

See how InfluxData built an AI-powered satellite fleet dashboard using InfluxDB 3 Enterprise. Head of Product Marketing Ryan Nelson demonstrates how the system handles high-cardinality telemetry and burst ingestion, detects anomalies in real time, and connects live operational data to Claude through the InfluxDB 3 MCP server. The result is faster fleet-wide analysis, reliable telemetry ingestion, and an integrated workflow for anomaly detection, investigation, and response.