Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on Monitoring for Websites, Applications, APIs, Infrastructure, and other technologies.

How to scale access control in Grafana Cloud

One of the primary reasons organizations adopt Grafana Cloud is to create a single pane of glass across the data they collect from self-hosted systems, cloud providers, and third-party platforms. Bringing those signals together enables richer correlations, reduces tool sprawl, and makes it easier for teams to understand what's happening across their environment. But as observability grows and becomes more centralized, access management becomes more important.

Runtime Aware PR Verifier | Lightrun

Lightrun's Or Golan demos the Runtime Aware PR Verifier, a new Lightrun product that simulates pull requests against live runtime behavior before you merge. Watch use Lightrun to simulate an AI-generated PR, identify the affected production flows, and uncover a hidden risk that static review would miss. Instead of only asking whether the code looks correct, Runtime Aware PR Verifier checks whether the change matches how your system actually behaves in production.

Site24x7 Free Training Series - Session 1: Introduction, Website Monitoring, RUM & DRA

Welcome to Day 1 of the Site24x7 Training Program! This is the first session in our 5-part training series covering every module of Site24x7. In this session, we introduce the platform and take a deep dive into: Website Monitoring Real User Monitoring (RUM) Digital Risk Analyser (DRA) Session 1 – Introduction & Website Monitoring, Real User Monitoring, Digital Risk Analyser.

The SolarWinds Customer Zero Story

In this SolarWinds Customer Zero story, team members share how they use SolarWinds products every day across observability, incident response, enterprise service management, log analytics, Kubernetes monitoring, and self-hosted infrastructure monitoring. Hear how internal teams serve as the first customer by testing real-world workflows, sending direct product feedback, and helping shape the platform through hands-on use.

DASH 2026 recap: Product news, sessions, and highlights

DASH 2026 brought thousands of engineers, builders, security professionals, and technology leaders to New York City for 2½ days focused on building, operating, and securing modern systems. Across hands-on sessions and more than 40 customer talks, teams shared how they’re tackling real-world challenges at scale with Datadog. On stage, the keynote set the direction for what’s next across observability, security, and AI, highlighting a shift toward more autonomous, AI-assisted operations.

From Alerting to Assurance: Why Proactive Operations Define Trust at Scale

There’s a difference between seeing a problem and preventing one is not a question of tooling. It is a question of operational posture. Across eleven operator interviews at Nexus Live, a consistent pattern emerged. Teams are not struggling because they lack visibility. They are struggling because visibility alone does not produce confidence. Alert floods, late root cause discovery, and 3am escalations have become normalized in hybrid environments. The result is not just fatigue.
Sponsored Post

Proactive error management: Collaborate effectively and work smarter with tags

Talking to many of our customers with different needs and use cases, one particular issue comes up all the time. When I'm seeing so many error groups in my app and so many error notifications in my inbox every day, it's easy to end up feeling overwhelmed. I want a more proactive system to alert me to which errors need attention and when, so that I can stop getting buried. Does this hit home? Then this article is written for you, the tech leads and the product managers who are on the front-line of issue prioritization.

Introducing AppSignal for Startups

Good monitoring shouldn't be a luxury for well-funded teams. Early-stage startups run the same production systems as everyone else, on a tighter budget. That's when clear observability earns its keep. Today we're launching AppSignal for Startups: an ongoing discount on the full AppSignal platform for early-stage teams, with a better deal for Y Combinator companies.

Stop Guessing Why Latency Spiked | Lightrun

Latency spikes are easy to detect. Understanding why they happened is the hard part. Gidi Freud explains how Lightrun helps engineers debug latency spikes by automatically capturing runtime context when a execution of code exceeds a defined threshold. Instead of only seeing that a method or code block was slow, you can capture local variables and source location from the exact execution that crossed the threshold.