Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on DevOps, CI/CD, Automation and related technologies.

What is APM Tracing?

APM tracing records the complete execution path of a request as it travels through your system, including database queries, external API calls, cache lookups, message queue events, and inter-service requests. Each step is captured with precise start and end timestamps, duration, and context such as service name, operation name, and relevant attributes. This lets you pinpoint where latency or errors originate without piecing together metrics and logs manually.

Building a DORA metrics Scorecard

There are a lot of ways to gauge the performance of your DevOps teams and the health of your software, but DORA metrics have emerged as the industry standard. If you aren’t familiar with DORA metrics, take a few minutes to read this comprehensive guide to understanding DORA metrics. DORA metrics were designed to offer a high-level, long-term view of how your teams are performing.

Self-Healing Networks and Sovereign AI: The Future of Global Connectivity

Bridging East and West in today’s digital landscape takes more than connectivity – it requires navigating regulations, geopolitics, and the rise of AI. In this episode of Uplink, Elena Chernykh, Head of Europe Enterprise Sales at CITIC Telecom CPC, joins Michael Reid to discuss how enterprises can thrive in a world where data sovereignty and AI demands are reshaping infrastructure.

Visualize Logs Alongside Metrics: Complete Observability for Slow MongoDB Operations

MongoDB’s strength of flexible schema and fast iteration can also hide costly queries until they surface as user-facing latency, replica lag, or spiky CPU. A handful of slow operations can impact the cache, starve other workloads, and cascade into timeouts across services. Monitoring slow queries gives you an early warning system for index gaps and query-plan regressions introduced by code deploys, schema changes, or shifting data shapes.

10 Ways to Optimize Data Center Operations

Running a data center efficiently is no small feat. From managing energy costs to preventing downtime, there's a lot that can go wrong—and a lot that can be optimized. Discover 10 actionable ways to enhance your data center operations, with practical tips on how Hyperview DCIM software can help you achieve these improvements more easily and effectively.

Why Cost Optimization Should Be More Like Pulling Levers, Not Using Scissors

The cloud, as we know it today, was created as recently as 2006. For most of its lifespan since then, companies have been throwing money at cloud services with abandon. The competitive edge gained by having the newest, best, and most powerful tools at their disposal made it worthwhile for companies to spend ever-increasing amounts without too much worry.

What is Automated Incident Response

While writing our 2024 recap, we found that teams handled over 2.2 million new incidents. Critical incidents alone tripled, increasing from 3,000 in 2023 to 9,200 in 2024. Dealing with such a large volume of incidents is not an easy task. And dealing with them manually is definitely not easy. Your valuable time goes into routine tasks like creating tickets, setting up war rooms, and notifying stakeholders. These keep you from fixing the actual problem.

AI's Impact on Developer Experience: GitLens Creator Eric Amodio on the Future of Coding

AI is reshaping how developers work, from enhanced autocomplete to agentic workflows. GitKraken CTO and GitLens creator Eric Amodio breaks down the current state of AI in development, potential risks of over-reliance, and where the industry is heading. Learn about the evolution from simple code completion to sophisticated agents, the challenges facing junior vs senior developers, and practical advice for leveraging AI tools effectively.

What is Single Pane of Glass Monitoring and How Can Enterprises Leverage It for Enhanced Visibility?

Large enterprises today grapple with increasingly complex IT environments - spanning multiple cloud services, hybrid infrastructures and countless applications. Exacerbated by technology silos, the sheer volumes of data generated in such environments can quickly overwhelm IT teams, impairing their ability to identify and respond to customer impacting issues before outages strike.