Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on DevOps, CI/CD, Automation and related technologies.

Sidecar or Agent for OpenTelemetry: How to Decide

Getting telemetry out of a distributed system isn’t the hard part. Getting it out cleanly, without noise, drop-offs, or odd performance side-effects — that’s where things get interesting. Before you worry about processors or storage costs, you need a clear plan for where the OTel Collector should run. Most teams narrow this down to two options: a sidecar that sits next to each service, or a node-level agent that handles data for everything running on the node. Both patterns are solid.

Jira Service Management (JSM) Review for Alerting (2025)

Atlassian is shutting down OpsGenie. New sales stopped on June 4, 2025, and the platform will be completely offline by April 5, 2027. As an OpsGenie user, you now face a critical decision: Migrate to Jira Service Management (JSM), Atlassian’s recommended path, or choose a different solution. And if you’re not sure JSM is the right fit for your team’s alerting needs, this review will help you decide. I signed up for JSM and put it through real-world testing.

Azure Cost Optimization: Best Practices for Cloud Solution Providers

In this episode, we explore practical Azure cost management strategies tailored for Cloud Solution Providers (CSPs). The conversation dives into cost visibility, optimization techniques, and billing transparency, helping CSPs improve margins and deliver more value to their customers. Featuring experts from West Coast, a leading CSP, including James Reed (Azure Sales Manager) and Mitchell G. (Azure Sales Specialist), along with Mike Stevenson, the discussion highlights real-world insights from the partner ecosystem.

CEO Diaries: Not All AI Talent Is Alike

If Meta’s (now halted) nine-figure AI talent poaching scheme was any indication, the AI talent market is pretty frothy. The number of AI-related job postings has roughly tripled since 2019, and the average salary has more than doubled (Bain). The race is on for companies to find the fastest, most sustainable routes to AI-driven business value; all companies, but especially software companies, are hotly pursuing racers. But despite what Zuckerberg & Co.

Open Source Cloud Orchestration Tools Compared

Before 2011, cloud infrastructure was still new. AWS had launched EC2 and S3 in 2006. But to deploy applications, engineers had to manually spin up servers, configure storage, and set up networking — all by hand or with custom scripts. There were early configuration management tools, such as Chef and Puppet, but those didn’t offer full cloud orchestration. Then in 2011, AWS launched AWS CloudFormation as the first major orchestration tool.

Building a Pregnancy App You Can Actually Trust!

We never talk about pregnancy in the workplace. Maybe it's time to change that. Tech is still male-dominated, which creates a ripple effect: poor maternity policies, overworked expecting developers, and privacy-invasive apps that fail the people who need them most. But what happens when a developer decides to solve this problem themselves? Rizel built a pregnancy app she could actually trust. Not because existing solutions didn't exist, but because they weren't built with real privacy, real needs, and real developer insight in mind.

RancherLive: Know before you Go - KubeCon Atlanta Edition

Join us for a special Know Before You Go online session all about getting ready for KubeCon + CloudNativeCon in Atlanta! Hosted by Orlin Vasilev, this live event will feature special guest Nick Eberts — a proud Atlantan, musician, father, and Project Manager at Google. Nick will share insider tips on making the most of your KubeCon experience — from navigating the conference to exploring the best of Atlanta’s music, food, and culture. Whether it’s your first time at KubeCon or your first visit to the city, this session will help you feel right at home.

Monitoring Chaos Experiments with New Relic Probe in Harness

New Relic probes in Harness Chaos Engineering let you automatically validate system performance against defined SLOs during chaos experiments, transforming subjective testing into objective, metrics-driven resilience validation. By querying New Relic metrics in real-time and comparing results against your success criteria, you can programmatically verify that your systems maintain acceptable performance levels even under failure conditions.