Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on DevOps, CI/CD, Automation and related technologies.

How to Implement Global View and High Availability for Prometheus

Ensuring that systems run reliably is a critical function of a site reliability engineer. A big part of that is collecting metrics, creating alerts and graph data. It’s of the utmost importance to gather system metrics, from several locations and services, and correlate them to understand system functionality as well as to support troubleshooting.

When to hire an Incident Commander

What comes to mind when you hear the term 'incident commander'? You are not alone if you think about fancy, tri-cornered hats, well-polished shoes, and a uniform weighed down by medals. The roles of incident commander, incident manager, or technical escalation manager have been typical in large organizations but are gaining popularity in smaller companies. For the purposes of this article, we will use the term 'incident commander,' but any of the above titles could work.

Rolling out Roles

We’ve been pretty lucky at incident.io to be able to avoid dealing with more complex authentication issues for quite a while, because we piggy-back on Slack to know who you are and which organisation you work in. Whole companies have been built around doing authentication and user profiles really well, so it was pretty neat to be able to avoid doing most of that work for so long!

Building Digital Platforms for Adaptive Resilience: Looking Inside Gartner Predicts 2022

With digital service expansion putting pressure on IT, 'Gartner Predicts 2022: Build Digital Platforms for Adaptive Resilience’ is a helpful guide for I&O leaders with their sights on 2025. If you looked up “tech trends” right now, how many search results would you expect to see? 100 million? 500 million? Think again. Between blog posts, research reports, and news articles, you’d actually find roughly 1.5 billion search results.

Hot Storage vs. Cold Storage

When it comes to data storage, all data isn’t equal. After all, the data you use daily doesn’t need the same level of protection or ease of access as long-term hot storage vs. cold storage backup. A large percentage of a business’ data remains unleveraged due to data management and security challenges, which highlights the need to implement a data storage strategy.

Shifting Left for DevSecOps Success

Not long ago, developers built applications with little awareness about security and compliance. Checking for vulnerabilities, misconfigurations and policy violations wasn’t their job. After creating a fully-functional application, they’d throw it over the proverbial fence, and a security team would evaluate it at some point – or maybe never. Those days are gone – due to three main shifts.

Top 12 Kubernetes Risks

What’s putting your K8s workloads at risk? You probably didn’t immediately think of memory and CPU resources—yet, these pose significant threats to cost and performance in your public cloud Kubernetes and OpenShift deployments. Learn about the top 12 K8s risks and how you can visualize the spread of risk in your containers deployment. You'll also hear a methodology for drilling down to individual misconfigurations and resolving them.

How to aggregate your Metrics using MetricFire

This article covers such a popular topic as using aggregation rules for metrics. We will learn why it is important to use aggregations and what tools exist for working with them. Also, we will explore all the benefits of using MetricFire's Hosted Graphite solution to store, process, analyze and monitor your metrics.

Introducing DevOps to the US Government - Part 2

In the first post in this series, I talked about the challenges for the US Government sector when attempting to introduce DevOps. The sector lags behind others such as Financial Services on every measure, yet the technical obstacles like a disruption to workflows and a lack of appropriate skills are the same.