Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on DevOps, CI/CD, Automation and related technologies.

Is your cloud data truly sovereign? The CLOUD Act & FISA 702 reality check

As UK public sector bodies, financial institutions, and enterprises accelerate cloud adoption, a pivotal question emerges: Who truly controls your data, and under which laws? With data breaches and regulatory scrutiny intensifying, storing data and workloads in a host country alone doesn't guarantee sovereignty. U.S.

Understanding GPUs for AI success: Insights from our panel discussion

This blog is based on the webinar, “Panel Discussion: Understanding the importance of GPUs for AI success”, you can watch the full recording by clicking here! Last week, we hosted a panel discussion surrounding the importance of GPUs for AI success that featured Kunal Kushwaha (Field CTO), Ben Norris (AI Engineer), and Kendall Miller (Strategic Business Development).

Tutorial: Visualize Your Puppet Data in Grafana with the Observability Data Connector

When you manage complex IT infrastructure, it becomes critical to use tooling to understand what’s happening across all of your systems in terms of performance, reliability, and compliance. Monitoring key indicators manually is simply no longer possible at that scale. Puppet has long been known as a solution for managing large environments and collecting a vast amount of data about your infrastructure, but accessing and visualizing that data in a meaningful way can be a challenge.

How To Sell Cloud Cost Optimization To Your CFO

You know you’re bleeding money in the cloud. Maybe not everywhere, but enough to feel it. Your engineers know it too. You’ve got idle resources humming away, AI workloads scaling like wildfire, and nobody can quite explain why last month’s bill jumped by 17%. So, you bring up the idea of investing in a cloud cost optimization product. Cue the skeptical glance from your CFO.

What is Java Performance Monitoring? [A Guide to DevOps Engineers]

You rolled out a Java application that worked fine in development. Fast, clean, no errors. However, once it went into production, things began to change. Suddenly, the app feels slow. CPU usage climbs without warning. Some users start getting timeouts. You check the dashboards, but nothing jumps out. You look through the logs, but it's mostly noise. And then the questions start coming in - "Is the JVM the problem?" If you've been in that situation, you're not alone.

What to expect in a Gremlin workshop

Gremlin workshops give your team hands-on training with Gremlin so they can get real results and dramatically improve your reliability. Full transcript:  The goal of our workshops is really to accelerate you and the team in your reliability journey. Whether you're starting out for the first time, or you're a more advanced user, this workshop is really designed for you to take you to the next level.

Automating Linux Disk Expansion with Resolve: Add & Extend VM Disks in Minutes!

Running into disk space issues on your Linux servers or virtual machines? In this step-by-step demo, we show how Resolve’s powerful automation platform can help you automatically add and expand disk space on Linux systems, eliminating manual processes, reducing human error, and improving operational efficiency. In this video, you’ll learn how to: Technologies Featured: Whether you're a system admin, IT operations engineer, or automation specialist, this demo highlights how to streamline critical disk management tasks that normally require elevated access and technical knowledge.

Automate Disk Space Management on Windows with Resolve

Struggling with managing disk space issues on your servers or virtual machines? See how you can use Resolve to automate disk space addition and expansion on Windows systems, saving time, reducing manual errors, and eliminating the need for high-level administrative access. In this video, you'll learn how Resolve automates the process of: Whether you’re a system admin, IT operations engineer, or automation enthusiast, this demo highlights how you can streamline infrastructure tasks using intelligent automation.

Use Telegraf Without the Prometheus Complexity

Every system needs observability. You need to know what your CPU, memory, disk, and network are doing, and maybe keep an eye on database query latency or Redis connection counts. But setting that up isn’t always simple. You start with a couple of shell scripts. Then come exporters. Then Prometheus. Before long, you’re managing scrape configs, tuning retention, and watching dashboards fail under load after two days of data.

Maximizing Peering Through Flow Analysis

Discover how to use flow data to pinpoint your most valuable traffic, identify missing peer opportunities, and make smarter peering choices across your internet exchanges. In previous peering blogs, we’ve shared how you can maximize the value of your connection to an IX by peering with the IX route servers, and identify and contact specific peers via bilateral sessions.