Operations | Monitoring | ITSM | DevOps | Cloud

How to Scale your AWS Infrastructure - Part 2

Welcome to the second post in a series of “How to Scale your AWS Infrastructure”. In the first post, we talked about horizontal scaling, autoscaling, CI/CD, infrastructure automation, containerization, etc. In this post, we will continue the discussion around databases, loose coupling, caching, CDN, etc. Let’s start the discussion with database scaling.

Podcast: Break Things on Purpose | Chris Martello: Day of Darkness

Dad jokes lead the way in this episode as we interview Chris Martello, manager of application performance at Cengage. Chris is a wearer of many testing hats, but his passion is chaos and breaking things on purpose. Chaos was a natural fit for Chris with his background as a middle school science teacher, so when he made the jump to tech chaos engineering was a natural fit.

What's new in Sysdig - March 2022

Welcome to another iteration of What’s New in Sysdig in 2022! The “What’s new in Sysdig” blog has fallen to me, Jason Donahue, for the month of March! I am a Solutions Engineer based in New Jersey and a member of the Sysdig US East Enterprise team since September, 2021. I have worn many hats in my career, from Networking to Systems Administration to Software Engineer.

New N-able Research: How are MSPs adapting to a rapidly changing environment?

Since the start of the pandemic, MSPs have faced extraordinary challenges—in many cases, carrying the burden of responsibility for ensuring their customers are able to continue operating in the face of an uncertain and constantly changing business environment. At the same time, they have also had to adapt their own ways of working in order to continue to operate and survive.

What's a fair compensation for being on-call?

For the vast majority of organisations, it’s necessary to have some form of round the clock cover to support the business. Whilst it’s most commonly a concern for engineering, it’s increasingly common to have folks from various disciplines available out-of-hours. Irrespective of role, compensating people fairly is an important factor of running a healthy and effective on-call system.

Video: How to configure and customize Grafana OnCall

Managing your on-call rotations just got a little less stressful. With Grafana 8.0, we introduced unified alerting, which centralizes alerting information into a single, searchable view. With the introduction of Grafana OnCall, an easy-to-use on-call management tool available in Grafana Cloud, you can now extend the alerting workflow in Grafana to ensure that the right notifications reach the right people at the right time using the right method.

Finding & Fixing Asymmetric Routing Issues

Asymmetric routing is a situation in which packets take one path to go from source to destination, but replies take a different path to return. Notice I called it a “situation” and not an “issue”? That’s because it’s not always a problem. It only becomes a problem where there’s something stateful in the path, like a NAT device or a firewall.

How to collect metrics and logs for NGINX using the OpsAgent

The Ops agent is Google’s recommended agent for collating your application’s telemetry data, and forwarding them to GCP for visualization, alerting and monitoring. The Ops agent collates logs and a metrics collector into one single powerhouse. Some of the key advantages of using the Ops agent are outlined below.