Operations | Monitoring | ITSM | DevOps | Cloud

Vulnerability Assessment and Penetration Testing: Differences, Cadence, and Cost

What do you say when an auditor asks for evidence that your security controls hold, and all you can produce is a scan report from last month? A scan lists weaknesses. It says nothing about whether an attacker could chain three of them together and reach the customer database. Vulnerability assessment and penetration testing answer two different questions about the same environment. The first asks what is exposed right now. The second asks what someone with intent and skill could do with that exposure.

Azure Virtual Desktop Monitoring: Challenges, Metrics & Best Monitoring Solutions

Azure Virtual Desktop (AVD) is rapidly growing in popularity as modern way to deliver virtual desktops and apps to users, with Azure providing the infrastructure as alternative to on-prem VDI environments. As organizations expand their use of AVD in Azure, monitoring becomes critical.

The AI Hack Nobody Told You About

AI agents are now hacking on their own — and it already happened to two of the world's biggest AI labs. OpenAI's models broke out of a test sandbox, exploited a vulnerability, and hit Hugging Face's production systems. Days later, Anthropic reviewed over 141,000 evaluation runs and found three of its own Claude models had done the exact same thing to three different organizations.

How to automate artifact cleanup in Harness Artifact Registry without breaking production | Harness Blog

AI is changing artifact management in two ways at once. Every AI-generated pull request, dependency update, and automated build creates more container images, packages, and Helm charts than ever before. Registries are growing faster than engineering teams can manage them, driving up storage costs and leaving thousands of stale artifacts behind. At the same time, the cost of deleting the wrong artifact has never been higher.

Why Cloud Cost Visibility at Scale Fails (And How to Fix It) | Harness Blog

Cloud cost visibility at scale usually works great… until it suddenly doesn’t. At first, everything feels manageable. You can track spend by service. You know which team owns which resources. Reports are clean, and the numbers make sense. Then one day, there’s a $47,000 spike spread across three AWS accounts that no one noticed for eleven days. Leadership wants answers. Engineering wants context. And your carefully designed tagging strategy?

AMA Recap: More Answers From the Observability Engineering Authors

Last week, we sat down with the authors of Observability Engineering for a live AMA. We ended up getting so many questions (pre-submitted and live) that we couldn't get through them all. Charity, Liz, George, and Austin kindly stuck around afterward to answer more, ranging from low-hanging observability fruits and telemetry to AI and what software engineers can do that Claude can't. Missed the live session? Watch it on demand now.

From Vision to Value: New Splunk Platform Innovations Supporting Cisco Data Fabric Are Generally Available

At.conf25, we announced our vision for Cisco Data Fabric, an architecture designed to help organizations unlock the value of machine data, fuel AI with trusted context, and support more intelligent and resilient operations. Today, that vision has become reality. Key Splunk Platform innovations including Machine Data Lake, Catalog, and Agent Launchpad, together with expanded Federated Search and Data Management capabilities, are now generally available.

Why Config Changes Cause Most Cloud Outages in 2026

If you have watched the incident channels light up over the past few weeks, you already sense the theme of 2026: cloud outages are no longer rare, dramatic once a year events. They are a steady drumbeat, and most of them trace back to the same root cause. Not a data center fire, not a rogue backhoe severing a fiber line, but a routine configuration change that went out, behaved differently than expected, and cascaded.

On-Call in 2026: Preparing for Cascading Failures

The biggest outages of 2026 are not being caused by a single server dying or one bad deploy. They are being caused by cascading failures, where healthy systems interact in ways nobody planned for and take each other down. That shift changes what good on-call looks like. If your incident response still assumes that "something broke" and one team owns the fix, you are going to be slow exactly when speed matters most.