Every business faces incidents, no matter how tight-knit or high-tech. Downtime, glitches, system failures, and security breaches are all part of online operations. So all companies must prepare to face such issues, including communicating them to key stakeholders. Take widespread data breaches, for example. If a breach occurs, a business might need to communicate with hundreds or thousands of stakeholders, including DevOps teams, affected accounts, investors, corporate leaders, and media outlets.
How do you track reliability in an organization with hundreds of engineers, dozens of daily production changes, and over 32 million monthly users? Even more, how do you do this in a way that's simple, presentable to executives, and doesn't dump a ton of extra work on to engineers' plates? Slack recently wrote about how they created the Service Delivery Index for Reliability (SDI-R), a simple yet comprehensive metric that became the basis for many of their reliability and performance indicators.
Ansible is a configuration management tool that helps you automatically deploy, manage, and configure software on your hosts. By turning manual workflows into automated processes, you can quicken your deployment lifecycle and ensure that all hosts are equipped with the proper configurations and tools. The Datadog collection is now available in both Ansible Galaxy and Ansible Automation Hub.