Skip to main content
banner-vector

Large Manufacturer Eliminates Manual Incident Response with Red Hat Ansible Automation Platform and Dynatrace

INDUSTRY

Manufacturing

A large manufacturer needed to reduce manual remediation toil, replace aging Puppet workflows, and connect their observability platform to their automation tooling. Arctiq delivered a working Dynatrace and Red Hat Ansible Automation Platform self-healing architecture in six weeks, with an enterprise governance framework built to scale to approximately 600 nodes across multiple plants.

Client Overview

The client is a Fortune 500 company operating dozens of manufacturing facilities across the United States.

The Challenge

Every time Dynatrace flagged a problem, whether a service degradation, disk pressure, or an unstable application, someone had to act on it. That meant pulling engineers from other work, introducing a window of exposure, and watching mean time to remediation climb.

The team was also carrying legacy Puppet workflows that had grown difficult to maintain as their infrastructure expanded. And when HashiCorp Terraform finished provisioning new infrastructure, the handoff to configuration management was still a manual step.

The client needed a platform that could close all three gaps at once: eliminate manual response for known incident classes, standardize configuration management without starting from scratch, and connect provisioning to configuration in a single pipeline, all with the governance controls a distributed manufacturing operation requires.

Our Solution

Arctiq designed and delivered a solution centered on Red Hat Ansible Automation Platform (AAP) 2.x with Event-Driven Ansible (EDA) at the core. The goal was to prove real business value quickly, not with a broad rollout, but with a tightly scoped, outcome-tied engagement delivered in six weeks.

The architecture connected Dynatrace to EDA via webhook. When Dynatrace fired a problem event, an EDA rulebook evaluated it against defined conditions and automatically triggered the appropriate AAP job template, with no human in the loop. This created a closed-loop remediation flow.

Alongside the self-healing use case, we validated a migration path off Puppet using an idempotent Ansible baseline role, and wired a Terraform pipeline to invoke AAP job templates via REST API for post-provision configuration.

Our Process

The engagement had a six-week delivery window, so we structured the work to prove value early and build governance in parallel.

  1. Confirmed access to non-production VMs, validated networking and firewall configurations, and connected AAP to Active Directory, Git, and the Dynatrace webhook before writing a single playbook

  2. Collaborated with the client team to select the first incident class for self-healing remediation and identify the pilot host group for Puppet replacement

  3. Built an EDA rulebook that received Dynatrace problem events and automatically triggered AAP job templates for service failures, disk thresholds, and application recycling scenarios

  4. Developed a custom Ansible role delivering idempotent baseline configuration, including users, packages, files, and services, and validated it against 5 to 10 target nodes that previously ran Puppet-managed configurations

  5. Implemented an integration between Terraform and AAP so newly provisioned infrastructure was automatically configured as part of the same pipeline run

  6. Configured Automation Hub as the client's curated content repository with approved collections and execution environments, and set up AD-integrated RBAC with team and role structures and full audit logging

  7. Closed with a live stakeholder demo of a real-time Dynatrace to EDA to AAP remediation run with execution logs, and delivered a findings report and sizing guide for the path to approximately 600 nodes

The Roadblocks

The engagement scope was clear from day one, but the environment introduced one complexity that required careful handling: integrating EDA with Dynatrace's webhook meant validating the full event-to-remediation chain in a non-production environment that didn't perfectly mirror production alerting conditions. We worked with the client's team to simulate representative incident classes and confirm the rulebook logic held before the stakeholder demo.

The bigger challenge was governance. Building RBAC, credential segregation, and Automation Hub content controls that would actually hold at scale, across multiple teams and eventually multiple plants, required more architectural rigor than a typical PoC. We treated it like a production foundation, not a prototype.

The Tool Stack

Red Hat Ansible Automation Platform 2.x (Controller, Event-Driven Ansible, Automation Hub): The core automation engine. Controller ran job templates and workflows, EDA handled event-driven remediation, and Automation Hub provided the curated content governance layer.

Dynatrace: The observability source. Problem events from Dynatrace's API triggered EDA rulebooks via webhook, closing the loop between detection and remediation.

Red Hat Enterprise Linux 8/9: Base OS for AAP nodes and target systems throughout the solution build.

HashiCorp Terraform: Existing provisioning pipeline. We wired it to invoke AAP job templates via REST API as a post-provision configuration step.

Active Directory: Integrated with AAP for SSO, LDAP-based RBAC, credential segregation, and audit trail.

Git: Source control for all AAP Projects and Execution Environments.

The Results

The client went from manual incident response to a closed-loop remediation architecture that detects, evaluates, and fixes covered incident classes without human intervention, in the time it takes a webhook to fire.

The Puppet replacement pilot gave the team a repeatable, idempotent Ansible baseline they can extend to the full node inventory without rebuilding from scratch for each host group. The Terraform integration eliminated the manual gap between provisioning and configuration. The governance framework, including AD/RBAC, Automation Hub, credential management, and audit logging, was designed to hold as the platform scales.

Key Results

  • Delivered a working Dynatrace, EDA, and AAP self-healing flow in three weeks
  • Validated Puppet replacement on 5 to 10 nodes with a reusable Ansible baseline role
  • Eliminated the manual post-provision configuration step through Terraform and AAP API integration
  • Deployed enterprise RBAC with AD/LDAP auth, team structures, and full audit trail
  • Produced architecture documentation and sizing guidance for expansion to approximately 600 nodes
  • Closed with a live stakeholder demonstration and accepted findings report

The Roadmap

At the conclusion of the engagement, the solution provided a foundation for broader adoption. The roadmap included extending EDA remediation to additional incident classes and onboarding more teams under the existing RBAC framework.

Longer-term priorities included migrating remaining Puppet-managed workloads to Ansible, scaling AAP toward the approximately 600-node target across multiple plants, and expanding Terraform and AAP pipeline integrations as provisioning patterns evolved.

Arctiq also established the automation architecture and governance framework needed to support that future expansion.