Building an AI-Powered Observability & RCA Platform on AWS for ExtraMiles

Use Case - AWS GenAI Implementation

Industry: Financial Services

Geography: India

Employee Size: 300+

Solution: AWS

About Customer

ExtraMiles is a leading FinTech organization providing digital financial services, payment processing, and transaction management solutions to a large customer base. The company operates business-critical applications that require high availability, real-time monitoring, and rapid incident resolution to ensure seamless customer experiences and regulatory compliance. As the organization expanded its digital services and transaction volumes, its technology landscape grew increasingly complex, spanning cloud-native applications, databases, APIs, and distributed infrastructure components.

Extramiles Logo

Challenges

Manual, Time-Consuming Incident Investigation

Engineers had to manually analyze logs, metrics, alarms, and traces across multiple sources, resulting in longer resolution times and higher operational overhead.

Lack of Automated Root Cause Analysis

Without automated RCA capabilities, teams struggled to quickly identify the underlying cause of application and infrastructure issues.

No Centralized Incident Knowledge

The absence of a centralized knowledge repository caused similar incidents to be investigated repeatedly, with operational knowledge siloed within individual teams and dependent on subject matter experts.

Limited Visibility & Scalability

Limited visibility into application dependencies, combined with growing transaction volumes and infrastructure complexity, made operations reactive and difficult to scale. This increased MTTR while challenging platform reliability and performance.

Our Approach

enreap designed and implemented an AI Observability & Application Root Cause Analysis (RCA) Platform on AWS for ExtraMiles, built to transform incident management from a manual, reactive process into an intelligent, automated one. The platform continuously collects and correlates metrics, logs, alarms, and traces from application and infrastructure components using Amazon CloudWatch and AWS X-Ray. Event-driven workflows orchestrated through Amazon EventBridge, AWS Lambda, and AWS Step Functions automatically analyze incidents, while Amazon Bedrock generates contextual root cause insights, supported by a centralized knowledge repository in Amazon DynamoDB and predictive analytics powered by Amazon SageMaker.

ExtraMiles Architecture Diagram

Our Solution

enreap implemented an end-to-end AI Observability & RCA platform, combining automated telemetry collection, event-driven workflows, Generative AI, and historical knowledge management to deliver intelligent incident detection, analysis, and resolution.

1. Observability Data Collection
Amazon CloudWatch was configured to collect infrastructure and application metrics, logs, and alarms, with AWS X-Ray enabled for distributed tracing. Centralized monitoring was established across Amazon EC2, Amazon EKS, ALB, and Amazon RDS workloads, with CloudWatch dashboards providing real-time operational visibility.

2. Event-Driven Incident Detection
Amazon EventBridge rules were configured to capture CloudWatch alarm events and automatically route infrastructure and application incidents for real-time processing without manual intervention.

3. Telemetry Normalization
Telemetry Normalizer Lambda function standardizes incoming metrics, logs, alarms, and trace information, enriching it with contextual operational data before preparing structured incident payloads for downstream workflows.

4. Workflow Orchestration
An Observability Router Workflow, built on AWS Step Functions, classifies incidents into infrastructure/CPU-related and application-level categories, automating the orchestration of RCA processes and AI analysis pipelines

5. AI-Powered Root Cause Analysis
Dedicated Lambda-based RCA engines — a CPU AI RCA Engine and an Application AI RCA Engine — integrate Amazon Bedrock Nova foundation models to generate contextual root cause assessments, impact analysis, and remediation recommendations, reducing dependency on manual troubleshooting and SME intervention.

6. Historical Incident Knowledge & RAG
An Incident Knowledge Base built on Amazon DynamoDB stores incident patterns, RCA results, resolutions, and remediation actions, enabling similarity-based incident matching before invoking AI models. Amazon Bedrock Knowledge Base adds Retrieval-Augmented Generation (RAG), allowing AI models to retrieve relevant operational knowledge and improve the accuracy of RCA recommendations.

7. Intelligent Cost Optimization
The platform checks historical incident matches before invoking Bedrock, reusing existing RCA findings whenever a similar incident is identified — reducing Generative AI inference costs while maintaining response quality.

8. Automated Notifications & Long-Term Archival
Amazon SNS delivers RCA summaries and remediation recommendations to operations teams in real time. DynamoDB Streams and a DynamoDB-to-S3 Archive Lambda function archive incident records and RCA history into Amazon S3, creating a scalable historical repository for analytics, compliance, and future learning.

9. Predictive AI & Operational Dashboards
Amazon SageMaker analyzes historical operational data to identify recurring patterns and generate early warning indicators for potential service disruptions, while CloudWatch dashboards provide end-to-end visibility into infrastructure health, application performance, incident trends, and RCA outcomes.

Business Outcome

1. Reduced Mean Time to Resolution (MTTR) by approximately 60–75% through automated incident detection and AI-powered RCA
2. Achieved up to 80% reduction in manual troubleshooting effort by automating log, metric, alarm, and trace correlation
3. Improved incident response time from 30–60 minutes to under 10 minutes for common operational issues
4. Enabled 40–50% faster resolution of recurring incidents through reuse of historical incident knowledge
5. Reduced dependency on subject matter experts by approximately 50%
6. Reduced repetitive operational activities by 70% through automated workflows
7. Established a centralized repository containing incident records, RCA findings, and remediation actions for continuous learning
8. Reduced Generative AI processing costs by approximately 25–35% through historical incident matching and knowledge reuse
9. Improved monitoring coverage, achieving near real-time observability across critical workloads
10. Implemented predictive analytics capable of identifying recurring operational patterns before business impact
11. Automated incident archival and knowledge retention, improving audit readiness