Skip to content
Research Article Open access CC BY 4.0

AI-Powered Auto-healing Agent for Real-time DataOps Pipeline Failures

Vandana Kollati

Asian Journal of Research in Computer Science · pp. 75–86 · Published 3 Oct 2026

10.9734/ajrcos/2026/v19i10919

Abstract

Modern DataOps environments are increasingly complex because of distributed continuous integration/continuous delivery (CI/CD) frameworks within lakehouse architectures. This scale amplifies failures arising from schema drift, transient network errors, and upstream data inconsistencies, resulting in prolonged downtime and increased engineering overhead. This study presents an auto-healing agent for real-time incident detection, root cause analysis (RCA), and transactionally safe remediation of routine failures. For complex cases, structured escalation preserves timely human oversight. The framework integrates open lakehouse technologies into a unified system: Apache Hudi supports near-real-time streaming signals; Apache Iceberg provides consistent time-travel snapshots for state comparison; and Delta Lake ensures safe ingestion and recovery without disruption. Detection employs an imbalance-aware weighted Random Forest classifier with class and sample weighting and a modified Gini impurity objective to address extreme class imbalance (~97.4%) in operational logs. On the Hadoop Distributed File System (HDFS) anomaly benchmark (>11 million logs; 37:1 imbalance), the model achieved a weighted F1-score of 85.74%, covering 468,279 detected anomalies. The policy prioritises anomaly recall (37.59%) over precision (4.61%) to reduce missed detections, directing 99.8% of cases to Level-2 escalation. Healing actions respect dependency paths through confidence-based checks before remediation to control risk propagation. Hybrid RCA combines rule-based analysis with simulated generative reasoning without relying on production LLMs. Operational key performance indicators (KPIs), including recovery time, indicate a 20–30% reduction in repetitive Level-1 tickets and a measurable improvement in mean time to recovery (MTTR) for automatable incidents. The results support a scalable, transparent, and safety-first framework for reliable DataOps pipeline operations.

Apache Hudi Apache iceberg DataOps Delta Lake Lakehouse architecture anomaly detection oot cause analysis weighted ensemble learning

Cited by 0

No indexed citations yet.

Article metrics

Real usage data collected on this platform.

1

Page views

0

PDF downloads

0

Outbound clicks

0

Citations

Views over time

Views by country

Approximate, from request IP at view time — not citizenship or institution. Countries with fewer than 5 views are grouped as "Other".

Traffic sources

Referring site, by host.

Views and downloads exclude known bots/crawlers. Citations combines this platform's own DOI-resolved index with each external source's own reported total — see Cited by above for individually listed citing works. Last refreshed 0 seconds ago.