
Architecting autonomous incident response for global payments.
Greenfield
System
Autonomous L2
Escalation Tier
Payments
Domain
Real-time
Processing
Overview
At Visa — the network behind hundreds of billions of dollars in annual payment volume — I work on the full-stack data pipeline and quantitative infrastructure team, architecting AIMS (Active Incident Management System) from a greenfield codebase. AIMS replaces slow, manual L2 escalation with an autonomous orchestration layer that detects, classifies, and routes production incidents in real time across Visa's payment infrastructure.
The core of the system is a proprietary routing engine paired with ML-based triage and real-time classification models that operate directly on Visa's production data lake payment telemetry. Instead of a human paging through dashboards, AIMS ingests high-volume event streams, correlates anomalies, infers likely root causes, and dispatches closed-loop remediation playbooks — dramatically compressing the window between detection and resolution.
I design high-throughput ingestion and enrichment pipelines over massive financial data lake volumes, including autonomous knowledge-graph construction that links services, dependencies, and historical incidents into queryable root-cause inference paths. On the platform side I build the web services and APIs that expose this intelligence to on-call engineers and downstream automation, and implement production-grade stream processing for real-time incident enrichment, anomaly correlation, and automated remediation workflows.
What I did
- —Architecting AIMS (Active Incident Management System) from greenfield on the full-stack data pipeline & quantitative infrastructure team — autonomous L2 incident orchestration replacing manual escalation with proprietary routing logic, AI-driven triage, and real-time classification on production Visa data lake payment telemetry
- —Designing high-throughput ingestion & enrichment pipelines over massive financial data lake volumes — autonomous knowledge-graph construction, remediation playbooks, root-cause inference paths, and web services for closed-loop incident response automation
- —Implementing production-grade stream processing with data lake access for real-time incident enrichment, anomaly correlation, and automated remediation workflows across payment infrastructure
- —Building the API and web-service layer that surfaces triage recommendations, incident context, and knowledge-graph inference to on-call engineers and downstream automation systems
Tech stack
Languages
Data & Streaming
ML & Intelligence
Platform
Connect with Dhruv Hegde
See more of Dhruv Hegde's work and background on LinkedIn, GitHub, and ResearchGate.