Architecting hyperscale data foundations and agentic AI systems for a world where data must not just scale—it must be certain.
For over 15 years, I've engineered data platforms that don't just scale—they think. Processing massive streams with sub-second latency, and bridging extreme data architecture with Agentic AI.
Building hyperscale infrastructure across industry leaders.
• Engineered a petabyte-scale streaming analytics framework using Apache Flink and Apache Pinot to modernize core enterprise contact-center reporting.
• Guaranteed sub-second query latency for more than 10,000 global customers.
• Reduced continuous compute costs by $150k-$200k monthly and cut pipeline processing time by 80% through optimized streaming ingestion patterns.
• Architected an agentic operational interface for Apache Pinot using LLMs and MCP, enabling cross-functional teams to query real-time streaming data through natural language.
• Developed a globally distributed Data Federation architecture across multiple regions, integrating batch and streaming pipelines into Apache Iceberg with GDPR-compliant PII masking.
• Leveraged the DeltaIO framework to build a customized Bring Your Own Compute (BYOC) infrastructure, enabling more than 800 enterprise customers to execute isolated analytical workloads.
• Led architectural review boards to establish enterprise-wide standards for streaming data platforms and AI integrations.
• Engineered and deployed a homegrown multi-cloud data lake ecosystem from the ground up to support Apple's petabyte-scale operational demands.
• Spearheaded the migration of massive data workloads from the on-premise Apple Private Cloud to an auto-scaling Amazon EKS infrastructure.
• Integrated Apache Spark and Apache Iceberg on Amazon S3 to analyze petabytes of data with strict ACID compliance.
• Designed an active-active, cross-region Disaster Recovery framework in AWS.
• Created a custom S3 data replication engine that bypassed native replication features, reducing disaster-recovery infrastructure costs by millions of dollars annually.
• Achieved an 80% reduction in cloud compute costs and a 4x improvement in query execution performance through rigorous Apache Iceberg table rewrite procedures.
• Mentored and upskilled mid-level and junior engineers in cloud-native data engineering and distributed systems best practices.
• Led the real-time, secure integration of India's Aadhaar National Biometric Identity system with legacy core-banking infrastructure for Andhra Bank.
• Designed and implemented RESTful web services using Spring Boot for banking and government-sector applications.
• Architected and developed an operational portal for NMPT to automate ground-level cargo handling and secure regulatory communications.
Formalized the R=W/L inflection point for petabyte-scale Lakehouse consistency.
Analyzing token economics and auth patterns for Agentic AI interfaces in analytical engines.
Layered architecture for 96% confidentiality in globally distributed network environments.
Achieved 96.8% privacy preservation and 94.5 Mbps secure throughput in distributed transfer.
Built a low-latency trading pipeline utilizing Apache Flink and Kafka. Engineered an ML inference engine using 5 regime learners.
Authored an open-source agent workflow for an iterative worker-reviewer cycle with subagent critiques.
Foundational tools for modern data platforms.
Fixed non-daemon threads blocking JVM shutdown in RenewableTlsUtils.
Merged configured and discovered provider models, unifying the API compatibility layer.