I am a Lead Data Engineer with more than nine years of experience building and operating large-scale data pipelines, streaming systems, and production-grade AI/ML solutions. I specialize in designing cloud-native data platforms that process high volumes of IoT, security, retail, and operational data.
I have deep hands-on expertise in PySpark, Apache Spark, Hadoop, Hive, HDFS, Kafka, and AWS Kinesis. My work includes developing reliable batch and near-real-time ETL pipelines, optimizing distributed workloads through partitioning and memory tuning, and implementing robust data models for analytics workloads.
I build machine learning, NLP, computer vision, anomaly detection, and predictive modeling solutions using TensorFlow, PyTorch, and scikit-learn. I also have experience integrating generative AI capabilities through LangChain, LangGraph, OpenAI API, AWS Bedrock, Google Vertex AI, RAG pipelines, and vector databases.
I am experienced in multi-cloud engineering across AWS, GCP, and Azure, including workflow orchestration with Airflow and deployment automation through CI/CD pipelines. I use Docker, Kubernetes, Terraform, Jenkins, and GitHub Actions to support scalable and repeatable delivery processes.
I place strong emphasis on data quality, reliability, security, and observability. I have implemented schema validation, reconciliation, SLA monitoring, RBAC, IAM controls, and monitoring solutions using CloudWatch, ELK Stack, Prometheus, Grafana, and New Relic.
Across my roles, I have led initiatives involving cloud-native IoT platforms, security analytics and SIEM pipelines, PCB automation, retail POS systems, and AI-powered document processing. I combine strong data engineering foundations with practical ML and platform engineering expertise to deliver reliable business-focused data products.