I am a Senior Data Engineer with 4+ years of experience designing, governing, and cost-optimizing production data platforms across Databricks, Spark, Delta Lake, Azure, and GCP.
I focus on building reliable, scalable, and cost-aware data systems. In my recent work, I led a ground-up platform rebuild for a global CPG client, defining standards for data quality testing, CI/CD, governance, observability, and maintainability.
I have strong hands-on experience with SQL, PySpark, incremental pipelines, dimensional modeling, and medallion architectures. I also work deeply with Databricks features such as Unity Catalog, Jobs API, cluster policies, and Delta Lake optimization patterns.
A major part of my work is reducing infrastructure cost without sacrificing reliability. I have redesigned streaming and batch workloads to dramatically lower cloud spend, including a pipeline cost reduction from USD 8,000 to USD 360 per month and a significant reduction in Azure monthly billing through workload isolation and right-sizing.
I enjoy solving complex data problems end to end, from ingestion and transformation to governance, monitoring, and analytics delivery. I have also built solutions involving web scraping, unstructured document extraction, and cloud automation for legal and operational use cases.
I am comfortable working in remote environments and collaborating with cross-functional teams. I value clean engineering practices, reproducibility, and practical architecture decisions that balance performance, maintainability, and cost.