I am a Data Engineer with more than three years of experience building and scaling data platforms for high-volume fintech, payment, e-commerce, and digital pharmacy environments. I specialize in the Hadoop ecosystem and use Java and Scala to develop reliable, scalable data solutions.
I design and support Data Lake architectures, Spark-based ETL and ELT pipelines, and near-real-time streaming systems. My experience includes Apache Spark, Kafka, HDFS, SQL, S3, Airflow, Kubernetes, Git, and Linux.
I have delivered measurable performance improvements in production data workflows. In one case, I refactored a critical reporting process from a two-to-four-hour runtime to five-to-ten minutes, enabling near-real-time business insights.
I also have experience migrating batch ETL workloads to streaming architectures. I migrated core daily jobs to Kafka and Spark Streaming pipelines, reducing data latency from 24 hours to under 10 minutes for live marketing dashboards.
I have worked on data reliability and governance capabilities, including HDFS table versioning, point-in-time queries, rollbacks, and real-time audit logging through the ELK stack. I am continuing my education in Computer Software Engineering while building toward more senior data platform engineering responsibilities.
I communicate effectively in English and Russian and have teaching experience in Python, data analysis, and mathematics. I am interested in remote data engineering roles where I can contribute to scalable data platforms and real-time analytics systems.