Senior Data & Machine Learning Engineer Capgemini
Designed and implemented a real-time fraud detection pipeline using Spark Structured Streaming, Kafka, and Postgres enrichment, enabling sub-second anomaly alerts and reducing fraudulent transaction losses by 17% in the first year. Operationalized ML models in SageMaker using MLflow for tracking and version control, integrating FastAPI APIs to serve batch and streaming inference with a 35% improvement in model deployment time. Re-engineered legacy pipelines into modular dbt projects with Airflow orchestration, reducing developer onboarding time by 40% and significantly improving DAG clarity for cross-team collaboration. Integrated S3-based data lakes with Redshift Spectrum to allow ad-hoc analytics without additional ETL processing, cutting query turnaround times by 50% for business teams. Built monitoring dashboards with Grafana and Prometheus to track Kafka queue lag, ML latency, and Spark job health, reducing incident resolution time by 25% through proactive alerting. Developed an internal Feature Store using DynamoDB and S3, ensuring feature parity between training and inference stages, eliminating 90% of feature drift cases. Mentored junior engineers on Terraform, Docker, reproducible Jupyter workflows, and experiment tracking best practices, raising team productivity and code quality. Partnered with compliance and governance teams to create ML explainability standards, including SHAP visualizations for regulatory audit readiness. Automated model retraining workflows with CI/CD integration, ensuring production models remain above 95% accuracy thresholds without downtime. Collaborated with data scientists and product managers to align ML feature engineering with evolving fraud prevention strategies, improving detection coverage for new fraud patterns.