ML Engineer & Data Scientist Ominimo Insurance
Reduced inference compute cost 4x while serving 10,000 concurrent quote requests by optimizing CatBoost and XGBoost models with TensorRT on GCP. Accelerated model retraining latency to under 24 hours by automating drift-triggered pipelines with GitHub Actions and GKE shadow deployments. Cut manual claims verification time from 14 days to under 48 hours by deploying a multimodal RAG system linking accident photos to audio transcripts. Validated pricing algorithms with statistical significance by designing an A/B/n experimentation framework.