Product Tech Lead ParsTech AI
Led MLOps and infrastructure automation by architecting Infrastructure as Code workflows using GitLab CI/CD and Ansible to automate GPU cluster provisioning, reducing setup time from days to hours. Migrated monolith to containerized AWS EC2 microservices with SQS task queuing ensuring 99.9% uptime. Established monitoring with Prometheus, Grafana, and CloudWatch for model drift and latency. Built GenAI pipeline processing multimodal inputs to auto-assign Jira issues, reducing diagnosis time from 8 hours to 20 minutes. Scaled ParsChat support ecosystem to handle 20,000+ daily messages for e-commerce and customer support automation. Optimized deployment of QWEN models using vLLM on NVIDIA RTX 5090 infrastructure achieving 60 messages per minute throughput. Re-engineered retrieval pipeline into graph-based architecture with intent-based routing to process up to 10 million characters with zero hallucination. Reduced chatbot delivery time from 1 week to 5 minutes and cut setup costs from $2.50 to $0.10 per user. Fine-tuned face recognition models on 49M images using NVIDIA A6000 GPUs achieving 99.97% accuracy. Optimized deep learning models for CPU-based edge architecture achieving real-time processing. Deployed services using FastAPI, gRPC, Celery, and Redis for scalable task management.