Senior Software Engineer Peer Consulting Resources Inc.
Designed and operated a Go-based real-time bid-response service for ad exchange traffic under strict latency requirements. Improved the request pipeline by overlapping feature retrieval and model inference, reducing p99 latency from approximately 180 ms to 95 ms. Built a TensorRT GPU inference layer with dynamic batching across concurrent model versions, increasing sustained throughput by about 35%.
Implemented multi-region resilience capabilities including request hedging, circuit breakers, regional caching, feature-store fallback behavior, and autoscaling. Resolved a batching-window production incident that reduced GPU inference queue wait by roughly 55%, led shadow-traffic and experiment-cohort rollouts for new model architectures, and scaled the fleet to approximately 2.8 million requests per second. Mentored engineers on distributed tracing and latency guardrails.