I am a principal-level software engineer with more than 20 years of experience building GPU-accelerated media pipelines, SDKs, and computer vision systems. My background spans major technology companies including AMD, Microsoft, Amazon, and Skype, where I have consistently worked on performance-critical products that move from research prototypes into production.
I specialize in C/C++ performance optimization, GPU programming with CUDA and ROCm/HIP, and modern machine learning engineering. My work also includes Transformers, RAG, LLMs, PyTorch, Docker, and CI/CD, with a strong focus on practical deployment and measurable performance gains.
At AMD, I led architecture and delivery for SmartAccess Video, a distributed GPU processing SDK that improved transcoding speed by more than 60%. I also owned profiling and optimization efforts across CPU and GPU bottlenecks, using tools such as GPUView, CodeXL, and VTune.
My experience includes technical leadership and cross-functional collaboration across enterprise integrations, media platforms, and real-time communication systems. I have worked on products and partnerships that supported large-scale launches, improved playback quality, and enabled hardware-accelerated implementations across multiple device ecosystems.
In consulting and independent work, I have delivered low-latency AR/VR streaming infrastructure, optimized video pipelines for gaming latency, redesigned scene analysis systems, and deployed automated content moderation solutions. These projects reflect my ability to combine systems engineering, applied AI, and product delivery.
More recently, I have built AI and ML projects such as autonomous job search agents, RAG assistants, and local-first computer vision tools. I continue to focus on agentic AI, LLM workflows, and production-ready ML systems while maintaining my core strength in GPU systems and media engineering.