About Me
I am a senior software engineer and AI engineer focused on building real-time, production-grade AI systems. My work centers on low-latency voice agents, multimodal retrieval, RAG pipelines, and distributed inference across healthcare, real estate, and enterprise platforms.
I have extensive experience architecting end-to-end systems that combine audio ingestion, embeddings, vector search, LLM orchestration, and scalable backend services. I enjoy solving hard problems around latency, reliability, observability, and evaluation, especially when the systems need to serve large user bases in production.
I have led development across Python, FastAPI, Django, Node.js, Go, React, Next.js, and cloud infrastructure on AWS and GCP. I have also worked deeply with tools and frameworks such as Whisper, Deepgram, ElevenLabs, LiveKit, WebRTC, LangGraph, LlamaIndex, FAISS, pgvector, Vertex AI, and Triton.
My background includes building clinician-facing applications, conversational search systems, multimodal AI pipelines, and high-throughput APIs. I have delivered solutions that improved grounding accuracy, reduced hallucinations, increased retrieval performance, and supported millions of users with strict latency requirements.
I value cross-functional collaboration and have worked closely with researchers, clinicians, product managers, and engineers to ship impactful AI features. I also enjoy mentoring engineers, improving engineering standards, and creating reproducible deployment and evaluation workflows.
I am open to roles where I can continue building advanced AI products, especially in areas involving voice AI, retrieval systems, applied machine learning infrastructure, and full-stack product engineering.
Notice Period: 1 week
Skills
PythonAWSJavaReactDockerAzureCI CDTypeScriptGCPTerraformSoftware EngineeringPostgreSQLGoMicroservicesData EngineeringETLRedisREST APIsAngularObservabilitySpring BootNext.jsDjangoBigQueryReduxTwilioCloud ArchitectureAccessibilityGovernancePerformance TuningfastAPIWebSocketsWebRTCAI EngineeringEvent Driven SystemsText To SpeechC#OnnxVector SearchFaissDeepgramTriton Inference ServerWhisperPgvectorHybrid SearchLiveKitRAGSynthetic Data Generation
Tech Stack & Tools
Application Utilities
Data Stores
Design
Experience
I architected and shipped a real-time AI voice agent platform using Python, FastAPI, Django, LiveKit, WebRTC, and Twilio on GCP. The system was designed for scalable, low-latency inference and bidirectional audio streaming.
I integrated Gemini 2.5 Pro, Vertex AI embeddings, Whisper, Deepgram STT, and ElevenLabs TTS into multimodal RAG and streaming inference pipelines. I also built interrupt-driven agent logic, turn-taking, barge-in detection, and context-aware state management for natural conversational flow.
I developed audio ingestion, preprocessing, batch inference, and evaluation workflows, and collaborated with clinicians, researchers, and product teams to improve accuracy and reliability. I also mentored junior engineers and helped drive cross-functional delivery of new AI capabilities.
I delivered full-stack AI applications using React, Next.js, TypeScript, Python, FastAPI, Django, Node.js, Redis, PostgreSQL, AWS, and GCP. My work supported internal AI platforms and large-scale user-facing retrieval systems.
I built and productionized multimodal, multilingual embedding-based retrieval models serving 20M+ users with low-latency and high-QPS requirements. I also designed cross-lingual contrastive learning pipelines, synthetic data generation workflows, and conversational search systems with query rewriting and RAG grounding.
I led observability and governance efforts using LangSmith, Langfuse, PostgreSQL, and custom telemetry dashboards, and I optimized FAISS indexes, Redis-backed serving, ONNX deployment, quantization, and Triton inference.
I led development of a clinical decision support platform that enabled natural language search across treatment guidance and patient insights. The work improved clinicians' ability to find relevant information quickly.
I built a large-scale document retrieval system indexing over 1M clinical documents using hybrid search, semantic ranking, vector retrieval, and contextual grounding. I also developed workflow automation services for retrieval, summarization, and compliance-driven processing.
I created scalable backend services and clinician-facing applications using Python, FastAPI, Go, React, Angular, TypeScript, and Redux, improving performance, accessibility, and workflow efficiency.
I built and optimized platform services on AWS using Node.js, Nest.js, Terraform, Java, Spring Boot, and .NET Core. My work improved deployment reliability, reduced downtime, and increased API performance.
I also developed responsive frontend applications with React, Angular, and Next.js, created reusable UI component libraries, and integrated applications with backend services through REST, GraphQL, and Apollo.
In addition, I helped design media-processing pipelines with Delta Lake, AWS Lambda, and FFmpeg, and I optimized Node.js services and AWS compute resources to reduce costs and improve rendering speed.
Education
Bachelor of Science, Computer Science
I earned a Bachelor of Science in Computer Science.
This professional hasn’t added portfolio projects yet.
This professional hasn’t listed any services yet.