I am an AI / LLM Engineer focused on building LLM features and shipping them to production. My work centers on RAG, semantic retrieval, cost-aware model routing, output control, and self-hosted or local inference.
I build systems that are designed to be reliable end-to-end. In my own production projects, I have implemented intent classification, prompt caching, persistent memory, and regression-tested routing between models such as Haiku and Sonnet.
I have six years of production background in support, QA, and monitoring, which gives me a strong operational mindset. That experience helps me think carefully about reliability, observability, and how AI systems behave in real-world environments.
I have also worked on client-facing AI integration projects, where I owned the LLM layer across multiple products. This included retrieval and embedding tuning, prompt design, model selection, API integration, and monitoring with tools like Langfuse and Grafana.
My recent work includes Telegram AI agents, job-search automation, and creative-writing routers, all running in production on infrastructure I manage myself. I am comfortable working across Python, FastAPI, PostgreSQL, React, Docker, and Linux, and I enjoy building practical systems that solve real product problems.
I am looking for an AI / LLM Engineer role where I can own LLM behavior and reliability end-to-end in a fast-moving product team. I am especially interested in RAG, routing, output control, evaluation pipelines, and the surrounding systems that make AI products dependable.