Suggested rewrite: Led a cross-functional initiative that improved [business outcome] by [measurable result], demonstrating experience relevant to this role...
Research Engineer – Inference
Review the role, location requirements, compensation details, and application process before deciding whether this opportunity fits your next career move.
- Salary
- Undisclosed
- Department
- Software Engineering
- Employment
- Full Time
- Experience
- Open level
- Published
- Apply before
- 7 Nov 2026
- Listing views
- 42
- Application actions
- 1
Make your next move.
Prepare your resume, explore your fit, and draft a cover letter for this opportunity.
The role, at a glance.
ElevenLabs is hiring a Research Engineer to productionize and optimize frontier AI models for low-latency, real-time serving. The role owns the path from research checkpoints through reliable, scalable inference infrastructure used by millions of users. Core work includes profiling bottlenecks and improving latency, throughput, and cost through quantization, distillation, KV-cache optimization, batching, custom kernels, and serving frameworks. Candidates need strong GPU and systems engineering capabilities, particularly with tools such as CUDA, Triton, TensorRT, vLLM, or SGLang. The environment is highly autonomous, fast-moving, and globally remote.
Role DNA
A quick view of the complexity, pace, ownership and collaboration implied by the job description.
Pace & Pressure
5/5Autonomy Level
5/5Communication Load
4/5Salary analysis
Estimated compensation compared with the broader US market for similar roles.
Core skills
Skills and capabilities most closely associated with this opportunity.
Sample interview questions
I would first establish a latency breakdown across request handling, tokenization, queueing, model execution, network transfer, and streaming response generation. I would then prioritize the dominant bottlenecks using targeted profiling, testing approaches such as dynamic batching, KV-cache reuse, optimized attention kernels, reduced precision, and improved scheduling while protecting output quality and reliability.
I would benchmark candidate quantization formats against a representative evaluation set and production-like hardware, measuring quality regression, latency, throughput, memory usage, and operational complexity. I would select the lowest-precision approach that meets defined quality and reliability thresholds, validate edge cases, and deploy progressively with monitoring and rollback capability.
I would monitor time to first token, inter-token latency, p50/p95/p99 end-to-end latency, requests and tokens per second, GPU utilization, memory use, queue depth, error rate, cache hit rate, and cost per request or token. I would segment these metrics by model version, hardware type, request shape, and traffic tier to identify actionable bottlenecks.
A strong answer would explain the initial symptom, the instrumentation used to isolate the constraint, the hypotheses tested, and the quantified impact of the selected fix. It should also cover safeguards such as regression benchmarks, canary deployment, and ongoing observability to ensure the improvement remained durable under real traffic.
I would build a standardized deployment path with reproducible packaging, automated correctness and performance benchmarks, hardware compatibility checks, versioned model artifacts, and clear promotion gates. Researchers should receive fast feedback on latency, throughput, memory, quality, and cost, while production releases use staged rollouts, monitoring, and rapid rollback controls.
About this role.
About ElevenLabs
ElevenLabs is an AI research and product company transforming how we interact with technology.
We launched in January 2023 with the first human-like AI voice model. Today, we serve millions of users and thousands of businesses – from fast-growing startups to large enterprises like Deutsche Telekom and Meta. Our investors are some of the world’s most prominent, including Andreessen Horowitz, ICONIQ Growth and Sequoia. We’ve raised $781M in funding and our last valuation was $22B – multiples of 11, always.
We have expanded from voice into three main platforms:
ElevenAgents enables businesses to deliver seamless and intelligent customer experiences, with the integrations, testing, monitoring, and reliability necessary to deploy voice and chat agents at scale.
ElevenCreative empowers creators and marketers to generate and edit speech, music, image, and video across 70+ languages.
ElevenAPI gives developers access to our leading AI audio foundational models.
Everything we do is the result of the creativity and commitment of our team – builders doing the best work of their lives. We are researchers, engineers, and operators. IOI medalists and ex-founders. If you want to work hard and create lasting positive impact, we want to hear from you.
How we work
High-velocity: Rapid experimentation, lean autonomous teams, and minimal bureaucracy.
Impact not job titles: We don’t have job titles. Instead, it’s about the impact you have. No task is above or beneath you.
AI first: We use AI to move faster with higher-quality results. We do this across the whole company—from engineering to growth to operations.
Excellence everywhere: Everything we do should match the quality of our AI models.
Global team: We prioritize your talent, not your location.
What we offer
Innovative culture: You’ll be part of a generational opportunity to define the trajectory of AI, surrounded by a team pushing the boundaries of what’s possible.
Growth paths: Joining ElevenLabs means joining a dynamic team with countless opportunities to drive impact – beyond your immediate role and responsibilities.
Learning & development: ElevenLabs proactively supports professional development through an annual discretionary stipend.
Social travel: We also provide an annual discretionary stipend to meet up with colleagues each year, however you choose.
Annual company offsite: Each year, we bring the entire team together in a new location – past offsites have included Croatia and Italy.
Co-working: If you’re not located near one of our main hubs, we offer a monthly co-working stipend.
About the role
We are looking for a Research Engineer to join the research team at ElevenLabs, focused on deploying and optimizing our frontier AI models in production. The quality of our models only matters if they can be served fast, reliably, and at scale. You will own the systems that turn research breakthroughs into real-time products used by millions. You will thrive in this role if you enjoy:
Deploying state-of-the-art models to production and owning the path from research checkpoint to serving infrastructure.
Optimizing inference performance across the stack, including latency, throughput, and cost, using techniques such as quantization, distillation, KV-cache optimization, batching strategies, and custom kernels.
Building and tuning high-performance serving systems for real-time, streaming workloads where every millisecond matters.
Creating tooling and infrastructure that lets researchers ship new models to production quickly, safely, and with confidence in their performance characteristics.
Requirements
We do not require any formal certifications or degrees. Instead, we are seeking enthusiastic engineers who can showcase solving impressively hard problems with artifacts such as past projects, designs, or GitHub contributions. Ideally, you bring:
Experience deploying and serving ML models in production, ideally for latency-sensitive or real-time applications.
Strong engineering skills in GPU programming and inference optimization (e.g., CUDA, Triton, TensorRT, or serving frameworks such as vLLM or SGLang).
The capacity to autonomously profile, diagnose, and eliminate bottlenecks across the serving stack, from model architecture to kernels to orchestration, and to build the tooling to measure it.
Location
This role is remote and can be executed globally. If you prefer, you can work from our offices in London, New York, San Francisco, and Warsaw.
#LI-Remote
We are an equal opportunity employer and do not discriminate on the basis of race, religion, national origin, gender, sexual orientation, age, veteran status, disability or other legally protected statuses.
Annual salary information is not provided for this position. Explore salary ranges for similar roles in our Salary Directory ›
This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.
Apply now.
Follow the employer’s application method and review Jobicy’s safety guidance before sharing personal information.
Continue on the employer website
Protect your personal information and never pay to secure an interview or job offer. .
