About this role.
Reka is seeking an experienced GPU Performance Engineer to improve the speed, efficiency, and scalability of large-model training and serving systems. The role combines Python/PyTorch development with low-level CUDA/C++ optimization, GPU profiling, and distributed compute orchestration on platforms such as Slurm or Kubernetes. The engineer will contribute directly to technical decisions and support post-training workflows including reinforcement learning and fine-tuning. This is a remote-first, full-time position available in the US, UK, and Singapore, working with a research-focused foundation-model team.
Role DNA
A quick view of the complexity, pace, ownership and collaboration implied by the job description.
Job Complexity
5/5Pace & Pressure
5/5Autonomy Level
5/5Communication Load
4/5Salary analysis
Estimated compensation compared with the broader US market for similar roles.
Core skills
Skills and capabilities most closely associated with this opportunity.
Sample interview questions
I would first establish a baseline with end-to-end throughput, GPU utilization, memory usage, kernel timing, data-loader metrics, and interconnect activity. I would use tools such as Nsight Systems/Compute and framework profilers to isolate whether the limiting factor is input loading, CPU overhead, kernel inefficiency, memory bandwidth, synchronization, or distributed communication, then validate each optimization against the baseline.
I have optimized workloads through kernel fusion, improved memory coalescing, reduced host-device transfers, mixed precision, better tensor layouts, stream concurrency, and minimizing synchronization points. I also assess occupancy, register pressure, shared-memory usage, and memory-bandwidth limits before changing kernels.
I would use distributed data parallelism or an appropriate sharding strategy, ensure deterministic rank and rendezvous configuration, and tune batch size, gradient accumulation, checkpointing, and data partitioning. I would monitor collective communication, network topology, stragglers, storage throughput, and failure recovery, then use profiling data to tune overlap between computation and communication.
I quantify the improvement using representative workloads and measure throughput, latency, GPU-hours saved, stability, and any effect on model quality. I prioritize changes with meaningful end-to-end gains, manageable maintenance cost, and low regression risk, supported by benchmarks and automated performance tests.
I would examine both training and inference components, since rollout generation, reward evaluation, data movement, and synchronization can dominate costs. I would optimize batching, sequence packing, memory management, distributed scheduling, and checkpoint behavior while preserving reproducibility, numerical stability, and experiment tracking.
We are seeking an experienced GPU Performance Engineer with a strong background in Python and large-scale model training. In this role, you will design and implement improvements to our training infrastructure and directly contribute to technical decisions that optimize performance of our models. You will also work on post-training processes, including reinforcement learning and fine-tuning. Furthermore, you will contribute to improving the efficiency and scalability of our model serving infrastructure.
Ideal Experience
Strong engineering skills with fluency in Python and PyTorch (or other frameworks).
Proven experience implementing and training large deep learning models.
Experience writing and debugging low-level GPU code (CUDA, C++).
Experience scaling up GPU jobs using large-scale compute clusters (e.g., Slurm or Kubernetes).
Demonstrated ability to analyze and optimize the performance of GPU-accelerated workloads, including profiling, identifying bottlenecks, and implementing performance tuning techniques.
Reka’s Mission
Reka’s mission is to build useful multimodal artificial intelligence and use it to empower organizations and businesses. We are a globally distributed foundation model startup, headquartered in the San Francisco Bay Area, California. Embracing a remote-first approach, our team brings together top talent from around the world. Our founding team, along with many of our team members, has contributed to numerous breakthroughs in AI over the past decade.
Why Reka?
An Elite Team: Collaborate with top-tier engineers, researchers, and operators from renowned organizations like Google DeepMind, Facebook AI Research (FAIR), and successful startups, driving innovation in AI technology.
Cutting-edge Infrastructure: Train state-of-the-art models leveraging the latest software and hardware, expanding the frontier of innovation in AI infrastructure development.
Inclusive and Open Culture: Thrive in an open and inclusive work environment that values diverse perspectives and fosters creativity.
Generous Benefits: Enjoy five weeks of paid leave to recharge, comprehensive healthcare benefits (including vision and dental), and additional perks that support your well-being.
Visa Support: We provide visa assistance, including H1B and OPT transfers, for US employees to ensure a smooth transition and support your career with us.
Annual salary information is not provided for this position. Explore salary ranges for similar roles in our Salary Directory ›
This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.









