[![Image]() Meet Jobicy Copilot — free AI autofill for job applications + remote job alerts ›](#)   [All remote jobs](https://jobicy.com/jobs.md)Open role[![Reka logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/4c8a0366-221.png)](https://jobicy.com/company/reka.md)Remote opportunity at[Reka](https://jobicy.com/company/reka.md)

# Member of Technical Staff (GPU Performance Engineer)

Review the role, location requirements, compensation details, and application process before deciding whether this opportunity fits your next career move.

[Apply for this job](#job-application)[View company](https://jobicy.com/company/reka.md)Share7 Sep 2026Published32Listing views2Application actions7 Oct 2026Apply before  Opportunity details

## About this role.

AI SummaryReka is seeking an experienced GPU Performance Engineer to improve the speed, efficiency, and scalability of large-model training and serving systems. The role combines Python/PyTorch development with low-level CUDA/C++ optimization, GPU profiling, and distributed compute orchestration on platforms such as Slurm or Kubernetes. The engineer will contribute directly to technical decisions and support post-training workflows including reinforcement learning and fine-tuning. This is a remote-first, full-time position available in the US, UK, and Singapore, working with a research-focused foundation-model team.

## Role DNA

A quick view of the complexity, pace, ownership and collaboration implied by the job description.

### Job Complexity

5/5EasyHard

### Pace & Pressure

5/5RelaxedFast-paced

### Autonomy Level

5/5GuidedFull ownership

### Communication Load

4/5IndependentCollaborative

AI insightThis is a highly specialized systems and ML infrastructure role requiring both deep-learning training expertise and low-level GPU performance engineering. Success depends on independently diagnosing complex distributed performance bottlenecks and translating findings into production-grade improvements.

## Salary analysis

Estimated compensation compared with the broader US market for similar roles.

Estimated job medianMarket rate$220,000US market range$170k–$300k0$330k

AI insightNo actual salary or compensation range is disclosed in the posting. Estimated US annual base-pay market range for an experienced GPU performance engineer focused on large-scale AI training infrastructure is $170,000–$300,000 USD, with an estimated midpoint of $220,000; equity and bonuses may materially increase total compensation at an AI startup.

## Core skills

Skills and capabilities most closely associated with this opportunity.

[Python](https://jobicy.com/jobs?search_keywords=Python.md)[PyTorch](https://jobicy.com/jobs?search_keywords=PyTorch.md)[CUDA](https://jobicy.com/jobs?search_keywords=CUDA.md)[C++](https://jobicy.com/jobs?search_keywords=C%2B%2B.md)[GPU performance optimization](https://jobicy.com/jobs?search_keywords=GPU%20performance%20optimization.md)[Deep learning](https://jobicy.com/jobs?search_keywords=Deep%20learning.md)[Distributed training](https://jobicy.com/jobs?search_keywords=Distributed%20training.md)[Kubernetes](https://jobicy.com/jobs?search_keywords=Kubernetes.md)[Slurm](https://jobicy.com/jobs?search_keywords=Slurm.md)[Reinforcement learning](https://jobicy.com/jobs?search_keywords=Reinforcement%20learning.md)

Sample interview questionsDescribe how you would investigate a large-model training job that is underutilizing GPUs.I would first establish a baseline with end-to-end throughput, GPU utilization, memory usage, kernel timing, data-loader metrics, and interconnect activity. I would use tools such as Nsight Systems/Compute and framework profilers to isolate whether the limiting factor is input loading, CPU overhead, kernel inefficiency, memory bandwidth, synchronization, or distributed communication, then validate each optimization against the baseline.

What techniques have you used to optimize CUDA or GPU-accelerated workloads?

I have optimized workloads through kernel fusion, improved memory coalescing, reduced host-device transfers, mixed precision, better tensor layouts, stream concurrency, and minimizing synchronization points. I also assess occupancy, register pressure, shared-memory usage, and memory-bandwidth limits before changing kernels.

How would you scale PyTorch training from a few GPUs to a large Slurm or Kubernetes cluster?

I would use distributed data parallelism or an appropriate sharding strategy, ensure deterministic rank and rendezvous configuration, and tune batch size, gradient accumulation, checkpointing, and data partitioning. I would monitor collective communication, network topology, stragglers, storage throughput, and failure recovery, then use profiling data to tune overlap between computation and communication.

How do you decide whether a performance optimization is worth deploying?

I quantify the improvement using representative workloads and measure throughput, latency, GPU-hours saved, stability, and any effect on model quality. I prioritize changes with meaningful end-to-end gains, manageable maintenance cost, and low regression risk, supported by benchmarks and automated performance tests.

What considerations are important when optimizing post-training workflows such as reinforcement learning or fine-tuning?

I would examine both training and inference components, since rollout generation, reward evaluation, data movement, and synchronization can dominate costs. I would optimize batching, sequence packing, memory management, distributed scheduling, and checkpoint behavior while preserving reproducibility, numerical stability, and experiment tracking.

We are seeking an experienced GPU Performance Engineer with a strong background in Python and large-scale model training. In this role, you will design and implement improvements to our training infrastructure and directly contribute to technical decisions that optimize performance of our models. You will also work on post-training processes, including reinforcement learning and fine-tuning. Furthermore, you will contribute to improving the efficiency and scalability of our model serving infrastructure.

Ideal Experience

*

Strong engineering skills with fluency in Python and PyTorch (or other frameworks).

*

Proven experience implementing and training large deep learning models.

*

Experience writing and debugging low-level GPU code (CUDA, C++).

*

Experience scaling up GPU jobs using large-scale compute clusters (e.g., Slurm or Kubernetes).

*

Demonstrated ability to analyze and optimize the performance of GPU-accelerated workloads, including profiling, identifying bottlenecks, and implementing performance tuning techniques.

Reka’s Mission
Reka’s mission is to build useful multimodal artificial intelligence and use it to empower organizations and businesses. We are a globally distributed foundation model startup, headquartered in the San Francisco Bay Area, California. Embracing a remote-first approach, our team brings together top talent from around the world. Our founding team, along with many of our team members, has contributed to numerous breakthroughs in AI over the past decade.

Why Reka?

*

An Elite Team: Collaborate with top-tier engineers, researchers, and operators from renowned organizations like Google DeepMind, Facebook AI Research (FAIR), and successful startups, driving innovation in AI technology.

*

Cutting-edge Infrastructure: Train state-of-the-art models leveraging the latest software and hardware, expanding the frontier of innovation in AI infrastructure development.

*

Inclusive and Open Culture: Thrive in an open and inclusive work environment that values diverse perspectives and fosters creativity.

*

Generous Benefits: Enjoy five weeks of paid leave to recharge, comprehensive healthcare benefits (including vision and dental), and additional perks that support your well-being.

*

Visa Support: We provide visa assistance, including H1B and OPT transfers, for US employees to ensure a smooth transition and support your career with us.

Show more

[Apply now >](https://jobicy.com/jobs/152767-member-of-technical-staff-gpu-performance-engineer.md)

>  Annual salary information is not provided for this position. Explore salary ranges for similar roles in our [Salary Directory ›](https://jobicy.com/salaries.md)

*

![Upload CV](data:image/svg+xml;base64,PHN2ZyB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciIHdpZHRoPSI2NSIgaGVpZ2h0PSI2NSIgZmlsbD0ibm9uZSIgeG1sbnM6dj0iaHR0cHM6Ly92ZWN0YS5pby9uYW5vIj48ZyBjbGlwLXBhdGg9InVybCgjQSkiPjxwYXRoIGQ9Ik0wIDBINjVWNjVIMFYwWiIgZmlsbD0iIzAyOWFlYiIvPjxnIGZpbGw9IiNmZmYiIHN0cm9rZT0iI2ZmZiIgc3Ryb2tlLXdpZHRoPSIyIj48cGF0aCBkPSJNMzMuMDQ5IDE1LjQ1NGExLjQzIDEuNDMgMCAwIDAtMi4wOTcgMGwtNy41NzkgOC4xNDdhMS4zOCAxLjM4IDAgMCAwIC4wOSAxLjk3MyAxLjQ0IDEuNDQgMCAwIDAgMi4wMDgtLjA4OGw1LjEwOS01LjQ5MnYyMC42MWExLjQxIDEuNDEgMCAwIDAgMS40MjEgMS4zOTdjLjc4NSAwIDEuNDIxLS42MjUgMS40MjEtMS4zOTd2LTIwLjYxbDUuMTA5IDUuNDkyYTEuNDQgMS40NCAwIDAgMCAyLjAwOC4wODggMS4zOCAxLjM4IDAgMCAwIC4wOS0xLjk3M2wtNy41NzktOC4xNDZ6TTE2Ljc2OSAzOC40YzAtLjc3My0uNjItMS40LTEuMzg1LTEuNFMxNCAzNy42MjcgMTQgMzguNHYuMTAybC4yMTUgNi4yMjljLjIyMyAxLjY4LjcwMSAzLjA5NSAxLjgxMyA0LjIxOHMyLjUxIDEuNjA3IDQuMTcyIDEuODMzYzEuNi4yMTggMy42MzYuMjE4IDYuMTYuMjE4aDExLjI4bDYuMTYtLjIxOGMxLjY2Mi0uMjI2IDMuMDYxLS43MDkgNC4xNzItMS44MzNzMS41ODktMi41MzggMS44MTMtNC4yMThDNTAgNDMuMTEzIDUwIDQxLjA1NSA1MCAzOC41MDNWMzguNGMwLS43NzMtLjYyLTEuNC0xLjM4NS0xLjRzLTEuMzg1LjYyNy0xLjM4NSAxLjRsLS4xOSA1Ljk1OGMtLjE4MiAxLjM3LS41MTUgMi4wOTUtMS4wMjYgMi42MTJzLTEuMjI4Ljg1My0yLjU4MyAxLjAzOGMtMS4zOTUuMTktMy4yNDMuMTkzLTUuODkzLjE5M0gyNi40NjJjLTIuNjUgMC00LjQ5OC0uMDAzLTUuODkzLS4xOTMtMS4zNTUtLjE4NC0yLjA3Mi0uNTIxLTIuNTgzLTEuMDM4cy0uODQ0LTEuMjQyLTEuMDI2LTIuNjEyYy0uMTg3LTEuNDEtLjE5MS0zLjI3OS0uMTkxLTUuOTU4eiIvPjwvZz48L2c+PGRlZnM+PGNsaXBQYXRoIGlkPSJBIj48cGF0aCBmaWxsPSIjZmZmIiBkPSJNMCAwaDY1djY1SDB6Ii8+PC9jbGlwUGF0aD48L2RlZnM+PC9zdmc+)

### Upload your resume now

To unlock remote work opportunities and be discovered by global employers.

This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.

Next step

## Apply now.

Follow the employer’s application method and review Jobicy’s safety guidance before sharing personal information.

Keep exploring

## Related remote jobs.

Matched by job category10 related opportunities[Software Engineering](https://jobicy.com/categories/engineering.md) [Browse all jobs](https://jobicy.com/jobs.md)
*
![Reka logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/4c8a0366-221.png)
Reka  Sep 7

### [Member of Technical Staff (Applied AI)](https://jobicy.com/jobs/152763-member-of-technical-staff-applied-ai.md)

As a Member of Technical Staff on Applied AI, you will: Productionize frontier AI models to solve complex real-world problems. Collaborate closely with researchers and other teammates on the latest…

*
![Gremlin logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2020/10/WRILS-201027082756-840509.jpg)
Gremlin  Sep 7

### [Senior Backend Engineer](https://jobicy.com/jobs/152754-senior-backend-engineer-4.md)

Today’s complex, fast-paced systems have become a minefield of reliability risks—any of which could cause an outage that costs millions and destroys customer confidence. That’s why high-availability teams use the…

*
![Playson logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/6f266d9c-221-1.png)
Playson  Sep 7

### [Senior Backend Engineer (Core Team)](https://jobicy.com/jobs/152752-senior-backend-engineer-core-team.md)

About the Role We’re looking for a Senior Backend Developer who thrives in a deeply hands-on environment and enjoys being close to the code on a daily basis. This role…

*
![Klaviyo logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/1ba9baff-221.jpeg)
Klaviyo  Sep 7

### [Software Engineer II (Mobile Developer – Android/Kotlin) – Mobile App Growth](https://jobicy.com/jobs/152750-software-engineer-ii-mobile-developer-android-kotlin-mobile-app-growth.md)

At Klaviyo, we value the unique backgrounds, experiences and perspectives each Klaviyo (we call ourselves Klaviyos) brings to our workplace each and every day. We believe everyone deserves a fair…

*
![Testlio logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/11024036-221.png)
Testlio  Sep 7

### [Test a Streaming Platform in France](https://jobicy.com/jobs/152737-test-a-streaming-platform-in-france.md)

Have you ever wished an app worked a little better? At Testlio, we help some of the world’s leading technology and consumer brands improve their apps and digital products by…

*
![SeatGeek logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/72d8f3a6-221.jpeg)
SeatGeek  Sep 7

### [Engineering Manager, SupportX](https://jobicy.com/jobs/152712-engineering-manager-supportx.md)

SeatGeek believes live events are powerful experiences that unite humans. With our technological savvy and fan-first attitude we’re simplifying and modernizing the ticketing industry. The SupportX team is dedicated to…

*
![Fleetio logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2023/07/e571b83c37e0ed5371754d319be16e76.jpeg)
Fleetio  Sep 7

### [Engineering Manager, Payments](https://jobicy.com/jobs/152709-engineering-manager-payments.md)

A little about us…Fleetio is a modern software platform that helps thousands of organizations worldwide manage their fleet operations. Transportation technology is a hot market, and we’re leading the charge…

*
![Dataiku logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2026/06/5a9d57bf-221.webp)
Dataiku  Sep 7

### [Fullstack Software Engineer – Core](https://jobicy.com/jobs/152707-fullstack-software-engineer-core.md)

Dataiku is the Platform for AI Success, the enterprise orchestration layer for building, deploying, and governing AI. In a single environment, teams design and operate analytics, machine learning, and AI…

*
![Maze logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/f53dfbf3-221.png)
Maze  Sep 7

### [Product Engineer (Full Stack)](https://jobicy.com/jobs/151461-product-engineer-full-stack.md)

Summary of the Role: As Product Engineer (Full Stack) at Maze, you’ll be the technical force behind our customer-facing product experience, owning the complete stack from UI to API while…

*
![Databricks logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2021/12/4e3f864ba9cf3c62d471e1c62414d098.jpg)
Databricks  Sep 7

### [AI Engineer – FDE (Forward Deployed Engineer) – U.S. Federal Sector](https://jobicy.com/jobs/152699-ai-engineer-fde-forward-deployed-engineer-u-s-federal-sector.md)

PLEASE NOTE: Due to federal contract requirements and client site access obligations, U.S. citizenship and eligibility for a U.S. government secret clearance are required to access classified information. The position…