About this role.
Temporal is seeking a senior, customer-facing technical engineer to help developers deploy, operate, and scale distributed Temporal environments. The role centers on diagnosing complex production and infrastructure issues across Kubernetes, cloud platforms, networking, observability, and automation tooling. The successful candidate will work directly with developers and engineering teams while translating recurring problems into platform improvements and self-service solutions. It requires at least six years of development experience, strong cloud-native operations expertise, and demonstrated customer-facing communication. Eligibility is limited to candidates residing in the United States Pacific time zone.
Role DNA
A quick view of the complexity, pace, ownership and collaboration implied by the job description.
Job Complexity
5/5Pace & Pressure
4/5Autonomy Level
4/5Communication Load
5/5Salary analysis
Estimated compensation compared with the broader US market for similar roles.
Core skills
Skills and capabilities most closely associated with this opportunity.
Sample interview questions
I would first define the affected workflows, time window, and user impact, then correlate application, Temporal, Kubernetes, and cloud metrics. I would inspect Prometheus and Grafana data for saturation, error rates, queue depth, pod restarts, network latency, and resource throttling, while reviewing traces and logs for the affected path. I would validate likely causes through targeted tests and document both the remediation and preventive monitoring.
I would assess the current architecture, operational risks, dependencies, and desired environments before defining reusable infrastructure modules. I would build an incremental plan using Terraform or the customer's preferred IaC tool, including versioned configuration, secrets management, networking, monitoring, and rollback procedures. I would validate the approach in a non-production environment, provide clear runbooks, and enable the customer team to maintain the solution independently.
I prioritize restoring safe service first by establishing impact, urgency, and a clear incident owner. During and after resolution, I capture evidence about root cause, workarounds, and recurring friction, then translate it into actionable feedback for product and engineering teams. This approach protects the customer experience while ensuring the issue can lead to better tooling, documentation, or platform behavior.
I would monitor availability, latency, error rates, throughput, queue depth, CPU and memory utilization, pod health, restart counts, autoscaling behavior, and storage and network performance. I would also track service-level indicators that reflect the user experience rather than relying only on infrastructure health. Alerts should be actionable, tied to meaningful thresholds, and supported by dashboards and runbooks.
I tailor the level of detail to the audience while keeping the facts consistent. For developers, I share the technical failure mode, evidence, mitigation, and implementation steps; for non-specialists, I explain customer impact, scope, risk, recovery status, and next actions in plain language. In both cases, I avoid speculation, communicate ownership and timelines, and provide a concise written follow-up.
Summary
The Senior Developer Success Engineer – West will be the frontline technical expert for our developer community. You will help users deploy and scale Temporal in cloud-native environments. You will also troubleshoot complex infrastructure issues, optimize performance, and develop automation solutions. This role is ideal for someone who thrives on solving distributed systems challenges and improving the developer experience.
What You’ll Do
Be a keen learner:
At Temporal, you’ll work with cloud-native, highly scalable infrastructure spanning AWS, GCP, Kubernetes, and microservices. You’ll gain deep expertise in container orchestration, networking, and observability while learning from complex, real-world customer use cases.
Our stack includes Go, Python, and Java, providing continuous opportunities to hone your programming skills in infrastructure automation, resilience engineering, and performance tuning.
Be a passionate problem solver:
If you enjoy tackling scalability, reliability, and troubleshooting challenges in distributed systems, you’ll thrive in this role.
As a Senior Developer Success Engineer, you’ll work directly with developers to debug complex infrastructure issues, optimize cloud performance, and enhance reliability for Temporal users.
You’ll develop observability solutions (Grafana, Prometheus), improve networking (load balancing, DNS, ingress/egress), and automate infrastructure operations (Terraform, IaC) to help customers run Temporal efficiently at scale.
Once ramped up, we expect you to independently drive technical solutions, whether debugging complex production issues or designing infrastructure best practices. Don’t worry—we have seasoned engineers and mentors to support you along the way!
Be a great communicator:
As a Senior Developer Success Engineer you will engage directly with developers, engineering teams, and product teams to understand infrastructure challenges and provide solutions that enhance scalability, performance, and reliability.
Your insights will influence platform improvements, from enhancing observability tooling to developing self-service infrastructure solutions that simplify troubleshooting (e.g., building diagnostic tools similar to Twilio’s Network Test).
You’ll serve as a bridge between developers and infrastructure, ensuring that reliability, performance, and developer experience remain top priorities as Temporal scales.
What You’ll Bring
Reside within the Pacific time zone (United States)
6+ years of experience as a developer, preferably fluent in one or more of the following languages: Python, Java, Golang, TypeScript.
Experience with deployment and managing medium to large-scale architectures (e.g., Kubernetes or Docker).
Experience with monitoring tools such as Prometheus and Grafana and troubleshooting performance and availability issues.
Minimum of one year experience in an internal or external customer-facing role.
Passion for helping others regardless of who they are or how they act.
Experience working with or as part of remote teams.
Strong written and verbal communication skills.
Seek to understand first, lead with data, and rely on facts.
Nice to Have
Previous experience in customer-facing positions such as a professional services consultant, solutions architect, customer engineer, etc.
Experience with security certificate management and implementation.
Ability to understand use cases and translate them into Temporal design decisions and architecture best practices.
Technologies: EKS, GKE, Kubernetes, Prometheus, Grafana, OpenTracing, Terraform/Ansible/CDK.
Compensation
The estimated pay range for this role is $140,000 – $180,000 based on experience and location.
Additionally, this role is eligible to participate in Temporal’s equity plan.
Temporal Technologies is an Equal Opportunity Employer. Temporal Technologies does not discriminate on the basis of race, religion, color, sex, gender identity, sexual orientation, age, non-disqualifying physical or mental disability, national origin, veteran status, or any other basis covered by appropriate law. All employment is decided on the basis of qualifications, merit, and business need. We embrace and celebrate differences and diversity.
Temporal is committed to providing access, equal opportunity, and reasonable accommodation for individuals with disabilities in employment, its services, programs, and activities. If you need to request a reasonable accommodation, please let your Recruiter know so we can assist.
This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.







