About this role.
This Staff DevOps Engineer role provides individual-contributor technical leadership for Nextiva's platform engineering practice and multi-cluster Kubernetes environment. The engineer will own platform architecture, GitOps workflows, internal developer-platform capabilities, service networking, reliability standards, and cloud cost optimization. The position requires deep production experience with Kubernetes, AWS, GCP, Terraform or Pulumi, observability tooling, and ML/AI infrastructure including GPU scheduling. It also calls for strong cross-functional influence, architecture review, incident leadership, and mentorship without direct management authority.
Role DNA
A quick view of the complexity, pace, ownership and collaboration implied by the job description.
Job Complexity
5/5Pace & Pressure
5/5Autonomy Level
5/5Communication Load
5/5Salary analysis
Estimated compensation compared with the broader US market for similar roles.
Core skills
Skills and capabilities most closely associated with this opportunity.
Sample interview questions
I would begin by defining workload classes, tenancy boundaries, reliability requirements, and regional placement needs. I would standardize cluster baselines, GitOps delivery, identity, policy enforcement, observability, and disaster-recovery practices, then use dedicated node pools, quotas, taints, and scheduling controls for GPU and ML workloads. The platform should expose self-service interfaces while retaining centrally governed security, cost, and reliability guardrails.
I use Git as the auditable source of truth for environment configuration and deploy changes through reconciliers such as Argo CD or Flux. I establish reusable templates, promotion workflows, policy checks, automated drift detection, and clear rollback procedures. This reduces manual changes, makes deployments repeatable, and gives teams a consistent path from application change to production.
I would first make costs visible by workload, team, cluster, and environment, then combine that data with utilization and SLO signals. Common improvements include right-sizing requests and limits, autoscaling, scheduled nonproduction shutdowns, storage lifecycle controls, reserved-capacity evaluation, and GPU utilization governance. I would prioritize changes through an engineering trade-off process so savings do not introduce unacceptable availability or performance risk.
I would identify the platform's critical user journeys, such as cluster access, deployment reconciliation, ingress, observability, and provisioning, then define measurable SLIs and SLOs for each. Alerting should focus on user impact and error-budget consumption rather than noisy infrastructure symptoms. After incidents, I would lead blameless postmortems, track corrective actions, and use recurring patterns to improve architecture and operational standards.
I would pair clear technical principles with an easier and safer developer experience than the alternatives. That means publishing reference implementations, paved-road templates, documentation, migration support, and metrics that demonstrate reliability, security, and delivery benefits. I would also involve representative teams early in design decisions and use their feedback to build practical standards with broad ownership.
Redefine the future of customer experiences. One conversation at a time.
At Nextiva, we’re reimagining how businesses connect, bringing together customer experience and team collaboration on a single, conversation centric platform. Powered by AI, driven by human innovation.
Our culture is forward thinking, customer obsessed and built on the belief that meaningful connections drive better business outcomes. Whether it’s through our signature Amazing Service®, the technology we create, or the experiences we cultivate, connection is at the core of who we are.
If you’re ready to collaborate with incredible people, make an impact, and help businesses everywhere deliver truly amazing experiences, this is where you belong.
While this role is open to remote candidates across Mexico, team members located within 80 kilometers of our Guadalajara office (Calle Amado Nervo 2200, Jardínes del Sol, 45050 Zapopan, Jal.) are expected to work onsite to support collaboration, speed, and execution.
Nextiva is looking for a Staff DevOps Engineer to lead the design and evolution of our platform engineering practice, with deep technical ownership of Kubernetes infrastructure.
This is an individual-contributor technical leadership role. You’ll help define how engineering teams build, deploy, and operate software by establishing platform architecture, standards, and paved roads that make infrastructure easier to consume without sacrificing reliability, security, or scalability.
Nextiva operates a globally available, redundant microservices architecture and positions reliability, API accessibility, and adaptability as key elements of its technology platform. This role will help evolve the infrastructure behind that environment while supporting the growing needs of engineering and AI/ML workloads.
What you’ll do
- Own the architecture and roadmap for our multi-cluster Kubernetes platform, including scaling, upgrades, and multi-tenancy.
- Establish GitOps-based deployment workflows using technologies such as Argo CD or Flux.
- Design and evolve an internal developer platform that provides engineering teams with self-service access to compute, environments, and observability.
- Architect service mesh, networking, and ingress strategies for reliable and secure communication across services.
- Define platform standards for GPU and ML workload scheduling and resource management on Kubernetes.
- Partner with AI/ML teams on infrastructure requirements for training and inference workloads.
- Drive capacity planning and cost optimization across Kubernetes and cloud infrastructure.
- Set technical direction and review architecture for platform-impacting changes across engineering.
- Define and own platform reliability through SLOs, incident-response leadership, and postmortems.
- Mentor experienced engineers and represent platform engineering in cross-organizational technical decisions.
- Influence teams toward common platform standards without relying on direct reporting authority.
What you bring
- Bachelor’s degree in Computer Science or a related field, or equivalent work experience.
- 8+ years of DevOps, platform, or infrastructure engineering experience.
- 5+ years of hands-on Kubernetes experience in production, including cluster architecture, upgrades, and multi-tenant environments.
- Strong hands-on experience operating cloud infrastructure across AWS and Google Cloud Platform (GCP).
- Strong GitOps experience with Argo CD or Flux.
- Strong Infrastructure-as-Code experience with Terraform or Pulumi.
- Deep understanding of container orchestration, networking, ingress, and service mesh technologies such as Istio, Linkerd, or Cilium.
- Experience with middleware and messaging technologies such as Nginx, Kafka, and Redis at scale.
- Experience with GPU scheduling, node pools, and resource quotas supporting ML/AI training or inference workloads in GKE and EKS.
- Experience designing internal developer platforms or paved-road tooling for engineering organizations.
- Strong Linux, networking, storage, and security fundamentals.
- Experience operating observability platforms such as Prometheus, Grafana, Datadog, or OpenTelemetry at platform scale.
- Excellent communication skills with the ability to influence technical direction across teams.
- Demonstrated ability to understand, calculate, forecast, and optimize cloud and Kubernetes infrastructure costs, including evaluating the cost implications of architecture, capacity, compute, and resource-allocation decisions.
- Experience with capacity planning and resource optimization, balancing performance, reliability, scalability, and infrastructure cost.
Additional experience that will help you succeed
- Kubernetes policy-as-code and security tooling such as OPA/Gatekeeper, Kyverno, or image-scanning technologies.
- AWS, GCP, Azure, CKA, or CKS certifications.
AI literacy
AI matters to this role primarily as an infrastructure and platform workload, not simply as an end-user productivity tool.
- Experience designing infrastructure capable of supporting ML/AI training or inference workloads.
- Understanding of GPU scheduling, resource allocation, node pools, quotas, reliability, and cost considerations for AI workloads.
- Ability to work with AI/ML engineering teams to translate workload requirements into scalable Kubernetes and cloud platform capabilities.
Nextiva DNA (Core Competencies)
Nextiva’s most successful team members share common traits and behaviors:
- Drives Results: Action-oriented problem solvers who quickly bring clarity and simplicity to ambiguity, challenge the status quo, and lead meaningful change; celebrating wins to fuel momentum. They act swiftly and pragmatically, learning and improving as they go.
- Critical Thinker: Data-driven, forward-thinking individuals who identify key drivers, anticipate risks, and deliver clear recommendations. They confidently leverage AI and automation to reduce friction, improve decision-making, and focus on higher-value work.
- Right Attitude: Collaborative, competitive, and resilient team players who jump in to solve tough problems, learn from setbacks, and foster a culture of service, respect, and care for customers and teammates.
Total Rewards
Our Total Rewards offerings are designed to allow Nexties to take care of themselves and their families so they can be their best, in and out of the office.
Our compensation packages are tailored to each role and candidate’s qualifications. We consider a wide range of factors, including skills, experience, training, and certifications, when determining compensation. We aim to offer competitive salaries or wages that reflect the value you bring to our team. Depending on the position, compensation may include base salary, hourly wages, incentives, or bonuses.
- Health 🍏– Major Health insurance for you, your legal partner, and children under 25 years, including Vision and Dental coverage
- Insurance 💼– Life Insurance (24 times your monthly salary)
- Work-Life Balance ⚖️ – Newly hired full-time employees at Nextiva receive 10 personal days before their first anniversary, 12 vacation days on their first anniversary, and 5 personal days annually thereafter in addition to vacation time
- Financial Security 💰– Enjoy a 30-day Christmas bonus, 50% vacation premium, company-matched food vouchers (1 UMA/month), and a 13% matched savings fund (capped at 1.3x annual UMA)
- Wellness 🤸– Employee Assistance Program and comprehensive wellness initiatives
- Growth 🌱– Access to ongoing learning and development opportunities and career advancement
At Nextiva, we’re committed to supporting our employees’ health, well-being, and professional growth. Join us and build a rewarding career!
#LI-AL1 #LI-REMOTE
Founded in 2008, Nextiva has grown into a global leader trusted by over 100,000 businesses and 1M+ users worldwide. Headquartered in Scottsdale, Arizona, and with teams across the globe, we’re the future of customer experience and team collaboration through our AI-powered, conversation-centric platform.
Want to see what life at Nextiva is all about? Connect with us on Instagram, Instagram MX, YouTube, LinkedIn, and the Nextiva Blog.
Annual salary information is not provided for this position. Explore salary ranges for similar roles in our Salary Directory ›
This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.









