About this role.
NetBox Labs is hiring a senior software engineer focused on the DevOps and platform engineering foundations of its on-premises NetBox Enterprise product. The role designs Kubernetes-based high-availability deployments, maintains a Kubernetes operator, and builds lifecycle-management tooling, APIs, and operational consoles. It requires strong Python, Linux, CI/CD, GitOps, Terraform, Kubernetes, Helm, and software supply-chain security expertise, with Rust or Go as valued additions. The engineer will also manage structured six-week releases and support customers and internal teams in diagnosing complex deployment issues. This is a US-remote, full-time senior role with substantial ownership of production-grade, self-managed infrastructure.
Role DNA
A quick view of the complexity, pace, ownership and collaboration implied by the job description.
Job Complexity
5/5Pace & Pressure
4/5Autonomy Level
5/5Communication Load
4/5Salary analysis
Estimated compensation compared with the broader US market for similar roles.
Core skills
Skills and capabilities most closely associated with this opportunity.
Sample interview questions
I would define a clear custom resource API that represents the desired product state, then implement reconciliation loops that are idempotent, observable, and safe to retry. The operator would validate prerequisites, manage dependencies in a deliberate order, use status conditions to expose progress and failures, and support versioned upgrade paths with preflight checks and rollback or recovery guidance. I would also add metrics, structured logs, and integration tests covering interrupted upgrades, degraded clusters, and configuration drift.
I would begin by identifying every external dependency, including container images, Helm charts, packages, licenses, and update metadata. I would provide a versioned offline bundle, documented image mirroring process, integrity checks, and a preflight utility that validates cluster capacity, DNS, storage, networking, and required images before installation. The deployment workflow should be reproducible and supported by diagnostic collection tools that customers can safely share during escalation.
I would separate fast validation, build, test, security scanning, artifact publishing, and release stages with explicit promotion criteria. Builds should be reproducible, dependencies pinned, secrets scoped minimally, artifacts signed or attested where possible, and vulnerability scanning enforced according to defined severity policies. I would use reusable workflows, caching carefully, protected environments, and clear failure reporting so engineers can quickly distinguish code, infrastructure, and dependency failures.
I would first establish impact, version, topology, recent changes, and the exact failed lifecycle stage. I would collect operator events, pod status, logs, resource definitions, network and storage signals, and relevant diagnostics through a structured support bundle while protecting customer-sensitive data. I would form and test hypotheses methodically, communicate status and next actions clearly, implement a safe mitigation, and then convert the root cause into improved validation, documentation, or automated remediation.
I would use a release plan with clear entry criteria, feature flags where appropriate, versioning rules, and early integration testing for upstream changes. During the build period, I would continuously assess risk and keep release artifacts and documentation current; during stabilization, I would prioritize regression testing, upgrade testing, vulnerability remediation, and customer-impacting defects. Decisions to defer work should be transparent and based on operational readiness rather than feature completeness alone.
NetBox Labs seeks a Senior Software Engineer with strong DevOps experience to build and drive NetBox Enterprise, our on-premises product suite. You’ll own installation, configuration, and lifecycle management for our customers, tackle complex technical challenges, proactively address deployment issues unique to on-prem environments, and collaborate closely with high-performing, cross-functional teams.
In this role you will:
Design, architect, and implement Kubernetes-based, high-availability (HA) on-premises solutions, including control plane applications, telemetry systems, air-gapped installations, and appliance offerings
Extend and maintain our Kubernetes operator, the core of how NetBox Enterprise installs, upgrades, and heals itself in customer environments
Develop and maintain our operational management console and associated tools, ensuring a seamless user experience for lifecycle management
Architect and implement robust multi-stage CI/CD pipelines using GitHub Actions and complementary DevOps technologies
Drive releases on our six-week cadence, four weeks of build followed by two weeks of stabilization, including versioning strategy and structured release management
Write Rust and Python across the stack, from operator internals to internal tooling and APIs
Build internal tooling and APIs to facilitate integration testing by upstream application teams, validating changes ahead of inclusion in on-premises releases
Develop comprehensive documentation and establish best practice guidelines for deployments
Take a weekly turn as Engineering Guardian, fielding questions and escalations from across the organization and working directly with Customer Success and customers to diagnose and resolve deployment and operational issues
Requirements:
5+ years in software engineering, platform engineering, or SRE, with proven experience writing robust, maintainable code
Demonstrated expertise with Kubernetes, Helm charts, and deployment automation, including complex multi-cluster and multi-region deployments
Hands-on experience building inside an AI-augmented development harness, including Claude Code and the workflows that make agentic tooling reliable
Strong Python skills for tooling, APIs, and automation
Hands-on experience with CI/CD systems (GitHub Actions), GitOps delivery (ArgoCD or FluxCD), and infrastructure-as-code (Terraform)
Strong knowledge of Linux systems, including system administration, troubleshooting, and networking
Solid understanding of software supply chain security, including vulnerability scanning and CVE management, with tools such as Trivy, Dependabot, and GitHub code scanning
Comfort working directly with customers to diagnose deployment and operational issues
Nice to Haves:
Rust or Golang programming experience.
Experience with virtual appliances and air-gapped installations
Experience deploying applications using Replicated
Familiarity with networking concepts or distributed system architecture
3+ years of experience in a B2B software startup or high-growth organization
Open-source contributions or project involvement
Join us at NetBox Labs to shape the future of on-premise infrastructure intelligence and innovation.
About NetBox Labs:
NetBox Labs helps companies build and manage complex networks. We help customers accelerate network automation by delivering open, composable products and supporting the network automation community.
NetBox Labs is the commercial steward of open source NetBox, the world’s most popular network source of truth, and Orb, the next-generation open source network observability platform. Our products include NetBox Enterprise, a fully supported self-managed NetBox with advanced features, and NetBox Cloud, a secure, scalable, and reliable SaaS edition of NetBox.
NetBox powers thousands of companies, and NetBox Labs is backed by investment from Notable Capital (formerly GGV), Grafana Labs CEO Raj Dutt, Flybridge, IBM, Salesforce Ventures, and Mango Capital.
Our culture and values:
We own and solve problems with high attention to detail.
Our open source contributors, users, customers & team are all part of our community. When our community wins, we win.
We prioritize simplicity and think twice before adding complexity
Clear communication helps keep our team aligned and collaborating smoothly.
NetBox Labs is proud to be an equal opportunity employer. We believe diverse teams build better software, and we welcome applicants of every race, color, religion, gender identity, sexual orientation, national origin, age, disability, and veteran status. If you need accommodation at any point in the process, just let us know.
This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.









