About this role.
Nebius is seeking a Senior Backend Software Engineer to join their Observability Platform team, building core services for logs, metrics, traces, and alerting at scale. The role requires 5+ years of experience, strong Golang skills, and expertise in distributed backend systems. Candidates with experience in open-source observability tools like Prometheus, Grafana, or ClickHouse are preferred. The position offers competitive compensation, career growth, and the opportunity to work on impactful AI projects in a fast-paced, collaborative environment. Nebius is a Nasdaq-listed company with a global presence and a focus on AI cloud infrastructure.
Role DNA
A quick view of the complexity, pace, ownership and collaboration implied by the job description.
Job Complexity
4/5Pace & Pressure
4/5Autonomy Level
5/5Communication Load
4/5Salary analysis
Estimated compensation compared with the broader US market for similar roles.
Core skills
Skills and capabilities most closely associated with this opportunity.
Cover letter sample
Dear Hiring Manager,
I am writing to express my enthusiasm for the Senior Backend Software Engineer (Observability) role at Nebius. With over 6 years of experience in building distributed backend systems and a strong proficiency in Golang, I am excited about the opportunity to contribute to your observability platform. My background includes designing high-volume telemetry ingestion pipelines and working with tools like Prometheus and Grafana, which aligns perfectly with your needs.
I thrive in fast-paced, collaborative environments and am passionate about creating reliable systems that empower engineers. At my previous role, I led the development of a scalable alerting system that reduced incident response time by 30%. I am eager to bring my expertise in observability and AI-assisted troubleshooting to Nebius and help shape the future of AI infrastructure.
Thank you for considering my application. I look forward to the possibility of discussing how I can contribute to your team.
Sincerely,
[Your Name]
Sample interview questions
I worked on a telemetry ingestion system handling millions of events per second. By implementing a multi-tiered caching layer and batching writes to ClickHouse, we reduced p99 latency by 60% and increased throughput by 3x. I also profiled bottlenecks using distributed tracing and optimized key data paths.
I would start with a rule evaluation engine that can scale horizontally, using a consistent hashing mechanism to distribute evaluation load. Alerts would be generated via a topic-based message queue and processed by a deduplication and aggregation service. I'd integrate with a time-series database for metric storage and use a state machine to handle alert lifecycle. Finally, I'd ensure the pipeline is resilient to failures by designing for eventual consistency and using idempotent notifications.
I have used Prometheus extensively for metric collection and alerting, including writing custom exporters and recording rules. For Grafana, I have built dashboards that provide real-time visibility into system health and business metrics. I also contributed to an open-source project that enhanced Prometheus's remote write performance.
First, I gather data from logs, metrics, and traces to identify the symptom and affected components. I then form hypotheses and isolate variables using canary deployments or targeted logging. I work closely with team members to share findings and leverage collective knowledge. For example, I once debugged a cascading failure by correlating latency spikes in traces with a misconfigured connection pool, and fixed it by adjusting pool settings.
I am excited by Nebius's focus on building a full-stack AI cloud platform and the opportunity to work on observability at scale. The company's emphasis on trust, ownership, and fast-moving innovation aligns with my values. I also appreciate the chance to collaborate with a global team of experts and contribute to open-source projects that advance the industry.
About Nebius:
Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.
Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.
The role
We are looking for a Senior Backend Software Engineer to join our Observability Platform team in Amsterdam or remotely. Building a reliable and profitable cloud platform is impossible without world-class observability. It spans every layer of the stack — from developer tooling and CI/CD pipelines to Kubernetes, managed services, databases, networking, and customer workloads. Our mission is to provide a unified ecosystem for logs, metrics, traces, alerting, and troubleshooting that enables engineers to understand and operate systems at scale. You will build and scale the core services behind this platform, working on high-volume telemetry ingestion, distributed storage, query engines, alerting pipelines, and new ways to help engineers make sense of operational data, including AI-assisted troubleshooting.
We expect you to have:
- 5+ years of professional software engineering experience
- Strong knowledge of Golang or willingness to quickly switch to it
- Experience building distributed backend systems
- Solid understanding of software reliability, scalability, and performance
- Ability to troubleshoot complex production issues
- Teamwork-oriented approach and strong communication skills
It would be an added bonus if you had:
- Experience building or contributing to observability platforms, telemetry systems, or related open-source projects such as Prometheus, Grafana, Loki, Jaeger, OpenTelemetry, VictoriaMetrics, Mimir, Tempo, or similar technologies
- Experience using ClickHouse in production
Benefits & Perks:
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams
What’s it like to work at Nebius:
Fast moving – Bold thinking – Constant growth – Meaningful impact – Trust and real ownership – Opportunity to shape the future of AI
Equal Opportunity Statement:
Nebius is an equal opportunity employer. We are committed to fostering an inclusive and diverse workplace and to providing equal employment opportunities in all aspects of employment. We do not discriminate on the basis of race, color, religion, sex (including pregnancy), national origin, ancestry, age, disability, genetic information, marital status, veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by applicable law.
Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire.
If you need accommodations during the application process, please let us know.
Annual salary information is not provided for this position. Explore salary ranges for similar roles in our Salary Directory ›
This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.









