All remote jobs
Open role
Remote opportunity atDune

Staff Software Engineer – Curated Data

Review the role, location requirements, compensation details, and application process before deciding whether this opportunity fits your next career move.

Published
25Listing views
1Application actions
24 Oct 2026Apply before
Opportunity details

About this role.

AI Summary

Dune is hiring a Staff Software Engineer to build the control plane and architecture for its curated onchain-data lifecycle. The role centers on dependency-aware orchestration, backfills, schema evolution, data correctness, alerting, and recovery across a large-scale data platform. It is a senior hands-on systems role requiring production backend engineering alongside deep data and streaming expertise. The engineer will turn ambiguous product needs into technical designs, break work into deliverable increments, and support a distributed remote team across East Coast US and Europe.

Role DNA

A quick view of the complexity, pace, ownership and collaboration implied by the job description.

Job Complexity

5/5
EasyHard

Pace & Pressure

4/5
RelaxedFast-paced

Autonomy Level

5/5
GuidedFull ownership

Communication Load

5/5
IndependentCollaborative
AI insightThis is a staff-level architecture role operating at significant scale, including thousands of interdependent models, multi-petabyte datasets, schema contracts, and stateful streaming systems. Success requires independent technical judgment, distributed-systems depth, and the ability to lead execution through ambiguity.

Salary analysis

Estimated compensation compared with the broader US market for similar roles.

Estimated job medianMarket rate
$210,000
US market range$175k–$260k
AI insightNo actual salary range is disclosed, so these are estimated US-market base-salary figures for a Staff Software Engineer specializing in backend data platforms and distributed systems. Estimated median base salary is $210,000 USD, with a typical market range of $175,000-$260,000 USD; equity may be meaningful but is not quantified in the posting.

Core skills

Skills and capabilities most closely associated with this opportunity.

Sample interview questions
How would you design an orchestration control plane for thousands of interdependent data models while supporting retries, backfills, and partial failures?

I would model the system as a versioned dependency graph with explicit run state, input/output contracts, idempotency keys, and durable execution metadata. The scheduler should distinguish transient failures from data-quality failures, execute only affected downstream partitions where possible, and make retries and backfills observable, bounded, and safe to resume.

Describe how you would manage a breaking schema change for a widely consumed curated dataset.

I would first identify consumers and define a migration contract, typically introducing a versioned schema or additive fields before removal. I would validate both old and new paths through automated compatibility checks, communicate deprecation timelines, monitor consumer adoption, and only retire the old contract after usage and data-quality criteria are met.

What production issues have you encountered with stateful stream processing, and how did you address them?

Common issues include late or duplicated events, skewed partitions, state growth, checkpoint failures, and schema incompatibilities. I address them with explicit event-time and watermark policies, idempotent sinks or deduplication, partitioning analysis, state TTL and compaction strategies, and tested recovery procedures using representative failure scenarios.

How do you design data-quality alerting that catches meaningful incidents without creating excessive on-call noise?

I begin with user-impacting invariants such as freshness, completeness, volume shifts, referential integrity, and reconciliation against trusted sources. Alerts should use sensible baselines, severity tiers, deduplication, and actionable context; I regularly review false positives and tune or remove signals that do not lead to useful action.

How would you approach an ambiguous request for a new curated blockchain dataset?

I would clarify the target users, decisions they need to make, correctness expectations, freshness requirements, and downstream interfaces. Then I would map source availability and transformation dependencies, propose whether the solution should be a model, service, or job, document the contracts and failure modes, and sequence an initial useful release before scaling the architecture.

This analysis is generated from the job description. Salary estimates, role characteristics and sample answers are guidance, not employer-provided facts.

About Dune

Dune’s mission is to make onchain finance observable. We’re the industry standard for onchain data: a blockchain data and intelligence provider that institutions, protocols, and analysts trust to understand the onchain world and the frontiers of finance. We deliver structured datasets, spanning stablecoins, RWAs, tokens, lending, trading, and more, from 130+ chains and counting, to 1,000+ industry leaders including Visa, WisdomTree, FINRA, the IMF, Bloomberg, Standard Chartered, Coinbase, Forbes, and the Financial Times.

We’re a tight knit team of ~50 hard working people, spread across Europe and eastern US timezones 🌍️, having an outsized impact on the industry. We take pride in being ambitious and doing world class work while staying humble and lighthearted. We believe in building open, verifiable data that lets individuals and institutions do deep research into ecosystems like Ethereum, Solana, Robinhood and many more.

We’re backed by some of the world’s best investors. In February 2022, we announced our Series B funding round led by Coatue and Union Square Ventures, an important milestone that let us double down on our mission.

If you like being challenged and want to work together with a brilliant team on an important mission, come join us.

Learn more about us:

Dune’s Vision

https://dune.com/blog/dune-vision

Values and working at Dune

https://dune.com/careers

About the role

Data Products builds and owns datasets end to end: from raw chain data through decoding to the 3000+ models and 4 petabytes we curate, share directly with customers, and replicate into their warehouses.

The role will focus on the lifecycle of building high quality data: orchestrating thousands of interdependent models, propagating schema changes without breaking downstream consumers, propagating corrections.
That is a software architecture problem in a data domain. This role is a hybrid: a backend engineer who thinks in systems and contracts, working on data.

You will be the engineer we hand ambiguous product requirements to, and will come back with a design, a sequence, and work the team can pick up, while building the hardest parts yourself.

In this role you will

  • Design and build the control plane for our curated data lifecycle: dependency-aware orchestration, backfills, restatements, retries, partial failure, and recovery

  • Decide, dataset by dataset, whether the answer is a model, a service or a job, and own that architecture through production

  • Design the contracts between ingestion and curation so a dataset can be reasoned about end to end

  • Build alerting and data quality signals that catch real problems and stay quiet otherwise, so on-call is about incidents rather than noise

  • Work across Go, Kotlin, Rust, Python and SQL, choosing the right tool rather than the familiar one

  • Break large problems into work other engineers can own, and sequence it so we ship something useful early

You might be a great fit if

  • You are a backend engineer who has gone deep on data systems, or a data engineer who became a strong software engineer. You ship production services, not only pipelines

  • You have built or materially extended orchestration and scheduling systems, and can explain precisely what breaks at scale and why

  • You have handled schema evolution and data correctness in a system with real consumers downstream, where a breaking change has a cost

  • You have built or operated stateful stream processing in production (Flink,Kafka Streams, Spark Structured Streaming, RisingWave, Materialize, Feldera)

  • You have strong SQL and modeling skills on large datasets, and an interest in how the query engine underneath actually executes your work

  • You have solid computer science fundamentals and distributed systems understanding

  • You debug independently and drive root cause analysis to a fix that holds

  • You use AI tools well enough that they have changed how you work, you understand their failure modes and dislike ai-slop.

  • You communicate clearly in writing and get the best out of a distributed team

Not required, but a plus

  • Deep experience with a transformation framework such as dbt or SQLMesh: specifically, having hit its limits and built beyond them

  • Data lake formats such as Parquet, Iceberg or Deltalog

  • Stateful stream processing in production (Flink, Kafka Streams, Spark Structured Streaming)

  • Experience at a company where the data is the product

Perks & Benefits

  • A competitive salary and equity package 🚀. Both salary and equity is top 25% of companies in the space

  • Our employee equity scheme has world-class employee-friendly terms with a heavily discounted strike price (~90%) and a 10-year exercise window

  • 5 weeks PTO + local public holidays (that can be swapped to suit you) 🏖

  • A fully remote-first approach 🧑‍💻 within a distributed team with flexible working hours; you structure your own day

  • Say goodbye to meeting overload! We believe in a healthy mix of async and sync work, so you can focus on what truly matters—no more wasted time on endless meetings!

  • Good health is important, so we offer private medical insurance, dental & vision as standard 🩺

  • We believe in paid parental leave 👶 to help you celebrate this important milestone, transition to your new life, and bond with your new baby. We offer 16 weeks to primary and 6 weeks to secondary caregivers, fully paid. Plus a 2-week part-time phased return at full pay to help you get used to your new (and slightly more complex!) schedule

  • Quarterly offsites in various exciting locations as a company or team to connect, work together and have fun (so far in Tuscany 🇮🇹 Berlin 🇩🇪 Austria 🇦🇹 and Athens 🇬🇷).

  • On top of this 👆each person gets a yearly travel allowance to connect and co-work with someone or a team of people for a few days.

  • An allowance for your at-home setup, to ensure you are happy, comfortable and productive. If you prefer a local co-working space, we’ll pay for your desk.

  • Work with some of the best people you’ll ever get to meet!

  • And of course, you get some awesome Dune swag! ✌️😎

——————————————————————————————

We are dedicated to building a diverse, inclusive, and authentic workplace, so if you’re excited about this role but your experience doesn’t align perfectly with every qualification in the job description, we encourage you to apply anyway. You may be just the right candidate for this or other roles.

#LI-Remote

Apply now >

This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.

Next step

Apply now.

Follow the employer’s application method and review Jobicy’s safety guidance before sharing personal information.

Did you apply?Let us know, and we’ll help you track your application.

Continue on the employer website

Protect your personal information and never pay to secure an interview or job offer. View safety guidance.

Log in to save
One quick step before you apply

Sign in to continue.

Sign in or create a free account to continue to the employer's application.

Applying is free. After signing in, return to this job and select Apply Now.
Add alert
Jobs Talent AI Tools Salaries
Menu