All remote jobs
Open role
Remote opportunity atSocure

Senior Data Engineer

Review the role, location requirements, compensation details, and application process before deciding whether this opportunity fits your next career move.

Published
58Listing views
2Application actions
22 Oct 2026Apply before
Opportunity details

About this role.

AI Summary

Socure is hiring a Senior Data Engineer for its Data Automation team to build scalable batch and streaming pipelines that support identity verification products, ML feature engineering, and analytics. The role owns ambiguous initiatives end to end, from architecture and implementation through deployment, monitoring, recovery, and documentation. Core requirements include Python or Scala, SQL, Apache Spark, AWS data services, data modeling, and production reliability practices. The engineer will collaborate with Data Science, Product, and Engineering while improving platform cost, performance, and operational automation.

Role DNA

A quick view of the complexity, pace, ownership and collaboration implied by the job description.

Job Complexity

5/5
EasyHard

Pace & Pressure

5/5
RelaxedFast-paced

Autonomy Level

5/5
GuidedFull ownership

Communication Load

4/5
IndependentCollaborative
AI insightThis is a senior, highly autonomous platform role requiring deep distributed-data expertise and ownership of complex production systems. The environment emphasizes rapid execution, ambiguous problem solving, reliability, and cross-functional technical leadership.

Salary analysis

Estimated compensation compared with the broader US market for similar roles.

Estimated job medianHighly competitive
$177,500
US market range$150k–$205k
AI insightThe disclosed annual base compensation range is USD 160,000–195,000, with a midpoint of USD 177,500. For a US-based senior data engineer focused on Spark, AWS, streaming, and data-platform ownership, an estimated US market range is approximately USD 150,000–205,000 annually; actual market pay varies by location, company stage, and total-equity package.

Core skills

Skills and capabilities most closely associated with this opportunity.

Sample interview questions
How would you design a scalable pipeline that supports both batch ingestion and near-real-time feature generation?

I would separate ingestion, storage, transformation, and serving concerns while using shared schemas and data-quality contracts. Batch workloads could use Spark on EMR with partitioned lake storage, while streaming events could be processed through Kafka or Kinesis with idempotent writes and checkpointing. I would add orchestration, monitoring, lineage, and backfill mechanisms from the outset so both paths remain reliable and operable.

Describe how you would tune a slow Apache Spark job processing a large, skewed dataset.

I would first inspect Spark UI metrics to identify whether the bottleneck is skew, shuffle volume, poor partition sizing, spills, or inefficient joins. Typical improvements include filtering and projecting early, repartitioning on appropriate keys, using broadcast joins where suitable, handling skewed keys with salting or adaptive query execution, and selecting efficient file formats such as Parquet. I would validate changes with representative production-scale data and track both runtime and infrastructure cost.

What practices would you use to make production data pipelines reliable?

I would implement data contracts, schema validation, freshness and completeness checks, idempotent processing, retries with bounded backoff, and clear dead-letter or quarantine paths. Operationally, I would provide actionable alerts, runbooks, dashboards for latency and failure rates, and tested procedures for backfills and recovery. CI/CD, automated tests, code review, and infrastructure-as-code would help prevent regressions before deployment.

How do you balance data-platform cost optimization with performance and reliability requirements?

I start by defining service-level expectations for latency, freshness, availability, and recovery time, then measure cost per workload and per unit of data processed. I optimize storage layout, retention, compute sizing, scheduling, autoscaling, and query patterns without weakening critical controls. Any change should be benchmarked, monitored after release, and reversible if it creates reliability or user-impact risks.

How would you work with Data Science and Product stakeholders when requirements are incomplete?

I would clarify the customer and business outcome, identify decisions that the data product must enable, and convert those into measurable technical requirements such as data sources, freshness, quality thresholds, and access patterns. I would propose an incremental design, document assumptions and trade-offs, and seek early feedback through lightweight prototypes. Regular communication on risks, dependencies, and delivery milestones keeps technical and non-technical stakeholders aligned.

This analysis is generated from the job description. Salary estimates, role characteristics and sample answers are guidance, not employer-provided facts.

Why Socure?

Socure is building the identity trust infrastructure for the digital economy — verifying 100% of good identities in real time and stopping fraud before it starts. The mission is big, the problems are complex, and the impact is felt by businesses, governments, and millions of people every day.

We hire people who want that level of responsibility. People who move fast, think critically, act like owners, and care deeply about solving customer problems with precision. If you want predictability or narrow scope, this won’t be your place. If you want to help build the future of identity with a team that holds a high bar for itself — keep reading.

About the Role

We are looking for a Senior Data Engineer to join our Data Automation team. You will play a critical role in designing and building scalable data platforms and pipelines that power Socure’s identity verification products and analytics. This role is ideal for someone who has a strong passion for solving real business problems with data, and combines deep hands-on data engineering expertise with strong ownership.

What You’ll Do

• Design and build batch and streaming data pipelines to support automated data ingestion, ML feature engineering and analytics across multiple product domains.

• Own end-to-end delivery of complex, ambiguous data initiatives, including architecture, implementation, testing, deployment, monitoring, and documentation.

• Develop and evolve the data platform to support large-scale data processing using modern cloud-native technologies.

• Automate data operations (validation, quality checks, alerting, backfills, and recovery workflows) to reduce manual effort and improve consistency.

• Optimize cost, performance, and reliability of data workloads.

• Partner closely with cross-functional teams (Data Science, Product, Engineering) to understand requirements, translate them into technical solutions.

• Evaluate and adopt new technologies (new processing engines, storage formats, orchestration tools, GenAI-assisted ingestion) to keep the platform modern and efficient.

What You Bring

• 5+ years of hands-on data engineering experience, building and maintaining production-grade data platforms and pipelines.

• Strong programming skills in general-purpose language (such as Python or Scala) for data processing, and SQL for data analytics.

• Deep experience with distributed data processing frameworks, such as Apache Spark, including performance tuning and optimization.

• Proven experience building data solutions using services on AWS (EMR, Lambda, s3, etc).

• Strong understanding of data modeling and data warehousing concepts, including partitioning, schema design for large-scale datasets.

• Experience operating and supporting production pipelines, including monitoring, alerting, incident response, and improving reliability over time.

• Solid foundation in software engineering practices, including version control, CI/CD, testing strategies, and code review.

• Strong communication and collaboration skills, with the ability to work effectively with both technical and non-technical stakeholders.

Preferred Qualifications

• Experience with streaming or near-real-time data processing (Kafka, Kinesis, etc).

• Hands-on experience with data orchestration tools (Airflow, Step Functions, etc).

• Familiarity with modern data platform patterns such as Data Lakehouse, Data Mesh, and large-scale data sharing across teams.

• Experience with prompt engineering using modern GenAI, Large Language Models (LLM).

• Experience mentoring other engineers and contributing to engineering-wide standards, best practices.

As a note; Socure cannot provide sponsorship now or in the future for this role.

Socure is an equal opportunity employer that values diversity in all its forms within our company. We do not discriminate based on race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.
If you need an accommodation during any stage of the application or hiring process—including interview or onboarding support—please reach out to your Socure recruiting partner directly.

Follow Us!

YouTube | LinkedIn | X (Twitter) | Facebook

Apply now >

This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.

Next step

Apply now.

Follow the employer’s application method and review Jobicy’s safety guidance before sharing personal information.

Did you apply?Let us know, and we’ll help you track your application.

Continue on the employer website

Protect your personal information and never pay to secure an interview or job offer. View safety guidance.

Log in to save
One quick step before you apply

Create your free account, then apply.

Build a more organized job search on Jobicy and continue to the employer's application when you're ready.

  • Never lose a promising opportunitySave roles and return to them from your dashboard.
  • See your entire search at a glanceTrack applications, stages and next steps in one place.
  • Get matched with relevant remote jobsChoose the alerts and digests that work for you.
Applying is free. The employer's application opens in a new tab.
Add alert
Jobs Talent AI Tools Salaries
Menu