About this role.
Socure is hiring a Senior Data Engineer for its Data Automation team to build scalable batch and streaming pipelines that support identity verification products, ML feature engineering, and analytics. The role owns ambiguous initiatives end to end, from architecture and implementation through deployment, monitoring, recovery, and documentation. Core requirements include Python or Scala, SQL, Apache Spark, AWS data services, data modeling, and production reliability practices. The engineer will collaborate with Data Science, Product, and Engineering while improving platform cost, performance, and operational automation.
Role DNA
A quick view of the complexity, pace, ownership and collaboration implied by the job description.
Job Complexity
5/5Pace & Pressure
5/5Autonomy Level
5/5Communication Load
4/5Salary analysis
Estimated compensation compared with the broader US market for similar roles.
Core skills
Skills and capabilities most closely associated with this opportunity.
Sample interview questions
I would separate ingestion, storage, transformation, and serving concerns while using shared schemas and data-quality contracts. Batch workloads could use Spark on EMR with partitioned lake storage, while streaming events could be processed through Kafka or Kinesis with idempotent writes and checkpointing. I would add orchestration, monitoring, lineage, and backfill mechanisms from the outset so both paths remain reliable and operable.
I would first inspect Spark UI metrics to identify whether the bottleneck is skew, shuffle volume, poor partition sizing, spills, or inefficient joins. Typical improvements include filtering and projecting early, repartitioning on appropriate keys, using broadcast joins where suitable, handling skewed keys with salting or adaptive query execution, and selecting efficient file formats such as Parquet. I would validate changes with representative production-scale data and track both runtime and infrastructure cost.
I would implement data contracts, schema validation, freshness and completeness checks, idempotent processing, retries with bounded backoff, and clear dead-letter or quarantine paths. Operationally, I would provide actionable alerts, runbooks, dashboards for latency and failure rates, and tested procedures for backfills and recovery. CI/CD, automated tests, code review, and infrastructure-as-code would help prevent regressions before deployment.
I start by defining service-level expectations for latency, freshness, availability, and recovery time, then measure cost per workload and per unit of data processed. I optimize storage layout, retention, compute sizing, scheduling, autoscaling, and query patterns without weakening critical controls. Any change should be benchmarked, monitored after release, and reversible if it creates reliability or user-impact risks.
I would clarify the customer and business outcome, identify decisions that the data product must enable, and convert those into measurable technical requirements such as data sources, freshness, quality thresholds, and access patterns. I would propose an incremental design, document assumptions and trade-offs, and seek early feedback through lightweight prototypes. Regular communication on risks, dependencies, and delivery milestones keeps technical and non-technical stakeholders aligned.
Why Socure?
Socure is building the identity trust infrastructure for the digital economy — verifying 100% of good identities in real time and stopping fraud before it starts. The mission is big, the problems are complex, and the impact is felt by businesses, governments, and millions of people every day.
We hire people who want that level of responsibility. People who move fast, think critically, act like owners, and care deeply about solving customer problems with precision. If you want predictability or narrow scope, this won’t be your place. If you want to help build the future of identity with a team that holds a high bar for itself — keep reading.
About the Role
We are looking for a Senior Data Engineer to join our Data Automation team. You will play a critical role in designing and building scalable data platforms and pipelines that power Socure’s identity verification products and analytics. This role is ideal for someone who has a strong passion for solving real business problems with data, and combines deep hands-on data engineering expertise with strong ownership.
What You’ll Do
• Design and build batch and streaming data pipelines to support automated data ingestion, ML feature engineering and analytics across multiple product domains.
• Own end-to-end delivery of complex, ambiguous data initiatives, including architecture, implementation, testing, deployment, monitoring, and documentation.
• Develop and evolve the data platform to support large-scale data processing using modern cloud-native technologies.
• Automate data operations (validation, quality checks, alerting, backfills, and recovery workflows) to reduce manual effort and improve consistency.
• Optimize cost, performance, and reliability of data workloads.
• Partner closely with cross-functional teams (Data Science, Product, Engineering) to understand requirements, translate them into technical solutions.
• Evaluate and adopt new technologies (new processing engines, storage formats, orchestration tools, GenAI-assisted ingestion) to keep the platform modern and efficient.
What You Bring
• 5+ years of hands-on data engineering experience, building and maintaining production-grade data platforms and pipelines.
• Strong programming skills in general-purpose language (such as Python or Scala) for data processing, and SQL for data analytics.
• Deep experience with distributed data processing frameworks, such as Apache Spark, including performance tuning and optimization.
• Proven experience building data solutions using services on AWS (EMR, Lambda, s3, etc).
• Strong understanding of data modeling and data warehousing concepts, including partitioning, schema design for large-scale datasets.
• Experience operating and supporting production pipelines, including monitoring, alerting, incident response, and improving reliability over time.
• Solid foundation in software engineering practices, including version control, CI/CD, testing strategies, and code review.
• Strong communication and collaboration skills, with the ability to work effectively with both technical and non-technical stakeholders.
Preferred Qualifications
• Experience with streaming or near-real-time data processing (Kafka, Kinesis, etc).
• Hands-on experience with data orchestration tools (Airflow, Step Functions, etc).
• Familiarity with modern data platform patterns such as Data Lakehouse, Data Mesh, and large-scale data sharing across teams.
• Experience with prompt engineering using modern GenAI, Large Language Models (LLM).
• Experience mentoring other engineers and contributing to engineering-wide standards, best practices.
As a note; Socure cannot provide sponsorship now or in the future for this role.
Socure is an equal opportunity employer that values diversity in all its forms within our company. We do not discriminate based on race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.
If you need an accommodation during any stage of the application or hiring process—including interview or onboarding support—please reach out to your Socure recruiting partner directly.
Follow Us!
YouTube | LinkedIn | X (Twitter) | Facebook
This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.








