About this role.
This is a mid-level Data Engineer contract role focused on building and scaling ETL/ELT pipelines for a unified business-data platform. The engineer will ingest, transform, normalize, enrich, and monitor third-party and proprietary firmographic datasets across two ecosystems. Core technologies include SQL, Python, dbt, Snowflake, and preferably Azure, with Spark, Airflow, and Azure Data Lake as valuable complementary experience. The role also involves entity matching, semantic models, data-quality controls, documentation, and close collaboration with product, architecture, analytics, and content teams. It is a six-month full-time remote contract in Poland with a potential extension.
Role DNA
A quick view of the complexity, pace, ownership and collaboration implied by the job description.
Job Complexity
4/5Pace & Pressure
4/5Autonomy Level
4/5Communication Load
4/5Salary analysis
Estimated compensation compared with the broader US market for similar roles.
Core skills
Skills and capabilities most closely associated with this opportunity.
Cover letter sample
Dear Hiring Team,
I am excited to apply for the Data Engineer role and contribute to building reliable, scalable data systems for a high-impact business data platform. My experience designing production ETL/ELT pipelines with SQL, Python, dbt, cloud data platforms, and robust monitoring practices aligns closely with your need for high-quality ingestion, transformation, and enrichment workflows.
I would bring a detail-oriented, ownership-driven approach to data modeling, entity resolution, data quality, and operational troubleshooting while collaborating effectively with product managers, architects, analysts, and content specialists. I am particularly motivated by the opportunity to support a modern, data-driven consulting product and help create a trusted global business directory.
Thank you for your consideration; I would welcome the opportunity to discuss how I can support the team’s data platform goals.
Sample interview questions
I would explain the source ingestion, transformation layers, orchestration approach, testing strategy, and observability setup. I would highlight idempotent loads, schema validation, data-quality checks, retry handling, alerting, and documented recovery procedures.
I would separate models into staging, intermediate, and marts or semantic layers. Staging models would standardize source fields, intermediate models would apply reusable business logic and entity preparation, and final models would expose governed metrics and dimensions, supported by tests, source freshness checks, and clear documentation.
I would begin with deterministic matching on reliable identifiers such as domain, registration number, or normalized contact attributes, then use probabilistic or ML-assisted matching for ambiguous records. I would retain match scores, lineage, survivorship rules, and a review path for low-confidence matches.
I would monitor pipeline completion, source freshness, row-count and volume anomalies, model-test results, warehouse costs, query failures, and end-to-end data latency. Alerts should be prioritized by business impact and linked to actionable runbooks and logs.
I would facilitate definition reviews to document the business meaning, grain, source of truth, quality expectations, and downstream impact of each dataset. I would version transformations, communicate changes early, and use dbt documentation and tests to keep definitions transparent and enforceable.
GT was founded in 2019 by a former Apple, Nest, and Google executive. GT’s mission is to connect the world’s best talent with product careers offered by high-growth companies in the UK, USA, Canada, Germany, and the Netherlands.
On behalf of the client, GT is looking for a Data Engineer who is interested in working with big data.
About the Client
Our client is a leading global management consultancy known for tackling some of the world’s most complex business challenges. With a focus on strategy, transformation, and performance improvement, the firm partners with major organizations across industries to drive lasting impact.
Recognized consistently as a top workplace, it combines deep industry expertise with a collaborative, innovative culture. Its centralized European hub plays a key role in supporting operations across the EMEA region, ensuring excellence and efficiency at scale.
About the Project & Role
The Data Engineer will play a critical role in developing, maintaining, and scaling the pipelines and data systems that power the client’s unified data platform. Working alongside data architects, product managers, and analysts, this role is focused on ingesting, transforming, and enriching firmographic data from multiple third-party and proprietary sources.
The engineer will support data operations across 2 ecosystems, contributing to the client’s mission of building the best business directory in the world and enabling differentiated, data-driven insights for consultants and clients.
Contract duration: 6 months, with the possibility of extension.
Key Responsibilities
Build Data Pipelines: Design and maintain robust, scalable ETL/ELT pipelines to ingest and process third-party and first-party datasets.
Data Quality & Enrichment: Apply transformation, normalization, and enrichment rules to ensure data consistency and usability.
Collaborate Across Teams: Work with product managers, data architects, and content experts to align data structure with business needs.
Operationalize Matching & Merging Logic: Support the implementation of data matching and entity resolution processes using AI/ML tools and proprietary frameworks.
Monitor & Troubleshoot Pipelines: Build alerts, logs, and metrics to ensure data flows remain healthy and issues are identified and resolved quickly.
dbt Development: Build and maintain data pipelines using dbt, developing models from raw data through staging and intermediate layers up to final semantic models for use in analytics and dashboards.
Documentation & Standards: Contribute to documentation, code quality standards, and internal best practices to ensure maintainability.
Ideal Candidate Profile
Experience: 4–8 years in data engineering, with experience in building production-grade data pipelines.
Tech Stack: Proficient in SQL, Python, and dbt, with strong experience in Snowflake. Additional experience with Spark, Airflow, and Azure Data Lake or similar technologies.
Cloud Platforms: Familiarity with Azure (preferred) or other major cloud platforms.
Data Engineering Best Practices: Understanding of data modeling, version control, CI/CD, and data governance principles.
Curiosity & Ownership: Proactive, detail-oriented, and eager to take ownership of projects and continuously improve systems.
Team Player: Comfortable working in a cross-functional environment and open to learning from and supporting teammates.
Why Join Us?
Join a fast-growing, high-impact team at the intersection of 2 key projects, building a data product core to client’s future.
Contribute to an ambitious effort to create the highest quality, most comprehensive business directory in the world.
Be part of a project, a startup-style group within the company that’s redefining how we deliver consulting through productization and data innovation.
Work with cutting-edge data tools, including AI/ML enrichment, semantic matching, and modern cloud-based infrastructure.
Annual salary information is not provided for this position. Explore salary ranges for similar roles in our Salary Directory ›
This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.




