About this role.
Mactores is seeking a Data Engineer Intern to contribute to AWS data-platform modernization projects under a mentor-oriented engineering team. The intern will write pipeline code using technologies such as Apache Spark or Apache Beam, apply ETL concepts with PySpark and SparkSQL, and collaborate with business and engineering stakeholders. The role emphasizes learning production data engineering practices, data quality decision-making, and agent-assisted delivery rather than independently leading architecture or cutovers. Preferred exposure includes AWS EMR, Apache Airflow, DataOps, and relevant cloud or big-data certifications.
Role DNA
A quick view of the complexity, pace, ownership and collaboration implied by the job description.
Job Complexity
3/5Pace & Pressure
4/5Autonomy Level
2/5Communication Load
4/5Salary analysis
Estimated compensation compared with the broader US market for similar roles.
Core skills
Skills and capabilities most closely associated with this opportunity.
Sample interview questions
I would first validate the source schema, file completeness, and data types during ingestion. I would use a layered approach: retain raw files, transform and standardize data in a cleaned layer, then publish curated tables with documented business logic. I would add quality checks for nulls, duplicates, referential integrity, and row-count anomalies, while making the pipeline idempotent so reruns do not create duplicate records.
Transformations, such as select, filter, and join, define a new DataFrame but are evaluated lazily. Actions, such as count, collect, write, or show, trigger Spark to execute the required computation. Understanding this distinction helps avoid unnecessary actions and supports more efficient pipeline design.
I would compare the current run with a known-good run, beginning with source volumes, schema changes, and pipeline row counts at each stage. I would review recent code and configuration changes, inspect join behavior and filtering logic, and identify whether records were lost, duplicated, or reclassified. After finding the cause, I would correct it, validate the result against expected business totals, and add a monitoring check to detect the issue earlier in future runs.
Airflow can orchestrate tasks in a pipeline by defining dependencies, schedules, retries, alerts, and execution parameters in a DAG. For example, it can coordinate source ingestion, Spark transformations, quality validation, and publication of a final dataset. It improves operational reliability by making task status, failures, and reruns visible and manageable.
I would state the impact clearly, explain the issue in business-oriented language, and avoid implying certainty before the investigation is complete. I would provide the current status, affected datasets or reports, mitigation steps, and a realistic next update time or delivery estimate. I would also confirm whether a temporary workaround or prior validated dataset can support the stakeholder's immediate need.
Mactores is the agent-native AWS modernization firm. Most modernization work doesn’t ship, it stalls in pilots, slips a year, or lands at three times the budget. We exist to ship it: production systems running, legacy retired, outcomes measured. Our delivery is built on Aedeon, the agent platform built by Mactores’ founders’ sister company, which absorbs the repetitive 60–70% of engagement work, discovery, dependency mapping, validation, test generation, that traditional consulting bills human hours against. Forward-deployed engineers own the rest: architecture, judgment, and cutover, on dates we commit to in the contract.
This internship is an apprenticeship in data platform modernization the pillar of our work where legacy warehouses get retired and data platforms reach production on AWS. You’ll write real pipeline code on real projects, working with business leads, analysts, and data scientists to understand the domain, then with engineers to build data products that make decisions better.
Here’s the honest framing: you’re learning the agent-native craft, not owning cutover. Agents handle much of the repetitive discovery and validation work that used to fill junior engineers’ days, which means your time goes further into Spark, ETL design, and understanding why data quality decisions matter to the business. If you care about the quality of the metrics a business runs on, and you want your solutions to scale to bigger questions, this is the seat.
What you will do?
- Write efficient code in the technology chosen for the project — Spark or Apache Beam, for example.
- Explore new technologies and learn new techniques to solve business problems creatively.
- Collaborate across engineering and business teams to build better data products and services.
- Deliver projects with the team and keep customers updated on time — shipping on schedule is a habit you’ll build early here.
What we are looking for?
- Exposure to Apache Spark.
- Exposure to ETL concepts using pySpark and SparkSQL.
- Exposure to SQL queries and stored procedures, and the appetite to work on challenging projects with a mentor-oriented leader.
You will be preferred if you have
- Prior experience in working on AWS EMR, Apache Airflow
- AWS Certified Big Data – Specialty certification, Azure Certification, Snowflake Certification
- Cloudera or Hortonworks Certified Big Data Engineer
- Understanding of DataOps Engineering
How we work?
Mactores delivers through agents plus forward-deployed engineers. Aedeon, the agent platform we deploy, does the repetitive majority source discovery, schema mapping, validation harnesses while engineers make the architecture and cutover calls. As an intern you work inside that model from day one: you’ll see how a data platform actually reaches production, and you’ll learn to work with agents as tooling rather than treating them as a threat or a magic trick. The engineers who grow fastest here are the ones who master that combination.
Compensation
Additional Information
Life at Mactores
We care about creating a culture that makes a real difference in the lives of every Mactorian. Our 10 Core Leadership Principles that honor Decision-making, Leadership, Collaboration, and Curiosity drive how we work.
1. Be one step ahead
2. Deliver the best
3. Be bold
4. Pay attention to the detail
5. Enjoy the challenge
6. Be curious and take action
7. Take leadership
8. Own it
9. Deliver value
10. Be collaborative
We would like you to read more details about the work culture on https://mactores.com/careers
The Path to Joining the Mactores Team
At Mactores, our recruitment process is structured around three distinct stages:
Pre-Employment Assessment:
A series of evaluations of your technical proficiency and suitability for the role.
Managerial Interview: The hiring manager engages with you in multiple discussions, 30 minutes to an hour each, covering technical skills, hands-on experience, leadership potential, and communication.
HR Discussion: During this 30-minute session, you’ll have the opportunity to discuss the offer and next steps with a member of the HR team.
Mactores provides equal opportunities in all employment practices. We don’t discriminate based on race, religion, gender, national origin, age, disability, marital status, military status, genetic information, or any other category protected by federal, state, and local laws. This applies to every part of the employment relationship, recruitment, compensation, promotions, transfers, disciplinary action, layoff, training, and social and recreational programs.
Note: Please answer as many questions as possible with this application to accelerate the hiring process.
Annual salary information is not provided for this position. Explore salary ranges for similar roles in our Salary Directory ›
This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.








