About this role.
This is a senior Reliability Engineer role at Meta, focused on asset management and reliability within data center facility operations. The position involves leading FMEA, developing maintenance strategies, driving corrective maintenance through root cause analysis, and implementing condition-based monitoring programs. The ideal candidate has 8+ years of experience in reliability engineering, expertise in RCM and FMEA, and proficiency with EAM solutions. This role offers an opportunity to work on critical infrastructure and collaborate across global teams.
Role DNA
A quick view of the complexity, pace, ownership and collaboration implied by the job description.
Job Complexity
4/5Pace & Pressure
4/5Autonomy Level
4/5Communication Load
5/5Salary analysis
Estimated compensation compared with the broader US market for similar roles.
Core skills
Skills and capabilities most closely associated with this opportunity.
Cover letter sample
Dear Hiring Manager,
I am excited to apply for the Data Center Facility Operations Reliability Engineer position at Meta. With over 8 years of experience in reliability engineering, including expertise in FMEA and RCM, I have a proven track record of optimizing asset performance and reducing downtime in critical environments. My background in data center operations and proficiency with EAM solutions align perfectly with the responsibilities outlined.
I have successfully led cross-functional teams to implement condition-based monitoring programs and drive continuous improvement in MTTR and MTBF metrics. I am confident that my technical skills and collaborative approach will contribute to Meta's mission of delivering reliable infrastructure.
Thank you for considering my application. I look forward to the opportunity to discuss how my experience can support your team.
Sincerely,
[Your Name]
Sample interview questions
In my previous role, I led FMEA sessions for critical cooling systems, identifying failure modes like pump seal failures. We implemented predictive maintenance and redesigned seals, reducing downtime by 30%. I also used the results to update our maintenance library and procedures.
I use risk-based prioritization, considering criticality, failure impact, and likelihood. I also factor in production risk models and mean time between failures (MTBF) data. For example, I once prioritized a generator issue over a cooling fan because it had higher business impact.
I have implemented vibration analysis on pumps and fans to detect bearing degradation early. I also used infrared thermography to identify hot spots in electrical panels. These sensors fed into our EAM system, triggering maintenance tasks automatically.
I established a governance framework where all maintenance procedure changes required approval from reliability and safety teams. I used a version control system in our EAM to track updates. For example, when we optimized a filter replacement schedule, I ensured proper documentation and training.
At my last job, I analyzed historical failure data for chillers and found a common root cause: improper refrigerant charge. I worked with the maintenance team to standardize charging procedures, which increased MTBF by 20% and reduced MTTR by 15% through better troubleshooting guides.
Meta is seeking an experienced and self-motivated Reliability Engineer to join our Asset Management & Reliability team within Facility Operations. This person will work within Facility Operations to identify and manage asset reliability risks and various stages of end-to-end asset lifecycle for the Data Center Operations. Managing stakeholders spread across time zones and key to the success of our individual projects and overall asset management, and reliability program. In this role, you will be expected to own and execute within your focus areas, contribute to scoping and roadmapping for reliability initiatives, and solve complex technical problems to ensure seamless operations across new and existing sites.ResponsibilitiesLead operational Failure Mode and Effects Analysis (FMEA) and establish maintenance strategies.* Manage and govern content for the Global Maintenance Library, ensuring rigorous Maintenance Change Management processes are followed.* Drive Corrective Maintenance initiatives through deep Failure Mode Analysis to identify root causes and implement preventive measures.* Evaluate regulatory and compliance requirements, providing governance and ensuring adherence across all reliability activities.* Support and execute Condition-Based Monitoring and Diagnostics programs, including testing, vibration analysis, infrared (IR), partial discharge (PD), and robotics applications.* Assess and optimize Mean Time To Repair (MTTR) and Mean Time Between Failures (MTBF) metrics to drive continuous improvement.* Contribute to continuous production risk models to prioritize resources and mitigate operational vulnerabilities.* Manage non-critical scope and support supplier scope management to ensure external partners meet reliability and quality standards.* Implement and optimize sensor-based condition monitoring and maintenance (e.g., infrared, temperature, pressure, flow, leak detection) to remotely monitor and assign maintenance tasks.* Support technical insourcing versus outsourcing decision-making for maintenance and reliability activities.* Establish Service Bill of Materials (BOM) setups and drive initial and continuous spares selection processes.* Develop and execute a comprehensive whole-equipment sparing strategy to minimize downtime.* Identify First-of-Kind (FOK) parts and establish Parts Factory Base Part Number (FBPN) setups.* Conduct Form-Fit-Function technical analysis to evaluate and approve replacement components.* Act as the primary liaison to IBOS and ISCE teams for managing replacement parts backlog, timeliness, tooling requirements, and auditing.* Identify reliability opportunities and contribute to engineering excellence within the team.QualificationsBachelor’s degree in Mechanical, Electrical, Reliability Engineering, or a similar technical discipline or 6+ years relevant industry experience will be considered in lieu of 8+ years relevant industry experience* 8+ years of experience in reliability engineering (related to electrical or mechanical cooling equipment) or related Engineering roles* Experienced in Reliability Centered Maintenance (RCM) and Failure Mode and Effects Analysis (FMEA) activities for maintenance, process, and equipment design optimization to meet reliability requirements* Proven ability to execute independently within defined scope and manage complex problems with guidance on ambiguous areas* Proficient in the usage of Enterprise Asset Management (EAM) solutions to extract data and develop meaningful insights* Knowledgeable of relevant ISO standards (ISO 14224, ISO 17359, ISO 55000)* Experience with project management and cross-functional collaboration* Ability to travel 25% domestically and internationally Experience with data center equipment such as critical cooling systems, generators, main switchboards, and network gear* Proficient in data analysis techniques that can include Process Control, Reliability modeling and prediction, Fault Tree Analysis, Weibull Tree Analysis, and Six Sigma (6σ) Methodology* Proficient in developing and executing comprehensive test plans for assets* Experience supporting technical initiatives and advocating for high-quality engineering standards* Certifications in Maintenance & Reliability such as CMRP, CRL, or CRE
Annual salary information is not provided for this position. Explore salary ranges for similar roles in our Salary Directory ›
This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.






