All remote jobs
Open role
Remote opportunity atBayesian Health

Infrastructure Engineer

Review the role, location requirements, compensation details, and application process before deciding whether this opportunity fits your next career move.

Published
34Listing views
2Application actions
18 Oct 2026Apply before
Opportunity details

About this role.

AI Summary

Bayesian Health is seeking an Infrastructure Engineer to build and operate scalable, fault-tolerant AWS infrastructure supporting a clinical AI and healthcare platform. The role owns infrastructure architecture, Kubernetes/EKS operations, Terraform modules, CI/CD workflows, observability, reliability, and cloud-cost optimization. It also requires close collaboration with security and cross-functional engineering teams to meet HIPAA, HITRUST, FDA, and client requirements for sensitive PHI/PII workloads. This is a senior hands-on startup role with responsibility for improving platform standards while designing an internal AI Ops capability for DevOps optimization.

Role DNA

A quick view of the complexity, pace, ownership and collaboration implied by the job description.

Job Complexity

5/5
EasyHard

Pace & Pressure

5/5
RelaxedFast-paced

Autonomy Level

5/5
GuidedFull ownership

Communication Load

4/5
IndependentCollaborative
AI insightThe position requires deep production expertise across AWS, Kubernetes, Terraform, databases, CI/CD, observability, and regulated healthcare security. Its startup context, enterprise scaling goals, and ownership of infrastructure standards create substantial technical ambiguity and operational responsibility.

Salary analysis

Estimated compensation compared with the broader US market for similar roles.

Estimated job medianMarket rate
$175,000
US market range$145k–$210k
AI insightNo salary was disclosed in the posting. For a US-remote Infrastructure/DevOps Engineer with 5+ years of AWS, Kubernetes, Terraform, CI/CD, and regulated-healthcare experience, the estimated annual base-salary median is $175,000 USD, with an estimated US market range of $145,000–$210,000 USD; this is a market estimate rather than an offer from Bayesian Health.

Core skills

Skills and capabilities most closely associated with this opportunity.

Sample interview questions
How would you design a cost-optimized and fault-tolerant AWS platform for a growing healthcare SaaS product?

I would begin by defining availability, recovery, performance, and compliance objectives for each workload. I would use multi-AZ managed services where appropriate, autoscaling Kubernetes workloads, infrastructure-as-code, backup and disaster-recovery testing, and clear service-level monitoring. Cost controls would include rightsizing, autoscaling policies, tagging, budget alerts, and regular review of storage, compute, and data-transfer usage.

Describe your approach to managing Kubernetes cluster bootstrapping and day-2 operations in EKS.

I use repeatable Terraform modules and versioned configuration to provision clusters, networking, IAM, node groups, add-ons, and baseline policies. For day-2 operations, I focus on upgrade runbooks, capacity planning, workload resource limits, secret management, observability, backup and restore validation, and incident response. I also document operational ownership and automate common maintenance tasks to reduce manual drift.

How would you build a CI/CD process that supports regulated change control requirements?

I would implement protected branches, mandatory peer reviews, automated tests, artifact versioning, and promotion through controlled environments. Each deployment should have auditable links to source changes, approvals, test results, and release records. I would favor automated, reversible deployments with clearly defined rollback procedures and separation between development, staging, and production access.

What practices would you use to protect PHI and PII in cloud infrastructure?

I would apply least-privilege IAM, encryption in transit and at rest, network segmentation, centralized audit logging, secure secrets management, and continuous vulnerability and configuration monitoring. I would ensure data access is traceable, limit production access through role-based controls, and partner with security stakeholders on threat modeling and compliance evidence. Regular backup validation and incident-response exercises are also essential.

Give an example of how you would troubleshoot a production performance issue affecting an application and its database.

I would first assess user impact and review service-level indicators, recent changes, application logs, infrastructure metrics, traces, and database performance data. I would isolate whether the bottleneck is compute, network, Kubernetes scheduling, connection saturation, slow queries, or an external dependency, then mitigate safely through rollback, scaling, traffic control, or query remediation. After recovery, I would lead a blameless review and implement monitoring, capacity, or architectural improvements to prevent recurrence.

This analysis is generated from the job description. Salary estimates, role characteristics and sample answers are guidance, not employer-provided facts.

In Brief

  • We’re a rapidly growing startup on a mission to make healthcare proactive by empowering physicians, nurses, and care team members with real-time data to save lives.

  • You will build and maintain the infrastructure for the Bayesian platform and develop CI/CD to enable other team members such as software engineers, data scientists, etc. to accelerate their development will drive expansion of our clinical AI/ML module offerings, health system enterprise-wide implementations, and revenue growth.

Who We Are

Bayesian Health’s mission is to improve patient outcomes by empowering clinicians with the insights they need to make the right decision for the right patient at the point-of-care. We’re a diverse team of clinicians, engineers, machine learning experts, product designers, and performance improvement leaders committed to enabling smarter, patient-specific care delivery through unlocking the power of data.

We’re funded by top tier tech and biotech investors: Obvious Ventures, Andreessen Horowitz, American Medical Association’s venture arm, Catalio Partners, and LifeForce Capital. Our company has won many awards; most recent recognitions include: Forbes AI Top 50, World Economic Forum Tech Pioneer, Time Best Inventions, BioTech AI Company of the Year.

Read more about our recent publication in Nature Medicine that associates our products with lives saved.

What you’ll do

As an Infrastructure Engineer, you will build and maintain the networking and infrastructure for the Bayesian platform and develop CI/CD pipelines to enable other team members such as software engineers, data scientists, etc. to accelerate their development. This role is crucial to drive expansion of our clinical AI/ML module offerings, health system enterprise-wide implementations, and revenue growth.

Responsibilities

  • Design cost-optimized, fault-tolerant infrastructure for scale: Propose enhancement to our infrastructure design to enable us to expand our client base and deploy new products on our platform while managing cloud costs and ensuring reliability.

  • Streamline development and deployment: Define a branching and promotion strategy that allow us to comply with the regulatory change control process. Build and maintain CI/CD pipelines using GitHub for automated testing and deployment.

  • Establish and evangelize infrastructure best practices: Create infrastructure guidelines and templates such as Terraform modules, and educate team members in leveraging them.

  • Infrastructure support and maintenance: Continuous monitoring of system performance and reliability, and apply software upgrades accordingly. Collaborate with other team members in troubleshooting infrastructure issues and optimize performance.

  • Secure infrastructure: Partner with SecOps engineer to implement security best practices complying with HIPAA, HITRUST, FDA, and client requirements.

  • AI Ops Platform Architecture: Architect and build a secure, internal AI Ops platform to safely host and manage AI/ML agents for infrastructure and DevOps optimization.

Minimum qualifications

  • 5+ years of experience building and operating production cloud infrastructure on AWS as a DevOps, Infrastructure, Site Reliability Engineer, or similar role.

  • Proficient with Kubernetes, preferably with EKS, including cluster bootstrapping and day-2 ops.

  • Strong operational knowledge of relational databases such as PostgreSQL/MySQL (backups, failover, performance tuning).

  • Deep expertise in Terraform (or equivalent IaC) and an eye for building clean and scalable modules.

  • Familiarity with observability tools, particularly the Datadog

  • Experience building infrastructure with sensitive data that contains PHI/PII.

  • Knowledge of CI/CD pipelines, preferably with CircleCI

  • Excellent communication skills and a proven ability to collaborate with cross-functional teams (e.g., engineering, data science) to translate requirements into robust technical solutions.

  • Experience handling ambiguity and uncertainty in a startup.

Preferred qualifications

  • Experience with using AI agents to optimize infrastructure management or DevOps workflows.

  • Experience with disaster recovery or business continuity plans.

  • Experience with multi account, multi cluster topologies.

  • Experience building systems in healthcare, life sciences, or similarly regulated industries

  • Chaos engineering or game-day facilitation.

  • Experience implementing and maintaining a GitOps framework.

Bayesian Health provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.

This policy applies to all terms and conditions of employment, including recruiting, hiring, placement, promotion, termination, layoff, recall, transfer, leaves of absence, compensation and training.

Apply now >

This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.

Next step

Apply now.

Follow the employer’s application method and review Jobicy’s safety guidance before sharing personal information.

Did you apply?Let us know, and we’ll help you track your application.

Continue on the employer website

Protect your personal information and never pay to secure an interview or job offer. View safety guidance.

Log in to save
One quick step before you apply

Create your free account, then apply.

Build a more organized job search on Jobicy and continue to the employer's application when you're ready.

  • Never lose a promising opportunitySave roles and return to them from your dashboard.
  • See your entire search at a glanceTrack applications, stages and next steps in one place.
  • Get matched with relevant remote jobsChoose the alerts and digests that work for you.
Applying is free. The employer's application opens in a new tab.
Add alert
Jobs Talent AI Tools Salaries
Menu