About this role.
Articul8 AI seeks a senior Infrastructure Engineer to design and operate scalable, reliable, secure, and cost-effective cloud infrastructure for enterprise GenAI products. The role works closely with engineering, research, product teams, and external partners to integrate infrastructure capabilities and improve product performance. Core responsibilities include automation, infrastructure-as-code, deployment, monitoring, security, performance analysis, code reviews, and mentoring. Candidates need at least seven years of relevant infrastructure, application, or consulting experience and a computer science or related degree. Preferred expertise spans AWS, Azure, GCP, Terraform, configuration management, Kubernetes, observability tooling, scripting, and CI/CD.
Role DNA
A quick view of the complexity, pace, ownership and collaboration implied by the job description.
Job Complexity
5/5Pace & Pressure
5/5Autonomy Level
5/5Communication Load
4/5Salary analysis
Estimated compensation compared with the broader US market for similar roles.
Core skills
Skills and capabilities most closely associated with this opportunity.
Cover letter sample
Dear Hiring Team,
I am excited to apply for the Infrastructure Engineer role at Articul8 AI. With extensive experience designing resilient cloud platforms, automating infrastructure delivery, and operating Kubernetes-based systems, I can help build secure, scalable foundations for enterprise GenAI products.
I bring a practical, outcome-focused approach to infrastructure as code, observability, performance optimization, and cross-functional engineering collaboration. I am particularly motivated by Articul8's commitment to reliable, explainable AI in regulated and high-stakes environments.
I would welcome the opportunity to contribute hands-on technical leadership and help advance your infrastructure capabilities in Brazil.
Sample interview questions
I would describe a platform built with multi-zone resiliency, stateless services where possible, automated failover, infrastructure as code, and service-level objectives. I would explain how I used load testing, capacity planning, monitoring, and incident reviews to validate reliability and improve latency.
I would start by defining measurable cost drivers such as compute utilization, storage, network egress, and idle capacity. Then I would use rightsizing, autoscaling, workload scheduling, reserved-capacity strategies where appropriate, and cost dashboards while protecting agreed reliability and performance SLOs.
I would use version-controlled Terraform modules, peer review, policy-as-code, remote state controls, automated plans, and CI/CD gates. I would separate environments, test changes in non-production first, and maintain clear rollback procedures for production deployments.
I would implement layered controls including least-privilege IAM, secrets management, encryption in transit and at rest, network segmentation, image and dependency scanning, audit logging, and continuous vulnerability remediation. For regulated workloads, I would also map controls to applicable compliance requirements and regularly test incident-response processes.
I would define actionable metrics, logs, traces, dashboards, and alerts aligned to user-impacting SLOs rather than alerting on every infrastructure event. I would use tools such as Prometheus, Grafana, and ELK to correlate symptoms across services, establish runbooks, and use post-incident reviews to reduce recurrence.
About us:
Articul8 was born from a simple belief: GenAI should work for the enterprise, not the other way around. Our platform — combining domain-specific models, autonomous agentic reasoning (ModelMesh™), reliable model evaluation (LLM-IQ™), and multimodal understanding — serves regulated industries such energy, semiconductor, finance, aerospace, supply chain, and more. Trusted by Fortune 500 enterprises, we bring together research, engineering, product, and domain expertise to deliver AI that meets the accuracy, explainability, and auditability standards that high-stakes environments demand.
Job Description:
Articul8 AI is seeking an exceptional Product/Software Engineer-Infrastructure to join us in shaping the future of Generative Artificial Intelligence (GenAI). We are looking for a Product/Software Engineer-Infrastructure with a proven track record of designing, and building scalable, reliable, and cost-effective clou-based products and solutions. As a member of our Product Technology team, you will play a pivotal role in shaping our infrastructure landscape, ensuring seamless integration, scalability, and reliability across our GenAI-driven products and services. You’ll collaborate with cross-functional teams to drive innovation, optimize performance, and foster growth. This position offers exciting opportunities to work closely with cross-functional teams and external partners to drive innovations in enterprise-grade GenAI.
Responsibilities:
Design, develop, test, deploy, maintain, and improve our cloud-based infrastructure, focusing on high availability, low latency, and cost-effectiveness.
Collaborate closely with engineering and research teams to integrate infrastructure components with product features, ensuring optimal system performance and user experience.
Be the subject matter expert in infrastructure when designing new products and introducing new technology to our existing product line.
Develop automation scripts, tools, and frameworks to streamline deployment, monitoring, and maintenance processes.
Implement robust security measures, adhering to industry standards and best practices, to safeguard sensitive data and prevent potential vulnerabilities.
Analyze performance metrics, identifying areas for optimization and proposing data-driven improvements.
Participate in code reviews, contribute to open-source projects, and mentor junior engineers.
Stay up-to-date with emerging trends, technologies, and methodologies, applying this knowledge to enhance our infrastructure capabilities.
Required Qualifications:
Professional experience: 7+ years of design, implementation, or consulting in applications and infrastructures experience.
Education: BSc degree in Computer Science, Engineering, or a related field.
Preferred Qualifications:
Technology stack:
Cloud Platforms: AWS, Azure, GCP
Infrastructure as Code (IaC): Terraform, CloudFormation, Ansile, etc.
Configuration Management: Ansible, Puppet, Chef
Container Orchestration: Kubernetes, Docker Swarm
Monitoring and Logging: Prometheus, Grafana, ELK Stack (Elasticsearch, Logstash, Kibana).
Programming/ Development: Python, Node.js, Bash, Go, Ruby, CI/CD pipelines, and version control.
Education: Master’s or PhD in Computer science or related technical fields.
Professional Attributes (Code42):
Practice Humility: You ask questions even when you think you know the answer. You seek feedback early, learn from anyone regardless of title, and treat every experiment — especially the failures — as data.
Bias for Outcomes: You measure your work by what changed, not what you tried. You ship results, not slide decks. When a deadline is real, you find a way.
Care Deeply: You treat every problem as yours to solve. You review your own work with the rigor you’d want from a reviewer. You help teammates without being asked.
Dare to Do the Impossible & Embrace Scarcity: You set goals that make you uncomfortable. When told something can’t be done, you find a way or a better question. Constraints sharpen your thinking, not slow it down.
Build a Better World: You believe AI should make things meaningfully better for real people. You hold yourself accountable not just for whether your model works, but for what it does in the world.
Annual salary information is not provided for this position. Explore salary ranges for similar roles in our Salary Directory ›
This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.









