About this role.
Articul8 is hiring a senior Infrastructure Engineer to build and operate scalable, reliable, secure, and cost-effective cloud infrastructure for enterprise GenAI products. The role partners closely with engineering and research teams to integrate infrastructure with product capabilities while improving availability, latency, observability, and deployment processes. Core work includes infrastructure automation, Kubernetes and cloud platform operations, IaC, CI/CD, monitoring, security controls, and performance optimization. The position requires at least seven years of relevant infrastructure, application, or consulting experience and includes technical leadership through code reviews, open-source contribution, and mentoring. The listed location is Brazil/Remote, despite the title referencing India.
Role DNA
A quick view of the complexity, pace, ownership and collaboration implied by the job description.
Job Complexity
5/5Pace & Pressure
4/5Autonomy Level
5/5Communication Load
4/5Salary analysis
Estimated compensation compared with the broader US market for similar roles.
Core skills
Skills and capabilities most closely associated with this opportunity.
Sample interview questions
I would begin by defining workload characteristics, availability targets, latency requirements, data sensitivity, and expected usage patterns. I would design multi-zone services with managed components where appropriate, autoscaling compute and Kubernetes workloads, resilient storage and networking, and clear disaster-recovery objectives. Cost controls would include rightsizing, autoscaling policies, environment lifecycle management, tagging, budget alerts, and ongoing utilization reviews. I would validate the design through load testing, failure testing, security reviews, and operational runbooks.
I use IaC to make environments repeatable, reviewable, and auditable. A strong approach is to create reusable modules, separate environments through controlled variables and state management, enforce pull-request reviews and automated plans, and apply policy checks before deployment. This reduces configuration drift, shortens provisioning time, and gives teams a reliable change history. I also document module interfaces and establish ownership so that shared platform components evolve safely.
I would instrument the platform across metrics, logs, traces, and user-facing service-level indicators. Prometheus and Grafana can provide infrastructure and application metrics, while centralized structured logging through an ELK-style stack supports troubleshooting and auditability. I would define actionable alerts based on symptoms such as error rate, latency, saturation, and availability rather than alerting on every low-level event. Dashboards, runbooks, alert ownership, and regular incident reviews would ensure observability leads to faster recovery and continuous improvement.
I treat security as a paved-road capability rather than a manual gate. I would provide secure defaults through IaC modules, least-privilege IAM templates, secrets management, image scanning, network controls, encryption, and automated policy checks in CI/CD. Teams can move quickly when these controls are easy to consume and failures provide clear remediation guidance. For sensitive GenAI workloads, I would also prioritize data classification, audit logging, access reviews, and threat modeling for new services.
I would use a structured process: establish the user impact and baseline, correlate metrics, logs, traces, and recent changes, then isolate likely constraints through targeted experiments. After implementing a measured fix, I would verify results against the original service-level objective and monitor for regressions. I would document the root cause, update runbooks and alerts, and share the learning with the team. This approach solves the immediate issue while reducing the likelihood and impact of recurrence.
About us:
Articul8 was born from a simple belief: GenAI should work for the enterprise, not the other way around. Our platform — combining domain-specific models, autonomous agentic reasoning (ModelMesh™), reliable model evaluation (LLM-IQ™), and multimodal understanding — serves regulated industries such energy, semiconductor, finance, aerospace, supply chain, and more. Trusted by Fortune 500 enterprises, we bring together research, engineering, product, and domain expertise to deliver AI that meets the accuracy, explainability, and auditability standards that high-stakes environments demand.
Job Description:
Articul8 AI is seeking an exceptional Product/Software Engineer-Infrastructure to join us in shaping the future of Generative Artificial Intelligence (GenAI). We are looking for a Product/Software Engineer-Infrastructure with a proven track record of designing, and building scalable, reliable, and cost-effective clou-based products and solutions. As a member of our Product Technology team, you will play a pivotal role in shaping our infrastructure landscape, ensuring seamless integration, scalability, and reliability across our GenAI-driven products and services. You’ll collaborate with cross-functional teams to drive innovation, optimize performance, and foster growth. This position offers exciting opportunities to work closely with cross-functional teams and external partners to drive innovations in enterprise-grade GenAI.
Responsibilities:
Design, develop, test, deploy, maintain, and improve our cloud-based infrastructure, focusing on high availability, low latency, and cost-effectiveness.
Collaborate closely with engineering and research teams to integrate infrastructure components with product features, ensuring optimal system performance and user experience.
Be the subject matter expert in infrastructure when designing new products and introducing new technology to our existing product line.
Develop automation scripts, tools, and frameworks to streamline deployment, monitoring, and maintenance processes.
Implement robust security measures, adhering to industry standards and best practices, to safeguard sensitive data and prevent potential vulnerabilities.
Analyze performance metrics, identifying areas for optimization and proposing data-driven improvements.
Participate in code reviews, contribute to open-source projects, and mentor junior engineers.
Stay up-to-date with emerging trends, technologies, and methodologies, applying this knowledge to enhance our infrastructure capabilities.
Required Qualifications:
Professional experience: 7+ years of design, implementation, or consulting in applications and infrastructures experience.
Education: BSc degree in Computer Science, Engineering, or a related field.
Preferred Qualifications:
Technology stack:
Cloud Platforms: AWS, Azure, GCP
Infrastructure as Code (IaC): Terraform, CloudFormation, Ansile, etc.
Configuration Management: Ansible, Puppet, Chef
Container Orchestration: Kubernetes, Docker Swarm
Monitoring and Logging: Prometheus, Grafana, ELK Stack (Elasticsearch, Logstash, Kibana).
Programming/ Development: Python, Node.js, Bash, Go, Ruby, CI/CD pipelines, and version control.
Education: Master’s or PhD in Computer science or related technical fields.
Professional Attributes (Code42):
Practice Humility: You ask questions even when you think you know the answer. You seek feedback early, learn from anyone regardless of title, and treat every experiment — especially the failures — as data.
Bias for Outcomes: You measure your work by what changed, not what you tried. You ship results, not slide decks. When a deadline is real, you find a way.
Care Deeply: You treat every problem as yours to solve. You review your own work with the rigor you’d want from a reviewer. You help teammates without being asked.
Dare to Do the Impossible & Embrace Scarcity: You set goals that make you uncomfortable. When told something can’t be done, you find a way or a better question. Constraints sharpen your thinking, not slow it down.
Build a Better World: You believe AI should make things meaningfully better for real people. You hold yourself accountable not just for whether your model works, but for what it does in the world.
Annual salary information is not provided for this position. Explore salary ranges for similar roles in our Salary Directory ›
This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.





