About this role.
Testlio is seeking a senior, client-facing AI Solutions Architect to define and deliver AI testing services for enterprise customers. The role combines hands-on AI/ML and software engineering expertise with ownership of service packaging, pricing input, evaluation methodology, and technical delivery. The architect will design testing strategies for LLMs, RAG systems, autonomous agents, and human-in-the-loop evaluation programs while supporting sales discovery and delivery escalations. Success requires translating complex, non-deterministic AI architectures into measurable quality criteria and clear business value. This is a fully remote role restricted to U.S. residents.
Role DNA
A quick view of the complexity, pace, ownership and collaboration implied by the job description.
Job Complexity
5/5Pace & Pressure
4/5Autonomy Level
5/5Communication Load
5/5Salary analysis
Estimated compensation compared with the broader US market for similar roles.
Core skills
Skills and capabilities most closely associated with this opportunity.
Sample interview questions
I would first define the business-critical user journeys, failure modes, and acceptable risk thresholds. I would then create a representative, versioned evaluation dataset and measure retrieval quality, groundedness, answer correctness, safety, latency, and task completion using deterministic checks, LLM judges, and targeted human review. Finally, I would establish release gates, production monitoring, and a feedback loop that turns observed failures into new regression cases.
I would model the agent as a system rather than test only final responses. This includes validating planning behavior, tool-selection accuracy, permission boundaries, state handling, recovery from failed actions, hallucinated tool calls, and harmful or costly loops. I would use scenario-based simulations, seeded edge cases, trace analysis, and human assessment for ambiguous outcomes, with explicit metrics for completion, policy compliance, cost, and escalation quality.
LLM judging is useful for high-volume, repeatable assessments where rubric quality can be calibrated against trusted human labels. I would use humans for nuanced domain decisions, safety-sensitive judgments, novel failure analysis, multilingual or cultural context, and validation of the judge itself. The strongest approach is typically a calibrated hybrid: automation handles scale while human experts adjudicate uncertainty and continuously improve the rubric and datasets.
I would avoid leading with model internals and instead explain the affected business workflow, likelihood, customer impact, and estimated cost of failure. I would present the evidence, such as scenario results and trend data, alongside a prioritized remediation plan and the decision required from leadership. This makes technical risk actionable while preserving enough detail for engineering teams to implement the solution.
I start with validated customer problems, target personas, outcomes, and differentiation rather than technology alone. I define clear service tiers, scope boundaries, staffing assumptions, deliverables, success metrics, and pricing logic, then pilot the offering with delivery and sales feedback. After measuring repeatability, customer value, margins, and implementation friction, I refine the playbook and enable teams with templates, training, and positioning materials.
Location: This is a 100% remote role open to those who reside in the United States.
About the job
Testlio’s fully managed crowdsourced testing platform, powered by our proprietary intelligence technology – LeoCore™, integrates expert, on-demand testers directly into your release process. Ship faster and more confidently with global coverage across 600,000+ devices, 800+ payment methods, 150+ countries, and 100+ languages. To learn more, visit testlio.com.
We are hiring a visionary and hands-on AI Solutions Architect to own the definition, packaging, technical design, and delivery methodology for our AI testing offerings. This role brings together deep practitioner expertise with a strong business perspective. You will design the testing approach for complex AI-powered applications (including agentic systems), and you will advance Testlio’s proprietary framework combining automated and human-in-the-loop evaluation while defining what Testlio sells as a service—setting the catalog, staffing models, and pricing structure. As the technical authority, you will decompose architectures into structured testing approaches and collaborate with sales and delivery to drive our AI business forward.
Why will you love this job?
- Evaluate at a Scale That Exists Nowhere Else: Every vendor has LLM-as-a-judge. Only Testlio can put thousands of vetted experts across 150 countries and 100+ languages behind a grading rubric. You will design evaluation systems that combine model graders and human experts.
- Innovate at the Frontier of QA: Shape the industry standards and operational playbooks for validating agentic AI, large language models, and the data pipelines behind them.
- High-Impact Consultative Visibility: Collaborate with global market leaders across diverse verticals. Interface directly with data scientists, machine learning engineers, and executive sponsors at market-leading global brands to architect their AI quality blueprints.
- Decompose Complex Ecosystems: Turn “unknown unknowns” into structured testing criteria: what the system must do, how it can fail, and what a failure costs.
- Drive company influence: Help build the pipeline, join sales calls and guide delivery on AI Testing opportunities, bringing your expertise to high-impact business decisions.
Why will you love being a part of Testlio?
- Great Culture: Testlio is a female-founded company, and half of our team identifies as women. We’re proud of our inclusive, purpose-driven culture where people genuinely enjoy collaborating. As part of our team (we call ourselves TestLions), you’ll help create exceptional digital experiences for our customers, while also contributing to our freelance network.
- Remote Work: Our culture is built around remote work. We’ve created systems to allow us to successfully work together asynchronously as a fully remote and globally distributed team. Testlio provides the tools and guidance for everyone to succeed in their careers in a fully remote setting. Our working style encourages everyone to make decisions, communicate effectively, and work at a sustainable pace.
- Investment in You: Your growth and well-being matter to us. You’ll have flexible paid time off—including national holidays, personal days, and sick days—plus stock options so you can grow with Testlio. We also provide a $300 annual learning stipend to support your personal and professional development.
- Winning Business: Testlio is growing, profitable, and cash-strong. We are leading our industry with exceptional clients who provide us with a high NPS score and a 4.7 rating on G2. Our business model is global, enterprise, and subscription-based. Several of our largest clients have been with us for 7+ years, and many spend $500K+/year with Testlio.
What would your day look like?
Strategic & Commercial Responsibilities
- Define and maintain the AI service catalog: offerings, tiers, packaging, and delivery methodology. Track, analyze, and present key metrics to stakeholder leadership as proof of quality, efficiency, and ROI.
- Vision Activation & Marketing: Imagine, drive, evangelize, and activate a viewpoint on the future of AI Testing. Produce white papers, articles, posts, speeches, and webinars.
- Sales Collaboration: Collaborate with sales and engagement teams to support pre-deal discovery, scoping, and articulating Testlio’s AI testing value proposition.
- Offering Lifecycle Management: Use market and competitive intelligence to refine offerings and supply Product Marketing with details for accurate positioning.
Collaboration on Hands-on Delivery
- System & Architectural Deconstruction: Execute comprehensive system-level breakdowns of client AI stack components—including autonomous agents, system prompts, foundation models, and RAG architectures—to formulate a tailored AI testing approach, inform test strategy, design evaluations, and provide technical advisory.
- Delivery Expertise: Act as an escalation point and subject-matter expert for Delivery, stepping in hands-on from scoping to review during complex or new client engagements.
- Technical Problem Solving: Solve technical roadblocks, stand up environments, and work on solutions that make AI testing and evaluation scalable and repeatable in strategic engagements.
- Framework Advancement: Advance Testlio’s proprietary AI testing framework by combining deterministic code graders, LLM-as-a-judge, and human-in-the-loop validation.
- Empower Internal Teams: Develop templates, curated training datasets, and comprehensive documentation to drive seamless, high-quality test execution.
What do you need to succeed?
Technical Skills
- Hands-on Data Science Background: Extensive practical experience developing, training, or fine-tuning machine learning models, NLP structures, or data pipelines.
- AI Evaluation Expertise: Deep familiarity and hands-on experience with modern LLM evaluation frameworks (e.g., DeepEval, RAGAs, etc) and metric-driven validation ecosystems.
- Agentic Framework Proficiency: Direct experience working with, deploying, or testing autonomous agent architectures, planning patterns, or multi-agent swarms (e.g., LangGraph, AutoGen, CrewAI).
- Strong Software Engineering Foundations: Strong software engineering foundations, including proficiency in dealing with complex non-deterministic and deterministic systems.
Human Skills
- Proven Client-Facing Experience: Solid track record successfully leading discovery calls, managing stakeholder workshops, and consulting on real-world enterprise architectures.
- Commercial & Strategic Acumen: Experience designing service offerings, understanding pricing dynamics, and formulating go-to-market strategies within professional services or SaaS.
- Consultative Communication: Exceptional ability to bridge technical data science outputs with clear, strategic customer business value, speaking confidently with both engineers and executives.
- Comfort with Ambiguity: High adaptability to navigate fast-evolving client layouts, unvetted sources, and the non-deterministic characteristics of AI application lifecycles.
- Mentorship Mindset: A passion for continuous learning, documentation, and sharing architectural best practices to uplift internal team members.
What is the application process?
At Testlio, we aim to hire individuals who are excited about their role, thrive in a fully remote environment, and have strong long-term potential with our team. Because we are a fully distributed company, our interview process includes conversations with several team members so you can get a well-rounded understanding of the role, the people you’ll work with, and how we collaborate. As a result, our interview process typically takes 3–4 weeks to complete.
Interview Process:
- Application
- Recruiter interview
- TestGorilla assessment
- ~ 4 Team and Stakeholder interviews
- Reference checks
- Offer & background check
Diversity and Inclusion
Testlio is an equal-opportunity employer deeply committed to creating an inclusive environment for people of all backgrounds and identities. We are female-founded, and half of our team members identify as women. For more information, see the DEI section of our website.
#LI-Remote
Annual salary information is not provided for this position. Explore salary ranges for similar roles in our Salary Directory ›
This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.





