About this role.
Gremlin is hiring a customer-facing Pre/Post Sales Solutions Architect to demonstrate its reliability management platform and guide customers through technical adoption. The role combines pre-sales discovery, demonstrations, webinars, proof-of-concepts, and technical workshops with post-sales architecture consulting and troubleshooting. The successful candidate will have deep hands-on infrastructure expertise across Kubernetes, Linux, cloud platforms, observability, CI/CD, service meshes, and APIs. This US-based remote role supports enterprise teams improving resilience through chaos engineering, reliability testing, and SRE/DevOps practices.
Role DNA
A quick view of the complexity, pace, ownership and collaboration implied by the job description.
Job Complexity
4/5Pace & Pressure
4/5Autonomy Level
4/5Communication Load
5/5Salary analysis
Estimated compensation compared with the broader US market for similar roles.
Core skills
Skills and capabilities most closely associated with this opportunity.
Sample interview questions
I would begin by defining business-critical services, reliability risks, success criteria, stakeholders, and guardrails. I would then select a small number of safe, high-value experiments, establish observability baselines, execute the tests collaboratively, and document findings and remediation actions. The POC would conclude with measurable results and a phased adoption plan tied to the customer's reliability objectives.
I would first map the application's dependencies, deployment topology, resource configuration, network paths, and existing monitoring signals. Next, I would review past incidents and validate hypotheses through controlled experiments such as pod failures, latency injection, or dependency disruption. I would recommend improvements such as health-check tuning, autoscaling, redundancy, timeout and retry policies, resource limits, and runbook updates based on the observed behavior.
I would frame chaos engineering as a controlled method for reducing uncontrolled outage risk. Rather than focusing on the experiment itself, I would connect it to business outcomes such as improved availability, lower incident cost, faster recovery, and stronger customer trust. I would emphasize progressive safeguards, clear rollback procedures, and measurable learning from each test.
I align early on the customer's business problem, technical environment, decision process, and evaluation timeline. During discovery, I ask technical questions that uncover reliability pain points and convert them into a solution narrative, while keeping the account executive informed of risks, stakeholders, and next steps. I also ensure commitments made during the sales process can be successfully delivered after the customer subscribes.
Before an experiment, I would use metrics, logs, traces, and service-level objectives to establish normal behavior and identify the most meaningful failure signals. During the test, I would monitor predefined indicators such as error rate, latency, saturation, and recovery time to determine whether resilience mechanisms behave as expected. The resulting data should drive targeted remediation and confirm whether subsequent experiments demonstrate improvement.
Today’s complex, fast-paced systems have become a minefield of reliability risks—any of which could cause an outage that costs millions and destroys customer confidence. That’s why high-availability teams use Gremlin to find and fix reliability risks before they become incidents. The Gremlin Reliability Platform helps software teams proactively monitor and test their systems for common reliability risks, build and enforce reliability standards, and automate their reliability practices organization-wide. As the industry leader in Chaos Engineering and reliability testing, we work with hundreds of the world’s largest organizations where high availability is non-negotiable.
About the Role of Pre/Post Sales Solutions Architect
Gremlin’s team is growing, and we’re seeking a passionate Solutions Architect to help prove the value of Reliability Management to customers. In this pre- and post-sales role, you will have the opportunity to demonstrate Gremlin Reliability Management and offer guidance on best practices for building reliable architectures. As customers convert to a paid subscription, you will advise on how to design and implement experiments to activate customers for their reliability journey.
In this role, you’ll get to:
- Demonstrate Gremlin in customer calls and webinars
- Partner with sales team to drive technical wins and grow Gremlin’s customer base
- Participate and lead proof-of-concepts with potential customers
- Educate potential customers on Reliability and Chaos Engineering
- Work with existing customers on technical projects and assist in troubleshooting
- Consult with customers on the resiliency of their applications and architecture, diagnose gaps and recommend solutions
- Participate in technical workshops and conferences
Collaborate with different functions of the company including Product Marketing, Support, and Engineering
We’ll expect you to have:
- 5+ years of experience as a Solution Architect in a tech company
- Excellent verbal and written communication skills
- Strong problem-solving skills
- Hands on experience with:
- Kubernetes Platforms, Managed and Unmanaged
- AKS, EKS, GKE
- OpenShift, Rancher
- Certified k8 Administrator is a plus
- Linux – Shell scripting, Certified Linux Administrator is a plus
- Container and Container Runtimes
- Operating Systems concepts (CPU, Memory, and networking)
- Kubernetes Platforms, Managed and Unmanaged
- Working knowledge of :
- Observability solutions – Application Performance Management
- Load Testing solutions (e.g JMeter, LoadRunner, Grafana K6)
- CI/CD and Automation Tools (e.g. Jenkins, Ansible)
- Service Mesh (e.g. Istio), REST APIs and related tools
- Familiarity with one or more programming Languages – Python, Java, Go
- Certification and experience with one or more public cloud providers including Amazon Web Services (AWS), Microsoft Azure, or Google Cloud Platform (GCP)
Bonus experience:
- Experience in a SRE or DevOps role resolving production outages
- Knowledge of modern DevOps and SRE tools
- Integration into ITSM Tools
*The role does not offer sponsorship employment benefits.
**If you don’t think you meet all of the criteria below but still are interested in the job, please apply. Nobody checks every box—we’re looking for candidates that are particularly strong in a few areas, and have some interest and capabilities in others.
Gremlin offers a competitive total rewards package, which includes:
- Base salary
- Equity
- Healthcare, dental, and vision benefits
- 401(k) with employer match.
- Variable compensation for specific roles.
Compensation is based on the candidate’s skills and qualifications.
About Gremlin:
Gremlin is a team of industry veterans and people eager to learn from one another. We set the standard for reliability and equip leading organizations with the mindset and expertise needed to drive reliability improvements that move the world forward. We’re backed by top-tier investors Index Ventures, Amplify Partners, and Redpoint Ventures. Our customers love us, and we’re thrilled to be a partner in their success.
What Do We Care About:
- We Care about our People
People are our critical differentiators. The company strives to treat our people with respect, empathy, and dignity. We expect that our people will treat each other similarly. In both cases, we will assume good intent. All are welcome at Gremlin. We know our differences make us stronger and that our best ideas and contributions can come from anyone at any level.
- We Care about Collaboration
Gremlin is strongest when we come together as one team with shared goals. Be the glue, not the glitter. But as a remote company, teamwork and collaboration won’t happen by accident. We approach every challenge as a shared challenge. We rely on each other for diverse perspectives and creative ideas. We celebrate our wins as a team.
- We Care about Results
Be high productivity, low drama. Results matter. To keep our pace, everyone owns the outcomes of their actions and takes action when needed. We reward speed over perfection. We empower each other to iterate and experiment.
You are welcome at Gremlin for who you are. The more voices and ideas we have represented in our business, the more we will all flourish, contribute, and build a more reliable internet. Gremlin is a place where everyone can grow and is encouraged. However you identify and whatever background you bring with you, please apply if this sounds like a role that would make you excited to come into work everyday. It’s in our differences that we will find the power to keep building a more reliable internet by building and designing tools used by the best companies in the world.
Visit our website to learn more – https://www.gremlin.com/about
Annual salary information is not provided for this position. Explore salary ranges for similar roles in our Salary Directory ›
This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.








