[All remote jobs](https://jobicy.com/jobs.md)[![Deepgram logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2020/12/WRILS-201212160008-406229.jpg)](https://jobicy.com/company/deepgram.md)[Deepgram](https://jobicy.com/company/deepgram.md)

# Software Test Engineer

Review the role, location requirements, compensation details, and application process before deciding whether this opportunity fits your next career move.

[Apply for this job](#job-application)[View company](https://jobicy.com/company/deepgram.md)ShareRemote from[USA](https://jobicy.com/job-region/usa.md)SalaryUSD 150k–220k / yrDepartment[QA & Testing](https://jobicy.com/categories/qa-testing.md)EmploymentFull TimeExperienceSeniorPublished8 Oct 2026Apply before7 Nov 2026Listing views28Application actions0Application toolkit

## Make your next move.

Prepare your resume, explore your fit, and draft a cover letter for this opportunity.

AI Summary

## The role, at a glance.

Deepgram is hiring a senior-level Software Test Engineer to develop scalable automated testing and evaluation infrastructure for voice AI products, APIs, SDKs, models, and data platforms. The role combines functional, integration, end-to-end, performance, reliability, API, browser, and data-quality testing, with strong emphasis on CI/CD release gates and production validation. The engineer will evaluate speech and other AI-powered behavior using regression coverage, human review, metrics, adversarial datasets, and reproducible test pipelines. Success requires backend scripting expertise, sound statistical judgment, and close collaboration across engineering, research, product, data, infrastructure, and QA. This is a fast-moving, high-autonomy role that expects active experimentation with AI tools and comfort with evolving priorities.

## Role DNA

A quick view of the complexity, pace, ownership and collaboration implied by the job description.

### Job Complexity

5/5EasyHard

### Pace & Pressure

5/5RelaxedFast-paced

### Autonomy Level

5/5GuidedFull ownership

### Communication Load

5/5IndependentCollaborative

AI insightThis role requires more than conventional QA automation: it involves creating test and evaluation systems for AI and real-time voice products while distinguishing meaningful model regressions from statistical noise. The scope spans multiple product surfaces, data workflows, CI/CD, load testing, and cross-functional technical leadership.

## Salary analysis

Estimated compensation compared with the broader US market for similar roles.

Estimated job medianHighly competitive$185,000US market range$145k–$210k0$231k

AI insightThe disclosed annual salary range is $150,000-$220,000 USD, with an offer midpoint of $185,000. For a US-based senior Software Test Engineer focused on AI evaluation, automation infrastructure, and real-time systems, the stated range is competitive and extends above a typical estimated market range of approximately $145,000-$210,000 annually.

## Core skills

Skills and capabilities most closely associated with this opportunity.

[Test automation](https://jobicy.com/jobs?search_keywords=Test%20automation.md)[QA engineering](https://jobicy.com/jobs?search_keywords=QA%20engineering.md)[Python](https://jobicy.com/jobs?search_keywords=Python.md)[API testing](https://jobicy.com/jobs?search_keywords=API%20testing.md)[CI/CD](https://jobicy.com/jobs?search_keywords=CICD.md)[End-to-end testing](https://jobicy.com/jobs?search_keywords=End-to-end%20testing.md)[AI model evaluation](https://jobicy.com/jobs?search_keywords=AI%20model%20evaluation.md)[Performance testing](https://jobicy.com/jobs?search_keywords=Performance%20testing.md)[Data quality](https://jobicy.com/jobs?search_keywords=Data%20quality.md)[Speech recognition](https://jobicy.com/jobs?search_keywords=Speech%20recognition.md)

Sample interview questionsHow would you design an automated regression framework for a speech-to-text API?I would define representative and adversarial audio datasets with versioned ground truth, then build API-level tests that capture transcripts, latency, error rates, and confidence-related outputs. The framework would calculate metrics such as WER by cohort, compare results to approved baselines with statistically appropriate thresholds, and publish results into CI dashboards with clear release-gate criteria.

How do you distinguish a true AI-model regression from normal evaluation noise?

I first ensure the evaluation data, environment, model version, and metric calculation are reproducible. I then analyze the size and consistency of the metric change across relevant cohorts, use confidence intervals or repeated runs where appropriate, inspect qualitative failures, and escalate only when the evidence exceeds agreed thresholds or creates meaningful customer impact.

Describe how you would add quality validation to a CI/CD pipeline for a service with batch and streaming workflows.

I would layer fast unit and contract checks early in the pipeline, followed by integration, API, and targeted end-to-end suites in ephemeral environments. Before deployment, I would run performance and regression gates against representative fixtures; after deployment, I would use canaries, monitoring, and rollback criteria to validate real streaming and batch behavior.

What information should a high-quality bug report contain for a model-powered product?

It should include a concise impact and severity statement, exact reproduction steps, version and environment details, inputs or securely referenced fixtures, parameters, expected and actual outcomes, timestamps, logs, traces, and relevant metric results. For model behavior, I would also identify the affected cohort or failure pattern and note reproducibility across runs.

How would you test data ingestion and processing workflows for downstream model evaluation?

I would validate schema conformance, completeness, integrity, provenance, deduplication, labeling consistency, leakage risks, and representativeness. Automated data-quality gates would block or flag invalid datasets, while reconciliation checks and sampled human review would verify that processed outputs remain fit for evaluation and downstream use.

Opportunity details

## About this role.

### Company Overview

Deepgram is the leading platform underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for speech-to-text (STT), text-to-speech (TTS), and building production-grade voice agents at scale. More than 200,000 developers and 1,300+ organizations build voice offerings that are ‘Powered by Deepgram’, including Twilio, Cloudflare, Sierra, Decagon, Vapi, Daily, Cresta, Granola, and Jack in the Box. Deepgram’s voice-native foundation models are accessed through cloud APIs or as self-hosted and on-premises software, with unmatched accuracy, low latency, and cost efficiency. Backed by a recent Series C led by leading global investors and strategic partners, Deepgram has processed over 50,000 years of audio and transcribed more than 1 trillion words. There is no organization in the world that understands voice better than Deepgram.

### Company Operating Rhythm

At Deepgram, we expect an AI-first mindset—AI use and comfort aren’t optional, they’re core to how we operate, innovate, and measure performance.

Every team member who works at Deepgram is expected to actively use and experiment with advanced AI tools, and even build your own into your everyday work. We measure how effectively AI is applied to deliver results, and consistent, creative use of the latest AI capabilities is key to success here. Candidates should be comfortable adopting new models and modes quickly, integrating AI into their workflows, and continuously pushing the boundaries of what these technologies can do.

Additionally, we move at the pace of AI. Change is rapid, and you can expect your day-to-day work to evolve just as quickly. This may not be the right role if you’re not excited to experiment, adapt, think on your feet, and learn constantly, or if you’re seeking something highly prescriptive with a traditional 9-to-5.

### The Opportunity

Deepgram is looking for a Software Test Engineer to design, build, and maintain automated test frameworks and exploratory test suites across our products, models, APIs, and data platforms. You enjoy breaking systems, probing edge cases, testing real-world and adversarial inputs, and automating repeatable validation so regressions are caught quickly.

You translate product requirements and model metrics into automated regression tests, evaluation pipelines, data-quality gates, load tests, and release criteria. You partner with QA, Research, Product, Data, and Engineering to plan testing, execute human and automated evaluations, support user acceptance testing, and communicate risks clearly.

When you find an issue, you provide precise reproduction steps, inputs, parameters, expected and actual results, and supporting data. What gets you excited? Building scalable automation that gives Deepgram confidence that its products, models, and data workflows work reliably for customers.

### What You’ll Do

*

Define and execute well-designed test plans across Deepgram’s products, APIs, SDKs, model-powered features, and data platforms, ensuring production software is robust, reliable, and performs well.

*

Design, build, and maintain automated test suites and frameworks for functional, integration, end-to-end, regression, API, browser, and service-level testing across batch and streaming workflows.

*

Translate product requirements and customer acceptance criteria into clear test strategies, repeatable test cases, and enforceable release gates.

*

Build and maintain representative, customer-focused, and adversarial test datasets, fixtures, and test environments that exercise real-world inputs, edge cases, failure modes, and system limits.

*

Validate model-powered behavior—including speech-to-text, text-to-speech, and other AI features—using appropriate metrics, expected outputs, human review, and regression coverage, while partnering with Research and model-evaluation specialists as needed.

*

Build testing infrastructure, including test harnesses, reusable scripts, test-data tooling, result-aggregation pipelines, dashboards, and visualizations that make quality signals easy to understand and act on.

*

Integrate automated tests, quality checks, canaries, and release validation into CI/CD so regressions are detected continuously rather than through manual testing alone.

*

Partner with Engineering, Product, Research, Data, Infrastructure, and DevOps to understand system behavior, dependencies, variations, performance limits, and deployment risks, and to establish appropriate test coverage.

*

Test data ingestion, processing, annotation, and quality-control workflows, validating data integrity, completeness, representativeness, deduplication, leakage, and downstream readiness.

*

Execute staging and production validation, load and reliability testing, cross-browser and customer-workflow testing, and user acceptance testing in partnership with internal stakeholders and customer QA teams.

*

Maintain and improve the test-case repository, automation coverage, test documentation, and release-readiness reporting so teams have a clear view of what was tested, what passed, and what remains risky.

*

Write precise, actionable bug reports with reproducible steps, inputs, parameters, expected and actual results, logs or artifacts, and clear severity; participate in triage and escalate issues when necessary.

*

Help raise the bar through code reviews, test-design reviews, technical discussions, and strong engineering, automation, and QA practices.

### What We’re Looking For

*

BS, MS, or PhD in Computer Science, AI, Applied Math, or a related field, or equivalent experience.

*

5+ years of professional software or QA engineering experience, with a track record of shipping test infrastructure or evaluation systems (senior candidates with significantly deeper experience welcome).

*

Solid backend/scripting experience in a language such as Python, Rust, Go, or similar.

*

Experience designing and building automated test pipelines, evaluation frameworks, or data-processing systems.

*

Strong analytical skills and comfort reasoning about metrics, thresholds, and statistical variation in results — able to distinguish real regressions from noise.

*

Ability to take charge of ambiguous technical challenges and communicate effectively across research, engineering, and product teams.

### Nice to Have / Ways to Stand Out

*

Hands-on experience testing or evaluating modern AI systems such as LLMs, RAG pipelines, agents, or multimodal models, including analyzing model behavior and failure modes.

*

Experience with voice, audio, speech recognition, or real-time systems, and familiarity with metrics such as WER, MOS, latency, and time-to-first-byte.

*

Experience building or improving test, evaluation, benchmarking, or ML infrastructure used by multiple teams or external users.

*

A strong appreciation for test and evaluation quality, including correctness, reproducibility, determinism, and consistency across environments.

*

Experience building test tooling for React Native, mobile applications, or other cross-platform environments that extends validation beyond the desktop.

*

Familiarity with cloud infrastructure, containers, ephemeral test environments, CI/CD systems, and monitoring tools such as Grafana, canaries, and anomaly detection.

*

Experience serving as a technical bridge across teams or platforms—including product, QA, evaluation, training, inference, data, or agent frameworks—with the communication skills to build alignment and influence decisions.

*

Prior involvement in open-source projects through contributions, reviews, maintenance, or community engagement.

*

Experience with voice, audio, speech recognition, or real-time systems, and familiarity with metrics like WER, MOS, or latency/TTFB.

*

Prior involvement in open-source projects, through contributions, reviews, maintenance, or community engagement.

*

Experience acting as a technical bridge across teams or platforms (evaluation, training, inference, agent frameworks), combining architectural understanding with clear communication and influence.

*

Familiarity with cloud infrastructure, containerized/ephemeral environments, and monitoring tooling (e.g. Grafana, canaries, anomaly detection).

Notice: We’re aware of individuals impersonating Deepgram recruiters. All legitimate Deepgram recruiting communication comes from an @[deepgram.com](http://deepgram.com) email address. If you’ve received a message claiming to be Deepgram, please forward it to careers@deepgram.com.

Show more

[Apply now >](https://jobicy.com/jobs/154872-software-test-engineer.md)

*

![Upload CV](data:image/svg+xml;base64,PHN2ZyB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciIHdpZHRoPSI2NSIgaGVpZ2h0PSI2NSIgZmlsbD0ibm9uZSIgeG1sbnM6dj0iaHR0cHM6Ly92ZWN0YS5pby9uYW5vIj48ZyBjbGlwLXBhdGg9InVybCgjQSkiPjxwYXRoIGQ9Ik0wIDBINjVWNjVIMFYwWiIgZmlsbD0iIzAyOWFlYiIvPjxnIGZpbGw9IiNmZmYiIHN0cm9rZT0iI2ZmZiIgc3Ryb2tlLXdpZHRoPSIyIj48cGF0aCBkPSJNMzMuMDQ5IDE1LjQ1NGExLjQzIDEuNDMgMCAwIDAtMi4wOTcgMGwtNy41NzkgOC4xNDdhMS4zOCAxLjM4IDAgMCAwIC4wOSAxLjk3MyAxLjQ0IDEuNDQgMCAwIDAgMi4wMDgtLjA4OGw1LjEwOS01LjQ5MnYyMC42MWExLjQxIDEuNDEgMCAwIDAgMS40MjEgMS4zOTdjLjc4NSAwIDEuNDIxLS42MjUgMS40MjEtMS4zOTd2LTIwLjYxbDUuMTA5IDUuNDkyYTEuNDQgMS40NCAwIDAgMCAyLjAwOC4wODggMS4zOCAxLjM4IDAgMCAwIC4wOS0xLjk3M2wtNy41NzktOC4xNDZ6TTE2Ljc2OSAzOC40YzAtLjc3My0uNjItMS40LTEuMzg1LTEuNFMxNCAzNy42MjcgMTQgMzguNHYuMTAybC4yMTUgNi4yMjljLjIyMyAxLjY4LjcwMSAzLjA5NSAxLjgxMyA0LjIxOHMyLjUxIDEuNjA3IDQuMTcyIDEuODMzYzEuNi4yMTggMy42MzYuMjE4IDYuMTYuMjE4aDExLjI4bDYuMTYtLjIxOGMxLjY2Mi0uMjI2IDMuMDYxLS43MDkgNC4xNzItMS44MzNzMS41ODktMi41MzggMS44MTMtNC4yMThDNTAgNDMuMTEzIDUwIDQxLjA1NSA1MCAzOC41MDNWMzguNGMwLS43NzMtLjYyLTEuNC0xLjM4NS0xLjRzLTEuMzg1LjYyNy0xLjM4NSAxLjRsLS4xOSA1Ljk1OGMtLjE4MiAxLjM3LS41MTUgMi4wOTUtMS4wMjYgMi42MTJzLTEuMjI4Ljg1My0yLjU4MyAxLjAzOGMtMS4zOTUuMTktMy4yNDMuMTkzLTUuODkzLjE5M0gyNi40NjJjLTIuNjUgMC00LjQ5OC0uMDAzLTUuODkzLS4xOTMtMS4zNTUtLjE4NC0yLjA3Mi0uNTIxLTIuNTgzLTEuMDM4cy0uODQ0LTEuMjQyLTEuMDI2LTIuNjEyYy0uMTg3LTEuNDEtLjE5MS0zLjI3OS0uMTkxLTUuOTU4eiIvPjwvZz48L2c+PGRlZnM+PGNsaXBQYXRoIGlkPSJBIj48cGF0aCBmaWxsPSIjZmZmIiBkPSJNMCAwaDY1djY1SDB6Ii8+PC9jbGlwUGF0aD48L2RlZnM+PC9zdmc+)

### Upload your resume now

To unlock remote work opportunities and be discovered by global employers.

This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.

Next step

## Apply now.

Follow the employer’s application method and review Jobicy’s safety guidance before sharing personal information.

Keep exploring

## Related remote jobs.

*
![Bayesian Health logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/08/f74593d8-221.png)
Bayesian Health Oct 8  New

### [Senior Design Quality Engineer](https://jobicy.com/jobs/154879-senior-design-quality-engineer.md)

Senior Design Quality Engineer In Brief We’re a rapidly growing startup on a mission to make healthcare proactive by empowering physicians, nurses, and care team members with real-time data to…

*
![Bloomreach logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/08/60790cd5-221.png)
Bloomreach Oct 8  New

### [Quality Assurance Engineer II](https://jobicy.com/jobs/154875-quality-assurance-engineer-ii.md)

Bloomreach is building the world’s premier agentic platform for personalization.We’re revolutionizing how businesses connect with their customers, building and deploying AI agents to personalize the entire customer journey. We’re taking…

*
![StackAdapt logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2022/01/bf5f2bb689c6f35f6cab5459f3eff869.jpg)
StackAdapt Oct 8  New

### [Senior Quality Engineer](https://jobicy.com/jobs/154871-senior-quality-engineer-3.md)

StackAdapt is the leading technology company that empowers marketers to reach, engage, and convert audiences with precision. With 465 billion automated optimizations per second, the AI-powered StackAdapt Marketing Platform seamlessly…

*
![Cribl logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/65d82c04-221.jpeg)
Cribl Oct 8  New

### [Sr. Engineering Manager, SDET](https://jobicy.com/jobs/154877-sr-engineering-manager-sdet.md)

Join the company that’s building the telemetry infrastructure for the AI era. At Cribl, we partner with IT and Security teams at many of the world’s biggest enterprises, including half…

*
![Testlio logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/11024036-221.png)
Testlio Oct 8  New

### [Freelance Software Tester With Smart TVs (Remote in Italy)](https://jobicy.com/jobs/154880-freelance-software-tester-with-smart-tvs-remote-in-italy.md)

Hi there! We are Testlio, a global software testing company with its own freelance network. Our freelancers test apps from companies across the globe. Testlio is a great place for…

*
![PadSplit logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2021/06/da525286873130ed0b95212896ef55d4.jpg)
PadSplit Oct 8  New

### [QA Engineer (Anywhere in Europe)](https://jobicy.com/jobs/154874-qa-engineer-anywhere-in-europe.md)

The Role We Need: PadSplit is hiring a QA Engineer to strengthen quality across our Marketplace pod, spanning Mobile (iOS/Android), Web, and backend Platform systems. As our product complexity increases,…

*
![PerfectServe logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2022/04/7e6f5669fc8c9ea0aa484ce7c3058599.jpg)
PerfectServe Oct 8  New

### [Junior QA Engineer – US Remote](https://jobicy.com/jobs/154878-junior-qa-engineer-us-remote.md)

What is PerfectServe? PerfectServe is a leading provider of clinical communication and physician scheduling solutions in the health IT space. The company was founded in 1997 and has grown steadily…

*
![Varicent logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/c86166a8-221.jpeg)
Varicent Oct 8  New

### [QA Analyst (Mexico, Remote)](https://jobicy.com/jobs/154873-qa-analyst-mexico-remote.md)

At Varicent, we’re not just transforming the Sales Performance Management (SPM) market—we’re redefining how organizations achieve revenue success. Our cutting-edge SaaS solutions empower revenue leaders globally to design smarter go-to-market…

*
![Testlio logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/11024036-221.png)
Testlio Oct 8  New

### [Test a Streaming Platform in France](https://jobicy.com/jobs/152737-test-a-streaming-platform-in-france.md)

Have you ever wished an app worked a little better? At Testlio, we help some of the world’s leading technology and consumer brands improve their apps and digital products by…

*
![Flex logo](https://jobicy.com/data/server-nyc0409/galaxy/mercury/2025/06/7bfdefe3-221.jpeg)
Flex Oct 7  New

### [Senior Software Development Engineer in Test (SDET)](https://jobicy.com/jobs/154770-senior-software-development-engineer-in-test-sdet.md)

Flex is a growth-stage, NYC headquartered FinTech company that is creating the best rent payment experience. It’s hard to believe that it’s 2026 and paying rent on time is expensive,…

[Browse all jobs](https://jobicy.com/jobs.md)