Qiang Zhang
Qiang Zhang

Staff AI Engineer

Actively looking · Member since 29 Aug 2026
Message
Location
Burlingame, United States
Desired salary
Unspecified
Work preference
Remote Only
Experience level
Lead

About

Professional summary

I am a Staff AI Engineer and Member of Technical Staff with more than 11 years of experience spanning foundation models, LLM post-training, model evaluation, distributed systems, and production AI platforms.

I specialize in improving frontier model quality across reasoning, coding, instruction following, structured generation, tool use, and API-oriented workflows. My work combines AI research rigor with practical engineering for reliable developer-facing products.

At OpenAI, I have contributed to Structured Outputs, o3-mini post-training, and GPT-4.1 research. I built large-scale evaluation infrastructure and research tooling that accelerated benchmark execution, regression analysis, failure investigation, and model-validation cycles.

Previously, I led architecture and reliability initiatives for high-volume distributed backend platforms at Stripe. I have extensive experience with highly available services, performance optimization, asynchronous processing, databases, observability, and cloud-native infrastructure.

I also bring applied machine learning experience from real-time dispatch and ETA prediction systems. I enjoy building scalable AI systems that connect research findings, robust evaluation, efficient deployment, and measurable product impact.

Skills

50 capabilities

Tech stack & tools

Working toolkit

Experience

Career history

Member of Technical Staff – Foundation Models & AI Systems OpenAI

Drive post-training and API model research for frontier foundation models across reasoning, coding, instruction following, and structured and tool-based generation. Contributed to Structured Outputs, o3-mini training and post-training, and GPT-4.1 research.

Architected evaluation infrastructure processing more than 100K benchmark samples and millions of generated tokens. Built Python tooling for experiment orchestration, benchmark execution, result aggregation, regression analysis, and failure analysis, reducing validation cycles by about 50% and manual evaluation effort by about 40%.

Staff Software Engineer – Distributed Systems & Platform Engineering Stripe

Led technical architecture for high-volume backend services, improving platform scalability by approximately 35% through service decomposition, workload balancing, and performance engineering. Directed architecture reviews and cross-team technical decisions for production services.

Optimized distributed services using Go, Java, Kafka, PostgreSQL, Redis, and Kubernetes. Introduced monitoring, automated recovery, and failure-isolation mechanisms that lowered recurring incidents while improving processing capacity and implementation efficiency.

Senior Software Engineer – Distributed Backend Systems Stripe

Delivered highly available backend services supporting millions of daily payment transactions with 99.99% availability across critical workloads. Improved service throughput through concurrency optimization, service decomposition, workload balancing, and backend performance tuning.

Refined database schemas, indexing, query execution, and caching strategies to reduce API latency. Expanded observability across metrics, logging, and distributed tracing, while strengthening automated testing, CI/CD, deployment validation, capacity planning, and reliability practices.

Software Engineer – Backend Infrastructure Stripe

Implemented backend services and APIs for payment workflows and internal platforms, increasing transaction-processing throughput. Created reusable service libraries and platform components that reduced implementation effort across recurring backend capabilities.

Tuned SQL queries, indexes, database access, and caching paths for high-traffic services. Introduced automated testing and continuous deployment practices that reduced production incidents and improved release consistency.

Data Scientist – Machine Learning Systems Luxe

Developed machine learning solutions for real-time dispatch and ETA prediction systems, improving prediction quality and operational efficiency. Applied statistical modeling and optimization to routing, resource allocation, and fleet-utilization problems.

Established end-to-end ML workflows covering data preparation, feature engineering, training, validation, and production integration. Analyzed prediction errors and operational outcomes and partnered with engineering and operations teams to integrate predictive models into real-time workflows.

Education

Learning history

University of California, Los Angeles (UCLA)

Master's Degree, Statistics

Nankai University

Bachelor's Degree, Math & Statistics

This professional hasn’t added portfolio projects yet.

This professional hasn’t listed any services yet.

People also viewed

All talent ›
Jobs Talent AI Tools Salaries
Menu