Christopher Wally
Christopher Wally

Senior Software Engineer

Open to offers · Member since 29 Jul 2026
Location
Hubbard, United States
Desired salary
Unspecified
Work preference
Remote Only
Experience level
Senior

About

Professional summary

I am a senior distributed systems engineer with more than nine years of experience building and operating real-time, low-latency infrastructure for high-throughput production systems. My work has included real-time advertising platforms, large-scale recommendation services, and backend APIs.

I specialize in Go-based services, distributed architectures, request routing, feature stores, caching, and performance optimization. I have designed systems that process millions of requests per second while meeting strict latency and reliability requirements.

I have extensive experience with machine learning serving infrastructure, including GPU inference, TensorRT, dynamic batching, model serving, shadow traffic validation, and canary rollouts. I collaborate closely with machine learning and product teams to launch new models safely.

I focus on operational excellence through distributed tracing, latency-percentile monitoring, incident response, capacity planning, autoscaling, and cost optimization. I enjoy diagnosing difficult production issues and improving reliability through practical engineering changes.

I have led migrations to microservices, implemented resilience mechanisms such as circuit breakers and backpressure controls, and improved deployment and testing practices. I also mentor engineers on observability, reliability, and architecture design.

Skills

20 capabilities

Tech stack & tools

Working toolkit

Data Stores

Languages & Frameworks

Monitoring

Experience

Career history

Senior Software Engineer Peer Consulting Resources Inc.

Designed and operated a Go-based real-time bid-response service for ad exchange traffic under strict latency requirements. Improved the request pipeline by overlapping feature retrieval and model inference, reducing p99 latency from approximately 180 ms to 95 ms. Built a TensorRT GPU inference layer with dynamic batching across concurrent model versions, increasing sustained throughput by about 35%.

Implemented multi-region resilience capabilities including request hedging, circuit breakers, regional caching, feature-store fallback behavior, and autoscaling. Resolved a batching-window production incident that reduced GPU inference queue wait by roughly 55%, led shadow-traffic and experiment-cohort rollouts for new model architectures, and scaled the fleet to approximately 2.8 million requests per second. Mentored engineers on distributed tracing and latency guardrails.

Software Engineer Meta

Developed Go services for a real-time recommendation API that selected and ranked candidates for millions of concurrent user sessions. Redesigned regional request routing to reduce average response time from approximately 60 ms to 38 ms, and built a feature store using precomputed embeddings cached at the edge for low-latency model access.

Introduced GPU-backed scoring for a high-traffic ranking model while retaining CPU fallback behavior. Implemented adaptive retry budgets and backpressure controls that lowered failed requests to roughly 0.6%, led a migration from a monolithic deployment to independently deployable microservices, and added distributed tracing and latency dashboards that reduced mean time to diagnose regressions by about 60%. Partnered with data science teams on offline validation and canary deployment processes.

Full Stack Developer Recode

Implemented Go REST API endpoints for an internal order-processing service, defining validation and response contracts used by multiple downstream teams. Improved a PostgreSQL-backed catalog service through schema changes and indexing designed to handle growing read volume while preserving write latency.

Stabilized a nightly batch pipeline through structured logging and retry logic, participated in on-call incident response for the checkout API, and documented findings in postmortems. Automated deployment with a scripted pipeline, reducing average release time from approximately 45 minutes to 20 minutes, and added integration tests for core payment workflows.

Education

Learning history

University of Florida

Bachelor's Degree, Computer Science

This professional hasn’t added portfolio projects yet.

This professional hasn’t listed any services yet.

Jobs Talent AI Tools Salaries
Menu