Senior AI Systems Engineer & Software Architect Self-employed
I architected a role-specialised multi-agent platform on an approximately 30,000-line Python runtime, using domain-driven boundaries, clean architectural layering, and strict per-agent tool-surface isolation.
I built and operate a 6-node self-hosted GPU inference fleet with 10 GPUs and roughly 496 GB VRAM, centred on a custom 4× RTX PRO 6000 Blackwell machine assembled by hand.
I also stood up local LLM serving on vLLM and SGLang with tensor parallelism and FP8/NVFP4 quantization, designed a role-replay evaluation harness, performed a REAP expert-prune of Qwen3.5-397B-A17B, and published the pruned model on Hugging Face.