Abhishek Jaiswal
Abhishek Jaiswal

AI Tech Lead

Open to offers · Member since 4 Oct 2026
Location
Bangalore, India
Desired salary
Unspecified
Work preference
Remote Only
Experience level
Lead

About

Professional summary

I am an AI Tech Lead specializing in production-grade large language model systems, agentic AI, retrieval-augmented generation, and model optimization. I lead the design and delivery of scalable AI platforms that improve reasoning quality, safety, latency, and operating cost.

I have extensive experience fine-tuning open-source language and multimodal models using QLoRA, LoRA, RLHF, DPO, and GRPO. My work includes reasoning distillation, preference optimization, quantization, distributed training, and high-throughput inference for enterprise applications.

I build multi-agent systems using LangGraph, CrewAI, LangChain, MCP servers, and custom orchestration layers. I focus on dependable enterprise capabilities such as prompt-injection protection, tenant isolation, policy enforcement, PII redaction, secure code sandboxes, and guardrail services.

I am experienced in MLOps and cloud infrastructure across AWS, GCP, Kubernetes, Terraform, MLflow, Docker, and GPU-serving stacks. I have delivered systems handling hundreds of thousands of daily requests while improving uptime, model deployment velocity, and inference economics.

My background also includes full-stack machine learning engineering with Python, FastAPI, React, TypeScript, Node.js, and REST APIs. I enjoy translating advanced AI research into measurable, production-ready products for financial intelligence, document automation, coding assistance, and private enterprise AI.

Skills

33 capabilities

Tech stack & tools

Working toolkit

Application Hosting

Languages & Frameworks

Libraries

Experience

Career history

AI Tech Lead Trianz

Led LLM reasoning distillation with DPO and GRPO optimization using Unsloth and Qwen3, achieving a 78% improvement on GSM8K mathematical reasoning benchmarks through a custom DSPy evaluation framework. Fine-tuned Gemma-4 for enterprise office-document generation using QLoRA and GRPO, reducing inference latency by 55% and cloud costs by 65% across more than 200,000 monthly automation requests.

Architected security, safety, and multi-tenant infrastructure for the Concierto multi-agent platform, including prompt-injection detection, MCP isolation, policy enforcement, and extensible guardrails. Led an enterprise GraphRAG coding assistant and dynamic MCP server platform, improving code retrieval accuracy, lowering developer resolution time, and reducing token and context overhead across 50+ production enterprise deployments.

Senior AI Engineer Streamingo.ai

Architected LLM serving infrastructure on GCP Vertex AI using vLLM and TensorRT-LLM, deploying 13B-parameter models at sub-200ms latency for over 500,000 requests per day. Delivered a 40% cost reduction compared with hosted APIs while maintaining 99.8% uptime over six months.

Orchestrated distributed training with DeepSpeed and FSDP on AWS SageMaker GPU clusters, reducing training time by 65% and scaling to 34B+ parameter models. Modernized the ML platform with Kubernetes, MLflow, and Terraform, reducing model deployment cycles from three weeks to two days and enabling significantly faster experimentation.

ML Engineer Deqode

Architected multi-agent financial-research systems using LangGraph and CrewAI, automating more than eight data sources and processing over 200,000 tasks monthly with 92% task completion accuracy. Built hybrid financial RAG solutions combining BM25 sparse retrieval and dense vector search for low-latency, high-accuracy answers across more than 500,000 financial documents.

Fine-tuned Qwen3 8B models on domain-specific financial corpora using QLoRA, NF4 quantization, and iterative RLHF, substantially improving entity extraction and reducing hallucinated figures. Deployed models through ONNX and TensorRT INT8 quantization, achieving 3.2x throughput improvement, sub-90ms p99 latency, and more than 500,000 daily inferences at lower cost.

Full Stack Engineer Humanix Technology

Developed and deployed classical machine learning models in Python using scikit-learn, XGBoost, and Random Forest for classification and regression applications. Built RNN and LSTM NLP models in PyTorch, including preprocessing, embedding, training, and evaluation pipelines.

Engineered full-stack ML applications with React, TypeScript, Node.js, Python, and REST APIs to expose real-time predictions in user workflows. Containerized microservices with Docker and established CI/CD pipelines with automated testing and zero-downtime deployments across staging and production.

Education

Learning history

Rajiv Gandhi Proudyogiki Vishwavidyalaya

Bachelor of Technology (B.Tech), Computer Science

Bhopal, India. CGPA: 8.8/10. Also completed an IIT Madras Elite certification in Programming, Data Structures & Algorithms.

This professional hasn’t added portfolio projects yet.

This professional hasn’t listed any services yet.

Jobs Talent AI Tools Salaries
Menu