Abilmansur Ashim
Abilmansur Ashim

ML / NLP Engineer

Open to offers · Member since 24 Sep 2026
Location
Almaty, Kazakhstan
Desired salary
2500–6000 USD/yearly
Work preference
Remote Only / Full Time
Experience level
Senior

About

Professional summary

I am an ML / NLP Engineer with over four years of experience building production-grade NLP, RAG, and LLM-agent systems. I work across the full machine learning lifecycle, from data preparation and evaluation design to inference optimization, deployment, monitoring, and iterative improvement.

I specialize in retrieval-augmented generation systems, including hybrid search with dense retrieval, BM25, reciprocal rank fusion, cross-encoder reranking, metadata filtering, and source-grounded answers. I have built document ingestion and indexing pipelines for technical documentation, tickets, reports, and knowledge bases.

I have hands-on experience fine-tuning Transformer models for text classification, named entity recognition, and domain retrieval. My work includes hard-negative mining, quality evaluation using Recall@10 and MRR@10, threshold calibration, quantization, and low-latency CPU inference.

I build reliable LLM agents with controlled tool calling, structured outputs, RBAC validation, human approval workflows, retries, fallbacks, and audit trails. I focus on making AI systems safe, observable, measurable, and useful in real production environments.

I use Python and modern ML tooling including PyTorch, Hugging Face, FastAPI, Qdrant, PostgreSQL, Redis, Docker, Kubernetes, and AWS. I am based in Almaty, Kazakhstan, and communicate professionally in Russian, Kazakh, and English.

Skills

20 capabilities

Tech stack & tools

Working toolkit

Analytics

Application Hosting

Data Stores

Languages & Frameworks

Libraries

Monitoring

Experience

Career history

ML / LLM Engineer Neuro.net

Develop and deploy an agentic support assistant that combines query routing, hybrid RAG, CRM and service-desk tool calling, and human approval for high-impact operations. The system automates approximately 35% of standard support requests.

Built hybrid retrieval over documentation and tickets using dense bi-encoder search, BM25, reciprocal rank fusion, cross-encoder reranking, ACL metadata filtering, and answer citations. Implemented controlled stateful workflows with Pydantic validation, RBAC checks, idempotency, retries, fallbacks, and audit trails.

Established offline evaluation using golden datasets, Recall@10, MRR@10, groundedness, answer relevance, escalation rate, and tool correctness. Delivered Dockerized services with CI/CD, health checks, and Prometheus/Grafana monitoring for latency, errors, token usage, caching, and retrieval quality.

NLP Engineer GRONICS

Built an end-to-end RAG assistant providing natural-language search and question answering for engineers and support teams across technical documentation, procedures, incident reports, and resolved cases. The production service handled approximately 12,000 requests per month.

Created reproducible ingestion and indexing pipelines for PDF, DOCX, wiki, and ticket data, including structure-aware extraction, cleaning, semantic chunking, metadata enrichment, deduplication, and incremental updates. Fine-tuned a domain bi-encoder with mined hard negatives, improving Recall@10 from 0.68 to 0.86 and MRR@10 from 0.54 to 0.74.

Implemented hybrid dense and BM25 retrieval with RRF and cross-encoder reranking. Deployed FastAPI services using Qdrant, PostgreSQL, Redis, and Docker, reducing median retrieval and generation latency from 3.8 seconds to 1.7 seconds through caching, batching, and asynchronous processing.

ML / NLP Engineer Fortech

Trained and deployed Transformer models for multi-class customer-request routing and named entity recognition from documents. Reduced manual routing of incoming requests by 42%.

Built reproducible training pipelines with data cleaning, temporal train/validation/test splits, class-imbalance handling, hyperparameter search, MLflow experiment tracking, and confusion-matrix error analysis. Improved Russian-language Transformer macro-F1 from 0.79 to 0.89 through mixed precision, dynamic padding, gradient accumulation, early stopping, and per-class threshold calibration.

Optimized CPU model inference by exporting models to ONNX, applying INT8 dynamic quantization, micro-batching, and FastAPI serving. Reduced p95 latency from 260 ms to 78 ms with less than one percentage point of quality loss, while adding Docker-based deployment, logging, Prometheus/Grafana monitoring, and drift checks.

Education

Learning history

Kazakh-British Technical University

M.Sc., Data Science (Computational Neuroscience)

Master of Science in Data Science with a specialization in Computational Neuroscience. Located in Almaty, Kazakhstan.

International Information Technology University

B.Sc., Data Science

Bachelor of Science in Data Science. Located in Almaty, Kazakhstan.

This professional hasn’t added portfolio projects yet.

This professional hasn’t listed any services yet.

Jobs Talent AI Tools Salaries
Menu