AI/LLM Evaluation Analyst Welocalize / Welo Data
I have evaluated more than 1,000 outputs for a confidential frontier AI lab program across increasingly complex evaluation phases. My work began with selecting best responses and identifying grammar and spelling issues in real user-model conversations.
I conduct multilingual search evaluation by authoring themed prompts, testing them across four models, and scoring source reliability, localization, and time correctness. I also design flawed-response prompts, step-by-step guided instructions, and rubrics for self-correction testing, while evaluating AI speech for naturalness, intelligibility, pronunciation, transcript alignment, and linguistic quality in Spanish and English.