All remote jobs
Open role
Remote opportunity atAble

AI Senior Engineer (Vision)

Review the role, location requirements, compensation details, and application process before deciding whether this opportunity fits your next career move.

Published
63Listing views
7Application actions
20 Oct 2026Apply before
Opportunity details

About this role.

AI Summary

This Senior AI Engineer role builds document-intelligence systems that combine computer vision, PDF parsing, and LLM orchestration. The engineer will develop pipelines for extracting structured information from complex visual documents, including layouts, charts, and tables. Core work includes LangChain or LangGraph agent workflows, multimodal model integration, prompt/control-flow design, and cost optimization. Strong Python ML experience, native PDF-processing expertise, and clear English communication are required. The role is fully remote for candidates within Latin America and expects a 40-hour workweek.

Role DNA

A quick view of the complexity, pace, ownership and collaboration implied by the job description.

Job Complexity

5/5
EasyHard

Pace & Pressure

4/5
RelaxedFast-paced

Autonomy Level

5/5
GuidedFull ownership

Communication Load

4/5
IndependentCollaborative
AI insightThis is a highly specialized senior role spanning multimodal AI, document parsing, agentic orchestration, and production cost/performance tradeoffs. Success requires independently solving ambiguous reliability challenges such as preserving PDF structure and reducing hallucinations on financial charts.

Salary analysis

Estimated compensation compared with the broader US market for similar roles.

Estimated job medianMarket rate
$190,000
US market range$165k–$230k
AI insightNo actual salary was disclosed; "Payments made in USD" only establishes payment currency. The figures shown are estimated annual US-market base-salary benchmarks for a senior AI/computer-vision engineer with LLM orchestration and document-intelligence expertise; actual LATAM compensation may differ substantially by country, engagement model, and benefits.

Core skills

Skills and capabilities most closely associated with this opportunity.

Sample interview questions
How would you design a pipeline to extract and interpret data from a complex financial PDF containing text, tables, charts, and mixed layouts?

I would first classify the document and extract native PDF elements with PyMuPDF or pdfplumber wherever possible, preserving page coordinates and reading order. For scanned or highly visual regions, I would use layout detection and targeted vision-language model calls, then normalize outputs into a schema with provenance links to page regions. I would validate critical fields using confidence thresholds, cross-checks, and human-review fallbacks for low-confidence financial values.

How do you manage context limits and reliability in LangChain or LangGraph agent workflows?

I define narrowly scoped graph nodes, structured inputs and outputs, and explicit state so each step has a clear responsibility. Rather than passing entire documents into a model, I retrieve only relevant chunks, table representations, and image regions while retaining metadata for traceability. I also implement retries, tool-result validation, guardrails, and evaluation datasets to identify failure modes before production release.

What techniques would you use to reduce hallucinations when models interpret charts or tables?

I would ground the model on extracted source data and require structured answers that cite page, chart, or table coordinates. I would use deterministic parsing for values when available, reserve visual models for interpretation tasks, and prompt the model to return uncertainty rather than infer missing information. Automated checks such as totals reconciliation, unit validation, and comparison against OCR or native text provide additional safeguards.

Describe how you would optimize the cost of a high-volume document-processing system using multimodal models.

I would segment documents and route each task to the least expensive model capable of meeting quality targets. Native text extraction, caching, batching, asynchronous processing, and selective high-resolution image rendering can greatly reduce expensive vision calls. I would measure per-document cost, latency, and extraction quality by document type, then use those metrics to tune routing thresholds and model choices.

How would you evaluate whether fine-tuning is preferable to prompt engineering for domain-specific financial documents?

I would begin with a representative benchmark set and establish a prompt-and-retrieval baseline using error categories such as chart misreads, layout mistakes, and incorrect field extraction. If recurring errors persist despite improved data representation, prompts, and validation, I would assess whether sufficient labeled examples exist for fine-tuning. The decision would compare measurable quality gains against annotation, training, maintenance, inference cost, and the ability to adapt to new document formats.

This analysis is generated from the job description. Salary estimates, role characteristics and sample answers are guidance, not employer-provided facts.

Back in 2012, we were a group of engineers and designers who decided we wanted to build things—so we did. Able started as an engineering and product hub building for a portfolio of early-stage startups. We built many relationships while developing products that were thoughtful, effective, and genuinely useful. But, since then, we’ve grown… and so has our ambition.

Now, we’re entering our next chapter—defined by applied AI. AI is a powerful force in the end-to-end software development cycle, and we’re creating practices that allow us to deliver software fast and more effectively than traditional approaches, creating meaningful value for our partners. Today, our builder mindset is driving us to become an AI-native organization across every function. We’re still evolving, and that’s part of the opportunity. If you want to build, learn, and tackle challenges alongside an ambitious team, let’s build together.

This position is 100% remote within LatAm.

What you’ll be doing

We are seeking someone who enjoys working at the cutting edge where Computer Vision meets Logic. You will be responsible for the “eyes” and the “brain” of our system—extracting complex data from visual documents and then orchestrating how that data is used by Large Language Models.

In short, someone who likes:

  • Unlocking Visual Data: Building pipelines that can “read” complex documents, understanding layout, charts, and visual context using Vision-Language Models (GPT-4V, Claude 3.5) and Layout Analysis.
  • Orchestrating Intelligence: Owning the application logic layer. You will use LangChain or LangGraph to build the agents and chains that query our data, reason about it, and generate responses.
  • Native PDF Handling: Handling the messy reality of PDF processing (PyMuPDF, layout parsing) to preserve structure before the AI even sees it.
  • Prompt Engineering & Logic: Crafting complex prompts and control flows to ensure models interpret financial charts and layouts accurately without hallucinating.
  • Cost & Scale: Applying a cost-optimization mindset (batch processing, model selection) to ensure our vision and orchestration layers are economically viable.

What we’re looking for

We want to work with people who have a passion for collaborating with their teams, building software while nurturing inclusive and respectful relationships with their coworkers. With the ones that are open about their shortcomings and what they do not know now, but remain eager to keep on growing and closing those gaps.

Ideally, they would also have:

  • LLM Orchestration (Must Have): Deep experience with LangChain, LangGraph, or similar frameworks. You know how to manage context windows, tool calling, and agentic workflows.
  • Multimodal AI Experience: Hands-on experience integrating state-of-the-art vision models (GPT-4V, Claude 3.5 Sonnet) and embedding models (CLIP).
  • Document Intelligence Specialist: Familiarity with specialized models (e.g., Donut, Pix2Struct) and tools like Unstructured.io or Docling.
  • PDF Processing Mastery: Mastery over tools like PyMuPDF or pdfplumber for native element extraction.
  • Python ML Stack: Strong proficiency in PyTorch or TensorFlow.

Nice-to-Have:

  • Fine-Tuning: Experience fine-tuning vision or language models, specifically to improve accuracy on domain-specific artifacts like financial charts or tables.
  • Domain Knowledge: Prior experience handling documents in the Real Estate or Finance sectors.

Able is powered by curious, thoughtful people who care about what they build and how they build it. We’re actively investing in our team through AI training, knowledge-sharing, and hands-on experimentation to ensure everyone grows alongside the technology.

This position is 100% remote within LatAm. Strong verbal and written communication skills in English are a requirement. As a team member, you can expect:

  • To work 40 hours per week, and be available during normal business hours as needed.
  • Payments made in USD.
  • 18 days of PTO per year, observance of local holidays, and an annual break between Christmas and New Years.
  • Wellness + Remote Stipend
  • AI Voucher

About Able

Able builds technology products in a portfolio model. We believe that people, teams, and processes are more important than the ideas themselves, so we’ve focused on bringing great people together, and investing in their growth.

We’ve built products in a variety of industries. Everything from media to finance to toys to healthcare. Sometimes we work with management teams to help their businesses grow faster or unlock value using technology. Other times we start or buy businesses outright. Each time, we look for opportunities to leverage technology built at the portfolio-level to drive value faster.

Able is committed to inclusion and diversity and is an equal-opportunity employer. All applicants will receive consideration without regard to race, color, religion, gender, gender identity, sexual orientation, national origin, disability, or veteran status.

This is but the beginning of a conversation we’d love to have with you.

Apply, and let’s get this adventure started!

Apply now >

This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.

Next step

Apply now.

Follow the employer’s application method and review Jobicy’s safety guidance before sharing personal information.

Did you apply?Let us know, and we’ll help you track your application.

Continue on the employer website

Protect your personal information and never pay to secure an interview or job offer. View safety guidance.

Log in to save
One quick step before you apply

Create your free account, then apply.

Build a more organized job search on Jobicy and continue to the employer's application when you're ready.

  • Never lose a promising opportunitySave roles and return to them from your dashboard.
  • See your entire search at a glanceTrack applications, stages and next steps in one place.
  • Get matched with relevant remote jobsChoose the alerts and digests that work for you.
Applying is free. The employer's application opens in a new tab.
Add alert
Jobs Talent AI Tools Salaries
Menu