About this role.
This Senior AI Engineer role builds document-intelligence systems that combine computer vision, PDF parsing, and LLM orchestration. The engineer will develop pipelines for extracting structured information from complex visual documents, including layouts, charts, and tables. Core work includes LangChain or LangGraph agent workflows, multimodal model integration, prompt/control-flow design, and cost optimization. Strong Python ML experience, native PDF-processing expertise, and clear English communication are required. The role is fully remote for candidates within Latin America and expects a 40-hour workweek.
Role DNA
A quick view of the complexity, pace, ownership and collaboration implied by the job description.
Job Complexity
5/5Pace & Pressure
4/5Autonomy Level
5/5Communication Load
4/5Salary analysis
Estimated compensation compared with the broader US market for similar roles.
Core skills
Skills and capabilities most closely associated with this opportunity.
Sample interview questions
I would first classify the document and extract native PDF elements with PyMuPDF or pdfplumber wherever possible, preserving page coordinates and reading order. For scanned or highly visual regions, I would use layout detection and targeted vision-language model calls, then normalize outputs into a schema with provenance links to page regions. I would validate critical fields using confidence thresholds, cross-checks, and human-review fallbacks for low-confidence financial values.
I define narrowly scoped graph nodes, structured inputs and outputs, and explicit state so each step has a clear responsibility. Rather than passing entire documents into a model, I retrieve only relevant chunks, table representations, and image regions while retaining metadata for traceability. I also implement retries, tool-result validation, guardrails, and evaluation datasets to identify failure modes before production release.
I would ground the model on extracted source data and require structured answers that cite page, chart, or table coordinates. I would use deterministic parsing for values when available, reserve visual models for interpretation tasks, and prompt the model to return uncertainty rather than infer missing information. Automated checks such as totals reconciliation, unit validation, and comparison against OCR or native text provide additional safeguards.
I would segment documents and route each task to the least expensive model capable of meeting quality targets. Native text extraction, caching, batching, asynchronous processing, and selective high-resolution image rendering can greatly reduce expensive vision calls. I would measure per-document cost, latency, and extraction quality by document type, then use those metrics to tune routing thresholds and model choices.
I would begin with a representative benchmark set and establish a prompt-and-retrieval baseline using error categories such as chart misreads, layout mistakes, and incorrect field extraction. If recurring errors persist despite improved data representation, prompts, and validation, I would assess whether sufficient labeled examples exist for fine-tuning. The decision would compare measurable quality gains against annotation, training, maintenance, inference cost, and the ability to adapt to new document formats.
Back in 2012, we were a group of engineers and designers who decided we wanted to build things—so we did. Able started as an engineering and product hub building for a portfolio of early-stage startups. We built many relationships while developing products that were thoughtful, effective, and genuinely useful. But, since then, we’ve grown… and so has our ambition.
Now, we’re entering our next chapter—defined by applied AI. AI is a powerful force in the end-to-end software development cycle, and we’re creating practices that allow us to deliver software fast and more effectively than traditional approaches, creating meaningful value for our partners. Today, our builder mindset is driving us to become an AI-native organization across every function. We’re still evolving, and that’s part of the opportunity. If you want to build, learn, and tackle challenges alongside an ambitious team, let’s build together.
This position is 100% remote within LatAm.
What you’ll be doing
We are seeking someone who enjoys working at the cutting edge where Computer Vision meets Logic. You will be responsible for the “eyes” and the “brain” of our system—extracting complex data from visual documents and then orchestrating how that data is used by Large Language Models.
In short, someone who likes:
- Unlocking Visual Data: Building pipelines that can “read” complex documents, understanding layout, charts, and visual context using Vision-Language Models (GPT-4V, Claude 3.5) and Layout Analysis.
- Orchestrating Intelligence: Owning the application logic layer. You will use LangChain or LangGraph to build the agents and chains that query our data, reason about it, and generate responses.
- Native PDF Handling: Handling the messy reality of PDF processing (PyMuPDF, layout parsing) to preserve structure before the AI even sees it.
- Prompt Engineering & Logic: Crafting complex prompts and control flows to ensure models interpret financial charts and layouts accurately without hallucinating.
- Cost & Scale: Applying a cost-optimization mindset (batch processing, model selection) to ensure our vision and orchestration layers are economically viable.
What we’re looking for
We want to work with people who have a passion for collaborating with their teams, building software while nurturing inclusive and respectful relationships with their coworkers. With the ones that are open about their shortcomings and what they do not know now, but remain eager to keep on growing and closing those gaps.
Ideally, they would also have:
- LLM Orchestration (Must Have): Deep experience with LangChain, LangGraph, or similar frameworks. You know how to manage context windows, tool calling, and agentic workflows.
- Multimodal AI Experience: Hands-on experience integrating state-of-the-art vision models (GPT-4V, Claude 3.5 Sonnet) and embedding models (CLIP).
- Document Intelligence Specialist: Familiarity with specialized models (e.g., Donut, Pix2Struct) and tools like Unstructured.io or Docling.
- PDF Processing Mastery: Mastery over tools like PyMuPDF or pdfplumber for native element extraction.
- Python ML Stack: Strong proficiency in PyTorch or TensorFlow.
Nice-to-Have:
- Fine-Tuning: Experience fine-tuning vision or language models, specifically to improve accuracy on domain-specific artifacts like financial charts or tables.
- Domain Knowledge: Prior experience handling documents in the Real Estate or Finance sectors.
Able is powered by curious, thoughtful people who care about what they build and how they build it. We’re actively investing in our team through AI training, knowledge-sharing, and hands-on experimentation to ensure everyone grows alongside the technology.
This position is 100% remote within LatAm. Strong verbal and written communication skills in English are a requirement. As a team member, you can expect:
- To work 40 hours per week, and be available during normal business hours as needed.
- Payments made in USD.
- 18 days of PTO per year, observance of local holidays, and an annual break between Christmas and New Years.
- Wellness + Remote Stipend
- AI Voucher
About Able
Able builds technology products in a portfolio model. We believe that people, teams, and processes are more important than the ideas themselves, so we’ve focused on bringing great people together, and investing in their growth.
We’ve built products in a variety of industries. Everything from media to finance to toys to healthcare. Sometimes we work with management teams to help their businesses grow faster or unlock value using technology. Other times we start or buy businesses outright. Each time, we look for opportunities to leverage technology built at the portfolio-level to drive value faster.
Able is committed to inclusion and diversity and is an equal-opportunity employer. All applicants will receive consideration without regard to race, color, religion, gender, gender identity, sexual orientation, national origin, disability, or veteran status.
This is but the beginning of a conversation we’d love to have with you.
Apply, and let’s get this adventure started!
Annual salary information is not provided for this position. Explore salary ranges for similar roles in our Salary Directory ›
This job listing has been manually reviewed by the Jobicy Trust & Safety Team for compliance with our posting guidelines, including verification of the company's legitimacy, accuracy of job details, clarity of remote work policy, and absence of misleading or fraudulent content.









