Senior Full-stack Engineer, LLM Applications (LATAM - Remote)
Remote
Punch List Digital · South, US
Fresh
Gig
Braintrust
Compensation
$70-$74/hr
Description
Sr. LLM Application Engineer
Punch List Digital (PLD) builds the platform that turns live client performance data and our marketing methodology into answers and reports. Strategists, client success, sales, and clients use it to understand performance across search visibility, paid media, analytics, and local search.
The platform has two surfaces:
- Chat — answers performance questions from live, multi-tenant client data, with every figure grounded in a query.
- Reports — executive-ready monthly client reports, generated section by section from our KPI calculations and methodology.
We're hiring a senior full-stack engineer to build both, across the backend, the front end, and the LLM layer that connects them. Our stack is Laravel and Vue.
What you'll build
- Report generation. The monthly report pipeline: pulling KPI outputs, assembling each section with the LLM grounded in our Knowledge Vault, and producing a client-ready document or PDF. Built as durable, observable, retryable background jobs.
- The chat experience. Streaming answers, conversation history, client and group scoping, exports, and feedback — from the interface through to the data layer.
- The answer pipeline. How a question or scheduled trigger becomes the right queries, the right context, and a grounded output. Prompts, tool definitions, structured outputs, and retrieval.
- Knowledge Vault retrieval. Ingestion and search across the curated knowledge base that grounds answers and reports.
- Integrations. The Anthropic API (tool-use, streaming, prompt caching), our internal KPI feed, and partner APIs such as Google Analytics, Google Ads, and SEMrush.
- Quality gates. Evaluation sets and sampled outputs checked against known-correct figures before any prompt, report, or feature change ships.
The work
- Shipping features end to end: the interface, the API, the prompt, and the tests.
- Choosing between prompting, tool-use, retrieval, and plain code for a given problem, and documenting the decision in a short design note.
- Designing multi-tenant scoping so one client's data never reaches another.
- Treating an incorrect figure in front of a client as the most expensive defect the system can produce, and designing grounding, validation, and review accordingly.
Qualifications
- 5+ years building web applications across backend and front end, with a modern backend framework, background jobs, and an established testing practice.
- Strong front-end skills in TypeScript with a modern component framework: component design, state management, streaming interfaces, and component testing.
- Production experience with LLM APIs: prompt design, tool-use, and structured outputs in customer-facing features, with prompts and schemas kept under version control and tested.
- Solid SQL: reporting queries, indexing, query-plan analysis, and judgment on when aggregation belongs in the database.
- Third-party API integration in production: OAuth and API-key authentication, pagination, rate limiting with backoff, and resilience to upstream schema changes.
- Sound judgment on data accuracy, and the standards to back it.
- Clear written communication for engineers, strategists, and leadership alike.
Preferred qualifications
- PHP and Laravel: queues and jobs, Eloquent, the service container, and PHPUnit or Pest.
- Vue 3 and Pinia.
- Production experience with Anthropic Claude.
- Retrieval-augmented generation in production: embeddings, vector search, pgvector or similar.
- PDF and document generation from templated content.
- Evaluation harnesses for LLM output.
- Multi-tenant data and access-control patterns.
- Familiarity with Google Analytics 4, Google Ads, SEMrush, Monday.com, or WhatConverts.
- Marketing or agency domain knowledge: SEO, paid media, local search, or answer-engine visibility.
How we work
- We write things down. Architectural choices get a short write-up of what we chose and why.
- We validate before we ship. Prompt and report changes are checked against known-correct figures first.
- We learn from review. Human corrections feed back into prompts, retrieval, and checks.
- Pace follows risk. Internal experiments move fast; anything on the path to a client gets more care.
How success is measured
Accuracy of chat answers and reports against known-correct figures, adoption by the teams and clients the platform serves, time returned to our strategists, and no incorrect figures reaching a client.
Hiring process
- Culture Index assessment (5 minutes) - short questionnaire to help determine culture fit
- Intro call (30 min). Background, fit, and your questions.
- Technical conversation (60 min). A walkthrough of a production system you built on an LLM API: how it worked, how it was tested, and how you verified its output.
- Working session with the team. A representative problem from our domain, worked through together.
- Offer.
- Contract
- Long
- Engagement
- Freelance
Skills & categories
LLMsEmbeddingspgvectorTypeScriptSQLVector DatabasesRAG
- Posted
- Sep 28, 2026
- Slots remaining
- 1
- First seen
- Sep 28, 2026
- Last seen
- Sep 29, 2026