About the product
We run an agentic marketplace for car buyers. Instead of forcing people through filter dropdowns, our users talk to an agent: “I need a 7 seater under R$120k that’s fuel-efficient and suitable for a family with two kids.” The agent researches, compares real inventory, explains trade-offs, and takes the person to a car they can actually buy.
Behind that conversation sits a fleet of models — routers, retrievers, rankers, extractors, summarizers, judges - each doing a specific job under a latency and cost budget. This role owns that fleet.
Responsibilities:
- Evaluation — the offline and online eval suites that tell us whether the agent is actually getting better: golden datasets from real traffic, calibrated LLM-as-judge rubrics, and eval as a release gate for every model, prompt, or pipeline change.
- Model strategy — deciding which model serves which step, and building the router that balances quality, latency, and cost. Benchmarking new releases against our own evals rather than vendor claims.
- Fine-tuning and adaptation — SFT, LoRA, and preference tuning on automotive-domain tasks when it genuinely beats better prompting or retrieval, plus the data flywheel that feeds it.
- The deep research pipeline — multi-step retrieval and synthesis across inventory, specs, pricing, reviews, and ownership cost, with grounding and citations we can trust.
- Serving and lightweight MLOps — inference services, embedding and index refresh jobs, versioning, staged rollout and rollback, and the observability to see quality and cost in production.
- Fallback and redundancy — multi-provider fallback chains, circuit breakers, and graceful degradation, so a slow or unavailable model never becomes a broken experience for the buyer.
- Cost-conscious infrastructure — owning cost per conversation and keeping the stack lean.