Ship production-grade AI.
I advise teams on the layer between a promising demo and a reliable product: agent platforms, retrieval quality, evals, and LLM cost. The same problems I solve at production scale — brought to your system.
Agent platforms & runtimes
Architecture and review for agent harnesses, tool authorization, memory, orchestration, and MCP developer platforms — the runtime layer your AI products stand on.
15M+ conversations / quarter on the runtime I architectedRetrieval & RAG quality
Diagnosing and fixing retrieval: hybrid search, reranking, layout-aware chunking, citation grounding, and document question answering that holds up under audit.
Evidence retrieval F1@10 raised from 86% to 97%Evals & release gates
Evaluation suites that gate releases: faithfulness, citation accuracy, hallucination rate, synthetic data with hard negatives, and regression harnesses your team can run on every change.
Release gates that kept an enterprise compliance product shippableDocument AI & cost engineering
Intelligent document processing, OCR routing, fine-tuning small models with LoRA/QLoRA, and cost-aware routing between open models and frontier APIs.
70% OCR cost reduction · 23+ legacy pipelines replaced by one modelThree ways to work together.
Every engagement starts with a free 20-minute intro call — book one directly. You describe the system and the problem, and we decide together whether I'm the right person. Fees are fixed and scoped before work begins.
Advisory session
A focused working session to unblock one decision: an architecture call, an eval strategy, a build-vs-buy question, a retrieval-quality diagnosis. You leave with a concrete recommendation and next steps.
Architecture & quality review
A deep dive into an existing system — agent runtime review, retrieval audit, eval-gap analysis — ending in a written, prioritized report your team can execute without me.
Ongoing advisory
Recurring access for teams building AI platforms: design reviews, roadmap input, eval and quality checkpoints, and help hiring and leveling up AI engineers.
Q::01What kinds of teams do you work with?
Teams shipping LLM products — from seed-stage startups to enterprise platform groups. The strongest fit is anyone building agents, retrieval/RAG, document intelligence, or the eval infrastructure around them.
Q::02Do you write code or only advise?
Both, depending on the engagement. Advisory sessions and reviews produce written recommendations, architecture diagrams, and eval designs. Scoped sprints can include hands-on prototyping and reference implementations.
Q::03How do fees work?
Fixed per engagement, agreed before work starts — after a free intro call to confirm fit and scope. No open-ended hourly billing.
Q::04Can you help us hire or level up our AI team?
Yes. Interview loops, evaluation rubrics for AI engineering roles, and mentoring for engineers moving from ML models into AI systems work.
Q::05Are you available for full-time roles?
For Staff/Principal AI platform roles, see the about page for my full track record, then reach out by email or LinkedIn.
Tell me what you're building.
Two or three sentences about your system and where it hurts is enough. I reply to every serious inquiry, usually within two working days.