About
I build the machinery around the model.
Calling an API is not the job. The job is retrieval that finds the right thing, output you can verify, failure you can see, and a number that tells you whether any of it is getting better.
What I do
What I actually do
- Focus
- Applied AI, production systems
- Experience
- 8 years engineering
- Languages
- Python, TypeScript, SQL
- Working style
- Solo, direct, opinionated
Retrieval architecture
The interesting problem is that vector search alone fails on exactly the terms people type: part numbers, policy references, API names. Fixing that means running lexical and exact-match channels alongside the embeddings and fusing them properly.
Agent systems that take real actions
A different discipline from agents that talk. Typed tools, structured outputs, a policy layer deciding what runs unsupervised, and a preview step before anything irreversible.
Evaluation
The rarest of the four. Golden datasets, ablation runs, recall and nDCG tracked per version. Most teams building on models cannot tell you whether last week's change helped that is a solvable problem rather than an inherent one.
The operational layer
Per-stage tracing so degradation is visible, idempotent writes so a retried webhook cannot double-charge anyone, token metering so nobody discovers the economics from an invoice.
Background
Where the discipline came from
Most of my career was inside healthcare payments and patient financing, where the data was the kind you cannot be casual with: cardholder details, medical history and personal information in one system, under the obligations that come with all three.
Least privilege, audit trails, redaction at the layer rather than by convention: those are habits from that period rather than requirements someone handed me later. It is also why I am comfortable putting AI near regulated data, and why the first thing I ask about a workflow is what the system must never be able to do.
Real stack extracted from every project
Languages, frameworks & tools.
Languages
- Python 3.10 – 3.12
- TypeScript 5.x
- SQL (PostgreSQL)
- JavaScript (ES modules)
Backend & APIs
- FastAPI + Uvicorn
- Flask + flask-cors
- Next.js 16 App Router
- asyncpg (raw, no ORM)
- Pydantic v2 + pydantic-settings
- Streamlit
- Typer + Rich CLI
AI / LLM
- Anthropic Claude (SDK)
- OpenAI GPT-4o / embeddings
- Google Gemini (@google/genai)
- Nova ADK (novalab-adk)
- fastembed (local, CPU)
- sentence-transformers (BGE)
- Cohere rerank
- ONNX Runtime
Retrieval & Search
- pgvector (HNSW + IVFFlat)
- tsvector full-text search
- pg_trgm (trigram / exact ID)
- RRF score fusion
- rank-bm25
- rapidfuzz fuzzy match
- tiktoken token budgeting
- MongoDB + Motor (async)
Frontend
- React 18 / 19
- Vite + SWC
- TailwindCSS
- Radix UI primitives
- shadcn/ui
- TanStack Query
- Zustand
- Recharts
Data & Infrastructure
- PostgreSQL + Supabase
- Docker + docker-compose
- Firebase Cloud Messaging
- LiveKit SIP (voice)
- n8n (automation)
- Vercel, AWS, GCP
- structlog + tenacity
- yfinance, pandas, NumPy
Document Intelligence
- pypdf + python-docx
- pdfplumber + reportlab
- openpyxl
- Puppeteer (web crawl)
- beautifulsoup4
- huggingface-hub
Quality & Eval
- pytest + pytest-asyncio
- ruff + mypy (strict)
- nDCG, recall@k, MRR
- Ablation tables
- Golden datasets (JSONL)
- Citation verification
Fit
Who I am useful to
Teams whose AI feature demos beautifully and misbehaves in production. Teams who need retrieval that handles exact identifiers as well as paraphrase. Teams who want an agent to do real work without handing it unsupervised access to their data. Teams near regulated information who need someone comfortable with that rather than nervous about it.
And teams who want to know whether the system is actually improving, rather than taking someone's word for it.
I work alone and prefer to. You talk to the person building the thing, and I will tell you when an idea cannot be implemented safely or measured properly. That is more useful than agreement, and it is usually cheaper than finding out later.