The model is rarely the hard part.
What makes AI systems difficult is everything around them: retrieval that has to be accurate, pipelines that survive a failing third-party API, extraction that gets validated before it reaches a database, and the measurement that tells you when quality has quietly slipped.
Four principles
Measured, not asserted
A golden dataset and ablation runs exist before tuning begins. Every stage of a pipeline has to justify its latency against a number, and a regression becomes visible the day it lands rather than the quarter after. Without this, tuning is guesswork and improvements are opinions.
Failure has to be visible
When a stage degrades, the response says so. A retrieval system that silently falls back to whatever it can find looks identical to one that works, and the difference is usually discovered by a customer. Every stage reports whether it ran cleanly, and that reaches the caller rather than a log nobody reads.
Preview before it writes
Anything irreversible is rendered as a proposed change and confirmed by a person. Writes are keyed so a retry cannot duplicate a record, a message or a charge. This is the difference between an agent a team trusts and one they switch off after a single bad afternoon.
Costs are designed in
Classification and extraction run on small fast models; only genuine reasoning reaches the expensive one. Caching is content-addressed, and usage is metered per tenant and per feature from the first release, not retrofitted once the invoice becomes alarming.
A typical engagement
Measurement is built in the first third, not the last. Built late it becomes a report nobody acts on, because by then the system is too entangled to attribute a regression to any single change.
The judgement call that matters most
Deciding what should not use a language model at all. If a competent person could write the logic down as a rule that stays stable, it belongs in code. Routing deterministic work through a model makes a solved problem probabilistic and unauditable at the same time, and it is the most common and most expensive mistake in applied AI.
Models earn their place where the input is unstructured and the ambiguity is genuine: reading a document to find what is missing, classifying a free-text reason code, drafting language, summarising a queue for a human. Everywhere else, ordinary software is faster, cheaper and correct every time.
How I work
Cadence
- Weekly working demo
- Written summary after each
- Blockers raised inside 48 hours
- Scope trades, never silent slippage
Code
- Typed, tested, documented
- Reviewed pull requests
- Someone else can pick it up
- Decision records for hard calls
Access
- Least privilege throughout
- Read-only against production by default
- Credentials stay in your systems
- Every action auditable
Handover
- Documentation written continuously
- Runbooks for incidents
- Known limitations written down
- Recorded walkthrough