Shashank

The model is rarely the hard part.

What makes AI systems difficult is everything around them: retrieval that has to be accurate, pipelines that survive a failing third-party API, extraction that gets validated before it reaches a database, and the measurement that tells you when quality has quietly slipped.

Four principles

Measured, not asserted

A golden dataset and ablation runs exist before tuning begins. Every stage of a pipeline has to justify its latency against a number, and a regression becomes visible the day it lands rather than the quarter after. Without this, tuning is guesswork and improvements are opinions.

Failure has to be visible

When a stage degrades, the response says so. A retrieval system that silently falls back to whatever it can find looks identical to one that works, and the difference is usually discovered by a customer. Every stage reports whether it ran cleanly, and that reaches the caller rather than a log nobody reads.

Preview before it writes

Anything irreversible is rendered as a proposed change and confirmed by a person. Writes are keyed so a retry cannot duplicate a record, a message or a charge. This is the difference between an agent a team trusts and one they switch off after a single bad afternoon.

Costs are designed in

Classification and extraction run on small fast models; only genuine reasoning reaches the expensive one. Caching is content-addressed, and usage is metered per tenant and per feature from the first release, not retrofitted once the invoice becomes alarming.

A typical engagement

WEEKS 1 TO 2 WEEKS 3 TO 6 WEEKS 7 ONWARD Map and separate Candidate workflows by impact and risk Deterministic rules split from AI tasks Source data defined per workflow Written up as decisions, not opinions Build the spine Evaluation harness and golden datasets Reusable execution layer Confidence thresholds and review gates One workflow running end to end Prove and harden Shadow mode against real traffic Production behind human approval Monitoring, cost and latency Documentation and handover weekly working demo, not a status report acceptance criteria agreed before work starts

Measurement is built in the first third, not the last. Built late it becomes a report nobody acts on, because by then the system is too entangled to attribute a regression to any single change.

The judgement call that matters most

Deciding what should not use a language model at all. If a competent person could write the logic down as a rule that stays stable, it belongs in code. Routing deterministic work through a model makes a solved problem probabilistic and unauditable at the same time, and it is the most common and most expensive mistake in applied AI.

Models earn their place where the input is unstructured and the ambiguity is genuine: reading a document to find what is missing, classifying a free-text reason code, drafting language, summarising a queue for a human. Everywhere else, ordinary software is faster, cheaper and correct every time.

How I work

Cadence

  • Weekly working demo
  • Written summary after each
  • Blockers raised inside 48 hours
  • Scope trades, never silent slippage

Code

  • Typed, tested, documented
  • Reviewed pull requests
  • Someone else can pick it up
  • Decision records for hard calls

Access

  • Least privilege throughout
  • Read-only against production by default
  • Credentials stay in your systems
  • Every action auditable

Handover

  • Documentation written continuously
  • Runbooks for incidents
  • Known limitations written down
  • Recorded walkthrough