Data analyst who ships production LLM systems. SQL and Python, with a statistics foundation (MITx MicroMasters in Statistics and Data Science). Came up through restaurant operations, then AP and invoice data.
I build things that make messy document data usable, and I write about where automated systems quietly get it wrong.
More data, fewer tokens A research agent that chose which queries to run burned 697K input tokens and up to 42 turns on a single diagnosis. Running every query up front and reasoning once cut that to 21K tokens and 24 seconds, and collapsed a 10x cost swing into a 1.4x band.
What it cost to put 15 restaurants on real food costing A line-item breakdown of a 2017 actual-vs-theoretical implementation, scored against which lines an AI agent removes today. The two failure exhibits are about clean text, high confidence, and a wrong number.
Recipe Flowchart — Dependency graphs from unstructured recipe text across 6 models and 3 providers, with a taxonomy-driven eval loop for a task that has no single right answer. Prompt iteration cut ordering errors 19 to 3; a tested LLM judge was dropped at 0.28 recall.
The Peeking Problem — Power analysis and Monte Carlo on a 90K-user A/B test. Daily peeking with early stopping inflates the false-positive rate from 5% to ~22%; a Pocock boundary corrects it.
Invoice Volume Forecast — Prophet model forecasting daily document intake 60 days out, with tuned weekly, quarterly, and holiday seasonality.
Price Tracker — Normalizing two federal data sources with different shapes and reporting cadences onto one weekly grid.

