A line-item breakdown of an actual-vs-theoretical implementation from 2017, and an assessment of which of those lines survive automation.
Rendered version: chodizzle.github.io/avt-implementation
I ran the implementation of actual-vs-theoretical food costing at Mendocino Farms from June 2017 to March 2020: CTUIT, PlateIQ, and Aloha across 15 locations and two commissaries, growing to roughly 23 stores by the time I left. I was hired into a role created for the project, and I was the only person on it full time. I still do CTUIT purchase-price review for a multi-unit client. This is what that project cost, line by line, and my assessment of which lines survive. Figures are reconstructed from memory. See the note on method at the end.
When I started, the company was around twelve stores with one commissary in Los Angeles. By the time I left it was about twenty-three stores across three regions, with a second commissary in Texas. NorCal's first store opened on December 8, 2017, the same winter the system went live. Texas opened in July 2019.
The stack: PlateIQ captured invoices at line-item level and fed purchases into CTUIT (Compeat Radar), which held inventory and the recipe tree and took depletion from Aloha. Operators counted weekly and entered waste. Out the other end came actual-vs-theoretical variance.
Before this, inventory was a spreadsheet. Proteins only.
- 400–500 recipes and sub-recipes, including yields on individual ingredients
- 200–300 inventory items
- 5,000+ vendor product identifiers, of which roughly 1,000–2,000 were real food items; the rest were services, supplies, and repairs that still had to be recognized as not-inventory and routed somewhere
- 2,000–3,000 POS objects mapped
- About 10 vendors: one broadliner, a few specialty, the rest miscellaneous
- Roughly 40 invoices per store per week, so about 600 a week company-wide
- Two commissaries, different physical layouts, real regional substitutions
These phases were not sequential blocks. I wove them: coordinating yield trials while entering recipes, mapping what I could when I could, testing the invoice flow in between. The calendar column records when the work happened, not a critical path.
| Phase | Calendar | My hours | Other people's time |
|---|---|---|---|
| Recipe validation | ~6 months, concurrent | ~500 | ~1 week each, across the R&D chef, commissary crew, and store crews |
| Inventory item creation | 1 week | 40 | — |
| Recipe entry, incl. variations | 2–3 weeks | 80–120 | — |
| VPI → inventory mapping and UOM conversion | 2–3 weeks | 80–120 | — |
| POS → recipe mapping | 1–2 weeks | 40–80 | — |
| Count sheet build, 15 locations | 2–3 weeks | 80–120 | KM time by phone and email; one 3-day trip to Texas |
| Commissary → store transfers | ~2 weeks | ~80 | — |
| PlateIQ setup: pack sizes, CTUIT export module, EDI onboarding | 1–2 weeks | 40–80 | PlateIQ engineering |
| Live validation, two count cycles | ~2 weeks | ~80 | store counting time |
| Aloha item file | — | 0 | 0 |
| Total | 6–7 months | ~1,100 |
~1,100 hours. Six to seven months. One dedicated headcount. Live that winter. Then about 8 hours a week, indefinitely, to keep it true.
The zero on Aloha is real. Someone had built that item file properly before I arrived, and CTUIT let me define a recipe for one unit of something and map it to a POS object with a multiplier: one ounce of chicken breast becomes extra chicken and double chicken at the right multiples. Two to three thousand objects, and it was the easiest part of the project. The line is worth reporting because it shows what this work costs when the upstream system is already clean.
If the job had been entering data and building maps, it would have been about two months.
It was not. Corporate's mandate was to validate: to double-check recipes accumulated over ten or fifteen years of operation. The recipe book existed and the discipline behind it was good. Versions needed some reconciling with kitchen teams. That was not the expensive part.
The expensive part was that a recipe written for a cook is not a recipe that can be costed. Stores worked from instructions like a two-ounce spoodle of crispy quinoa. That is a good instruction, and the kitchens executed it consistently. However, a spoodle is a volume tool, crispy quinoa's density depends on how it fried, and the commissary yields and packs it by weight. The true grams per spoodle was a number that existed nowhere, because nobody had ever needed it.
That is the 500 hours. It was not typing. It was standing in kitchens during morning prep watching people cook, running yields with the chef and with line cooks, and converting a document written for execution into one that supports arithmetic.
My stopping rule was three consecutive consistent yields. On high-cost or high-variance items that usually meant five or six trials. On twice-cooked pork belly it took about ten. The chain is marinate, braise, cool, cut, fry, bag in its own juices, ship, reheat at the store, serve: nine steps, losses at most of them, and it crosses from the commissary into a store I was not standing in. A fried grain item (cook, cool, deep fry, season, drain) cost me a similar stretch, because it loses moisture, absorbs fat, gains seasoning, then drains, in that order.
Count sheets were the other thing that did not compress. Every location got its own, built in shelf-to-sheet order off the inventory items, in whatever unit made counting fast and hard to get wrong: pounds for expensive proteins, quarts for vinaigrettes because a cambro has graduations you can read from across a walk-in, discrete containers for pre-portioned items. Some of that I did by phone with kitchen managers. For Texas I got on a plane: a day and a half in the commissary running yields and building its sheet, then most of another day at the nearest store, half a day's drive away.
Most stores landed inside ±2% variance, and food cost came down about two points across the company, off a baseline that was already disciplined.
The work also outlived the software. Mendocino later moved from CTUIT to CrunchTime, and as far as I know that transition was not painful, because the expensive part had already been paid: validated yields, a correct recipe tree, working maps, count sheets that matched real shelves. The system was replaceable. The model of the operation was not.
Here is what I would expect an agent to take off that table today. These are estimates from experience rather than measurements, and they are the softest numbers in this report.
| Line item | Then | Now | Verdict |
|---|---|---|---|
| Recipe validation | ~500 | ~350 | Trials do not move; targeting and record-keeping do |
| Inventory item creation | 40 | ~10 | Collapses |
| Recipe entry | ~100 | ~15 | Collapses |
| VPI → inventory mapping | ~100 | ~35 | Mostly, with the exceptions below |
| POS → recipe mapping | ~60 | ~10 | Collapses |
| Count sheet build | ~100 | ~40 | Item list and units collapse; shelf order does not |
| Commissary transfers | ~80 | ~72 | Mechanics collapse; the costing decision does not |
| PlateIQ setup | ~60 | ~15 | Mostly |
| Live validation | ~80 | ~25 | Anomaly catching collapses; the phone call does not |
| Total | ~1,120 | ~570 |
About half. Fifteen locations, two commissaries, four to five hundred recipes, cut from roughly seven months to something like three and a half.
Bucketing the non-inventory VPIs. Of five thousand-plus vendor product identifiers, only one to two thousand were food. The rest were services, supplies, and repairs that still had to be recognized and routed. That is classification work, and a model does it well.
Vendor-side duplicates. PlateIQ was young in 2017 and would hand me five variations of one real item from one vendor, so I built five VPI mappings pointing at one inventory item. Deduplicating near-identical vendor descriptions is precisely what semantic matching is for.
Recipe entry, and POS mapping. Structured extraction from documents into a recipe tree was weeks of typing. CTUIT was already good at the POS side (fuzzy match plus that one-unit-and-a-multiplier pattern), so it was the cheapest line on my table, and it goes to nearly nothing.
Maintenance. The eight hours a week collapses harder than any of it. Most of that was auditing and chasing corrections with operators (why does your beef say 3,000 pounds, you mean 300), plus a weekly scorecard for corporate. The pattern is to match a store against its own history, then make a phone call. I do a version of this work now, reviewing purchase pricing for a multi-unit client on CTUIT, and I have automated most of the audit pass with an LLM-driven tool. It works.
Two exhibits, both from the mapping work, which is the same work the model is best at.
Avocados. SoCal stores reliably got 40–60 count cases. NorCal could only get 60–70. The vendor did print them as separate SKUs, so the data was not lying. However, the names were close enough to be confusing, and the correct handling runs directly against the system's own default.
The entire premise of VPI mapping is consolidation: many vendor identities collapse onto one real thing you count. My own fan-out was five to seven to one. That policy is right almost every time, and two similar avocado SKUs is exactly the shape it wants to merge. Merging them is the obvious, defensible, wrong answer: they are different physical units with different per-each yields going to different regions.
The fix was not in the mapping. It was in the schema: two inventory items, forked recipes everywhere an avocado appeared, and region-gated POS mappings so each store hit the right version. It took me a while to work that out. I was green, and the failure was mine before it was anyone's. Once it was in, avocado variance worked everywhere. A GM who had come over from a large casual-dining chain told me he had never seen anyone get that right before. I do not think that is because it is hard to fix. I think it is because it is hard to notice.
The signal is detectable. Same base item, different pack or count spec, and the two never appear at the same store. All of that is sitting in invoice history. A system could flag it: these look like one item, their pack specs differ, they never co-occur by region, so do not merge, ask. That is a rule someone writes after getting it wrong once.
The spoodle. This is a different failure. The avocado data was fine and the modeling decision was hard. Here the data is clean, complete, and wrong. A model reading two-ounce spoodle of crispy quinoa records two ounces, with no reason to doubt it, and is silently off on every plate that goes out. Multiply that across a menu built in scoops, ladles, portions, and pans.
Neither of these is ambiguity. Ambiguity is the easy case: the model is uncertain, it flags, a human resolves it. Both of these produce clean text, high confidence, and a wrong number, and a wrong number does not raise an error. It quietly reports variance that is not real, or hides variance that is. Calibrated confidence does not help, because confidence is computed from the text and the text is fine.
The countermeasure is not a better model. It is a rule, and rules of this kind are written by people who have already been wrong once. A description that appears in more than one region with a different pack count should not be silently merged. A recipe unit that names a tool rather than a measure, spoodle, scoop, ladle, ring, is not a quantity until someone captures it physically. An ingredient whose yield spans a fry, a braise, or a drain is unvalidated until someone runs it. I know each of those because I got it wrong first, not because it can be derived from the corpus.
Validation was about 45% of my project. On these numbers it becomes roughly 60%, and it keeps climbing, because everything getting cheaper sits on the other side of the ledger.
The job changes from input to audit. That is an improvement: auditing is higher-leverage work than typing, and I would have taken that trade in 2017 without thinking about it.
An audit needs an auditor.
What remains is the part I spent 500 hours on, and it remains in a more concentrated form than it started.
Sampling cannot be inferred. Three consecutive consistent yields; five or six trials on a high-variance item, about ten on pork belly. That number is not in any document and cannot be estimated from one. It is a property of how much the process actually varies, which you learn by running it repeatedly. A system can tell me which recipes probably need trials. It cannot run them, and it cannot tell me when I have run enough.
Losses compound down the tree, and the tree crosses organizations. A bad yield at step two of nine silently poisons every number above it. Pork belly's last few losses happen in a kitchen the commissary never sees, in liquid that makes shipped weight different from usable product. You can model that chain once you know the numbers. Getting the numbers means standing there.
Physical divergence is not documented anywhere. Two commissaries, same company, same menu, real differences in substitutions and completely different layouts. I found that out by flying there. No document described the difference, because to everyone working in either building there was nothing to describe.
Live and true are different dates. Commissary items transferred to stores priced at breakeven, which meant every store recipe built on a commissary sub-recipe carried an understated cost. Fixing that was a second project: timing production runs, allocating overhead across recipes at high/mid/low, eventually pushing labor into the commissary recipes themselves. A few items came back with costs I did not expect. The system had been reporting variance for months before that, and the variance had been confidently incomplete the whole time.
On counting. We counted in whatever unit made counting fast and hard to get wrong. Those unit choices were not a data schema. They were the accuracy mechanism: pick units a tired person can judge at 6am, then build the conversions behind them. Count accuracy is now advertised in the high nineties, and the figure that determines what such a number means is what it was measured against. A scale, or another human count. Having reconciled a large number of counts, I know how far apart those two answers can be.
I no longer have access to the records. Hours, calendar spans, and object counts are recollections of a project that ran from 2017 to 2018, reconstructed in 2026, and should be read as order-of-magnitude rather than exact. Where my memory gives a range, I have printed the range.
Absolute cost percentages, vendor pricing, recipes, yields, and inventory sheets are omitted throughout. They are not mine to publish. The food cost improvement is given as a relative change only.
The projections in Part 2 are estimates from experience. They have not been measured against a real automated onboarding, and I would revise them after observing one.