1. Hard to plan
Hitting 1,800 calories and 120 g of protein means adding up every label. AI chatbots just guess.
“Where did 420 calories come from?” The chatbot made it up.
The problem, who it’s for, where it sits among existing tools, how it works in five levels, what building it taught me, what it will never do, and what’s still missing.
Every week, people who eat to a goal spend an hour or so on the same chores: adding up calories from labels, checking what’s in the fridge, writing a grocery list, and searching a store’s website for each item.
CartMyMeals plans your week around your goals, skips what you already have, and shows how every number was worked out. Then it fills your own Amazon or Walmart cart.
I wanted one path from goal to cart.
Hitting 1,800 calories and 120 g of protein means adding up every label. AI chatbots just guess.
“Where did 420 calories come from?” The chatbot made it up.
Some of it is at home. The rest waits for a store run, or an order you put off or forget.
“I have the plan, but not the ingredients until Saturday.”
So you swap in what you have, like pasta for chicken, and the week drifts off target.
“I swapped in pasta and fell 30 g short on protein.”
Three people I designed for, each mapped to a feature that serves them. Hypotheses Being checked in a private beta of about 150 people.
1,800 calories and 120 g protein a day, without a spreadsheet.
“I spend Sunday adding up labels, and I still don’t trust the numbers.”
Every number computed from USDA data, with the math shown.
One menu the whole family eats, sized for the kids.
“By the time I’ve checked every label and built the cart, dinner is late.”
Kids’ portions and allergies, and the whole list in one Walmart cart.
Enough protein, without wasting what’s already at home.
“Half my lentils expire because I forget I bought them.”
Vegan filters, and a week planned around the pantry first.
Willingness to pay Hypothesis Strong interest in anything that saves one to two hours a week and removes nutritional guesswork. Not yet tested.
Trackers get the numbers right after you’ve eaten. Chatbots plan ahead but recall their numbers. Meal kits ignore what’s in your fridge. No one plans ahead with verifiable numbers and ends at a cart.
Profile in, seven-day plan out. The model invents meals and recalls their calories. Kept deliberately, as a live benchmark: its figures disagree with Level 2’s by roughly 15% on calories and 17% on protein.
The model may only pick recipe ids from a shortlist. Every calorie is arithmetic in a database view: 60 g oats × 380 kcal/100 g = 228 kcal. Live for visitors: start with one sentence and the week is arranged to your wishes, then change it in your own words (“swap beef dinners for tofu”). Claude only picks from meals the rules allow, and any id outside them is dropped.
Need minus have, in grams, across three unit systems (grams, packages, shoppable units). Each gap is matched to a real product by a scorer; weak matches are declined and shown, never guessed. The cart is built and handed over at checkout, on the demo store or as one link that fills the shopper’s own Walmart cart. The same tools are offered to AI apps as an MCP server, so a week can be planned and shopped from inside Claude (cartmymeals.com/mcp).
Claude already has five jobs, each with a model chosen by comparison: Sonnet reads a household sentence and receipts, arranges the week and turns requests into changes; Haiku lists a dish’s groceries; rules check each one. Still to build: one controller over them, the two approval gates as real steps, and a visible run trace of which model ran, what it cost and where it waited for a person.
A small model trained on logged Keep / Swap / Skip decisions, reranking candidates without loosening the guardrails. Waiting on enough real decisions; they are logged from day one and cannot be backfilled.
Next.js 16 and React 19 on Vercel · TypeScript · Postgres on Neon · Walmart prices via SerpApi · Shopify Storefront GraphQL · Claude Haiku 4.5 and Sonnet 5.5 · MCP (Streamable HTTP) · USDA FoodData Central
8 grocery integrations built or evaluated. Two are live for shoppers, who pick the store: Amazon (through Amazon Fresh or Whole Foods) or Walmart, each taking the whole list; ordering sits behind one interface (a shortfall list in, a cart out), so the rest can drop in without touching the planner.
Anyone can try it, with no account:
By default the week is chosen by retrieval and rules, with no model call, so it’s instant. Ask CartMyMeals to arrange or change it and it picks only from the recipes the rules retrieved; every number is still computed. It also reads what people write or photograph, into choices they check.

The meals repeat way too often. I have the same lunch almost every day.
Beta testerFive breakfasts, six lunches and a different dinner each night, and a note when choices are too narrow to avoid repeats.
I would never eat chicken breast at breakfast.
Beta testerNo meat or fish at breakfast by default, and a choice of sweet or savory.
I would never buy Atlantic salmon, only wild caught.
Beta testerOrganic, grass-fed, free-range and wild-caught swap the products that go in the cart.
The meals were very different from what my family eats.
Beta testerCuisines they love, meats they eat (no beef, no pork), and 54 more recipes, from dal makhani to chicken tinga tacos. Indian dishes appear only for those who pick Indian.
Ask for favorites.
Beta testerFavorite dishes repeat every week. Any dish, recipe link or YouTube video can go anywhere in the week, with its ingredients in the cart.
Items added to my Walmart cart were not the correct amounts: 12 broccoli crowns.
Beta testerQuantities are rounded to what stores sell, and anything not found is listed so nothing goes missing quietly.
Checkout tools such as complete_checkout are never wired in. The agent can stage a cart; only a person can pay.
Some cart APIs are write-only and can’t be read back, so every payload is recorded and shown, and test environments are labeled as test.
Below a confidence threshold the match is declined, the reason is shown, and the choice goes back to the person.
Nothing goes on a shopping list until you approve the week. Whatever CartMyMeals reads or changes (a sentence, a dish, a receipt, a request) comes back as choices you can see and undo, and its use is capped per visitor per day.
“Soy” removes tofu, edamame and soy sauce, not just the word. A child’s allergy applies to everyone’s meals, since the family eats together. The allergy eval checks this: 20 of 20.
If no meal fits your diet and allergies, it tells you rather than bending a rule. A pantry line it can’t read, or a “box” with no known weight, is flagged, not guessed.
It picks recipe ids from a retrieved shortlist and never writes a calorie. An id outside that list is dropped and reported; every figure is computed from gram weights.
Replace and Add only offer meals that already passed your diet and allergy filters, and the server re-checks every choice, so an edited request can’t bring a meal back in.
Walmart’s link adds to whatever is already in your cart and can’t empty it, so the page tells you to empty it first rather than letting old items slip through.
| Metric | Target | Where it stands |
|---|---|---|
| Weekly planning and ordering time | From ~60 minutes to under 10 | Not yet measured |
| Nutrition figures produced by a model | Zero | Met by design Every figure is computed from gram weights, using USDA values for 51 of 99 ingredients and standard food-composition tables for the rest. |
| Product match precision | Over 90% of matches accepted first time; no unflagged wrong substitutes | Partly built In the offline eval every stocked ingredient matched the right product and none matched a wrong one; acceptance by real shoppers not yet measured |
| From list to a real retailer’s cart | One click, no retyping | Met for Amazon and Walmart The whole list goes to the shopper’s cart in one click: Walmart’s add-to-cart link, or Amazon Fresh and Whole Foods through Amazon’s ingredients page. Instacart still goes item by item |
| Food waste | Less leftover food each week, by netting out what is on hand | Partly built Pantry netting built; waste not yet measured |
| Eval | What it asks | Result |
|---|---|---|
| Diet and style rules safety check | Across every diet, style, gluten-free and protein combination, does any meal or Replace option break the rules? | 80 of 80 |
| Allergies and dislikes safety check | When someone types an allergen in "Avoid", does any meal or Replace option still contain it? | 20 of 20 |
| Nutrition fit | Does the week land within 15% of the calorie target, and within 10% of the protein goal when one is set? | 35 of 36 |
| Pantry reading | Does typed pantry text become the right ingredient and amount, and is anything unknown left unrecognised rather than guessed? 0 non-food or unknown items were wrongly matched | 42 of 42 |
| Cart matching safety check | For every ingredient the store stocks, is the right product put in the cart? Is anything put in the cart that is wrong? 0 wrong products in the cart, 0 held back for a person to choose, 5 not stocked | 92 of 92 |
Since Oct 1, visitors built 115 carts and sent 10 to Amazon and 12 to Walmart. These measure the product. A private beta with real shoppers is running now: see what they told me below.
ProblemSomeone avoiding soy still got tofu. Avoiding milk still gave yogurt.
FixedAn allergy now covers every food made from it.
Problem“Almond milk” in the pantry was read as almonds.
FixedThe pantry reader tells them apart.
Problem“2 cans black beans” was read as dried beans.
FixedCanned beans count as cooked beans.
ProblemA vegetarian week included a soup made with chicken broth.
FixedBroth now counts as meat.
ProblemA gluten-free week included bread crumbs.
FixedBread crumbs now count as gluten.
ProblemThe shopping list showed a different tortilla than the one added to the cart.
FixedThe list now shows exactly what goes in the cart.
ProblemA cheaper model looked as good at reading receipts, but read 10 oz of spinach as 10 g.
FixedThe model test now checks amounts, and receipts stay on Sonnet.
The same test inputs went to three Claude models for each AI task: accuracy, then seconds and cost per call. The rule: the cheapest model that scores as well as the best. Bigger wasn’t better on most tasks.
| Task | Haiku 4.5 | Sonnet 5.5 | Opus 5.5 |
|---|---|---|---|
| Describe household | 98%1.2 s · 0.17¢ | 98%2.2 s · 0.39¢In use | 100%2.5 s · 0.81¢ |
| Suggest ingredients | 90%1.2 s · 0.13¢In use | 93%1.9 s · 0.29¢ | 93%2.2 s · 0.59¢ |
| Read receipts | 92%2.4 s · 0.14¢ | 100%2.4 s · 0.34¢In use | 100%3.5 s · 0.91¢ |
| Arrange week | 100%6.3 s · 0.59¢ | 98%8.4 s · 1.80¢In use | 93%11.8 s · 3.93¢ |
Small samples (10 sentences, 10 dishes, 5 receipts, 3 weeks), and the week test checks safety and variety, not taste. The first version of this test only checked that receipt items were found, and Haiku looked as good as Opus. Checking amounts too showed Haiku reading ounces as grams (10 oz of spinach as 10 g), so receipts stay on Sonnet. Haiku lists groceries, where it matched the larger models; Sonnet arranges the week until a test of week quality, not just safety and variety, says otherwise.