In the previous article, I treated recipe reconstruction as constrained optimization. Ingredient quantities were unknown variables, the nutrition label supplied a target, and an optimizer returned one formulation that matched it.
That answered a practical question: can we build a recipe consistent with published data? It did not establish that recovered recipe was unique.
One best fit hides three things:
- another formulation may match almost as well;
- decimal precision comes from solver, not necessarily evidence;
- compensation between ingredients remains invisible.
If optimizer returns 19.47% sunflower oil, we know that 19.47% works in one fitted formulation. We do not know whether oil is constrained around that value or can move several percentage points while other ingredients compensate.
This time, method changes in four steps. We keep forward nutrition model. We give label values tolerances instead of treating them as exact equations. We combine nutritional fit with ingredient order and declared percentages. Finally, we sample many compatible recipes and summarize them as intervals and correlations.
The output is no longer one supposed recipe. It is a map of what the label constrains and what it leaves unresolved.
Data from the jar
The case study is a chocolate-hazelnut spread. Its ingredient list is:
- Sugar
- Hazelnuts, 22%
- Sunflower oil
- Skimmed milk powder
- Low-fat cocoa powder, 2.5%
- Cocoa butter
- Vanilla extract
- Sunflower lecithin
The nutrition declaration per 100 g is:
| Nutrient | Label value | Model tolerance |
|---|---|---|
| Energy | 558 kcal | 10 kcal |
| Fat | 35 g | 0.5 g |
| Saturated fat | 4.2 g | 0.3 g |
| Carbohydrates | 53 g | 0.5 g |
| Sugars | 51 g | 0.5 g |
| Fiber | 2.6 g | 0.3 g |
| Protein | 6.5 g | 0.3 g |
| Salt | 0.10 g | 0.03 g |
The label values are observed data. The tolerances are modelling choices used in this case study. They account for rounding, ingredient variability and mismatch between representative ingredient specifications and actual factory inputs. Changing them changes posterior ranges.
From a recipe to nutrition
Before running problem backwards, we need calculation in forward direction.
Representative hazelnut specification used here contains 60.75 g fat per 100 g. With hazelnuts fixed at 22% of recipe, their fat contribution is:
The same multiplication applies to every ingredient and every nutrient. Put ingredient compositions into matrix and recipe percentages into vector :
is calculated nutrition per 100 g. Each column of describes one ingredient. Multiplying a column by its percentage gives that ingredient’s contribution; summing columns gives finished nutrition.
Move oil, keep total mass fixed
Sugar compensates every oil change so recipe remains at 100%.
Try increasing oil. Experiment removes same amount of sugar so recipe remains at 100%. Fat rises while sugars fall. One formulation change moves several label values simultaneously, which is why multiple nutrients help constrain inverse problem.
Turn model around
Now and the observed nutrition are known, while the recipe is unknown:
A candidate still has to be a valid recipe:
The ingredient list adds more information. In the European Union, ingredients are generally listed in descending order by weight when used, subject to rules and exceptions in Regulation (EU) No 1169/2011, Article 18. Labels governed by another jurisdiction need the corresponding rule.
For this product:
Two declared percentages anchor sequence:
These numbers immediately imply , , and .
For example, vector below satisfies mass balance, order and both declarations:
It is a feasible recipe, not a recovered recipe. Nutritional fit still has to be evaluated.
Why one optimum is not enough
The previous method asks for a point that minimizes discrepancy:
The optimizer may return oil at 19.47%. That number identifies one minimum of the selected objective under the selected constraints. It does not show how quickly fit deteriorates around the minimum.
Sunflower oil and cocoa butter both supply fat and energy. Raising one while lowering other can preserve total fat within label tolerance. Milk powder affects protein, carbohydrates and sugars together, so changes there may be partly compensated elsewhere.
The forward calculation preserves recipe information. The reverse calculation cannot recreate information the label never recorded:
The question therefore becomes: which recipes remain plausible, and how strongly does the evidence constrain each quantity?
Bayesian formulation
Instead of returning only , we describe posterior distribution:
contains structural label information: mass balance, positivity, order, exact declarations and optional bounds.
The prior is uniform across recipes satisfying those constraints and zero outside them. It does not prefer 18% oil over 19% oil merely because one number looks more natural.
The likelihood scores nutritional agreement. For nutrient with tolerance :
Take fat target of 35 g with tolerance 0.5 g. Candidate predicting 34.9 g has standardized error:
Candidate predicting 37 g is four tolerance units away:
The second candidate receives far less support from the fat observation. The same calculation is added across enabled nutrients.
Explore recipe space
The tool uses adaptive random-walk Metropolis. It starts from a feasible recipe and proposes a mass transfer between two non-fixed ingredients. A proposal might move oil from 19.0% to 19.3% while sugar moves from 46.0% to 45.7%. Total remains 100%.
Proposals breaking order, bounds or exact declarations are rejected. Remaining proposals are accepted according to posterior score. Several seeded chains run independently; warm-up adapts proposal size.
The retained samples give:
- median quantity for each ingredient;
- 90% credible interval;
- correlations between quantities;
- predicted nutrition;
- acceptance and between-chain diagnostics.
A credible interval is conditional on the entered ingredient data, tolerances, order, exact declarations and process assumptions. It does not cover errors outside the model.
What label identifies
Every ingredient receives distribution rather than isolated estimate. Width matters as much as centre.
Hazelnuts are exactly identifiable because jar states 22%. Oil can be more tightly constrained because it contributes strongly to fat and energy, sits below hazelnuts, and differs from cocoa butter through saturated fat. Milk powder influences protein, carbohydrates and sugars simultaneously.
Vanilla and lecithin sit at low concentrations. Moving either may barely affect rounded nutrition panel. Solver could return 0.31%, but label may not justify two decimal places.
Narrow posterior interval means observations constrain quantity strongly under model. Wide interval means many quantities remain compatible. Low identifiability is information about evidence, not sampler failure.
Remove one piece of evidence
The declared 22% hazelnut quantity removes one degree of freedom and fixes a major source of fat, fiber and protein. The interactive experiment runs the same model twice: once with the declaration, once after removing it.
Remove one piece of evidence
Run sampler with or without declared 22% hazelnut quantity. Other label evidence stays unchanged.
Compare hazelnut interval and rest of formulation. Removing one exact declaration enlarges feasible region and lets other ingredients absorb more nutritional contribution. This shows difference between direct evidence and indirect inference from nutrition.
Correlations are part of answer
Individual intervals do not show how ingredients compensate. Sunflower oil and cocoa butter provide clearest example. If plausible samples raise oil while lowering cocoa butter, quantities have negative posterior correlation.
See what ingredients trade off
Sampler ranks quantity pairs by absolute posterior correlation. Strong correlation signals substitution label cannot fully resolve.
Strong correlation means label constrains combined contribution more tightly than split between ingredients. Single optimized recipe hides that geometry.
Similar trade-offs can appear between sugar and milk powder through sugars and carbohydrates, or among low-dose ingredients whose effect sits below label precision.
Limits of model
Ingredient compositions in matrix are fixed inputs. Hazelnut fat value of 60.75 g per 100 g remains 60.75 throughout sampling. Model does not invent uncertainty around supplier or database value. Recipe uncertainty can exist even with perfectly known because inverse problem itself is not unique.
Ingredient specifications could become distributions in a larger model, but that requires defensible measurements or ranges. Adding arbitrary uncertainty would not make result more honest.
Process assumptions matter too. Main case treats spread as direct mixture with no mass loss. Tool supports fixed per-ingredient water loss as limited process approximation.
Suppose 100 g input retains 35 g fat but loses 5 g water. Finished mass is 95 g, so declared fat becomes:
V1 does not infer evaporation, fat loss, volatile loss or reaction kinetics. A sampler quantifies uncertainty inside model; it cannot repair missing process physics.
What this changes
The jar provides enough information to reject many recipes. It fixes hazelnuts and cocoa, bounds neighboring ingredients through order, and constrains nutrient contributions through eight label values.
It still does not disclose one unique formulation. Several ingredient combinations can remain compatible, especially when ingredients share nutritional roles or contribute too little to move rounded label.
Difference from first article is therefore not merely another solver:
| Optimization | Inference |
|---|---|
| Returns one best point | Returns plausible distribution |
| Hides trade-offs | Exposes correlations |
| Precision follows solver | Precision follows evidence and assumptions |
| Builds candidate formulation | Measures identifiability |
Finding one recipe that fits is useful. Showing what else fits, and why, is stronger conclusion.
You can reproduce case study or enter your own ingredients in Recipe Inference from Nutrition Labels.
References
- European Parliament and Council. Regulation (EU) No 1169/2011 on provision of food information to consumers.
- U.S. Department of Agriculture. FoodData Central. Representative nutrient records for hazelnuts, sunflower oil, skimmed milk powder, cocoa powder, cocoa butter and vanilla extract.
- TraceGains Gather. Sunflower lecithin supplier specification.