NutriIngredientNet v3

Itemised, energy-consistent nutrition from one overhead RGB photo (Nutrition5k, official RGB split). v3 adds a direct totals head and reconciles the itemised breakdown to it, buying top-down accuracy without giving up the audit trail. Scaling all grams by one factor preserves the Atwater identity exactly (verified numerically in §12).

Test results (official RGB split, calibrated T=1.02, thr=0.68, TTA=True)

  • Detection: P 0.61 / R 0.62 / F1 0.62 (freq-prior F1 0.22)
  • Grams MAE: 18.7 g (oracle-presence mean-gram baseline 23.5 g)
  • Per-ingredient kcal MAE (reconciled, TPs): 25.8 kcal
  • Dish kcal MAE by estimator (vs T2 real vocab-covered): hard 118.3 soft 77.1 direct 84.5 recon 83.8
  • Best estimator: soft | T3 (paper-comparable) 76.5 kcal
    • reference: Nutrition5k RGB-only 2D-direct baseline = 70.6 kcal (26.1%)
    • previous version (v2, bottom-up only) = 117.8 kcal (46.6%)
  • OOV ceiling: 7.7 kcal | cost of exact energy consistency: +0.0 kcal
  • Conformal set (lambda=0.25): target recall >= 90%, empirical 0.848, avg set size 11.6
  • Gram intervals: +/- 0.81 x mean_gram, empirical coverage 0.876

Honest limitations

  • Fixed camera geometry. Nutrition5k's camera height never changes, which is the only reason apparent size maps to grams. On an uncalibrated phone photo, scale is unrecoverable from a single view — this model is NOT validated for in-the-wild phone photos.
  • Nutrition = grams x a fixed per-gram table; the network does detection + portion only. That is what buys exact energy consistency and auditability.
  • 164-ingredient vocabulary; out-of-vocab food is invisible (8 kcal ceiling).
  • Conformal guarantees assume val/test exchangeability, which the official split does not perfectly satisfy; empirical numbers above.
  • Depth-as-privileged-information head: enabled (deleted at inference).
Downloads last month
16
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support