Geometry-Aware Calorie Estimation from a Single Image
This repository implements a novel three-stage pipeline for estimating nutritional values—such as calorie, mass, fat, carbs, and protein—from a single RGB food image. By integrating semantic segmentation, monocular depth estimation, and RGB-D feature fusion, the system provides geometry-aware calorie predictions that are both scalable and practical for real-world dietary applications.
Our Technical Report is here!👻
- Backbone: Mask R-CNN (ResNet-50 + FPN)
- Dataset: FoodSeg103
- Purpose: Isolate food items from the background to provide object-level granularity.

- Architecture: PWCNet-based optical flow + lightweight DepthNet
- Supervision: Scale-aligned triangulation from dense optical flow
- Datasets: Pretrained on NYUv2, fine-tuned on Nutrition5K
- Output: Food-specific depth maps, even from monocular input

- Fusion Backbone: Dual ResNet-101 + Feature Pyramid Network + CBAM + Non-local attention
- Purpose: Predict five types of nutritional values (calories, mass, fat, carbohydrate, protein) from RGB and depth images.
- Output: Calories, Mass, Fat, Carbs, Protein

- Wefine-tunedaMaskR-CNNmodelontheFoodSeg103 dataset to improve food segmentation performance.
- We propose a novel depth estimation approach by in- tegrating optical flow pathways with depth prediction, trained in a self-supervised manner.
- Wesystematicallyintegratedsegmentationanddepthes- timation outputs into the nutrition prediction pipeline, enabling fast and accurate inference from a single RGB image.
- We constructed a new training set using our predicted depth maps and refined segmentation masks on the Nu- trition5k dataset, which significantly improved nutrition estimation in real-world scenarios.
- We developed a lightweight system demo to support practical use and showcase the full pipeline.

