Back to Search View Original Cite This Article

Abstract

<jats:p>Background: Dietary assessment is the cornerstone of clinical management and research studies evaluating diet and health. Traditional methods such as food diaries and 24-hour recalls can be burdensome, prone to recall bias, and difficult to adhere to. Image-based dietary assessment using vision-language models (VLMs) offers a potential solution. Objective: Our goal was to benchmark state-of-the-art VLMs for automated food recognition, weight estimation, and calorie estimation using Google's Nutrition5k dataset. Methods: We evaluated 3,229 food images using ten approaches: proprietary VLMs (Gemini 2.0 Flash, 2.5 Flash, 3.0 Flash, and 3.1 Flash Lite; GPT-4o, GPT-4o-mini, and GPT-5 Mini; and Claude Haiku 4.5), an open-source VLM (Qwen2-VL-7B), and a commercial food recognition API (FatSecret). We assessed calorie and weight estimation using Lin's Concordance Correlation Coefficient (CCC) and component detection using Jaccard similarity. Results: Gemini 3.0 Flash achieved the best calorie estimation (CCC 0.767, MAE 80.7 kcal), while Gemini 3.1 Flash Lite offered very comparable accuracy (CCC 0.754) with the highest ingredient recognition (Jaccard 0.655) at the lowest cost among top-performing models ($0.59/1K images). Among earlier-generation models, Gemini 2.0 Flash remained competitive (CCC 0.742, Jaccard 0.621) at a fraction of the cost ($0.10/1K images). A human validation study in which four annotators reviewed 440 images revealed systematic omissions in the original Nutrition5k labels. After correction, the extrapolated ingredient-overlap score for Gemini 2.0 Flash increased from 0.62 to an estimated 0.82, suggesting that raw Jaccard scores substantially underestimate true model performance. Conclusions: Current VLMs can perform automated dietary assessment with reasonable accuracy from single overhead photographs. Our results inform model selection for dietary assessment applications and highlight remaining challenges in calorie estimation and component detection for complex, multi-item meals.</jats:p>

Show More

Keywords

flash using estimation gemini dietary

Related Articles

PORE

About

Connect