Can machines detect ultra-processed foods? A head-to-head evaluation of large language models using NOVA classification
International Journal of Food Science and Technology, cilt.61, sa.2, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 61 Sayı: 2
- Basım Tarihi: 2026
- Doi Numarası: 10.1093/ijfood/vvag142
- Dergi Adı: International Journal of Food Science and Technology
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, BIOSIS, Chemical Abstracts Core, Compendex, Food Science & Technology Abstracts, INSPEC, Directory of Open Access Journals, Academic Search Ultimate (EBSCO), Natural Science Collection (ProQuest), Biomedical Reference Collection: Corporate Edition (EBSCO), Engineering Source (EBSCO), Materials Science & Engineering Collection (ProQuest), Technology Collection (ProQuest)
- Anahtar Kelimeler: food labelling, large language models, NOVA classification, supermarket foods, ultra-processed foods
- İstanbul Gelişim Üniversitesi Adresli: Evet
Özet
This cross-sectional study compared three large language models (LLMs) (Grok 4.1, Gemini 3, and ChatGPT 5.2) in classifying ultra-processed foods (UPF) using best-selling products from leading supermarket chains covering 53.2% of the national market. Of 3,001 products, 2,920 with complete ingredient information were included; two trained dietitians assigned NOVA groups as the reference standard. In the reference classification, 74.3% of products were UPF. Under the baseline prompt, all models underestimated UPF prevalence compared with the reference standard (p < .001). ChatGPT 5.2 yielded the highest binary UPF detection performance (accuracy: 69.01%; sensitivity: 59.01%; specificity: 98.00%; F1: 73.88%). Prompt sensitivity analyses revealed that a minimal prompt substantially outperformed the detailed baseline for most models (Gemini 3 F1: 94.20%; ChatGPT 5.2 F1: 92.62%), though Grok 4.1's gain reflected a specificity trade-off (sensitivity: 97.79%; specificity: 21.09%). Low inter-run agreement (κ: 0.01–0.27) indicated sensitivity to model updates, supporting prompt calibration and human oversight. These findings suggest that off-the-shelf LLMs require prompt calibration and human oversight before UPF surveillance workflows.