Can machines detect ultra-processed foods? A head-to-head evaluation of large language models using NOVA classification


BAYRAM H. M., ARSLAN S., ÖZTÜRKCAN S. A.

International Journal of Food Science and Technology, cilt.61, sa.2, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 61 Sayı: 2
  • Basım Tarihi: 2026
  • Doi Numarası: 10.1093/ijfood/vvag142
  • Dergi Adı: International Journal of Food Science and Technology
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, BIOSIS, Chemical Abstracts Core, Compendex, Food Science & Technology Abstracts, INSPEC, Directory of Open Access Journals, Academic Search Ultimate (EBSCO), Natural Science Collection (ProQuest), Biomedical Reference Collection: Corporate Edition (EBSCO), Engineering Source (EBSCO), Materials Science & Engineering Collection (ProQuest), Technology Collection (ProQuest)
  • Anahtar Kelimeler: food labelling, large language models, NOVA classification, supermarket foods, ultra-processed foods
  • İstanbul Gelişim Üniversitesi Adresli: Evet

Özet

This cross-sectional study compared three large language models (LLMs) (Grok 4.1, Gemini 3, and ChatGPT 5.2) in classifying ultra-processed foods (UPF) using best-selling products from leading supermarket chains covering 53.2% of the national market. Of 3,001 products, 2,920 with complete ingredient information were included; two trained dietitians assigned NOVA groups as the reference standard. In the reference classification, 74.3% of products were UPF. Under the baseline prompt, all models underestimated UPF prevalence compared with the reference standard (p < .001). ChatGPT 5.2 yielded the highest binary UPF detection performance (accuracy: 69.01%; sensitivity: 59.01%; specificity: 98.00%; F1: 73.88%). Prompt sensitivity analyses revealed that a minimal prompt substantially outperformed the detailed baseline for most models (Gemini 3 F1: 94.20%; ChatGPT 5.2 F1: 92.62%), though Grok 4.1's gain reflected a specificity trade-off (sensitivity: 97.79%; specificity: 21.09%). Low inter-run agreement (κ: 0.01–0.27) indicated sensitivity to model updates, supporting prompt calibration and human oversight. These findings suggest that off-the-shelf LLMs require prompt calibration and human oversight before UPF surveillance workflows.