Results
To contextualize model performance, we first examined how well individuals’ self-assessments aligned with RDs expert evaluations, measured by the percentage of decisions that matched across all nutritional goals in our dataset, i.e., not only the four goals selected for subsequent analysis. Overall, participants’ assessments aligned with assessments by expert RDs on 57.6% of meals, with alignment rates ranging from 34.8% to 88.9% depending on the goal ( Figure 2 ).
The lowest levels of alignment (below 50% accuracy) were observed in four of the nutritional goals that are more complex or ambiguous or require estimating portion sizes, such as one forth carb (34.8%), vegetable fat (39.1%), reduced portions (44.0%), and choose plant proteins (45.8%). Moderate agreement (accuracy 50%−70%) occurred for goals such as half fruits and vegetables (51.6%), lean protein (55.7%), and drink water (62.6%), indicating that while individuals generally understood these nutritional goals, some confusion or difficulty remained. At the higher end, about 80% or more of individuals’ self-assessments aligned with RDs for nutritional goals such as no added sugar (79.9%), low fat dairy (80.7%), vegetables (88.2%), and fruit (88.9%).
Given the substantial variability and inconsistency across different goals in self-assessments, especially for goals requiring inference, we next evaluated whether machine learning models could achieve accurate and consistent classification performance.
To address Q1 and Q2—“Can ML accurately classify meals logged in free-text format as meeting or not meeting specified nutritional goals” and “How does ML classification accuracy compare to accuracy of such an assessment by individuals who recorded the meals”, we compared individuals’ self-assessments performance (on the test set) to those of ML algorithms trained on RDs assessments.
Results show that even without enrichment, ML algorithms can accurately classify goal alignment. Specifically, our models reached 0.765–0.827 accuracy and 0.750–0.799 F1 score for drink water goal, 0.801–0.841 accuracy and 0.539–0.730 F1 score for lean proteins goal, 0.729–0.759 accuracy and 0.342–0.603 F1 score for half fruits and vegetables goal, 0.726–0.784 accuracy and 0.000–0.328 F1 score for one forth carbohydrates goal ( Figure 3 and Supplementary Table 7 ). The full set of metrics (accuracy, precision, recall, and F1-score) is reported to provide a comprehensive evaluation, as different use cases may place greater importance on precision or recall.
Notably, all four combinations of ML methods and embedding techniques (logistic regression with TF-IDF, logistic regression with BERT embeddings, multilayer perceptron with TF-IDF, and multilayer perceptron with BERT embeddings) consistently achieved robust and accurate performance, with model-specific performance markers shown within each blue bar in Figure 3 .
Compared to individuals’ self-assessments, ML algorithms achieved higher decision accuracy across the four goals, even without any enrichment of individuals’ recorded meal descriptions ( Figure 3 ). For each goal, ML predictions (blue bars) consistently outperformed individuals’ self-assessments (red bars), as measured by accuracy of ML classifications compared to gold standard (RDs assessments). For the drink water and lean proteins goals, ML methods showed substantial improvements in classification accuracy compared to individuals’ self-assessments, rising from approximately 0.50 to around 0.80 (0.726 to 0.841). For the half fruits and vegetables and one forth carbohydrates goals, ML methods still outperformed self-assessments, though the gains were more modest, with accuracy increasing by less than 0.10.
To address Q3, we evaluated the potential of the different input enrichment techniques to further improve the performance of ML algorithms in assisting individuals’ assessments on their nutritional goals. This analysis showed 3 key findings ( Figure 4 , Table 1 , and Supplementary Table 7 ):
ML with enrichment achieves higher performance
Overall, introduction of enrichment techniques improved the performance of ML methods, across all nutritional goals. Results revealed that combining MLP TF-IDF with FoodOn achieved the highest accuracy of 0.90 (F1 score 0.76), while the LR BERT with Macronutrients Text achieved the highest F1 score of 0.83 (accuracy 0.84). Different enrichment techniques show different performance across different goals
For the drink water goal, LR BERT with Macronutrients Text achieved the highest prediction accuracy of 0.8437. Food Entities, Food Entities & Macronutrients Text, Food Entities & Parsed Ingredients, Macronutrients Text, Parsed Ingredients, and Parsed Ingredients & Macronutrients Text all performed well, with most ML methods exceeding 0.80 in prediction accuracy.
ML with enrichment achieves higher performance
Overall, introduction of enrichment techniques improved the performance of ML methods, across all nutritional goals. Results revealed that combining MLP TF-IDF with FoodOn achieved the highest accuracy of 0.90 (F1 score 0.76), while the LR BERT with Macronutrients Text achieved the highest F1 score of 0.83 (accuracy 0.84).
Different enrichment techniques show different performance across different goals
For the drink water goal, LR BERT with Macronutrients Text achieved the highest prediction accuracy of 0.8437. Food Entities, Food Entities & Macronutrients Text, Food Entities & Parsed Ingredients, Macronutrients Text, Parsed Ingredients, and Parsed Ingredients & Macronutrients Text all performed well, with most ML methods exceeding 0.80 in prediction accuracy.
When predicting the lean protein goal, LR BERT with Food Entities & Parsed Ingredients reached the highest accuracy of 0.846. Food Entities & Parsed Ingredients, Macronutrients Numeric & No Enrichment, Parsed Ingredients, and No Enrichment demonstrated strong overall prediction performance (>0.80 for all ML methods), followed by Food Entities, Food Entities & Macronutrients Text (>0.80 in 3 out of 4 ML methods), compared to other enrichment techniques.
For the nutritional goal of half fruit and vegetables , MLP TF-IDF with Food Entities & FoodOn achieved the highest accuracy of 0.8143. Only Macronutrients Text & FoodOn (0.8122), Parsed Ingredients & Macronutrients Text & FoodOn (0.8062), FoodOn (0.8059), and Food Entities & Macronutrients Text & FoodOn (0.803), all with MLP TF-IDF, exceeded 0.80 in accuracy. Other enrichment techniques across different ML methods performed relatively poorly.
Regarding the nutritional goal of one fourth carbohydrates , MLP TF-IDF with FoodOn attained the highest accuracy of 0.9017. Other than that, MLP TF-IDF with only Food Entities & Macronutrients Text & FoodOn, Macronutrients Text & FoodOn, Parsed Ingredients & FoodOn, and Parsed Ingredients & Macronutrients Text & FoodOn perform relatively better (>0.80) than all other enrichment techniques. Of note, this nutritional goal exhibited lowest precision, recall, and F1 score compared with other nutritional goals, and MLP demonstrated better F1 score compared to LR classifiers (see Supplementary Table 7 ).
Finally, to address RQ4, we compared performance across different enrichment techniques and classification algorithms. Despite the differences described above, some enrichment techniques consistently delivered higher accuracy across multiple ML algorithms. Specifically, Parsed Ingredients and Food Entities & Parsed Ingredients showed strong overall performance, with a high chance of achieving above 0.80 accuracy (7 out of 16 combinations of ML methods and nutritional goals).
Several enrichment approaches, including No Enrichment, Food Entities, Food Entities & Macronutrients Text, Macronutrients Numeric & No Enrichment, and Parsed Ingredients & Macronutrients Text & FoodOn , each had 6 out of 16 methods that surpassed the 0.80 accuracy threshold. However, Parsed Ingredients & FoodOn had only 3 out of 16 methods that exceeded the 0.80 accuracy threshold.
On the other hand, Food Entities & Macronutrients Text enrichment techniques exhibited the best universal prediction performance among different ML algorithms, with 4 out of 16 combinations of ML methods and nutritional goals ranked as top 3 for accuracy. Macronutrients Text is the second optimal option, with 3 out of 16 combinations achieving the top 3 ranking methods in terms of accuracy.
Materials
Our dataset contains meal records collected from two previous studies by the research team. The first study was conducted in 2018–2019 and included a pilot assessment of a chatbot for personalized coaching in type 2 diabetes (T2 Coach) [ 30 ]. The second study was conducted in 2017–2019 and examined users’ engagement with a novel mHealth app for diabetes self-management, Platano [ 31 ]. In each study, participants logged meals using a standard interface that captured: (1) meal type (breakfast, lunch, dinner, or snack), (2) a free-text meal title, (3) a free-text description of the meal, and (4) a photograph of the meal captured with their smartphone. In contrast to past research that already examined approaches to meal analysis using images [ 32 ], in this study, we specifically focus on using free-text user-entered descriptions.
Both studies included features for setting nutritional goals for healthy eating; the goals were developed by a team of Registered Dietitians (RDs) and consistent with recommendations by the United States Department of Agriculture (USDA) MyPlate [ 33 ]. Examples of such goals include “Choose lean proteins” and “Drink water with your meals”; see a description of the full list of nutritional goals in Supplementary Table 1 .
Participants chose a nutritional goal from the offered list and assess each newly logged meal on its alignment with their selected goal. In pilot testing, participants were given multiple goals simultaneously; however, feedback indicated this was overwhelming, and each meal in the dataset used in this study was assessed against a single nutritional goal. In addition to these self-assessments, a team of six registered dietitians used a unified protocol to evaluate each logged meal and assign a binary label of whether the meal met the user’s goal or not. Each RD underwent training with tutorials, practice sessions, and trial tests, achieving at least 90% accuracy against a gold-standard reference set before contributing to the main annotations. The original dataset consisted of 5,284 unique meals recorded in English and Spanish from 114 different individuals recruited from a low-income, ethnically diverse community in a major US city ( Supplementary Table 2 ). To avoid errors due to translation, we excluded meals recorded in Spanish; the final dataset included all meals recorded in English (3,169 meals recorded by 84 unique participants). Although meal photos were collected, they were not included in this analysis because image capture is not consistently available across mobile apps and the quality of images in our dataset varied (e.g., poor lighting).
We selected four nutritional goals (out of 12 available in T2 Coach and Platano): “Drink water with your meals” (short name drink water , 431 labeled meals), “Choose lean proteins” (short name lean protein , 498 labeled meals), “Make half of your meal fruits and vegetables” (short name half fruits and vegetable , 369 labeled meals), and “Make ¼ of your meal carbohydrates” (short name one fourth carbohydrates , 514 labeled meals) ( Supplementary Table 3 ). We chose these goals because they differ in semantic complexity and corresponding self-assessment effort. Specifically, drink water simply requires the presence of water in the meal (or the word “water” in the meal description). Lean protein is more complex, because it requires some inference of which foods are considered lean proteins, and the words “lean protein” are unlikely to be present in the meal description. Both goals are qualitative in nature as they do not consider quantities of particular foods. In contrast, the other two goals are quantitative; furthermore, both require additional inference to identify which items included in meal descriptions can suggest the presence of “fruits and vegetables” and “carbohydrates”.
We converted all text to lowercase and applied a preprocessing pipeline to the meal records’ free-text fields that corrected misspellings, filtered out stop words defined by Python’s Natural Language Toolkit (NLTK) [ 34 ], and removed irrelevant characters (i.e., punctuation, numbers, abbreviations, etc.) ( Supplementary Table 4 ). Using this cleaned data, we concatenated the meal title and meal description into a single text input. This raw text (referred to as “No Enrichment”) was used as our baseline.
We leveraged Nutritionix, an NLP toolkit (Syndigo, Inc) [ 35 ], for computing nutritional content from free-text input. Nutritionix maintains a database of over 1,200,00 food items, which includes entries for restaurant and packaged foods, along with their respective nutritional information. Nutritionix publishes an API that allows users to extract food entities from inputted text data, learn the ingredient composition of a food entity, and query an entity’s nutritional information. Example outputs from the Nutritionix API can be found in Supplementary Table 5 .
We utilized FoodOn, a food ontology developed by agencies across Canada, United Kingdom, and United States to standardize food concept usage in domains such as nutrition and food composition [ 12 ]. FoodOn is structured using concept relationships ranging from food categories and products, to packaging and preservation processes. Since FoodOn is built upon vocabularies derived from the intersection of science and food regulation domains, it is a useful public ontology to tackle challenges in precision nutrition problem spaces. We illustrate in Figure 1 the ontological representation of the FoodOn concept “corn flakes.”
We use the Nutritionix API to determine the food entities (food ingredients) contained in the meal record title and description. We also retrieve the component ingredients contained within each food entity (parsed ingredients) using the Nutritionix API (i.e., the entity “salad” contains ingredients such as “spinach” and “iceberg lettuce”). The Nutritionix API also provides the macronutrient information (calories, carbs, dietary fiber, fat, and protein) contained in each meal record (referred to as macronutrients text). Using the macronutrients text data, we engineer a feature that captures the normalized, numeric sums of a meal record’s macronutrient content (macronutrients numeric). For each food entity identified by the Nutritionix API, we use FoodOn ontology to retrieve the corresponding ontological hierarchy and concepts, and concatenate these concepts into a text representation (referred to as FoodOn).
Based on this plethora of information, we created various combinations of enrichment techniques. Supplementary Table 6 is an example to demonstrate how each technique enriches a sample free-text meal record (i.e., “fried whiting sandwich wheat bread bottle water”).
For each enrichment technique, we concatenate the meal title and meal description with enriched text and represent them computationally via word embeddings using TF-IDF and BERT. Term Frequency Inverse Document Frequency (TF-IDF) is a statistical method that reflects a word’s importance within a document relative to a corpus [ 36 ]. It is used to transform the text into a high-dimensional sparse feature vector based on weighted term frequencies. Bidirectional Encoder Representations from Transformers (BERT) is a pre-trained language model that captures contextual relationships between words by considering both left and right context. It is pre-trained to capture semantic meanings of words more accurately [ 37 ]. BERT-base were applied to convert meal records text into numerical representations. Specifically, we use the [CLS] token at the start of each meal description, to signal the beginning of the sentence, and extract the BERT embedding from the final hidden state of our BERT tokenizer to generate the meal embeddings, without further fine-tuning. We considered both TF-IDF and BERT to compare frequency-based lexical features with semantic embeddings, allowing us to evaluate whether contextualized representations provide improved performance for classifying nutritional goal alignment.
Next, we trained two supervised binary classifiers on the text embeddings to predict whether each meal aligned with a given nutritional goal. We selected two representative approaches: a simple, linear decision rule-based classifier (Logistic Regression, LR) and a more advanced, neural network-based nonlinear classifier (Multilayer Perceptron, MLP). The LR is a supervised machine learning algorithm for binary classification, which predicts the probability of an outcome being equal to 1 or 0 [ 38 ]. The MLP is a feed-forward neural network with one or more hidden layers, which uses ReLU activations, and is trained via stochastic gradient descent for optimal classification performance [ 39 ]. For all models, we employ an early stopping criteria tolerance of 0.0001 and allow for 10,000 maximum iterations for each model’s solver to converge. We utilize a learning rate of 0.001 for the multilayer perceptron classifier models.
To evaluate the performance of the classification algorithms, we used accuracy (ratio of correctly predicted observations to the number of total observations), precision (ratio of true positive predictions to the total predicted positives), recall (ratio of true positive predictions to all actual positives), and F1 score (harmonic mean of the precision and recall). For each metric, we report the mean and standard deviation across multiple runs to reflect variability and robustness in model performance. While we report all four metrics, we highlight accuracy in the main text because it is widely used and easily interpretable across domains. However, we note that the relative importance of different metrics depends on the specific application context and class distribution. Data were split into 80% (training) and 20% (testing) sets. All the analyses were conducted using Python 3.9.
Discussion
In this study, we demonstrated that ML methods, when combined with domain-specific enrichment, can classify whether meals meet specific nutritional goals using PGHD such as free-text meal records. These models significantly outperformed individuals’ self-assessments, achieving an average accuracy score of 0.8 compared to 0.5 for participants (see Figure 3 ).
Overall, our findings revealed three key trends. First, ML models, even without enrichment, outperform individuals’ self-assessments across different nutritional goals. Second, using different enrichment techniques resulted in considerable improvements in classification accuracy, with improvement gains reaching 0.719–0.902 and an average accuracy of 0.785. These improvements were particularly substantial for complex nutritional goals, such as the one fourth carbohydrates goal. Third, different enrichment techniques resulted in different accuracy gains across different goals; however, several enrichment techniques showed consistently good performance. For example, Food Entities & Macronutrients Text -based enrichment showed the strongest universal prediction performance among enrichment combinations across multiple nutritional goals, while Food Entities & Parsed Ingredients also exhibited consistently good performance. Overall, these findings show that unstructured PGHD (e.g., free-form meal logs) with ML analysis and enrichment can be effectively used to make important inferences regarding nutritional choices.
These findings extend previous investigations concerning the use of PGHD in health and wellness. While the challenges of PGHD including noise, inconsistency, and lack of standardization are well established in the literature, most prior work has focused on integrating structured or semi-structured PGHD with electronic health records (EHR) to support retrospective phenotyping or population-level analyses. For example, prior work has demonstrated the potential of ML to derive phenotypes from self-tracked free-text symptom data in endometriosis [ 40 ] and used NLP to extract social and behavioral determinants of health from free-text EHR narratives [ 41 ]. Similarly, free-text data from social media has been classified to infer mental health status, such as depression, based on language patterns [ 42
43 ]. While these efforts center on diagnostic inference or structured phenotyping, our approach diverges in both objective and granularity: we transform naturalistic, meal-level free-text logs into interpretable signals that can enable data-driven self-management interventions. In doing so, we illustrate how task-specific representation learning, when grounded in domain knowledge, can help bridge the gap between noisy, self-reported data and actionable support for everyday health behaviors.
Real-time assessment of meal-goal alignment bridges a critical gap in precision nutrition by enabling timely, patient-facing feedback and behavioral guidance based on free-text meal logs. In practice, this capability can be embedded in mobile applications, providing patients with an always-available tool for real-time self-monitoring and feedback at the point of decision. Accurate assessment of alignment between an individual’s meals and their nutritional goals is fundamental to their ability to follow a healthful diet [ 44 ]. Yet previous research showed that such self-assessment can suffer from low accuracy as individuals often underestimate portion sizes in their meals and overestimate consumption of fruits and vegetables [ 45 ]. This is particularly the case for individuals with low health and nutritional literacy [ 46 ]. Our approach to assessing meal-goal alignment can provide these individuals with useful feedback and improve their chances of reaching their goals, supporting more effective dietary self-management.
Beyond supporting day-to-day behavior change, such systems can be integrated into nutritional therapy workflows and linked with electronic health records (EHRs), providing clinicians with a more complete view of dietary patterns. Traditional patient meal logs are often time-consuming to review and prone to inaccuracies [ 47 ]. In contrast, ML-based classification generates structured summaries of dietary behaviors, enabling dietitians to tailor interventions more effectively and support long-term care.
Our study revealed substantial variability in classification performance across nutritional goals, highlighting both the complexity of certain tasks and the importance of task-specific design. The alignment between participants’ self-assessments and registered dietitian (RD) evaluations varied widely across nutritional goals, ranging from 34.8% for the one fourth carbohydrates goal to 88.9% for the half fruits and vegetable goal. Similarly, machine learning model performance and the impact of enrichment techniques differed by goal. For example, the Food Entities & Macronutrients Text enrichment yielded strong accuracy for drink water (0.8437 with BERT) and lean protein (0.8446 with BERT) but was less effective for “half fruits and vegetables” (0.7435 with TF-IDF) and particularly inconsistent for “one-fourth carbohydrates,” where many enrichment methods failed to improve over the baseline without enrichment. These results highlight that no single enrichment approach performs optimally across all goal types, suggesting that tailoring models to the complexity and structure of specific dietary goals may be necessary to maximize classification accuracy. This variability raises important questions regarding the trade-offs between cross-goal generalizability and goal-specific optimization for solution design. Generalizable solutions are easier to scale but will likely lead to lower accuracy in assessment of meal-goal alignment. In contrast, goal-specific solutions, though less scalable and requiring additional re-training for each newly added goal, enable tailored feature engineering, model tuning, and enrichment strategies that can substantially boost performance for individual goals. A key direction for future work lies in hybrid modeling strategies: can we develop modular or multitask architectures that retain the scalability of general models while incorporating goal-specific nuances? Addressing this question will be critical for building flexible, accurate, and interpretable nutrition-feedback systems at scale.
Finally, in this work, we focused on the more traditional ML approaches to meal classification based on free-text meal descriptions. Given the increasing accuracy and availability of Large Language Models (LLMs), they present a plausible alternative to other text-based types of analysis. In fact, previous work has already explored using LLMs to enrich meal descriptions [ 48 ]. However, in our setting, LLMs face several important challenges including high computational cost, instability over time and sensitivity to drift, domain-specific hallucinations, and unresolved privacy concerns when handling sensitive patient data. Nonetheless, exploring the integration of LLM-based methods represents a valuable direction for future work.
Despite these promising results, there are several limitations. First, the quality of PGHD is dependent on individuals’ ability and willingness to accurately log dietary details. Differences in nutritional knowledge may have influenced how participants assessed meal-goal alignment, and future work should include validated measures of nutritional literacy to better capture its potential moderating role in participants’ self-assessments. Second, our model was trained exclusively on English free-text entries, so its adaptability to multilingual populations requires further exploration. Third, excluding meal images from analysis ensured generalizability across apps and avoided issues with poor image quality; integrating image data with text represents an important direction for future work. Fourth, the Nutritionix API provides default ingredient lists for generic food categories such as “salad,” that may not accurately capture meal-specific variation and introduce noise into enrichment features. Fifth, the simplification of the classification into a binary task helped to assess methodological validity but reduced ecological validity as real-world nutritional management often requires balancing multiple nutritional goals. Finally, some of the choices made during data pre-processing, such as removal of abbreviations (e.g., “tsp” and “tbsp”) may have impacted classification for quantitatively-oriented goals.
Conclusions
This study highlights the potential of applying machine learning methods, particularly when enriched by domain-specific knowledge (i.e., ontologies and nutritional composition), in reliably transforming free-text meal logs into actionable feedback on whether individuals’ meals have met their nutritional goals. Compared to individuals’ self-assessment, ML models achieved significantly higher accuracy and consistency. While no single enrichment technique was optimal across all nutritional goal types, the use of Parsed Ingredients , Food Entities , and Macronutrients Text with BERT or TF-IDF embeddings proved beneficial for complex nutritional goals, which help individuals make more informed dietary choices. These findings underscore the potential of AI-enhanced PGHD analysis to advance evidence-based patient-centered nutrition guidance, particularly important for underserved populations with diabetes who might benefit most from accessible and tailored nutrition support.
Introduction
Patient-generated health data (PGHD), captured by patients through wearable and mobile devices, offers unique, context-rich insights that complement traditional clinical data. PGHD serves as a critical source of information for evidence-based healthcare, with promising potential to benefit high-quality and patient-centered precision care [ 1 ]. The use of PGHD can support clinical outcome assessments, enhance diagnosis and management of chronic diseases, reduce patient burden, and promote health equity [ 2 – 6 ].
Despite its potential, effectively utilizing PGHD for clinical purposes remains challenging [ 5 ], particularly in the case of the emerging field of precision nutrition. In this domain, PGHD includes meal logs and dietary preferences and restrictions, recorded as unstructured free-text entries. These data can be noisy, sparse, inconsistent, and variable in language, thus complicating standardization and integration with clinical data [ 7 ]. Mobile apps for diet tracking, which often rely on unstructured text, suffer from similar problems and can produce idiosyncratic, error-prone logs lacking standardization [ 8 ]. Identifying patterns in such free-text PGHD is essential for nutritional assessment and intervention, yet remains challenging even for expert registered dietitians (RDs) [ 9
10 ]. Common approaches to addressing these challenges include structured digital logging through food databases, which produce data more amenable to computational analysis. However, these approaches are time-consuming, error-prone, and burdensome for patients [ 11 ]. These limitations underscore the need for innovative solutions to effectively analyze free-text PGHD.
Advancements in machine learning (ML) and natural language processing (NLP) offer a promising solution to address these challenges. For example, these methods can help to automatically and cost-effectively pre-process unstructured text and integrate heterogeneous data sources. Furthermore, enriching patient-entered records with domain-specific knowledge in the form of ontologies and nutritional databases can further facilitate computational analysis of nutritional records. For example, FoodOn ontology has been used to improve classification and interpretability, providing a promising path for robust analysis of free-text dietary PGHD [ 12 ]. This enhanced analysis enables personalized dietary recommendation [ 13 ], timely glucose control [ 14 ], and effective weight management.
However, while current research demonstrates the successful application of NLP and ML in nutrition data analysis, their implementation for unstructured text in precision nutrition remains under-explored. Previous studies have primarily focused on structured or semi-structured sources, such as nutrition labels and recipes, linking food descriptions with nutrient fact tables to enable classification and predict nutrition quality and clinical outcomes. For instance, NLP and ML techniques have effectively automated food and recipe classification, ultra-processed food classification, and nutrition quality prediction, using unstructured text input (e.g., recipe or ingredient text) with structured nutrition fact tables [ 15 – 20 ]. Additionally, ML and generative artificial intelligence (AI) models have been applied to analyze dietary patterns and build personalized nutrition recommendation systems based on structured databases [ 21 – 26 ]. However, no study to date has utilized ML or NLP to address the specific challenges posed by unstructured PGHD data, such as free-text meal records. [ 27 ]. Yet, the potential of ML-NLP-driven analysis of unstructured data extends beyond nutrition, including applications to mental health (e.g., patient-reported mood and stress) and disease phenotyping, detection and prevention [ 28
29 ].
This study addresses this gap by evaluating the extent to which ML models, combined with domain-specific enrichment techniques and NLP methods for processing text, can accurately classify meals on their alignment with different nutritional goals using unstructured meal log text. We chose this classification task because it is an essential task in many nutritional interventions but can be challenging for non-experts. Here, we frame the task as a binary classification—whether a meal meets a nutritional goal or not—to establish methodological validity, demonstrating feasibility and laying the foundation for more complex models that approach the classification task in a more nuanced non-binary way. Based on a dataset of over 3000 meal records collected with a mobile app by 84 individuals in a diverse, low-income U.S. community, we compare model predictions with gold-standard assessments from registered dietitians. Our specific research questions include the following:
Q1: Can ML accurately classify meals logged in free-text format as meeting or not meeting specified nutritional goals? Q2: How does ML classification accuracy compare to accuracy of such assessment by individuals who recorded the meals? Q3: Can classification accuracy be improved with addition of different text enrichment techniques and different types of embeddings? Q4: Are there particular types of enrichment techniques/embeddings that lead to improved classification accuracy across nutritional goals and classification methods?
Q1: Can ML accurately classify meals logged in free-text format as meeting or not meeting specified nutritional goals?
Q2: How does ML classification accuracy compare to accuracy of such assessment by individuals who recorded the meals?
Q3: Can classification accuracy be improved with addition of different text enrichment techniques and different types of embeddings?
Q4: Are there particular types of enrichment techniques/embeddings that lead to improved classification accuracy across nutritional goals and classification methods?
Supplementary Material
Supplementary materials are available at Journal of the American Medical Informatics Association online.
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.