White paper  /  Independent research

The Future of Personalised Nutrition

How AI, wearables and continuous health data are reshaping preventive health

Health technology  ·  Nutrition  ·  Preventive health  ·  Evidence synthesis

September 2026 3,250 words 14 min read Not commissioned by any client
Abstract study of stacked continuous data streams, used as cover artwork

Abstract

Personalised nutrition rests on a finding that is real and repeatedly replicated. People differ substantially in how they respond to identical food. Whether that variation can be measured precisely enough, and acted on reliably enough, to improve health in people who are not ill is a separate question, and the evidence answering it is thinner than the market implies. This paper synthesises peer-reviewed evidence on four converging technologies, continuous glucose monitoring, consumer wearables, algorithmic dietary recommendation and longitudinal multi-sensor data, and separates what the research establishes from what the category asserts. It draws deliberately on null results, methodological critiques and trials whose findings run against the commercial narrative. The conclusion is narrower than the sector's marketing and is not dismissive of it. These tools measure some things well, plausibly shift some behaviour, and have not yet been shown to beat well-delivered conventional dietary advice under a fair comparison. The binding constraint is not sensor quality. It is trial design.

Key findings

  1. Between-person variation is real. Personalising on it has not yet proved its worth.

    Interpersonal differences in postprandial response are large and well replicated. But the four-arm trial of 1,607 adults built to isolate that question found no added benefit from phenotypic, or phenotypic plus genotypic, information over dietary personalisation alone.

  2. Wearables change behaviour modestly, and physiology barely at all.

    A synthesis of 39 reviews covering 163,992 participants found about 1,800 extra steps per day, with small to negligible effects on blood pressure, cholesterol and HbA1c. One 24-month randomised trial found adding a wearable to a behavioural programme produced less weight loss, not more.

  3. The glucose-monitoring evidence base is largely not about healthy people.

    In the most recent meta-analysis of continuous glucose monitoring as a behaviour change tool, only 3 of 25 randomised trials studied populations without diabetes.

  4. Device accuracy is uneven, and unevenly distributed.

    Heart rate is measured well. Energy expenditure and sleep staging are not. Accuracy appears to vary with skin tone in ways the validation literature has not been designed to resolve.

  5. The mechanism the category depends on is the least tested link in it.

    That continuous feedback changes what people actually eat is plausible, widely assumed and supported by very few controlled measurements.

1. The question behind the market

The argument for personalised nutrition is easy to state and hard to test. Dietary guidance has historically been issued at population scale, in the form of averages. Averages conceal variation. If individuals differ in how they metabolise the same meal, then advice calibrated to the individual should outperform advice calibrated to the population, and continuous measurement plus machine learning now makes that calibration technically possible at consumer scale.

Each step in that chain is plausible. The question this paper asks is which steps have been demonstrated, and under what conditions. That distinction matters commercially as well as scientifically. A company whose marketing rests on a step the literature has not yet taken is carrying a risk that grows as regulators, clinicians and informed consumers read the same papers.

Two framings need separating before any evidence is assessed. Technologies developed for people with diagnosed disease and technologies sold to people without it are not the same proposition, even when the hardware is identical. A sensor validated for insulin dosing in type 1 diabetes has been validated against a clinical decision with a clear endpoint. The same sensor worn by a metabolically healthy adult produces a data stream against which no comparable endpoint has been agreed. Most of the confusion in this category lives in that gap.

2. The technology landscape

Four layers stack to produce a personalised recommendation, and each degrades the one above it.

Sensing. Continuous glucose monitors, photoplethysmography-based heart rate and heart rate variability, accelerometry, temperature and, in research settings, microbiome sequencing and metabolomics.

Self-report capture. Logged meals, photographed plates, barcode scans. This layer is where most of the nutritional signal enters, and it is the weakest.

Inference. Models that map inputs to a predicted response or a score. These range from published, peer-reviewed algorithms to proprietary systems whose internals are not disclosed.

Delivery. The app, the score, the nudge. This is the layer that has to change behaviour, and the layer least often measured.

Table 1. What each layer measures reliably, and what the evidence currently supports.
TechnologyMeasures wellMeasures poorlyEvidence status
Continuous glucose monitoring Interstitial glucose trend over time Anything with an agreed clinical meaning in a normoglycaemic adult Strong in diabetes. Sparse outside it
Wrist-worn optical sensors Heart rate, with median error under 5% in six of seven tested devices during cycling Energy expenditure, with no tested device under 20% error Accuracy well characterised. Health benefit modest
Accelerometry and sleep tracking Sleep versus wake Sleep stage classification Behaviour change real but small
Self-reported dietary logging Foods and eating occasions as the user records them Absolute energy and nutrient intake Systematic under-reporting well documented
Algorithmic recommendation Predicting measured postprandial responses within the cohorts studied Demonstrating that acting on them improves outcomes Emerging. Comparator design is the open problem

3. What the evidence supports

3.1 Interpersonal variation in dietary response is real and large

The foundational result came from an 800-person Israeli cohort in which participants wore continuous glucose monitors while 46,898 real-life meals were logged across 5,435 days. A machine-learning model integrating blood parameters, anthropometrics, physical activity, dietary habits and gut microbiota predicted individual postprandial glucose responses better than carbohydrate or calorie counting, was validated in an independent 100-person cohort, and was tested in a blinded randomised trial of 26 further participants.1

Large inter-individual variability has also been reported in a different population. PREDICT 1, conducted in 1,002 twins and unrelated healthy adults in the United Kingdom, measured responses to standardised identical meals and reported a population coefficient of variation, standard deviation over mean, of 103 percent for blood triglyceride, 68 percent for glucose and 59 percent for insulin. Genetic variants contributed modestly to prediction, 9.5 percent for glucose, which is itself a commercially significant result because it implies that measurement adds information a genotype-only test would miss.2

3.2 Self-monitoring produces a modest but genuine behavioural effect

The most comprehensive synthesis of wearable activity trackers pooled 39 systematic reviews and meta-analyses covering 163,992 participants and 390 unique component studies. Its findings were approximately 1,800 extra steps per day, 40 minutes per day more walking, and reductions of approximately 1 kilogram in body weight.3

That is a real effect and it should be reported with its own caveats, which the authors supply. Most of the included reviews were graded low or critically low in confidence under AMSTAR 2, most trials were conducted in high-income countries, and effects on blood pressure, cholesterol and glycosylated haemoglobin were small to negligible.3 Movement changed. Physiology largely did not.

3.3 Glucose feedback improves glycaemic control where it has been tested

A 2024 systematic review and meta-analysis of continuous glucose monitoring as a behaviour change tool pooled 25 randomised controlled trials with 2,996 participants. CGM-based feedback reduced HbA1c by 0.28 percent, with a 95 percent confidence interval from 0.15 to 0.42, and increased time in range by 7.4 percent.4 The effect is genuine. Section 4 addresses who it came from.

3.4 A structured personalised programme can outperform a leaflet

An 18-week randomised trial of an app-based personalised dietary programme in 347 adults reported a significant reduction in triglycerides against standard dietary guidance, a mean difference of −0.13 mmol/L, and no significant difference in LDL cholesterol, the co-primary outcome. Secondary outcomes favoured the programme, including 2.46 kilograms of body weight and 2.35 centimetres of waist circumference.5 It is the strongest commercial-programme trial in the category, and it is also the clearest illustration of the comparator problem set out below.

4. Where the evidence falls short

4.1 The comparator problem

This is the central methodological weakness of the field, and it is not obscure. When a personalised programme with an app, a sensor, a food log and a support community is compared against a leaflet, a positive result cannot distinguish the personalisation from the attention.

One trial was designed specifically to isolate that distinction. Food4Me randomised 1,607 adults across seven European countries into four arms, control advice based on population guidelines, personalised advice based on dietary intake alone, the same plus phenotype, and the same plus phenotype and genotype. Personalised advice outperformed the control on several dietary measures. The authors reported no evidence that including phenotypic, or phenotypic plus genotypic, information enhanced the effectiveness of the personalised advice.6

That result was published in 2017, and this review found no trial of comparable design reporting a different result since. Absence of such a trial is not proof that personalisation adds nothing, and it does mean the claim remains untested at that level of rigour. It should be the first thing any company in this category is asked about, because Food4Me isolates precisely the thing being sold.

The critique extends to the individual trials. Independent commentary on the 18-week personalised programme trial, led by academic dietitian Nicola Guess at the University of Oxford, argued that the comparison was between a full support programme and a leaflet, that only the intervention arm was required to log food in real time, which is itself a behavioural intervention, and that the absence of blinding left expectation effects uncontrolled.7 None of that makes the trial worthless. It makes the effect unattributable to personalisation specifically.

Similar scrutiny has been applied to the foundational glucose-prediction work. In a commentary in the European Journal of Clinical Nutrition, Thomas Wolever argued that much of the variation attributed to differences between people was in fact variation within the same person across occasions, and that the model was never compared against established clinical methods such as fasting glucose or HbA1c.8 Whether or not one accepts every point, the objection about within-person variability is structural. A measurement that varies substantially on repeated occasions in the same individual cannot support a precise individual recommendation.

4.2 Measurement quality, and what it does to everything downstream

Device accuracy is not uniform across what devices claim to measure. In a controlled study of 60 adults wearing seven wrist devices, median heart rate error was between 2.0 and 6.8 percent, and six of seven devices achieved under 5 percent during cycling. No device achieved an energy expenditure error below 20 percent, and errors ranged from 27.4 to 92.6 percent.9 A calorie-balance recommendation built on that output is built on noise.

Sleep shows the same pattern. Against polysomnography in 34 healthy adults, seven consumer devices detected sleep with sensitivity of at least 0.93, but specificity for detecting wake ranged from 0.18 to 0.54, and sleep stage classification was inconsistent across devices.10 Knowing that someone slept is well supported. Telling them how much deep sleep they had is not.

The largest problem is upstream of any sensor. A systematic review of 59 studies in 6,298 free-living adults comparing self-reported dietary assessment against doubly labelled water found that the majority reported significant under-reporting of energy intake, with the degree of under-reporting highly variable even within a single method.11 Every algorithm in this category ingests self-reported food data. A model can be excellent and still inherit that error.

4.3 Endpoints that do not exist yet

In the CGM meta-analysis above, most participants had type 2 diabetes and only 3 of the 25 included trials, 12 percent, were conducted in populations without diabetes. The authors state that further research in such populations is needed and that outcomes other than HbA1c, including glycaemic variability and behaviour change, may be more appropriate there.4 HbA1c is a weak endpoint in a person whose HbA1c is already normal.

Nor is there an agreed standard for interpreting a healthy person's glucose trace. Time in range and similar targets were developed for diabetes management, and the need for interpretation guidance for non-diabetic CGM reports remains an open question in the field rather than a settled one.12 Any in-app score applied to such a trace is therefore a product design decision presented in the visual language of a measurement.

Regulators have drawn one line here explicitly. In February 2024 the FDA stated that it has not authorised, cleared or approved any smartwatch or smart ring intended to measure or estimate blood glucose values on its own, and warned that use of such devices could produce inaccurate readings.13

4.4 Bias, and the difference between a proven problem and an unanswered question

Optical heart rate sensing depends on light reflected through skin, which makes skin tone a plausible source of measurement bias. The honest summary of the evidence is that this is unresolved rather than established. A systematic review of 10 studies covering 469 participants found that four reported significantly reduced accuracy in darker-skinned individuals, four found no effect and two were mixed, and that only six of the ten studies reported Fitzpatrick skin type at all. The authors concluded that preliminary evidence is inconclusive and that larger studies with objective stratification by skin tone are needed.14 The controlled seven-device study cited above did find higher error associated with darker skin tone, alongside higher BMI and larger wrist circumference.9

The defensible statement is therefore narrow and still serious. Validation studies have not been designed to answer the question, which means a company cannot currently demonstrate that its device performs equally across skin tones. Combined with the finding that most wearable trials have been run in high-income countries,3 the evidence base underpinning preventive claims is drawn from a narrower population than the products are sold to.

4.5 The barely tested mechanism

The category's theory of change is that continuous feedback alters eating and activity behaviour, which in time alters health. The middle step is the least measured. In the 25-trial meta-analysis, only 4 studies evaluated the effect of CGM on dietary change and only 5 evaluated physical activity, too few and too inconsistent to support a pooled analysis of the behavioural pathway.4

There is also a direct counter-example, and it is a clean one. In a 24-month randomised trial of 471 adults aged 18 to 35 with a body mass index between 25 and under 40, both arms received identical diet and physical activity prescriptions. The only difference was the self-monitoring method, a website in one arm and a wearable device with web interface in the other. The wearable arm lost less weight, 3.5 kilograms against 5.9 kilograms, a difference of 2.4 kilograms. The authors concluded that among young adults with a BMI between 25 and under 40, adding a wearable technology device to a standard behavioural intervention resulted in less weight loss over 24 months.15 One trial does not settle a field, and this one tested a narrow age band under one specific programme, but a result running against the category's core assumption deserves quoting rather than omitting.

5. Regulation, privacy and claims risk

Two governance facts shape what a company in this space can responsibly say and store.

The first is the wellness boundary. Products that avoid making medical claims generally sit outside the regulatory pathway that would require them to demonstrate clinical performance. That is commercially convenient and it has a cost, because a product that has never been assessed against a clinical endpoint cannot later claim one. The FDA's February 2024 communication is a marker of where the boundary is currently being enforced.13

The second is that most consumer health data is not covered by the framework people assume protects it. Health apps and connected devices that draw data from multiple sources generally fall outside the HHS rule and instead under the Federal Trade Commission's Health Breach Notification Rule, which requires notification to consumers and to the FTC when identifiable health data is disclosed or acquired without authorisation, with civil penalties which that 2021 notice put at up to 43,792 US dollars per violation per day, a ceiling adjusted annually.16 Continuous physiological data is among the most sensitive categories a consumer company can hold, and it is held under a lighter regime than clinical data of equivalent sensitivity.

For a company building in this category, claims discipline is therefore not a compliance tax. It is one of the few available competitive signals in a market where every competitor is making the strongest claim it thinks it can survive.

6. What follows in practice

For companies building products. Separate the two evidence bases in all external communication. The science on interpersonal variation is strong. The science on whether acting on it beats conventional advice is not, and marketing that fuses them inherits the weaker one's exposure. State the population and the outcome whenever a trial is cited. Label interpretive features, scores, grades and spike alerts as product logic rather than measurement, until interpretation standards exist.

For clinicians and health services. The useful question about a patient arriving with device data is not whether the device is accurate in general but whether it is accurate for the specific quantity being acted on. Heart rate and sleep-versus-wake are usable. Energy expenditure, sleep staging and glucose scores in normoglycaemic adults are not, on current evidence.

For individuals. On the evidence reviewed here, the mechanism most likely to be doing the work is self-monitoring itself, a behavioural technique that predates every technology in this paper. That is not a reason to dismiss the tools. It is a reason to be sceptical of the premium attached to the personalisation layer specifically.

7. Outlook

The most consequential development in the field is publicly funded. In January 2022 the National Institutes of Health awarded 170 million US dollars over five years for Nutrition for Precision Health, which aims to enrol 10,000 participants drawn from the All of Us Research Program and to develop algorithms predicting individual responses to food and dietary patterns, explicitly acknowledging that how microbiome, metabolism, nutritional status, genetics and environment interact remains poorly understood.17

Four things would change the assessment in this paper. Randomised trials in people without diagnosed disease, with equalised support and contact time between arms. Behavioural primary endpoints, measured directly rather than inferred. Durability data showing whether any change survives removal of the device. And validation work establishing what a given pattern in a healthy adult's trace actually predicts, which would convert the central feature of these products from a design choice into a measurement.

None of that is exotic. All of it is expensive, and most of it is unattractive to fund for a company whose current claims are already selling.

8. Conclusion

How far can these technologies take personalised preventive health today? Further than sceptics allow and considerably less far than the category claims.

Three things are established. People differ substantially in their responses to identical food. Continuous self-monitoring produces modest, genuine increases in physical activity. Glucose-guided feedback improves glycaemic control in populations with diabetes.

One thing is not established, and it is the thing being sold. This review found no trial of comparable design showing that biological personalisation improves outcomes beyond what careful conventional dietary advice achieves when delivered with the same attention. That is an absence of evidence rather than evidence of absence, and it is the gap a buyer should price in. Until a trial equalises support between arms and measures behaviour directly, the most parsimonious explanation for the category's positive results remains self-monitoring and engagement rather than personalisation.

That conclusion is commercially awkward and it is also an opportunity. The company that funds the fair comparison, and reports it honestly whichever way it falls, will own the only defensible evidence claim in a market currently competing on assertion.

Method and limitations

This paper was researched and written independently in September 2026. It was not commissioned, sponsored or reviewed by any company, and no organisation named in it was consulted. The author holds no financial interest in any product or company mentioned.

Sources were identified by searching the peer-reviewed literature and official regulatory publications for each technology in scope, and then deliberately searching again for null results, methodological critiques and trials reporting effects contrary to the commercial narrative. Where a trade publication is cited, it is cited only as the record of a named expert's public criticism and is labelled as such. Statistics were traced to the original publication and checked against it rather than taken from secondary reporting.

Three limitations should be stated. First, this is a narrative synthesis rather than a systematic review, so it carries selection risk that a registered protocol would reduce. Second, the field moves quickly and several cited trials are recent enough that replication is pending. Third, the author is a food systems researcher and writer, not a clinician, and nothing here constitutes clinical guidance. Claims with clinical consequence should be reviewed by a qualified clinician before publication or use.

References

  1. Zeevi, D., Korem, T., Zmora, N. et al. (2015) Personalized nutrition by prediction of glycemic responses. Cell, 163(5), 1079 to 1094. doi.org/10.1016/j.cell.2015.11.001
  2. Berry, S.E. et al. (2020) Human postprandial responses to food and potential for precision nutrition. Nature Medicine, 26. doi.org/10.1038/s41591-020-0934-0
  3. Ferguson, T., Olds, T., Curtis, R. et al. (2022) Effectiveness of wearable activity trackers to increase physical activity and improve health: a systematic review of systematic reviews and meta-analyses. The Lancet Digital Health, 4(8), e615 to e626. doi.org/10.1016/S2589-7500(22)00111-X
  4. Richardson, K.M., Jospe, M.R., Bohlen, L.C. et al. (2024) The efficacy of using continuous glucose monitoring as a behaviour change tool in populations with and without diabetes: a systematic review and meta-analysis of randomised controlled trials. International Journal of Behavioral Nutrition and Physical Activity, 21, 145. doi.org/10.1186/s12966-024-01692-6
  5. Bermingham, K.M. et al. (2024) Effects of a personalized nutrition program on cardiometabolic health: a randomized controlled trial. Nature Medicine, 30. doi.org/10.1038/s41591-024-02951-6
  6. Celis-Morales, C. et al. (2017) Effect of personalized nutrition on health-related behaviour change: evidence from the Food4Me European randomized controlled trial. International Journal of Epidemiology, 46(2), 578. doi.org/10.1093/ije/dyw186
  7. NutraIngredients (2024) Zoe hails personalized nutrition trial success, results come under scrutiny. Cited as the published record of criticism made publicly by Dr Nicola Guess, University of Oxford, not as independent scientific evidence.
  8. Wolever, T.M.S. (2016) Personalized nutrition by prediction of glycaemic responses: fact or fantasy? European Journal of Clinical Nutrition, 70. doi.org/10.1038/ejcn.2016.31
  9. Shcherbina, A., Mattsson, C.M. et al. (2017) Accuracy in wrist-worn, sensor-based measurements of heart rate and energy expenditure in a diverse cohort. Journal of Personalized Medicine, 7(2), 3. doi.org/10.3390/jpm7020003
  10. Chinoy, E.D., Cuellar, J.A. et al. (2021) Performance of seven consumer sleep-tracking devices compared with polysomnography. SLEEP, 44(5), zsaa291. doi.org/10.1093/sleep/zsaa291
  11. Burrows, T.L., Ho, Y.Y., Rollo, M.E. and Collins, C.E. (2019) Validity of dietary assessment methods when compared to the method of doubly labeled water: a systematic review in adults. Frontiers in Endocrinology, 10, 850. doi.org/10.3389/fendo.2019.00850
  12. Boston University Chobanian & Avedisian School of Medicine (2025) Guidelines needed for interpreting continuous glucose monitoring reports in those without diabetes.
  13. US Food and Drug Administration (2024) Do not use smartwatches or smart rings to measure blood glucose levels. FDA Safety Communication, 21 February 2024.
  14. Koerber, D., Khan, S., Shamsheri, T., Kirubarajan, A. and Mehta, S. (2023) Accuracy of heart rate measurement with wrist-worn wearable devices in various skin tones: a systematic review. Journal of Racial and Ethnic Health Disparities, 10. doi.org/10.1007/s40615-022-01446-9
  15. Jakicic, J.M. et al. (2016) Effect of wearable technology combined with a lifestyle intervention on long-term weight loss: the IDEA randomized clinical trial. JAMA, 316(11). doi.org/10.1001/jama.2016.12858
  16. Federal Trade Commission (2021) FTC warns health apps and connected device companies to comply with Health Breach Notification Rule. 15 September 2021.
  17. National Institutes of Health (2022) NIH awards $170 million for precision nutrition study. 20 January 2022.
More

Other work

All writing

Food systems

The fertiliser gap Africa cannot close with rate alone

9 min read

Sustainability

What 640 field comparisons say about growing two crops on one field

10 min read

Public health

Hidden hunger is a supply problem before it is a diet problem

8 min read

Work with me

Bring me a question that deserves a real answer.

cypriankimomo@gmail.com
Start a brief LinkedIn ORCID