top of page

Augmented health: the potential of data in healthcare

  • Jun 12
  • 6 min read

AI-Assisted Self-Diagnosis


One of the most visible transformations in the relationship between citizens and healthcare is the growing use of large language models (LLMs) to obtain clinical information and, increasingly, attempt self-diagnosis.

Today, healthcare is one of the most common ways in which people use ChatGPT.

More than 230 million people worldwide ask ChatGPT questions related to health and wellbeing every week.

A growing proportion of users turn to AI tools to obtain medical information before contacting healthcare professionals. Although figures vary by country, the pattern is consistent: AI has become a “first point of consultation”.


This phenomenon is not new. Even before the era of LLMs, many people relied on “Dr Google”. What has changed is the nature of the interaction: from a passive search on Google to an active dialogue with a model that responds in natural language, interprets symptoms, and suggests differential diagnoses. The data on accuracy are both revealing and concerning. The study Can Public LLMs be used for Self-Diagnosis of Medical Conditions? (2024) compared the performance of GPT-4 and Gemini across 10,000 self-diagnosis cases, achieving accuracy rates of 63% and only 6%, respectively. This is an enormous gap, demonstrating the inconsistency of the tools currently available to the public.


At the same time, the study Human-AI collectives most accurately diagnose clinical vignettes (2025) showed that the best results in clinical diagnosis are achieved neither by AI models operating alone nor by physicians working independently, but by hybrid collectives (groups that combine the analytical capabilities of LLMs with the clinical reasoning of doctors), outperforming either party on its own.

The benefits of AI-assisted self-diagnosis are real: democratisation of access to medical information, a reduction in unnecessary consultations for minor complaints, and lower geographical and economic barriers to accessing health guidance. However, the risks are equally significant. The rapid integration of LLMs into healthcare raises global concerns regarding the quality, clarity, and robustness of generated responses, particularly the risk of medical misinformation.


The 2025 study Medical Misinformation in AI-Assisted Self-Diagnosis (EvalPrompt) warns that the inevitable impact of these models as self-diagnosis tools, and their role in the spread of clinical misinformation, have not yet been properly assessed or regulated.




Ilustração isométrica em tons de roxo sobre telessaúde e monitorização médica. Do lado esquerdo, um médico com estetoscópio surge de um computador portátil rodeado por ícones flutuantes de saúde. Do lado direito, um smartwatch exibe dados biométricos interligados por linhas brilhantes à interface do médico.

Wearables: the explosion of data


The continuous collection of health data began to scale with devices such as Fitbit (2009) and, later, the Apple Watch (2015), which popularised heart-rate sensors, ECG capabilities, and physical activity monitoring. The wearable health-monitoring phenomenon truly took off with the mass-market launch of these consumer-friendly smartwatches and fitness bands.

Since then, growth has been exponential. In 2026, the global wearables market is valued at US$238 billion, with projections suggesting it will exceed US$900 billion by 2035, representing a compound annual growth rate of 16.35%. According to IDC, shipments in 2025 surpassed 600 million devices, including smartwatches, fitness bands, smart hearing devices, and augmented reality equipment.


In terms of adoption, one in three global consumers owns at least one wearable device. In countries such as India (57%), China (53%), and the United Kingdom (52%), more than half of the population uses a wearable. Most users rely on these devices for real-time health monitoring. By 2029, the global number of smartwatch users could exceed 740 million.


What these devices generate is an unprecedented wealth of continuous biometric data: heart rate, sleep patterns, oxygen saturation, electrocardiograms, body temperature, physical activity, and stress markers. The integration of these data with hospital systems and AI platforms creates unique opportunities for the early detection of anomalies, personalised therapies, and the remote monitoring of chronic patients, an area in which wearables are expected to save the healthcare sector more than US$200 billion over the coming decades.



Biases and incomplete data


The quality of the health data used to train artificial intelligence models is one of the most serious and least visible challenges within the medical AI ecosystem. Historical clinical data, including electronic health records (EHRs), reflect decades of inequalities in access to healthcare. As a result, models trained on these data inherit and amplify existing disparities.

The 2024 study Accounting for Bias in Medical Data, based on data from two leading American institutions, found that the rate of clinical testing among white patients is 4.5% higher than among Black patients with the same clinical profile, including age, sex, presenting complaints, and emergency triage status.

The consequences of this reality for AI models are direct and potentially dangerous. When patients from certain ethnic groups are systematically under-tested, the health data available to train algorithms underestimate the prevalence and severity of disease within those groups. Black patients are effectively assumed to be healthier in training datasets, leading resulting models to underestimate disease burden within that population. A particularly striking example involved an algorithm widely used by US hospitals to predict the need for specialised care. The model was trained using healthcare expenditure as a proxy for illness, and because Black patients have historically generated lower healthcare costs despite experiencing the same level of illness, the result was significant bias in care allocation.


In 2024, the Yale School of Medicine published an in-depth analysis entitled Bias in medical AI: Implications for clinical decision-making, concluding that bias can emerge at any stage of a model’s lifecycle: during data collection, annotation and labelling, model development and evaluation, deployment, and even scientific publication. The underrepresentation of women in clinical trials, rural populations in hospital databases, and ethnic and linguistic minorities across virtually all reference datasets creates a systemic risk that cannot be resolved simply by collecting more data. It requires a fundamental shift in how data are collected, processed, and governed.



Data Governance in Health


Faced with challenges of quality, interoperability, and bias, data governance emerges as a necessary, yet frequently overlooked, condition for AI in healthcare to fulfil its potential. Data governance defines who can access which data, for what purpose, in what format, and under what safeguards. In healthcare, where data are simultaneously highly sensitive and of immense clinical and scientific value, achieving this balance is particularly complex.


The UK Biobank model is often cited as an international benchmark for good data governance in healthcare. With rigorous access control mechanisms for researchers, shortened approval timelines that accelerate the research cycle, and a portfolio of thousands of peer-reviewed publications, the UK Biobank demonstrates that it is possible to share health data at scale without compromising participant ethics or privacy.

Data governance is not merely a technical or legal issue; it is a matter of trust. Without citizens’ trust in how their health data are used, the entire edifice of AI-driven medicine is put at risk.



From mass medicine to personalised medicine


The most promising frontier of AI in healthcare is personalised, or precision, medicine: the ability to tailor diagnosis, prevention, and treatment to the unique profile of each patient. This approach has become possible thanks to the convergence of three factors: the sharp decline in the cost of genomic sequencing; the availability of large volumes of structured clinical data; and the capacity of machine-learning algorithms to identify patterns within multidimensional datasets of a complexity far beyond human cognitive capabilities.

In 2024, an AI model known as DeepDRA was developed, capable of predicting with high accuracy whether a particular medication will be effective for an individual patient.

The integration of genomic data with clinical information derived from EHRs and wearables represents the next stage of this transformation.



The integration of genomic data with clinical data from EHRs and wearables represents the next stage of this revolution. Research published in 2025 shows that combining multi-omics data (genomics, transcriptomics, epigenomics, proteomics, and metabolomics) with AI models enables a far more comprehensive understanding of individual responses to drugs.




Ilustração vetorial em tons de roxo e rosa mostrando um robô humanoide que segura uma prancheta e interage com um ecrã. À frente, encontram-se três tubos de ensaio de laboratório e, do lado direito, uma mulher analisa gráficos e dados de um perfil numa interface digital.

An Ecosystem of Data Serving Healthcare


Health data are, in many respects, the new fuel of artificial intelligence, and the quality, breadth, and governance of these data will determine whether AI in healthcare becomes a force for democratisation and equity or an amplifier of existing inequalities. The opportunities are extraordinary: from informed self-diagnosis to personalised medicine, from wearables to the prediction of hospital readmissions. Yet the risks are equally real: algorithmic biases that perpetuate injustice, incomplete data that generate flawed decisions, and interoperability that remains far from ideal.



The path forward requires investment in robust data governance, diversification of training datasets to include underrepresented populations, ensuring that clinical AI remains a decision-support tool rather than a replacement for human judgement, and educating citizens to use digital health technologies critically and responsibly. Artificial intelligence has the potential to become the greatest transformation in medicine since the discovery of antibiotics, but only if it is built upon data foundations that are robust, ethical, and genuinely representative.



We are facing a structural shift: health data have ceased to be merely clinical records and have become an active infrastructure of intelligence. The future of medicine will depend less on the quantity of data and more on the ability to integrate, interpret, and manage it in an ethical and efficient manner.

 

Published in Marketeer magazine

 

 
 
Mão segurando lâmpada
bottom of page