What ChatGPT can say about the heart – medically
A test of the natural language model suggests that it can’t replace doctors yet
JUST over two decades ago, the advent of Google changed how we learnt and discovered the world, with knowledge now available just by typing a question in a search bar. In the past year, there appears to be another game changer: ChatGPT.
This machine learning model can interpret written input and generate new text, based on the large datasets on which its neutral network is trained.
While still evolving, ChatGPT has given the layperson a tool for creative writing, essay writing, prompt writing and code writing. With its user-friendly interface and largely coherent and comprehensive answers, it has made inroads in multiple domains – including the healthcare industry.
Putting ChatGPT to the test
The excitement that ChatGPT generated in the medical fraternity has been great. Put through the United States Medical Licensing Exam by a group of researchers, it managed to perform at or near the passing threshold of 60 per cent accuracy.
This study generated plenty of interest, potentially boosting ChatGPT’s credibility to members of the public who might use it to seek health information.
Our research team in National Heart Centre Singapore decided to evaluate how well ChatGPT can explain common procedures and conditions to patients.
We decided to assess the answers provided by ChatGPT about coronary angiography, a relatively common procedure used to evaluate heart arteries for blockages.
We found that its answers were generally presented in a systematic fashion, covering most of the major areas required. The language used was easy to understand, avoiding medical terminology unfamiliar to laypersons without clinical experience.
Yet there were also areas of genuine concern.
First, while infrequent, there were significant factual inaccuracies. Serious mistakes included inaccurately recommending who should undergo coronary angiography to evaluate their arteries.
For example, it recommended that all patients who had experienced a stroke should do so – which is definitely not the case. While a stroke is a risk factor for heart disease, it is not itself an indication – particularly in an acute setting where the blood thinners required for an angiogram may cause bleeding.
ChatGPT also omitted the most important reason someone should do an angiogram: in the case of a heart attack.
Second, ChatGPT seemed inflexible in its recommendations beyond the line of questioning. Its answers only focused on the topic, without providing other relevant and crucial considerations that would typically be brought up in a consultation.
For example, when asked about general causes of chest pain, ChatGPT merely provided causes related to cardiac concerns. In contrast, a healthcare professional would have considered a broader range of possibilities, including respiratory or musculoskeletal issues.
Such responses may lead persons without prior knowledge to ignore other causes – which could be dangerous
Third, all the responses were generic and failed to take into account the specific circumstances of the individual. With an increasing focus on personalised medicine, this shortcoming may become more glaring in the future.
ChatGPT’s shortcomings can be attributed to several reasons.
First, errors have also been seen in many other generated responses from ChatGPT, dubbed “hallucinations”. This occurs when the generated texts are semantically or syntactically sound, but incorrect or nonsensical in content.
Unfortunately, to the reader, the eloquent presentation may cause these mistaken answers to sound trustworthy. It would be unwise to rely on ChatGPT without fact-checking.
Second, the model’s inability to be flexible in recommendations may be due to the scoping of the topic.
This is only acceptable if people are just looking to increase their knowledge; it is likely counter-productive for evaluation, which needs to cover other concerns.
Third, natural language artificial intelligence (AI) models are limited by their data inputs. The latest developments might not be in the data sets on which the software was trained. In the ever-evolving medical landscape, such omissions may cause users to receive outdated information.
No heart for conversation
In another example, our research team looked at the advice that ChatGPT gave to patients with advanced heart failure, and had similar findings.
We also found that while ChatGPT could provide factual data, it was not suitable for topics requiring more empathetic approaches, such as discussions surrounding end-of-life issues.
Some have suggested that these shortcomings may be overcome with rapid development and modification of the underlying mechanics of ChatGPT.
But recent work by researchers from Stanford University and UC Berkeley suggests a change in the behaviour of ChatGPT over time, with fluctuating performance.
Interestingly, their study revealed that the GPT-4 performed significantly worse on math problems and generating code compared to its predecessor, GPT-3.5.
On the whole, while thought-provoking, ChatGPT’s performance remains insufficient to replace a healthcare provider’s role in delivering personalised advice and health management.
Still, as natural language AI models improve and progress, it will be exciting to see how physicians may incorporate these seamlessly into clinical practice, to improve the care delivered to patients.
Dr Samuel Koh is senior resident and associate professor Jonathan Yap is a consultant at the department of cardiology, National Heart Centre Singapore
TRENDING NOW
Grab CEO’s wife Chloe Tong on life with Anthony Tan and finding her purpose
What role can Japan play in Asean’s future?
He built the Vingroup empire. Now South-east Asia’s richest man is handing some key roles to his sons
Asean’s challenge is to become resilient against global geopolitics: former Indonesia trade minister