Illustration of Babylon logo

Babylon Health: Improving the experience of speaking to a health chatbot


user research analytics ux strategy

About Babylon Health

Babylon Health provides private primary healthcare services that are accessible via an app. They operate globally with a slightly different business model in each country. A service free for everybody is the AI powered triaging tool, also known as symptom checker.

Screencrab
Babylon's telemedicine services include consultations with GPs and a free symptom checker.

The symptom checker is a chatbot that aims to signpost users to the most appropriate service for the symptoms they present.

How does it work? A user is prompted to provide their age and sex and the symptoms they are concerned about, followed by a series of questions to include or exclude certain conditions. At the end of the conversation, based on a probabilistic data model, the chatbot provides a number of explanations for possible causes of the symptoms and recommended levels of care.

My role

I was hired to help the symptom checker tribe with a number of projects aiming to improve the product experience. It included a combined research/design role in a squad that explored improvements to the length and the quality of the automated symptom check conversation. Following this, I worked with a squad that aimed to improve the diagnostic accuracy of a completed symptom check. I worked with clinicians, data scientists, and engineers on internal tooling that would help clinicians to easily run flow simulations with modified data and algorithms so that the impact of modifications could be demonstrated and measured prior to release.

”Ulf ‘s brilliant. I worked closely with him on a complex AI and machine learning problem with lots of stakeholder pressure. And he executed it flawlessly, cooly navigating diverse user needs (doctors, data scientists, senior academics, global users of varying demographic profiles) to plot out a solution.”

Logan Smith , Senior AI Product Manager at Babylon Health

Improving the quality of the question flow

User research had identified two areas that compromised the experience of the symptom checker: Firstly, some of the question-answer flows were excessively long. In some instances users were asked 40 questions or more, at which points most people abandoned the chat. Secondly, some questions were perceived as either irrelevant for the presented symptom or they were deemed inappropriate, such as questions with regards to sexual behaviour. Related to this, users felt that sensitive questions should not be asked too early in the flow.

The underlying algorithm that picked a question from a catalogue of 6000 available questions did not allow for manual intervention based on user research insights. So the squad started experimenting with manually curated question flows for the most frequently presented symptoms. This gave clinicians a greater level of control with regards to flow length, question types, and order of questions.

In order to ensure user insights were well considered, I worked closely together with doctors when they chose appropriate questions and defined the order and decision branches of a flow. I then prepared a number of comparative perception studies to test the difference in experience between the different models with a total of 240 participants. After completing the question flows of the chatbot, users were asked to rate their experience with regards to questions and outcome of the symptom checker.

As expected, versions with manually curated question flows performed significantly better with regards to perceived flow length and question relevance and the new models were released initially to the UK and South Asian markets and then globally.

With valuable data from the perception studies at hand, I did some further digging and analysed the correlation of results between score categories and the overall experience of the Symptom Checker. I could reveal a surprising insight that paved the way for the next piece of work: The experience of the outcome of the chat (i.e. presumed conditions or diseases) correlates stronger with the overall experience than the experience of the questions or the flow of questions.

Screengrab of data showing that experience of results correlate stronger than experience of questions

In other words: Users seem to be more forgiving for an imperfect question experience than for results or advice they feel to be unreasonable. Our focus on improving question flow and relevance was bound to have limited impact on the overall experience if outcome accuracy and advice would continue to perform poorly.

Improving outcome accuracy

Improving the quality of the outcome was now high on the agenda. Outcome refers to both the list of probable causes for the symptoms and the care recommendations at the end of the flow. Whereas care recommendations were under constant scrutiny to ensure compliance with regulations, the diagnostic part of the outcome had not yet received much product love. From user feedback and internal testing we already knew that differentials, i.e. probable diseases or conditions given a user’s presented symptoms, demographic, and risk factors, were frequently off the mark of what a user felt to be likely and appropriate.

Accelerating feedback from experimenting with data

The biggest hurdle for improving outcome accuracy was the effort and complexity of having to generate a new working model after any modification to data or logic. There was no nimble way to test the impact of changes and to ensure that modifications designed to improve outcome accuracy don’t compromise the integrity and safety of the product. Creating a feedback loop that provides quick and reliable results from experiments was made top priority for our squad. Although the outcome of this project would ultimately improve the results for users of the chatbot, the users that I had to cater for in this internal piece of work were primarily doctors and data scientists.

To kick off, I prepared a spreadsheet template to systematically capture and examine the data of chat conversations with problematic results, which were then analysed by doctors. This allowed us to identify patterns and areas of improvement that we should focus on. At the same time I started sketching out a digital solution to record and analyse problem cases and to run experiments with modified data. A software engineer and a data scientist worked on the technical engine to run these experiments on individual cases and simulate the model’s performance on a cluster of 3000 representative cases.

Screenshot of the tool that would allow doctors to analyse and experiment with the underlying logic for a given chatbot conversation. Here, a doctor could add new associations (links) between a symptom and a disease to the model and see how this would affect the outcome.

Crowd-sourcing medical intelligence

One of the suspected culprits for diagnostic inaccuracy was the data set that describes the probability of an occurring symptom to a given disease, so called marginals. The values were collected from in-house doctors who, if possible, would find relevant data in literature or otherwise would give their informed estimate. The marginals applied to the model were then calculated from the values of two or three doctors.

A close inspection revealed some fairly unrealistic marginals, probably due to different medical backgrounds and interpretation of the circumstances. To improve data collection in a rigorous and statistically justifiable way, we explored alternatives how to get data from a much larger cohort of doctors, also from outside the company.

Together with data scientists, I designed an app that –rather than specifying the exact probability of a disease to cause a symptom– would calculate data points from relative probabilities (see picture below). It turned out that doctors found it much easier to vote on “What is more likely to cause xyz?” as opposed to provide exact values.

mobile phone screen
These mock-ups illustrate experiments with different ways to establishing the probability of a disease to cause a given symptom.
mobile phone screen
'This or that' questions turned out to be most successful in establishing marginals from a large and diverse group of medical professionals.

Troublemaker questions

’This or that’ questions turned out to be most successful in establishing marginals from a large and diverse group of medical professionals.

Apart from missing or flawed data, we also had identified problems on a more human level, i.e. miscommunication between the chatbot and the user. We could identify a range of questions in the chat that either users found confusing or open to interpretation.

For example, when asked about ‘sweating a lot at night’, many users agreed, but the kind of night sweat that this question was referring to was much more extreme than what many people experience on a regular basis. Agreeing to night sweat here could easily put you in a group at risk of Parkinson’s disease.

The action we took was a thorough review of these questions with doctors and copy writers. And I provided a systematic classification system for the 6000 questions that would allow the team to interrogate their perfomance across different categories.

Blurred screengrab of a spreadsheet
This new classification system allowed all 6000 questions to be tagged and categorised. It would help clinicians and product people to analyse their performance across different categories.

Loose ends

I left Babylon Health at the peak of the Covid pandemic when squads were re-organised to meet the demands of changing priorities.

There was much more work to be done: Experiments to be carried out, loose ends to be tied up, and products to be built, iterated and turned into working tools.

We had uncovered key issues and established a clear way forward to tackle diagnostic inaccuracies and ultimately to improve the user experience of the symptom checker for millions of users world wide.

© 2026 Ulf Krautmacher
Created using Astrofy