Uploading vs. building codelists: different approaches for different research needs
- Posted:
- Written by:
- Categories:
As part of the Research Software Innovation team’s work to examine how people build and use ICD-10 codelists in OpenCodelists, we noticed an interesting pattern.
We found that codelists uploaded into OpenCodelists often capture fewer relevant clinical events compared to lists built directly in the tool.
In one analysis, 51% fewer primary diagnosis events and 84% fewer secondary diagnosis events were captured in the Hospital Episode Statistics (HES) data.
This appeared to stem from a higher rate of excluded child codes in uploaded codelists, compared to built codelists. We thought this was likely to be similar across other hierarchical coding systems and datasets.
In Fig. 1 we can see an example of those differences. In the left image, we can see that a researcher interested in COVID-19 clinical events may upload a codelist containing the SNOMED-CT code for ‘Fever’ but not the child code ‘Cough with fever’. In this case, ‘Cough with fever’ will be excluded from the codelist unless the researcher manually clicks the plus sign to include it – put simply, it remains faithful to the codelist as-uploaded.
Conversely, in the right image, if the researcher used the builder and included the same parent code ‘Fever’, OpenCodelists would in this case automatically include additional child codes such as ‘Cough with fever’, unless the researcher chose to exclude them. In other words, it involves a more inclusive approach.

Fig. 1. The differences in default behaviour when uploading (left) and building (right) a codelist; uploading takes a faithful approach, and the builder takes an inclusive approach.
We hypothesised that users didn’t intend for uploaded codelists to be less inclusive, and that the upload process as-is might be leading users to overlook the need to explicitly handle child codes.
Our goal was to help users ensure that their final codelist reflected their intended handling of child codes.
What we did
Over one week, we ran remote usability tests, comparing the current upload process with two prototypes that introduced earlier prompts and clear visual cues about child codes. Fig. 2 shows one of these prototypes.

Fig. 2. One version of the prototype, where we explained that some potential child codes weren’t included in the CSV file, gave users the option to choose how to handle them at the point of file upload, and showed the affected codes and their child codes more visually, in a hierarchical list with clear icons.
What we learned
We noticed several themes across the usability tests:
- Researchers rarely start with a clear plan for handling child codes; intent forms after seeing the codes in context of the hierarchy, i.e. in the visual decision tree in OpenCodelists.
- It’s hard to prompt questions about intent earlier, as explanations in text aren’t an adequate replacement for the visual hierarchy.
- Uploading and building serve different research needs. Differences in included codes may reflect intentional use of pre-curated lists, rather than user error.
We’d like to understand the third theme in particular before concluding that a specific design change is required.
Why this matters
We didn’t come away with a single clear answer, but we did learn something important: decisions about including child codes happen later than we expected, and depend heavily on seeing the code hierarchy.
We also now think that uploading and building serve genuinely different research needs, so staying faithful to the uploaded codelist as-is (the current default behaviour of OpenCodelists), instead of automatically including child codes, may still be the correct approach in many cases. That’s something we’d need to explore further before proposing and making any changes to OpenCodelists.
We intentionally explore new ideas quickly (this piece took two weeks) and, in this case, the approach helped us avoid making any unnecessary changes to OpenCodelists, and helped us learn new information about our users that we didn’t know before.
Overall, the work made us rethink our initial assumptions; uploading isn’t the same as building.
Want to share your thoughts?
OpenCodelists supports a wide range of research, and no single pathway fits every context. We’d like to know more about how researchers use our tools in practice, so that we can test our assumptions and make development decisions based on the evidence.
If you’re someone who uploads codelists to OpenCodelists, we’d love to know:
What leads you to upload a codelist rather than build one directly within OpenCodelists?
Email us at: team@opensafely.org