Why characterise workload in UX?
Characterising the workload associated with a product, an interface, or a service is useful in different contexts. During a creation, optimisation, or redesign effort, for example. This characterisation makes it possible to identify elements with a strong impact on usability (in the sense of the ISO 9241-113 standard). To do this, we cross it with verbal feedback from users and performance markers (e.g. success rate).

What approach should you take to characterise workload in UX?
It would be tempting to use physiological indicators (e.g. pupil dilation) to assess workload. The same goes for biomechanics (e.g. muscle activation level). After all, the mix of technology and signals plotted on a screen is reassuring, with plenty of marketing appeal. But beyond the cost of the kit involved, the expertise and time needed to analyse the data make these approaches cumbersome. That sits awkwardly with any genuinely efficient use in UX.

Besides, the level of precision that actually helps when creating or redesigning a project usually applies to the task as a whole. Subjective tools such as questionnaires remain a sensible investment when you weigh time against results. The NASA-TLX questionnaire is one of these tools.
Lessons learned with the NASA-TLX
For those who would like to know everything about the NASA-TLX, we invite you to read this fine explanatory article. For the others, just remember that the NASA-TLX is a questionnaire made up of 6 components of workload. These are: mental demand, physical demand, temporal demand, self-rated performance, effort, and frustration. Each item is presented in the form of scales ranging from “low” to “high”. That is, 21 unnumbered graduations, translating into scores from 0 to 20.
And in practical terms, what can we say about this questionnaire?

Administration
The original method provides for administration in 2 steps. It is possible to do without the weighting step (see the “analysis” section).
Self-assessment
Following the weighting phase, users place a mark or a cross on the graduations to self-assess their experience. This step is repeated after each task (or sequence of tasks) carried out. The total filling-in time rarely exceeds 2 minutes, but there are a few points to watch out for:
-
users must place their marks / crosses on the graduations, and not between the graduations. The risk? Going back to a 20-point scale (and therefore a score ranging from 0 to 19);
-
the “temporal demand” item is often misunderstood. By re-describing it out loud as “time pressure” to achieve the objectives, users no longer have any doubt and answer easily;
-
the “effort” item is often misunderstood, because it is perceived as redundant with the “mental demand / physical demand / temporal demand” items. It is worth indicating that this item corresponds to a general feeling, to make it easier to understand;
-
users often make the mistake of rating good self-assessed performance by placing their marks towards the right-hand end of the scale. An explanation for this? Counterintuitively, the scale is built from left to right, from a “0” for “good performance” towards a “20” for “poor performance”. And this is not a scale-construction error! Indeed, from the point of view of cognitive load, the labels at the right-hand ends all correspond to high workloads. So be vigilant and make sure that your users have properly grasped this subtlety.

Weighting
The items are presented in pairs: the participant must indicate which of the 2 items prevails in their experience of the workload associated with the task. Let’s take a chess game as an example. The first pair presented will be “mental demand vs physical demand”. The user must indicate whether the workload felt during the chess game was associated more with mental demand or with physical effort. A pretty clear answer, right? Yes, except that it is far less intuitive when it comes to driving a Formula 1 car! Hence the value of this phase.
The same process is then repeated for each possible pair of items: “mental demand vs temporal demand”, then “mental demand vs self-rated performance”, and so on. In total, the user makes 15 comparisons. This phase will make it possible to weight the scores of each item during the analysis phase.

It is worth noting that in the original version, the weighting was carried out after the self-assessment by dimension. Since then, references can be found indicating the reverse order. We’ll let you make up your own mind!
Analysis
If you do not administer the questionnaire on a computer (or tablet/smartphone), you will have to record the scores of the respective scales yourself on the paper questionnaires. A tip to save time: print the questionnaires with scales 20 cm long. You will only have to run a 20 cm ruler over them to read off the scores marked by the users, instead of counting the graduations. A questionnaire is thus analysed in 1 minute, while minimising errors (and reducing the workload 😅).
What about the scores obtained for each item: should they be added up? Averaged? Kept separate? Two parts to the answer:
-
The weighting phase described earlier makes it possible to assign a weight to each item. This weight can be considered either 1) as a way of prioritising the items relative to one another when the items are interpreted independently, or 2) as a way of weighting the scores obtained per item in the calculation of an overall workload value. In this second case, the overall score adds up the item scores each multiplied by their weight (the number of times the item was chosen as the most appropriate to describe the workload associated with the task). The total number is divided by the sum of the weights (15) to arrive at a score out of 20, then multiplied by 5 to arrive at a score out of 100.
-
It has also been reported that the 6 items were correlated with one another, which leads the author of the analysis to think that the 6 items probably measure the same underlying process. Although it is therefore wise to remain cautious and until proven otherwise, it is possible to use the items independently and without weighting (saving time and easing the analysis), with a view to interpreting the results.

Interpretation
There is currently no threshold value above which it is possible to state that a task induces too high a workload. As a result, it makes sense to pair the numerical results with the verbatims and the performance markers, so that together they inform your thinking about workload.
For example:
- Intuitively, if the NASA-TLX score is high and associated with poor performance on the task. One option could be to reduce the task’s workload in order to move towards better performance;
- conversely, if the NASA-TLX score is low and associated with poor performance on the task, one option could be to make the task more demanding in order to raise the workload and reach better levels of engagement with it.
It is also worth remembering that such a score is useful in A/B testing campaigns: although we are not able to state whether projects A and B are too impactful / not impactful enough in terms of load, we are able to say whether one is more so than the other by comparing the scores.
One way around the difficulty caused by the current absence of a threshold value would be to compare the task under study with a reference condition. This would amount to A-B testing, where A would be a reference task and B the task to evaluate.

To wrap up
The majority of the points raised earlier can be overcome by applying the test on a tablet / smartphone / computer. However, it has been shown that the paper and digital versions do not report strictly the same results. So make sure you keep the same measurement methods across users, across tasks, and across test sessions!
On the Akiani side, we have been using this questionnaire for quite a while now, particularly in our neuroergonomics activities (e.g. for cognitive training in e-sports players) but also in UX (e.g. for designing interfaces for the autonomous vehicle).
So, will you give it a go?
References
- Laussu, J. (2018). Charge de travail et ergonomie : histoire et mobilisation d’une notion. Revue des conditions de travail.
- Leplat, J. (1977). Les facteurs déterminant la charge de travail : rapport introductif. Le Travail Humain, 40:2.
- www.iso.org/fr/standard/63500
- measuringu.com/nasa-tlx/
- Hart, SG (2006). NASA-Task Load Index (NASA-TLX); 20 years later. Human Factors and Ergonomics Society Annual Meeting Proceedings, 5:9