All Blogs
WearableQA: your AI health coach should be allowed to use a calculator
WearableQA found Python access raised GPT-5.4 accuracy from 51.2% to 71.3%. What AI health coaches need from calculations and personal history.
We give language models enormous wearable histories, then ask them to do arithmetic in prose. It is an odd arrangement. Nobody designing a normal analytics product would remove the calculation engine because the interface had become more conversational.
WearableQA, published on September 4, 2026, gives this habit a useful stress test. The benchmark contains 4,084 multiple-choice questions grounded in 200 people's wearable records, blood biomarkers, and demographics, with up to 500 days of daily measurements per person. It evaluates reasoning about measurements as well as health interpretation. Its results describe benchmark performance; they do not clinically validate a coaching product. WearableQA paper.
One experiment deserves the attention of anyone building an AI health assistant. Giving GPT-5.4 Python access raised overall accuracy from 51.2% to 71.3%, a gain of 20.1 percentage points. Switching between ordinary text formats produced much smaller changes. Allowing the agent to choose additional representations did not improve on the fixed-representation tool setup. Tool-use results, Table 3.
There is a satisfying lack of glamour to the implication. Before buying another week of prompt engineering, give the system a way to calculate.
An AI health coach needs retrieval, calculation and explanation
Suppose a user asks how their sleep changed over the last month. A sensible product has several jobs to do: find the relevant records, define the comparison periods, handle missing nights, calculate a result, and explain it.
Make each step inspectable. The paper tests computational access; a product still needs its own evaluation of the retrieval, calculation and explanation that surround it.
Start by making the calculation reproducible. Specify whether you are comparing calendar months, consecutive four-week periods, or a recent window against a longer baseline. Record which nights were included. A sleep average with three recorded nights should carry different context from one with twenty-eight.
Then let the model explain the output in language someone would choose to read. It can describe the difference, ask for useful context, or acknowledge that the record is too sparse. Those are valuable abilities. They become more valuable when the number underneath them is dependable.
Keep the receipt
For each generated observation, retain the input window, source metrics, calculation version, missing-data treatment, and resulting values. A user need not see an engineering log in the main interface. They should be able to open a concise explanation of what the observation was based on.
This also makes internal review considerably less miserable. If an answer looks strange, you can determine whether the problem came from retrieval, computation, or explanation. Otherwise everyone gathers around a paragraph and guesses which part of the machine had a bad afternoon.
The distinction matters for evaluation. Test calculations against known results. Test retrieval against records that should and should not be included. Test explanations for invented causes, overstated certainty, and claims that the computed result does not support.
One aggregate quality score can hide all of those failures.
A personal baseline needs actual history
WearableQA also examines what happens when longitudinal context is removed. In the Claude Opus 4.6 ablation, accuracy on history-dependent questions fell from 75.9% to 39.8% when historical context was withheld. Window-dependent accuracy barely changed. The researchers separately found that some models handled familiar signal relationships better than user-specific associations present in the measurements. Context and prior-knowledge analyses.
For a builder, our reading is straightforward: knowing a plausible story about sleep and activity is insufficient evidence for a story about this person's sleep and activity.
A useful assistant should be comfortable saying it has not established an association. We would rather have a coach ask one good question than deliver an elegant explanation assembled from general knowledge and a vaguely relevant chart.
This is where wearable data infrastructure earns its keep. Clean units, explicit sources, useful history, and dependable calculations give the conversational layer something solid to work with. Terra's AI health-data interface describes our approach to normalising measurements, building structured summaries and retaining health history for AI products. That preparation belongs underneath the conversation; each product still needs to evaluate its own answers.
Keep the conversation. Keep the curiosity. Let the calculator do its job.