Two weeks ago, I attended the Alzheimer's Association International Conference (AAIC) in London, arguably the most significant conference for Alzheimer’s disease research. This was my second AAIC meeting— the first was in 2016. A decade later, we have multiple approved treatments tau-targeting therapies moving down the pipeline, trial designs increasingly shifting to early disease stages, and, of course, talks of AI around every corner.
Alzheimer’s research seems to be moving confidently into the future — except for the endpoints. And the endpoints are how we know whether any of these therapeutic innovations are actually working.
In a packed auditorium during a session on trial readouts, researchers from leading developers such as Biogen presented their early-phase results. The clinical endpoints, across the board, were the same old ADAS-Cog, CDR-SB, etc… Not one person mentioned a digital or computerized cognitive test. I felt a sense of déjà vu, as I had assumed we would have made progress over the measurement problem we were arguing about a decade ago.
That’s the part that stayed with me on the flight home. Not that progress is hard — of course it’s hard — but that in the one place where I was sure we’d have moved forward, we hadn’t. And, as I discovered later, we’ve actually gone backwards.
We finally found a job for the robots: giving the same test we’ve given since 1984
Instruments older than the industry that uses them
The workhorse cognitive endpoint in Alzheimer’s trials, the ADAS-Cog, was published in 1984 — the same year Apple released the original Macintosh.[1] The MMSE, still everywhere as a screening and secondary measure, was developed in 1975; it is, quite literally, as old as Microsoft.[2] The CDR dates to the early 1980s.[3] These are the tools we still use today - a rater with a clipboard and a stopwatch - to detect whether a drug is bending the course of a neurodegenerative disease.
I don’t say this to sneer at conventional instruments. They were good science for their time, and they’ve earned their place in clinical development. But they were not built for detecting the slow, subtle, early change that disease-modifying therapies are intended to produce, and certainly not for the critical shift to early-stage disease treatment, before significant cognitive impairment occurs. They are episodic, they are rater-dependent, and they are blunt exactly where we now need them to be sharp.
And none of this is a recent realization. The shortcomings of these instruments were understood decades ago — and that recognition is exactly why academics set out to build something better. Computerized cognitive testing was already an active field by the late 1980s. CANTAB, a touchscreen battery, was created at Cambridge by Trevor Robbins and Barbara Sahakian expressly to catch Alzheimer’s earlier than the clinic scales could.[4] Around the same time, Keith Wesnes built the Cognitive Drug Research system, among the first computerized batteries used in drug trials.[5] A decade later, in 1999, David Darby and Paul Maruff founded Cogstate in Melbourne.[6] These are not obscure tools. Collectively they underpin cognitive research cited across many thousands of peer-reviewed studies.
I recall they were already being used as secondary endpoints in many trials a decade ago. So how is it that, 10 years later, I’m seeing no progress at all?
I pulled the numbers — and we went backwards
Maybe my snapshot at AAIC was biased. So with the help of Claude, I dug into ClinicalTrials.gov to look at the data. We pulled every phase 2 and 3 Alzheimer’s trial over the last twenty years — more than a thousand of them — and looked at what they actually used to measure cognition.[7] The ADAS-Cog still headlines about one in five trials as a primary endpoint; MMSE, ADCS-ADL, NPI, and CDR still fill the secondary positions they’ve filled for decades.
If anything has changed, it’s that the ADAS-Cog as a primary endpoint has fallen off over the last five years. However, it was not replaced by anything objective, but rather by other scales: the CDR-Sum of Boxes, the iADRS, the PACC.
So where are the digital cognitive tests? They turn up as secondary endpoints in a small minority of trials — around five percent. And, to my surprise, their adoption has fallen over the last decade. It climbed to a peak around 2011–2015, at approximatley eight percent of trials, then declined to roughly four or five percent today. Read that again, because I had to: there was more computerized cognitive testing in Alzheimer’s trials a decade ago than there is now — in the very years where the digital-biomarker field was supposed to be taking off.
Use of digital or computerized cognitive testing in phase 2/3 Alzheimer’s trials, by study-start period. Source: author’s analysis of ClinicalTrials.gov.
Digital transformation of clinical trials — an oxymoron?
Clinical development has always been a digital laggard. The latest digital transitions —EDC and eCOA, which merely change the form of data capture from paper and pencil to electronic forms,ook the better part of two decades to reach majority adoption. And healthcare sits near the bottom of virtually every cross-industry ranking of digital maturity; McKinsey’s Industry Digitization Index has placed it in the lowest tier since the measure was created.[8]
It isn’t that we don’t need the transformation. The economics of drug development are under real strain. Pharmaceutical R&D productivity has declined so steadily, and for so long, that it has its own name — Eroom’s Law, Moore’s Law spelled backwards: the number of new drugs approved per billion dollars of R&D has roughly halved every nine years for decades.[9] The returns on late-stage pipelines have slid toward — and by some measures below — the cost of capital.[10] If measurement is the bottleneck, and I believe it is, then endpoints aren’t a side issue. They sit close to the center of the productivity problem.
Is AI going to save us?
AI is booming across every industry, and clinical trials are no exception. New biotechs are launched on the promise of AI-driven drug discovery; clinical operations lean on AI for everything from site selection to workflow. But endpoints — the most critical element of the whole enterprise, the thing we rely on to decide whether a drug actually works in a human — are the same ones we inherited decades ago. This raises an almost absurd question: will the future of our industry be deploying robots and large language models to administer and score a word-list test from 1984 (the initial cartoon image)?
The cost of that stalemate is not abstract. Insensitive endpoints are the reason our trials are so enormous, so long, and so expensive. They are why a real but modest drug effect can vanish into rater noise — and why we may have discarded therapies that were quietly working. Every year we spend defending the familiar is a year of patients enrolled in studies that struggle to see what they were built to see.
I left AAIC frustrated, but not hopeless. As a scientist, I spend my days on exactly the unglamorous work the field keeps deferring: establishing the measurement properties of digital endpoints — reliability, validity, sensitivity to change — and we are doing this pre-competitively in our Digital Endpoint Collaboration for Outcomes DEvelopment (DECODE) initiatives, so no single company has to carry the risk alone.[11] The measurement science is ready. The instruments exist. What’s missing is the collective decision to validate them together and give regulators something they can say yes to.
But we also need the industry’s key stakeholders to think differently, and to ask themselves the hard question: does my endpoint strategy reflect the best measurement science of 2026 — or of 1984?
The trial figures above come from my own analysis of phase 2/3 Alzheimer’s disease trials registered on ClinicalTrials.gov; they are a quick descriptive pull, not a peer-reviewed study, and the digital-adoption numbers are a text-based lower bound. But the direction of travel is hard to miss.
Sources
[1] Rosen WG, Mohs RC, Davis KL. A new rating scale for Alzheimer’s disease. Am J Psychiatry. 1984;141(11):1356–1364.
[2] Folstein MF, Folstein SE, McHugh PR. “Mini-mental state”: a practical method for grading the cognitive state of patients for the clinician. J Psychiatr Res. 1975;12(3):189–198.
[3] Hughes CP, Berg L, Danziger WL, et al. A new clinical scale for the staging of dementia. Br J Psychiatry. 1982;140:566–572; Morris JC. The Clinical Dementia Rating (CDR). Neurology. 1993;43(11):2412–2414.
[4] Sahakian BJ, Owen AM. Computerized assessment in neuropsychiatry using CANTAB: discussion paper. J R Soc Med. 1992 Jul;85(7):399-402.
[5] Wesnes KA, McNamara C, Annas P (2016). Norms for healthy adults aged 18–87 years for the Cognitive Drug Research System: An automated set of tests of attention, information processing and memory for use in clinical trials. Journal of Psychopharmacology, 30(3):263–272.
[6] Maruff P, Thomas E, Cysique L, Brew B, Collie A, Snyder P, Pietrzak RH. Validity of the CogState brief battery: relationship to standardized tests and sensitivity to cognitive impairment in mild traumatic brain injury, schizophrenia, and AIDS dementia complex. Arch Clin Neuropsychol. 2009 Mar;24(2):165-78. doi: 10.1093/arclin/acp010. Epub 2009 Mar 25.
[7] Author’s analysis of phase 2/3 Alzheimer’s disease trials on ClinicalTrials.gov (2005–2025).
[8] Manyika J, et al. Digital America: a tale of the haves and have-mores. McKinsey Global Institute; December 2015 (Industry Digitization Index ranks healthcare near the bottom).
[9] Scannell JW, Blanckley A, Boldon H, Warrington B. Diagnosing the decline in pharmaceutical R&D efficiency. Nat Rev Drug Discov. 2012;11(3):191–200.
[10] Deloitte Centre for Health Solutions. Measuring the return from pharmaceutical innovation (annual report series).
[11] https://ametris.com/decode/