Somewhere on your phone there is probably an app that promises to make you sharper. It flashes up a grid of tiles, asks you to remember where the fish were hiding, and rewards you with a rising score and an agreeable chime. Three weeks in, the score has climbed a long way. The question cognitive psychologists have been arguing about for the best part of two decades is what, exactly, that climbing number represents.

Two very different kinds of getting better

The research literature leans on a useful pair of terms. Near transfer describes improvement that spreads to tasks closely resembling the one you practised: train on a digit-span exercise and you will probably do better on a slightly different digit-span exercise. Far transfer describes improvement that spreads to abilities you never trained at all, and eventually to life away from the screen. Following a complicated lecture. Holding a recipe in your head. Remembering where you left the car.

Near transfer is easy to demonstrate and nobody disputes it: people get better at things they practise. Far transfer is where the entire commercial proposition lives, and it is where the evidence becomes thin and quarrelsome.

One of the cleanest early tests appeared in 2010, when Adrian Owen and colleagues published a large online trial in Nature, run in collaboration with a BBC science programme. More than eleven thousand people trained for six weeks on reasoning, memory, planning or attention tasks. They improved on everything they practised. On untrained benchmark measures of cognition, the training groups were indistinguishable from a control group who had simply spent the same weeks answering general-knowledge questions using a search engine.

Why so many studies disagree with each other

Plenty of smaller papers do report far transfer, which is why the debate has never quite resolved. A 2016 review by Daniel Simons and colleagues in Psychological Science in the Public Interest went through that literature with a set of methodological criteria and explained why so much of it is hard to lean on. The recurring problem is the control condition. If the comparison group is put on a waiting list and asked to do nothing, then by the end of the study the trained group differs from them in two ways rather than one: they practised, and they knew they were the ones being helped. Expectation on its own is enough to nudge scores on cognitive tests, so a design without an active control cannot tell the two apart.

Small samples compound the problem: when a study has forty participants and a dozen outcome measures, something will clear significance by luck, and that something reaches the press release. The reviewers' summary was blunt in its shape: strong evidence that training improves the trained task, moderate evidence for closely related tasks, and very little credible evidence that any of it carries over into everyday cognitive performance.

The letter, and the counter-letter

In October 2014 the Stanford Center on Longevity, working with the Max Planck Institute for Human Development in Berlin, published a statement signed by around seventy researchers in cognitive psychology and neuroscience. Their position was that the scientific literature did not support claims that commercial brain games improve general cognitive performance in daily life or hold off cognitive decline and brain disease, and that marketing in the sector was frequently exaggerated. The statement objected in particular to advertising that traded on the anxieties of people watching themselves grow older.

Within weeks a rival open letter appeared, signed by a comparable number of credentialled scientists, arguing that the consensus statement was too dismissive of genuine findings, particularly in older adults. It was hosted on a website funded by firms in the industry, and several prominent signatories had commercial interests in training products. Disclosure of that sort does not make an argument wrong, and some of the counter-letter's points about specific trials were fair. It does tell you how to read the disagreement: this was not a settled field being disturbed by cranks, but a genuinely contested one in which some of the loudest voices had something to sell.

What a regulator decided

The commercial question got a sharper answer in January 2016, when the US Federal Trade Commission settled with Lumos Labs, the company behind Lumosity. According to the FTC's announcement, the company had advertised that its games would improve performance at school, at work and in sport, delay age-related cognitive decline, and protect against mild cognitive impairment, dementia and Alzheimer's disease, along with reducing impairment linked to stroke, traumatic brain injury, post-traumatic stress disorder, attention deficit disorder and the cognitive side effects of chemotherapy. The agency's charge was that none of this was adequately substantiated. It also alleged the company failed to disclose that some glowing testimonials on its site had been gathered through prize competitions.

The settlement recorded a judgment of fifty million dollars, suspended on payment of two million, and required the company to make cancelling an auto-renewing subscription straightforward. It is worth being precise about what this established. The FTC did not rule that the games have no effect. It ruled that the advertising had run a very long way ahead of the evidence.

The trial that came closest to a real result

The most serious attempt to test cognitive training properly was not an app at all. ACTIVE, short for Advanced Cognitive Training for Independent and Vital Elderly, randomised 2,802 adults aged between 65 and 94 into one of three training programmes, targeting memory, reasoning or speed of processing, or into a control group. Training ran to about ten supervised sessions over five or six weeks, with boosters for some participants.

A decade later, George Rebok and colleagues reported the long-term follow-up. The reasoning and speed-of-processing groups were still outperforming controls on the specific ability they had trained, ten years after a few weeks of practice, which is a striking piece of durability. The memory group's advantage had faded. All three training groups reported less difficulty with everyday activities such as handling finances and managing medication than the control group did, but that outcome came from self-report, and objectively measured everyday performance largely failed to separate the groups.

Read carefully, ACTIVE describes a real, long-lasting, narrow effect, produced by supervised human-led training in older adults, with an ambiguous relationship to daily life. That is a genuinely encouraging result. It is not the same thing as a subscription app making a healthy thirty-year-old cleverer.

What the evidence does support

If the goal is a brain that holds up well, the interventions with the strongest backing are stubbornly unglamorous. Regular aerobic exercise has the best randomised evidence of anything in the field, though even there the measured effects on cognition are modest rather than transformative. Sleep is not optional maintenance; it is when the day's learning gets consolidated, and chronic short sleep degrades memory and attention in ways no puzzle game will offset. Managing blood pressure, blood sugar and smoking protects the vasculature that supplies the brain. Correcting hearing loss appears to matter, partly because unaddressed deafness pulls people out of conversation.

Education, cognitively demanding work and an active social life are all associated with better cognitive ageing, under the general heading of cognitive reserve. Honesty requires a caveat here: much of that evidence is observational, and causation can run backwards. People in the earliest, undetectable stages of decline may withdraw from demanding work and busy social lives, which would produce the same statistical pattern without the intervention doing anything. The associations are consistent enough to act on, but they are softer than the headlines suggest.

A fair verdict on the apps

None of this makes brain training apps a swindle. They are pleasant, well-designed, mildly demanding games that impose structure on ten minutes of a day and give a satisfying sense of progress. Some of them, particularly speed-of-processing tasks in older users, sit closer to the evidence than others. If you enjoy them, there is no reason to stop.

The honest framing is simply narrower than the marketing. You are buying entertainment with a small cognitive workout attached, not insurance against dementia. The real cost of the overclaiming is not the subscription fee but the substitution: an hour of tile-matching that displaces an hour of walking, a conversation, or a decent night's sleep, all of which have better evidence behind them and none of which come with a leaderboard.