TL;DR
- Ego depletion, the willpower-as-fuel theory decision fatigue was popularly built on, largely failed to replicate: a 23-lab preregistered replication of 2,141 people found an effect indistinguishable from zero.
- The narrower claim that decision quality declines across a long run of consequential choices is supported by field data from parole rulings, antibiotic prescribing and cancer screening orders, all observational and all pointing the same way.
- A separate and newer line of work implicates glutamate accumulation in the lateral prefrontal cortex, a mechanism that does not require the willpower-resource metaphor to be correct.
What the original studies claimed
Decision fatigue as most people understand it descends from one research programme. In 1998, Baumeister, Bratslavsky, Muraven and Tice published a set of experiments in the Journal of Personality and Social Psychology arguing that self-control draws on a single limited resource. In the best known, participants made to eat radishes while a plate of warm cookies sat in front of them gave up sooner on an unsolvable puzzle. The inference was that resisting the cookies had spent something the puzzle then lacked.
That paper named the effect ego depletion and set out the strength model: one resource, shared across acts of willpower, drained by use and restored by rest. A metabolic story was layered on top: blood glucose as the fuel self-control burns. A 2010 meta-analysis in Psychological Bulletin pooled roughly two hundred experiments and reported a medium-sized effect, d = 0.62. That number is why the idea left the journals for textbooks and management training, and why popular decision fatigue inherited the parts that were only ever metaphor - the tank that empties, willpower as a muscle, the glucose fix.
The replications that failed
Then the field checked its own work. In 2014, Carter and McCullough reanalysed the evidence for small-study effects - the footprint left when small, imprecise experiments report large results - and found bias severe enough that the apparent effect might be spurious. A follow-up series by the same authors reached the same verdict.
The direct tests were designed to be decisive. A preregistered multilab replication in Perspectives on Psychological Science ran one standardised depletion protocol across 23 laboratories with 2,141 participants. The pooled effect was d = 0.04, with a confidence interval that comfortably contained zero. Michael Hagger led both that replication and the 2010 meta-analysis it undercut: the correction came from inside the research programme.
A larger attempt followed in Psychological Science in 2021 - 36 laboratories, 3,531 participants, a design agreed in advance, with Roy Baumeister among the co-authors. The confirmatory result was again non-significant at d = 0.06, and the Bayesian analysis favoured the null over an informed prior by roughly four to one.
The glucose account fared no better, and the way it failed is instructive. Molden and colleagues showed in Psychological Science in 2012 that exerting self-control did not increase carbohydrate metabolism, and that rinsing the mouth with a carbohydrate solution and spitting it out - delivering no fuel to the brain - restored performance about as well as swallowing it. Whatever the sugar was doing, it was signalling rather than fuelling.
Key idea
Asked whether the willpower-tank version of decision fatigue is real, the honest answer is that the evidence does not support it. Nothing measurable runs out.
That is an awkward thing for a company in this category to publish, which is most of the reason to publish it. Much of the decision fatigue research quoted in supplement marketing is this literature, cited at its 2010 confidence rather than its 2021 confidence. Telling a durable finding from a marketed one is a skill worth having on its own.
What held up when the lab paradigm did not
Something did survive, and it is narrower than the theory built on top of it. Set willpower-as-substance aside and ask a plainer question: does the quality of consequential decisions decline across a long, uninterrupted run of them? For that, the strongest evidence was never the lab task but field data from people deciding real things.
The canonical case is Danziger, Levav and Avnaim-Pesso's 2011 analysis in PNAS of parole rulings by Israeli judges. Favourable rulings began a session at roughly 65 percent and fell towards zero by its end, then jumped back after a food break.
That study is also the best available lesson in reading such data. Weinshall-Margel and Shapard replied in PNAS that case ordering was not random - the board cleared one prison's cases before breaking, and unrepresented prisoners tended to be heard late - so the pattern could be an artefact of scheduling. The original authors answered that it persisted with representation controlled. Andreas Glöckner then showed in Judgment and Decision Making in 2016 that the headline drop is implausibly large and partly reproducible by scheduling alone, because favourable hearings run longer. The direction is probably real; the magnitude is nothing like 65 percent to zero.
Clinical medicine offers more useful replications, because the same within-session pattern appears in independent health systems with unrelated outcomes. Linder and colleagues reported in JAMA Internal Medicine in 2014 that primary care clinicians grew more likely to prescribe antibiotics for respiratory infections as clinic sessions wore on. Hsiang and colleagues reported in JAMA Network Open in 2019 that ordering of guideline-recommended cancer screening fell with later appointment times.
"The claim that survived is not that willpower runs out. It is that a long sequence of consequential decisions degrades the ones at the end of it."
These studies are observational too, and their limits are not decorative. Late appointments are not randomly assigned: clinics run behind, patient mix shifts, and clinicians are hungry and ordinarily tired. What makes the pattern hard to wave away is that it recurs across unrelated settings and outcomes, always in the same direction - which is why it surfaces most consistently in clinical judgment under sustained load.
The neurochemical evidence for a real mechanism
The third claim is the newest, and a different argument rather than a rescue of the old one. In 2022, Wiehler and colleagues published a study in Current Biology using magnetic resonance spectroscopy to ask what accumulates in the brain over a working day of demanding cognitive work. One group did hard versions of cognitive control tasks, the other easy versions for the same duration.
Only the hard-work group showed rising glutamate in the lateral prefrontal cortex, the region most associated with cognitive control, and glutamate was the only metabolite to show the predicted pattern. That group also shifted its choices towards smaller immediate rewards over larger delayed ones, the behavioural signature of weakened control. The authors' framing: fatigue may arise not because something is depleted but because taxing a region leads to metabolites accumulating within it.
That does not require willpower to be a substance, a tank or a budget. Nothing runs out. Something builds up. Those are different claims with different remedies: a depletion account predicts refuelling helps, an accumulation account predicts clearance does. The glucose experiments argue against the first; the spectroscopy work is a reason to take the second seriously. More on how spectroscopy captures this and glutamate's role in decision quality.
The caveats are those of any young literature. Spectroscopy measures metabolite concentrations indirectly, samples are small, and correlation is not cause. This is an active research question, not a settled mechanism.
Where our own trial fits
My own work sits inside this third claim, and its limits belong before its result. We ran a randomised, placebo-controlled crossover trial in Phoenix, Arizona, published in Frontiers in Nutrition in 2025. Twenty-three healthy adults each completed two 13-hour sessions of sustained competitive load, ten matches per session, separated by a seven-day washout, taking Numin in one arm and placebo in the other. The placebo arm declined significantly in cognitive efficiency after about four hours; the Numin arm held baseline to the end, with 43% fewer decision errors and zero adverse events.
Twenty-three participants is a small sample. Competitive e-sports players under 13 hours of load are a specific population, a crossover design does not make a single trial conclusive, and Numin designed, funded and published this study. The right weight for a company's own single trial is suggestive pending independent replication, not settled. What would raise our confidence is what we would ask of anyone else: preregistered replications by independent groups, at larger samples. The mechanism is on the Numin science page.
Where the science actually stands
Three claims, three verdicts. The popular willpower framing is weak and should be retired. The observation that decision quality degrades over a long day of consequential choices is well supported, most credibly by field data whose direction is consistent even where its limits are real. The mechanism is an open question, with glutamate accumulation in the lateral prefrontal cortex the most promising candidate.
Numin is built on the third claim, not the first. It supports the brain's natural glutamate clearance system, so it is designed to prevent a decline from your own baseline rather than push performance above it. It has no caffeine and no stimulants and does not manipulate dopamine or norepinephrine to force more output: closer to keeping a processor from overheating than to overclocking it. Five named ingredients, one job each, no proprietary blend. A 20-pack is $54 on subscription, one sachet a day, after lunch.
One boundary: this is a post about an evidence base, not about symptoms. Cognitive decline that is severe, present from the morning, or new for you is a clinical question, not a decision fatigue question, and belongs with a doctor.
If you came here asking whether decision fatigue is real, the question packs two claims together and they carry very different weights of evidence. The afternoon decline in your judgment is measurable. The tank is a metaphor that did not survive replication. We would rather build on the part of the science that holds up, and say so where it does not.
Frequently asked questions
Is decision fatigue a real thing or just an excuse?
Was ego depletion debunked?
Is willpower a limited resource?
What evidence is there that decision fatigue is real?
Is Numin's clinical trial independent?
More from the Numin journal
Keep reading
Decision Fatigue
Decision Fatigue Is Real. Here's Why we Spent Four Years Building Something For It.
The honest story of building Numin — a clinically tested supplement for decision fatigue. Nine months in market, more content than companies ten times our size, and still fighting to...
Read article
Decision Fatigue
Why Decision Fatigue Doesn't Care How Good Your Day Was
Slept well, good mood, no real stress, and you still snapped, stalled, or dropped the ball. Here's why a good day doesn't guarantee a good decision.
Read article
Decision Fatigue
Why the Last Candidate Rarely Gets the Job
Interview five people back to back and the fifth is assessed by a different brain than the first. What research on sequential evaluation reveals about hiring decisions and decision fatigue.
Read article