Our Approach · The History · Act V

2005 to 2010 · Free Energy

Karl Friston

The theoretician who cast the brain as a prediction engine bound to minimize surprise

Karl Friston is the architect of the free energy principle, the 2006 claim that any system which persists must act to minimize the surprise of its own sensations. His active inference framework recast perception as inference and motor commands as proprioceptive predictions, and it priced regulation in one currency: prediction error, paid in metabolic energy. That price on mismatch is what his work gives the Unified Model of Tone.

Fportrait
forthcoming

Born

12 July 1959, York, England

Key paper

A free energy principle for the brain, 2006, Journal of Physiology-Paris 100, pages 70 to 87

Core claim

Free energy is an upper bound on surprise, so minimizing it minimizes surprise

Honours

Fellow of the Royal Society 2006, Weldon Memorial Prize 2013

THE CLAIM

The brain is a prediction engine, and prediction error is its currency

Karl Friston proposed that any system which persists must minimize a quantity called variational free energy, a mathematical upper bound on surprise. He set the argument out with James Kilner and Lee Harrison in A free energy principle for the brain, published in 2006 in the Journal of Physiology-Paris, volume 100, pages 70 to 87 (Friston 2006). The claim is deliberately broad. It is meant to hold for a cell, a cortex, a spinal segment, a whole person. Friston was not describing one brain region or one reflex. He was describing what it takes for anything to keep its shape while the world keeps changing around it.

Ask what that commits you to. If a nervous system exists to keep surprise low, then perception is not reception. It is inference. The senses do not hand over the world. They hand over evidence, and the system guesses at what produced it. Every guess carries an error, and that error is the only quantity the system can actually work with. Friston named the error, gave it units, and showed that the same units cover perception, movement, attention and learning. That is the move. It changes what a clinician is looking at when hands go on a body, because it turns the difference between expectation and sensation into a measurable load rather than a figure of speech.

FIRST INSIGHT

A boy turned over a log and saw a gradient, not a purpose

Friston dates his own thinking to the age of eight, under an overturned log, in one of the hot English summers of the 1960s. He watched woodlice scatter and realized they were not seeking darkness at all. They were simply moving faster when warmed by the sun. He tells the story himself in The history of the future of the Bayesian brain, in NeuroImage volume 62, page 1230 (Friston 2012). He says he drew on that notion for the next forty years, while learning natural selection, information theory, machine learning and statistical thermodynamics.

Why should a childhood observation matter to a theory of the brain? Because it fixes the grammar of everything that follows. Behavior that looks purposeful can fall straight out of a simple rule about rates and gradients. No intention is required, and no homunculus is needed to hold the intention. Carry that grammar upward and a guarding muscle, a held breath, a shortened stride and an avoided rotation stop looking like decisions. They start looking like a system descending a slope it cannot see. That reading is available to anyone who works with living tissue, and it is more forgiving than the alternatives, because it stops treating a protective pattern as a character flaw.

LINEAGE

The instruments came first, then the theory, then the people who forced it

Before he wrote a line about free energy, Friston built the software that most of brain imaging runs on. He invented statistical parametric mapping, the framework for asking where in a brain volume a signal has genuinely changed. His group developed voxel-based morphometry for comparing tissue across whole brains without drawing regions by hand. In 2003 he introduced dynamic causal modeling with Lee Harrison and Will Penny, in NeuroImage volume 19, pages 1273 to 1302 (Friston 2003). That method estimates how one region drives another, rather than merely which regions light up. The order matters. Friston reached theory through measurement, which is why his theory keeps asking what would be measured.

The people are documented too. He traces perception as inference to Hermann von Helmholtz, whose work on physiological optics he cites at 1860 and 1866. In 1990 Semir Zeki telephoned him and drew him and Richard Frackowiak into an informal functional integration club with Horace Barlow, Graeme Mitchison and Peter Foldiak. Barlow supplied the principle of minimum redundancy, published in 1961, and with it the idea that the nervous system optimizes something countable. At the Neurosciences Institute under the Nobel laureate Gerald Edelman, Friston put mathematics on a whiteboard. In his own account he was rusticated to the library for six months, to read from Charles Darwin to Ernst Mayr. Later, at the Gatsby Computational Neuroscience Unit in London, Geoffrey Hinton showed him that the variational methods Richard Feynman had used in statistical physics could evaluate the evidence for a model. Friston wrote his thoughts to Hinton in a letter, never heard back, and published them as the free energy principle about four years later.

The free energy principle states that systems change to decrease their free energy.

Karl Friston, James Kilner and Lee Harrison · A free energy principle for the brain, 2006, page 71

THE QUANTITY

Free energy is a bound on surprise, not a fuel gauge

Friston is explicit that his free energy is an information quantity and not a thermodynamic one. In The free-energy principle: a unified brain theory?, published in Nature Reviews Neuroscience volume 11 in 2010 (Friston 2010), he defines it in the margin. Free energy is a measure that bounds, by being greater than, the surprise on sampling some data given a generative model. In the body of the same paper he states that it is an information theoretic quantity, as opposed to a thermodynamic quantity. Hold onto that sentence, because the confusion is everywhere. Friston is not claiming the brain minimizes calories.

Why a bound at all? Because a system cannot compute its own surprise. To do that it would have to know every way the world might have produced its sensations, and it does not. Free energy is the quantity it can compute from its own model, and it always sits above surprise, so pushing the bound down drags surprise down with it. Friston makes the biology vivid on the first page of the review, with a fish out of water as the example of a state that is surprising both emotionally and mathematically. Entropy, in this language, is nothing more than the long-run average of surprise. Staying alive means keeping that average low, which is why physiology reduces so cleanly to homeostasis.

TWO MOVES

Perception changes the prediction, action changes the sensation, and Friston calls action active inference

A prediction error can be reduced in exactly two ways. Change the prediction to fit the sensation, which is perception. Or change the sensation to fit the prediction, which is action. Friston calls the second one active inference, and it is the hinge of the entire framework. In the 2010 review he gives the everyday version: feeling our way in darkness, we anticipate what we might touch next and then try to confirm those expectations. Reaching is a hypothesis. The fingertips are the experiment.

Notice what this does to the old boundary between sensing and moving. It dissolves it. Perception and action become two directions of one descent, one adjusting the model, the other adjusting the sample. Ask what that implies about a body that has stopped moving well. A person who no longer rotates their neck is not simply avoiding pain. They are running a policy that keeps incoming evidence inside the range their model already expects, and that policy is doing exactly what the mathematics says it should. The restriction is not a failure of the system. It is the system succeeding at the wrong problem.

The free-energy principle suggests that we should sample the world to ensure our predictions become a self-fulfilling prophecy.

Karl Friston, Jean Daunizeau, James Kilner and Stefan Kiebel · Action and behavior: a free-energy formulation, Biological Cybernetics 102, 2010, page 232

PROPRIOCEPTION

Motor commands are predictions about what the body should feel

In 2013 Rick Adams, Stewart Shipp and Karl Friston argued that descending signals from motor cortex are not commands to muscles but predictions of what those muscles should feel. The paper is Predictions not commands: active inference in the motor system, in Brain Structure and Function volume 218, pages 611 to 643 (Adams and Shipp 2013). Their starting puzzle is anatomical. Projections descending from motor cortex look, in their laminar and physiological signature, far more like the backward connections of visual cortex than like forward driving connections. Motor cortex is also agranular, missing the input layer you would expect in a region that issues commands. Both facts fall into place if the descending traffic is prediction rather than instruction.

The circuit they describe is specific and testable. Proprioceptive input arrives from muscle spindles by way of Ia and II afferents, from Golgi tendon organs by way of Ib afferents, and from articular and cutaneous receptors. A descending proprioceptive prediction meets that input at the spinal cord, and the difference is a prediction error. Alpha motor neurons fire to produce the shortening the prediction requires. Gamma motor neurons set the sensitivity or gain of the spindles themselves. The classical reflex arc, the same loop a tendon hammer tests at the knee, becomes the machinery that resolves proprioceptive prediction error. In their phrasing, peripheral proprioceptive prediction errors are or become motor commands.

This explanation appeals to active inference, in which higher cortical levels send descending proprioceptive predictions, rather than motor commands.

Rick Adams, Stewart Shipp and Karl Friston · Predictions not commands: active inference in the motor system, Brain Structure and Function 218, 2013, page 611

Friston and the model

What the body predicts, beneath everything else, is itself

The Unified Model of Tone takes active inference for what it proposes and for what supports it: the nervous system runs on prediction, and prediction error is costly. Friston did not propose tone, and nothing in his work should be read as endorsing it. The step past his work is the model's own. What the body predicts, beneath everything else, is itself, and the variable it predicts is tone. Descending commands are predictions of what the body should feel, and the spinal reflex arc resolves the difference, the circuit Adams, Shipp and Friston published in 2013 (Adams and Shipp 2013). The body's account of itself and the body's actual state must be kept in correspondence, and the work of keeping them so has a price.

The price is countable. Neural information processing is expensive at the level of the individual signal: the neuroenergetics literature prices a single bit carried across a chemical synapse in thousands of ATP molecules (Laughlin 1998). The brain carries about a fiftieth of the body's mass and consumes roughly a fifth of its energy (Raichle and Gusnard 2002). Information in this system is bought rather than given. When predictions are chronically violated by corrupted sensory input, the brain raises cortical firing rates and recruits extra processing networks. A nervous system in good tone is cheap to run, because its predictions match its reality. A nervous system in poor tone is expensive, because the mismatch consumes resources the body needed for something else.

The model states one energetic prediction here and offers it as a prediction rather than a result in hand. Two stimuli matched in intensity and differing only in predictability should differ in metabolic cost, and the difference should scale with the size of the mismatch rather than the strength of the stimulus. A metabolic premium on unpredictability, scaling with the size of the mismatch, confirms this account at its foundation. The measured cortical changes Heidi Haavik reported after spinal manipulation sit inside the same accounting: a change in the input is a change in what the brain must spend to process it.

The free energy principle is a contested account, and the model is not built to depend on it. If active inference is superseded, the surviving claim is the one the model makes for itself: the mismatch between a body and its own account of itself is expensive, and the expense is clinical. The claim that the deepest prediction is the body itself, and that its name is tone, belongs to the model. The extension is the model's. The mathematics is his.

PRECISION

Confidence is a gain control, and gain has clinical consequences

Precision is the confidence a nervous system assigns to a signal, and in the neuronal implementation it is a gain. Friston builds it into the framework as a weighting on prediction errors, plausibly carried by neuromodulators and by the excitability of the neurons that encode error. Attention, in this account, is the optimization of precision rather than a separate faculty. Two nervous systems with identical predictions and identical raw errors will behave completely differently if they weight those errors differently.

This is where amplitude enters the model without any hand waving. A system can be wrong about what it expects, or it can be right and simply too confident. Overweighted proprioceptive precision makes a small mismatch feel enormous and worth guarding against, and guarding generates more mismatch. Underweighted precision lets genuine signals go unregistered until the tissue has taken real damage. Anyone who has watched two people with near-identical imaging findings behave in completely different ways has watched precision at work. Frequency and amplitude are separate variables in every oscillating system, and a nervous system is no exception.

INWARD

The same loop runs on the viscera, and that is where feeling comes from

Anil Seth and Karl Friston extended active inference to the interior of the body, where autonomic reflexes carry out descending predictions. Their paper, Active interoceptive inference and the emotional brain, appeared in Philosophical Transactions of the Royal Society B, volume 371, issue 1708, article 20160007, in 2016 (Seth and Friston 2016). Their formulation is that bodily states are regulated by autonomic reflexes that are enslaved by descending predictions from deep generative models of our internal and external milieu. The word enslaved is theirs, and it is chosen precisely.

Follow the implication. If autonomic reflexes are error-correcting loops, then heart rate, gut motility, vasomotor tone and the depth of a breath are all being steered by expectation rather than by fixed dials. Seth and Friston write that emotion becomes an attribute of any representation that generates interoceptive predictions. They also argue that the balance between homoeostatic reflex and goal-directed allostatic behavior rests on the confidence, meaning the precision, placed in deeper expectations about how we will behave. Feeling and regulation stop being separate departments. They are one loop, sampled at different depths.

BOUNDARY

A Markov blanket is the smallest definition of a self

A Markov blanket is the statistical boundary that separates the internal states of a system from everything outside it. Friston argues that having one is nearly enough to look alive. He made the case in Life as we know it, published in the Journal of the Royal Society Interface, volume 10, article number 20130475, in 2013 (Friston 2013). The paper offers a heuristic proof and a simulation of a primordial soup. The argument runs from geometry. If coupling among an ensemble of systems is mediated by short-range forces, remote systems must be conditionally independent, and that independence induces a blanket.

The consequence he draws is startling, and he states it as a heuristic rather than a theorem. Anything with such a blanket will appear to minimize a free energy functional of the states of its blanket, which is the same quantity optimized in Bayesian inference. So it will appear to model and act on its world to preserve its functional and structural integrity, leading to homoeostasis and a simple form of autopoiesis. Friston is not claiming that a cell holds beliefs. He is claiming that a cell with a membrane behaves exactly as though it does, and that the behavior needs no further explanation than the boundary itself.

OBJECTIONS

The strongest objection is the dark room, and it has an answer

The best known objection to the free energy principle is the dark room, and Friston answered it in print. If a creature only wants to minimize surprise, why does it not find a dark corner and stay there forever? Animals plainly do not. Play and exploration are core features of many life forms. Friston, Christopher Thornton and Andy Clark took the question head on in Free-energy minimization and the dark-room problem, in Frontiers in Psychology volume 3, article 130, in 2012 (Friston and Clark 2012). They wrote it as a dialogue between a philosopher and a physicist. Their answer is that surprise is defined relative to a model and never in the abstract. A dark room affords low surprise only to an agent that evolution or development has shaped to predict and inhabit it. For a human being, prolonged darkness and stillness are among the most surprising states available.

Two further points, stated plainly. The free energy principle is a principle rather than a hypothesis, and critics are right that the principle itself resists direct falsification. What can be tested are the process theories built on it: predictive coding, active inference, precision weighting, and the motor anatomy that Adams and Shipp described. Friston also did not invent predictive coding and has never said he did, crediting David Mumford in 1992 (Mumford 1992), Rajesh Rao and Dana Ballard in 1999 (Rao and Ballard 1999), and Helmholtz long before either. His contribution was to put perception, action, attention and learning under a single quantity. For the story of tone, that single quantity is the gift. Friston gave the healing arts a unit for something they had only ever described, because after him the resting tension a body holds can be read as a prediction it has not yet revised.

What the record shows

Karl Friston's prediction engine in seven dated findings

  • 2003. Dynamic causal modeling, built with Lee Harrison and Will Penny in NeuroImage volume 19, pages 1273 to 1302, estimates how one brain region drives another (Friston 2003). Friston reached theory through measurement, and his theory keeps asking what would be measured.
  • 2006. A free energy principle for the brain, with James Kilner and Lee Harrison, Journal of Physiology-Paris volume 100, pages 70 to 87, states the principle in one line: systems change to decrease their free energy (Friston 2006).
  • 2010. The Nature Reviews Neuroscience review, volume 11, defines free energy as an information theoretic quantity, as opposed to a thermodynamic quantity, that bounds the surprise on sampling data given a model. The brain is minimizing mismatch before it is minimizing anything else (Friston 2010).
  • 2013. Adams, Shipp and Friston, Brain Structure and Function volume 218, pages 611 to 643: descending motor signals are proprioceptive predictions, and the spinal reflex arc is the machinery that resolves the error (Adams and Shipp 2013).
  • 2013. Life as we know it, Journal of the Royal Society Interface volume 10: anything wrapped in a Markov blanket behaves as though it models its world to preserve its own integrity. That is homoeostasis stated as statistics (Friston 2013).
  • 2016. With Anil Seth, Philosophical Transactions of the Royal Society B volume 371: autonomic reflexes are enslaved by descending predictions, so heart rate, gut motility and the depth of a breath are steered by expectation (Seth and Friston 2016).
  • Thousands of ATP molecules. That is the neuroenergetics price of one bit crossing a chemical synapse, in a brain that is a fiftieth of the body's mass and a fifth of its energy budget. A standing prediction error is a standing metabolic bill (Laughlin 1998).

Questions people ask

What is the free energy principle in plain language?

It says that anything which keeps existing must act to keep its sensations close to what it expects. Friston, Kilner and Harrison put it in one line in 2006: systems change to decrease their free energy. Free energy is a computable quantity that always sits above surprise, so a system that lowers the bound lowers surprise as well, without ever needing to measure surprise directly.

Is this the same free energy as in thermodynamics?

No. In his 2010 review in Nature Reviews Neuroscience, Friston states that it is an information theoretic quantity, as opposed to a thermodynamic quantity. It measures the mismatch between a model and the evidence, not heat and not metabolic fuel. The two ideas share mathematics inherited from statistical physics, and they share a name. They do not share a meaning, and confusing them is the most common error made about his work.

Did Karl Friston invent predictive coding?

No, and he does not claim to have. He credits David Mumford in 1992 and Rajesh Rao and Dana Ballard in Nature Neuroscience in 1999, and he traces perception as inference back to Hermann von Helmholtz in the nineteenth century. What Friston added was the argument that perception, action, attention and learning are all descents on one quantity, together with the anatomy needed to test it.

What does the free energy principle mean for hands-on care?

It gives resting tension a job description. Under active inference, descending signals are proprioceptive predictions and the classical reflex arc resolves the error, so a change in muscle tone is a change in what the nervous system expects. That is why clean sensory evidence, delivered well, can alter tension without anyone forcing it. It is also why this library treats tone as the variable that everything else is measured against.

What did Karl Friston give the Unified Model of Tone?

A currency. Friston established that the nervous system runs on prediction and that prediction error is costly, priced in the neuroenergetics literature at thousands of ATP molecules for one bit crossing a synapse. The model adds its own step: what the body predicts, beneath everything else, is itself, and the variable it predicts is tone. Friston did not propose tone. The model takes his accounting and names what the account is kept on.