Résumé de lecture, chapitre par chapitre — rapporte ce que le livre dit.
THIS DOCUMENT IS NOT A CITATION SOURCE — a reading summary, ungraded.
C'est un résumé de lecture, non coté, qui rapporte ce que le livre dit. Il ne porte aucun poids probant et ne fonde aucune affirmation. Toute citation reste ancrée sur le .txt du livre, jamais sur ce résumé.
Gershman (2019) - What Does the Free Energy Principle Tell Us About the Brain| Title | What does the free energy principle tell us about the brain? |
| Author | Samuel J. Gershman (Department of Psychology & Center for Brain Science, Harvard) |
| Type | Theoretical article ("ORIGINAL ARTICLE — Theory") |
| Funding | Alfred P. Sloan Foundation (research fellowship — attested in the file) |
| Corpus source | library/39_Predictive_processing/_articles/Gershman (2019) - What Does the Free Energy Principle Tell Us About the Brain.txt |
| Text sha256 (16) | cd3c4d48283cb6a7 |
| Length | 4,712 words |
| Structure | Abstract · 7 numbered sections · Conclusions · References |
| Table of contents | ✅ READ in the file (all 7 headers are in the text) |
Method. The automated reading pipeline declared a segmentation fallback ("NO CHAPTERS — insufficient detection"): the script could not isolate the sections. They do exist in the text, and were re-read by hand — numbered headers 1 INTRODUCTION … 8 CONCLUSIONS. This summary follows that real structure, not the fallback. The 6 quotations produced were re-grepped against the .txt: 6/6 confirmed.
Gershman seeks neither to validate nor to refute Friston's free energy principle (FEP): he deconstructs it to find out what it predicts that is distinctive. His methodological thesis opens the article: the FEP "does not have a fixed set of distinctive claims; it makes different claims under different sets of assumptions." This is not objectionable provided one can verify those assumptions in each application and thus render the theoretical claims falsifiable.
The thread is a series of conditional equivalences. Without restriction of the variational family, minimizing free energy is equivalent to exact Bayesian inference — so, for passive observations, the FEP's predictions are indistinguishable from those of the Bayesian brain hypothesis. Predictive coding, often presented as the empirical core of the FEP, is not a generic consequence of it: it emerges only under certain restrictions (Laplace/Gaussian approximation) and a specific optimization scheme. In the active setting (the agent influences its observations), active inference is equivalent to an information-gain policy when the posterior is exact and observations deterministic; otherwise it induces risk aversion. Finally, planning as inference (setting utility equal to log prior probability) reads as a notational variant of Bayesian decision theory, so long as utilities and prior preferences coincide.
The epistemic status that emerges is the point the foyer's boundary retains: a principle cannot be refuted the way a theory can. Gershman does not say the FEP is false; he says that, as a principle, it does not by itself state a set of falsifiable empirical claims — that is a status, not a verdict. The empirical credit of predictive coding therefore does not transfer to the FEP "in general."
The FEP "states, in a single sentence, that the brain seeks to minimize surprise" — "the most ambitious theory of the brain available today," which claims to subsume predictive coding, efficient coding, Bayesian inference, and optimal control theory. It is this very generality that is the problem: the underlying assumptions are malleable (generative models, approximations, neural implementations vary across applications). Gershman announces a systematic deconstruction, and lays down the usage boundary for the whole foyer: variability of claims is acceptable if one can render them falsifiable application by application. He treats fairly two adversarial objections — including the one that says distinguishing the claims of a unifying theory is futile.
"This is not necessarily a bad thing, provided we can verify these assumptions in any particular application and thus render the theoretical assumptions falsifiable." (§1)
"Another qualm with this approach is based on the argument that FEP is not a theory at all, in the sense that a theory constitutes a set of falsifiable claims about empirical phenomena." (§1 — reported objection)
The prerequisite is laid out "in terms more familiar to neuroscientists," and is equivalent to the FEP under certain conditions. Here the brain infers hidden states of the world from noisy/ambiguous observations, via a prior and a likelihood. Gershman notes at once that the hypothesis, while widely supported, has empirical departures, sometimes rationalizable through approximate-inference algorithms. This is where the equivalence of the next section is prepared.
"For example, the hidden state might be the orientation of a line segment on the surface of an object, and the prior might be a distribution that favors cardinal over oblique orientations [5]." (§2)
The FEP's basic move: convert Bayesian inference into an optimization problem. Gershman establishes the identity log p(o) = D[q(s)‖p(s|o)] − F[q(s)], where the variational free energy F is **the negative of the evidence lower bound (the usual term in machine learning). The central consequence: without restriction of the family q, minimizing free energy is equivalent to exact inference — hence to the Bayesian brain hypothesis. From which the warning follows: if FEP = Bayes, its predictions can no longer be distinguished** from any other asymptotically correct algorithm.
"If FEP = Bayes, then we cannot distinguish its predictions from other asymptotically correct inference algorithms" (§3)
It is by restricting q that one gains distinctive predictions — and makes the problem tractable. Gershman reviews the mean-field approximation (the posterior factorizes, producing errors that are sometimes discernible in human behavior), the Gaussian approximation, and the Laplace approximation — the last of these having "intriguing implications for the neurobiological implementation." The message: testable predictions come not from the principle, but from the restrictions one adds to it.
"One challenge facing applications of the Gaussian approximation is that the free energy is not, in general, tractable (except in the case where the exact posterior is Gaussian)." (§4)
The pivotal section for the boundary. Predictive coding has a long history (signal processing, Barlow's efficient coding); Friston and colleagues showed how to derive it within free-energy minimization, map it onto microcircuits, and apply it to motor control. But Gershman insists: predictive coding is not a generic consequence of the FEP — it presupposes the Laplace approximation and a particular optimization scheme. Only under these assumptions does the FEP make predictions that go beyond the general Bayesian hypothesis and "have received ample empirical support."
"It is important to emphasize that predictive coding is not a generic consequence of FEP, or even of FEP with a specific approximation family." (§5)
The agent now acts (policy π) to influence its observations, minimizing expected free energy. Gershman demonstrates the equivalence: when the approximate posterior is exact and observations are deterministic functions of actions, minimizing expected free energy is equivalent to maximizing information gain — the policy already studied in standard Bayesian treatments of information acquisition. When observations are stochastic, active inference additionally induces a risk aversion absent from the pure information-gain policy.
"We showed that FEP is equivalent to Bayesian information gain only under the special case of an exact posterior and deterministic outcomes in the future." (§6)
The further conceptual step: setting the utility of an outcome equal to its log prior probability (u(o) = log p(o|π), the "prior preference"). Gershman first notes the strangeness (highly probable events can have low utility, such as being born into poverty), then defuses it: this is "potentially Bayesian decision theory in disguise" — a notational variant, as long as utilities and probabilities coincide (which free-energy theorists stipulate). The FEP makes a distinctive prediction only where they do not coincide.
"This also arises naturally in Bayesian decision theory applied to sequential decision problems and hence is not a distinctive prediction." (§8, restating the point from §7)
Gershman summarizes himself, in five points the boundary can cite without paraphrasing: (1) for passive observations and an unrestricted variational family, the FEP's predictions are indistinguishable from the Bayesian hypothesis; (2) predictive coding is not a generic consequence of the FEP; (3) in the active setting, active inference = information-gain policy under exact posterior + deterministic observations, and induces risk aversion otherwise; (4) utility = log probability ⇒ planning as inference, distinguishable from utility maximization only when the two do not coincide; (5) utilities = prior preferences ⇒ value on information gain, which "also arises naturally in Bayesian decision theory and hence is not a distinctive prediction." Finally he bounds his own scope: these messages do not exhaust the uses of the FEP (self-organization, niche construction) — the article focused on what is central to neuroscience.
"Predictive coding is not a generic consequence of FEP; it arises only under certain restrictions of the variational family and a specific choice of optimization scheme." (§8)
"When utilities are interpreted as prior preferences, FEP places value on information gain. This also arises naturally in Bayesian decision theory applied to sequential decision problems and hence is not a distinctive prediction." (§8)
It does not say the FEP is refuted. Gershman says a principle cannot be refuted the way a theory can — "In this sense, a principle cannot be falsified through the study of empirical phenomena." This is an epistemic status, not a verdict of falsity. Writing "the FEP is false/refuted" contradicts the source.
It does not say predictive coding is false. It says predictive coding does not follow generically from the FEP — so one cannot transfer to the FEP "in general" the empirical credit of predictive coding. The two propositions are distinct.
It provides no efficacy data. No figure of clinical, behavioral, or applied performance is produced here: this is an article of theoretical clarification. Drawing an empirical evidence level from it would be a category error.
It is not an outside skeptic. Gershman is internal to the field and favorable to the FEP's unifying utility ("make the elegant synthesis offered by FEP more accessible"). This strengthens the value of the boundary; it does not weaken it.
Segmentation: the automated pipeline declared fallback (script unable to isolate the sections). The 7 sections + Conclusions were re-read by hand in the .txt (numbered headers present). The summary follows the real structure, flagged as hand-read and not reconstructed.
Anti-fabrication control: the 6 distinct quotations produced (excluding the take-home-message restatements) were re-grepped against the .txt. 6/6 confirmed.
Declared limits: the equations and figures are known only through the text that comments on them. The references (44 entries) were not treated as content.
Role in the foyer: source of kb/39_Predictive_processing/65 (evidence boundary). This summary re-grades nothing; it documents the source of the boundary.