Special thanks to Kevin Baker, Leif Weatherby, and Ben Recht for multiple exchanges about the writings of Ivor A. Richards, his idea of “feedforward,” and related ideas that inspired this post.
Claude Shannon wrote in a 1959 paper: “[There is] a duality between past and future and the notions of control and knowledge. Thus we may have knowledge of the past but cannot control it; we may control the future but have no knowledge of it.”1 More than a hundred years before Shannon, Søren Kierkegaard expressed essentially the same thought in his journal:
It is perfectly true, as philosophers say, that life must be understood backwards. But they forget the other proposition, that it must be lived forwards. And if one thinks over that proposition it becomes more and more evident that life can never really be understood in time simply because at no particular moment can I find the necessary resting-place from which to understand it—backwards.
To a control theorist, this passage illustrates the complementarity of feedback and feedforward. As we will see next, this complementarity has fascinating intellectual and political history, touching on ideas in control engineering, theories of decision-making, cybernetics, and even some Stalin-era academic politics in the Soviet Union. We will begin with the latter.
In 1939, an article with the unassuming title “Theory and design of automatic regulators” was published in the inaugural issue of Avtomatika i Telemekhanika (Automation and Remote Control), the first Soviet journal devoted to the emerging field of control engineering. The author of the article was Georgii V. Shchipanov, an aviation engineer who had recently joined the newly created Institut Avtomatiki i Telemekhaniki (Institute of Automation and Remote Control) in Moscow as a member of the research staff. The first three sentences of the abstract read: “A problem of automatic regulation is formulated in this article. The role of the regulator attached to a machine or to a process is established. The main problem for the regulator is to compensate for the influence of perturbing forces on the parameter being regulated.”
This sounds innocuous, even somewhat dry and boring by today’s standards, and it is certainly something any control engineer would take for granted. The technical content of Shchipanov’s paper is also fairly standard (correcting for the fact that he did not have a lot of formal mathematical training): The effect of control signals and disturbances on the system was modeled using a system of linear differential equations with constant coefficients, and the problem of compensation was to design (or synthesize) the overall system in such a way that a specifically designated output variable remained insensitive (or invariant) to arbitrary disturbances. Despite Shchipanov’s lack of formal mathematical training, there were several technical and conceptual innovations in his work. The first one was the idea that control systems could be synthesized or designed by reasoning from the desired general behavior to a particular system realization by interconnection of various standard building blocks or components. The second one was that one could reject disturbances by using them to directly actuate the control input (the French engineer and mathematician Jean-Victor Poncelet had a similar idea in 1829, but his attempt to implement it to stabilize the angular velocity of a steam engine proved unsuccessful). This was in contrast to the usual, indirect approach based on output error feedback, where the controller is actuated by the error signal, obtained by comparing the output to the desired reference. The third one was Shchipanov’s intuition that, for any system to have such an invariance property for a designated output, the overall design would have to incorporate multiple feedback loops.
These ideas were ahead of their time. However, the reception by the Soviet scientific and technical establishment was hostile. A critical review by another control engineer, L. N. Mikhailov, appeared in 1939 in Vestnik Inzhenerov i Tehknikov (Bulletin of Engineers and Technicians), followed by scathing reports in Izvestiya Akademii Nauk SSSR (Notices of the Academy of Sciences of the USSR) in 1940 by two prominent mathematicians, Sergei L. Sobolev (of the “Sobolev space” fame) and Felix R. Gantmacher. Shchipanov’s response was also published, with a follow-up by Sobolev. Neither Sobolev nor Gantmacher pulled their punches. Sobolev, reviewing Shchipanov’s earlier monographs on the design of aviation equipment and gyroscopes, called them “scientifically and mathematically illiterate.” Gantmacher, who is well-known to control engineers for his excellent 1953 text Theory of Matrices, was similarly unsparing and wrote that Shchipanov’s article was “erroneous from start to finish and based on a completely fantastical idea of the author on an ‘ideal universal regulator’ (something akin to a perpetuum mobile).” Moreover, both Sobolev and Gantmacher attacked not only Shchipanov’s work but also other scientific publications that were coming out of the Institute of Automation and Remote Control and took its director, Victor S. Kulebakin, to task for allowing such low-quality work to take place under his leadership. (Shchipanov’s polemical disposition didn’t help matters much; for example, at the end of his original publication he confidently stated that, because of the obvious superiority of his invariance approach, all other designs of automatic regulators must be deemed unsuitable.)
The debate over Shchipanov’s work, and over the work of the institute in general, quickly turned political. In 1941, a front page article in Pravda about the state of Soviet science and industry called out a “pseudoscientific, absurd theory in the field of automatic control.” This accusation was leveled on the basis of a letter sent to Pravda by a group of scientists concerned about the attempts at the institute to “develop, by means of mathematical speculations, a fantastical ‘universal and ideal regulator’.” The editorial called on the Soviet Academy of Sciences to investigate the matter and to impose appropriate corrective measures. In the same year, an article with the title “Pseudoscientific works at the Institute of Automation and Remote Control,” published in the influential Bolshevik journal, condemned several publications in Automation and Remote Control, including the papers by Shchipanov, Kulebakin, and the eminent mathematician Nikolai N. Luzin, himself the target of an earlier smear campaign in 1936 that resulted in his dismissal from the Steklov Mathematical Institute. Having narrowly escaped imprisonment, Luzin was hired as a researcher at the Institute of Automation and Remote Control, and his article on the theory of matrix differential equations was, in part, an attempt to defend Shchipanov’s work from the critics even though he did not cite it explicitly. The article in Bolshevik is also remarkable for its list of authors—in addition to Sobolev and Gantmacher, it included other prominent researchers in mathematics and control, such as A. Vinter, C. Khristianovich, and I. Voznesenskii. The report of the specially appointed committee of the Soviet Academy of Sciences ordered a complete stop to any further work on Shchipanov’s invariance principle. Only the beginning of World War II helped Shchipanov and Kulebakin avoid imprisonment or worse; Kulebakin was removed from his position as institute director and Shchipanov was transferred to the Institute of Aviation.
The causes underlying this “anti-Shchipanov campaign” are not entirely clear. Historian of technology Christopher Bissell speculates about two possible explanations.2 One has to do with a covert conflict between rival factions in the Soviet control community, the Moscow one associated with Kulebakin and the Leningrad one that coalesced at the Central Boiler and Turbine Institute around Ivan N. Voznesenskii (one of the biggest critics of Shchipanov and a co-author of the 1941 Bolshevik article). Something like this also played out in 1936 in the Luzin affair, where a group of ambitious young mathematicians thought that they could co-opt the formidable machinery of the Soviet state to get rid of an older distinguished scientist who, they felt, had a stifling influence on Soviet mathematics. The other one, according to Bissell, may have been philosophical:
It may also be that there are echoes, in the vehement attack on the academic work, of the well-known anti-idealist movement in the Soviet Union, which over the years had criticized much “bourgeois” science—including relativity and quantum theory—for its absence of a philosophical basis in Marxist dialectical materialism. (The repeated use of terms, such as “ideal and universal” in the criticism renders this interpretation tempting, and “idealism” was certainly part of the criticism leveled at Luzin just a few years earlier.) Furthermore, one might even be inclined to view the Shchipanov Affair at least partly a contest between a traditional paradigm of scientific analysis, rooted in physics and mathematics (some of the severest critics came from this background), and an emerging design culture and systems approach of control engineering and related subject areas, in which highly idealized and even noncausal models may usefully be employed.
This paradoxical aspect of Shchipanov’s ideas may have contributed to the negative reception of his work by the peers. The paradox consists in the following: In a feedback control system, the very possibility of regulation is predicated on the ability to measure the deviation of the output from the reference. If there is no deviation (as mandated by Shchipanov’s compensation condition), then the controller does not receive any information and thus cannot function as intended. This intuition, however, is misleading because Shchipanov’s designs were of feedforward type, and the apparent violation of causality could be explained either by the presence of a sensor that actually measures the disturbance signal or by some sort of a direct physical coupling between the disturbance and the control input.
One particularly important way in which Shchipanov was ahead of his time was his view of control systems as designed behaviors. In his 1939 paper, he wrote that
automatic control refers to a complex of measures pertaining to certain dynamical properties and behaviors of machines (or processes) and artificial alteration of these properties. … Dynamical properties … of machines and processes are characterized by differential equations and depend primarily on the structure of these equations. …
With the help of a regulator—a new system attached to the existing one—one can address the question of altering the dynamical properties of machines and processes. … Thus, the problem of automatic control is about how many differential equations should be added to the equations for the given machine; how these equations must be related to one another; and to what the terms of these additional equations, i.e., the forces entering them, correspond constructively.
This viewpoint, which is commonplace now, was too abstract for the state of control theory at the time, even dangerously so in the face of rigid Stalinist dogma that ruled over everything in Shchipanov’s social and professional milieu.3 Shchipanov provided several concrete examples of how his ideas could be implemented in mechanical systems based on his experience as a designer of aviation equipment; however, these examples were not completely convincing because of particular limitations of mechanical devices of that era. In hindsight, his ideas were much better suited for realization in electronic control systems.
It was primarily the feedforward architecture of his “ideal compensators” that made many of his critics uneasy; they saw in it some mysterious violation of causality. This is, actually, a nontrivial point. Despite the unfortunate turn of phrase “circular causality” often used in connection with feedback, it is easy to see why causes precede effects in a feedback system: At each time instant, the current control input is determined by the measured deviation of the controlled variable from the setpoint. Shchipanov’s approach does not rely on measurements of the output; what information does the controller use then?
This question, as it turns out, has a precise answer. According to a very general formulation by Hans Witsenhausen, the problem of control system design amounts to the choice of a control law, i.e., the function that specifies the control input to be applied at each time instant based on the currently available information subject to given constraints.4 This information may include the past and current disturbances, past and current outputs, and past control inputs. A straightforward inductive argument shows that any system variable can be represented by a function of initial conditions, disturbances, and control inputs only. We then say, following Witsenhausen, that a feedforward control architecture is one where the data available to the controller at each time instant depend only on the disturbances, but not on the control inputs that were applied in the past. Such feedforward architectures are often used as building blocks in more complicated arrangements, where their outputs are used as prediction signals in some surrounding feedback loops. Think about a decision-maker operating in a complex environment, such as a small investor in the stock market, whose actions affect only his beliefs about the environment’s expected future behavior, but not the environment itself.
In fact, such an economic interpretation of feedforward in terms of maximizing expected utility was given by Václav Beneš, a Princeton-educated logician and philosopher who had a second career as an applied mathematician at Bell Labs.5 Analyzing a special case “in which decisions affect only the performance criterion, and not the trajectory of a random dynamical system,” he acknowledged a referee for suggesting that
the situation of the small investor in the stock market can be represented approximately by the kind of setup considered here. In this case the vector of prices of stocks on the market forms the stochastic process x(.) in question; over this the small investor has no control. The control variables are the amounts of money the investor has invested in each stock. The performance index is the sum over all the stocks of the integral of the product of the amount invested (in the stock in question) times the rate at which the price is decreasing. The construction to be given would show that an optimal investment policy exists, and that it is obtained by choosing the control which minimizes the conditional expectation of the performance rate with respect to the investor’s information. This corresponds, not surprisingly, to placing money in the stocks with greatest expected growth based on the facts known to the investor, which is what investors generally try to do.
The key point here is that the state of the market is invariant relative to the actions of such a small investor. The only thing the investor is in control of is his own beliefs about the state of the market given his preferences. There is a curious dialectic at play here because the market can be seen as a feedback controller acting on the investor, while the investor acts as a feedforward controller converting his experience into anticipation of the market’s future behavior. The appellation “small” used by Beneš is not accidental here—the talk of maximizing expected utility immediately brings to mind Jimmie Savage’s formalization of Bayesian rationality in terms of acts that map states of the decision-maker’s world to consequences of relevance to the decision-maker. Savage encloses the entire process in what he calls a “small world,” i.e., one in which it is possible to look before leaping.
Savage makes the distinction between “small worlds,” where all contingencies can be accounted for and modeled in advance, and “large worlds,” where some genuine surprises can occur, in Chapter 2 of his 1954 book The Foundations of Statistics. He then revisits this concept in Chapter 5 on utility:
Allusion was made in the penultimate paragraph of §2.5 to the practical necessity of confining attention to, or isolating, relatively simple situations in almost all applications of the theory of decision developed in this book. As was mentioned there, I find it difficult to say with any completeness how such isolated situations are actually arrived at and justified. …
Making an extreme idealisation, which has in principle guided the whole argument of this book thus far, a person has only one decision to make in his whole life. He must, namely, decide how to live, and this he might in principle do once and for all. Though many, like myself, have found the concept of overall decision stimulating, it is certainly highly unrealistic and in many contexts unwieldy. Any claim to realism made by this book—or indeed by almost any theory of personal decision of which I know—is predicated on the idea that some of the individual decision situations into which actual people tend to subdivide the single grand decision do recapitulate in microcosm the mechanism of the idealized grand decision. One application of the theory of utility to overall decision has, however, been attempted by Milton Friedman.
With this rhetorical flourish, Savage is articulating the full force of the Kierkegaardian double bind: the consequences of our acts must be understood backwards, but the acts can only be decided forwards. He proposes to embed the small-world consequences as acts in a grand world. A grand world is not the same as a large world where Knightian uncertainty reigns supreme, it is simply a massive random environment that cannot be moved by isolated acts of isolated small-world actors. A small decision-maker can only act on the grand world in a feedforward manner—think of Beneš’ small investor. As Savage puts it, “a small-world consequence is a grand-world act.” The consequences, for a small investor, are history-derived plans regarding future investments. They become acts in a grand world (viz., the market) with grand-world consequences that flow back as feedback signals to the small investor. With a bit of hand-waving, we can even view Shchipanov’s compensator as a small-world actor embedded in the grand world of a more complex system consisting of multiple feedback loops.
Neither Savage nor Beneš used the term “feedforward,” although Savage definitely heard it used in 1951. He was one of the regular participants of the Macy Conferences on cybernetics that were held at the Beekman Hotel at 575 Park Avenue in New York. The 1951 conference featured a talk by the literary critic and theorist of rhetoric Ivor A. Richards. In this talk, entitled “Communication between men: the meaning of language,” Richards introduced the term “feedforward” as a complement, or even a prerequisite, to the usual “circular and feedback mechanisms” studied in cybernetics.6 Instead of giving a precise definition of feedforward, Richards motivated it through an ingenious use of circularity and feedback inherent in language (the use of the inverted quotes »…« was his):
Perhaps this thing on which I want to put the spotlight will be considered to be included in some ingenious way under the word »feedback.« But what I am going to stress stands in an obvious and superficial opposition to »feedback,« and it will, in certain frames of thought, be given nearly, if not quite so much, importance, and sometimes more importance than feedback itself in certain connections. It is certainly as circular. You have no doubt fed forward enough to see that what I am going to talk about from now on is feedforward. I am going to try to suggest its importance in describing how language works and, above all, in determining how languages may best be learned.
Feedforward, in Richards’ telling, has to do with arranging things in anticipation of the general shape in which one’s immediate future will unfold, so that appropriate feedback control mechanisms could be set in place, waiting to be actuated by the appropriate error signals if and when they come. This entails anticipating the effect, if not the exact realization, of disturbances before they happen. One of the examples Richards gave had to do with teaching children to read. He argued that good pedagogical practice would involve recognizing various heuristics and biases inherent in visual perception and preventing them from acting as disturbances to the process of learning in its initial stages. This would mean, for example, that the teacher should not introduce the letters “p”, “b”, “d”, and “q” at the same time since they are images of one another under rotations and reflections, and the visual system’s tendency to factor out such symmetries would cause unnecessary confusion:
The child couldn’t live life unless he saw a knife, say, as a knife, no matter which way up it was. It is bad technique to make a sudden transformation to script, in which it is all-important whether the u is upside down – or is it the n that is upside down? We penalize the bright child by setting a whole set of bogus traps for him in the script we begin to teach him. They don’t belong to the subject. They just betray him through his biological smartness.
The importance of invariance in pattern recognition has been pointed out in many places — e.g., in Wiener’s Cybernetics, in Pitts and McCulloch’s 1947 paper “How we know universals: the perception of auditory and visual forms,” and in Marvin Minsky’s 1961 paper “Steps toward artificial intelligence.” Biological visual pattern recognition systems have evolved to identify objects reliably despite uniform changes in size, position, and orientation. And yet, this obviously useful evolutionary adaptation acts as a disturbance when one is learning to read. Hence, Richards’ suggestion of staggering the introduction of such letters to counteract this disturbance effect can be seen as a feedforward control strategy that factors out invariances in low-level visual perception in order to induce higher-order invariances in letter and word recognition. (A curious side effect of this is the loss of robustness in such precisely crafted control systems: Most people would have a hard time trying to read the mirror image of a printed sentence.)
Richards recapitulates these ideas in a 1968 essay called “The secret of ‘feedforward,’” where he writes that feedforward
is the reciprocal, the necessary condition of what the cybernetics and automation people call “feedback.”
Whatever we may be doing, some sort of preparation for, some design arrangement for one sort of outcome rather than another is part of our activity. This may be conscious, as an expectancy—or unconscious, as a mere assumption. If we are walking downstairs, a readiness in the advanced leg (but indeed in our whole body) to meet something solid under its toe is needed if we are to continue. Usually, on the stairs, this feedforward is fulfilled. There is confirmatory feedback at the end of each step cycle—the foot finds the expected, the presupposed footing. Compare pitch-dark and broad daylight as to the degree of awareness we may have of our feedforward. If the feedback does not come, if it is falsifying and not verifying, we have to do something else and rather quickly. The point is that feedforward is a needed prescription or plan for a feedback, to which the actual feedback may or may not conform.
Evidently, feedforward is a product of former experience: a selective reflection of what has been relevant in similar activity in our past.
He goes on:
Feedforward is, as I see it, highly various. At one end of the scale, it can be a highly articulate examinable process, the sort of thing known as scientific hypothesis waiting to be okayed or destroyed by evidence. At the other, it may be hardly cognized or embodied at all, even in the vaguest schematic image. It can be no more than a readiness to be surprised or disturbed by one kind of event rather than by another: green-lighted or red-lighted by it. … [B]illions of hierarchically systematic cycles, through which we live and move and have our being, are guided in all they do for one another by concord or discord in their feedforward-feedback. And it is perhaps a reasonable suggestion that much in what we call ourselves and admit to be “us” includes these billions of concordant cycles.
This has a distinctly Hayekian ring to it—recall Hayek’s idea of the primacy of the abstract, namely that “the dispositions for a kind of action possessing certain properties comes first and the particular action is determined by the superimposition of many such dispositions.” Compare this with what Richards wrote in Poetries and Sciences (a 1970 re-issue with commentary of his 1926 book Science and Poetry):
We should picture the mind as a system of very delicately poised balances, a system which so long as we are in health is constantly growing. Every situation we come into disturbs some of these balances to some degree. The ways in which they swing back to a new equipoise are the impulses with which we respond to the situation. And the chief balances in the system are our chief interests.
We have now fed forward all the way to the end of the essay. In proper Soviet fashion, Shchipanov was “rehabilitated” in 1959, six years after his death at the age of 50 of a chronic respiratory illness. Researchers were now officially permitted to cite his work and to build on it. There were several interesting follow-up studies of his invariance principle at the Institute of Automation and Remote Control. For example, Maxim Braverman and Lev Rozonoer investigated structural stability of Shchipanov-type systems. A given property of a dynamical system is structurally stable if it is insensitive to small perturbations of system parameters. Rozonoer showed, in particular, that Shchipanov-type “invariant systems” are not structurally stable in the presence of small delays in the feedforward path from the disturbance to the control actuated by it. (Incidentally, Rozonoer was one of the authors of The Method of Potential Functions in the Theory of Machine Learning, a pioneering 1970 text that introduced kernel methods and advocated the use of stochastic gradient descent in machine learning. The other two authors were Emmanuil Braverman, Maxim Braverman’s father, and Mark Aizerman, one of the giants of Soviet control theory and a decorated World War II veteran who, as a young man, had enough courage and integrity to write a letter to Otto Schmidt, the Vice-President of the Academy of Sciences of the USSR, offering a rigorously reasoned yet impassioned defense of Shchipanov’s ideas.)
The importance of anticipation and the interplay between feedforward and feedback was also recognized early on by several Soviet psychologists and neurophysiologists, such as Nikolai Bernstein and Pyotr Anokhin. In the early 1950s their ideas were officially condemned as dangerous deviations from Pavlov’s thought (to be “rehabilitated” later, of course). Their work has been vindicated in modern neuroscience that seamlessly incorporates such control-theoretic concepts as the internal model principle, model predictive control, sliding mode control, and, indeed, Shchipanov-type disturbance-actuated control. In a way, Shchipanov in control and Bernstein and Anokhin in neurophysiology were arguing for the importance of small world models in a grand world, long before the term “world models” became a fashionable buzzword in AI. These researchers led their lives forwards, so that we can now use the benefit of feedback and appreciate their contributions from our vantage point.
Claude E. Shannon, "Coding theorems for a discrete source with a fidelity criterion," IRE International Convention Records, vol. 7, pp. 142--163, 1959.
Christopher C. Bissell, “Control engineering in the former USSR: Some ideological aspects of the early years,” IEEE Control Systems Magazine, vol. 19, pp. 117-116, 1999.
The delicious irony here is that the oft-repeated quote “Philosophers have only interpreted the world, in various ways; the point, however, is to change it” from Marx’s Theses on Feuerbach is control engineering in a nutshell.
Hans S. Witsenhausen, “Separation of estimation and control in discrete time systems,” Proceedings of the IEEE, vol. 59, no. 11, pp. 1557-1566, 1971.
Václav E. Beneš, “Existence of optimal strategies based on specified information, for a class of stochastic decision problems,” SIAM Journal of Control, vol. 8, no. 2, pp. 179-188, 1970.
To the best of my knowledge, Richards’ coinage was completely original, independent of the use of this term in engineering contexts.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.