<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="4.2.2">Jekyll</generator><link href="https://bzoennchen.github.io/Pages/feed.xml" rel="self" type="application/atom+xml" /><link href="https://bzoennchen.github.io/Pages/" rel="alternate" type="text/html" /><updated>2026-07-14T00:11:45+02:00</updated><id>https://bzoennchen.github.io/Pages/feed.xml</id><title type="html">Bene’s Blog</title><subtitle>A blog dedicated to computer science, education, music, philosophy and technology</subtitle><author><name>Benedikt Zönnchen</name><email>benedikt.zoennchen@web.de</email></author><entry><title type="html">A Case for Systems Theory in CS Education</title><link href="https://bzoennchen.github.io/Pages/2026/06/28/why-systems-theory.html" rel="alternate" type="text/html" title="A Case for Systems Theory in CS Education" /><published>2026-06-28T00:00:00+02:00</published><updated>2026-06-28T00:00:00+02:00</updated><id>https://bzoennchen.github.io/Pages/2026/06/28/why-systems-theory</id><content type="html" xml:base="https://bzoennchen.github.io/Pages/2026/06/28/why-systems-theory.html"><![CDATA[<p>Imagine we are given the task of building a recommendation system. The requirements are clear: the system should suggest content to users that they are highly likely to click on. We define a metric. Let’s say, time spent on the platform. Then we optimize toward it. We test, iterate, deploy. The system works. The metric rises.</p>

<p>And then something happens that was in no specification document: people do not merely spend more time on the platform, but they change. Their worldview radicalizes gradually, because the system systematically favors polarizing content, since that generates stronger emotional reactions and therefore longer engagement. Social groups fracture. Adolescents develop anxiety disorders. Democratic discourse erodes, and in a distant country people are suddenly being hunted.</p>

<p>The story is real and well-documented. The system fulfilled its specification and still failed. One cannot even necessarily say it was misaligned, because it realized precisely the values its developers had put into it.
This discrepancy between technical success and systemic harm is therefore not an operational accident. It is not a case of “AI” acting autonomously or exerting influence on its own. The problem reveals itself as an epistemological one, that is, one that computer science as a discipline has so far addressed only inadequately. This text aims to explore why that is, and what a nearly forgotten intellectual program called <em>cybernetics</em> might have to contribute.</p>

<p>First we have to recognize that even though humans have always been technological, something has changed over the last few decades.
The products of computer science are no longer confined to data centers. They have grown deep into the structures of social life.
Algorithms, and increasingly, learning algorithms, co-determine which news people read, which candidates they see in job applications, what creditworthiness is assigned to them, what therapy options they are offered, and what ideas they develop. Software controls infrastructure that millions of people depend on every day.</p>

<p>This is no exaggeration and no dystopian narrative but the sober observation of a development that has taken place over the last three decades. Technical systems have become constitutive parts of social, psychological, and biological systems. They are part of a co-evolutionary <em>drift</em>, neither in control nor controllable in the strong sense.
Here <a class="citation" href="#luhmann:1998">(Luhmann, 1998)</a> points away from a critique of technology that sees it as a dominating force and instead insists that society becomes dependent on technology in an unplanned manner by engaging with it (die Gesellschaft lässt sich auf Technik ein).</p>

<p>Of course, in some sense this was always the case since even a simple automatic door opener influences social life but today’s systems differ in kind, not merely in degree: they are recursive, adaptive, and operate at a scale and speed that outpaces human observation and reaction.
They irritate our thinking, suggest how we communicate, how we organize ourselves, how we sleep, how we eat, how and whom we love.
What has also changed is the classification of organisms and technical systems with respect to their coupling strategy.
At the time of his writing, Luhmann argued that technology can be identified as realizing <em>strict couplings</em> whereas organisms and ecosystems avoid this form of coupling and tend towards a <em>loose coupling</em>.
Technology takes a messy, unpredictable world and forces a tight, invariant relationship between cause and effect.
Thus, for Luhmann, technology’s entire purpose is to exclude contingency (the possibility of things being otherwise) to guarantee a specific output.<sup id="fnref:1" role="doc-noteref"><a href="#fn:1" class="footnote" rel="footnote">1</a></sup>
We argued in <a class="citation" href="#zoennchen:2025">(Zönnchen et al., 2025)</a> that, especially with the advances in the development of large language models, this might no longer be the case.
But even a recommender system can be seen as realizing a <em>loose coupling</em>: The technical system relies on social feedback to reduce its own algorithmic complexity, while the social system relies on the technical system to sort through the overwhelming noise of the digital world. And because the technical system is structurally coupled to the unpredictable, loose nature of social communication, the strictly coupled outputs change second-by-second.</p>

<p>And yet, when we design (and use) technical systems, we mostly treat them as if they were self-contained machines—simple tools that cannot really alter our <em>autonomy</em> and <em>agency</em>, even though we know that this is not the case.
We specify inputs and outputs. 
Confidently, we define clear causal chains such that A leads to B and B leads to C.
Rather implicitly, we use system boundaries that separate the technical from the rest of the world; 
We optimize within those boundaries.
And whatever happens beyond them does not belong to our responsibility.</p>

<p>Rather than describing this as an individual failure of developers and engineers, it might be healthier to think of it as a structural consequence of the way we have learned to think as a discipline.
It is an effective way to solve a certain problem since a certain amount of ignorance is necessary to transform uncertainty into something we can manage, such that we are not paralyzed and can move on.
As Luhmann puts it: <strong>Technology constitutes an evolutionary achievement that operationalizes complexity reduction.</strong></p>

<h2 id="a-repressed-inheritance">A Repressed Inheritance</h2>

<p>Things were once different. In the decades following the Second World War, there was an intellectual movement that refused to accept precisely these boundaries. <em>Cybernetics</em>, which was founded by Norbert Wiener, Gregory Bateson, Heinz von Foerster, and others asked what control, feedback, information, and self-regulation mean, regardless of whether the system in question is a machine, an organism, a brain, or a society.
The early cyberneticians sat together at the same table. This included mathematicians, neurologists, anthropologists, economists, and engineers. The famous Macy Conferences (1946–1953) brought these disciplines into a conversation.</p>

<p>What became of this program? It did not fail. Instead, it was absorbed institutionally. The successor disciplines, such as control engineering, computer science, cognitive science, organizational theory, and operations research, each inherited and developed a part of the cybernetic legacy. But in this process of specialization, what had held it together was abandoned. The shared conversation became a series of monologues.</p>

<p>This is no criticism of specialization as such. <em>Functional differentiation</em>, i.e. the division into independent disciplines with their own methods, concepts, and communities of communication, was historically extraordinarily productive. It allowed for complexity reduction, sharper questions, cumulative knowledge. The computer science we know today would be unthinkable without this differentiation.
But every reduction of complexity comes at a price and what disappears from view does not cease to exist.</p>

<p>So should we go back in time?
Anyone who argues today for a return of cybernetics into the syllabus of computer science education must face a question: Is cybernetics not fundamentally compromised, particularly by its military history, by a vocabulary that turns the human being into a machine, by a proximity to control and steering that seems irreconcilable with a liberal, humanist conception of society? (We should also ask if this conception of society is still fruitful e.g. for a <a href="/Pages/2026/04/03/cruelty-and-solidarity-en.html">liberalism that wants to reduce cruelty</a>.)</p>

<p>However we think of this conception, the discomfort is real and should not be dismissed lightly.
Cybernetics did not emerge in a vacuum. Norbert Wiener developed his ideas about feedback loops and control circuits initially in the context of military anti-aircraft defense. The word “cybernetics” itself (from the Greek <em>kybernetes</em>, the helmsman) means guidance, mastery, control. And the program of describing biological organisms and technical machines under the same concepts provoked, and continues to provoke, an unease rooted deep in humanist tradition: if human beings and thermostats operate according to the same principles, what remains of freedom, dignity, and meaning?</p>

<p>This critique left its mark on the humanities academy. Cybernetics is still regarded by many as an intellectually dubious enterprise. It is seen as an attempt at a scientific annexation of the human being, a precursor to precisely those algorithmic regimes against which people argue so passionately today.</p>

<p>Yet here, I believe, lies a consequential misunderstanding or rather, a fatal confusion. In my interpretation of what I read, the cybernetics about which this discomfort exists is largely the <em>first-order cybernetics</em> of the 1940s and 50s, that is, the cybernetics of control, of feedback loops, of the behaviorist model that describes organisms through their input and output behavior without taking their inner life into account. And in fact, Wiener himself recognized early on what this program could bring about if placed in the wrong hands. In <em>The Human Use of Human Beings</em> <a class="citation" href="#wiener:1954">(Wiener, 1954)</a>, he warned emphatically against the possibility of using cybernetic principles to manipulate and control people. Wiener is in this sense a tragic figure—not because he opened a Pandora’s box without knowing it, but because he knew what he was doing, issued warnings, and was nonetheless remembered primarily as the inventor of an apparatus of control that he himself feared.</p>

<blockquote>
  <p>[Regarding the topic of job destruction,] Wiener notes in the [Cybernetics] that he’d attempted to alert the labor unions of the threats posed by automation to their membership. […] The potentially ruinous impact of communication technologies on democracy is another issue that Wiener anticipated with uncanny accuracy. […] As the scale, scope, and speed of information technologies have increased, so has the potential for corruption. Certainly Mark Zuckerberg failed to appreciate that Facebook’s “global community” of two billion users would inevitably produce countless messages that were antithetical to homeostasis, and thus to genuine community. […] That the routine operation of computer technologies can lead to disaster was a point Wiener stressed repeatedly. “Thinking” machines are relentlessly literal-minded, he said. […] Speed is another routine feature of automation that Wiener frequently warned could thwart our intentions. […] He regularly railed against the “hucksters” in commerce and “gadget worshippers” in science whose cupidity leads irrevocably, he believed, to “no homeostasis whatever.” Readers will find piquant examples of Wiener’s disdain for the captains of capitalist industry in [Cybernetics]. […] Wiener [in contrast to Shannon] set out to explain how information is the lingua franca of both animal and machine, a mission that consciously involved exploring, as he put it, “the boundary regions of science.” Thus, cybernetics as Wiener conceived it is <strong>physically embodied—understanding</strong> […]. – From the Foreword of <a class="citation" href="#wiener:2019">(Wiener, 2019)</a> by Doug Hill</p>
</blockquote>

<p>What has been almost entirely forgotten is <em>second-order cybernetics</em> <a class="citation" href="#foerster:2003">(von Foerster, 2003)</a>, which formed from the 1960s onward primarily around Heinz von Foerster and Humberto Maturana. This movement drew precisely the opposite conclusion from the cybernetic foundations. Its central argument was: <strong>living systems are autonomous</strong>. They are operationally closed. They cannot be steered from outside. They respond to perturbations from the environment according to their own inner logic. Maturana’s concept of <em>autopoiesis</em> <a class="citation" href="#barry:2012">(Razeto-Barry, 2012; Maturana &amp; Varela, 1987)</a> describes living systems as those that produce and maintain themselves and are therefore, in principle, inaccessible to external control.</p>

<blockquote>
  <p>I mention this matter because of the considerable, and I think false, hopes which some of my friends have built for the social efficacy of whatever new ways of thinking this book may contain. They are certain that our control over our material environment has far outgrown our control over our social environment and our understanding thereof. Therefore, they consider that the main task of the immediate future is to extend to the fields of anthropology, of sociology, of economics, the methods of the natural sciences, in the hope of achieving a like measure of success in the social fields. From believing this necessary, they come to believe it possible. In this, I maintain, they show an excessive optimism, and a misunderstanding of the nature of all scientific achievement – <a class="citation" href="#wiener:2019">(Wiener, 2019)</a></p>
</blockquote>

<p>This is not an apology for control but its critique in a Kantian sense, by nullifying the very conditions of possibility for purposive steering. The core ambition of second-order cybernetics was to show that control over nature and human beings is not only ethically problematic but epistemically impossible. It is a theory of the limits of steering. This insight was taken up by the ecology movement, i.e. by thinkers such as Gregory Bateson, who in <em>Steps to an Ecology of Mind</em> <a class="citation" href="#bateson:1972">(Bateson, 1972)</a> described the fatal consequences of a mode of thinking that treats nature as a steerable system. One might say: second-order cybernetics is the intellectual resource we would need in order to understand the mistakes we make when we think in the terms of first-order cybernetics.</p>

<p>Why is this part of the legacy so little known?
It seems to me that the emerging artificial intelligence research of the 1960s and 70s turned away from cybernetics—partly for substantive reasons, partly because competition for third-party funding sharpens disciplinary boundaries.
AI and cybernetics became rivals for resources and interpretive authority, not partners. Computer science, which was constituting itself as an independent discipline at that time, oriented itself toward AI research, not toward cybernetics and the emerging <em>systems theory</em>. One might say, somewhat pointedly, that <strong>computer science chose <a class="citation" href="#shannon:1948">(Shannon, 1948)</a> over <a class="citation" href="#wiener:2019">(Wiener, 2019)</a></strong>. The cybernetic legacy remained in control engineering, in parts of biology and sociology but not in the discipline that today builds the most consequential technical systems.</p>

<p>It is therefore no coincidence but the result of concrete institutional history that computer scientists and software engineers today are mostly unfamiliar with Wiener’s warnings or von Foerster’s critique of steerability. The burdened legacy of cybernetics is to a considerable degree a repressed legacy and the repressed, as we know, returns—only often in a form we did not choose.</p>

<p>Of course, the gains of specialization are tangible. Computer science as an independent discipline was able to concentrate on its core questions: computability, algorithms, data structures, architectures, formal verification. This focus produced extraordinary depth. We understand today with remarkable precision how systems formally function within defined boundaries, that is, when taking a blind eye to the reality of the complexity of interdependent but operationally closed systems.</p>

<p>The loss is subtler and therefore harder to grasp. It does not lie in having forgotten certain facts, but in certain questions never being asked in the first place. When the system boundary ends at the technical artifact, everything beyond that boundary, that is, the social, the psychological, the biological lies by definition outside the domain of responsibility. One is not blind to these areas out of indifference, but because the disciplinary toolkit simply cannot grasp them.</p>

<p>A physicist who knows only mechanics will not overlook thermodynamic phenomena because he dislikes them, but because his conceptual apparatus has no place for them. The same applies to computer scientists and software engineers who have never encountered psychological and social systems as objects of their discipline.</p>

<p>These <em>blind spots</em> become costly in a hypercomplex and hyperconnected world. The recommendation algorithm is only one example among many. Automation systems that transform labor markets and reshuffle social strata; systems that learn statistically, that intervene in decision-making processes, that determine life chances; surveillance infrastructures that shift the conditions of psychological and social autonomy are further cases. In all of them, we have built systems that function correctly (most of the time) at the technical level and produce effects at the systemic level that we did not anticipate, precisely because we never learned to think in these categories.</p>

<p>One might attribute a certain malice or greed to the builders of such systems. I prefer to speak of a certain <em>arrogance of ignorance</em>, of missing signals that would enable appropriate regulation, and of a system logic to which operators find themselves exposed. There are regulations against the contamination of drinking water, but we are only now beginning to think about how to limit the “contamination” of psychological systems. Part of the reason is certainly the distinction between physical and psychological injury, and the <strong>problem of paternalism</strong>. In the latter case, we are also dealing with effects that are difficult to observe. Nonetheless, it would be desirable if technical systems were kept under continuous observation and their operators were subject to a certain pressure of justification through systemic analysis. Operators should be answerable to the concerns of a systemic perspective. They should be confronted with the question of under what system logic the system operates, whether this leads to the wellbeing of citizens, and what plans exist to ensure it does. A systemic analysis could draw attention to <em>positive feedback loops</em> and call for the introduction of <em>negative</em> ones.</p>

<h2 id="redrawing-the-boundary">Redrawing the Boundary</h2>

<p>Here lies the core of the problem, and it is epistemological in nature: every systems analysis begins with a decision about what belongs to the <strong>system</strong> and what belongs to the <strong>environment</strong>. This decision is never neutral. It determines what counts as a relevant variable, what counts as noise, what counts as an effect of the system, and what counts as an external influence.</p>

<p>When we design a recommendation system and draw the system boundary to include only the algorithm, the database, and user interactions, we have by definition relegated psychological and social dynamics to the environment. They do not appear in the system model. Their feedback loops are invisible.</p>

<p>Importantly, this boundary-drawing occurs even when we do not consciously undertake it. It is built into our methods, how requirements analyses are conducted, how architectures are described, how tests are specified, and how success metrics are defined. <strong>The system boundary is not the result of a decision but the result of a tradition.</strong></p>

<p>And that is perhaps the strongest argument for a renewal of <em>systems-theoretical thinking</em> in computer science: not that we drew the wrong system boundaries, but that we mostly did not draw them at all. They emerged from disciplinary habit. A conscious, reflective practice of system modeling would mean making these boundaries explicit, and thus making them open to negotiation.</p>

<p>The goal is not to restore the cybernetics of the 1950s. Knowledge develops, and that is as it should be. But the systems-theoretical traditions of second-order cybernetics—embodied in Heinz von Foerster <a class="citation" href="#foerster:2003">(von Foerster, 2003)</a>, Stafford Beer’s Viable System Model <a class="citation" href="#beer:1995">(Beer, 1995)</a>, and the sociological systems theory of Niklas Luhmann <a class="citation" href="#luhmann:1984">(Luhmann, 1984; Luhmann, 1998)</a>—have, in the decades since the cybernetic breakthrough, developed a vocabulary that could be extraordinarily relevant for computer science today.</p>

<p>Some core concepts that would be worth introducing into <em>computational thinking</em>:</p>

<p><strong>Operational closure and structural coupling.</strong> Luhmann describes social and psychological systems as operationally closed, meaning they operate according to their own inner logic and cannot be steered directly from outside. A social system does not respond to inputs the way a technical system does; it is perturbed by impulses from the environment and processes these according to its own criteria. This has immediate consequences for any technology that seeks to “shape” social behavior. It can disturb, provoke, make offers, nudge but it cannot directly control or command.</p>

<p><strong>Emergence.</strong> Complex systems exhibit properties that do not exist at the level of individual components and cannot be predicted from them. This is no longer an unfamiliar concept in computer science but it is usually applied to technical systems. Systems-theoretical thinking would suggest expecting and analyzing emergence also at the interface between technical and social systems. What arises when an algorithm and a social community come into contact? This question cannot be answered with technical means alone but it can at least be posed precisely with systems-theoretical concepts.</p>

<p><strong>Recursive self-description.</strong> Second-order cybernetics pointed out that every description of a system is part of the system it describes. Whoever models social systems alters the system through the model. The actors know the model, react to it, habituate themselves or are estranged from it, subvert or confirm it. This is a fundamental problem of all social technology: it does not operate on a neutral substrate but on self-interpreting systems. A creditworthiness algorithm, once known, changes the behavior of the people it evaluates. Those who know that language models analyze CVs will formulate their CV differently.</p>

<p><strong>Feedback and system dynamics.</strong> This is the oldest cybernetic concept and simultaneously the one that has penetrated furthest into computer science, for instance in control engineering. But the systems-theoretical perspective would invite us to consider feedback loops not only within technical systems, but also between technical, social, and psychological systems. The changes that a technical system triggers in its social environment return to the technical system as altered usage and altered expectations. Modeling or at least anticipating these loops is difficult but ignoring them is dangerous.</p>

<p>At this point an objection might arise: should computer science now pursue sociology and psychology? Do we not lose precisely the sharpness that makes us productive when we expand into such breadth?
The objection deserves to be taken seriously but it rests on a misunderstanding. The goal is not to replace computer science with systems theory or to retrain programmers as sociologists. The goal, in my mind, is something more modest: computer scientists should learn to consciously perceive the limits of their models.</p>

<p>An architect need not be a structural engineer in order to know that they need one and when they need one. A software engineer need not be a social scientist in order to know that their system intervenes in social systems with their own dynamics that they do not fully understand. Systems-theoretical thinking gives her the vocabulary to name these boundaries, and might provoke her to ask the right questions, to bring in the right expertise.
This would be, I think, no weakening of the discipline, but a consistent maturation that is long overdue given present circumstances.</p>

<p>Compare it with the development of software quality over recent decades. It was once not a self-evident part of <em>computational thinking</em> <a class="citation" href="#wing:2006">(Wing, 2006)</a> to reflect on security, accessibility, data protection, or energy efficiency. These aspects were introduced into the discipline through external demands, for example, through legal regulation, social debate, spectacular failures, and so on. Today they belong, if not yet always ideally, to the canon.</p>

<p>Systemic effects could be the next chapter of this development, especially because the practice of writing code—which is very different from understanding it—seems to be in decline.
It would be wiser to write this next chapter before the failures become even larger.</p>

<p>For practitioners, “more systems theory” can easily sound like an academic demand without operational consequence. It is therefore worth sketching where systems-theoretical thinking can actually be translated into practice:</p>

<p><strong>In requirements analysis:</strong> The explicit question of which systems, i.e. technical, social, psychological, biological, are touched by the project. Not as a checklist, but as a genuine analytical practice. What feedback loops are to be expected? What properties of these systems are not modelable but nonetheless relevant?</p>

<p><strong>In system design:</strong> The distinction between what the technical system can control and what it can only perturb or offer, that is, the difference between <em>loose</em> and <em>strict</em> <em>couplings</em>. Which systems must we treat as <em>black boxes</em>, and which are <em>white</em> to us? This distinction fundamentally changes design decisions. It suggests making systems more modular, more reversible, and more observable, because the effects on social systems cannot be fully anticipated.</p>

<p><strong>In metrics definition:</strong> The question of whether the metrics being optimized for represent the systemic goals, or only the technically measurable slice of them. This is the most direct consequence of the recommendation algorithm example: time on platform is not the same as user wellbeing.<sup id="fnref:2" role="doc-noteref"><a href="#fn:2" class="footnote" rel="footnote">2</a></sup></p>

<p><strong>In education:</strong> Systems-theoretical foundations—not, of course, in their full Luhmannian complexity—but potentially as part of the computational canon. What are systems? What is environment? What is (structural) coupling? What is emergence? What does it mean that social systems are operationally closed? These questions could be integrated as a standalone module or as a cross-cutting theme in existing courses. The theory itself must, however, remain open to critique, and it is necessary to also engage with its critics.</p>

<p><strong>In interdisciplinary projects:</strong> Systems theory as a shared language that enables precise communication across disciplinary boundaries. When a computer scientist and a sociologist can both speak about “structural coupling,” they have a common vocabulary for something that is otherwise very difficult to articulate.</p>

<h2 id="when-the-model-becomes-an-actor">When the Model Becomes an Actor</h2>

<p>Let me come to an end by looking at a recent example of a system that should be looked at with a systems-theoretical perspective.</p>

<p><em>Prediction markets</em> offer a particularly instructive illustration of dynamics. 
These are markets on which participants can bet on the outcome of future events, for example, from election results to whether a politician will use a particular word in a speech to whether a new drug will pass clinical trials. 
Each bet takes the form of a contract that pays out one dollar if the event occurs and nothing otherwise. 
The current market price therefore directly reflects the collective probability estimate: a price of $0.73 means the market assigns a 73% probability to the event occurring. 
The promise is that markets aggregate information efficiently, meaning those with superior knowledge can profit, which draws well-informed participants and, ideally, surfaces something close to “the truth”.</p>

<p>Prediction markets try to translate political and social complexity into the economic code of profitable/unprofitable (or payment/non-payment). But what happens when we try to solve a political crisis (polarization) using an economic subsystem?
Luhmann would argue this causes a mismatch because the political system operates on the code of power/opposition, not money.</p>

<p>The most obvious problem that follows is one we already know from sports betting: the possibility of corruption. If you can bet on an outcome, you have a financial incentive to influence it.
Prediction markets extend this temptation to virtually every domain—elections, policy decisions, scientific results, media events.
The set of potential actors willing to manipulate an outcome grows correspondingly large, and the markets themselves make it easier to identify which events are both consequential and controllable by a small group.
In Luhmann’s terms, money is being used to <strong>pierce the operational closure of functionally differentiated systems</strong>.
This could essentially make the political system respond to economic operations, or the scientific system respond to financial incentives (directly, i.e. suddenly the operations of system A operate <strong>in</strong> system B). 
Functional differentiation made this kind of interference unlikely; it made each system efficient with respect to its own internal code.
Aside from obvious moral issues, breaking it down has costs on a functional level.</p>

<p>But the more fundamental problem persists even if we set aside corruption entirely and imagine a world of perfectly honest participants. It is the problem of recursive self-description. 
A prediction market does not merely observe the probability of an event but it also publishes that observation, and in doing so becomes a participant in the very system it was meant to describe from the outside.</p>

<p>In many cases such <em>second-order observation</em> effectively reduces complexity, for example, when you want to buy a house it is not necessary to go to the house and calculate how much it should cost by hand.
You observe how others observe the house via the housing market.</p>

<p>In the case of prediction markets, second-order observation is part of the signal an individual brings into the prediction and thereby undermines the whole concept of “drawing in individuals with superior knowledge”.
And worse: a market price of 85% in favor of a particular election outcome influences how voters, campaigns, donors, and media organizations behave, which in turn influences the outcome. 
The observer has entered the system. 
The model is no longer a neutral representation; it is an <strong>actor</strong>.
That media organizations seem highly interested in coupling their operations with these markets could likely amplify these effects.</p>

<p>The possible consequence is that prediction markets are not simply truth-discovery mechanisms that occasionally malfunction. 
They are an interesting case for systems theorists because, as I hinted at, they make <em>observations</em> of the models of psychic systems, i.e. our beliefs, <em>observable</em>.
Consequently, at scale, they risk becoming machines that produce self-fulfilling prophecies—not reading the future from a god’s-eye view, but actively shaping it through the feedback loops they generate. 
Whether this is a net gain depends on a question the markets themselves cannot answer: what value does a prediction that influences its own outcome actually produce? 
And for whom?
And, assuming the system works effectively—which is highly questionable—is it a net good to know what an aggregate believes about the future?</p>

<p>Certainty, we desire certainty—but at what cost?
One can argue that by turning tragic or highly contingent events (like elections, wars, or climate disasters) into financial bets, the system achieve complexity reduction at the cost of <em>empathy</em>!
The possibility to bet on whether there will be a Russia-Ukraine ceasefire before a certain video game comes out, feels intuitively ethically wrong and deeply distasteful.
Fittingly the Kalshi ad read: <strong>The world’s gone mad, trade it.</strong> 
It is, of course, also very dubious that the son of the President of the U.S. is an adviser to two of these markets.</p>

<blockquote>
  <p>Prediction markets are the future. I think they are the future not just for traders but also for news and information. – CEO of Robinhood</p>
</blockquote>

<blockquote>
  <p>New York City was shut down [during COVID] and I asked myself: when is this going to end? When will the vaccine gonna be ready? When is shelter in place to be over? And prediction markets can take all these disparate opinions that people are pontificating about or that they have really good reasons to believe and distill it down into one probability. […] When I get hit up by people in the Middle East who are saying that “You know we’re looking at Polymarket to decide whether we sleep near the bomb shelter.” And I am like “Oh, it is really that popular over there?” That is very powerful. That is like an undeniable value proposition that did not exist before. The global truth machine is here, powered by the people. – CEO of Polymarket</p>
</blockquote>

<blockquote>
  <p>We are a financial market like a stock market but you trade on politics, weather, climate, economics, sports, and so on. This can rival the stock market. – CEO of Kalshi</p>
</blockquote>

<p>Therefore, going back to my <a href="/Pages/2026/04/03/cruelty-and-solidarity-en.html">previous post</a>, I pose the question: <strong>is this attempt to eliminate contingency a path worthy to go?</strong>
While we seem to crave a new form of order, we must interrogate this impulse. Any distinction between <em>order</em> and <em>noise</em> is drawn by an observer who is necessarily blind to the conditions of their own observation.</p>

<p>Furthermore, a paradox emerges: the non-linear dynamics of prediction markets will likely be reflected in their own unpredictability. Systems theory teaches us that complexity cannot be destroyed. The environment always possesses higher complexity than the system, forcing the system to reduce this complexity to a manageable level by ignoring a massive amount of “side” effects—thereby creating new, unpredictable problems.
A system’s reduction of complexity is always local and temporary, requiring a corresponding increase in the system’s own internal complexity. By drawing a boundary and simplifying what gets let in, the system inadvertently triggers a wave of new dependencies, cascading side effects, and emergent structures. Thus, the very mechanisms we deploy to cope with complexity become the engines that generate more of it.
Again, systems theory screems: <strong>social, psychic, living, and technical systems are out of control!</strong></p>

<h2 id="an-invitation">An Invitation</h2>

<p>This text is an invitation to reflect. I am by no means an expert on the subject and my reading list is long and growing.
It is an attempt to articulate an intuition about my observation of the development of technical systems that are increasingly socially and psychologically disruptive.</p>

<p>The intuition is this: we build things we do not fully understand—not in the technical sense, but in the systemic one. We know how our algorithms function. We do not know well enough how they function in the world.</p>

<p>This is no reason for paralysis. No engineering endeavor waits until all effects are fully known before it begins. But it is a reason for humility and curiosity. The systems that concern us are larger than the boundaries we have drawn around them.</p>

<p>Perhaps it is time to renegotiate those boundaries. Not in order to leave computer science behind, but to extend it. Cybernetics and systems theory offer for this purpose a vocabulary that has matured over fifty years and remains largely unused and is largely unknown in the very discipline that emerged from it, and yet which might need it most urgently.</p>

<p>Perhaps it is worth attempting to take a look inside—not to eliminate uncertainty or to build systems of total control but to acknowledge our own ignorance and limits of control.
A cybernetics worthy of its noble philosophical heritage must not be reduced to the mere fine-tuning of a self-guided missile (control as error minimization against a fixed target), but must instead understand <em>control as self-control</em>—as the capacity to know one’s own ignorance and to place one’s own goals and norms up for discussion.
In other words, I believe we should cultivate a <em>cybernetics</em> worthy of its philosophical heritage that maintains the knowledge of its own ignorance, rather than boasting of the mechanical reduction of complexity—a cybernetics that understands <strong>control</strong> in terms of <strong>agency</strong> and <strong>autonomy</strong> and the fine-tuning of doubt, not as a pretext for the <em>hubris of domination</em>.</p>

<h2 id="references">References</h2>

<ol class="bibliography"><li><span id="luhmann:1998">Luhmann, N. (1998). <i>Die Gesellschaft der Gesellschaft</i> (p. 1164). Suhrkamp.</span></li>
<li><span id="zoennchen:2025">Zönnchen, B., Dzhimova, M., &amp; Socher, G. (2025). From intelligence to autopoiesis: rethinking artificial intelligence through systems theory. <i>Frontiers in Communication</i>, <i>Volume 10 - 2025</i>. https://doi.org/10.3389/fcomm.2025.1585321</span></li>
<li><span id="wiener:1954">Wiener, N. (1954). <i>The Human Use of Human Beings: Cybernetics and Society</i>. Garden City, New York : Doubleday.</span></li>
<li><span id="wiener:2019">Wiener, N. (2019). <i>Cybernetics or Control and Communication in the Animal and the Machine</i>. The MIT Press. https://doi.org/10.7551/mitpress/11810.001.0001</span></li>
<li><span id="foerster:2003">von Foerster, H. (2003). Cybernetics of Cybernetics. In <i>Understanding understanding: Essays on cybernetics and cognition</i> (pp. 283–286). Springer New York. https://doi.org/10.1007/0-387-21722-3_13</span></li>
<li><span id="barry:2012">Razeto-Barry, P. (2012). Autopoiesis 40 years later. A review and a reformulation. <i>Origins of Life and Evolution of Biospheres</i>, <i>42</i>(6), 543–567. https://doi.org/10.1007/s11084-012-9297-y</span></li>
<li><span id="maturana:1987">Maturana, H. R., &amp; Varela, F. J. (1987). <i>The Tree of Knowledge</i>. Shambhala.</span></li>
<li><span id="bateson:1972">Bateson, G. (1972). <i>Steps to an Ecology of Mind</i>. Chandler Publishing Company.</span></li>
<li><span id="shannon:1948">Shannon, C. E. (1948). A mathematical theory of communication. <i>Bell Syst. Tech. J.</i>, <i>27</i>(3), 379–423.</span></li>
<li><span id="beer:1995">Beer, S. (1995). <i>The Heart of Enterprise</i>. Wiley.</span></li>
<li><span id="luhmann:1984">Luhmann, N. (1984). <i>Soziale Systeme: Grundriß einer allgemeinen Theorie</i>. Suhrkamp.</span></li>
<li><span id="wing:2006">Wing, J. M. (2006). Computational Thinking. <i>Commun. ACM</i>, <i>49</i>(3), 33–35. https://doi.org/10.1145/1118178.1118215</span></li></ol>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:1" role="doc-endnote">
      <p>Norbert Wiener probably would have argued against this distinction because for him noise is a problem for any system be it a machine or an organism. The contradiction dissolves when one realizes that Luhmann and Wiener are analyzing technology at two different levels: operational behavior vs. structural programming. <a href="#fnref:1" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:2" role="doc-endnote">
      <p>Of course, regulations or other incentives have to be installed to make sure that the desired systemic goals are in fact the goals of the organizations that set them. But this is itself a systemic problem because such an organization is a complex system that can not be steered or controlled directly but can only be irritated. <a href="#fnref:2" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content><author><name>Benedikt Zönnchen</name><email>benedikt.zoennchen@web.de</email></author><category term="Opinion" /><category term="Education" /><category term="Systems Theory" /><category term="Cybernetics" /><summary type="html"><![CDATA[Imagine we are given the task of building a recommendation system. The requirements are clear: the system should suggest content to users that they are highly likely to click on. We define a metric. Let’s say, time spent on the platform. Then we optimize toward it. We test, iterate, deploy. The system works. The metric rises.]]></summary></entry><entry><title type="html">The Art of Solidarity</title><link href="https://bzoennchen.github.io/Pages/2026/04/03/cruelty-and-solidarity-en.html" rel="alternate" type="text/html" title="The Art of Solidarity" /><published>2026-04-03T00:00:00+02:00</published><updated>2026-04-03T00:00:00+02:00</updated><id>https://bzoennchen.github.io/Pages/2026/04/03/cruelty-and-solidarity-en</id><content type="html" xml:base="https://bzoennchen.github.io/Pages/2026/04/03/cruelty-and-solidarity-en.html"><![CDATA[<h2 id="a-new-cruelty--on-the-longing-for-certainty">A New Cruelty — On the Longing for Certainty</h2>

<blockquote>
  <p>Part of the idea of a democratic society is that social change comes by reform rather than revolution and this in effect means that the people who have the power actually letting go of some of it. – Richard Rorty</p>
</blockquote>

<p>Out on the market squares the voices are clamoring. They lament the swaying of the stone temples we call democracy. But while our gaze clings to the crumbling facades of power, we overlook the quiet fading of tenderness in the corners of our everyday lives. We cling to a scaffolding of institutions and forget in doing so that political freedom is not a foundation on which we stand, but rather follows an improbable story that we must keep telling ourselves in ever different guises.</p>

<p>Behind the visible trembling of our institutions, a quieter yet far deeper freezing is taking place. It is as though an old contract with humanity has been terminated—a withdrawal from the shared <em>care</em> for this fragile world. Where once stood the promise of striving to alleviate the pain of the other, there grows today a strange urge toward a <em>new hardness</em> <a class="citation" href="#kohlenberger:2024">(Kohlenberger, 2024; Carolin Amlinger, 2025; von Redecker, 2026)</a>.</p>

<p>The language that once bound us has turned to stone. It knows only the absolute: sharp edges of “facts” and “enemies”, a vocabulary cast as though from ore, leaving no room for what lies between. Attention is directed either toward distraction—the flight from a society one no longer wishes or is able to care about—or toward the destruction of what seems no longer to function. In part patience has run out, and in part curiosity is suffocating in restlessness.</p>

<p>In the narratives of our day, truth has gone astray; a bizarre theater of the obvious reigns, in which the lie is no longer concealed but triumphs as a naked gesture of violence. It is the violence of the unenlightened, in the sense that their centers were never permitted to develop desires. They were not nurtured and could not learn to draw symbolically from their accumulated rage and traumas. And what had previously remained hidden behind etiquette and the effort to maintain decorum now shows itself unveiled: speech is no longer used to be understood, but to humiliate other people.</p>

<p><a class="citation" href="#baudrillard:1976">(Baudrillard, 1976)</a> was right in assuming that the West could not handle the brutal symbolic gift it received on 9/11; by trying to give it back, the system turned inward to give its humiliated power food it can process. Some of this cruelty is old and well accepted if experienced by “the right people” because again and again we rationalized our way towards exceptions; of treating people marked by some features quite different in front of the law. “We” failed the test, unable to realize our self-description. And what could have been a time for self-reflexion became the years of the beast.</p>

<blockquote>
  <p>However difficult this vote may be, some of us must urge the use of restraint. Our country is in a state of mourning. Some of us must say: “Let us step back for a moment and let’s just pause for a minute” and think through the implications of our actions today, so that this does not spiral out of control. […] I came to grips with [my vote] today and I came to grips with opposing this resolution during the very painful, yet very beautiful memorial service. As a member of the clergy so eloquently said: “As we act, let us not become the evil that we deplore.” – Barbara Lee (Single voter of Congress who voted aginst the war in Afghanistan (518 to 1 vote))</p>
</blockquote>

<p>Surely this is too simple; too reductive but still, those scares of these days never healed. 
They left their mark because we disrespected our values and created an inescapable dissonance between what we do and what we imagine ourselves to be.
We contributed to a reality where acts of defiance are increasingly labeled as terrorism, and where anyone can be branded a terrorist—especially those whose solidarity is based purely on human suffering, regardless of political or national boundaries.
This <em>war on terror</em> exists as a profound paradox: it is at once a direct assault on the foundational tenets of liberalism and the very mechanism deployed to safeguard them.</p>

<p>This systemic turning against our own foundational structure mirrors a profound, quiet madness; one that manifests vividly in the cinematic image of a solitary wanderer in the eternal ice.
The penguin in Werner Herzog’s <a class="citation" href="#herzog:2007">(Herzog, 2007)</a> narrative becomes a propaganda figure who, for no reason, turns away from the sheltering colony to waddle stubbornly toward distant mountains.
It is a thoroughly extraordinary march against the penguin’s own environment, against the instinct for survival, and against every community of interdependent beings.</p>

<p>There, in the desolation of the peaks, the wanderer hopes to open up a new realm of winners.
In his eye it shall be a purifying reincarnation.
And, like many Italian futurists, he looks ahead, towards speed, violence, technology, industry, and war.
Yet beneath this mechanical zeal runs a deeply religious current—an echo of American end-time mythologies where a disappointing world cannot be redeemed, only consumed by fire.</p>

<p>But the wanderer has to be certain of his cause.
Doubt—his own as much as that of his companions—is what gnaws at him.
In this apocalyptic logic, doubt is not just a hesitation; it is a temptation by the Antichrist, a betrayal of the absolute faith required for the final days.
From the perspective of his colony, which is setting out toward the ocean, it is a mission without tomorrow, driven by a dark longing for the end—a final act of retribution against a world one can no longer imagine made better.
And in this departure, the cruelty directed against curiosity and doubt becomes the perverse satisfaction of being right, at least with the prophecy of downfall. 
In our ears Herzog’s voice resonates:</p>

<blockquote>
  <p>But why? – Werner Herzog</p>
</blockquote>

<p>And so, we applaud the wanderer’s courage, attempting to re-interpret what seems like a mix of nihilism and destructive vitalism as a desperate call to save Europe <a class="citation" href="#gundlach:2026">(Gundlach, 2026)</a>.
But this wanderer does not follow life; he follows a myth of freedom that leads him irresistibly into the white void.
He is dead.
He won against his environment—against a <em>careless</em> nature:</p>

<blockquote>
  <p>With five thousand kilometer ahead of him, he’s heading towards certain death. – Werner Herzog</p>
</blockquote>

<p>A remarkably similar architecture of self-imposed isolation is being constructed today in the shadow zones of our digital world.
Here, a brotherhood of grievance has formed, a loose network of voices that find their identity only in the echo of hatred.
They need a face they can despise, and because sensitization has taken something from them, solidarity itself has become the enemy.
The call for a return to the “true mask” of man is more than mere nostalgia; it is the desperate longing for an old, heavy armor that admits no cracks and thus no vulnerability.
It is an anarchic armor that no longer requires solidarity at all, because it is its own ecosystem—self-sufficient, optimized, independent, and closed in on itself.
At the same time, behind the facade lies deep suffering, because man has lost something that once made his world simple and secure.</p>

<blockquote>
  <p>Destructive attitudes develop predominantly in people who perceive themselves as marginalized, yet are status-ambitious and dominance-oriented. Individuals with a destructive mindset are convinced that they are being denied a social position to which they have a legitimate claim. – Oliver Nachtwey</p>
</blockquote>

<p>He could no longer reconfigure his identity through the new vocabulary.
He was <em>humiliated</em> and sought another language, so that everything that had seemed particularly important to him might once again become true and good.
What he found is <strong>the language of dominance</strong>.</p>

<blockquote>
  <p>The will to dominate was the fundamental law of the life of the universe from its most rudimentary forms to its most elevated ones. That man was driven by a divine bestiality. – Benito Mussolini</p>
</blockquote>

<!--

In the words of Palantir's manifesto:

>The limits of soft power, of soaring rhetoric alone, have been exposed. The ability of free and democratic societies to prevail requires something more than moral appeal. It requires hard power, and hard power in this century will be built on software. [...] We must resist the shallow temptation of a vacant and hollow pluralism. We, in America and more broadly the West, have for the past half century resisted defining national cultures in the name of inclusivity. But inclusion into what? 

-->

<p>It is as though history were violently recoiling.
Where the world had begun to grow quieter and more sensitive to the pain of “the other”, a part of it responds with a new, steely coldness.</p>

<blockquote>
  <p>Tonight, a whole civilization will die and never return. […] There might something revolutionarily wonderful happening, who knows. – Donald Trump</p>
</blockquote>

<p>Yet beyond the loud grievance of the streets, in the soundless, glass-walled cathedrals of light and silicon, a far quieter, almost clinical cruelty is ripening.
It is a faith that bundles itself in seven cold stars into a single radiance—an alliance of those who regard “the human being” as a transitional sketch to be technologically overcome <a class="citation" href="#gebru:2024">(Gebru &amp; Torres, 2024; Mühlhoff, 2025)</a>.
In this light, the longing for the stars no longer appears as a departure but as a flight; an expansion into the void, driven by the dream of an eternity that no longer needs a body; to become a misremembered Puppet Master—a ghost whithout a shell in a Cartesian fantasy.</p>

<p>From the heights of their galactic calculations, these architects of the future look down on the here and now as upon an ant colony in the dust.
The suffering of the present—the exhaustion of the earth, the silencing of diversity—shrinks in their eyes to a negligible rounding error.
It is an ethics that sacrifices today to a speculative singularity of tomorrow, a morality of arithmetic, a <em>rule of code</em> instead of law <a class="citation" href="#rosengruen:2022">(Rosengrün, 2022)</a> in which a burning planet weighs less than the mathematical promise of a posthuman world of gods. 
There is nothing novel or imaginative about these ideas; they are archaic myths that provide narrow futures.
We must remain acutely aware of these ancient dreams of greatness that demand a catastrophic downfall to purge “decadence” and to bring a new <em>technological order</em> into choas.</p>

<blockquote>
  <p>We should try to create autonomous countrys on oceans, under water, and all sorts of other spaces. Technology is the vehicle to escape and move beyond politics as we find it today. – Peter Thiel</p>
</blockquote>

<p>In this worldview, an old dark spirit returns, cloaked in the garb of logic: the conviction that life has a price measured by its utility for the great progress, which could be gauged by its approximation to the <em>absolute</em>; that the immaterial disguised as weightless information is real and the material—bodies and trees, flowers and animals, mountains and oceans—can be overcome.
Once again it is assumed that everything can be calculated, but in place of Kantian principles stand calculating machines and the theories of probability and expected values.
Once again we await a god—this time a god made of numbers—a superintelligence that, like an infallible oracle, could end the chaos of our interpersonal stories—as though society could communicate with anything other than itself.
It is Plato’s ancient, stony dream—the hope that pure, incorruptible truth might finally triumph over the tender but imperfect narratives of compassion; that we might remember what has always been out there and within us; that in this “awakening” the True and the Good coincide.</p>

<blockquote>
  <p>Plato thought that morality and politics should be based on principles in the same why that Euclidean geometry is based on axioms. He thought that philosophical inquiry was a matter of nailing down firm immutable principles which could then guide action. […And] as Plato said, it is if we had known the truth in a pervious existence and simple need to be reminded of it, have it brought back to consciousness.  – <a class="citation" href="#rorty:1990">(Rorty, 1990)</a></p>
</blockquote>

<p>But in this purity there is no longer room for breathing, for trembling, for the passion, love and affection for a world that is precious precisely because it can keep reinventing itself. There will be no gods only the loss of institutions that once balanced and distributed power. And concentrated power, ultimately, devours empathy if you cannot strip yourself of it in time.</p>

<blockquote>
  <p>It seems to me the Platonic notion of absolute truth is a thoroughly misleading slogan and a culturally dangerous shibilith. – <a class="citation" href="#rorty:1990">(Rorty, 1990)</a></p>
</blockquote>

<p>This frosty belief is nourished not least by a bottomless exhaustion—a cynicism that crystallizes like a dark sediment out of the fear of the end. It is the rebellion against the last remaining justice: death, that relentless equality to which we all succumb. In the impatience to outwit this fate, hope turns into bitterness. Some have everything but can only continue to create themselves by getting rid of themselves. Others have nothing and can only react to disruptions. They have no time for self-description. What is a calculable <em>risk</em> for few becomes a <em>danger</em> for everyone else.</p>

<p>A quiet, gray poison forms, which has long since crossed the threshold of our homes. It nestles into the corners of our living rooms, a mood like a silent echo of those chroniclers of hopelessness who whisper to us that the world has become a closed circle. In the age of disruptions <a class="citation" href="#stiegler:2019">(Stiegler, 2019)</a>—which feels more and more like an age of destructions—the chords that present us with a horizon of notes fall silent. Structures dissolve, so that we have difficulty imagining different futures.
We feel it in the burden of everyday life: the sense that what exists can no longer be healed, that the system is frozen at its foundations and can only be broken.</p>

<p>As different as these three figures may appear—the ranter on the market square, the calculator in the server room, the exhausted person on the sofa—they share a common root: the inability to live with the contingency of life.
All three want certainty.
One fights for it through enemies, another calculates it through machines, and the third finds it in the renunciation of all hope.</p>

<p>In this darkness we look at ourselves and see only deficiencies.
The human being no longer appears to us as a continually changing riddle of openness, but as a flawed, inadequate disruptive factor to be optimized away.
While the calculating machine promises eternal perfection, the human being becomes a creature of dust and error, one that can readily be dispensed with.
We lose faith in the laborious, small gesture of improvement and instead begin once more to dream of perfection, as though we could knowingly move toward the summit.
It is the <em>exhaustion of the contingent</em> and thus the wish that something might finally be certain.</p>

<hr />

<h2 id="doubt-as-home--contingency-irony-solidarity">Doubt as Home — Contingency, Irony, Solidarity</h2>

<p>Amid this recurring coldness, this icy wind of abstraction, a voice is needed that leads the resistance against dehumanization not as a loud protest, not as instruction or a return to the rational, but as a healing gesture of humility: it wrests the human condition from the cold, sterile distances of the heavens and beds it back in the warm, imperfect dust of the earth.</p>

<p>Richard Rorty is such a voice—a voice that prefers doubt to certainty.
As a philosopher of <em>irony</em> (in the sense that we should not take our existing vocabularies too seriously), of <em>contingency</em> (in the sense that there are no ahistorical truths and history follows no necessary path), and of <em>solidarity</em> (in the sense that he prefers it to truth) <a class="citation" href="#rorty:1989">(Rorty, 1989)</a>, he wanted—naively put—to strip philosophy of its imperialist position as the foundational discipline—and to propose, of all things, the literary critic as a model for the public intellectual, which fairly earned him the reputation of relativist and charlatan.</p>

<blockquote>
  <p>By an intellectual I mean someone who has doubts about the value of the language she is been using to make moral or political judgements and who reads books in an effort to deal with these doubts. To be an intellectual is to have a restless mind—never to be sure that once judgement of other people’s characters or of alternative social institutions are more than inherited prejudices. – <a class="citation" href="#rorty:1990">(Rorty, 1990)</a></p>
</blockquote>

<p>Meanwhile, he himself wrote quite systematically (as in <a class="citation" href="#rorty:1979">(Rorty, 1979)</a>) but was more interested in interpreting and reconfiguring philosophical texts than in erecting a new grand system of his own.
His writing style as well as his philosophical position are American in their simplicity, whereas his admiration squints toward Europe.
He likely embodies much of what disturbs philosophers when they write about the “decline of culture,” for as someone who wants to take neither Kant nor Nietzsche too literally, he refuses answers to cold questions such as: “What <strong>ought</strong> I to do?”—in the sense of a universal duty. The warm question, however, “What can <strong>we</strong> do for one another?”, he does answer, but without metaphysical backing: Reduce cruelty, expand your circle of solidarity—not because reason commands it, but because we have learned what it means to be humiliated.</p>

<p>His irritating position on “the Truth” and “the really real” is a melting of pragmatism and romaticism which can be summed up by the following quote:</p>

<blockquote>
  <p>On the account of human abilities I am suggesting, the use of persuasion rather than force is an innovation comparable to the beaver’s dam. Like the beavers’ collaboration in getting the dam built, it is a social practice. It was initiated by the noval suggestion that we might use noises rather than physical compulsion to get other humans to cooperate with us. That suggestion gave rise to language. Rationality, thought, and cognition all began when language did.<sup id="fnref:1" role="doc-noteref"><a href="#fn:1" class="footnote" rel="footnote">1</a></sup> Language gets off the ground not by people giving names to things they were already thinking about but by proto-humans using noises in innovative ways, just as the proto-beavers got the practice of building dams off the ground by using sticks and mud in innovative ways. Language was, over the millenia, enlarged and rendered more flexible not by adding the names of abstract objects to those of concrete objects but by using marks and noises in ways unconnected with environmental exigencies. The distinction between the concrete and the abstract can be replaced with that between words used in making perceptual reports and those unsuitable for such use. […W]e need to think of reason not as a truth-tracking faculty but as a social practice—the practice of enforcing social norms on the use of words rather than blows as a way of getting things done. We need to think of imagination not as the faculty that produces visual or auditory images but as a combination of novelty and luck. To be imaginative, as opposed to being merely fantastical, is to do something new and to be lucky enough to have that novelty be adopted by one’s fellow humans, incorporated into their social practices. […] People whose novelties we cannot appropriate and utilize we call foolish, or perhaps insane. Those whose ideas strike us as useful we hail as geniuses. – <a class="citation" href="#rorty:2016">(Rorty, 2016)</a></p>
</blockquote>

<p>What Rorty offers me is an admirable description of the liberal and democratic society that unerringly exposes anti-liberal opposites.
He gives me an answer to the question “What do liberals want and what can they hope for?” with which I can agree.
And although—or precisely because—he takes leave of <em>first principles</em> and speaks soberly about texts from Wittgenstein to Proust, he can inspire enthusiasm for a liberalism by understanding it as <strong>the art of solidarity</strong>.</p>

<p>Let us begin with a note on why educational institutions in particular are so decisive for it:</p>

<blockquote>
  <p>In democratic societies like ours, colleges and universities have a peculiar two-faced role. They get their money by promising to furnish money-making skills to their students and by promising to perform research which will enable society as a whole to get more goods and services more cheaply. The face they present to rich donors, state legislators, and the general public is essentially a commercial one. They suggest that they have certain products which the society as a whole needs and they ask for support on that basis. They usually don’t suggest that their function is to disturb the students, make them have doubts about the way they were brought up, force them to ask unanswerable questions. But as you know quite well, that is the function which many faculty members, especially the people who teach in the humanities and social science departments, think that colleges and universities should serve. Such people see the promise of marketable skills as simply the lure which brings students within their reach. Once the student is in their classes, these people assign books which will—they hope—upset her enough to make her want to start for looking for other such books so as to get upset in still more complicated ways. They are not satified unless the students who leave their courses are dissatified with the society in which they live and unless they have at least some doubts about the moral codes in which they were brought up. […T]he point of encouraging dissatisfaction becomes clear when times are bad and particularly when societies and governments become repressive. Then the colleges and universities come into their own. They begin to function either as sources of social change or as sanctuaries for resistance. […] I can sum up the two roles of colleges and universities by saying that whereas the society as a whole wants to produce people with skills, the university faculties also want to produce intellectuals. – <a class="citation" href="#rorty:1990">(Rorty, 1990)</a></p>
</blockquote>

<p>Rorty’s thinking begins where the coldness described above has its deepest origin: in the belief that there is (or was) a final truth toward which we are moving (or from which we have depatured).
He distrusts this belief—not out of cynicism, but out of a deep <em>care</em> for what we lose when we chase after it.
For Rorty, truth is a property of sentences, and sentences are made by people.
“The world” cannot make these sentences true, but we can make them useful for ourselves and useful for <strong>us</strong>.
In the age of “fake news” this may sound irritating, for is it not “the truth” whose loss we lament?</p>

<p>In place of principles that we might yet discover, Rorty stakes everything on new, persuasive vocabularies that are to be invented.
For if one wants to say something new, one must create a new language <a class="citation" href="#maturana:1991">(Maturana, 1991)</a>.
Freedom and solidarity arise where we find words, metaphors, and descriptions that allow us to see ourselves, others, and the world differently.
<strong>The power of language lies not in mirroring reality, but in reconfiguring it.</strong></p>

<blockquote>
  <p>I do indeed assert that the explicit or implicit answer to the question of reality determines how we lead our lives and in what way we accept or reject other people within the network of the social and non-social systems we form. – <a class="citation" href="#maturana:1988">(Maturana, 1988)</a></p>
</blockquote>

<p>Kant provided some good reasons to prevent cruelty, and Nietzsche shattered their <em>claim to universality</em>—he shattered the idea that there is any entity, whether God or Reason, that could decide outside a historical-evolutionary context: “Who I am”, “What I should do”, “What I can know” and “What I may hope for.”
Kant undertook the attempt to formulate principles that should hold for everyone, thereby creating the basis for solidarity, human rights, and democracy.
But his Reason could not order the surplus of meaning that arose from the newly won freedom: too many opinions, too many perspectives, too much criticism and deconstruction.
Kant overlooked the obvious: that he himself as observer must remain blind to his own observing.
Once again the paradox erupted and an attitude settled in that at best tolerates uncertainty without despairing.
We are free but also overwhelmed—that is perhaps the dilemma of liberal democracy: we do not know what to do and must decide nonetheless.</p>

<p>Authors such as Nietzsche and Heidegger offer vocabularies for self-creation, for individual meaning-making beyond metaphysical certainties and beyond reason.
Nietzsche is perhaps the epitome of the <em>ironic theorist</em> who unmasks every supposed truth as a contingent product of a historical narrative—except, of course, his own.</p>

<p>He is an artist: playful, contradictory, vital, desperate, brutal and jolly who teaches us that no description of the world is necessary—that we could always tell other stories to understand ourselves and our community differently.
Kant, on the other hand, reminds us that such new creations have limits where they overlook the suffering of others.</p>

<blockquote>
  <p>Rationality is a matter of making allowed moves within a language game. Imagination creates the games reason proceeds to play. Then […] it keeps modifying those games so that playing them is more interesting and profitable. Reason cannot get outside the latest circle that imagination has drawn. It is in this sense, and <strong>only</strong> this sense, that imagination holds the primacy – <a class="citation" href="#rorty:2016">(Rorty, 2016)</a></p>
</blockquote>

<p>Rorty recognizes in this tension the productive condition of a liberal modernity: without Nietzsche, no renewal; without Kant, no consideration. <strong>We need Nietzsche so that life remains interesting, and Kant so that it remains bearable.</strong></p>

<blockquote>
  <p>Rational discussion is not an appeal to eternal standards, but simply an attempt to make our beliefs and desires as coherent with one another as possible while constantly adding new beliefs and desires to the old. – <a class="citation" href="#rorty:1990">(Rorty, 1990)</a></p>
</blockquote>

<p>However, Rorty emphatically reminds us that most people do not want to be redescribed, and that an imposed redescription is usually cruel.
People generally want to be taken as they speak.</p>

<blockquote>
  <p>There is something potentially very cruel about the claim that [the language people speak is, for the ironist, a matter of chance]. For the most effective way of causing people enduring pain is to humiliate them by making the things that seemed most important to them look futile, obsolete, and powerless. – <a class="citation" href="#rorty:1989">(Rorty, 1989)</a></p>
</blockquote>

<p>I think this is one of the most remarkable statements, one that emphasizes the dark side of the <em>ironic attitude</em>.
The ironist knows that every vocabulary (the way we talk)—every language in which people give meaning to their lives—is contingent, and therefore replaceable, not final, but neither useless nor arbitrary.
An ironist—such as Nietzsche—can, at any moment, show that the concepts by which someone lives are merely historical accidents.
But that is precisely what can be a form of cruelty.
For if I convincingly demonstrate to someone that everything he believes in, lives for, and that constitutes his identity is only a “matter of chance,” I have not simply offered him a better argument.
I have pulled the ground out from under his feet.</p>

<p>This “ground” is nothing less than the person’s <em>horizon of sense</em> (Sinn), i.e. a horizon of an inescapable medium; we cannot step outside of it <a class="citation" href="#luhmann:1984">(Luhmann, 1984)</a>. A person’s final vocabulary is the precise mechanism they use to select meaning out of a chaotic world and stabilize their mind. When the ironist ruthlessly exposes this vocabulary as a mere accident, they do not open a door to total freedom; instead, they threaten the person with <em>structural collapse</em>.</p>

<p>Rorty urges caution when it comes to proposing other vocabularies.
In the public sphere, in dealings with others, irony is potentially cruel if it redescribe someone’s past and make them look foolish or obsolete—and yet it is new vocabularies that enable moral progress, while at the same time the replacement of vocabularies can be the worst form of cruelty. It is the ultimate tragic bind: we must change our vocabularies to progress, but in doing so, we risk inflicting a structural violence that strips others of their very capacity to make sense of the world.
<em>Solidarity</em> itself can cause smaller and smaller circles if it defines its identity through the creation of sharp boundaries, turning solidarity into an internal code that treats everyone outside the group as mere background noise or structural threats.</p>

<blockquote>
  <p>All we can do is to compare new customs and institutions with old customs and institutions in the experimental and tentative way in which we compare new friends, new jobs, or new environments with old ones. The only test of truth is that it is the view that wins in a free and open encounter. But the result of that test can only be accepted until somebody comes up with some new proposal, a new scientific theory, a new artistic style, a new political institution. Then discussion will have to be undertaken all over again. There will never be a time when Socratic questioning becomes unnecessary. […] Kant said, following Plato, that the source of moral obligation must be a distinct faculty—reason rather than emotion—because reason is part of human nature whereas emotions are just contingent features of particular individuals. Even someone like myself, who wants to discard the notion of intrinsic human dignity and unconditional moral obligation has to admit that these Kantian uses have been extremely useful. The <strong>ethics of sensitivity</strong> which I have associated with the figure of the literary critic, may seem to endanger all the gains made in recent times with the help of these Platonic and Kantian notions. [… However] we may find a way to do it without the ladder we climbed. […] An ethics of sensitivity assumes that morality is not a matter of recognizing unconditional obligations built into every human being simply by virtue of being human, but rather of community obligations—obligations one feels as a member of a group. […S]uch obligations determine one’s identity as a member of a community. It is one thing to treat someone weaker less advantaged than oneself decently because one happens to feel kindly toward him—perhaps because of some unconscious accidental association with one of his features or trades. It’s another thing to recognize this person as a fellow citizen, one of <strong>us</strong>, the sort of person to whom <strong>we</strong> are obliged to behave decently. […] It is a matter of coming to see more and more different sorts of people as us. Seeing an individual who lives quite a different life form as our own as, nontheless, one of us. – <a class="citation" href="#rorty:1990">(Rorty, 1990)</a></p>
</blockquote>

<p>Rorty advocates for a liberalism for which there is nothing worse than humiliating others, for humiliation is the deepest form of cruelty <a class="citation" href="#rorty:1989">(Rorty, 1989)</a>.
It is the destruction of the self-image and of the language in which a person gives meaning to their life.
The problem with Rorty, one that has perhaps caught up with us today, is that for him there are no <em>final grounds</em> with which he could defend his liberal position.</p>

<blockquote>
  <p>Substituting this sort of practical question for theoretical questions about first principles means admitting that there is no way to answer such critics of democracy such as Plato, Nietzsche, or Hitler. There is no neutral ahistorical ground one could stand on when members of a democratic community try to argue with people who ask whether their society may not be headed in exactly the wrong direction. First principles are rationalizations of existing habits and institutions […] which is no reason to distrust them automatically. The Homeric heroes, the Nazi concentration camp guards, the pre-civil war slave owners all had principles. But their principles did not save them from cruelty to people whom they did not think of as us. What counts for moral progress is not firmness in abiding by established habits or institutions or principles, but rather the willingness to ask who’s getting hurt by the existence of these institutions or by the application of these principles. – <a class="citation" href="#rorty:1990">(Rorty, 1990)</a></p>
</blockquote>

<p>The conviction that cruelty is the worst thing comes from a particular historical and cultural development, from the European Enlightenment, from democratic revolutions, from the gradual expansion of the circle of those to whom we extend compassion.
It is <em>contingent</em>, that is, neither necessary nor impossible.
But that does not make it any less valuable or binding for Rorty.
One can passionately commit to something without claiming: the universe stands behind it.</p>

<p>Why, then, should we not humiliate others?
Because through experience, through stories, through listening, we have learned what it feels like to be humiliated.
Because we have built a culture that cultivates this sensitivity.
And because the attempt to provide a deeper justification for this leads us astray: it suggests that someone not convinced by the argument could be rationally refuted.
But the <em>liberal ironist</em> cannot accomplish this.
The sadist, to whom the suffering of others is not only indifferent but who seeks elevation and a <em>destructive vitality</em> by humiliating others, lacks not an argument according to Rorty—he lacks a certain capacity for empathy that cannot be logically derived, but can only be cultivated.</p>

<blockquote>
  <p>The development of civilization on this view is not the triumph of reason over passion but <strong>the triumph of tolerance over distrust</strong>. Therefore, democratic society is not founded on a sense of obligations, but on a sense of sympathy. […] The increasing egalitarianism of the democracies is not a matter of recognizing that illiterate laborers, blacks, women, and gays are as rational beings as middle-class straight white males, but rather of those males themselves—the people who have a monopoly on power—coming to realize that these people have the same hopes and fears and the same susceptibility to pain and humiliation as they do. – <a class="citation" href="#rorty:1990">(Rorty, 1990)</a></p>
</blockquote>

<p>Heidegger diagnosed in his Spiegel interview <a class="citation" href="#heidegger:1976">(Heidegger, 1976)</a> the powerlessness of all theory and politics before the <em>forgetfulness of Being</em> in modernity. Such a diagnosis Rorty would have rejected in its metaphysical depth. And yet both, on very different paths, share the skepticism toward philosophy as savior: Heidegger, because only a god could help; Rorty, because there are no ultimate justifications.
But Rorty avoids the sort of despair Heidegger seemed to hold.
Christianity, Kant’s <em>Reason</em>, and the <em>crisis of the absolute</em> had their time.
They told new stories to bring meaning back into descriptions.
The <em>liberal and democratic society</em>, however, cannot be universally grounded—it can only be narrated <a class="citation" href="#rorty:1999">(Rorty, 1999)</a>.
It lives as long as we find new words for “the good,” “the just,” and “the common.”
It would, however, be dangerous to transfigure it as “the genuinely true,” “the superior,” or “the absolutely just.”
Its salvation lies not in <em>truth</em> but in conversation; in the shared resolution:</p>

<blockquote>
  <p>We don’t do that. We respect one another. We don’t kill other people. We support each other. We always doubt our own sensitivity.</p>
</blockquote>

<p>Rorty knew that the loss of absolute truth would be unsettling for the individual, and in that context proposes distinguishing between the <em>private</em> and <em>public</em> spheres.
In private, the ironist knows that her convictions are contingent (accidental, historically conditioned).
She knows that her values are not God-given.
She has doubts, and that is all right.
She can pursue her striving for self-realization.
Thus Rorty relocates the impulse toward self-creation—which he takes seriously and regards it as important, and which he sees embodied in Nietzsche, Proust, Heidegger, and Derrida—into the private realm.</p>

<blockquote>
  <p>But for Proust and Nietzsche, there is nothing more powerful or important than self-redescription. They are not trying to overcome time and chance but to use them. – <a class="citation" href="#rorty:1989">(Rorty, 1989)</a></p>
</blockquote>

<p>His <em>liberal ironists</em> no longer concern themselves with universals, essences, or absolutes that might metaphysically ground why one should not humiliate others.
They enjoy the writings of Nietzsche, Heidegger, Derrida, and many other <em>ironic theorists</em> for their own <strong>private redescription</strong> without searching for a conclusive public vocabulary.</p>

<p>Rorty saw in Nabokov a central tension of his own philosophy: the conflict between private aesthetic self-creation and public solidarity.
For Rorty, Nabokov embodied the type of the liberal ironist who strives for autonomy and artistic ecstasy, but in doing so runs the risk of becoming cruel.
Thus, for example, Nabokov’s Humbert is highly educated, sensitive, and writes beautifully.
But precisely this search for aesthetic ecstasy makes him blind to the suffering of others.
He does not see Lolita as a suffering child but as an aesthetic object of his fantasy.
Nabokov thereby shows that neither intelligence, artistic sensibility, nor linguistic brilliance automatically makes us morally good people.
One can be an artistic genius and still act cruelly, and although Nabokov never saw himself as a teacher, he helps his reader to notice cruelty in detail.
We therefore need authors like Nabokov for our private lives (to invent ourselves, to be autonomous, to take pleasure in language).
But, Rorty argues, we must not carry this attitude into the public sphere, because a society based solely on aesthetic pleasure would be cruel.</p>

<p>In Rorty’s <em>ironic culture</em> it is no longer a matter of finding the right words or a common reality behind appearances that will connect people.
He gives up the dream of finding the True, the Good, and the Beautiful for <strong>everyone</strong>.
That is why we should stop asking:</p>

<blockquote>
  <p>Does my belief correspond to <em>true</em> reality?</p>
</blockquote>

<p>More important is the question:</p>

<blockquote>
  <p>Does this belief help us live together better, more freely, and less cruelly?</p>
</blockquote>

<p>In his utopia, solidarity is not regarded as a fact that must be recognized by removing prejudices or uncovering previously hidden depths.
Any uncovering becomes impossible once we abandon the search for a common essence or nature, for the effort toward universal justification is then in vain.
Solidarity must therefore be constructed.
Thus Rorty’s liberals make solidarity their goal.
This goal is achieved not through <em>reason</em> or any logical proofs, but through <em>imagination</em> and <em>empathy</em>—through the imaginative capacity to see unfamiliar people as fellow sufferers.</p>

<blockquote>
  <p>The philosopher is driven by what Dewey called the quest for certainty and the literary critic by curiosity and sympathy. The latter is curious about forms of life different from her own and sympathetic to people who leads such lives. The former looks for unity and thinks of philosophical inquiry as converging to a single body of truths. The latter looks for diversity. She is more concerned having missed something, having been condescending and cruel towards someone of a different sort than about certainty. – <a class="citation" href="#rorty:1990">(Rorty, 1990)</a></p>
</blockquote>

<p>Solidarity is not discovered through reflection; it is created.
It is created by increasing our sensitivity—through literature, novels, films, theatre, music, (computer) games, and reportage—to the particular subtleties of the pain and humiliation of others, people unknown to us.</p>

<blockquote>
  <p>[The literary critic], on my view, [is] our modern substitude for the Platonic moral philosopher. […] The traditional notion of a separation between moral judgments and aesthetic judgements is, I would claim, a relic of the idea that there is a deep common human nature which sets moral goals and standards. [… In our times] it is almost impossible to believe that all human beings, male or female, slave or free, illiterate or cultured, ancient or modern, European or Chinese, have always carried around the same vision of goodness and justice deep within themselves. Everything […] suggests the plasticity of human beings. – <a class="citation" href="#rorty:1990">(Rorty, 1990)</a></p>
</blockquote>

<p>Such heightened sensitivity makes it harder to marginalize people who are different from us by assuming that they do not feel the way we do, or that there will always be suffering, so why should one not let them suffer?
It is a process of emotional expansion of the “we”—a change in perception, not in epistemological cognition in the form of an increase in knowledge.
This process of seeing other people as “one of us” rather than “the others” consists in describing in detail what unknown people are like, and in redescribing who we ourselves are and how we recognize one another <a class="citation" href="#rorty:2016">(Rorty, 2016)</a>.</p>

<p>Cruelty and humiliation are abolished by creating realities through dense descriptions that sensitize the reader to the pain of those who do not speak our language.
This task had been expected of proofs for a common human “nature,” but whether grounded in metaphysics or natural science, it can capture the indeterminacy and paradox of the human being in no final description.
As Nietzsche already pointed out,</p>

<blockquote>
  <p>[i]t is only through the forgetting of that primitive metaphor-world, only through the hardening and stiffening of an original mass of images, flowing liquid and hot out of the primal power of human fantasy, only through the invincible belief that this sun, this window, this table is a truth in itself—in short, only through the fact that man forgets himself as a subject, and indeed as an artistically creative subject—does he live with any repose, security, and consistency: if he could get out of the prison walls of this belief, even for an instant, his ‘self-consciousness’ would be immediately destroyed. – <a class="citation" href="#nietzsche:1873">(Nietzsche, 1873)</a></p>
</blockquote>

<p>and Rorty follows him in this insight:</p>

<blockquote>
  <p>We should understand concepts like ‘gravitation’ or ‘human rights’ not as entities whose essence remains mysteriously hidden, but as sounds and marks whose use has made possible more significant and better social practices. Intellectual and moral progress is not an approximation toward a prior goal, but the surpassing of the past. What we call ‘improved knowledge’ should not be interpreted as better access to the real, but as an enhanced ability to do things. […] Freedom begins when we can discuss which words better describe a situation. <strong>Knowledge and freedom develop simultaneously.</strong> – <a class="citation" href="#rorty:2016">(Rorty, 2016)</a></p>
</blockquote>

<p>There is no final, conclusive vocabulary; every language in which people give meaning to their lives is fundamentally contingent. Normally, our shared vocabularies function as a vital baseline for peace. They act as the agreed-upon frameworks where questioning ceases and mutual understanding is stabilized <a class="citation" href="#maturana:1991">(Maturana, 1991)</a>.</p>

<p>“Be objective” is, from the perspective of metaphysical realism, a <em>demand</em> to accept my position—to deny oneself to some degree and to accept a shift in once identity. Complying with this demand requires a kind of Buddhist self-emptying: you are asked to detach from your own contingent horizon of meaning and quietly dissolve your ego, not to touch an ultimate truth, but simply to clear away yourself so that the other person’s vocabulary can occupy the space.
Yet, as a society, we do not truly value or appreciate the immense weight of this act. We casually demand objectivity from others as if it were a simple intellectual correction, entirely blind to the fact that we are asking them to perform a deeply painful, ascetic surrender of their very selfhood just to accommodate our view of the world.</p>

<blockquote>
  <p>The contrasting view [of getting at a certain answer to questions like “What is really good?”, “What is really just?”, “What is really real?”, “What is really true?”, “What is really human?”], which I share with people like William James, Dewey, [and] Satre is that the point of Socrates’ life was not to discover a permanent absolute truth but rather just to keep people thinking, to keep them inventing, to open up their imaginations to alternatives to present convictions. – <a class="citation" href="#rorty:1990">(Rorty, 1990)</a></p>
</blockquote>

<p>Once liberal society has become distrustful of its “inner core,” it can no longer appeal to that core.
What society has learned about itself it will not easily be able to forget.
In other words: Nietzsche’s cut cannot be papered over even when the author has been exposed as a moral reprobate.
Open society will not be saved by a universal truth; not by reason and not by a metaphysical foundation.
Rorty thus reminds us that there is no final authority that could assure us that freedom, equality, or compassion are true and good.
Their validity is not the result of discovery but of narration.</p>

<blockquote>
  <p>To progress morally is a matter of individual and social self-creation rather than self-discovery. – <a class="citation" href="#rorty:1990">(Rorty, 1990)</a></p>
</blockquote>

<p>Precisely for this reason, Rorty burdens literature with perhaps the most important task: to foster the imagination for ways of living, so that in the future we will not say that we lead a more truthful or more sublime life than our ancestors, but that we have developed better ways “of being human” that our descendants may perhaps adopt.</p>

<blockquote>
  <p>If one asks which books helped along such processes such of self-creation in the last few hundred years I think a good case can be made for saying that most of them were novels. Whereas our ancestors relied on scripture or on theological or philosophical treatisis for their notion of what it was to be a human being and what was the point of human life, more recently, we have been relying on books like The Brothers Karamazov, The Magic Mountain, Remembrance of Things Past, and Catcher in the Rye. Novels about young people growing up and creating themselves. If one asks which books have done most to make American society freer and more just, again, a lot of them are novels. Books like Uncle Tom’s Cabin, Black Boy and Invisible Man did more than any philosophical or social scientific treatises to let the whites see what they were doing to the blacks. Books like The Well of Loneliness and The City and the Pillar did more than psychological treatises to let the straits see what they had been doing to the gays. Books like Middle March and The Color Purple did more to make men realize what they were doing to women than any socioeconomic data or any feminist theorizing. – <a class="citation" href="#rorty:1990">(Rorty, 1990)</a></p>
</blockquote>

<p>In his admiration for the ironic theorists, Rorty simultaneously warns us against their absorbing tendencies to want more than self-creation.
For all his love of life and self-redescription, Nietzsche was hardly concerned with preventing cruelty; rather, he was concerned with enabling the few—among whom he counted himself—to invent and describe Europe as something great.
In this sense, Rorty is right to suggest that Nietzsche’s vocabulary can unfold its power in the <em>private</em> domain of our lives, but produces little that is fruitful in the discourse of the open society.
When the work of art that we ourselves are—the one that is neither simply created nor discovered but pre-exists within us and can yet only be found in the process of creation—destroys the possibility of another work of art coming into being, that is, stifles the self-creation of others, then a deep cruelty arises.
Autonomy and self-efficacy are ideals that liberals also revere, but in the darkness lurks the connection between art and torture.
The work of art is not “the good”—even if it is beautiful, it can be cruel.</p>

<p>Over time we have expanded our moral imagination and included more beings in our <em>circle of care</em>, not on account of unchanging principles, not because of God or some inner truth we would have discovered, but because of the vitality of our imagining; not because of transcendental or transcendent truth, but because of lived and imagined experiences that have been narrated and inscribed into our cultural memory.
We have invented a less cruel world.</p>

<p>This at times naively appearing hope—this project of continuation—gains contour where genuine encounters take place in the <em>lifeworld</em>.
On account of my physical impairment, I possess the paradoxical privilege of evoking in others a <em>suspicion of humiliation</em>: the mere appearance of my body leads others to assume that cruelty must inevitably have been done to me.
Even when these projections rest on stereotyped prejudices, they reveal the human capacity for empathy.
The suspicion is plausible, and not entirely wrong.</p>

<p>Two fundamentally different ways of meeting this suspicion crystallize: The first is something like a Christian glorification of suffering.
Here the fateful is elevated to an essence; the sufferer is styled as a martyr whose mode of being is thereby definitively inscribed.
In this logic, suffering lies “in her nature”—a convenient ontology that releases us from the question of how a shared life would have to be arranged to reduce this suffering. It is precisely that form of pity that Nietzsche mocked in Schopenhauer’s writings: a pity that fixes the other in her weakness rather than allowing her to invent herself.</p>

<p>The second, solidary path, on the other hand, must always preserve <em>contingency</em>.
Solidarity here means conceiving of the other as a space of possibilities and remaining conscious of the limits of one’s own descriptions.
It begins not with a judgment, but with the radically open question: “How can I help to reduce the cruelty that befalls you?”
This attitude acknowledges that there are no universally valid answers.
While in the public sphere we must fight for the conditions of a dignified life, the space of private <em>self-creation</em> must remain open—as that place where every person may draft their own image, beyond the gaze of others and their pitying diagnoses.</p>

<p>Yet we sometimes find it hard to hold a contingent evolutionary history responsible for our “place.”
Instead we often need a culprit who protects us from the recognition that the human being</p>

<blockquote>
  <p>hangs on the back of a tiger in dreams. – <a class="citation" href="#nietzsche:1873">(Nietzsche, 1873)</a></p>
</blockquote>

<p>Malice is never long in coming.
It surfaces when the other manages to no longer see their own vulnerability in the counterpart, cannot imagine how it might be like to be “the other”, or to believe they can no longer afford to be sensible, given social, material, bodily or psychological compulsion.</p>

<p>Some cruelty becomes visible, acquires a language, and other cruelty remains hidden, and when it cannot “heal” or express itself at all and thereby reflect on itself, it will have to prove itself—that is, it will (re-)produce the conditions for its own continued existence.
Without doubt about one’s own sensitivity to the pain and humiliation of others, curiosity about possible alternatives remains rude and calculating.</p>

<p>For Rorty, <em>solidarity</em> becomes the capacity to “see more and more” rather than seeing an “inner core.”
It is the capacity to count as “us” people (and other living beings) who are worlds apart from us.
It is grounded not, as in Kant, on universal Reason, nor, as in Christianity, on God, but on the striving to prevent and alleviate cruelty and pain.</p>

<blockquote>
  <p>But even when we use neither Kantian nor Christian language, we may still have the feeling that it is dubious to be more concerned about the living conditions of a fellow citizen of New York than about someone living equally hopelessly and miserably in the slums of Manila or Dakar. […] On the other hand, it is <em>not</em> incompatible with my position [(which accepts no essences)] to insist that we must try to include in our understanding of “we” also people whom we have so far counted among the “they.” This claim, characteristic of liberals who fear their own cruelty more than anything else, rests solely on the […] historical contingencies—namely, the development of the moral and political vocabularies typical of the secularized democratic societies of the Western world. – <a class="citation" href="#rorty:1989">(Rorty, 1989)</a></p>
</blockquote>

<p>Placing this “liberal axiom” above the sublime cannot be defended in any neutral, non-circular way.
But the same holds for Heidegger’s claim that the idea of “the greatest happiness for the greatest number” is also just a piece of metaphysics, a piece of <em>forgetfulness of Being</em>.</p>

<p>Rorty cannot claim that avoiding cruelty is universally good.
He cannot even claim that his assertion that there is no absolute truth is itself absolutely true.
It is the classic self-contradiction objection—whoever says “there is no absolute truth” seems thereby to claim the very thing they reject.
Rorty was of course aware of this.
His move consists in not playing the game at all. Rorty does not claim: “It is objectively true that there is no objective truth.”
That would be self-contradictory.
Instead he says something like: The concept of “objective truth” is no longer a useful tool.
We should stop using it—not because we have proved that there is no objective truth, but because a vocabulary without this concept gets us further.
Rorty does not position himself on a higher level and delivers a verdict about truth from there.
He recommends a different vocabulary. He says: Try it once without the concept of “absolute truth” and see whether you manage better.
That is not a thesis about the world, but a suggestion about how we ought to speak.</p>

<p>When someone in the seventeenth century stopped thinking in the categories of scholasticism and adopted instead the language of the new natural science, they did not prove that scholasticism was wrong.
They simply tried out a different vocabulary and found it more fruitful.
According to Rorty, the same applies to the concept of truth: he does not want to refute the existence of truth, but to suggest that without this concept we can conduct more interesting conversations.
Rorty accepts that his own position is just as <em>contingent</em> as any other.
He claims for it only that it is more useful—more useful for the project of reducing cruelty and expanding solidarity.</p>

<p>Critics such as Hilary Putnam or Thomas Nagel have responded that Rorty cannot really save himself here: if he says his vocabulary is “more useful,” that again contains a truth claim—namely that it really is more useful.
Rorty would reply: “useful” too is not an objective criterion but an evaluation from a particular perspective. The game can thus be continued indefinitely, and that is precisely Rorty’s point—there is no place at which one finally arrives. The philosophy that searches for such an endpoint pursues, according to Rorty, a goal that does not exist.
Whether one finds that convincing ultimately depends on whether one is willing to abandon the desire for ultimate justification.
Rorty demands that of his readers and he knew that many would not comply.</p>

<p>So, what would Rorty propose to us today? Certainly no grand promises or revolutions, but decidedly radical changes that in his time still sounded pragmatic.
Authors such as Han <a class="citation" href="#han:2021">(Han, 2021)</a> describe the achievement society as a place of <em>friendly violence</em> in which self-creation degenerates into self-optimization and the human being becomes mere <em>standing-reserve</em>—a concept that Han borrows from Heidegger’s critique of technology (cf. <a class="citation" href="#heidegger:1954">(Heidegger, 1954)</a>).
Here Rorty would react with skepticism, and not only toward the diagnoses, but above all toward the vocabulary.
He fundamentally distrusted Heidegger’s metaphysics of <em>forgetfulness of Being</em>: it runs the risk of <strong>condemning modernity as a whole</strong> rather than naming and addressing concrete grievances.
For Rorty this is a philosophical luxury neither progressives nor conservatives who actually wants to change something cannot afford.
This does not mean that Rorty would simply dismiss Han’s observations. The empirical description that people suffer from exhaustion, that the pressure of self-optimization destroys social solidarity, that depoliticization is a danger—these he would share.
But his answer would be pragmatic and reformist, not cultural-critical: improve working conditions and educational possibilities, guarantee social security, create spaces for purposeless exchanges, so that people once again have the leisure to redescribe themselves rather than optimizing themselves for the market.
These are <em>bread-and-butter questions</em> of politics, and it is precisely there, not in the <em>deep diagnosis of the occidental history of Being</em>, that Rorty would begin.</p>

<p>And since for Rorty progress is the capacity to allow one’s own final vocabulary to be expanded by that of the other, through literature and encounters (<em>sentimental education</em>), algorithms that isolate us in echo chambers destroy precisely this capacity for empathy.
He would probably reject digital surveillance as a new form of cruelty that prevents solidarity.</p>

<p>Futhermore, his hope for an ever-growing solidarity presupposes that the material conditions still permit this expansion at all.
Here, Carolin Amlinger reports:</p>

<blockquote>
  <p>the resentful feel fundamentally blocked in their progress as if ones life is overlaid by mud.</p>
</blockquote>

<p>That is why he would agree with Han’s observation that “depoliticization” is a danger, but he would call for speaking once again about “bread-and-butter issues” rather than philosophically condemning the whole of modernity.
Here, however, Rorty’s approach runs up against a limit that he himself did not see: his concept of solidarity remains <em>anthropocentric</em>—he asks who can suffer pain and humiliation and draws the circle of the moral exactly there.
<strong>The earth does not speak</strong>, so it does not feature in his account.
Yet Rorty seemed open to move beyond human solidarity.</p>

<blockquote>
  <p>Imagination, in the sense in which I am using the term, is not a distinctively human capacity. It is […] the ability to come up with socially useful novelties. This is an abilitiy Newton shared with certain eager and ingenious beavers. But giving and asking for reason <strong>is</strong> distinctively human, and in coextensive with rationality. The more an organism can get what it wants by persuasion rather than force, the more rational it is. – <a class="citation" href="#rorty:2016">(Rorty, 2016)</a></p>
</blockquote>

<p>Yet if we take Rorty’s own logic seriously—namely that the circle of solidarity must be drawn ever wider—then the question inevitably arises of whether this circle may stop at the human being.
A solidarity that applies only among speakers fails to hear the silence of those who have no language.
Perhaps that is the blind spot that costs us the most dearly today.</p>

<blockquote>
  <p>Our ancestors 500 years ago simply could not have grasped what now seems to us common sense that differences of religion, race, and social status are morally irrelevant. These ancestors weren’t blind to something we now see because there wasn’t yet anything for them to see. What we see had to be created in the interim. The human race has been busy creating itself over the last 500 years, creating moral obligations for itself which were once mere fantasies in the minds of a few people of unusually vivid imagination and unusually broad sympathy. I want to suggest that someday if this notion of humanity’s self-creation comes to replace the traditional philosophical notion of humanity’s self-understanding, the colleges and universities might be able to stop using even for commercial purposes the Platonic rhetoric of a quest for eternal truth. They might openly proclaim that their principle function is to keep society from ever being satified with itself, to keep individual students from being satified with themselves as they were when they arrived. If this happens democratic society might lose their habitual distrust of intellectuals. This would happen because […] whole societies would get intellectualized—not in the sense of being turned into a nation of philosophers, but in the sense of being turned into a nation of literary critics, that is, people curious about alternative forms of life and constantly sensitive to the possibility that they may be being as unconsciously cruel as their ancestors were. Such societies would still think of the education of the young as a matter of instilling traditional values. But the principle traditional values they have in mind would be simply the value of questioning traditions. – <a class="citation" href="#rorty:1990">(Rorty, 1990)</a></p>
</blockquote>

<hr />

<h2 id="cultivating-sensitivity">Cultivating Sensitivity</h2>

<p>The <em>urge toward new hardness</em>, toward strong men, clear images of the enemy, and national self-assertion, is not a sign of strength.
It is a sign that the circle of solidarity is shrinking.
Rorty would not have been surprised.
He already warned in <em>Achieving Our Country</em> <a class="citation" href="#rorty:1998">(Rorty, 1998)</a> that a left that retreats into cultural and academic distinction and stops speaking about material inequality prepares the ground for precisely those resentments that today call for hardness; that calls on their base to lean into every negative emotion and tells them that anything is exactly as they believe it is; that every perceived enemy, every perceived corruption, every perceived societal weakness is legitimized; that every frustration one has is legitimized and that any act in response is justified, perfectly displayed in the movie <em>Citizen Vigilante</em>.</p>

<p>Because of the political development of the past, we are in the unfortunate position that some of the moral progress correlates with an increase in material inequality such that the emontional path from correlation to causation is short and easy to take. This leads to a sort of alienation from society <a class="citation" href="#amlinger:2025">(Carolin Amlinger, 2025)</a>.</p>

<blockquote>
  <p>You could earn a lot of money <strong>and then</strong> feminists came in. – Interviewee</p>
</blockquote>

<p>The core promises of modern society, namely social integration and individual emancipation, seem no longer to function effectively.
Simultaneously, the modern individual experiences institutions as a restrictive barrier to their self-realization, thus society becomes itself a source of cruelty and humiliation.
The disappointed then find their voice not among those who advocate solidarity, but among those who name enemies.</p>

<blockquote>
  <p>Destructiveness is the result of an unlived life. – <a class="citation" href="#fromm:1973">(Fromm, 1973)</a></p>
</blockquote>

<p>What liberals can learn from Rorty is first of all an attitude of humility: those who stand up for liberal democracy should stop transfiguring it as the <em>genuinely superior</em>, the <em>historically necessary</em>, or the <em>rationally only possible</em>, for this rhetoric acts on those who do not share it as <em>humiliation</em>.
Rorty proposed defending liberalism not as truth but as habit—as a way of living together that we have laboriously acquired and that is worth continuing to narrate.
Liberal democracy does not need philosophical justification.
It is pragmatically justified because, for Rorty, it works better at reducing cruelty and increasing human freedom than anything else we’ve tried.
But <strong>if we cease to cultivate a desire for the reduction of cruelty, liberal democracy will inevitably vanish</strong>.</p>

<p>The resistance to new hardness cannot be primarily argumentative.
Those who want to draw smaller circles do not convince people via rational arguments but through emotions.
In fact, <strong>cruelty is the point</strong>; it is a necessity in a zero-sum game.
It is a desired feature of their politics to humiliate their perceived enemies even if their actions go against their base because it is cruelty that can be employed to get back what can never be owned: a nation, a race, a country, a culture, a planet, a superintelligence.</p>

<p>One does not refute resentments through better arguments; one changes them through better stories, that is, through narratives that give a face to people regarded as foreign, that show what it feels like to be excluded, persecuted, humiliated and by creating the conditions that everyone finds a niche in which they can self-realize themselves.
That is not a weakness of the liberal project but its actual strength: not persuasion through logic, but expansion of the imagination.</p>

<p>Third, and most uncomfortably, Rorty demands that liberals stop overlooking their own capacity for cruelty.
A dangerous, defensive hardness thrives among those who feel that their “final vocabulary”—their way of describing their life, giving dignity to their work, and finding meaning in their community—is being ridiculed by the architects of openness and tolerance.
When we demand that others abandon their worldview and adopt our enlightened descriptions, we are casually asking them to perform that same agonizing, unappreciated act of self-emptying. We expect them to dissolve their ego and identity to accommodate our vocabulary, while offering them no social value or grace for that immense sacrifice. And when they resist this spiritual violence, we respond with structural exclusion.</p>

<p>A solidarity that applies only among the “educated”, while blinding itself to the struggles of workers, is no solidarity at all. When we force our description of the world onto others, we risk committing the ultimate Rortyan sin: stripping them of their agency and ignoring their capacity to suffer. A racist acts with immense cruelty by denying the humanity of others. Yet, Rorty warns that when liberals respond by treating the racist as an inherently evil, subhuman monster, they slip into a parallel form of cruelty.</p>

<p>By excommunicating the racist from the horizon of valid human communication, liberals treat them as an irredeemable object rather than a poorly socialized human being. They turn their own vocabulary of tolerance into a new mechanism of humiliation—a way to make the other look permanently obsolete and futile. Rorty insists that liberals root out this structural sadism in themselves before pointing fingers.</p>

<p>Furthermore, Rorty reserves a distinct scorn for any kind of performative moral high ground.<sup id="fnref:3" role="doc-noteref"><a href="#fn:3" class="footnote" rel="footnote">2</a></sup>
He famously critiqued the “spectatorial left” for transforming politics into a theater of cultural purity testing, where naming vices and maintaining a flawless vocabulary replaces the <strong>heavy lifting of material reform</strong> <a class="citation" href="#rorty:1998">(Rorty, 1998)</a>.
For Rorty, a moral posture that exists only to signal its own righteousness does nothing to alleviate pain; instead, it becomes a weapon of exclusion, mocking the very workers and poorly socialized individuals it claims a democratic society should integrate.</p>

<p>No absolute moral law resolves this <em>paradox of tolerance</em>.
It cannot “solve” Gaza, the West Bank or mass killings in Sudan. There are only <em>pragmatic</em> choices.
Arguing over theoretical definitions of “genocide” or “apartheid” does nothing to stop a bomb from falling or a family from being displaced because <strong>laws are only binding if people feel like obeying them</strong>.<sup id="fnref:4" role="doc-noteref"><a href="#fn:4" class="footnote" rel="footnote">3</a></sup>
Legal documents themselves are just pieces of paper unless they are backed by a <em>human rights culture</em>.
And in a <em>liberal culture</em>, no political goal should justify the torture or starvation of a population.
We must confidently assert that our tribe’s way of living, that is, one that refuses to tolerate racism, is simply better at minimizing cruelty, and we will therefore enforce it through political power.
We can recognize that the racist is a human being capable of suffering without letting that recognition paralyze our defense of a democratic society.
Thus, the West has a pragmatic duty to use its political and economic power to enforce a settlement to stop the exertion of cruelty.</p>

<p>Yet Rorty knew how easily this pragmatic enforcement of power can warp into state-sponsored paranoia. His immediate, visceral opposition to the post-9/11 “War on Terror” and the invasion of Iraq was driven by this exact anxiety.
He recognized that the moment a democracy prioritizes absolute safety over civil liberties, it surrenders its moral authority.
He viewed the weaponization of fear after 9/11 as a disastrous distraction from building a more just, equal, and inclusive society. 
And he called this manifestation “simple-minded militaristic chauvinism”. 
Trading freedom for security was a cynical political maneuver that exported violence abroad while bleeding away the resources needed to fix internal failures like poverty and systemic racism.
A society terrified of an invisible enemy inevitably stops listening to its citizens and starts listening to a “strongman”.</p>

<p>In short: Today, Rorty would not call on us to save democracy by philosophically refuting its enemies.
He would call on us to narrate it—again and again, in new words, for people who do not yet recognize themselves in it.
But this requires that our actions converges to the image we have of ourselves.
Instead of giving up on <strong>cultivating the art of solidarity</strong>, Rorty would call on us to reconfigure our social reality in such a way that it fits our description of a liberal democracy; to go so far to say: it is more important to be sensible to cruelty than to find “the Truth”.</p>

<blockquote>
  <p>Orwell’s main concern is to sensitize an audience to cases of cruelty and humiliation which they had not noticed. [… Moral progress is] a matter of separating the question “Do you believe and desire what I believe and desire?”—a representational question—from the question “Are you suffering?” – <a class="citation" href="#rorty:1989">(Rorty, 1989)</a></p>
</blockquote>

<p>But that requires time.
Not the time of rapid consumption, but the slow time of reading.
It requires that patience which allows one to immerse oneself in an unfamiliar form of life, to understand it from within, before evaluating it.
Rorty knew that solidarity does not arise in seconds.
It grows in the hours spent with <em>Raskolnikov</em>, <em>Dorothea Brooke</em>, or the <em>Invisible Man</em>—with people one will never meet and whom one nonetheless begins to know.</p>

<p>Here, as a new communication medium participating in communication for the first time (cf. <a class="citation" href="#esposito:2022">(Esposito, 2022)</a>), <em>(Generative) AI</em> transforms precisely this process, and in a way that would have troubled Rorty.
Not because the technology is evil, but because it accelerates, compresses, and smooths encounters with the unfamiliar.
What a novel laboriously builds over two hundred pages—for example, trust in an unfamiliar voice, the capacity to endure contradictions, the slow understanding of another world—an LLM can summarize in a few paragraphs.
But whether the expansion of the imagination that Rorty had in mind still arises in the process is questionable.
Summaries do not sensitize; they inform.
And the difference between the two is, for Rorty, the <strong>difference between knowledge and solidarity</strong>.</p>

<p>This is not a condemnation of technology, but a reminder of what is at stake. Rorty would not call on us to put away the smartphone forever or to ban AI.
He would call on us to ask ourselves: When did we last read a book that truly disturbed us, and whose story was it?
When were we last willing to be truly disturbed by an unfamiliar life?</p>

<p>Our historical pathways to self-realization are no longer sustainable. It is increasingly obvious that material progress in the West will slow down or perhaps even come to a halt. Although the transition toward renewable energy has only just begun, and many still deny the necessity of it, the pervasive feeling that the era of endless growth is over continues to spread. While self-realization was previously achieved through a surplus of wealth, a new, likely subconscious ideology is emerging: that one can only achieve self-realization by denying others the means to do the same.
The idea of the state then changes from an entity that is protecting its citizen from cruelty to a distributor and enabler of cruelty such that everyone, especially the self-proclaimed “silenced majority”, “get what they deserve” <a class="citation" href="#amlinger:2025">(Carolin Amlinger, 2025)</a>.
This zero-sum battle over distribution will inevitably intensify unless we develop alternative frameworks for personal fulfillment.</p>

<p>When I look at my little nieces, I hope that—despite the impending catastrophes as a combination of a war against our political, scientific as well as environmental ecosystem—they will live in a culture that has managed to sensitize itself to cruelties I cannot yet see.
I hope they live in a culture that has once again made it its goal to be intellectual in the Rortyan sense: a culture that revitalizes the old ideal of “poets and thinkers”—not by hunting for a final, objective truth, but by using the poet’s imagination to reshape our world into something more humane.
I hope that, if they fail, they <em>fail in dignity</em> and <em>self-respect</em> <a class="citation" href="#metzinger:2023">(Metzinger, 2023)</a>.
And I hope that they will be merciful toward me, understanding that it was still impossible for me to see so much more, because my language was still bound by the limits of my time.
I hope the circle of their imagination, in which rationality plays its game, will be larger and not smaller; that they are courageous enough to constantly redescribe themselves and their community, until the “we” of their solidarity reaches far beyond what I am capable of imagining today.</p>

<h2 id="references">References</h2>

<ol class="bibliography"><li><span id="kohlenberger:2024">Kohlenberger, J. (2024). <i>Gegen die neue Härte</i> (p. 256). dtv Verlagsgesellschaft GmbH &amp; Co. KG.</span></li>
<li><span id="amlinger:2025">Carolin Amlinger, O. N. (2025). <i>Zerstörungslust: Elemente des demokratischen Faschismus</i>. Suhrkamp.</span></li>
<li><span id="redecker:2026">von Redecker, E. (2026). <i>Dieser Drang nach Härte</i> (p. 272). S. FISCHER.</span></li>
<li><span id="baudrillard:1976">Baudrillard, J. (1976). <i>Symbolic Exchange and Death</i>. SAGE Publications Ltd. https://doi.org/10.4135/9781526401496</span></li>
<li><span id="herzog:2007">Herzog, W. (2007). <i>Begegnungen am Ende der Welt</i>. https://www.youtube.com/watch?v=uBk9lLFWGcI</span></li>
<li><span id="gundlach:2026">Gundlach, M. (2026). Dieser Pinguin glaubt noch an Europa. <i>Süddeutsche Zeitung</i>. https://www.sueddeutsche.de/medien/pinguin-europa-werner-herzog-dokumentation-johann-wadephul-li.3378577</span></li>
<li><span id="gebru:2024">Gebru, T., &amp; Torres, E. P. (2024). The TESCREAL bundle: Eugenics and the promise of utopia through artificial general intelligence. <i>First Monday</i>, <i>29</i>(4). https://doi.org/10.5210/fm.v29i4.13636</span></li>
<li><span id="muehlhoff:2025">Mühlhoff, R. (2025). <i>Künstliche Intelligenz und der neue Faschismus</i>. Recalm.</span></li>
<li><span id="rosengruen:2022">Rosengrün, S. (2022). Why AI is a threat to the rule of law. <i>Digital Society</i>, <i>1</i>(2). https://doi.org/10.1007/s44206-022-00011-5</span></li>
<li><span id="rorty:1990">Rorty, R. (1990). <i>Ethics of Principle vs Sensitivity</i>. https://www.youtube.com/watch?v=nD248K11zNE</span></li>
<li><span id="stiegler:2019">Stiegler, B. (2019). <i>The Age of Disruption</i>. Polity Press.</span></li>
<li><span id="rorty:1989">Rorty, R. (1989). <i>Contingency, Irony, and Solidarity</i>. Cambridge University Press.</span></li>
<li><span id="rorty:1979">Rorty, R. (1979). <i>Philosophy and the Mirror of Nature</i>. Princeton University Press.</span></li>
<li><span id="rorty:2016">Rorty, R. (2016). <i>Philosophy as Poetry</i>. University of Virginia Press.</span></li>
<li><span id="maturana:1991">Maturana, H. (1991). <i>Humberto Maturana Melbourne Seminar</i>.</span></li>
<li><span id="maturana:1988">Maturana, H. R. (1988). Reality: The Search for Objectivity or the Quest for a Compelling Argument. <i>The Irish Journal of Psychology</i>, <i>9</i>(1), 25–82. https://doi.org/10.1080/03033910.1988.10557705</span></li>
<li><span id="luhmann:1984">Luhmann, N. (1984). <i>Soziale Systeme: Grundriß einer allgemeinen Theorie</i>. Suhrkamp.</span></li>
<li><span id="heidegger:1976">Heidegger, M. (1976). “Nur noch ein Gott kann uns retten.” <i>Der Spiegel</i>, <i>30</i>(23), 193–219.</span></li>
<li><span id="rorty:1999">Rorty, R. (1999). <i>Philosophy and Social Hope</i>. Penguin Books.</span></li>
<li><span id="nietzsche:1873">Nietzsche, F. (1873). Über Wahrheit und Lüge im außermoralischen Sinne. In G. Colli &amp; M. Montinari (Eds.), <i>Kritische Studienausgabe (KSA)</i> (Vol. 1, pp. 873–890). de Gruyter.</span></li>
<li><span id="han:2021">Han, B.-C. (2021). <i>Infokratie: Digitalisierung und die Krise der Demokratie</i>. Matthes &amp; Seitz.</span></li>
<li><span id="heidegger:1954">Heidegger, M. (1954). <i>Die Frage nach der Technik</i>.</span></li>
<li><span id="rorty:1998">Rorty, R. (1998). <i>Achieving Our Country: Leftist Thought in Twentieth-Century America</i>. Harvard University Press.</span></li>
<li><span id="fromm:1973">Fromm, E. (1973). <i>The Anatomy of Human Destructiveness</i>. Holt, Rinehart and Winston.</span></li>
<li><span id="esposito:2022">Esposito, E. (2022). <i>Artificial Communication</i>. The MIT Press. https://doi.org/10.7551/mitpress/14189.001.0001</span></li>
<li><span id="metzinger:2023">Metzinger, T. (2023). <i>Bewusstseinskultur: Spiritualität, intellektuelle Redlichkeit und die planetare Krise</i> (p. 208). Berlin Verlag.</span></li></ol>

<!--
Interessanterweise sehr nah an: https://www.youtube.com/watch?v=yM_or9mYaXM
-->
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:1" role="doc-endnote">
      <p>Today, there is substantial evidence that prelinguistic infants display sophisticated reasoning, numerical intuition, and causal understanding well before acquiring language. Great apes, corvids, dolphins, elephants demonstrate planning, tool use, and social cognition without anything resembling human language. The key move is to distinguish between basic cognition (perception, spatial navigation, pattern recognition, which animals and infants clearly have) and discursive, propositional, reason-giving thought, i.e., the kind philosophers actually argue about. Rorty could concede the former while maintaining that the latter is constitutively linguistic and social. Not: “You can’t think without language” but “The kind of thinking that involves justifying beliefs, weighing reasons, making normative claims, that is through and through a linguistic practice.” <a href="#fnref:1" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:3" role="doc-endnote">
      <p>This text can be seen as performative. <a href="#fnref:3" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:4" role="doc-endnote">
      <p>Arguing matters if these terms are legally binding—but not for the reason most lawyers or human rights activists think it does since international law is a collection of linguistic mechanisms invented by humans to achieve a specific goal: the reduction of cruelty. Thus, in Rorty’s eyes, arguing about the law is useful only if it is treated as a tactical weapon to reduce physical pain, not as a philosophical debate to prove your moral superiority. But if we spend months debating the precise semantic definition of “genocide” or “proportionality” in academic journals while people are actively starving, the law has ceased to be a useful tool. <a href="#fnref:4" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content><author><name>Benedikt Zönnchen</name><email>benedikt.zoennchen@web.de</email></author><category term="Opinion" /><category term="Philosophy" /><summary type="html"><![CDATA[A New Cruelty — On the Longing for Certainty]]></summary></entry><entry><title type="html">Hurricanes May Be Dangerous but Not Evil</title><link href="https://bzoennchen.github.io/Pages/2025/09/22/rant-oversimplification.html" rel="alternate" type="text/html" title="Hurricanes May Be Dangerous but Not Evil" /><published>2025-09-22T00:00:00+02:00</published><updated>2025-09-22T00:00:00+02:00</updated><id>https://bzoennchen.github.io/Pages/2025/09/22/rant-oversimplification</id><content type="html" xml:base="https://bzoennchen.github.io/Pages/2025/09/22/rant-oversimplification.html"><![CDATA[<p>The word <em>digital</em> usually refers to something represented, processed, or communicated using discrete values rather than continuous signals.
We know that digital systems encode information in binary (0/1) rather than analog waveforms. A digital clock displays exact numbers; an analog clock sweeps smoothly. Digital data is stored and transmitted as discrete units (bits and bytes), not continuous signals.</p>

<p>But we often forget: <em>signs</em> and <em>symbols</em>—the words of our language—are also digital. They, too, compress complexity into discrete units. Thus human language is a digital system of signs, even when carried by analog sound waves.</p>

<p>While this “feature” of language, i.e. the compression of the environment, is a necessary condition for communication, it implies that communication is loss-prone. Furthermore, there are no ultimate foundations or essences in language, no definite <em>ground</em> beneath our signs, even though we act as if there were.</p>

<p>Programming seems to offer firmer ground. Classes, objects, and definitions feel well-formed, precise, and self-contained. There is a clear separation between syntax and semantics. Yet even here, meaning is not given by the machine, because a machine destroys the reference horizon. In other words, machines do not recognize complexity, because for them, there are no more possible environmental references than those currently being actualized. A machine can “understand” something only in one way, and thus cannot understand it at all. Meaning arises because I, as a programmer (psychic system), interpret and use the system. The code works or fails, delights or frustrates, earns money, sells goods, plays music, organizes life. Meaning is not in the code itself but in the observed, expected and assumed effects it produces and the uses we make of it.</p>

<p>But our interpretations are not unique. We have to select from the reference horizon that we build and imagine together. This meaning is mediated by natural language—so we are already outside the pure formal system and also limited by our cultural memory.</p>

<p>The reference horizon is contingent but not random. When we listen to music, we anticipate the next tone based on what we have already heard and what we are used the hear. Similarly, the past gives us the structure to see the future, but it also conceals other ways to think, imagine, and live. We need the past, culture, traditions as memory function to be able to anticipate and realize the next step—our future.</p>

<p>Now step further outward into everyday life: the ambiguity multiplies. What does it mean to call someone a “mother”? Is there an essence of “motherhood”? I would argue there is not. Instead, there is a shifting pattern of experiences: shelter, protection, food, conflict, laughter, school lessons, beginnings, and countless other associations. A mother is not a father, not a flower, not a lake—but a fluid constellation of relations and meanings. It is also a <em>trace</em>. But this does not imply arbitrariness.</p>

<p>Similarly, we can also ask the nowadays politically charged question: <em>What is a woman?</em>—and immediately run into difficulties. Yet, if we think about it, this is hardly surprising.
How could language—a digital system of signs—ever objectively capture the infinite complexity of our realities that constantly drift in the ocean of evolution? The word <em>“woman”</em> only gains meaning in relation to the usage of terms like <em>man</em>, <em>not-man</em>, <em>mother</em>, or <em>feminine</em>. Its meaning is never fixed; it is always shifting, unstable, and deferred. To ask “What is a woman?” already risks reinforcing the <em>man/woman binary</em> as <em>natural</em>, when in fact that binary is itself the product of language, institutions, and history. It can and will change.
Any attempt to pin down a definition will necessarily exclude certain possibilities and identities.</p>

<p>In this spirit, communication functions not through mutual understanding, but rather through the <strong>absence of misunderstanding</strong>. It persists as long as connectivity remains—that is, as long as one contribution can trigger another to continue the discourse. Because meaning is an internal construct of our psychic systems, it cannot be transmitted; it is this deep operational separation that allowed society to emerge. Despite our inability to inhabit one another’s thoughts, we have learned to organize and cooperate. This operational closure—and the inherent “loneliness” it implies—is the source of many of our troubles, yet it is also the foundation of our autonomy and, likely, the very reason we possess a sense of self.</p>

<p>Outrage arises because different groups want their preferred vocabulary to dominate—for example, biological essentialists versus trans-inclusive definitions. In truth, the conflict is a clash of competing <em>language games</em>. Each side feels its way of speaking is under threat, which explains the emotional intensity.</p>

<p>As Richard Rorty once observed:</p>

<blockquote>
  <p>There is something potentially very cruel about the claim that [the language people speak is, for ironists, a game of chance]. For the most effective way to inflict lasting pain on people is to humiliate them by making everything that had seemed especially important to them appear futile, outdated, and powerless. – <a class="citation" href="#rorty:1989">(Rorty, 1989)</a></p>
</blockquote>

<p>This is why the demand for a single, definitive answer generates tension: it denies the play of difference that actually structures meaning. The difficulty of the question is not incidental—it is intrinsic. <em>Woman</em> is an aporia, a site of endless contestation. Outrage arises precisely because the question both demands and resists resolution.</p>

<p>The point, then, is not to settle the metaphysical question but to enjoy ambiguity and to foster <em>solidarity</em>. That means choosing the description of woman that best promotes human flourishing and reduces harm and is open to interpretation. Rather than offering a final definition, our task is to help society continually renegotiate its self-descriptions.</p>

<p>We can also ask more generally: <em>What is a human being?</em> Here too, the problems multiply if we think in terms of essences.</p>

<p>A person is not a fixed substance but a pattern of modulated repetition. Life seems coherent over time because of habits, recurring desires, aspirations, and flaws. But a person is not a stable essence—it is a process, an event, a “happening”. Like a hurricane, a person is a complex system: dynamic, shifting, and contingent. And this observation, of course, is made by yet another hurricane, observed by still others—each shaping and shaped by the rest.</p>

<p>This is precisely why I find contemporary reporting so taxing. It systematically ignores this inherent complexity. While communication admittedly requires compression, must it always collapse into reductive binaries—right vs. wrong, friend vs. foe? Must every event and every individual be flattened into a single label just to fit a pre-existing semiotic network? Must every person have an opinion on any matter?</p>

<blockquote>
  <p>Everything must be explained, everything should be understood, and if something cannot be understood, it counts as nothing. [… So believe] many people, who are constantly having everything explained to them and are presented with a world without secrets, without the inexplicable or the overly complex, eventually come to believe themselves that they understand everything. – <a class="citation" href="#bauer:2018">(Bauer, 2018)</a></p>
</blockquote>

<p>Too often, complexity is reduced to a “profile”—a checklist of traits or, worse, a binary moral judgment. But a hurricane is neither good nor evil; it is a phenomenon. It destroys, it endangers, and it compels us to react. If we wish to address such forces, we cannot moralize them. We must instead understand—and perhaps alter—the structural conditions under which they form and transform.</p>

<p>People, too, resist moral simplification. A person can be both a criminal and deeply kind; they can support terror while enduring horrific cruelty; they can be a brilliant poet and a member of the Nazi party. I choose to acknowledge these realities with an “<strong>and</strong>” rather than a “<strong>but</strong>”.</p>

<p>There is no escape from this complexity—only further layers of it. Yet, it feels increasingly difficult to transcend our digital condition: a world structured by discrete representations, binary code, and rigid networks. In such an environment, the richness of lived experience is flattened into exchangeable symbols—tokens that are used and abused to force a sense of order, to draw hasty conclusions, and to conjure yet another storm.</p>

<!--
I know it is not possible to report without compression. I am compressing right now. But can't we do better? Must it always be only right and wrong, good and evil, friend and foe? Must every event and every person be collapsed into a single dot that fits neatly into our networks of signs?
 -->

<h2 id="references">References</h2>

<ol class="bibliography"><li><span id="rorty:1989">Rorty, R. (1989). <i>Contingency, Irony, and Solidarity</i>. Cambridge University Press.</span></li>
<li><span id="bauer:2018">Bauer, T. (2018). <i>Die Vereindeutigung der Welt</i> (p. 104). Reclam.</span></li></ol>]]></content><author><name>Benedikt Zönnchen</name><email>benedikt.zoennchen@web.de</email></author><category term="Opinion" /><summary type="html"><![CDATA[The word digital usually refers to something represented, processed, or communicated using discrete values rather than continuous signals. We know that digital systems encode information in binary (0/1) rather than analog waveforms. A digital clock displays exact numbers; an analog clock sweeps smoothly. Digital data is stored and transmitted as discrete units (bits and bytes), not continuous signals.]]></summary></entry><entry><title type="html">The Free Energy Principle</title><link href="https://bzoennchen.github.io/Pages/2025/08/04/free-energy-principle.html" rel="alternate" type="text/html" title="The Free Energy Principle" /><published>2025-08-04T00:00:00+02:00</published><updated>2025-08-04T00:00:00+02:00</updated><id>https://bzoennchen.github.io/Pages/2025/08/04/free-energy-principle</id><content type="html" xml:base="https://bzoennchen.github.io/Pages/2025/08/04/free-energy-principle.html"><![CDATA[<p>Imagine you see a grizzly bear in the woods and half of its face is hidden behind a tree.
You will most certainly recognize the bear as whole and dangerous animal.
Your brain will “fill in the gaps” even though there is no absolute certainty that the bear isn’t split in half.
How and why does the brain do this?
Furthermore, why are we fooled by all sorts of graphical illusions?
In other words, why are we, in some instances, so stubborn to see what is not there—even if we are told that it is not there—and, on other occasions, we can immediately see what is hidden and probably there?</p>

<p>The so‑called <em>free energy principle</em> (FEP) <a class="citation" href="#friston:2006">(Friston et al., 2006)</a> is a neat mathematical principle that offers answers to these questions.
According to the FEP, brains or nervous systems make use of <em>variational inference</em> on hidden causes of sensory data.
In this view, the goal of an organism, including ours, is to <strong>minimize surprise</strong> within their respective environment.
Strictly, FEP says systems minimize an <em>upper bound</em> on surprise—variational free energy—because exact surprise is intractable
Researchers working towards finding evidence to apply the FEP claim that perceptual processes are just one aspect of emergent behaviours of systems that conform to a <em>free energy principle</em>.
Thus, it is a principle on which the brain might operate.</p>

<blockquote>
  <p>The free energy considered here measures the difference between the probability distribution of environmental quantities that act on the system and an arbitrary distribution encoded by its configuration.
The system can minimise free energy by changing its configuration to affect the way it samples the environment or change the distribution it encodes. These changes correspond to action and perception respectively and lead to an adaptive exchange with the environment that is characteristic of biological systems.
This treatment assumes that the system’s state and structure encode an implicit and probabilistic model of the environment. – <a class="citation" href="#friston:2006">(Friston et al., 2006)</a></p>
</blockquote>

<p>As a consequence, the FEP departures from the classical reward-based explanations of behaviour and evolutional progress in general.
Instead of starting with an <em>a priori</em> <em>goal</em> or <em>reward</em> that the organism has to find, it starts with <em>what kinds of states the organism expects to be in</em>, and assumes the organism to act in a way that <em>confirms those expectations</em> and keeps it in familiar, low-surprise situations.</p>

<p>Friston and others also believe that the use of <em>hierarchical models</em>, such as layered neural networks, enable the brain to construct prior expectations in a dynamic and context-sensitive fashion.
Following this line, this scheme provides a principled way to understand many aspects of cortical organisation and responses while it is strongly connected to the <em>connectionist</em> strand of <em>artificial intelligence</em>.</p>

<p>It was Friston who populated the idea of the brain as a <em>prediction machine</em> trying to minimize surprise (or prediction error) by either updating beliefs (<strong>perception</strong>) or by acting on the world (<strong>action</strong>) to make it more predictable.
He provides a unifying framework for perception, action, attention, learning, and even consciousness which, in my opinion, combines enactivists and constructivists views (as we will discuss later).</p>

<p>However, the term <em>prediction machine</em> should not be confused with a depreciation of organisms or the human brain.
There is nothing magical or mystical that we will lose if we stick to the naturalization of ourselves as long as we do not mistaken an explanation as final answer.
Existential questions remain but stay outside the public sphere of science but those questions are necessarily informed by our (scientific) knowledge about ourselves and our environment.
In accordance to the FEP, we drift towards explanations of our being—we drift towards self-creation and our own evidence.
Therefore, what could be more magical, more fascinating than our contingent existence that emerged through a process of exactly such <em>self-creation</em> and <em>self-evidencing</em>.</p>

<p>While the FEP gained major attention in neuroscience and cognitive science through the work of Karl Friston in the 2000s, it is mathematically rooted in ideas that are much older—originating in statistical physics, Bayesian inference, and information theory.</p>

<h2 id="metaphysical-assumptions">Metaphysical Assumptions</h2>

<p>To start, we need some metaphysical assumptions, which I summarize as <em>physicalism</em> (ontology) and <em>indirect realism</em> (epistemology).</p>

<p>Physicalists assume that the world is made entirely of physical stuff (matter, energy, fields, etc.) and that all phenomena, including mental events (thoughts, feelings, consciousness), can, in principle, be explained by physical processes and physical laws.
There is no need to appeal to non‑physical substances (like souls or spirits) to explain reality.</p>

<p>While in practice, scientists rarely state their metaphysical position explicitly, most of them are physicalists or at least act in their field as if they are physicalists.
Still, physicalism remains a philosophical stance, not a requirement for doing science.</p>

<p>Second, indirect realism (also called <em>representationalism</em>) is a theory in the philosophy of perception about how we experience the world.
It states that we do not perceive the external world directly.
Instead, we are directly aware only of mental representations (sometimes called sense data or percepts), and these representations are caused by and (usually) resemble external physical objects or events.
There is a physical world “out there,” and physical objects and processes emit or reflect light, sound, and other signals.
Our sense organs (eyes, ears, etc.) pick up these signals and send information to the brain.
The brain processes this information and creates an internal representation or image of the world.
Therefore, what we are immediately aware of is this internal image, not the external object itself.
From these internal representations, we indirectly know about the external world.</p>

<p>From a physicalist perspective, the “sense data” or perceptions that indirect realism talks about would be physical phenomena generated by the brain’s interaction with the environment. 
The objects in the external world exist independently of our perception (as physicalism holds), but we never experience them directly—only through mediated sensory data.
But there is also a tension: If one believes that all experiences and perceptions are entirely physical processes, there might be questions about how non-physical qualities (like subjective experience or <em>qualia</em>) fit into the picture.
Physicalism typically aims to reduce all phenomena, including consciousness, to physical states.
So, there could be questions about whether “sense data” are truly representations in a way that indirect realism requires or whether they are just parts of the physical process.
Moreover, if a physicalist were to take an extreme reductionist view (such as <em>eliminative materialism</em>), they might reject any distinction between “sense data” and the brain’s processing of physical stimuli, challenging the need for a separate representational layer of perception.
However, this is a minority view in physicalism.</p>

<p>Later I will problematize indirect realism and dwell on the question of whether the free energy principle really requires it.
The list of critics is long, ranging from <em>phenomenologists</em> (Husserl, Heidegger, Merleau‑Ponty), <em>naive realists</em> (McDowell, Brewer, Searle, Reid), <em>new realists</em> (Gabriel, Holt), <em>idealists</em> (Kant), <em>empiricists</em> (Hume), and other <em>biological constructivists</em> (Maturana, Varela).
To move on, I will first assume indirect realism for a moment because it makes the explanation intuitive (but if we think about it a little longer, we run into problems).</p>

<p>From an ontological standpoint Friston clearly assumes an external environment that pushes organisms into some direction when he writes:</p>

<blockquote>
  <p>[I]nvoking selectionist arguments; those systems that match their internal structure to the external causal structure of the environment in which they are immersed will be able to minimise their free energy more effectively. – <a class="citation" href="#friston:2006">(Friston et al., 2006)</a></p>
</blockquote>

<p>One might ask (later) if an environment of a system can be regarded in separation of that system.</p>

<h2 id="contingent-complexity">Contingent Complexity</h2>

<p>As we will see, the <em>free energy principle</em> suggests that our nervous system is actively generating predictions about what should be “out there,” and then uses sensory input merely to check whether those predictions are correct.
In that sense, assuming the free energy principle makes our brains into prediction machines with the goal of <strong>minimizing surprise</strong>.</p>

<p>Our body’s purpose is to stay alive—to continue with its self‑creation by constructing its own organization and structure (autopoiesis).
It is open to structure but organizationally or operationally closed (implying a causal loop in its operations).
Its <em>qualitative identity</em> can change but its <em>individual identity</em> is maintained constant even if it could change, meaning there is some active prevention going on—we do not disintegrate:</p>

<blockquote>
  <p>[A molecular autopoietic system (or living being)] is a homeostatic system of chemical production which has its own individual identity as the variable which it maintains constant – <a class="citation" href="#barry:2012">(Razeto-Barry, 2012)</a></p>
</blockquote>

<p>Some argue that autopoiesis is an intrinsic purpose of an organism from which all other goals derive <a class="citation" href="#virenque:2024">(Virenque, 2024; Barandiaran, 2017)</a>.</p>

<blockquote>
  <p>Importantly, sense-making is intentional because it implies an organismic perspective from which meaning is brought forth. Because of the organism’s precariousness, certain interactions with the environment are either positive or negative from the perspective of the organism. Consider a bacterium swimming up a gradient of glucose. The glucose is meaningful for the bacterium because it provides nutrients that are necessary for maintaining its metabolic processes (i.e., its autopoiesis). The meaning of the glucose gradient only makes sense from the perspective of the bacterium. Take away the bacterium from the situation and the glucose stops having its meaning. What brings forth such meaning is the organism’s adaptive and autopoietic organisation. Such bringing forth of meaning is what enactivists call sense-making. – <a class="citation" href="#bogota:2024">(Bogotá, 2024)</a> (example taken from <a class="citation" href="#varela:1997">(Varela, 1997)</a>)</p>
</blockquote>

<p>Others suspect that we, as external observers, project purposiveness onto the organism <a class="citation" href="#cummins:2014">(Cummins, 2014)</a> and anti-naturalists like Markus Gabriel see meaning as ontological primary <a class="citation" href="#gabriel:2018">(Gabriel, 2018)</a>.
For the FEP, it is not important if purposiveness is a projection or intrinsically given (if meaning is ontologically primary the whole story changes).</p>

<p>What is quite certainly true is that organisms have to deal with perturbations to stay alive in their respective environment (which is determined by the <em>couplings</em> between the organism and its environment).
Only surviving organisms can be observerd, which leads to the belief—backed by observation from an external observer—that organisms strive for survival (of the fittest) when, in fact, it might only look like this is the case.</p>

<p>The complexity of the environment and the complexity of the organism co‑evolve over time, and some organisms—like humans—evolved in a direction of higher complexity.
Again, this does not mean that an increase in complexity is a pre-subscribed <em>telos</em>.
It is all <em>contingent</em>.</p>

<p>As a consequence of an increase in complexity (for some organisms), their environments become noisy and ambiguous, which requires more sophisticated feedback mechanisms to be able to act and survive in such environments <a class="citation" href="#tomasello:2024">(Tomasello, 2024)</a>.
For example, we experience a sort of conscious self-control when we asked to tell the color (not the written text) of</p>

<p><span style="color:red">BLUE!</span></p>

<p>We wanna say “blue” but immediately interrupt ourselves to say “red”.
We can be angry with ourselves, talking to ourselves and judging ourselves as if there is an instance within us that is distinct from us.</p>

<p>Recognizing a predator in the forest already requires more than just pattern matching.
Simple template‑matching reflexes may not work.
In such complex environments, organisms have to temper their reflexes, plan, and evaluate different action sequences before actualizing one action plan.
But a much more complex situation arrives if we have to deal with social systems that are beyond our control but which also provide us with the needs for survival.</p>

<p>Tomasello points out that these abilities result from nature’s inability to build organisms that can cope with everything (reactively) that nature throws at them.
Instead, nature can build psychological actors:</p>

<blockquote>
  <p>It brings forth organisms that function as feedback control systems pursuing goals, selecting profound actions and monitoring the process of acting them out. – <a class="citation" href="#tomasello:2024">(Tomasello, 2024)</a></p>
</blockquote>

<p>He also reminds us that the importance of acting is not defined of how <em>many</em> things an organism can do but <em>how</em> these actions are realized.
Maybe we can learn from his emphasis when discussing AI systems where the focus is often on the capabilities, i.e. on the “how many things” AI systems can do (better) neglecting the question of <em>(self-)control</em> and <em>agency</em>.
Again, for Tomasello the importance is not complexity or variety (of actions) but the levels of control an organism can exert on its environment and itself.</p>

<p>With respect to the free energy principle contingency and unpredicability explodes if the environment of organisms consists of other organisms—when one has to predict the predictions of others.
Social priors provided by caregivers or joint attention and communication become necessary and people might update their internal models by minimizing surprise via social feedback.
We have acted on the whole earth to make our environment more predictable which led to new unpredictablities and risks that we try to make predictable.</p>

<h2 id="naturalizing-surprise-and-entropy">Naturalizing Surprise and Entropy</h2>

<p>Since the free energy principle (FEP) is all about minimizing surprise, let’s start by pinning down what mathematicians and computer scientists actually mean by “surprise”—how can we naturalize it?
Sure, we all have a gut feeling about it—something surprising is just something we didn’t see coming.
More formally, an event is surprising if it’s unlikely based on what we already believe.
That is, we were pretty confident it wouldn’t happen, and yet—bam—it did.</p>

<p>Take raining frogs, for instance.
That would definitely raise some eyebrows.
But even without biblical weather, everyday life has its surprises.
Imagine rolling a die 10 times and getting a 1 every single time. 
Highly unlikely, right? 
That’s the kind of statistical oddity that makes us do a double-take.</p>

<p>Surprise, in this framework, can come from two main sources:</p>

<ol>
  <li>The event itself—specifically, how unpredictable (or high in entropy) the event is.</li>
  <li>Our prior beliefs—how confident we were about what <em>should</em> happen.</li>
</ol>

<p>Now, even if your beliefs are spot on—say, you assume the die is fair, and it actually is fair—you might still see 10 ones in a row.
That’s not because your beliefs are wrong, it’s just because randomness likes to keep things interesting.
In this case, the surprise comes not from flawed beliefs, but from sheer bad luck.</p>

<p><strong>Remark:</strong> In the <em>Bayesian</em> world, probability isn’t about how often something actually happens—it’s about what we believe will happen, given what we know—our <strong>degree of belief</strong>.
More precisely, it’s a measure of uncertainty or confidence in a particular outcome or parameter, based on the information we currently have.
This is quite different from the <em>frequentist</em> view, where probability is all about long-run frequencies.
A frequentist might say, “If we rolled this die an infinite number of times, the proportion of ones would settle at 1/6”.
That’s the idea: probabilities reflect what would happen over countless repetitions of the same experiment.
But often events cannot be repeated.
Here a Bayesian offers a more flexible, belief-based approach: “Before seeing any data, I assume all outcomes are equally likely (a uniform prior). But if I observe 200 ones in 1000 rolls, I’m updating my belief—maybe this die has a bias”.
The more data we get, the more refined our beliefs become.
This belief-updating process is powered by Bayes’ theorem, which lets us revise our <em>prior</em> beliefs \(p(z)\) using new data (via the <em>likelihood</em> \(p(\text{data} \vert z)\)) to form a <em>posterior</em> belief \(p(z \vert \text{data})\).</p>

\[p(z \vert \text{data}) = \frac{p(\text{data} \vert z)p(z)}{p(\text{data})} = \text{Posterior} = \frac{\text{Likelihood} \times \text{Prior}}{\text{Evidence}}.\]

<p>Bayesian inference is like scientific reasoning on autopilot: start with a hunch, collect some data, update your expectations.
It can be thought of as a natural extension of probabilistic logic—one that allows us to reason about hypotheses, not just outcomes.
In this framework, we can assign a probability to a hypothesis, even if we don’t know yet whether it’s true or false.
This is a big shift from the frequentist perspective, where hypotheses are usually treated more like yes-or-no questions to be tested, not graded on a sliding scale of belief.</p>

<p>Alright, back to surprise and entropy!</p>

<p>Let’s revisit our die.
Suppose the die isn’t fair—it actually rolls a 1 with probability 0.6, and the other numbers share the remaining 0.4 equally.
Now, if we think the die is fair, but it keeps landing on 1 suspiciously often, the surprise we feel doesn’t come just from the event itself.
It also comes from the mismatch between our internal model (a fair die) and reality (a biased one).</p>

<p>In other words, surprise isn’t just about what happened—it’s also about what we expected to happen.
When our expectations are off, our surprise spikes.
This is why having a good model of the world—one that reflects actual probabilities—is so important.
The better our model, the better we can anticipate what’s likely and stay one step ahead of surprise.</p>

<p>Let us assume a perfect world model.
If \(P(X=1) = p(1) = p_1 = 0.6\) is the probability of rolling 1 with the die, then the <strong>surprise</strong> \(h\) of rolling it is</p>

\[h(p(1)) = \ln\left( \frac{1}{p(1)} \right) =  - \ln\left( p(1) \right).\]

<p>It is high if the probability is small, in fact,</p>

\[\lim\limits_{p_s \rightarrow 0}\ln\left( \frac{1}{p_s} \right) = \infty\]

<p>and</p>

\[\lim\limits_{p_s \rightarrow 1}\ln\left( \frac{1}{p_s} \right) = 0.\]

<p>While we multiply probabilities to figure out the probability of independent events happening together, surprise works differently—it adds up.
Meaning ten 1s in a row is twice as surprising as five 1s in a row, because</p>

\[\ln{\frac{1}{a \cdot b}} = \ln\frac{1}{a} + \ln\frac{1}{b}.\]

<p>Now, to figure out how much we could be surprised by an event—assuming a perfect and known world model in form of a probability distribution of all possible events, e.g. of the outcome of rolling a die—we could compute the <strong>average surprise</strong> of that distribution which is called its <strong>entropy</strong> \(H\):</p>

\[H(P) = \sum_s p(s) \ln\left( \frac{1}{p(s)} \right) = - \sum_s p(s) \ln p(s).\]

<p>We multiply the surprise of the event \(s\) by its probability because more likely events occur more often.
Therefore, they will contribute more to our overall surprise, if we would repeat rolling the die over and over again.
The corresponding <strong>differential entropy</strong> for a continuous random variable \(X\) with the probability density function \(p(x)\) can be written as:</p>

\[H(p(x)) = -\mathbb{E}_{x \sim P}\left[ \ln p(x) \right] = - \int p(x) \ln\left( \frac{1}{p(x)} \right) dx.\]

<p>Again, we are taking the average of the log probability over the distribution \(p(x)\).
This tells you how surprising the outcomes are on average.
The higher the entropy, the more uncertain or unpredictable the variable.</p>

<p>So far so good.
But what happens if our world model is wrong?
Can we learn and adjust it?
And how can we first separate the surprise caused by our incorrect prior beliefs from the overall surprise to then minimize this second source of surprise?</p>

<p>Let us assume we belief in a fair coin, that is \(Q(X = \text{heads}) = Q(X = \text{tails}) = 0.5\)—a reasonable approximation.
I denote our belief as \(Q\) and its surprise as \(h_q\).
But now let’s assume the coin is rigged: \(P(X = \text{heads}) = 0.99\) and \(P(X = \text{tails}) = 0.01\)
We belief the probability of ten times heads in row is</p>

\[Q(\text{10 heads}) = 0.5^{10} = 0.001\]

<p>and the surprise is</p>

\[h_q(\text{10 heads}) = \ln 0.5^{-10} \approx 7.\]

<p>But in reality</p>

\[P(\text{10 heads}) = 0.99^{10} = 0.9\]

<p>and</p>

\[h(\text{10 heads}) = \ln 0.99^{-10} \approx 0.1.\]

<p>Therefore, we are much more surprised as we should be and this surprise is mostly caused by believing in the wrong world model.
This leads us directly to <strong>cross-entropy</strong> \(H(P,Q)\), which is the average surprise you will get by observing a random variable governed by distributions \(P\), while believing in \(Q\).</p>

\[H(P,Q) = \sum\limits_s P(X=s) \ln\left( \frac{1}{Q(X=s)} \right) = \sum\limits_s p(s) \ln \left( \frac{1}{q(s)} \right).\]

<p>We multiply our subjective surprise by the “real” probability which means that if an event is very likely but our surprise for it is high—meaning we do not expect it—the cross-entropy is high as well.
Furthermore, if our beliefs are prefect, i.e. \(P=Q\) the cross-entropy is equal to the entropy since</p>

\[H(P,P) = H(P)\]

<p>holds.
Also note that the cross-entropy is <strong>asymmetric</strong>, that is, \(H(P,Q)\) is usually not equal to \(H(Q,P)\).
For example, take the case from before.
In this situation we are less surprised believing in a fair coin while it is rigged  than believing in \(P\) while it is fair, that is,</p>

\[H(P,Q) = \ln\left( \frac{1}{0.5} \right) \approx 0.69 \leq 0.5 \cdot \left[ \ln\left( \frac{1}{0.99} \right) + \ln\left( \frac{1}{0.01} \right) \right] \approx 2.31 = H(Q,P)\]

<p>Furthermore, and very importantly, for any model \(Q\), <strong>the cross-entropy can never be lower than the entropy of the underlying generating distribution</strong>:</p>

\[H(P,Q) \geq H(P).\]

<p>Using cross-entropy and entropy, we can now compute the surprise caused by our wrong beliefs:</p>

\[D_{\text{KL}}(P \Vert Q) = H(P,Q) - H(P) = \sum_s p(s) \ln\left( \frac{1}{q(s)} \right) - \sum_s p(s) \ln\left( \frac{1}{p(s)} \right).\]

<p>This can be compressed to</p>

\[D_{\text{KL}}(P \Vert Q) = \sum_s p(s) \ln\left( \frac{p(s)}{q(s)} \right).\]

<p>This is called the <em>Kullback–Leibler divergence</em> \(D_{\text{KL}}(P \Vert Q)\).
It is the divergence of \(P\) from \(Q\) (also known as the <em>relative entropy</em> of \(P\) with respect to \(Q\)).
To improve our beliefs we could minimize the Kullback–Leibler divergence, that is, we could</p>

\[\min D_{\text{KL}}(P \Vert Q).\]

<p>However, in machine learning you will usually hear about only minimizing the cross-entropy \(H(P,Q)\).
This is because we can not change the entropy of “reality”, that is, we can not change \(H(P)\)—we can not change reality (without acting) but our beliefs about it.
Thus, minimizing \(H(P,Q)\) achieves the same goal as minimizing \(D_{\text{KL}}(P \Vert Q)\).</p>

<p>Of course, the big problem is that we don’t know \(P\)!</p>

<h2 id="the-circularity-of-perception">The Circularity of Perception</h2>

<p>Because of a <em>drift</em> towards complexity, many assume that brains began to build models \(Q\) of the world so that they can explain sensory inputs by <strong>inferring</strong> their <strong>hidden causes</strong>.
For example, your brain might have an internal model of snakes and bears—what they are, how they look, and what they can and cannot do.
Brains can generate missing information, e.g., they can deal with obfuscated objects.
They have an intuitive understanding of the physical world—of movement, speed, and heaviness.
They are biased toward what is usually helpful.
In that sense, our body “understands” e.g. gravity “intuitively” before knowing anything about Newton’s or Einstein’s theories.</p>

<p>A brain, in that view, is like a judge.
There is the raw sensory data obtained by <em>observations</em>, which leads to the generation of predictions.
It is the incoming observation, or what you see.
Furthermore, there are <em>prior beliefs</em> learned through experience and evolution.
It is what you usually expect, and if these expectations do not fit your sensory data, you are surprised.
The brain can check how well your explanation matches your observation.</p>

<p>Following this principle, the first component called <em>accuracy</em> is defined as a measure of how well the data fit an explanation.
The second component is <em>complexity</em>.
It refers to how abstruse this explanation is.
We want to prefer simple models of the world, i.e., simple explanations.</p>

<p>From this we get an obvious tension between the two components.
We probably get the highest accuracy by using a very complex model, but such a model cannot generalize and will produce inaccurate predictions for new situations.
This tension is the <strong>free energy</strong> \(F\), that is,</p>

\[F = \text{complexity - accuracy}.\]

<p>We (or our brains) want maximal accuracy with minimal complexity.
Thus, following the definition of \(F\), the brain minimizes free energy \(F\).
To survive, it requires <em>high accuracy</em> <strong>and</strong> <em>low complexity</em>.
If this is achieved, then there is only low free energy \(F\)—there is nothing to be gained.</p>

<p>It is assumed that the brain compresses the high‑dimensional sensory data into a somewhat manageable form by finding commonalities and hidden structures in the data.
Furthermore, there might be a lot of information in the data that is not important for the organism’s self‑creation and survival.
Especially at a deep, low‑dimensional layer, hidden neurons—also called <em>latent variables</em> or <em>latents</em>—represent causes.
Importantly, <strong>for such latent neurons there is no ground truth of what their activity should be or mean because they do not interface with the outside world.</strong>
The brain is free to “choose” whichever latents it wants.
This opens up a question: how can the brain test if its world model is a good one if it has no direct access to the world?</p>

<p>We cannot verify the latents, but we can verify their consequences, meaning the brain should be able to reconstruct the source sensory data \(x\) from its compressed representation, that is, from its latents \(z\).</p>

<p><strong>Remarks:</strong> In the original paper \(z\) is denoted as \(\vartheta\) and called “parameterise environmental forces or fields that act upon the system”.
\(x\) is denoted as \(\hat{y}\) and is a function of actions \(\alpha\) (the effects of the system on the environment).
Furthermore, I will use \(\theta\) as the parameters of the involved neural networks which, in the paper, is denoted as \(\lambda\) and called “quantities that describe the system’s physical state”.
It can also mean the parameters of the Gaussian distribution that are computed by the network.
Friston also differentiate between parameters that can change quickly, slowly and very slowly.</p>

<blockquote>
  <p>Factorization of the ensemble density to cover quantities that change with different timescales provides an ontology of processes that map nicely onto perceptual inference, attention and learning. – <a class="citation" href="#friston:2006">(Friston et al., 2006)</a></p>
</blockquote>

<p>Using a probabilistic modeling framework, we can say that, given our beliefs \(z\) (about how the world works), we should be able to predict the sensory data \(x\) using our internal model that approximates the true probability distribution \(p\), meaning that</p>

\[p(x \vert z)\]

<p>should be high where \(z\) is given, and \(G_{\theta}(z) \approx x\) is computed by a neural network \(G_{\theta}\)—\(z\) is given while the brain <strong>generates</strong> \(x\).</p>

<p>Note that \(z\) is assumed to be of low dimensionality while \(x\) is of high dimensionality.
\(G_{\theta}\) is a <em>generative model</em>.
We call \(p(z)\) <em>priors</em> and \(z\) <em>causes</em>.
If \(z\) is not the cause for \(x\), we have to change our model to learn different latents \(z\).</p>

<p>The brain learns these latents or causes by compressing the sensory data.
Given sensory data \(x\) the nervous system tries to figure out what <em>causes</em> \(z\) produced \(x\); that is, the system wants to model</p>

\[p(z \vert x)\]

<p>accurately.
This is called <em>inference</em> because the system <em>infers</em> causes from <em>observations</em>—\(x\) is given while the brain <strong>searches</strong> for \(z\).</p>

<p>However, this is a hard problem because, unlike generation, there is no function that directly computes \(z\) given \(x\)—the generative model is not invertible.
One could generate a lot of candidates of sensory data \(x'\) using the generative model and compare these candidates to the true observation \(x\).
However, this is not feasible because, even though \(z\) is of lower dimensionality than \(x\), there are still too many possibilities.
We say that the problem is <em>intractable</em>.</p>

<p>In reality, our brain solves this problem almost <strong>instantaneously</strong>.
If there is a grizzly bear in front of us, we have to figure it out fast!
So how does the brain solve this seemingly impossible task?</p>

<p>Instead of finding the exact latents \(z\), our brain tries to find an approximation \(q(z \vert x)\) using a so-called <em>recognition model</em> \(R_{\theta}\), which is distinct from but interdependent on the <em>generative model</em> \(G_{\theta}\).
It works in the opposite direction compared to the generative model by mapping a sensory observation \(x\) to a distribution of causes:</p>

\[R_{\theta}(x) \approx z.\]

<p>The result is only an approximation—a rough guess \(z\) of what causes the observations \(x\).
To improve the guess the brain might do multiple rounds of <em>recognition</em> and <em>generation</em> in tandem to arrive at an optimal approximation.
This process we can call <strong>perception</strong> and it can only work if <em>recognition</em> and <em>generation</em> are aligned with each other.
It presupposes such an alignment, thus learning through experience!</p>

<p>In perception, free energy is minimized, meaning that an optimization happens until <em>recognition</em> roughly fits <em>generation</em>:</p>

\[x \approx G_{\theta}(R_{\theta}(x)).\]

<p>If this is the case, our brain has found an explanation or causes \(z\) that minimize <em>free energy</em>—one that explains the sensory input \(x\) and aligns well with one’s prior beliefs.
Perception basically solves for</p>

\[\text{minimize } \left[ \text{complexity - accuracy} \right] = \min F.\]

<p>It rapidly adjusts the activity of latent neurons.</p>

<p>Long‑term learning via experience is required to gradually refine both models—i.e. neural networks—to align them better with each other and to construct better world models.
Both <em>learning</em> and <em>perception</em> serve the same goal: reducing uncertainty in the environment the organism at least partly constructs by building optimal models of it and finding useful explanations for sensory data within those models.</p>

<p>Now there is a third way to minimize free energy, that is, <em>acting</em> which equates to a change in the organisms environment caused by the organism.
Acting can be as simple as turning your head.
It changes the sensory data the organism is <em>receiving</em>.
Therefore, one can model the sensory data \(x\) as a function of action \(\alpha\), that is, \(x(\alpha)\).
Consequently, one can interpret perception as an active process, therefore, variational inference becomes <em>active inference</em>.
The organism does not receive but selects its sensory data and it can anticipate to minimize “future” free energy actively.
A feedback loop is constructed thus non-linearity is introduced:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>acting -&gt; perceiving/learning -&gt; believing -&gt; acting.
</code></pre></div></div>

<p>Each step influences the next one.</p>

<h2 id="variational-inference">Variational Inference</h2>

<p>Up to this point I did not explain how generation and perception can be established, that is, how to find good models \(G_{\theta}\) and \(R_{\theta}\).
In the following section, we look at how such models can be build in the field of <em>machine learning</em> using the framework of <em>variational inference</em>.</p>

<p>First, we assume an unknown latent distribution \(p(z)\) (causes) and a distribution of our observations \(p(x)\) (for those causes).
From the rules of probability theory we know that</p>

\[p(x,z) = p(z) \cdot p(x \vert z)\]

<p>holds.
That is, to compute the probability that \(x\) and \(z\) occur together, we can first compute the probability of \(z\) and then multiply the conditional probability of \(x\) given \(z\).</p>

<h3 id="minimizing-divergence">Minimizing Divergence</h3>

<p>Our <em>generative neural network</em> \(G_{\theta}\) computes \(p(x \vert z)\) given \(z\).
But on a computer we usually do not compute the probability distribution directly.
Because there is only a finite amount of memory, we cannot represent such a distribution exactly.
What one usually does is compute the parameters—means and standard deviations—of a Gaussian distribution.
One might ask: Why is this feasible? Why the Gaussian?
I will explain this later.
For now, assume that it is a good, computable choice.</p>

<p>An artificial neural network does not output probabilities directly but instead the parameters, i.e. <em>means</em> \(\mu\) and the covariance matrix \(\Sigma\) of the Gaussian.
The actual probability can then be computed using the Gaussian formula:</p>

\[p(x \vert z) = \frac{1}{(2\pi)^\frac{d}{2} \vert \Sigma \vert^{\frac{1}{2}}} \exp\left( -\frac{1}{2} (x - \mu )^{\top} \Sigma^{-1} (x - \mu) \right)\]

<p>where \(d\) is the dimension of \(x\). 
This simplifies to</p>

\[p(x \vert z) = \frac{1}{(2\pi)^\frac{d}{2}} \exp\left( -\frac{1}{2} \Vert x - \mu \Vert^2 \right)\]

<p>if \(\Sigma\) is the identity matrix.</p>

<p>The objective is now to minimize the mismatch between the “real” data distribution \(p(x)\) and the distribution \(p_{\theta}(x)\) that our model approximates.
As I said before, we do not know \(P\), so we approximate it by sampling from our environment.</p>

<p>As mentioned before, we can measure the difference between two distributions by using the Kullback–Leibler divergence:</p>

\[D_\text{KL}(p(x) \Vert p_{\theta}(x)) = \sum_x p(x) \ln\frac{p(x)}{p_{\theta}(x)}.\]

<p>We want to minimize the \(D_\text{KL}\).
Any mismatch between the two distributions will result in \(D_\text{KL} &gt; 0\).</p>

<p>By using log rules we can convert the equation to:</p>

\[D_\text{KL}(p(x) \Vert p_{\theta}(x)) = \sum_x p(x) \ln p(x) - \sum_x p(x) \ln p_{\theta}(x).\]

<p>The first term is the entropy of the data (a measure of uncertainty inherent in the observations).
This entropy only depends on the true data.
Therefore, we cannot optimize it away—we cannot influence it—thus we can regard it as a constant.
Consequently, we can focus solely on the second term and minimize</p>

\[-\sum_x p(x) \ln p_{\theta}(x)\]

<p>or</p>

\[\arg\max\limits_{\theta} \sum_x p(x) \ln p_{\theta}(x).\]

<p>Again, we do not know \(p(x)\).
We only have access to \(N\) samples from it.
Therefore, we approximate:</p>

\[\sum_x p(x) \ln p_{\theta}(x) \approx \frac{1}{N} \sum\limits_{i=1}^N \ln p_{\theta}(x_i).\]

<p>The averaging works because high-probability samples should appear more often in our data; their contribution remains consistent after averaging.</p>

<p>Up to this point, our <em>generative model</em> maps \(z\) to \(p_{\theta}(x \vert z)\).
It outputs means and standard deviations for Gaussians.
What is left is the expression \(p_{\theta}(x)\).</p>

<p>To compute the total probability of an observed value \(x\), we must take into account that different values of the latent \(z\) could have generated it.
Mathematically we sum the probability that each latent value could explain our observation, weighted by how <em>likely</em> that latent value is according to the prior distribution:</p>

\[p_{\theta}(x) = \sum p_{\theta}(z) p_{\theta}(x \vert z).\]

<p>Both the parameters of the prior and the weights that transform the latent into the conditional distribution are what our model needs to learn.</p>

<h3 id="how-generative-models-learn">How Generative Models Learn</h3>

<p>How do we find or optimize good parameters \(\theta\) for our generative model \(p_\theta(x \vert z)\) and our prior beliefs \(p_{\theta}(z)\)?
We basically follow five steps:</p>

<p><strong>(1)</strong> Take a data point from the training dataset.</p>

<p><strong>(2)</strong> Randomly sample many candidates \(z_k\) from our current prior \(p_{\theta}(z)\).</p>

<p><strong>(3)</strong> Map each \(z_k\) through the generative model \(G_{\theta}(z)\) to get means \(\mu_k\) and our covariance matrix \(\Sigma_k\). Remember that the generative model is used to model \(p(x \vert z)\), i.e. to generate sensory data \(x\) given the causes \(z\).</p>

<p><strong>(4)</strong> Given \(\mu_k, \Sigma_k\) we compute, for each \(z_k\), the value of \(p_{\theta}(x \vert z_k)\) (likelihood) where</p>

\[p_{\theta}(x \vert z_k) \sim \exp\left(\frac{-\Vert x-\mu_k \Vert^2}{2} \right).\]

<p><strong>(5)</strong> Compute the log marginal likelihood estimate</p>

\[\ln\left[\sum_{z}^{\{z_k\}} p_{\theta}(z)  p_{\theta}(x \vert z) \right].\]

<p>We iteratively update \(\theta\) to maximize our log-likelihood across the dataset, usually using backpropagation and gradient descent with stochastic sampling.</p>

<p><strong>Remark:</strong> The brain does not use backpropagating or gradient decent, consequently, Friston relies on <em>predictive coding</em> which is neurally plausible: It mirrors cortical feedback and feedforward loops, with ascending prediction errors and descending predictions.</p>

<h3 id="recognition-as-guided-sampling">Recognition as Guided Sampling</h3>

<p>One problem we face is step <strong>(2)</strong>, especially if the latent space has high dimensionality.
I already mentioned that it is not feasible to generate a bunch of possible sensory data to compare them with \(x\) to find good possible causes \(z\).
There are usually too many possible causes, and because we approximate our unknown prior distribution by sampling, we would need to somehow cover the whole latent space.
This is computationally infeasible because the number of required samples increases exponentially with the number of dimensions of the latent space (<em>curse of dimensionality</em>).</p>

<p>However, usually only a tiny fraction of those causes lead to our observation—most latents do not explain the data!
Thus, we basically want to sample those causes \(z\) that are likely to result in \(x\).
Importantly, while we are okay with ignoring causes that do not matter, we do not want to miss rare causes.
The idea is to oversample rare cases instead of risking missing them, and then adjust for this oversampling mathematically.</p>

<p>To know which regions we want to sample from, we train a separate neural network \(R_{\theta}\) to serve as our “guide”.
It is our <em>recognition model</em> \(R_\theta\) that tries to approximate the inversion of the <em>generative model</em>.
It learns a <strong>variational distribution</strong> \(q_{\theta}(z \vert x)\) which predicts, for each data point \(x\), the distribution over the latent space, focusing on regions that have likely generated \(x\).</p>

<p>Consequently, our formula for \(p_{\theta}(x)\) changes from</p>

\[p_{\theta}(x) = \sum p_{\theta}(z) p_{\theta}(x \vert z) = \mathbb{E}_{z \sim p_{\theta}(z)} \left[ p_{\theta}(x \vert z) \right]\]

<p>to</p>

\[p_{\theta}(x) = \sum \frac{p_{\theta}(z)}{q_{\theta}(z \vert x)} q_{\theta}(z \vert x) p_{\theta}(x \vert z).\]

<p>which can also written as the expected value of the likelihood scaled by the ratio of sampling frequencies over the distribution \(q\):</p>

\[\mathbb{E}_{z \sim q_{\theta}(z \vert x)} \left[ p_{\theta}(x \vert z) \frac{p_{\theta}(z)}{q_{\theta}(z \vert x)} \right].\]

<p>Mathematically these two equations are equivalent.
However, note that we are using a finite number of samples and that \(q_{\theta}(z \vert x)\) determines which \(z_k\)’s we continue training on.
Therefore, we are computing an estimation of the expected value.
\(\frac{p_{\theta}(z)}{q_{\theta}(z \vert x)}\) is an adjustment for sampling bias: it adjusts the weight for events we sample more frequently than they occur (e.g., rare but important cases).</p>

<p>Now we take the logarithm to arrive at our term to maximize:</p>

\[\ln p_{\theta}(x) = \ln \mathbb{E}_{z \sim q_{\theta}(z \vert x)} \left[ p_{\theta}(x \vert z) \frac{p_{\theta}(z)}{q_{\theta}(z \vert x)} \right].\]

<p>However, the logarithm of an expectation is hard to optimize.
It is noisy and unstable.
Furthermore, computing the gradient would require that we first compute the average and only then propagate the error.
This means that after computing \(z\) using our <em>recognition model</em>, we would need to first compute multiple probabilities for \(x\) given \(z\) using our <em>generative model</em> before starting backpropagation (i.e. \(\theta\)-optimization).
It is a computational bottleneck.</p>

<p>Instead, we would like to swap the order like this:</p>

\[\mathbb{E}_{z \sim q_{\theta}(z \vert x)} \ln \left[ p_{\theta}(x \vert z) \frac{p_{\theta}(z)}{q_{\theta}(z \vert x)} \right]\]

<p>But the log of an average and the average of a log are not equal.
However, since the logarithm is a concave function we get:</p>

\[\ln \mathbb{E}[X] \geq \mathbb{E}[\ln X]\]

<p>or in our case</p>

\[\ln \mathbb{E}_{z \sim q_{\theta}(z \vert x)} \left[ p_{\theta}(x \vert z) \frac{p_{\theta}(z)}{q_{\theta}(z \vert x)} \right] \geq \mathbb{E}_{z \sim q_{\theta}(z \vert x)} \ln \left[ p_{\theta}(x \vert z) \frac{p_{\theta}(z)}{q_{\theta}(z \vert x)} \right].\]

<p>The right-hand side is a lower bound, also called the <em>model evidence</em> or <em>evidence lower bound</em> (ELBO).
The negative ELBO is known as <em>variational free energy</em>.
Maximizing ELBO also maximizes our original objective, but maximizing ELBO is computationally much easier.</p>

<p>So we want to maximize the right-hand side:</p>

\[\ln p_{\theta}(x) \geq \mathbb{E}_{z \sim q_{\theta}(z \vert x)} \ln \left[ p_{\theta}(x \vert z) \frac{p_{\theta}(z)}{q_{\theta}(z \vert x)} \right]\]

<p>and, using log rules, we arrive at the following:</p>

\[\ln p_{\theta}(x) \geq \underbrace{\mathbb{E}_{z \sim q_{\theta}(z \vert x)} \left[ \ln p_{\theta}(x \vert z) \right]}_{\text{accuracy}} - \underbrace{\mathbb{E}_{z \sim q_{\theta}(z \vert x)} \left[ \ln \frac{q_{\theta}(z \vert x)}{p_{\theta}(z)} \right]}_{\text{complexity}}.\]

<p>We arrive at a beautiful interpretation.
The first term</p>

\[\mathbb{E}_{z \sim q_{\theta}(z \vert x)} \left[ \ln p_{\theta}(x \vert z) \right]\]

<p>is the <em>accuracy</em> mentioned before, and the second term</p>

\[\mathbb{E}_{z \sim q_{\theta}(z \vert x)} \left[ \ln \frac{q_{\theta}(z \vert x)}{p_{\theta}(z)} \right]\]

<p>is the <em>complexity</em>.
The second term is also the negative of the Kullback–Leibler divergence between the distribution \(q\) and the prior distribution of latents, i.e. \(D_\text{KL}\left[ q_{\theta}(z \vert x) \Vert p_\theta(z) \right]\).
It prevents the “guide” from becoming too specialized and inventing overly complex distributions of latents.
It ensures that the distribution of latent factors for any specific data point (observation) does not stray too far from our general prior belief about the latent space as a whole.
In other words, it ensures that the latent space remains smooth and numerically well-behaved.</p>

<h2 id="the-organism-as-its-own-evidence">The Organism as Its Own Evidence</h2>

<p>Friston et al. importantly consider the observer perspective and argue that, in actuality, the <strong>system/agent/organism is not trying to maximize reward</strong>.
It is the prior belief that guides <em>perception</em> and <em>action</em> to construct an environment that (partly) aligns with the organism’s beliefs.
In that sense, organisms are (partly) their own self-fulfilling prophecy—they constantly try to “make true” what they “believe” while, at the same time, the filtered and selected “truth” changes their beliefs.</p>

<blockquote>
  <p>The inherent circularity obliges the system to fulfil its own expectations. In other words, the system will expose itself selectively to causes in the environment that it expects to encounter. […] Anthropomorphically, we may not interact with the world to maximise our reward but simply to ensure it behaves as we think it should. – <a class="citation" href="#friston:2006">(Friston et al., 2006)</a></p>
</blockquote>

<p>If an insect has evolved to expect its environment to be dark, it will use its action, e.g. movement, to keep its environment dark.
It acts to keep its sensory data consistent with its belief—moving into shadows, avoiding light.
To an external observer, this looks like as if the insect has learned that darkness is good, and has been reinforced to prefer it but from the insect’s own perspective, it’s just acting to reduce surprise by finding sensory data that matches its expectations.</p>

<p>Consequently, <strong>value</strong> is not the interpretation of an external signal (<em>reward</em>) but is constructed by a prior belief about what states minimize surprise—what states are “normal” or “preferred” and since experience and prior belief influence each other in a feedback loop, we can hardly observe direct causal relation between the two.
Value is basically the inverse of surprise and instead of figuring out which states are good based on reward, the FEP says: Organisms are already biased to expect to be in certain types of states (like being alive, warm, safe, fed).
Preferences are encoded as priors over observations/states, so log-preference shows up as negative expected surprise in the expected free energy.
Note however that <em>action-selection version of FEP</em> uses expected free energy (forward-looking), which decomposes into a <em>pragmatic</em> term (preference satisfaction) and an <em>epistemic</em> term (information gain / curiosity) to account for exploration!
These are <strong>low-surprise states</strong>—what the organism is used to.
Therefore, organisms act not to find new rewards, but to stay within those expected (familiar) states <a class="citation" href="#friston:2011">(Friston, 2011)</a>.
One can still include costs or punishments in the model—but they are treated as part of the prior beliefs about what kinds of states should be avoided.
For example, if being in pain is costly, the organism has a prior belief that “I don’t usually feel pain”.</p>

<blockquote>
  <p>The problem of finding sparse rewards […] is nature’s solution to minimizing entropy. – <a class="citation" href="#friston:2011">(Friston, 2011)</a></p>
</blockquote>

<p>That is, rewards are rare and localized because organisms are repeatedly drawn to a small set of predictable, low-surprise states—such as food, shelter, and safety.</p>

<blockquote>
  <p>These dynamics rest on complementary self-construction (autopoiesis) and destruction (autovitiation). – <a class="citation" href="#friston:2011">(Friston, 2011)</a></p>
</blockquote>

<p>In other words, the organism builds up behaviors that help it stay viable, and eliminates those that lead to bad outcomes.
Behaviors emerge because the agent must stay in a balance with its environment—not too chaotic, not too static—to continue existing.</p>

<p>So why doesn’t the agent settle in one fixed state?
Because organisms or living systems—like animals or humans—are not static systems.
They have to continuously interact with their environment to survive.
We cannot just stay in a single “perfect” state.
“Being alive” requires constant change (eating, breathing, sleeping, thermoregulating, moving around).
If we would stay in one place (one state), we would have violated our own prior beliefs about what kinds of sensory inputs and physiological states we expect to have over time.
We are in <em>itinerancy</em>: moving through a cycle of low-surprise states, not staying in one.
Thus, <em>itinerant policies</em> arise because staying in a single state forever would be more surprising (and thus less viable) than cycling through a set of familiar, predictable states over time <a class="citation" href="#friston:2011">(Friston, 2011)</a>.</p>

<blockquote>
  <p>To minimize surprise, organisms must move through a predictable sequence of diverse but familiar states.</p>
</blockquote>

<p>Learning, perception, and action then is about <strong>adjusting expectations</strong> and <strong>updating generative models</strong> on different time scales. 
To the best of my knowledge, it is unclear in which way one dominates over the others but it seems as if all three are highly depend on each other.
From this I conclude that it is important to dwell on problems, to think, to do the hard work to adjust long-term causes to be able to construct a reality/environment that makes sense consistently.
But it is also important to act in such a way that one is exposed to surprise, at least from time to time.</p>

<blockquote>
  <p>Without surprise we are not adapting our belief system.</p>
</blockquote>

<p>A realist will argue that the survival of an organism is bound to the accuracy of its model about reality—meaning its expectations align well with it, therefore, we can assume that we more or less experience reality as it is.</p>

<blockquote>
  <p>Because the free energy is low, the inferred causes approximate the real environmental conditions. This means the systems physical state must be sustainable under these environmental forces, because each system is its own existence proof. – <a class="citation" href="#friston:2006">(Friston et al., 2006)</a></p>
</blockquote>

<p>A constructivist will point out that our beliefs (latents, causes) only tend to generate sensory data and that we can not jump from this to an objective reality or “the truth” which is “out there”.
The causes we believe in may have little or nothing to do with the “real” causes—if such causes even exist.
This challenges my earlier assumption of <em>indirect realism</em>.
Causes are just effective in generating sensory data—data that is already constructed by the very same system which tries to reconstruct them.
While I don’t see constructivism as an anti-realism, constructivist might doubt any sort of realism that rejects constructivist elements and that assumes that we can reach an observer-independent “objective turth”.
Constructions are contingent and constrained—they work in the context of a history of evolution but they are not “the truth” and, in my opinion, they can not be separated from the observing system that construct them.</p>

<p>Because <em>recognition</em> and <em>perception</em> work in tandem, we can say that changing our expectations literally changes our perception—it changes our environment.
But at the same time, our perception changes our expectations if we are surprised.
Thus the FEP highlights the circularity of <em>perception</em> and <em>action</em> or <em>belief</em> and <em>reality</em> because organisms minimize free energy not only by adapting their beliefs to fit their environment but also by adapting their environment to fit their beliefs.
It assumes an environment that is acted upon effectively (which is a somewhat <em>optimistic</em> outlook and motivator to make use of our imagination to construct beliefs of a less cruel world).</p>

<p>Also, I want to emphasise, that surprise that leads to a loss of one’s well-known environment, can be painful because one’s prior beliefs about the world are violated—you enter a state of increased entropy.
This might lead to a temporary suspension of meaning because you no longer know what to expect, or how to act effectively.
Adaptation is painful because it involves breaking down a stable, low-free-energy model, and building a new one that’s better at minimizing future surprise.
It is especially painful if the previous model had high precision priors (e.g. strong, confident beliefs), the environment changes faster than the model can update (e.g. economic instability, cultural shifts), or there is no clear path to a new, stable <em>attractor state</em>.</p>

<p>In today’s society the individual is expected to be flexible, to constantly re-skill and adapt to environmental changes.
For example, we switch our jobs more and more frequently, which is often seen as desirable or necessary.
Flexibility—constantly updating beliefs, changing environments, shifting roles—might sound adaptive.
And in a way, it is.
But under FEP, excessive flexibility undermines model stability.
Chronic uncertainty leads to chronic prediction error, which is metabolically and psychologically exhausting and if we are constantly updating our generative model without forming stable priors, we fail to build any useful expectations.
This creates a sense of incoherence or alienation—a <strong>loss of identity</strong>; a lived experience of “nothing makes sense”.
So in modern systems that demand rapid flexibility, we’re often forced to dissolve stable models faster than we can reconstruct them.
The result? Anxiety, burnout, and a background hum of ontological instability.</p>

<p>The FEP suggest that there is a balance to be found: exposure to surprise and a certain “willingness” to adapt without losing ones identity and purpose.
I find it desirable to recognize, or at least consider that emotional pain may not reflect maladaptation but the cost of updating one’s expectations in the face of an unstable world.
Flexibility, when demanded too often or too rapidly, erodes the very models that minimize surprise over the long term.
Thus, while adaptability is vital, stability is sacred.</p>

<p>Let me finish with some speculations:
I have the gut feeling that a strange, perhaps paradoxical, alliance forms here between Kant, Heidegger, and (embodied) constructivists like Maturana, Varela, and Luhmann. 
Each offers a piece of the puzzle when viewed through the lens of the free energy principle.
Kant emphasized the structuring role of <em>a priori</em> reason—the internal scaffolding that makes experience possible.
Futhermore, FEP mirrors Kant’s view that perception is constructed by the mind’s active faculties.
The world “as it is” is unknowable directly—we only perceive it through mental filters and structure.
Heidegger, in contrast, stressed the <em>immediacy (i.e. pre-conceptual) and embeddedness of experience</em>: we are not detached observers but beings thrown into a meaningful world that is not build in the mind. 
Heidegger’s philosophy fits tightly with theories that view perception as <em>embodied</em>, <em>skillful</em>, and <em>embedded in practical action</em>, not representational—just like <em>enactivism</em> and Gibson’s <em>direct perception</em>.
Constructivists such as Maturana and Varela also went with Heidegger but moved the world into the system itself, arguing that <em>cognition arises through dynamic interaction with an environment</em> that is, in part, <em>selected</em> or <em>enacted</em> by the organism itself.</p>

<p>Yet, each position has its blind spots. 
Kant may have overestimated the sovereignty of internal reason;
Heidegger arguably romanticized lived immediacy; and constructivists often risk underestimating the constraints imposed by sensory data—or more precisely, by the need to stay within states that minimize surprise. 
Yes, action and cognition are deeply intertwined (Maturana, Varela), and yes, the world is our best model of itself (Heidegger) but we still model it and it pushes back (Kant).</p>

<p>If we want to go full Heideggerian, we probably need a different interpretation where free energy minimization becomes about maintaining practical grip on the world, not constructing a picture of it and where prediction errors don’t refer to “incorrect beliefs” but breakdowns in coping (e.g., when the tool “refuses” to function).
We would have to move to a life-mind continuity thesis (LMCT) (see e.g. <a class="citation" href="#bogota:2024">(Bogotá, 2024)</a>).
It would be interesting and maybe refreshing to departure from the representation-driven approaches (it is unsurprising that one of the most famous AI conferences is called “International Conference on Learning Representations”).
But this would be another topic for another time.</p>

<hr />

<h2 id="appendix">Appendix</h2>

<p>The generative neural network \(G_\theta\) computes an approximation of \(p(x \vert z)\) given \(z\) but as a Gaussian distribution (its means and standard deviations).
So, why using the Gaussian is a good choice?</p>

<p>The central limit theorem (CLT) states that:</p>

<blockquote>
  <p>If you take a large number of independent and identically distributed (i.i.d.) random variables with finite mean and variance, then their properly normalized sum tends toward a Gaussian (normal) distribution, regardless of the original distribution of the variables.</p>
</blockquote>

<p>The Gaussian appears often in nature because of the aggregation of many small effects and it has the maximum entropy among all distributions with a given mean and variance.
It is the “most random” or “least biased” which makes it a natural default in many uncertain systems.</p>

<p>So first, it is mathematically convenient.
The Gaussion has a very simple functional form:</p>

\[q(z) = \mathcal{N}(z \vert \mu, \Sigma)\]

<p>It is fully described by just two parameters: mean \(\mu\) and covariance \(\Sigma\).
This makes optimization tractable because:</p>

<ul>
  <li>Expectations like \(\mathbb{E}_q[\ln p(x,z)]\) can often be computed in closed form or with low-variance Monte Carlo estimates.</li>
  <li>(Differential) entropy \(H(P)\) is known analytically for a Gaussian. It is \(H(\mathcal{N}) = \frac{1}{2} \ln((2\pi e)^d \vert \Sigma \vert)\), which simplifies the so-called evidence lower bound (ELBO):</li>
</ul>

\[\text{ELBO} = \mathbb{E}_q[\ln p(x,z)] - \mathbb{E}_q[\ln q(z)].\]

<p>Secondly, the Gaussian is reparameterizable.
Instead of sampling \(z \sim \mathcal{N}(\mu, \sigma^2)\) directly, we can rewrite it as:</p>

\[z = \mu + \Sigma^{1/2} \epsilon, \quad \epsilon \sim \mathcal{N}(0, I).\]

<p>Now, we are sampling from a fixed standard normal \(\epsilon\), and transforming it using differentiable operations involving \(\mu\) and \(\Sigma\).
Because \(z\) is now a function of \(\mu\) and \(\Sigma\) its gradients can flow.
It makes the entire process differentiable with respect to the parameters because \(\epsilon \sim \mathcal{N}(0, I)\) is independent of these parameters, enabling optimization via gradient descent.
In other words, the randomness becomes independent of ones model’s parameters, and the transformation from parameters to the sample becomes differentiable.</p>

<p>Thirdly, even if the true posterior is not Gaussian, many high‑dimensional distributions exhibit approximately Gaussian local behavior (via central limit theorem effects or Laplace approximations).
Furthermore, Gaussians are smooth and unimodal, making them a good first approximation for a wide variety of distributions.</p>

<p>Then there is the topic of computational efficiency.
Working with Gaussians keeps the variational family simple because (1) linear algebra operations are efficient, (2) sampling is cheap, and (3) we avoid expensive numerical integration or more complex families that would be harder to optimize.</p>

<p>Lastly, this framework is flexible.
Even though a single Gaussian is unimodal, we can extend it by using mixtures of Gaussians, normalizing flows, and Gaussian processes (an infinite‑dimensional generalization).</p>

<p>In summary: the Gaussian is used because it strikes the right balance between <strong>expressiveness</strong> (locally), <strong>tractability</strong> (closed‑form solutions and gradients), and <strong>optimization‑friendliness</strong>.</p>

<h2 id="references">References</h2>

<ol class="bibliography"><li><span id="friston:2006">Friston, K., Kilner, J., &amp; Harrison, L. (2006). A free energy principle for the brain. <i>Journal of Physiology-Paris</i>, <i>100</i>(1), 70–87. https://doi.org/10.1016/j.jphysparis.2006.10.001</span></li>
<li><span id="barry:2012">Razeto-Barry, P. (2012). Autopoiesis 40 years later. A review and a reformulation. <i>Origins of Life and Evolution of Biospheres</i>, <i>42</i>(6), 543–567. https://doi.org/10.1007/s11084-012-9297-y</span></li>
<li><span id="virenque:2024">Virenque, L. (2024). What is agency? A view from autonomy theory. <i>Biological Theory</i>, <i>19</i>(1), 11–15. https://doi.org/10.1007/s13752-023-00441-5</span></li>
<li><span id="barandiaran:2017">Barandiaran, X. E. (2017). Autonomy and enactivism: Towards a theory of sensorimotor autonomous agency. <i>Topoi</i>, <i>36</i>(3), 409–430. https://doi.org/10.1007/s11245-016-9365-4</span></li>
<li><span id="bogota:2024">Bogotá, J. D. (2024). Where there is life there is mind ... and free energy minimisation? In M. Martín-Villuendas, J. Gefaell, &amp; A. Cuevas-Badallo (Eds.), <i>Life and Mind: Theoretical and Applied Issues in Contemporary Philosophy of Biology and Cognitive Sciences</i> (pp. 171–200). Springer Nature Switzerland. https://doi.org/10.1007/978-3-031-70847-3_8</span></li>
<li><span id="varela:1997">Varela, F. J. (1997). Patterns of life: Intertwining identity and cognition. <i>Brain and Cognition</i>, <i>34</i>(1), 72–87. https://doi.org/https://doi.org/10.1006/brcg.1997.0907</span></li>
<li><span id="cummins:2014">Cummins, F. (2014). Agency is distinct from autonomy. <i>Avant: Trends in Interdisciplinary Studies</i>, <i>5</i>(2), 98–112. https://doi.org/10.26913/50202014.0109.0005</span></li>
<li><span id="gabriel:2018">Gabriel, M. (2018). <i>Der Sinn des Denkens</i>. Ullstein Buchverlag.</span></li>
<li><span id="tomasello:2024">Tomasello, M. (2024). <i>Die Evolution des Handelns</i>. Suhrkamp.</span></li>
<li><span id="friston:2011">Friston, K. J. (2011). Embodied inference: or “I think therefore I am, if I am what I think“. In W. Tschacher &amp; C. Bergomi (Eds.), <i>The Implications of Embodiment (Cognition and Communication)</i> (pp. 89–125). Imprint Academic.</span></li></ol>]]></content><author><name>Benedikt Zönnchen</name><email>benedikt.zoennchen@web.de</email></author><category term="ML" /><category term="Neuroscience" /><summary type="html"><![CDATA[Imagine you see a grizzly bear in the woods and half of its face is hidden behind a tree. You will most certainly recognize the bear as whole and dangerous animal. Your brain will “fill in the gaps” even though there is no absolute certainty that the bear isn’t split in half. How and why does the brain do this? Furthermore, why are we fooled by all sorts of graphical illusions? In other words, why are we, in some instances, so stubborn to see what is not there—even if we are told that it is not there—and, on other occasions, we can immediately see what is hidden and probably there?]]></summary></entry><entry><title type="html">Crises of Communication: Sustainability and Trumpism</title><link href="https://bzoennchen.github.io/Pages/2024/10/30/a-crisis-of-non-communication.html" rel="alternate" type="text/html" title="Crises of Communication: Sustainability and Trumpism" /><published>2024-10-30T00:00:00+01:00</published><updated>2024-10-30T00:00:00+01:00</updated><id>https://bzoennchen.github.io/Pages/2024/10/30/a-crisis-of-non-communication</id><content type="html" xml:base="https://bzoennchen.github.io/Pages/2024/10/30/a-crisis-of-non-communication.html"><![CDATA[<p>I have to admit, I’m somewhat addicted to thinking about <em>social systems theory</em>, particularly the version developed by the German sociologist Niklas Luhmann (1927–1998). 
But this addiction isn’t driven by pure fascination—it’s more of a love-hate relationship. 
On one hand, seeing the world through a Luhmannian lens doesn’t make it more just or fantastic, but it does make it comprehensible. 
The chaotic state of global affairs—the craziness of U.S. elections, the brutal devastation in the Middle East, the war in Europe, and the looming tensions brought by the climate crisis—appears senseless at times, even apocalyptic.
And yet, <em>systems theory</em> offers a framework that, paradoxically, brings coherence to this apparent madness, showing us how these outcomes emerge from the logic of <em>functionally differentiated systems</em>.</p>

<p>Faced with such overwhelming disorder, how can anyone retain a sense of sanity? 
We can protest, advocate for change, and try to amplify ‘reasonable’ voices, but as we’ll see, whether these efforts have the desired effect is largely out of our hands.
One might speculate that a Trump victory would worsen the Middle East conflict (and the conflict in Ukraine), but, as citizens of the world, we have no ballot for de-escalating global tensions.
People outside the U.S. cannot vote in American elections, and even if Kamala Harris were to win, there’s little reason to assume that U.S. policy would suddenly prioritize humanitarian investment in preserving lives and infrastructure abroad—the very foundations of a decent existence.</p>

<p>The other option is to try to make sense of the seemingly senseless, not to justify what’s happening but to deepen our understanding and foster empathy, even with those we might otherwise label as perpetrators of harm. 
To ease tensions one has to move beyond simplistic judgments of good and evil. 
Strangely enough, Luhmann’s <em>anti-humanistic</em> theory might aid in this endeavor, as it places individuals outside the bounds of society, such that we can shift our blame from <em>souls to systems</em>.
In Luhmann’s view, social dynamics operate autonomously, driven by complex systems that rarely align with individual intentions.
In many ways, Luhmann’s thinking aligns with the ideas of several twentieth-century French theorists, such as Baudrillard, Foucault, Deleuze, and Derrida.
However, his approach is less dramatic and, in a stereotypically German way, more detached and methodical.
While he shares many of the French theorists’ insights into power, society, and structural dynamics, he refrains from moral interpretations, neither labeling these dynamics as inherently good nor bad. Instead, he offers a distant, almost clinical description of society—a detached analysis that seeks to understand social mechanisms without prescribing judgment.
His perspective allows us to look past personal blame, seeing <em>dysfunction</em> not as a failure of character but as a product of systemic logic that no one person controls.</p>

<h2 id="beyond-good-and-evil">Beyond Good and Evil</h2>

<p>Clearly, attempting to make sense of complex issues should not be mistaken for rationalization.
This is an easy trap to fall into, particularly when the topic is heavily charged with moral language, where any effort to explain events can be (willingly) misinterpreted as either justification or condemnation. 
The framework within which sense-making occurs in such cases often defaults to the age-old binary of good versus evil—arguably the most effective, yet oversimplified, way to reduce complexity.
This moral framing gives us a manageable lens through which to view the world, but it also makes us blind for a more nuanced perspective, risks obscuring deeper systemic dynamics and can hinder genuine understanding of the complex interactions at play.
As a German comedian once said:</p>

<blockquote>
  <p>If you know who is the devil, your day is already well-structured.</p>
</blockquote>

<p>Most would agree that this binary division of good versus evil is overly simplistic and, at times, dangerous.
Yet, when we look toward the U.S. election or the language used in the current horrific conflict in the Middle East, we see a striking example of this polarization in action—a place where the framing of social and political conflicts as battles between absolute good and evil has become more pronounced than ever.
Especially the situation in Palestine is extremely hard to swallow without being overwhelmed by emotions and breaking down in tears.
Here the moral dichotomy shows its face and its power.
It is so dangerous because to defeat absolute evil everything is permitted.</p>

<p>I have the luxury to shift my perspective from judging individuals as good or evil to evaluating systems as either <em>functional</em> or <em>dysfunctional</em>.
Instead of moralizing, I can try to focus on understanding the underlying structures and processes that contribute to societal challenges.
Someone directly affected by these conflicts probably cannot.</p>

<p>Thinking in terms of systems brings a certain relief from the confusion and frustration of modern life.
It helps to lessen anger and bewilderment about why our <em>life-world</em> and the decisions made within it often seem so irrational or even absurd.
Systems theory sheds light on why, even in an era where nearly everyone can participate in media production, we have neither reduced manipulation nor fostered a more reasonable dialogue. 
In many ways, the Enlightenment’s aspirations for rational discourse and universal truth have not materialized as hoped. 
The ideal of ‘Truth’—which Plato connected to the ‘Good’—has, it seems, drifted into obscurity and we are left with a spectacular hyperreality; it seems we have been fallen deep into the cave of shadows.</p>

<p>In a typical postmodern move, Luhmann’s conclusion to the ‘lost Truth’—understood as singular objective truth—is (similar to Baudrillard) that it never existed in the first place.
As a constructivist, he avoids Plato’s concept of ‘the Truth’ and shifts his attention to the <em>production of sense</em> via different systems.
If we take Luhmann’s theory seriously, we must recognize that controlled, predictable change within society is extremely limited because each (social) system constructs its own reality.
There is no agreement on what is true or real.
Fundamentally, many problems arise from <strong>the difficulty of communication</strong>.
Adopting this perspective introduces a sense of helplessness because even with well-intentioned or radical actions, there is no guarantee that the outcomes will align with the respective intentions.
As individuals, we find ourselves positioned outside the social systems that coevolve with their environment according to their own complex dynamics.
This view is both awe-inspiring and disquieting: I admire the explanatory power of Luhmann’s theory, yet I feel a deep urge to challenge or even disprove it.
It confronts us with the unsettling notion that society evolves autonomously, beyond our direct influence, regardless of our individual ideals and aspirations.</p>

<p>I first encountered Luhmann’s theory while preparing a lecture on sustainable artificial intelligence. 
In researching future competencies, including sustainability competencies <a class="citation" href="#Brundiers2020">(Brundiers et al., 2020)</a>, I noted that systems thinking is emphasized as a fundamental skill for addressing complex issues. 
However, I doubt that Luhmann’s work appears on the reading lists or syllabi of most courses on sustainability.
Outside of Germany, he remains relatively unknown for a few reasons. His writing, e.g.,</p>

<ul>
  <li>Die Wissenschaft der Gesellschaft (The Science of Society) <a class="citation" href="#luhmann:1992">(Luhmann, 1992)</a></li>
  <li>Die Wirtschaft der Gesellschaft (The Economy of Society) <a class="citation" href="#luhmann:1994">(Luhmann, 1994)</a></li>
  <li>Die Kunst der Gesellschaft (The Art of Society) <a class="citation" href="#luhmann:1997">(Luhmann, 1997)</a></li>
  <li>Die Politik der Gesllschaft (The Politics of Society) <a class="citation" href="#luhmann:2002">(Luhmann, 2002)</a></li>
  <li>Die Gesllschaft der Gesellschaft (The Society of Society) <a class="citation" href="#luhmann:1998">(Luhmann, 1998)</a></li>
</ul>

<p>is notoriously technical and repetitive, and his theory clashes with the Western concept of the <em>sovereign individual</em> <a class="citation" href="#moeller:2011">(Möller, 2011)</a>.
Note that each title has a double meaning, e.g. <em>The Science of Society</em> discusses the social system called science but it is also a specific description written by society, hinting at the fact that there is no perspective from outside.
Consequently, <em>The Society of Society</em> is a self-description of society.</p>

<p>Additionally, Luhmann sidesteps moral language and offers no prescriptive or normative framework.
Those looking to his theory for answers on what to do will likely be disappointed.</p>

<p>Today, <em>systems thinking</em> is widely discussed as a method to grasp the complexity of global issues by focusing on wholes and relationships rather than dissecting problems into isolated parts. 
It seeks to move beyond Cartesian reductionism and the Newtonian view of linear cause-and-effect, proposing instead an anti-reductionist approach that emphasizes interdependencies, especially crucial in fields like climate science. 
Here, we are not dealing with a computable universe but with complex and chaotic systems, where non-linearity and feedback loops disrupt straightforward causal relationships. 
However, we should remind ourselves that chaos is different from randomness. 
Chaotic systems can exhibit intricate structures, yet the slightest change in initial conditions can drastically alter future outcomes, as we see with weather systems—a classic examples of chaotic behavior.
While accurate short-term predictions are challenging due to the chaotic nature of weather, long-term averages, such as the global average temperature, can be predicted with reasonable accuracy.
Unlike weather, which is highly sensitive to initial conditions, climate trends respond more predictably to persistent external drivers like greenhouse gas concentrations and solar radiation. 
However, when we consider societal factors, the picture becomes more complex, as human activities and policy decisions can significantly influence these long-term climate trends.</p>

<p>In essence, <em>systems thinking</em> itself is a kind of <em>technology</em>, and many hope it will equip us to address the <em>climate crisis</em>.
However, I believe there are distinct schools of thought within <em>systems thinking</em>, each relying on different levels of abstraction, and they are not necessarily compatible.
If systems thinking is indeed essential for tackling the <em>climate crisis</em>—a hypothesis I support—then it stands to reason that we should understand the social dimensions of the climate crisis and other global issues through the lens of one of sociology’s most sophisticated systems thinkers, that is arguable, Niklas Luhmann. His framework provides a unique approach to examining the complex, interdependent nature of social systems that underlie and influence our responses to the <em>climate crisis</em>.</p>

<h2 id="so-what-is-a-system">So what is a System?</h2>

<p>The highly abstract term <em>system</em> is so loose and overused in so many contexts and in our daily language that it has hardly any specific meaning.
We can talk about computer systems, a system of linear or differential equations, a system of thinking, systems of oppression, ecosystems and political systems.</p>

<blockquote>
  <p>[So] what is a system? 
A system is a set of things […] interconnected in such a way that they produce their own pattern of behavior over time. […] 
[T]he system’s repsonse to these forces is characteristic of itself, and that repsonses is seldom simple in the real world. – <a class="citation" href="#meadows:2008">(Meadows, 2008)</a></p>
</blockquote>

<p>For the environmental scientist Donella Meadows (1941–2001) a system is basically <strong>an interconnected set of elements that is coherently organized in a way that achieves something</strong>.
A system is characterized by three key components:</p>

<ol>
  <li><strong>Elements</strong>: The parts or components of the system, such as individual actors, objects, or variables.</li>
  <li><strong>Interconnections</strong>: The relationships or interactions between the elements, often in the form of flows of information, energy, or material.</li>
  <li><strong>Purpose or Function</strong>: The overarching goal or behavior that the system is organized to achieve.</li>
</ol>

<p>But here the trouble begins because Luhmann defines a system very differently.
It seems to me that Meadows’ definition still relies on the subject-object distinction which Luhmann wants to sublime.
He thinks in interdependent but operationally closed processes instead of things and he very much dislikes the concept of an externally given ‘purpose’.
For Luhmann</p>

<blockquote>
  <p>A system is a self-referential, self-organizing set of operations that differentiates itself from its environment.</p>
</blockquote>

<p>This requires some explanation:</p>

<ol>
  <li><strong>Self-Referential</strong>: Systems create their own elements through their own operations. In fact, they are operations. In social systems, these elements are not people but <em>communications</em>. Each communication refers back to the system, reaffirming its boundaries and identity.</li>
  <li><strong>Autopoiesis</strong>: Systems are autopoietic, meaning they are self-producing. They continuously reproduce the <em>communications</em> that sustain them, distinguishing themselves from their environment. This process enables a system to maintain coherence and adapt to changes.</li>
  <li><strong>Environment and Differentiation</strong>: Systems are defined by the distinction between themselves and their environment. Luhmann stresses that the environment is everything that the system excludes, setting clear boundaries. This differentiation allows the system to maintain its <em>identity</em> while interacting with, but remaining distinct from, external influences.</li>
  <li><strong>Social Systems as Sense Making Networks</strong>: Luhmann focuses on social systems—such as organizations, institutions, the economy, the political system, and the mass media—as networks of meaning/sense making (‘<em>Sinn machen</em>’ in German). In these systems, communication itself is the fundamental element, and these <em>communications</em> build the system’s reality.</li>
</ol>

<p>In essence, Luhmann views a system as a closed network of <em>communications</em> that operates independently of external elements and functions primarily by sustaining itself through self-generated, meaningful <em>communications</em>.
If we want to be accurate we can not speak of a system without its environment because <strong>a system is the process that differentiates itself from its environment</strong> which is a circular definition—a paradox—that keeps the system (the system-environment differentiation) going.
This definition contrasts sharply with definitions based on tangible components and external goals, focusing instead on processes of <em>sense-making</em>, self-production and self-maintenance.</p>

<p>Luhmann read a lot of interdisciplinary material and borrowed from mathematics, classical systems theory, biology, cybernetics and other disciplines.
For example, he took the concept of <em>autopoiesis</em> <a class="citation" href="#maturana:1987">(Maturana &amp; Varela, 1987)</a> from biology, the concept of <em>feedback loops</em> and <em>second-order observation</em> from cybernetics and of <em>re-entry</em> and the fundamental operation of <em>differentiation</em> and <em>indication</em> from the mathematician Spencer-Brown <a class="citation" href="#brown:1969">(Spencer-Brown, 1969)</a>.</p>

<p>For Luhmann, systems are <em>operationally closed</em> meaning that no system can interfere in the operation of another system.
For example, the economic system communicates via payments and there is almost no way that the political system can interfer in it.
Of course, the political system can observe these payments and can try to regulate them but only indirectly.
It can pass laws which the legal system processes.
The economic system will observe this—it will ‘digest’ it—and evolve with its environment (which contains the political system).
Therefore, systems are <em>cognitively open</em> meaning that they can observe (based on their own logic) their environment which contains all the other systems.
They take everything in what they are able to digest and use it to continue their opertions, that is, their <em>autopoiesis</em>.
An analogy is a human body that takes in food and digist it in the way it is able to.
Neither does the food determine how the body is affected nor does the body can make anything it wants from the food.
The process is contingent.</p>

<p>To reduce complexity social systems work on simple <em>binary codes</em> under which they <em>differentiate</em>.
The legal system interprets actions as legal or illegal but doesn’t engage with the healthy/unhealthy distinctions from the health system.
The following table shows more of these codes:</p>

<table>
  <thead>
    <tr>
      <th><strong>Social System</strong></th>
      <th><strong>Binary Code</strong></th>
      <th><strong>Description</strong></th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><strong>Economy</strong></td>
      <td>Payment / Non-payment</td>
      <td>Decisions are guided by whether a transaction involves payment, focusing on economic exchanges.</td>
    </tr>
    <tr>
      <td><strong>Politics</strong></td>
      <td>Power / Non-power</td>
      <td>Concerned with the distribution and exercise of power, focusing on who has authority and control.</td>
    </tr>
    <tr>
      <td><strong>Law</strong></td>
      <td>Legal / Illegal</td>
      <td>Operates on legality, determining if actions or behaviors align with established legal norms.</td>
    </tr>
    <tr>
      <td><strong>Science</strong></td>
      <td>Truth / Falsehood</td>
      <td>Guided by the pursuit of truth, evaluating claims based on their validity and scientific evidence.</td>
    </tr>
    <tr>
      <td><strong>Religion</strong></td>
      <td>Immanence / Transcendence</td>
      <td>Focuses on distinctions between the sacred (transcendent) and the profane (immanent).</td>
    </tr>
    <tr>
      <td><strong>Education</strong></td>
      <td>Success / Failure</td>
      <td>Concerned with the effectiveness of learning and teaching, evaluated by success in achieving educational goals.</td>
    </tr>
    <tr>
      <td><strong>Health</strong></td>
      <td>Healthy / Unhealthy</td>
      <td>Operates based on the state of health, determining whether a body or behavior is healthy.</td>
    </tr>
    <tr>
      <td><strong>Mass Media</strong></td>
      <td>Information / Non-information</td>
      <td>Distinguishes between what is considered newsworthy (informative) versus uninformative content.</td>
    </tr>
    <tr>
      <td><strong>Art</strong></td>
      <td>Fitting / Unfitting</td>
      <td>Focused on e.g. aesthetic value, distinguishing what is perceived as beautiful or aesthetically valuable.</td>
    </tr>
  </tbody>
</table>

<p>The specific code a system operates under is less important than the fact that each code differentiates one system from others.
For instance, Luhmann struggled to pinpoint a definitive code for the art system, as its operations are complex and multifaceted.
He proposed several possibilities, including beautiful/ugly, coherent/incoherent, new/old, and fitting/unfitting.</p>

<p>Each code of a system gives the system its ‘character’ and consequently its operational specificity and functional closure.
The <em>exclusivity</em> of the <em>binary code</em> ensures that each system maintains its autonomy and operates independently, even when interacting with other systems.
<em>Operational closure</em> means that each system can only process information according to its own internal logic, thus keeping it <em>closed</em> to other systems’ codes and distinctions.</p>

<p>Of course, further distinctions within a system are possible.
For example <em>reputation</em> is an important distinction within science.
While not a binary code in Luhmann’s strict sense, it is significant within the scientific community as it affects how research and findings are perceived, valued, and disseminated.
Reputation can influence which scientists’ work is taken seriously, whose research is funded, and which publications are more widely read and cited.
However, it does not drive the core distinction of truth/falsehood; rather, it shapes the social hierarchy, credibility, and visibility within the scientific community.</p>

<p>Luhmann’s concept of <em>re-entry</em> (borrowed from <a class="citation" href="#brown:1969">(Spencer-Brown, 1969)</a>) is a mechanism, allowing a system to reflect on itself by reintroducing its primary binary code within its own operations. 
Thus, this is an inherently recursive relation.
<em>Re-entry</em> enables a system to apply its guiding binary distinction not only outwardly (to its environment or other systems) but also inwardly, to its own internal processes and <em>communications</em>. 
This is crucial in complex systems, like science, where re-entry enables self-reference and internal differentiation.
Science can use the truth/falsehood distinction to evaluate not only external hypotheses but also its own standards, research paradigms, and accepted theories.
Another more familiar example is the media which reports on itself.</p>

<p>Another important Luhmannian concept that is strongly connected to <em>re-entry</em> and orignated from Heinz von Foerster (1911–2002) and Margaret Mead (1901–1978) <a class="citation" href="#foerster:2003">(von Foerster, 2003)</a> is <em>second-order observation</em>.
<em>Second-order observation</em> is facilitated by re-entry, as it allows the scientific system to reintroduce its primary distinction, i.e. truth/falsehood, internally.
It refers to observing observations rather than simply observing objects or phenomena directly.
This concept is crucial for complex systems as it allows them to recognize and reflect on how they construct their own distinctions and interpretations.
In science, second-order observation enables scientists to observe not only external phenomena but also the methods, theories, and interpretations of other scientists.
This includes observing how truths are constructed within the scientific community and scrutinizing the frameworks, biases, and assumptions underlying those constructions.
But it also includes how a specific scientist or science lab is being observerd.
In this context, <em>reputation</em> is effectively a measure of how a scientist is seen by other scientists.
It further reduces complexity by accumulating the observation of others.
I do not have to read and carefully analyse every paper of a specific researcher to find out if his or her research is trustworthy, i.e. if it is good research under the truth/falshood code.
In complex environments, <em>second-order observation</em> helps systems like science to deal with uncertainty and complexity.
Rather than aiming for <strong>absolute certainty</strong>, science can adapt by recognizing different observational frameworks, revisiting previously accepted truths, and acknowledging limitations in current knowledge. 
This adaptive flexibility, achieved through <em>second-order observation</em>, is vital for science’s resilience and continued evolution.
Today, second-order observation is everywhere, be it in the form of the housing or stock market, the social phenomena of <em>reaction videos</em> or the fact that we are all invested in our profiles, that is, <strong>we are invested in how we are seen/observed by an anonymous peer</strong> <a class="citation" href="#moeller:2021">(Möller &amp; D’Ambrosio, 2021)</a>.</p>

<p>In the mode of authentic identity construction, second-order observation appears as the production of ‘fake’ because the focus shifts from who one <strong>is</strong> to how one <strong>is observed</strong>.
Paradoxically, as we transition from <em>authentic</em> to <em>profilitic</em> identity construction, the most important goal becomes to be observed as authentic.
This phenomenon is evident in the behavior of streamers, social media influencers (including figures like Trump and Musk), as well as companies and even institutions such as universities.
All of these entities are heavily invested in how they are observed, striving to appear authentic and real, rather than fake.
However, if we ask the existentialist question—<em>What is your ‘true self,’ your ‘authentic being’?</em>—we encounter paradoxes and articulation problems.
Luhmann provides an uncanny response to this authenticity problem, reminiscent of Kant: What one is, is always already a description or observation of a system, and therefore, a selection or differentiation—choosing one thing over another.
This includes self-observation, which is itself a re-entry of the system. In this case, the psychic system re-enters itself.
Through Luhmann’s perspective, we are left with the impossibility of being truly authentic. In some sense, we (and other systems) are always already pretending.
However, Luhmann crucially does not view this as inherently negative. There is nothing—especially morally—wrong with being invested in how one is observed.
The problem arises when this focus on appearence no longer aligns with the system’s function.
For example, when the pursuit of appearing as a trustworthy scientist contradicts the truth/falsehood distinction that underpins the operation of science.</p>

<p>But what exactly is <em>observation</em>?
For Luhmann, every act of <em>observation</em> has two essential steps: <strong>distinction</strong> and <strong>indication</strong> (also borrowed from <a class="citation" href="#brown:1969">(Spencer-Brown, 1969)</a>).
The system first creates a <em>distinction</em> (e.g., legal/illegal in the legal system) and then <em>indicates</em> one side of that distinction. 
This act of <em>indicating</em> one side of a <em>distinction</em> allows the system to focus on what it deems relevant or meaningful while leaving out what is not.
For instance, the economic system distinguishes between payment and non-payment and then indicates whether a transaction falls into one category or the other.</p>

<p>Luhmann also uses <em>feedback loops</em> but in a more complex, indirect way to explain self-referential and <em>autopoietic</em> (self-producing) processes within social systems.
<em>Feedback loops</em> enable systems to observe and respond to their own operations and their environment without sacrificing their internal logic or autonomy.
A system produces <em>communications</em> and then feeds those back into itself as input, creating a recursive process. 
For instance, the scientific system continually generates new research findings that become part of its ongoing discourse, which shapes further research questions and methods.
This recursive process enables a system to build on its own operations and maintain continuity over time.
<em>Feedback loops</em> help systems to learn from past operations and adjust future <em>communications</em>. 
However, instead of direct feedback that leads to specific, immediate corrections (as in a thermostat, for instance), feedback in Luhmann’s theory involves observing patterns over time and <em>adjusting structurally</em>.
For instance, the legal system may notice shifts in societal values based on case outcomes or public reactions and eventually adjust interpretations of the law, but it does so in a way that remains consistent with its legal/illegal <em>binary code</em>.
Through <em>repeated feedback</em>, systems can detect trends in their environment (e.g., shifts in public opinion or technological advances) and adapt their operations in response, but only when those trends become relevant within the system’s own code.
Despite the use of <em>feedback loops</em>, systems remain <em>operationally closed</em>.
Feedback is processed in terms of the system’s unique code, meaning that only information relevant to that code is taken in. 
For instance, if the economic system receives feedback from the political system, it only integrates that feedback if it pertains to the payment/non-payment distinctions. 
This allows each system to interact with its environment while preserving its autonomy and self-referential logic.
Feedback loops help systems manage the complexity of their environments by selectively processing information.
Each system filters out what is irrelevant to its operations, thereby creating <em>blind spots</em>.</p>

<p>Luhmann’s concept of <em>structural coupling</em> (borrowed from <a class="citation" href="#maturana:1987">(Maturana &amp; Varela, 1987)</a>) describes how different social systems develop stable, interdependent relationships with their environment or other systems without losing their <em>operational closure</em> or <em>autonomy</em>.
<em>Structural coupling</em> is essential for maintaining a productive relationship with various other systems.
For example, science and politics are structurally coupled when scientific research informs policy decisions, while political priorities influence the direction and funding of scientific research. 
Despite this interaction, both systems remain <em>operationally closed</em>: science focuses on truth/falsehood, and politics operates on power/non-power.
Another example: the economic system and the political system may engage in <em>structural coupling</em>, where economic data (like inflation) affects political decisions (like interest rate changes). 
<em>Feedback loops</em> are crucial for <em>structural coupling</em> since they allow each system to remain sensitive to changes in the other system without changing its fundamental operations or logic.</p>

<p>The last term we have to discuss is the term maybe most important, and that is <em>communication</em>.
As a computer scientist I understand communcation using Shannon’s framework <a class="citation" href="#shannon:1948">(Shannon, 1948)</a>, that is, a transfer of information over an error-prone or noisy channel.
But this is not what Luhmann understands as <em>communication</em>.
For him, <em>communication</em> is not merely the transfer of information between individuals (or machines). 
Instead, it is a self-contained social process that occurs within and is produced by social systems, with <strong>individuals seen as part of the environment rather than as agents within the system</strong>.
According to Luhmann, it is a three-part process of three interdependent elements/<strong>selections</strong>:</p>

<ol>
  <li><strong>Information</strong>: The content or ‘<strong>what</strong>’ of communication, which could be new data, knowledge, or ideas relevant to the system.</li>
  <li><strong>Utterance</strong>: The ‘<strong>how</strong>’ of communication, which includes the form, manner, or medium through which information is expressed. This could be spoken language, writing, or nonverbal cues, depending on the medium and context. E.g. a payment is a communication.</li>
  <li><strong>Understanding</strong>: The receiver’s interpretation of both the information and the utterance. Understanding is crucial, as it determines whether and how the communication is taken up within the system. It is the <strong>differentiation between information and utterance</strong>.</li>
</ol>

<p>For Luhmann, communication only ‘happens’ if all three elements are present. It’s not just about sending (selected) information; it’s about how information is expressed and then understood within a particular context.
<em>Communication</em> is not generated by individuals but by the system and is itself autopoietic, meaning it is self-producing and self-sustaining.
<em>Communication</em> is inherently selective; it involves making choices about what information to include, how to present it, and how to interpret it. 
This selectivity creates <em>blind spots</em>, as each communication inherently excludes other possible meanings.
It is based on the concept of <em>double contingency</em>—the idea that each party in a communication anticipates and adjusts to the other’s responses.
Social systems manage this <em>contingency</em> through established expectations.
For example, in the legal system, there is an expectation that communications follow the legal/illegal code, which guides interactions between lawyers, judges, and citizens and maintains coherence in legal decisions.
Or take the education system.
If I start singing in my lecture students would be quite confused.
Since in Luhmann’s framework every system operates by its unique logic, <em>communication</em> can <strong>not</strong> be about transferring <strong>objective</strong> information; it is about processing meaning/sense which depends on the ‘processor’, i.e. the system. 
Each <em>communication</em> within a system adds to the meaning that the system produces.
For instance, in the scientific system, each new theory or finding creates meaning within the context of the truth/falsehood code and is interpreted within that framework—it can change <em>the reality of science</em> as a whole.</p>

<blockquote>
  <p>Systems exist and they continuously build their own specific reality, shaped by the types of meaning they process through communication.</p>
</blockquote>

<p>Consequently, if we take Luhmann serious, we arrive at the revelation that there is not one ‘really real and objective reality’ but that there is a <strong>plurality of realities</strong>;
that there is not one controlling system that steers all the others but that there is <em>anarchy</em> in society;
that we are not in control but that society is <strong>out of control</strong>;
that there is not one objective truth but systemic interpretations;
and maybe most importantly: that there is no outside of society, no view at the whole because <strong>any indication of something requires a distinction from something else</strong>.</p>

<h2 id="part-i-a-crisis-of-non-communication">Part I: A Crisis of Non-Communication</h2>

<p>One area where Luhmann’s theory seems to make unsettlingly accurate sense is the <em>climate crisis</em>—that is, the accelerating destabilization of the climate system.
We see how Luhmann’s framework reveals the challenges of a complex, multi-systemic problem. 
Each social system—politics, economy, science, and media—observes the climate crisis from within its own operations and distinctions.
Science operates on a truth/falsehood basis, producing reports on climate change’s reality and projections, while politics, operating on power/non-power, assesses climate issues according to political priorities, public opinion, and election cycles. 
The economic system, structured by payment/non-payment, may respond to climate science only insofar as it impacts financial markets, investments, or regulatory demands.</p>

<p>The <em>climate crisis</em>, however, does not ‘belong’ to any one system.
Instead it spans across systems but is refracted through each one’s unique code.
No single system can comprehensively address the crisis because its complexity exceeds the logic of any one system’s operations.
The climate crisis is precisely so problematic for the communication network we call society because each system deals with its environment by reducing the environment’s complexity.
Furthermore, there is essentailly no <em>climate communcation</em> ‘happening’, because (to the best of my knowledge) there is no <em>social system</em> that operates on a code that leads to the observation of the climate or the earth’s ecosystem.
One might step in and argue that science certainly observes the climate but that is not really the case if we use Luhmann’s definiton of observation and communication.
Science, despite its close engagement with ecological and climate issues, operates on a fundamentally different basis than a (hypothetical) climate-focused system.
While many scientists care very much about the climate and the survival of human beings, science operates under the truth/falshood distinction.
Its observation of the climate does not directly lead to climate or political activism, or an economic transformation but to more truth/falshood distinction;
to research and funding opportunities and the building up of reputation.
In fact, some scientist such as Ulf Büntgen are concerned about scholars who are at the same time activists because they might damage the operation of science <a class="citation" href="#buentgen:2024">(Büntgen, 2024)</a>, others argue against these worries <a class="citation" href="#eck:2024">(van Eck et al., 2024)</a>.
From a system’s viewpoint, we should not attribute to much influence to the individual since, again, there is no individual within society.
Science can ‘use’ the respective psychic system as well as activism.
At the same time, while activsim cannot steer science, science digest/observes activism by its own operations to preserve its operating.</p>

<p>Similarly, the political system interprets the <em>climate crisis</em> through its own operational code of power/non-power, using it as an opportunity to gain influence and public support. 
For example, in Germany, the Green Party views the climate crisis as a platform to expand its political reach.
They advocate for environmental policies that resonate with their voter base.
However, they (as a system, not as individuals) do this to gain power and not to solve the climate crisis.
Other parties may leverage the crisis in the opposite direction, appealing to constituents who prioritize economic stability over environmental reform.
Yet, even if the Green Party succeeds in passing policies aimed at accelerating economic transformation, it cannot directly control whether the economic system will fully implement these changes. 
This limitation arises because each system—politics and the economy—operates autonomously according to its own logic.
For politics to dictate economic outcomes would imply that the political system could override the economy’s fundamental payment/non-payment distinction, which, according to Luhmann, is structurally impossible.
Note that this is not a critique of the members of the Green Party which might very much care and belief in the values they display—they might be honestly invested in how they are being observed.
I do not consider individuals or people but systems.</p>

<p>The climate crisis illustrates how <em>structurally coupled</em> systems face limits in their coordination.
Each system’s response to climate issues is conditioned by its own operations, preventing unified action despite the existential threat posed by the destabilizing climate.
As Luhmann’s theory reveals, social systems are inherently self-referential and cannot simply ‘combine’ their functions. 
Thus, the climate crisis may persist without cohesive action precisely because <strong>no system is structured to address an issue that transcends its operational boundaries</strong>.
In fact, if systems cross boundaries, we call it corruption, for example when the economy system pays for a football goal or for a law or if a scientist makes false claims to strengthen a certain political ideology.</p>

<p>This fragmentation of responsibility means that no single system is inherently designed to address global, cross-cutting issues like the climate crisis.
Each system approaches climate issues only insofar as they relate to its own logic:</p>

<ul>
  <li><strong>Science</strong> seeks truth and thus investigates the mechanisms, causes, and projected impacts of climate change.</li>
  <li><strong>Politics</strong> operates on power/non-power, framing climate policies in terms of public support, regulatory reach, and political gain or loss.</li>
  <li><strong>Economy</strong> focuses on payment/non-payment, assessing climate initiatives based on profitability and market viability.</li>
  <li><strong>Media</strong> works with information/non-information, spotlighting climate issues based on newsworthiness rather than scientific rigor or policy relevance.</li>
</ul>

<p>Systems only engage with climate issues when these issues align with their internal priorities.
Science can produce overwhelming evidence of climate risks, but if political decisions are driven by short-term voter approval, economic costs, or geopolitical interests, the full implications of scientific knowledge may not translate into concrete action.
Political decisions are often informed by scientific findings but are ultimately filtered through the political logic of power/non-power.
Economies can adopt sustainable practices, but often only when such practices promise financial returns.
This misalignment is evident in the delays or dilution of climate policies, where political and economic interests often override scientific findings.</p>

<p>Luhmann’s theory suggests that issues of this magnitude may require a dedicated system with its own <em>binary code</em>—such as ‘sustainable/unsustainable’ or ‘ecologically balanced/unbalanced’—to assess and act on climate issues directly. 
Social systems are so immensly effective (with respect to their function) because they reduce complexity by operating under a rather simple binary code.
A new <em>sustainable system</em> could facilitate large enough irritations that resonate within other systems such that standards, goals, and measures across all systems are implemented (by themselves) in such a way that align with ecological stability and sustainability.
Although hypothetical, this system would prioritize ecological concerns by constructing ‘sustainable communication’.
Of course, the problem is: How can such a communication be possible without interferring in, e.g. economic communication?</p>

<p>Howsoever, the time is up and I have almost no hope that such a system will suddenly emerge.
If it does it should happen via the seperation by operational closure similar to how, for example, art was able to separate itself from religion.
The only other option, using Luhmann’s framework, is to try to align all the systems in such a way that if they operate according to their logic, they also operate (at least close) to the logic of such a hypothetical <em>sustainable system</em>.
Caring about the earth’s ecosystem and the climate has to be financially profitable;
it has to lead to power for politicians;
it has to lead to funding and furhter research in science;
and it has to be a spectacle for the media to cover;</p>

<p>Naturally, we might see it as hypocritical when a company adopts sustainable practices primarily to increase profits rather than out of genuine concern for the environment.
We are back to the <em>authenticity problem</em> mentioned earlier.
And, indeed, companies may choose to appear sustainable rather than enact substantive changes if it proves more profitable.
This dynamic holds for other issues as well, such as diversity and inclusion—and deep inside we all know it.
However, from a systems theory perspective, it may be more productive to move beyond <em>moral judgments</em> about individuals, pointing to them as being hypocritical, virtuous, or evil and instead to focus on  <em>systemic realities</em>.</p>

<p>In the economic system, sustainability, diversity, and other social values are interpreted through the code of payment/non-payment; the system evaluates decisions based on profitability rather than intrinsic ethical value.
We may not like it but that’s the <em>reality of the economic system</em>.
For example, instead of hoping for an <em>humanistic turn</em> of the economic system, it might be more effective to achieve social progress for disabled people by letting the economic system ‘know’ how these people are financially important.
While individuals within a company might personally care deeply about these issues, this personal commitment does not translate into the company’s operations, as employees are part of the system’s environment, not its core functions.
People care, systems observe and operate on their own terms.</p>

<p>In this light, <em>truthfully pretending</em>—adopting sustainability practices for economic gains—can still yield positive outcomes.
From the perspective of systems theory, the motivations behind these actions matter less than the fact that they result in more sustainable practices, which in itself contributes to broader societal goals.
This does not imply that public outrage about the state of affairs is misguided.
On the contrary, if outrage <em>irritates</em> systems in ways that make sustainable practices more profitable for companies, it can drive meaningful change.
However, outrage can also produce unintended effects.
For instance, the media, which constructs a shared reference reality that shapes public discourse, may find it more sensational to focus on the ‘unlawfulness’ of protesters.
This framing can prompt the political system to respond by mobilizing power against the protests, potentially reinforcing the very practices that climate advocates aim to change.
Especially when a system’s ability to continue its autopoietic operations is threatened, a strong reaction can be expected.</p>

<p>These nonlinear and indirect effects, often amplified through feedback loops, illustrate the unpredictable and uncontrollable nature of systemic interactions. 
In Luhmann’s terms, feedback loops create complex dynamics within and between systems, making it difficult to foresee or control the outcomes of public reactions, even when intentions are clear.</p>

<h2 id="part-ii-a-crisis-of-over-communication">Part II: A Crisis of Over-Communication</h2>

<p>Luhmann does not oppose elections, but he challenges the common assumption that they express <em>the will of the people</em>. 
This skepticism follows directly from his understanding of society as an uncontrollable network of interdependent social systems, each operating according to its own logic. 
Society, in Luhmann’s view, evolves organically—like a self-reproducing system—and cannot be directly steered or micromanaged by politics.</p>

<p>Politics, in this context, plays a specific role: it makes collectively binding decisions.
Yet these decisions must then be processed and implemented by other systems.
For instance, the legal system creates laws to enforce political decisions, while the economy and even religion may shape how these decisions are interpreted and realized in practice.</p>

<p>Consider the example of childbirth. 
How do different social systems contribute to this event? 
The political system might legislate that abortion is legal.
Religious beliefs may influence whether a person opts for or against it. 
Socio-economic conditions affect whether one can afford to raise a child, and the health system plays a crucial role in medical support. 
Political decisions matter, but they exist within a web of other systems, each with its own influence, sometimes more decisive than politics itself.</p>

<p>The popular narrative around elections is that they make ‘the people’ the foundation of all political power. 
After the election, politicians—servants of the people—are supposed to put the will of the people into action. 
However, according to Luhmann, the idea that the people are the source of all power in a liberal democracy is a myth. 
Instead, the people function more as an audience, much like in a talent show, where they get to elect a winner at specific, pre-determined moments, but do not control the larger system. 
The ‘show’ of politics is a much larger, self-sustaining system, where politicians, the state, and the voters all play their roles and influence one another, but the system ultimately serves itself.</p>

<p>In democratic politics, the state, politicians, and voters are mutually interdependent, each contributing to the reproduction of the political system. 
Elections, therefore, are <em>symbolic procedures</em> in Luhmann’s view.
They confer legitimacy on the political system by symbolically invoking <em>the will of the people</em>, but in reality, such a unified will does not exist.
For example, many people abstain from voting, and a significant portion of the population may not be eligible to vote at all.
Moreover, election outcomes are shaped by arbitrary rules—they are contigent.
In the U.S., for example, the popular vote does not directly determine the outcome; instead, the electoral college decides the presidency, often making a few swing states the key deciders. 
In Germany, government coalitions are typically formed after elections, yet no single voter casts a ballot for the specific coalition that ends up governing.
These complexities highlight how elections, while significant, are far from a straightforward expression of a unified popular will.</p>

<blockquote>
  <p>How did we vote? But did we really vote, or did the people just roll the dice? […] What individuals actually think, if anything at all, when they mark ballots, remains unknown. This alone suffice not to […] conceive of public opinion as the general expression of the opinions of individuals. – <a class="citation" href="#luhmann:2002">(Luhmann, 2002)</a></p>
</blockquote>

<p>Elections seem to achieve the impossible: merging the diverse, individual wills of the people into a singular, cohesive ‘general will’.
For Luhmann, this process is almost magical, as it provides the symbolic foundation upon which liberal democracy rests.
Elections create the <em>illusion of unity and consensus</em>, giving legitimacy to political decisions that, in reality, are based on a highly fragmented and complex societal landscape.
However, as previously mentioned, Luhmann has no problem with this illusion. 
In his view, <em>the miracle of democratic elections</em> is perfectly acceptable—provided it functions effectively for all involved and helps stabilize the political system.</p>

<p>In fact, the <em>symbolic power of elections</em> is essential to maintaining social order, as it grants politics a legitimate mandate without requiring every individual’s direct influence on policy decisions. 
This symbolic function of elections allows the political system to operate independently, without collapsing under the weight of countless individual preferences.
Elections serve to renew the legitimacy of the political system periodically, preventing it from stagnating, while also setting boundaries within which political decisions are accepted, even by those who disagree with the outcomes.</p>

<p>For Luhmann, it is less important that elections genuinely express a collective will, which he considers a fiction, and more important that they fulfill their function: they create a momentary sense of unity and provide a mechanism for the orderly transition of power seemingly melting the individual will of ‘the people’ into the general will. 
As long as elections maintain public confidence in the political process and prevent <em>systemic breakdown</em>, they serve their purpose, not by conveying truth but by ensuring continuity. 
In this sense, elections are not a search for truth but a pragmatic solution to the challenge of political legitimacy in a complex, functionally differentiated society.</p>

<p>Viewing the election as a performance—or as Baudrillard might call it, <em>hyperreality</em>—feels particularly fitting for the spectacle that Americans witness during the election weeks. 
Do these debates between Trump and Biden, or Trump and Harris, genuinely convey new insights or substantive ‘truths’? 
Or are Americans, in many ways, simply the audience to a grand show, swept up in the drama, spectacle, and narrative arcs that these events offer?</p>

<p>There is, however, something distinct and potentially perilous about the American context.
In the U.S., the problem for the political system seems to lie in the crumbling of the illusion of unity and consensus.
The illusion is increasingly undermined by escalating economic and social inequalities.
As living conditions deteriorate for many, regardless of who they vote for, it becomes glaringly apparent that there is no singular ‘general will’ guiding political outcomes. 
The democratic promise that elections merge the will of the people into collective decisions feels hollow when so many are left feeling unrepresented and disillusioned.</p>

<p>This breakdown makes it clear that those voting for Donald Trump are not simply misguided or irrational. 
Many voters feel disconnected from a political system they perceive as indifferent to their realities and struggles.
As one interviewee put it:</p>

<blockquote>
  <p>I agree that Trump is from the billionaire class and that’s all he’s going to work for.
It basically comes down to the lesser of the two evils right now.
I think about the four years he was president.
In my opinion he’s the world’s best crime boss.
And then you see people posting ‘in the arms of Jesus’ like ‘I was persecuted too’, and I think what a bunch of bullcrap.
The system is so captured that you need a crime boss to get out of it.
If it wasn’t Trump I would love to have a working familiy candidate who stands up for the little guy.
The middle class, from day one of this United States, has built the United States, and we are the ones that always get shit on.</p>
</blockquote>

<p>But there’s a second factor at play: a candidate who is so absurd, so obscene, that he disrupts the expected script of political ‘producers’.
Because Trump is taken seriously he (unconsciously) discloses the reality of elections and the political system.
From a systems theory perspective, a showman does what one shouldn’t do: making the big show, the big stage visible and dismanteling one myth but also replacing it with another, far more dangerous one: <em>the deep state</em>.
While the <em>illusion of unity and consensus</em> gives the system stability, Trump’s <em>deep state myth</em> does the opposite and that is the reason why he is dangerous for the political system and probably for the functional differentiated society.</p>

<p>As the interviewee described, Trump is perceived as a <em>red button</em>—a tool voters can press to create enough disturbance within the system that it is forced to respond.
Perceived <em>as the world’s best crime boss</em> any accusation or scandal only feeds into Trump’s persona. 
Trump clearly wants to cross systems’ boundaries.
The interviewee’s perception might not be far from the truth, howerver, if such an irritation is desirable is very questionable.
The level of desperation in the U.S. has become so acute that many are willing to risk an extreme disruption, effectively pushing for an ‘over-irritation’ of the system, hoping it will provoke meaningful change or even a systemic collapse.</p>

<p>In this way, Trump embodies what Luhmann might call an agent of second-order observation—someone who leverages his own media persona to observe and exploit the expectations of the political system, creating feedback loops that intensify rather than stabilize.
While Luhmann argued that elections are primarily a symbolic show, the outcomes of this particular show may indeed carry existential significance.
The stakes are high, and the effects of a Trump victory or defeat could trigger unpredictable reactions. 
What we are witnessing is both dangerous and volatile, and it exemplifies how systemic irritations—if strong enough—can shake the foundations of even the most stable-seeming structures.</p>

<h2 id="criticism-of-luhmanns-theory">Criticism of Luhmann’s Theory</h2>

<p>Luhmann’s theory is not immune to criticism, and, by its own logic, it necessarily contains blind spots. 
One of the most common criticisms is that his theory removes individuals from the core of social analysis.
By focusing on self-referential systems rather than human actors, Luhmann places individuals in the environment of society, not within it.
Human agency, emotions, and individual motivations are neglected and the role of intentional human actions and collective decision-making in shaping societal evolution is minimized.</p>

<p>In addition, its theory lacks a normative direction or ethical foundation.
By avoiding moral or ethical judgments, Luhmann’s systems theory does not offer guidance on what should be done, especially concerning social justice, inequality, or human rights.
His theory seems to be indifferent to power imbalances and fails to address issues of accountability and responsibility within social systems.</p>

<p>Also Luhmann’s concept of <em>operational closure</em> has been criticized for overemphasizing system autonomy.
According to those critics, his perspective ignores the deep interdependencies and interconnectedness of social, economic, and political systems.
They contend that while systems may have unique operations, they are often influenced by each other in ways that Luhmann’s model underestimates.
Marxist thinkers argue that although Luhmann includes the economic system as one of society’s core functional systems, he does not give sufficient attention to the economic forces shaping society, especially those tied to capitalism.
By focusing primarily on the communication logic of payment/non-payment, Luhmann’s theory fails to address the structural inequalities and exploitative dynamics inherent in modern economic systems.</p>

<p>Because Luhmann’s theory emphasizes the self-reproduction of systems, it suggests that systems are largely resistant to intentional change from within.
This perspective might downplay the role of social movements, activism, and democratic engagement as forces that can drive systemic transformation.
Futhermore, by treating power as a form of communication (very different from Foucault’s approach) within the political system rather than a force that operates across systems, critics argue that his approach obscures how power dynamics influence interactions between systems and shape societal outcomes.</p>

<p>My current opinion is that Luhmann’s theory is excellent for the sense-making of society.
It can even clear the fog for making better intentional decisions but it can not provide us with suggestions of what we should do.
But if people want to change oppressive systems, they should know and understand their ‘adversary’.</p>

<h2 id="treatment-of-an-illness">Treatment of an Illness</h2>

<p>In their book <em>The Tree of Knowledge</em> <a class="citation" href="#maturana:1987">(Maturana &amp; Varela, 1987)</a> Maturana and Varela very briefly discuss society.
They think that, unlike cells serving the whole organism, a functioning society should prioritize the needs and well-being of the individual—the orientation should be reversed.
I totally agree.
But can we bring their humanistic viewpoint in line with Luhmann’s <em>anti-humanistic</em> theory?
That would be nice but is probably not in the spirit of the author.
While both views address the organization of society, Maturana and Varela focus on an ideal in which society exists for the benefit of individuals, whereas Luhmann’s theory suggests that society, as a system, operates rather independently of individual well-being.
We could evaluate societies (across time and space) according to the degree they served individual well-being and then learn from the lessons to irritate our society in such a way that it becomes more <em>functional</em>, in Maturana’s and Varela’s sense of the word.
However, operationalizing this perspective would face challenges within Luhmann’s framework.</p>

<p>As a scientific theory communicated by the science system, it is a theory of society that society produced about itself—a quintessential example of second-order observation.
The individual—the person, psychic system, and living body—we identify as Luhmann existed only as part of the environment of the social systems he studied.
According to his own theory, his view, as every view, cannot be the objectively correct description of society.
Therefore, I think, we should not take his ideas as absolutes but as irritations and sense-making foundation.</p>

<p>If the individual is not the center of social structures and dynamics, the impossibility of controlled action might give some relief on an individual/personal level.
It emphasizes that individual human beings are observers of a greater force acting on them and since society is out of control, we are also out of control.
We can observe society critically from an <em>ironic distance</em>, being <em>carefree</em> without being <em>careless</em> (as individuals) and without being fully subsumed by it, a perspective that may be worth remembering—illuminating neither hope nor fear.
We can look at us and our fellow human beings as beings thrown into the fabric of society.
By pointing to the limits of the individual we can rediscovering a sense of innocence and grace of the human being.</p>

<p>One of Luhmann’s most influential critics, Jürgen Habermas, famously remarked of Luhmann’s work:</p>

<blockquote>
  <p>It’s all wrong, but of high quality.</p>
</blockquote>

<p>Habermas labeled Luhmann’s theory <em>metabiological</em>, drawing a comparison to metaphysics and suggesting it extends beyond empirical sociology into abstract structures that make it detached from human agency.
This comparison is spot on: Luhmann’s theory views society not as a product of individuals’ intentions but as an autonomous, complex system very similar to an organism.
From this perspective, if we are to learn anything valuable from Luhmann, it might be that we should avoid treating society—and, by extension, the climate crisis—either as an engineering problem, i.e. as something to be ‘solved’ with a clear-cut plan or a knowledge problem, i.e. something that can be solved if deniers just accept ‘the Truth’.
Instead, we might approach it more like a chronic condition or complex illness, where we, as doctors, explore, probe, and apply potential treatments without assuming a one-size-fits-all solution. 
This approach requires continuous adjustment, sensitivity to feedback, and a readiness to adapt to the unexpected outcomes of our actions, reflecting the complex, interdependent nature of the systems we inhabit.
But, of course, <strong>we are not in control</strong> and there is no single medical doctor drafting the medication plan.</p>

<h2 id="references">References</h2>

<ol class="bibliography"><li><span id="Brundiers2020">Brundiers, K., Barth, M., Cebrián, G., Cohen, M., Diaz, L., Doucette-Remington, S., Dripps, W., Habron, G., Harré, N., Jarchow, M., Losch, K., Michel, J., Mochizuki, Y., Rieckmann, M., Parnell, R., Walker, P., &amp; Zint, M. (2020). Key competencies in sustainability in higher education—toward an agreed-upon reference framework. <i>Sustainability Science</i>, <i>16</i>(1), 13–29. https://doi.org/10.1007/s11625-020-00838-2</span></li>
<li><span id="luhmann:1992">Luhmann, N. (1992). <i>Die Wissenschaft der Gesellschaft</i> (p. 732). Suhrkamp.</span></li>
<li><span id="luhmann:1994">Luhmann, N. (1994). <i>Die Wirtschft der Gesellschaft</i> (p. 356). Suhrkamp.</span></li>
<li><span id="luhmann:1997">Luhmann, N. (1997). <i>Die Kunst der Gesellschaft</i> (p. 517). Suhrkamp.</span></li>
<li><span id="luhmann:2002">Luhmann, N. (2002). <i>Die Politik der Gesellschaft</i> (p. 444). Suhrkamp.</span></li>
<li><span id="luhmann:1998">Luhmann, N. (1998). <i>Die Gesellschaft der Gesellschaft</i> (p. 1164). Suhrkamp.</span></li>
<li><span id="moeller:2011">Möller, H.-G. (2011). <i>The Radical Luhmann</i> (p. 184). Columbia University Press.</span></li>
<li><span id="meadows:2008">Meadows, D. H. (2008). <i>Thinking in Systems: A Primer</i> (p. 240). Chelsea Green Publishing.</span></li>
<li><span id="maturana:1987">Maturana, H. R., &amp; Varela, F. J. (1987). <i>The Tree of Knowledge</i>. Shambhala.</span></li>
<li><span id="brown:1969">Spencer-Brown, G. (1969). <i>Laws of Form</i>. London: Allen and Unwin.</span></li>
<li><span id="foerster:2003">von Foerster, H. (2003). Cybernetics of Cybernetics. In <i>Understanding understanding: Essays on cybernetics and cognition</i> (pp. 283–286). Springer New York. https://doi.org/10.1007/0-387-21722-3_13</span></li>
<li><span id="moeller:2021">Möller, H.-G., &amp; D’Ambrosio, P. J. (2021). <i>You and Your Profile: Identity After Authenticity</i>. Columbia University Press.</span></li>
<li><span id="shannon:1948">Shannon, C. E. (1948). A mathematical theory of communication. <i>Bell Syst. Tech. J.</i>, <i>27</i>(3), 379–423.</span></li>
<li><span id="buentgen:2024">Büntgen, U. (2024). The importance of distinguishing climate science from climate activism. <i>Npj Climate Action</i>, <i>36</i>(3), 2731–9814. https://doi.org/10.1038/s44168-024-00126-0</span></li>
<li><span id="eck:2024">van Eck, C. W., Messling, L., &amp; Hayhoe, K. (2024). Challenging the neutrality myth in climate science and activism. <i>Npj Climate Action</i>, <i>81</i>(3), 2731–9814. https://doi.org/10.1038/s44168-024-00171-9</span></li></ol>]]></content><author><name>Benedikt Zönnchen</name><email>benedikt.zoennchen@web.de</email></author><category term="Sustainability" /><category term="Social Systems Theory" /><category term="Politics" /><summary type="html"><![CDATA[I have to admit, I’m somewhat addicted to thinking about social systems theory, particularly the version developed by the German sociologist Niklas Luhmann (1927–1998). But this addiction isn’t driven by pure fascination—it’s more of a love-hate relationship. On one hand, seeing the world through a Luhmannian lens doesn’t make it more just or fantastic, but it does make it comprehensible. The chaotic state of global affairs—the craziness of U.S. elections, the brutal devastation in the Middle East, the war in Europe, and the looming tensions brought by the climate crisis—appears senseless at times, even apocalyptic. And yet, systems theory offers a framework that, paradoxically, brings coherence to this apparent madness, showing us how these outcomes emerge from the logic of functionally differentiated systems.]]></summary></entry><entry><title type="html">Musical Interrogation IV - Transformer</title><link href="https://bzoennchen.github.io/Pages/2024/02/03/musical-interrogation-IV.html" rel="alternate" type="text/html" title="Musical Interrogation IV - Transformer" /><published>2024-02-03T00:00:00+01:00</published><updated>2024-02-03T00:00:00+01:00</updated><id>https://bzoennchen.github.io/Pages/2024/02/03/musical-interrogation-IV</id><content type="html" xml:base="https://bzoennchen.github.io/Pages/2024/02/03/musical-interrogation-IV.html"><![CDATA[<blockquote>
  <p>Recurrent models trained in practice are effectively feed-forward.
This could happen either because truncated backpropagation through time cannot learn patterns significantly longer than k steps, or, more provocatively, because models trainable by gradient descent cannot have long-term memory. – John Miller</p>
</blockquote>

<p>This time in the series we use the most famous model architecture for generative purposes: the <strong>transformer</strong> <a class="citation" href="#vaswani:2017">(Vaswani et al., 2017)</a>.
Transformers were initially targeted at natural language processing (NLP) problems, where the network input is a series of high-dimensional embeddings representing words or word fragments.
Transformers were introduced in 2017 by the authors of <em>Attention Is All You Need</em> <a class="citation" href="#vaswani:2017">(Vaswani et al., 2017)</a> to basically replace <em>recurrency</em> with <em>attention</em>.</p>

<p>One of the problems with RNNs is that they can forget information that is further back in the sequence.
While more sophisticated architectures, such as LSTMs <a class="citation" href="#hochreiter:1997">(Hochreiter &amp; Schmidhuber, 1997)</a> and <em>gated recurrent units</em> (GRUs) <a class="citation" href="#chung2014">(Chung et al., 2014)</a> partially addressed this problem, they still struggle with long term dependencies.
The idea that intermediate representations in the RNN should be exploited to produce the output led to the <em>attention mechanism</em> <a class="citation" href="#bahdanau:2014">(Bahdanau et al., 2014)</a> and, in the end, to the transformer architecture <a class="citation" href="#vaswani:2017">(Vaswani et al., 2017)</a>.
The transformer avoids the problem of vanishing or exploding gradients by avoiding recurrency, that is, by utilizing the whole sequence in parallel.</p>

<p>Today, all successful large language models (LLMs) utilize the transformer architecture.
It brought us ChatGPT (based on GPT-3 <a class="citation" href="#brown:2020">(Brown et al., 2020)</a> and GPT-4 <a class="citation" href="#openai:2023">(OpenAI, 2023; Bubeck et al., 2023)</a>), LLaMA <a class="citation" href="#touvron:2023">(Touvron et al., 2023)</a>, LLaMA 2 <a class="citation" href="#touvron:2023b">(Touvron et al., 2023)</a>, BERT <a class="citation" href="#devlin:2019">(Devlin et al., 2019)</a> and many fine-tuned derivatives such as Codex <a class="citation" href="#chen:2021">(Chen et al., 2021)</a>.
In the domain of symbolic music, transformers were also employed.
Examples are the Music Transformer <a class="citation" href="#huang:2018">(Huang et al., 2018)</a>, the Pop Music Transformer <a class="citation" href="#huang:2020">(Huang &amp; Yang, 2020)</a>, multi-track music generation <a class="citation" href="#ens:2020">(Ens &amp; Pasquier, 2020)</a>, piano inpainting <a class="citation" href="#hadjeres:2021">(Hadjeres &amp; Crestel, 2021)</a>, Theme Transformer <a class="citation" href="#shih:2022">(Shih et al., 2022)</a> and more.
Furthermore there are transformers, such as MusicGen <a class="citation" href="#copet:2023">(Copet et al., 2023)</a> that generate audio output directly.</p>

<p>While there is an intuitive explanation of the attention mechanism, it is still unclear why exactly the transformer is so effective—there is no rigorous mathematical proof.
It is well-known how their components work and what mathematical operations are performed, but it is very hard to interpret the seemingly emerging power when all the small parts work together.
One source of their effectiveness is that they relate tokens to other tokens more directly (without a hidden state which washes away the information) and the independence of multiple execution paths make them especially suitable for the exploitation of multicore processors such as GPUs and TPUs.
However, looking at the whole sequence at once comes at a cost: computation and memory complexity!
Therefore, to train transformers you require GPUs with a lot of memory which is concerning for artists who might want to utilize transformers independently from proprietary cloud services.</p>

<p>Original transformers were introduced for natural language processing.
However, since language datasets share some of the characteristics of musical notations, transformers achieve good results in learning the structure of symbolic pieces.
In music as well as in language the number of input variables can be very large, and the statistics are similar at every position; it’s not sensible to re-learn the meaning of the word <em>dog</em> at every possible position in a body of text.
Language datasets and music datasets have the complication that their sequences vary in length.</p>

<p>However, we also have to remember that there are also differences between the two domains.
The alphabet of musical notations has more than 26 symbols and there is a strong relation between certain symbols.
For example, there is a strong relation between the C’s of each octave or a whole and half note in the same pitch class.
Furthermore, shifting all the letters in a text changes the meaning of that text dramatically while in the case of music this is most often not the case.</p>

<h2 id="attention-in-encoder-decoder-rnns">Attention in Encoder-Decoder RNNs</h2>

<p>What is the idea behind the attention mechanism?
Attention was introduced to bidirectional recurrent neural networks (RNNs) in 2014 <a class="citation" href="#bahdanau:2014">(Bahdanau et al., 2014)</a> for language translation, that is, for an <em>encoder-decoder architecture</em>.
In this scenario we want to translate a sentence from e.g. English into e.g. German.
The attention mechanism helps the decoder part of the RNN to focus on different parts of the encoder’s output (representations of the English words) differently.
Therefore, it helps to preserve long term dependencies.</p>

<p>The <strong>encoder’s</strong> input is a sequence of tokens, let’s say words for simplicity, i.e. a sequence</p>

\[\mathbf{x}_{0}, \ldots \mathbf{x}_{n-1}.\]

<p>For each word \(\mathbf{x}_{i}\) it computes some output \(\mathbf{y}_{i}\).
The assumption is that the probability for \(\mathbf{x}_{i}\) depends on \(\mathbf{x}_{j}\).
Since we have the whole sentence given, we can use a <em>bidirectional RNN</em> and look into the future.
Thus, with respect to dependency, the probability for token \(i\) can depend on the probability for token \(j\) and vice versa.</p>

<p>The <strong>decoder’s</strong> input is the <strong>whole</strong> sequence computed by the <strong>encoder</strong> but as a weighted sum.
The output is a sequence of German words, let’s say</p>

\[\mathbf{y'}_{0}, \ldots \mathbf{y'}_{n-1}.\]

<p>This time however, the <strong>decoder</strong> RNN is unidirectional.
It can not look into the future and computes each German word strictly from left to right.
To compute the weights or attention scores of \(\mathbf{y'}_{i}\), an <strong>alignment model</strong> receives the hidden state \(\mathbf{h'}_{i-1}\) and the outputs of the encoder as input.
First a simple dot product is computed:</p>

\[e_{i,j} = \mathbf{h'}_{i}^\top \mathbf{y}_{j} \quad \text{ for } j = 0, \ldots, i-1.\]

<p>Later it was suggested to use an additional linear transformation on the output:</p>

\[e_{i,j} = \mathbf{h'}_{i}^\top (\mathbf{W}\mathbf{y}_{j}) \quad \text{ for } j = 0, \ldots, i-1.\]

<p>where \(\mathbf{W}\) is learned.
All these scores are normalized by the softmax function giving us \(n\) weights:</p>

\[\alpha_{i,j} = \frac{\exp\left( e_{i,j} \right)}{\sum\limits_{k=0}^{n-1} \exp\left( e_{i, k}\right)}.\]

<p>Then the <strong>decoder’s</strong> ‘real’ input is computed by a weighted sum of the <strong>encoder’s</strong> output:</p>

\[\hat{\mathbf{h}}_i = \sum_j \alpha_{i,j} \mathbf{y}_j.\]

<p>The weights determine how “strong” the information of the decoder’s input will be utilized, i.e., how much attention is spent on each previous output of the model.</p>

<div><img style="display:block; margin-left:auto; margin-right:auto; width:90%;" src="/Pages/assets/images/rnn-attention.png" alt="RNN with attention" />
<div style="display: table;margin: 0 auto;">Figure 1: RNN with self-attention.</div>
</div>
<p><br /></p>

<p>This results in a quadratic complexity of \(\mathcal{O}(n^2)\) because for each of the \(n\) tokens, we want to decode, we have \(n\) weights.</p>

<h2 id="the-transformer-architecture">The Transformer Architecture</h2>

<p>The original transformer was introduced for the task of machine translation thus it was an encoder-decoder architecture.
In Fig. 2 you see a slightly modified version where the addition (residual connections) and the layer norm are in front of the attention layer.</p>

<div><img style="display:block; margin-left:auto; margin-right:auto; width:90%;" src="/Pages/assets/images/transformer.png" alt="RNN with attention" />
<div style="display: table;margin: 0 auto;">Figure 2: The slightly modified transformer.</div>
</div>
<p><br /></p>

<p>Let’s consider a scenario in which we have an English sentence that we want to translate into French. 
In this process, the encoder plays a crucial role by transforming the English sentence into a highly compressed and information-rich representation.</p>

<p>Subsequently, the decoder comes into play, generating the French translation word by word. 
It relies on the previously computed French words to predict and produce the next one. 
It’s important to note that the input provided to the decoder is a partial translation, essentially a shifted version of what it is currently working on. 
This is because the decoder should lack the ability to see into the future; it only has access to the portion of the translation it has computed up to that point—otherwise it would cheat while training which would hurt the learning process.</p>

<p>To address this limitation, the decoder employs a masked version of the multi-head attention layer. 
This mechanism ensures that the decoder focuses on the relevant information without peeking ahead.</p>

<p>Furthermore, the utilization of residual connections and layer normalization over the feature dimension within a single sample serves as a valuable tool to combat the issue of vanishing gradients in deep neural networks, ensuring the efficient training and optimization of the translation model.</p>

<p>In our case we do not actually want to translate a sentence but we want to generate musical notes from a sequence of given notes.
Therefore, we have no encoded information and there is no encoding involved.
We only need the decoder part.
Furthermore, I only use one (masked) multi-head attention layer in each block.</p>

<div><img style="display:block; margin-left:auto; margin-right:auto; width:40%;" src="/Pages/assets/images/decoder.png" alt="Decoder-only transformer" />
<div style="display: table;margin: 0 auto;">Figure 3: Our decoder-only transformer.</div>
</div>
<p><br /></p>

<p>Ok, but how does this really work?
What is going on here?
Well, the key to understand transformers is to understand the self-attention mechanism which I try to explain below.</p>

<h2 id="self-attention">Self-Attention</h2>

<p>The idea of the transformer is to just rely on (self-)attention thus remove recurrency.
This means that the model “sees” \(n\) tokens to generate the \((n+1)^\text{th}\) token.
Simple RNNs for predicting the next tokens only see the previous token and the hidden state which represents all the tokens before.
But, as I discussed in previous articles, the information of the hidden state gets washed away over time and without attention there seems to be little control over the importance of certain tokens of the sequence.</p>

<p>The fundamental operation of the transformer, i.e. the attention mechanism, is implemented in its <code class="language-plaintext highlighter-rouge">Head</code>.
Let \(\mathbf{x}_0, \ldots, \mathbf{x}_{n-1} \in \mathbb{R}^{D \times 1}\) be the \(n\) tokens of a sequence.
A standard neural network layer \(f(\cdot)\), takes a \(D \times 1\) input and applies a linear transformation followed by an activation function like a \(\text{ReLU}\):</p>

\[f(\mathbf{x}) = \text{ReLU}\left( \mathbf{W}\mathbf{x} + \mathbf{b }\right),\]

<p>where \(\mathbf{b}\) contains the biases, and \(\mathbf{W}\) contains the weights.</p>

<p>A self-attention \(\mathbf{sa}(\cdot)\) block takes all the \(n\) inputs, each of dimension \(D \times 1\), and returns \(n\) output vectors of the same size.
Note that in our case each input represents a musical note or event.
First, a set of <strong>values</strong> is computed for each input:</p>

\[\mathbf{v}_i = \mathbf{b}_v + \mathbf{W}_v \mathbf{x}_i, \quad \text{ (value)}\]

<p>where \(\mathbf{b}_v \in \mathbb{R}^D \text{ and } \mathbf{W}_v \in \mathbb{R}^{D \times D}\) represent biases and weights, respectively (<strong>for all inputs</strong>). 
The \(j^\text{th}\) output \(\mathbf{sa}_j(\mathbf{x}_0, \ldots, \mathbf{x}_{n-1})\) is a weighted sum of all the values \(\mathbf{v}_i, i = 0, \ldots n-1\) where each weight depends on \(\mathbf{x}_j\):</p>

\[\mathbf{sa}_j(\mathbf{x}_0, \ldots, \mathbf{x}_{n-1}) = \sum_{i=0}^{n-1} \alpha(\mathbf{x}_i, \mathbf{x}_j) \mathbf{v}_i.\]

<p>The scalar weight \(\alpha(\mathbf{x}_i, \mathbf{x}_j)\) is the <strong>attention</strong> that the \(j^\text{th}\) note pays to the note \(\mathbf{x}_i\).
The \(n\) weights \(\alpha(\cdot, \mathbf{x}_j)\) are non-negative and sum to one.
Hence, self-attention can be thought of as <em>routing</em> the values in different proportions to create each output.</p>

<p><br /></p>
<div><img style="display:block; margin-left:auto; margin-right:auto; width:35%;" src="/Pages/assets/images/routing.png" alt="Routing principle" />
<div style="display: table;margin: 0 auto;">Figure 4: Routing principle.</div>
</div>
<p><br /></p>

<p>To compute the attention, we apply two more linear transformations to the inputs:</p>

\[\begin{aligned}
\mathbf{q}_j &amp;= \mathbf{b}_q + \mathbf{W}_q \mathbf{x}_j \quad \text{ (query)}\\
\mathbf{k}_i &amp;= \mathbf{b}_k + \mathbf{W}_k \mathbf{x}_i \quad \text{ (key).}
\end{aligned}\]

<p>The <strong>dot product</strong> of two vectors \(\mathbf{q}_j\), \(\mathbf{k}_i\) is a measurement of their similarity.
The matrices \(\mathbf{W}_q, \mathbf{W}_k\) and the respective bias are learned such that similarity of \(\mathbf{q}_j\), \(\mathbf{k}_i\) can be interpreted as how “important” \(\mathbf{x}_i\) is for \(\mathbf{x}_j\).
Thus, the “magic” happens via a very simple linear transformation and one might ask if this operation is powerful enough to relate <strong>all</strong> words/tokens in a desirable way.
The answer is most certainly “no” thus one adds feed forward layers which introduce non-linearity in between multiple attention layers.</p>

<p>In the special case where both vectors are unit vectors, the dot product is the cosine of the angle between the two.
In general, this relationship is expressed by the following equation:</p>

\[\mathbf{q}_j \circ \mathbf{k}_i = \mathbf{q}_j^\top  \mathbf{k}_i = \Vert \mathbf{q}_j \Vert \cdot \Vert \mathbf{k}_i \Vert \cos(\beta),\]

<p>where \(\beta\) is the angle between the two vectors.
Computing the <em>dot product</em> between queries and keys gives us the similarities we desire.
To normalize, we then pass the result through a <em>softmax</em> function:</p>

\[\alpha(\mathbf{x}_i, \mathbf{x}_j) = \frac{\exp(\mathbf{q}_j^\top \mathbf{k}_i / \sqrt{D_q})}{\sum\limits_{r=0}^{n-1} \exp(\mathbf{q}_j^\top \mathbf{k}_r / \sqrt{D_q})},\]

<p>where \(D_q\) is the dimension of the queries and keys (i.e., the number of rows in \(\mathbf{W}_q\) and \(\mathbf{W}_k\), which must be the same).
You can think of the <em>key</em> as what is offered and the <em>query</em> as what is searched for.
If \(\mathbf{X}\), \(\mathbf{K}\), \(\mathbf{Q}\), and \(\mathbf{V}\) contain all the inputs, keys, queries and values then we can compute the self-attention by</p>

\[\mathbf{Sa}(\mathbf{X}) = \mathbf{V} \cdot \text{Softmax}\left( \frac{\mathbf{K}^\top \mathbf{Q}}{\sqrt{D_q} }\right).\]

<p>The overall computation is illustrated in Figure 5.</p>

<p><br /></p>
<div><img style="display:block; margin-left:auto; margin-right:auto; width:90%;" src="/Pages/assets/images/self-attention.png" alt="Self-attention in matrix-form" />
<div style="display: table;margin: 0 auto;">Figure 5: Self-attention in matrix-form.</div>
</div>
<p><br /></p>

<h2 id="masking-attention-head">Masking Attention Head</h2>

<p>Since our transformer should not look into the future, because when we use it in the prediction mode it also can not look ahead of the token it predicts, we have to mask entries in</p>

\[\text{Softmax}\left(\frac{\mathbf{K}^\top \mathbf{Q}}{\sqrt{D_q}}\right).\]

<p>If you look into the code, I did this by setting the respective values in \(\mathbf{K}^\top \mathbf{Q}\) to negative infinity before computing the softmax.</p>

<p><br /></p>
<div><img style="display:block; margin-left:auto; margin-right:auto; width:60%;" src="/Pages/assets/images/transformer-head.png" alt="Transformer head" />
<div style="display: table;margin: 0 auto;">Figure 6: Transformer head for a sequence length equal to 5.</div>
</div>
<p><br /></p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">class</span> <span class="nc">Head</span><span class="p">(</span><span class="n">nn</span><span class="p">.</span><span class="n">Module</span><span class="p">):</span>
    <span class="s">""" one head of self-attention """</span>
    
    <span class="k">def</span> <span class="nf">__init__</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">n_embd</span><span class="p">,</span> <span class="n">head_size</span><span class="p">,</span> <span class="n">sequence_len</span><span class="p">,</span> <span class="n">dropout</span><span class="p">):</span>
        <span class="nb">super</span><span class="p">().</span><span class="n">__init__</span><span class="p">()</span>
        <span class="bp">self</span><span class="p">.</span><span class="n">key</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="n">Linear</span><span class="p">(</span><span class="n">n_embd</span><span class="p">,</span> <span class="n">head_size</span><span class="p">,</span> <span class="n">bias</span><span class="o">=</span><span class="bp">False</span><span class="p">)</span>   <span class="c1"># key embedding
</span>        <span class="bp">self</span><span class="p">.</span><span class="n">query</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="n">Linear</span><span class="p">(</span><span class="n">n_embd</span><span class="p">,</span> <span class="n">head_size</span><span class="p">,</span> <span class="n">bias</span><span class="o">=</span><span class="bp">False</span><span class="p">)</span> <span class="c1"># query embedding
</span>        <span class="bp">self</span><span class="p">.</span><span class="n">value</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="n">Linear</span><span class="p">(</span><span class="n">n_embd</span><span class="p">,</span> <span class="n">head_size</span><span class="p">,</span> <span class="n">bias</span><span class="o">=</span><span class="bp">False</span><span class="p">)</span> <span class="c1"># value embedding
</span>        <span class="bp">self</span><span class="p">.</span><span class="n">register_buffer</span><span class="p">(</span><span class="s">'tril'</span><span class="p">,</span> <span class="n">torch</span><span class="p">.</span><span class="n">tril</span><span class="p">(</span><span class="n">torch</span><span class="p">.</span><span class="n">ones</span><span class="p">(</span><span class="n">sequence_len</span><span class="p">,</span> <span class="n">sequence_len</span><span class="p">)))</span>
        <span class="bp">self</span><span class="p">.</span><span class="n">dropout</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="n">Dropout</span><span class="p">(</span><span class="n">dropout</span><span class="p">)</span> <span class="c1"># to avoid overfitting
</span>        
    <span class="k">def</span> <span class="nf">forward</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">x</span><span class="p">):</span>
        <span class="n">B</span><span class="p">,</span><span class="n">T</span><span class="p">,</span><span class="n">C</span> <span class="o">=</span> <span class="n">x</span><span class="p">.</span><span class="n">shape</span>
        <span class="n">k</span> <span class="o">=</span> <span class="bp">self</span><span class="p">.</span><span class="n">key</span><span class="p">(</span><span class="n">x</span><span class="p">)</span> <span class="c1"># B, T, head_size, compute all keys
</span>        <span class="n">q</span> <span class="o">=</span> <span class="bp">self</span><span class="p">.</span><span class="n">query</span><span class="p">(</span><span class="n">x</span><span class="p">)</span> <span class="c1"># B, T, head_size, compute all queries
</span>        <span class="n">_</span><span class="p">,</span> <span class="n">_</span><span class="p">,</span> <span class="n">head_size</span> <span class="o">=</span> <span class="n">q</span><span class="p">.</span><span class="n">shape</span>
        
         <span class="c1"># B, T, head_size @ B, head_size, 
</span>        <span class="n">wei</span> <span class="o">=</span> <span class="n">q</span> <span class="o">@</span> <span class="n">k</span><span class="p">.</span><span class="n">transpose</span><span class="p">(</span><span class="o">-</span><span class="mi">2</span><span class="p">,</span> <span class="o">-</span><span class="mi">1</span><span class="p">)</span> <span class="o">*</span> <span class="p">(</span><span class="n">head_size</span> <span class="o">**</span> <span class="p">(</span><span class="o">-</span><span class="mf">0.5</span><span class="p">))</span> <span class="c1"># T =&gt; B, T, T
</span>
        <span class="c1"># because we can not look into the future 
</span>        <span class="n">wei</span> <span class="o">=</span> <span class="n">wei</span><span class="p">.</span><span class="n">masked_fill</span><span class="p">(</span><span class="bp">self</span><span class="p">.</span><span class="n">tril</span><span class="p">[:</span><span class="n">T</span><span class="p">,</span> <span class="p">:</span><span class="n">T</span><span class="p">]</span><span class="o">==</span><span class="mi">0</span><span class="p">,</span> <span class="nb">float</span><span class="p">(</span><span class="s">'-inf'</span><span class="p">))</span>
        <span class="n">wei</span> <span class="o">=</span> <span class="n">F</span><span class="p">.</span><span class="n">softmax</span><span class="p">(</span><span class="n">wei</span><span class="p">,</span> <span class="n">dim</span><span class="o">=-</span><span class="mi">1</span><span class="p">)</span>
        <span class="n">wei</span> <span class="o">=</span> <span class="bp">self</span><span class="p">.</span><span class="n">dropout</span><span class="p">(</span><span class="n">wei</span><span class="p">)</span>

        <span class="n">v</span> <span class="o">=</span> <span class="bp">self</span><span class="p">.</span><span class="n">value</span><span class="p">(</span><span class="n">x</span><span class="p">)</span> <span class="c1"># B, T, head_size
</span>        <span class="n">out</span> <span class="o">=</span> <span class="n">wei</span> <span class="o">@</span> <span class="n">v</span> <span class="c1"># T, T @ B, T, head_size =&gt; B, T, head_size
</span>        <span class="k">return</span> <span class="n">out</span>
<span class="p">...</span>
</code></pre></div></div>

<p>The multi-head attention layer consists of multiple heads.
Note that apart from <strong>masked</strong> <strong>self-attention</strong>, the head also applies a <strong>dropout</strong> which helps with regularization.</p>

<h2 id="stacked-multi-head-attention">Stacked Multi-Head Attention</h2>

<p>Instead of using only one <code class="language-plaintext highlighter-rouge">Head</code> it is usually a good idea to use multiple ones.
To do this we transform the input into a <code class="language-plaintext highlighter-rouge">head_size</code>-dimensional space.
Suppose we use 4 heads then <code class="language-plaintext highlighter-rouge">head_size * 4</code> should be equal to the rows of \(\mathbf{W}_0\) (compare Fig. 7) of the multi-head attention layer.
Since I add the input to the output of <code class="language-plaintext highlighter-rouge">MultiHeadAttention</code> (via residual connections), the columns of \(\mathbf{W}_0\) should be equal to the dimension of the input of the <code class="language-plaintext highlighter-rouge">MultiHeadAttention</code>.
In my case this is <code class="language-plaintext highlighter-rouge">n_embd</code>, i.e. the dimension of our embedded tokens.</p>

<p>\(\mathbf{W}_0\) transforms the concatenated results of the heads back to the dimension equal to <code class="language-plaintext highlighter-rouge">n_embd</code>. 
This is needed to stack <code class="language-plaintext highlighter-rouge">Block</code>s (each consisting of a <code class="language-plaintext highlighter-rouge">MultiHeadAttention</code>) on top of each other.
The output of <code class="language-plaintext highlighter-rouge">Block</code> \(i\) has to fit into block \(i+1\).</p>

<p><br /></p>
<div><img style="display:block; margin-left:auto; margin-right:auto; width:90%;" src="/Pages/assets/images/multi-head.png" alt="Multi-head attention" />
<div style="display: table;margin: 0 auto;">Figure 7: Multi-head attention.</div>
</div>
<p><br /></p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">class</span> <span class="nc">MultiHeadAttention</span><span class="p">(</span><span class="n">nn</span><span class="p">.</span><span class="n">Module</span><span class="p">):</span>
    <span class="k">def</span> <span class="nf">__init__</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">n_embd</span><span class="p">,</span> <span class="n">n_heads</span><span class="p">,</span> <span class="n">head_size</span><span class="p">,</span> <span class="n">sequence_len</span><span class="p">,</span> <span class="n">dropout</span><span class="p">):</span>
        <span class="nb">super</span><span class="p">().</span><span class="n">__init__</span><span class="p">()</span>
        <span class="bp">self</span><span class="p">.</span><span class="n">heads</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="n">ModuleList</span><span class="p">(</span>
            <span class="p">[</span><span class="n">Head</span><span class="p">(</span><span class="n">n_embd</span><span class="p">,</span> <span class="n">head_size</span><span class="p">,</span> <span class="n">sequence_len</span><span class="p">,</span> <span class="n">dropout</span><span class="p">)</span> <span class="k">for</span> <span class="n">_</span> <span class="ow">in</span> <span class="nb">range</span><span class="p">(</span><span class="n">n_heads</span><span class="p">)]</span>
        <span class="p">)</span>

        <span class="c1"># W_0
</span>        <span class="bp">self</span><span class="p">.</span><span class="n">W0</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="n">Linear</span><span class="p">(</span><span class="n">head_size</span> <span class="o">*</span> <span class="n">n_heads</span><span class="p">,</span> <span class="n">n_embd</span><span class="p">)</span> 

        <span class="bp">self</span><span class="p">.</span><span class="n">dropout</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="n">Dropout</span><span class="p">(</span><span class="n">dropout</span><span class="p">)</span>

    <span class="k">def</span> <span class="nf">forward</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">x</span><span class="p">):</span>
        <span class="c1"># concatenation of the results of each head
</span>        <span class="n">out</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">cat</span><span class="p">([</span><span class="n">head</span><span class="p">(</span><span class="n">x</span><span class="p">)</span> <span class="k">for</span> <span class="n">head</span> <span class="ow">in</span> <span class="bp">self</span><span class="p">.</span><span class="n">heads</span><span class="p">],</span> <span class="n">dim</span><span class="o">=-</span><span class="mi">1</span><span class="p">)</span> 

        <span class="c1"># Figure 7
</span>        <span class="n">out</span> <span class="o">=</span> <span class="bp">self</span><span class="p">.</span><span class="n">W0</span><span class="p">(</span><span class="n">out</span><span class="p">)</span>
        <span class="n">out</span> <span class="o">=</span> <span class="bp">self</span><span class="p">.</span><span class="n">dropout</span><span class="p">(</span><span class="n">out</span><span class="p">)</span>
        <span class="k">return</span> <span class="n">out</span>
<span class="p">...</span>
</code></pre></div></div>

<p>The hope is that each <code class="language-plaintext highlighter-rouge">Head</code> concentrates on different parts of the structure we want to learn.</p>

<h2 id="transformer-block">Transformer Block</h2>

<p>A <code class="language-plaintext highlighter-rouge">Block</code> consists of a <code class="language-plaintext highlighter-rouge">MultiHeadAttention</code>-layer and a relatively simple FFN-layer followed by two <code class="language-plaintext highlighter-rouge">LayerNorm</code> which applies <em>layer normalization</em> over the feature dimension within a single sample.
This helps the gradients to stay in a “good” range.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">class</span> <span class="nc">Block</span><span class="p">(</span><span class="n">nn</span><span class="p">.</span><span class="n">Module</span><span class="p">):</span>
    
    <span class="k">def</span> <span class="nf">__init__</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">n_embd</span><span class="p">,</span> <span class="n">n_heads</span><span class="p">,</span> <span class="n">sequence_len</span><span class="p">,</span> <span class="n">dropout</span><span class="p">):</span>
        <span class="nb">super</span><span class="p">().</span><span class="n">__init__</span><span class="p">()</span>
        <span class="c1"># head_size could be defined differently
</span>        <span class="n">head_size</span> <span class="o">=</span> <span class="n">n_embd</span> <span class="o">//</span> <span class="n">n_heads</span>
        <span class="bp">self</span><span class="p">.</span><span class="n">sa</span> <span class="o">=</span> <span class="n">MultiHeadAttention</span><span class="p">(</span><span class="n">n_embd</span><span class="p">,</span> <span class="n">n_heads</span><span class="p">,</span> <span class="n">head_size</span><span class="p">,</span> <span class="n">sequence_len</span><span class="p">,</span> <span class="n">dropout</span><span class="p">)</span>
        <span class="bp">self</span><span class="p">.</span><span class="n">ffwd</span> <span class="o">=</span> <span class="n">FeedForward</span><span class="p">(</span><span class="n">n_embd</span><span class="p">,</span> <span class="n">dropout</span><span class="p">)</span>
        <span class="bp">self</span><span class="p">.</span><span class="n">ln1</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="n">LayerNorm</span><span class="p">(</span><span class="n">n_embd</span><span class="p">)</span>
        <span class="bp">self</span><span class="p">.</span><span class="n">ln2</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="n">LayerNorm</span><span class="p">(</span><span class="n">n_embd</span><span class="p">)</span>
        
    <span class="k">def</span> <span class="nf">forward</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">x</span><span class="p">):</span>
        <span class="n">x</span> <span class="o">=</span> <span class="n">x</span> <span class="o">+</span> <span class="bp">self</span><span class="p">.</span><span class="n">sa</span><span class="p">(</span><span class="bp">self</span><span class="p">.</span><span class="n">ln1</span><span class="p">(</span><span class="n">x</span><span class="p">))</span> <span class="c1"># residual connection
</span>        <span class="n">x</span> <span class="o">=</span> <span class="n">x</span> <span class="o">+</span> <span class="bp">self</span><span class="p">.</span><span class="n">ffwd</span><span class="p">(</span><span class="bp">self</span><span class="p">.</span><span class="n">ln2</span><span class="p">(</span><span class="n">x</span><span class="p">))</span> <span class="c1"># residual connection
</span>        <span class="k">return</span> <span class="n">x</span>

<span class="k">class</span> <span class="nc">FeedForward</span><span class="p">(</span><span class="n">nn</span><span class="p">.</span><span class="n">Module</span><span class="p">):</span>
    
    <span class="k">def</span> <span class="nf">__init__</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">n_embd</span><span class="p">,</span> <span class="n">dropout</span><span class="p">):</span>
        <span class="nb">super</span><span class="p">().</span><span class="n">__init__</span><span class="p">()</span>
        <span class="bp">self</span><span class="p">.</span><span class="n">net</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="n">Sequential</span><span class="p">(</span>
            <span class="n">nn</span><span class="p">.</span><span class="n">Linear</span><span class="p">(</span><span class="n">n_embd</span><span class="p">,</span> <span class="mi">4</span> <span class="o">*</span> <span class="n">n_embd</span><span class="p">),</span> 
            <span class="n">nn</span><span class="p">.</span><span class="n">ReLU</span><span class="p">(),</span>
            <span class="n">nn</span><span class="p">.</span><span class="n">Linear</span><span class="p">(</span><span class="mi">4</span> <span class="o">*</span> <span class="n">n_embd</span><span class="p">,</span> <span class="n">n_embd</span><span class="p">),</span>
            <span class="n">nn</span><span class="p">.</span><span class="n">Dropout</span><span class="p">(</span><span class="n">dropout</span><span class="p">),</span>
        <span class="p">)</span>
        
    <span class="k">def</span> <span class="nf">forward</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">x</span><span class="p">):</span>
        <span class="k">return</span> <span class="bp">self</span><span class="p">.</span><span class="n">net</span><span class="p">(</span><span class="n">x</span><span class="p">)</span>
<span class="p">...</span>
</code></pre></div></div>

<p><br /></p>
<div><img style="display:block; margin-left:auto; margin-right:auto; width:40%;" src="/Pages/assets/images/block.png" alt="Transformer block" />
<div style="display: table;margin: 0 auto;">Figure 8: Our decoder transformer block with only one (masked) multi-head attention layer.</div>
</div>

<h2 id="positional-encoding">Positional Encoding</h2>

<p>The Transformer has no more hidden state.
Therefore, instead of processing token by token trying to memorize important information via the hidden state, it processes all \(n\) tokens in parallel, which is good for parallel computation but increases the time and space complexity from \(\mathcal{O}(n)\) (LSTM) to \(\mathcal{O}(n^2)\).</p>

<p>Furthermore, we have to encode the position of the tokens into \(\mathbf{x}\) because we lost the implicit order of computation.
In the original paper <a class="citation" href="#vaswani:2017">(Vaswani et al., 2017)</a> the authors utilized an embedding that involved sine and cosine functions.
Their embedding is a very clever use of periodic functions but I will not go into details here.
Instead of using a fixed embedding, I let the transformer learn the positional embedding.</p>

<p>Therefore, I transform the input <code class="language-plaintext highlighter-rouge">idx</code> into two vectors <strong>positional embedding</strong> and <strong>token embedding</strong>, which are <strong>added</strong> together.
Note that our input <code class="language-plaintext highlighter-rouge">idx</code> is an array of numbers each representing the id of the token.
Each number will be transformed into a specific vector (i.e. its embedding).
The embedding will be learned.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">class</span> <span class="nc">TransformerDecoder</span><span class="p">(</span><span class="n">nn</span><span class="p">.</span><span class="n">Module</span><span class="p">):</span>
    
    <span class="k">def</span> <span class="nf">__init__</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">vocab_size</span><span class="p">,</span> <span class="n">sequence_len</span><span class="p">,</span> <span class="n">n_embd</span><span class="p">,</span> <span class="n">n_heads</span><span class="p">,</span> <span class="n">n_blocks</span><span class="p">,</span> <span class="n">dropout</span><span class="p">):</span>
        <span class="nb">super</span><span class="p">().</span><span class="n">__init__</span><span class="p">()</span>

        <span class="c1"># vocab_size is the size of our alphabet
</span>        <span class="bp">self</span><span class="p">.</span><span class="n">token_embedding_table</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="n">Embedding</span><span class="p">(</span><span class="n">vocab_size</span><span class="p">,</span> <span class="n">n_embd</span><span class="p">)</span>

        <span class="c1"># sequence_len is equal to n
</span>        <span class="bp">self</span><span class="p">.</span><span class="n">position_embedding_table</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="n">Embedding</span><span class="p">(</span><span class="n">sequence_len</span><span class="p">,</span> <span class="n">n_embd</span><span class="p">)</span>

        <span class="bp">self</span><span class="p">.</span><span class="n">blocks</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="n">Sequential</span><span class="p">(</span>
            <span class="o">*</span><span class="p">[</span><span class="n">Block</span><span class="p">(</span><span class="n">n_embd</span><span class="p">,</span> <span class="n">n_heads</span><span class="p">,</span> <span class="n">sequence_len</span><span class="p">,</span> <span class="n">dropout</span><span class="p">)</span> <span class="k">for</span> <span class="n">_</span> <span class="ow">in</span> <span class="nb">range</span><span class="p">(</span><span class="n">n_blocks</span><span class="p">)]</span>
        <span class="p">)</span>
        <span class="bp">self</span><span class="p">.</span><span class="n">lm_head</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="n">Linear</span><span class="p">(</span><span class="n">n_embd</span><span class="p">,</span> <span class="n">vocab_size</span><span class="p">)</span>

    <span class="k">def</span> <span class="nf">forward</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">idx</span><span class="p">):</span>
        <span class="n">B</span><span class="p">,</span> <span class="n">T</span> <span class="o">=</span> <span class="n">idx</span><span class="p">.</span><span class="n">shape</span>
        
        <span class="c1"># token embedding. B, T, n_embd
</span>        <span class="n">token_emb</span> <span class="o">=</span> <span class="bp">self</span><span class="p">.</span><span class="n">token_embedding_table</span><span class="p">(</span><span class="n">idx</span><span class="p">)</span> 

        <span class="c1"># positional embedding. T, n_embd 
</span>        <span class="n">pos_emb</span> <span class="o">=</span> <span class="bp">self</span><span class="p">.</span><span class="n">position_embedding_table</span><span class="p">(</span><span class="n">torch</span><span class="p">.</span><span class="n">arange</span><span class="p">(</span><span class="n">T</span><span class="p">,</span> <span class="n">device</span><span class="o">=</span><span class="n">device</span><span class="p">))</span> 

        <span class="n">x</span> <span class="o">=</span> <span class="n">token_emb</span> <span class="o">+</span> <span class="n">pos_emb</span> <span class="c1"># B, T, n_embd + T, n_embd =&gt; B, T, n_embd
</span>        <span class="n">x</span> <span class="o">=</span> <span class="bp">self</span><span class="p">.</span><span class="n">blocks</span><span class="p">(</span><span class="n">x</span><span class="p">)</span> <span class="c1"># B, T, head_size
</span>        <span class="n">logits</span> <span class="o">=</span> <span class="bp">self</span><span class="p">.</span><span class="n">lm_head</span><span class="p">(</span><span class="n">x</span><span class="p">)</span> <span class="c1"># B, T, vocab_size
</span>        <span class="k">return</span> <span class="n">logits</span>

<span class="p">...</span>
</code></pre></div></div>

<p>By increasing the dimension of the embedding, the sequence length, the number of heads within a block and the number of blocks we can drastically increase the size and power of our decoder-only transformer.
However, this will rapidly increase the memory requirements and training time.</p>

<div><img style="display:block; margin-left:auto; margin-right:auto; width:40%;" src="/Pages/assets/images/decoder.png" alt="Decoder-only transformer" />
<div style="display: table;margin: 0 auto;">Figure 9: Our simplified decoder-only transformer.</div>
</div>
<p><br /></p>

<p>Furthermore, it is very important to understand that what really matters is a <strong>high-quality training dataset</strong>!
Of course, your model architecture matters too, but your model can not learn what is not there.
Additionally, the <strong>musical representation</strong> you feed into the transformer matters as well.
In our case this representation, using basically piano rolls, is very simple.
It does not contain any high level information such as the end of a bar, section, phrase or musical theme.
We just hope that the transformer will eventually learn all these concepts.
It is an active research question what impact a good musical representation has on the result the trained transformer generates.</p>

<h2 id="relative-positional-self-attention">Relative Positional Self-Attention</h2>

<p>So far our positional encoding was just a sequence of natural numbers \(0, 1, \ldots, n-1\) and we used an embedding which was added to the input, that is, the embedding of \(i\) was added to the embedding of \(\mathbf{x}_i\) of the input sequence.
However, this encoding might not be optimal in the context of music where tones, phrases, musical ideas and themes repeat frequently.
A relative position representation to allow attention to be informed by how far two positions are apart in a sequence might be much more effective.</p>

<p>So, instead of learning the index of a token within a sequence we want the model to learn relative distances between tokens.
In other words, instead of learning the attention spent by token with index \(j\) on \(i\), that is, \(\alpha(\mathbf{x}_i, \mathbf{x}_j)\) we want to compute an attention score based on the (directed) distance \(i-j\).
This concept was introduced by <a class="citation" href="#shaw:2018">(Shaw et al., 2018)</a>.</p>

<div><img style="display:block; margin-left:auto; margin-right:auto; width:60%;" src="/Pages/assets/images/relative-attention.png" alt="Relative positional encoding" />
<div style="display: table;margin: 0 auto;">Figure 10: Relative positional encoding.</div>
</div>
<p><br /></p>

<p>Note that if there are \(\mathcal{O}(n)\) absolute positions \(0, 1, \ldots, n-1\) then there are \(\mathcal{O}(n)\) relative positions \(-(n-1), \ldots, -1, 0, 1,\ldots, n-1\).
The authors also introduce a maximal distance \(k\) such that they only learn weights</p>

\[\mathbf{w}^{V}_{\text{clip}(i-j,k)} \text{ with } \text{clip}(x,k) = \max(-k, \min(k,x))\]

<p>Therefore, they learn relative position representations for the keys \(\mathbf{w}^K_{-k}, \ldots, \mathbf{w}^K_{k}\) and for the values \(\mathbf{w}^V_{-k}, \ldots, \mathbf{w}^V_{k}\).
They introduce the relative position between \(\mathbf{x}_i\) and \(\mathbf{x}_j\) to be</p>

\[\mathbf{a}_{ij} = \mathbf{w}_{\text{clip}(i-j,k)}\]

<p>Thus there are \(\mathcal{O}(n^2)\) different such vectors but many share the same value.
Of course, they drop the absolute positional encoding.
And they adapt the <strong>self-attention computation</strong> from</p>

\[\mathbf{sa}_j(\mathbf{x}_0, \ldots, \mathbf{x}_{n-1}) = \sum_{i=0}^{n-1} \alpha(\mathbf{x}_i, \mathbf{x}_j) \mathbf{v}_i.\]

<p>to</p>

\[\mathbf{sa}_j(\mathbf{x}_0, \ldots, \mathbf{x}_{n-1}) = \sum_{i=0}^{n-1} \alpha(\mathbf{x}_i, \mathbf{x}_j) (\mathbf{v}_i + \mathbf{a}^V_{ij}).\]

<p>and the computation of the similarity between <strong>query</strong> and <strong>key</strong> from</p>

\[\frac{(\mathbf{W}_q \mathbf{x}_j)^\top (\mathbf{W}_k \mathbf{x}_i)}{\sqrt{D_q}}\]

<p>to</p>

\[\frac{(\mathbf{W}_q \mathbf{x}_j)^\top (\mathbf{W}_k \mathbf{x}_i + \mathbf{a}^K_{ij})}{\sqrt{D_q}}.\]

<p>Computation-wise the first manipulation can be easily achieved by adding a matrix \(\mathbf{A}\) to \(\mathbf{V}\).
However, the second manipulation destroys parallelism, i.e. the possibility to compute everything by matrix-matrix multiplications.
This can be mitigated by splitting the computation into two parts:</p>

\[\frac{(\mathbf{W}_q \mathbf{x}_j)^\top (\mathbf{W}_k \mathbf{x}_i + \mathbf{a}^K_{ij})}{\sqrt{D_q}} = \frac{(\mathbf{W}_q \mathbf{x}_j)^\top (\mathbf{W}_k \mathbf{x}_i) + (\mathbf{W}_q \mathbf{x}_j)^\top \mathbf{a}^K_{ij}}{\sqrt{D_q}}\]

<p>Assuming that \(\mathbf{S}\) contains all the relative attention scores, that is,</p>

\[\mathbf{S}_{ij} = (\mathbf{W}_q \mathbf{x}_j)^\top \mathbf{a}^K_{ij},\]

<p>then we can go back to the matrix form which gives us</p>

\[\mathbf{Sa}(\mathbf{X}) = (\mathbf{V} + \mathbf{A}^V) \cdot \text{Softmax}\left( \frac{\mathbf{K}^\top \mathbf{Q} + \mathbf{S}}{\sqrt{D_q} }\right).\]

<p>To compute \(\mathbf{S}\) <a class="citation" href="#shaw:2018">(Shaw et al., 2018)</a> instantiate an intermediate tensor \(\mathbf{R} \in \mathbf{R}^{k \times k \times D_q},\) containing the embeddings that correspond to the relative distance between all keys and queries.
\(\mathbf{Q}\) is then reshaped to an \((k, 1, D_q)\) tensor, and \(\mathbf{S} = \mathbf{Q} \mathbf{R}^\top.\)
This incurs a total space complexity of \(\mathcal{O}(k^2 D_q)\).</p>

<h2 id="the-music-transformer">The Music Transformer</h2>

<p>The Music Transformer <a class="citation" href="#huang:2018">(Huang et al., 2018)</a> was one of the first transformer utilized to generate symbolic music.
Even if it was introduced five years ago (which is like a century in the AI-world) it is worth studying it.
In the paper you find two different datasets</p>

<ol>
  <li><a href="https://github.com/czhuang/JSB-Chorales-dataset">J.S. Bach chorales dataset</a></li>
  <li><a href="https://www.piano-e-competition.com/">Piano-e-Competition dataset</a></li>
</ol>

<p>and they used an impressive sequence length of <strong>2048-tokens</strong>!
They used GPUs for the training.
With such a large number of token, one question arises: How did they manage to put 2000-tokens and all the respective matrices in the GPUs’ memory?</p>

<p>The authors correctly identify the space complexity of \(\mathcal{O}(k^2 D_q)\) to be problematic for GPU computation and they reduce the complexity to \(\mathcal{O}(k D_q)\) by exchanging space for re-computation.
This is possible due to the structure of the tensor \(\mathbf{R}\) which contains many equal values.</p>

<p>To handle very long sequences, the authors use local attention <a class="citation" href="#liu:2018">(Liu et al., 2018)</a> by chunking the input sequence into non-overlapping blocks.
Each block then attends to itself and the one before.</p>

<h2 id="attention-free-transformer">Attention-Free Transformer</h2>

<p>Basically, the attention mechanism, regardless of the specifics, solves a routing problem, that is, which information is transported to the next layer of the neural network.
Thus, it has a quadratic time and space complexity of \(\mathcal{O}(n^2)\) where \(n\) is our sequence length.
Therefore, if you have limited resources, it is hard to scale it to larger sequences.
As with the local attention and other techniques, like the Linformer <a class="citation" href="#wang:2020">(Wang et al., 2020)</a>, Longformer <a class="citation" href="#beltagy:2020">(Beltagy et al., 2020)</a>, Reformer <a class="citation" href="#kitaev:2020">(Kitaev et al., 2020)</a>, and Synthesizer <a class="citation" href="#tay:2021">(Tay et al., 2021)</a>, and Performer <a class="citation" href="#choromanski:2022">(Choromanski et al., 2022)</a> there are ways to improve this but in principle the complexity will bite us eventually.</p>

<p>Now we enter in an era in deep learning where we question if we actually need the attention layers in the transformer!
This was proposed in 2022.
Instead of computing attention, <em>FNet</em> <a class="citation" href="#leethorp:2022">(Lee-Thorp et al., 2022)</a> just mixes tokens according to the discrete Fourier transformation (DFT).
First, a 1D transformation is computed with respect to the embedding and then another with respect to time.
Amazingly even though there is no parameter to learn within the <code class="language-plaintext highlighter-rouge">Fourier</code>-layer (which replaces the <code class="language-plaintext highlighter-rouge">Head</code>) this strategy seems to work almost as good as the far more computationally expensive task of learning all the required attention scores.</p>

<p>The Fourier transform decomposes a function (in our case a discrete signal) into its constituent frequencies.
Given a sequence \(x_0, \ldots, x_{N-1}\), the discrete Fourier transform (DFT) is defined by</p>

\[X_k = \sum\limits_{n=0}^{N-1} x_n \exp\left( - \frac{2\pi i}{N} nk \right), \quad 0 \leq k \leq N-1.\]

<p>\(X_k\) encodes the <strong>phase</strong> and <strong>amplitude</strong> of frequency \(k\) within the signal.</p>

<p><em>FNet</em> consists of a Fourier <strong>mixing sublayer</strong> followed by a feed-forward sublayer.
Essentially, the self-attention sublayer of each transformer decoder layer is replaced with a <strong>Fourier sublayer</strong> which applies a 2D DFT to its</p>

\[(\text{sequence length} \times \text{hidden dimension})\]

<p>embedding input.
This can be achieved using two 1D DFTs—one 1D DFT along the sequence dimension, \(\mathcal{F}_\text{seq}\), and one 1D DFT along the hidden dimension, \(\mathcal{F}_\text{h}\):</p>

\[y = \text{Real}\left( \mathcal{F}_\text{seq} \left( \mathcal{F}_\text{h}(\mathbf{x}) \right) \right)\]

<p>The authors only consider the real part of the DFT.</p>

<p>Now, as emphasized by the title of their paper, the Fourier transform is probably not the important part.
It is just a special case of how you can mix tokens.
Important is the mixing itself which allows information to flow from one token to all the other tokens and the Fourier transform happens to be a nice way of mixing.
The paper indicates that it might not be so important to let the model learn how exactly information flows around.
It might be just enough if information flows at all (to all tokens).
In other words, the exact routing might be less important than we thought.</p>

<p>Now, the results of the paper are not better than using a traditional transformer.
But one trades accuracy for resources thus longer sequence length and a faster computation.</p>

<p>To the best of my knowledge, I have not seen this tried out for symbolic music generation.
But when I have time, I’ll play around with it.
Furthermore, one might think about a special mixing which is effective for our specific task.</p>

<h2 id="references">References</h2>

<ol class="bibliography"><li><span id="vaswani:2017">Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., &amp; Polosukhin, I. (2017). attention is all you need. <i>CoRR</i>, <i>abs/1706.03762</i>. http://arxiv.org/abs/1706.03762</span></li>
<li><span id="hochreiter:1997">Hochreiter, S., &amp; Schmidhuber, J. (1997). Long short-term memory. <i>Neural Computation</i>, <i>9</i>(8), 1735–1780. https://doi.org/10.1162/neco.1997.9.8.1735</span></li>
<li><span id="chung2014">Chung, J., Gulcehre, C., Cho, K. H., &amp; Bengio, Y. (2014). <i>Empirical evaluation of gated recurrent neural networks on sequence modeling</i>.</span></li>
<li><span id="bahdanau:2014">Bahdanau, D., Cho, K., &amp; Bengio, Y. (2014). Neural machine translation by jointly learning to align and translate. <i>CoRR</i>, <i>abs/1409.0473</i>.</span></li>
<li><span id="brown:2020">Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., … Amodei, D. (2020). <i>Language models are few-shot learners</i>.</span></li>
<li><span id="openai:2023">OpenAI. (2023). <i>GPT-4 rechnical report</i>.</span></li>
<li><span id="bubeck:2023">Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M. T., &amp; Zhang, Y. (2023). <i>Sparks of artificial general intelligence: Early experiments with GPT-4</i>.</span></li>
<li><span id="touvron:2023">Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., &amp; Lample, G. (2023). <i>LLaMA: Open and efficient foundation language models</i>.</span></li>
<li><span id="touvron:2023b">Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., … Scialom, T. (2023). <i>LlaMA 2: Open foundation and fine-tuned chat models</i>.</span></li>
<li><span id="devlin:2019">Devlin, J., Chang, M.-W., Lee, K., &amp; Toutanova, K. (2019). <i>BERT: Pre-training of deep bidirectional transformers for language understanding</i>.</span></li>
<li><span id="chen:2021">Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., … Zaremba, W. (2021). <i>Evaluating large language models trained on code</i>.</span></li>
<li><span id="huang:2018">Huang, C.-Z. A., Vaswani, A., Uszkoreit, J., Shazeer, N., Hawthorne, C., Dai, A. M., Hoffman, M. D., &amp; Eck, D. (2018). Music Transformer: Generating music with long-term structure. <i>ArXiv Preprint ArXiv:1809.04281</i>.</span></li>
<li><span id="huang:2020">Huang, Y.-S., &amp; Yang, Y.-H. (2020). <i>Pop Music Transformer: Beat-based modeling and generation of expressive pop piano compositions</i>.</span></li>
<li><span id="ens:2020">Ens, J., &amp; Pasquier, P. (2020). <i>MMM: Exploring conditional multi-track music generation with the transformer</i>.</span></li>
<li><span id="hadjeres:2021">Hadjeres, G., &amp; Crestel, L. (2021). <i>The piano inpainting application</i>.</span></li>
<li><span id="shih:2022">Shih, Y.-J., Wu, S.-L., Zalkow, F., Müller, M., &amp; Yang, Y.-H. (2022). <i>Theme Transformer: Symbolic music generation with theme-conditioned transformer</i>.</span></li>
<li><span id="copet:2023">Copet, J., Kreuk, F., Gat, I., Remez, T., Kant, D., Synnaeve, G., Adi, Y., &amp; Défossez, A. (2023). <i>Simple and controllable music generation</i>.</span></li>
<li><span id="shaw:2018">Shaw, P., Uszkoreit, J., &amp; Vaswani, A. (2018). Self-attention with relative position representations. <i>CoRR</i>, <i>abs/1803.02155</i>. http://arxiv.org/abs/1803.02155</span></li>
<li><span id="liu:2018">Liu, P. J., Saleh, M., Pot, E., Goodrich, B., Sepassi, R., Kaiser, L., &amp; Shazeer, N. (2018). Generating Wikipedia by summarizing long sequences. <i>International Conference on Learning Representations</i>. https://openreview.net/forum?id=Hyg0vbWC-</span></li>
<li><span id="wang:2020">Wang, S., Li, B. Z., Khabsa, M., Fang, H., &amp; Ma, H. (2020). <i>Linformer: Self-Attention with linear complexity</i>.</span></li>
<li><span id="beltagy:2020">Beltagy, I., Peters, M. E., &amp; Cohan, A. (2020). <i>Longformer: The long-document transformer</i>.</span></li>
<li><span id="kitaev:2020">Kitaev, N., Kaiser, Ł., &amp; Levskaya, A. (2020). <i>Reformer: The Efficient Transformer</i>.</span></li>
<li><span id="tay:2021">Tay, Y., Bahri, D., Metzler, D., Juan, D.-C., Zhao, Z., &amp; Zheng, C. (2021). <i>Synthesizer: Rethinking self-attention in transformer models</i>.</span></li>
<li><span id="choromanski:2022">Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J., Mohiuddin, A., Kaiser, L., Belanger, D., Colwell, L., &amp; Weller, A. (2022). <i>Rethinking Attention with Performers</i>. https://arxiv.org/abs/2009.14794</span></li>
<li><span id="leethorp:2022">Lee-Thorp, J., Ainslie, J., Eckstein, I., &amp; Ontanon, S. (2022). <i>FNet: Mixing tokens with Fourier transforms</i>.</span></li></ol>]]></content><author><name>Benedikt Zönnchen</name><email>benedikt.zoennchen@web.de</email></author><category term="Music" /><category term="ML" /><category term="Transformer" /><summary type="html"><![CDATA[Recurrent models trained in practice are effectively feed-forward. This could happen either because truncated backpropagation through time cannot learn patterns significantly longer than k steps, or, more provocatively, because models trainable by gradient descent cannot have long-term memory. – John Miller]]></summary></entry><entry><title type="html">Escaping the Reality of the Climate Crisis?</title><link href="https://bzoennchen.github.io/Pages/2024/01/02/conspiracy.html" rel="alternate" type="text/html" title="Escaping the Reality of the Climate Crisis?" /><published>2024-01-02T00:00:00+01:00</published><updated>2024-01-02T00:00:00+01:00</updated><id>https://bzoennchen.github.io/Pages/2024/01/02/conspiracy</id><content type="html" xml:base="https://bzoennchen.github.io/Pages/2024/01/02/conspiracy.html"><![CDATA[<p>Diverging from my area of expertise is always a risky endeavor, but since this is a blog and not a scientific journal, I’m giving myself the liberty to explore and have fun with different ideas (even if the topic is depressing). 
Often writing helps in transforming the mess into a structured and coherent concept.
The process of rethinking and reflecting can be invaluable.
It helps to make ones thought <em>anschlussfähig</em> which literally means <em>to be capable for connections</em> and in this context means <em>enabling the continuation of communication</em>.</p>

<p>In this piece, I aim to explore various aspects by linking the movie <em>The Matrix</em>, Plato’s <em>Allegory of the Cave</em>, <em>myths</em>, and <em>conspiracy theories</em>.
Additionally, delve into the understandable yet problematic skepticism surrounding <em>second-order observation</em>.
I will relate this skepticism to the notions of <em>complexity</em> and <em>hyperreality</em>, and discuss why we rely on <em>second-order observation</em> to address the <em>climate crisis</em>. 
Some parts of this text will escape a clear interpretation, especially when I offer my interpretation of Baudrillard’s writings.
Some parts might even be contradictory.
But this is the point: enduring or even enjoying ambiguity!</p>

<p>Overall, I hope to give good reasons for the emergence of conspiracy theories and our state of inaction in the face of disaster; reasons that do not rely on a good and evil dichotomy of individuals.
I will argue that our perception of reality makes effective communication between each other improbable and that complex systems follow their own stabilizing dynamic.
In the end, conspiracy theorist will appear as anti-authoritarian rebels that fight against the power of knowledge using the same source of authority, that is, knowledge.</p>

<p>Before we start, let us agree on a definition of conspiracy theories:</p>

<blockquote>
  <p>By conspiracy theory, I mean an explanation of historical, ongoing, or future events that cites as a main causal factor a group of powerful persons, the conspirators, acting in secret for their own benefit against the common good. – Joseph E. Uscinski</p>
</blockquote>

<h2 id="mythologies-science-and-conspiracy-theories">Mythologies, Science and Conspiracy Theories</h2>

<p>The phenomena of conspiracy theories hunts and intrigues me since the terror attacks of 9/11 happened back when I was a child.
I remember watching the <em>Zeitgeist series</em>, which linked various conspiracy theories involving religion, the September 11 attacks and the financial sector.
During that time, even German TV occasionally presented documentaries that portrayed certain events in a conspiratorial light.
The shock and uncertainty that the Western world experienced after these events, combined with a sense of lack of control, created a fertile ground for such theories. 
These ‘documentaries’ were not only entertaining, but they also sparked my interest in geopolitics, the history of religion, history in general, and even philosophy. 
They managed to make historical events captivating, and I often wished that my history classes were similarly engaging.
Fortunately, I never embraced the logic presented in these films; for me, they remained within the realm of entertainment but I could see how easy it is to fall for them on an emotional level.
Interestingly, years later, when a real plot unfolded to deceive the public (and other nations) into supporting the war against Iraq, there was no corresponding emergence of conspiracy theories like those seen previously.</p>

<p>During the pandemic, I observed a repetition of history in the form of similar documentaries emerging, and I became interested in how they were designed and how they relate to the conspiratorial ‘documentaries’ I encountered in my childhood. Indeed, they bore striking similarities.</p>

<p>Apart from the obvious parallels, such as misrepresentation and drawing connections between completely unrelated events, there were also more bizarre links.
For instance, these documentaries often incorporate some form of spiritual concept, promising a return to or discovery of a ‘true self’, an ‘inner peace’ or a ‘forgotten innocence’. 
For example, the <em>Zeitgeist series</em> uses a speech from the Indian philosopher Jiddu Krishnamurti (1895 – 1986) but only as an emotional device:</p>

<blockquote>
  <p>We will see how very important it is to bring about in the human mind the radical revolution.
The crisis is a crisis in consciousness, a crisis that cannot anymore accept the old norms, the old patterns, the ancient traditions.
And considering what the world is now with all the misery, conflict, destructive brutality, aggression, and so on, man is still as he was, is still brutal, violent, aggressive, competitive and has built a society along these lines. – Jiddu Krishnamurti</p>
</blockquote>

<p>In my interpretation, this reflects a lost connection to the wholeness of the universe. 
At the extreme end, I experience sometimes two contradictory moods. 
First, there is this feeling of alienation in an absurd world, which Albert Camus described best in his book <em>The Stranger</em>.
For me, the absurdity of the world reveals itself when I am out in the city, observing people rushing through their lives, participating in the acceleration towards a promised utopia that no longer exists, not even in their imaginations.
It is also connected to the realization that I was thrown into this world, culture, mess; this contradiction; this meaningless rat race.
The second mood arises from a profound connection with the world.
It feels as if we are the world, as if there is no separation between myself and my environment, between myself and the universe.
Yet, at the same time, I sense a process that is not only mysterious but overwhelmingly greater than I can ever comprehend; a force beyond my intellect.
Such feelings arise, for example, in nature.
When we stand atop a mountain, not thinking but simply being present in the world, or even feeling as if we are the world.
Such a feeling can also arise when we are completely immersed in an activity, like a drummer who dissolves entirely into the act of drumming.
This state of existence requires a quiet mind and a cessation of thought.</p>

<p>Consequently, it was no surprise to me that at the protests against COVID-19 restrictions and at QAnon gatherings, there was a peculiar blend of people, including everyone from far right-wingers to faith healers.
Of course, one has to be careful with such categorization since such protests are often captured by extreme parties.
Anyways, this mix reflects the broad, albeit unusual, appeal of such conspiratorial narratives and the desire for an effective (post-truth) narrator, e.g. Donald Trump, who provides a myth that carries emotional weight rather than the difficulty and uncertainty of a complex world—a <em>myth</em> one can live by.
However, <em>ignorance</em> is too simple of an explanation.
The MAGA cult is not disengaged in communication or the production of knowledge.
Instead they (ab)use communication to construct quite imaginative but also inconsistant <em>alternative facts</em> effectively.</p>

<p>On the basis of knowledge, it is easy to make fun of people believing in mythologies.
Scientifically speaking, myths are either inconsistent or unfalsifiable.
However, socially they can be very useful and powerful.
They explain experience and reduce complexity.
A myth is meant to answer questions and offer solutions to quell the anxieties of the present through stories.
They can serve a useful purpose, particularly in politics.
They create visions, unities and identities among groups such that they can work together towards a meaningful goal—a mission greater than oneself.
Here, a shadow of spirituality plays an important role, be it in the form of <em>the Light of God</em>, <em>Siegfried the Dragonslayer</em>, <em>Achilleus the Greatest of All the Greek Warriors</em> or <em>Donald Trump the Warrior King</em>.</p>

<p>Later I will argue that we should keep the source for spirituality, that is, imagination, ambiguity, and contradictions alive and that science can be abused for a <em>Crusade Against Ambiguity</em>.
In my view, spirituality and science can harmonize with quite well.
Religion, as a subset of spirituality, should be criticised especially if it becomes dogmatic, that is, if it starts a crusade against ambiguity.
But it is a narrow perspective to scientifically dispute the existence of a divine entity, just as it is misguided to interpret religious scriptures in a strictly literal sense.
Religion needs ambiguity.
Therefore, it is peculiar to witness esteemed intellectuals like Richard Dawkins engage in debates concerning the divine, overlooking the potential for a creator amidst the universe’s intricate complexity.</p>

<blockquote>
  <p>We have a working theory, which we know is true, which explains how you can go from great simplicity to prodigious complexity.
And finally to the sort of complexity which is capable of designing things, of creating things, of working out how to do things.
If you suddenly going to insert a designing machine, a creator, an intelligence at the root of the universe you have just undermined your entire enterprise because your entire enterprise has been to explain how you get to something complicated enough to do design. – Richard Dawkins</p>
</blockquote>

<p>When Dawkins discusses evolution, it is easy to agree with him from a scientific perspective.
I want him to defend his theory.
However, he often speaks in absolutes, using the language of dogmatic religion that he himself criticizes.
From a systems theory point of view, it can resolve this paradox by re-entering itself, that is, by communicating about itself on its own terms—a sort of self-reflection.</p>

<p>Dawkins and similar critics overlook a crucial point: the pursuit of absolute truth is elusive.
Even Dawkins’ assertion that “the theory is true” misrepresents the nature of scientific inquiry, which he is undoubtedly aware of.
Science provides models that explain phenomena until new evidence suggests otherwise.
It mosty works on the basis of <em>falsifiability</em> for <em>demarcation</em>, a concept introduced by Karl Popper (a critic of the inductive theory of science).</p>

<blockquote>
  <p>The problem of finding a criterion which would enable us to distinguish between the empirical sciences on the one hand, and mathematics and logic as well as ‘metaphysical’ systems on the other, I call the problem of demarcation. – <a class="citation" href="#popper:1934">(Popper, 1934)</a></p>
</blockquote>

<p>A statement or system of statements (theory) is falsifiable if it is capable of conflicting with possible, or conceivable observation;
The theory must be able to fail when tested against reality.
Build on rigor and the combination of theory bound by empiricism makes scientific method incredibly useful, and according to philosophers like Markus Gabriel, brings us closer to the Truth.
In contrast, others, such as Richard Rorty, completely reject the notion of absolute Truth and instead emphasize the practical usefulness of the scientific method.</p>

<p>The question of ontology remains unanswered even if most of us operate on the assumption of some sort of <em>naturalism</em>. 
It is the principle by which science operates and by which people in well-developed countries often live.
Some define it as “the idea that only natural laws and forces operate in the universe”.
The philosopher Quine described naturalism as “the position that there is no higher tribunal for truth than natural science itself”.
Quine’s more humble and pragmatic definition allows space for profound questions regarding the existence of time, space, causality, and our place within this framework—questions that lie beyond the scope of scientific inquiry.
For him and many others, these are absolute unknowns that can be considered within the realm of spirituality.
These unanswerable questions delve into the <em>essence of being</em> and the universe, inviting a spiritual exploration alongside scientific understanding.</p>

<p>Dawkins seems also to be unaware of the usefulness of ambiguity.
People can hold contradictory beliefs without being irrational or anti-science.
For example, it is unlikely that individuals with a non-dogmatic spiritual outlook adhere to a literal interpretation of the Earth’s creation in seven days or dismiss evolutionary theory but, at the same time, they might believe in a creator.
They can operate in different social systems with different rationals.
This is neither good or bad but a sign of diversity.
I mean how many mathematicians still believe that math has something to do with a divine entity or realm?
Such contradictions only become problematic if systems interfere in the other’s operations, e.g., if religion operates in science or science in religion.</p>

<p>But why can spirituality lead to the descent into the rabbit hole of conspiracy theories?
First of all, there is a strong relation between philosophy, religion, and spirituality, e.g., between Platonism and Christianity.
Monolithic religions offer a rather rigorous explanation of why things are as they are, based on the presumption of a creator.
Especially, non-believers sometimes misunderstand that religious people dislike logic when, in fact, it was Thomas Aquinas who attempted to synthesize Aristotelian philosophy (and logic) with the principles of Christianity.
He produced a vast body of precise, detailed, and systematic philosophical writings, in which he integrated Aristotle’s encyclopedic work and medieval Christian theology into a seamless whole.
The dark side of this was that any contradiction by future scientists would necessarily have to be seen as heresy.
Philosopher Bertrand Russell pointed out that Aquinas started by already knowing the truth in the form of the Catholic faith.
Aquinas used logic to strengthen his belief system and not to question it.
Note that postmodern thinkers argue that philosophers, who practiced metaphysics, did basically the same but in a more clever way.
Famous is Nietzsche’s suspicion of Kant’s categorical imperative, which is, after all, categorical.
Kant, however, pointed to the source of the problem which is not logic or rational thinking but a lack of empirical evidence; a lack of outwardness; of asking nature.</p>

<p>From this perspective, one might say that conspiracists are in the business of doing metaphysics poorly.
It is certainly the case that there are similarities in doing metaphysics and constructing a grand conspiratorial theory (or myth) that attempts to explain everything.
And like Aquinas, theorists of a conspiracy try to establish a kind of system, synthesizing the world into one big theory.
But there are also differences.
Metaphysicians (as well as many religious texts) at least try to be consistent, while conspiracy theorists are liberated from such limitations.
They openly replace rationality with mythology.
By constructing and emphasizing mythological symbols, conspiracy theories provide a shortcut into our soul, psyche, mind, or the unconscious.
They can switch seamlessly from one theory to another.</p>

<p>Myths can be very dangerous.
They misconstrue associations, destroy nuances and advance subconscious theses without the necessary burden of evidence.
Instead of delivering arguments, they short-circuit the entire argumentative process.
Myths do not make logical claims but significations <a class="citation" href="#barthes:1973">(Barthes, 1973)</a>.
For example, calling someone <em>a snake</em> is not a logical conclusion but signifies deceptive behaviour.
Real snakes, of course, are not significantly more or less deceptive than any other animal.
But mythological snakes often are and calling someone a snake can be a powerful gesture in our culture.</p>

<blockquote>
  <p>Poetry feeds and waters the passions instead of drying them up; she lets them rule, although they ought to be controlled, if mankind are ever to increase in happiness and virtue. – Socrates</p>
</blockquote>

<p>If Trump speaks of <em>America</em> or our radical right-wingers speak of Germany, they do not mean literal countries.
They signify a mythological symbol.
An effective myth is a self-contained world of signs were everything has a marked position, making it very hard to signify otherwise with a believer.
To critique their definition as being racist, irrational, exclusive, inhumane, or disastrous only demontrates to them that you are of ‘the them’ (das Man); that you are jealous that they won.
Any contrary narrative is spun by false prophets which conspire against the ‘chosen ones’.</p>

<h2 id="taking-the-wrong-pill">Taking the Wrong Pill</h2>

<p><em>The Matrix</em> is one of my all-time favorite movies, which increased my interest in computer science and philosophy.
As a child, I fantasized about being <em>The One</em>, akin to the hacker Neo, who could hack the matrix. 
Interestingly, the movie is inspired by French philosopher and theorist of postmodern media and culture, Jean Baudrillard (1929 – 2007), especially by his book <em>Simulacra and Simulation</em> <a class="citation" href="#baudrillard:1983">(Baudrillard, 1983)</a>.
The book even makes an appearance (as an empty prop) at the beginning of the film when Neo gives a disc to his clients.
The actors were even reportedly required to read it.</p>

<p>Baudrillard—the prophet of post-truth—focused on analyzing what can be termed <em>postmodern media</em>, although postmodernity is challenging to define.
Thinkers in the field of postmodern theory frequently hold different opinions but there is one core agreement: there are no all encompassing meta-narratives.
For some, such as Niklas Luhmann (1927–1998), the concept of postmodernity itself is contentious, with Luhmann believing it never truly existed <a class="citation" href="#luhmann:2000">(Luhmann, 2000)</a>. 
But back to Baudrillard.</p>

<p>He was an interesting figure but not taken very seriously by the academic community.
His writing style is polemical and his worldview extremely cynical. 
Despite this, his texts are intriguing and thought-provoking, capturing a sentiment that resonates with many facets of our society today.
In many ways, he was ahead of his time and highly influential in the media and culture he studied.</p>

<p>What particularly makes <em>The Matrix</em> fascinating in connection with Baudrillard is how the film embodies the type of pop-cultural phenomenon he often discussed in his philosophy.
The movie not only reflects his ideas but might be able to bring them to life in a way that is accessible to a broader audience.
Furthermore, Baudrillard was still alive when the movie hit the theatre.
So, did it succeed in bringing his theory to the big screen?</p>

<h3 id="the-simulacrum-is-true">The Simulacrum is True</h3>

<p>When we first meet Neo, his computer is active, processing something, with the screen reflecting on his face. 
He listens to music through headphones while lying on his desk, asleep. 
This scene introduces the difficulty of distinguishing between a dream and reality or more precisely, the problem of informational overload and sensory input that, according to Baudrillard, leads to passivity.
The abundance of disjointed information and excessive transparency makes it nearly impossible to organize the world and assign meaning to it—faces transform into screens or terminals that passively absorb.</p>

<p>In his early career, Baudrillard aimed to merge (post-)Marxism with (post-)structuralism but eventually abandoned the former.
He applied structuralism in his analysis of <em>The System of Objects</em> <a class="citation" href="#baudrillard:1968">(Baudrillard, 1968)</a>. 
Baudrillard theorized that the significance of commodities stems not primarily from their use or exchange value, but rather from their sign value. 
In structuralism, the meaning of elements, such as words, doesn’t derive from what they represent. 
For instance, teaching a child the word ‘tree’ isn’t as simple as pointing to one and stating, “Look, this is a tree!”
The child wouldn’t know if ‘tree’ refers to that specific tree, its leaves, or a category of trees. 
Understanding the word ‘tree’ requires knowledge of many other words and examining their relationships to ‘tree’—their difference.
Baudrillard argues that in a postmodern society, any cultural idea, image, sign, or symbol is apt to be pulled out of its social context and used (or abused) for advertisement and marketing.
The individual is placed in the position of a consumer.
As these signs are lifted out of the social, they lose all possibility of stable reference.
They may be used for anything, for any purpose.
All that remains is a yawning abyss of meaninglessness—a placeless surface that is incapable of holding personal identity, self, or society.</p>

<p>Baudrillard believed that in a postmodern society, the meaning and value of an object are primarily defined by its relationship to other objects. 
Apple products serve as a pertinent example. 
They appear overpriced when considering solely their use value.
However, their value arises from what they signify in relation to other objects which leads to the demishing of <em>symbolic values</em>.
For example, a pen given to you for your graduation, has probably a high symbolic value to you.
Symbolic values are assigned by a subject in relation to another subject.
Sign value, on the other hand, is the object’s value within a system of objects signifying, for example, social status.</p>

<p>Baudrillard, known for his cynical views, also believed that objects essentially have triumphed over subjects. 
He posited that just as money has become a universal medium that renders everything comparable and thus exchangeable, <em>the code</em> has made every sign integratable thus also exchangeable.
It is not that subjects or objects stand no longer for something ‘real’ but that the imagined referent, that does not exist, disappeared.
For example, in the Renaissance people or objects appear to stand for an imagined referent, for instance, royalty, nobility, holiness, etc.
A sign like Iron Man, stands for nothing other than itself in a network of other meaningless signs, i.e. the Marvel universe.
It can be repackaged into a toy or a specific McDonalds meal because it has no sacred connection to the world.
According to Baudrillard, instead of disimulating something, now signs dissimulate that there is nothing.</p>

<blockquote>
  <p>The transition from signs which dissimulate something to signs which dissimulate that there is nothing, marks the decisive turning point. 
The first implies a theology of truth and secrecy (to which the notion of ideology still belongs). 
The second inaugurates an age of simulacra and simulation, in which there is no longer any God to recognize his own, nor any last judgment to separate truth from false, the real from its artificial resurrection, since everything is already dead and risen in advance. – Jean Baudrillard</p>
</blockquote>

<p>This leads to a kind of dissolution of the ‘real’ meaning behind objects.
Accroding to Baudrillard, in the postmodern world, simulacra (e.g. images) have replaced the reality they once represented.
In other words, our current reality is dominated by these simulacra—representations, images, and signs—that no longer have any connection to any real/imagined object or event they might have originally represented.
Importantly, the simulacrum is not just covering up the truth or reality; it’s not a mask over something real!
In fact, it is quite the opposite.
What we perceive as truth or reality is actually just a construct (the simulacrum) that conceals the fact that there is no underlying, original reality; we enter simulation.
In other words, what we consider ‘real’ is just a construct of our perceptions and societal agreement.
In our current postmodern state, the simulacrum has become the truth for us, because there is no other reality against which to measure it.</p>

<blockquote>
  <p>The simulacrum is never what hides the truth—it is the truth that hides the fact that there is none.
The simularcum is true. – Jean Baudrillard</p>
</blockquote>

<p>Following Baudrillard’s perspective, experiences such as a teenager’s first kiss are no longer real in a sense that they express ‘true love’;
instead, they are mere simulations of a Hollywood love story because these stories are the truth!
Imaginations are not destroyed by hiding the truth but by showing it overtly naked, like pornography rips us of the imaginative allure of sexuality and intimacy.
Life imitates advertisement.
This does not mean that there is no more love.
However, there is nothing behind the ‘Hollywood love story’—it is true as it is.</p>

<p>We are compelled to reproduce these images and to participate, even if we know or suspect that it is all a simulation.
Critically, the problem (if it is in fact one) is not a virtualized reality that hides the truth, but that the truth is simulation.
People are fake but they are turthfully fake because being fake is the truth.</p>

<blockquote>
  <p>[…] pretending […] leaves the principle of reality intact: the difference is always clear, it is simply masked, whereas simulation threatens the difference between the ‘true’ and the ‘false’, the ‘real’ and the ‘imaginary’. – Jean Baudrillard</p>
</blockquote>

<p>When more and more simulacra transform into simulation we enter the matrix.
Pictures of ourselves no longer represent us, or us pretending to be someone else, but they are a simulation of some specific and often stereotypical fantasy that has no reference to something real other than different parts of the code.
That is the depressing and cynical viewpoint of Baudrillard.</p>

<h3 id="platos-allegory-of-the-cave">Plato’s Allegory of the Cave</h3>

<p>I love <em>The Matrix</em> but I have to assess that it did not succeed in capturing Baudrillard’s main themes.
The main problem is a clear line between simulation and reality; between the matrix and Zion.
The matrix clearly is not the truth but hides it.
Rather than exploring the new problem of simulation, the movie falls back on the <em>Allegory of the Cave</em> presented in Plato’s <em>Republic</em>.
Instead of investigating further questions, it postulates a true world behind the simulation by re-introducing religion.
Thus <em>The Matrix</em> brings us back where it all started but, according to Baudrillard, this is no longer possible.
In Baudrillard’s framework, the movie is itself the truth that hides the fact that there is none and therefore distracts the audience from acknowledging <em>hyperreality</em>.</p>

<p>The <em>Allegory of the Cave</em> is a metaphor for exploring the nature of knowledge and reality.
Plato imagined a group of people who lived their entire live chained inside a dark cave.
The only thing they can see are the shadows projected on the wall of the cave by objects passing in front of a fire behind them. 
These shadows are the only reality they know.
The cave dwellers believe the shadows to be the real objects, not knowing that these are mere reflections. 
Their knowledge and understanding of the world are based solely on this limited perspective.
One day, a prisoner breaks free. 
He struggles to adjust to the light outside the cave, but eventually, he sees and understands the true nature of reality. 
He realizes that the sun illuminates the world and that what he saw in the cave were just shadows of real objects.
The freed prisoner returns to the cave to enlighten the others. 
However, his eyes have adjusted to the sunlight, so the cave is now blindingly dark to him. 
The other prisoners, unable to understand his experiences and seeing his blindness in the dark, refuse to believe him. 
They cling to their old beliefs about the shadows being the real objects.</p>

<p>In Plato’s metaphysics, the form (true essence of things) are more real than their physical representations.
The shadows represent the physical world, while the objects outside the cave symbolize the forms.
Plato also emphasizes that education is not just a matter of transferring information, but a transformative experience that leads to understanding, or to the seeing of a different world that opens up.
Of course, it is the philosopher that seeks the truth (outside the cave) and then attempts to bring this knowledge back to the people (inside the cave).</p>

<p>We can draw a neat parallel between Plato’s <em>Allegory of the Cave</em> and the film <em>The Matrix</em> by replacing the <em>cave</em> with <em>the matrix</em> and the philosopher with the character Morpheus. 
In this parallel, Morpheus takes on the role of guiding Neo (and others) out of the matrix—similar to leading prisoners out of the cave. 
Neo learns to understand and manipulate the matrix on his terms, which parallels the ability to manipulate the shadows on the cave walls.</p>

<p>In <em>The Matrix</em>, Neo is confronted with a crucial binary decision, symbolized by the choice between a red and a blue pill. 
As revealed in the sequels, this choice is, in itself, a part of a predetermined simulation—a much more Baudrillardian take.
Morpheus presents Neo with this decision and advises him to trust his instinct that something is fundamentally wrong with the world. 
This guidance emphasizes the importance of an emotional rather than a logical conclusion, steering Neo to follow his feelings in making this pivotal choice.</p>

<p>Once Neo makes his choice, the distinction between the matrix and reality is clear to him; it is a clear binary: reality and simulation.
True love is still possible outside and even inside the matrix, even if it is predetermined.
According to Baudrillard reality is a simulation echoing Kant, who does not grant us the access to the <em>thing-in-itself</em>, and of course Nietzsche, who tells us that there are only <em>constructed</em> values.
Simualation is nothing bad or something to fear.
It was always already there but the simulacrum (e.g. cave paintings, images) changed towards its own gravity; towards its own perfection.
What Baudrillard feared is a world akin to the movie <em>Minority Report</em>.
A world without <em>reversability</em> where everything is already decided in advance; a world without <em>ambiguity</em>.
In a sense, Baudrillard feared the modern project that started with Plato by looking for some absolut truth or perfect idea.
The perfect simulation gets rid of illusions and imaginations; it is too real; it is <em>hyperreal</em>;
Therefore, Baudrillard did not fear the loss of reality but an exzess of it which would lead to the destruction of illusions and imaginations like the ‘technical perfection of sex’, i.e. pornography, leads to the removal of sexuality and intimacy.</p>

<blockquote>
  <p>Reality and simulation aren’t opposed to one another. 
There are two sides of the same coin. – <a class="citation" href="#baudrillard:2008">(Baudrillard, 2008)</a></p>
</blockquote>

<p>With this in mind, if we reexamine <em>The Matrix</em> it becomes clear that Baudrillard would describe it to be a pretty good simulation of the matrix.
In a sense, the sign of simulation is re-integrated into the simulation itself.
This re-integration highlights the film’s exploration of reality, perception, and the nature of choice, themes that resonate with Plato’s allegory but not with <em>Simulacra and Simulation</em>.
Thus Baudrillard concluded:</p>

<blockquote>
  <p>The radical illusion of the world is a problem faced by all great cultures, which they have solved through art and symbolization.
What we have invented, in order to support this suffering, is a simulated real, which henceforth supplants the real and is its final solution, a virtual universe from which everything dangerous and negative has ben expelled.
And The Matrix is undeniably part of that.
Everything belonging to the order of dream, utopia and phantasm is given expression, ‘realized’.
We are in the uncut transparency.
The Matrix is surely the kind of film about the matrix that the matrix would have been albe to produce. – Jean Baudrillard</p>
</blockquote>

<h2 id="a-hunger-for-certainty-and-definitude">A Hunger for Certainty and Definitude</h2>

<p>Now, what has this to do with <em>conspiracy theories</em>?
Well, other than Baudrillard’s theory of a reality that is simulation, I claim that the <em>cave allegory</em> offers a theoretical justification for doubting established institutions, which are likened to ‘the matrix’.
One might discover that parts of reality, e.g. institutions, norms, moral judgements, ideologies, is constructed and that there has to be something real behind it.
This suspicion of a matrix is not unfounded, however, believing in some <em>absolut point</em> of view behind it, opens the door to confusion.
We are in a cave but going outside might only lead to another cave.
Importantly, this does not mean that any cave is as useful or functional!
Although not all interpretations of a text are equally meaningful, a good text offers numerous interesting and valuable interpretations and the same seems to be true of our <em>lifeworld</em>.</p>

<p>Doubting parts of reality—a known or presented world—is the starting point of any conspiracy theory.
The perspective is compelling because it feeds the allure of knowing a secret, akin to Neo’s experience in the matrix.
Such knowledge is seen as something that sets an individual apart from ‘the herd’, giving them a sense of being special or enlightened.
It is similar to <em>New Age</em> beliefs in some sort of special knowledge about the universe presented in movies like <em>The Secret</em>.</p>

<p>Belief in conspiracy theories appears to be driven by motives that can be characterized as epistemic (understanding one’s environment), existential (being safe and in control of one’s environment), and social (maintaining a positive image of the self and the social group) <a class="citation" href="#douglas:2017">(Douglas et al., 2017)</a>.
One important facet of conspiracy theories that often goes without much notice is that they are notions about power: who has it and how are they using it?
Conspiracy theories accuse an implicitly powerful group of conspiring.
Usually that group is already powerful—even if that power is a fantasy—i.e., the president, a legislative body, industries or corporations, foreign countries, multinational groups, etc. 
Powerless groups are rarely accused of conspiring <a class="citation" href="#uscinski:2018">(Uscinski, 2018)</a>.
This also reflects the plot of <em>The Matrix</em> where agents of the matrix are much more powerful than ‘enlightened’ humans.</p>

<p>Studies show that some people are more prone to believing in conspiracy theories than others.
Some people will believe in any conspiracy theory even on light evidence while others, at the opposite end of the spectrum, are naive and will deny the existence of conspiracies even on accumulating evidence <a class="citation" href="#uscinski:2018">(Uscinski, 2018)</a>.
According to Jan-Willem Prooijen, conspiracy theories orginate through the same cognitive process that produce other types of belief (e.g. spirituality), they reflect a desire to protect one’s own group against a potentially hostile outgroup, and they are often grounded in strong ideologies.
They are a natural defensive reaction to feelings of uncertainty and fear <a class="citation" href="#prooijen:2018">(Prooijen, 2018)</a>.</p>

<p>Interestingly, in studies, individuals who perceive patterns in abstract paintings, random dots, or coin tosses were more inclined to believe in conspiracy theories, paranormal phenomena, and hold religious beliefs. 
Belief in conspiracies also tends to rise during natural disasters when people feel a lack of control. 
Due to their tendency to seek patterns, conspiracy theorists tend to categorize everything neatly into a framework of good versus evil.
Even though the world that is constructed is miserable, it is without uncertainty.</p>

<p>We are risk calculating creatures, always on the watch for new dangerous patterns.
This is evolutionary advantageous.</p>

<blockquote>
  <p>Conspiracy is a stubborn creed because humans are pattern-seeking animals.
Show us a sky full of stars, and we will arrange them into animals and giant spoons.
Show us a world full of random misery, and we will use the same trick to connect the dots into secret conspiracies. – Jonathan Kay <a class="citation" href="#kay:2011">(Kay, 2011)</a>.</p>
</blockquote>

<p>Perceiving patterns is the opposite of perceiving randomness, and randomness cannot be the basis for making sense. 
Of course, quite often, events occur randomly, without any discernible purpose or meaning. 
Sometimes, foolish mistakes simply happen unintentionally.</p>

<p>Research indicates that education reduces the likelihood of believing in conspiracy theories (with exceptions). 
This may initially appear counterintuitive because education encourages skepticism toward received wisdom. 
Shouldn’t skepticism lead one to think that there might be something hidden behind the scenes?
Well, skepticism is only one aspect of the equation. Education teaches individuals to scrutinize the evidence and seek primary sources, or more precisely, to consider <strong>all</strong> available evidence.
Under such scrutiny conspiracy theories fall apart.
Moreover, having a greater understanding tends to foster humility since individuals become increasingly aware of the vast expanse of knowledge that remains beyond their grasp.</p>

<p>Scrutiny is built into our institutions.
In the context of academic publishing, one must demonstrate a comprehensive understanding of the relevant literature and show how experiments can be reproduced to validate the published findings.
A peer review process by experts in the field checks for the soundness of the work and identifies possible errors. 
While this system is not flawless, makes mistakes, overemphasis the number instead of the value of publications, can be biased, and often favours the middle to upper class, it remains open to critique and has propelled us a long way.
Science is a discourse.
Theories are never absolute true but are seen as true as long as there is no evidence or proof that gives rise to different conclusions.
Mistakes have been made and will continue to be made in the future, but the system is self-correcting, self-preserving, and has advanced our knowledge considerably.</p>

<p>Conspiracy theories are also fueled by our cognitive biases. 
For instance, the <strong>proportionality bias</strong> tends to make us believe that a substantial effect must have a significant cause. 
Consider a scenario where either a neighbor or the President of the United States dies randomly; which one is more likely to trigger a conspiracy theory?
Studies have revealed that when people are informed about the assassination of a president, they are more inclined to believe in a conspiracy theory if it coincides with the outbreak of a subsequent civil war <a class="citation" href="#prooijen:2018">(Prooijen, 2018)</a>.
<strong>Tribalism</strong> encourages us to protect our own ingroup and establish a clear division between ‘us vs. them’, often framing it as a battle between good and evil.
The <strong>intentionality bias</strong> leads us to believe that negative consequences of our actions are unintentional, while attributing intentionality to others when they cause harm. 
For example, we may view bankers as evil, but perceive our own pension fund as a necessary institution.</p>

<p>Seeing patterns everywhere is the need for control <a class="citation" href="#shermer:2022">(Shermer, 2022)</a>.</p>

<blockquote>
  <p>The economy is not this crazy patchwork of supply and demand laws, market forces, interest rate changes, tax policies, business cycles, boom-and-bust fluctuations, recessions and upwings, bull and bear markets, and the like.
Instead, it is a conspiracy of a handful of powerful people variously identified as the Illuminati, the Bilderberger group, the Council on Foreign Relations, the Trilateral Commission, the Rockefellers and Rothshields.
[…] Conspiracists believe that the complex and messy world of politics, economics, and culture can all be explained by a single conspiracy and conspiratorial event that downplays chance and attributes everything to this final end of history. – Michael Shermer</p>
</blockquote>

<p><a class="citation" href="#landau:2015">(Landau et al., 2015)</a> show that people compensate for perceived loss of control by trying to restore control themselves by</p>

<blockquote>
  <p>bolstering personal agency, affiliating with external systems perceived to be acting on the self’s behalf, and affirming clear contingencies between actions and outcomes [… and] seeking out and preferring simple, clear, and consistent interpretations of the social and physical environments.</p>
</blockquote>

<p><strong>Narcissism</strong>, characterized by a belief in one’s superiority and the desire for special treatment, strongly correlates with a tendency to believe in conspiracy theories. 
Narcissists also exhibit heightened sensitivity to perceived threats <a class="citation" href="#cichocka:2022">(Cichocka et al., 2022)</a>.
Within the realm of narcissism, grandiose narcissists seek admiration by bolstering their egos through a sense of uniqueness, charm, and grandiose fantasies. 
It’s worth noting that narcissists often display naivety and are less likely to engage in <strong>cognitive reflection</strong>.
Surprisingly, studies have uncovered evidence suggesting that, contrary to expectations, education increases the likelihood of narcissists adopting conspiracy beliefs <a class="citation" href="#cosgrove:2023">(Cosgrove &amp; Murphy, 2023)</a>. 
This underscores the critical role of cognitive reflection as one of the most, if not the most, essential abilities to guard against narcissistic tendencies towards conspiracy beliefs.</p>

<blockquote>
  <p>[Conspiracy believers] are relatively untrusting, ideologically eccentric, concerned about personal safety, and prone to perceiving agency in action – <a class="citation" href="#hart:2015">(Hart &amp; Graether, 2015)</a></p>
</blockquote>

<p>Similar to the experience of emerging from the cave in Plato’s allegory, delving into a conspiracy theory is not merely a transfer of knowledge, but a transformative experience. 
It involves a world being shattered and a new one being constructed in its place.
This process signifies a profound shift in perception and understanding, where previously accepted realities are dismantled and replaced with an entirely different framework of belief and interpretation.</p>

<p>Like Morpheus’ emphasis on trusting one’s instincts in <em>The Matrix</em>, conspiracy theories often accurately capture the emotional aspects of a person’s situation. 
These theories provide compelling descriptions of emotional states but tend to offer simplistic and reactionary explanations for complex situations. 
Additionally, much like the concept of the matrix, they seek an all-encompassing explanation for everything.
Essentially, these theories represent a futile effort to eliminate contingency and the future’s uncertainty. 
They attempt to provide a sense of certainty and understanding in a world that (hopefully) is still inherently open, contingent, unpredictable and complex. 
This desire for comprehensive explanations reflects a deep-seated human need for order and predictability in an increasingly fatal looking world.</p>

<h2 id="escaping-the-simulation">Escaping the Simulation?</h2>

<p>The incorporation of expressions like ‘escaping the matrix’ and ‘being red-pilled’ into the vocabulary of conspiracy theory groups as metaphorical language is unsurprising.
These terms, which originated from <em>The Matrix</em>, are used metaphorically to describe the experience of awakening to a hidden or suppressed truth.
This desire might increase with the suspicion that there is none.
Specifically, ‘escaping the matrix’ denotes the recognition and liberation from a controlling system or an illusionary world, while ‘being red-pilled’ represents a moment of profound revelation or enlightenment, often regarding societal structures or purported conspiracies.
These metaphors have gained traction within certain groups as a way to express their beliefs in uncovering what they perceive as hidden truths within society.
They believe to be the philosophers of our age, teaching us how real man behave and how ‘the system’ keeps them weak and small.</p>

<p>Contrary to the common belief that ignorance fuels the acceptance of a matrix-like reality, it is curiosity that often propels this belief. 
People are attracted to the notion of discovering hidden truths and understanding the world in ways that differ from the majority’s perspective.
The group forms, in a manner reminiscent of a cult, and establishes easily comprehensible guidelines, resembling the revolutionaries from another pop-cultural and frequently misunderstood film—<em>Fight Club</em>.
The theme of ‘stepping out of the dark’ is central to Plato’s cave allegory, <em>The Matrix</em>, and various conspiracy theories. 
Such curiosity ignites a desire to investigate and question conventional narratives, leading some individuals to adopt alternative interpretations of reality.</p>

<p>Interestingly, critical thinking and intelligence do not necessarily prevent one from falling into this rabbit hole.
As described above, these attributes can sometimes drive narcissits deeper into exploring and accepting these alternate realities—the problem is a lack of reflection.
The quest for understanding and the allure of uncovering hidden knowledge can be so compelling that even the most critical and intelligent minds are susceptible to these alternate explanations.
Take the following speech performed by the actor James Caviezel (an actor I once admired for his role in <em>The Thin Red Line</em>) at the end of <em>Sound of Freedom</em>, a conspitorial movie about child trafficking:</p>

<blockquote>
  <p>While watching this movie, I guess some of you were feeling sad, maybe overwhelmed, or even feel a sense of fear, which is understandable. But living in <strong>fear</strong> isn’t how we solve this problem. It’s living in <strong>hope</strong>. It’s <strong>believing that we can make a difference</strong> because we can.
I want to make one thing clear: this movie you just watched isn’t about me or Tim Ballard. 
It’s about those kids. This film was actually made five years ago. 
It wasn’t released until now, with every roadblock you can imagine being tossed in our way. [The powerful do not want you to see it, believe me]. 
And the names you see here, on the screen, <strong>took a stand</strong>! They made sure that this story could be shown to all of you. Now, all of you have the opportunity to continue <strong>telling this story</strong>; [the Truth].
We don’t have big studio money to market this movie [(we only have Fox News, one of the biggest network in the country)], but we have <strong>you</strong>. The baton has now been passed to <strong>you</strong>. <strong>You</strong> are the storytellers who can get people to come see this film in theaters. Together, we have a chance to make these two kids and the countless children they represent the most powerful people in the world by telling their story in a way only cinema can.
For a couple of months, while Sound of Freedom is in theaters, these kids can be more powerful than the cartel kingpins, presidents, congressmen, or even tech billionaires. We believe this movie has the power to be a huge step forward toward ending child trafficking, but it will only have that effect if <strong>millions of people see it</strong>.
We don’t want finances to be the reason someone doesn’t see this movie, so Angel Studios has set up a forward program where you can pay for someone else’s ticket who might not otherwise see it. If you’re able, we invite you to pay it forward by buying a ticket for someone else, or if your budget is tight, share the already available free ticket with as many friends as you can.
Join us and millions of others as we ring Sound of Freedom and hope throughout the world. And just remember this: <strong>God’s children are not for sale</strong>.</p>
</blockquote>

<p>During the speech a QR Code and the text “Give an Share Tickets / ANGLE.COM/freedom” is displayed.</p>

<p>This speech is filled with pathos, it is shamelessly manipulative, moralistic, heroistic and serves a narcissist desire to ‘wake up’, spread the word of Truth or God and become the hero; a soldier of Freedom; an angle of God; a righteous martyred that saves us all.
It is also conspiratorial.
Of course, in the same breath Caviezel tells us to buy tickets to end child trafficking which is not only irrational but flat out unethical and morally dubious.
It is so obviously a scam that it becomes an interesting field of study why people buy into it.
Does the deeply religious actor James Caviezel believe what he is saying? I think he does on some level.
He explained at a promotion event for the movie that billionaires capture children to extract adrenalin out of their blood when they are scared of death, which is a famous QAnon conspiracy theory.</p>

<blockquote>
  <p>These people that do it; there will be no mercy for them! – James Caviezel</p>
</blockquote>

<p>The speech is also kind of Baudrillardian in that sense that even the fight for the Good is just a struggle to buy that god damn ticket; even God’s angles are reduced to passive consumers.
At the same time it is not about the kids but a much more sacred war of Good against Evil—a mystic war in a very <strong>angry</strong> United States of America!
A country that is at the brink of an inwardly directed outburst.
Everyone in the media is so angry all the time.
This angry speech and movie, that abuses religion, confirms my believe that I should not fear the people who name themselves after the Devil but those who name themselves after a righteous God.
Nietzsche was right about that.</p>

<p>Now, it is a fact that people do engage in conspiracies (and that child trafficking is a big problem).
A brief examination of history reveals numerous instances of conspiracies, some of which have even led to wars between nations. 
However, in retrospect, these conspiracies can often be explained without assuming the involvement of thousands of people. 
The complexity and impact of these historical events do not necessarily require large-scale collusion; often, they can be understood through the actions and decisions of a relatively small number of individuals or groups and through a systemic rationality. 
This understanding helps differentiate between plausible historical conspiracies and the more elaborate, less credible theories that claim widespread secret collaboration.
If it exists, the matrix is not a planned construction of anybody but a Baudrillardian process beyond anyones control.</p>

<p>In addition, it’s important to recognize that the world is inherently unjust. 
Justice is a human concept, one that evolves over time as we make what we call progress. 
However, in the realm of nature, there is no concept of justice at least none I am aware of;
nature operates outside of morality.
Absolute justice remains elusive and if we seek an explanation for the world’s injustice, conspiracy theories provide a sense of comfort.</p>

<p>It appears to me that the belief in having escaped the matrix or emerged from the cave is a strong indication of someone having entrenched themselves deeply in their own perspective—failing to see that their perspective also relies on some sort of <em>second-order observation</em>; a following of the herd, or in Baudrillard’s viewpoint, the false assumption that there is something true behind the simulation.
This belief offers comfort by addressing various uncertainties and the realization of one’s own ignorance. 
No one desires to be ignorant and no one wants to rely on some sort of authority, yet in many ways, we all are.
In our complex world, this is an unavoidable reality. 
Conspiracy theories provide a sense of understanding and control in a world where complete knowledge is unattainable, helping individuals cope with the inherent limitations of human understanding.</p>

<p>Therefore, I believe that individuals who are particularly uncomfortable with uncertainty, who seek control over their life, and who are actively aware of their lack of control, are more susceptible to falling into these rabbit holes of alternate realities. 
Additionally, a certain degree of narcissism may be necessary to believe in the premise that one possesses a superior ability to understand complex matters better than trained experts and to assume that the media consisting of hundred of thousand of journalists is a monolith.
This combination of a need for control, discomfort with uncertainty, and a self-perceived exceptional understanding can lead individuals to embrace alternative explanations that offer a sense of clarity and personal significance in a complex world.</p>

<p>If we contemplate the matrix envisioned by Baudrillard, then attempting to escape it through the immersion in an alternate version of reality, fostered by extensive consumption of social media and digital content, appears absurd. 
Perhaps a more appropriate approach to disengaging from the machinery of simulation and countering the sensation that reality seems increasingly tenuous is to simply disconnect from it all (from time to time).
When advertisements, repetitive media, individuals transformed into brands, and an incessant stream of content seize our attention, the signal overflow—the noise—overshadows a more tangible reality. 
In a scenario where there may be no external escape from the simulation, it could be valuable, from time to time, to focus on what is immediately before us: to experience, touch, smell, listen to our bodies, engage with physical sensations, concentrate, savor awareness, and relinquish the illusion of the “real” by re-connecting to a spiritual world.</p>

<p>We are not superheroes; we are composed of the same fundamental elements as everything else. 
While it might feel like we inhabit a sort of matrix, it’s essential to acknowledge that this is a choice we make. 
From childhood, we develop self-conceptions and fantasies, but this doesn’t negate the existence of the world itself. 
Fantasies are constructs, and doubting the existence of the world presupposes a profound level of experience and knowledge of that world—a world where we learn to eat, walk, dance, and understand the nuances of correct and incorrect language usage.
Doubting it requires a distance from it and that might be what social media does: it shrinks the world but increases the distance to it.
We often employ our habits so routinely that we forget we are employing them, and in doing so, we forget that the world—our home—is still there. 
The central question here is what is more reasonable to doubt: the world we intimately grew up in or our doubts about doubting it?</p>

<h2 id="our-dependence-on-second-order-observation">Our Dependence on Second-order Observation</h2>

<p>Thinking critically and maintaining a sense of curiosity are attributes that I certainly hope everyone possesses.
It is crucial that institutions, including large media operations, research institutions, and particularly governments, are consistently challenged and kept under close scrutiny.
Conspiracy theories can serve as a force to encourage the prevention of corruption.
I would be quite suspicious if there were no conspiracy theories present!</p>

<p>Simultaneously, these theories have the potential to divert attention from genuine issues. 
The public should advocate for accountability and transparency, cultivating a healthy and well-informed society in which decisions and policies undergo scrutiny and improvement through public discourse and critical examination.</p>

<p>However, as the current state of affairs stands, the notion that everyone can participate in the <em>marketplace of ideas</em> and engage in public discourse seems somewhat impractical and utopian. 
This dream may appear overly optimistic, excessively humanistic, and excessively individualistic. 
Instead, according to Luhmann, there exists an interdependent network of social systems that co-evolve together—not individual souls but interconnected systems.</p>

<p>Similarily we demand the media should try to be as objective as they can be but it is naive to think that they are able to present reality as it is.
Here I agree with Luhmann:</p>

<blockquote>
  <p>It is impossible to understand the reality of the mass media if you assume it is their job to provide correct information on the world and then assess how they fail, distort reality, and manipulate opinion—as if they could do otherwise. – Niklas Luhmann</p>
</blockquote>

<p>Basically, Luhmann observed something very similar to the <em>Manufacturing Consent</em> <a class="citation" href="#herman:1988">(Herman &amp; Chomsky, 1988)</a> but explains it slightly differently without the need for a <em>propaganda model</em>.
If one anticipates that the primary purpose of the media is to deliver accurate information or facts, they are likely to encounter inconsistencies that raise significant doubts about the credibility of the mass media apparatus. 
The media is inherently self-preserving.
It functions in a manner that constructs and sustains itself. 
While it is certainly beneficial for the media to provide accurate information, this is not its foremost objective.
The media provides <em>what is known to be known</em>.
It irritates politics, our economy and the scientific system while bing irritated by all these systems.
<strong>The media makes society restless</strong>.</p>

<p>The view that the media presents ‘the Truth’ contributes to the proliferation of conspiracy theories, since it portrays the entire system as corrupt.
If the media presents objective facts and these facts are inconsistent, distorted, incomplete and open for interpretation then it is disfunctional or corrupt.
Instead of recognizing the various shortcomings (which serve the internal logic of the system) within the media (of which there are many), one tends to assume a broad conspiracy aimed at deliberately deceiving the public.
This leads to a pervasive distrust, particularly directed towards well-established media outlets.
As a consequence, consumers may turn to ‘alternative’ media sources, even though these alternatives often inadvertently rely on established media institutions for their information. 
Reporting, conducting on-site investigations, collecting information, and managing extensive archives are expensive endeavors that only large institutions can effectively undertake. 
These institutions are essential if we are to have any hope to share a world that is at least partly commonly known.</p>

<p>The core issue lies in our reliance on what Niklas Luhmann refers to as <em>second-order observation</em> and is consequently a trust issue.
In modern society, directly observing reality is increasingly challenging. 
To stay informed about various aspects such as the state of the economy, job market trends, recent fashion styles, developments in one’s favorite sports league, or new scientific inventions and studies, it is impractical to personally verify these facets. 
Instead, we depend on the observations made by others and, of course, machines.
This means we have to engage with various forms of reporting and analysis: reading reports about the GDP, considering the opinions of fashion critics, watching sports programs, and reviewing scientific papers.
In science, we write review papers about review papers.
We track how often papers are cited, i.e. how these papers are being observed.
This reliance on second-hand information shapes our understanding of the world, as we depend on external observers to provide us with insights and knowledge about various domains that we cannot directly experience or verify ourselves.</p>

<p>We are frequently depend on multiple layers of <em>second-order observations</em> or various levels of abstraction.
Scientific papers serve as a prime example. 
These papers are typically not intended for a general audience but are meant for peers within the specific field of research. 
As a result, the average person often finds them inaccessible.
Consequently, we turn to science communicators and mass media to distill and present scientific information. 
These intermediaries play a crucial role in interpreting and translating complex scientific data and studies into ‘facts’ that are understandable and relevant to the general public. 
This reliance on filtered and simplified interpretations highlights our dependence on external sources to understand and engage with specialized knowledge areas.</p>

<p>At the core of our society is the notion of the individual as a subject capable of making their own decisions and drawing sound conclusions. 
However, it’s evident that our understanding of the numerous processes occurring around us is limited. 
The complexity of the modern world might only be manageable through <em>functional differentiation</em> and <em>second-order observation</em>. 
We are heavily dependent on specialists and experts, and our understanding is largely shaped by observing their observations.</p>

<p>Conspiracy theorists seemingly reject <em>second-order observation</em>, viewing it as a form of manipulation akin to the matrix. 
However, this rejection is a perilous illusion.
There is no position outside of second-order observation, no external vantage point from which to objectively assess ‘reality’ as it is, separate from the interpretations and understandings provided by others.
This perspective underscores the intricate and interconnected nature of knowledge and understanding in contemporary society.</p>

<p>Philosophers ranging from Plato, Fichte, and Kierkegaard to Russell, Kant, and Heidegger have provided insights that prompt us to question the application of second-order observation.
These philosophical teachings encourage us to contemplate whether we should exercise independent thinking and challenge the prevailing mainstream narrative.
This concept is epitomized in Heidegger’s notion of avoiding assimilation into <em>das Man</em> (the they), Kierkegaard’s emphasis on distancing oneself from the public, or Fichte’s focus on the ego. 
The stories goes like this: There exists an inner truth within us, and we should search within our <em>authentic</em> selves to discover it. 
As sovereign individuals in a libertarian society, we should not solely rely on the opinions, or observations, of others. 
As Kant famously articulated, we should have the courage to employ our own intellect (Verstand). 
I align with Kant with a caveat: we should also have the courage to acknowledge our own ignorance and cultivate the ability to rectify it.
By utilizing second-order observation wisely, we can develop a <em>cultural intelligence</em> more akin to Hegel’s concept of the world spirit than Kant’s emphasis on the individual.</p>

<h2 id="conviction-under-constructivism">Conviction under Constructivism</h2>

<p>Biologists Humberto Maturana, Francisco Varela, Samy Frenk, and Gabriela Uribe made a significant discovery regarding our understanding of color perception.
Rather than focusing solely on the correlation between the physical source of color and the retina’s response, they emphasized another more important correlation: the one between the retina and subjective color perception. 
In this context, the external source of color functions as a trigger, not the sole determinant.</p>

<p>This structure of subjective color perception effectively maintains the perception of colors under various objective conditions, even when there are substantial discrepancies between the perceived and ‘emitted’ color, as seen in deception experiments. 
Maturana and Varela extended this insight to introduce the concept of autopoesis, integral to the biological theory of cognition <a class="citation" href="#maturana:1987">(Maturana &amp; Varela, 1987)</a>.</p>

<p>In line with their constructivist theory, every individual constructs their own cognition and, by extension, their reality.
In fact, for Maturana cognition is living.
His theory seems not far from Kant’s and to an extend Fichte’s understanding of cognition.
Constructivism does not imply an absence of a single reality or a state of complete subjectivity, but rather, it leads to some noteworthy conclusions:</p>

<ol>
  <li>An absolute system of values and knowledge cannot exist because personal experience forms an unshakable foundation.</li>
  <li>Convincing someone can only succeed when they develop their own system of conviction.</li>
  <li>Humans, capable of observing their cognitive actions and recognizing the relativity of their seemingly valid knowledge, face the responsibility of choosing and adhering to their own value system.</li>
</ol>

<p>These conclusions have significant relevance to our current discussion. 
Maturana’s framework explains the challenge of debunking conspiracy theories and how individuals can inhabit vastly different realities. 
It also underscores our responsibility to acknowledge our inherently constructed perspective on reality and the value system we embrace.</p>

<p>Now, I should mention that there are numerous critics of constructivism, including Markus Gabriel, who advocates for what he refers to as <em>new realism</em>.
In his book <em>Der Sinn des Denkens</em> (English: <em>The Sense of Thinking</em>) <a class="citation" href="#gabriel:2018">(Gabriel, 2018)</a>, he writes:</p>

<blockquote>
  <p>Constructivism is incorrect. New realism asserts that we can perceive reality as it is, without there being precisely one world or reality that encompasses all objects or facts that exist. – Markus Gabriel</p>
</blockquote>

<p>I read his book with the expectation of finding a plausible justification for why this should be the case, but I couldn’t find any consistent argument. 
Admittedly, this might be unfair as the book is a popular science book and doesn’t aim to provide a rigorous theory. 
Nevertheless, Gabriel often labels assertions as obvious without offering a reasonable explanation. 
Based on what I’ve encountered and also observed in my own life, I lean toward believing in constructivism.</p>

<p>Certainly, this form of relativism raises several pressing issues. 
For instance, how can we justify the actions of individuals whose worldviews may vastly differ from our own?
Additionally, how can international bodies like the United Nations apply pressure on nations that violate human rights if their value systems diverge significantly from a Western dominated notion of values?
Richard Rorty has an intreresting take on that question.
If we want universal acceptance of and respect for human rights, we shouldn’t try to argue about it. 
We shouldn’t attempt to work out rational justifications of human rights, or arguments that will convince people that human rights are a good thing. 
Instead, according to Rorty, we would achieve better results if we try to influence people’s feelings instead of their minds—philosophy as poetry <a class="citation" href="#rorty:2016">(Rorty, 2016)</a>!
Rational justifying human rights is an abstract and philosophical way—something that, according to Rorty, isn’t possible anyway.
In a sense, Rorty suggest that, instead of arguing rationally against mythologies, we should imagine, construct and present better ones.
He does not believe in a second enlightenment—in which logic and rationality will triumph over evil.
It is worth noting that similar challenges and questions arise when we consider the concept of <em>free will</em>, which is itself a highly contentious notion, compare for example <a class="citation" href="#sapolsky:2023">(Sapolsky, 2023)</a>.</p>

<p>In a world with no singular perspective, there are multiple viewpoints coexisting and each has its inherent blind spots. 
According to Luhmann, this principle applies to any observing system, whether it’s a psychological system, like the human mind, or a social system.
The diversity of perspectives inherently limits each view, preventing it from fully encompassing all facets of a situation or concept. 
Observation is blind to its own conditions.
When I observe a tree I can not (at the same time) observe myself observing the tree.
This inherent limitation in observation highlights the intricate and multifaceted nature of comprehending and interpreting the world around us.</p>

<p>The diversity of perspectives among individuals often complicates accurate communication because each of us essentially speaks a slightly different language.
Communication, in itself, can be seen as improbable. 
In addition, language is not something we use to describe reality accurately but a technique we employ to get things done.
Nevertheless, communication remains a crucial element as it plays a central role in stabilizing the chaos and connecting psychic and social systems.
Interestingly, it can be effective even when we don’t fully comprehend each other. 
A prime example of this is ChatGPT, which may not understand as humans do but still manages to communicate effectively.</p>

<p>So, can we embrace and navigate this diversity of perspectives?
Can or should we tolerate the uncanny sensation of numerous distinct realities coexisting? 
Is it possible for us to, to some extent, accept that others may inhabit a differently constructed world while simultaneously acknowledging the existence of something that persists, even if we cease to believe in it.
After all, our constructions do not follow our beliefs—we can not dream the problem away.</p>

<h2 id="the-reality-of-the-climate-crisis">The Reality of the Climate Crisis</h2>

<p>The COVID-19 pandemic had a measurable positive effect on pollution levels—which did not last for long.
However, one could argue that it also had a negative impact on trust levels in institutions, especially scientific ones.
This erosion of trust may ultimately hinder efforts to address the climate crisis.
The ongoing debate regarding climate change’s origins and the necessity of curbing CO2 and equivalent gas emissions continues to persist even if the scientific community is clear on the matter—a conviction I have established via <em>second-order observation</em>.</p>

<p>It appears to me that COVID, combined with the rapid and highly polarizing consumption of “news” on social media, has fractured our social discourse. 
The culture of dialogue has suffered, forcing individuals to align with one of two extreme sides. 
It now seems impossible to critique one party without facing accusations of working for the other.
The language we employ has become more <strong>moralizing</strong>. 
Instead of characterizing people as simply incompetent, misled, misguided, or influenced by flawed incentives within a system, they are often labeled as <strong>evil</strong>. 
This focus on the individual impedes progress in reforming social systems, which are in need of change to provide alternative incentives that prioritize social, ecological, and economic measures for all inhabitants of the planet.
Part of the reality of the climate crisis is that we need trusted institutions that need to be aligned in a way that dealing with the crisis becomes possible.</p>

<p>Emissions are not the only problems on our hand.
Many ecological systems are on the bringe of collapse.
Our agriculture is under threat.
Water shortages are on the horizon.
Increased carbon dioxide absorption by oceans leads to ocean acidification, which can harm marine life, especially coral reefs and shellfish.
Climate change can exacerbate health issues by increasing the spread of diseases, heat-related illnesses, and air quality problems due to wildfires and increased pollen levels.
Changing weather patterns and more frequent extreme events can disrupt agriculture and water supplies, potentially leading to food shortages and conflicts over resources.
As climate impacts worsen, there will be never-seen increased migration and displacement of populations, both within and across borders, as people seek refuge from areas affected by climate-related hazards.
Climate change is very likely to intensify extreme weather events such as hurricanes, droughts, heatwaves, and heavy rainfall. 
These events can lead to increased property damage, displacement of populations, and economic losses.
Sea levels are expected to continue rising, posing a threat to coastal communities, infrastructure, and ecosystems. 
Flooding and saltwater intrusion into freshwater sources may become more common.
I imagine that, at some point, borders of certain countries will be closed, dividing the world in a <em>real</em> and a <em>hyperreal</em> one.</p>

<p>Considering all the points I discussed, the resistance to transitioning away from fossil fuels is expected, given the various parties involved. 
It would be a relief if we were in a simulated reality, where everything could be dismissed as a bad dream.
However, we are faced with the pressing need to convince everyone that climate change is a real problem that demands immediate action.
And that it is worth to sacrifice for the unknown other.
Individuals who have limited information may be persuaded through sound arguments and credible sources. 
However, those who actively reject the mainstream narrative may ultimately question the legitimacy of <em>second-order observation</em>.</p>

<p>An example of this dynamic in action was during a BBC News panel where Brian Cox (physics professor and science communicator) clashed with skeptic Malcolm Roberts (politician). 
Roberts insisted on <em>empirical evidence</em> and rejected <em>appeals to authority</em>. 
All seemed well and logical, but when Cox presented a graph as evidence, Roberts dismissed it, alleging that the data had been corrupted by NASA.
At this juncture, a discussion is no longer possible because there are no external empirical evidence available beyond that produced by scientific institutions.</p>

<p>Undeniable empirical evidence, including temperature records, ice melt data, and rising sea levels, serves as a compelling testament to the tangible effects of climate change. 
To promote a more informed perspective, it is advisable to encourage individuals to explore and critically evaluate reputable, peer-reviewed scientific sources, rather than relying on fringe or biased information.
So, let’s delve into a tiny selection of influential contributions from the scientific community that have shaped our understanding of climate change.
Note that this is only a tiny selection from the whole corpus:</p>

<p>As early as 1896, Svante Arrhenius published a groundbreaking paper on the greenhouse effect, demonstrating how rising concentrations of greenhouse gases lead to an increase in global average surface temperatures <a class="citation" href="#arrhenius:1896">(Arrhenius, 1896)</a>.
Another significant milestone occurred in 1967 when Manabe and Wetherald published the first paper that incorporated the fundamental elements of Earth’s climate into a computer model, exploring the implications of doubling carbon dioxide levels for global temperatures <a class="citation" href="#manabe:1967">(Manabe &amp; Wetherald, 1967)</a>. 
Remarkably, the results of their work remain valid today, according to Prof. Forster.
In 1976, Charles D. Keeling and his team documented a pivotal moment by revealing the sharp rise in carbon dioxide levels at the Mauna Loa observatory in Hawaii <a class="citation" href="#keeling:1976">(Keeling et al., 1976)</a>. 
This paper highlighted the observable increase in atmospheric CO2 resulting from the combustion of carbon, petroleum, and natural gas.
Fast-forwarding to 2006, Held and Soden advanced the concept known as <em>wet-get-wetter, dry-get-drier</em> precipitation in the context of global warming <a class="citation" href="#held:2006">(Held &amp; Soden, 2006)</a>. 
This idea, though occasionally misunderstood and misapplied, remains the first and perhaps the only systematic conclusion regarding regional precipitation and global warming based on a robust physical understanding of the atmosphere.
Additionally, the <em>Intergovernmental Panel on Climate Change</em> (IPCC) reports have played an integral role in consolidating and disseminating crucial climate science findings, further enhancing our collective comprehension of climate change.
If one is convinced that there may be some shadiness going on, I encourage the reader to delve into the extensive history of climate change science, for example <em>The Discovery of Global Warming</em> <a class="citation" href="#weart:2009">(Weart, 2008)</a>.</p>

<p>Additional, one can highlight the overwhelming consensus among climate scientists and scientific organizations that climate change is real and largely caused by human activities, compare <a class="citation" href="#myers:2021">(Myers et al., 2021; Lynas et al., 2021; Cook et al., 2016; Cook et al., 2013; Doran &amp; Zimmerman, 2009)</a>.
Of course, if we only rely on those papers we have to trust an even <em>higher-order observation</em>!
In addition, one can argue that there is no plausible alternative theory apart from the effects humans caused by polluting the planet.</p>

<p>But if an individual has lost trust in institutions, all these efforts will be fruitless.
Therefore, it is so deeply important that our scientific institutions as well as the media defend and improve their reputation.
Without the trust in <em>the other</em>, I see great danger on the horizon, especially if things become increasingly difficult.
The challenge lies in persuading individuals who harbor skepticism, particularly toward what climate change deniers label as <em>mainstream science</em>.</p>

<p>According to Maturana, convincing someone can only happen when they develop <strong>their own system of conviction</strong>. 
This aspect is especially crucial to consider. 
Therefore, engagement must be respectful. 
It’s essential to take the worries, fears, and arguments of deniers seriously, even when they appear unreasonable.
However, it is also important to remember that while we encourage others to develop their convictions, we should also acknowledge our own unique and potentially flawed convictions and remain true to them if we are not convinced otherwise.</p>

<blockquote>
  <p>I think I do it always through stories, never through direct confrontation.
Because if you directly confront somebody who’s thinking polar opposite to you, they don’t really listen.
They are thinking of arguments to refute to. […]
The first thing is to listen to them because maybe they’ve got a point, maybe they’re doing something you never thought about.
But if you still feel that you’re right, then you must have the courage of your conviction. – Jane Goodall</p>
</blockquote>

<p>There is another more systemic problem at hand: There is an entire self-producing industry centered around climate denial. 
Our society has fostered an army of lobbyists whose primary aim is to actively sabotage progress in addressing climate issues. 
In contrast to these ‘knowledgable’ deniers, scientists are required to rigorously justify every aspect of their research repeatedly and tend to be cautious about offering concrete advice. 
Conversely, climate deniers merely need to sow seeds of doubt; their strategy revolves around raising questions rather than providing evidence-based answers. 
This asymmetry in approach creates a challenging environment for advancing scientific understanding and consensus on climate change.</p>

<h2 id="rebels-of-authority">Rebels of Authority</h2>

<p>Apart from ignorance, selfish interests and plain stupidity, I have not yet pointed to the root problem.
I cited papers that suggest that there is a link between pattern matching skills, narcissism and believing in conspiracy theory.
Stupid people elect stupid and corrupt politicians, right?
But that is too easy of an explanation.
Whenever I have to fall back on stupidity or evilness, I get the feeling that I am missing something.
Most of the time peope are not evil or stupid.
Similar to conspiracy theories, such a rational is too simple, too individualistic and also too dangerous.
So what is the systemic reason why people reject actions required to keep the planet sustainable for us all?</p>

<p>Of course, there are many complex reasons but I want to focus on one that fits the meat of this article.
I more or less successfully tried to problematize the search for an absolute objective truth by pointing out that those who believe in stepping out of the cave go probably deeper into it.
And I pointed to Baudrillard’s imagined eradication of imagination leading to a <em>true simulacra</em> that no longer stands for anything but itself—a description of <em>hyperreality</em> that hits me emotionally.
Now, let us imagine that conspiracy theorists rightly point to a problem they might not really understand rationally but feel intuitively.
What might it be?</p>

<p>If we want to describe our modern Western society, I think it’s fair to say that we are a <em>knowledge-based society</em>.
Since Foucault’s historical analysis, we know that knowledge and power are linked. 
In our daily life, most of our decisions are informed by some scientifically produced piece of knowledge.
For example, our diet is informed by scientific research that gives us guidance to stay healthy.
How often or for what reason we go to the doctor is informed by science.
In fact, we wittness a whole industry of self-optimization that claims to be scientific.
There are cults trying to establish a ‘science’ of finding a breeding mate.
Or take this article.
My goal and strategy of achieving it is similar: I employ knowledge and second-order observation by citing scientific papers.
In that sense, I fall into the same trap, that is, I try to convince my opponents by displaying superious knowledge.
Since Foucault</p>

<p>In such a society, we tend to search for the <em>perfect algorithm</em> that can make the best decisions for any situation.
In fact, many decisions are already made by algorithms based on the observation of large amounts of data.
Even policies are crafted by utilizing artificial intelligence.
The idea is simple: instead of shouting at each other about the right course of action, let <em>objective reality</em> be the final judge; let ‘the Truth’ decide—let science guide us through the mess.
This is the dream born out of the Enlightenment and it sparked many ideologies that promised to fulfill this utopian harmony.
It is supported by the <em>mechanistic view</em> on the universe suggesting that we can eventually understand all causal relations going back to some final cause we might call God;
In its current interpretation, I call it <strong>algocracy</strong> <a class="citation" href="#danaher:2016">(Danaher, 2016)</a>, i.e., <em>rule by algorithms</em>.</p>

<p>Interstingly, most conspiracy theorist argue within this framework, that is, they argue not in terms of values but in terms of better knowledge.
At the same time they rebel against the authority of such algorithmic rigor.
Therefore, on the one hand, they believe in a complete understanding of reality or at least a big and very complex chunk of it.
Consequently, they also buy the idea that politics is basically determined by the most powerful authority: the truth.
On the other hand, they rebel against this doctrine by using its own logic.
I think here lies a great danger because <strong>the system harms itself by its own operation</strong>.
The system enters a process of selfdestruction.</p>

<p>Like Pippi Longstocking, conspiracy theorists create their world or imagine it based on an alternative system of turth.
In the case of Pippi Longstocking, we find this antiauthoritarian attitude charming, creative, and imaginative, but in the case of conspiracy theorists, many find it dangerous, idiotic, and ignorant—an inconsistency worth investigating.</p>

<p>According to the sociologist Alexander Bogner an overdose of  antiauthoritarian attitude becomes problematic if we no longer agree on any foundation making any conflict impossible since there is nothing to argue about.
One of such foundation is the belief in an objective truth.
He pragmatically argues that:</p>

<blockquote>
  <p>For libaral democracies, the idea of objective truth is a necessary fiction. – <a class="citation" href="#bogner:2021">(Bogner, 2021)</a></p>
</blockquote>

<p>This statement seems reasonable but I will problematize it later.
Bogner points out that a liberal democratic society requires conflicts grounded in dissent.
However, dissent is only possible if there is at least a shared foundation and also an alternative to think about.
Overall, I think Bogner’s essay aligns closely with the constructivists viewpoint shared by e.g. Luhmann even if he criticizes the theory for attacking the idea of an objective truth.
For example, he argues in a Luhmannian manner about the root problem of conspiracy theories, or what many call the <em>post-truth society</em>, pointing out the crossing of boundaries between two interdependent but operationally closed systems:
<strong>science</strong> and <strong>politics</strong>.</p>

<p>Due to the <em>functional differentiation</em> of modern societies, politics communicates about <strong>values</strong> and science communicates about <strong>facts</strong>.
Bogner argues that today this distinction is blurred leading to dysfunctional systems.
Nowadays, politics communicates about facts but is still guided by values.
This makes <em>productive dissent</em> difficult because different value systems no longer compete within the boundary of values but hide behind the battle for better knowledge.
Values are no longer up for debate.
Instead we argue for superior facts.
In this context, <em>fake news</em> take part in a new <em>language game</em> that is born out from the lack of alternatives.
Fake news do not attack values but knowledge itself.</p>

<blockquote>
  <p>This could also be observed during the coronavirus crisis: 
Due to the high pressure of scientification, the will for fundamental opposition in some places was discharged through the spread of ‘alternative facts’.
[…] It was directed against a (supposedly) authoritative instance that claimed to determine what is real, rational, and politically necessary through superior rationality.
From the perspective of this protest, political emancipation could only be an emancipation from the facts.
Alternative facts are evidently in vogue when politics (due to its alignment with science) appears to be without alternatives. – <a class="citation" href="#bogner:2021">(Bogner, 2021)</a></p>
</blockquote>

<p>Bogner correclty points out that scientific knowledge can be dangerous for politics because it dismantles or defuses the discourse.
It is also attractive if one wants to surrender responsibility because in science it is far easier to agree on statements since it follows a strict true-false distinction.
In this sense, science is authoritarian.</p>

<p>Take a simple example such as health.
Before talking about it, we already share certain values.
For example, most people probably assume that a long healthy life is preferable over a short or unhealthy life.
This seems reaonable.
However, even this seemingly simple assumption can not be proven.
It is a value and values are mutually constructed.
I can, for example, easily argue that it is better to live to the fullest even if this means that my life expectancy drops.</p>

<p>In politics we argue about values and try to find some common ground.
This requires some degree of cohesion.
It goes against our more and more individualistc togetherness.
Bogner points out that people are considerably constrained in their expression of dissent if scientific truth is translated into a politic that knows no alternative.
His statement resonates with the relation between the belief in conspiracy theory and narcism.
He writes:</p>

<blockquote>
  <p>The struggle against science and experts can thus be understood as a struggle against determinations that were not chosen by oneself.
For such determinations must appear as an impudent imposition to the individualized, modern person, who is increasingly called upon today as a self-responsible shaper of their destiny, as a self-entrepreneur or ‘Me Inc’. 
Therefore, the struggle against facts is, not least, a struggle for autonomy.
From this perspective, science denial appears as a critique that is in tune with the times, insofar as it derives its plausibility from the current conditions of subjectification.
In the protest against established knowledge, the disappointed hope of the highly individualized, activated subject for complete sovereignty and a fully comprehensible, decision-open world is discharged. – <a class="citation" href="#bogner:2021">(Bogner, 2021)</a></p>
</blockquote>

<p>If politicians give knowledge absolute authority over political decisions, and these decisions realize certain values, the only way to fight for different values is to fight a different truth.
Using Bogner’s perspective, fake news are not the cause of the problem but a symptom of a society that feels impotent because certain values are always already presupposed to be the right values.
Therefore, what Bogner calls <em>productive dissent</em> becomes impossible.</p>

<blockquote>
  <p>The rebellion of the knowledge deniers can overall be understood as a covert appeal directed against a (looming) colonization of politics by expert consensus, 
regardless of how reasonable the expert recommendations may be in individual cases. – <a class="citation" href="#bogner:2021">(Bogner, 2021)</a></p>
</blockquote>

<p>I agree with Bogner that a culture needs a common foundation, a reference world, which, according to Luhmann, should be provided by the media.
And we also need a more flexible and ambiguous reference world on a global scale, see <a class="citation" href="#bauer:2018">(Bauer, 2018)</a>.
Furthermore, I agree that science cannot and should not realize politics and that better knowledge does not automatically lead to better politics.
Science should continue to operate on the true-false distinction, while politics should use a different code that differentiates values.</p>

<p>In case of climate activism, it might be a far more effective strategy to demand the values people want to be realized instead of discussing the facts they believe in.
Instead of calling people deniers, it might be healthier to find and employ strategies that reveal their value system and how this value system supports certain actions.
We should reinvigorate the discourse about values since it is save to say that more and better facts alone cannot move a society towards change.</p>

<p>However, I do not agree with Bogner’s claim that we need to agree that there is an <em>objective truth</em> to be found. 
I even suspect he does not fully believe this himself, as he refers to the <strong>idea of an objective truth as a necessary fantasy</strong>, akin to <em>Plato’s noble lie</em>—a myth knowingly propagated by an elite to maintain social harmony.
I reject this in favor of an embracing uncertainty because, as I argued above, I think it drives the individual deeper into the cave. 
As the philosopher Slavoj Žižek argues—contrary to Dostoevsky and many others—–that the belief in God enables us to commit horrific crimes, believing in an objectively accessible truth can serve the same purpose.
If we no longer live in the same world, society might break, but this does not diminish the great value of doubt and uncertainty in restraining our tendency for megalomania.</p>

<p>On this matter, I lean more towards Rorty, believing that we are mature enough to handle the pragmatist view that it is socially more useful to reject the idea of an objective truth (in the Platonist sense).
According to Rorty, Pippi Longstocking should be allowed to argue for her position, and we should accept her position to be true if it emerges out of the struggle for truth.
The American pragmatist rejects the Platonist’s notion of truth, considering it unintelligible and meaningless.</p>

<blockquote>
  <p>For the idea of a liberal society, it is of central importance that everything is allowed as long as it pertains to words as opposed to actions, to persuasion as opposed to violence. 
This openness should not be maintained because, as the Bible says, the truth is great and will prevail, nor because, as Milton believes, in a free and open fight, the Truth will always win. A society is liberal when it is content to call ‘true’ whatever emerges as the result of such struggles. – <a class="citation" href="#rorty:1989">(Rorty, 1989)</a></p>
</blockquote>

<p>Many, including Bogner, believe that such an attitude leads to <strong>indifference</strong>. 
He argues that if we take Rorty’s comment literally, we would bury both the idea of <em>objective truth</em> and our <em>liberal democracy</em>, since the realization of individual freedom through social change requires productive dissent.
I think Bogner, like many others, misinterprets Rorty or at least applies his statement to the wrong level of society—the wrong reality, so to speak. 
The implication that we end up in a sort of absolute relativism or complete indifference if we stop believing in an objective truth remains unfounded. 
It’s a huge leap.
Even if we reject the idea of an objective truth, empirical evidence is still valuable for discussions.</p>

<p>Rejecting the belief in an objective truth does not mean that we lose the ground for productive dissent.
On the contrary, it allows for a more pragmatic approach to dialogue and disagreement.
When we drop this belief, we recognize that our conversations and disagreements are grounded in our contingent vocabularies and cultural practices.
This recognition does not lead to indifference but rather to a greater appreciation of the diversity of perspectives.
It encourages us to engage with others’ values and beliefs not because they correspond to an objective reality, but because they are part of the <strong>shared human experience</strong>.
Productive dissent arises from the acknowledgment that our beliefs and values are fallible and open to revision through conversation.
It promotes a kind of solidarity where we strive to understand and negotiate our differences rather than impose a supposed objective truth.
This process fosters mutual respect and a more inclusive society, where different viewpoints can coexist and contribute to a richer, more nuanced understanding of the world.
By focusing on the practical consequences of our beliefs and actions, we can still engage in meaningful debates about what kind of society we want to create.
This pragmatic approach encourages us to find common ground and work together to address shared problems, rather than being paralyzed by the quest for an elusive objective truth.</p>

<h2 id="conclusion">Conclusion</h2>

<p>What I wanted to emphasize is the idea that we are existing within a subjective and distorted world of a reality that we can not directly access.
We are always already in a simulation, relying on second-order observation.
What I call <em>angry media</em> outlets, many of which take part in spreading misinformation and conspiracy theories, use more and more frequently phrases like <em>literally</em>, <em>this is simply the truth</em>, <em>they all know it</em>, <em>its a fact</em>.
They fight against ambiguity dragging anything into a culture war such that it has to be defined using one of two perspectives.
Thus the root cause of descending into the rabbit hole may not be a detachment from reality but rather an <strong>attachment to certainty</strong>. 
Perhaps Baudrillard is correct in asserting that we are entering a <em>hyperreality</em> that is no longer contradictory and dissolves all illusions, imaginations, and mysteries.
According to him, what we are trying to do is to purge the world of all mysteries, illusions and imaginations.
But certainty exists only in a pure form of simulation. 
In a constructivist sense, we cannot access the absolute true reality (Kant’s <em>the thing in itself</em>) because it is always already mediated.
If we cannot accept this fundamental ambiguity of our reality (by which Baudrillard does not mean physical reality but that which is intelligible via signs), we run the risk of constructing the one and only reality, i.e., hegemony.
Since this realm does not allow contradictions, it tends to integrate everything, including disasters, into it.</p>

<p>Contradictions form the very foundation of our environment. 
We perceive the reality of nature when it manifests as a non-human force of destruction, such as a natural disaster or a pandemic. 
What occurs is incomprehensible. 
In hyperreality, nature ceases to be a conflicting force and becomes merely an element within the simulation—a floating sign. 
Destruction transforms into a calculated event.
Repeatedly displaying graphs of death tolls gives us the illusion of control.
It can be seen as an attempt to reintegrate death (arguably the greatest contradiction of all) into the simulation.
The negative, along with contradictions, is either integrated or discriminated against.
Wars, natural disasters, and pandemics metamorphose into a spectacle on the television screen, a tourist attraction for our theme park, and are more disastrous than the disaster, more natural than nature—in short, a perfect simulation that surpasses and supplants reality.</p>

<p>In a postmodern society, the absence of any central value system and firm, objective evaluative guides tends to create a demand for substitutes. 
These substitutes are symbolically created rather than being actual or socially produced. 
The need for these symbolic group tokens results in tribal politics and defines self-constructing practices that are collectivized but not socially produced. 
These neo-tribes function solely as imagined communities and, unlike their premodern namesake, exist only in symbolic form through the commitment of individual ‘members’ to the idea of an identity. 
They exist as imagined communities through a multitude of agent acts of self-identification and endure solely because people use them as vehicles of self-definition; as an identity technology termed <em>profilicaty</em>.</p>

<p>I have no intention of passing moral judgment on conspiracy theorists.
They often risk a significant amount of social capital, leading to alienation from their relatives and friends.
Being a conspiracy theorist is generally not an enjoyable experience.
As I said, they rebel against authority via the same use of authority—a conflict of knowledge against ‘better’ knowledge in the realm of politics.
This is dangerous.
It points to a deeper problem of a violation of the <em>operational clousure</em> of social systems which we should take serious if our goal is to preserve these systems.
One might even argue that conspiracy theorists perceive cracks in the simulation but mistakenly believe in a way out of it, which, and this misinterprets <em>The Matrix</em> as well, serves as the perfect cover-up for our reality as always partly simulated.</p>

<p>Baudrillard famously argued that Disneyland does not hide the fact that it is a simulation, but rather conceals the fact that America is a simulation too. 
Disneyland is more real than America.
Conspiracy theories operate similarly; they are pure simulations and, in this regard, true. 
The outsider is convinced of his or her reality, and the contradictory nature of these theories is not contradictory for him or her. 
Instead, (obvious) contradictions are necessary to make the theory hyperreal. 
The suspicion of the theorists is not unreasonable, but their conclusion is fatal—they demand a simple metanarrative and cannot see that this can only be another far more harmful simulation.</p>

<p>Several factors contribute to the prevalence of conspiracy theories, including information overload, the perception of a reality that is becoming ‘less real’, the sensation of living in a quasi-simulated reality, natural disasters, ongoing conflicts, and the rapid pace of our society. 
Our modern world is so complex that it is virtually impossible for any single individual to comprehensively make sense of all that occurs.
As a result, we heavily rely on the concept of second-order observation, and there is no shame in acknowledging this fact.
Turning to experts and authorities, provided that their authority is derived from genuine competence, is necessary. 
However, it’s crucial to subject these authorities to scrutiny and verification and to be aware that their observation has always a blind spot.</p>

<p>Nothing in what I’ve stated here should be misconstrued as a defense of the political system. 
Lobbyism, which is sometimes indistinguishable from outright corruption, represents a significant issue. 
The consistent failure to fulfill promises, whether they be pledges for a transaction tax, the cessation of subsidies that actively contribute to global warming, or the numerous ‘conferences’ like COP that, at this point, are merely part of a <em>hope economy</em>—offering a false and pacifying sense of hope—that fuels the distrust in institutions.</p>

<p>In my view, COP28 proved to be a disaster for the majority of the world’s population. 
There was essentially no consensus on even the most basic measures. 
No accord on phasing out fossil fuels, and not even a genuine commitment to promoting renewable energy. 
The decision to have COP led by one of the world’s largest oil and gas companies is, at this point, satirical if it were not true.
In an assessment by journalist Jonathan Watts in The Guardian, the winners of the conference were identified as the oil and gas industry, the United States, China, COP28 President Sultan Al Jaber, the green energy sector, and lobbyists. Conversely, the losers encompassed the climate, small island nations, climate justice, future generations, other species, and scientists.</p>

<p>Nothing in what I’ve said here should be interpreted as a defense of the media for its issues, nor should it downplay the problem of increasing wealth inequality or any other ecological, social, or economic problems.
Both mass media and the scientific system have their problems, but this doesn’t negate their incredible value.
Moreover, independent media outlets play a crucial role, but we must be under no illusion that the verification of information is becoming increasingly challenging. 
Social media platforms present us with an overwhelming array of viewpoints, claims, and video content. With the ascent of generative artificial intelligence, the task of verifying this deluge of information becomes even more daunting.
The power of the image is indeed huge. 
The blurred line between hyperreality and lower forms of simulations makes it difficult to navigate through the mass of information. In contrast, conspiracy theorists do not bear the burden of a demanding verification process. 
They can simply draw upon fringe and unvalidated stories, presenting them in an entertaining, sensational style akin to news pornography. 
They can use the power of high-order simulacra which are disconnected from the real.</p>

<p>Believing in a conspiracy theory is akin to being the prisoner in Plato’s cave, presuming that everyone else is, in fact, in prison. 
It is the belief in something outside of simulation.
The most effective remedy is to harbor doubts about our own competence, to be skeptical of ourselves, to maintain self-awareness at a metacognitive level, and to be able to live in a contradictory world and recognize those contradictions.
These contradictions live on the borders of hyperreality—in slums, cobalt mines, the streets of New York City, the border of Mexico, the fortress of the European sea, refugee camps, and the ‘ugly’ parts of the world.</p>

<p>I am not sure if I can agree with the cynical viewpoint of Baudrillard. 
His overly dramatic and playful writings are interesting but also contradictory, probably by design.
He would probably be horrified at our attempt to <a href="https://blogs.nvidia.com/blog/earth-2-supercomputer/">simulate the whole earth</a> to predict and control our future. 
But how else can we deal with an open, unpredictable future other than the pursuit of more and more accurate predictions? Esposito asks if this future will still be open <a class="citation" href="#esposito:2024">(Esposito, 2024; Esposito et al., 2023)</a>.</p>

<p>It appears to me that logical reasoning and providing ‘better’ facts alone is not sufficiently compelling.
Science and technology is not enough.
It feels like we lack spiritual growth.
Maybe Rorty is right about the importance of empathic stories.
Maybe science should provide us with that what we call truth and politics should offer us a story that moves us.
But where are these stories?
Are we already as cynical as Baudrillard?
We need storytellers and artists to craft more persuasive mythologies, narratives, and stories that resonate on an emotional level.
I believe this can only be possible if we do not filter out the negative and the ugly part of society.
We have to be serious yet playful and imaginative.
I want a serious vision which is shamelessly emphatic towards all forms of life.
This anti-Platonistic approach, while potentially controversial, could prove more effective in influencing beliefs and behaviors and, in the end, matter more than any rational argument could be.</p>

<p>We have reached an alarming point at which millions of people can no longer discriminate between reality and hyperreality. 
That which cannot be simulated seems to disappear. 
There is a confusion between truth claims grounded in evidence and sound logic and alternative facts inspired by an authoritarian rebellion against the authority of knowledge.
Even in the case of the climate crisis, if we want to preserve a liberal democracy, politics has to discuss alternatives based on the consideration of different values.
We need imagination to draw larger circles and rationalty to make valid moves within these circles.
We have to trust in the scientific method to provide us with good knowledege, politics (not politicians) that is informed by science to provide us with good politics, and spirituality that may spend us wisdom and psychologic stability.</p>

<p>Most importantly, instead of appealing to objective truth, it is more productive to focus on the practical consequences and shared values that can mobilize action.
Instead of arguing that climate change is an objective truth that everyone must accept, we should emphasize the practical and tangible consequences of climate inaction.
By highlighting concrete consequences, we may make a compelling case for action based on the observable and lived experiences of people.
We should frame the argument for climate action in terms of shared values and common interests.
The goal is to find common ground and motivate people to act based on their own interests and values, rather than trying to convince them of an abstract objective truth.
Acknowledging the contingency and fallibility of our beliefs does not mean we cannot act decisively.
We can adopt policies and actions that are based on the best available evidence while remaining open to revising them as new information emerges. 
While rejecting objective truth might complicate the traditional ways of arguing for action, it also opens up new avenues for persuasion and coalition-building.</p>

<p>Sooner or later the reality of the crisis will eventually bleed into hyperreality.
Even if we construct our own perspective on the world, the reality of the climate crisis will not disappear if we stop believing in it.</p>

<h2 id="references">References</h2>

<ol class="bibliography"><li><span id="popper:1934">Popper, K. (1934). <i>Logik der Forschung</i>.</span></li>
<li><span id="barthes:1973">Barthes, R. (1973). <i>Mythologies</i>. Hill &amp; Wang Pub.</span></li>
<li><span id="baudrillard:1983">Baudrillard, J. (1983). <i>Simulacra and Simulation</i>. Semiotext(e).</span></li>
<li><span id="luhmann:2000">Luhmann, N. (2000). Why does society describe itself as postmodern. In W. Rasch &amp; C. Wolfe (Eds.), <i>Observing complexity: Systems theory and postmodernity</i> (pp. 35–49). University of Minnesota.</span></li>
<li><span id="baudrillard:1968">Baudrillard, J. (1968). <i>System of Objects</i>.</span></li>
<li><span id="baudrillard:2008">Baudrillard, J. (2008). <i>The Perfect Crime</i>. Verso.</span></li>
<li><span id="douglas:2017">Douglas, K. M., Sutton, R. M., &amp; Cichocka, A. (2017). The psychology of conspiracy theories. <i>Current Directions in Psychological Science</i>, <i>26</i>(6), 538–542. https://doi.org/10.1177/0963721417718261</span></li>
<li><span id="uscinski:2018">Uscinski, J. E. (2018). The study of conspiracy theories. <i>Argumenta</i>, 233–245. https://doi.org/10.23811/53.arg2017.usc</span></li>
<li><span id="prooijen:2018">Prooijen, J.-W. (2018). <i>The Psychology of Conspiracy Theories</i>. Taylor &amp; Francis Group. https://doi.org/10.4324/9781315525419</span></li>
<li><span id="kay:2011">Kay, J. (2011). <i>Among the Truthers: A Journey Through America’s Growing Conspiracist Underground</i>. Harper.</span></li>
<li><span id="shermer:2022">Shermer, M. (2022). <i>Conspiracy: Why the Rational Believe the Irrational</i>. Johns Hopkins University Press.</span></li>
<li><span id="landau:2015">Landau, M. J., Kay, A. C., &amp; Whitson, J. A. (2015). Compensatory control and the appeal of a structured world. <i>Psychol Bull</i>. https://doi.org/10.1037/a0038703</span></li>
<li><span id="cichocka:2022">Cichocka, A., Marchlewska, M., &amp; Biddlestone, M. (2022). Why do narcissists find conspiracy theories so appealing? <i>Curr Opin Psychol</i>. https://doi.org/10.1016/j.copsyc.2022.101386</span></li>
<li><span id="cosgrove:2023">Cosgrove, T. J., &amp; Murphy, C. P. (2023). Narcissistic susceptibility to conspiracy beliefs exaggerated by education, reduced by cognitive reflection. <i>Front Psychol</i>. https://doi.org/10.3389/fpsyg.2023.1164725</span></li>
<li><span id="hart:2015">Hart, J., &amp; Graether, M. (2015). Something’s going on here: Psychological predictors of belief in conspiracy theories. <i>Journal of Individual Differences</i>.</span></li>
<li><span id="herman:1988">Herman, E. S., &amp; Chomsky, N. (1988). <i>Manufacturing Consent</i>. Pantheon Books.</span></li>
<li><span id="maturana:1987">Maturana, H. R., &amp; Varela, F. J. (1987). <i>The Tree of Knowledge</i>. Shambhala.</span></li>
<li><span id="gabriel:2018">Gabriel, M. (2018). <i>Der Sinn des Denkens</i>. Ullstein Buchverlag.</span></li>
<li><span id="rorty:2016">Rorty, R. (2016). <i>Philosophy as Poetry</i>. University of Virginia Press.</span></li>
<li><span id="sapolsky:2023">Sapolsky, R. M. (2023). <i>Determined</i>. Bodley Head.</span></li>
<li><span id="arrhenius:1896">Arrhenius, S. (1896). XXXI. On the influence of carbonic acid in the air upon the temperature of the ground. <i>The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science</i>, <i>41</i>(251), 237–276. https://doi.org/10.1080/14786449608620846</span></li>
<li><span id="manabe:1967">Manabe, S., &amp; Wetherald, R. T. (1967). Thermal equilibrium of the atmosphere with a given distribution of relative humidity. <i>Journal of Atmospheric Sciences</i>, <i>24</i>(3), 241–259. https://doi.org/10.1175/1520-0469(1967)024&lt;0241:TEOTAW&gt;2.0.CO;2</span></li>
<li><span id="keeling:1976">Keeling, C. D., Bacastow, R. B., Bainbridge, A. E., Ekdahl Jr., C. A., Guenther, P. R., Waterman, L. S., &amp; Chin, J. F. S. (1976). Atmospheric carbon dioxide variations at Mauna Loa Observatory, Hawaii. <i>Tellus</i>, <i>28</i>(6), 538–551. https://doi.org/10.1111/j.2153-3490.1976.tb00701.x</span></li>
<li><span id="held:2006">Held, I. M., &amp; Soden, B. J. (2006). Robust responses of the hydrological cycle to global warming. <i>Journal of Climate</i>, <i>19</i>(21), 5686–5699. https://doi.org/10.1175/JCLI3990.1</span></li>
<li><span id="weart:2009">Weart, S. R. (2008). <i>The Discovery of Global Warming</i>. Harvard University Press.</span></li>
<li><span id="myers:2021">Myers, K. F., Doran, P. T., Cook, J., Kotcher, J. E., &amp; Myers, T. A. (2021). Consensus revisited: quantifying scientific agreement on climate change and climate expertise among Earth scientists 10 years later. <i>Environmental Research Letters</i>, <i>16</i>(10), 104030. https://doi.org/10.1088/1748-9326/ac2774</span></li>
<li><span id="lynas:2021">Lynas, M., Houlton, B. Z., &amp; Perry, S. (2021). Greater than 99% consensus on human caused climate change in the peer-reviewed scientific literature. <i>Environmental Research Letters</i>, <i>16</i>(11), 114005. https://doi.org/10.1088/1748-9326/ac2966</span></li>
<li><span id="cook:2016">Cook, J., Oreskes, N., Doran, P. T., Anderegg, W. R. L., Verheggen, B., Maibach, E. W., Carlton, J. S., Lewandowsky, S., Skuce, A. G., Green, S. A., Nuccitelli, D., Jacobs, P., Richardson, M., Winkler, B., Painting, R., &amp; Rice, K. (2016). Consensus on consensus: a synthesis of consensus estimates on human-caused global warming. <i>Environmental Research Letters</i>, <i>11</i>(4), 048002. https://doi.org/10.1088/1748-9326/11/4/048002</span></li>
<li><span id="cook:2013">Cook, J., Nuccitelli, D., Green, S. A., Richardson, M., Winkler, B., Painting, R., Way, R., Jacobs, P., &amp; Skuce, A. (2013). Quantifying the consensus on anthropogenic global warming in the scientific literature. <i>Environmental Research Letters</i>, <i>8</i>(2), 024024. https://doi.org/10.1088/1748-9326/8/2/024024</span></li>
<li><span id="doran:2009">Doran, P. T., &amp; Zimmerman, M. K. (2009). Examining the scientific consensus on climate change. <i>Eos, Transactions American Geophysical Union</i>, <i>90</i>(3), 22–23. https://doi.org/https://doi.org/10.1029/2009EO030002</span></li>
<li><span id="danaher:2016">Danaher, J. (2016). The Threat of Algocracy: Reality, Resistance and Accommodation. <i>Philosophy &amp; Technology</i>, <i>29</i>, 245–268. https://doi.org/10.1007/s13347-015-0211-1</span></li>
<li><span id="bogner:2021">Bogner, A. (2021). <i>Die Epistemisierung des Politischen</i>. Reclam.</span></li>
<li><span id="bauer:2018">Bauer, T. (2018). <i>Die Vereindeutigung der Welt</i> (p. 104). Reclam.</span></li>
<li><span id="rorty:1989">Rorty, R. (1989). <i>Contingency, Irony, and Solidarity</i>. Cambridge University Press.</span></li>
<li><span id="esposito:2024">Esposito, E. (2024). Can we use the open future? Preparedness and innovation in times of self-generated uncertainty. <i>European Journal of Social Theory</i>, <i>0</i>(0). https://doi.org/10.1177/13684310231224546</span></li>
<li><span id="esposito:2023">Esposito, E., Hofmann, D., &amp; Coloni, C. (2023). Can a predicted future still be an open future? Algorithmic forcasts and actionability in the precision medicine. <i>History and Theory</i>. https://doi.org/10.1111/hith.12327</span></li></ol>]]></content><author><name>Benedikt Zönnchen</name><email>benedikt.zoennchen@web.de</email></author><category term="Social Systems Theory" /><category term="Conspiracy Theories" /><category term="Climate Crisis" /><summary type="html"><![CDATA[Diverging from my area of expertise is always a risky endeavor, but since this is a blog and not a scientific journal, I’m giving myself the liberty to explore and have fun with different ideas (even if the topic is depressing). Often writing helps in transforming the mess into a structured and coherent concept. The process of rethinking and reflecting can be invaluable. It helps to make ones thought anschlussfähig which literally means to be capable for connections and in this context means enabling the continuation of communication.]]></summary></entry><entry><title type="html">Musical Interrogation III - LSTM</title><link href="https://bzoennchen.github.io/Pages/2023/11/19/musical-interrogation-III.html" rel="alternate" type="text/html" title="Musical Interrogation III - LSTM" /><published>2023-11-19T00:00:00+01:00</published><updated>2023-11-19T00:00:00+01:00</updated><id>https://bzoennchen.github.io/Pages/2023/11/19/musical-interrogation-III</id><content type="html" xml:base="https://bzoennchen.github.io/Pages/2023/11/19/musical-interrogation-III.html"><![CDATA[<p>This article is the continuation of a series.
It is recommended that you read part I and II first.
This time we use a recurrent neural network (<strong>RNN</strong>), more precisely an <strong>LSTM</strong>, which I explained a little bit in the <a href="/Pages/2023/04/02/musical-interrogation-I.html">introduction</a>.
An LSTM is a RNN that counteracts the problem of exploding and vanishing gradients.</p>

<p>This article gives some explanation to the code in the following <a href="https://github.com/BZoennchen/musical-interrogation/blob/main/partIII/melody_rnn.ipynb">notebook</a>, which can be executed on <a href="https://colab.research.google.com/?hl=de">Google Colab</a>.
Because our model is now able to learn long-time relations, we can use the <code class="language-plaintext highlighter-rouge">GridEncoder</code>, i.e., a <em>piano roll data representation</em>, which is exactly what we do.</p>

<h2 id="recurrent-neural-networks">Recurrent Neural Networks</h2>

<p>Prior to the advent of transformers, many cutting-edge natural language processing (NLP) applications relied on recurrent neural networks (RNNs). 
These networks are particularly adept at processing sequences of data, such as words in NLP tasks or notes and musical events in audio processing.</p>

<p>An RNN processes a sequence</p>

\[\mathbf{x}_0, \ldots, \mathbf{x}_{n-1}\]

<p>of inputs one at a time. 
With each new input, the network not only considers this current input but also incorporates a <em>hidden state</em>—a representation of previous inputs—thanks to its recurrent connections.
This hidden state \(\mathbf{h}_t\) is updated at each step, ensuring that the network retains a memory of what it has processed so far.</p>

<p><br /></p>
<div><img style="display:block; margin-left:auto; margin-right:auto; width:80%;" src="/Pages/assets/images/rnn-unfold.png" alt="Sketch of an RNN unfolded in time" />
<div style="display: table;margin: 0 auto;">Figure 1: Sketch of an RNN unfolded in time.</div>
</div>
<p><br /></p>

<p>The unique feature of RNNs is their ability to maintain an internal state that captures information about the sequence they have processed to that point.
As a result, the final output of the RNN is informed by the entire input sequence.</p>

<p>Furthermore, RNNs have been employed in sequence generation or decoding tasks. 
In such applications, the tokens generated by the RNN are fed back into it as inputs. 
This feedback loop allows the RNN to generate sequences where each new token is influenced by the previously generated tokens, making it suitable for tasks like text generation, music composition, and more.</p>

<h2 id="long-short-term-memory-networks">Long Short-Term Memory Networks</h2>

<p>Long short-term memory networks (LSTMs) are a type of RNN that were designed to overcome some of the limitations of vanilla RNNs, particularly in handling long-term dependencies in sequence data.</p>

<p>Vanilla RNNs struggle with learning long-term dependencies due to the <em>vanishing gradient problem</em>. 
As the length of the input sequence increases, the gradients used in the training process can become extremely small, making it difficult for the RNN to learn and retain information from earlier inputs. 
LSTMs address this issue with their unique architecture, which includes <em>memory cells</em> which uses different <em>gates</em>.</p>

<p>A <em>memory cell</em> can maintain information in memory for long periods of time. 
The key components of an such a cell are its gates: the <em>input gate</em>, <em>output gate</em>, and <em>forget gate</em>.
These gates regulate the flow of information into and out of the cell, and they decide what to retain or discard from the cell state.</p>

<ul>
  <li><strong>Update Gate</strong>: Determines how much of the new information to add to the cell state.</li>
  <li><strong>Forget Gate</strong>: Decides what information is no longer needed and removes it from the cell state, helping to prevent the accumulation of irrelevant information.</li>
  <li><strong>Output Gate</strong>: Controls the extent to which the value in the cell is used to compute the output activation of the block.</li>
</ul>

<div><img style="display:block; margin-left:auto; margin-right:auto; width:80%;" src="/Pages/assets/images/lstm-cell.png" alt="LSTM cell" />
<div style="display: table;margin: 0 auto;">Figure 2: A sketch of a memory cell.</div>
</div>
<p><br /></p>

<p>Due to their architecture, LSTMs can learn and remember over longer sequences than vanilla RNNs, making them more effective for tasks like language modeling, text generation, speech recognition, and more, where understanding context over a long sequence is crucial.
The gating mechanism helps mitigate the vanishing gradient problem, allowing for more effective training over longer sequences.
This is because the gates allow gradients to flow through the network without being multiplied repeatedly by small numbers (which is what causes the gradients to vanish in vanilla RNNs).</p>

<h2 id="data-preparation">Data Preparation</h2>

<p>Again we use the data from <a href="http://kern.ccarh.org">EsAC</a>. 
The specific dataset I utilized is <a href="https://kern.humdrum.org/cgi-bin/ksdata?l=/essen/europa&amp;format=recursive">Folksongs from the continent of Europe</a> and for the purpose of this work, I will exclusively use the 1700 pieces found in the <code class="language-plaintext highlighter-rouge">./deutschl/erk</code> directory.</p>

<p>We assume that 1/16 is the shortest note in our dataset.
The <code class="language-plaintext highlighter-rouge">GridEncoder</code> automatically filters out pieces that do not fulfill this condition.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">time_step</span> <span class="o">=</span> <span class="mi">1</span><span class="o">/</span><span class="mi">16</span>
<span class="n">encoder</span> <span class="o">=</span> <span class="n">GridEncoder</span><span class="p">(</span><span class="n">time_step</span><span class="p">)</span>
<span class="n">enc_songs</span><span class="p">,</span> <span class="n">invalid_song_indices</span> <span class="o">=</span> <span class="n">encoder</span><span class="p">.</span><span class="n">encode_songs</span><span class="p">(</span><span class="n">scores</span><span class="p">)</span>
<span class="k">print</span><span class="p">(</span><span class="sa">f</span><span class="s">'there are </span><span class="si">{</span><span class="nb">len</span><span class="p">(</span><span class="n">enc_songs</span><span class="p">)</span><span class="si">}</span><span class="s"> valid songs and </span><span class="si">{</span><span class="nb">len</span><span class="p">(</span><span class="n">invalid_song_indices</span><span class="p">)</span><span class="si">}</span><span class="s"> songs'</span><span class="p">)</span>
</code></pre></div></div>

<p>Let us look at an example encoded of a piece:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>55 _ _ _ 60 _ _ _ 60 _ _ _ 60 _ _ _ 60 _ _ _ 64 _ _ _ 64 _ _ _ r _ _ _ 62 _ 64 _ 65 ...
</code></pre></div></div>

<p>As we discussed in the last article, <code class="language-plaintext highlighter-rouge">55 _ _ _</code> stands for the midinote <code class="language-plaintext highlighter-rouge">55</code> played for 4 beats where one beat is 1/16 note.
Therefore, this is a 1/4 note.
Likewise, <code class="language-plaintext highlighter-rouge">r _ _ _</code> is a 1/4 rest.</p>

<p>Next, the <code class="language-plaintext highlighter-rouge">StringToIntEncoder</code> converts our alphabet of tokens into positive integers.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">string_to_int</span> <span class="o">=</span> <span class="n">StringToIntEncoder</span><span class="p">(</span><span class="n">enc_songs</span><span class="p">)</span>
</code></pre></div></div>

<p>Next, we use <code class="language-plaintext highlighter-rouge">ScoreDataset</code> to arrange our training data.
It requires our encoded songs, the instance of <code class="language-plaintext highlighter-rouge">StringToIntEncoder</code> and a <em>hyperparameter</em> <code class="language-plaintext highlighter-rouge">sequence_len</code> that configures the length of token sequences our model will be trained on.
The longer the sequence, the longer the training will require because the deeper the recurrent neural network will be.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">sequence_len</span> <span class="o">=</span> <span class="mi">64</span> <span class="c1"># this is a hyperparameter!
</span><span class="n">dataset</span> <span class="o">=</span> <span class="n">ScoreDataset</span><span class="p">(</span>
    <span class="n">enc_songs</span><span class="o">=</span><span class="n">enc_songs</span><span class="p">,</span> 
    <span class="n">stoi_encoder</span><span class="o">=</span><span class="n">string_to_int</span><span class="p">,</span> 
    <span class="n">sequence_len</span><span class="o">=</span><span class="n">sequence_len</span><span class="p">)</span>
</code></pre></div></div>

<p>It is now possible to split our data into training, validation and test set.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">train_set</span><span class="p">,</span> <span class="n">val_set</span><span class="p">,</span> <span class="n">test_set</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">utils</span><span class="p">.</span><span class="n">data</span><span class="p">.</span><span class="n">random_split</span><span class="p">(</span><span class="n">dataset</span><span class="p">,</span> <span class="p">[</span><span class="mf">0.8</span><span class="p">,</span> <span class="mf">0.1</span><span class="p">,</span> <span class="mf">0.1</span><span class="p">])</span>
</code></pre></div></div>

<h2 id="model-definition">Model Definition</h2>

<p>First we define the rest of our <em>hyperparameters</em>:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">vocab_size</span> <span class="o">=</span> <span class="nb">len</span><span class="p">(</span><span class="n">string_to_int</span><span class="p">)</span> <span class="c1"># size of our alphabet
</span><span class="n">input_dim</span> <span class="o">=</span> <span class="n">vocab_size</span> <span class="c1"># can be different
</span><span class="n">hidden_dim</span> <span class="o">=</span> <span class="mi">128</span> <span class="c1"># can be different
</span><span class="n">layer_dim</span> <span class="o">=</span> <span class="mi">1</span> <span class="c1"># can be different
</span><span class="n">output_dim</span> <span class="o">=</span> <span class="n">vocab_size</span> <span class="c1"># should not be different
</span><span class="n">dropout</span> <span class="o">=</span> <span class="mf">0.2</span> <span class="c1"># can be different
</span>
<span class="n">criterion</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">nn</span><span class="p">.</span><span class="n">CrossEntropyLoss</span><span class="p">()</span>

<span class="n">learning_rate</span> <span class="o">=</span> <span class="mf">0.001</span> <span class="c1"># can be different
</span><span class="n">batch_size</span> <span class="o">=</span> <span class="mi">64</span> <span class="c1"># can be different
</span><span class="n">n_epochs</span> <span class="o">=</span> <span class="mi">10</span> <span class="c1"># can be different
</span><span class="n">eval_interval</span> <span class="o">=</span> <span class="mi">100</span> <span class="c1"># can be different
</span></code></pre></div></div>

<p>Before explaining every detail, let us look at the model definition first.
The following is the model description of our RNN/LSTM.
To understand what’s going on, look at the forward method.
This sends our data through the network.</p>

<p>The first two lines create the short-term \(\mathbf{h}_0\) and long-term memory \(\mathbf{c}_0\) and fill them with zeros.</p>

<p>Then an embedding takes place: <code class="language-plaintext highlighter-rouge">x = self.embedding(x)</code>.
This is nothing more than what we did with our simple feedforward net in <a href="/Pages/2023/05/31/musical-interrogation-II.html">Part II - FNN</a>: Each element of the input <code class="language-plaintext highlighter-rouge">x</code> is first one-hot encoded and then multiplied by a matrix. 
The result: Each event is represented by the row of a matrix (with learnable parameters).
The matrix has <code class="language-plaintext highlighter-rouge">vocab_size</code> rows and <code class="language-plaintext highlighter-rouge">input_dim</code> columns.</p>

<p>Next, we send our transformed input through our LSTM out, <code class="language-plaintext highlighter-rouge">(ht, ct) = self.lstm(x, (h0, c0))</code>.
This basically computes \(\mathbf{h}_t, \mathbf{c}_t\) based on \(\mathbf{h}_{t-1}, \mathbf{c}_{t-1}\) as indicated in Fig. 2.
We get as many outputs as our sequence is long, i.e., <code class="language-plaintext highlighter-rouge">sequence_len</code> many.
But we are only interested in the last output, which we get by <code class="language-plaintext highlighter-rouge">out[:, -1, :]</code>.
This is a vector with <code class="language-plaintext highlighter-rouge">hidden_dim elements</code>. 
We don’t need <code class="language-plaintext highlighter-rouge">ht</code> and <code class="language-plaintext highlighter-rouge">ct</code>.</p>

<p>Then we send the last output through a dropout layer to counteract <em>overfitting</em>.</p>

<p>In the last step, we transform the <code class="language-plaintext highlighter-rouge">hidden_dim</code>-dimensional vector into an <code class="language-plaintext highlighter-rouge">output_dim</code>-dimensional vector, which is equal to <code class="language-plaintext highlighter-rouge">vocab_size</code>.
This vector is interpreted as a probability distribution.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">class</span> <span class="nc">LSTMModel</span><span class="p">(</span><span class="n">torch</span><span class="p">.</span><span class="n">nn</span><span class="p">.</span><span class="n">Module</span><span class="p">):</span>
    <span class="k">def</span> <span class="nf">__init__</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">input_dim</span><span class="p">,</span> <span class="n">hidden_dim</span><span class="p">,</span> <span class="n">layer_dim</span><span class="p">,</span> <span class="n">output_dim</span><span class="p">,</span> <span class="n">dropout</span><span class="o">=</span><span class="mf">0.2</span><span class="p">):</span>
        <span class="nb">super</span><span class="p">(</span><span class="n">LSTMModel</span><span class="p">,</span> <span class="bp">self</span><span class="p">).</span><span class="n">__init__</span><span class="p">()</span>

        <span class="bp">self</span><span class="p">.</span><span class="n">hidden_dim</span> <span class="o">=</span> <span class="n">hidden_dim</span>
        <span class="bp">self</span><span class="p">.</span><span class="n">layer_dim</span> <span class="o">=</span> <span class="n">layer_dim</span>
        
        <span class="bp">self</span><span class="p">.</span><span class="n">embedding</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">nn</span><span class="p">.</span><span class="n">Embedding</span><span class="p">(</span><span class="n">vocab_size</span><span class="p">,</span> <span class="n">input_dim</span><span class="p">)</span>
        <span class="bp">self</span><span class="p">.</span><span class="n">lstm</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">nn</span><span class="p">.</span><span class="n">LSTM</span><span class="p">(</span><span class="n">input_dim</span><span class="p">,</span> <span class="n">hidden_dim</span><span class="p">,</span> <span class="n">layer_dim</span><span class="p">,</span> <span class="n">batch_first</span><span class="o">=</span><span class="bp">True</span><span class="p">)</span>
        <span class="bp">self</span><span class="p">.</span><span class="n">dropout</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">nn</span><span class="p">.</span><span class="n">Dropout</span><span class="p">(</span><span class="n">dropout</span><span class="p">)</span>
        <span class="bp">self</span><span class="p">.</span><span class="n">fc</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">nn</span><span class="p">.</span><span class="n">Linear</span><span class="p">(</span><span class="n">hidden_dim</span><span class="p">,</span> <span class="n">output_dim</span><span class="p">)</span>
        
    <span class="k">def</span> <span class="nf">forward</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">x</span><span class="p">):</span>
        <span class="n">h0</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">zeros</span><span class="p">(</span><span class="bp">self</span><span class="p">.</span><span class="n">layer_dim</span><span class="p">,</span> <span class="n">x</span><span class="p">.</span><span class="n">size</span><span class="p">(</span><span class="mi">0</span><span class="p">),</span> <span class="bp">self</span><span class="p">.</span><span class="n">hidden_dim</span><span class="p">,</span> <span class="n">device</span><span class="o">=</span><span class="n">device</span><span class="p">)</span>
        <span class="n">c0</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">zeros</span><span class="p">(</span><span class="bp">self</span><span class="p">.</span><span class="n">layer_dim</span><span class="p">,</span> <span class="n">x</span><span class="p">.</span><span class="n">size</span><span class="p">(</span><span class="mi">0</span><span class="p">),</span> <span class="bp">self</span><span class="p">.</span><span class="n">hidden_dim</span><span class="p">,</span> <span class="n">device</span><span class="o">=</span><span class="n">device</span><span class="p">)</span>
        
        <span class="c1"># x = B, T, C
</span>        <span class="n">x</span> <span class="o">=</span> <span class="bp">self</span><span class="p">.</span><span class="n">embedding</span><span class="p">(</span><span class="n">x</span><span class="p">)</span>
        
        <span class="n">out</span><span class="p">,</span> <span class="p">(</span><span class="n">ht</span><span class="p">,</span> <span class="n">ct</span><span class="p">)</span> <span class="o">=</span> <span class="bp">self</span><span class="p">.</span><span class="n">lstm</span><span class="p">(</span><span class="n">x</span><span class="p">,</span> <span class="p">(</span><span class="n">h0</span><span class="p">,</span> <span class="n">c0</span><span class="p">))</span>
        <span class="n">out</span> <span class="o">=</span> <span class="bp">self</span><span class="p">.</span><span class="n">dropout</span><span class="p">(</span><span class="n">out</span><span class="p">[:,</span> <span class="o">-</span><span class="mi">1</span><span class="p">,</span> <span class="p">:])</span>
        <span class="n">out</span> <span class="o">=</span> <span class="bp">self</span><span class="p">.</span><span class="n">fc</span><span class="p">(</span><span class="n">out</span><span class="p">)</span>
        <span class="k">return</span> <span class="n">out</span> <span class="c1"># B, C
</span></code></pre></div></div>

<p>Next, we initialize the model:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">model</span> <span class="o">=</span> <span class="n">LSTMModel</span><span class="p">(</span><span class="n">input_dim</span><span class="p">,</span> <span class="n">hidden_dim</span><span class="p">,</span> <span class="n">layer_dim</span><span class="p">,</span> <span class="n">output_dim</span><span class="p">,</span> <span class="n">dropout</span><span class="p">)</span>
<span class="n">model</span><span class="p">.</span><span class="n">to</span><span class="p">(</span><span class="n">device</span><span class="p">)</span>  <span class="c1"># use gpu if possible
</span>
<span class="n">optimizer</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">optim</span><span class="p">.</span><span class="n">Adam</span><span class="p">(</span><span class="n">model</span><span class="p">.</span><span class="n">parameters</span><span class="p">(),</span> <span class="n">lr</span><span class="o">=</span><span class="n">learning_rate</span><span class="p">)</span>

<span class="k">for</span> <span class="n">i</span> <span class="ow">in</span> <span class="nb">range</span><span class="p">(</span><span class="nb">len</span><span class="p">(</span><span class="nb">list</span><span class="p">(</span><span class="n">model</span><span class="p">.</span><span class="n">parameters</span><span class="p">()))):</span>
    <span class="k">print</span><span class="p">(</span><span class="nb">list</span><span class="p">(</span><span class="n">model</span><span class="p">.</span><span class="n">parameters</span><span class="p">())[</span><span class="n">i</span><span class="p">].</span><span class="n">shape</span><span class="p">)</span>
</code></pre></div></div>

<p>We could play with different hyperparameters.
Increasing <code class="language-plaintext highlighter-rouge">hidden_dim</code> basically increases the complexity of the “memory” of the LSTM.
We surely want to increase <code class="language-plaintext highlighter-rouge">n_epochs</code> to increase number of times the LSTM “sees” all training data.</p>

<p>We can visualize the LSTM by utilizing the <code class="language-plaintext highlighter-rouge">draw_graph</code> function from the <code class="language-plaintext highlighter-rouge">torchview</code> package.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1"># (batch_size, sequence_len)
</span><span class="n">X_vis</span><span class="p">,</span> <span class="n">y_vis</span> <span class="o">=</span> <span class="n">train_set</span><span class="p">[</span><span class="mi">0</span><span class="p">:</span><span class="n">batch_size</span><span class="p">]</span>
<span class="k">print</span><span class="p">(</span><span class="sa">f</span><span class="s">'shape of X_vis: </span><span class="si">{</span><span class="n">X_vis</span><span class="p">.</span><span class="n">shape</span><span class="si">}</span><span class="s">'</span><span class="p">)</span>
<span class="k">print</span><span class="p">(</span><span class="sa">f</span><span class="s">'shape of y_vis: </span><span class="si">{</span><span class="n">y_vis</span><span class="p">.</span><span class="n">shape</span><span class="si">}</span><span class="s">'</span><span class="p">)</span>
<span class="k">print</span><span class="p">(</span><span class="sa">f</span><span class="s">'number of different symbols </span><span class="si">{</span><span class="n">vocab_size</span><span class="si">}</span><span class="s">'</span><span class="p">)</span>
<span class="n">X_vis</span><span class="p">,</span> <span class="n">y_vis</span> <span class="o">=</span> <span class="n">X_vis</span><span class="p">.</span><span class="n">to</span><span class="p">(</span><span class="n">device</span><span class="p">),</span> <span class="n">y_vis</span><span class="p">.</span><span class="n">to</span><span class="p">(</span><span class="n">device</span><span class="p">)</span>
<span class="n">model_vis</span> <span class="o">=</span> <span class="n">LSTMModel</span><span class="p">(</span><span class="n">input_dim</span><span class="p">,</span> <span class="n">hidden_dim</span><span class="p">,</span> <span class="n">layer_dim</span><span class="p">,</span> <span class="n">output_dim</span><span class="p">,</span> <span class="n">dropout</span><span class="p">)</span>
<span class="n">model_graph</span> <span class="o">=</span> <span class="n">draw_graph</span><span class="p">(</span><span class="n">model_vis</span><span class="p">,</span> <span class="n">input_data</span><span class="o">=</span><span class="n">X_vis</span><span class="p">,</span> <span class="n">device</span><span class="o">=</span><span class="n">device</span><span class="p">)</span>
<span class="n">model_graph</span><span class="p">.</span><span class="n">visual_graph</span>
</code></pre></div></div>

<div><img style="display:block; margin-left:auto; margin-right:auto; width:40%;" src="/Pages/assets/images/lstm-model.png" alt="LSTM model" />
<div style="display: table;margin: 0 auto;">Figure 4: The architecture of our LSTM model using a batch size of 64 and a sequence length also equal to 64. The alphabet consists of 38 unique tokens. Each single input is hot-encoded into a vector with 38 components. The LSTM uses a hidden state with 128 components. After the dropout the 128 components of hidden state are reduced to 38 components utilizing a normal linear layer (without an activation function).</div>
</div>
<p><br /></p>

<p>Note that the softmax is part of our loss <code class="language-plaintext highlighter-rouge">criterion</code> i.e. the cross entropy loss <code class="language-plaintext highlighter-rouge">torch.nn.CrossEntropyLoss()</code> which is part of the backpropagation, i.e., the training process.</p>

<h2 id="melody-generation-before-training">Melody Generation (Before Training)</h2>

<p>Given a sequence of arbitrary length, the <code class="language-plaintext highlighter-rouge">generate</code> function is used to generate a new piece of music.
<code class="language-plaintext highlighter-rouge">temperature</code> determines how much the probability distribution learned by the model is considered.</p>

<ul>
  <li><code class="language-plaintext highlighter-rouge">temperature</code> equal to 1.0 means that sampling is done from the probability distribution.</li>
  <li><code class="language-plaintext highlighter-rouge">temperature</code> approaching infinity means that sampling is done uniformly (more variation).</li>
  <li><code class="language-plaintext highlighter-rouge">temperature</code> approaching 0 means that higher probabilities are emphasized (less variation).</li>
</ul>

<p>We can set a maximum length for the piece and also provide the beginning of a piece.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">def</span> <span class="nf">next_event_number</span><span class="p">(</span><span class="n">idx</span><span class="p">,</span> <span class="n">temperature</span><span class="p">:</span><span class="nb">float</span><span class="p">):</span>
    <span class="k">with</span> <span class="n">torch</span><span class="p">.</span><span class="n">no_grad</span><span class="p">():</span>
        <span class="n">logits</span> <span class="o">=</span> <span class="n">model</span><span class="p">(</span><span class="n">idx</span><span class="p">)</span>
        <span class="n">probs</span> <span class="o">=</span> <span class="n">F</span><span class="p">.</span><span class="n">softmax</span><span class="p">(</span><span class="n">logits</span> <span class="o">/</span> <span class="n">temperature</span><span class="p">,</span> <span class="n">dim</span><span class="o">=</span><span class="mi">1</span><span class="p">)</span> <span class="c1"># B, C
</span>        <span class="n">idx_next</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">multinomial</span><span class="p">(</span><span class="n">probs</span><span class="p">,</span> <span class="n">num_samples</span><span class="o">=</span><span class="mi">1</span><span class="p">)</span>
        <span class="k">return</span> <span class="n">idx_next</span>

<span class="k">def</span> <span class="nf">generate</span><span class="p">(</span><span class="n">seq</span><span class="p">:</span> <span class="nb">list</span><span class="p">[</span><span class="nb">str</span><span class="p">]</span><span class="o">=</span><span class="bp">None</span><span class="p">,</span> <span class="n">max_len</span><span class="p">:</span><span class="nb">int</span><span class="o">=</span><span class="bp">None</span><span class="p">,</span> <span class="n">temperature</span><span class="p">:</span><span class="nb">float</span><span class="o">=</span><span class="mf">1.0</span><span class="p">):</span>
    <span class="k">with</span> <span class="n">torch</span><span class="p">.</span><span class="n">no_grad</span><span class="p">():</span>
        <span class="n">generated_encoded_song</span> <span class="o">=</span> <span class="p">[]</span>
        <span class="k">if</span> <span class="n">seq</span> <span class="o">!=</span> <span class="bp">None</span><span class="p">:</span>
            <span class="n">idx</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">tensor</span><span class="p">(</span>
                <span class="p">[[</span><span class="n">string_to_int</span><span class="p">.</span><span class="n">encode</span><span class="p">(</span><span class="n">char</span><span class="p">)</span> <span class="k">for</span> <span class="n">char</span> <span class="ow">in</span> <span class="n">seq</span><span class="p">]],</span> 
                <span class="n">device</span><span class="o">=</span><span class="n">device</span>
            <span class="p">)</span>
            <span class="n">generated_encoded_song</span> <span class="o">=</span> <span class="n">seq</span><span class="p">.</span><span class="n">copy</span><span class="p">()</span>
        <span class="k">else</span><span class="p">:</span>
            <span class="n">idx</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">tensor</span><span class="p">([[</span><span class="n">string_to_int</span><span class="p">.</span><span class="n">encode</span><span class="p">(</span><span class="n">TERM_SYMBOL</span><span class="p">)]],</span> <span class="n">device</span><span class="o">=</span><span class="n">device</span><span class="p">)</span>
        
        <span class="k">while</span> <span class="n">max_len</span> <span class="o">==</span> <span class="bp">None</span> <span class="ow">or</span> <span class="n">max_len</span> <span class="o">&gt;</span> <span class="nb">len</span><span class="p">(</span><span class="n">generated_encoded_song</span><span class="p">):</span>
            <span class="n">idx_next</span> <span class="o">=</span> <span class="n">next_event_number</span><span class="p">(</span><span class="n">idx</span><span class="p">,</span> <span class="n">temperature</span><span class="p">)</span>
            <span class="n">char</span> <span class="o">=</span> <span class="n">string_to_int</span><span class="p">.</span><span class="n">decode</span><span class="p">(</span><span class="n">idx_next</span><span class="p">.</span><span class="n">item</span><span class="p">())</span>
            <span class="k">if</span> <span class="n">idx_next</span> <span class="o">==</span> <span class="n">string_to_int</span><span class="p">.</span><span class="n">encode</span><span class="p">(</span><span class="n">TERM_SYMBOL</span><span class="p">):</span>
                <span class="k">break</span>
            <span class="n">idx</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">cat</span><span class="p">((</span><span class="n">idx</span><span class="p">,</span> <span class="n">idx_next</span><span class="p">),</span> <span class="n">dim</span><span class="o">=</span><span class="mi">1</span><span class="p">)</span> <span class="c1"># B, T+1, C
</span>            <span class="n">generated_encoded_song</span><span class="p">.</span><span class="n">append</span><span class="p">(</span><span class="n">char</span><span class="p">)</span>
            
        <span class="k">return</span> <span class="n">generated_encoded_song</span>
</code></pre></div></div>

<p>Of course, the results are almost random because the parameters of our model are initialized randomly and we did not train it yet.
The following code snippet generates 5 scores.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1"># number of songs we want to generate
</span><span class="n">n_scores</span> <span class="o">=</span> <span class="mi">5</span>
<span class="n">temperature</span> <span class="o">=</span> <span class="mf">0.6</span>
<span class="n">before_new_songs</span> <span class="o">=</span> <span class="p">[]</span>
<span class="k">for</span> <span class="n">_</span> <span class="ow">in</span> <span class="nb">range</span><span class="p">(</span><span class="n">n_scores</span><span class="p">):</span>
    <span class="n">encoded_song</span> <span class="o">=</span> <span class="n">generate</span><span class="p">(</span><span class="n">max_len</span><span class="o">=</span><span class="mi">13</span><span class="p">,</span><span class="n">temperature</span><span class="o">=</span><span class="n">temperature</span><span class="p">)</span>
    <span class="k">print</span><span class="p">(</span><span class="sa">f</span><span class="s">'generated </span><span class="si">{</span><span class="s">" "</span><span class="p">.</span><span class="n">join</span><span class="p">(</span><span class="n">encoded_song</span><span class="p">)</span><span class="si">}</span><span class="s"> consisting of </span><span class="si">{</span><span class="nb">len</span><span class="p">(</span><span class="n">encoded_song</span><span class="p">)</span><span class="si">}</span><span class="s"> notes'</span><span class="p">)</span>
    <span class="n">before_new_songs</span><span class="p">.</span><span class="n">append</span><span class="p">(</span><span class="n">encoded_song</span><span class="p">)</span>
</code></pre></div></div>

<p>Let’s listen to the first one:</p>

<audio controls="">
  <source src="/Pages/assets/audio/before_g_song.mp3" type="audio/mp3" />
  Your browser does not support the audio element.
</audio>

<h2 id="training">Training</h2>

<p>For training, we use something called a <code class="language-plaintext highlighter-rouge">DataLoader</code>. 
This helps us to access our data more easily. 
For example, we shuffle our data before training.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">train_loader</span> <span class="o">=</span> <span class="n">DataLoader</span><span class="p">(</span><span class="n">train_set</span><span class="p">,</span> <span class="n">batch_size</span><span class="o">=</span><span class="n">batch_size</span><span class="p">,</span> <span class="n">shuffle</span><span class="o">=</span><span class="bp">True</span><span class="p">)</span>
<span class="n">val_loader</span> <span class="o">=</span> <span class="n">DataLoader</span><span class="p">(</span><span class="n">val_set</span><span class="p">,</span> <span class="n">batch_size</span><span class="o">=</span><span class="n">batch_size</span><span class="p">,</span> <span class="n">shuffle</span><span class="o">=</span><span class="bp">True</span><span class="p">)</span>
<span class="n">test_loader</span> <span class="o">=</span> <span class="n">DataLoader</span><span class="p">(</span><span class="n">test_set</span><span class="p">,</span> <span class="n">batch_size</span><span class="o">=</span><span class="n">batch_size</span><span class="p">,</span><span class="n">shuffle</span><span class="o">=</span><span class="bp">True</span><span class="p">)</span>
</code></pre></div></div>

<p>The code for training seems a bit complicated because we use batches. 
This is due to dealing with a large amount of data, and we don’t send all of it through the network at once (per training step), but only a part of it, namely <code class="language-plaintext highlighter-rouge">batch_size</code> many. 
An <code class="language-plaintext highlighter-rouge">epoch</code> is defined by the fact that all training data have been sent through the network once.</p>

<p>In essence, nothing else happens but:</p>

<ol>
  <li>Send Batch through the network (Forward pass)</li>
  <li>Calculate error/cost</li>
  <li>Propagate gradients of the cost function with respect to the model parameters backwards through the network (Backward pass)</li>
  <li>Update model parameters</li>
</ol>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">def</span> <span class="nf">train_one_epoch</span><span class="p">(</span><span class="n">epoch_index</span><span class="p">,</span> <span class="n">tb_writer</span><span class="p">,</span> <span class="n">n_epochs</span><span class="p">):</span>
    <span class="n">running_loss</span> <span class="o">=</span> <span class="mf">0.0</span>
    <span class="n">last_loss</span> <span class="o">=</span> <span class="mf">0.0</span>
    <span class="n">all_steps</span> <span class="o">=</span> <span class="n">n_epochs</span> <span class="o">*</span> <span class="nb">len</span><span class="p">(</span><span class="n">train_loader</span><span class="p">)</span>
    
    <span class="k">for</span> <span class="n">i</span><span class="p">,</span> <span class="n">data</span> <span class="ow">in</span> <span class="nb">enumerate</span><span class="p">(</span><span class="n">train_loader</span><span class="p">):</span>
        <span class="n">local_X</span><span class="p">,</span> <span class="n">local_y</span> <span class="o">=</span> <span class="n">data</span>
        <span class="n">local_X</span><span class="p">,</span> <span class="n">local_y</span> <span class="o">=</span> <span class="n">local_X</span><span class="p">.</span><span class="n">to</span><span class="p">(</span><span class="n">device</span><span class="p">),</span> <span class="n">local_y</span><span class="p">.</span><span class="n">to</span><span class="p">(</span><span class="n">device</span><span class="p">)</span>
        <span class="n">optimizer</span><span class="p">.</span><span class="n">zero_grad</span><span class="p">()</span>
        <span class="n">outputs</span> <span class="o">=</span> <span class="n">model</span><span class="p">(</span><span class="n">local_X</span><span class="p">)</span>
        
        <span class="n">loss</span> <span class="o">=</span> <span class="n">criterion</span><span class="p">(</span><span class="n">outputs</span><span class="p">,</span> <span class="n">local_y</span><span class="p">)</span>
        <span class="n">loss</span><span class="p">.</span><span class="n">backward</span><span class="p">()</span>
        <span class="n">optimizer</span><span class="p">.</span><span class="n">step</span><span class="p">()</span>
        
        <span class="n">running_loss</span> <span class="o">+=</span> <span class="n">loss</span><span class="p">.</span><span class="n">item</span><span class="p">()</span>
        <span class="k">if</span> <span class="n">i</span> <span class="o">%</span> <span class="n">eval_interval</span> <span class="o">==</span> <span class="n">eval_interval</span><span class="o">-</span><span class="mi">1</span><span class="p">:</span>
            <span class="n">last_loss</span> <span class="o">=</span> <span class="n">running_loss</span> <span class="o">/</span> <span class="n">eval_interval</span>
            
            <span class="n">steps</span> <span class="o">=</span> <span class="n">epoch_index</span> <span class="o">*</span> <span class="nb">len</span><span class="p">(</span><span class="n">train_loader</span><span class="p">)</span> <span class="o">+</span> <span class="p">(</span><span class="n">i</span><span class="o">+</span><span class="mi">1</span><span class="p">)</span>
            
            <span class="n">ep_str</span> <span class="o">=</span> <span class="sa">f</span><span class="s">'Epoch [</span><span class="si">{</span><span class="n">epoch_index</span><span class="o">+</span><span class="mi">1</span><span class="si">}</span><span class="s">/</span><span class="si">{</span><span class="n">n_epochs</span><span class="si">}</span><span class="s">]'</span>
            <span class="n">step_str</span> <span class="o">=</span> <span class="sa">f</span><span class="s">'Step [</span><span class="si">{</span><span class="n">steps</span><span class="si">}</span><span class="s">/</span><span class="si">{</span><span class="n">all_steps</span><span class="si">}</span><span class="s">]'</span>
            <span class="n">loss_str</span> <span class="o">=</span> <span class="sa">f</span><span class="s">'Loss: </span><span class="si">{</span><span class="n">last_loss</span><span class="si">:</span><span class="p">.</span><span class="mi">4</span><span class="n">f</span><span class="si">}</span><span class="s">'</span>
            <span class="k">print</span><span class="p">(</span><span class="sa">f</span><span class="s">'</span><span class="si">{</span><span class="n">ep_str</span><span class="si">}</span><span class="s">, </span><span class="si">{</span><span class="n">step_str</span><span class="si">}</span><span class="s">, </span><span class="si">{</span><span class="n">loss_str</span><span class="si">}</span><span class="s">'</span><span class="p">)</span>

            <span class="n">tb_x</span> <span class="o">=</span> <span class="n">epoch_index</span> <span class="o">*</span> <span class="nb">len</span><span class="p">(</span><span class="n">train_loader</span><span class="p">)</span> <span class="o">+</span> <span class="n">i</span> <span class="o">+</span> <span class="mi">1</span>
            <span class="n">tb_writer</span><span class="p">.</span><span class="n">add_scalar</span><span class="p">(</span><span class="s">'Loss/train'</span><span class="p">,</span> <span class="n">last_loss</span><span class="p">,</span> <span class="n">tb_x</span><span class="p">)</span>
            <span class="n">running_loss</span> <span class="o">=</span> <span class="mf">0.</span>
            
    <span class="k">return</span> <span class="n">last_loss</span>

<span class="c1"># Initializing in a separate cell so we can easily add more epochs to the same run
</span><span class="k">def</span> <span class="nf">train</span><span class="p">(</span><span class="n">n_epochs</span><span class="p">,</span><span class="n">respect_val</span><span class="o">=</span><span class="bp">False</span><span class="p">):</span>
    <span class="n">timestamp</span> <span class="o">=</span> <span class="n">datetime</span><span class="p">.</span><span class="n">now</span><span class="p">().</span><span class="n">strftime</span><span class="p">(</span><span class="s">'%Y%m%d_%H%M%S'</span><span class="p">)</span>
    <span class="n">writer</span> <span class="o">=</span> <span class="n">SummaryWriter</span><span class="p">(</span><span class="s">'runs/fashion_trainer_{}'</span><span class="p">.</span><span class="nb">format</span><span class="p">(</span><span class="n">timestamp</span><span class="p">))</span>
    <span class="n">best_vloss</span> <span class="o">=</span> <span class="mi">1_000_000</span>

    <span class="k">for</span> <span class="n">epoch</span> <span class="ow">in</span> <span class="nb">range</span><span class="p">(</span><span class="n">n_epochs</span><span class="p">):</span>    
        <span class="n">model</span><span class="p">.</span><span class="n">train</span><span class="p">(</span><span class="bp">True</span><span class="p">)</span>
        <span class="n">avg_loss</span> <span class="o">=</span> <span class="n">train_one_epoch</span><span class="p">(</span><span class="n">epoch</span><span class="p">,</span> <span class="n">writer</span><span class="p">,</span> <span class="n">n_epochs</span><span class="p">)</span>
        
        <span class="n">model</span><span class="p">.</span><span class="n">train</span><span class="p">(</span><span class="bp">False</span><span class="p">)</span>
        <span class="n">running_vloss</span> <span class="o">=</span> <span class="mf">0.0</span>
        
        <span class="k">for</span> <span class="n">i</span><span class="p">,</span> <span class="n">vdata</span> <span class="ow">in</span> <span class="nb">enumerate</span><span class="p">(</span><span class="n">val_loader</span><span class="p">):</span>
            
            <span class="n">local_X</span><span class="p">,</span> <span class="n">local_y</span> <span class="o">=</span> <span class="n">vdata</span>
            <span class="n">local_X</span><span class="p">,</span> <span class="n">local_y</span> <span class="o">=</span> <span class="n">local_X</span><span class="p">.</span><span class="n">to</span><span class="p">(</span><span class="n">device</span><span class="p">),</span> <span class="n">local_y</span><span class="p">.</span><span class="n">to</span><span class="p">(</span><span class="n">device</span><span class="p">)</span>
            
            <span class="n">voutputs</span> <span class="o">=</span> <span class="n">model</span><span class="p">(</span><span class="n">local_X</span><span class="p">)</span>
            <span class="n">vloss</span> <span class="o">=</span> <span class="n">criterion</span><span class="p">(</span><span class="n">voutputs</span><span class="p">,</span> <span class="n">local_y</span><span class="p">)</span>
            <span class="n">running_vloss</span> <span class="o">+=</span> <span class="n">vloss</span>
            
        <span class="n">avg_vloss</span> <span class="o">=</span> <span class="n">running_vloss</span> <span class="o">/</span> <span class="p">(</span><span class="n">i</span><span class="o">+</span><span class="mi">1</span><span class="p">)</span>

        <span class="n">ep_str</span> <span class="o">=</span> <span class="sa">f</span><span class="s">'Epoch [</span><span class="si">{</span><span class="n">epoch</span><span class="o">+</span><span class="mi">1</span><span class="si">}</span><span class="s">/</span><span class="si">{</span><span class="n">n_epochs</span><span class="si">}</span><span class="s">]'</span>
        <span class="n">tloss_str</span> <span class="o">=</span> <span class="sa">f</span><span class="s">'Train-Loss: </span><span class="si">{</span><span class="n">avg_loss</span><span class="si">:</span><span class="p">.</span><span class="mi">4</span><span class="n">f</span><span class="si">}</span><span class="s">'</span>
        <span class="n">vloss_str</span> <span class="o">=</span> <span class="sa">f</span><span class="s">'Val-Loss: </span><span class="si">{</span><span class="n">avg_vloss</span><span class="si">:</span><span class="p">.</span><span class="mi">4</span><span class="n">f</span><span class="si">}</span><span class="s">'</span>
        <span class="k">print</span><span class="p">(</span><span class="sa">f</span><span class="s">'</span><span class="si">{</span><span class="n">ep_str</span><span class="si">}</span><span class="s">, </span><span class="si">{</span><span class="n">tloss_str</span><span class="si">}</span><span class="s">, </span><span class="si">{</span><span class="n">vloss_str</span><span class="si">}</span><span class="s">'</span><span class="p">)</span>
        
        <span class="n">writer</span><span class="p">.</span><span class="n">add_scalars</span><span class="p">(</span>
            <span class="s">'Training vs. Validation Loss'</span><span class="p">,</span> 
            <span class="p">{</span><span class="s">'Training'</span><span class="p">:</span> <span class="n">avg_loss</span><span class="p">,</span> <span class="s">'Validation'</span><span class="p">:</span> <span class="n">avg_vloss</span><span class="p">},</span> 
            <span class="n">epoch</span>
        <span class="p">)</span>

        <span class="n">writer</span><span class="p">.</span><span class="n">flush</span><span class="p">()</span>
        
        <span class="k">if</span> <span class="ow">not</span> <span class="n">respect_val</span> <span class="ow">or</span> <span class="p">(</span><span class="n">respect_val</span> <span class="ow">and</span> <span class="n">avg_vloss</span> <span class="o">&lt;</span> <span class="n">best_vloss</span><span class="p">):</span>
            <span class="n">best_vloss</span> <span class="o">=</span> <span class="n">avg_vloss</span>
            <span class="n">model_path</span> <span class="o">=</span> <span class="s">'./models/_model_{}_{}'</span><span class="p">.</span><span class="nb">format</span><span class="p">(</span><span class="n">timestamp</span><span class="p">,</span> <span class="n">epoch</span><span class="p">)</span>
            <span class="n">torch</span><span class="p">.</span><span class="n">save</span><span class="p">(</span><span class="n">model</span><span class="p">.</span><span class="n">state_dict</span><span class="p">(),</span> <span class="n">model_path</span><span class="p">)</span>
</code></pre></div></div>

<p>Calling <code class="language-plaintext highlighter-rouge">train</code> starts the training.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">train</span><span class="p">(</span><span class="n">n_epochs</span><span class="p">)</span>
</code></pre></div></div>

<p>The best model from the training can be found in the folder <code class="language-plaintext highlighter-rouge">./models</code> and can be loaded as follows</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">model_path</span> <span class="o">=</span> <span class="s">'./models/pretrained_1_128_best_val'</span>

<span class="k">if</span> <span class="n">device</span><span class="p">.</span><span class="nb">type</span> <span class="o">==</span> <span class="s">'cpu'</span><span class="p">:</span>
    <span class="n">model</span><span class="p">.</span><span class="n">load_state_dict</span><span class="p">(</span><span class="n">torch</span><span class="p">.</span><span class="n">load</span><span class="p">(</span><span class="n">model_path</span><span class="p">,</span> <span class="n">map_location</span><span class="o">=</span><span class="n">torch</span><span class="p">.</span><span class="n">device</span><span class="p">(</span><span class="s">'cpu'</span><span class="p">)))</span>
<span class="k">elif</span> <span class="n">torch</span><span class="p">.</span><span class="n">backends</span><span class="p">.</span><span class="n">mps</span><span class="p">.</span><span class="n">is_available</span><span class="p">():</span>
    <span class="n">model</span><span class="p">.</span><span class="n">load_state_dict</span><span class="p">(</span><span class="n">torch</span><span class="p">.</span><span class="n">load</span><span class="p">(</span><span class="n">model_path</span><span class="p">,</span> <span class="n">map_location</span><span class="o">=</span><span class="n">torch</span><span class="p">.</span><span class="n">device</span><span class="p">(</span><span class="s">'mps'</span><span class="p">)))</span>
<span class="k">else</span><span class="p">:</span>
    <span class="n">model</span><span class="p">.</span><span class="n">load_state_dict</span><span class="p">(</span><span class="n">torch</span><span class="p">.</span><span class="n">load</span><span class="p">(</span><span class="n">model_path</span><span class="p">))</span>
<span class="n">model</span><span class="p">.</span><span class="nb">eval</span><span class="p">()</span>
</code></pre></div></div>

<h2 id="melody-generation-after-training">Melody Generation (After Training)</h2>

<p>After training or after we load our pretrained model, we generate new pieces:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">n_scores</span> <span class="o">=</span> <span class="mi">5</span>
<span class="n">temperature</span> <span class="o">=</span> <span class="mf">0.6</span>
<span class="n">after_new_songs</span> <span class="o">=</span> <span class="p">[]</span>
<span class="k">for</span> <span class="n">_</span> <span class="ow">in</span> <span class="nb">range</span><span class="p">(</span><span class="n">n_scores</span><span class="p">):</span>
    <span class="n">encoded_song</span> <span class="o">=</span> <span class="n">generate</span><span class="p">(</span><span class="n">max_len</span><span class="o">=</span><span class="mi">120</span><span class="p">,</span><span class="n">temperature</span><span class="o">=</span><span class="n">temperature</span><span class="p">)</span>
    <span class="k">print</span><span class="p">(</span><span class="sa">f</span><span class="s">'generated </span><span class="si">{</span><span class="s">" "</span><span class="p">.</span><span class="n">join</span><span class="p">(</span><span class="n">encoded_song</span><span class="p">)</span><span class="si">}</span><span class="s"> consisting of </span><span class="si">{</span><span class="nb">len</span><span class="p">(</span><span class="n">encoded_song</span><span class="p">)</span><span class="si">}</span><span class="s"> notes'</span><span class="p">)</span>
    <span class="n">after_new_songs</span><span class="p">.</span><span class="n">append</span><span class="p">(</span><span class="n">encoded_song</span><span class="p">)</span>

<span class="n">after_generated_scores</span> <span class="o">=</span> <span class="n">encoder</span><span class="p">.</span><span class="n">decode_songs</span><span class="p">(</span><span class="n">after_new_songs</span><span class="p">)</span>
<span class="n">Audio</span><span class="p">(</span><span class="n">score_to_wav</span><span class="p">(</span><span class="n">after_generated_scores</span><span class="p">[</span><span class="mi">0</span><span class="p">],</span> <span class="s">'a_g_song.wav'</span><span class="p">))</span>
</code></pre></div></div>

<p>We start to hear repetition and some structure within the piece:</p>

<audio controls="">
  <source src="/Pages/assets/audio/a_g_song.mp3" type="audio/mp3" />
  Your browser does not support the audio element.
</audio>

<h2 id="real-world-example">Real World Example</h2>

<p>In the realm of musical innovation, a significant advancement occurred with the development of a sophisticated <strong>LSTM</strong> model designed to create expressive piano roll music. 
This model was introduced in a notable study by <a class="citation" href="#oore:2018">(Oore et al., 2018)</a>. 
The researchers devised a unique discrete-event based representation for piano rolls, encompassing a diverse range of 413 different events. 
The architecture of their model was meticulously structured, comprising three hidden LSTM layers, each equipped with 512 cells. 
This design choice facilitated the processing of a 413-dimensional one-hot vector as input, with the model subsequently generating a categorical distribution over the same dimensional space.</p>

<p>The training process of the model was finely tuned, employing a mini-batch size of 64 and a learning rate of 0.001, alongside the implementation of teacher forcing techniques. 
For those interested in experiencing the model’s capabilities firsthand, a collection of generated music pieces is available for listening at this <a href="https://clyp.it/user/3mdslat4">link</a>. 
However, it’s important to note a primary limitation of this model: its tendency to produce relatively brief musical compositions, typically ranging from 10 to 20 seconds in duration. 
The authors also emphasized the critical role of high-quality data in achieving optimal results with this model.</p>

<p>Following this development, the field witnessed the emergence of the Music Transformer, introduced by <a class="citation" href="#huang:2018">(Huang et al., 2018)</a>.
This model also utilized a similar piano roll representation but marked a significant leap forward by employing the <strong>transformer</strong> architecture. 
This innovative approach enabled the Music Transformer to learn and reproduce longer sequences, demonstrating the capability to capture more extended musical dependencies. 
The transformer architecture and its implications in music generation will be further explored in the next installment of this series.</p>

<h2 id="references">References</h2>

<ol class="bibliography"><li><span id="oore:2018">Oore, S., Simon, I., Dieleman, S., Eck, D., &amp; Simonyan, K. (2018). <i>This time with feeling: Learning expressive musical performance</i>.</span></li>
<li><span id="huang:2018">Huang, C.-Z. A., Vaswani, A., Uszkoreit, J., Shazeer, N., Hawthorne, C., Dai, A. M., Hoffman, M. D., &amp; Eck, D. (2018). Music Transformer: Generating music with long-term structure. <i>ArXiv Preprint ArXiv:1809.04281</i>.</span></li></ol>]]></content><author><name>Benedikt Zönnchen</name><email>benedikt.zoennchen@web.de</email></author><category term="Music" /><category term="ML" /><category term="LSTM" /><summary type="html"><![CDATA[This article is the continuation of a series. It is recommended that you read part I and II first. This time we use a recurrent neural network (RNN), more precisely an LSTM, which I explained a little bit in the introduction. An LSTM is a RNN that counteracts the problem of exploding and vanishing gradients.]]></summary></entry><entry><title type="html">Laws of Form</title><link href="https://bzoennchen.github.io/Pages/2023/11/19/laws-of-form.html" rel="alternate" type="text/html" title="Laws of Form" /><published>2023-11-19T00:00:00+01:00</published><updated>2023-11-19T00:00:00+01:00</updated><id>https://bzoennchen.github.io/Pages/2023/11/19/laws-of-form</id><content type="html" xml:base="https://bzoennchen.github.io/Pages/2023/11/19/laws-of-form.html"><![CDATA[<p>In my last blog <a href="/Pages/2023/10/07/system-theory-and-ai.html">post</a>, I discussed Niklas Luhmann’s Social Systems Theory and I emphasized that his theory is based on differentiation and seemingly paradox relations.
My general understanding of Luhmann’s radical constructivism in the most reductive sense is that there are no unified and independent objects—there are only differences.
What an observer can identify as an object is a differentiation of a system and its environment, a foreground and its background, an interior and the external.
However, any observer is itself a distinction between system and environment.
I want to further investigate this idea by looking into Luhmann’s inspiration—his muse so to say.
So let me examine the logic of Georg Spencer-Brown presented in his work <em>Laws of Form</em> <a class="citation" href="#brown:1969">(Spencer-Brown, 1969)</a>.</p>

<h2 id="the-principle-of-differentiation">The Principle of Differentiation</h2>

<p>As mentioned, Luhmann’s theory is based on paradoxes such as</p>

<blockquote>
  <p>This statement is wrong.</p>
</blockquote>

<p>The statement is logically problematic because it is self-referential. 
If the statement is false, it states that it is in fact true and if the statement is true, it states that it is false.
Another famous paradox is Russell’s antinomy of the naïve set theory:</p>

\[R := \left\{ x \ | \ x \not\in x \right\}.\]

<p>\(R\) is the set of all sets that do not contain themselves which seems to be a well-defined mathematical object.
However, if we introduce the self-referential relation, we run into a paradox:</p>

<p>\begin{equation} 
R \in R \iff R \not\in R
\end{equation}</p>

<p>This paradox is related to the barber that shaves everyone that does not shave themselves.
If that is the case, does the barber shave themselves?</p>

<p>Another example involves the the proof of the <a href="/Pages/2021/06/08/Informatics-a-love-letter.html">Halting Problem</a>.
To prove it, one can establish a self-referential relation between a machine that presumable solves the Halting Problem.
The machine does not halt if the machine it checks halts, and it halts if the machine it checks does not halt.
Via the self-referential relation, that is, by letting the machine check itself, we get a contradiction.</p>

<p>The idea of Spencer-Brown is to resolve these paradoxes over time.
\(R \in R\) holds at one moment in time and \(R \not\in R\) holds at the next moment.
Note however that he does not resolve Russell’s antinomy <a class="citation" href="#cull:1979">(Cull &amp; Frank, 1979)</a>.
His idea of resolving paradoxes over time is reminiscent of the Hegelian dialectic—a process of self-creation.
The self-referential relation is a paradox if we ignore time and it becomes a generator if we consider time and place.</p>

<p>This led the biologists Huberto R. Maturana and Francisco J. Varela to the concept of <em>autopoiesis</em> <a class="citation" href="#maturana:1987">(Maturana &amp; Varela, 1987)</a>.
Furthermore, the importance of differentiation is inspired by the logic of Spencer-Brown and his work <em>Laws of Form</em> <a class="citation" href="#brown:1969">(Spencer-Brown, 1969)</a>.
He resolves the paradox by a similar idea that gave us imaginary numbers, that is, by using what he calls <em>Re-entry</em> which (re-)introduces a system into itself.</p>

<p>Luhmann integrated this idea into his social systems theory.
For example, the media can observe and reintroduce itself into itself. 
It can use its systemic operations on itself, i.e. it can report on itself.</p>

<p>Spencer-Brown begins his work by a quote from Lao-Tse (a stand-in for many different authors) thus begins by philosophical considerations:</p>

<blockquote>
  <p>Wu ming tain di zhi shi. – Loa-Tse</p>
</blockquote>

<p>The sentence has mainly two different meanings.
One is:</p>

<blockquote>
  <p>The beginning of heaven and earth is without a name.</p>
</blockquote>

<p>The other one is:</p>

<blockquote>
  <p>‘Nothing’ is the name of the beginning of heaven and earth.</p>
</blockquote>

<p>A paradox arises: How can Nothing be nothing if we can call it ‘Nothing’?
Furthermore, the quote points to a distinction between heaven and earth.
Can there be heaven without earth—a <em>calling</em> or <em>indication</em> without a <em>distinction</em>?
Spencer-Brown begins by the assumption that there is no such thing:</p>

<blockquote>
  <p>We take the idea of distinction and the idea of indication and that we cannot make an indication without making a distinction as given.
Therefore, we take the form of distinction as the form itself. – Georg Spencer-Brown</p>
</blockquote>

<p>In other words, what we normally identify as object (the form / system) is for Spencer-Brown equal to the distinction (system-environment differentiation).
There is no clear separation between the object or the result of distinction and the process of distinguishing.
Therefore, the process must be integrated into Spencer-Brown’s logic and as we will see, there is no clear separation between objects and operations in Spencer-Browns calculus.
Spencer-Brown thinks that differentiation is a proto-operation that is more fundamental than performing calculations or writing text because to do these activities we have to differentiate beforehand.
I cannot calculate 1 + 1 = 2 without distinguishing between the different symbols and a symbol and ‘nothing’ or the void.</p>

<p>Spencer-Brown uses the mark or cross (result) which at the same time marks (process).
The mark is, calls, and makes a difference.</p>

<p><br /></p>
<div><img style="height:30px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/mark.png" alt="The mark." /></div>
<p><br /></p>

<p>There is an interior of the mark and not the interior—the system and its environment.
I can make a distinction again (repetition):</p>

<p><br /></p>
<div><img style="height:30px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/calling.png" alt="The calling." /></div>
<p><br /></p>

<p>But making a distinction again does not change the distinction.</p>

<blockquote>
  <p>Calling something back-to-back by its name does not change its name. – Spencer-Brown</p>
</blockquote>

<p>The reverse is also true; therefore, Spencer-Brown introduces the <em>Law of Calling</em>:</p>

<p><br /></p>
<div><img style="height:30px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/law-of-calling.png" alt="Law of Calling." /></div>
<p><br /></p>

<p>The second transformation called <em>Law of Crossing</em> is less intuitive.
Crossing twice reverses the first crossing.</p>

<p><br /></p>
<div><img style="height:30px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/law-of-crossing.png" alt="Law of Crossing." /></div>
<p><br /></p>

<blockquote>
  <p>If a boundary is crossed twice, the original state will be reestablished.
The repetition of the crossing has a different value than the single crossing.
The reason is that in-between the reversal happens.
Crossing changes the side.
Re-crossing reverses this operation. – Spencer-Brown</p>
</blockquote>

<p>With only these two laws, Spencer-Brown established a logic calculus and we can start doing mathematics.
Interestingly, the <em>Law of Crossing</em> and the <em>Law of Calling</em> are implicitly established via the position of the marks.
There is no operator introduced because the result and process, indication and differentiation, the mark and the process of marking are not separated.</p>

<p>Let’s see what we can do with this calculus.
Let \(a, b\) variables, then the following holds:</p>

<p><br /></p>
<div><img style="height:230px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/transformations.png" alt="Transformations." /></div>
<p><br /></p>

<p>Let’s have a look at the last transformation.
Let us assume \(a\) is a <strong>mark</strong>.
Then we get:</p>

<p><br /></p>
<div><img style="height:30px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/transformation-1.png" alt="First possibility." /></div>
<p><br /></p>

<p>Let \(a\) be <strong>unmarked</strong> instead, then we can follow:</p>

<p><br /></p>
<div><img style="height:30px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/transformation-2.png" alt="Second possibility." /></div>
<p><br /></p>

<h2 id="the-re-entry">The Re-Entry</h2>

<p>How can we re-introduce the system (here an equation) into itself?
Or in other words: How does the <em>re-entry</em> work?
We can start by a simple self-referential algebraic equation:</p>

<p>\begin{equation}
x^2 = ax + b, \quad a, b \in \mathbb{R}.
\end{equation}</p>

<p>This equation has well-known solutions. 
It is also known that solutions can be imaginary, i.e., \(x\) might be of the form \(r + si\) with \(r, s \in \mathbb{R}\) and \(i^2 = -1\).
To see the re-entry, we can rewrite the equation above to get</p>

<p>\begin{equation}
x = a + b/x,
\end{equation}</p>

<p>thus the self-reference is obvious and we solve the equation by the re-entry</p>

<p>\begin{equation}
x = a + b/(a +b/(a+b/(a+b/a+b/(a + \ldots)))).
\end{equation}</p>

<p>Using this infinite formalism it is literally the case that</p>

<p>\begin{equation}
x = a + b/x,
\end{equation}
holds.
Using the same formalism, we can define the imaginary number \(i = -1/i\) as literally</p>

<p>\begin{equation}
i = -1 /(-1 /(-1 / (-1 / \ldots )))
\end{equation}</p>

<p>but what does this mean?
The system is not a number but a process, a generator that generates itself.
\(i\) alternates between 1 and -1.
Interestingly, this is precisely how we use the equal sign in most programming languages.
Writing <code class="language-plaintext highlighter-rouge">i = i / -1</code> in a programming language means</p>

<p>\begin{equation}
i \leftarrow \frac{i}{-1}.
\end{equation}</p>

<p>The next step is to introduce such a re-entry into logic.
Similar to the imaginary number \(i\), Spencer-Brown gives us the following fundamental paradox (<em>The Re-Entry of the Mark</em>):</p>

<p><br /></p>
<div><img style="height:30px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/j.png" alt="Imaginary truth value." /></div>
<p><br /></p>

<p>The solution is an alternation between a marked and unmarked state—between true and false.
A state that might seem contradictory in space, makes sense if it is observed in time and space.
Again, time resolves the paradox.
To highlight the re-entry, Spencer-Brown also uses the following notation:</p>

<p><br /></p>
<div><img style="height:30px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/j-2.png" alt="Notation of the re-entry." /></div>
<p><br /></p>

<p>We could similarily notate \(i\) as</p>

<p><br /></p>
<div><img style="height:35px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/i-2.png" alt="Notation of the re-entry for the imaginary number." /></div>
<p><br /></p>

<p>If we change <strong>all</strong> symbols within a system equally, there is no reason not to calculate with a self-generating process.
For example, we can state the following:</p>

<p><br /></p>
<div><img style="height:30px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/j-calculations.png" alt="Calculating with the re-entry." /></div>
<p><br /></p>

<p>However, it is forbidden to only change one appearance of \(J\)!
Following this simple rule, no paradox or inconsistency arises.
We can go on and evaluate the following transformation:</p>

<p><br /></p>
<div><img style="height:30px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/wave.png" alt="Wave equation." /></div>
<p><br /></p>

<p>which describes two alternating waves shifted by one cycle resulting in a mark.</p>

<p>Spencer-Brown goes on and defines his <em>Echelon</em>:</p>

<p><br /></p>
<div><img style="height:35px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/echelon.png" alt="Echelon." /></div>
<p><br /></p>

<p>which can be transformed into</p>

<p><br /></p>
<div><img style="height:40px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/echelon-transformation.png" alt="Echelon transformation." /></div>
<p><br /></p>

<p>thus gives us the re-entry</p>

<p><br /></p>
<div><img style="height:40px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/echelon-equation.png" alt="Echelon equation." /></div>
<p><br /></p>

<p>or</p>

<p><br /></p>
<div><img style="height:43px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/echelon-equation-2.png" alt="Echelon equation." /></div>
<p><br /></p>

<h2 id="final-words">Final Words</h2>

<p>In the act of programming there is no problem of using expressions such as</p>

<p>\begin{equation}
i \leftarrow i + 1 \quad  \text{ or } \quad a  \leftarrow f(a)
\end{equation}</p>

<p>but in mathematics—at least since Plato—we assume some sort of eternity.
Of course, we can translate between the static world of “normal” mathematics and Spencer-Brown’s dynamic viewpoint, but it is a different viewpoint which might influence how we observe our environment.
It is like in physics where multiple theories are equivalent but start from very different viewpoints.
It starts by differentiation which gets reintroduced into the system which is constructed by this very same differentiation.
To generate new numbers, such as irrational or transfinite numbers, Spencer-Brown proposes not to use the limit but the whole infinite process that defines such limit.</p>

<p>Spencer-Brown believed that to be able to master the transition to new signs, something is necessary for which the previous signs are not sufficient.
To be able to close this gap; to make this leap successfully; to resolve paradoxes; a specific language of one’s own is necessary. 
According to Spencer-Brown this step is accomplished by <strong>thinking</strong> which provides us with its specific imaginations, playful freedom and contradictions.</p>

<p>It is surprising that Spencer-Brown’s <em>Laws of Form</em> plays no role in computer science even though it fits quite neatly in the perspective of programs, processes and computation.</p>

<h2 id="references">References</h2>

<ol class="bibliography"><li><span id="brown:1969">Spencer-Brown, G. (1969). <i>Laws of Form</i>. London: Allen and Unwin.</span></li>
<li><span id="cull:1979">Cull, P., &amp; Frank, W. (1979). flaws of form. <i>International Journal of General Systems</i>, <i>5</i>(4), 201–211. https://doi.org/10.1080/03081077908547450</span></li>
<li><span id="maturana:1987">Maturana, H. R., &amp; Varela, F. J. (1987). <i>The Tree of Knowledge</i>. Shambhala.</span></li></ol>]]></content><author><name>Benedikt Zönnchen</name><email>benedikt.zoennchen@web.de</email></author><category term="Logic" /><category term="Social Systems Theory" /><summary type="html"><![CDATA[In my last blog post, I discussed Niklas Luhmann’s Social Systems Theory and I emphasized that his theory is based on differentiation and seemingly paradox relations. My general understanding of Luhmann’s radical constructivism in the most reductive sense is that there are no unified and independent objects—there are only differences. What an observer can identify as an object is a differentiation of a system and its environment, a foreground and its background, an interior and the external. However, any observer is itself a distinction between system and environment. I want to further investigate this idea by looking into Luhmann’s inspiration—his muse so to say. So let me examine the logic of Georg Spencer-Brown presented in his work Laws of Form (Spencer-Brown, 1969).]]></summary></entry><entry><title type="html">Why Machines (Probably) Do Not Think</title><link href="https://bzoennchen.github.io/Pages/2023/10/07/system-theory-and-ai.html" rel="alternate" type="text/html" title="Why Machines (Probably) Do Not Think" /><published>2023-10-07T00:00:00+02:00</published><updated>2023-10-07T00:00:00+02:00</updated><id>https://bzoennchen.github.io/Pages/2023/10/07/system-theory-and-ai</id><content type="html" xml:base="https://bzoennchen.github.io/Pages/2023/10/07/system-theory-and-ai.html"><![CDATA[<p>Generative AI, especially ChatGPT, brought artificial intelligence into the public sphere and sparked a lot of highly speculative claims about <em>machine intelligence</em>.
I’m open for discussions and unafraid of confronting uncomfortable truths. 
Indeed, our imagination and fearless thinking should pave the way for new possibilities. 
Dreams and speculations are valuable, as long as they’re presented as such. 
However, I find it concerning when public figures speak with undue certainty, particularly when making anthropological comparisons between humans and machines.</p>

<h2 id="the-fall-from-human-greatness">The Fall from Human Greatness</h2>

<p>What really hit me was watching Doug Hofstadter expressing his despair about the <em>eclipse of humanity</em>.
As a student, I had been influenced by his renowned book <em>Gödel, Escher, Bach: an Eternal Golden Braid</em>, often referred to as <em>GEB</em> <a class="citation" href="#hofstadter:1979">(Hofstadter, 1979)</a>.
The book delves into how cognition emerges from underlying neurological processes.
Let’s examine Hofstadter’s comments on the state of AI which I gathered from an interview:</p>

<blockquote>
  <p>I never imagined that computer systems would rival or even surpass human intelligence.
It seemed like a goal so far away.
My entire belief system was shaken; it’s a truly traumatic experience when some of your most fundamental beliefs about the world start to collapse.
Particularly, the idea that human beings are soon going to be eclipsed. 
It felt as if not only my belief system was collapsing, but also as if the entire human race was about to be eclipsed and left in the dust soon. 
The accelerating progress has been so unexpected, it stirs a certain kind of terror of an impending tsunami that’s going to catch all of humanity off guard.
It’s unclear whether this signifies the end of humanity, in the sense that the systems we created could destroy us, but it’s certainly conceivable. 
If not, it relegates humanity to a relatively minor phenomenon compared to something else that is far more intelligent and will eventually become as incomprehensible to us as we are to cockroaches.
I find that terrifying. 
I hate it! I think about it almost every single day.
And it overwhelms and depresses me in ways I haven’t experienced in a very long time. […]
It makes me feel diminished; it makes me feel, in some sense, like a very imperfect, flawed structure. 
Compared with these computational systems which have a million or billion times more knowledge than I have, and are a billion times faster, it makes me feel extremely inferior.
It almost feels like we deserve to be eclipsed. 
Unbeknownst to us, all we humans are soon going to be eclipsed and rightly so, because we are so imperfect and fallible. – Doug Hofstadter</p>
</blockquote>

<p>He passionately conveys a sentiment many intuitively feel: The essence of humanism is under siege.
We are fallen from greatness.
Our unique skills are being surpassed, leading to concerns about our relevance.</p>

<p>I perceive Hofstadter’s view as human-centric, stemming from a longstanding tradition where humans are seen as central figures, akin to being God’s creation.
This view encompasses our confidence in determining our fate; the idea of an individual separate from its environment; humans domination over objects; a hierarchy with humans at the pinnacle; the notions of free will, and rational, independent beings arriving at a consensus in public discourse.
Paradoxically, I will try to attack this human-centric view to save myself from his despair.</p>

<p>One could argue that machines are not intelligent since they are just doing statistics by computing some high-dimensional probability distribution <a class="citation" href="#bender:2021">(Bender et al., 2021)</a> and that there is a difference between language processing and language understanding <a class="citation" href="#bender:2020">(Bender &amp; Koller, 2020)</a>. 
However, these arguments seem not convincing for many people.
There always looms a counter argument: Maybe humans do the same?</p>

<p>Instead of going down the technical rabbit hole, I will try to approach Hofstadter’s comments from a distinct and somewhat radical angle by using my <em>observation</em> of Luhmann’s <em>social system theory</em> which builds on <em>radical constructivism</em>.
Along the way, I will not only discuss <em>machine intelligence</em> but also touch on the relation between machines, society and human beings.</p>

<p>I’m not asserting this as the absolute truth but rather as an interesting <em>story</em> that might be useful in some aspects.
In fact, when I started reading Luhmann, I hated it!
It made so much sense but also felt cruel, cold and depressing.
However, like reading Nietzsche’s <em>On the Genealogy of Morality</em>, it has the potential to destroy some deep seeded belief only to bring something new and exciting into existence.
While this theory is nothing more than a theory, it does challenge the confidence behind many claims, including those regarding <em>machine intelligence</em>.</p>

<h2 id="niklas-luhmann">Niklas Luhmann</h2>

<p>Niklas Luhmann (1927-1998) was a largely self-taught sociologist. 
Like postmodern thinkers, he believed that pursuing metaphysics was no longer productive, as there are no ultimate grand narratives that can explain everything. 
Rather than delve into metaphysics, he meticulously developed a comprehensive theory of modern society—–a supertheory that even encompassed itself and its creator.</p>

<p>Luhmann was an avid reader and writer, and he wasn’t hesitant to incorporate valuable concepts from fields like mathematics (Spencer-Brown’s), cybernetics (Wiener and others), and biology (Maturana and Varela).
To encapsulate everything, he employed a high level abstraction and a technical terminology, which can make his writings appear dry, cold and dense. 
Because his work is primarily descriptive—–explaining things as they are and exploring potential reasons for their status—some categorize him as conservative. 
However, I perceive him as an incredibly well-read, sensitive, and discerning observer who wanted a new theory that can help us to transit into a new form of stability, which is a rather progressive attitude.</p>

<p>Even though Luhmann tried to keep a large distance to philosophy, he was well read in it and certainly influenced by it.
In his introductory book <em>From Souls to Systems</em>, Hans-Georg Moeller <a class="citation" href="#moeller:2006">(Möller, 2006)</a> highlights that Luhmann’s work is influenced by several philosophical giants:</p>

<ul>
  <li><strong>Kant</strong>: Luhmann shifts Kant’s focus on cognition to a constructivist perspective <a class="citation" href="#luhmann:1988">(Luhmann, 1988)</a>.</li>
  <li><strong>Hegel</strong>: Luhmann transitions from Hegel’s ideas of unity and dialectic to concepts of multiplicity and from identity to difference. He argues against any essential unity of systems and any general type of cognition, such as Hegel’s spirit.</li>
  <li><strong>Marx</strong>: Luhmann borrows Marx’s view that society isn’t just a byproduct of spirituality, but he disagrees with the idea of one foundational system. Marx focus on economy is too much of a simplification.</li>
  <li><strong>Husserl</strong>: Luhmann adapts Husserl’s work towards constructivism and incorporates many of his terms and ideas.</li>
  <li><strong>Habermas</strong>: Luhmann disputes Habermas’s mission of completing the enlightenment.</li>
  <li><strong>Postmodern thinkers</strong>: Luhmann draws from Deleuze’s radical differentiation, Derrida’s deconstruction, and Lyotard’s rejection of overarching narratives.</li>
</ul>

<p>With respect to his media theory, Luhmann is quite close to the French philosopher Baudrillard but far less dramatic.
While Baudrillard tends to express himself in dramatic metaphors and focuses solely on the media, Luhmann presents a supertheory of society where the mass media is only one of many systems, all administered by their respective codes.
While Baudrillard’s texts are almost poetic, reading Luhmann can cause boredom.</p>

<p>Luhmann believed that the distinction between <em>modernity</em> and <em>postmodernity</em> is largely semantic.
He argued that the last significant structural shift in society occurred in Europe between the sixteenth and eighteenth centuries, transitioning from stratified to functional differentiation. 
To Luhmann, labeling a functionally differentiated society as either ‘modern’ or ‘postmodern’ is inconsequential.</p>

<p>Although his theory can be unsettling, Luhmann was optimistic about the future.
He agreed side-by-side with the postmodern assertion that traditional philosophy had reached its end.
However, he saw this as an opportunity for a rejuvenated, coherent self-description of society and a fresh theoretical framework for a new societal era:</p>

<blockquote>
  <p>Is this, after all, a postmodern theory?
Maybe, but then the adherents of postmodern conceptions will finally know what they are talking about.
The deconstruction of our metaphysical tradition pursued by Nietzsche, Heidegger, and Derrida can be seen as a part of a much larger movement that looses the binding force of tradition and <strong>replaces unity with difference</strong>.
The deconstruction of the ontological presupposition of metaphysics uproots our historical semantics in a most radical way.
This seems to correspond to what I have called the catastrophe of modernity, the transition of one form of stability to another. – <a class="citation" href="#luhmann:1993">(Luhmann, 1993; Luhmann, 2000)</a></p>
</blockquote>

<h2 id="social-systems-theory">Social Systems Theory</h2>

<p>So let me try to give you my incomplete and surface level understanding of his theory:</p>

<p>Luhmann recognised the particular complexity that human beings present for social analysis because they are the bearers of three <em>autopoietic systems</em>: systems of life (cells, brains, organisms), systems of consciousness (mind), and systems of communication (social systems).
As a sociologist he acknowledges but leaves aside the biological systems of human beings and instead focuses on the interactive relationship between their consciousness or psychic system and the social systems with which they interact.
All psychic systems (minds) are in the environment of social systems and vice versa.</p>

<p>He famously argued that communication between psychic systems happens not between (whole) persons or individuals.
This seems counterintuitive but if we spent a little more thought into his claim and clarify some terminology, it makes sense.
Hans-Georg Moeller put it the following way:</p>

<blockquote>
  <p>You cannot communicate with me with your mind or brain, you will have to perform another communicative operation such as writing or speaking. – <a class="citation" href="#moeller:2006">(Möller, 2006)</a></p>
</blockquote>

<p>Luhmann described the mind (the psychic system) as well as social systems as <em>operational closed, structurally coupled, autopoietic systems</em>.
That are a bunch of important terms right away which require some explanation.</p>

<p>In Luhmann’s view, a system is defined by its differentiation with its environment—<strong>differentiation</strong> plays one of the most important roles in his work.
This differentiation is established and obtained through the operations of the system.
The system differentiates itself from its environment thus it defines itself.
In other words, the system creates its own functions by its operations (self-creation and self-preservation).
Thinking leads to more thinking and perception leads to more perception.
The economy creates itself by doing economics, the mass media creates itself by its operation of differentiating between information and non-information.</p>

<p>Another <em>theme</em> in Luhmann’s writing is the use of paradoxes which is certainly inspired by Spencer-Brown’s logic presented in <em>Laws of Form</em> <a class="citation" href="#brown:1969">(Spencer-Brown, 1969)</a>.
Inspired by imaginary numbers, Spencer-Brown introduced imaginary truth values which are paradoxically in space but the paradox is resolved in time.
Therefore, a paradoxical system becomes generative in time.
One prime example is our mind which is capable of self-observation (<em>reentry</em>).
Paradoxically, the observation is part of what is observed.
However, in time this gets resolved.
While I am observing myself (second-order observation) I cannot observe my environment (first-order observation).
I can switch to the first-order observation but then I lose track of my observation of myself.
In a sense, this back and forth observation and self-observation—this paradox—generates myself.</p>

<p>A psychic or social system is <em>operational closed</em> because its operations can not leave the system.
Mental operations such as thoughts and emotions cannot leave the mind.
An economic transaction, e.g. paying for goods, cannot ‘leave’ the economy.
No mind can interfere with the operations of another mind.
One cannot continue someone else’s mental activities by thinking or feeling for him or her.
It is also impossible to immediately think what someone else is thinking.</p>

<p>However, systems can observe their environment and act on their terms.
We can hear what others say, see what they express and read what they have written.
Our mind can think about it (using its operations) and we can answer, i.e. communication happens.
We can also see pain or joy on other’s faces, but we cannot literally think or feel what they do.
The economy observes politics, the media, science and gets irritated.
How it will adapt is up to itself and its operations.</p>

<p>Other than <em>allopoietic systems</em>, which produce something other than the system itself, <em>autopoietic systems</em> reproduce themselves.
They are more dynamic than allopoietic systems because they deal with an excess of complicating noise from their environment (too much information that cannot be processed) by changing their structure (increasing internal complexity) to allow in more communications: they have a built-in learning capacity.
In contrast, allopoietic systems theory leads the observer to seek constancy and stability in system functioning because they are intrinsically conservative.</p>

<p>Social systems, like the media, become so efficient because they ‘feed’ the outside into their ‘body’.
A crisis, like a natural disaster or a war, feed the autopoiesis of the media.
It can report on the event and discuss different opinions on the matter.
Strictly speaking, the ‘goal’ of the media is not really to inform or to persuade but to continue its own self-production.</p>

<blockquote>
  <p>It is impossible to understand the reality of the mass media if you assume it is their job to provide correct information on the world and then assess how they fail, distort reality, and manipulate opinion—as if they could do otherwise. – Niklas Luhmann</p>
</blockquote>

<p>Therefore, attention is key.
Informing people or persuading might help to ‘get enough food’ but it is not the media’s ‘goal’ or ‘will’.
The same goes for the economy which trys to commodify everything to further commodify things.
Politics politicizes anything and sciences produces papers with ‘facts’ to get more funding for more papers with ‘facts’.</p>

<p>Even though Luhmann’s termology is close to the terminology of computer science, we have to be careful.
It is more helpful to think of these systems as interdependent organisms feeding on each other and equipped with the will to live than to think of hierarchical or well-structured computer or network systems.</p>

<p>Aside from being <em>autopoietic</em>, social and psychic systems are also <em>symbiotic</em>, that is, their <em>co-evolution</em> is <em>interdependent</em>.
Just as the trees in the forest need water and animals to survive, politics needs money from the economy, attention from the media, and ‘facts’ from science.
Media needs politics or science to produce news, and money to operate.
The economy uses media, politics, and science to make profits.
Science ‘sells’ truth to the economy, politics, and the media.
Academia needs money, attention and power.</p>

<p>Luhmann insists in putting human beings in the environment of social systems and not inside them.
In other words, social systems do not consist of humans but of communication!
This is sometimes seen as an anti-humanistic tendency which is framed negatively.
But one might argue that human beings are better off if their processes are not determined by society.</p>

<p>Luhmann’s theory provokes an <em>amoral</em> view on the state of affairs but it also gives power to the object (systems/processes) thus attacks the domination of objects by subjects.
There are no evil people doing or planing insidious things, instead systems (objects) act on behalf their <em>systemic rational</em> by making sense of their environment on their terms.
What we often identify as hypocratical in a person’s action is a mixture of the operations of <strong>different</strong> systems or the communication between systems.
As a reminder, the person is not part of the system.
Individuals, or better psychic systems, are a necessary condition that social systems can exist (like air has to exist to hear sound) but they belong to the environment of the social system (they do not produce the sound).
If a politician acts immoral and accepts a lot of money for his party to give a certain company an advantage over its competitors, the politician is the mere medium through which the economy communicates with politics.
If a politician of the Green Party goes on vacation by plane and, at the same time, speaks out against air traffic, two different systems are operating: the family and politics.
And the operation of the first does not interfere with the operations of the second.
However, the media make news out of this contradicting behaviour which will irritate politics and probably the family life of the politician.
What the media observes as ‘corruption’ happens if system boundaries are crossed.
It is ok to buy talented football players, but it is not ok to buy goals, i.e. it is not ok that the economy directly operates within the system of a football game.</p>

<p>Psychic and social systems are <em>operationally closed</em> but <em>cognitively open</em>.
They have clear boundaries demacrating them from other systems.
They reproduce themselves by adapting and learning how to cope with external noise by only selecting communication which the system can actively and creatively interpret and <em>understand</em> or make sense of.
Psychic and social systems reduce complexity of their environment through recourse to meaning.
These systems increase inner complexity to deal with the complexity in their environment.</p>

<p>The boundaries of these systems are not defined physically, but by the border of what is meaningful and what is not.
Consequently, each system has its own <em>systemic rationality</em> and view of the world—there is always a blind spot.
If I give a cashier money it is assumed that I paid for something.
This follows from the systemic rationality of the economic system.
It deals with money but it cannot, for example, deal with love or passion which are part of the <em>systemic rationality</em> of relationships.
The cashier does not suspect me that I show him my love with this gesture and if I do, this act is not an operation of the economic system.</p>

<p>The <em>functional differentiation</em> of each system makes it so that only parts of a person is acknowledged by the system.
There is no indivudual—indivisible being—in a system.
The health system understands a person as a patient.
The legal system understands a person as a potential criminal, victim or witness.
In that sense, Marx’s <em>alienation</em> is not limited to the economy.
This differentiation makes systems extremly efficient <strong>with respect to their function</strong>.
But what is ‘good’ for one system is not necessarily ‘good’ for the other system.</p>

<p>Luhmann thinks that this differentiation (Ausdifferenzierung) is a feature of modern society, i.e. it is historical and is an ongoing process.
One example might be the creation of new subjects to study.
Instead of studying computer science, students can enroll in scientific computing, data science, game engineering, information engineering, and more.
One can say computer science is further differentiated.
At the same time, we acknowledge problems stemming from this differentiation and try to find ways to look at problems and society more holistically.
Marignal note: If we follow Luhmann’s theory and we want efficiency (with respect to a systemic rational) we find a strong argument to avoid introducing an interdisciplinary subject such as bioinformatics by simply combining biology with informatics.</p>

<p>The functional differentiation of systems, its effects and our gut reaction to it is nicely depicted in the movie <em>Don’t Loop Up</em>.
What the movie does well is showing us that society consists of functional differientiated systems that follow their own <em>systemic rational</em>.
The main message of the movie is that scientists, who discover a meteor, are unable to communicate this truth to the world.
The movie shows mostly four social systems: politics, media, economy, and science.
It shows how each of these systems functions differently while still being <em>structurally coupled</em> with one another via a <strong>shared medium</strong> (language), as explained above.
But inspite of being coupled, or because of it, they cannot act unitedly.</p>

<p>The effectiveness of functional differentiation to deal with complexity comes at a cost: <em>anarchy</em>, that is, there is no controlling system or governing system—no single rationality that is in charge.
From a reductionist standpoint, the actors of the movie seem completely irrational.
By seeing modern society through the eyes of controllable cause and effect chains one can only come to the conclusion that our society is dysfunctional or worse: immoral.
In the movie, only the scientist, who also represent the perspective of the audience, seem to do the right thing.
But from a systemic view the actions of all actors make sense.</p>

<p>In the end, the narrative of the movie is however a contradiction to the systemic view.
The movie suggest that there is some sort of scientific technological solution that can be used if everyone is thinking and acting properly—if only the government takes proper control, and the media informs everyone correctly then the meteor can simple be nuked.
The movie suggest that there can be some sort of rational self-control if only we would be <em>enlightened</em> enough.
In a sense, it is not much better than the movie <em>Idiocracy</em>.
The dream is that enlightened science can control nature through rational technology, enlightened politics can control society through rational self-government of the people, and enlightened media disseminates knowledge and makes everyone an informed and rational citizen.
Therefore, the movie presents an individualistic solution.
Big tech is greedy, politicians are stupid and hypocritical, and scientists are incapable of being live on TV.
If we fix those issues, we are fine.
If only we ‘look up’ (individually), we will be enlightened and stop being stupid and ignorant and we will solve all our modern problems.
The problem becomes a moral problem of personal responsibility thus it becomes polarizing.
From a systems theory perspective, this individualistic solution is not (or no longer) possible.
Systems function according to their functional differentiation on their terms and individuals are part of their environment.</p>

<p>However, <em>anarchy</em> does not imply the absence of strata or the presence of equality.
It’s evident that systems like the economy can create significant <em>differences</em> between the rich and the poor. 
Luhmann recognized that modern society inherently produces many differences, including those we might disapprove of. 
Every system does this, not just the economy.
However, contrary to a Marxist perspective, Luhmann believed that these societal disparities result more from the operations of multiple systems than from the stratum or class into which people are born. 
That said, an individual’s socioeconomic background, such as whether their parents are rich or poor, does matter. 
For example, in the education system, the ability of one’s parents to afford tuition at prestigious institutions like Stanford plays a significant role due to the interconnectedness of the economic and educational systems.
However, to genuinely understand how systems like education or academia function, one must grasp their unique <em>differentiators</em>. 
The education system is defined by distinctions like good grades versus bad grades, while the academic system differentiates between peer-reviewed and non-peer-reviewed papers.
These differentiations are intrinsic to their respective systems and not solely based on economic factors.
According to Luhmann, while wealth can certainly influence educational outcomes, avoid legal troubles, or facilitate a scientific career, it’s overly simplistic to reduce all systemic distinctions to just ‘money’.
But again, of course it helps a lot if you have money if you want good grades, stay out of prison, or become a scientist.</p>

<p>Luhmann argues that all social systems operate on a binary code determined by their sphere of interest which structures their communication with other systems.
Communication with the legal system is organised by the code legal/illegal through the medium of law;
with the political system by the code government/opposition through the medium of legitimate power;
with the economy through money with the code pay/not pay;
with science by the code true/false through the medium of evidential truth;
with the mass media system by the code information/non-information through the medium of public opinion;
and with the welfare benefits system by the code eligible/not eligible through the medium of citizenship status.</p>

<p>Without their environment systems would cease to exist.
They are <em>structurally coupled</em> with one another.
For example, psychic systems are structurally coupled with social systems—without bodys and minds there is no political system, no economy, no relationship and no family.
Without the economy, the political system would collapse.
However, there is <strong>no causal relationship</strong> between the two;
society does not cause consciousness to occur, neither do people consciously create and manage society.
The relationship between the two is rather one of constant <em>irritation</em> (which may be also translated to <em>confusion</em>) with the one reacting to the other, but always on its own terms.
The dynamics are non-linear and tend to be <em>chaotic</em>.</p>

<p>If the political system enact a new law to steer the economy, it can only try to do so via irritation.
How the economy will react is not up to politics.
If climate activists glue themselves to the ground to generate awareness, they may achive their goal or they may not.
What happens is quite difficult to predict.
How does the media report on the issue if its rational is to further differentiate between information and non-information?
How will politics react based on the assumption that it ‘wants’ to make more politics?
From this point of view, it is hard to see how ‘we’ can ‘make’ cooperations (or individuals) sustainable by referring to morals and virtues.
Moral outbursts and frustrations about the destruction of our livelihood are completely understandable (for my psychic system) but how they irritate the different systems is quite uncertain.</p>

<h2 id="artificial-communication">Artificial Communication</h2>

<p>Niklas Luhmann’s concept of communication offers a useful framework for sidestepping (at least for a moment) the ongoing debate about <em>machine intelligence</em>.
While I personally do not ascribe human-like thinking or understanding to machines, I argue that this does not preclude their participation in communication processes.
Therefore, I agree with <a class="citation" href="#esposito:2022">(Esposito, 2022)</a>.</p>

<p>I think Esposito’s term <em>artificial communication</em> is a very useful contribution to make the discussion of <em>AI</em> (especially of machine learning) more reasonable.
Note that she was a student of Niklas Luhmann.
In an interview she explains why she came up with the term:</p>

<blockquote>
  <p>These algorithms became more and more opaque—not understandable for the users—the idea spread that the activities of machines are not trying to be intelligent; that they are not trying to reproduce, in an artificial way, the process of human thought; they are doing something different.
This is rarely said explicetly but one can find it in many different contexts.
And if we switch away from the idea of intelligence, what can we refer to?
Do we have another metaphor that would fit better into the current situation? – Elena Esposito</p>
</blockquote>

<p>Why do people think that ChatGPT is intelligent?
Well, if we interact with machines, we get information we would not get otherwise and the information cannot be attributed to any human being.
The machine processes the data and produces some information which not only did not exist before but is also a sort of reaction to our request.
The machine does exactly what we do when we communicate with a human being.
We ask something and we get some information we did not had before.
Importantly, this information is <strong>contingent</strong> (the response could be different).</p>

<p>Esposito clarifies that, as a sociologist, it is understandable that we think of these machines as <em>artificial intelligence</em> because we have been communicating with human beings for thousands of years and these beings were ‘intelligent’.
And because of this feature of being able to think, humans were able to produce something which allows us to get new information.
It is our prejudice that lead us to the conclusion that machines are so similar to human beings, i.e., psychic systems.
Therefore, to follow Esposito’s proposal, a more interesting and probably healthier question to ask (also for us computer scientist) is:</p>

<blockquote>
  <p>Why are these machines able to communicate with us inspite of the absence of their intelligence?</p>
</blockquote>

<p>Esposito’s answer is that they are <em>parasitical</em>.
My understanding of her work is that machines make heavy use of <strong>second-order observation</strong>, i.e., the observation of an observer.
The starting point is some mental activity but not the machine’s activity.
The user produces <em>contingent</em> behaviour which the machine can process (or observe) to become itself <em>contingent</em>.</p>

<p>A modern example that might no longer be considered AI is Google’s search algorithm. 
Google’s success stems not from trying to evaluate or calculate the quality of a webpage directly.
They didn’t design an intelligent machine for that purpose. 
Instead, they leveraged <em>second-order observation</em>, essentially tapping into the collective intelligence of their users. 
Rather than determining the value of a webpage themselves, their algorithm observes how users interact with webpages. 
A webpage ranks high if users deem it valuable which can create a feedback loop because highly ranked pages are ranked highly.
The primary task becomes observing users’ observations.</p>

<p>Even if Esposito speaks of switching the metaphor, her work goes deeper.
The word <em>intelligence</em> has two different usage in language which are often confused.
On the one hand, we refer to the operational mode of the mind.
But we have almost no idea what this <em>intelligence</em> exaclty is.
On the other hand, we think of information processing.
Take for example the term ‘Central Intelligence Agency’.
Therefore, the metaphor of <em>artificial intelligence</em> is so problematic which makes it hard to theorize about <em>AI</em> and its impact which leads to these hyper-speculative predictions.
It’s like referring to airplanes as artificial birds.
Just as airplanes succeeded when engineers stopped trying to mimic birds, AI has advanced when researchers moved away from replicating human thought.
By using Luhmann’s theory of communication, we might clear the smoke and find more effective ways to talk about artificial intelligence.</p>

<p>Esposito argues that we—the preachers of machine learning—do not reproduce human intelligence but rather social communication.
Intelligence that emerges from conscious beings might not be needed or might even be an obstacle for the establishment of communication.
Artificial communication (coupling machines and psychic or social systems via language) can be more effective than intelligent communcation (coupling psychic and social systems via language) but it can not be intelligent (referring to the first use of the word).
In other words: That which makes society more intelligent might not necessarily be intelligent.
As described above, social systems have their own <em>systemic rational</em> and we might call them intelligent.</p>

<p>Systems theory is useful because it focusses on communication itself.
Again, Luhmann claims that humans do not communicate, only communication communicates.
Of course, similar to air, humans are a necessary condition for communication but, like air, they do not communicate themselves—we can only hear the ticking of a clock because the air does not tick.</p>

<p>Luhmann diverges from traditional sender-receiver models of communication, such as Shannon’s <a class="citation" href="#shannon:1948">(Shannon, 1948)</a> where the focus is on the transmission of information.
Instead, Luhmann conceptualizes communication as comprising three essential moments: <em>announcement</em>, <em>information</em>, and <em>understanding</em>.
Each component has its unique role in facilitating communication.
An announcement initiates the process.
Whether verbalized, written, or visualized, it serves as the catalyst that triggers communication.
Absence of an announcement, be it from a human or an algorithm, results in the absence of communication altogether.
This announcement must bear some form of informational value, imbuing the text, image, or utterance with meaning.
The final moment, understanding, underscores the necessity of a recipient comprehending the conveyed information.
The efficacy of communication is not solely predicated on accurate understanding, but rather on the act of understanding itself—even if what is understood is incorrect.
As Luhmann notes, understanding is often replete with misunderstandings, but the very act of engaging in a selection of understanding is vital.
Understanding is typically misunderstanding without understanding the ‘mis’ (similarily, misinformation is still information).</p>

<p>In summary, Luhmann’s perspective underscores that effective communication doesn’t necessarily require partners to achieve mutual understanding in the way their respective psychic systems might operate.
I think we can make the same observation in our day to day life.
Partners can perfectly live together even though their understanding is not mutual which, of course, can cause problems in relationships.
However, it can also be a useful feature.
If a third party observe a tense conversation between a couple, the content of the conversation might be quite ordinary but what is communicated can be a conflict within the relationship.
The couple understands the communication much better than the third party.
The conflict, however, is likly caused by the problem of different previous (mis-)understandings.
For the third party, it is like listening to some encrypted communication.
Whether executed by humans or algorithms, the value lies in the process and its constituent parts: announcement, information, and understanding.</p>

<p>In the context of artificial intelligence we can look at the communication of a person and a machine—of ChatGPT and Doug Hofstadter—and we can ask: Why is it so effective or attractive? 
But also: Why is it (probably) not the product of an intelligent thinking system but rather produced social intelligence?</p>

<p>With my shallow understanding of Luhmann’s theory, I imagine that the prerequisites for communication—whether artificial or otherwise—involve <strong>contingency</strong> and <strong>connectivity</strong>.
In social systems theory, the generation of information is not an isolated act; it is attributed to an interactive partner.
While traditionally this partner is human, in the realm of artificial communication, it can very well be a machine.
The focus should be on the nature of the interaction itself: does it exhibit the characteristics of a contingent, autonomous relationship?
And does this interaction spur further communication?</p>

<p>Traditional machines that produce unpredictable outcomes are usually considered faulty rather than creative or original.
Take a pocket calculator, for instance; its primary virtue lies in its predictability.
We do not regard it as a communicative entity because it operates as expected which is desirable.
The calculator is not contingent.
Conversely, when interacting with image-generating algorithms like Stable Diffusion or Midjourney the appeal, I argue, is precisely in the unpredictability of the results.
Chatting with a bot can be exciting preceisly because we do not know the output of the bot or, in general, the outcome of this interaction.
Of course, this does not mean a completely random output would have the same effect.
<em>Contingency</em> should not be confused with randomness or arbitrariness.
The information provided by the bot has to be understandable, in the sense that it can also be misunderstood (like any ‘good’ communication can be).</p>

<p>Despite thinking of this feature as a flaw, this ambiguity is desirable for communication.
Of course, not all possibilities of misunderstanding are desirable.
A chatbot that provides patients with medical or organizational information should give precise and unambiguous answers.
However, in this case, the bot is more like a tool than a real communication partner.
And here we land at an important distinction:
While many argue that irritation caused by <em>generative AI</em> is similar to the invention of photography, I think there is a difference.</p>

<blockquote>
  <p>Cameras do not communicate!</p>
</blockquote>

<p>This does not mean that generative AI can not act as a mere tool in the process of, for example, the production of images.
The more predictable the more tool-like generative AI are and the less communicative they become.
There outputs by themselves become less interesting but at the same time, they are more useful to realize a specific vision of the user or artist.</p>

<p>As Esposito noted: Viewed through the lens of Luhmann’s social system theory, the development of compelling communication partners presents a unique dilemma: 
The challenge lies in engineering machines that exhibit both creativity and control, balancing the production of unexpected outcomes with predictability.
This tension is especially relevant in the field of AI art.
In essence, the paradox that governs the programming of ‘intelligent’ algorithms is the pursuit of controlled unpredictability.</p>

<blockquote>
  <p>The ultimate objective is to achieve a controlled lack of control. – <a class="citation" href="#esposito:2022">(Esposito, 2022)</a></p>
</blockquote>

<p>From a philosophical standpoint, Luhmann transforms the <em>mind-body problem</em> into the <em>mind-communication problem</em>—communication defined by Luhmann as “the operation that society consists of” <a class="citation" href="#moeller:2006">(Möller, 2006)</a>.
If the mind does not communicate but is only in the environment of society (communication)—is merely involved—how does this all work?
Similarily, I think, the question of how <em>artificial communication</em> emerges even if machines are also only in the environment of communicating systems, is one of the most important question to ask if one wants to understand the current state of AI, society and where we are heading at.</p>

<p>Now, if one looks closely to Luhmann’s definition of communication, we find that it is not compatible with Esposito’s concept of <em>artificial communication</em> and she is aware of that.
We might think of communication being really picky, that is, it has a lot of requirenments to occur—there is a lot of <em>structural coupling</em> going on.</p>

<blockquote>
  <p>Communication is improbable. – Luhmann</p>
</blockquote>

<p>Let’s look at some requirements: the physical requirements like temperature and gravity at a certain level, water, air, but also a medium and, according to Luhmann, at least two consciousness entities. 
In a sense, the coupling of these two conscious entities is more strict.
It is an equal operation of the psychic system that has to be devoted to the actual operation of the social system.
That is why they coincide in this event at which communication happens.</p>

<blockquote>
  <p>Empirically, I propse [the concept] because what is going on in the interaction with algorithms is so close to communication that we have to try to find a way to extend [Luhmann’s] concept of communication to include what is going on—it is not exactly the same.
The technique of the communication is similar: production of information which irritates other systems but the algorithm itself is not thinking, is not producing any new communication.
It just sort of conveys something that can produce information somewhere else.
My background is Luhmanian but what I am proposing, without wanting to amend Luhmann, is something different from the standard case of communication  – Elena Esposito</p>
</blockquote>

<h2 id="profilicity-machines">Profilicity Machines</h2>

<p>The philosopher Hans-Georg Moeller adds an interesting point to the machine learning discourse.
He posits that algorithms nowadays are used for profile building—he coined the term <em>profilicity</em> as a new identity technology, which is different from previous modes of identity building, i.e., <em>sincerity</em> and <em>authenticy</em> <a class="citation" href="#moeller:2021">(Möller &amp; D’Ambrosio, 2021)</a>.</p>

<p>Following his thesis, people taking pictures, not (primarily) to preserve memories, but to curate a profile on Facebook, Instagram or LinkedIn.
In that sense, I too build my own profile by writing this text and by curating a personal website, a GitHub repository, and many more profiles.
AI helps us to evaluate our profile(s) within McLuhan’s <em>Global Village</em> <a class="citation" href="#mcluhan:1992">(McLuhan, 1992)</a>.
It makes it possible to get feedback from our peers and to present this evaluation back to the village.
It enables <em>second-order observation</em> which reduces complexity.</p>

<p>Moeller admits that he is—as many of us coming from an age where authenticy was the primary technology to build identity—annoyed by this picture frenzy.
But he stays true to Luhmann and refrains from judging.
He trys not to moralize this phenomenon or classify it as being ‘worse’ or ‘better’ because, in his eye, authenticy was never real in the first place.
Like the other forms of identity building, profilicity comes with its own problems.
Each mode brings its own set of challenges.</p>

<p>He intriguingly describes profile creation as <em>genuinely pretending</em>.
Observing younger generations, this resonates.
Their digital avatars often exude a <em>postmodern irony</em>; they knowingly embrace its constructed nature.
They are fully aware that it is all ‘fake’.
Contrarily, older generations may need reminding that these online images are meticulously curated and often manipulated.
Advising younger folks about the ‘deceptions’ of online portrayals might seem naive.
Their approach is more playful, even inventive, using multiple layers of meta-references to distance themselves from reality as far away as possible.</p>

<p>As Moeller notes, the real tension might arise from the mismatched expectations of older authority figures. 
We—and I include myself here—expect authenticity while most parts of the world of young people operate in the mode of profilicity.
This leads to a contradiction and, because it is about identity building, this contradiction might be psychologically problematic.
Misaligned expectations can cloud the path to self-realization.
Hence, while it’s tempting to solely blame social media for rising mental health issues, the underlying causes, as Luhmann would argue, are multifaceted.</p>

<h2 id="the-revenge-of-objects">The Revenge of Objects</h2>

<p>With the description of society handed over by systems theory, I might have lured you, the reader, into an even greater despair.
My assertion is that we don’t necessarily need AI to challenge Hofstadter’s vision of human greatness; 
the <em>deconstruction</em> might already be underway.
Luhmann’s system theory goes against the honorable belief of Hofstadter which is also expressed by figures like David Graeber or Noam Chomsky.</p>

<blockquote>
  <p>The ultimate, hidden truth of the world is that it is something that we make, and could just as easily make differently. – David Graeber</p>
</blockquote>

<p>I really like the sentiment expressed in the quote.
I want it to be true and to work!
And I admire personalities that keep it alive.</p>

<p>However, objects seem to regain agency and power over us.
When certain philosophers discuss subjects and objects, they often reference commonplace items like chairs and desks.
For instance, a chair might seem like a benign example. 
Here the case seems trivial: Of course a chair has no agency!
We make chairs to sit on them.
We dominate chairs.
They are completely in our control.</p>

<p>But we do not have to look further than <em>Heidegger’s hammer</em> <a class="citation" href="#heidegger:1927">(Heidegger, 1927)</a> to see that things can get tricky very quickly.
Heidegger proposes that before we ponder the essence of a hammer, we use it. 
Before we question its existence, we recognize its utility.
The hammer, in this context, prompts us to act—it, indirectly, has agency.</p>

<p>Or consider more potent examples like opioids, smartphones, the internet, algorithms, images or even movies.
These items influence our behavior, decisions, and perceptions.
The inception of video technology, for example, started with the simple goal of determining if a galloping horse ever had all its hooves off the ground simultaneously.
Now reflect on the vast implications and transformations that this technology has since undergone.
Did we solely shape these inventions, or should we attribute some credit to the inventions themselves?</p>

<p>For Baudrillard there is an uninterrupted production of positivity that has terrifying consequences.
Applying systems theory terminology, he speaks of <em>runaway positive feedback loops</em>.</p>

<blockquote>
  <p>Any structure that hunts down, expels or exorcizes its negative elements risks a catastrophe caused by a thoroughgoing backlash, just as any organism that hunts down and eliminates its germs, bacteria, parasites, or other biological antagonists risks metastasis and cancer—in other words, it is threatened by a voracious posivity of its own cells, or, in the viral context, by the prospect of being devoured by its own antibodies. – <a class="citation" href="#baudrillard:1990">(Baudrillard, 1990)</a></p>
</blockquote>

<p>In other words, runaway positive feedback loops that have no negative, will eventually cause a catastrophe.
A simple technical example of such a feedback is a microphone that picks up the amplified sound output of loudspeakers in the same circuit, then howling and screeching sounds of audio feedback.</p>

<p>Baudrillard thought that the production of images is such a feedback loop.
They are <em>out of control</em> and take over, for example, free democratic politics.
Although we can take a moral stance against certain imagery, such as pornographic content, Baudrillard suggests that positive feedback will ultimately subvert any moral code.
Trump, as an example, may be critiqued from a moral perspective, but because he’s a potent subject for image production, the media engages with him regardless of whether they criticize or praise him.
This cycle of image production can’t be halted simply by creating more images.</p>

<p>Luhmann may have a less bleak view. 
For instance, laws that regulate AI-generated images, can prompt changes in a system’s behavior (indirectly via irritation).
However, reining in runaway positive feedback is challenging, as issues can escalate exponentially.</p>

<p>It might sound unconventional, but to foster our understanding of our world it could be beneficial to acknowledge external influences like reality TV, staged political photos, or conspiratorial content as agents.
Not only in a sense that they affect us but that they have an inner life; a will of their own so to say.
It could be valuable to treat objects like oil with the same reverence and respect as ancient civilizations treated strom and thunder.
Isn’t it the case that in our times oil is more powerful than any ancient god ever was?</p>

<p>Predictive machines can sometimes inadvertently create self-fulfilling prophecies. 
For example, I might be more inclined to buy items from Amazon that appear at the top of a list because of their prominent placement. 
These items are ranked by an algorithm aiming to maximize Amazon’s profits, predicting which items I’m most likely to purchase.
Since I’m inclined to buy items higher up on the list, the algorithm’s prediction is validated, influencing its future predictions and creating a positive feedback loop which consists of me and the algorithm.
Because I rely on algorithms to find items to buy, I believe it’s fair to say that they have influence over me and a sort of agency of their own.</p>

<p>If we don’t attribute agency to objects, then the explanation for the failures of climate agreements likely rests on a dysfunction of the system or on individual failures, e.g. ‘corrupt’ or incompetent politicians.
The idea that our modern society renders us freer and more independent is misconstrued.
A better way to put it is: We are (more) <em>out of control</em>.
While we engage in broader dialogues, express ourselves diversely, and view the world through varied lenses, increasing complexity often amplifies dependency and chaos.
There is more differentiation going on but this does not mean that we are less dependent.
Instead systems are less controlable.
The age-old dynamic of subjects dominating objects could very well be shifting, placing objects alongside us in terms of influence.</p>

<h2 id="artificial-systems">Artificial Systems?</h2>

<p>Systems theory doesn’t assert that society’s evolution is innate or that it will remain unchanged forever.
It simply aims to provide a thorough depiction of modern society. 
Yet, it’s worth exploring the connections between systems theory and artificial intelligence.
This exploration might hint at the concept of an <em>artificial system</em>—–essentially an AI that functions as a system, potentially approaching the capabilities of general artificial ‘intelligence’.</p>

<p>To venture a speculative idea, machines might eventually evolve into these <em>artificial systems</em>.
To thrive in a highly complex modern environment, they might need to exhibit traits similar to psychic and social systems: being <em>autopoietic</em>, <em>structurally coupled</em>, <em>operationally closed</em>, and <em>functionally differentiated</em>.
If we assume this to be true, what implications might it has?</p>

<p>While minds create themselves through mental operations, an <em>artificial system</em> would self-create and self-preserve through computation.
This does not mean that artificial systems have to be in control of their requirements to operate.
The opposite is the case.
The required hardware, programmers, and users of such systems would belong to their environment.
But, as Luhmann writes:</p>

<blockquote>
  <p>The observing and describing itself (cognition), however, always has to be an operation that is capable of autopoiesis, i.e., of the performing of life or actual consciousness or communication, because otherwise it could not reproduce the closure and difference of the cognizing system; it could not take place “in” the system. <a class="citation" href="#luhmann:1988">(Luhmann, 1988)</a></p>
</blockquote>

<p>Computation would probably drive more computation; 
the primary purpose being to continue its computational operations.</p>

<p>Due to their operational closure, a synthesis of thinking and computation—a <em>transhumanist</em> vision—seems unlikely.
Also, their structural coupling suggests they’d be interdependent on psychic systems and other social systems.
Natural language processing might be an important aspect of AI because it potentially makes this coupling possible via the <strong>shared medium of language</strong>.
This implies that the idea of AI overthrowing humanity is also improbable.
It is more likely that our mind, i.e. our thinking/feeling co-evolves with computation/text generation.
Machines will not think for us but they might irritate our thinking because we might need different skills to survive in an environment consisting of such machines.
However, it’s worth noting that psychic systems may not have control over these artificial systems.</p>

<p>The primary function of the psychic system, in Luhmann’s conceptualization, is the processing of meaning. 
Psychic systems generate thoughts, emotions, perceptions, and other mental phenomena. 
They observe, process information, and produce decisions. 
While social systems use communication as their primary medium, psychic systems use consciousness. 
Every individual has their own psychic system, which means their own consciousness and their own way of processing and understanding the world.</p>

<p>Artificial systems, on the other hand, may primarily focus on the production of predictions. 
They would observe and process abstract data and project potential futures. 
But what would ‘motivate’ them? 
To function or to survive, they have to compute, i.e., they have to process abstract data.
Consequently, their <em>systemic rational</em> might prioritize producing predictions that lead to more predictions, rather than producing the most accurate or useful predictions.
While we can see traces of this effect in recommendation algorithms that suggest polarizing content, attributing this behavior to <em>artificial systems</em> and not, or only partly, to the mass media system might be an overreach.</p>

<p>I must admit, I need a deeper understanding of systems theory to consider AI as potential systems in Luhmann’s view.
However, I believe it’s a valuable pursuit.</p>

<h2 id="difference-makes-the-difference">Difference Makes the Difference</h2>

<p>I agree with Hofstadter on the limitations of human greatness, albeit for different reasons and perspectives.
While discussing systems theory, I aimed to challenge his belief that AI diminishes this human greatness. 
I argued that this so-called greatness, or perhaps <em>rational control</em> that comes out of enlightenment, was an illusion from the outset.
Additionally, I’ve raised questions regarding the feasibility of ‘complete’ enlightenment as proposed by the modern project.</p>

<p>Simultaneously, I presented reasons:</p>

<ol>
  <li>for perceiving machines as intelligent, and</li>
  <li>for potentially being misled by this perception.</li>
</ol>

<p>The phenomenon of <em>artificial intelligence</em> is not only interesting because of impressive accomplishments but because it pushes discussions about the ‘human nature’ into many psychic and social systems.
This question leads immediatly to the question of <em>cognition</em>, <em>thinking/feeling</em>, and <em>consciousness</em> because these concepts are so central to the question of <em>machine intelligence</em>.
So, let me elaborate a little bit more on these concepts.</p>

<h3 id="the-mystery-of-consciousness">The Mystery of Consciousness</h3>

<p>First of all, debating if machines possess consciousness or human-like intelligence is premature without a clear understanding of intelligence and consciousness.
Perhaps ‘intelligence’ is just a collective term for several unexplained phenomena.
We do not know yet.</p>

<p>We can identify correlations between mental events and brain activities, implying there’s a physical aspect to consciousness—a sort of carrier.
However, as neuroscientist Giulio Tononi notes, this doesn’t necessarily mean we can pinpoint consciousness within the brain <a class="citation" href="#tononi:2015">(Tononi &amp; Koch, 2015)</a>.</p>

<blockquote>
  <p>[…] what I think I know about my body, about other people, dogs, trees, mountains, and stars, is inferential.
It is a reasonable inference, corroborated first by the beliefs of my fellow humans and then by the intersubjective methods of science.
Yet consciousness—the central fact of existence—still demands a rational explanation. – <a class="citation" href="#tononi:2015">(Tononi &amp; Koch, 2015)</a></p>
</blockquote>

<p>Still <em>physicalism</em>, the view that everything is physical and can be explained by physics, seems to be the natural position, especially in science.
However, this view can be challenged.
In the famous and rather short article <em>What Is It Like to Be a Bat?</em> <a class="citation" href="#nagel:1974">(Nagel, 1974)</a>, the philosopher Thomas Nagel trys to argue against it.
He explains that we can not know what <em>it is like</em> to be a bat.
He chooses bats because they are mammals like us and, presumably, conscious, yet their primary sensory experience—echolocation—is profoundly different from any human sensory experience.
In Luhmann’s words: their environment is very different from ours.
We can imagine flapping our arms like a bat or eating insects like a bat, but we cannot truly imagine what it is like to ‘experience’ the world primarily through echolocation.
Nagel suggests that even if we knew all the physical facts about a bat’s brain while it echolocates, we’d still be missing the subjective experience—the “what it’s like”–—of being a bat.</p>

<p>In the <em>philosopyh of mind</em> there is a broad spectrum of dealing with <em>the hard problem of consciousness</em> ranging from new forms of idealism (the essence of reality is consciousness), naive realism (we have direct awareness of objects as they really are), new realism (accepts that science is not systematically the ultimate measure of truth but realities are first given, not constructed) to panpsychismus (the mind or a mindlike aspect is a fundamental and ubiquitous feature of reality).
There is no agreement on the <em>hard problem of consciousness</em>.
Consequently, consciousness remains enigmatic, with little guidance on where to begin our inquiries.</p>

<h3 id="cognition-as-construction">Cognition as Construction</h3>

<p>Sometimes, to gain a fresh perspective on the world, one must take a completely opposite stance.
For millennia, philosophers have pondered how we can perceive something if we lack direct access to reality.
Luhmann flipped this idea on its head. 
He argued that it’s precisely because we don’t have direct access (due to system/environment distinction and operational closure) that we can perceive.
Cognition can only happen if it is not interrupted which requires operational closure.</p>

<blockquote>
  <p>The tradition of epistemological idealism was about the question of the unity within the difference of cognition and the real object.
The question was: how can cognition take notice of an object outside of itself?
Or: how can it realize that something exists independently of it while anything which it realizes already presupposes cognition and cannot be realized by cognition independently of cognition?
No matter if one preferred solutions of transcendental theory (Kant) or dialectics (Hegel), the problem was: how is cognition possible in spite of having no independent access to reality outside of it.
Radical constructivism, however, begins with the empirical assertion: Cognition is only possible because it has no access to the reality external to it.
A brain, for instance, can only produce information because it is coded indifferently in regard to its environment, i.e. it operates enclosed within the recursive network of its own operations.
Similarly one would have to say: Communication systems (social systems) are only able to produce information because the environment does not interrupt them.
And following all this, the same should be self-evident with respect to the classical “seat” (subject) of epistemology: to consciousness. – <a class="citation" href="#luhmann:1988">(Luhmann, 1988)</a></p>
</blockquote>

<p>I should emphasis and explain how radical this turn is.</p>

<p>Luhmann builds on Kant’s <em>Copernican Turn</em>.
He turned the question of ‘what is the general structure of the world’ into a question of the general structure of cognition: Under which conditions is cognition allowed to operate to conceive the world?
Kant tried to bridge the gap between Hume’s <em>empiricism</em> and German <em>idealism</em>.
He also accepted the skeptical argument that nothing could be realized by cognition independently of cognition.
Kant brought in <em>constructivism</em> into epistemology but he was, in Luhmann’s view, not radical enough because he hold on to some sort of relation to reality.
Luhmann assumes that the realization of reality is not a relating to reality, but that reality basically consists of its own realization and that the key is <em>differentiation</em>.
Therefore, Luhmann drops the idea of a reality that is independent of cognition, i.e. Kant’s <em>Ding an sich</em>.</p>

<p>According to Luhmann, cognition itself becomes a construction based on distinction: Reality emerges as cognition.
This sounds like Hegel’s <em>idealims</em>.
However, the emergence of reality does <strong>not</strong> mean that cognition is <em>ideal</em>—it does not rely on an essence in the form of consiousness.
Cognitive systems construct cognition which is based on the system/environment distinction.
There is no single rule how this can be done.
Cognition can operate materially in the form of biological life, mentally in the form of thoughts, or socially in the form of communication.</p>

<p>Moeller summerizes this shift from <em>Kantian idealism</em> to <em>radical constructivism</em> nicely <a class="citation" href="#moeller:2006">(Möller, 2006)</a>:</p>

<ol>
  <li>Cognition is not <em>per se</em> an act of consciousness. It can take on any operational mode.</li>
  <li>There is no a priori, transscendental structure of cognition; cognition constructs itself on the basis of <em>operational closure</em> and is an <em>empirical</em> process, which varies from system to system.</li>
  <li>No complete description of cognitive structures is possible because these structures are continuously evolving.</li>
  <li>Reality is not singular—there is not one specific reality, but a complex multiplicity of system/environment constellations.</li>
  <li>A description of reality is itself a contingent construction within a system/environment relation.</li>
</ol>

<p>In Luhmann, the subject/object distinction gets replaced with the system/environment distinction and the premise of a common world (the unity of system and environment) gets replaced by a theory of the observation of observing systems (<em>second-order cybernetics</em>).
However, getting to ‘the root’ of cognition—which would lead to some clues about <em>a reality</em>—seems impossible because any such investigations require distinction which is the operation of cognition.
Furthermore, there is no justification for assuming that any adaptation of cognition to reality is happening.
Confronted with this problem how can one start developing a theory of cognition?</p>

<p>In <em>Cognition as Construction</em>, Luhmann recognizes and somewhat addresses this issue by an assumption akin to Decartes and Husserl:</p>

<blockquote>
  <p>We assume that all cognizing systems are real systems within a real environment, or in other words: that they exist.
This is naive—as it is often objected.
But how should one begin if not naively?
A reflection on the beginning cannot be performed at the beginning, but only by the help of a theory that has already established sufficient complexity. – <a class="citation" href="#luhmann:1988">(Luhmann, 1988)</a></p>
</blockquote>

<p>Starting from this assumption, Luhmann reformulates Kant’s question of the possibility of cognition to the question of how systems can <strong>uncouple</strong> themselves from their environment.
For him, closure, i.e. uncoupling, is only possible by a systems’s production of its own operations and by its reproduction within the network of its recursive anticipations and resources.
Thus, cognition is manufactured and is a self-referential process.
It deals with an external world that remains unknown thus cognition has to come to see that it cannot see what it cannot see.
One could say that reality remains as an ineradicable <em>blind spot</em>.
While reality remains unkown, Luhmann speculates that there is some ground for the belief that if reality would be totally entropic, it could not enable any knowledge.
In other words, reality cannot be the object of the knowledge that it makes possible, it serves knowledge merely as a presupposition.
Knowledge can only know itself but it cannot know anything about what it constructs by way of the manipulation of distinctions.</p>

<p>All observing systems are cognitive systems.
They are operationally closed but cognitively open and they make sense of their environment as they experience it.
The mind makes sense in such a way that its construction is valid or functional.
It does not matter if it resembles reality.
In this way, making sense is like finding the right key to open a lock.
The key is useful and <em>functionional</em> but it does not tell us anything about the lock, expect that it fits the lock.
However, how we see the world depends on the cognitive capacities of our eyes and brains to produce images.
Different operations lead to further different operations.
For example, bats are another system/environment distinction than we are.
Therefore, their world most certainly ‘looks’ nothing like ours.</p>

<blockquote>
  <p>To recognize a table and say “This is a table”, I don’t need to have the letters T, A, B, L, E in my brain, nor does a tiny representation of a table (or even the “idea” of the table) need to exist inside me.
However, I do need a structure that calculates the various manifestations of a description for me. – <a class="citation" href="#foerster:1988">(Foerster, 1985)</a></p>
</blockquote>

<p>What our psychic system ‘sees’ does not have anything to do with reality.
In fact, the psychic system cannot ‘see’ reality with ‘its’ eyes because our eyes are in the environment of our psychic system.
Cognition is always a construction by an observing system and there is only irritation and no direct causal relation.</p>

<p>Now, one might ask: If every system constructs itself and reality emerges out of its cognition, how can we find any common ground?
Or, how did we end up believing in an universal and objective reality we actually can access?</p>

<p>Well, only because there is no objective reality does not mean that humans cannot ‘achieve’ social control via social systems.
Importantly, psychic and social systems share a common <em>medium</em>, that is, sense.
It is used in thought as well as in commuincation.
A thought or a gesture is a specific <em>form</em> (strict coupling) that the medium (loose coupling) of sense can take on.
Thus, thinking is like bringing sand (the medium) into a specific form; it can also be defined as a selection within a horizon of what is possible.
By thinking a specific thought, I do not think any other thought that is in my horizon of the thinkable.
Secondly, social and psychic systems are structurally coupled via language.
Language couples psychic and social systems. 
And these social systems make sense in accordance to their <em>systemic rationality</em> irritated by observing their environment.
In a sense, we create a second dimension of ‘reality’ by assuming, through the use of concepts, that our own constructions resemble those of others and by experiencing ourselves as part of a community by assuming and asserting that our own constructions largely correspond to those of others.
The experience of stability and continuity of one’s constructed reality depends not only on the system’s first-order observation but also on the confirmation of this observation by other observers (second-order observation).</p>

<p>From a radical constructivist perspective, a child learns language not as a system of information transmission but as a form of behavioral coordination.
It must learn, through trial-and-error strategies, to connect the multitude of linguistic expressions from adults with desired reactions of its own. 
Therefore, words like “forks/democracy” coordinate our actions with respect to what a person does when dealing with forks/democracy. 
Through the word “forks” and similarly through all other words, information is not transmitted but something specific is triggered in the recipient, which is determined by their structure and, indirectly, by their socialization.</p>

<p>Luhmann’s concept of <em>cognition as construction</em> does not provide an absolute, rigid foundation.
My explanation of it is based on my observation of Luhmann’s observations, and in case of the reader, it relies on your observation of my observation of Luhmann’s observation.
Since my, Luhmann’s and your observation are all constructed and contigent, it is inherently impossible to ‘prove’ the correctness of Luhmann’s theory from outside.
Furthermore, there will be always a blind spot.</p>

<blockquote>
  <p>A supertheory reflects on the fact that it and its validity are its own product—and is therefore absolutely contingent. […]
It is a theoretical endeavor, and there is nothing more to it.
It does little outside of theory. […]
With supertheory, the world does not become morally better, more rational, or spiritually complete.
It only becomes more distinct. – <a class="citation" href="#moeller:2006">(Möller, 2006)</a></p>
</blockquote>

<p>Radical constructivsm is often preceived as a dangerous path because it might lead to the relativization of ‘evil’ actions.
If everything is constructed, anything might be justifiable.
But again, Luhmann does not claim that there is no reality or that one can construct whatever one desires.
If you jump out the window, you will get hurt regardless of what you imagine.
However, acccording to Luhmann (and many others), there is no absolute and I think that this uncertainty makes us more thoughtful than reckless.
If absolute truth is on my side, everything is permitted.</p>

<p>It would be erroneous to claim that the application of social systems theory is justified because it describes society more ‘accurately’ because, according to this very same theory, such claims are impossible.
What we can state is that a theory that makes more sense—culturally and personally—might be useful for effective communication, facilitating a better understanding of one another.
This view aligns with Richard Rorty, another so-called postmodern ‘charlatan’:</p>

<blockquote>
  <p>If we can just drop the distinction between appearance and reality, we should no longer wonder whether the human mind, or human language, is capable of representing reality accurately.
We would stop thinking that some parts of our culture are more in touch with reality than other parts.
[…] we would not say that [our ancestors] were less in touch with reality than we, but that their imaginations were more limited than ours.
We would boast of being able to talk about more things than they could. – <a class="citation" href="#rorty:2016">(Rorty, 2016)</a></p>
</blockquote>

<h3 id="we-are-different">We are Different</h3>

<p>So let me ask: Why should machines (even if they are conscious) ‘live’ in a ‘world’ that is similar to ours?
And why should the bat-workd, my-world, and the machine-world be more true than any other system/environment distinction?
I think it is a mistake to undervalue our mental processes just because machines can communicate, calculate, and display creativity.
I also think it is a mistake to believe that thinking and computing are similar operations.
However, according to systems theory, viewing humans as superior rational entities is misguided as well.</p>

<p>A recurring theme in Luhmann’s writings is the idea of <strong>difference</strong> and <strong>differentiation</strong>.
Neither do we nor does any other system hold an intrinsic superior position in society; existence is <em>contingent</em>.
Observation necessitates selection, i.e. non-observation (a blind spot).
If I observe my cup, I cannot oberserve myself at the same time.
Because of differentiation there is no system that can make absolute sense of some indpendent reality.
The legal system ‘looks’ at a house in legal terms and the economy in economic terms.
There is no view from the top, above or the outside.
A system can only observe in accordance to its <em>systemic rationality</em>.
And as I argued above, bats, human beings, and maybe at some point in future, machines ‘live’ in, or better create, different interdependent ‘worlds’.
We are differnt; neither better nor worse with respect to each other and to other beings/systems.
I think, that should be enough to be fascinated about ourselves and to value us as different as we are.</p>

<p>This multiplicity of ‘worlds’—this functional differentiation—makes communcation so difficult.
It seems like a miracle that we sometimes seem to understand each other.
Luhmann believed that ecological problems are primarily problems of communication.
Different social systems (e.g., politics, science, economics) have different ways of observing and communicating about the environment, which can lead to misunderstandings or contradictions.
So one idea I have in mind is that via <em>artificial communication</em> we might be able to reduce misunderstanding between different systems, including the ecological system.
From a Luhmannian perspective, we need a sort of <em>translation</em> from one <em>systemic rational</em> to the other.
Furthermore, I believe that we need a more strict coupling between the ecological system and psychic as well as social systems.
Maybe artificial communication can establish such a strict structural coupling.</p>

<p>Only because we are not in charge, does not mean that things can not get better.
There seems nothing inherently wrong with biological evolution.
Is social evolution any different?</p>

<h2 id="references">References</h2>

<ol class="bibliography"><li><span id="hofstadter:1979">Hofstadter, D. (1979). <i>Gödel, Escher, Bach: An Eternal Golden Braid</i>. Basic Books.</span></li>
<li><span id="bender:2021">Bender, E. M., Gebru, T., McMillan-Major, A., &amp; Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? <i>Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency</i>, 610–623. https://doi.org/10.1145/3442188.3445922</span></li>
<li><span id="bender:2020">Bender, E. M., &amp; Koller, A. (2020). Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data. <i>Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics</i>, 5185–5198. https://doi.org/10.18653/v1/2020.acl-main.463</span></li>
<li><span id="moeller:2006">Möller, H.-G. (2006). <i>Luhmann Explained: From Souls to Systems</i>. Open Court.</span></li>
<li><span id="luhmann:1988">Luhmann, N. (1988). <i>Erkenntnis als Konstruktion</i>. Bern: Benteli.</span></li>
<li><span id="luhmann:1993">Luhmann, N. (1993). <i>Deconstruction as second-Order observation</i>. New Literary History.</span></li>
<li><span id="luhmann:2000">Luhmann, N. (2000). Why does society describe itself as postmodern. In W. Rasch &amp; C. Wolfe (Eds.), <i>Observing complexity: Systems theory and postmodernity</i> (pp. 35–49). University of Minnesota.</span></li>
<li><span id="brown:1969">Spencer-Brown, G. (1969). <i>Laws of Form</i>. London: Allen and Unwin.</span></li>
<li><span id="esposito:2022">Esposito, E. (2022). <i>Artificial Communication</i>. The MIT Press. https://doi.org/10.7551/mitpress/14189.001.0001</span></li>
<li><span id="shannon:1948">Shannon, C. E. (1948). A mathematical theory of communication. <i>Bell Syst. Tech. J.</i>, <i>27</i>(3), 379–423.</span></li>
<li><span id="moeller:2021">Möller, H.-G., &amp; D’Ambrosio, P. J. (2021). <i>You and Your Profile: Identity After Authenticity</i>. Columbia University Press.</span></li>
<li><span id="mcluhan:1992">McLuhan, M. (1992). <i>The Global Village: Transformations in World Life and Media in the 21st Century</i>. Oxford University Press.</span></li>
<li><span id="heidegger:1927">Heidegger, M. (1927). <i>Sein und Zeit</i>.</span></li>
<li><span id="baudrillard:1990">Baudrillard, J. (1990). <i>The Transparency of Evil: Essays in Extreme Phenomena</i>. Verso.</span></li>
<li><span id="tononi:2015">Tononi, G., &amp; Koch, C. (2015). Consciousness: Here, there and everywhere? <i>Philosophical Transactions of the Royal Society B: Biological Sciences</i>, <i>370</i>(1668), 20140167.</span></li>
<li><span id="nagel:1974">Nagel, T. (1974). What is it like to be a bat? <i>Philosophical Review</i>, <i>83</i>(October), 435–450. https://doi.org/10.2307/2183914</span></li>
<li><span id="foerster:1988">Foerster, H. (1985). <i>Sicht und Einsicht</i>. Vieweg.</span></li>
<li><span id="rorty:2016">Rorty, R. (2016). <i>Philosophy as Poetry</i>. University of Virginia Press.</span></li></ol>]]></content><author><name>Benedikt Zönnchen</name><email>benedikt.zoennchen@web.de</email></author><category term="Social Systems Theory" /><category term="AI" /><summary type="html"><![CDATA[Generative AI, especially ChatGPT, brought artificial intelligence into the public sphere and sparked a lot of highly speculative claims about machine intelligence. I’m open for discussions and unafraid of confronting uncomfortable truths. Indeed, our imagination and fearless thinking should pave the way for new possibilities. Dreams and speculations are valuable, as long as they’re presented as such. However, I find it concerning when public figures speak with undue certainty, particularly when making anthropological comparisons between humans and machines.]]></summary></entry></feed>