<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="4.2.2">Jekyll</generator><link href="https://bzoennchen.github.io/Pages/feed.xml" rel="self" type="application/atom+xml" /><link href="https://bzoennchen.github.io/Pages/" rel="alternate" type="text/html" /><updated>2026-09-04T15:27:59+02:00</updated><id>https://bzoennchen.github.io/Pages/feed.xml</id><title type="html">Bene’s Blog</title><subtitle>A blog dedicated to computer science, education, music, philosophy and technology</subtitle><author><name>Benedikt Zönnchen</name><email>benedikt.zoennchen@web.de</email></author><entry><title type="html">Cognition as Construction</title><link href="https://bzoennchen.github.io/Pages/2026/08/19/cognition-as-construction-en.html" rel="alternate" type="text/html" title="Cognition as Construction" /><published>2026-08-19T00:00:00+02:00</published><updated>2026-08-19T00:00:00+02:00</updated><id>https://bzoennchen.github.io/Pages/2026/08/19/cognition-as-construction-en</id><content type="html" xml:base="https://bzoennchen.github.io/Pages/2026/08/19/cognition-as-construction-en.html"><![CDATA[<h2 id="0-toward-a-new-epistemology">0. Toward a New Epistemology</h2>

<p>The aim of the following text is to bring clarity to Niklas Luhmann’s essay <em>Erkenntnis als Konstruktion</em> (<em>Cognition as Construction</em>) <a class="citation" href="#luhmann:1988">(Luhmann, 1988)</a>.
We try to untangle one of modern systems theory’s most radical propositions: that cognition succeeds not despite being completely cut off from an external reality, but precisely <strong>because</strong> of it.
By doing so, I will do my best to translate Luhmann’s dense terminology into an accessible roadmap, exploring how <em>operationally closed systems</em>—from minds to social communication—construct their own worlds through <em>distinctions</em>, <em>blind spots</em>, and <em>structural couplings</em> rather than mirroring a pre-given truth.</p>

<p>Since I am interested in applying his theory to questions concerning artificial intelligence, I will sprinkle in my own remarks from time to time.</p>

<p>Whether you are curious about the philosophy of science, artificial intelligence, or simply how systems manage to make sense of an otherwise unstructured environment, this step-by-step breakdown offers hopefully a fresh, humbling perspective on what it actually means to “know”, to “<em>make sense</em> of the world”, and to live in a <em>structural drift</em> in an <em>environment</em> of interdependet but also autonomous <em>systems</em>.</p>

<p>If you want to follow along, I am using the German Reclam edition from the <em>Great Papers Soziologie</em> series, working through the chapters step by step.
Each headline of this article (except the last one) corresponds to chapter in the book.</p>

<p>Naturally, the core intellectual achievement—as we attribute it usually—belongs entirely to Luhmann and the authors whose work he integrated into his theory.
Apart from relating his work to “artifical systems”, my contribution, if any, lies simply in the attempt to follow his reasoning.
As a computer scientist familiar with the \(\mathcal{P} \stackrel{?}{=} \mathcal{NP}\) problem, I find that tracing a complex thought (\(\mathcal{P}\)) is considerably easier than originating it (\(\mathcal{NP}\)).
As we will see, information is not transmitted to the reader but constructed within “her”.
Thus, all interpretations are inevitably misinterpretations; they can never perfectly capture an author’s actual thoughts, as even the author is constrained by the limits of communication and thought is not communication.
Whether my interpretation will, in turn, succeed in stimulating coherent thoughts in the reader’s mind remains to be seen.</p>

<p><em>Erkenntnis als Konstruktion</em> originated as a lecture Luhmann delivered at the Kunstmuseum Bern on October 23, 1988.
It was published later that same year as a standalone volume by Verlag Benteli, in a series edited by Gerhard J. Lischka.
It is often said that to understand Luhmann’s “supertheory” of social systems, there is no bypassing this foundational text. It serves as the starting point for a theory aiming to comprehend society <em>as</em> communication <em>through</em> communication. As the back cover of the small edition aptly puts it:</p>

<blockquote>
  <p>Luhmann in a nutshell.</p>
</blockquote>

<p>Luhmann can be seen as a contiental philosopher who was never taken serious by philosophy departments.
And therein lies my interest regarding this text, that is, not in his broader sociology and more in the novel epistemology he promises—a framework that connects to the idealist tradition (Kant, Fichte, Hegel) while simultaneously threatening to supersede it.</p>

<p>What I find most compelling is how Luhmann detaches cognition from the human subject.
This detachment is heartbreaking but also fuels my interest in finding a theory and a “suitable vocabulary” (to borrow from Richard Rorty) for discussing not only what we currently call “artificial intelligence,” but also our ongoing ecological crises and catastrophes.
As someone without formal training in any philosophical tradition, I leave the final judgment on the success of this undertaking to others.</p>

<h2 id="1-from-retreat-to-radical-closure">1. From Retreat to Radical Closure</h2>

<p>Luhmann voices doubts about the true radicality of the <em>radical constructivism</em> <a class="citation" href="#glasersfeld:1995">(Glasersfeld, 1995)</a> of his time, precisely because it feels the need to inflate itself as <em>radical</em>.
Such an emphatic attribute immediately raises the suspicion that certain theoretical inconsistencies are being concealed.
As an analogy, he points to Kant’s critique of Descartes’s “problematic idealism.”
This Cartesian stance is not nearly as innocently “open” as it appears; according to Kant, the consciousness of my own existence in time (the very thing Descartes takes as immediately certain) already presupposes the perception of something persistent outside of me. Inner experiential consciousness, in other words, is only possible on the presupposition of an outer one.</p>

<p>Luhmann does not choose this reference by accident.
He sees the old Cartesian figure resurfacing in radical constructivism, which repeats the maneuver by declaring its own operations (whether cognitive or systemic) to be accessible and given, while treating the status of an independent reality as undecidable or epistemologically irrelevant.
Instead of starting from the <em>cogito</em>, it starts from self-observing, self-referentially closed systems.
Yet the underlying gesture remains identical: the inside is securely accessible, while the outside remains “problematic.”</p>

<p>So, what is actually new about radical constructivism if it merely rehashes an issue Kant had already settled?
Luhmann’s answer—which also serves as his own ambitious theoretical agenda—is that constructivism must do its homework.
It must demonstrate exactly how the <strong>operative closure</strong> of the system is possible.</p>

<p>In doing so, however, there can be no backtracking.
Cognition must remain exclusively constructive.
No direct relation to an external reality may be presupposed—which, of course, does not imply that no such reality exists. 
It is ultimately up to the reader, the audience, and perhaps society at large to “decide” whether Luhmann succeeded in this task.</p>

<p>Returning to Kant, Luhmann accuses him of making precisely the same kind of (much-discussed) theoretical retreat as many constructivists seem to make.
In the <em>Transcendental Aesthetic</em>, Kant put forward a thesis that, measured against philosophical tradition, was quite radical: space and time are not properties of <em>things-in-themselves</em>, but pure, subjective forms of intuition. 
This has a far-reaching consequence.
Everything that exists in space, including outer objects—the “outer world” in the common sense—possesses only empirical reality, but transcendental ideality. In other words, spatial things are merely appearances to a subject equipped with our specific sensory apparatus; they are not things-in-themselves.
Strictly speaking, this already constitutes a form of idealism regarding the outer world, albeit a different and subtler one than Berkeley’s. (On Berkeley cf. <a class="citation" href="#downing:2004">(Downing, 2004)</a>.)</p>

<p>In the first edition of the <em>Critique of Pure Reason</em>, Kant treated this proximity to idealism relatively openly in the <em>Fourth Paralogism</em>. 
There, the existence of outer objects appeared as something that—unlike the immediately certain inner consciousness—first had to be inferred, rendering it ultimately less certain.
Structurally, this stance aligned closely with Descartes’s problematic idealism.
This formulation, however, earned Kant fierce criticism: the notorious Feder-Garve review of 1782 accused his transcendental idealism of being ultimately indistinguishable from the dogmatic idealism of Berkeley, who denied the existence of matter altogether.</p>

<p>It was precisely this dogmatism that Kant had sought to overcome.
He aimed to abandon transcendent speculations in favor of investigating the conditions of the possibility—meaning, for instance, that rather than attempting to prove God’s existence, he asked why we possess or articulate a representation of God in the first place.
This marks Kant’s famous Copernican turn: the shift from the transcendent to the <strong>transcendental</strong>.</p>

<p>Reacting to this backlash in his second edition, Kant largely struck the <em>Fourth Paralogism</em> and inserted a new section, the <em>Refutation of Idealism</em>.
In this precise passage, he explicitly distances himself from both Descartes’s problematic idealism and Berkeley’s dogmatic idealism, arguing instead that the consciousness of my own existence in time already presupposes the perception of something persistent outside of me.
Outer experience, therefore, is at least as immediately certain as inner experience.</p>

<p>With this maneuver, Kant attempted to defuse the tension inherent in the distinction between the <em>empirically real</em> and the <em>transcendentally ideal</em><sup id="fnref:1" role="doc-noteref"><a href="#fn:1" class="footnote" rel="footnote">1</a></sup>.
Whether he actually succeeded in resolving this consistently, or merely asserted a solution, remains a matter of dispute today.
Luhmann’s judgment here is pointed: this revision was not a genuine solution, but rather an “unclearly” executed weakening of an originally sharper, more idealist position.
For Luhmann, it represents a capitulation to external pressure rather than a fully realized theoretical result.</p>

<p>Kant’s transcendental philosophy might still be regarded today as the final word in epistemology; 
perhaps we will simply get no closer to solving the riddle.
Faced with this limit, one might be tempted to capitulate, close the file, and conclude that epistemology cannot fully resolve its own problems, pivoting instead to empirical research (such as cognitive science or the psychology of perception).</p>

<p>Luhmann, however, signals that he does not consider this capitulation inevitable.
Instead, he prepares the ground for his own project: developing a systems-theoretical, constructivist epistemology that genuinely advances these perennial problems (inside/outside, subject/world, certainty/doubt).
Rather than retreating or resigning, Luhmann proposes a radical reformulation based on the concept of operationally closed, observing systems—one that avoids merely redrawing the old Cartesian battle lines between inner and outer.</p>

<p>Since Kant (and continuing through Fichte, Hegel, and others), the fundamental problem has been this: cognition and the real object it seeks to know are two distinct things (<em>difference</em>).
Yet, there must somehow be a relation and thus a <em>unity</em> between the two; otherwise, cognition would be mere illusion.
Within this tradition, epistemology investigates this <em>unity of the difference</em>.
How can these two domains be held together when the theoretical starting point is their absolute separation? How can a cognizing consciousness confirm or establish something that is not itself a part of that consciousness? 
This is the classic <em>subject-object problem</em> in its most acute form.</p>

<p>Whether one takes Kant’s transcendental-philosophical route (where the conditions for the possibility of cognition lie within the subject itself, in its forms of intuition and categories) or Hegel’s dialectical route (where subject and object mediate one another through an unfolding historical process)—in both cases, the lack of direct access to reality is treated as an obstacle to be overcome.
The operative word here is “<strong>although</strong>”: the closedness of cognition is viewed as a problem, a structural deficiency that cognition must somehow transcend to achieve its aim.</p>

<p>The decisive turn of radical constructivism—one that Luhmann readily adopts—is a shift from “<strong>although</strong>” to “<strong>because</strong>.”
What Kant and Hegel viewed as a deficit to be overcome becomes the very precondition of cognition.
The premise is no longer that cognition succeeds despite its closure from the world, but rather that it succeeds only <strong>because</strong> it is closed.</p>

<p>Luhmann calls this an “empirical finding,” arguing that when we look at our environment, we observe <em>operationally closed systems</em>.
This is a surprising choice of words regarding the topic and our “naiv” starting point.
These “empirical findings” are heavily laden with presuppositions; they only be thought of as “real” through the very theory they are supposed to ground.
The logic is undeniably circular, and anyone who finds such circularity inadmissible might as well set the text aside here.
Furthermore, by claiming an empirical basis, Luhmann makes far less rigorous demands than Kant.
He offers no necessary, <em>a priori</em> validity.
Instead, his findings are explicitly fallible and revisable, making no claim to ultimate, independent certainty.
While this may disappoint anyone seeking the <strong>Absolute</strong>, it underscores that Luhmann’s theoretical edifice is built not on pure philosophical speculation, but on the observation of real-world phenomena (like brains or communication).</p>

<p>Drawing on <a class="citation" href="#maturana:1987">(Maturana &amp; Varela, 1987)</a>, Luhmann offers examples of <em>autopoiesis</em> and <em>operational closure</em>.
A brain, for instance, does not “see” light or “hear” sound directly.
Neurons respond only to the electrochemical impulses of other neurons, never directly to the outside world.
Figuratively speaking, the brain is blind and deaf to its environment, operating exclusively within its own network.
Strikingly, communication (<strong>!</strong>) operates the same way: it can only connect to further communication, never directly to physical or psychic events.
The environment—whether a consciousness, a brain, or a natural event—cannot “join the conversation.”
It can trigger communication, but it can never become part of it. (Luhmann famoulsy concludes that humans belong to the environment of social systems.)</p>

<p>The same logic applies to consciousness, traditionally viewed as the “subject” of cognition <a class="citation" href="#luhmann:1985">(Luhmann, 1985)</a>.
Psychic systems are closed loops; thoughts connect only to other thoughts, never directly to the world or to another mind.
Consciousness can produce information (by drawing internal distinctions) only because it is wired to be <em>environment-indifferent</em>.
Again, it succeeds not despite its insulation, but because of it—closure prevents the environment from constantly interfering.</p>

<p>Through this abstraction—and by admitting communication into the ranks of <em>observing systems</em>—Luhmann entirely recasts the classical metaphor of the subject as the “seat” of cognition.
He does not refute Kant (which would exceed the ambitions of systems theory) but rather revises him.
The core of this revision strips consciousness of its privileged status as the transcendental site of cognition, demoting it to just one of several (structural) equal autopoietic systems.
This shift has far-reaching consequences for what “cognizing” within a consciousness actually means.</p>

<p>However, Luhmann cautions that even with the shift from “although” to “because,” radical constructivism has not been fully thought through. 
It remains a program rather than a finished solution. 
The truly compelling question is no longer whether systems are closed, but how this closure and decoupling occur in the first place. This poses a dynamic research question, not a static assertion.</p>

<p>Having delivered this liberating theoretical blow, Luhmann opens up his research program.
Previously, theories of the subject were hobbled by the problem of introspection: if my only access to consciousness is my own, how can I know how others experience the world?
This required inferring another’s inner experience from one’s own, without any direct access.
The classical, typically Kantian, way out was to assume that all minds function according to shared principles (e.g., Kant’s universal categories and forms of intuition).
By introspecting, one could ostensibly generalize how every consciousness orders reality and I would argue that in everyday life we operate under a Kantian worldview.</p>

<p>While this secured an understanding of the “other,” it came at a steep price: the required presupposition of a shared world (Kant’s <em>thing-in-itself</em>) to which all subjects relate.
Without shared categories, perfectly private, internally consistent “worlds” would be possible, and objectivity would collapse.
But holding onto a shared world makes it impossible to seriously claim that every cognizing system constructs its own environment in radical isolation.
The presupposition of a shared reality fundamentally contradicts the idea of radical closure.
In other words, there is now way to know how <em>it is like to be</em> a social system.</p>

<p>Here, Luhmann indirectly suggests that the radical constructivism of his era was still too idealist, too Kantian, and ultimately not radical enough.
He insists on following this path to its absolute end; otherwise, the theory will collapse under its own contradictions.</p>

<p>He then considers the obvious alternative: instead of starting with the subject, one could treat the cognizing being purely as an object of scientific description—as a physical system, a biological organism, a psychological mind, or a sociological unit.
While sufficient for many empirical projects, this approach fails an epistemology focused on the question of closure.
Simply describing a cognizing being as an object with internal processes tacitly presupposes its closure rather than explaining it.
It describes what happens inside the object, but not how or why the object decouples from its environment to begin with.
The core problem is simply bypassed.</p>

<p>This leads to the central theoretical move of Luhmann’s systems theory: since neither subject-theory nor object-theory works, a new foundational distinction is required.
He replaces subject/object with <strong>system/environment</strong>.
This is far more than a mere swap of vocabulary, as the new terms carry entirely different implications
Drawing on George Spencer-Brown’s calculus of forms and the concept of <strong>re-entry</strong> <a class="citation" href="#brown:1969">(Spencer-Brown, 1969)</a>, Luhmann argues that the system/environment split, like subject/object, begins with a distinction.
But just as the classical subject held a representation of the object within itself, Luhmann posits that the system/environment distinction “re-enters” one of its own sides (the system).
The system observes both itself and its environment using a distinction it has drawn itself.
While this retains a structural similarity to the classical problem, it transcends it.
It is neither purely subject-oriented (the system is not simply consciousness) nor purely object-oriented (the environment is not just an objective reality passively depicted).
By proposing that system and environment are constructed through the very act of distinguishing between them, Luhmann simultaneously critiques and integrates both traditions.</p>

<p>Two concrete gains emerge from this new distinction. First, the question “How is closure possible?” can be reframed specifically as the <strong>differentiation (Ausdifferenzierung) of systems from their environment</strong>, opening the door for new systems-theoretical research programs.
Second, the old presupposition of a shared, objective world can be discarded in favor of a theory of <em>second-order observation</em>.
Following Heinz von Foerster’s second-order cybernetics <a class="citation" href="#foerster:2003">(von Foerster, 2003)</a>, a system no longer observes “the world” directly; instead, it observes how other systems observe.
This allows us to conceptualize closed systems in a radically different way, free from the crutch of a shared reality.
These systems can mutually observe one another as observers, despite having no direct access to each other. 
Ultimately, this leaves us with the pressing question: <strong>How is observation via closure possible?</strong></p>

<h2 id="2-an-environment-without-distinctions">2. An Environment Without Distinctions</h2>

<p>We must emphasize once again that Luhmann does not claim there is nothing outside of cognition.
That would amount to a <em>naive solipsism</em>, a position from which he strictly distances himself.
Instead, he posits as a starting point that cognizing systems (brains, consciousness, communication systems) genuinely exist within an equally real environment.</p>

<blockquote>
  <p>We proceed on the assumption that all cognizing systems are real systems in a real environment, in other words: that they exist. – <a class="citation" href="#luhmann:1988">(Luhmann, 1988)</a></p>
</blockquote>

<p>This is a <strong>realist profession of faith</strong>, a detail often overlooked when constructivism is crudely equated with the claim that “there is no reality.”
One should note, however, that this assertion does not presuppose direct access to that “reality.”
It must therefore be strictly distinguished from abandoning the concept of operational closure—we do not retreat!</p>

<p>Luhmann anticipates the predictable objection: asserting such a claim of existence (without access) simply as an ungrounded presupposition seems philosophically unsatisfying, or even “<em>naive</em>,” since it is not secured by epistemology.
Yet Luhmann uses this to highlight the fundamental dilemma of all cognition: the beginning of any theory is necessarily ungrounded, because grounding it would require a theory that does not yet exist.
This is the well-known <em>problem of ultimate grounding</em> (cf. the “Münchhausen trilemma”): one cannot ask for justifications infinitely; one must start somewhere.
Naivety at the outset is thus not just unavoidable, but constitutive of any theorizing.</p>

<p>The justification for this starting point can only be delivered retroactively (though it must indeed be delivered eventually), using the very theory that was developed from it.
This represents a circular—but for Luhmann, legitimate—structure: a theory grounds its own premises retroactively, <strong>once it has become complex enough to observe itself</strong>.<sup id="fnref:2" role="doc-noteref"><a href="#fn:2" class="footnote" rel="footnote">2</a></sup>
This strategy, though often conveniently forgotten, is practiced in mathematics as well.
There, mathematicians feel compelled to escape circularity through axiomatic construction and hierarchical levels, inevitably resulting in <em>incompleteness</em> (cf. <a class="citation" href="#goedel:1931">(Gödel, 1931; Turing, 1937; Bücker, 2024)</a><sup id="fnref:3" role="doc-noteref"><a href="#fn:3" class="footnote" rel="footnote">3</a></sup>). 
Luhmann’s remarks here offer a foretaste of the systems-theoretical principle of <em>self-reference</em>, as well as the frequently criticized difficulty of <em>falsifying</em> his—in this sense, Hegelian—supertheory.</p>

<p>Luhmann then reformulates the leading question from the previous section.
Instead of asking “How does cognition function?”, one must now ask, “How does a system close itself off operationally from its environment?”
Von Foerster’s play on words, “closure through enclosure,” captures the point perfectly: a system becomes closed (relative to its environment) by “enclosing” itself—that is, by drawing its own boundary and recursively referring back into itself.</p>

<p>Merely formulating this question carries an implicit realization: a system cannot decouple arbitrarily or effortlessly, as this requires highly selective, rigorous conditions to be met.
To frame it in Kantian terms, we are asking know about <strong>the conditions of the possibility of decoupling</strong>.
Decoupling, therefore, is no trivial matter; it is a highly presupposition-laden achievement!</p>

<p>Luhmann anticipates another central misunderstanding of constructivism: the assumption that if a system is “cut off” from its environment and operates solely upon itself, it can “do whatever it wants.”
This often leads to the naive accusation of pure relativism—the idea that “anything goes.”
Luhmann explicitly rejects this.
Citing operationally closed systems across various levels—biological (a cell), immunological (an immune system), neural (a nervous system), psychic (a consciousness), and social (a communication system)—he argues that decoupling does not produce arbitrariness.
On the contrary, it produces the exact opposite: strict order and constraint.
From the perspective of an outside observer, as soon as a system closes itself operationally, its subsequent operations become sharply restricted.
The closure itself dictates how the system can proceed, resulting in structural determination and path-dependency rooted in the system’s own recursive history.</p>

<p>In the real world, arbitrariness or randomness in the strict sense simply does not exist.
Everything that happens is conditioned by preceding states and structures.
There is no “causal closure”!
When someone attributes arbitrariness to a system, it is usually a symptom of insufficient observation; look closer, and the hidden regularities and constraints governing the seemingly erratic behavior will reveal themselves.
Arbitrariness, therefore, is never a property of the observed system itself.
Instead, it points to a deficit—or a specific standpoint—of the observer, who has not (yet) looked closely enough.
Consequently, the concept functions as an invitation to second-order observation: rather than observing the system, one must observe the (first-order) observer who is attributing the arbitrariness in the first place.</p>

<p>This leads directly to the answer to the original question: <strong>closure occurs when a system generates its own operations</strong>, which then relate to one another in a forward- and backward-referential recursive network.
This is the abstracted core of <em>autopoiesis</em>: operations generate further operations, referring back to past ones and anticipating future ones within a closed cycle.
Crucially, the distinction between system and environment is not given in advance; it is produced through the operating itself.
By operating recursively, the system actively creates its own boundary. The boundary is thus a result of the process, not a precondition for it.</p>

<p>Luhmann then highlights a surprising convergence between two vastly different theoretical traditions: Maturana’s biological concept of autopoiesis<sup id="fnref:4" role="doc-noteref"><a href="#fn:4" class="footnote" rel="footnote">4</a></sup> (self-producing, self-reproducing systems) and Lyotard’s philosophy-of-language concepts, such as “phrase” (the individual speech act or sentence), “enchaînement” (the linking of sentences to one another), and “différend” (Lyotard’s term for a dispute that cannot be resolved because the parties lack a shared rule of judgment or idiom).
Both thinkers arrive independently at the same fundamental idea: systems and discourses constitute themselves through the recursive linking of their own elements, rather than through reference to something external.<sup id="fnref:5" role="doc-noteref"><a href="#fn:5" class="footnote" rel="footnote">5</a></sup>
Luhmann emphasizes, however, that systems theory expresses this shared core most clearly.
The principle of operational closure holds without exception, even for cognitive systems—they can never operate “outward,” but only ever within their own boundaries.</p>

<p>This raises the question of whether every autopoietic operation (e.g., even the simple metabolism of a cell) should be termed “<em>cognition</em>,” or whether cognition is a specific subclass of autopoietic operations requiring more precise delimitation.
This leads to a comparison between Maturana’s position and Luhmann’s own.
Maturana’s solution was to postulate a unified (“congruent”) concept of cognition—one acknowledging that autopoietic systems operate “blindly” (without any internal representation or depiction of the environment) while nevertheless existing within a domain of interactions.</p>

<p>To visualize this, picture a single cell, such as a bacterium.
It is operationally closed: it does not “see” its environment and has no inner representation of it.
Nevertheless, it reacts to specific environmental stimuli—moving along a sugar gradient (chemotaxis), adjusting its metabolism to temperature fluctuations, or responding to chemical signals from other cells.
The totality of these possible interactions—all the environmental perturbations to which the bacterium can react with a structure-preserving change that maintains its organization—is what Maturana calls the cell’s <em>domain of interactions</em>.
What counts as a “recognizable,” connectable perturbation for the bacterium (i.e., something it can react to without losing its autopoiesis) depends entirely on its own internal structure and its organization.
The exact same environment would produce a completely different domain of interactions for a differently structured cell.</p>

<p>Through ongoing <em>structural coupling</em> with its surroundings, the system’s internal structure changes over time (a process Maturana terms “<em>structural drift</em>”), which in turn alters the future perturbations it is capable of processing.
In simple cells, this domain is small; with the evolutionary development of nervous systems, it becomes vastly larger and more highly differentiated. (It is precisely on this foundation that Maturana later builds his explanation of language and social coordination, framing them as “<em>consensual domains</em>” between structurally coupled nervous systems, cf. <a class="citation" href="#maturana:2000">(Maturana, 2000)</a>.)</p>

<p>For Maturana, it is crucial that this domain of interactions is something an outside observer establishes by describing the correlation between environmental events and the system’s reactions.
The system itself “knows” nothing of this and does not “represent” this domain to itself; it simply reacts blindly within the limits dictated by its structure.
Following these premises, if one asks where the categorical differences in cognition lie between highly complex and rather simple organisms, Maturana concludes that there are none.
One must concede that all life is cognition.
Thus, Maturana defines cognition exceedingly broadly: practically all living activity is “cognitive,” leading to the equation <strong>life = cognition</strong>.</p>

<p>Maturana then carves out a narrower concept: the “<em>observer</em>,” defined specifically by the possession of language.
While cognition (in Maturana’s broad sense) applies to all living beings, the “observer” is tied exclusively to linguistic capacity—essentially restricting it to humans or linguistically coordinated social systems.</p>

<p>Luhmann explicitly distances himself from Maturana here, defining both terms differently.
First, he wants to construe “cognizing” more narrowly, so that not every autopoietic operation automatically qualifies as cognition.
Second, he defines the “observer” not through language, but through the fundamental operations of distinguishing and indicating<sup id="fnref:6" role="doc-noteref"><a href="#fn:6" class="footnote" rel="footnote">6</a></sup>.
Again drawing directly on George Spencer-Brown’s calculus of forms in <em>Laws of Form</em> <a class="citation" href="#brown:1969">(Spencer-Brown, 1969)</a>, Luhmann posits that observing means drawing a <em>distinction</em> and <em>indicating</em> one of its sides.
This is a deliberately more formal, language-independent definition—one potentially applicable to non-linguistic systems as well (such as cells, the immune system, or perhaps AI systems?<sup id="fnref:7" role="doc-noteref"><a href="#fn:7" class="footnote" rel="footnote">7</a></sup>).
Yet, as the next sentence suggests, this formalization serves to narrow, rather than widen, the concept of cognition.</p>

<p>This reformulation of <em>cognizing</em> and <em>observing</em> is what first makes it possible to conceive of <em>social systems</em> as <em>cognizing systems</em>.
For me, it also opens the door to reflecting on the conditions of the possibility of technical systems.
By freeing cognition from biology (life as the sole possible space of cognition) and allowing for non-linguistic observation, Luhmann specifies the act of observing much more precisely, even if the definition remains highly abstract and nearly devoid of content.
Are “cognizing” and “observing,” then, just two terms for the exact same act?
This is not yet entirely clear at this point, though it certainly appears so.
Luhmann closes the section with an announcement, postponing the resolution of this conceptual fork in the road to later sections.
We are left in suspense as to what theoretical yield will come from choosing distinguishing/indicating over language as the definitional basis of the observer.</p>

<p>Building on this foundation (where the observer = a system that distinguishes and indicates), Luhmann now defines cognition operationally.
It consists of two interlinked operations: <em>observing</em> (distinguishing + indicating) and <em>describing</em> (understood as the recording or fixing of observations in some form).
Cognition is thus not a static state or a possession, but an ongoing activity.
Crucially, this definition explicitly includes the reflexive level.
A system can not only observe “the world,” but it can also observe the act of observing itself (<em>second-order observation</em>; cf. <a class="citation" href="#foerster:2003">(von Foerster, 2003)</a>).
Likewise, it can describe already existing descriptions (meta-description).
From the outset, cognition is structured to be capable of recursion, never confined to a single, first-order level.
The mechanism is simple but absolute: one draws a distinction (e.g., hot/cold, true/false, system/environment) and marks one of the two sides to indicate the one being referred to.
Both steps are inseparable.
Without a distinction, there can be no indication, and any indication is entirely dependent on the underlying distinction (one can only indicate one side of a given boundary, not just point to anything at random).</p>

<p>This definition of observing is kept deliberately broad so that it applies equally to all three types of autopoietic systems in Luhmann’s framework: living systems (operating via life, e.g., cells), psychic systems (operating via consciousness, e.g., thoughts and feelings), and social systems (operating via communication).
In this context, “indifferent” means that the concept of observing makes no distinction among these three levels.
It applies to all of them with equal, formal validity, because it demands only the logical structure of distinguishing and indicating, remaining completely indifferent to the specific material substrate.</p>

<p>At first glance, this inclusion of communication among observing systems makes Luhmann’s theory look suspiciously like <em>functionalism</em>—the idea from cognitive science that cognitive processes are <em>substrate-indifferent</em> and can run on any “hardware,” whether a biological brain, a computer, or a society.
But this is only true if one defines cognition in such an abstract way that it is no longer bound to minds.
The reverse is not the case, i.e., <strong>we cannot follow that cognition is conscious or that all cognizing systems think</strong>.
The mode of operation (thinking, communicating, biochemical reactions) differ and define a non-trasferable boundary.</p>

<p><strong>Remark:</strong> In my opinion, this precise distinction is exactly what is missing in much of the contemporary discourse surrounding <em>artificial intelligence</em>, where both the humanities and computer science often lack theoretical precision.
The humanities frequently remain trapped in a rigidly humanistic paradigm, dismissing AI systems as mere tools or instruments simply because they lack a human psyche.
Computer scientists, conversely, regularly fall into the functionalist trap, anthropomorphizing their creations by confusing statistical computation with actual “thinking.”
Luhmann’s framework offers a way out of this deadlock.
It allows us to seriously investigate whether an AI system might cognize—in the formal sense of drawing distinctions and indicating them within an operationally closed network—without ever needing to assert that it thinks, feels, or possesses consciousness.
The operational medium of the machine (algorithmic computation) remains strictly separated from the psychic medium of thought, preventing both the reduction of the machine to a simple tool and the romantic illusion that it has a mind.</p>

<p>This concept of cognition remains entirely neutral regarding how observations are stored or recorded—that is, regarding the specific form of a system’s “memory.”
We can point to concrete examples of this variety: in living or immunological systems, memory takes the form of biochemical storage (e.g., immunological memory or neural imprinting); in social systems, it might consist of written texts (books, files, and documents serve as the “memory” of society or of specific organizational subsystems).
The function—the recording of observations—remains identical, while the medium is radically different.</p>

<p>Crucially, however, observing and describing are not abstract, “free-floating” processes.
They must be realized as concrete operations within the relevant type of system—as a literal act of living, an actual, currently occurring act of consciousness, or a genuine event of communication.
An observation only reproduces a system’s closure and its system/environment boundary if it is an authentic operation of the system itself.
Only then does it truly “belong” to the network of that system’s recursive operations.</p>

<blockquote>
  <p>But the observing and describing must always themselves be a possible autopoietic operation, that is, an act of living, or an actual act of consciousness, or communication, for otherwise it would not reproduce the closedness and difference of the cognizing system, that is, it would not take place “within” the system. – <a class="citation" href="#luhmann:1988">(Luhmann, 1988)</a></p>
</blockquote>

<p>If it were anything else, the observation would not take place “within” the system at all; it would lie outside it, effectively non-existent for that specific system.<sup id="fnref:8" role="doc-noteref"><a href="#fn:8" class="footnote" rel="footnote">8</a></sup></p>

<p>From this premise, Luhmann introduces two important qualifications to forestall overinterpretation:
(1) Not every operation of a system is an observational operation.
A consciousness does not exclusively produce observing thoughts, nor does a communication system consist solely of observational communication.
Observing and describing form a subset of possible operations, not their totality.
(2) Even those operations that genuinely are observations need not be observable or describable exclusively as such.
Whether a given operation appears as an “observation” depends entirely on the perspective of a (further) observer.
Under different circumstances, the exact same operation could be described in entirely different terms (e.g., purely physiologically, purely energetically, purely computationally?, or purely in terms of social structure) without ever needing to be recognized as an “observation.”
“Being an observation,” therefore, is not an absolute, objectively visible property; it is inherently observer-relative.
If it were otherwise, the theory would collapse under its own contradictions.
Consequently, we must constantly keep the <em>principle of second-order observation</em> in view—a demanding task, since we cannot operate simultaneously in the first and second order.
Yet, this principle runs through the entire text.</p>

<p>How, then, should we define “the environment”? Luhmann draws a radical consequence from his preceding definition (cognizing = distinguishing + indicating): if cognition operates this way, and if we simultaneously insist on operational closure, a very specific understanding of a system’s environment emerges automatically:</p>

<blockquote>
  <p>A system’s environment contains no distinctions. – <a class="citation" href="#luhmann:1988">(Luhmann, 1988)</a></p>
</blockquote>

<p>The environment “as it is” does not inherently possess categories, oppositions, or divisions. It is simply unstructured facticity.
Distinctions—such as hot/cold, true/false, or this/that—are never properties of the environment itself; they are always something an observer <em>brings to</em> it.
Even concepts like “otherness” (the idea that one thing differs from another) and “possibility” (the idea that something could be otherwise) already presuppose a distinction: namely, between the actual and that from which it differs, or between the actual and the possible.
Without distinction, there are neither alternatives nor contrasts.
The environment is pure, incomparable, undifferentiated actuality; it does not “exist” in contrast to something else, it simply occurs.</p>

<p>Furthermore, even the seemingly objective statement that “there are other observers out there” is not a direct reading of the environment.
It, too, requires a distinction—namely, separating events one designates as “observing” from those one does not.
Recognizing other observers is thus an act of one’s own distinguishing.
In sum: there is nothing “simply observable,” because everything observable is always already a product of the system’s own observational operations.
This holds self-referentially even when observing other observers.
Second-order observation is thus itself an achievement of the observing system, not a mere perception of something pre-given.</p>

<p>This amounts to a direct assault on the <em>correspondence theory of cognition</em> (the classical idea that cognition “depicts” reality).
Because every cognitive content relies on a distinction (marking something as “this” and not “that”), and because the environment itself contains no distinctions, there can structurally be nothing in the environment that “corresponds” to what cognition produces.
Correspondence in the classical sense is impossible—not because cognition is unreliable, but because “correspondence” presupposes a structure that simply does not exist on the side of the environment.
Even “things” and “events,” as delimited, individuated units distinct from one another, do not exist in the environment itself.
Individuation (delimiting something as this specific thing) is always an achievement of the observer.</p>

<p>What follows is a reflection on the “naive beginning” of the theory’s own starting point.
Even the concept of “the environment” does not exist independently.<sup id="fnref:9" role="doc-noteref"><a href="#fn:9" class="footnote" rel="footnote">9</a></sup>
It is a strictly relational concept, one that only makes sense when paired with a “system” (an environment is always the environment of something).
Therefore, there is no “environment in itself,” only ever an “environment-for-a-particular-system.”<sup id="fnref:10" role="doc-noteref"><a href="#fn:10" class="footnote" rel="footnote">10</a></sup></p>

<p>Luhmann carries this consequence through to its logical end: <strong>apart from cognition</strong> (that is, apart from the observing operation of distinguishing), there are no “systems,” because the concept of a “system” is itself only one side of a distinction (system/environment) drawn by an observer.
The parenthetical remark in his text serves as a deliberate callback to the beginning of the previous section, where Luhmann had “naively” asserted: “there are systems.”
Now, he reflexively unmasks his own starting assertion: it is itself already an act of cognition, a drawn distinction, rather than a direct reading of reality.
This is not a contradiction, but the consistent application of the theory to itself (self-reference)—exactly what was promised in the first section.
The “naive” beginning can only be reflected upon retroactively, using the very theory built upon it.
Even Luhmann’s primary instrument—the system/environment distinction—is not a description of a pre-given reality, but is itself an operation of cognition, a tool that enables and guides cognition in the first place.</p>

<p>From everything established so far, it follows neither that the environment is “not real” (<em>idealism</em>/<em>anti-realism</em>), nor that “nothing” exists outside the system (<em>solipsism</em>).
Luhmann deliberately distances his position from both extremes.
If we were to assert that “there is nothing outside,” we would already be operating within cognition (since the statement relies on the distinction between nothing/something).
This treats “nothing” grammatically as a substantive, a nameable entity, even though “nothing” is precisely meant to designate the absence of all nameable things.
The result is a performative self-contradiction, and the inference fails for subtle logical reasons.
The solipsistic claim meets the same fate: precisely because it is an act of cognition, it is subject to the same limitations as any other cognition and can never claim to “correspond” directly to reality.
Neither the realist nor the solipsistic extreme escapes the <em>closedness of cognition</em>; both are merely further constructions, not privileged glimpses of reality “as it really is.”</p>

<p>Furthermore, even the seemingly most all-encompassing, “objective” concepts—reality, matter, the world—remain bound to distinctions the moment they are utilized by cognition.
These total concepts attempt to name the unity that a distinction presupposes (the overarching “form” from which the distinction carves out its two sides).
In an ironic allusion to Hegel’s concept of “<em>Spirit</em>” as the unity of differences, Luhmann refers to this as the “spirit” of the distinction: the ideal overarching context that spans both sides.
Even the concept of “<em>ultimate reality</em>” cannot escape this closedness.
It, too, can only be grasped using a specific distinction (in this case, system/environment), and therefore remains just as internal to the system as any other cognitive content.</p>

<p>Metaphorically speaking, the leading distinction a system currently employs acts as its “<em>blind spot</em>.”
It makes seeing and observing possible, but cannot itself be seen in that very moment (just as the eye cannot see itself while looking).
If one were to try to observe the active distinction, one would need a new, supplementary distinction to do so, which would immediately become the new, invisible blind spot. 
The blind spot can therefore never be eliminated; it can only be shifted.
This results in an infinite regress that precludes any absolute, presuppositionless observing.
(This is structurally related to the self-contradiction of classical epistemology discussed earlier, but here it is accepted as the necessary structure of all observation rather than lamented as a problem.)</p>

<p>In the vocabulary of George Spencer-Brown’s <em>Laws of Form</em> <a class="citation" href="#brown:1969">(Spencer-Brown, 1969)</a>, this dynamic is articulated as follows: prior to any distinction, there is the <em>unmarked space</em>—a state of pure, undivided possibility, devoid of any marking.
Every observation acts as a “cut” through this undivided space, leaving one side marked (indicated) and the other unmarked.
The undivided wholeness is, so to speak, wounded or broken open by the act of distinguishing.
Observation demands this cut as a prerequisite (one cannot observe without distinguishing), yet every observation is precisely what executes this cut in the first place (the boundary does not pre-exist; it arises through the act of observing).
The cut is simultaneously a <em>precondition</em> and a <em>construction</em>—or, in short, a <strong>presupposed constructing</strong>.
The process itself produces the difference.</p>

<p><em>Operational epistemology</em> turns away from the question of what cognition is (its essence or its classical conditions of possibility) and asks instead what kind of operation realizes.
Cognition is treated as just one operation among others—one that can be distinguished from mere metabolism, an unconscious reflex, or arbitrary communication.
This is a deliberately sober, demystifying approach, even if the concept of an “operation” itself is admittedly maximally empty.</p>

<p>The reader then might ask: But what about “the Truth”?
Crucially, the decisive criterion for whether cognition “happens” does not depend on whether it is correct. It depends solely on whether the system can continue its own self-reproduction (<em>autopoiesis</em>) by means of this operation because otherwise it would disappear and we could not observe it.
At its most fundamental level, then, <strong>cognition has absolutely nothing to do with truth</strong>—nothing to do with “cognizing correctly”!
For the mere functioning of autopoiesis, it is entirely immaterial whether a cognitive operation is true, false, or produces truth.
Both true and false operations function equally well to keep the system running.
Cognition is thus defined functionally, not normatively.
<strong>It is a purely operative capacity to keep going, not a seal of quality</strong>.</p>

<p>This can be grounded empirically at the neurobiological level: there is no physiological difference between a brain state carrying a true belief and one carrying a false belief.
Neurons do not “know” whether what they are currently processing is correct.
Similarly, consciousness<sup id="fnref:11" role="doc-noteref"><a href="#fn:11" class="footnote" rel="footnote">11</a></sup> and communication possess no separate “channels” or distinct mechanisms for the true versus the false.
The exact same attentional achievements, grammar, and sentence structures carry true statements just as effectively as false ones. A false sentence is not “built” any differently—grammatically, communicatively, or psychically—than a true one.</p>

<p>Precisely this indifference explains a common everyday phenomenon: the fact that errors can appear to us as truths at all (when they occur, they do not “feel” any different from truths).
This points to the actual epistemological challenge.
The primary question is not “How does one arrive at truth?” but rather “How does one, after the fact, recognize and eliminate errors?”
This is a marked shift from the classical framing.
Because the system’s bare operations are inherently indifferent to true/untrue, a <em>binary code</em> (true/untrue) must be imposed on this indifferent activity subsequently and additionally.
The code is not a natural property of the operations; it is a secondary, artificially instituted supplementary structure. 
This leads to a decisive follow-up question: if the code is not there organically, who or what introduces it?</p>

<p>Luhmann’s answer is direct: the observers themselves!
It is the observer who introduces the distinction true/untrue.
The code is not contained in the operations, but is <strong>imposed</strong> upon them from an observational vantage point.
However, observing is itself just another operation.
Therefore, in the exact instant it occurs, it cannot simultaneously grasp itself using the very distinction it is currently employing.
While I am judging something as true or false, I cannot, in the same simultaneous act, judge whether my judging itself is true or false. 
That would require a further, subsequent observational step in time.</p>

<p>Let us apply this to a classic logical-philosophical problem—familiar from debates over theories of truth in Frege or discussions of the “redundancy theory of truth”—namely, the difference between the simple assertion “A is the case” and the intensified assertion “It is true that A is the case.”
The latter arises only through an additional observation: a <em>second-order observation</em> that observes the primary observation and marks it as “true.”
The primary observation does not do this; it merely distinguishes “A” from “not-A,” without simultaneously qualifying itself as true or false.
<strong>Truth, then, is structurally always a second-order phenomenon</strong>.</p>

<p>Here, Luhmann alludes to classical logical strategies for avoiding paradoxes of self-reference (such as Russell’s <em>theory of types</em> or Tarski’s <em>distinction between object language and metalanguage</em>, which attempt to “dissolve” the liar’s paradox by strictly separating levels of statement).
This level-based solution, however, does not genuinely solve the problem; it merely displaces it.
The exact same question inevitably resurfaces regarding the distinction between the levels themselves: is this metalevel distinction itself true or false?
Does it belong to the object level or the metalevel?
The paradox is not dissolved; it is simply reproduced one level up.
It is hidden away, allowing the logician to (seemingly) carry on undisturbed.
This is a hallmark of Luhmann’s critique of purely logical attempts to solve problems that are, at their core, empirical and operative.<sup id="fnref:12" role="doc-noteref"><a href="#fn:12" class="footnote" rel="footnote">12</a></sup></p>

<p>Luhmann offers an alternative proposal: instead of trying to logically “resolve away” the paradox, we should investigate empirically how real cognitive systems (brains, consciousness, science as a social system) concretely organize self-observation to continuously detect and neutralize the errors they produce.
The goal is not to logically forbid the paradox, but to temporalize it.
A system does not have to resolve <em>self-reference</em> in a single, timeless logical act (as a formal derivation would require).
Instead, it can handle it distributed over time, across successive operations.
The question thus becomes: how do systems keep going despite the fundamental impossibility of marking errors as errors in the exact instant they are produced?
The theorist, here and elsewhere, turns out to be surprisingly empirical.</p>

<p>The solution Luhmann offers—which lays the groundwork for his famous theory of the binary codes of societal function systems (science: true/untrue; law: lawful/unlawful; economy: payment/non-payment)—lies in the concept of binary coding.
This artificially instituted two-value code (e.g., true/untrue) allows the system to sort its own operations retroactively and systematically, without ever having to solve the fundamental paradox of self-observation. 
The code works with the paradox rather than dissolving it.
Luhmann calls this <em>de-paradoxification</em> (<em>Entparadoxierung</em>).<sup id="fnref:13" role="doc-noteref"><a href="#fn:13" class="footnote" rel="footnote">13</a></sup>
The paradox is no longer a paralyzing logical dead-end, but a productive structure kept in constant motion.</p>

<h2 id="3-the-observer-within-the-system">3. The Observer Within the System</h2>

<p>Luhmann introduces a conceptual toolkit: if an epistemologist wishes to observe systems that in turn observe their own observing (i.e., engage in at least second-order observation), a highly refined set of distinctions is required to avoid theoretical confusion. He enumerates five:</p>

<ol>
  <li>“Operation” is the umbrella term (encompassing every autopoietic event: a thought, a communication, a metabolic process). “Observation” is a specific subtype: the operation of distinguishing and indicating. This produces a circularity: observation is at once a kind of operation <strong>and</strong> the very means by which one distinguishes what counts as an “operation” in the first place. One thus defines the part (observation) using the exact same tool that constitutes the whole (operation). Luhmann defuses this pragmatically: this circularity is only disturbing if one is already operating on the level of second-order observation; on a simple, first-order level, it carries no weight. An analogy illustrates this: the fact that we can only define language through language does not bother us while we are speaking. The circularity only becomes a problem when a linguist or philosopher of language tries to define “language” conceptually—that is, the moment they enter a metalevel of reflection.</li>
  <li>Every observer draws their own system/environment distinction (their own “system reference”—the standpoint from which what counts as system and what counts as environment is decided). A first-order observer has one such reference; a second-order observer (observing the first) has a different, distinct system reference. Because these two references are not identical, establishing how they differ would require a third-order observer to compare both from the outside. This reveals the infinite regress of observational levels implied earlier: every level has its own blind spot, which only the next-higher level can make visible—though that higher level inevitably introduces a new blind spot of its own.</li>
  <li>Whether an observation is an “observation of another” (observing something external) or “self-observation” (a system observing itself) can only be established if one already knows—has already distinguished—what belongs to the system (inside/self) and what belongs to the environment (outside/other). This distinction is therefore logically presupposed.</li>
  <li>When I observe another observer, I can focus on two entirely different things: either the content of their observation (what they are looking at) or the form of their observation (precisely what they themselves cannot see: their leading distinction, their blind spot). This is the truly interesting core of second-order observation—not checking what the other sees, but discerning how they see it (and what necessarily escapes them in the process).</li>
  <li>The truth code is only one possible form of a system’s (self-)observation. There are other codes and forms of observation (e.g., reputation, lawful/unlawful, payment/non-payment) that must not be confused with the truth code.</li>
</ol>

<p>Luhmann then establishes a strict, almost normative standard: the honorary title of “<em>constructivist</em>” is earned only by a theory that rigorously works through this entire web of distinctions (operation/observation, orders of observation, self-/other-observation, content/form, code/other codes) without repressing the inevitable paradoxes, but rather resolving them.
The decisive point is his concluding clause: cognition is not traced back to any ultimate “ground” (as in classical foundational philosophy, e.g., Descartes’ cogito or an Archimedean point), but only to ever further distinctions—distinctions of distinctions. 
There is thus <strong>no ground</strong> on which the regress comes to an end, only a <strong>web of differences</strong> that mutually account for one another. 
This is an explicit <strong>rejection of any form of foundationalism</strong>.</p>

<p>As long as cognition is explained biologically (via the autopoiesis of living systems, such as neurobiology) or psychologically (via the autopoiesis of consciousness), the epistemologist can maintain the illusion of standing outside the object under investigation.
My own brain or consciousness is a distinct, individual specimen separate from the one currently being studied.
I can observe a foreign brain “from the outside,” while my own continues to operate elsewhere, fundamentally uninvolved.
The only price of this external position is conceding a similarity of conditions: my brain, too, functions according to the same general biochemical laws.
Yet, as an individual specimen, I remain separate from the observed case.</p>

<p>If, however, cognition is construed <em>sociologically</em>—as a function of communication and an operation of social systems—the situation changes fundamentally.
Unlike brains or consciousnesses, of which there are many separate specimens, there is only one society, a single all-encompassing system of the autopoiesis of communication.
Every communication, including that of the epistemologist formulating the theory, belongs to this exact same system.<sup id="fnref:14" role="doc-noteref"><a href="#fn:14" class="footnote" rel="footnote">14</a></sup></p>

<p>Luhmann offers us an image: an experimenter observing rats in a maze typically stands safely outside it.
Kant, for instance, could write about the subject from an external vantage point (as a consciousness observing from the outside).
With a sociological concept of cognition, this external position vanishes.
The theorist becomes just another rat in the maze of society and communication they are attempting to describe. 
They can no longer retreat to an uninvolved observer’s standpoint outside the system; they must reflect on the position within the system from which they are observing the other participants.
The problem is not simply that cognition is described by means of communication, but that this very communication takes place within the cognizing system it is attempting to describe. 
This description is therefore necessarily <strong>incomplete</strong>, yet the resulting paradox can unfold productively.
(This is also why Luhmann called his magnus opus <em>The society of society</em> (<em>Die Gesellschaft der Gesellschaft</em>) <a class="citation" href="#luhmann:1998">(Luhmann, 1998)</a> because (1) it is a description or a theory about society, i.e., society as the gramatical object and (2) a theory produced by society as the gramatical subject.)</p>

<p>Compared to the biological or psychological cases, it is no longer enough to concede a mere commonality of conditions.
We have now arrived at the <em>unity</em> of the system itself: the theorist (or perhaps better, the theory) and the object of investigation are parts of the exact same communication system.
Any attempt to present oneself as “standing outside” (e.g., treating science as a special domain floating above the rest of society) must instead be accounted for as an internal systemic differentiation—a division occurring within the one society, rather than a genuine externality.</p>

<p>This is why Luhmann insists that constructivism only becomes truly <em>radical</em> when epistemology is conducted sociologically, because it must finally include itself.
The theorist can no longer secure a privileged position on the outside (as they secretly still could in the biological or psychological variants).
They must acknowledge that their own theory is simply a further communication within the very system it describes. This is the ultimate, most consistent form of self-reference toward which the entire text has been driving.</p>

<p>If we turn this distinction back onto consciousness itself—which might particularly interest us, since “we,” in this case, are the rat within the self-observing psychic system—the following applies: because Luhmann, unlike Kant, believes consciousness has no built-in, guaranteed commonality with other consciousnesses, the observable stability and coordination among millions of closed minds must come from somewhere.
For Luhmann, this <strong>explanatory burden is carried exclusively by communication</strong>.
In order for communication to achieve this—to produce stable, self-reproducing coordination across many different, insulated minds over long periods of time—it must itself be a self-producing, autopoietic system, not merely a loose collection of unconnected speech acts.
Only as an independent, closed system can communication take on the monumental role that, for Kant, was played by shared cognitive categories.</p>

<h2 id="4-the-theological-precursor">4. The Theological Precursor</h2>

<p>Despite all its radicality, constructivism remains exactly what it set out to be from the outset: an empirical theory, not a metaphysical or transcendental one.
But if it is merely a sober, empirical framework, why does it strike us (in the West) as so shockingly novel—so “<em>radical</em>”?
To answer this, Luhmann argues, we must look at history.
We have to examine the conceptual space that such a theory has historically been strictly forbidden from occupying in the history of thought.</p>

<p>Luhmann’s thesis is striking: no philosophical epistemology has ever ventured this far (with Hegel’s Logic bracketed as a possible exception, given its intensive engagement with self-reference, contradiction, and absolute negativity).
The reason is that the conceptual “place” where one would have had to think about radical undifferentiatedness—about that which lies beyond all distinction—had historically been monopolized by theology, not philosophy.
In negative theology, God is not merely beyond ordinary distinctions (large/small, warm/cold);
He is beyond every possible meta-distinction, even the distinction between “being distinguished” and “not being distinguished.”
In this sense, theology had long anticipated the <em>unmarked space</em>.</p>

<p>Within this theological framework, God cannot simply be designated as “the other” (relative to creation). 
Doing so would reduce “him” to a distinguishable object among others, subjecting “him” to the apparatus of distinction.
Instead, God is the “not-other”—the <strong>very condition of the possibility of all distinguishing</strong>, which cannot itself be positioned as an “other” without contradicting itself.
In God, all pairs of opposites used in comparative human thought (bigger/smaller, faster/slower) perfectly coincide.
Because God exists beyond every relative determination, the greatest and the smallest, the fastest and the slowest, are identical in “him”.
There are simply no standards left by which “bigger” or “smaller” could even be distinguished.</p>

<p>As radical and mystical as this sounds, it had to remain compatible with official, dogmatically binding Christian doctrine.
Pure philosophical mysticism was insufficient; the theory had to remain tethered to the Church.
Consequently, God had to be definable simultaneously as a person and a Trinity (three persons in one God), <strong>and</strong> as the wholly undifferentiated, “secret” (hidden) essence of all things.</p>

<p>For Nicholas of Cusa (Cusanus), the theological solution to the problem of cognition ran as follows: the things of the world are a “contractio” of God.
They represent a “contraction,” or finitization, of the infinite, undivided God into finite, distinguishable, created entities.
Through this process, God—though entirely unknowable in “himself”—makes “himself” indirectly knowable through “his” creation.
For human beings, truth consists in the correspondence between our own cognitive distinctions and the (God-created) distinctions inherent in things themselves.
At its core, this is a theologically grounded correspondence theory of truth.</p>

<p>However, this unleashes a genuine theological dilemma.
On one hand, Scripture promises the redeemed the “visio Dei”—the beatifying, immediate vision of God in heaven (beatitudo).
On the other hand, theology must stubbornly maintain that the divine essence remains strictly incomprehensible in itself (“divinam essentiam per se incomprehensibilem esse”); otherwise, God would be reduced to a finite, fully graspable object.
To salvage both claims, theologians had to attribute a capacity for self-observation to God (without which he would not be a personal, self-conscious being—a theological necessity for the Trinity, where the divine persons know and love one another). 
Yet, they could not push this capacity for observation so far that it mirrored the devil. In this tradition, the devil was considered the “boldest observer of God,” whose ultimate sin was the presumptuous, overreaching attempt to fully grasp God and thereby become “his” equal.</p>

<p>To solve this dilemma, Luhmann argues, medieval theology effectively had to practice a highly advanced form of <em>second-order cybernetics</em>!
It required a carefully differentiated observing of observers: of the elect (electi, the blessed who are permitted to see God), of the devil (who illegitimately attempts to “observe” God), and of God “himself” (who must observe “himself”).
Structurally, this is the exact same problem Heinz von Foerster would reformulate cybernetically centuries later.</p>

<p>This theological solution, however, drifted into dangerous proximity to a heterodox, almost heretical consequence: it implied that God needs creation, and perhaps even the damnation of the devil, in order to observe “himself” and achieve full self-consciousness.
God would no longer be the classically self-sufficient deity resting entirely within “himself”;
He would depend on an “other” (creation, the fallen devil) to act as a mirror.
This implication was so delicate that Cusanus himself, according to Luhmann, warned against putting these writings into the hands of unprepared readers.
It was simply too dangerous and too easily misunderstood.</p>

<p>This is why Luhmann insists that <em>radical constructivism</em> finds its true intellectual precursor not in the history of philosophy (Descartes, Locke, Kant), but in the history of theology.
More precisely, it is rooted in that technically advanced, conceptually rigorous theology (like that of Cusanus) which pushed to the absolute limits of what the theological institution could bear (courting suspicion of heresy and incomprehensibility, as evidenced by the warnings to “unprepared minds”).</p>

<p>One must distinguish, once again, the distinction of distinctions (observing the leading distinctions that other observers use) from the radically undifferentiated itself.
That undifferentiated horizon—which was once called “<em>God</em>”—is today, in Luhmann’s conceptual vocabulary, called “<em>world</em>” (when distinguishing system and environment) or “<em>reality</em>” (when distinguishing object and cognition).
In Luhmann’s theory, “world” and “reality” play the exact same structural role that God played in medieval theology: they serve as an unobservable, distinction-free horizon that every single distinction already presupposes, yet which can never be captured by a distinction itself.
The only difference is one of theoretical economy: where theology required a massive, highly complex dogmatic edifice to manage this undifferentiated horizon (the Trinity, the visio Dei, the problem of the devil), modern systems theory simply deploys the much leaner, depersonalized concepts of “world” and “reality.” 
They perform the identical structural function, entirely free of the theological burden.</p>

<p>Combining this historical insight with the earlier claim that cognition is fundamentally indifferent to truth, we arrive dangerously close to a Rortian pragmatism.
If truth is not a mirror of nature, then the medieval theological solution is structurally no “less true” or “more wrong” than our current scientific paradigms.
Richard Rorty famously summarized this epistemological shift by arguing that as human thought evolves, we do not get closer to the absolute Truth; rather, we develop new vocabularies that simply allow us to do more things.
Luhmann translates this pragmatic insight into the language of systems theory: through evolution, a cognizing system does not become more accurately aligned with its environment, because the environment remains an unreachable, distinctionless horizon.
What the system achieves instead is greater internal complexity.
By building distinctions upon distinctions, the system vastly expands its own internal states and possible operations, thereby achieving a higher degree of decoupling, autonomy, and ultimately, operational freedom.</p>

<p>Ultimately, I would argue that adopting such a view fuels humility and doubt—which serve as potent antidotes to cruelty.
When we abandon the illusion that our cognitive distinctions grant us direct, privileged access to an absolute Truth, we lose the dogmatic certainty that has historically justified imposing our worldviews on others.
Recognizing that our highest truths are complex, contingent, and fundamentally internal constructions fosters a deep theoretical modesty.
In this sense, acknowledging the operational closure of cognition is not a sentence to an intellectual prison, but rather an invitation to an empathetic and intellectually humble solidarity.</p>

<!-- TODO -->

<h2 id="5-medium-and-form">5. Medium and Form</h2>

<p>Now that we have clarified the formal definition of <em>observing</em> (distinguishing + indicating), a pressing empirical question arises: how can these two components fuse into one single, coherent process?
Luhmann emphasizes once again that <strong>a system must meet very narrow, highly selective conditions to execute this complex double operation</strong>.
Observing is no trivial achievement.
One conjecture is that for <em>sense-making</em> (Sinn-machende) systems (consciousness, communication) succeeds because for them it is just barely possible to perceive two things simultaneously as a unified whole.
Think of figure-ground perception: one sees the figure and the ground simultaneously, as a single coherent image, rather than as two separate acts of perception.</p>

<p>It is also possible that only sufficiently complex systems can, over time, amplify minute differences—such as conspicuous fluctuations in their own oscillating movements—into massive effects.
Cybernetics calls this deviation amplification <em>(positive feedback</em>), while linguistics observes a related phenomenon in hypercorrection (when speakers overcorrect to match a perceived norm, thereby overshooting the mark).
For themselves, operations are timeless.
They happen and disappear in an instant.
So it is the recursive linking what turns a sequence of momentary, otherwise-unrelated events into something with duration, expectation, and memory.
Eigenvalues <a class="citation" href="#foerster:2003b">(von Foerster, 2003)</a> can emerge through recursive operations, that is, small, unstable, momentary differences get built up, through recursive self-reference, into stable structures.
Crucially, this temporal amplification process also presupposes operational closure.
It requires an “own time”—an internal temporal rhythm specific to the system’s own operations, running parallel to the environment, which undoubtedly continues to exist simultaneously, albeit not in the same rhythm.
A system has its own time (Luhmann sometimes calls this “Eigenzeit”) precisely because, and only because, it’s operationally closed. (The connection to <a class="citation" href="#husserl:1928">(Husserl, 1928)</a> seems strong here.)</p>

<p>This internal rhythm, in turn, requires memory to perform two distinct functions.
First, it must conduct an ongoing <strong>consistency check</strong>, evaluating whether new impressions align with currently activated and relevant structures.
Second, it requires a schema that prevents emerging contradictions from registering as a paralyzing logical scandal.
Instead, memory “pulls apart” these contradictions into spatial or temporal differences (e.g., resolving a contradiction by concluding “it was like that earlier, but now it is different”).
Here again, we see the exact same mechanism of <strong>de-paradoxification-through-temporalization</strong> discussed earlier.</p>

<p>At this point, however, Luhmann reins himself in.
All of this empirical detail merely provides an ever more precise description of how evolutionarily improbable, yet nonetheless possible, cognizing systems actually are.
But it does not yield any deeper philosophical clarification.
One could certainly differentiate this line of research further, asking: does it make a difference whether the capacity for distinction is grounded biochemically (life), psychically (consciousness), or communicatively (or perhaps even computationally?)—and if so, what exactly is that difference?</p>

<p>Yet, as fascinating as this research agenda would be, it cannot illuminate the core philosophical question: <strong>the relation between cognition and object</strong>.
Empirical disciplines like neurobiology can tell us much about the real, material constitution of cognitive operations.
But they can tell us absolutely nothing about the reality of the environment—that which these operations must presuppose outside themselves as unknown and fundamentally unknowable.</p>

<p>This boundary never shifts, no matter how refined empirical inquiry becomes.
Why? Because empirical research is itself always just another operation of a cognizing system; it is never a transparent window to the outside.
Picture the most advanced neurobiology imaginable: high-resolution imaging, single-cell recordings, sophisticated computational models.
All of this produces data.
This data, however, is the product of an entire chain of operations: measuring instruments, statistical evaluations, theoretical interpretive frameworks, and scientific conventions dictating what counts as a “significant signal.”
Every single one of these operations is itself an act of distinguishing and indicating, and therefore itself already cognition.
Consequently, research only ever produces more cognition about the operations (more descriptions, more models, more data)—never a vantage point that steps completely outside cognition to check whether these descriptions actually correspond to something unconstructed.</p>

<p>This asymmetry therefore remains firmly in place, no matter how far empirical inquiry advances.
System operations (neurons firing, patterns of communication, psychic events) are concrete events that can be observed, measured, and described; they belong entirely to the system side and are therefore empirically researchable.
The environment, by contrast, is defined precisely as that which is not the system.
As soon as one attempts to <strong>positively capture it</strong>, the result automatically becomes part of the system’s own internal construction—it is no longer “the environment in itself.”
Better empirics, therefore, do not shift the boundary between the two.
They merely fill the domain of system operations with ever greater detail, without ever managing to cross the boundary itself.</p>

<p>Luhmann’s systems-theoretical translation of this dynamic is that the environment—in contrast to the system’s highly mobile operations—<em>appears</em> as something persistent, which in turn allows for recurrence and repetition. 
However, identifying something as “the same, recurring thing” is already an achievement of the cognizing system; it is not simply a given property of the environment.
Here, Luhmann accuses Kant of conflating two different claims.
Kant asserted that persistence in the environment is a condition for the actual <strong>existence</strong> of the subject in time (a strong metaphysical claim).
In reality, it is at most a condition for the subject being able to <strong>cognize</strong> or identify its own existence in time (a much weaker, purely epistemological claim).
Kant argues imprecisely here, sliding illegitimately from the weaker claim into the stronger one.</p>

<p>Luhmann reverses this relationship: it might be that the environment is not the persistent element at all. 
Rather, the cognizing system maintains a constancy (insisting on “the same thing”) even after its actual relationship with the environment has fundamentally changed.
The object simply “has to put up with” being designated in the same way, despite having altered.
Language, for instance, allows us to use a constant, unchanging word (like “movement”) to refer to something that is by its very nature inconstant (since movement is change itself). The constancy resides entirely in the linguistic sign, not in the signified.
The system, therefore, does not have to change to the same extent as the environment it references; it can deploy a stable linguistic tool to indicate change without changing alongside it.</p>

<p>A self-differentiating cognizing system enters a state that exists simultaneously with the environment, yet no longer runs in the same rhythm.
This decoupling, Luhmann suggests, is only possible if the environment itself exhibits temporal breaks or discontinuities against which the system can offset its own rhythm.
It requires deviation—and something can only deviate in relation to a baseline.
If the environment were temporally entirely smooth and continuous, lacking any interruptions or caesuras, there would be absolutely nothing against which the system could mark its own tempo as “different.”</p>

<p>Luhmann then turns to the early work of the Austrian-American psychologist Fritz Heider (1896–1988)—specifically his 1926/27 essay “Ding und Medium” (“Thing and Medium”), a text that has received scant attention in academic epistemology.
In it, Heider describes the real (physical) conditions that make perception at a distance possible in the first place.
According to Heider, the physical outer world relies on a difference between <em>loosely coupled</em> elements (like air molecules, which move relatively independently) and <em>firmly coupled</em> structures (like a specific sound or tone, which holds a stable form).
What is decisive is the difference itself.
If the air itself constantly made noise, or if light itself were directly visible (rather than merely illuminating other things), noise-free perception would be impossible.
In other words, the medium itself must remain “silent” and inconspicuous so that distinct forms can become perceptible against its background.</p>

<p>Much like the dynamic of the <em>blind spot</em>, the <em>medium/form distinction</em> makes perception possible only if it remains unperceived in the act itself.
The difference is the very condition that must vanish for its result to become visible.
This <em>latent</em> structure is only revealed when a theory moves to the level of second-order cybernetics, observing the perceiving observer rather than merely perceiving alongside them.</p>

<p>Luhmann typically follows this with a generalization of the <em>medium/form distinction</em>.
The individual sounds or letters of a language (loosely coupled, “granular”) serve as the medium from which sentences (firm forms) are constructed.
Similarly, money serves as the medium from which concrete prices (forms) are forged.
Under certain conditions, forms can themselves become the medium for a higher level.
Words, which are forms relative to letters, become the medium from which sentences are formed.
A rhythm or sequence of clicks (medium) becomes a unified rhythm (form), and an accelerated rhythm (medium) can become a sustained tone (form).
At each new level of complexity, the underlying difference is once again rendered invisible.</p>

<p>Here, however, Luhmann pulls us back again.
While generalizing this concept demonstrates its theoretical reach, it also risks leading us away from actual epistemology (much like the earlier biological/psychic detour).
Therefore, he does not pursue it further in this text.
The decisive takeaway remains the basic assumption: without a physically anchored medium/form difference, no cognizing system could develop in the first place.
It would remain forever confined to bare, immediate contact events at its own boundary, devoid of any spatiotemporal distance to the environment—and thus entirely lacking the capacity for distanced, mediated perception.</p>

<p>Crucially, one must add that the medium is never exhausted by the creation of forms.
It must regenerate (just as the air is not “used up” when a sound is made, or money remains available for circulation after a price is set).
Furthermore, the form is always stronger and more assertive than the medium—Luhmann hints that a not-yet-fully-understood “secret” rationality may hide within this dynamic, though he leaves it unexplored here.
Finally, what acts as a mere “medium” on one level (invisible and latent) can itself become visible as a form if a sufficiently refined medium of observation is introduced. Air molecules, for instance, are merely an invisible medium to the naked ear; but under an electron microscope or a particle detector, they become observable forms in their own right.</p>

<!-- TODO -->

<p>Taking this logic to its absolute conclusion leads directly to quantum physics!
Where classical physics tacitly assumed that the measuring instrument itself always remains a neutral, invisible medium—it does not ‘disturb’ what is measured, it simply shows what the object is ‘truly’ like, independently of the measurement.
But because quantum mechanics utilizes measuring instruments so refined that earlier “media” become observable forms themselves (cf. the measurement problem and the observer effect<sup id="fnref:15" role="doc-noteref"><a href="#fn:15" class="footnote" rel="footnote">15</a></sup>), it ultimately describes little more than physicists observing physicists.
Short-wavelength light inevitably carries a great deal of momentum per photon—short wavelength means high energy, which means a high momentum for the individual light quantum—and when this high-energy photon bounces off the electron to make its position visible, it inevitably imparts a substantial, no longer negligible kick to the electron’s momentum.
The more one refines the measuring instrument to determine the position more precisely (shorter wavelength), the more strongly and inevitably you disturb the electron’s momentum in the process.
Physics itself is forced to theorize the observation situation as part of its own subject matter; it can no longer return to a neutral, uninvolved standpoint.
It is a theory existing purely at the level of second-order cybernetics.</p>

<blockquote>
  <p>In a sense, <em>hard sciences</em> becomes really <em>hard</em> when it pushes to its most outer boundary to hit a problem that <em>soft sciences</em> have to deal with from the beginning.<sup id="fnref:15:1" role="doc-noteref"><a href="#fn:15" class="footnote" rel="footnote">15</a></sup></p>
</blockquote>

<p>Correspondingly, it describes reality as fundamentally indeterminable (as seen in Heisenberg’s uncertainty principle or the collapse of the wave function).
Luhmann immediately defuses this ontological panic, however: this does not necessarily mean that reality itself is indeterminable!
It simply means that the observing-of-observing—the measuring and the predicting of measurements—produces new forms that turn into new media.
This claim rests entirely on the recursive application of the medium/form logic; it is not a metaphysical proof of absolute, objective indeterminacy.</p>

<p>Interestingly, this recursive <em>form-becomes-medium</em> structure is not just a theoretical abstraction.
We experiment with it elsewhere, such as in modern, self-reflexive poetry, which consciously plays with forms that turn themselves into media.
Yet, even though forms can recursively become media, this does not imply that the world’s self-observation could ever manage entirely <strong>without</strong> a latent medium/form difference.
The blind spot wanders, but it never completely disappears.</p>

<p>Luhmann’s conclusion is stark: while cognition absolutely requires a structurally suitable environment (one featuring medium/form differences, temporal discontinuities, and so on), we cannot infer that cognition therefore adapts itself to reality in the sense of increasing alignment or correspondence.
Even less tenable is the optimistic approach of older cybernetics, which attempted to explain improved performance (better cognition) and environmental adaptation (better correspondence with reality) using one and the same model.
For Luhmann, conflating internal systemic complexity with external alignment is an epistemological dead end.</p>

<blockquote>
  <p>In any case, scientific research, seen in the context of ecology, rather gives the opposite impression. The deviation from what appears to be given keeps increasing, since cognition, in ever bolder leaps, corrects itself.<sup id="fnref:16" role="doc-noteref"><a href="#fn:16" class="footnote" rel="footnote">16</a></sup> – <a class="citation" href="#luhmann:1988">(Luhmann, 1988)</a></p>
</blockquote>

<p>Science does not converge gently upon a fixed truth; rather, it drifts ever further, and ever more boldly away from any assumed starting point. 
Cognition projects its own distinctions onto a reality that is itself entirely devoid of distinctions.
In doing so, it seizes a freedom that is nowhere “provided for” or guaranteed by reality itself—<strong>a radically self-produced, rather than externally legitimated freedom</strong>.</p>

<p>Even the attempt to label this freedom as “causeless” (spontaneous or uncaused) would merely be another judgment regarding causation and attribution.
It would be just one more act of cognition, entirely unable to escape the inescapable loop of its own constructedness.</p>

<p>Therefore, if cognition does not aim at correspondence with a pre-given reality, but functions instead as an ongoing process of self-reinforcing deviation, a final, forward-looking question emerges: what kind of order (rather than what truth) can still be achieved at all within such a process?</p>

<h2 id="6-world-reality-meaning">6. World, Reality, Meaning</h2>

<p>Reality, insofar as it remains unknowable, would permit no cognition at all if it were entirely <em>entropic</em> (in the thermodynamic sense of being completely disordered, structureless, and uniform).
The medium/form differences and temporal discontinuities discussed earlier presuppose precisely that the environment is not entirely structureless.</p>

<p>But this is exactly what cognition cannot directly formulate.
Even if we suspect that the environment requires a certain structure to make cognition possible, cognition cannot project this insight “outward” as a distinction of its own.
Any such formulation would simply be another internal achievement of the system, not a genuine capture of the environment itself.
The attempt to name one’s own external conditions from the “outside” collapses immediately into just another internal construction.
Cognition is completely and exclusively distinction-based construction, and nothing more.
Structurally, it knows nothing outside itself that could ever correspond to it.</p>

<p>Nevertheless, we can suspect—though never truly know—that in this “outside” (which cognition marks via the distinction between self-reference and other-reference), there really are structural conditions for cognition: temporal and material discontinuities, differences in speed, and varied structural couplings.
In other words, exactly the phenomena we have just been discussing (medium/form, Eigen-time).
The operative word here is “suspect”: it remains pure speculation, not secured knowledge.
To sharpen the point: precisely because these real conditions may exist, cognition cannot use them directly as its own distinctions.
Renouncing direct access to reality is the absolute precondition for operational closure. If cognition were to “use” these conditions directly—claiming to grasp them immediately—it would undercut its own closure entirely.</p>

<p>Does one, then, need special, “distinction-free” concepts—concepts not definable through the usual pairing of opposites, and therefore necessarily paradoxical (like “world”)—in order to manage this unreachable boundary?
Historically, the concept of God took on exactly this role, “absorbing” the paradox, and for some, this may remain the most satisfying solution to this day.
Luhmann cautiously and non-committally offers a secular alternative.
With a playful sidelong glance at the Christian Trinity, he proposes three concepts that structurally fulfill this exact function: “<em>World</em>” names the unity behind the difference system/environment; “<em>reality</em>” names the unity behind the difference cognition/object; and “<em>meaning/sense</em>” (Sinn) names the unity behind the difference actuality/possibility.</p>

<p>None of these three are mere objects alongside other objects; rather, they are names for the presupposed wholeness from which a given distinction first carves out its two sides.
These concepts are characterized by the fact that they cannot be negated from the outside.
Any attempt to do so automatically falls back into them. Whoever asserts “there is no world” must perform this statement somewhere, and that somewhere is inevitably, once again, the world.
One cannot negate the world from outside the world.
Similarly, asserting “there is no reality” is itself a real statement, a real speech act.
The negation performatively confirms the very thing it seeks to deny.
Finally, the statement “meaning (Sinn) does not exist” must, in order to be understood at all, itself be meaningful (Sinn haben).
And even if it really “makes no sense” (in the colloquial sense of being nonsensical), it still makes sense precisely in the sense that it makes no sense!</p>

<p>Either way, the negation cannot escape the concept.
In sum, these totalizing concepts cannot be defined like ordinary concepts through contrast with an opposite.
“World vs. non-world” does not work, nor does “meaningful vs. meaningless” at this absolute level.
They can only be defined via the specific distinction whose underlying unity they designate.</p>

<p>But not just any distinction is suited to serve as the basis for such a totalizing boundary concept; only a very select few are. This once again underscores the evolutionary improbability of cognition.
Moreover, these concepts can only be derived from within cognition itself, never from an external standpoint.</p>

<p>All three primary distinctions (<em>system/environment</em>, <em>cognition/object</em>, <em>actuality/possibility</em>) are strictly asymmetrical.
The phenomenon of <em>re-entry</em> becomes starkly apparent here: a distinction “re-enters” one of its two sides when that side reproduces the original distinction within itself.
Luhmann’s crucial point is that for all three distinctions, this re-entry works on only one side, never on both.
Only the system (not the environment) can use “world” as its orienting concept; only the system can hold both itself and its environment together “in view,” thereby folding the original distinction back into itself.
The environment possesses no operations of its own to achieve this.
Similarly, only cognition (not the object) can use “reality” as an overarching concept; the idea that reality encompasses both sides arises solely within the performance of cognition, never independently of it. 
Finally, meaning (Sinn) functions exclusively on the side of actuality.
Only an operation that has actually been executed can point toward a horizon of further possibilities (whether real, conceptual, or purely fictional).
Bare possibility itself has no operations that could establish such a reference.</p>

<p>Structurally, all three concepts serve the exact same purpose: the resolution of the paradox that something must simultaneously be One (the unity) and Two (the difference).
However, this paradox does not exist “objectively”; it is only ever perceived as a paradox by an observer. 
If one evaluates a theory not by its specific content but by its function (in this case, paradox-resolution), one can begin to ask about functionally equivalent alternatives.
This functional equivalence is precisely what allows Luhmann to replace “God” with “world,” “reality,” and “meaning,” without necessarily claiming that the secular concepts refute the theological one.
Both solve the exact same structural task, merely by different means.
If one construes the paradox of observation personally rather than abstractly—that is, if one posits the paradox itself as an observing subject—one lands directly back at the question of God, exactly as Cusanus did.
Luhmann leaves both paths open, refusing to dogmatically commit to either.</p>

<p>He reassures us, however, that all this highly abstract paradox-wringing does not interfere with the level of everyday operating.
This should be intuitively clear: a consciousness or a communication system simply distinguishes, indicates, observes, and describes without ever stumbling over these philosophical paradoxes in daily life.
If we did, our everyday existence would be exhausting to the point of paralysis!
At the level of pure operation, there is nothing deeper to say than: “it happens when it happens.”
No metaphysical foundation is required to simply keep going.</p>

<p>It is only when one wishes to grasp, distinguish, and conceptually understand this occurrence that one must switch into the position of the (second-order) observer.
And that, Luhmann states in closing, is the actual, enduring task of epistemology: not merely to co-operate, but to observingly describe the nature of operating itself.</p>

<h2 id="7-language-as-coupling">7. Language as Coupling</h2>

<p>Whatever conditions must be met for cognition, the proof that they are fulfilled lies simply in the fact that cognition actually takes place.
It has nothing extra to prove.
It simply does what it does, and that alone suffices as proof of its own possibility.
The question of whether cognition is possible at all is thus settled trivially: it is, because it happens. 
The genuinely interesting question concerns the enhancement of cognition—how cognitive achievements can become more complex and extensive; and, increasingly today, <strong>the ecological question of whether this continuous enhancement remains environmentally sustainable</strong>.</p>

<p>Classical theories (such as Piaget’s “assimilation,” traditional representation theories, or adaptation models) inherently built environmental compatibility into the very concept of cognition.
They assumed that <strong>cognition, by its very nature, inevitably tended toward a “fit” with the environment</strong>. Even certain cybernetic theories retain this optimistic <em>notion of adaptation</em>—precisely the “evolutionary optimism” that Luhmann sharply criticized earlier.
Instead of asking about adaptation, Luhmann asks how a closed system builds up its own internal complexity.
Enhancement, therefore, occurs exclusively from within, not through a better fit with the outside world.</p>

<p>At this juncture, it is natural to think of language.
As discussed, Maturana ties his concept of the observer directly to linguistic capacity.
Ernst von Glasersfeld (another central figure of radical constructivism) similarly views linguistic research as the empirical proving ground for the entire theory.
This alliance seems fitting, given that linguistics, ever since Ferdinand de Saussure, has largely abandoned the idea that signs refer directly to external things (outer reference).
Saussure’s structuralism demonstrates—and today’s large language models empirically suggest much of the same—that the value of a linguistic sign arises purely differentially.
A sign is defined exclusively by its relation to other signs within the closed system of language, not by direct reference to the external world.</p>

<p>This <em>structuralist</em> insight acts as a powerful linguistic anticipation of the core constructivist idea.
Yet this seemingly convenient alliance conceals a fundamental problem: cognitive operations are entirely different depending on the specific type of system carrying them out.
To remain theoretically precise, one must strictly distinguish between consciousness (which thinks), communication (which communicates)—and, we might add today, computing machines (which compute).</p>

<blockquote>
  <p>Precisely this ([convenient alliance of the constructivist basic idea]), however, conceals a problem. For the operations of cognition are, depending on the kind of system that carries them out, entirely different. One must distinguish between psychic and social systems, between consciousness as it currently operates and communication. – <a class="citation" href="#luhmann:1988">(Luhmann, 1988)</a></p>
</blockquote>

<p>All of these systems—consciousness, communication, and perhaps even algorithmic computation—can utilize language.
They use it for thinking, for communicating, and potentially for computing.
For all of them, the high degree of complexity we are familiar with, only becomes possible through language in the first place.
Nevertheless, they remain operationally closed and entirely separate systems.
<strong>The shared use of language by no means fuses them into a single system</strong>.
There is absolutely no operative overlap.
A word occurring within a thought connects to other thoughts according to the internal connective rules of consciousness.
That exact same word, when deployed in a communication, connects to other communications following an entirely different set of rules.
A thought cannot directly connect to a communication, nor can a communication directly connect to a thought. 
Thus, the exact same linguistic material functions within each system according to its own strictly incompatible conditions of connection.</p>

<p>According to Luhmann, one must neither ignore nor underestimate language, but one must realize that language is not a system that carries out cognition as a real operation.
In fact, <strong>it is not a system at all</strong>.
Rather, <strong>it provides the structural coupling between these separate systems</strong>.
That is its function—no more, and no less.
Language possesses its own medium (sounds, optical signs), which is then formed into words and, subsequently, into sentences. Once again, we encounter a recursive medium/form logic.
This highly specific differentiation makes language available to the participating systems as a <strong>jointly usable resource</strong>.
Consciousness (in linguistic thinking), communication (in successive sentences), and perhaps computation (in linguistic processing) can each extract and forge their own separate forms from it.</p>

<p>This represents a central theoretical break.
As soon as one defines systems strictly through their own boundary-drawing operations—as Luhmann does—rather than through a vague sense of “belonging together,” the convenient equation of linguistic theory with constructivism (as seen in Maturana and von Glasersfeld) can no longer be maintained.
Language is not a third system; it is merely a coupling mechanism between two (or, perhaps soon, three) operationally closed domains.</p>

<p>Nevertheless, this coupling function remains absolutely central.
Language fascinates consciousness, binding its attention to a conspicuous repertoire of acoustic or written forms.
It ensures that communication keeps going while simultaneously keeping enough consciousness activated to carry that communication forward.
It certainly constrains the degrees of freedom of consciousness during a communicative act, but never completely.
One can still perceive incidental things in the background, entertain unspoken side-thoughts, and, crucially, deliberately lie using language (or, in the case of computation, produce “hallucinations”?).
This proves that structural coupling never leads to complete fusion.<sup id="fnref:17" role="doc-noteref"><a href="#fn:17" class="footnote" rel="footnote">17</a></sup>
Consciousness remains strictly independent, even in the very midst of communication.</p>

<p>Before the invention of writing, communication was entirely dependent on the (often overestimated) memory of the participating consciousnesses just to be continued.
Conversely, without the capacity to imagine thoughts linguistically (whether in sound or script), consciousness would remain hopelessly bound to what is immediately and presently perceived.
It would be bound to such a degree, in fact, that Luhmann asks whether, in that case, one could still speak of “consciousness” in the full sense at all.</p>

<p>Luhmann’s conclusion is clear: a coupling between consciousness and communication that allows for a continuous growth in complexity can only be explained via language.
Importantly, however, language itself does not “speak”; it is not an independently acting system.
It is always the concrete systems—consciousness, communication, and perhaps computing machines—that use it.
Language merely provides the medium/form difference.
The concrete realization of cognition is subject to many further constraints that cannot be explained linguistically, but only psychologically (for consciousness), sociologically (for communication), or perhaps information-theoretically (for computing machines).
These constraints concern, above all, the conditions of autopoietic closure itself and their internal consequences.
For these questions, one needs psychology, sociology, and computer science—not just linguistics.</p>

<p>If one replaces the traditional, object-like definition of systems (viewed as particularly densely interconnected clusters of things) with the fundamental difference between system and environment, an entirely new theoretical architecture emerges.
The central questions shift: Which operations close off a system? (rather than: what things hang together?). And: What form does the connection take (now newly conceived as structural coupling) once this closure has already been achieved?
The old, vague concept of “connection” is not discarded, but made precise.
It becomes a result of the system’s boundary-drawing, no longer its defining precondition.
This paradigm shift has far-reaching, still barely foreseeable consequences; we are, truly, only at the beginning (bearing in mind that Luhmann wrote this in 1988).
Cognition is possible because systems operatively close themselves off at the level of their distinguishing and indicating, thereby rendering themselves indifferent to the excluded environment.</p>

<p>Finally, Luhmann offers a few closing pointers to guard against ultimate misunderstandings.
First, the insight into operative closure does not mean that cognition is “unreal” (this is a firm rejection of any nihilistic or illusionist reading of constructivism).
Second, it does not imply that there can be absolutely no correspondences between a system’s differentiating operations and the environment.
If that were the case, the system would lose all purchase on its environment and continually dissolve into it, rendering cognition impossible from the start.</p>

<p>I would even add that we, as living and psychic systems, are so deeply at home in our environment because we are not a mere assembly of parts—not engineered within and for a specific environment.
We are not only the result of a long process of differentiation from what we call ‘home,’ but its very identity. To word it a little more mystically: we are a certain cut of the whole we can only observe by cutting it.</p>

<p>Luhmann then deliberately closes the text not with a radical, solipsistic declaration that “there is nothing but construction,” but with a finely balanced position: there is no direct correspondence in the classical, mirroring sense, but there is also no complete arbitrariness or detachment from reality. 
There is only just enough <strong>structural fit</strong> for the system to assert and sustain itself as a system against its environment.
A system is closed with respect to its operations, but it remains structurally coupled, physically embedded, and constantly susceptible to irritation by its environment. 
This very coupling presupposes a certain ‘fit’.</p>

<h2 id="8-reality-as-inconsistency-solution">8. Reality as Inconsistency Solution</h2>

<p>To summarize Luhmann’s <em>operational constructivism</em>:
We assume that systems and their respective environments are real, meaning that they genuinely exist.
We concede that this assumption can never be definitively proven from the inside, because there is no direct access to the environment.
Anything entering cognition is entirely constructed by cognition—it is a self-generated performance (Eigenleistung) of the system.
Yet the fact that our knowledge—including our observation of living, psychic, and social systems—is constructed, mediated, and fallible does not imply that what we thereby gain knowledge of is unreal.</p>

<p>Asking then about the conditions of the possibility for operational closure leads one to highly selective and improbable mechanisms, most notably <em>autopoiesis</em>.
All existing systems must continuously reproduce their own operations; otherwise, they would dissolve back into their environment.</p>

<p>Assuming (some) systems observe, it seems logical that (some) systems can observe observations.
For example, my mind can observe itself.
I can think about my thoughts; second-order observation, in this case, can thus be assumed.
I might also be able to observe the observations of a social system, that is, of something that lies on the unmarked side of my re-entry.</p>

<p>We can conclude that there is empirical evidence that psychic and social systems use second-order observation, but like any such evidence, it is already cognition that comes to such conclusion.
If I indicate something that I observe as “an observer”, I am applying my own distinction of what counts as an observation.
But if treating such an event as an “observation” allows my own system to successfully continue its autopoiesis, handle irritations, and remain connectable, then the categorization is functionally validated within my system.</p>

<p>We can then ask why (some) systems develop the ability to observe observations.
Before doing that, we should clarify that not all systems are observing systems.
While complex systems (like consciousness and communication) utilize second-order observation to manage their boundaries, simpler autopoietic systems—such as biological cells—maintain their closure through blind, structural couplings and biochemical reactions without ever observing observations.
Importantly, we departed from Maturana’s claim that life is cognition.</p>

<blockquote>
  <p>So first the system produces a difference of system and environment, and then it learns to control its own body and not the environment to make a difference in the system. So cognition then becomes a secondary achievement in a sense, tied to a specific operation which, I think, is that of making a distinction and indicating one side and not the other. It’s an explosion of possibilities, if you always have the whole world present in your distinctions. – Luhmann (in <a class="citation" href="#hayles:1995">(Hayles et al., 1995)</a>)</p>
</blockquote>

<p>First the system produces its own difference.
That is autopiesis and operational closure but <em>not</em> an act performed by anyone or anything, but an effect.
It is only after closure that Luhmann wants to locate distinguishing-and-indicating in the Spencer-Brown sense: selecting one side (the body) as the reference point for continuing operations, orienting itself by it, using the distinction rather than merely being its effect.
Furthermore, we depart from the claim that observation is linguistic—according to Luhmann, it is pre-linguistic.<sup id="fnref:18" role="doc-noteref"><a href="#fn:18" class="footnote" rel="footnote">18</a></sup></p>

<p>Over many years, psychic and social systems irritated each other in a <em>structural drift</em> (evolution).
In the case of minds, a biological organism reached a level of systemic complexity high enough that an internal, self-referential loop of consciousness emerged.
While the mind is operationally closed, it is structurally coupled with its biological substrate (the brain/body) and its environment.
The physical body and its nervous system absorb environmental perturbations, which irritate the closed psychic system, prompting it to generate new thoughts.</p>

<p>Social systems and minds became highly dependent on each other, and a structural coupling co-evolved via language.
This prevented minds from remaining entirely trapped in immediate, momentary perceptions; language acts as a structural coupling mechanism between separate psychic systems and social communication.
In general, systems decouple from their environment not by stepping outside of themselves, but by building an internal, recursive network of operations (e.,g. thoughts, communication) so complex that it becomes indifferent to the environment—relying entirely on its own self-reproducing operations to make sense of whatever irritations leak through its boundaries.</p>

<p>No individual operation aims at this decoupling; it is an evolutionary, non-intentional byproduct of autopoiesis.
Thus, for any system, this is a highly improbable event.</p>

<p>We can surmise that minds and social systems observe other systems (which themselves are able to observe) in their environment to cope with the unformatted, unstructured complexity of their surroundings.
Ultimately, advanced systems rely on these observational practices to reproduce their internal structures and prevent themselves from collapsing into their environment.
Because every leading distinction creates an inherent blind spot (just as an eye cannot see itself seeing), a system that could neither reflect on its own operations nor observe the observations of others would be entirely unable to navigate or compensate for its own blindness.
In that sense, observation leads to second-order observation.</p>

<p>Systems are able to cope so effectively with their environment precisely because they are closed off.
The system does not need to know what the environment “truly” is in order to react to it.
Environmental events act as physical or biological perturbations (irritations) that trigger internal operations within the system.
The system responds to these triggers using its own internal structures.
And while the system is operationally closed, it is structurally coupled with its environment.</p>

<p>Again, this means the system and its environment have co-evolved a history of mutual compatibility.
In that sense, <strong>we as psychic system are deeply at home in our environment!</strong>
Over time, only those systems whose internal structures remain compatible enough with their environment continue to reproduce; incompatible ones simply cease to exist.</p>

<p>For a system to function well, its operations do not need to “correspond” to objective reality.
They only need to be functionally successful enough to keep the system alive and reproducing.
As we discussed, true and false thoughts function equally well to keep a consciousness running—the operational network does not require truth to operate.
Instead of getting closer to an absolute truth, a successful system builds up massive internal complexity.
Through second-order observation, error-handling, and language as a coupling mechanism, systems build elaborate internal models that allow them to navigate, anticipate, and manage their environment efficiently—all entirely within their own closed loops.
There is no mirror of reality required, only a sufficient <em>structural fit</em> for the system to assert and sustain itself as a system against its environment—to keep autopoiesis going.</p>

<p>What systems then encounter as resistance, cannot be reality that resists because resistance is internal.
Luhmann notes:</p>

<blockquote>
  <p>I think we should not abandon [Kant’s] idea of resistance, but we should relocate it into the system. It is the result of resolving an internal conflict—the result of the system’s operations resisting the operations of the same system. –  <a class="citation" href="#luhmann:1984">(Luhmann, 1984)</a></p>
</blockquote>

<p>And in a discussion with Katherine Hayles he points out that</p>

<blockquote>
  <p>[t]hen, if you use for a moment the idea that reality is tested by resistance—that’s Kant—how can you have external resistance if you cannot cross the boundary of the system with your own operations? You cannot touch the environment with your brain, and even if you touch it you feel something here [points to his head] and not there, and you make an external reality just to explain that you feel something here [points again] and not in other places on your body. So, finally, it’s always an internal calculation; otherwise, you should simply refuse the term ‘operational closure’. But if we have operational closure, we have to construct every resistance to the operations of a system against the operations of the same system. And reality then is just a form—or, to say it in other terms, things or objects outside are simply a form in which you take into account the resolution of internal conflicts. – Luhmann (in <a class="citation" href="#hayles:1995">(Hayles et al., 1995)</a>)</p>
</blockquote>

<p>Reality—which, in the old tradition, was the invisible side of a thing (<em>res</em>)—now emerges if you have inconsistency in your <em>operations</em>.
It is just the acceptance of solutions for inconsistency problems—just what a system calls “reality” when a contradiction-handling operation succeeds.</p>

<h2 id="literature">Literature</h2>

<ol class="bibliography"><li><span id="luhmann:1988">Luhmann, N. (1988). <i>Erkenntnis als Konstruktion</i>. Bern: Benteli.</span></li>
<li><span id="glasersfeld:1995">Glasersfeld, E. von. (1995). <i>Radical Constructivism: A Way of Knowing and Learning</i>. Falmer Press.</span></li>
<li><span id="downing:2004">Downing, L. (2004). George Berkeley. In E. N. Zalta &amp; U. Nodelman (Eds.), <i>The Stanford Encyclopedia of Philosophy</i> (Spring 2011). Metaphysics Research Lab, Stanford University. https://plato.stanford.edu/archives/spr2026/entries/berkeley/</span></li>
<li><span id="maturana:1987">Maturana, H. R., &amp; Varela, F. J. (1987). <i>The Tree of Knowledge</i>. Shambhala.</span></li>
<li><span id="luhmann:1985">Luhmann, N. (1985). Die Autopoiesis des Bewußtseins. <i>Soziale Welt</i>, 402–446.</span></li>
<li><span id="brown:1969">Spencer-Brown, G. (1969). <i>Laws of Form</i>. London: Allen and Unwin.</span></li>
<li><span id="foerster:2003">von Foerster, H. (2003). Cybernetics of Cybernetics. In <i>Understanding understanding: Essays on cybernetics and cognition</i> (pp. 283–286). Springer New York. https://doi.org/10.1007/0-387-21722-3_13</span></li>
<li><span id="goedel:1931">Gödel, K. (1931). Über formal unentscheidbare Sätze der Principia Mathematica und verwandter Systeme I. <i>Monatshefte Für Mathematik Und Physik</i>, <i>38</i>(1), 173–198. https://doi.org/10.1007/BF01700692</span></li>
<li><span id="turing:1937">Turing, A. M. (1937). On computable numbers, with an application to the Entscheidungsproblem. <i>Proceedings of the London Mathematical Society</i>, <i>s2-42</i>(1), 230–265. https://doi.org/https://doi.org/10.1112/plms/s2-42.1.230</span></li>
<li><span id="bucker:2024">Bücker, S. (2024). Gödels Unvollständigkeitssätze. In S. Bücker (Ed.), <i>Über die Möglichkeit von Freiheit im Kalkül: Ein Vergleich von Kants Freiheitsbeweis und Gödels Unvollständigkeitssätzen</i> (pp. 9–29). Springer Fachmedien Wiesbaden. https://doi.org/10.1007/978-3-658-46386-1_2</span></li>
<li><span id="friston:2010">Friston, K. (2010). The free-energy principle: a unified brain theory? <i>Nature Reviews Neuroscience</i>, <i>11</i>(2), 127–138.</span></li>
<li><span id="wittgenstein:1921">Wittgenstein, L. (1921). <i>Logisch-philosophische Abhandlung</i>.</span></li>
<li><span id="harnad:1990">Harnad, S. (1990). The symbol grounding problem. <i>Physica D: Nonlinear Phenomena</i>, <i>42</i>(1), 335–346. https://doi.org/10.1016/0167-2789(90)90087-6</span></li>
<li><span id="maturana:2000">Maturana, H. R. (2000). <i>Biologie der Realität</i>. Suhrkamp Verlag.</span></li>
<li><span id="kuntze:2009">Schönwälder-Kuntze, T., Wille, K., &amp; Hölscher, T. (2009). <i>George Spencer Brown: Eine Einführung in die "Laws of Form"</i> (2nd ed., pp. X, 314). VS Verlag für Sozialwissenschaften. https://doi.org/10.1007/978-3-531-91964-5</span></li>
<li><span id="gabriel:2013">Gabriel, M. (2013). <i>Warum es die Welt nicht gibt</i>. Ullstein.</span></li>
<li><span id="luhmann:1998">Luhmann, N. (1998). <i>Die Gesellschaft der Gesellschaft</i> (p. 1164). Suhrkamp.</span></li>
<li><span id="foerster:2003b">von Foerster, H. (2003). Objects: Tokens for (Eigen-)Behaviors. In H. von Foerster (Ed.), <i>Understanding Understanding: Essays on Cybernetics and Cognition</i> (pp. 261–271). Springer New York. https://doi.org/10.1007/0-387-21722-3_11</span></li>
<li><span id="husserl:1928">Husserl, E. (1928). Vorlesungen zur Phänomenologie des inneren Zeitbewusstseins. <i>Jahrbuch Für Philosophie Und Phänomenologische Forschung</i>, <i>9</i>, 367–498.</span></li>
<li><span id="zoennchen:2025">Zönnchen, B., Dzhimova, M., &amp; Socher, G. (2025). From intelligence to autopoiesis: rethinking artificial intelligence through systems theory. <i>Frontiers in Communication</i>, <i>Volume 10 - 2025</i>. https://doi.org/10.3389/fcomm.2025.1585321</span></li>
<li><span id="hayles:1995">Hayles, K., Luhmann, N., Rasch, W., Knodt, E., &amp; Wolfe, C. (1995). Theory of a Different Order: A Conversation with Katherine Hayles and Niklas Luhmann. <i>Cultural Critique</i>, <i>31</i>, 7–36. https://doi.org/10.2307/1354443</span></li>
<li><span id="luhmann:1984">Luhmann, N. (1984). <i>Soziale Systeme: Grundriß einer allgemeinen Theorie</i>. Suhrkamp.</span></li></ol>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:1" role="doc-endnote">
      <p>Transcendentally ideal means: space, time, and, mediated through them, all objects of experience (the appearances) are not properties of things as they exist in themselves, independently of any consciousness. Rather, they are forms that the cognizing subject itself contributes to experience, that is, conditions without which no experience would be possible for us at all. “Ideal” here means: dependent on the subject, not a property of a subject-independent reality. Empirically real, by contrast, means: within experience, that is, empirically considered, space, time, and the objects within them are entirely real. They are not an illusion, not a mere subjective fancy, but objective, intersubjectively verifiable, causally effective. A tree, a stone, a planet are real, objective objects within experience, not private phantasms. Because for Kant space and time are contributed by the subject itself (transcendentally ideal), we can know with certainty and a priori that every possible object of experience will conform to them. And because these forms are then consistently applied to all possible objects of experience, the resulting empirical world is thoroughly real and objective (empirically real), not merely private or illusory. The <strong>thing in itself</strong> thus falls entirely outside this whole scheme. It is neither empirically real (it is not located in space and time at all, is not an object of possible experience) nor transcendentally ideal (by definition it does not depend on the subject, but is supposed to exist independently of any cognitive achievement). Precisely for this reason it remains unknowable for Kant: it falls outside both categories that, for Kant, are what make knowability possible in the first place. <a href="#fnref:1" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:2" role="doc-endnote">
      <p>One can, I think, already draw a bridge here to modern theories in cognitive science (cf. <a class="citation" href="#friston:2010">(Friston, 2010)</a>), specifically in the sense that nervous systems, once they reach a sufficient level of complexity, observe and validate their own distinctions—proving themselves, as it were, self-consciously. <a href="#fnref:2" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:3" role="doc-endnote">
      <p>Gödel proved that the formal systems of mathematicians are incapable of fully resolving mathematical decision problems beyond a certain level of complexity. Analogously, Wittgenstein argued in his <em>Tractatus Logico-Philosophicus</em> <a class="citation" href="#wittgenstein:1921">(Wittgenstein, 1921)</a> that it is impossible to speak about that which lies behind what is said. <a href="#fnref:3" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:4" role="doc-endnote">
      <p>Maturana strongly objected to abstracting the concept of autopoiesis in this manner, insisting it should only be applied to <em>living systems</em>. <a href="#fnref:4" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:5" role="doc-endnote">
      <p>Given the empirically measured performance of current language models, one might wonder whether this performance could serve as empirical evidence for the closure of language itself. That is, whether the so-called <em>grounding problem</em> <a class="citation" href="#harnad:1990">(Harnad, 1990)</a> might be resolved through structural coupling, and whether this closure provides a clue as to why communication with essentially incomprehensible machines works at all. As we will see, however, Luhmann rejects the idea that language itself can be classified as a system. <a href="#fnref:5" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:6" role="doc-endnote">
      <p>In German, many—including Luhmann—translate this as “Bezeichnen.” However, according to <a class="citation" href="#kuntze:2009">(Schönwälder-Kuntze et al., 2009)</a>, “Hinweisen” is likely more accurate to Spencer-Brown’s original intent. <a href="#fnref:6" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:7" role="doc-endnote">
      <p>This is precisely why this theory interests me: it offers a framework for searching for further cognition-performing systems. In particular, it allows us to ask whether technical systems, such as AI, can truly cognize, and what rigorous preconditions would have to be met. <a href="#fnref:7" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:8" role="doc-endnote">
      <p>This condition is central to the question of whether AI systems can cognize. It is not enough that a distinction and an indication merely occur; rather, this act must take place “within” the system, belonging to the network of recursive operations that (re-)produce the system’s boundary. Otherwise, one could just as easily classify a thermostat as a cognizing system. <a href="#fnref:8" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:9" role="doc-endnote">
      <p>Here, a confrontation with the works of Markus Gabriel would be interesting, as he actively pushes back against constructivism. He arrives at the conclusion ‘that the world does not exist’ <a class="citation" href="#gabriel:2013">(Gabriel, 2013)</a>, though—motivated by Frege—he does so through a rather formal-logical argumentation, in the style of a set-theoretic paradox. Because ‘the world’ does not exist, there are instead countless fields of sense (Sinnfelder) that exist really, objectively, and mind-independently. In other words, for Gabriel, the rejection of the one big world clears the way for an overflowing multiplicity of actually existing things (numbers, fictional characters, social facts, physical objects—everything is real, just each within its own specific field). Luhmann, by contrast, draws the exact opposite, explicitly constructivist consequence from the very same structural observation. Gabriel postulates a realism of being, while Luhmann postulates a realism of doing. <a href="#fnref:9" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:10" role="doc-endnote">
      <p>One hears an echo here of the tradition of <em>enactivism</em> founded by Varela. <a href="#fnref:10" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:11" role="doc-endnote">
      <p>Like Kant, Luhmann strictly distinguishes between the brain and consciousness! <a href="#fnref:11" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:12" role="doc-endnote">
      <p>Within a formal calculus, classical logical strategies may yield mathematically correct solutions. However, they fail to explain how an actually operating system (e.g., a consciousness, a communication system, or a scientific discipline) navigates this problem during its ongoing operations. The theory of types is an artificial restriction on language introduced from the outside—a logician’s retrospective stipulation—rather than a model of what a brain or a communication network actually does in the moment to keep operating despite the paradox. Furthermore, the level-distinction itself must be drawn, subjecting it to the same infinite regress: does the distinction between object language and metalanguage belong to the object level or the metalevel? The paradox is merely shifted higher, never truly resolved. <a href="#fnref:12" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:13" role="doc-endnote">
      <p>Luhmann likely derives this from the <em>Laws of Form</em> <a class="citation" href="#brown:1969">(Spencer-Brown, 1969)</a>, where Spencer-Brown demonstrates how a form that “re-enters” itself yields not a static logical contradiction, but a dynamic oscillation between two values across successive moments in time. Luhmann maps this exact temporal pattern onto psychic and social systems. <a href="#fnref:13" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:14" role="doc-endnote">
      <p>It has to be shown that communication, unlike consciousnesses or organisms, knows no natural, principled boundary. In other words, that any communication is in principle connectable to any other communication (via translation, trade, diplomacy, media), and that there is thus no insurmountable structural separation. <a href="#fnref:14" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:15" role="doc-endnote">
      <p>“Soft here does not mean imprecise, unrigorous, or merely subjective opinion—on the contrary, quantum mechanics is perhaps the most precisely experimentally confirmed theory in the entire history of science. <a href="#fnref:15" class="reversefootnote" role="doc-backlink">&#8617;</a> <a href="#fnref:15:1" class="reversefootnote" role="doc-backlink">&#8617;<sup>2</sup></a></p>
    </li>
    <li id="fn:16" role="doc-endnote">
      <p>The distinction “better cognition” versus “better adaptation” already points toward a consideration of ecological communication, and toward the problem of viewing anthropogenic climate change accordingly, of counteracting it, and of adapting to it. <a href="#fnref:16" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:17" role="doc-endnote">
      <p>At this point, I would like to point to our discussion—conducted elsewhere—of the loose and tight coupling of technical machines <a class="citation" href="#zoennchen:2025">(Zönnchen et al., 2025)</a>. The core argument is this: large language models appear to enter into a loose coupling with psychic and social systems, whereas ordinary computing machines are strictly and tightly coupled! Here, then, lies a crucial, observable difference. <a href="#fnref:17" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:18" role="doc-endnote">
      <p>Negation is not built into the act of distinguishing itself; it is a product of language specifically, and it is there for a functional reason: it keeps the system open. Negation, in a Luhmannian sense, is a social technology for preventing communication from being forced toward one predetermined result, and the identity of the reference has to be secured before the yes/no coding can do its work. <a href="#fnref:18" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content><author><name>Benedikt Zönnchen</name><email>benedikt.zoennchen@web.de</email></author><category term="Systems Theory" /><category term="Philosophy" /><category term="AI" /><summary type="html"><![CDATA[0. Toward a New Epistemology]]></summary></entry><entry><title type="html">A Case for Systems Theory in CS Education</title><link href="https://bzoennchen.github.io/Pages/2026/06/28/why-systems-theory.html" rel="alternate" type="text/html" title="A Case for Systems Theory in CS Education" /><published>2026-06-28T00:00:00+02:00</published><updated>2026-06-28T00:00:00+02:00</updated><id>https://bzoennchen.github.io/Pages/2026/06/28/why-systems-theory</id><content type="html" xml:base="https://bzoennchen.github.io/Pages/2026/06/28/why-systems-theory.html"><![CDATA[<p>Imagine we are given the task of building a recommendation system. The requirements are clear: the system should suggest content to users that they are highly likely to click on. We define a metric. Let’s say, time spent on the platform. Then we optimize toward it. We test, iterate, deploy. The system works. The metric rises.</p>

<p>And then something happens that was in no specification document: people do not merely spend more time on the platform, but they change. Their worldview radicalizes gradually, because the system systematically favors polarizing content, since that generates stronger emotional reactions and therefore longer engagement. Social groups fracture. Adolescents develop anxiety disorders. Democratic discourse erodes, and in a distant country people are suddenly being hunted.</p>

<p>The story is real and well-documented. The system fulfilled its specification and still failed. One cannot even necessarily say it was misaligned, because it realized precisely the values its developers had put into it.
This discrepancy between technical success and systemic harm is therefore not an operational accident. It is not a case of “AI” acting autonomously or exerting influence on its own. The problem reveals itself as an epistemological one, that is, one that computer science as a discipline has so far addressed only inadequately. This text aims to explore why that is, and what a nearly forgotten intellectual program called <em>cybernetics</em> might have to contribute.</p>

<p>First we have to recognize that even though humans have always been technological, something has changed over the last few decades.
The products of computer science are no longer confined to data centers. They have grown deep into the structures of social life.
Algorithms, and increasingly, learning algorithms, co-determine which news people read, which candidates they see in job applications, what creditworthiness is assigned to them, what therapy options they are offered, and what ideas they develop. Software controls infrastructure that millions of people depend on every day.</p>

<p>This is no exaggeration and no dystopian narrative but the sober observation of a development that has taken place over the last three decades. Technical systems have become constitutive parts of social, psychological, and biological systems. They are part of a co-evolutionary <em>drift</em>, neither in control nor controllable in the strong sense.
Here <a class="citation" href="#luhmann:1998">(Luhmann, 1998)</a> points away from a critique of technology that sees it as a dominating force and instead insists that society becomes dependent on technology in an unplanned manner by engaging with it (die Gesellschaft lässt sich auf Technik ein).</p>

<p>Of course, in some sense this was always the case since even a simple automatic door opener influences social life but today’s systems differ in kind, not merely in degree: they are recursive, adaptive, and operate at a scale and speed that outpaces human observation and reaction.
They irritate our thinking, suggest how we communicate, how we organize ourselves, how we sleep, how we eat, how and whom we love.
What has also changed is the classification of organisms and technical systems with respect to their coupling strategy.
At the time of his writing, Luhmann argued that technology can be identified as realizing <em>strict couplings</em> whereas organisms and ecosystems avoid this form of coupling and tend towards a <em>loose coupling</em>.
Technology takes a messy, unpredictable world and forces a tight, invariant relationship between cause and effect.
Thus, for Luhmann, technology’s entire purpose is to exclude contingency (the possibility of things being otherwise) to guarantee a specific output.<sup id="fnref:1" role="doc-noteref"><a href="#fn:1" class="footnote" rel="footnote">1</a></sup>
We argued in <a class="citation" href="#zoennchen:2025">(Zönnchen et al., 2025)</a> that, especially with the advances in the development of large language models, this might no longer be the case.
But even a recommender system can be seen as realizing a <em>loose coupling</em>: The technical system relies on social feedback to reduce its own algorithmic complexity, while the social system relies on the technical system to sort through the overwhelming noise of the digital world. And because the technical system is structurally coupled to the unpredictable, loose nature of social communication, the strictly coupled outputs change second-by-second.</p>

<p>And yet, when we design (and use) technical systems, we mostly treat them as if they were self-contained machines—simple tools that cannot really alter our <em>autonomy</em> and <em>agency</em>, even though we know that this is not the case.
We specify inputs and outputs. 
Confidently, we define clear causal chains such that A leads to B and B leads to C.
Rather implicitly, we use system boundaries that separate the technical from the rest of the world; 
We optimize within those boundaries.
And whatever happens beyond them does not belong to our responsibility.</p>

<p>Rather than describing this as an individual failure of developers and engineers, it might be healthier to think of it as a structural consequence of the way we have learned to think as a discipline.
It is an effective way to solve a certain problem since a certain amount of ignorance is necessary to transform uncertainty into something we can manage, such that we are not paralyzed and can move on.
As Luhmann puts it: <strong>Technology constitutes an evolutionary achievement that operationalizes complexity reduction.</strong></p>

<h2 id="a-repressed-inheritance">A Repressed Inheritance</h2>

<p>Things were once different. In the decades following the Second World War, there was an intellectual movement that refused to accept precisely these boundaries. <em>Cybernetics</em>, which was founded by Norbert Wiener, Gregory Bateson, Heinz von Foerster, and others asked what control, feedback, information, and self-regulation mean, regardless of whether the system in question is a machine, an organism, a brain, or a society.
The early cyberneticians sat together at the same table. This included mathematicians, neurologists, anthropologists, economists, and engineers. The famous Macy Conferences (1946–1953) brought these disciplines into a conversation.</p>

<p>What became of this program? It did not fail. Instead, it was absorbed institutionally. The successor disciplines, such as control engineering, computer science, cognitive science, organizational theory, and operations research, each inherited and developed a part of the cybernetic legacy. But in this process of specialization, what had held it together was abandoned. The shared conversation became a series of monologues.</p>

<p>This is no criticism of specialization as such. <em>Functional differentiation</em>, i.e. the division into independent disciplines with their own methods, concepts, and communities of communication, was historically extraordinarily productive. It allowed for complexity reduction, sharper questions, cumulative knowledge. The computer science we know today would be unthinkable without this differentiation.
But every reduction of complexity comes at a price and what disappears from view does not cease to exist.</p>

<p>So should we go back in time?
Anyone who argues today for a return of cybernetics into the syllabus of computer science education must face a question: Is cybernetics not fundamentally compromised, particularly by its military history, by a vocabulary that turns the human being into a machine, by a proximity to control and steering that seems irreconcilable with a liberal, humanist conception of society? (We should also ask if this conception of society is still fruitful e.g. for a <a href="/Pages/2026/04/03/cruelty-and-solidarity-en.html">liberalism that wants to reduce cruelty</a>.)</p>

<p>However we think of this conception, the discomfort is real and should not be dismissed lightly.
Cybernetics did not emerge in a vacuum. Norbert Wiener developed his ideas about feedback loops and control circuits initially in the context of military anti-aircraft defense. The word “cybernetics” itself (from the Greek <em>kybernetes</em>, the helmsman) means guidance, mastery, control. And the program of describing biological organisms and technical machines under the same concepts provoked, and continues to provoke, an unease rooted deep in humanist tradition: if human beings and thermostats operate according to the same principles, what remains of freedom, dignity, and meaning?</p>

<p>This critique left its mark on the humanities academy. Cybernetics is still regarded by many as an intellectually dubious enterprise. It is seen as an attempt at a scientific annexation of the human being, a precursor to precisely those algorithmic regimes against which people argue so passionately today.</p>

<p>Yet here, I believe, lies a consequential misunderstanding or rather, a fatal confusion. In my interpretation of what I read, the cybernetics about which this discomfort exists is largely the <em>first-order cybernetics</em> of the 1940s and 50s, that is, the cybernetics of control, of feedback loops, of the behaviorist model that describes organisms through their input and output behavior without taking their inner life into account. And in fact, Wiener himself recognized early on what this program could bring about if placed in the wrong hands. In <em>The Human Use of Human Beings</em> <a class="citation" href="#wiener:1954">(Wiener, 1954)</a>, he warned emphatically against the possibility of using cybernetic principles to manipulate and control people. Wiener is in this sense a tragic figure—not because he opened a Pandora’s box without knowing it, but because he knew what he was doing, issued warnings, and was nonetheless remembered primarily as the inventor of an apparatus of control that he himself feared.</p>

<blockquote>
  <p>[Regarding the topic of job destruction,] Wiener notes in the [Cybernetics] that he’d attempted to alert the labor unions of the threats posed by automation to their membership. […] The potentially ruinous impact of communication technologies on democracy is another issue that Wiener anticipated with uncanny accuracy. […] As the scale, scope, and speed of information technologies have increased, so has the potential for corruption. Certainly Mark Zuckerberg failed to appreciate that Facebook’s “global community” of two billion users would inevitably produce countless messages that were antithetical to homeostasis, and thus to genuine community. […] That the routine operation of computer technologies can lead to disaster was a point Wiener stressed repeatedly. “Thinking” machines are relentlessly literal-minded, he said. […] Speed is another routine feature of automation that Wiener frequently warned could thwart our intentions. […] He regularly railed against the “hucksters” in commerce and “gadget worshippers” in science whose cupidity leads irrevocably, he believed, to “no homeostasis whatever.” Readers will find piquant examples of Wiener’s disdain for the captains of capitalist industry in [Cybernetics]. […] Wiener [in contrast to Shannon] set out to explain how information is the lingua franca of both animal and machine, a mission that consciously involved exploring, as he put it, “the boundary regions of science.” Thus, cybernetics as Wiener conceived it is <strong>physically embodied—understanding</strong> […]. – From the Foreword of <a class="citation" href="#wiener:2019">(Wiener, 2019)</a> by Doug Hill</p>
</blockquote>

<p>What has been almost entirely forgotten is <em>second-order cybernetics</em> <a class="citation" href="#foerster:2003">(von Foerster, 2003)</a>, which formed from the 1960s onward primarily around Heinz von Foerster and Humberto Maturana. This movement drew precisely the opposite conclusion from the cybernetic foundations. Its central argument was: <strong>living systems are autonomous</strong>. They are operationally closed. They cannot be steered from outside. They respond to perturbations from the environment according to their own inner logic. Maturana’s concept of <em>autopoiesis</em> <a class="citation" href="#barry:2012">(Razeto-Barry, 2012; Maturana &amp; Varela, 1987)</a> describes living systems as those that produce and maintain themselves and are therefore, in principle, inaccessible to external control.</p>

<blockquote>
  <p>I mention this matter because of the considerable, and I think false, hopes which some of my friends have built for the social efficacy of whatever new ways of thinking this book may contain. They are certain that our control over our material environment has far outgrown our control over our social environment and our understanding thereof. Therefore, they consider that the main task of the immediate future is to extend to the fields of anthropology, of sociology, of economics, the methods of the natural sciences, in the hope of achieving a like measure of success in the social fields. From believing this necessary, they come to believe it possible. In this, I maintain, they show an excessive optimism, and a misunderstanding of the nature of all scientific achievement – <a class="citation" href="#wiener:2019">(Wiener, 2019)</a></p>
</blockquote>

<p>This is not an apology for control but its critique in a Kantian sense, by nullifying the very conditions of possibility for purposive steering. The core ambition of second-order cybernetics was to show that control over nature and human beings is not only ethically problematic but epistemically impossible. It is a theory of the limits of steering. This insight was taken up by the ecology movement, i.e. by thinkers such as Gregory Bateson, who in <em>Steps to an Ecology of Mind</em> <a class="citation" href="#bateson:1972">(Bateson, 1972)</a> described the fatal consequences of a mode of thinking that treats nature as a steerable system. One might say: second-order cybernetics is the intellectual resource we would need in order to understand the mistakes we make when we think in the terms of first-order cybernetics.</p>

<p>Why is this part of the legacy so little known?
It seems to me that the emerging artificial intelligence research of the 1960s and 70s turned away from cybernetics—partly for substantive reasons, partly because competition for third-party funding sharpens disciplinary boundaries.
AI and cybernetics became rivals for resources and interpretive authority, not partners. Computer science, which was constituting itself as an independent discipline at that time, oriented itself toward AI research, not toward cybernetics and the emerging <em>systems theory</em>. One might say, somewhat pointedly, that <strong>computer science chose <a class="citation" href="#shannon:1948">(Shannon, 1948)</a> over <a class="citation" href="#wiener:2019">(Wiener, 2019)</a></strong>. The cybernetic legacy remained in control engineering, in parts of biology and sociology but not in the discipline that today builds the most consequential technical systems.</p>

<p>It is therefore no coincidence but the result of concrete institutional history that computer scientists and software engineers today are mostly unfamiliar with Wiener’s warnings or von Foerster’s critique of steerability. The burdened legacy of cybernetics is to a considerable degree a repressed legacy and the repressed, as we know, returns—only often in a form we did not choose.</p>

<p>Of course, the gains of specialization are tangible. Computer science as an independent discipline was able to concentrate on its core questions: computability, algorithms, data structures, architectures, formal verification. This focus produced extraordinary depth. We understand today with remarkable precision how systems formally function within defined boundaries, that is, when taking a blind eye to the reality of the complexity of interdependent but operationally closed systems.</p>

<p>The loss is subtler and therefore harder to grasp. It does not lie in having forgotten certain facts, but in certain questions never being asked in the first place. When the system boundary ends at the technical artifact, everything beyond that boundary, that is, the social, the psychological, the biological lies by definition outside the domain of responsibility. One is not blind to these areas out of indifference, but because the disciplinary toolkit simply cannot grasp them.</p>

<p>A physicist who knows only mechanics will not overlook thermodynamic phenomena because he dislikes them, but because his conceptual apparatus has no place for them. The same applies to computer scientists and software engineers who have never encountered psychological and social systems as objects of their discipline.</p>

<p>These <em>blind spots</em> become costly in a hypercomplex and hyperconnected world. The recommendation algorithm is only one example among many. Automation systems that transform labor markets and reshuffle social strata; systems that learn statistically, that intervene in decision-making processes, that determine life chances; surveillance infrastructures that shift the conditions of psychological and social autonomy are further cases. In all of them, we have built systems that function correctly (most of the time) at the technical level and produce effects at the systemic level that we did not anticipate, precisely because we never learned to think in these categories.</p>

<p>One might attribute a certain malice or greed to the builders of such systems. I prefer to speak of a certain <em>arrogance of ignorance</em>, of missing signals that would enable appropriate regulation, and of a system logic to which operators find themselves exposed. There are regulations against the contamination of drinking water, but we are only now beginning to think about how to limit the “contamination” of psychological systems. Part of the reason is certainly the distinction between physical and psychological injury, and the <strong>problem of paternalism</strong>. In the latter case, we are also dealing with effects that are difficult to observe. Nonetheless, it would be desirable if technical systems were kept under continuous observation and their operators were subject to a certain pressure of justification through systemic analysis. Operators should be answerable to the concerns of a systemic perspective. They should be confronted with the question of under what system logic the system operates, whether this leads to the wellbeing of citizens, and what plans exist to ensure it does. A systemic analysis could draw attention to <em>positive feedback loops</em> and call for the introduction of <em>negative</em> ones.</p>

<h2 id="redrawing-the-boundary">Redrawing the Boundary</h2>

<p>Here lies the core of the problem, and it is epistemological in nature: every systems analysis begins with a decision about what belongs to the <strong>system</strong> and what belongs to the <strong>environment</strong>. This decision is never neutral. It determines what counts as a relevant variable, what counts as noise, what counts as an effect of the system, and what counts as an external influence.</p>

<p>When we design a recommendation system and draw the system boundary to include only the algorithm, the database, and user interactions, we have by definition relegated psychological and social dynamics to the environment. They do not appear in the system model. Their feedback loops are invisible.</p>

<p>Importantly, this boundary-drawing occurs even when we do not consciously undertake it. It is built into our methods, how requirements analyses are conducted, how architectures are described, how tests are specified, and how success metrics are defined. <strong>The system boundary is not the result of a decision but the result of a tradition.</strong></p>

<p>And that is perhaps the strongest argument for a renewal of <em>systems-theoretical thinking</em> in computer science: not that we drew the wrong system boundaries, but that we mostly did not draw them at all. They emerged from disciplinary habit. A conscious, reflective practice of system modeling would mean making these boundaries explicit, and thus making them open to negotiation.</p>

<p>The goal is not to restore the cybernetics of the 1950s. Knowledge develops, and that is as it should be. But the systems-theoretical traditions of second-order cybernetics—embodied in Heinz von Foerster <a class="citation" href="#foerster:2003">(von Foerster, 2003)</a>, Stafford Beer’s Viable System Model <a class="citation" href="#beer:1995">(Beer, 1995)</a>, and the sociological systems theory of Niklas Luhmann <a class="citation" href="#luhmann:1984">(Luhmann, 1984; Luhmann, 1998)</a>—have, in the decades since the cybernetic breakthrough, developed a vocabulary that could be extraordinarily relevant for computer science today.</p>

<p>Some core concepts that would be worth introducing into <em>computational thinking</em>:</p>

<p><strong>Operational closure and structural coupling.</strong> Luhmann describes social and psychological systems as operationally closed, meaning they operate according to their own inner logic and cannot be steered directly from outside. A social system does not respond to inputs the way a technical system does; it is perturbed by impulses from the environment and processes these according to its own criteria. This has immediate consequences for any technology that seeks to “shape” social behavior. It can disturb, provoke, make offers, nudge but it cannot directly control or command.</p>

<p><strong>Emergence.</strong> Complex systems exhibit properties that do not exist at the level of individual components and cannot be predicted from them. This is no longer an unfamiliar concept in computer science but it is usually applied to technical systems. Systems-theoretical thinking would suggest expecting and analyzing emergence also at the interface between technical and social systems. What arises when an algorithm and a social community come into contact? This question cannot be answered with technical means alone but it can at least be posed precisely with systems-theoretical concepts.</p>

<p><strong>Recursive self-description.</strong> Second-order cybernetics pointed out that every description of a system is part of the system it describes. Whoever models social systems alters the system through the model. The actors know the model, react to it, habituate themselves or are estranged from it, subvert or confirm it. This is a fundamental problem of all social technology: it does not operate on a neutral substrate but on self-interpreting systems. A creditworthiness algorithm, once known, changes the behavior of the people it evaluates. Those who know that language models analyze CVs will formulate their CV differently.</p>

<p><strong>Feedback and system dynamics.</strong> This is the oldest cybernetic concept and simultaneously the one that has penetrated furthest into computer science, for instance in control engineering. But the systems-theoretical perspective would invite us to consider feedback loops not only within technical systems, but also between technical, social, and psychological systems. The changes that a technical system triggers in its social environment return to the technical system as altered usage and altered expectations. Modeling or at least anticipating these loops is difficult but ignoring them is dangerous.</p>

<p>At this point an objection might arise: should computer science now pursue sociology and psychology? Do we not lose precisely the sharpness that makes us productive when we expand into such breadth?
The objection deserves to be taken seriously but it rests on a misunderstanding. The goal is not to replace computer science with systems theory or to retrain programmers as sociologists. The goal, in my mind, is something more modest: computer scientists should learn to consciously perceive the limits of their models.</p>

<p>An architect need not be a structural engineer in order to know that they need one and when they need one. A software engineer need not be a social scientist in order to know that their system intervenes in social systems with their own dynamics that they do not fully understand. Systems-theoretical thinking gives her the vocabulary to name these boundaries, and might provoke her to ask the right questions, to bring in the right expertise.
This would be, I think, no weakening of the discipline, but a consistent maturation that is long overdue given present circumstances.</p>

<p>Compare it with the development of software quality over recent decades. It was once not a self-evident part of <em>computational thinking</em> <a class="citation" href="#wing:2006">(Wing, 2006)</a> to reflect on security, accessibility, data protection, or energy efficiency. These aspects were introduced into the discipline through external demands, for example, through legal regulation, social debate, spectacular failures, and so on. Today they belong, if not yet always ideally, to the canon.</p>

<p>Systemic effects could be the next chapter of this development, especially because the practice of writing code—which is very different from understanding it—seems to be in decline.
It would be wiser to write this next chapter before the failures become even larger.</p>

<p>For practitioners, “more systems theory” can easily sound like an academic demand without operational consequence. It is therefore worth sketching where systems-theoretical thinking can actually be translated into practice:</p>

<p><strong>In requirements analysis:</strong> The explicit question of which systems, i.e. technical, social, psychological, biological, are touched by the project. Not as a checklist, but as a genuine analytical practice. What feedback loops are to be expected? What properties of these systems are not modelable but nonetheless relevant?</p>

<p><strong>In system design:</strong> The distinction between what the technical system can control and what it can only perturb or offer, that is, the difference between <em>loose</em> and <em>strict</em> <em>couplings</em>. Which systems must we treat as <em>black boxes</em>, and which are <em>white</em> to us? This distinction fundamentally changes design decisions. It suggests making systems more modular, more reversible, and more observable, because the effects on social systems cannot be fully anticipated.</p>

<p><strong>In metrics definition:</strong> The question of whether the metrics being optimized for represent the systemic goals, or only the technically measurable slice of them. This is the most direct consequence of the recommendation algorithm example: time on platform is not the same as user wellbeing.<sup id="fnref:2" role="doc-noteref"><a href="#fn:2" class="footnote" rel="footnote">2</a></sup></p>

<p><strong>In education:</strong> Systems-theoretical foundations—not, of course, in their full Luhmannian complexity—but potentially as part of the computational canon. What are systems? What is environment? What is (structural) coupling? What is emergence? What does it mean that social systems are operationally closed? These questions could be integrated as a standalone module or as a cross-cutting theme in existing courses. The theory itself must, however, remain open to critique, and it is necessary to also engage with its critics.</p>

<p><strong>In interdisciplinary projects:</strong> Systems theory as a shared language that enables precise communication across disciplinary boundaries. When a computer scientist and a sociologist can both speak about “structural coupling,” they have a common vocabulary for something that is otherwise very difficult to articulate.</p>

<h2 id="when-the-model-becomes-an-actor">When the Model Becomes an Actor</h2>

<p>Let me come to an end by looking at a recent example of a system that should be looked at with a systems-theoretical perspective.</p>

<p><em>Prediction markets</em> offer a particularly instructive illustration of dynamics. 
These are markets on which participants can bet on the outcome of future events, for example, from election results to whether a politician will use a particular word in a speech to whether a new drug will pass clinical trials. 
Each bet takes the form of a contract that pays out one dollar if the event occurs and nothing otherwise. 
The current market price therefore directly reflects the collective probability estimate: a price of $0.73 means the market assigns a 73% probability to the event occurring. 
The promise is that markets aggregate information efficiently, meaning those with superior knowledge can profit, which draws well-informed participants and, ideally, surfaces something close to “the truth”.</p>

<p>Prediction markets try to translate political and social complexity into the economic code of profitable/unprofitable (or payment/non-payment). But what happens when we try to solve a political crisis (polarization) using an economic subsystem?
Luhmann would argue this causes a mismatch because the political system operates on the code of power/opposition, not money.</p>

<p>The most obvious problem that follows is one we already know from sports betting: the possibility of corruption. If you can bet on an outcome, you have a financial incentive to influence it.
Prediction markets extend this temptation to virtually every domain—elections, policy decisions, scientific results, media events.
The set of potential actors willing to manipulate an outcome grows correspondingly large, and the markets themselves make it easier to identify which events are both consequential and controllable by a small group.
In Luhmann’s terms, money is being used to <strong>pierce the operational closure of functionally differentiated systems</strong>.
This could essentially make the political system respond to economic operations, or the scientific system respond to financial incentives (directly, i.e. suddenly the operations of system A operate <strong>in</strong> system B). 
Functional differentiation made this kind of interference unlikely; it made each system efficient with respect to its own internal code.
Aside from obvious moral issues, breaking it down has costs on a functional level.</p>

<p>But the more fundamental problem persists even if we set aside corruption entirely and imagine a world of perfectly honest participants. It is the problem of recursive self-description. 
A prediction market does not merely observe the probability of an event but it also publishes that observation, and in doing so becomes a participant in the very system it was meant to describe from the outside.</p>

<p>In many cases such <em>second-order observation</em> effectively reduces complexity, for example, when you want to buy a house it is not necessary to go to the house and calculate how much it should cost by hand.
You observe how others observe the house via the housing market.</p>

<p>In the case of prediction markets, second-order observation is part of the signal an individual brings into the prediction and thereby undermines the whole concept of “drawing in individuals with superior knowledge”.
And worse: a market price of 85% in favor of a particular election outcome influences how voters, campaigns, donors, and media organizations behave, which in turn influences the outcome. 
The observer has entered the system. 
The model is no longer a neutral representation; it is an <strong>actor</strong>.
That media organizations seem highly interested in coupling their operations with these markets could likely amplify these effects.</p>

<p>The possible consequence is that prediction markets are not simply truth-discovery mechanisms that occasionally malfunction. 
They are an interesting case for systems theorists because, as I hinted at, they make <em>observations</em> of the models of psychic systems, i.e. our beliefs, <em>observable</em>.
Consequently, at scale, they risk becoming machines that produce self-fulfilling prophecies—not reading the future from a god’s-eye view, but actively shaping it through the feedback loops they generate. 
Whether this is a net gain depends on a question the markets themselves cannot answer: what value does a prediction that influences its own outcome actually produce? 
And for whom?
And, assuming the system works effectively—which is highly questionable—is it a net good to know what an aggregate believes about the future?</p>

<p>Certainty, we desire certainty—but at what cost?
One can argue that by turning tragic or highly contingent events (like elections, wars, or climate disasters) into financial bets, the system achieve complexity reduction at the cost of <em>empathy</em>!
The possibility to bet on whether there will be a Russia-Ukraine ceasefire before a certain video game comes out, feels intuitively ethically wrong and deeply distasteful.
Fittingly the Kalshi ad read: <strong>The world’s gone mad, trade it.</strong> 
It is, of course, also very dubious that the son of the President of the U.S. is an adviser to two of these markets.</p>

<blockquote>
  <p>Prediction markets are the future. I think they are the future not just for traders but also for news and information. – CEO of Robinhood</p>
</blockquote>

<blockquote>
  <p>New York City was shut down [during COVID] and I asked myself: when is this going to end? When will the vaccine gonna be ready? When is shelter in place to be over? And prediction markets can take all these disparate opinions that people are pontificating about or that they have really good reasons to believe and distill it down into one probability. […] When I get hit up by people in the Middle East who are saying that “You know we’re looking at Polymarket to decide whether we sleep near the bomb shelter.” And I am like “Oh, it is really that popular over there?” That is very powerful. That is like an undeniable value proposition that did not exist before. The global truth machine is here, powered by the people. – CEO of Polymarket</p>
</blockquote>

<blockquote>
  <p>We are a financial market like a stock market but you trade on politics, weather, climate, economics, sports, and so on. This can rival the stock market. – CEO of Kalshi</p>
</blockquote>

<p>Therefore, going back to my <a href="/Pages/2026/04/03/cruelty-and-solidarity-en.html">previous post</a>, I pose the question: <strong>is this attempt to eliminate contingency a path worthy to go?</strong>
While we seem to crave a new form of order, we must interrogate this impulse. Any distinction between <em>order</em> and <em>noise</em> is drawn by an observer who is necessarily blind to the conditions of their own observation.</p>

<p>Furthermore, a paradox emerges: the non-linear dynamics of prediction markets will likely be reflected in their own unpredictability. Systems theory teaches us that complexity cannot be destroyed. The environment always possesses higher complexity than the system, forcing the system to reduce this complexity to a manageable level by ignoring a massive amount of “side” effects—thereby creating new, unpredictable problems.
A system’s reduction of complexity is always local and temporary, requiring a corresponding increase in the system’s own internal complexity. By drawing a boundary and simplifying what gets let in, the system inadvertently triggers a wave of new dependencies, cascading side effects, and emergent structures. Thus, the very mechanisms we deploy to cope with complexity become the engines that generate more of it.
Again, systems theory screems: <strong>social, psychic, living, and technical systems are out of control!</strong></p>

<h2 id="an-invitation">An Invitation</h2>

<p>This text is an invitation to reflect. I am by no means an expert on the subject and my reading list is long and growing.
It is an attempt to articulate an intuition about my observation of the development of technical systems that are increasingly socially and psychologically disruptive.</p>

<p>The intuition is this: we build things we do not fully understand—not in the technical sense, but in the systemic one. We know how our algorithms function. We do not know well enough how they function in the world.</p>

<p>This is no reason for paralysis. No engineering endeavor waits until all effects are fully known before it begins. But it is a reason for humility and curiosity. The systems that concern us are larger than the boundaries we have drawn around them.</p>

<p>Perhaps it is time to renegotiate those boundaries. Not in order to leave computer science behind, but to extend it. Cybernetics and systems theory offer for this purpose a vocabulary that has matured over fifty years and remains largely unused and is largely unknown in the very discipline that emerged from it, and yet which might need it most urgently.</p>

<p>Perhaps it is worth attempting to take a look inside—not to eliminate uncertainty or to build systems of total control but to acknowledge our own ignorance and limits of control.
A cybernetics worthy of its noble philosophical heritage must not be reduced to the mere fine-tuning of a self-guided missile (control as error minimization against a fixed target), but must instead understand <em>control as self-control</em>—as the capacity to know one’s own ignorance and to place one’s own goals and norms up for discussion.
In other words, I believe we should cultivate a <em>cybernetics</em> worthy of its philosophical heritage that maintains the knowledge of its own ignorance, rather than boasting of the mechanical reduction of complexity—a cybernetics that understands <strong>control</strong> in terms of <strong>agency</strong> and <strong>autonomy</strong> and the fine-tuning of doubt, not as a pretext for the <em>hubris of domination</em>.</p>

<h2 id="references">References</h2>

<ol class="bibliography"><li><span id="luhmann:1998">Luhmann, N. (1998). <i>Die Gesellschaft der Gesellschaft</i> (p. 1164). Suhrkamp.</span></li>
<li><span id="zoennchen:2025">Zönnchen, B., Dzhimova, M., &amp; Socher, G. (2025). From intelligence to autopoiesis: rethinking artificial intelligence through systems theory. <i>Frontiers in Communication</i>, <i>Volume 10 - 2025</i>. https://doi.org/10.3389/fcomm.2025.1585321</span></li>
<li><span id="wiener:1954">Wiener, N. (1954). <i>The Human Use of Human Beings: Cybernetics and Society</i>. Garden City, New York : Doubleday.</span></li>
<li><span id="wiener:2019">Wiener, N. (2019). <i>Cybernetics or Control and Communication in the Animal and the Machine</i>. The MIT Press. https://doi.org/10.7551/mitpress/11810.001.0001</span></li>
<li><span id="foerster:2003">von Foerster, H. (2003). Cybernetics of Cybernetics. In <i>Understanding understanding: Essays on cybernetics and cognition</i> (pp. 283–286). Springer New York. https://doi.org/10.1007/0-387-21722-3_13</span></li>
<li><span id="barry:2012">Razeto-Barry, P. (2012). Autopoiesis 40 years later. A review and a reformulation. <i>Origins of Life and Evolution of Biospheres</i>, <i>42</i>(6), 543–567. https://doi.org/10.1007/s11084-012-9297-y</span></li>
<li><span id="maturana:1987">Maturana, H. R., &amp; Varela, F. J. (1987). <i>The Tree of Knowledge</i>. Shambhala.</span></li>
<li><span id="bateson:1972">Bateson, G. (1972). <i>Steps to an Ecology of Mind</i>. Chandler Publishing Company.</span></li>
<li><span id="shannon:1948">Shannon, C. E. (1948). A mathematical theory of communication. <i>Bell Syst. Tech. J.</i>, <i>27</i>(3), 379–423.</span></li>
<li><span id="beer:1995">Beer, S. (1995). <i>The Heart of Enterprise</i>. Wiley.</span></li>
<li><span id="luhmann:1984">Luhmann, N. (1984). <i>Soziale Systeme: Grundriß einer allgemeinen Theorie</i>. Suhrkamp.</span></li>
<li><span id="wing:2006">Wing, J. M. (2006). Computational Thinking. <i>Commun. ACM</i>, <i>49</i>(3), 33–35. https://doi.org/10.1145/1118178.1118215</span></li></ol>
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:1" role="doc-endnote">
      <p>Norbert Wiener probably would have argued against this distinction because for him noise is a problem for any system be it a machine or an organism. The contradiction dissolves when one realizes that Luhmann and Wiener are analyzing technology at two different levels: operational behavior vs. structural programming. <a href="#fnref:1" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:2" role="doc-endnote">
      <p>Of course, regulations or other incentives have to be installed to make sure that the desired systemic goals are in fact the goals of the organizations that set them. But this is itself a systemic problem because such an organization is a complex system that can not be steered or controlled directly but can only be irritated. <a href="#fnref:2" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content><author><name>Benedikt Zönnchen</name><email>benedikt.zoennchen@web.de</email></author><category term="Opinion" /><category term="Education" /><category term="Systems Theory" /><category term="Cybernetics" /><summary type="html"><![CDATA[Imagine we are given the task of building a recommendation system. The requirements are clear: the system should suggest content to users that they are highly likely to click on. We define a metric. Let’s say, time spent on the platform. Then we optimize toward it. We test, iterate, deploy. The system works. The metric rises.]]></summary></entry><entry><title type="html">The Art of Solidarity</title><link href="https://bzoennchen.github.io/Pages/2026/04/03/cruelty-and-solidarity-en.html" rel="alternate" type="text/html" title="The Art of Solidarity" /><published>2026-04-03T00:00:00+02:00</published><updated>2026-04-03T00:00:00+02:00</updated><id>https://bzoennchen.github.io/Pages/2026/04/03/cruelty-and-solidarity-en</id><content type="html" xml:base="https://bzoennchen.github.io/Pages/2026/04/03/cruelty-and-solidarity-en.html"><![CDATA[<blockquote>
  <p>Part of the idea of a democratic society is that social change comes by reform rather than revolution and this in effect means that the people who have the power actually letting go of some of it. – Richard Rorty</p>
</blockquote>

<h2 id="a-new-cruelty--on-the-longing-for-certainty">A New Cruelty — On the Longing for Certainty</h2>

<p>Out on the market squares the voices are clamoring. They lament the swaying of the stone temples we call democracy. But while our gaze clings to the crumbling facades of power, we overlook the quiet fading of tenderness in the corners of our everyday lives. We cling to a scaffolding of institutions and forget in doing so that political freedom is not a foundation on which we stand, but rather follows an improbable story that we must keep telling ourselves in ever different guises.</p>

<p>Behind the visible trembling of our institutions, a quieter yet far deeper freezing is taking place. It is as though an old contract with humanity has been terminated—a withdrawal from the shared <em>care</em> for this fragile world. Where once stood the promise of striving to alleviate the pain of the other, there grows today a strange urge toward a <em>new hardness</em> <a class="citation" href="#kohlenberger:2024">(Kohlenberger, 2024; Carolin Amlinger, 2025; von Redecker, 2026)</a>.</p>

<p>The language that once bound us has turned to stone. It knows only the absolute: sharp edges of “facts” and “enemies”, a vocabulary cast as though from ore, leaving no room for what lies between. Attention is directed either toward distraction—the flight from a society one no longer wishes or is able to care about—or toward the destruction of what seems no longer to function. In part patience has run out, and in part curiosity is suffocating in restlessness.</p>

<p>In the narratives of our day, truth has gone astray; a bizarre theater of the obvious reigns, in which the lie is no longer concealed but triumphs as a naked gesture of violence. It is the violence of the unenlightened, in the sense that their centers were never permitted to develop desires. They were not nurtured and could not learn to draw symbolically from their accumulated rage and traumas. And what had previously remained hidden behind etiquette and the effort to maintain decorum now shows itself unveiled: speech is no longer used to be understood, but to humiliate other people.</p>

<p><a class="citation" href="#baudrillard:1976">(Baudrillard, 1976)</a> was right in assuming that the West could not handle the brutal symbolic gift it received on 9/11; by trying to give it back, the system turned inward to give its humiliated power food it can process. Some of this cruelty is old and well accepted if experienced by “the right people” because again and again we rationalized our way towards exceptions; of treating people marked by some features quite different in front of the law. “We” failed the test, unable to realize our self-description. And what could have been a time for self-reflexion became the years of the beast.</p>

<blockquote>
  <p>However difficult this vote may be, some of us must urge the use of restraint. Our country is in a state of mourning. Some of us must say: “Let us step back for a moment and let’s just pause for a minute” and think through the implications of our actions today, so that this does not spiral out of control. […] I came to grips with [my vote] today and I came to grips with opposing this resolution during the very painful, yet very beautiful memorial service. As a member of the clergy so eloquently said: “As we act, let us not become the evil that we deplore.” – Barbara Lee (Single voter of Congress who voted aginst the war in Afghanistan (518 to 1 vote))</p>
</blockquote>

<p>Surely this is too simple; too reductive but still, those scares of these days never healed. 
They left their mark because we disrespected our values and created an inescapable dissonance between what we do and what we imagine ourselves to be.
We contributed to a reality where acts of defiance are increasingly labeled as terrorism, and where anyone can be branded a terrorist—especially those whose solidarity is based purely on human suffering, regardless of political or national boundaries.
This <em>war on terror</em> exists as a profound paradox: it is at once a direct assault on the foundational tenets of liberalism and the very mechanism deployed to safeguard them.</p>

<p>This systemic turning against our own foundational structure mirrors a profound, quiet madness; one that manifests vividly in the cinematic image of a solitary wanderer in the eternal ice.
The penguin in Werner Herzog’s <a class="citation" href="#herzog:2007">(Herzog, 2007)</a> narrative becomes a propaganda figure who, for no reason, turns away from the sheltering colony to waddle stubbornly toward distant mountains.
It is a thoroughly extraordinary march against the penguin’s own environment, against the instinct for survival, and against every community of interdependent beings.</p>

<p>There, in the desolation of the peaks, the wanderer hopes to open up a new realm of winners.
In his eye it shall be a purifying reincarnation.
And, like many Italian futurists, he looks ahead, towards speed, violence, technology, industry, and war.
Yet beneath this mechanical zeal runs a deeply religious current—an echo of American end-time mythologies where a disappointing world cannot be redeemed, only consumed by fire.</p>

<p>But the wanderer has to be certain of his cause.
Doubt—his own as much as that of his companions—is what gnaws at him.
In this apocalyptic logic, doubt is not just a hesitation; it is a temptation by the Antichrist, a betrayal of the absolute faith required for the final days.
From the perspective of his colony, which is setting out toward the ocean, it is a mission without tomorrow, driven by a dark longing for the end—a final act of retribution against a world one can no longer imagine made better.
And in this departure, the cruelty directed against curiosity and doubt becomes the perverse satisfaction of being right, at least with the prophecy of downfall. 
In our ears Herzog’s voice resonates:</p>

<blockquote>
  <p>But why? – Werner Herzog</p>
</blockquote>

<p>And so, we applaud the wanderer’s courage, attempting to re-interpret what seems like a mix of nihilism and destructive vitalism as a desperate call to save Europe <a class="citation" href="#gundlach:2026">(Gundlach, 2026)</a>.
But this wanderer does not follow life; he follows a myth of freedom that leads him irresistibly into the white void.
He is dead.
He won against his environment—against a <em>careless</em> nature:</p>

<blockquote>
  <p>With five thousand kilometer ahead of him, he’s heading towards certain death. – Werner Herzog</p>
</blockquote>

<p>A remarkably similar architecture of self-imposed isolation is being constructed today in the shadow zones of our digital world.
Here, a brotherhood of grievance has formed, a loose network of voices that find their identity only in the echo of hatred.
They need a face they can despise, and because sensitization has taken something from them, solidarity itself has become the enemy.
The call for a return to the “true mask” of man is more than mere nostalgia; it is the desperate longing for an old, heavy armor that admits no cracks and thus no vulnerability.
It is an anarchic armor that no longer requires solidarity at all, because it is its own ecosystem—self-sufficient, optimized, independent, and closed in on itself.
At the same time, behind the facade lies deep suffering, because man has lost something that once made his world simple and secure.</p>

<blockquote>
  <p>Destructive attitudes develop predominantly in people who perceive themselves as marginalized, yet are status-ambitious and dominance-oriented. Individuals with a destructive mindset are convinced that they are being denied a social position to which they have a legitimate claim. – Oliver Nachtwey</p>
</blockquote>

<p>He could no longer reconfigure his identity through the new vocabulary.
He was <em>humiliated</em> and sought another language, so that everything that had seemed particularly important to him might once again become true and good.
What he found is <strong>the language of dominance</strong>.</p>

<blockquote>
  <p>The will to dominate was the fundamental law of the life of the universe from its most rudimentary forms to its most elevated ones. That man was driven by a divine bestiality. – Benito Mussolini</p>
</blockquote>

<!--

In the words of Palantir's manifesto:

>The limits of soft power, of soaring rhetoric alone, have been exposed. The ability of free and democratic societies to prevail requires something more than moral appeal. It requires hard power, and hard power in this century will be built on software. [...] We must resist the shallow temptation of a vacant and hollow pluralism. We, in America and more broadly the West, have for the past half century resisted defining national cultures in the name of inclusivity. But inclusion into what? 

-->

<p>It is as though history were violently recoiling.
Where the world had begun to grow quieter and more sensitive to the pain of “the other”, a part of it responds with a new, steely coldness.</p>

<blockquote>
  <p>Tonight, a whole civilization will die and never return. […] There might something revolutionarily wonderful happening, who knows. – Donald Trump</p>
</blockquote>

<p>Yet beyond the loud grievance of the streets, in the soundless, glass-walled cathedrals of light and silicon, a far quieter, almost clinical cruelty is ripening.
It is a faith that bundles itself in seven cold stars into a single radiance—an alliance of those who regard “the human being” as a transitional sketch to be technologically overcome <a class="citation" href="#gebru:2024">(Gebru &amp; Torres, 2024; Mühlhoff, 2025)</a>.
In this light, the longing for the stars no longer appears as a departure but as a flight; an expansion into the void, driven by the dream of an eternity that no longer needs a body; to become a misremembered Puppet Master—a ghost whithout a shell in a Cartesian fantasy.</p>

<p>From the heights of their galactic calculations, these architects of the future look down on the here and now as upon an ant colony in the dust.
The suffering of the present—the exhaustion of the earth, the silencing of diversity—shrinks in their eyes to a negligible rounding error.
It is an ethics that sacrifices today to a speculative singularity of tomorrow, a morality of arithmetic, a <em>rule of code</em> instead of law <a class="citation" href="#rosengruen:2022">(Rosengrün, 2022)</a> in which a burning planet weighs less than the mathematical promise of a posthuman world of gods. 
There is nothing novel or imaginative about these ideas; they are archaic myths that provide narrow futures.
We must remain acutely aware of these ancient dreams of greatness that demand a catastrophic downfall to purge “decadence” and to bring a new <em>technological order</em> into choas.</p>

<blockquote>
  <p>We should try to create autonomous countrys on oceans, under water, and all sorts of other spaces. Technology is the vehicle to escape and move beyond politics as we find it today. – Peter Thiel</p>
</blockquote>

<p>In this worldview, an old dark spirit returns, cloaked in the garb of logic: the conviction that life has a price measured by its utility for the great progress, which could be gauged by its approximation to the <em>absolute</em>; that the immaterial disguised as weightless information is real and the material—bodies and trees, flowers and animals, mountains and oceans—can be overcome.
Once again it is assumed that everything can be calculated, but in place of Kantian principles stand calculating machines and the theories of probability and expected values.
Once again we await a god—this time a god made of numbers—a superintelligence that, like an infallible oracle, could end the chaos of our interpersonal stories—as though society could communicate with anything other than itself.
It is Plato’s ancient, stony dream—the hope that pure, incorruptible truth might finally triumph over the tender but imperfect narratives of compassion; that we might remember what has always been out there and within us; that in this “awakening” the True and the Good coincide.</p>

<blockquote>
  <p>Plato thought that morality and politics should be based on principles in the same why that Euclidean geometry is based on axioms. He thought that philosophical inquiry was a matter of nailing down firm immutable principles which could then guide action. […And] as Plato said, it is if we had known the truth in a pervious existence and simple need to be reminded of it, have it brought back to consciousness.  – <a class="citation" href="#rorty:1990">(Rorty, 1990)</a></p>
</blockquote>

<p>But in this purity there is no longer room for breathing, for trembling, for the passion, love and affection for a world that is precious precisely because it can keep reinventing itself. There will be no gods only the loss of institutions that once balanced and distributed power. And concentrated power, ultimately, devours empathy if you cannot strip yourself of it in time.</p>

<blockquote>
  <p>It seems to me the Platonic notion of absolute truth is a thoroughly misleading slogan and a culturally dangerous shibilith. – <a class="citation" href="#rorty:1990">(Rorty, 1990)</a></p>
</blockquote>

<p>This frosty belief is nourished not least by a bottomless exhaustion—a cynicism that crystallizes like a dark sediment out of the fear of the end. It is the rebellion against the last remaining justice: death, that relentless equality to which we all succumb. In the impatience to outwit this fate, hope turns into bitterness. Some have everything but can only continue to create themselves by getting rid of themselves. Others have nothing and can only react to disruptions. They have no time for self-description. What is a calculable <em>risk</em> for few becomes a <em>danger</em> for everyone else.</p>

<p>A quiet, gray poison forms, which has long since crossed the threshold of our homes. It nestles into the corners of our living rooms, a mood like a silent echo of those chroniclers of hopelessness who whisper to us that the world has become a closed circle. In the age of disruptions <a class="citation" href="#stiegler:2019">(Stiegler, 2019)</a>—which feels more and more like an age of destructions—the chords that present us with a horizon of notes fall silent. Structures dissolve, so that we have difficulty imagining different futures.
We feel it in the burden of everyday life: the sense that what exists can no longer be healed, that the system is frozen at its foundations and can only be broken.</p>

<p>As different as these three figures may appear—the ranter on the market square, the calculator in the server room, the exhausted person on the sofa—they share a common root: the inability to live with the contingency of life.
All three want certainty.
One fights for it through enemies, another calculates it through machines, and the third finds it in the renunciation of all hope.</p>

<p>In this darkness we look at ourselves and see only deficiencies.
The human being no longer appears to us as a continually changing riddle of openness, but as a flawed, inadequate disruptive factor to be optimized away.
While the calculating machine promises eternal perfection, the human being becomes a creature of dust and error, one that can readily be dispensed with.
We lose faith in the laborious, small gesture of improvement and instead begin once more to dream of perfection, as though we could knowingly move toward the summit.
It is the <em>exhaustion of the contingent</em> and thus the wish that something might finally be certain.</p>

<hr />

<h2 id="doubt-as-home--contingency-irony-solidarity">Doubt as Home — Contingency, Irony, Solidarity</h2>

<p>Amid this recurring coldness, this icy wind of abstraction, a voice is needed that leads the resistance against dehumanization not as a loud protest, not as instruction or a return to the rational, but as a healing gesture of humility: it wrests the human condition from the cold, sterile distances of the heavens and beds it back in the warm, imperfect dust of the earth.</p>

<p>Richard Rorty is such a voice—a voice that prefers doubt to certainty.
As a philosopher of <em>irony</em> (in the sense that we should not take our existing vocabularies too seriously), of <em>contingency</em> (in the sense that there are no ahistorical truths and history follows no necessary path), and of <em>solidarity</em> (in the sense that he prefers it to truth <sup id="fnref:1" role="doc-noteref"><a href="#fn:1" class="footnote" rel="footnote">1</a></sup>) <a class="citation" href="#rorty:1989">(Rorty, 1989)</a>, he wanted—naively put—to strip philosophy of its imperialist position as the foundational discipline—and to propose, of all things, the literary critic as a model for the public intellectual, which fairly earned him the reputation of relativist and charlatan.</p>

<blockquote>
  <p>By an intellectual I mean someone who has doubts about the value of the language she is been using to make moral or political judgements and who reads books in an effort to deal with these doubts. To be an intellectual is to have a restless mind—never to be sure that once judgement of other people’s characters or of alternative social institutions are more than inherited prejudices. – <a class="citation" href="#rorty:1990">(Rorty, 1990)</a></p>
</blockquote>

<p>Meanwhile, he himself wrote quite systematically (as in <a class="citation" href="#rorty:1979">(Rorty, 1979)</a>) but was more interested in interpreting and reconfiguring philosophical texts than in erecting a new grand system of his own.
His writing style as well as his philosophical position are American in their simplicity, whereas his admiration squints toward Europe.
He likely embodies much of what disturbs philosophers when they write about the “decline of culture,” for as someone who wants to take neither Kant nor Nietzsche too literally, he refuses answers to cold questions such as: “What <strong>ought</strong> I to do?”—in the sense of a universal duty. The warm question, however, “What can <strong>we</strong> do for one another?”, he does answer, but without metaphysical backing: Reduce cruelty, expand your circle of solidarity—not because reason commands it, but because we have learned what it means to be humiliated.</p>

<p>His irritating position on “the Truth” and “the really real” is a melting of pragmatism and romaticism which can be summed up by the following quote:</p>

<blockquote>
  <p>On the account of human abilities I am suggesting, the use of persuasion rather than force is an innovation comparable to the beaver’s dam. Like the beavers’ collaboration in getting the dam built, it is a social practice. It was initiated by the noval suggestion that we might use noises rather than physical compulsion to get other humans to cooperate with us. That suggestion gave rise to language. Rationality, thought, and cognition all began when language did.<sup id="fnref:2" role="doc-noteref"><a href="#fn:2" class="footnote" rel="footnote">2</a></sup> Language gets off the ground not by people giving names to things they were already thinking about but by proto-humans using noises in innovative ways, just as the proto-beavers got the practice of building dams off the ground by using sticks and mud in innovative ways. Language was, over the millenia, enlarged and rendered more flexible not by adding the names of abstract objects to those of concrete objects but by using marks and noises in ways unconnected with environmental exigencies. The distinction between the concrete and the abstract can be replaced with that between words used in making perceptual reports and those unsuitable for such use. […W]e need to think of reason not as a truth-tracking faculty but as a social practice—the practice of enforcing social norms on the use of words rather than blows as a way of getting things done. We need to think of imagination not as the faculty that produces visual or auditory images but as a combination of novelty and luck. To be imaginative, as opposed to being merely fantastical, is to do something new and to be lucky enough to have that novelty be adopted by one’s fellow humans, incorporated into their social practices. […] People whose novelties we cannot appropriate and utilize we call foolish, or perhaps insane. Those whose ideas strike us as useful we hail as geniuses. – <a class="citation" href="#rorty:2016">(Rorty, 2016)</a></p>
</blockquote>

<p>What Rorty offers me is an admirable description of the liberal and democratic society that unerringly exposes anti-liberal opposites.
He gives me an answer to the question “What do liberals want and what can they hope for?” with which I can agree.
And although—or precisely because—he takes leave of <em>first principles</em> and speaks soberly about texts from Wittgenstein to Proust, he can inspire enthusiasm for a liberalism by understanding it as <strong>the art of solidarity</strong>.</p>

<p>Let us begin with a note on why educational institutions in particular are so decisive for it:</p>

<blockquote>
  <p>In democratic societies like ours, colleges and universities have a peculiar two-faced role. They get their money by promising to furnish money-making skills to their students and by promising to perform research which will enable society as a whole to get more goods and services more cheaply. The face they present to rich donors, state legislators, and the general public is essentially a commercial one. They suggest that they have certain products which the society as a whole needs and they ask for support on that basis. They usually don’t suggest that their function is to disturb the students, make them have doubts about the way they were brought up, force them to ask unanswerable questions. But as you know quite well, that is the function which many faculty members, especially the people who teach in the humanities and social science departments, think that colleges and universities should serve. Such people see the promise of marketable skills as simply the lure which brings students within their reach. Once the student is in their classes, these people assign books which will—they hope—upset her enough to make her want to start for looking for other such books so as to get upset in still more complicated ways. They are not satified unless the students who leave their courses are dissatified with the society in which they live and unless they have at least some doubts about the moral codes in which they were brought up. […T]he point of encouraging dissatisfaction becomes clear when times are bad and particularly when societies and governments become repressive. Then the colleges and universities come into their own. They begin to function either as sources of social change or as sanctuaries for resistance. […] I can sum up the two roles of colleges and universities by saying that whereas the society as a whole wants to produce people with skills, the university faculties also want to produce intellectuals. – <a class="citation" href="#rorty:1990">(Rorty, 1990)</a></p>
</blockquote>

<p>Rorty’s thinking begins where the coldness described above has its deepest origin: in the belief that there is (or was) a final truth toward which we are moving (or from which we have depatured).
He distrusts this belief—not out of cynicism, but out of a deep <em>care</em> for what we lose when we chase after it.
For Rorty, truth is a property of sentences, and sentences are made by people.
“The world” cannot make these sentences true, but we can make them useful for ourselves and useful for <strong>us</strong>.
In the age of “fake news” this may sound irritating, for is it not “the truth” whose loss we lament?</p>

<p>In place of principles that we might yet discover, Rorty stakes everything on new, persuasive vocabularies that are to be invented.
For if one wants to say something new, one must create a new language <a class="citation" href="#maturana:1991">(Maturana, 1991)</a>.
Freedom and solidarity arise where we find words, metaphors, and descriptions that allow us to see ourselves, others, and the world differently.
<strong>The power of language lies not in mirroring reality, but in reconfiguring it.</strong></p>

<blockquote>
  <p>I do indeed assert that the explicit or implicit answer to the question of reality determines how we lead our lives and in what way we accept or reject other people within the network of the social and non-social systems we form. – <a class="citation" href="#maturana:1988">(Maturana, 1988)</a></p>
</blockquote>

<p>Kant provided some good reasons to prevent cruelty, and Nietzsche shattered their <em>claim to universality</em>—he shattered the idea that there is any entity, whether God or Reason, that could decide outside a historical-evolutionary context: “Who I am”, “What I should do”, “What I can know” and “What I may hope for.”
Kant undertook the attempt to formulate principles that should hold for everyone, thereby creating the basis for solidarity, human rights, and democracy.
But his Reason could not order the surplus of meaning that arose from the newly won freedom: too many opinions, too many perspectives, too much criticism and deconstruction.
Kant overlooked the obvious: that he himself as observer must remain blind to his own observing.
Once again the paradox erupted and an attitude settled in that at best tolerates uncertainty without despairing.
We are free but also overwhelmed—that is perhaps the dilemma of liberal democracy: we do not know what to do and must decide nonetheless.</p>

<p>Authors such as Nietzsche and Heidegger offer vocabularies for self-creation, for individual meaning-making beyond metaphysical certainties and beyond reason.
Nietzsche is perhaps the epitome of the <em>ironic theorist</em> who unmasks every supposed truth as a contingent product of a historical narrative—except, of course, his own.</p>

<p>He is an artist: playful, contradictory, vital, desperate, brutal and jolly who teaches us that no description of the world is necessary—that we could always tell other stories to understand ourselves and our community differently.
Kant, on the other hand, reminds us that such new creations have limits where they overlook the suffering of others.</p>

<blockquote>
  <p>Rationality is a matter of making allowed moves within a language game. Imagination creates the games reason proceeds to play. Then […] it keeps modifying those games so that playing them is more interesting and profitable. Reason cannot get outside the latest circle that imagination has drawn. It is in this sense, and <strong>only</strong> this sense, that imagination holds the primacy – <a class="citation" href="#rorty:2016">(Rorty, 2016)</a></p>
</blockquote>

<p>Rorty recognizes in this tension the productive condition of a liberal modernity: without Nietzsche, no renewal; without Kant, no consideration. <strong>We need Nietzsche so that life remains interesting, and Kant so that it remains bearable.</strong></p>

<blockquote>
  <p>Rational discussion is not an appeal to eternal standards, but simply an attempt to make our beliefs and desires as coherent with one another as possible while constantly adding new beliefs and desires to the old. – <a class="citation" href="#rorty:1990">(Rorty, 1990)</a></p>
</blockquote>

<p>However, Rorty emphatically reminds us that most people do not want to be redescribed, and that an imposed redescription is usually cruel.
People generally want to be taken as they speak.</p>

<blockquote>
  <p>There is something potentially very cruel about the claim that [the language people speak is, for the ironist, a matter of chance]. For the most effective way of causing people enduring pain is to humiliate them by making the things that seemed most important to them look futile, obsolete, and powerless. – <a class="citation" href="#rorty:1989">(Rorty, 1989)</a></p>
</blockquote>

<p>I think this is one of the most remarkable statements, one that emphasizes the dark side of the <em>ironic attitude</em>.
The ironist knows that every vocabulary (the way we talk)—every language in which people give meaning to their lives—is contingent, and therefore replaceable, not final, but neither useless nor arbitrary.
An ironist—such as Nietzsche—can, at any moment, show that the concepts by which someone lives are merely historical accidents.
But that is precisely what can be a form of cruelty.
For if I convincingly demonstrate to someone that everything he believes in, lives for, and that constitutes his identity is only a “matter of chance,” I have not simply offered him a better argument.
I have pulled the ground out from under his feet.</p>

<p>This “ground” is nothing less than the person’s <em>horizon of sense</em> (Sinn), i.e. a horizon of an inescapable medium; we cannot step outside of it <a class="citation" href="#luhmann:1984">(Luhmann, 1984)</a>. A person’s final vocabulary is the precise mechanism they use to select meaning out of a chaotic world and stabilize their mind. When the ironist ruthlessly exposes this vocabulary as a mere accident, they do not open a door to total freedom; instead, they threaten the person with <em>structural collapse</em>.</p>

<p>Rorty urges caution when it comes to proposing other vocabularies.
In the public sphere, in dealings with others, irony is potentially cruel if it redescribe someone’s past and make them look foolish or obsolete—and yet it is new vocabularies that enable moral progress, while at the same time the replacement of vocabularies can be the worst form of cruelty. It is the ultimate tragic bind: we must change our vocabularies to progress, but in doing so, we risk inflicting a structural violence that strips others of their very capacity to make sense of the world.
<em>Solidarity</em> itself can cause smaller and smaller circles if it defines its identity through the creation of sharp boundaries, turning solidarity into an internal code that treats everyone outside the group as mere background noise or structural threats.</p>

<blockquote>
  <p>All we can do is to compare new customs and institutions with old customs and institutions in the experimental and tentative way in which we compare new friends, new jobs, or new environments with old ones. The only test of truth is that it is the view that wins in a free and open encounter. But the result of that test can only be accepted until somebody comes up with some new proposal, a new scientific theory, a new artistic style, a new political institution. Then discussion will have to be undertaken all over again. There will never be a time when Socratic questioning becomes unnecessary. […] Kant said, following Plato, that the source of moral obligation must be a distinct faculty—reason rather than emotion—because reason is part of human nature whereas emotions are just contingent features of particular individuals. Even someone like myself, who wants to discard the notion of intrinsic human dignity and unconditional moral obligation has to admit that these Kantian uses have been extremely useful. The <strong>ethics of sensitivity</strong> which I have associated with the figure of the literary critic, may seem to endanger all the gains made in recent times with the help of these Platonic and Kantian notions. [… However] we may find a way to do it without the ladder we climbed. […] An ethics of sensitivity assumes that morality is not a matter of recognizing unconditional obligations built into every human being simply by virtue of being human, but rather of community obligations—obligations one feels as a member of a group. […S]uch obligations determine one’s identity as a member of a community. It is one thing to treat someone weaker less advantaged than oneself decently because one happens to feel kindly toward him—perhaps because of some unconscious accidental association with one of his features or trades. It’s another thing to recognize this person as a fellow citizen, one of <strong>us</strong>, the sort of person to whom <strong>we</strong> are obliged to behave decently. […] It is a matter of coming to see more and more different sorts of people as us. Seeing an individual who lives quite a different life form as our own as, nontheless, one of us. – <a class="citation" href="#rorty:1990">(Rorty, 1990)</a></p>
</blockquote>

<p>Rorty advocates for a liberalism for which there is nothing worse than humiliating others, for humiliation is the deepest form of cruelty <a class="citation" href="#rorty:1989">(Rorty, 1989)</a>.
It is the destruction of the self-image and of the language in which a person gives meaning to their life.
The problem with Rorty, one that has perhaps caught up with us today, is that for him there are no <em>final grounds</em> with which he could defend his liberal position.</p>

<blockquote>
  <p>Substituting this sort of practical question for theoretical questions about first principles means admitting that there is no way to answer such critics of democracy such as Plato, Nietzsche, or Hitler. There is no neutral ahistorical ground one could stand on when members of a democratic community try to argue with people who ask whether their society may not be headed in exactly the wrong direction. First principles are rationalizations of existing habits and institutions […] which is no reason to distrust them automatically. The Homeric heroes, the Nazi concentration camp guards, the pre-civil war slave owners all had principles. But their principles did not save them from cruelty to people whom they did not think of as us. What counts for moral progress is not firmness in abiding by established habits or institutions or principles, but rather the willingness to ask who’s getting hurt by the existence of these institutions or by the application of these principles. – <a class="citation" href="#rorty:1990">(Rorty, 1990)</a></p>
</blockquote>

<p>The conviction that cruelty is the worst thing comes from a particular historical and cultural development, from the European Enlightenment, from democratic revolutions, from the gradual expansion of the circle of those to whom we extend compassion.
It is <em>contingent</em>, that is, neither necessary nor impossible.
But that does not make it any less valuable or binding for Rorty.
One can passionately commit to something without claiming: the universe stands behind it.</p>

<p>Why, then, should we not humiliate others?
Because through experience, through stories, through listening, we have learned what it feels like to be humiliated.
Because we have built a culture that cultivates this sensitivity.
And because the attempt to provide a deeper justification for this leads us astray: it suggests that someone not convinced by the argument could be rationally refuted.
But the <em>liberal ironist</em> cannot accomplish this.
The sadist, to whom the suffering of others is not only indifferent but who seeks elevation and a <em>destructive vitality</em> by humiliating others, lacks not an argument according to Rorty—he lacks a certain capacity for empathy that cannot be logically derived, but can only be cultivated.</p>

<blockquote>
  <p>The development of civilization on this view is not the triumph of reason over passion but <strong>the triumph of tolerance over distrust</strong>. Therefore, democratic society is not founded on a sense of obligations, but on a sense of sympathy. […] The increasing egalitarianism of the democracies is not a matter of recognizing that illiterate laborers, blacks, women, and gays are as rational beings as middle-class straight white males, but rather of those males themselves—the people who have a monopoly on power—coming to realize that these people have the same hopes and fears and the same susceptibility to pain and humiliation as they do. – <a class="citation" href="#rorty:1990">(Rorty, 1990)</a></p>
</blockquote>

<p>Heidegger diagnosed in his Spiegel interview <a class="citation" href="#heidegger:1976">(Heidegger, 1976)</a> the powerlessness of all theory and politics before the <em>forgetfulness of Being</em> in modernity. Such a diagnosis Rorty would have rejected in its metaphysical depth. And yet both, on very different paths, share the skepticism toward philosophy as savior: Heidegger, because only a god could help; Rorty, because there are no ultimate justifications.
But Rorty avoids the sort of despair Heidegger seemed to hold.
Christianity, Kant’s <em>Reason</em>, and the <em>crisis of the absolute</em> had their time.
They told new stories to bring meaning back into descriptions.
The <em>liberal and democratic society</em>, however, cannot be universally grounded—it can only be narrated <a class="citation" href="#rorty:1999">(Rorty, 1999)</a>.
It lives as long as we find new words for “the good,” “the just,” and “the common.”
It would, however, be dangerous to transfigure it as “the genuinely true,” “the superior,” or “the absolutely just.”
Its salvation lies not in <em>truth</em> but in conversation; in the shared resolution:</p>

<blockquote>
  <p>We don’t do that. We respect one another. We don’t kill other people. We support each other. We always doubt our own sensitivity.</p>
</blockquote>

<p>Rorty knew that the loss of absolute truth would be unsettling for the individual, and in that context proposes distinguishing between the <em>private</em> and <em>public</em> spheres.
In private, the ironist knows that her convictions are contingent (accidental, historically conditioned).
She knows that her values are not God-given.
She has doubts, and that is all right.
She can pursue her striving for self-realization.
Thus Rorty relocates the impulse toward self-creation—which he takes seriously and regards it as important, and which he sees embodied in Nietzsche, Proust, Heidegger, and Derrida—into the private realm.</p>

<blockquote>
  <p>But for Proust and Nietzsche, there is nothing more powerful or important than self-redescription. They are not trying to overcome time and chance but to use them. – <a class="citation" href="#rorty:1989">(Rorty, 1989)</a></p>
</blockquote>

<p>His <em>liberal ironists</em> no longer concern themselves with universals, essences, or absolutes that might metaphysically ground why one should not humiliate others.
They enjoy the writings of Nietzsche, Heidegger, Derrida, and many other <em>ironic theorists</em> for their own <strong>private redescription</strong> without searching for a conclusive public vocabulary.</p>

<p>Rorty saw in Nabokov a central tension of his own philosophy: the conflict between private aesthetic self-creation and public solidarity.
For Rorty, Nabokov embodied the type of the liberal ironist who strives for autonomy and artistic ecstasy, but in doing so runs the risk of becoming cruel.
Thus, for example, Nabokov’s Humbert is highly educated, sensitive, and writes beautifully.
But precisely this search for aesthetic ecstasy makes him blind to the suffering of others.
He does not see Lolita as a suffering child but as an aesthetic object of his fantasy.
Nabokov thereby shows that neither intelligence, artistic sensibility, nor linguistic brilliance automatically makes us morally good people.
One can be an artistic genius and still act cruelly, and although Nabokov never saw himself as a teacher, he helps his reader to notice cruelty in detail.
We therefore need authors like Nabokov for our private lives (to invent ourselves, to be autonomous, to take pleasure in language).
But, Rorty argues, we must not carry this attitude into the public sphere, because a society based solely on aesthetic pleasure would be cruel.</p>

<p>In Rorty’s <em>ironic culture</em> it is no longer a matter of finding the right words or a common reality behind appearances that will connect people.
He gives up the dream of finding the True, the Good, and the Beautiful for <strong>everyone</strong>.
That is why we should stop asking:</p>

<blockquote>
  <p>Does my belief correspond to <em>true</em> reality?</p>
</blockquote>

<p>More important is the question:</p>

<blockquote>
  <p>Does this belief help us live together better, more freely, and less cruelly?</p>
</blockquote>

<p>In his utopia, solidarity is not regarded as a fact that must be recognized by removing prejudices or uncovering previously hidden depths.
Any uncovering becomes impossible once we abandon the search for a common essence or nature, for the effort toward universal justification is then in vain.
Solidarity must therefore be constructed.
Thus Rorty’s liberals make solidarity their goal.
This goal is achieved not through <em>reason</em> or any logical proofs, but through <em>imagination</em> and <em>empathy</em>—through the imaginative capacity to see unfamiliar people as fellow sufferers.</p>

<blockquote>
  <p>The philosopher is driven by what Dewey called the quest for certainty and the literary critic by curiosity and sympathy. The latter is curious about forms of life different from her own and sympathetic to people who leads such lives. The former looks for unity and thinks of philosophical inquiry as converging to a single body of truths. The latter looks for diversity. She is more concerned having missed something, having been condescending and cruel towards someone of a different sort than about certainty. – <a class="citation" href="#rorty:1990">(Rorty, 1990)</a></p>
</blockquote>

<p>Solidarity is not discovered through reflection; it is created.
It is created by increasing our sensitivity—through literature, novels, films, theatre, music, (computer) games, and reportage—to the particular subtleties of the pain and humiliation of others, people unknown to us.</p>

<blockquote>
  <p>[The literary critic], on my view, [is] our modern substitude for the Platonic moral philosopher. […] The traditional notion of a separation between moral judgments and aesthetic judgements is, I would claim, a relic of the idea that there is a deep common human nature which sets moral goals and standards. [… In our times] it is almost impossible to believe that all human beings, male or female, slave or free, illiterate or cultured, ancient or modern, European or Chinese, have always carried around the same vision of goodness and justice deep within themselves. Everything […] suggests the plasticity of human beings. – <a class="citation" href="#rorty:1990">(Rorty, 1990)</a></p>
</blockquote>

<p>Such heightened sensitivity makes it harder to marginalize people who are different from us by assuming that they do not feel the way we do, or that there will always be suffering, so why should one not let them suffer?
It is a process of emotional expansion of the “we”—a change in perception, not in epistemological cognition in the form of an increase in knowledge.
This process of seeing other people as “one of us” rather than “the others” consists in describing in detail what unknown people are like, and in redescribing who we ourselves are and how we recognize one another <a class="citation" href="#rorty:2016">(Rorty, 2016)</a>.</p>

<p>Cruelty and humiliation are abolished by creating realities through dense descriptions that sensitize the reader to the pain of those who do not speak our language.
This task had been expected of proofs for a common human “nature,” but whether grounded in metaphysics or natural science, it can capture the indeterminacy and paradox of the human being in no final description.
As Nietzsche already pointed out,</p>

<blockquote>
  <p>[i]t is only through the forgetting of that primitive metaphor-world, only through the hardening and stiffening of an original mass of images, flowing liquid and hot out of the primal power of human fantasy, only through the invincible belief that this sun, this window, this table is a truth in itself—in short, only through the fact that man forgets himself as a subject, and indeed as an artistically creative subject—does he live with any repose, security, and consistency: if he could get out of the prison walls of this belief, even for an instant, his ‘self-consciousness’ would be immediately destroyed. – <a class="citation" href="#nietzsche:1873">(Nietzsche, 1873)</a></p>
</blockquote>

<p>and Rorty follows him in this insight:</p>

<blockquote>
  <p>We should understand concepts like ‘gravitation’ or ‘human rights’ not as entities whose essence remains mysteriously hidden, but as sounds and marks whose use has made possible more significant and better social practices. Intellectual and moral progress is not an approximation toward a prior goal, but the surpassing of the past. What we call ‘improved knowledge’ should not be interpreted as better access to the real, but as an enhanced ability to do things. […] Freedom begins when we can discuss which words better describe a situation. <strong>Knowledge and freedom develop simultaneously.</strong> – <a class="citation" href="#rorty:2016">(Rorty, 2016)</a></p>
</blockquote>

<p>There is no final, conclusive vocabulary; every language in which people give meaning to their lives is fundamentally contingent. Normally, our shared vocabularies function as a vital baseline for peace. They act as the agreed-upon frameworks where questioning ceases and mutual understanding is stabilized <a class="citation" href="#maturana:1991">(Maturana, 1991)</a>.</p>

<p>“Be objective” is, from the perspective of metaphysical realism, a <em>demand</em> to accept my position—to deny oneself to some degree and to accept a shift in once identity. Complying with this demand requires a kind of Buddhist self-emptying: you are asked to detach from your own contingent horizon of meaning and quietly dissolve your ego, not to touch an ultimate truth, but simply to clear away yourself so that the other person’s vocabulary can occupy the space.
Yet, as a society, we do not truly value or appreciate the immense weight of this act. We casually demand objectivity from others as if it were a simple intellectual correction, entirely blind to the fact that we are asking them to perform a deeply painful, ascetic surrender of their very selfhood just to accommodate our view of the world.</p>

<blockquote>
  <p>The contrasting view [of getting at a certain answer to questions like “What is really good?”, “What is really just?”, “What is really real?”, “What is really true?”, “What is really human?”], which I share with people like William James, Dewey, [and] Satre is that the point of Socrates’ life was not to discover a permanent absolute truth but rather just to keep people thinking, to keep them inventing, to open up their imaginations to alternatives to present convictions. – <a class="citation" href="#rorty:1990">(Rorty, 1990)</a></p>
</blockquote>

<p>Once liberal society has become distrustful of its “inner core,” it can no longer appeal to that core.
What society has learned about itself it will not easily be able to forget.
In other words: Nietzsche’s cut cannot be papered over even when the author has been exposed as a moral reprobate.
Open society will not be saved by a universal truth; not by reason and not by a metaphysical foundation.
Rorty thus reminds us that there is no final authority that could assure us that freedom, equality, or compassion are true and good.
Their validity is not the result of discovery but of narration.</p>

<blockquote>
  <p>To progress morally is a matter of individual and social self-creation rather than self-discovery. – <a class="citation" href="#rorty:1990">(Rorty, 1990)</a></p>
</blockquote>

<p>Precisely for this reason, Rorty burdens literature with perhaps the most important task: to foster the imagination for ways of living, so that in the future we will not say that we lead a more truthful or more sublime life than our ancestors, but that we have developed better ways “of being human” that our descendants may perhaps adopt.</p>

<blockquote>
  <p>If one asks which books helped along such processes such of self-creation in the last few hundred years I think a good case can be made for saying that most of them were novels. Whereas our ancestors relied on scripture or on theological or philosophical treatisis for their notion of what it was to be a human being and what was the point of human life, more recently, we have been relying on books like The Brothers Karamazov, The Magic Mountain, Remembrance of Things Past, and Catcher in the Rye. Novels about young people growing up and creating themselves. If one asks which books have done most to make American society freer and more just, again, a lot of them are novels. Books like Uncle Tom’s Cabin, Black Boy and Invisible Man did more than any philosophical or social scientific treatises to let the whites see what they were doing to the blacks. Books like The Well of Loneliness and The City and the Pillar did more than psychological treatises to let the straits see what they had been doing to the gays. Books like Middle March and The Color Purple did more to make men realize what they were doing to women than any socioeconomic data or any feminist theorizing. – <a class="citation" href="#rorty:1990">(Rorty, 1990)</a></p>
</blockquote>

<p>In his admiration for the ironic theorists, Rorty simultaneously warns us against their absorbing tendencies to want more than self-creation.
For all his love of life and self-redescription, Nietzsche was hardly concerned with preventing cruelty; rather, he was concerned with enabling the few—among whom he counted himself—to invent and describe Europe as something great.
In this sense, Rorty is right to suggest that Nietzsche’s vocabulary can unfold its power in the <em>private</em> domain of our lives, but produces little that is fruitful in the discourse of the open society.
When the work of art that we ourselves are—the one that is neither simply created nor discovered but pre-exists within us and can yet only be found in the process of creation—destroys the possibility of another work of art coming into being, that is, stifles the self-creation of others, then a deep cruelty arises.
Autonomy and self-efficacy are ideals that liberals also revere, but in the darkness lurks the connection between art and torture.
The work of art is not “the good”—even if it is beautiful, it can be cruel.</p>

<p>Over time we have expanded our moral imagination and included more beings in our <em>circle of care</em>, not on account of unchanging principles, not because of God or some inner truth we would have discovered, but because of the vitality of our imagining; not because of transcendental or transcendent truth, but because of lived and imagined experiences that have been narrated and inscribed into our cultural memory.
We have invented a less cruel world.</p>

<p>This at times naively appearing hope—this project of continuation—gains contour where genuine encounters take place in the <em>lifeworld</em>.
On account of my physical impairment, I possess the paradoxical privilege of evoking in others a <em>suspicion of humiliation</em>: the mere appearance of my body leads others to assume that cruelty must inevitably have been done to me.
Even when these projections rest on stereotyped prejudices, they reveal the human capacity for empathy.
The suspicion is plausible, and not entirely wrong.</p>

<p>Two fundamentally different ways of meeting this suspicion crystallize: The first is something like a Christian glorification of suffering.
Here the fateful is elevated to an essence; the sufferer is styled as a martyr whose mode of being is thereby definitively inscribed.
In this logic, suffering lies “in her nature”—a convenient ontology that releases us from the question of how a shared life would have to be arranged to reduce this suffering. It is precisely that form of pity that Nietzsche mocked in Schopenhauer’s writings: a pity that fixes the other in her weakness rather than allowing her to invent herself.</p>

<p>The second, solidary path, on the other hand, must always preserve <em>contingency</em>.
Solidarity here means conceiving of the other as a space of possibilities and remaining conscious of the limits of one’s own descriptions.
It begins not with a judgment, but with the radically open question: “How can I help to reduce the cruelty that befalls you?”
This attitude acknowledges that there are no universally valid answers.
While in the public sphere we must fight for the conditions of a dignified life, the space of private <em>self-creation</em> must remain open—as that place where every person may draft their own image, beyond the gaze of others and their pitying diagnoses.</p>

<p>Yet we sometimes find it hard to hold a contingent evolutionary history responsible for our “place.”
Instead we often need a culprit who protects us from the recognition that the human being</p>

<blockquote>
  <p>hangs on the back of a tiger in dreams. – <a class="citation" href="#nietzsche:1873">(Nietzsche, 1873)</a></p>
</blockquote>

<p>Malice is never long in coming.
It surfaces when the other manages to no longer see their own vulnerability in the counterpart, cannot imagine how it might be like to be “the other”, or to believe they can no longer afford to be sensible, given social, material, bodily or psychological compulsion.</p>

<p>Some cruelty becomes visible, acquires a language, and other cruelty remains hidden, and when it cannot “heal” or express itself at all and thereby reflect on itself, it will have to prove itself—that is, it will (re-)produce the conditions for its own continued existence.
Without doubt about one’s own sensitivity to the pain and humiliation of others, curiosity about possible alternatives remains rude and calculating.</p>

<p>For Rorty, <em>solidarity</em> becomes the capacity to “see more and more” rather than seeing an “inner core.”
It is the capacity to count as “us” people (and other living beings) who are worlds apart from us.
It is grounded not, as in Kant, on universal Reason, nor, as in Christianity, on God, but on the striving to prevent and alleviate cruelty and pain.</p>

<blockquote>
  <p>But even when we use neither Kantian nor Christian language, we may still have the feeling that it is dubious to be more concerned about the living conditions of a fellow citizen of New York than about someone living equally hopelessly and miserably in the slums of Manila or Dakar. […] On the other hand, it is <em>not</em> incompatible with my position [(which accepts no essences)] to insist that we must try to include in our understanding of “we” also people whom we have so far counted among the “they.” This claim, characteristic of liberals who fear their own cruelty more than anything else, rests solely on the […] historical contingencies—namely, the development of the moral and political vocabularies typical of the secularized democratic societies of the Western world. – <a class="citation" href="#rorty:1989">(Rorty, 1989)</a></p>
</blockquote>

<p>Placing this “liberal axiom” above the sublime cannot be defended in any neutral, non-circular way.
But the same holds for Heidegger’s claim that the idea of “the greatest happiness for the greatest number” is also just a piece of metaphysics, a piece of <em>forgetfulness of Being</em>.</p>

<p>Rorty cannot claim that avoiding cruelty is universally good.
He cannot even claim that his assertion that there is no absolute truth is itself absolutely true.
It is the classic self-contradiction objection—whoever says “there is no absolute truth” seems thereby to claim the very thing they reject.
Rorty was of course aware of this.
His move consists in not playing the game at all. Rorty does not claim: “It is objectively true that there is no objective truth.”
That would be self-contradictory.
Instead he says something like: The concept of “objective truth” is no longer a useful tool.
We should stop using it—not because we have proved that there is no objective truth, but because a vocabulary without this concept gets us further.
Rorty does not position himself on a higher level and delivers a verdict about truth from there.
He recommends a different vocabulary. He says: Try it once without the concept of “absolute truth” and see whether you manage better.
That is not a thesis about the world, but a suggestion about how we ought to speak.</p>

<p>When someone in the seventeenth century stopped thinking in the categories of scholasticism and adopted instead the language of the new natural science, they did not prove that scholasticism was wrong.
They simply tried out a different vocabulary and found it more fruitful.
According to Rorty, the same applies to the concept of truth: he does not want to refute the existence of truth, but to suggest that without this concept we can conduct more interesting conversations.
Rorty accepts that his own position is just as <em>contingent</em> as any other.
He claims for it only that it is more useful—more useful for the project of reducing cruelty and expanding solidarity.</p>

<p>Critics such as Hilary Putnam or Thomas Nagel have responded that Rorty cannot really save himself here: if he says his vocabulary is “more useful,” that again contains a truth claim—namely that it really is more useful.
Rorty would reply: “useful” too is not an objective criterion but an evaluation from a particular perspective. The game can thus be continued indefinitely, and that is precisely Rorty’s point—there is no place at which one finally arrives. The philosophy that searches for such an endpoint pursues, according to Rorty, a goal that does not exist.
Whether one finds that convincing ultimately depends on whether one is willing to abandon the desire for ultimate justification.
Rorty demands that of his readers and he knew that many would not comply.</p>

<p>So, what would Rorty propose to us today? Certainly no grand promises or revolutions, but decidedly radical changes that in his time still sounded pragmatic.
Authors such as <a class="citation" href="#han:2021">(Han, 2021)</a> describe the achievement society as a place of <em>friendly violence</em> in which self-creation degenerates into self-optimization and the human being becomes mere <em>standing-reserve</em>—a concept that Han borrows from Heidegger’s critique of technology (cf. <a class="citation" href="#heidegger:1954">(Heidegger, 1954)</a>).
Here Rorty would react with skepticism, and not only toward the diagnoses, but above all toward the vocabulary.
He fundamentally distrusted Heidegger’s metaphysics of <em>forgetfulness of Being</em>: it runs the risk of <strong>condemning modernity as a whole</strong> rather than naming and addressing concrete grievances.
For Rorty this is a philosophical luxury neither progressives nor conservatives who actually wants to change something cannot afford.
This does not mean that Rorty would simply dismiss Han’s observations. The empirical description that people suffer from exhaustion, that the pressure of self-optimization destroys social solidarity, that depoliticization is a danger—these he would share.
But his answer would be pragmatic and reformist, not cultural-critical: improve working conditions and educational possibilities, guarantee social security, create spaces for purposeless exchanges, so that people once again have the leisure to redescribe themselves rather than optimizing themselves for the market.
These are <em>bread-and-butter questions</em> of politics, and it is precisely there, not in the <em>deep diagnosis of the occidental history of Being</em>, that Rorty would begin.</p>

<p>And since for Rorty progress is the capacity to allow one’s own final vocabulary to be expanded by that of the other, through literature and encounters (<em>sentimental education</em>), algorithms that isolate us in echo chambers destroy precisely this capacity for empathy.
He would probably reject digital surveillance as a new form of cruelty that prevents solidarity.</p>

<p>Futhermore, his hope for an ever-growing solidarity presupposes that the material conditions still permit this expansion at all.
Here, Carolin Amlinger reports:</p>

<blockquote>
  <p>the resentful feel fundamentally blocked in their progress as if ones life is overlaid by mud.</p>
</blockquote>

<p>That is why he would agree with Han’s observation that “depoliticization” is a danger, but he would call for speaking once again about “bread-and-butter issues” rather than philosophically condemning the whole of modernity.
Here, however, Rorty’s approach runs up against a limit that he himself did not see: his concept of solidarity remains <em>anthropocentric</em>—he asks who can suffer pain and humiliation and draws the circle of the moral exactly there.
<strong>The earth does not speak</strong>, so it does not feature in his account.
Yet Rorty seemed open to move beyond human solidarity.</p>

<blockquote>
  <p>Imagination, in the sense in which I am using the term, is not a distinctively human capacity. It is […] the ability to come up with socially useful novelties. This is an abilitiy Newton shared with certain eager and ingenious beavers. But giving and asking for reason <strong>is</strong> distinctively human, and in coextensive with rationality. The more an organism can get what it wants by persuasion rather than force, the more rational it is. – <a class="citation" href="#rorty:2016">(Rorty, 2016)</a></p>
</blockquote>

<p>Yet if we take Rorty’s own logic seriously—namely that the circle of solidarity must be drawn ever wider—then the question inevitably arises of whether this circle may stop at the human being.
A solidarity that applies only among speakers fails to hear the silence of those who have no language.
Perhaps that is the blind spot that costs us the most dearly today.</p>

<blockquote>
  <p>Our ancestors 500 years ago simply could not have grasped what now seems to us common sense that differences of religion, race, and social status are morally irrelevant. These ancestors weren’t blind to something we now see because there wasn’t yet anything for them to see. What we see had to be created in the interim. The human race has been busy creating itself over the last 500 years, creating moral obligations for itself which were once mere fantasies in the minds of a few people of unusually vivid imagination and unusually broad sympathy. I want to suggest that someday if this notion of humanity’s self-creation comes to replace the traditional philosophical notion of humanity’s self-understanding, the colleges and universities might be able to stop using even for commercial purposes the Platonic rhetoric of a quest for eternal truth. They might openly proclaim that their principle function is to keep society from ever being satified with itself, to keep individual students from being satified with themselves as they were when they arrived. If this happens democratic society might lose their habitual distrust of intellectuals. This would happen because […] whole societies would get intellectualized—not in the sense of being turned into a nation of philosophers, but in the sense of being turned into a nation of literary critics, that is, people curious about alternative forms of life and constantly sensitive to the possibility that they may be being as unconsciously cruel as their ancestors were. Such societies would still think of the education of the young as a matter of instilling traditional values. But the principle traditional values they have in mind would be simply the value of questioning traditions. – <a class="citation" href="#rorty:1990">(Rorty, 1990)</a></p>
</blockquote>

<hr />

<h2 id="cultivating-sensitivity">Cultivating Sensitivity</h2>

<p>The <em>urge toward new hardness</em>, toward strong men, clear images of the enemy, and national self-assertion, is not a sign of strength.
It is a sign that the circle of solidarity is shrinking.
Rorty would not have been surprised.
He already warned in <em>Achieving Our Country</em> <a class="citation" href="#rorty:1998">(Rorty, 1998)</a> that a left that retreats into cultural and academic distinction and stops speaking about material inequality prepares the ground for precisely those resentments that today call for hardness; that calls on their base to lean into every negative emotion and tells them that anything is exactly as they believe it is; that every perceived enemy, every perceived corruption, every perceived societal weakness is legitimized; that every frustration one has is legitimized and that any act in response is justified. It is an attitude that is displayed in the movie <em>Citizen Vigilante</em>.</p>

<p>Because of the political development of the past, we are in the unfortunate position that some of the moral progress correlates with an increase in material inequality such that the emontional path from correlation to causation is short and easy to take. This leads to a sort of alienation from society <a class="citation" href="#amlinger:2025">(Carolin Amlinger, 2025)</a>.</p>

<blockquote>
  <p>You could earn a lot of money <strong>and then</strong> feminists came in. – Interviewee</p>
</blockquote>

<p>The core promises of modern society, namely social integration and individual emancipation, seem no longer to function effectively.
Simultaneously, the modern individual experiences institutions as a restrictive barrier to their self-realization, thus society becomes itself a source of cruelty and humiliation.
The disappointed then find their voice not among those who advocate solidarity, but among those who name enemies.</p>

<blockquote>
  <p>Destructiveness is the result of an unlived life. – <a class="citation" href="#fromm:1973">(Fromm, 1973)</a></p>
</blockquote>

<p>What liberals can learn from Rorty is first of all an attitude of humility: those who stand up for liberal democracy should stop transfiguring it as the <em>genuinely superior</em>, the <em>historically necessary</em>, or the <em>rationally only possible</em>, for this rhetoric acts on those who do not share it as <em>humiliation</em>.
Rorty proposed defending liberalism not as truth but as habit—as a way of living together that we have laboriously acquired and that is worth continuing to narrate.
Liberal democracy does not need philosophical justification.
It is pragmatically justified because, for Rorty, it works better at reducing cruelty and increasing human freedom than anything else we’ve tried.
But <strong>if we cease to cultivate a desire for the reduction of cruelty, liberal democracy will inevitably vanish</strong>.</p>

<p>The resistance to new hardness cannot be primarily argumentative.
Those who want to draw smaller circles do not convince people via rational arguments but through emotions.
In fact, <strong>cruelty is the point</strong>; it is a necessity in a zero-sum game.
It is a desired feature of their politics to humiliate their perceived enemies even if their actions go against their base because it is cruelty that can be employed to get back what can never be owned: a nation, a race, a country, a culture, a planet, a superintelligence.</p>

<p>One does not refute resentments through better arguments; one changes them through better stories, that is, through narratives that give a face to people regarded as foreign, that show what it feels like to be excluded, persecuted, humiliated and by creating the conditions that everyone finds a niche in which they can self-realize themselves.
That is not a weakness of the liberal project but its actual strength: not persuasion through logic, but expansion of the imagination.</p>

<p>Third, and most uncomfortably, Rorty demands that liberals stop overlooking their own capacity for cruelty.
A dangerous, defensive hardness thrives among those who feel that their “final vocabulary”—their way of describing their life, giving dignity to their work, and finding meaning in their community—is being ridiculed by the architects of openness and tolerance.
When we demand that others abandon their worldview and adopt our enlightened descriptions, we are casually asking them to perform that same agonizing, unappreciated act of self-emptying. We expect them to dissolve their ego and identity to accommodate our vocabulary, while offering them no social value or grace for that immense sacrifice. And when they resist this spiritual violence, we respond with structural exclusion.</p>

<p>A solidarity that applies only among the “educated”, while blinding itself to the struggles of workers, is no solidarity at all. When we force our description of the world onto others, we risk committing the ultimate Rortyan sin: stripping them of their agency and ignoring their capacity to suffer. A racist acts with immense cruelty by denying the humanity of others. Yet, Rorty warns that when liberals respond by treating the racist as an inherently evil, subhuman monster, they slip into a parallel form of cruelty.</p>

<p>By excommunicating the racist from the horizon of valid human communication, liberals treat them as an irredeemable object rather than a poorly socialized human being. They turn their own vocabulary of tolerance into a new mechanism of humiliation—a way to make the other look permanently obsolete and futile. Rorty insists that liberals root out this structural sadism in themselves before pointing fingers.</p>

<p>Furthermore, Rorty reserves a distinct scorn for any kind of performative moral high ground.<sup id="fnref:3" role="doc-noteref"><a href="#fn:3" class="footnote" rel="footnote">3</a></sup>
He famously critiqued the “spectatorial left” for transforming politics into a theater of cultural purity testing, where naming vices and maintaining a flawless vocabulary replaces the <strong>heavy lifting of material reform</strong> <a class="citation" href="#rorty:1998">(Rorty, 1998)</a>.
For Rorty, a moral posture that exists only to signal its own righteousness does nothing to alleviate pain; instead, it becomes a weapon of exclusion, mocking the very workers and poorly socialized individuals it claims a democratic society should integrate.</p>

<p>No absolute moral law resolves this <em>paradox of tolerance</em>.
It cannot “solve” Gaza, the West Bank or mass killings in Sudan. There are only <em>pragmatic</em> choices.
Arguing over theoretical definitions of “genocide” or “apartheid” does nothing to stop a bomb from falling or a family from being displaced because <strong>laws are only binding if people feel like obeying them</strong>.<sup id="fnref:4" role="doc-noteref"><a href="#fn:4" class="footnote" rel="footnote">4</a></sup>
Legal documents themselves are just pieces of paper unless they are backed by a <em>human rights culture</em>.
And in a <em>liberal culture</em>, no political goal should justify the torture or starvation of a population.
We must confidently assert that our tribe’s way of living, that is, one that refuses to tolerate racism, is simply better at minimizing cruelty, and we will therefore enforce it through political power.
We can recognize that the racist is a human being capable of suffering without letting that recognition paralyze our defense of a democratic society.
Thus, the West has a pragmatic duty to use its political and economic power to enforce a settlement to stop the exertion of cruelty.</p>

<p>Yet Rorty knew how easily this pragmatic enforcement of power can warp into state-sponsored paranoia. His immediate, visceral opposition to the post-9/11 “War on Terror” and the invasion of Iraq was driven by this exact anxiety.
He recognized that the moment a democracy prioritizes absolute safety over civil liberties, it surrenders its moral authority.
He viewed the weaponization of fear after 9/11 as a disastrous distraction from building a more just, equal, and inclusive society. 
And he called this manifestation “simple-minded militaristic chauvinism”. 
Trading freedom for security was a cynical political maneuver that exported violence abroad while bleeding away the resources needed to fix internal failures like poverty and systemic racism.
A society terrified of an invisible enemy inevitably stops listening to its citizens and starts listening to a “strongman”.</p>

<p>In short: Today, Rorty would not call on us to save democracy by philosophically refuting its enemies.
He would call on us to narrate it—again and again, in new words, for people who do not yet recognize themselves in it.
But this requires that our actions converges to the image we have of ourselves.
Instead of giving up on <strong>cultivating the art of solidarity</strong>, Rorty would call on us to reconfigure our social reality in such a way that it fits our description of a liberal democracy; to go so far to say: it is more important to be sensible to cruelty than to find “the Truth”.</p>

<blockquote>
  <p>Orwell’s main concern is to sensitize an audience to cases of cruelty and humiliation which they had not noticed. [… Moral progress is] a matter of separating the question “Do you believe and desire what I believe and desire?”—a representational question—from the question “Are you suffering?” – <a class="citation" href="#rorty:1989">(Rorty, 1989)</a></p>
</blockquote>

<p>But that requires time.
Not the time of rapid consumption, but the slow time of reading.
It requires that patience which allows one to immerse oneself in an unfamiliar form of life, to understand it from within, before evaluating it.
Rorty knew that solidarity does not arise in seconds.
It grows in the hours spent with <em>Raskolnikov</em>, <em>Dorothea Brooke</em>, or the <em>Invisible Man</em>—with people one will never meet and whom one nonetheless begins to know.</p>

<p>Here, as a new communication medium participating in communication for the first time (cf. <a class="citation" href="#esposito:2022">(Esposito, 2022)</a>), <em>(Generative) AI</em> transforms precisely this process, and in a way that would have troubled Rorty.
Not because the technology is evil, but because it accelerates, compresses, and smooths encounters with the unfamiliar.
What a novel laboriously builds over two hundred pages—for example, trust in an unfamiliar voice, the capacity to endure contradictions, the slow understanding of another world—an LLM can summarize in a few paragraphs.
But whether the expansion of the imagination that Rorty had in mind still arises in the process is questionable.
Summaries do not sensitize; they inform.
And the difference between the two is, for Rorty, the <strong>difference between knowledge and solidarity</strong>.</p>

<p>This is not a condemnation of technology, but a reminder of what is at stake.
Rorty would not call on us to put away the smartphone forever or to ban AI.
He would call on us to ask ourselves: When did we last read a book that truly disturbed us, and whose story was it?
When were we last willing to be truly disturbed by an unfamiliar life?</p>

<p>This returns us to the epigraph with which this essay opened: democracy, taken seriously, demands that the power now concentrating in the development, deployment, and governance of AI systems be dispersed back into the many hands it came from. 
The paternalistic warnings issued by the architects of these systems, i.e., that a super-human intelligence might destroy humanity, deserve our suspicion, not because catastrophe is unthinkable, but because such warnings conveniently justify the very consolidation of power they claim to guard against. 
A more sober formulation might read: a concentration of power will inevitably increase the cruelty exercised by both men and machine.
History offers a familiar precedent.
The anxieties now voiced about artificial intelligence echo, almost word for word, the anxieties once voiced about writing, and later about the printing press: that if everyone were permitted to write, everyone would eventually write what those already in power did not wish to see written; that the chaos loosed by universal literacy would corrode the structures on which order rested.
And they were right. 
It did corrode them. 
Yet it was precisely this corrosion that redistributed the power Rorty’s liberalism seems, on the whole, to welcome: the slow multiplication of vocabularies, of voices entitled to redescribe themselves and be heard.</p>

<p>Our historical pathways to self-realization are no longer sustainable. It is increasingly obvious that material progress in the West will slow down or perhaps even come to a halt. Although the transition toward renewable energy has only just begun, and many still deny the necessity of it, the pervasive feeling that the era of endless growth is over continues to spread. While self-realization was previously achieved through a surplus of wealth, a new, likely subconscious ideology is emerging: that one can only achieve self-realization by denying others the means to do the same.
The idea of the state then changes from an entity that is protecting its citizen from cruelty to a distributor and enabler of cruelty such that everyone, especially the self-proclaimed “silenced majority”, “get what they deserve” <a class="citation" href="#amlinger:2025">(Carolin Amlinger, 2025)</a>.
This zero-sum battle over distribution will inevitably intensify unless we develop alternative frameworks for personal fulfillment.</p>

<p>When I look at my little nieces, I hope that—despite the impending catastrophes as a combination of a war against our political, scientific as well as environmental ecosystem—they will live in a culture that has managed to sensitize itself to cruelties I cannot yet see.
I hope they live in a culture that has once again made it its goal to be intellectual in the Rortyan sense: a culture that revitalizes the old ideal of “poets and thinkers”—not by hunting for a final, objective truth, but by using the poet’s imagination to reshape our world into something more humane.
I hope that, if they fail, they <em>fail in dignity</em> and <em>self-respect</em> <a class="citation" href="#metzinger:2023">(Metzinger, 2023)</a>.
And I hope that they will be merciful toward me, understanding that it was still impossible for me to see so much more, because my language was still bound by the limits of my time.
I hope the circle of their imagination, in which rationality plays its game, will be larger and not smaller; that they are courageous enough to constantly redescribe themselves and their community, until the “we” of their solidarity reaches far beyond what I am capable of imagining today.</p>

<h2 id="references">References</h2>

<ol class="bibliography"><li><span id="kohlenberger:2024">Kohlenberger, J. (2024). <i>Gegen die neue Härte</i> (p. 256). dtv Verlagsgesellschaft GmbH &amp; Co. KG.</span></li>
<li><span id="amlinger:2025">Carolin Amlinger, O. N. (2025). <i>Zerstörungslust: Elemente des demokratischen Faschismus</i>. Suhrkamp.</span></li>
<li><span id="redecker:2026">von Redecker, E. (2026). <i>Dieser Drang nach Härte</i> (p. 272). S. FISCHER.</span></li>
<li><span id="baudrillard:1976">Baudrillard, J. (1976). <i>Symbolic Exchange and Death</i>. SAGE Publications Ltd. https://doi.org/10.4135/9781526401496</span></li>
<li><span id="herzog:2007">Herzog, W. (2007). <i>Begegnungen am Ende der Welt</i>. https://www.youtube.com/watch?v=uBk9lLFWGcI</span></li>
<li><span id="gundlach:2026">Gundlach, M. (2026). Dieser Pinguin glaubt noch an Europa. <i>Süddeutsche Zeitung</i>. https://www.sueddeutsche.de/medien/pinguin-europa-werner-herzog-dokumentation-johann-wadephul-li.3378577</span></li>
<li><span id="gebru:2024">Gebru, T., &amp; Torres, E. P. (2024). The TESCREAL bundle: Eugenics and the promise of utopia through artificial general intelligence. <i>First Monday</i>, <i>29</i>(4). https://doi.org/10.5210/fm.v29i4.13636</span></li>
<li><span id="muehlhoff:2025">Mühlhoff, R. (2025). <i>Künstliche Intelligenz und der neue Faschismus</i>. Recalm.</span></li>
<li><span id="rosengruen:2022">Rosengrün, S. (2022). Why AI is a threat to the rule of law. <i>Digital Society</i>, <i>1</i>(2). https://doi.org/10.1007/s44206-022-00011-5</span></li>
<li><span id="rorty:1990">Rorty, R. (1990). <i>Ethics of Principle vs Sensitivity</i>. https://www.youtube.com/watch?v=nD248K11zNE</span></li>
<li><span id="stiegler:2019">Stiegler, B. (2019). <i>The Age of Disruption</i>. Polity Press.</span></li>
<li><span id="rorty:1989">Rorty, R. (1989). <i>Contingency, Irony, and Solidarity</i>. Cambridge University Press.</span></li>
<li><span id="rorty:1979">Rorty, R. (1979). <i>Philosophy and the Mirror of Nature</i>. Princeton University Press.</span></li>
<li><span id="rorty:2016">Rorty, R. (2016). <i>Philosophy as Poetry</i>. University of Virginia Press.</span></li>
<li><span id="maturana:1991">Maturana, H. (1991). <i>Humberto Maturana Melbourne Seminar</i>.</span></li>
<li><span id="maturana:1988">Maturana, H. R. (1988). Reality: The Search for Objectivity or the Quest for a Compelling Argument. <i>The Irish Journal of Psychology</i>, <i>9</i>(1), 25–82. https://doi.org/10.1080/03033910.1988.10557705</span></li>
<li><span id="luhmann:1984">Luhmann, N. (1984). <i>Soziale Systeme: Grundriß einer allgemeinen Theorie</i>. Suhrkamp.</span></li>
<li><span id="heidegger:1976">Heidegger, M. (1976). “Nur noch ein Gott kann uns retten.” <i>Der Spiegel</i>, <i>30</i>(23), 193–219.</span></li>
<li><span id="rorty:1999">Rorty, R. (1999). <i>Philosophy and Social Hope</i>. Penguin Books.</span></li>
<li><span id="nietzsche:1873">Nietzsche, F. (1873). Über Wahrheit und Lüge im außermoralischen Sinne. In G. Colli &amp; M. Montinari (Eds.), <i>Kritische Studienausgabe (KSA)</i> (Vol. 1, pp. 873–890). de Gruyter.</span></li>
<li><span id="han:2021">Han, B.-C. (2021). <i>Infokratie: Digitalisierung und die Krise der Demokratie</i>. Matthes &amp; Seitz.</span></li>
<li><span id="heidegger:1954">Heidegger, M. (1954). <i>Die Frage nach der Technik</i>.</span></li>
<li><span id="rorty:1998">Rorty, R. (1998). <i>Achieving Our Country: Leftist Thought in Twentieth-Century America</i>. Harvard University Press.</span></li>
<li><span id="fromm:1973">Fromm, E. (1973). <i>The Anatomy of Human Destructiveness</i>. Holt, Rinehart and Winston.</span></li>
<li><span id="esposito:2022">Esposito, E. (2022). <i>Artificial Communication</i>. The MIT Press. https://doi.org/10.7551/mitpress/14189.001.0001</span></li>
<li><span id="metzinger:2023">Metzinger, T. (2023). <i>Bewusstseinskultur: Spiritualität, intellektuelle Redlichkeit und die planetare Krise</i> (p. 208). Berlin Verlag.</span></li></ol>

<!--
Interessanterweise sehr nah an: https://www.youtube.com/watch?v=yM_or9mYaXM
-->
<div class="footnotes" role="doc-endnotes">
  <ol>
    <li id="fn:1" role="doc-endnote">
      <p>I see similarities to the late Heinz von Foersters who tried, based on his constructivism, to motivate a switch from <strong>truth</strong> into <strong>trust</strong>. As soon as one speaks of truth one tends to dominate other perspektives and consequently tend to rule over others. Where there is truth there is also a liar. <a href="#fnref:1" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:2" role="doc-endnote">
      <p>Today, there is substantial evidence that prelinguistic infants display sophisticated reasoning, numerical intuition, and causal understanding well before acquiring language. Great apes, corvids, dolphins, elephants demonstrate planning, tool use, and social cognition without anything resembling human language. The key move is to distinguish between basic cognition (perception, spatial navigation, pattern recognition, which animals and infants clearly have) and discursive, propositional, reason-giving thought, i.e., the kind philosophers actually argue about. Rorty could concede the former while maintaining that the latter is constitutively linguistic and social. Not: “You can’t think without language” but “The kind of thinking that involves justifying beliefs, weighing reasons, making normative claims, that is through and through a linguistic practice.” <a href="#fnref:2" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:3" role="doc-endnote">
      <p>This text can be seen as performative. <a href="#fnref:3" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
    <li id="fn:4" role="doc-endnote">
      <p>Arguing matters if these terms are legally binding—but not for the reason most lawyers or human rights activists think it does since international law is a collection of linguistic mechanisms invented by humans to achieve a specific goal: the reduction of cruelty. Thus, in Rorty’s eyes, arguing about the law is useful only if it is treated as a tactical weapon to reduce physical pain, not as a philosophical debate to prove your moral superiority. But if we spend months debating the precise semantic definition of “genocide” or “proportionality” in academic journals while people are actively starving, the law has ceased to be a useful tool. <a href="#fnref:4" class="reversefootnote" role="doc-backlink">&#8617;</a></p>
    </li>
  </ol>
</div>]]></content><author><name>Benedikt Zönnchen</name><email>benedikt.zoennchen@web.de</email></author><category term="Opinion" /><category term="Philosophy" /><summary type="html"><![CDATA[Part of the idea of a democratic society is that social change comes by reform rather than revolution and this in effect means that the people who have the power actually letting go of some of it. – Richard Rorty]]></summary></entry><entry><title type="html">Hurricanes May Be Dangerous but Not Evil</title><link href="https://bzoennchen.github.io/Pages/2025/09/22/rant-oversimplification.html" rel="alternate" type="text/html" title="Hurricanes May Be Dangerous but Not Evil" /><published>2025-09-22T00:00:00+02:00</published><updated>2025-09-22T00:00:00+02:00</updated><id>https://bzoennchen.github.io/Pages/2025/09/22/rant-oversimplification</id><content type="html" xml:base="https://bzoennchen.github.io/Pages/2025/09/22/rant-oversimplification.html"><![CDATA[<p>The word <em>digital</em> usually refers to something represented, processed, or communicated using discrete values rather than continuous signals.
We know that digital systems encode information in binary (0/1) rather than analog waveforms. A digital clock displays exact numbers; an analog clock sweeps smoothly. Digital data is stored and transmitted as discrete units (bits and bytes), not continuous signals.</p>

<p>But we often forget: <em>signs</em> and <em>symbols</em>—the words of our language—are also digital. They, too, compress complexity into discrete units. Thus human language is a digital system of signs, even when carried by analog sound waves.</p>

<p>While this “feature” of language, i.e. the compression of the environment, is a necessary condition for communication, it implies that communication is loss-prone. Furthermore, there are no ultimate foundations or essences in language, no definite <em>ground</em> beneath our signs, even though we act as if there were.</p>

<p>Programming seems to offer firmer ground. Classes, objects, and definitions feel well-formed, precise, and self-contained. There is a clear separation between syntax and semantics. Yet even here, meaning is not given by the machine, because a machine destroys the reference horizon. In other words, machines do not recognize complexity, because for them, there are no more possible environmental references than those currently being actualized. A machine can “understand” something only in one way, and thus cannot understand it at all. Meaning arises because I, as a programmer (psychic system), interpret and use the system. The code works or fails, delights or frustrates, earns money, sells goods, plays music, organizes life. Meaning is not in the code itself but in the observed, expected and assumed effects it produces and the uses we make of it.</p>

<p>But our interpretations are not unique. We have to select from the reference horizon that we build and imagine together. This meaning is mediated by natural language—so we are already outside the pure formal system and also limited by our cultural memory.</p>

<p>The reference horizon is contingent but not random. When we listen to music, we anticipate the next tone based on what we have already heard and what we are used the hear. Similarly, the past gives us the structure to see the future, but it also conceals other ways to think, imagine, and live. We need the past, culture, traditions as memory function to be able to anticipate and realize the next step—our future.</p>

<p>Now step further outward into everyday life: the ambiguity multiplies. What does it mean to call someone a “mother”? Is there an essence of “motherhood”? I would argue there is not. Instead, there is a shifting pattern of experiences: shelter, protection, food, conflict, laughter, school lessons, beginnings, and countless other associations. A mother is not a father, not a flower, not a lake—but a fluid constellation of relations and meanings. It is also a <em>trace</em>. But this does not imply arbitrariness.</p>

<p>Similarly, we can also ask the nowadays politically charged question: <em>What is a woman?</em>—and immediately run into difficulties. Yet, if we think about it, this is hardly surprising.
How could language—a digital system of signs—ever objectively capture the infinite complexity of our realities that constantly drift in the ocean of evolution? The word <em>“woman”</em> only gains meaning in relation to the usage of terms like <em>man</em>, <em>not-man</em>, <em>mother</em>, or <em>feminine</em>. Its meaning is never fixed; it is always shifting, unstable, and deferred. To ask “What is a woman?” already risks reinforcing the <em>man/woman binary</em> as <em>natural</em>, when in fact that binary is itself the product of language, institutions, and history. It can and will change.
Any attempt to pin down a definition will necessarily exclude certain possibilities and identities.</p>

<p>In this spirit, communication functions not through mutual understanding, but rather through the <strong>absence of misunderstanding</strong>. It persists as long as connectivity remains—that is, as long as one contribution can trigger another to continue the discourse. Because meaning is an internal construct of our psychic systems, it cannot be transmitted; it is this deep operational separation that allowed society to emerge. Despite our inability to inhabit one another’s thoughts, we have learned to organize and cooperate. This operational closure—and the inherent “loneliness” it implies—is the source of many of our troubles, yet it is also the foundation of our autonomy and, likely, the very reason we possess a sense of self.</p>

<p>Outrage arises because different groups want their preferred vocabulary to dominate—for example, biological essentialists versus trans-inclusive definitions. In truth, the conflict is a clash of competing <em>language games</em>. Each side feels its way of speaking is under threat, which explains the emotional intensity.</p>

<p>As Richard Rorty once observed:</p>

<blockquote>
  <p>There is something potentially very cruel about the claim that [the language people speak is, for ironists, a game of chance]. For the most effective way to inflict lasting pain on people is to humiliate them by making everything that had seemed especially important to them appear futile, outdated, and powerless. – <a class="citation" href="#rorty:1989">(Rorty, 1989)</a></p>
</blockquote>

<p>This is why the demand for a single, definitive answer generates tension: it denies the play of difference that actually structures meaning. The difficulty of the question is not incidental—it is intrinsic. <em>Woman</em> is an aporia, a site of endless contestation. Outrage arises precisely because the question both demands and resists resolution.</p>

<p>The point, then, is not to settle the metaphysical question but to enjoy ambiguity and to foster <em>solidarity</em>. That means choosing the description of woman that best promotes human flourishing and reduces harm and is open to interpretation. Rather than offering a final definition, our task is to help society continually renegotiate its self-descriptions.</p>

<p>We can also ask more generally: <em>What is a human being?</em> Here too, the problems multiply if we think in terms of essences.</p>

<p>A person is not a fixed substance but a pattern of modulated repetition. Life seems coherent over time because of habits, recurring desires, aspirations, and flaws. But a person is not a stable essence—it is a process, an event, a “happening”. Like a hurricane, a person is a complex system: dynamic, shifting, and contingent. And this observation, of course, is made by yet another hurricane, observed by still others—each shaping and shaped by the rest.</p>

<p>This is precisely why I find contemporary reporting so taxing. It systematically ignores this inherent complexity. While communication admittedly requires compression, must it always collapse into reductive binaries—right vs. wrong, friend vs. foe? Must every event and every individual be flattened into a single label just to fit a pre-existing semiotic network? Must every person have an opinion on any matter?</p>

<blockquote>
  <p>Everything must be explained, everything should be understood, and if something cannot be understood, it counts as nothing. [… So believe] many people, who are constantly having everything explained to them and are presented with a world without secrets, without the inexplicable or the overly complex, eventually come to believe themselves that they understand everything. – <a class="citation" href="#bauer:2018">(Bauer, 2018)</a></p>
</blockquote>

<p>Too often, complexity is reduced to a “profile”—a checklist of traits or, worse, a binary moral judgment. But a hurricane is neither good nor evil; it is a phenomenon. It destroys, it endangers, and it compels us to react. If we wish to address such forces, we cannot moralize them. We must instead understand—and perhaps alter—the structural conditions under which they form and transform.</p>

<p>People, too, resist moral simplification. A person can be both a criminal and deeply kind; they can support terror while enduring horrific cruelty; they can be a brilliant poet and a member of the Nazi party. I choose to acknowledge these realities with an “<strong>and</strong>” rather than a “<strong>but</strong>”.</p>

<p>There is no escape from this complexity—only further layers of it. Yet, it feels increasingly difficult to transcend our digital condition: a world structured by discrete representations, binary code, and rigid networks. In such an environment, the richness of lived experience is flattened into exchangeable symbols—tokens that are used and abused to force a sense of order, to draw hasty conclusions, and to conjure yet another storm.</p>

<!--
I know it is not possible to report without compression. I am compressing right now. But can't we do better? Must it always be only right and wrong, good and evil, friend and foe? Must every event and every person be collapsed into a single dot that fits neatly into our networks of signs?
 -->

<h2 id="references">References</h2>

<ol class="bibliography"><li><span id="rorty:1989">Rorty, R. (1989). <i>Contingency, Irony, and Solidarity</i>. Cambridge University Press.</span></li>
<li><span id="bauer:2018">Bauer, T. (2018). <i>Die Vereindeutigung der Welt</i> (p. 104). Reclam.</span></li></ol>]]></content><author><name>Benedikt Zönnchen</name><email>benedikt.zoennchen@web.de</email></author><category term="Opinion" /><summary type="html"><![CDATA[The word digital usually refers to something represented, processed, or communicated using discrete values rather than continuous signals. We know that digital systems encode information in binary (0/1) rather than analog waveforms. A digital clock displays exact numbers; an analog clock sweeps smoothly. Digital data is stored and transmitted as discrete units (bits and bytes), not continuous signals.]]></summary></entry><entry><title type="html">The Free Energy Principle</title><link href="https://bzoennchen.github.io/Pages/2025/08/04/free-energy-principle.html" rel="alternate" type="text/html" title="The Free Energy Principle" /><published>2025-08-04T00:00:00+02:00</published><updated>2025-08-04T00:00:00+02:00</updated><id>https://bzoennchen.github.io/Pages/2025/08/04/free-energy-principle</id><content type="html" xml:base="https://bzoennchen.github.io/Pages/2025/08/04/free-energy-principle.html"><![CDATA[<p>Imagine you see a grizzly bear in the woods and half of its face is hidden behind a tree.
You will most certainly recognize the bear as whole and dangerous animal.
Your brain will “fill in the gaps” even though there is no absolute certainty that the bear isn’t split in half.
How and why does the brain do this?
Furthermore, why are we fooled by all sorts of graphical illusions?
In other words, why are we, in some instances, so stubborn to see what is not there—even if we are told that it is not there—and, on other occasions, we can immediately see what is hidden and probably there?</p>

<p>The so‑called <em>free energy principle</em> (FEP) <a class="citation" href="#friston:2006">(Friston et al., 2006)</a> is a neat mathematical principle that offers answers to these questions.
According to the FEP, brains or nervous systems make use of <em>variational inference</em> on hidden causes of sensory data.
In this view, the goal of an organism, including ours, is to <strong>minimize surprise</strong> within their respective environment.
Strictly, FEP says systems minimize an <em>upper bound</em> on surprise—variational free energy—because exact surprise is intractable
Researchers working towards finding evidence to apply the FEP claim that perceptual processes are just one aspect of emergent behaviours of systems that conform to a <em>free energy principle</em>.
Thus, it is a principle on which the brain might operate.</p>

<blockquote>
  <p>The free energy considered here measures the difference between the probability distribution of environmental quantities that act on the system and an arbitrary distribution encoded by its configuration.
The system can minimise free energy by changing its configuration to affect the way it samples the environment or change the distribution it encodes. These changes correspond to action and perception respectively and lead to an adaptive exchange with the environment that is characteristic of biological systems.
This treatment assumes that the system’s state and structure encode an implicit and probabilistic model of the environment. – <a class="citation" href="#friston:2006">(Friston et al., 2006)</a></p>
</blockquote>

<p>As a consequence, the FEP departures from the classical reward-based explanations of behaviour and evolutional progress in general.
Instead of starting with an <em>a priori</em> <em>goal</em> or <em>reward</em> that the organism has to find, it starts with <em>what kinds of states the organism expects to be in</em>, and assumes the organism to act in a way that <em>confirms those expectations</em> and keeps it in familiar, low-surprise situations.</p>

<p>Friston and others also believe that the use of <em>hierarchical models</em>, such as layered neural networks, enable the brain to construct prior expectations in a dynamic and context-sensitive fashion.
Following this line, this scheme provides a principled way to understand many aspects of cortical organisation and responses while it is strongly connected to the <em>connectionist</em> strand of <em>artificial intelligence</em>.</p>

<p>It was Friston who populated the idea of the brain as a <em>prediction machine</em> trying to minimize surprise (or prediction error) by either updating beliefs (<strong>perception</strong>) or by acting on the world (<strong>action</strong>) to make it more predictable.
He provides a unifying framework for perception, action, attention, learning, and even consciousness which, in my opinion, combines enactivists and constructivists views (as we will discuss later).</p>

<p>However, the term <em>prediction machine</em> should not be confused with a depreciation of organisms or the human brain.
There is nothing magical or mystical that we will lose if we stick to the naturalization of ourselves as long as we do not mistaken an explanation as final answer.
Existential questions remain but stay outside the public sphere of science but those questions are necessarily informed by our (scientific) knowledge about ourselves and our environment.
In accordance to the FEP, we drift towards explanations of our being—we drift towards self-creation and our own evidence.
Therefore, what could be more magical, more fascinating than our contingent existence that emerged through a process of exactly such <em>self-creation</em> and <em>self-evidencing</em>.</p>

<p>While the FEP gained major attention in neuroscience and cognitive science through the work of Karl Friston in the 2000s, it is mathematically rooted in ideas that are much older—originating in statistical physics, Bayesian inference, and information theory.</p>

<h2 id="metaphysical-assumptions">Metaphysical Assumptions</h2>

<p>To start, we need some metaphysical assumptions, which I summarize as <em>physicalism</em> (ontology) and <em>indirect realism</em> (epistemology).</p>

<p>Physicalists assume that the world is made entirely of physical stuff (matter, energy, fields, etc.) and that all phenomena, including mental events (thoughts, feelings, consciousness), can, in principle, be explained by physical processes and physical laws.
There is no need to appeal to non‑physical substances (like souls or spirits) to explain reality.</p>

<p>While in practice, scientists rarely state their metaphysical position explicitly, most of them are physicalists or at least act in their field as if they are physicalists.
Still, physicalism remains a philosophical stance, not a requirement for doing science.</p>

<p>Second, indirect realism (also called <em>representationalism</em>) is a theory in the philosophy of perception about how we experience the world.
It states that we do not perceive the external world directly.
Instead, we are directly aware only of mental representations (sometimes called sense data or percepts), and these representations are caused by and (usually) resemble external physical objects or events.
There is a physical world “out there,” and physical objects and processes emit or reflect light, sound, and other signals.
Our sense organs (eyes, ears, etc.) pick up these signals and send information to the brain.
The brain processes this information and creates an internal representation or image of the world.
Therefore, what we are immediately aware of is this internal image, not the external object itself.
From these internal representations, we indirectly know about the external world.</p>

<p>From a physicalist perspective, the “sense data” or perceptions that indirect realism talks about would be physical phenomena generated by the brain’s interaction with the environment. 
The objects in the external world exist independently of our perception (as physicalism holds), but we never experience them directly—only through mediated sensory data.
But there is also a tension: If one believes that all experiences and perceptions are entirely physical processes, there might be questions about how non-physical qualities (like subjective experience or <em>qualia</em>) fit into the picture.
Physicalism typically aims to reduce all phenomena, including consciousness, to physical states.
So, there could be questions about whether “sense data” are truly representations in a way that indirect realism requires or whether they are just parts of the physical process.
Moreover, if a physicalist were to take an extreme reductionist view (such as <em>eliminative materialism</em>), they might reject any distinction between “sense data” and the brain’s processing of physical stimuli, challenging the need for a separate representational layer of perception.
However, this is a minority view in physicalism.</p>

<p>Later I will problematize indirect realism and dwell on the question of whether the free energy principle really requires it.
The list of critics is long, ranging from <em>phenomenologists</em> (Husserl, Heidegger, Merleau‑Ponty), <em>naive realists</em> (McDowell, Brewer, Searle, Reid), <em>new realists</em> (Gabriel, Holt), <em>idealists</em> (Kant), <em>empiricists</em> (Hume), and other <em>biological constructivists</em> (Maturana, Varela).
To move on, I will first assume indirect realism for a moment because it makes the explanation intuitive (but if we think about it a little longer, we run into problems).</p>

<p>From an ontological standpoint Friston clearly assumes an external environment that pushes organisms into some direction when he writes:</p>

<blockquote>
  <p>[I]nvoking selectionist arguments; those systems that match their internal structure to the external causal structure of the environment in which they are immersed will be able to minimise their free energy more effectively. – <a class="citation" href="#friston:2006">(Friston et al., 2006)</a></p>
</blockquote>

<p>One might ask (later) if an environment of a system can be regarded in separation of that system.</p>

<h2 id="contingent-complexity">Contingent Complexity</h2>

<p>As we will see, the <em>free energy principle</em> suggests that our nervous system is actively generating predictions about what should be “out there,” and then uses sensory input merely to check whether those predictions are correct.
In that sense, assuming the free energy principle makes our brains into prediction machines with the goal of <strong>minimizing surprise</strong>.</p>

<p>Our body’s purpose is to stay alive—to continue with its self‑creation by constructing its own organization and structure (autopoiesis).
It is open to structure but organizationally or operationally closed (implying a causal loop in its operations).
Its <em>qualitative identity</em> can change but its <em>individual identity</em> is maintained constant even if it could change, meaning there is some active prevention going on—we do not disintegrate:</p>

<blockquote>
  <p>[A molecular autopoietic system (or living being)] is a homeostatic system of chemical production which has its own individual identity as the variable which it maintains constant – <a class="citation" href="#barry:2012">(Razeto-Barry, 2012)</a></p>
</blockquote>

<p>Some argue that autopoiesis is an intrinsic purpose of an organism from which all other goals derive <a class="citation" href="#virenque:2024">(Virenque, 2024; Barandiaran, 2017)</a>.</p>

<blockquote>
  <p>Importantly, sense-making is intentional because it implies an organismic perspective from which meaning is brought forth. Because of the organism’s precariousness, certain interactions with the environment are either positive or negative from the perspective of the organism. Consider a bacterium swimming up a gradient of glucose. The glucose is meaningful for the bacterium because it provides nutrients that are necessary for maintaining its metabolic processes (i.e., its autopoiesis). The meaning of the glucose gradient only makes sense from the perspective of the bacterium. Take away the bacterium from the situation and the glucose stops having its meaning. What brings forth such meaning is the organism’s adaptive and autopoietic organisation. Such bringing forth of meaning is what enactivists call sense-making. – <a class="citation" href="#bogota:2024">(Bogotá, 2024)</a> (example taken from <a class="citation" href="#varela:1997">(Varela, 1997)</a>)</p>
</blockquote>

<p>Others suspect that we, as external observers, project purposiveness onto the organism <a class="citation" href="#cummins:2014">(Cummins, 2014)</a> and anti-naturalists like Markus Gabriel see meaning as ontological primary <a class="citation" href="#gabriel:2018">(Gabriel, 2018)</a>.
For the FEP, it is not important if purposiveness is a projection or intrinsically given (if meaning is ontologically primary the whole story changes).</p>

<p>What is quite certainly true is that organisms have to deal with perturbations to stay alive in their respective environment (which is determined by the <em>couplings</em> between the organism and its environment).
Only surviving organisms can be observerd, which leads to the belief—backed by observation from an external observer—that organisms strive for survival (of the fittest) when, in fact, it might only look like this is the case.</p>

<p>The complexity of the environment and the complexity of the organism co‑evolve over time, and some organisms—like humans—evolved in a direction of higher complexity.
Again, this does not mean that an increase in complexity is a pre-subscribed <em>telos</em>.
It is all <em>contingent</em>.</p>

<p>As a consequence of an increase in complexity (for some organisms), their environments become noisy and ambiguous, which requires more sophisticated feedback mechanisms to be able to act and survive in such environments <a class="citation" href="#tomasello:2024">(Tomasello, 2024)</a>.
For example, we experience a sort of conscious self-control when we asked to tell the color (not the written text) of</p>

<p><span style="color:red">BLUE!</span></p>

<p>We wanna say “blue” but immediately interrupt ourselves to say “red”.
We can be angry with ourselves, talking to ourselves and judging ourselves as if there is an instance within us that is distinct from us.</p>

<p>Recognizing a predator in the forest already requires more than just pattern matching.
Simple template‑matching reflexes may not work.
In such complex environments, organisms have to temper their reflexes, plan, and evaluate different action sequences before actualizing one action plan.
But a much more complex situation arrives if we have to deal with social systems that are beyond our control but which also provide us with the needs for survival.</p>

<p>Tomasello points out that these abilities result from nature’s inability to build organisms that can cope with everything (reactively) that nature throws at them.
Instead, nature can build psychological actors:</p>

<blockquote>
  <p>It brings forth organisms that function as feedback control systems pursuing goals, selecting profound actions and monitoring the process of acting them out. – <a class="citation" href="#tomasello:2024">(Tomasello, 2024)</a></p>
</blockquote>

<p>He also reminds us that the importance of acting is not defined of how <em>many</em> things an organism can do but <em>how</em> these actions are realized.
Maybe we can learn from his emphasis when discussing AI systems where the focus is often on the capabilities, i.e. on the “how many things” AI systems can do (better) neglecting the question of <em>(self-)control</em> and <em>agency</em>.
Again, for Tomasello the importance is not complexity or variety (of actions) but the levels of control an organism can exert on its environment and itself.</p>

<p>With respect to the free energy principle contingency and unpredicability explodes if the environment of organisms consists of other organisms—when one has to predict the predictions of others.
Social priors provided by caregivers or joint attention and communication become necessary and people might update their internal models by minimizing surprise via social feedback.
We have acted on the whole earth to make our environment more predictable which led to new unpredictablities and risks that we try to make predictable.</p>

<h2 id="naturalizing-surprise-and-entropy">Naturalizing Surprise and Entropy</h2>

<p>Since the free energy principle (FEP) is all about minimizing surprise, let’s start by pinning down what mathematicians and computer scientists actually mean by “surprise”—how can we naturalize it?
Sure, we all have a gut feeling about it—something surprising is just something we didn’t see coming.
More formally, an event is surprising if it’s unlikely based on what we already believe.
That is, we were pretty confident it wouldn’t happen, and yet—bam—it did.</p>

<p>Take raining frogs, for instance.
That would definitely raise some eyebrows.
But even without biblical weather, everyday life has its surprises.
Imagine rolling a die 10 times and getting a 1 every single time. 
Highly unlikely, right? 
That’s the kind of statistical oddity that makes us do a double-take.</p>

<p>Surprise, in this framework, can come from two main sources:</p>

<ol>
  <li>The event itself—specifically, how unpredictable (or high in entropy) the event is.</li>
  <li>Our prior beliefs—how confident we were about what <em>should</em> happen.</li>
</ol>

<p>Now, even if your beliefs are spot on—say, you assume the die is fair, and it actually is fair—you might still see 10 ones in a row.
That’s not because your beliefs are wrong, it’s just because randomness likes to keep things interesting.
In this case, the surprise comes not from flawed beliefs, but from sheer bad luck.</p>

<p><strong>Remark:</strong> In the <em>Bayesian</em> world, probability isn’t about how often something actually happens—it’s about what we believe will happen, given what we know—our <strong>degree of belief</strong>.
More precisely, it’s a measure of uncertainty or confidence in a particular outcome or parameter, based on the information we currently have.
This is quite different from the <em>frequentist</em> view, where probability is all about long-run frequencies.
A frequentist might say, “If we rolled this die an infinite number of times, the proportion of ones would settle at 1/6”.
That’s the idea: probabilities reflect what would happen over countless repetitions of the same experiment.
But often events cannot be repeated.
Here a Bayesian offers a more flexible, belief-based approach: “Before seeing any data, I assume all outcomes are equally likely (a uniform prior). But if I observe 200 ones in 1000 rolls, I’m updating my belief—maybe this die has a bias”.
The more data we get, the more refined our beliefs become.
This belief-updating process is powered by Bayes’ theorem, which lets us revise our <em>prior</em> beliefs \(p(z)\) using new data (via the <em>likelihood</em> \(p(\text{data} \vert z)\)) to form a <em>posterior</em> belief \(p(z \vert \text{data})\).</p>

\[p(z \vert \text{data}) = \frac{p(\text{data} \vert z)p(z)}{p(\text{data})} = \text{Posterior} = \frac{\text{Likelihood} \times \text{Prior}}{\text{Evidence}}.\]

<p>Bayesian inference is like scientific reasoning on autopilot: start with a hunch, collect some data, update your expectations.
It can be thought of as a natural extension of probabilistic logic—one that allows us to reason about hypotheses, not just outcomes.
In this framework, we can assign a probability to a hypothesis, even if we don’t know yet whether it’s true or false.
This is a big shift from the frequentist perspective, where hypotheses are usually treated more like yes-or-no questions to be tested, not graded on a sliding scale of belief.</p>

<p>Alright, back to surprise and entropy!</p>

<p>Let’s revisit our die.
Suppose the die isn’t fair—it actually rolls a 1 with probability 0.6, and the other numbers share the remaining 0.4 equally.
Now, if we think the die is fair, but it keeps landing on 1 suspiciously often, the surprise we feel doesn’t come just from the event itself.
It also comes from the mismatch between our internal model (a fair die) and reality (a biased one).</p>

<p>In other words, surprise isn’t just about what happened—it’s also about what we expected to happen.
When our expectations are off, our surprise spikes.
This is why having a good model of the world—one that reflects actual probabilities—is so important.
The better our model, the better we can anticipate what’s likely and stay one step ahead of surprise.</p>

<p>Let us assume a perfect world model.
If \(P(X=1) = p(1) = p_1 = 0.6\) is the probability of rolling 1 with the die, then the <strong>surprise</strong> \(h\) of rolling it is</p>

\[h(p(1)) = \ln\left( \frac{1}{p(1)} \right) =  - \ln\left( p(1) \right).\]

<p>It is high if the probability is small, in fact,</p>

\[\lim\limits_{p_s \rightarrow 0}\ln\left( \frac{1}{p_s} \right) = \infty\]

<p>and</p>

\[\lim\limits_{p_s \rightarrow 1}\ln\left( \frac{1}{p_s} \right) = 0.\]

<p>While we multiply probabilities to figure out the probability of independent events happening together, surprise works differently—it adds up.
Meaning ten 1s in a row is twice as surprising as five 1s in a row, because</p>

\[\ln{\frac{1}{a \cdot b}} = \ln\frac{1}{a} + \ln\frac{1}{b}.\]

<p>Now, to figure out how much we could be surprised by an event—assuming a perfect and known world model in form of a probability distribution of all possible events, e.g. of the outcome of rolling a die—we could compute the <strong>average surprise</strong> of that distribution which is called its <strong>entropy</strong> \(H\):</p>

\[H(P) = \sum_s p(s) \ln\left( \frac{1}{p(s)} \right) = - \sum_s p(s) \ln p(s).\]

<p>We multiply the surprise of the event \(s\) by its probability because more likely events occur more often.
Therefore, they will contribute more to our overall surprise, if we would repeat rolling the die over and over again.
The corresponding <strong>differential entropy</strong> for a continuous random variable \(X\) with the probability density function \(p(x)\) can be written as:</p>

\[H(p(x)) = -\mathbb{E}_{x \sim P}\left[ \ln p(x) \right] = - \int p(x) \ln\left( \frac{1}{p(x)} \right) dx.\]

<p>Again, we are taking the average of the log probability over the distribution \(p(x)\).
This tells you how surprising the outcomes are on average.
The higher the entropy, the more uncertain or unpredictable the variable.</p>

<p>So far so good.
But what happens if our world model is wrong?
Can we learn and adjust it?
And how can we first separate the surprise caused by our incorrect prior beliefs from the overall surprise to then minimize this second source of surprise?</p>

<p>Let us assume we belief in a fair coin, that is \(Q(X = \text{heads}) = Q(X = \text{tails}) = 0.5\)—a reasonable approximation.
I denote our belief as \(Q\) and its surprise as \(h_q\).
But now let’s assume the coin is rigged: \(P(X = \text{heads}) = 0.99\) and \(P(X = \text{tails}) = 0.01\)
We belief the probability of ten times heads in row is</p>

\[Q(\text{10 heads}) = 0.5^{10} = 0.001\]

<p>and the surprise is</p>

\[h_q(\text{10 heads}) = \ln 0.5^{-10} \approx 7.\]

<p>But in reality</p>

\[P(\text{10 heads}) = 0.99^{10} = 0.9\]

<p>and</p>

\[h(\text{10 heads}) = \ln 0.99^{-10} \approx 0.1.\]

<p>Therefore, we are much more surprised as we should be and this surprise is mostly caused by believing in the wrong world model.
This leads us directly to <strong>cross-entropy</strong> \(H(P,Q)\), which is the average surprise you will get by observing a random variable governed by distributions \(P\), while believing in \(Q\).</p>

\[H(P,Q) = \sum\limits_s P(X=s) \ln\left( \frac{1}{Q(X=s)} \right) = \sum\limits_s p(s) \ln \left( \frac{1}{q(s)} \right).\]

<p>We multiply our subjective surprise by the “real” probability which means that if an event is very likely but our surprise for it is high—meaning we do not expect it—the cross-entropy is high as well.
Furthermore, if our beliefs are prefect, i.e. \(P=Q\) the cross-entropy is equal to the entropy since</p>

\[H(P,P) = H(P)\]

<p>holds.
Also note that the cross-entropy is <strong>asymmetric</strong>, that is, \(H(P,Q)\) is usually not equal to \(H(Q,P)\).
For example, take the case from before.
In this situation we are less surprised believing in a fair coin while it is rigged  than believing in \(P\) while it is fair, that is,</p>

\[H(P,Q) = \ln\left( \frac{1}{0.5} \right) \approx 0.69 \leq 0.5 \cdot \left[ \ln\left( \frac{1}{0.99} \right) + \ln\left( \frac{1}{0.01} \right) \right] \approx 2.31 = H(Q,P)\]

<p>Furthermore, and very importantly, for any model \(Q\), <strong>the cross-entropy can never be lower than the entropy of the underlying generating distribution</strong>:</p>

\[H(P,Q) \geq H(P).\]

<p>Using cross-entropy and entropy, we can now compute the surprise caused by our wrong beliefs:</p>

\[D_{\text{KL}}(P \Vert Q) = H(P,Q) - H(P) = \sum_s p(s) \ln\left( \frac{1}{q(s)} \right) - \sum_s p(s) \ln\left( \frac{1}{p(s)} \right).\]

<p>This can be compressed to</p>

\[D_{\text{KL}}(P \Vert Q) = \sum_s p(s) \ln\left( \frac{p(s)}{q(s)} \right).\]

<p>This is called the <em>Kullback–Leibler divergence</em> \(D_{\text{KL}}(P \Vert Q)\).
It is the divergence of \(P\) from \(Q\) (also known as the <em>relative entropy</em> of \(P\) with respect to \(Q\)).
To improve our beliefs we could minimize the Kullback–Leibler divergence, that is, we could</p>

\[\min D_{\text{KL}}(P \Vert Q).\]

<p>However, in machine learning you will usually hear about only minimizing the cross-entropy \(H(P,Q)\).
This is because we can not change the entropy of “reality”, that is, we can not change \(H(P)\)—we can not change reality (without acting) but our beliefs about it.
Thus, minimizing \(H(P,Q)\) achieves the same goal as minimizing \(D_{\text{KL}}(P \Vert Q)\).</p>

<p>Of course, the big problem is that we don’t know \(P\)!</p>

<h2 id="the-circularity-of-perception">The Circularity of Perception</h2>

<p>Because of a <em>drift</em> towards complexity, many assume that brains began to build models \(Q\) of the world so that they can explain sensory inputs by <strong>inferring</strong> their <strong>hidden causes</strong>.
For example, your brain might have an internal model of snakes and bears—what they are, how they look, and what they can and cannot do.
Brains can generate missing information, e.g., they can deal with obfuscated objects.
They have an intuitive understanding of the physical world—of movement, speed, and heaviness.
They are biased toward what is usually helpful.
In that sense, our body “understands” e.g. gravity “intuitively” before knowing anything about Newton’s or Einstein’s theories.</p>

<p>A brain, in that view, is like a judge.
There is the raw sensory data obtained by <em>observations</em>, which leads to the generation of predictions.
It is the incoming observation, or what you see.
Furthermore, there are <em>prior beliefs</em> learned through experience and evolution.
It is what you usually expect, and if these expectations do not fit your sensory data, you are surprised.
The brain can check how well your explanation matches your observation.</p>

<p>Following this principle, the first component called <em>accuracy</em> is defined as a measure of how well the data fit an explanation.
The second component is <em>complexity</em>.
It refers to how abstruse this explanation is.
We want to prefer simple models of the world, i.e., simple explanations.</p>

<p>From this we get an obvious tension between the two components.
We probably get the highest accuracy by using a very complex model, but such a model cannot generalize and will produce inaccurate predictions for new situations.
This tension is the <strong>free energy</strong> \(F\), that is,</p>

\[F = \text{complexity - accuracy}.\]

<p>We (or our brains) want maximal accuracy with minimal complexity.
Thus, following the definition of \(F\), the brain minimizes free energy \(F\).
To survive, it requires <em>high accuracy</em> <strong>and</strong> <em>low complexity</em>.
If this is achieved, then there is only low free energy \(F\)—there is nothing to be gained.</p>

<p>It is assumed that the brain compresses the high‑dimensional sensory data into a somewhat manageable form by finding commonalities and hidden structures in the data.
Furthermore, there might be a lot of information in the data that is not important for the organism’s self‑creation and survival.
Especially at a deep, low‑dimensional layer, hidden neurons—also called <em>latent variables</em> or <em>latents</em>—represent causes.
Importantly, <strong>for such latent neurons there is no ground truth of what their activity should be or mean because they do not interface with the outside world.</strong>
The brain is free to “choose” whichever latents it wants.
This opens up a question: how can the brain test if its world model is a good one if it has no direct access to the world?</p>

<p>We cannot verify the latents, but we can verify their consequences, meaning the brain should be able to reconstruct the source sensory data \(x\) from its compressed representation, that is, from its latents \(z\).</p>

<p><strong>Remarks:</strong> In the original paper \(z\) is denoted as \(\vartheta\) and called “parameterise environmental forces or fields that act upon the system”.
\(x\) is denoted as \(\hat{y}\) and is a function of actions \(\alpha\) (the effects of the system on the environment).
Furthermore, I will use \(\theta\) as the parameters of the involved neural networks which, in the paper, is denoted as \(\lambda\) and called “quantities that describe the system’s physical state”.
It can also mean the parameters of the Gaussian distribution that are computed by the network.
Friston also differentiate between parameters that can change quickly, slowly and very slowly.</p>

<blockquote>
  <p>Factorization of the ensemble density to cover quantities that change with different timescales provides an ontology of processes that map nicely onto perceptual inference, attention and learning. – <a class="citation" href="#friston:2006">(Friston et al., 2006)</a></p>
</blockquote>

<p>Using a probabilistic modeling framework, we can say that, given our beliefs \(z\) (about how the world works), we should be able to predict the sensory data \(x\) using our internal model that approximates the true probability distribution \(p\), meaning that</p>

\[p(x \vert z)\]

<p>should be high where \(z\) is given, and \(G_{\theta}(z) \approx x\) is computed by a neural network \(G_{\theta}\)—\(z\) is given while the brain <strong>generates</strong> \(x\).</p>

<p>Note that \(z\) is assumed to be of low dimensionality while \(x\) is of high dimensionality.
\(G_{\theta}\) is a <em>generative model</em>.
We call \(p(z)\) <em>priors</em> and \(z\) <em>causes</em>.
If \(z\) is not the cause for \(x\), we have to change our model to learn different latents \(z\).</p>

<p>The brain learns these latents or causes by compressing the sensory data.
Given sensory data \(x\) the nervous system tries to figure out what <em>causes</em> \(z\) produced \(x\); that is, the system wants to model</p>

\[p(z \vert x)\]

<p>accurately.
This is called <em>inference</em> because the system <em>infers</em> causes from <em>observations</em>—\(x\) is given while the brain <strong>searches</strong> for \(z\).</p>

<p>However, this is a hard problem because, unlike generation, there is no function that directly computes \(z\) given \(x\)—the generative model is not invertible.
One could generate a lot of candidates of sensory data \(x'\) using the generative model and compare these candidates to the true observation \(x\).
However, this is not feasible because, even though \(z\) is of lower dimensionality than \(x\), there are still too many possibilities.
We say that the problem is <em>intractable</em>.</p>

<p>In reality, our brain solves this problem almost <strong>instantaneously</strong>.
If there is a grizzly bear in front of us, we have to figure it out fast!
So how does the brain solve this seemingly impossible task?</p>

<p>Instead of finding the exact latents \(z\), our brain tries to find an approximation \(q(z \vert x)\) using a so-called <em>recognition model</em> \(R_{\theta}\), which is distinct from but interdependent on the <em>generative model</em> \(G_{\theta}\).
It works in the opposite direction compared to the generative model by mapping a sensory observation \(x\) to a distribution of causes:</p>

\[R_{\theta}(x) \approx z.\]

<p>The result is only an approximation—a rough guess \(z\) of what causes the observations \(x\).
To improve the guess the brain might do multiple rounds of <em>recognition</em> and <em>generation</em> in tandem to arrive at an optimal approximation.
This process we can call <strong>perception</strong> and it can only work if <em>recognition</em> and <em>generation</em> are aligned with each other.
It presupposes such an alignment, thus learning through experience!</p>

<p>In perception, free energy is minimized, meaning that an optimization happens until <em>recognition</em> roughly fits <em>generation</em>:</p>

\[x \approx G_{\theta}(R_{\theta}(x)).\]

<p>If this is the case, our brain has found an explanation or causes \(z\) that minimize <em>free energy</em>—one that explains the sensory input \(x\) and aligns well with one’s prior beliefs.
Perception basically solves for</p>

\[\text{minimize } \left[ \text{complexity - accuracy} \right] = \min F.\]

<p>It rapidly adjusts the activity of latent neurons.</p>

<p>Long‑term learning via experience is required to gradually refine both models—i.e. neural networks—to align them better with each other and to construct better world models.
Both <em>learning</em> and <em>perception</em> serve the same goal: reducing uncertainty in the environment the organism at least partly constructs by building optimal models of it and finding useful explanations for sensory data within those models.</p>

<p>Now there is a third way to minimize free energy, that is, <em>acting</em> which equates to a change in the organisms environment caused by the organism.
Acting can be as simple as turning your head.
It changes the sensory data the organism is <em>receiving</em>.
Therefore, one can model the sensory data \(x\) as a function of action \(\alpha\), that is, \(x(\alpha)\).
Consequently, one can interpret perception as an active process, therefore, variational inference becomes <em>active inference</em>.
The organism does not receive but selects its sensory data and it can anticipate to minimize “future” free energy actively.
A feedback loop is constructed thus non-linearity is introduced:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>acting -&gt; perceiving/learning -&gt; believing -&gt; acting.
</code></pre></div></div>

<p>Each step influences the next one.</p>

<h2 id="variational-inference">Variational Inference</h2>

<p>Up to this point I did not explain how generation and perception can be established, that is, how to find good models \(G_{\theta}\) and \(R_{\theta}\).
In the following section, we look at how such models can be build in the field of <em>machine learning</em> using the framework of <em>variational inference</em>.</p>

<p>First, we assume an unknown latent distribution \(p(z)\) (causes) and a distribution of our observations \(p(x)\) (for those causes).
From the rules of probability theory we know that</p>

\[p(x,z) = p(z) \cdot p(x \vert z)\]

<p>holds.
That is, to compute the probability that \(x\) and \(z\) occur together, we can first compute the probability of \(z\) and then multiply the conditional probability of \(x\) given \(z\).</p>

<h3 id="minimizing-divergence">Minimizing Divergence</h3>

<p>Our <em>generative neural network</em> \(G_{\theta}\) computes \(p(x \vert z)\) given \(z\).
But on a computer we usually do not compute the probability distribution directly.
Because there is only a finite amount of memory, we cannot represent such a distribution exactly.
What one usually does is compute the parameters—means and standard deviations—of a Gaussian distribution.
One might ask: Why is this feasible? Why the Gaussian?
I will explain this later.
For now, assume that it is a good, computable choice.</p>

<p>An artificial neural network does not output probabilities directly but instead the parameters, i.e. <em>means</em> \(\mu\) and the covariance matrix \(\Sigma\) of the Gaussian.
The actual probability can then be computed using the Gaussian formula:</p>

\[p(x \vert z) = \frac{1}{(2\pi)^\frac{d}{2} \vert \Sigma \vert^{\frac{1}{2}}} \exp\left( -\frac{1}{2} (x - \mu )^{\top} \Sigma^{-1} (x - \mu) \right)\]

<p>where \(d\) is the dimension of \(x\). 
This simplifies to</p>

\[p(x \vert z) = \frac{1}{(2\pi)^\frac{d}{2}} \exp\left( -\frac{1}{2} \Vert x - \mu \Vert^2 \right)\]

<p>if \(\Sigma\) is the identity matrix.</p>

<p>The objective is now to minimize the mismatch between the “real” data distribution \(p(x)\) and the distribution \(p_{\theta}(x)\) that our model approximates.
As I said before, we do not know \(P\), so we approximate it by sampling from our environment.</p>

<p>As mentioned before, we can measure the difference between two distributions by using the Kullback–Leibler divergence:</p>

\[D_\text{KL}(p(x) \Vert p_{\theta}(x)) = \sum_x p(x) \ln\frac{p(x)}{p_{\theta}(x)}.\]

<p>We want to minimize the \(D_\text{KL}\).
Any mismatch between the two distributions will result in \(D_\text{KL} &gt; 0\).</p>

<p>By using log rules we can convert the equation to:</p>

\[D_\text{KL}(p(x) \Vert p_{\theta}(x)) = \sum_x p(x) \ln p(x) - \sum_x p(x) \ln p_{\theta}(x).\]

<p>The first term is the entropy of the data (a measure of uncertainty inherent in the observations).
This entropy only depends on the true data.
Therefore, we cannot optimize it away—we cannot influence it—thus we can regard it as a constant.
Consequently, we can focus solely on the second term and minimize</p>

\[-\sum_x p(x) \ln p_{\theta}(x)\]

<p>or</p>

\[\arg\max\limits_{\theta} \sum_x p(x) \ln p_{\theta}(x).\]

<p>Again, we do not know \(p(x)\).
We only have access to \(N\) samples from it.
Therefore, we approximate:</p>

\[\sum_x p(x) \ln p_{\theta}(x) \approx \frac{1}{N} \sum\limits_{i=1}^N \ln p_{\theta}(x_i).\]

<p>The averaging works because high-probability samples should appear more often in our data; their contribution remains consistent after averaging.</p>

<p>Up to this point, our <em>generative model</em> maps \(z\) to \(p_{\theta}(x \vert z)\).
It outputs means and standard deviations for Gaussians.
What is left is the expression \(p_{\theta}(x)\).</p>

<p>To compute the total probability of an observed value \(x\), we must take into account that different values of the latent \(z\) could have generated it.
Mathematically we sum the probability that each latent value could explain our observation, weighted by how <em>likely</em> that latent value is according to the prior distribution:</p>

\[p_{\theta}(x) = \sum p_{\theta}(z) p_{\theta}(x \vert z).\]

<p>Both the parameters of the prior and the weights that transform the latent into the conditional distribution are what our model needs to learn.</p>

<h3 id="how-generative-models-learn">How Generative Models Learn</h3>

<p>How do we find or optimize good parameters \(\theta\) for our generative model \(p_\theta(x \vert z)\) and our prior beliefs \(p_{\theta}(z)\)?
We basically follow five steps:</p>

<p><strong>(1)</strong> Take a data point from the training dataset.</p>

<p><strong>(2)</strong> Randomly sample many candidates \(z_k\) from our current prior \(p_{\theta}(z)\).</p>

<p><strong>(3)</strong> Map each \(z_k\) through the generative model \(G_{\theta}(z)\) to get means \(\mu_k\) and our covariance matrix \(\Sigma_k\). Remember that the generative model is used to model \(p(x \vert z)\), i.e. to generate sensory data \(x\) given the causes \(z\).</p>

<p><strong>(4)</strong> Given \(\mu_k, \Sigma_k\) we compute, for each \(z_k\), the value of \(p_{\theta}(x \vert z_k)\) (likelihood) where</p>

\[p_{\theta}(x \vert z_k) \sim \exp\left(\frac{-\Vert x-\mu_k \Vert^2}{2} \right).\]

<p><strong>(5)</strong> Compute the log marginal likelihood estimate</p>

\[\ln\left[\sum_{z}^{\{z_k\}} p_{\theta}(z)  p_{\theta}(x \vert z) \right].\]

<p>We iteratively update \(\theta\) to maximize our log-likelihood across the dataset, usually using backpropagation and gradient descent with stochastic sampling.</p>

<p><strong>Remark:</strong> The brain does not use backpropagating or gradient decent, consequently, Friston relies on <em>predictive coding</em> which is neurally plausible: It mirrors cortical feedback and feedforward loops, with ascending prediction errors and descending predictions.</p>

<h3 id="recognition-as-guided-sampling">Recognition as Guided Sampling</h3>

<p>One problem we face is step <strong>(2)</strong>, especially if the latent space has high dimensionality.
I already mentioned that it is not feasible to generate a bunch of possible sensory data to compare them with \(x\) to find good possible causes \(z\).
There are usually too many possible causes, and because we approximate our unknown prior distribution by sampling, we would need to somehow cover the whole latent space.
This is computationally infeasible because the number of required samples increases exponentially with the number of dimensions of the latent space (<em>curse of dimensionality</em>).</p>

<p>However, usually only a tiny fraction of those causes lead to our observation—most latents do not explain the data!
Thus, we basically want to sample those causes \(z\) that are likely to result in \(x\).
Importantly, while we are okay with ignoring causes that do not matter, we do not want to miss rare causes.
The idea is to oversample rare cases instead of risking missing them, and then adjust for this oversampling mathematically.</p>

<p>To know which regions we want to sample from, we train a separate neural network \(R_{\theta}\) to serve as our “guide”.
It is our <em>recognition model</em> \(R_\theta\) that tries to approximate the inversion of the <em>generative model</em>.
It learns a <strong>variational distribution</strong> \(q_{\theta}(z \vert x)\) which predicts, for each data point \(x\), the distribution over the latent space, focusing on regions that have likely generated \(x\).</p>

<p>Consequently, our formula for \(p_{\theta}(x)\) changes from</p>

\[p_{\theta}(x) = \sum p_{\theta}(z) p_{\theta}(x \vert z) = \mathbb{E}_{z \sim p_{\theta}(z)} \left[ p_{\theta}(x \vert z) \right]\]

<p>to</p>

\[p_{\theta}(x) = \sum \frac{p_{\theta}(z)}{q_{\theta}(z \vert x)} q_{\theta}(z \vert x) p_{\theta}(x \vert z).\]

<p>which can also written as the expected value of the likelihood scaled by the ratio of sampling frequencies over the distribution \(q\):</p>

\[\mathbb{E}_{z \sim q_{\theta}(z \vert x)} \left[ p_{\theta}(x \vert z) \frac{p_{\theta}(z)}{q_{\theta}(z \vert x)} \right].\]

<p>Mathematically these two equations are equivalent.
However, note that we are using a finite number of samples and that \(q_{\theta}(z \vert x)\) determines which \(z_k\)’s we continue training on.
Therefore, we are computing an estimation of the expected value.
\(\frac{p_{\theta}(z)}{q_{\theta}(z \vert x)}\) is an adjustment for sampling bias: it adjusts the weight for events we sample more frequently than they occur (e.g., rare but important cases).</p>

<p>Now we take the logarithm to arrive at our term to maximize:</p>

\[\ln p_{\theta}(x) = \ln \mathbb{E}_{z \sim q_{\theta}(z \vert x)} \left[ p_{\theta}(x \vert z) \frac{p_{\theta}(z)}{q_{\theta}(z \vert x)} \right].\]

<p>However, the logarithm of an expectation is hard to optimize.
It is noisy and unstable.
Furthermore, computing the gradient would require that we first compute the average and only then propagate the error.
This means that after computing \(z\) using our <em>recognition model</em>, we would need to first compute multiple probabilities for \(x\) given \(z\) using our <em>generative model</em> before starting backpropagation (i.e. \(\theta\)-optimization).
It is a computational bottleneck.</p>

<p>Instead, we would like to swap the order like this:</p>

\[\mathbb{E}_{z \sim q_{\theta}(z \vert x)} \ln \left[ p_{\theta}(x \vert z) \frac{p_{\theta}(z)}{q_{\theta}(z \vert x)} \right]\]

<p>But the log of an average and the average of a log are not equal.
However, since the logarithm is a concave function we get:</p>

\[\ln \mathbb{E}[X] \geq \mathbb{E}[\ln X]\]

<p>or in our case</p>

\[\ln \mathbb{E}_{z \sim q_{\theta}(z \vert x)} \left[ p_{\theta}(x \vert z) \frac{p_{\theta}(z)}{q_{\theta}(z \vert x)} \right] \geq \mathbb{E}_{z \sim q_{\theta}(z \vert x)} \ln \left[ p_{\theta}(x \vert z) \frac{p_{\theta}(z)}{q_{\theta}(z \vert x)} \right].\]

<p>The right-hand side is a lower bound, also called the <em>model evidence</em> or <em>evidence lower bound</em> (ELBO).
The negative ELBO is known as <em>variational free energy</em>.
Maximizing ELBO also maximizes our original objective, but maximizing ELBO is computationally much easier.</p>

<p>So we want to maximize the right-hand side:</p>

\[\ln p_{\theta}(x) \geq \mathbb{E}_{z \sim q_{\theta}(z \vert x)} \ln \left[ p_{\theta}(x \vert z) \frac{p_{\theta}(z)}{q_{\theta}(z \vert x)} \right]\]

<p>and, using log rules, we arrive at the following:</p>

\[\ln p_{\theta}(x) \geq \underbrace{\mathbb{E}_{z \sim q_{\theta}(z \vert x)} \left[ \ln p_{\theta}(x \vert z) \right]}_{\text{accuracy}} - \underbrace{\mathbb{E}_{z \sim q_{\theta}(z \vert x)} \left[ \ln \frac{q_{\theta}(z \vert x)}{p_{\theta}(z)} \right]}_{\text{complexity}}.\]

<p>We arrive at a beautiful interpretation.
The first term</p>

\[\mathbb{E}_{z \sim q_{\theta}(z \vert x)} \left[ \ln p_{\theta}(x \vert z) \right]\]

<p>is the <em>accuracy</em> mentioned before, and the second term</p>

\[\mathbb{E}_{z \sim q_{\theta}(z \vert x)} \left[ \ln \frac{q_{\theta}(z \vert x)}{p_{\theta}(z)} \right]\]

<p>is the <em>complexity</em>.
The second term is also the negative of the Kullback–Leibler divergence between the distribution \(q\) and the prior distribution of latents, i.e. \(D_\text{KL}\left[ q_{\theta}(z \vert x) \Vert p_\theta(z) \right]\).
It prevents the “guide” from becoming too specialized and inventing overly complex distributions of latents.
It ensures that the distribution of latent factors for any specific data point (observation) does not stray too far from our general prior belief about the latent space as a whole.
In other words, it ensures that the latent space remains smooth and numerically well-behaved.</p>

<h2 id="the-organism-as-its-own-evidence">The Organism as Its Own Evidence</h2>

<p>Friston et al. importantly consider the observer perspective and argue that, in actuality, the <strong>system/agent/organism is not trying to maximize reward</strong>.
It is the prior belief that guides <em>perception</em> and <em>action</em> to construct an environment that (partly) aligns with the organism’s beliefs.
In that sense, organisms are (partly) their own self-fulfilling prophecy—they constantly try to “make true” what they “believe” while, at the same time, the filtered and selected “truth” changes their beliefs.</p>

<blockquote>
  <p>The inherent circularity obliges the system to fulfil its own expectations. In other words, the system will expose itself selectively to causes in the environment that it expects to encounter. […] Anthropomorphically, we may not interact with the world to maximise our reward but simply to ensure it behaves as we think it should. – <a class="citation" href="#friston:2006">(Friston et al., 2006)</a></p>
</blockquote>

<p>If an insect has evolved to expect its environment to be dark, it will use its action, e.g. movement, to keep its environment dark.
It acts to keep its sensory data consistent with its belief—moving into shadows, avoiding light.
To an external observer, this looks like as if the insect has learned that darkness is good, and has been reinforced to prefer it but from the insect’s own perspective, it’s just acting to reduce surprise by finding sensory data that matches its expectations.</p>

<p>Consequently, <strong>value</strong> is not the interpretation of an external signal (<em>reward</em>) but is constructed by a prior belief about what states minimize surprise—what states are “normal” or “preferred” and since experience and prior belief influence each other in a feedback loop, we can hardly observe direct causal relation between the two.
Value is basically the inverse of surprise and instead of figuring out which states are good based on reward, the FEP says: Organisms are already biased to expect to be in certain types of states (like being alive, warm, safe, fed).
Preferences are encoded as priors over observations/states, so log-preference shows up as negative expected surprise in the expected free energy.
Note however that <em>action-selection version of FEP</em> uses expected free energy (forward-looking), which decomposes into a <em>pragmatic</em> term (preference satisfaction) and an <em>epistemic</em> term (information gain / curiosity) to account for exploration!
These are <strong>low-surprise states</strong>—what the organism is used to.
Therefore, organisms act not to find new rewards, but to stay within those expected (familiar) states <a class="citation" href="#friston:2011">(Friston, 2011)</a>.
One can still include costs or punishments in the model—but they are treated as part of the prior beliefs about what kinds of states should be avoided.
For example, if being in pain is costly, the organism has a prior belief that “I don’t usually feel pain”.</p>

<blockquote>
  <p>The problem of finding sparse rewards […] is nature’s solution to minimizing entropy. – <a class="citation" href="#friston:2011">(Friston, 2011)</a></p>
</blockquote>

<p>That is, rewards are rare and localized because organisms are repeatedly drawn to a small set of predictable, low-surprise states—such as food, shelter, and safety.</p>

<blockquote>
  <p>These dynamics rest on complementary self-construction (autopoiesis) and destruction (autovitiation). – <a class="citation" href="#friston:2011">(Friston, 2011)</a></p>
</blockquote>

<p>In other words, the organism builds up behaviors that help it stay viable, and eliminates those that lead to bad outcomes.
Behaviors emerge because the agent must stay in a balance with its environment—not too chaotic, not too static—to continue existing.</p>

<p>So why doesn’t the agent settle in one fixed state?
Because organisms or living systems—like animals or humans—are not static systems.
They have to continuously interact with their environment to survive.
We cannot just stay in a single “perfect” state.
“Being alive” requires constant change (eating, breathing, sleeping, thermoregulating, moving around).
If we would stay in one place (one state), we would have violated our own prior beliefs about what kinds of sensory inputs and physiological states we expect to have over time.
We are in <em>itinerancy</em>: moving through a cycle of low-surprise states, not staying in one.
Thus, <em>itinerant policies</em> arise because staying in a single state forever would be more surprising (and thus less viable) than cycling through a set of familiar, predictable states over time <a class="citation" href="#friston:2011">(Friston, 2011)</a>.</p>

<blockquote>
  <p>To minimize surprise, organisms must move through a predictable sequence of diverse but familiar states.</p>
</blockquote>

<p>Learning, perception, and action then is about <strong>adjusting expectations</strong> and <strong>updating generative models</strong> on different time scales. 
To the best of my knowledge, it is unclear in which way one dominates over the others but it seems as if all three are highly depend on each other.
From this I conclude that it is important to dwell on problems, to think, to do the hard work to adjust long-term causes to be able to construct a reality/environment that makes sense consistently.
But it is also important to act in such a way that one is exposed to surprise, at least from time to time.</p>

<blockquote>
  <p>Without surprise we are not adapting our belief system.</p>
</blockquote>

<p>A realist will argue that the survival of an organism is bound to the accuracy of its model about reality—meaning its expectations align well with it, therefore, we can assume that we more or less experience reality as it is.</p>

<blockquote>
  <p>Because the free energy is low, the inferred causes approximate the real environmental conditions. This means the systems physical state must be sustainable under these environmental forces, because each system is its own existence proof. – <a class="citation" href="#friston:2006">(Friston et al., 2006)</a></p>
</blockquote>

<p>A constructivist will point out that our beliefs (latents, causes) only tend to generate sensory data and that we can not jump from this to an objective reality or “the truth” which is “out there”.
The causes we believe in may have little or nothing to do with the “real” causes—if such causes even exist.
This challenges my earlier assumption of <em>indirect realism</em>.
Causes are just effective in generating sensory data—data that is already constructed by the very same system which tries to reconstruct them.
While I don’t see constructivism as an anti-realism, constructivist might doubt any sort of realism that rejects constructivist elements and that assumes that we can reach an observer-independent “objective turth”.
Constructions are contingent and constrained—they work in the context of a history of evolution but they are not “the truth” and, in my opinion, they can not be separated from the observing system that construct them.</p>

<p>Because <em>recognition</em> and <em>perception</em> work in tandem, we can say that changing our expectations literally changes our perception—it changes our environment.
But at the same time, our perception changes our expectations if we are surprised.
Thus the FEP highlights the circularity of <em>perception</em> and <em>action</em> or <em>belief</em> and <em>reality</em> because organisms minimize free energy not only by adapting their beliefs to fit their environment but also by adapting their environment to fit their beliefs.
It assumes an environment that is acted upon effectively (which is a somewhat <em>optimistic</em> outlook and motivator to make use of our imagination to construct beliefs of a less cruel world).</p>

<p>Also, I want to emphasise, that surprise that leads to a loss of one’s well-known environment, can be painful because one’s prior beliefs about the world are violated—you enter a state of increased entropy.
This might lead to a temporary suspension of meaning because you no longer know what to expect, or how to act effectively.
Adaptation is painful because it involves breaking down a stable, low-free-energy model, and building a new one that’s better at minimizing future surprise.
It is especially painful if the previous model had high precision priors (e.g. strong, confident beliefs), the environment changes faster than the model can update (e.g. economic instability, cultural shifts), or there is no clear path to a new, stable <em>attractor state</em>.</p>

<p>In today’s society the individual is expected to be flexible, to constantly re-skill and adapt to environmental changes.
For example, we switch our jobs more and more frequently, which is often seen as desirable or necessary.
Flexibility—constantly updating beliefs, changing environments, shifting roles—might sound adaptive.
And in a way, it is.
But under FEP, excessive flexibility undermines model stability.
Chronic uncertainty leads to chronic prediction error, which is metabolically and psychologically exhausting and if we are constantly updating our generative model without forming stable priors, we fail to build any useful expectations.
This creates a sense of incoherence or alienation—a <strong>loss of identity</strong>; a lived experience of “nothing makes sense”.
So in modern systems that demand rapid flexibility, we’re often forced to dissolve stable models faster than we can reconstruct them.
The result? Anxiety, burnout, and a background hum of ontological instability.</p>

<p>The FEP suggest that there is a balance to be found: exposure to surprise and a certain “willingness” to adapt without losing ones identity and purpose.
I find it desirable to recognize, or at least consider that emotional pain may not reflect maladaptation but the cost of updating one’s expectations in the face of an unstable world.
Flexibility, when demanded too often or too rapidly, erodes the very models that minimize surprise over the long term.
Thus, while adaptability is vital, stability is sacred.</p>

<p>Let me finish with some speculations:
I have the gut feeling that a strange, perhaps paradoxical, alliance forms here between Kant, Heidegger, and (embodied) constructivists like Maturana, Varela, and Luhmann. 
Each offers a piece of the puzzle when viewed through the lens of the free energy principle.
Kant emphasized the structuring role of <em>a priori</em> reason—the internal scaffolding that makes experience possible.
Futhermore, FEP mirrors Kant’s view that perception is constructed by the mind’s active faculties.
The world “as it is” is unknowable directly—we only perceive it through mental filters and structure.
Heidegger, in contrast, stressed the <em>immediacy (i.e. pre-conceptual) and embeddedness of experience</em>: we are not detached observers but beings thrown into a meaningful world that is not build in the mind. 
Heidegger’s philosophy fits tightly with theories that view perception as <em>embodied</em>, <em>skillful</em>, and <em>embedded in practical action</em>, not representational—just like <em>enactivism</em> and Gibson’s <em>direct perception</em>.
Constructivists such as Maturana and Varela also went with Heidegger but moved the world into the system itself, arguing that <em>cognition arises through dynamic interaction with an environment</em> that is, in part, <em>selected</em> or <em>enacted</em> by the organism itself.</p>

<p>Yet, each position has its blind spots. 
Kant may have overestimated the sovereignty of internal reason;
Heidegger arguably romanticized lived immediacy; and constructivists often risk underestimating the constraints imposed by sensory data—or more precisely, by the need to stay within states that minimize surprise. 
Yes, action and cognition are deeply intertwined (Maturana, Varela), and yes, the world is our best model of itself (Heidegger) but we still model it and it pushes back (Kant).</p>

<p>If we want to go full Heideggerian, we probably need a different interpretation where free energy minimization becomes about maintaining practical grip on the world, not constructing a picture of it and where prediction errors don’t refer to “incorrect beliefs” but breakdowns in coping (e.g., when the tool “refuses” to function).
We would have to move to a life-mind continuity thesis (LMCT) (see e.g. <a class="citation" href="#bogota:2024">(Bogotá, 2024)</a>).
It would be interesting and maybe refreshing to departure from the representation-driven approaches (it is unsurprising that one of the most famous AI conferences is called “International Conference on Learning Representations”).
But this would be another topic for another time.</p>

<hr />

<h2 id="appendix">Appendix</h2>

<p>The generative neural network \(G_\theta\) computes an approximation of \(p(x \vert z)\) given \(z\) but as a Gaussian distribution (its means and standard deviations).
So, why using the Gaussian is a good choice?</p>

<p>The central limit theorem (CLT) states that:</p>

<blockquote>
  <p>If you take a large number of independent and identically distributed (i.i.d.) random variables with finite mean and variance, then their properly normalized sum tends toward a Gaussian (normal) distribution, regardless of the original distribution of the variables.</p>
</blockquote>

<p>The Gaussian appears often in nature because of the aggregation of many small effects and it has the maximum entropy among all distributions with a given mean and variance.
It is the “most random” or “least biased” which makes it a natural default in many uncertain systems.</p>

<p>So first, it is mathematically convenient.
The Gaussion has a very simple functional form:</p>

\[q(z) = \mathcal{N}(z \vert \mu, \Sigma)\]

<p>It is fully described by just two parameters: mean \(\mu\) and covariance \(\Sigma\).
This makes optimization tractable because:</p>

<ul>
  <li>Expectations like \(\mathbb{E}_q[\ln p(x,z)]\) can often be computed in closed form or with low-variance Monte Carlo estimates.</li>
  <li>(Differential) entropy \(H(P)\) is known analytically for a Gaussian. It is \(H(\mathcal{N}) = \frac{1}{2} \ln((2\pi e)^d \vert \Sigma \vert)\), which simplifies the so-called evidence lower bound (ELBO):</li>
</ul>

\[\text{ELBO} = \mathbb{E}_q[\ln p(x,z)] - \mathbb{E}_q[\ln q(z)].\]

<p>Secondly, the Gaussian is reparameterizable.
Instead of sampling \(z \sim \mathcal{N}(\mu, \sigma^2)\) directly, we can rewrite it as:</p>

\[z = \mu + \Sigma^{1/2} \epsilon, \quad \epsilon \sim \mathcal{N}(0, I).\]

<p>Now, we are sampling from a fixed standard normal \(\epsilon\), and transforming it using differentiable operations involving \(\mu\) and \(\Sigma\).
Because \(z\) is now a function of \(\mu\) and \(\Sigma\) its gradients can flow.
It makes the entire process differentiable with respect to the parameters because \(\epsilon \sim \mathcal{N}(0, I)\) is independent of these parameters, enabling optimization via gradient descent.
In other words, the randomness becomes independent of ones model’s parameters, and the transformation from parameters to the sample becomes differentiable.</p>

<p>Thirdly, even if the true posterior is not Gaussian, many high‑dimensional distributions exhibit approximately Gaussian local behavior (via central limit theorem effects or Laplace approximations).
Furthermore, Gaussians are smooth and unimodal, making them a good first approximation for a wide variety of distributions.</p>

<p>Then there is the topic of computational efficiency.
Working with Gaussians keeps the variational family simple because (1) linear algebra operations are efficient, (2) sampling is cheap, and (3) we avoid expensive numerical integration or more complex families that would be harder to optimize.</p>

<p>Lastly, this framework is flexible.
Even though a single Gaussian is unimodal, we can extend it by using mixtures of Gaussians, normalizing flows, and Gaussian processes (an infinite‑dimensional generalization).</p>

<p>In summary: the Gaussian is used because it strikes the right balance between <strong>expressiveness</strong> (locally), <strong>tractability</strong> (closed‑form solutions and gradients), and <strong>optimization‑friendliness</strong>.</p>

<h2 id="references">References</h2>

<ol class="bibliography"><li><span id="friston:2006">Friston, K., Kilner, J., &amp; Harrison, L. (2006). A free energy principle for the brain. <i>Journal of Physiology-Paris</i>, <i>100</i>(1), 70–87. https://doi.org/10.1016/j.jphysparis.2006.10.001</span></li>
<li><span id="barry:2012">Razeto-Barry, P. (2012). Autopoiesis 40 years later. A review and a reformulation. <i>Origins of Life and Evolution of Biospheres</i>, <i>42</i>(6), 543–567. https://doi.org/10.1007/s11084-012-9297-y</span></li>
<li><span id="virenque:2024">Virenque, L. (2024). What is agency? A view from autonomy theory. <i>Biological Theory</i>, <i>19</i>(1), 11–15. https://doi.org/10.1007/s13752-023-00441-5</span></li>
<li><span id="barandiaran:2017">Barandiaran, X. E. (2017). Autonomy and enactivism: Towards a theory of sensorimotor autonomous agency. <i>Topoi</i>, <i>36</i>(3), 409–430. https://doi.org/10.1007/s11245-016-9365-4</span></li>
<li><span id="bogota:2024">Bogotá, J. D. (2024). Where there is life there is mind ... and free energy minimisation? In M. Martín-Villuendas, J. Gefaell, &amp; A. Cuevas-Badallo (Eds.), <i>Life and Mind: Theoretical and Applied Issues in Contemporary Philosophy of Biology and Cognitive Sciences</i> (pp. 171–200). Springer Nature Switzerland. https://doi.org/10.1007/978-3-031-70847-3_8</span></li>
<li><span id="varela:1997">Varela, F. J. (1997). Patterns of life: Intertwining identity and cognition. <i>Brain and Cognition</i>, <i>34</i>(1), 72–87. https://doi.org/https://doi.org/10.1006/brcg.1997.0907</span></li>
<li><span id="cummins:2014">Cummins, F. (2014). Agency is distinct from autonomy. <i>Avant: Trends in Interdisciplinary Studies</i>, <i>5</i>(2), 98–112. https://doi.org/10.26913/50202014.0109.0005</span></li>
<li><span id="gabriel:2018">Gabriel, M. (2018). <i>Der Sinn des Denkens</i>. Ullstein Buchverlag.</span></li>
<li><span id="tomasello:2024">Tomasello, M. (2024). <i>Die Evolution des Handelns</i>. Suhrkamp.</span></li>
<li><span id="friston:2011">Friston, K. J. (2011). Embodied inference: or “I think therefore I am, if I am what I think“. In W. Tschacher &amp; C. Bergomi (Eds.), <i>The Implications of Embodiment (Cognition and Communication)</i> (pp. 89–125). Imprint Academic.</span></li></ol>]]></content><author><name>Benedikt Zönnchen</name><email>benedikt.zoennchen@web.de</email></author><category term="ML" /><category term="Neuroscience" /><summary type="html"><![CDATA[Imagine you see a grizzly bear in the woods and half of its face is hidden behind a tree. You will most certainly recognize the bear as whole and dangerous animal. Your brain will “fill in the gaps” even though there is no absolute certainty that the bear isn’t split in half. How and why does the brain do this? Furthermore, why are we fooled by all sorts of graphical illusions? In other words, why are we, in some instances, so stubborn to see what is not there—even if we are told that it is not there—and, on other occasions, we can immediately see what is hidden and probably there?]]></summary></entry><entry><title type="html">Crises of Communication: Sustainability and Trumpism</title><link href="https://bzoennchen.github.io/Pages/2024/10/30/a-crisis-of-non-communication.html" rel="alternate" type="text/html" title="Crises of Communication: Sustainability and Trumpism" /><published>2024-10-30T00:00:00+01:00</published><updated>2024-10-30T00:00:00+01:00</updated><id>https://bzoennchen.github.io/Pages/2024/10/30/a-crisis-of-non-communication</id><content type="html" xml:base="https://bzoennchen.github.io/Pages/2024/10/30/a-crisis-of-non-communication.html"><![CDATA[<p>I have to admit, I’m somewhat addicted to thinking about <em>social systems theory</em>, particularly the version developed by the German sociologist Niklas Luhmann (1927–1998). 
But this addiction isn’t driven by pure fascination—it’s more of a love-hate relationship. 
On one hand, seeing the world through a Luhmannian lens doesn’t make it more just or fantastic, but it does make it comprehensible. 
The chaotic state of global affairs—the craziness of U.S. elections, the brutal devastation in the Middle East, the war in Europe, and the looming tensions brought by the climate crisis—appears senseless at times, even apocalyptic.
And yet, <em>systems theory</em> offers a framework that, paradoxically, brings coherence to this apparent madness, showing us how these outcomes emerge from the logic of <em>functionally differentiated systems</em>.</p>

<p>Faced with such overwhelming disorder, how can anyone retain a sense of sanity? 
We can protest, advocate for change, and try to amplify ‘reasonable’ voices, but as we’ll see, whether these efforts have the desired effect is largely out of our hands.
One might speculate that a Trump victory would worsen the Middle East conflict (and the conflict in Ukraine), but, as citizens of the world, we have no ballot for de-escalating global tensions.
People outside the U.S. cannot vote in American elections, and even if Kamala Harris were to win, there’s little reason to assume that U.S. policy would suddenly prioritize humanitarian investment in preserving lives and infrastructure abroad—the very foundations of a decent existence.</p>

<p>The other option is to try to make sense of the seemingly senseless, not to justify what’s happening but to deepen our understanding and foster empathy, even with those we might otherwise label as perpetrators of harm. 
To ease tensions one has to move beyond simplistic judgments of good and evil. 
Strangely enough, Luhmann’s <em>anti-humanistic</em> theory might aid in this endeavor, as it places individuals outside the bounds of society, such that we can shift our blame from <em>souls to systems</em>.
In Luhmann’s view, social dynamics operate autonomously, driven by complex systems that rarely align with individual intentions.
In many ways, Luhmann’s thinking aligns with the ideas of several twentieth-century French theorists, such as Baudrillard, Foucault, Deleuze, and Derrida.
However, his approach is less dramatic and, in a stereotypically German way, more detached and methodical.
While he shares many of the French theorists’ insights into power, society, and structural dynamics, he refrains from moral interpretations, neither labeling these dynamics as inherently good nor bad. Instead, he offers a distant, almost clinical description of society—a detached analysis that seeks to understand social mechanisms without prescribing judgment.
His perspective allows us to look past personal blame, seeing <em>dysfunction</em> not as a failure of character but as a product of systemic logic that no one person controls.</p>

<h2 id="beyond-good-and-evil">Beyond Good and Evil</h2>

<p>Clearly, attempting to make sense of complex issues should not be mistaken for rationalization.
This is an easy trap to fall into, particularly when the topic is heavily charged with moral language, where any effort to explain events can be (willingly) misinterpreted as either justification or condemnation. 
The framework within which sense-making occurs in such cases often defaults to the age-old binary of good versus evil—arguably the most effective, yet oversimplified, way to reduce complexity.
This moral framing gives us a manageable lens through which to view the world, but it also makes us blind for a more nuanced perspective, risks obscuring deeper systemic dynamics and can hinder genuine understanding of the complex interactions at play.
As a German comedian once said:</p>

<blockquote>
  <p>If you know who is the devil, your day is already well-structured.</p>
</blockquote>

<p>Most would agree that this binary division of good versus evil is overly simplistic and, at times, dangerous.
Yet, when we look toward the U.S. election or the language used in the current horrific conflict in the Middle East, we see a striking example of this polarization in action—a place where the framing of social and political conflicts as battles between absolute good and evil has become more pronounced than ever.
Especially the situation in Palestine is extremely hard to swallow without being overwhelmed by emotions and breaking down in tears.
Here the moral dichotomy shows its face and its power.
It is so dangerous because to defeat absolute evil everything is permitted.</p>

<p>I have the luxury to shift my perspective from judging individuals as good or evil to evaluating systems as either <em>functional</em> or <em>dysfunctional</em>.
Instead of moralizing, I can try to focus on understanding the underlying structures and processes that contribute to societal challenges.
Someone directly affected by these conflicts probably cannot.</p>

<p>Thinking in terms of systems brings a certain relief from the confusion and frustration of modern life.
It helps to lessen anger and bewilderment about why our <em>life-world</em> and the decisions made within it often seem so irrational or even absurd.
Systems theory sheds light on why, even in an era where nearly everyone can participate in media production, we have neither reduced manipulation nor fostered a more reasonable dialogue. 
In many ways, the Enlightenment’s aspirations for rational discourse and universal truth have not materialized as hoped. 
The ideal of ‘Truth’—which Plato connected to the ‘Good’—has, it seems, drifted into obscurity and we are left with a spectacular hyperreality; it seems we have been fallen deep into the cave of shadows.</p>

<p>In a typical postmodern move, Luhmann’s conclusion to the ‘lost Truth’—understood as singular objective truth—is (similar to Baudrillard) that it never existed in the first place.
As a constructivist, he avoids Plato’s concept of ‘the Truth’ and shifts his attention to the <em>production of sense</em> via different systems.
If we take Luhmann’s theory seriously, we must recognize that controlled, predictable change within society is extremely limited because each (social) system constructs its own reality.
There is no agreement on what is true or real.
Fundamentally, many problems arise from <strong>the difficulty of communication</strong>.
Adopting this perspective introduces a sense of helplessness because even with well-intentioned or radical actions, there is no guarantee that the outcomes will align with the respective intentions.
As individuals, we find ourselves positioned outside the social systems that coevolve with their environment according to their own complex dynamics.
This view is both awe-inspiring and disquieting: I admire the explanatory power of Luhmann’s theory, yet I feel a deep urge to challenge or even disprove it.
It confronts us with the unsettling notion that society evolves autonomously, beyond our direct influence, regardless of our individual ideals and aspirations.</p>

<p>I first encountered Luhmann’s theory while preparing a lecture on sustainable artificial intelligence. 
In researching future competencies, including sustainability competencies <a class="citation" href="#Brundiers2020">(Brundiers et al., 2020)</a>, I noted that systems thinking is emphasized as a fundamental skill for addressing complex issues. 
However, I doubt that Luhmann’s work appears on the reading lists or syllabi of most courses on sustainability.
Outside of Germany, he remains relatively unknown for a few reasons. His writing, e.g.,</p>

<ul>
  <li>Die Wissenschaft der Gesellschaft (The Science of Society) <a class="citation" href="#luhmann:1992">(Luhmann, 1992)</a></li>
  <li>Die Wirtschaft der Gesellschaft (The Economy of Society) <a class="citation" href="#luhmann:1994">(Luhmann, 1994)</a></li>
  <li>Die Kunst der Gesellschaft (The Art of Society) <a class="citation" href="#luhmann:1997">(Luhmann, 1997)</a></li>
  <li>Die Politik der Gesllschaft (The Politics of Society) <a class="citation" href="#luhmann:2002">(Luhmann, 2002)</a></li>
  <li>Die Gesllschaft der Gesellschaft (The Society of Society) <a class="citation" href="#luhmann:1998">(Luhmann, 1998)</a></li>
</ul>

<p>is notoriously technical and repetitive, and his theory clashes with the Western concept of the <em>sovereign individual</em> <a class="citation" href="#moeller:2011">(Möller, 2011)</a>.
Note that each title has a double meaning, e.g. <em>The Science of Society</em> discusses the social system called science but it is also a specific description written by society, hinting at the fact that there is no perspective from outside.
Consequently, <em>The Society of Society</em> is a self-description of society.</p>

<p>Additionally, Luhmann sidesteps moral language and offers no prescriptive or normative framework.
Those looking to his theory for answers on what to do will likely be disappointed.</p>

<p>Today, <em>systems thinking</em> is widely discussed as a method to grasp the complexity of global issues by focusing on wholes and relationships rather than dissecting problems into isolated parts. 
It seeks to move beyond Cartesian reductionism and the Newtonian view of linear cause-and-effect, proposing instead an anti-reductionist approach that emphasizes interdependencies, especially crucial in fields like climate science. 
Here, we are not dealing with a computable universe but with complex and chaotic systems, where non-linearity and feedback loops disrupt straightforward causal relationships. 
However, we should remind ourselves that chaos is different from randomness. 
Chaotic systems can exhibit intricate structures, yet the slightest change in initial conditions can drastically alter future outcomes, as we see with weather systems—a classic examples of chaotic behavior.
While accurate short-term predictions are challenging due to the chaotic nature of weather, long-term averages, such as the global average temperature, can be predicted with reasonable accuracy.
Unlike weather, which is highly sensitive to initial conditions, climate trends respond more predictably to persistent external drivers like greenhouse gas concentrations and solar radiation. 
However, when we consider societal factors, the picture becomes more complex, as human activities and policy decisions can significantly influence these long-term climate trends.</p>

<p>In essence, <em>systems thinking</em> itself is a kind of <em>technology</em>, and many hope it will equip us to address the <em>climate crisis</em>.
However, I believe there are distinct schools of thought within <em>systems thinking</em>, each relying on different levels of abstraction, and they are not necessarily compatible.
If systems thinking is indeed essential for tackling the <em>climate crisis</em>—a hypothesis I support—then it stands to reason that we should understand the social dimensions of the climate crisis and other global issues through the lens of one of sociology’s most sophisticated systems thinkers, that is arguable, Niklas Luhmann. His framework provides a unique approach to examining the complex, interdependent nature of social systems that underlie and influence our responses to the <em>climate crisis</em>.</p>

<h2 id="so-what-is-a-system">So what is a System?</h2>

<p>The highly abstract term <em>system</em> is so loose and overused in so many contexts and in our daily language that it has hardly any specific meaning.
We can talk about computer systems, a system of linear or differential equations, a system of thinking, systems of oppression, ecosystems and political systems.</p>

<blockquote>
  <p>[So] what is a system? 
A system is a set of things […] interconnected in such a way that they produce their own pattern of behavior over time. […] 
[T]he system’s repsonse to these forces is characteristic of itself, and that repsonses is seldom simple in the real world. – <a class="citation" href="#meadows:2008">(Meadows, 2008)</a></p>
</blockquote>

<p>For the environmental scientist Donella Meadows (1941–2001) a system is basically <strong>an interconnected set of elements that is coherently organized in a way that achieves something</strong>.
A system is characterized by three key components:</p>

<ol>
  <li><strong>Elements</strong>: The parts or components of the system, such as individual actors, objects, or variables.</li>
  <li><strong>Interconnections</strong>: The relationships or interactions between the elements, often in the form of flows of information, energy, or material.</li>
  <li><strong>Purpose or Function</strong>: The overarching goal or behavior that the system is organized to achieve.</li>
</ol>

<p>But here the trouble begins because Luhmann defines a system very differently.
It seems to me that Meadows’ definition still relies on the subject-object distinction which Luhmann wants to sublime.
He thinks in interdependent but operationally closed processes instead of things and he very much dislikes the concept of an externally given ‘purpose’.
For Luhmann</p>

<blockquote>
  <p>A system is a self-referential, self-organizing set of operations that differentiates itself from its environment.</p>
</blockquote>

<p>This requires some explanation:</p>

<ol>
  <li><strong>Self-Referential</strong>: Systems create their own elements through their own operations. In fact, they are operations. In social systems, these elements are not people but <em>communications</em>. Each communication refers back to the system, reaffirming its boundaries and identity.</li>
  <li><strong>Autopoiesis</strong>: Systems are autopoietic, meaning they are self-producing. They continuously reproduce the <em>communications</em> that sustain them, distinguishing themselves from their environment. This process enables a system to maintain coherence and adapt to changes.</li>
  <li><strong>Environment and Differentiation</strong>: Systems are defined by the distinction between themselves and their environment. Luhmann stresses that the environment is everything that the system excludes, setting clear boundaries. This differentiation allows the system to maintain its <em>identity</em> while interacting with, but remaining distinct from, external influences.</li>
  <li><strong>Social Systems as Sense Making Networks</strong>: Luhmann focuses on social systems—such as organizations, institutions, the economy, the political system, and the mass media—as networks of meaning/sense making (‘<em>Sinn machen</em>’ in German). In these systems, communication itself is the fundamental element, and these <em>communications</em> build the system’s reality.</li>
</ol>

<p>In essence, Luhmann views a system as a closed network of <em>communications</em> that operates independently of external elements and functions primarily by sustaining itself through self-generated, meaningful <em>communications</em>.
If we want to be accurate we can not speak of a system without its environment because <strong>a system is the process that differentiates itself from its environment</strong> which is a circular definition—a paradox—that keeps the system (the system-environment differentiation) going.
This definition contrasts sharply with definitions based on tangible components and external goals, focusing instead on processes of <em>sense-making</em>, self-production and self-maintenance.</p>

<p>Luhmann read a lot of interdisciplinary material and borrowed from mathematics, classical systems theory, biology, cybernetics and other disciplines.
For example, he took the concept of <em>autopoiesis</em> <a class="citation" href="#maturana:1987">(Maturana &amp; Varela, 1987)</a> from biology, the concept of <em>feedback loops</em> and <em>second-order observation</em> from cybernetics and of <em>re-entry</em> and the fundamental operation of <em>differentiation</em> and <em>indication</em> from the mathematician Spencer-Brown <a class="citation" href="#brown:1969">(Spencer-Brown, 1969)</a>.</p>

<p>For Luhmann, systems are <em>operationally closed</em> meaning that no system can interfere in the operation of another system.
For example, the economic system communicates via payments and there is almost no way that the political system can interfer in it.
Of course, the political system can observe these payments and can try to regulate them but only indirectly.
It can pass laws which the legal system processes.
The economic system will observe this—it will ‘digest’ it—and evolve with its environment (which contains the political system).
Therefore, systems are <em>cognitively open</em> meaning that they can observe (based on their own logic) their environment which contains all the other systems.
They take everything in what they are able to digest and use it to continue their opertions, that is, their <em>autopoiesis</em>.
An analogy is a human body that takes in food and digist it in the way it is able to.
Neither does the food determine how the body is affected nor does the body can make anything it wants from the food.
The process is contingent.</p>

<p>To reduce complexity social systems work on simple <em>binary codes</em> under which they <em>differentiate</em>.
The legal system interprets actions as legal or illegal but doesn’t engage with the healthy/unhealthy distinctions from the health system.
The following table shows more of these codes:</p>

<table>
  <thead>
    <tr>
      <th><strong>Social System</strong></th>
      <th><strong>Binary Code</strong></th>
      <th><strong>Description</strong></th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><strong>Economy</strong></td>
      <td>Payment / Non-payment</td>
      <td>Decisions are guided by whether a transaction involves payment, focusing on economic exchanges.</td>
    </tr>
    <tr>
      <td><strong>Politics</strong></td>
      <td>Power / Non-power</td>
      <td>Concerned with the distribution and exercise of power, focusing on who has authority and control.</td>
    </tr>
    <tr>
      <td><strong>Law</strong></td>
      <td>Legal / Illegal</td>
      <td>Operates on legality, determining if actions or behaviors align with established legal norms.</td>
    </tr>
    <tr>
      <td><strong>Science</strong></td>
      <td>Truth / Falsehood</td>
      <td>Guided by the pursuit of truth, evaluating claims based on their validity and scientific evidence.</td>
    </tr>
    <tr>
      <td><strong>Religion</strong></td>
      <td>Immanence / Transcendence</td>
      <td>Focuses on distinctions between the sacred (transcendent) and the profane (immanent).</td>
    </tr>
    <tr>
      <td><strong>Education</strong></td>
      <td>Success / Failure</td>
      <td>Concerned with the effectiveness of learning and teaching, evaluated by success in achieving educational goals.</td>
    </tr>
    <tr>
      <td><strong>Health</strong></td>
      <td>Healthy / Unhealthy</td>
      <td>Operates based on the state of health, determining whether a body or behavior is healthy.</td>
    </tr>
    <tr>
      <td><strong>Mass Media</strong></td>
      <td>Information / Non-information</td>
      <td>Distinguishes between what is considered newsworthy (informative) versus uninformative content.</td>
    </tr>
    <tr>
      <td><strong>Art</strong></td>
      <td>Fitting / Unfitting</td>
      <td>Focused on e.g. aesthetic value, distinguishing what is perceived as beautiful or aesthetically valuable.</td>
    </tr>
  </tbody>
</table>

<p>The specific code a system operates under is less important than the fact that each code differentiates one system from others.
For instance, Luhmann struggled to pinpoint a definitive code for the art system, as its operations are complex and multifaceted.
He proposed several possibilities, including beautiful/ugly, coherent/incoherent, new/old, and fitting/unfitting.</p>

<p>Each code of a system gives the system its ‘character’ and consequently its operational specificity and functional closure.
The <em>exclusivity</em> of the <em>binary code</em> ensures that each system maintains its autonomy and operates independently, even when interacting with other systems.
<em>Operational closure</em> means that each system can only process information according to its own internal logic, thus keeping it <em>closed</em> to other systems’ codes and distinctions.</p>

<p>Of course, further distinctions within a system are possible.
For example <em>reputation</em> is an important distinction within science.
While not a binary code in Luhmann’s strict sense, it is significant within the scientific community as it affects how research and findings are perceived, valued, and disseminated.
Reputation can influence which scientists’ work is taken seriously, whose research is funded, and which publications are more widely read and cited.
However, it does not drive the core distinction of truth/falsehood; rather, it shapes the social hierarchy, credibility, and visibility within the scientific community.</p>

<p>Luhmann’s concept of <em>re-entry</em> (borrowed from <a class="citation" href="#brown:1969">(Spencer-Brown, 1969)</a>) is a mechanism, allowing a system to reflect on itself by reintroducing its primary binary code within its own operations. 
Thus, this is an inherently recursive relation.
<em>Re-entry</em> enables a system to apply its guiding binary distinction not only outwardly (to its environment or other systems) but also inwardly, to its own internal processes and <em>communications</em>. 
This is crucial in complex systems, like science, where re-entry enables self-reference and internal differentiation.
Science can use the truth/falsehood distinction to evaluate not only external hypotheses but also its own standards, research paradigms, and accepted theories.
Another more familiar example is the media which reports on itself.</p>

<p>Another important Luhmannian concept that is strongly connected to <em>re-entry</em> and orignated from Heinz von Foerster (1911–2002) and Margaret Mead (1901–1978) <a class="citation" href="#foerster:2003">(von Foerster, 2003)</a> is <em>second-order observation</em>.
<em>Second-order observation</em> is facilitated by re-entry, as it allows the scientific system to reintroduce its primary distinction, i.e. truth/falsehood, internally.
It refers to observing observations rather than simply observing objects or phenomena directly.
This concept is crucial for complex systems as it allows them to recognize and reflect on how they construct their own distinctions and interpretations.
In science, second-order observation enables scientists to observe not only external phenomena but also the methods, theories, and interpretations of other scientists.
This includes observing how truths are constructed within the scientific community and scrutinizing the frameworks, biases, and assumptions underlying those constructions.
But it also includes how a specific scientist or science lab is being observerd.
In this context, <em>reputation</em> is effectively a measure of how a scientist is seen by other scientists.
It further reduces complexity by accumulating the observation of others.
I do not have to read and carefully analyse every paper of a specific researcher to find out if his or her research is trustworthy, i.e. if it is good research under the truth/falshood code.
In complex environments, <em>second-order observation</em> helps systems like science to deal with uncertainty and complexity.
Rather than aiming for <strong>absolute certainty</strong>, science can adapt by recognizing different observational frameworks, revisiting previously accepted truths, and acknowledging limitations in current knowledge. 
This adaptive flexibility, achieved through <em>second-order observation</em>, is vital for science’s resilience and continued evolution.
Today, second-order observation is everywhere, be it in the form of the housing or stock market, the social phenomena of <em>reaction videos</em> or the fact that we are all invested in our profiles, that is, <strong>we are invested in how we are seen/observed by an anonymous peer</strong> <a class="citation" href="#moeller:2021">(Möller &amp; D’Ambrosio, 2021)</a>.</p>

<p>In the mode of authentic identity construction, second-order observation appears as the production of ‘fake’ because the focus shifts from who one <strong>is</strong> to how one <strong>is observed</strong>.
Paradoxically, as we transition from <em>authentic</em> to <em>profilitic</em> identity construction, the most important goal becomes to be observed as authentic.
This phenomenon is evident in the behavior of streamers, social media influencers (including figures like Trump and Musk), as well as companies and even institutions such as universities.
All of these entities are heavily invested in how they are observed, striving to appear authentic and real, rather than fake.
However, if we ask the existentialist question—<em>What is your ‘true self,’ your ‘authentic being’?</em>—we encounter paradoxes and articulation problems.
Luhmann provides an uncanny response to this authenticity problem, reminiscent of Kant: What one is, is always already a description or observation of a system, and therefore, a selection or differentiation—choosing one thing over another.
This includes self-observation, which is itself a re-entry of the system. In this case, the psychic system re-enters itself.
Through Luhmann’s perspective, we are left with the impossibility of being truly authentic. In some sense, we (and other systems) are always already pretending.
However, Luhmann crucially does not view this as inherently negative. There is nothing—especially morally—wrong with being invested in how one is observed.
The problem arises when this focus on appearence no longer aligns with the system’s function.
For example, when the pursuit of appearing as a trustworthy scientist contradicts the truth/falsehood distinction that underpins the operation of science.</p>

<p>But what exactly is <em>observation</em>?
For Luhmann, every act of <em>observation</em> has two essential steps: <strong>distinction</strong> and <strong>indication</strong> (also borrowed from <a class="citation" href="#brown:1969">(Spencer-Brown, 1969)</a>).
The system first creates a <em>distinction</em> (e.g., legal/illegal in the legal system) and then <em>indicates</em> one side of that distinction. 
This act of <em>indicating</em> one side of a <em>distinction</em> allows the system to focus on what it deems relevant or meaningful while leaving out what is not.
For instance, the economic system distinguishes between payment and non-payment and then indicates whether a transaction falls into one category or the other.</p>

<p>Luhmann also uses <em>feedback loops</em> but in a more complex, indirect way to explain self-referential and <em>autopoietic</em> (self-producing) processes within social systems.
<em>Feedback loops</em> enable systems to observe and respond to their own operations and their environment without sacrificing their internal logic or autonomy.
A system produces <em>communications</em> and then feeds those back into itself as input, creating a recursive process. 
For instance, the scientific system continually generates new research findings that become part of its ongoing discourse, which shapes further research questions and methods.
This recursive process enables a system to build on its own operations and maintain continuity over time.
<em>Feedback loops</em> help systems to learn from past operations and adjust future <em>communications</em>. 
However, instead of direct feedback that leads to specific, immediate corrections (as in a thermostat, for instance), feedback in Luhmann’s theory involves observing patterns over time and <em>adjusting structurally</em>.
For instance, the legal system may notice shifts in societal values based on case outcomes or public reactions and eventually adjust interpretations of the law, but it does so in a way that remains consistent with its legal/illegal <em>binary code</em>.
Through <em>repeated feedback</em>, systems can detect trends in their environment (e.g., shifts in public opinion or technological advances) and adapt their operations in response, but only when those trends become relevant within the system’s own code.
Despite the use of <em>feedback loops</em>, systems remain <em>operationally closed</em>.
Feedback is processed in terms of the system’s unique code, meaning that only information relevant to that code is taken in. 
For instance, if the economic system receives feedback from the political system, it only integrates that feedback if it pertains to the payment/non-payment distinctions. 
This allows each system to interact with its environment while preserving its autonomy and self-referential logic.
Feedback loops help systems manage the complexity of their environments by selectively processing information.
Each system filters out what is irrelevant to its operations, thereby creating <em>blind spots</em>.</p>

<p>Luhmann’s concept of <em>structural coupling</em> (borrowed from <a class="citation" href="#maturana:1987">(Maturana &amp; Varela, 1987)</a>) describes how different social systems develop stable, interdependent relationships with their environment or other systems without losing their <em>operational closure</em> or <em>autonomy</em>.
<em>Structural coupling</em> is essential for maintaining a productive relationship with various other systems.
For example, science and politics are structurally coupled when scientific research informs policy decisions, while political priorities influence the direction and funding of scientific research. 
Despite this interaction, both systems remain <em>operationally closed</em>: science focuses on truth/falsehood, and politics operates on power/non-power.
Another example: the economic system and the political system may engage in <em>structural coupling</em>, where economic data (like inflation) affects political decisions (like interest rate changes). 
<em>Feedback loops</em> are crucial for <em>structural coupling</em> since they allow each system to remain sensitive to changes in the other system without changing its fundamental operations or logic.</p>

<p>The last term we have to discuss is the term maybe most important, and that is <em>communication</em>.
As a computer scientist I understand communcation using Shannon’s framework <a class="citation" href="#shannon:1948">(Shannon, 1948)</a>, that is, a transfer of information over an error-prone or noisy channel.
But this is not what Luhmann understands as <em>communication</em>.
For him, <em>communication</em> is not merely the transfer of information between individuals (or machines). 
Instead, it is a self-contained social process that occurs within and is produced by social systems, with <strong>individuals seen as part of the environment rather than as agents within the system</strong>.
According to Luhmann, it is a three-part process of three interdependent elements/<strong>selections</strong>:</p>

<ol>
  <li><strong>Information</strong>: The content or ‘<strong>what</strong>’ of communication, which could be new data, knowledge, or ideas relevant to the system.</li>
  <li><strong>Utterance</strong>: The ‘<strong>how</strong>’ of communication, which includes the form, manner, or medium through which information is expressed. This could be spoken language, writing, or nonverbal cues, depending on the medium and context. E.g. a payment is a communication.</li>
  <li><strong>Understanding</strong>: The receiver’s interpretation of both the information and the utterance. Understanding is crucial, as it determines whether and how the communication is taken up within the system. It is the <strong>differentiation between information and utterance</strong>.</li>
</ol>

<p>For Luhmann, communication only ‘happens’ if all three elements are present. It’s not just about sending (selected) information; it’s about how information is expressed and then understood within a particular context.
<em>Communication</em> is not generated by individuals but by the system and is itself autopoietic, meaning it is self-producing and self-sustaining.
<em>Communication</em> is inherently selective; it involves making choices about what information to include, how to present it, and how to interpret it. 
This selectivity creates <em>blind spots</em>, as each communication inherently excludes other possible meanings.
It is based on the concept of <em>double contingency</em>—the idea that each party in a communication anticipates and adjusts to the other’s responses.
Social systems manage this <em>contingency</em> through established expectations.
For example, in the legal system, there is an expectation that communications follow the legal/illegal code, which guides interactions between lawyers, judges, and citizens and maintains coherence in legal decisions.
Or take the education system.
If I start singing in my lecture students would be quite confused.
Since in Luhmann’s framework every system operates by its unique logic, <em>communication</em> can <strong>not</strong> be about transferring <strong>objective</strong> information; it is about processing meaning/sense which depends on the ‘processor’, i.e. the system. 
Each <em>communication</em> within a system adds to the meaning that the system produces.
For instance, in the scientific system, each new theory or finding creates meaning within the context of the truth/falsehood code and is interpreted within that framework—it can change <em>the reality of science</em> as a whole.</p>

<blockquote>
  <p>Systems exist and they continuously build their own specific reality, shaped by the types of meaning they process through communication.</p>
</blockquote>

<p>Consequently, if we take Luhmann serious, we arrive at the revelation that there is not one ‘really real and objective reality’ but that there is a <strong>plurality of realities</strong>;
that there is not one controlling system that steers all the others but that there is <em>anarchy</em> in society;
that we are not in control but that society is <strong>out of control</strong>;
that there is not one objective truth but systemic interpretations;
and maybe most importantly: that there is no outside of society, no view at the whole because <strong>any indication of something requires a distinction from something else</strong>.</p>

<h2 id="part-i-a-crisis-of-non-communication">Part I: A Crisis of Non-Communication</h2>

<p>One area where Luhmann’s theory seems to make unsettlingly accurate sense is the <em>climate crisis</em>—that is, the accelerating destabilization of the climate system.
We see how Luhmann’s framework reveals the challenges of a complex, multi-systemic problem. 
Each social system—politics, economy, science, and media—observes the climate crisis from within its own operations and distinctions.
Science operates on a truth/falsehood basis, producing reports on climate change’s reality and projections, while politics, operating on power/non-power, assesses climate issues according to political priorities, public opinion, and election cycles. 
The economic system, structured by payment/non-payment, may respond to climate science only insofar as it impacts financial markets, investments, or regulatory demands.</p>

<p>The <em>climate crisis</em>, however, does not ‘belong’ to any one system.
Instead it spans across systems but is refracted through each one’s unique code.
No single system can comprehensively address the crisis because its complexity exceeds the logic of any one system’s operations.
The climate crisis is precisely so problematic for the communication network we call society because each system deals with its environment by reducing the environment’s complexity.
Furthermore, there is essentailly no <em>climate communcation</em> ‘happening’, because (to the best of my knowledge) there is no <em>social system</em> that operates on a code that leads to the observation of the climate or the earth’s ecosystem.
One might step in and argue that science certainly observes the climate but that is not really the case if we use Luhmann’s definiton of observation and communication.
Science, despite its close engagement with ecological and climate issues, operates on a fundamentally different basis than a (hypothetical) climate-focused system.
While many scientists care very much about the climate and the survival of human beings, science operates under the truth/falshood distinction.
Its observation of the climate does not directly lead to climate or political activism, or an economic transformation but to more truth/falshood distinction;
to research and funding opportunities and the building up of reputation.
In fact, some scientist such as Ulf Büntgen are concerned about scholars who are at the same time activists because they might damage the operation of science <a class="citation" href="#buentgen:2024">(Büntgen, 2024)</a>, others argue against these worries <a class="citation" href="#eck:2024">(van Eck et al., 2024)</a>.
From a system’s viewpoint, we should not attribute to much influence to the individual since, again, there is no individual within society.
Science can ‘use’ the respective psychic system as well as activism.
At the same time, while activsim cannot steer science, science digest/observes activism by its own operations to preserve its operating.</p>

<p>Similarly, the political system interprets the <em>climate crisis</em> through its own operational code of power/non-power, using it as an opportunity to gain influence and public support. 
For example, in Germany, the Green Party views the climate crisis as a platform to expand its political reach.
They advocate for environmental policies that resonate with their voter base.
However, they (as a system, not as individuals) do this to gain power and not to solve the climate crisis.
Other parties may leverage the crisis in the opposite direction, appealing to constituents who prioritize economic stability over environmental reform.
Yet, even if the Green Party succeeds in passing policies aimed at accelerating economic transformation, it cannot directly control whether the economic system will fully implement these changes. 
This limitation arises because each system—politics and the economy—operates autonomously according to its own logic.
For politics to dictate economic outcomes would imply that the political system could override the economy’s fundamental payment/non-payment distinction, which, according to Luhmann, is structurally impossible.
Note that this is not a critique of the members of the Green Party which might very much care and belief in the values they display—they might be honestly invested in how they are being observed.
I do not consider individuals or people but systems.</p>

<p>The climate crisis illustrates how <em>structurally coupled</em> systems face limits in their coordination.
Each system’s response to climate issues is conditioned by its own operations, preventing unified action despite the existential threat posed by the destabilizing climate.
As Luhmann’s theory reveals, social systems are inherently self-referential and cannot simply ‘combine’ their functions. 
Thus, the climate crisis may persist without cohesive action precisely because <strong>no system is structured to address an issue that transcends its operational boundaries</strong>.
In fact, if systems cross boundaries, we call it corruption, for example when the economy system pays for a football goal or for a law or if a scientist makes false claims to strengthen a certain political ideology.</p>

<p>This fragmentation of responsibility means that no single system is inherently designed to address global, cross-cutting issues like the climate crisis.
Each system approaches climate issues only insofar as they relate to its own logic:</p>

<ul>
  <li><strong>Science</strong> seeks truth and thus investigates the mechanisms, causes, and projected impacts of climate change.</li>
  <li><strong>Politics</strong> operates on power/non-power, framing climate policies in terms of public support, regulatory reach, and political gain or loss.</li>
  <li><strong>Economy</strong> focuses on payment/non-payment, assessing climate initiatives based on profitability and market viability.</li>
  <li><strong>Media</strong> works with information/non-information, spotlighting climate issues based on newsworthiness rather than scientific rigor or policy relevance.</li>
</ul>

<p>Systems only engage with climate issues when these issues align with their internal priorities.
Science can produce overwhelming evidence of climate risks, but if political decisions are driven by short-term voter approval, economic costs, or geopolitical interests, the full implications of scientific knowledge may not translate into concrete action.
Political decisions are often informed by scientific findings but are ultimately filtered through the political logic of power/non-power.
Economies can adopt sustainable practices, but often only when such practices promise financial returns.
This misalignment is evident in the delays or dilution of climate policies, where political and economic interests often override scientific findings.</p>

<p>Luhmann’s theory suggests that issues of this magnitude may require a dedicated system with its own <em>binary code</em>—such as ‘sustainable/unsustainable’ or ‘ecologically balanced/unbalanced’—to assess and act on climate issues directly. 
Social systems are so immensly effective (with respect to their function) because they reduce complexity by operating under a rather simple binary code.
A new <em>sustainable system</em> could facilitate large enough irritations that resonate within other systems such that standards, goals, and measures across all systems are implemented (by themselves) in such a way that align with ecological stability and sustainability.
Although hypothetical, this system would prioritize ecological concerns by constructing ‘sustainable communication’.
Of course, the problem is: How can such a communication be possible without interferring in, e.g. economic communication?</p>

<p>Howsoever, the time is up and I have almost no hope that such a system will suddenly emerge.
If it does it should happen via the seperation by operational closure similar to how, for example, art was able to separate itself from religion.
The only other option, using Luhmann’s framework, is to try to align all the systems in such a way that if they operate according to their logic, they also operate (at least close) to the logic of such a hypothetical <em>sustainable system</em>.
Caring about the earth’s ecosystem and the climate has to be financially profitable;
it has to lead to power for politicians;
it has to lead to funding and furhter research in science;
and it has to be a spectacle for the media to cover;</p>

<p>Naturally, we might see it as hypocritical when a company adopts sustainable practices primarily to increase profits rather than out of genuine concern for the environment.
We are back to the <em>authenticity problem</em> mentioned earlier.
And, indeed, companies may choose to appear sustainable rather than enact substantive changes if it proves more profitable.
This dynamic holds for other issues as well, such as diversity and inclusion—and deep inside we all know it.
However, from a systems theory perspective, it may be more productive to move beyond <em>moral judgments</em> about individuals, pointing to them as being hypocritical, virtuous, or evil and instead to focus on  <em>systemic realities</em>.</p>

<p>In the economic system, sustainability, diversity, and other social values are interpreted through the code of payment/non-payment; the system evaluates decisions based on profitability rather than intrinsic ethical value.
We may not like it but that’s the <em>reality of the economic system</em>.
For example, instead of hoping for an <em>humanistic turn</em> of the economic system, it might be more effective to achieve social progress for disabled people by letting the economic system ‘know’ how these people are financially important.
While individuals within a company might personally care deeply about these issues, this personal commitment does not translate into the company’s operations, as employees are part of the system’s environment, not its core functions.
People care, systems observe and operate on their own terms.</p>

<p>In this light, <em>truthfully pretending</em>—adopting sustainability practices for economic gains—can still yield positive outcomes.
From the perspective of systems theory, the motivations behind these actions matter less than the fact that they result in more sustainable practices, which in itself contributes to broader societal goals.
This does not imply that public outrage about the state of affairs is misguided.
On the contrary, if outrage <em>irritates</em> systems in ways that make sustainable practices more profitable for companies, it can drive meaningful change.
However, outrage can also produce unintended effects.
For instance, the media, which constructs a shared reference reality that shapes public discourse, may find it more sensational to focus on the ‘unlawfulness’ of protesters.
This framing can prompt the political system to respond by mobilizing power against the protests, potentially reinforcing the very practices that climate advocates aim to change.
Especially when a system’s ability to continue its autopoietic operations is threatened, a strong reaction can be expected.</p>

<p>These nonlinear and indirect effects, often amplified through feedback loops, illustrate the unpredictable and uncontrollable nature of systemic interactions. 
In Luhmann’s terms, feedback loops create complex dynamics within and between systems, making it difficult to foresee or control the outcomes of public reactions, even when intentions are clear.</p>

<h2 id="part-ii-a-crisis-of-over-communication">Part II: A Crisis of Over-Communication</h2>

<p>Luhmann does not oppose elections, but he challenges the common assumption that they express <em>the will of the people</em>. 
This skepticism follows directly from his understanding of society as an uncontrollable network of interdependent social systems, each operating according to its own logic. 
Society, in Luhmann’s view, evolves organically—like a self-reproducing system—and cannot be directly steered or micromanaged by politics.</p>

<p>Politics, in this context, plays a specific role: it makes collectively binding decisions.
Yet these decisions must then be processed and implemented by other systems.
For instance, the legal system creates laws to enforce political decisions, while the economy and even religion may shape how these decisions are interpreted and realized in practice.</p>

<p>Consider the example of childbirth. 
How do different social systems contribute to this event? 
The political system might legislate that abortion is legal.
Religious beliefs may influence whether a person opts for or against it. 
Socio-economic conditions affect whether one can afford to raise a child, and the health system plays a crucial role in medical support. 
Political decisions matter, but they exist within a web of other systems, each with its own influence, sometimes more decisive than politics itself.</p>

<p>The popular narrative around elections is that they make ‘the people’ the foundation of all political power. 
After the election, politicians—servants of the people—are supposed to put the will of the people into action. 
However, according to Luhmann, the idea that the people are the source of all power in a liberal democracy is a myth. 
Instead, the people function more as an audience, much like in a talent show, where they get to elect a winner at specific, pre-determined moments, but do not control the larger system. 
The ‘show’ of politics is a much larger, self-sustaining system, where politicians, the state, and the voters all play their roles and influence one another, but the system ultimately serves itself.</p>

<p>In democratic politics, the state, politicians, and voters are mutually interdependent, each contributing to the reproduction of the political system. 
Elections, therefore, are <em>symbolic procedures</em> in Luhmann’s view.
They confer legitimacy on the political system by symbolically invoking <em>the will of the people</em>, but in reality, such a unified will does not exist.
For example, many people abstain from voting, and a significant portion of the population may not be eligible to vote at all.
Moreover, election outcomes are shaped by arbitrary rules—they are contigent.
In the U.S., for example, the popular vote does not directly determine the outcome; instead, the electoral college decides the presidency, often making a few swing states the key deciders. 
In Germany, government coalitions are typically formed after elections, yet no single voter casts a ballot for the specific coalition that ends up governing.
These complexities highlight how elections, while significant, are far from a straightforward expression of a unified popular will.</p>

<blockquote>
  <p>How did we vote? But did we really vote, or did the people just roll the dice? […] What individuals actually think, if anything at all, when they mark ballots, remains unknown. This alone suffice not to […] conceive of public opinion as the general expression of the opinions of individuals. – <a class="citation" href="#luhmann:2002">(Luhmann, 2002)</a></p>
</blockquote>

<p>Elections seem to achieve the impossible: merging the diverse, individual wills of the people into a singular, cohesive ‘general will’.
For Luhmann, this process is almost magical, as it provides the symbolic foundation upon which liberal democracy rests.
Elections create the <em>illusion of unity and consensus</em>, giving legitimacy to political decisions that, in reality, are based on a highly fragmented and complex societal landscape.
However, as previously mentioned, Luhmann has no problem with this illusion. 
In his view, <em>the miracle of democratic elections</em> is perfectly acceptable—provided it functions effectively for all involved and helps stabilize the political system.</p>

<p>In fact, the <em>symbolic power of elections</em> is essential to maintaining social order, as it grants politics a legitimate mandate without requiring every individual’s direct influence on policy decisions. 
This symbolic function of elections allows the political system to operate independently, without collapsing under the weight of countless individual preferences.
Elections serve to renew the legitimacy of the political system periodically, preventing it from stagnating, while also setting boundaries within which political decisions are accepted, even by those who disagree with the outcomes.</p>

<p>For Luhmann, it is less important that elections genuinely express a collective will, which he considers a fiction, and more important that they fulfill their function: they create a momentary sense of unity and provide a mechanism for the orderly transition of power seemingly melting the individual will of ‘the people’ into the general will. 
As long as elections maintain public confidence in the political process and prevent <em>systemic breakdown</em>, they serve their purpose, not by conveying truth but by ensuring continuity. 
In this sense, elections are not a search for truth but a pragmatic solution to the challenge of political legitimacy in a complex, functionally differentiated society.</p>

<p>Viewing the election as a performance—or as Baudrillard might call it, <em>hyperreality</em>—feels particularly fitting for the spectacle that Americans witness during the election weeks. 
Do these debates between Trump and Biden, or Trump and Harris, genuinely convey new insights or substantive ‘truths’? 
Or are Americans, in many ways, simply the audience to a grand show, swept up in the drama, spectacle, and narrative arcs that these events offer?</p>

<p>There is, however, something distinct and potentially perilous about the American context.
In the U.S., the problem for the political system seems to lie in the crumbling of the illusion of unity and consensus.
The illusion is increasingly undermined by escalating economic and social inequalities.
As living conditions deteriorate for many, regardless of who they vote for, it becomes glaringly apparent that there is no singular ‘general will’ guiding political outcomes. 
The democratic promise that elections merge the will of the people into collective decisions feels hollow when so many are left feeling unrepresented and disillusioned.</p>

<p>This breakdown makes it clear that those voting for Donald Trump are not simply misguided or irrational. 
Many voters feel disconnected from a political system they perceive as indifferent to their realities and struggles.
As one interviewee put it:</p>

<blockquote>
  <p>I agree that Trump is from the billionaire class and that’s all he’s going to work for.
It basically comes down to the lesser of the two evils right now.
I think about the four years he was president.
In my opinion he’s the world’s best crime boss.
And then you see people posting ‘in the arms of Jesus’ like ‘I was persecuted too’, and I think what a bunch of bullcrap.
The system is so captured that you need a crime boss to get out of it.
If it wasn’t Trump I would love to have a working familiy candidate who stands up for the little guy.
The middle class, from day one of this United States, has built the United States, and we are the ones that always get shit on.</p>
</blockquote>

<p>But there’s a second factor at play: a candidate who is so absurd, so obscene, that he disrupts the expected script of political ‘producers’.
Because Trump is taken seriously he (unconsciously) discloses the reality of elections and the political system.
From a systems theory perspective, a showman does what one shouldn’t do: making the big show, the big stage visible and dismanteling one myth but also replacing it with another, far more dangerous one: <em>the deep state</em>.
While the <em>illusion of unity and consensus</em> gives the system stability, Trump’s <em>deep state myth</em> does the opposite and that is the reason why he is dangerous for the political system and probably for the functional differentiated society.</p>

<p>As the interviewee described, Trump is perceived as a <em>red button</em>—a tool voters can press to create enough disturbance within the system that it is forced to respond.
Perceived <em>as the world’s best crime boss</em> any accusation or scandal only feeds into Trump’s persona. 
Trump clearly wants to cross systems’ boundaries.
The interviewee’s perception might not be far from the truth, howerver, if such an irritation is desirable is very questionable.
The level of desperation in the U.S. has become so acute that many are willing to risk an extreme disruption, effectively pushing for an ‘over-irritation’ of the system, hoping it will provoke meaningful change or even a systemic collapse.</p>

<p>In this way, Trump embodies what Luhmann might call an agent of second-order observation—someone who leverages his own media persona to observe and exploit the expectations of the political system, creating feedback loops that intensify rather than stabilize.
While Luhmann argued that elections are primarily a symbolic show, the outcomes of this particular show may indeed carry existential significance.
The stakes are high, and the effects of a Trump victory or defeat could trigger unpredictable reactions. 
What we are witnessing is both dangerous and volatile, and it exemplifies how systemic irritations—if strong enough—can shake the foundations of even the most stable-seeming structures.</p>

<h2 id="criticism-of-luhmanns-theory">Criticism of Luhmann’s Theory</h2>

<p>Luhmann’s theory is not immune to criticism, and, by its own logic, it necessarily contains blind spots. 
One of the most common criticisms is that his theory removes individuals from the core of social analysis.
By focusing on self-referential systems rather than human actors, Luhmann places individuals in the environment of society, not within it.
Human agency, emotions, and individual motivations are neglected and the role of intentional human actions and collective decision-making in shaping societal evolution is minimized.</p>

<p>In addition, its theory lacks a normative direction or ethical foundation.
By avoiding moral or ethical judgments, Luhmann’s systems theory does not offer guidance on what should be done, especially concerning social justice, inequality, or human rights.
His theory seems to be indifferent to power imbalances and fails to address issues of accountability and responsibility within social systems.</p>

<p>Also Luhmann’s concept of <em>operational closure</em> has been criticized for overemphasizing system autonomy.
According to those critics, his perspective ignores the deep interdependencies and interconnectedness of social, economic, and political systems.
They contend that while systems may have unique operations, they are often influenced by each other in ways that Luhmann’s model underestimates.
Marxist thinkers argue that although Luhmann includes the economic system as one of society’s core functional systems, he does not give sufficient attention to the economic forces shaping society, especially those tied to capitalism.
By focusing primarily on the communication logic of payment/non-payment, Luhmann’s theory fails to address the structural inequalities and exploitative dynamics inherent in modern economic systems.</p>

<p>Because Luhmann’s theory emphasizes the self-reproduction of systems, it suggests that systems are largely resistant to intentional change from within.
This perspective might downplay the role of social movements, activism, and democratic engagement as forces that can drive systemic transformation.
Futhermore, by treating power as a form of communication (very different from Foucault’s approach) within the political system rather than a force that operates across systems, critics argue that his approach obscures how power dynamics influence interactions between systems and shape societal outcomes.</p>

<p>My current opinion is that Luhmann’s theory is excellent for the sense-making of society.
It can even clear the fog for making better intentional decisions but it can not provide us with suggestions of what we should do.
But if people want to change oppressive systems, they should know and understand their ‘adversary’.</p>

<h2 id="treatment-of-an-illness">Treatment of an Illness</h2>

<p>In their book <em>The Tree of Knowledge</em> <a class="citation" href="#maturana:1987">(Maturana &amp; Varela, 1987)</a> Maturana and Varela very briefly discuss society.
They think that, unlike cells serving the whole organism, a functioning society should prioritize the needs and well-being of the individual—the orientation should be reversed.
I totally agree.
But can we bring their humanistic viewpoint in line with Luhmann’s <em>anti-humanistic</em> theory?
That would be nice but is probably not in the spirit of the author.
While both views address the organization of society, Maturana and Varela focus on an ideal in which society exists for the benefit of individuals, whereas Luhmann’s theory suggests that society, as a system, operates rather independently of individual well-being.
We could evaluate societies (across time and space) according to the degree they served individual well-being and then learn from the lessons to irritate our society in such a way that it becomes more <em>functional</em>, in Maturana’s and Varela’s sense of the word.
However, operationalizing this perspective would face challenges within Luhmann’s framework.</p>

<p>As a scientific theory communicated by the science system, it is a theory of society that society produced about itself—a quintessential example of second-order observation.
The individual—the person, psychic system, and living body—we identify as Luhmann existed only as part of the environment of the social systems he studied.
According to his own theory, his view, as every view, cannot be the objectively correct description of society.
Therefore, I think, we should not take his ideas as absolutes but as irritations and sense-making foundation.</p>

<p>If the individual is not the center of social structures and dynamics, the impossibility of controlled action might give some relief on an individual/personal level.
It emphasizes that individual human beings are observers of a greater force acting on them and since society is out of control, we are also out of control.
We can observe society critically from an <em>ironic distance</em>, being <em>carefree</em> without being <em>careless</em> (as individuals) and without being fully subsumed by it, a perspective that may be worth remembering—illuminating neither hope nor fear.
We can look at us and our fellow human beings as beings thrown into the fabric of society.
By pointing to the limits of the individual we can rediscovering a sense of innocence and grace of the human being.</p>

<p>One of Luhmann’s most influential critics, Jürgen Habermas, famously remarked of Luhmann’s work:</p>

<blockquote>
  <p>It’s all wrong, but of high quality.</p>
</blockquote>

<p>Habermas labeled Luhmann’s theory <em>metabiological</em>, drawing a comparison to metaphysics and suggesting it extends beyond empirical sociology into abstract structures that make it detached from human agency.
This comparison is spot on: Luhmann’s theory views society not as a product of individuals’ intentions but as an autonomous, complex system very similar to an organism.
From this perspective, if we are to learn anything valuable from Luhmann, it might be that we should avoid treating society—and, by extension, the climate crisis—either as an engineering problem, i.e. as something to be ‘solved’ with a clear-cut plan or a knowledge problem, i.e. something that can be solved if deniers just accept ‘the Truth’.
Instead, we might approach it more like a chronic condition or complex illness, where we, as doctors, explore, probe, and apply potential treatments without assuming a one-size-fits-all solution. 
This approach requires continuous adjustment, sensitivity to feedback, and a readiness to adapt to the unexpected outcomes of our actions, reflecting the complex, interdependent nature of the systems we inhabit.
But, of course, <strong>we are not in control</strong> and there is no single medical doctor drafting the medication plan.</p>

<h2 id="references">References</h2>

<ol class="bibliography"><li><span id="Brundiers2020">Brundiers, K., Barth, M., Cebrián, G., Cohen, M., Diaz, L., Doucette-Remington, S., Dripps, W., Habron, G., Harré, N., Jarchow, M., Losch, K., Michel, J., Mochizuki, Y., Rieckmann, M., Parnell, R., Walker, P., &amp; Zint, M. (2020). Key competencies in sustainability in higher education—toward an agreed-upon reference framework. <i>Sustainability Science</i>, <i>16</i>(1), 13–29. https://doi.org/10.1007/s11625-020-00838-2</span></li>
<li><span id="luhmann:1992">Luhmann, N. (1992). <i>Die Wissenschaft der Gesellschaft</i> (p. 732). Suhrkamp.</span></li>
<li><span id="luhmann:1994">Luhmann, N. (1994). <i>Die Wirtschft der Gesellschaft</i> (p. 356). Suhrkamp.</span></li>
<li><span id="luhmann:1997">Luhmann, N. (1997). <i>Die Kunst der Gesellschaft</i> (p. 517). Suhrkamp.</span></li>
<li><span id="luhmann:2002">Luhmann, N. (2002). <i>Die Politik der Gesellschaft</i> (p. 444). Suhrkamp.</span></li>
<li><span id="luhmann:1998">Luhmann, N. (1998). <i>Die Gesellschaft der Gesellschaft</i> (p. 1164). Suhrkamp.</span></li>
<li><span id="moeller:2011">Möller, H.-G. (2011). <i>The Radical Luhmann</i> (p. 184). Columbia University Press.</span></li>
<li><span id="meadows:2008">Meadows, D. H. (2008). <i>Thinking in Systems: A Primer</i> (p. 240). Chelsea Green Publishing.</span></li>
<li><span id="maturana:1987">Maturana, H. R., &amp; Varela, F. J. (1987). <i>The Tree of Knowledge</i>. Shambhala.</span></li>
<li><span id="brown:1969">Spencer-Brown, G. (1969). <i>Laws of Form</i>. London: Allen and Unwin.</span></li>
<li><span id="foerster:2003">von Foerster, H. (2003). Cybernetics of Cybernetics. In <i>Understanding understanding: Essays on cybernetics and cognition</i> (pp. 283–286). Springer New York. https://doi.org/10.1007/0-387-21722-3_13</span></li>
<li><span id="moeller:2021">Möller, H.-G., &amp; D’Ambrosio, P. J. (2021). <i>You and Your Profile: Identity After Authenticity</i>. Columbia University Press.</span></li>
<li><span id="shannon:1948">Shannon, C. E. (1948). A mathematical theory of communication. <i>Bell Syst. Tech. J.</i>, <i>27</i>(3), 379–423.</span></li>
<li><span id="buentgen:2024">Büntgen, U. (2024). The importance of distinguishing climate science from climate activism. <i>Npj Climate Action</i>, <i>36</i>(3), 2731–9814. https://doi.org/10.1038/s44168-024-00126-0</span></li>
<li><span id="eck:2024">van Eck, C. W., Messling, L., &amp; Hayhoe, K. (2024). Challenging the neutrality myth in climate science and activism. <i>Npj Climate Action</i>, <i>81</i>(3), 2731–9814. https://doi.org/10.1038/s44168-024-00171-9</span></li></ol>]]></content><author><name>Benedikt Zönnchen</name><email>benedikt.zoennchen@web.de</email></author><category term="Sustainability" /><category term="Social Systems Theory" /><category term="Politics" /><summary type="html"><![CDATA[I have to admit, I’m somewhat addicted to thinking about social systems theory, particularly the version developed by the German sociologist Niklas Luhmann (1927–1998). But this addiction isn’t driven by pure fascination—it’s more of a love-hate relationship. On one hand, seeing the world through a Luhmannian lens doesn’t make it more just or fantastic, but it does make it comprehensible. The chaotic state of global affairs—the craziness of U.S. elections, the brutal devastation in the Middle East, the war in Europe, and the looming tensions brought by the climate crisis—appears senseless at times, even apocalyptic. And yet, systems theory offers a framework that, paradoxically, brings coherence to this apparent madness, showing us how these outcomes emerge from the logic of functionally differentiated systems.]]></summary></entry><entry><title type="html">Musical Interrogation IV - Transformer</title><link href="https://bzoennchen.github.io/Pages/2024/02/03/musical-interrogation-IV.html" rel="alternate" type="text/html" title="Musical Interrogation IV - Transformer" /><published>2024-02-03T00:00:00+01:00</published><updated>2024-02-03T00:00:00+01:00</updated><id>https://bzoennchen.github.io/Pages/2024/02/03/musical-interrogation-IV</id><content type="html" xml:base="https://bzoennchen.github.io/Pages/2024/02/03/musical-interrogation-IV.html"><![CDATA[<blockquote>
  <p>Recurrent models trained in practice are effectively feed-forward.
This could happen either because truncated backpropagation through time cannot learn patterns significantly longer than k steps, or, more provocatively, because models trainable by gradient descent cannot have long-term memory. – John Miller</p>
</blockquote>

<p>This time in the series we use the most famous model architecture for generative purposes: the <strong>transformer</strong> <a class="citation" href="#vaswani:2017">(Vaswani et al., 2017)</a>.
Transformers were initially targeted at natural language processing (NLP) problems, where the network input is a series of high-dimensional embeddings representing words or word fragments.
Transformers were introduced in 2017 by the authors of <em>Attention Is All You Need</em> <a class="citation" href="#vaswani:2017">(Vaswani et al., 2017)</a> to basically replace <em>recurrency</em> with <em>attention</em>.</p>

<p>One of the problems with RNNs is that they can forget information that is further back in the sequence.
While more sophisticated architectures, such as LSTMs <a class="citation" href="#hochreiter:1997">(Hochreiter &amp; Schmidhuber, 1997)</a> and <em>gated recurrent units</em> (GRUs) <a class="citation" href="#chung2014">(Chung et al., 2014)</a> partially addressed this problem, they still struggle with long term dependencies.
The idea that intermediate representations in the RNN should be exploited to produce the output led to the <em>attention mechanism</em> <a class="citation" href="#bahdanau:2014">(Bahdanau et al., 2014)</a> and, in the end, to the transformer architecture <a class="citation" href="#vaswani:2017">(Vaswani et al., 2017)</a>.
The transformer avoids the problem of vanishing or exploding gradients by avoiding recurrency, that is, by utilizing the whole sequence in parallel.</p>

<p>Today, all successful large language models (LLMs) utilize the transformer architecture.
It brought us ChatGPT (based on GPT-3 <a class="citation" href="#brown:2020">(Brown et al., 2020)</a> and GPT-4 <a class="citation" href="#openai:2023">(OpenAI, 2023; Bubeck et al., 2023)</a>), LLaMA <a class="citation" href="#touvron:2023">(Touvron et al., 2023)</a>, LLaMA 2 <a class="citation" href="#touvron:2023b">(Touvron et al., 2023)</a>, BERT <a class="citation" href="#devlin:2019">(Devlin et al., 2019)</a> and many fine-tuned derivatives such as Codex <a class="citation" href="#chen:2021">(Chen et al., 2021)</a>.
In the domain of symbolic music, transformers were also employed.
Examples are the Music Transformer <a class="citation" href="#huang:2018">(Huang et al., 2018)</a>, the Pop Music Transformer <a class="citation" href="#huang:2020">(Huang &amp; Yang, 2020)</a>, multi-track music generation <a class="citation" href="#ens:2020">(Ens &amp; Pasquier, 2020)</a>, piano inpainting <a class="citation" href="#hadjeres:2021">(Hadjeres &amp; Crestel, 2021)</a>, Theme Transformer <a class="citation" href="#shih:2022">(Shih et al., 2022)</a> and more.
Furthermore there are transformers, such as MusicGen <a class="citation" href="#copet:2023">(Copet et al., 2023)</a> that generate audio output directly.</p>

<p>While there is an intuitive explanation of the attention mechanism, it is still unclear why exactly the transformer is so effective—there is no rigorous mathematical proof.
It is well-known how their components work and what mathematical operations are performed, but it is very hard to interpret the seemingly emerging power when all the small parts work together.
One source of their effectiveness is that they relate tokens to other tokens more directly (without a hidden state which washes away the information) and the independence of multiple execution paths make them especially suitable for the exploitation of multicore processors such as GPUs and TPUs.
However, looking at the whole sequence at once comes at a cost: computation and memory complexity!
Therefore, to train transformers you require GPUs with a lot of memory which is concerning for artists who might want to utilize transformers independently from proprietary cloud services.</p>

<p>Original transformers were introduced for natural language processing.
However, since language datasets share some of the characteristics of musical notations, transformers achieve good results in learning the structure of symbolic pieces.
In music as well as in language the number of input variables can be very large, and the statistics are similar at every position; it’s not sensible to re-learn the meaning of the word <em>dog</em> at every possible position in a body of text.
Language datasets and music datasets have the complication that their sequences vary in length.</p>

<p>However, we also have to remember that there are also differences between the two domains.
The alphabet of musical notations has more than 26 symbols and there is a strong relation between certain symbols.
For example, there is a strong relation between the C’s of each octave or a whole and half note in the same pitch class.
Furthermore, shifting all the letters in a text changes the meaning of that text dramatically while in the case of music this is most often not the case.</p>

<h2 id="attention-in-encoder-decoder-rnns">Attention in Encoder-Decoder RNNs</h2>

<p>What is the idea behind the attention mechanism?
Attention was introduced to bidirectional recurrent neural networks (RNNs) in 2014 <a class="citation" href="#bahdanau:2014">(Bahdanau et al., 2014)</a> for language translation, that is, for an <em>encoder-decoder architecture</em>.
In this scenario we want to translate a sentence from e.g. English into e.g. German.
The attention mechanism helps the decoder part of the RNN to focus on different parts of the encoder’s output (representations of the English words) differently.
Therefore, it helps to preserve long term dependencies.</p>

<p>The <strong>encoder’s</strong> input is a sequence of tokens, let’s say words for simplicity, i.e. a sequence</p>

\[\mathbf{x}_{0}, \ldots \mathbf{x}_{n-1}.\]

<p>For each word \(\mathbf{x}_{i}\) it computes some output \(\mathbf{y}_{i}\).
The assumption is that the probability for \(\mathbf{x}_{i}\) depends on \(\mathbf{x}_{j}\).
Since we have the whole sentence given, we can use a <em>bidirectional RNN</em> and look into the future.
Thus, with respect to dependency, the probability for token \(i\) can depend on the probability for token \(j\) and vice versa.</p>

<p>The <strong>decoder’s</strong> input is the <strong>whole</strong> sequence computed by the <strong>encoder</strong> but as a weighted sum.
The output is a sequence of German words, let’s say</p>

\[\mathbf{y'}_{0}, \ldots \mathbf{y'}_{n-1}.\]

<p>This time however, the <strong>decoder</strong> RNN is unidirectional.
It can not look into the future and computes each German word strictly from left to right.
To compute the weights or attention scores of \(\mathbf{y'}_{i}\), an <strong>alignment model</strong> receives the hidden state \(\mathbf{h'}_{i-1}\) and the outputs of the encoder as input.
First a simple dot product is computed:</p>

\[e_{i,j} = \mathbf{h'}_{i}^\top \mathbf{y}_{j} \quad \text{ for } j = 0, \ldots, i-1.\]

<p>Later it was suggested to use an additional linear transformation on the output:</p>

\[e_{i,j} = \mathbf{h'}_{i}^\top (\mathbf{W}\mathbf{y}_{j}) \quad \text{ for } j = 0, \ldots, i-1.\]

<p>where \(\mathbf{W}\) is learned.
All these scores are normalized by the softmax function giving us \(n\) weights:</p>

\[\alpha_{i,j} = \frac{\exp\left( e_{i,j} \right)}{\sum\limits_{k=0}^{n-1} \exp\left( e_{i, k}\right)}.\]

<p>Then the <strong>decoder’s</strong> ‘real’ input is computed by a weighted sum of the <strong>encoder’s</strong> output:</p>

\[\hat{\mathbf{h}}_i = \sum_j \alpha_{i,j} \mathbf{y}_j.\]

<p>The weights determine how “strong” the information of the decoder’s input will be utilized, i.e., how much attention is spent on each previous output of the model.</p>

<div><img style="display:block; margin-left:auto; margin-right:auto; width:90%;" src="/Pages/assets/images/rnn-attention.png" alt="RNN with attention" />
<div style="display: table;margin: 0 auto;">Figure 1: RNN with self-attention.</div>
</div>
<p><br /></p>

<p>This results in a quadratic complexity of \(\mathcal{O}(n^2)\) because for each of the \(n\) tokens, we want to decode, we have \(n\) weights.</p>

<h2 id="the-transformer-architecture">The Transformer Architecture</h2>

<p>The original transformer was introduced for the task of machine translation thus it was an encoder-decoder architecture.
In Fig. 2 you see a slightly modified version where the addition (residual connections) and the layer norm are in front of the attention layer.</p>

<div><img style="display:block; margin-left:auto; margin-right:auto; width:90%;" src="/Pages/assets/images/transformer.png" alt="RNN with attention" />
<div style="display: table;margin: 0 auto;">Figure 2: The slightly modified transformer.</div>
</div>
<p><br /></p>

<p>Let’s consider a scenario in which we have an English sentence that we want to translate into French. 
In this process, the encoder plays a crucial role by transforming the English sentence into a highly compressed and information-rich representation.</p>

<p>Subsequently, the decoder comes into play, generating the French translation word by word. 
It relies on the previously computed French words to predict and produce the next one. 
It’s important to note that the input provided to the decoder is a partial translation, essentially a shifted version of what it is currently working on. 
This is because the decoder should lack the ability to see into the future; it only has access to the portion of the translation it has computed up to that point—otherwise it would cheat while training which would hurt the learning process.</p>

<p>To address this limitation, the decoder employs a masked version of the multi-head attention layer. 
This mechanism ensures that the decoder focuses on the relevant information without peeking ahead.</p>

<p>Furthermore, the utilization of residual connections and layer normalization over the feature dimension within a single sample serves as a valuable tool to combat the issue of vanishing gradients in deep neural networks, ensuring the efficient training and optimization of the translation model.</p>

<p>In our case we do not actually want to translate a sentence but we want to generate musical notes from a sequence of given notes.
Therefore, we have no encoded information and there is no encoding involved.
We only need the decoder part.
Furthermore, I only use one (masked) multi-head attention layer in each block.</p>

<div><img style="display:block; margin-left:auto; margin-right:auto; width:40%;" src="/Pages/assets/images/decoder.png" alt="Decoder-only transformer" />
<div style="display: table;margin: 0 auto;">Figure 3: Our decoder-only transformer.</div>
</div>
<p><br /></p>

<p>Ok, but how does this really work?
What is going on here?
Well, the key to understand transformers is to understand the self-attention mechanism which I try to explain below.</p>

<h2 id="self-attention">Self-Attention</h2>

<p>The idea of the transformer is to just rely on (self-)attention thus remove recurrency.
This means that the model “sees” \(n\) tokens to generate the \((n+1)^\text{th}\) token.
Simple RNNs for predicting the next tokens only see the previous token and the hidden state which represents all the tokens before.
But, as I discussed in previous articles, the information of the hidden state gets washed away over time and without attention there seems to be little control over the importance of certain tokens of the sequence.</p>

<p>The fundamental operation of the transformer, i.e. the attention mechanism, is implemented in its <code class="language-plaintext highlighter-rouge">Head</code>.
Let \(\mathbf{x}_0, \ldots, \mathbf{x}_{n-1} \in \mathbb{R}^{D \times 1}\) be the \(n\) tokens of a sequence.
A standard neural network layer \(f(\cdot)\), takes a \(D \times 1\) input and applies a linear transformation followed by an activation function like a \(\text{ReLU}\):</p>

\[f(\mathbf{x}) = \text{ReLU}\left( \mathbf{W}\mathbf{x} + \mathbf{b }\right),\]

<p>where \(\mathbf{b}\) contains the biases, and \(\mathbf{W}\) contains the weights.</p>

<p>A self-attention \(\mathbf{sa}(\cdot)\) block takes all the \(n\) inputs, each of dimension \(D \times 1\), and returns \(n\) output vectors of the same size.
Note that in our case each input represents a musical note or event.
First, a set of <strong>values</strong> is computed for each input:</p>

\[\mathbf{v}_i = \mathbf{b}_v + \mathbf{W}_v \mathbf{x}_i, \quad \text{ (value)}\]

<p>where \(\mathbf{b}_v \in \mathbb{R}^D \text{ and } \mathbf{W}_v \in \mathbb{R}^{D \times D}\) represent biases and weights, respectively (<strong>for all inputs</strong>). 
The \(j^\text{th}\) output \(\mathbf{sa}_j(\mathbf{x}_0, \ldots, \mathbf{x}_{n-1})\) is a weighted sum of all the values \(\mathbf{v}_i, i = 0, \ldots n-1\) where each weight depends on \(\mathbf{x}_j\):</p>

\[\mathbf{sa}_j(\mathbf{x}_0, \ldots, \mathbf{x}_{n-1}) = \sum_{i=0}^{n-1} \alpha(\mathbf{x}_i, \mathbf{x}_j) \mathbf{v}_i.\]

<p>The scalar weight \(\alpha(\mathbf{x}_i, \mathbf{x}_j)\) is the <strong>attention</strong> that the \(j^\text{th}\) note pays to the note \(\mathbf{x}_i\).
The \(n\) weights \(\alpha(\cdot, \mathbf{x}_j)\) are non-negative and sum to one.
Hence, self-attention can be thought of as <em>routing</em> the values in different proportions to create each output.</p>

<p><br /></p>
<div><img style="display:block; margin-left:auto; margin-right:auto; width:35%;" src="/Pages/assets/images/routing.png" alt="Routing principle" />
<div style="display: table;margin: 0 auto;">Figure 4: Routing principle.</div>
</div>
<p><br /></p>

<p>To compute the attention, we apply two more linear transformations to the inputs:</p>

\[\begin{aligned}
\mathbf{q}_j &amp;= \mathbf{b}_q + \mathbf{W}_q \mathbf{x}_j \quad \text{ (query)}\\
\mathbf{k}_i &amp;= \mathbf{b}_k + \mathbf{W}_k \mathbf{x}_i \quad \text{ (key).}
\end{aligned}\]

<p>The <strong>dot product</strong> of two vectors \(\mathbf{q}_j\), \(\mathbf{k}_i\) is a measurement of their similarity.
The matrices \(\mathbf{W}_q, \mathbf{W}_k\) and the respective bias are learned such that similarity of \(\mathbf{q}_j\), \(\mathbf{k}_i\) can be interpreted as how “important” \(\mathbf{x}_i\) is for \(\mathbf{x}_j\).
Thus, the “magic” happens via a very simple linear transformation and one might ask if this operation is powerful enough to relate <strong>all</strong> words/tokens in a desirable way.
The answer is most certainly “no” thus one adds feed forward layers which introduce non-linearity in between multiple attention layers.</p>

<p>In the special case where both vectors are unit vectors, the dot product is the cosine of the angle between the two.
In general, this relationship is expressed by the following equation:</p>

\[\mathbf{q}_j \circ \mathbf{k}_i = \mathbf{q}_j^\top  \mathbf{k}_i = \Vert \mathbf{q}_j \Vert \cdot \Vert \mathbf{k}_i \Vert \cos(\beta),\]

<p>where \(\beta\) is the angle between the two vectors.
Computing the <em>dot product</em> between queries and keys gives us the similarities we desire.
To normalize, we then pass the result through a <em>softmax</em> function:</p>

\[\alpha(\mathbf{x}_i, \mathbf{x}_j) = \frac{\exp(\mathbf{q}_j^\top \mathbf{k}_i / \sqrt{D_q})}{\sum\limits_{r=0}^{n-1} \exp(\mathbf{q}_j^\top \mathbf{k}_r / \sqrt{D_q})},\]

<p>where \(D_q\) is the dimension of the queries and keys (i.e., the number of rows in \(\mathbf{W}_q\) and \(\mathbf{W}_k\), which must be the same).
You can think of the <em>key</em> as what is offered and the <em>query</em> as what is searched for.
If \(\mathbf{X}\), \(\mathbf{K}\), \(\mathbf{Q}\), and \(\mathbf{V}\) contain all the inputs, keys, queries and values then we can compute the self-attention by</p>

\[\mathbf{Sa}(\mathbf{X}) = \mathbf{V} \cdot \text{Softmax}\left( \frac{\mathbf{K}^\top \mathbf{Q}}{\sqrt{D_q} }\right).\]

<p>The overall computation is illustrated in Figure 5.</p>

<p><br /></p>
<div><img style="display:block; margin-left:auto; margin-right:auto; width:90%;" src="/Pages/assets/images/self-attention.png" alt="Self-attention in matrix-form" />
<div style="display: table;margin: 0 auto;">Figure 5: Self-attention in matrix-form.</div>
</div>
<p><br /></p>

<h2 id="masking-attention-head">Masking Attention Head</h2>

<p>Since our transformer should not look into the future, because when we use it in the prediction mode it also can not look ahead of the token it predicts, we have to mask entries in</p>

\[\text{Softmax}\left(\frac{\mathbf{K}^\top \mathbf{Q}}{\sqrt{D_q}}\right).\]

<p>If you look into the code, I did this by setting the respective values in \(\mathbf{K}^\top \mathbf{Q}\) to negative infinity before computing the softmax.</p>

<p><br /></p>
<div><img style="display:block; margin-left:auto; margin-right:auto; width:60%;" src="/Pages/assets/images/transformer-head.png" alt="Transformer head" />
<div style="display: table;margin: 0 auto;">Figure 6: Transformer head for a sequence length equal to 5.</div>
</div>
<p><br /></p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">class</span> <span class="nc">Head</span><span class="p">(</span><span class="n">nn</span><span class="p">.</span><span class="n">Module</span><span class="p">):</span>
    <span class="s">""" one head of self-attention """</span>
    
    <span class="k">def</span> <span class="nf">__init__</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">n_embd</span><span class="p">,</span> <span class="n">head_size</span><span class="p">,</span> <span class="n">sequence_len</span><span class="p">,</span> <span class="n">dropout</span><span class="p">):</span>
        <span class="nb">super</span><span class="p">().</span><span class="n">__init__</span><span class="p">()</span>
        <span class="bp">self</span><span class="p">.</span><span class="n">key</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="n">Linear</span><span class="p">(</span><span class="n">n_embd</span><span class="p">,</span> <span class="n">head_size</span><span class="p">,</span> <span class="n">bias</span><span class="o">=</span><span class="bp">False</span><span class="p">)</span>   <span class="c1"># key embedding
</span>        <span class="bp">self</span><span class="p">.</span><span class="n">query</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="n">Linear</span><span class="p">(</span><span class="n">n_embd</span><span class="p">,</span> <span class="n">head_size</span><span class="p">,</span> <span class="n">bias</span><span class="o">=</span><span class="bp">False</span><span class="p">)</span> <span class="c1"># query embedding
</span>        <span class="bp">self</span><span class="p">.</span><span class="n">value</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="n">Linear</span><span class="p">(</span><span class="n">n_embd</span><span class="p">,</span> <span class="n">head_size</span><span class="p">,</span> <span class="n">bias</span><span class="o">=</span><span class="bp">False</span><span class="p">)</span> <span class="c1"># value embedding
</span>        <span class="bp">self</span><span class="p">.</span><span class="n">register_buffer</span><span class="p">(</span><span class="s">'tril'</span><span class="p">,</span> <span class="n">torch</span><span class="p">.</span><span class="n">tril</span><span class="p">(</span><span class="n">torch</span><span class="p">.</span><span class="n">ones</span><span class="p">(</span><span class="n">sequence_len</span><span class="p">,</span> <span class="n">sequence_len</span><span class="p">)))</span>
        <span class="bp">self</span><span class="p">.</span><span class="n">dropout</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="n">Dropout</span><span class="p">(</span><span class="n">dropout</span><span class="p">)</span> <span class="c1"># to avoid overfitting
</span>        
    <span class="k">def</span> <span class="nf">forward</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">x</span><span class="p">):</span>
        <span class="n">B</span><span class="p">,</span><span class="n">T</span><span class="p">,</span><span class="n">C</span> <span class="o">=</span> <span class="n">x</span><span class="p">.</span><span class="n">shape</span>
        <span class="n">k</span> <span class="o">=</span> <span class="bp">self</span><span class="p">.</span><span class="n">key</span><span class="p">(</span><span class="n">x</span><span class="p">)</span> <span class="c1"># B, T, head_size, compute all keys
</span>        <span class="n">q</span> <span class="o">=</span> <span class="bp">self</span><span class="p">.</span><span class="n">query</span><span class="p">(</span><span class="n">x</span><span class="p">)</span> <span class="c1"># B, T, head_size, compute all queries
</span>        <span class="n">_</span><span class="p">,</span> <span class="n">_</span><span class="p">,</span> <span class="n">head_size</span> <span class="o">=</span> <span class="n">q</span><span class="p">.</span><span class="n">shape</span>
        
         <span class="c1"># B, T, head_size @ B, head_size, 
</span>        <span class="n">wei</span> <span class="o">=</span> <span class="n">q</span> <span class="o">@</span> <span class="n">k</span><span class="p">.</span><span class="n">transpose</span><span class="p">(</span><span class="o">-</span><span class="mi">2</span><span class="p">,</span> <span class="o">-</span><span class="mi">1</span><span class="p">)</span> <span class="o">*</span> <span class="p">(</span><span class="n">head_size</span> <span class="o">**</span> <span class="p">(</span><span class="o">-</span><span class="mf">0.5</span><span class="p">))</span> <span class="c1"># T =&gt; B, T, T
</span>
        <span class="c1"># because we can not look into the future 
</span>        <span class="n">wei</span> <span class="o">=</span> <span class="n">wei</span><span class="p">.</span><span class="n">masked_fill</span><span class="p">(</span><span class="bp">self</span><span class="p">.</span><span class="n">tril</span><span class="p">[:</span><span class="n">T</span><span class="p">,</span> <span class="p">:</span><span class="n">T</span><span class="p">]</span><span class="o">==</span><span class="mi">0</span><span class="p">,</span> <span class="nb">float</span><span class="p">(</span><span class="s">'-inf'</span><span class="p">))</span>
        <span class="n">wei</span> <span class="o">=</span> <span class="n">F</span><span class="p">.</span><span class="n">softmax</span><span class="p">(</span><span class="n">wei</span><span class="p">,</span> <span class="n">dim</span><span class="o">=-</span><span class="mi">1</span><span class="p">)</span>
        <span class="n">wei</span> <span class="o">=</span> <span class="bp">self</span><span class="p">.</span><span class="n">dropout</span><span class="p">(</span><span class="n">wei</span><span class="p">)</span>

        <span class="n">v</span> <span class="o">=</span> <span class="bp">self</span><span class="p">.</span><span class="n">value</span><span class="p">(</span><span class="n">x</span><span class="p">)</span> <span class="c1"># B, T, head_size
</span>        <span class="n">out</span> <span class="o">=</span> <span class="n">wei</span> <span class="o">@</span> <span class="n">v</span> <span class="c1"># T, T @ B, T, head_size =&gt; B, T, head_size
</span>        <span class="k">return</span> <span class="n">out</span>
<span class="p">...</span>
</code></pre></div></div>

<p>The multi-head attention layer consists of multiple heads.
Note that apart from <strong>masked</strong> <strong>self-attention</strong>, the head also applies a <strong>dropout</strong> which helps with regularization.</p>

<h2 id="stacked-multi-head-attention">Stacked Multi-Head Attention</h2>

<p>Instead of using only one <code class="language-plaintext highlighter-rouge">Head</code> it is usually a good idea to use multiple ones.
To do this we transform the input into a <code class="language-plaintext highlighter-rouge">head_size</code>-dimensional space.
Suppose we use 4 heads then <code class="language-plaintext highlighter-rouge">head_size * 4</code> should be equal to the rows of \(\mathbf{W}_0\) (compare Fig. 7) of the multi-head attention layer.
Since I add the input to the output of <code class="language-plaintext highlighter-rouge">MultiHeadAttention</code> (via residual connections), the columns of \(\mathbf{W}_0\) should be equal to the dimension of the input of the <code class="language-plaintext highlighter-rouge">MultiHeadAttention</code>.
In my case this is <code class="language-plaintext highlighter-rouge">n_embd</code>, i.e. the dimension of our embedded tokens.</p>

<p>\(\mathbf{W}_0\) transforms the concatenated results of the heads back to the dimension equal to <code class="language-plaintext highlighter-rouge">n_embd</code>. 
This is needed to stack <code class="language-plaintext highlighter-rouge">Block</code>s (each consisting of a <code class="language-plaintext highlighter-rouge">MultiHeadAttention</code>) on top of each other.
The output of <code class="language-plaintext highlighter-rouge">Block</code> \(i\) has to fit into block \(i+1\).</p>

<p><br /></p>
<div><img style="display:block; margin-left:auto; margin-right:auto; width:90%;" src="/Pages/assets/images/multi-head.png" alt="Multi-head attention" />
<div style="display: table;margin: 0 auto;">Figure 7: Multi-head attention.</div>
</div>
<p><br /></p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">class</span> <span class="nc">MultiHeadAttention</span><span class="p">(</span><span class="n">nn</span><span class="p">.</span><span class="n">Module</span><span class="p">):</span>
    <span class="k">def</span> <span class="nf">__init__</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">n_embd</span><span class="p">,</span> <span class="n">n_heads</span><span class="p">,</span> <span class="n">head_size</span><span class="p">,</span> <span class="n">sequence_len</span><span class="p">,</span> <span class="n">dropout</span><span class="p">):</span>
        <span class="nb">super</span><span class="p">().</span><span class="n">__init__</span><span class="p">()</span>
        <span class="bp">self</span><span class="p">.</span><span class="n">heads</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="n">ModuleList</span><span class="p">(</span>
            <span class="p">[</span><span class="n">Head</span><span class="p">(</span><span class="n">n_embd</span><span class="p">,</span> <span class="n">head_size</span><span class="p">,</span> <span class="n">sequence_len</span><span class="p">,</span> <span class="n">dropout</span><span class="p">)</span> <span class="k">for</span> <span class="n">_</span> <span class="ow">in</span> <span class="nb">range</span><span class="p">(</span><span class="n">n_heads</span><span class="p">)]</span>
        <span class="p">)</span>

        <span class="c1"># W_0
</span>        <span class="bp">self</span><span class="p">.</span><span class="n">W0</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="n">Linear</span><span class="p">(</span><span class="n">head_size</span> <span class="o">*</span> <span class="n">n_heads</span><span class="p">,</span> <span class="n">n_embd</span><span class="p">)</span> 

        <span class="bp">self</span><span class="p">.</span><span class="n">dropout</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="n">Dropout</span><span class="p">(</span><span class="n">dropout</span><span class="p">)</span>

    <span class="k">def</span> <span class="nf">forward</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">x</span><span class="p">):</span>
        <span class="c1"># concatenation of the results of each head
</span>        <span class="n">out</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">cat</span><span class="p">([</span><span class="n">head</span><span class="p">(</span><span class="n">x</span><span class="p">)</span> <span class="k">for</span> <span class="n">head</span> <span class="ow">in</span> <span class="bp">self</span><span class="p">.</span><span class="n">heads</span><span class="p">],</span> <span class="n">dim</span><span class="o">=-</span><span class="mi">1</span><span class="p">)</span> 

        <span class="c1"># Figure 7
</span>        <span class="n">out</span> <span class="o">=</span> <span class="bp">self</span><span class="p">.</span><span class="n">W0</span><span class="p">(</span><span class="n">out</span><span class="p">)</span>
        <span class="n">out</span> <span class="o">=</span> <span class="bp">self</span><span class="p">.</span><span class="n">dropout</span><span class="p">(</span><span class="n">out</span><span class="p">)</span>
        <span class="k">return</span> <span class="n">out</span>
<span class="p">...</span>
</code></pre></div></div>

<p>The hope is that each <code class="language-plaintext highlighter-rouge">Head</code> concentrates on different parts of the structure we want to learn.</p>

<h2 id="transformer-block">Transformer Block</h2>

<p>A <code class="language-plaintext highlighter-rouge">Block</code> consists of a <code class="language-plaintext highlighter-rouge">MultiHeadAttention</code>-layer and a relatively simple FFN-layer followed by two <code class="language-plaintext highlighter-rouge">LayerNorm</code> which applies <em>layer normalization</em> over the feature dimension within a single sample.
This helps the gradients to stay in a “good” range.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">class</span> <span class="nc">Block</span><span class="p">(</span><span class="n">nn</span><span class="p">.</span><span class="n">Module</span><span class="p">):</span>
    
    <span class="k">def</span> <span class="nf">__init__</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">n_embd</span><span class="p">,</span> <span class="n">n_heads</span><span class="p">,</span> <span class="n">sequence_len</span><span class="p">,</span> <span class="n">dropout</span><span class="p">):</span>
        <span class="nb">super</span><span class="p">().</span><span class="n">__init__</span><span class="p">()</span>
        <span class="c1"># head_size could be defined differently
</span>        <span class="n">head_size</span> <span class="o">=</span> <span class="n">n_embd</span> <span class="o">//</span> <span class="n">n_heads</span>
        <span class="bp">self</span><span class="p">.</span><span class="n">sa</span> <span class="o">=</span> <span class="n">MultiHeadAttention</span><span class="p">(</span><span class="n">n_embd</span><span class="p">,</span> <span class="n">n_heads</span><span class="p">,</span> <span class="n">head_size</span><span class="p">,</span> <span class="n">sequence_len</span><span class="p">,</span> <span class="n">dropout</span><span class="p">)</span>
        <span class="bp">self</span><span class="p">.</span><span class="n">ffwd</span> <span class="o">=</span> <span class="n">FeedForward</span><span class="p">(</span><span class="n">n_embd</span><span class="p">,</span> <span class="n">dropout</span><span class="p">)</span>
        <span class="bp">self</span><span class="p">.</span><span class="n">ln1</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="n">LayerNorm</span><span class="p">(</span><span class="n">n_embd</span><span class="p">)</span>
        <span class="bp">self</span><span class="p">.</span><span class="n">ln2</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="n">LayerNorm</span><span class="p">(</span><span class="n">n_embd</span><span class="p">)</span>
        
    <span class="k">def</span> <span class="nf">forward</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">x</span><span class="p">):</span>
        <span class="n">x</span> <span class="o">=</span> <span class="n">x</span> <span class="o">+</span> <span class="bp">self</span><span class="p">.</span><span class="n">sa</span><span class="p">(</span><span class="bp">self</span><span class="p">.</span><span class="n">ln1</span><span class="p">(</span><span class="n">x</span><span class="p">))</span> <span class="c1"># residual connection
</span>        <span class="n">x</span> <span class="o">=</span> <span class="n">x</span> <span class="o">+</span> <span class="bp">self</span><span class="p">.</span><span class="n">ffwd</span><span class="p">(</span><span class="bp">self</span><span class="p">.</span><span class="n">ln2</span><span class="p">(</span><span class="n">x</span><span class="p">))</span> <span class="c1"># residual connection
</span>        <span class="k">return</span> <span class="n">x</span>

<span class="k">class</span> <span class="nc">FeedForward</span><span class="p">(</span><span class="n">nn</span><span class="p">.</span><span class="n">Module</span><span class="p">):</span>
    
    <span class="k">def</span> <span class="nf">__init__</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">n_embd</span><span class="p">,</span> <span class="n">dropout</span><span class="p">):</span>
        <span class="nb">super</span><span class="p">().</span><span class="n">__init__</span><span class="p">()</span>
        <span class="bp">self</span><span class="p">.</span><span class="n">net</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="n">Sequential</span><span class="p">(</span>
            <span class="n">nn</span><span class="p">.</span><span class="n">Linear</span><span class="p">(</span><span class="n">n_embd</span><span class="p">,</span> <span class="mi">4</span> <span class="o">*</span> <span class="n">n_embd</span><span class="p">),</span> 
            <span class="n">nn</span><span class="p">.</span><span class="n">ReLU</span><span class="p">(),</span>
            <span class="n">nn</span><span class="p">.</span><span class="n">Linear</span><span class="p">(</span><span class="mi">4</span> <span class="o">*</span> <span class="n">n_embd</span><span class="p">,</span> <span class="n">n_embd</span><span class="p">),</span>
            <span class="n">nn</span><span class="p">.</span><span class="n">Dropout</span><span class="p">(</span><span class="n">dropout</span><span class="p">),</span>
        <span class="p">)</span>
        
    <span class="k">def</span> <span class="nf">forward</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">x</span><span class="p">):</span>
        <span class="k">return</span> <span class="bp">self</span><span class="p">.</span><span class="n">net</span><span class="p">(</span><span class="n">x</span><span class="p">)</span>
<span class="p">...</span>
</code></pre></div></div>

<p><br /></p>
<div><img style="display:block; margin-left:auto; margin-right:auto; width:40%;" src="/Pages/assets/images/block.png" alt="Transformer block" />
<div style="display: table;margin: 0 auto;">Figure 8: Our decoder transformer block with only one (masked) multi-head attention layer.</div>
</div>

<h2 id="positional-encoding">Positional Encoding</h2>

<p>The Transformer has no more hidden state.
Therefore, instead of processing token by token trying to memorize important information via the hidden state, it processes all \(n\) tokens in parallel, which is good for parallel computation but increases the time and space complexity from \(\mathcal{O}(n)\) (LSTM) to \(\mathcal{O}(n^2)\).</p>

<p>Furthermore, we have to encode the position of the tokens into \(\mathbf{x}\) because we lost the implicit order of computation.
In the original paper <a class="citation" href="#vaswani:2017">(Vaswani et al., 2017)</a> the authors utilized an embedding that involved sine and cosine functions.
Their embedding is a very clever use of periodic functions but I will not go into details here.
Instead of using a fixed embedding, I let the transformer learn the positional embedding.</p>

<p>Therefore, I transform the input <code class="language-plaintext highlighter-rouge">idx</code> into two vectors <strong>positional embedding</strong> and <strong>token embedding</strong>, which are <strong>added</strong> together.
Note that our input <code class="language-plaintext highlighter-rouge">idx</code> is an array of numbers each representing the id of the token.
Each number will be transformed into a specific vector (i.e. its embedding).
The embedding will be learned.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">class</span> <span class="nc">TransformerDecoder</span><span class="p">(</span><span class="n">nn</span><span class="p">.</span><span class="n">Module</span><span class="p">):</span>
    
    <span class="k">def</span> <span class="nf">__init__</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">vocab_size</span><span class="p">,</span> <span class="n">sequence_len</span><span class="p">,</span> <span class="n">n_embd</span><span class="p">,</span> <span class="n">n_heads</span><span class="p">,</span> <span class="n">n_blocks</span><span class="p">,</span> <span class="n">dropout</span><span class="p">):</span>
        <span class="nb">super</span><span class="p">().</span><span class="n">__init__</span><span class="p">()</span>

        <span class="c1"># vocab_size is the size of our alphabet
</span>        <span class="bp">self</span><span class="p">.</span><span class="n">token_embedding_table</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="n">Embedding</span><span class="p">(</span><span class="n">vocab_size</span><span class="p">,</span> <span class="n">n_embd</span><span class="p">)</span>

        <span class="c1"># sequence_len is equal to n
</span>        <span class="bp">self</span><span class="p">.</span><span class="n">position_embedding_table</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="n">Embedding</span><span class="p">(</span><span class="n">sequence_len</span><span class="p">,</span> <span class="n">n_embd</span><span class="p">)</span>

        <span class="bp">self</span><span class="p">.</span><span class="n">blocks</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="n">Sequential</span><span class="p">(</span>
            <span class="o">*</span><span class="p">[</span><span class="n">Block</span><span class="p">(</span><span class="n">n_embd</span><span class="p">,</span> <span class="n">n_heads</span><span class="p">,</span> <span class="n">sequence_len</span><span class="p">,</span> <span class="n">dropout</span><span class="p">)</span> <span class="k">for</span> <span class="n">_</span> <span class="ow">in</span> <span class="nb">range</span><span class="p">(</span><span class="n">n_blocks</span><span class="p">)]</span>
        <span class="p">)</span>
        <span class="bp">self</span><span class="p">.</span><span class="n">lm_head</span> <span class="o">=</span> <span class="n">nn</span><span class="p">.</span><span class="n">Linear</span><span class="p">(</span><span class="n">n_embd</span><span class="p">,</span> <span class="n">vocab_size</span><span class="p">)</span>

    <span class="k">def</span> <span class="nf">forward</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">idx</span><span class="p">):</span>
        <span class="n">B</span><span class="p">,</span> <span class="n">T</span> <span class="o">=</span> <span class="n">idx</span><span class="p">.</span><span class="n">shape</span>
        
        <span class="c1"># token embedding. B, T, n_embd
</span>        <span class="n">token_emb</span> <span class="o">=</span> <span class="bp">self</span><span class="p">.</span><span class="n">token_embedding_table</span><span class="p">(</span><span class="n">idx</span><span class="p">)</span> 

        <span class="c1"># positional embedding. T, n_embd 
</span>        <span class="n">pos_emb</span> <span class="o">=</span> <span class="bp">self</span><span class="p">.</span><span class="n">position_embedding_table</span><span class="p">(</span><span class="n">torch</span><span class="p">.</span><span class="n">arange</span><span class="p">(</span><span class="n">T</span><span class="p">,</span> <span class="n">device</span><span class="o">=</span><span class="n">device</span><span class="p">))</span> 

        <span class="n">x</span> <span class="o">=</span> <span class="n">token_emb</span> <span class="o">+</span> <span class="n">pos_emb</span> <span class="c1"># B, T, n_embd + T, n_embd =&gt; B, T, n_embd
</span>        <span class="n">x</span> <span class="o">=</span> <span class="bp">self</span><span class="p">.</span><span class="n">blocks</span><span class="p">(</span><span class="n">x</span><span class="p">)</span> <span class="c1"># B, T, head_size
</span>        <span class="n">logits</span> <span class="o">=</span> <span class="bp">self</span><span class="p">.</span><span class="n">lm_head</span><span class="p">(</span><span class="n">x</span><span class="p">)</span> <span class="c1"># B, T, vocab_size
</span>        <span class="k">return</span> <span class="n">logits</span>

<span class="p">...</span>
</code></pre></div></div>

<p>By increasing the dimension of the embedding, the sequence length, the number of heads within a block and the number of blocks we can drastically increase the size and power of our decoder-only transformer.
However, this will rapidly increase the memory requirements and training time.</p>

<div><img style="display:block; margin-left:auto; margin-right:auto; width:40%;" src="/Pages/assets/images/decoder.png" alt="Decoder-only transformer" />
<div style="display: table;margin: 0 auto;">Figure 9: Our simplified decoder-only transformer.</div>
</div>
<p><br /></p>

<p>Furthermore, it is very important to understand that what really matters is a <strong>high-quality training dataset</strong>!
Of course, your model architecture matters too, but your model can not learn what is not there.
Additionally, the <strong>musical representation</strong> you feed into the transformer matters as well.
In our case this representation, using basically piano rolls, is very simple.
It does not contain any high level information such as the end of a bar, section, phrase or musical theme.
We just hope that the transformer will eventually learn all these concepts.
It is an active research question what impact a good musical representation has on the result the trained transformer generates.</p>

<h2 id="relative-positional-self-attention">Relative Positional Self-Attention</h2>

<p>So far our positional encoding was just a sequence of natural numbers \(0, 1, \ldots, n-1\) and we used an embedding which was added to the input, that is, the embedding of \(i\) was added to the embedding of \(\mathbf{x}_i\) of the input sequence.
However, this encoding might not be optimal in the context of music where tones, phrases, musical ideas and themes repeat frequently.
A relative position representation to allow attention to be informed by how far two positions are apart in a sequence might be much more effective.</p>

<p>So, instead of learning the index of a token within a sequence we want the model to learn relative distances between tokens.
In other words, instead of learning the attention spent by token with index \(j\) on \(i\), that is, \(\alpha(\mathbf{x}_i, \mathbf{x}_j)\) we want to compute an attention score based on the (directed) distance \(i-j\).
This concept was introduced by <a class="citation" href="#shaw:2018">(Shaw et al., 2018)</a>.</p>

<div><img style="display:block; margin-left:auto; margin-right:auto; width:60%;" src="/Pages/assets/images/relative-attention.png" alt="Relative positional encoding" />
<div style="display: table;margin: 0 auto;">Figure 10: Relative positional encoding.</div>
</div>
<p><br /></p>

<p>Note that if there are \(\mathcal{O}(n)\) absolute positions \(0, 1, \ldots, n-1\) then there are \(\mathcal{O}(n)\) relative positions \(-(n-1), \ldots, -1, 0, 1,\ldots, n-1\).
The authors also introduce a maximal distance \(k\) such that they only learn weights</p>

\[\mathbf{w}^{V}_{\text{clip}(i-j,k)} \text{ with } \text{clip}(x,k) = \max(-k, \min(k,x))\]

<p>Therefore, they learn relative position representations for the keys \(\mathbf{w}^K_{-k}, \ldots, \mathbf{w}^K_{k}\) and for the values \(\mathbf{w}^V_{-k}, \ldots, \mathbf{w}^V_{k}\).
They introduce the relative position between \(\mathbf{x}_i\) and \(\mathbf{x}_j\) to be</p>

\[\mathbf{a}_{ij} = \mathbf{w}_{\text{clip}(i-j,k)}\]

<p>Thus there are \(\mathcal{O}(n^2)\) different such vectors but many share the same value.
Of course, they drop the absolute positional encoding.
And they adapt the <strong>self-attention computation</strong> from</p>

\[\mathbf{sa}_j(\mathbf{x}_0, \ldots, \mathbf{x}_{n-1}) = \sum_{i=0}^{n-1} \alpha(\mathbf{x}_i, \mathbf{x}_j) \mathbf{v}_i.\]

<p>to</p>

\[\mathbf{sa}_j(\mathbf{x}_0, \ldots, \mathbf{x}_{n-1}) = \sum_{i=0}^{n-1} \alpha(\mathbf{x}_i, \mathbf{x}_j) (\mathbf{v}_i + \mathbf{a}^V_{ij}).\]

<p>and the computation of the similarity between <strong>query</strong> and <strong>key</strong> from</p>

\[\frac{(\mathbf{W}_q \mathbf{x}_j)^\top (\mathbf{W}_k \mathbf{x}_i)}{\sqrt{D_q}}\]

<p>to</p>

\[\frac{(\mathbf{W}_q \mathbf{x}_j)^\top (\mathbf{W}_k \mathbf{x}_i + \mathbf{a}^K_{ij})}{\sqrt{D_q}}.\]

<p>Computation-wise the first manipulation can be easily achieved by adding a matrix \(\mathbf{A}\) to \(\mathbf{V}\).
However, the second manipulation destroys parallelism, i.e. the possibility to compute everything by matrix-matrix multiplications.
This can be mitigated by splitting the computation into two parts:</p>

\[\frac{(\mathbf{W}_q \mathbf{x}_j)^\top (\mathbf{W}_k \mathbf{x}_i + \mathbf{a}^K_{ij})}{\sqrt{D_q}} = \frac{(\mathbf{W}_q \mathbf{x}_j)^\top (\mathbf{W}_k \mathbf{x}_i) + (\mathbf{W}_q \mathbf{x}_j)^\top \mathbf{a}^K_{ij}}{\sqrt{D_q}}\]

<p>Assuming that \(\mathbf{S}\) contains all the relative attention scores, that is,</p>

\[\mathbf{S}_{ij} = (\mathbf{W}_q \mathbf{x}_j)^\top \mathbf{a}^K_{ij},\]

<p>then we can go back to the matrix form which gives us</p>

\[\mathbf{Sa}(\mathbf{X}) = (\mathbf{V} + \mathbf{A}^V) \cdot \text{Softmax}\left( \frac{\mathbf{K}^\top \mathbf{Q} + \mathbf{S}}{\sqrt{D_q} }\right).\]

<p>To compute \(\mathbf{S}\) <a class="citation" href="#shaw:2018">(Shaw et al., 2018)</a> instantiate an intermediate tensor \(\mathbf{R} \in \mathbf{R}^{k \times k \times D_q},\) containing the embeddings that correspond to the relative distance between all keys and queries.
\(\mathbf{Q}\) is then reshaped to an \((k, 1, D_q)\) tensor, and \(\mathbf{S} = \mathbf{Q} \mathbf{R}^\top.\)
This incurs a total space complexity of \(\mathcal{O}(k^2 D_q)\).</p>

<h2 id="the-music-transformer">The Music Transformer</h2>

<p>The Music Transformer <a class="citation" href="#huang:2018">(Huang et al., 2018)</a> was one of the first transformer utilized to generate symbolic music.
Even if it was introduced five years ago (which is like a century in the AI-world) it is worth studying it.
In the paper you find two different datasets</p>

<ol>
  <li><a href="https://github.com/czhuang/JSB-Chorales-dataset">J.S. Bach chorales dataset</a></li>
  <li><a href="https://www.piano-e-competition.com/">Piano-e-Competition dataset</a></li>
</ol>

<p>and they used an impressive sequence length of <strong>2048-tokens</strong>!
They used GPUs for the training.
With such a large number of token, one question arises: How did they manage to put 2000-tokens and all the respective matrices in the GPUs’ memory?</p>

<p>The authors correctly identify the space complexity of \(\mathcal{O}(k^2 D_q)\) to be problematic for GPU computation and they reduce the complexity to \(\mathcal{O}(k D_q)\) by exchanging space for re-computation.
This is possible due to the structure of the tensor \(\mathbf{R}\) which contains many equal values.</p>

<p>To handle very long sequences, the authors use local attention <a class="citation" href="#liu:2018">(Liu et al., 2018)</a> by chunking the input sequence into non-overlapping blocks.
Each block then attends to itself and the one before.</p>

<h2 id="attention-free-transformer">Attention-Free Transformer</h2>

<p>Basically, the attention mechanism, regardless of the specifics, solves a routing problem, that is, which information is transported to the next layer of the neural network.
Thus, it has a quadratic time and space complexity of \(\mathcal{O}(n^2)\) where \(n\) is our sequence length.
Therefore, if you have limited resources, it is hard to scale it to larger sequences.
As with the local attention and other techniques, like the Linformer <a class="citation" href="#wang:2020">(Wang et al., 2020)</a>, Longformer <a class="citation" href="#beltagy:2020">(Beltagy et al., 2020)</a>, Reformer <a class="citation" href="#kitaev:2020">(Kitaev et al., 2020)</a>, and Synthesizer <a class="citation" href="#tay:2021">(Tay et al., 2021)</a>, and Performer <a class="citation" href="#choromanski:2022">(Choromanski et al., 2022)</a> there are ways to improve this but in principle the complexity will bite us eventually.</p>

<p>Now we enter in an era in deep learning where we question if we actually need the attention layers in the transformer!
This was proposed in 2022.
Instead of computing attention, <em>FNet</em> <a class="citation" href="#leethorp:2022">(Lee-Thorp et al., 2022)</a> just mixes tokens according to the discrete Fourier transformation (DFT).
First, a 1D transformation is computed with respect to the embedding and then another with respect to time.
Amazingly even though there is no parameter to learn within the <code class="language-plaintext highlighter-rouge">Fourier</code>-layer (which replaces the <code class="language-plaintext highlighter-rouge">Head</code>) this strategy seems to work almost as good as the far more computationally expensive task of learning all the required attention scores.</p>

<p>The Fourier transform decomposes a function (in our case a discrete signal) into its constituent frequencies.
Given a sequence \(x_0, \ldots, x_{N-1}\), the discrete Fourier transform (DFT) is defined by</p>

\[X_k = \sum\limits_{n=0}^{N-1} x_n \exp\left( - \frac{2\pi i}{N} nk \right), \quad 0 \leq k \leq N-1.\]

<p>\(X_k\) encodes the <strong>phase</strong> and <strong>amplitude</strong> of frequency \(k\) within the signal.</p>

<p><em>FNet</em> consists of a Fourier <strong>mixing sublayer</strong> followed by a feed-forward sublayer.
Essentially, the self-attention sublayer of each transformer decoder layer is replaced with a <strong>Fourier sublayer</strong> which applies a 2D DFT to its</p>

\[(\text{sequence length} \times \text{hidden dimension})\]

<p>embedding input.
This can be achieved using two 1D DFTs—one 1D DFT along the sequence dimension, \(\mathcal{F}_\text{seq}\), and one 1D DFT along the hidden dimension, \(\mathcal{F}_\text{h}\):</p>

\[y = \text{Real}\left( \mathcal{F}_\text{seq} \left( \mathcal{F}_\text{h}(\mathbf{x}) \right) \right)\]

<p>The authors only consider the real part of the DFT.</p>

<p>Now, as emphasized by the title of their paper, the Fourier transform is probably not the important part.
It is just a special case of how you can mix tokens.
Important is the mixing itself which allows information to flow from one token to all the other tokens and the Fourier transform happens to be a nice way of mixing.
The paper indicates that it might not be so important to let the model learn how exactly information flows around.
It might be just enough if information flows at all (to all tokens).
In other words, the exact routing might be less important than we thought.</p>

<p>Now, the results of the paper are not better than using a traditional transformer.
But one trades accuracy for resources thus longer sequence length and a faster computation.</p>

<p>To the best of my knowledge, I have not seen this tried out for symbolic music generation.
But when I have time, I’ll play around with it.
Furthermore, one might think about a special mixing which is effective for our specific task.</p>

<h2 id="references">References</h2>

<ol class="bibliography"><li><span id="vaswani:2017">Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., &amp; Polosukhin, I. (2017). attention is all you need. <i>CoRR</i>, <i>abs/1706.03762</i>. http://arxiv.org/abs/1706.03762</span></li>
<li><span id="hochreiter:1997">Hochreiter, S., &amp; Schmidhuber, J. (1997). Long short-term memory. <i>Neural Computation</i>, <i>9</i>(8), 1735–1780. https://doi.org/10.1162/neco.1997.9.8.1735</span></li>
<li><span id="chung2014">Chung, J., Gulcehre, C., Cho, K. H., &amp; Bengio, Y. (2014). <i>Empirical evaluation of gated recurrent neural networks on sequence modeling</i>.</span></li>
<li><span id="bahdanau:2014">Bahdanau, D., Cho, K., &amp; Bengio, Y. (2014). Neural machine translation by jointly learning to align and translate. <i>CoRR</i>, <i>abs/1409.0473</i>.</span></li>
<li><span id="brown:2020">Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., … Amodei, D. (2020). <i>Language models are few-shot learners</i>.</span></li>
<li><span id="openai:2023">OpenAI. (2023). <i>GPT-4 rechnical report</i>.</span></li>
<li><span id="bubeck:2023">Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M. T., &amp; Zhang, Y. (2023). <i>Sparks of artificial general intelligence: Early experiments with GPT-4</i>.</span></li>
<li><span id="touvron:2023">Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., &amp; Lample, G. (2023). <i>LLaMA: Open and efficient foundation language models</i>.</span></li>
<li><span id="touvron:2023b">Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., … Scialom, T. (2023). <i>LlaMA 2: Open foundation and fine-tuned chat models</i>.</span></li>
<li><span id="devlin:2019">Devlin, J., Chang, M.-W., Lee, K., &amp; Toutanova, K. (2019). <i>BERT: Pre-training of deep bidirectional transformers for language understanding</i>.</span></li>
<li><span id="chen:2021">Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., … Zaremba, W. (2021). <i>Evaluating large language models trained on code</i>.</span></li>
<li><span id="huang:2018">Huang, C.-Z. A., Vaswani, A., Uszkoreit, J., Shazeer, N., Hawthorne, C., Dai, A. M., Hoffman, M. D., &amp; Eck, D. (2018). Music Transformer: Generating music with long-term structure. <i>ArXiv Preprint ArXiv:1809.04281</i>.</span></li>
<li><span id="huang:2020">Huang, Y.-S., &amp; Yang, Y.-H. (2020). <i>Pop Music Transformer: Beat-based modeling and generation of expressive pop piano compositions</i>.</span></li>
<li><span id="ens:2020">Ens, J., &amp; Pasquier, P. (2020). <i>MMM: Exploring conditional multi-track music generation with the transformer</i>.</span></li>
<li><span id="hadjeres:2021">Hadjeres, G., &amp; Crestel, L. (2021). <i>The piano inpainting application</i>.</span></li>
<li><span id="shih:2022">Shih, Y.-J., Wu, S.-L., Zalkow, F., Müller, M., &amp; Yang, Y.-H. (2022). <i>Theme Transformer: Symbolic music generation with theme-conditioned transformer</i>.</span></li>
<li><span id="copet:2023">Copet, J., Kreuk, F., Gat, I., Remez, T., Kant, D., Synnaeve, G., Adi, Y., &amp; Défossez, A. (2023). <i>Simple and controllable music generation</i>.</span></li>
<li><span id="shaw:2018">Shaw, P., Uszkoreit, J., &amp; Vaswani, A. (2018). Self-attention with relative position representations. <i>CoRR</i>, <i>abs/1803.02155</i>. http://arxiv.org/abs/1803.02155</span></li>
<li><span id="liu:2018">Liu, P. J., Saleh, M., Pot, E., Goodrich, B., Sepassi, R., Kaiser, L., &amp; Shazeer, N. (2018). Generating Wikipedia by summarizing long sequences. <i>International Conference on Learning Representations</i>. https://openreview.net/forum?id=Hyg0vbWC-</span></li>
<li><span id="wang:2020">Wang, S., Li, B. Z., Khabsa, M., Fang, H., &amp; Ma, H. (2020). <i>Linformer: Self-Attention with linear complexity</i>.</span></li>
<li><span id="beltagy:2020">Beltagy, I., Peters, M. E., &amp; Cohan, A. (2020). <i>Longformer: The long-document transformer</i>.</span></li>
<li><span id="kitaev:2020">Kitaev, N., Kaiser, Ł., &amp; Levskaya, A. (2020). <i>Reformer: The Efficient Transformer</i>.</span></li>
<li><span id="tay:2021">Tay, Y., Bahri, D., Metzler, D., Juan, D.-C., Zhao, Z., &amp; Zheng, C. (2021). <i>Synthesizer: Rethinking self-attention in transformer models</i>.</span></li>
<li><span id="choromanski:2022">Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J., Mohiuddin, A., Kaiser, L., Belanger, D., Colwell, L., &amp; Weller, A. (2022). <i>Rethinking Attention with Performers</i>. https://arxiv.org/abs/2009.14794</span></li>
<li><span id="leethorp:2022">Lee-Thorp, J., Ainslie, J., Eckstein, I., &amp; Ontanon, S. (2022). <i>FNet: Mixing tokens with Fourier transforms</i>.</span></li></ol>]]></content><author><name>Benedikt Zönnchen</name><email>benedikt.zoennchen@web.de</email></author><category term="Music" /><category term="ML" /><category term="Transformer" /><summary type="html"><![CDATA[Recurrent models trained in practice are effectively feed-forward. This could happen either because truncated backpropagation through time cannot learn patterns significantly longer than k steps, or, more provocatively, because models trainable by gradient descent cannot have long-term memory. – John Miller]]></summary></entry><entry><title type="html">Escaping the Reality of the Climate Crisis?</title><link href="https://bzoennchen.github.io/Pages/2024/01/02/conspiracy.html" rel="alternate" type="text/html" title="Escaping the Reality of the Climate Crisis?" /><published>2024-01-02T00:00:00+01:00</published><updated>2024-01-02T00:00:00+01:00</updated><id>https://bzoennchen.github.io/Pages/2024/01/02/conspiracy</id><content type="html" xml:base="https://bzoennchen.github.io/Pages/2024/01/02/conspiracy.html"><![CDATA[<p>Diverging from my area of expertise is always a risky endeavor, but since this is a blog and not a scientific journal, I’m giving myself the liberty to explore and have fun with different ideas (even if the topic is depressing). 
Often writing helps in transforming the mess into a structured and coherent concept.
The process of rethinking and reflecting can be invaluable.
It helps to make ones thought <em>anschlussfähig</em> which literally means <em>to be capable for connections</em> and in this context means <em>enabling the continuation of communication</em>.</p>

<p>In this piece, I aim to explore various aspects by linking the movie <em>The Matrix</em>, Plato’s <em>Allegory of the Cave</em>, <em>myths</em>, and <em>conspiracy theories</em>.
Additionally, delve into the understandable yet problematic skepticism surrounding <em>second-order observation</em>.
I will relate this skepticism to the notions of <em>complexity</em> and <em>hyperreality</em>, and discuss why we rely on <em>second-order observation</em> to address the <em>climate crisis</em>. 
Some parts of this text will escape a clear interpretation, especially when I offer my interpretation of Baudrillard’s writings.
Some parts might even be contradictory.
But this is the point: enduring or even enjoying ambiguity!</p>

<p>Overall, I hope to give good reasons for the emergence of conspiracy theories and our state of inaction in the face of disaster; reasons that do not rely on a good and evil dichotomy of individuals.
I will argue that our perception of reality makes effective communication between each other improbable and that complex systems follow their own stabilizing dynamic.
In the end, conspiracy theorist will appear as anti-authoritarian rebels that fight against the power of knowledge using the same source of authority, that is, knowledge.</p>

<p>Before we start, let us agree on a definition of conspiracy theories:</p>

<blockquote>
  <p>By conspiracy theory, I mean an explanation of historical, ongoing, or future events that cites as a main causal factor a group of powerful persons, the conspirators, acting in secret for their own benefit against the common good. – Joseph E. Uscinski</p>
</blockquote>

<h2 id="mythologies-science-and-conspiracy-theories">Mythologies, Science and Conspiracy Theories</h2>

<p>The phenomena of conspiracy theories hunts and intrigues me since the terror attacks of 9/11 happened back when I was a child.
I remember watching the <em>Zeitgeist series</em>, which linked various conspiracy theories involving religion, the September 11 attacks and the financial sector.
During that time, even German TV occasionally presented documentaries that portrayed certain events in a conspiratorial light.
The shock and uncertainty that the Western world experienced after these events, combined with a sense of lack of control, created a fertile ground for such theories. 
These ‘documentaries’ were not only entertaining, but they also sparked my interest in geopolitics, the history of religion, history in general, and even philosophy. 
They managed to make historical events captivating, and I often wished that my history classes were similarly engaging.
Fortunately, I never embraced the logic presented in these films; for me, they remained within the realm of entertainment but I could see how easy it is to fall for them on an emotional level.
Interestingly, years later, when a real plot unfolded to deceive the public (and other nations) into supporting the war against Iraq, there was no corresponding emergence of conspiracy theories like those seen previously.</p>

<p>During the pandemic, I observed a repetition of history in the form of similar documentaries emerging, and I became interested in how they were designed and how they relate to the conspiratorial ‘documentaries’ I encountered in my childhood. Indeed, they bore striking similarities.</p>

<p>Apart from the obvious parallels, such as misrepresentation and drawing connections between completely unrelated events, there were also more bizarre links.
For instance, these documentaries often incorporate some form of spiritual concept, promising a return to or discovery of a ‘true self’, an ‘inner peace’ or a ‘forgotten innocence’. 
For example, the <em>Zeitgeist series</em> uses a speech from the Indian philosopher Jiddu Krishnamurti (1895 – 1986) but only as an emotional device:</p>

<blockquote>
  <p>We will see how very important it is to bring about in the human mind the radical revolution.
The crisis is a crisis in consciousness, a crisis that cannot anymore accept the old norms, the old patterns, the ancient traditions.
And considering what the world is now with all the misery, conflict, destructive brutality, aggression, and so on, man is still as he was, is still brutal, violent, aggressive, competitive and has built a society along these lines. – Jiddu Krishnamurti</p>
</blockquote>

<p>In my interpretation, this reflects a lost connection to the wholeness of the universe. 
At the extreme end, I experience sometimes two contradictory moods. 
First, there is this feeling of alienation in an absurd world, which Albert Camus described best in his book <em>The Stranger</em>.
For me, the absurdity of the world reveals itself when I am out in the city, observing people rushing through their lives, participating in the acceleration towards a promised utopia that no longer exists, not even in their imaginations.
It is also connected to the realization that I was thrown into this world, culture, mess; this contradiction; this meaningless rat race.
The second mood arises from a profound connection with the world.
It feels as if we are the world, as if there is no separation between myself and my environment, between myself and the universe.
Yet, at the same time, I sense a process that is not only mysterious but overwhelmingly greater than I can ever comprehend; a force beyond my intellect.
Such feelings arise, for example, in nature.
When we stand atop a mountain, not thinking but simply being present in the world, or even feeling as if we are the world.
Such a feeling can also arise when we are completely immersed in an activity, like a drummer who dissolves entirely into the act of drumming.
This state of existence requires a quiet mind and a cessation of thought.</p>

<p>Consequently, it was no surprise to me that at the protests against COVID-19 restrictions and at QAnon gatherings, there was a peculiar blend of people, including everyone from far right-wingers to faith healers.
Of course, one has to be careful with such categorization since such protests are often captured by extreme parties.
Anyways, this mix reflects the broad, albeit unusual, appeal of such conspiratorial narratives and the desire for an effective (post-truth) narrator, e.g. Donald Trump, who provides a myth that carries emotional weight rather than the difficulty and uncertainty of a complex world—a <em>myth</em> one can live by.
However, <em>ignorance</em> is too simple of an explanation.
The MAGA cult is not disengaged in communication or the production of knowledge.
Instead they (ab)use communication to construct quite imaginative but also inconsistant <em>alternative facts</em> effectively.</p>

<p>On the basis of knowledge, it is easy to make fun of people believing in mythologies.
Scientifically speaking, myths are either inconsistent or unfalsifiable.
However, socially they can be very useful and powerful.
They explain experience and reduce complexity.
A myth is meant to answer questions and offer solutions to quell the anxieties of the present through stories.
They can serve a useful purpose, particularly in politics.
They create visions, unities and identities among groups such that they can work together towards a meaningful goal—a mission greater than oneself.
Here, a shadow of spirituality plays an important role, be it in the form of <em>the Light of God</em>, <em>Siegfried the Dragonslayer</em>, <em>Achilleus the Greatest of All the Greek Warriors</em> or <em>Donald Trump the Warrior King</em>.</p>

<p>Later I will argue that we should keep the source for spirituality, that is, imagination, ambiguity, and contradictions alive and that science can be abused for a <em>Crusade Against Ambiguity</em>.
In my view, spirituality and science can harmonize with quite well.
Religion, as a subset of spirituality, should be criticised especially if it becomes dogmatic, that is, if it starts a crusade against ambiguity.
But it is a narrow perspective to scientifically dispute the existence of a divine entity, just as it is misguided to interpret religious scriptures in a strictly literal sense.
Religion needs ambiguity.
Therefore, it is peculiar to witness esteemed intellectuals like Richard Dawkins engage in debates concerning the divine, overlooking the potential for a creator amidst the universe’s intricate complexity.</p>

<blockquote>
  <p>We have a working theory, which we know is true, which explains how you can go from great simplicity to prodigious complexity.
And finally to the sort of complexity which is capable of designing things, of creating things, of working out how to do things.
If you suddenly going to insert a designing machine, a creator, an intelligence at the root of the universe you have just undermined your entire enterprise because your entire enterprise has been to explain how you get to something complicated enough to do design. – Richard Dawkins</p>
</blockquote>

<p>When Dawkins discusses evolution, it is easy to agree with him from a scientific perspective.
I want him to defend his theory.
However, he often speaks in absolutes, using the language of dogmatic religion that he himself criticizes.
From a systems theory point of view, it can resolve this paradox by re-entering itself, that is, by communicating about itself on its own terms—a sort of self-reflection.</p>

<p>Dawkins and similar critics overlook a crucial point: the pursuit of absolute truth is elusive.
Even Dawkins’ assertion that “the theory is true” misrepresents the nature of scientific inquiry, which he is undoubtedly aware of.
Science provides models that explain phenomena until new evidence suggests otherwise.
It mosty works on the basis of <em>falsifiability</em> for <em>demarcation</em>, a concept introduced by Karl Popper (a critic of the inductive theory of science).</p>

<blockquote>
  <p>The problem of finding a criterion which would enable us to distinguish between the empirical sciences on the one hand, and mathematics and logic as well as ‘metaphysical’ systems on the other, I call the problem of demarcation. – <a class="citation" href="#popper:1934">(Popper, 1934)</a></p>
</blockquote>

<p>A statement or system of statements (theory) is falsifiable if it is capable of conflicting with possible, or conceivable observation;
The theory must be able to fail when tested against reality.
Build on rigor and the combination of theory bound by empiricism makes scientific method incredibly useful, and according to philosophers like Markus Gabriel, brings us closer to the Truth.
In contrast, others, such as Richard Rorty, completely reject the notion of absolute Truth and instead emphasize the practical usefulness of the scientific method.</p>

<p>The question of ontology remains unanswered even if most of us operate on the assumption of some sort of <em>naturalism</em>. 
It is the principle by which science operates and by which people in well-developed countries often live.
Some define it as “the idea that only natural laws and forces operate in the universe”.
The philosopher Quine described naturalism as “the position that there is no higher tribunal for truth than natural science itself”.
Quine’s more humble and pragmatic definition allows space for profound questions regarding the existence of time, space, causality, and our place within this framework—questions that lie beyond the scope of scientific inquiry.
For him and many others, these are absolute unknowns that can be considered within the realm of spirituality.
These unanswerable questions delve into the <em>essence of being</em> and the universe, inviting a spiritual exploration alongside scientific understanding.</p>

<p>Dawkins seems also to be unaware of the usefulness of ambiguity.
People can hold contradictory beliefs without being irrational or anti-science.
For example, it is unlikely that individuals with a non-dogmatic spiritual outlook adhere to a literal interpretation of the Earth’s creation in seven days or dismiss evolutionary theory but, at the same time, they might believe in a creator.
They can operate in different social systems with different rationals.
This is neither good or bad but a sign of diversity.
I mean how many mathematicians still believe that math has something to do with a divine entity or realm?
Such contradictions only become problematic if systems interfere in the other’s operations, e.g., if religion operates in science or science in religion.</p>

<p>But why can spirituality lead to the descent into the rabbit hole of conspiracy theories?
First of all, there is a strong relation between philosophy, religion, and spirituality, e.g., between Platonism and Christianity.
Monolithic religions offer a rather rigorous explanation of why things are as they are, based on the presumption of a creator.
Especially, non-believers sometimes misunderstand that religious people dislike logic when, in fact, it was Thomas Aquinas who attempted to synthesize Aristotelian philosophy (and logic) with the principles of Christianity.
He produced a vast body of precise, detailed, and systematic philosophical writings, in which he integrated Aristotle’s encyclopedic work and medieval Christian theology into a seamless whole.
The dark side of this was that any contradiction by future scientists would necessarily have to be seen as heresy.
Philosopher Bertrand Russell pointed out that Aquinas started by already knowing the truth in the form of the Catholic faith.
Aquinas used logic to strengthen his belief system and not to question it.
Note that postmodern thinkers argue that philosophers, who practiced metaphysics, did basically the same but in a more clever way.
Famous is Nietzsche’s suspicion of Kant’s categorical imperative, which is, after all, categorical.
Kant, however, pointed to the source of the problem which is not logic or rational thinking but a lack of empirical evidence; a lack of outwardness; of asking nature.</p>

<p>From this perspective, one might say that conspiracists are in the business of doing metaphysics poorly.
It is certainly the case that there are similarities in doing metaphysics and constructing a grand conspiratorial theory (or myth) that attempts to explain everything.
And like Aquinas, theorists of a conspiracy try to establish a kind of system, synthesizing the world into one big theory.
But there are also differences.
Metaphysicians (as well as many religious texts) at least try to be consistent, while conspiracy theorists are liberated from such limitations.
They openly replace rationality with mythology.
By constructing and emphasizing mythological symbols, conspiracy theories provide a shortcut into our soul, psyche, mind, or the unconscious.
They can switch seamlessly from one theory to another.</p>

<p>Myths can be very dangerous.
They misconstrue associations, destroy nuances and advance subconscious theses without the necessary burden of evidence.
Instead of delivering arguments, they short-circuit the entire argumentative process.
Myths do not make logical claims but significations <a class="citation" href="#barthes:1973">(Barthes, 1973)</a>.
For example, calling someone <em>a snake</em> is not a logical conclusion but signifies deceptive behaviour.
Real snakes, of course, are not significantly more or less deceptive than any other animal.
But mythological snakes often are and calling someone a snake can be a powerful gesture in our culture.</p>

<blockquote>
  <p>Poetry feeds and waters the passions instead of drying them up; she lets them rule, although they ought to be controlled, if mankind are ever to increase in happiness and virtue. – Socrates</p>
</blockquote>

<p>If Trump speaks of <em>America</em> or our radical right-wingers speak of Germany, they do not mean literal countries.
They signify a mythological symbol.
An effective myth is a self-contained world of signs were everything has a marked position, making it very hard to signify otherwise with a believer.
To critique their definition as being racist, irrational, exclusive, inhumane, or disastrous only demontrates to them that you are of ‘the them’ (das Man); that you are jealous that they won.
Any contrary narrative is spun by false prophets which conspire against the ‘chosen ones’.</p>

<h2 id="taking-the-wrong-pill">Taking the Wrong Pill</h2>

<p><em>The Matrix</em> is one of my all-time favorite movies, which increased my interest in computer science and philosophy.
As a child, I fantasized about being <em>The One</em>, akin to the hacker Neo, who could hack the matrix. 
Interestingly, the movie is inspired by French philosopher and theorist of postmodern media and culture, Jean Baudrillard (1929 – 2007), especially by his book <em>Simulacra and Simulation</em> <a class="citation" href="#baudrillard:1983">(Baudrillard, 1983)</a>.
The book even makes an appearance (as an empty prop) at the beginning of the film when Neo gives a disc to his clients.
The actors were even reportedly required to read it.</p>

<p>Baudrillard—the prophet of post-truth—focused on analyzing what can be termed <em>postmodern media</em>, although postmodernity is challenging to define.
Thinkers in the field of postmodern theory frequently hold different opinions but there is one core agreement: there are no all encompassing meta-narratives.
For some, such as Niklas Luhmann (1927–1998), the concept of postmodernity itself is contentious, with Luhmann believing it never truly existed <a class="citation" href="#luhmann:2000">(Luhmann, 2000)</a>. 
But back to Baudrillard.</p>

<p>He was an interesting figure but not taken very seriously by the academic community.
His writing style is polemical and his worldview extremely cynical. 
Despite this, his texts are intriguing and thought-provoking, capturing a sentiment that resonates with many facets of our society today.
In many ways, he was ahead of his time and highly influential in the media and culture he studied.</p>

<p>What particularly makes <em>The Matrix</em> fascinating in connection with Baudrillard is how the film embodies the type of pop-cultural phenomenon he often discussed in his philosophy.
The movie not only reflects his ideas but might be able to bring them to life in a way that is accessible to a broader audience.
Furthermore, Baudrillard was still alive when the movie hit the theatre.
So, did it succeed in bringing his theory to the big screen?</p>

<h3 id="the-simulacrum-is-true">The Simulacrum is True</h3>

<p>When we first meet Neo, his computer is active, processing something, with the screen reflecting on his face. 
He listens to music through headphones while lying on his desk, asleep. 
This scene introduces the difficulty of distinguishing between a dream and reality or more precisely, the problem of informational overload and sensory input that, according to Baudrillard, leads to passivity.
The abundance of disjointed information and excessive transparency makes it nearly impossible to organize the world and assign meaning to it—faces transform into screens or terminals that passively absorb.</p>

<p>In his early career, Baudrillard aimed to merge (post-)Marxism with (post-)structuralism but eventually abandoned the former.
He applied structuralism in his analysis of <em>The System of Objects</em> <a class="citation" href="#baudrillard:1968">(Baudrillard, 1968)</a>. 
Baudrillard theorized that the significance of commodities stems not primarily from their use or exchange value, but rather from their sign value. 
In structuralism, the meaning of elements, such as words, doesn’t derive from what they represent. 
For instance, teaching a child the word ‘tree’ isn’t as simple as pointing to one and stating, “Look, this is a tree!”
The child wouldn’t know if ‘tree’ refers to that specific tree, its leaves, or a category of trees. 
Understanding the word ‘tree’ requires knowledge of many other words and examining their relationships to ‘tree’—their difference.
Baudrillard argues that in a postmodern society, any cultural idea, image, sign, or symbol is apt to be pulled out of its social context and used (or abused) for advertisement and marketing.
The individual is placed in the position of a consumer.
As these signs are lifted out of the social, they lose all possibility of stable reference.
They may be used for anything, for any purpose.
All that remains is a yawning abyss of meaninglessness—a placeless surface that is incapable of holding personal identity, self, or society.</p>

<p>Baudrillard believed that in a postmodern society, the meaning and value of an object are primarily defined by its relationship to other objects. 
Apple products serve as a pertinent example. 
They appear overpriced when considering solely their use value.
However, their value arises from what they signify in relation to other objects which leads to the demishing of <em>symbolic values</em>.
For example, a pen given to you for your graduation, has probably a high symbolic value to you.
Symbolic values are assigned by a subject in relation to another subject.
Sign value, on the other hand, is the object’s value within a system of objects signifying, for example, social status.</p>

<p>Baudrillard, known for his cynical views, also believed that objects essentially have triumphed over subjects. 
He posited that just as money has become a universal medium that renders everything comparable and thus exchangeable, <em>the code</em> has made every sign integratable thus also exchangeable.
It is not that subjects or objects stand no longer for something ‘real’ but that the imagined referent, that does not exist, disappeared.
For example, in the Renaissance people or objects appear to stand for an imagined referent, for instance, royalty, nobility, holiness, etc.
A sign like Iron Man, stands for nothing other than itself in a network of other meaningless signs, i.e. the Marvel universe.
It can be repackaged into a toy or a specific McDonalds meal because it has no sacred connection to the world.
According to Baudrillard, instead of disimulating something, now signs dissimulate that there is nothing.</p>

<blockquote>
  <p>The transition from signs which dissimulate something to signs which dissimulate that there is nothing, marks the decisive turning point. 
The first implies a theology of truth and secrecy (to which the notion of ideology still belongs). 
The second inaugurates an age of simulacra and simulation, in which there is no longer any God to recognize his own, nor any last judgment to separate truth from false, the real from its artificial resurrection, since everything is already dead and risen in advance. – Jean Baudrillard</p>
</blockquote>

<p>This leads to a kind of dissolution of the ‘real’ meaning behind objects.
Accroding to Baudrillard, in the postmodern world, simulacra (e.g. images) have replaced the reality they once represented.
In other words, our current reality is dominated by these simulacra—representations, images, and signs—that no longer have any connection to any real/imagined object or event they might have originally represented.
Importantly, the simulacrum is not just covering up the truth or reality; it’s not a mask over something real!
In fact, it is quite the opposite.
What we perceive as truth or reality is actually just a construct (the simulacrum) that conceals the fact that there is no underlying, original reality; we enter simulation.
In other words, what we consider ‘real’ is just a construct of our perceptions and societal agreement.
In our current postmodern state, the simulacrum has become the truth for us, because there is no other reality against which to measure it.</p>

<blockquote>
  <p>The simulacrum is never what hides the truth—it is the truth that hides the fact that there is none.
The simularcum is true. – Jean Baudrillard</p>
</blockquote>

<p>Following Baudrillard’s perspective, experiences such as a teenager’s first kiss are no longer real in a sense that they express ‘true love’;
instead, they are mere simulations of a Hollywood love story because these stories are the truth!
Imaginations are not destroyed by hiding the truth but by showing it overtly naked, like pornography rips us of the imaginative allure of sexuality and intimacy.
Life imitates advertisement.
This does not mean that there is no more love.
However, there is nothing behind the ‘Hollywood love story’—it is true as it is.</p>

<p>We are compelled to reproduce these images and to participate, even if we know or suspect that it is all a simulation.
Critically, the problem (if it is in fact one) is not a virtualized reality that hides the truth, but that the truth is simulation.
People are fake but they are turthfully fake because being fake is the truth.</p>

<blockquote>
  <p>[…] pretending […] leaves the principle of reality intact: the difference is always clear, it is simply masked, whereas simulation threatens the difference between the ‘true’ and the ‘false’, the ‘real’ and the ‘imaginary’. – Jean Baudrillard</p>
</blockquote>

<p>When more and more simulacra transform into simulation we enter the matrix.
Pictures of ourselves no longer represent us, or us pretending to be someone else, but they are a simulation of some specific and often stereotypical fantasy that has no reference to something real other than different parts of the code.
That is the depressing and cynical viewpoint of Baudrillard.</p>

<h3 id="platos-allegory-of-the-cave">Plato’s Allegory of the Cave</h3>

<p>I love <em>The Matrix</em> but I have to assess that it did not succeed in capturing Baudrillard’s main themes.
The main problem is a clear line between simulation and reality; between the matrix and Zion.
The matrix clearly is not the truth but hides it.
Rather than exploring the new problem of simulation, the movie falls back on the <em>Allegory of the Cave</em> presented in Plato’s <em>Republic</em>.
Instead of investigating further questions, it postulates a true world behind the simulation by re-introducing religion.
Thus <em>The Matrix</em> brings us back where it all started but, according to Baudrillard, this is no longer possible.
In Baudrillard’s framework, the movie is itself the truth that hides the fact that there is none and therefore distracts the audience from acknowledging <em>hyperreality</em>.</p>

<p>The <em>Allegory of the Cave</em> is a metaphor for exploring the nature of knowledge and reality.
Plato imagined a group of people who lived their entire live chained inside a dark cave.
The only thing they can see are the shadows projected on the wall of the cave by objects passing in front of a fire behind them. 
These shadows are the only reality they know.
The cave dwellers believe the shadows to be the real objects, not knowing that these are mere reflections. 
Their knowledge and understanding of the world are based solely on this limited perspective.
One day, a prisoner breaks free. 
He struggles to adjust to the light outside the cave, but eventually, he sees and understands the true nature of reality. 
He realizes that the sun illuminates the world and that what he saw in the cave were just shadows of real objects.
The freed prisoner returns to the cave to enlighten the others. 
However, his eyes have adjusted to the sunlight, so the cave is now blindingly dark to him. 
The other prisoners, unable to understand his experiences and seeing his blindness in the dark, refuse to believe him. 
They cling to their old beliefs about the shadows being the real objects.</p>

<p>In Plato’s metaphysics, the form (true essence of things) are more real than their physical representations.
The shadows represent the physical world, while the objects outside the cave symbolize the forms.
Plato also emphasizes that education is not just a matter of transferring information, but a transformative experience that leads to understanding, or to the seeing of a different world that opens up.
Of course, it is the philosopher that seeks the truth (outside the cave) and then attempts to bring this knowledge back to the people (inside the cave).</p>

<p>We can draw a neat parallel between Plato’s <em>Allegory of the Cave</em> and the film <em>The Matrix</em> by replacing the <em>cave</em> with <em>the matrix</em> and the philosopher with the character Morpheus. 
In this parallel, Morpheus takes on the role of guiding Neo (and others) out of the matrix—similar to leading prisoners out of the cave. 
Neo learns to understand and manipulate the matrix on his terms, which parallels the ability to manipulate the shadows on the cave walls.</p>

<p>In <em>The Matrix</em>, Neo is confronted with a crucial binary decision, symbolized by the choice between a red and a blue pill. 
As revealed in the sequels, this choice is, in itself, a part of a predetermined simulation—a much more Baudrillardian take.
Morpheus presents Neo with this decision and advises him to trust his instinct that something is fundamentally wrong with the world. 
This guidance emphasizes the importance of an emotional rather than a logical conclusion, steering Neo to follow his feelings in making this pivotal choice.</p>

<p>Once Neo makes his choice, the distinction between the matrix and reality is clear to him; it is a clear binary: reality and simulation.
True love is still possible outside and even inside the matrix, even if it is predetermined.
According to Baudrillard reality is a simulation echoing Kant, who does not grant us the access to the <em>thing-in-itself</em>, and of course Nietzsche, who tells us that there are only <em>constructed</em> values.
Simualation is nothing bad or something to fear.
It was always already there but the simulacrum (e.g. cave paintings, images) changed towards its own gravity; towards its own perfection.
What Baudrillard feared is a world akin to the movie <em>Minority Report</em>.
A world without <em>reversability</em> where everything is already decided in advance; a world without <em>ambiguity</em>.
In a sense, Baudrillard feared the modern project that started with Plato by looking for some absolut truth or perfect idea.
The perfect simulation gets rid of illusions and imaginations; it is too real; it is <em>hyperreal</em>;
Therefore, Baudrillard did not fear the loss of reality but an exzess of it which would lead to the destruction of illusions and imaginations like the ‘technical perfection of sex’, i.e. pornography, leads to the removal of sexuality and intimacy.</p>

<blockquote>
  <p>Reality and simulation aren’t opposed to one another. 
There are two sides of the same coin. – <a class="citation" href="#baudrillard:2008">(Baudrillard, 2008)</a></p>
</blockquote>

<p>With this in mind, if we reexamine <em>The Matrix</em> it becomes clear that Baudrillard would describe it to be a pretty good simulation of the matrix.
In a sense, the sign of simulation is re-integrated into the simulation itself.
This re-integration highlights the film’s exploration of reality, perception, and the nature of choice, themes that resonate with Plato’s allegory but not with <em>Simulacra and Simulation</em>.
Thus Baudrillard concluded:</p>

<blockquote>
  <p>The radical illusion of the world is a problem faced by all great cultures, which they have solved through art and symbolization.
What we have invented, in order to support this suffering, is a simulated real, which henceforth supplants the real and is its final solution, a virtual universe from which everything dangerous and negative has ben expelled.
And The Matrix is undeniably part of that.
Everything belonging to the order of dream, utopia and phantasm is given expression, ‘realized’.
We are in the uncut transparency.
The Matrix is surely the kind of film about the matrix that the matrix would have been albe to produce. – Jean Baudrillard</p>
</blockquote>

<h2 id="a-hunger-for-certainty-and-definitude">A Hunger for Certainty and Definitude</h2>

<p>Now, what has this to do with <em>conspiracy theories</em>?
Well, other than Baudrillard’s theory of a reality that is simulation, I claim that the <em>cave allegory</em> offers a theoretical justification for doubting established institutions, which are likened to ‘the matrix’.
One might discover that parts of reality, e.g. institutions, norms, moral judgements, ideologies, is constructed and that there has to be something real behind it.
This suspicion of a matrix is not unfounded, however, believing in some <em>absolut point</em> of view behind it, opens the door to confusion.
We are in a cave but going outside might only lead to another cave.
Importantly, this does not mean that any cave is as useful or functional!
Although not all interpretations of a text are equally meaningful, a good text offers numerous interesting and valuable interpretations and the same seems to be true of our <em>lifeworld</em>.</p>

<p>Doubting parts of reality—a known or presented world—is the starting point of any conspiracy theory.
The perspective is compelling because it feeds the allure of knowing a secret, akin to Neo’s experience in the matrix.
Such knowledge is seen as something that sets an individual apart from ‘the herd’, giving them a sense of being special or enlightened.
It is similar to <em>New Age</em> beliefs in some sort of special knowledge about the universe presented in movies like <em>The Secret</em>.</p>

<p>Belief in conspiracy theories appears to be driven by motives that can be characterized as epistemic (understanding one’s environment), existential (being safe and in control of one’s environment), and social (maintaining a positive image of the self and the social group) <a class="citation" href="#douglas:2017">(Douglas et al., 2017)</a>.
One important facet of conspiracy theories that often goes without much notice is that they are notions about power: who has it and how are they using it?
Conspiracy theories accuse an implicitly powerful group of conspiring.
Usually that group is already powerful—even if that power is a fantasy—i.e., the president, a legislative body, industries or corporations, foreign countries, multinational groups, etc. 
Powerless groups are rarely accused of conspiring <a class="citation" href="#uscinski:2018">(Uscinski, 2018)</a>.
This also reflects the plot of <em>The Matrix</em> where agents of the matrix are much more powerful than ‘enlightened’ humans.</p>

<p>Studies show that some people are more prone to believing in conspiracy theories than others.
Some people will believe in any conspiracy theory even on light evidence while others, at the opposite end of the spectrum, are naive and will deny the existence of conspiracies even on accumulating evidence <a class="citation" href="#uscinski:2018">(Uscinski, 2018)</a>.
According to Jan-Willem Prooijen, conspiracy theories orginate through the same cognitive process that produce other types of belief (e.g. spirituality), they reflect a desire to protect one’s own group against a potentially hostile outgroup, and they are often grounded in strong ideologies.
They are a natural defensive reaction to feelings of uncertainty and fear <a class="citation" href="#prooijen:2018">(Prooijen, 2018)</a>.</p>

<p>Interestingly, in studies, individuals who perceive patterns in abstract paintings, random dots, or coin tosses were more inclined to believe in conspiracy theories, paranormal phenomena, and hold religious beliefs. 
Belief in conspiracies also tends to rise during natural disasters when people feel a lack of control. 
Due to their tendency to seek patterns, conspiracy theorists tend to categorize everything neatly into a framework of good versus evil.
Even though the world that is constructed is miserable, it is without uncertainty.</p>

<p>We are risk calculating creatures, always on the watch for new dangerous patterns.
This is evolutionary advantageous.</p>

<blockquote>
  <p>Conspiracy is a stubborn creed because humans are pattern-seeking animals.
Show us a sky full of stars, and we will arrange them into animals and giant spoons.
Show us a world full of random misery, and we will use the same trick to connect the dots into secret conspiracies. – Jonathan Kay <a class="citation" href="#kay:2011">(Kay, 2011)</a>.</p>
</blockquote>

<p>Perceiving patterns is the opposite of perceiving randomness, and randomness cannot be the basis for making sense. 
Of course, quite often, events occur randomly, without any discernible purpose or meaning. 
Sometimes, foolish mistakes simply happen unintentionally.</p>

<p>Research indicates that education reduces the likelihood of believing in conspiracy theories (with exceptions). 
This may initially appear counterintuitive because education encourages skepticism toward received wisdom. 
Shouldn’t skepticism lead one to think that there might be something hidden behind the scenes?
Well, skepticism is only one aspect of the equation. Education teaches individuals to scrutinize the evidence and seek primary sources, or more precisely, to consider <strong>all</strong> available evidence.
Under such scrutiny conspiracy theories fall apart.
Moreover, having a greater understanding tends to foster humility since individuals become increasingly aware of the vast expanse of knowledge that remains beyond their grasp.</p>

<p>Scrutiny is built into our institutions.
In the context of academic publishing, one must demonstrate a comprehensive understanding of the relevant literature and show how experiments can be reproduced to validate the published findings.
A peer review process by experts in the field checks for the soundness of the work and identifies possible errors. 
While this system is not flawless, makes mistakes, overemphasis the number instead of the value of publications, can be biased, and often favours the middle to upper class, it remains open to critique and has propelled us a long way.
Science is a discourse.
Theories are never absolute true but are seen as true as long as there is no evidence or proof that gives rise to different conclusions.
Mistakes have been made and will continue to be made in the future, but the system is self-correcting, self-preserving, and has advanced our knowledge considerably.</p>

<p>Conspiracy theories are also fueled by our cognitive biases. 
For instance, the <strong>proportionality bias</strong> tends to make us believe that a substantial effect must have a significant cause. 
Consider a scenario where either a neighbor or the President of the United States dies randomly; which one is more likely to trigger a conspiracy theory?
Studies have revealed that when people are informed about the assassination of a president, they are more inclined to believe in a conspiracy theory if it coincides with the outbreak of a subsequent civil war <a class="citation" href="#prooijen:2018">(Prooijen, 2018)</a>.
<strong>Tribalism</strong> encourages us to protect our own ingroup and establish a clear division between ‘us vs. them’, often framing it as a battle between good and evil.
The <strong>intentionality bias</strong> leads us to believe that negative consequences of our actions are unintentional, while attributing intentionality to others when they cause harm. 
For example, we may view bankers as evil, but perceive our own pension fund as a necessary institution.</p>

<p>Seeing patterns everywhere is the need for control <a class="citation" href="#shermer:2022">(Shermer, 2022)</a>.</p>

<blockquote>
  <p>The economy is not this crazy patchwork of supply and demand laws, market forces, interest rate changes, tax policies, business cycles, boom-and-bust fluctuations, recessions and upwings, bull and bear markets, and the like.
Instead, it is a conspiracy of a handful of powerful people variously identified as the Illuminati, the Bilderberger group, the Council on Foreign Relations, the Trilateral Commission, the Rockefellers and Rothshields.
[…] Conspiracists believe that the complex and messy world of politics, economics, and culture can all be explained by a single conspiracy and conspiratorial event that downplays chance and attributes everything to this final end of history. – Michael Shermer</p>
</blockquote>

<p><a class="citation" href="#landau:2015">(Landau et al., 2015)</a> show that people compensate for perceived loss of control by trying to restore control themselves by</p>

<blockquote>
  <p>bolstering personal agency, affiliating with external systems perceived to be acting on the self’s behalf, and affirming clear contingencies between actions and outcomes [… and] seeking out and preferring simple, clear, and consistent interpretations of the social and physical environments.</p>
</blockquote>

<p><strong>Narcissism</strong>, characterized by a belief in one’s superiority and the desire for special treatment, strongly correlates with a tendency to believe in conspiracy theories. 
Narcissists also exhibit heightened sensitivity to perceived threats <a class="citation" href="#cichocka:2022">(Cichocka et al., 2022)</a>.
Within the realm of narcissism, grandiose narcissists seek admiration by bolstering their egos through a sense of uniqueness, charm, and grandiose fantasies. 
It’s worth noting that narcissists often display naivety and are less likely to engage in <strong>cognitive reflection</strong>.
Surprisingly, studies have uncovered evidence suggesting that, contrary to expectations, education increases the likelihood of narcissists adopting conspiracy beliefs <a class="citation" href="#cosgrove:2023">(Cosgrove &amp; Murphy, 2023)</a>. 
This underscores the critical role of cognitive reflection as one of the most, if not the most, essential abilities to guard against narcissistic tendencies towards conspiracy beliefs.</p>

<blockquote>
  <p>[Conspiracy believers] are relatively untrusting, ideologically eccentric, concerned about personal safety, and prone to perceiving agency in action – <a class="citation" href="#hart:2015">(Hart &amp; Graether, 2015)</a></p>
</blockquote>

<p>Similar to the experience of emerging from the cave in Plato’s allegory, delving into a conspiracy theory is not merely a transfer of knowledge, but a transformative experience. 
It involves a world being shattered and a new one being constructed in its place.
This process signifies a profound shift in perception and understanding, where previously accepted realities are dismantled and replaced with an entirely different framework of belief and interpretation.</p>

<p>Like Morpheus’ emphasis on trusting one’s instincts in <em>The Matrix</em>, conspiracy theories often accurately capture the emotional aspects of a person’s situation. 
These theories provide compelling descriptions of emotional states but tend to offer simplistic and reactionary explanations for complex situations. 
Additionally, much like the concept of the matrix, they seek an all-encompassing explanation for everything.
Essentially, these theories represent a futile effort to eliminate contingency and the future’s uncertainty. 
They attempt to provide a sense of certainty and understanding in a world that (hopefully) is still inherently open, contingent, unpredictable and complex. 
This desire for comprehensive explanations reflects a deep-seated human need for order and predictability in an increasingly fatal looking world.</p>

<h2 id="escaping-the-simulation">Escaping the Simulation?</h2>

<p>The incorporation of expressions like ‘escaping the matrix’ and ‘being red-pilled’ into the vocabulary of conspiracy theory groups as metaphorical language is unsurprising.
These terms, which originated from <em>The Matrix</em>, are used metaphorically to describe the experience of awakening to a hidden or suppressed truth.
This desire might increase with the suspicion that there is none.
Specifically, ‘escaping the matrix’ denotes the recognition and liberation from a controlling system or an illusionary world, while ‘being red-pilled’ represents a moment of profound revelation or enlightenment, often regarding societal structures or purported conspiracies.
These metaphors have gained traction within certain groups as a way to express their beliefs in uncovering what they perceive as hidden truths within society.
They believe to be the philosophers of our age, teaching us how real man behave and how ‘the system’ keeps them weak and small.</p>

<p>Contrary to the common belief that ignorance fuels the acceptance of a matrix-like reality, it is curiosity that often propels this belief. 
People are attracted to the notion of discovering hidden truths and understanding the world in ways that differ from the majority’s perspective.
The group forms, in a manner reminiscent of a cult, and establishes easily comprehensible guidelines, resembling the revolutionaries from another pop-cultural and frequently misunderstood film—<em>Fight Club</em>.
The theme of ‘stepping out of the dark’ is central to Plato’s cave allegory, <em>The Matrix</em>, and various conspiracy theories. 
Such curiosity ignites a desire to investigate and question conventional narratives, leading some individuals to adopt alternative interpretations of reality.</p>

<p>Interestingly, critical thinking and intelligence do not necessarily prevent one from falling into this rabbit hole.
As described above, these attributes can sometimes drive narcissits deeper into exploring and accepting these alternate realities—the problem is a lack of reflection.
The quest for understanding and the allure of uncovering hidden knowledge can be so compelling that even the most critical and intelligent minds are susceptible to these alternate explanations.
Take the following speech performed by the actor James Caviezel (an actor I once admired for his role in <em>The Thin Red Line</em>) at the end of <em>Sound of Freedom</em>, a conspitorial movie about child trafficking:</p>

<blockquote>
  <p>While watching this movie, I guess some of you were feeling sad, maybe overwhelmed, or even feel a sense of fear, which is understandable. But living in <strong>fear</strong> isn’t how we solve this problem. It’s living in <strong>hope</strong>. It’s <strong>believing that we can make a difference</strong> because we can.
I want to make one thing clear: this movie you just watched isn’t about me or Tim Ballard. 
It’s about those kids. This film was actually made five years ago. 
It wasn’t released until now, with every roadblock you can imagine being tossed in our way. [The powerful do not want you to see it, believe me]. 
And the names you see here, on the screen, <strong>took a stand</strong>! They made sure that this story could be shown to all of you. Now, all of you have the opportunity to continue <strong>telling this story</strong>; [the Truth].
We don’t have big studio money to market this movie [(we only have Fox News, one of the biggest network in the country)], but we have <strong>you</strong>. The baton has now been passed to <strong>you</strong>. <strong>You</strong> are the storytellers who can get people to come see this film in theaters. Together, we have a chance to make these two kids and the countless children they represent the most powerful people in the world by telling their story in a way only cinema can.
For a couple of months, while Sound of Freedom is in theaters, these kids can be more powerful than the cartel kingpins, presidents, congressmen, or even tech billionaires. We believe this movie has the power to be a huge step forward toward ending child trafficking, but it will only have that effect if <strong>millions of people see it</strong>.
We don’t want finances to be the reason someone doesn’t see this movie, so Angel Studios has set up a forward program where you can pay for someone else’s ticket who might not otherwise see it. If you’re able, we invite you to pay it forward by buying a ticket for someone else, or if your budget is tight, share the already available free ticket with as many friends as you can.
Join us and millions of others as we ring Sound of Freedom and hope throughout the world. And just remember this: <strong>God’s children are not for sale</strong>.</p>
</blockquote>

<p>During the speech a QR Code and the text “Give an Share Tickets / ANGLE.COM/freedom” is displayed.</p>

<p>This speech is filled with pathos, it is shamelessly manipulative, moralistic, heroistic and serves a narcissist desire to ‘wake up’, spread the word of Truth or God and become the hero; a soldier of Freedom; an angle of God; a righteous martyred that saves us all.
It is also conspiratorial.
Of course, in the same breath Caviezel tells us to buy tickets to end child trafficking which is not only irrational but flat out unethical and morally dubious.
It is so obviously a scam that it becomes an interesting field of study why people buy into it.
Does the deeply religious actor James Caviezel believe what he is saying? I think he does on some level.
He explained at a promotion event for the movie that billionaires capture children to extract adrenalin out of their blood when they are scared of death, which is a famous QAnon conspiracy theory.</p>

<blockquote>
  <p>These people that do it; there will be no mercy for them! – James Caviezel</p>
</blockquote>

<p>The speech is also kind of Baudrillardian in that sense that even the fight for the Good is just a struggle to buy that god damn ticket; even God’s angles are reduced to passive consumers.
At the same time it is not about the kids but a much more sacred war of Good against Evil—a mystic war in a very <strong>angry</strong> United States of America!
A country that is at the brink of an inwardly directed outburst.
Everyone in the media is so angry all the time.
This angry speech and movie, that abuses religion, confirms my believe that I should not fear the people who name themselves after the Devil but those who name themselves after a righteous God.
Nietzsche was right about that.</p>

<p>Now, it is a fact that people do engage in conspiracies (and that child trafficking is a big problem).
A brief examination of history reveals numerous instances of conspiracies, some of which have even led to wars between nations. 
However, in retrospect, these conspiracies can often be explained without assuming the involvement of thousands of people. 
The complexity and impact of these historical events do not necessarily require large-scale collusion; often, they can be understood through the actions and decisions of a relatively small number of individuals or groups and through a systemic rationality. 
This understanding helps differentiate between plausible historical conspiracies and the more elaborate, less credible theories that claim widespread secret collaboration.
If it exists, the matrix is not a planned construction of anybody but a Baudrillardian process beyond anyones control.</p>

<p>In addition, it’s important to recognize that the world is inherently unjust. 
Justice is a human concept, one that evolves over time as we make what we call progress. 
However, in the realm of nature, there is no concept of justice at least none I am aware of;
nature operates outside of morality.
Absolute justice remains elusive and if we seek an explanation for the world’s injustice, conspiracy theories provide a sense of comfort.</p>

<p>It appears to me that the belief in having escaped the matrix or emerged from the cave is a strong indication of someone having entrenched themselves deeply in their own perspective—failing to see that their perspective also relies on some sort of <em>second-order observation</em>; a following of the herd, or in Baudrillard’s viewpoint, the false assumption that there is something true behind the simulation.
This belief offers comfort by addressing various uncertainties and the realization of one’s own ignorance. 
No one desires to be ignorant and no one wants to rely on some sort of authority, yet in many ways, we all are.
In our complex world, this is an unavoidable reality. 
Conspiracy theories provide a sense of understanding and control in a world where complete knowledge is unattainable, helping individuals cope with the inherent limitations of human understanding.</p>

<p>Therefore, I believe that individuals who are particularly uncomfortable with uncertainty, who seek control over their life, and who are actively aware of their lack of control, are more susceptible to falling into these rabbit holes of alternate realities. 
Additionally, a certain degree of narcissism may be necessary to believe in the premise that one possesses a superior ability to understand complex matters better than trained experts and to assume that the media consisting of hundred of thousand of journalists is a monolith.
This combination of a need for control, discomfort with uncertainty, and a self-perceived exceptional understanding can lead individuals to embrace alternative explanations that offer a sense of clarity and personal significance in a complex world.</p>

<p>If we contemplate the matrix envisioned by Baudrillard, then attempting to escape it through the immersion in an alternate version of reality, fostered by extensive consumption of social media and digital content, appears absurd. 
Perhaps a more appropriate approach to disengaging from the machinery of simulation and countering the sensation that reality seems increasingly tenuous is to simply disconnect from it all (from time to time).
When advertisements, repetitive media, individuals transformed into brands, and an incessant stream of content seize our attention, the signal overflow—the noise—overshadows a more tangible reality. 
In a scenario where there may be no external escape from the simulation, it could be valuable, from time to time, to focus on what is immediately before us: to experience, touch, smell, listen to our bodies, engage with physical sensations, concentrate, savor awareness, and relinquish the illusion of the “real” by re-connecting to a spiritual world.</p>

<p>We are not superheroes; we are composed of the same fundamental elements as everything else. 
While it might feel like we inhabit a sort of matrix, it’s essential to acknowledge that this is a choice we make. 
From childhood, we develop self-conceptions and fantasies, but this doesn’t negate the existence of the world itself. 
Fantasies are constructs, and doubting the existence of the world presupposes a profound level of experience and knowledge of that world—a world where we learn to eat, walk, dance, and understand the nuances of correct and incorrect language usage.
Doubting it requires a distance from it and that might be what social media does: it shrinks the world but increases the distance to it.
We often employ our habits so routinely that we forget we are employing them, and in doing so, we forget that the world—our home—is still there. 
The central question here is what is more reasonable to doubt: the world we intimately grew up in or our doubts about doubting it?</p>

<h2 id="our-dependence-on-second-order-observation">Our Dependence on Second-order Observation</h2>

<p>Thinking critically and maintaining a sense of curiosity are attributes that I certainly hope everyone possesses.
It is crucial that institutions, including large media operations, research institutions, and particularly governments, are consistently challenged and kept under close scrutiny.
Conspiracy theories can serve as a force to encourage the prevention of corruption.
I would be quite suspicious if there were no conspiracy theories present!</p>

<p>Simultaneously, these theories have the potential to divert attention from genuine issues. 
The public should advocate for accountability and transparency, cultivating a healthy and well-informed society in which decisions and policies undergo scrutiny and improvement through public discourse and critical examination.</p>

<p>However, as the current state of affairs stands, the notion that everyone can participate in the <em>marketplace of ideas</em> and engage in public discourse seems somewhat impractical and utopian. 
This dream may appear overly optimistic, excessively humanistic, and excessively individualistic. 
Instead, according to Luhmann, there exists an interdependent network of social systems that co-evolve together—not individual souls but interconnected systems.</p>

<p>Similarily we demand the media should try to be as objective as they can be but it is naive to think that they are able to present reality as it is.
Here I agree with Luhmann:</p>

<blockquote>
  <p>It is impossible to understand the reality of the mass media if you assume it is their job to provide correct information on the world and then assess how they fail, distort reality, and manipulate opinion—as if they could do otherwise. – Niklas Luhmann</p>
</blockquote>

<p>Basically, Luhmann observed something very similar to the <em>Manufacturing Consent</em> <a class="citation" href="#herman:1988">(Herman &amp; Chomsky, 1988)</a> but explains it slightly differently without the need for a <em>propaganda model</em>.
If one anticipates that the primary purpose of the media is to deliver accurate information or facts, they are likely to encounter inconsistencies that raise significant doubts about the credibility of the mass media apparatus. 
The media is inherently self-preserving.
It functions in a manner that constructs and sustains itself. 
While it is certainly beneficial for the media to provide accurate information, this is not its foremost objective.
The media provides <em>what is known to be known</em>.
It irritates politics, our economy and the scientific system while bing irritated by all these systems.
<strong>The media makes society restless</strong>.</p>

<p>The view that the media presents ‘the Truth’ contributes to the proliferation of conspiracy theories, since it portrays the entire system as corrupt.
If the media presents objective facts and these facts are inconsistent, distorted, incomplete and open for interpretation then it is disfunctional or corrupt.
Instead of recognizing the various shortcomings (which serve the internal logic of the system) within the media (of which there are many), one tends to assume a broad conspiracy aimed at deliberately deceiving the public.
This leads to a pervasive distrust, particularly directed towards well-established media outlets.
As a consequence, consumers may turn to ‘alternative’ media sources, even though these alternatives often inadvertently rely on established media institutions for their information. 
Reporting, conducting on-site investigations, collecting information, and managing extensive archives are expensive endeavors that only large institutions can effectively undertake. 
These institutions are essential if we are to have any hope to share a world that is at least partly commonly known.</p>

<p>The core issue lies in our reliance on what Niklas Luhmann refers to as <em>second-order observation</em> and is consequently a trust issue.
In modern society, directly observing reality is increasingly challenging. 
To stay informed about various aspects such as the state of the economy, job market trends, recent fashion styles, developments in one’s favorite sports league, or new scientific inventions and studies, it is impractical to personally verify these facets. 
Instead, we depend on the observations made by others and, of course, machines.
This means we have to engage with various forms of reporting and analysis: reading reports about the GDP, considering the opinions of fashion critics, watching sports programs, and reviewing scientific papers.
In science, we write review papers about review papers.
We track how often papers are cited, i.e. how these papers are being observed.
This reliance on second-hand information shapes our understanding of the world, as we depend on external observers to provide us with insights and knowledge about various domains that we cannot directly experience or verify ourselves.</p>

<p>We are frequently depend on multiple layers of <em>second-order observations</em> or various levels of abstraction.
Scientific papers serve as a prime example. 
These papers are typically not intended for a general audience but are meant for peers within the specific field of research. 
As a result, the average person often finds them inaccessible.
Consequently, we turn to science communicators and mass media to distill and present scientific information. 
These intermediaries play a crucial role in interpreting and translating complex scientific data and studies into ‘facts’ that are understandable and relevant to the general public. 
This reliance on filtered and simplified interpretations highlights our dependence on external sources to understand and engage with specialized knowledge areas.</p>

<p>At the core of our society is the notion of the individual as a subject capable of making their own decisions and drawing sound conclusions. 
However, it’s evident that our understanding of the numerous processes occurring around us is limited. 
The complexity of the modern world might only be manageable through <em>functional differentiation</em> and <em>second-order observation</em>. 
We are heavily dependent on specialists and experts, and our understanding is largely shaped by observing their observations.</p>

<p>Conspiracy theorists seemingly reject <em>second-order observation</em>, viewing it as a form of manipulation akin to the matrix. 
However, this rejection is a perilous illusion.
There is no position outside of second-order observation, no external vantage point from which to objectively assess ‘reality’ as it is, separate from the interpretations and understandings provided by others.
This perspective underscores the intricate and interconnected nature of knowledge and understanding in contemporary society.</p>

<p>Philosophers ranging from Plato, Fichte, and Kierkegaard to Russell, Kant, and Heidegger have provided insights that prompt us to question the application of second-order observation.
These philosophical teachings encourage us to contemplate whether we should exercise independent thinking and challenge the prevailing mainstream narrative.
This concept is epitomized in Heidegger’s notion of avoiding assimilation into <em>das Man</em> (the they), Kierkegaard’s emphasis on distancing oneself from the public, or Fichte’s focus on the ego. 
The stories goes like this: There exists an inner truth within us, and we should search within our <em>authentic</em> selves to discover it. 
As sovereign individuals in a libertarian society, we should not solely rely on the opinions, or observations, of others. 
As Kant famously articulated, we should have the courage to employ our own intellect (Verstand). 
I align with Kant with a caveat: we should also have the courage to acknowledge our own ignorance and cultivate the ability to rectify it.
By utilizing second-order observation wisely, we can develop a <em>cultural intelligence</em> more akin to Hegel’s concept of the world spirit than Kant’s emphasis on the individual.</p>

<h2 id="conviction-under-constructivism">Conviction under Constructivism</h2>

<p>Biologists Humberto Maturana, Francisco Varela, Samy Frenk, and Gabriela Uribe made a significant discovery regarding our understanding of color perception.
Rather than focusing solely on the correlation between the physical source of color and the retina’s response, they emphasized another more important correlation: the one between the retina and subjective color perception. 
In this context, the external source of color functions as a trigger, not the sole determinant.</p>

<p>This structure of subjective color perception effectively maintains the perception of colors under various objective conditions, even when there are substantial discrepancies between the perceived and ‘emitted’ color, as seen in deception experiments. 
Maturana and Varela extended this insight to introduce the concept of autopoesis, integral to the biological theory of cognition <a class="citation" href="#maturana:1987">(Maturana &amp; Varela, 1987)</a>.</p>

<p>In line with their constructivist theory, every individual constructs their own cognition and, by extension, their reality.
In fact, for Maturana cognition is living.
His theory seems not far from Kant’s and to an extend Fichte’s understanding of cognition.
Constructivism does not imply an absence of a single reality or a state of complete subjectivity, but rather, it leads to some noteworthy conclusions:</p>

<ol>
  <li>An absolute system of values and knowledge cannot exist because personal experience forms an unshakable foundation.</li>
  <li>Convincing someone can only succeed when they develop their own system of conviction.</li>
  <li>Humans, capable of observing their cognitive actions and recognizing the relativity of their seemingly valid knowledge, face the responsibility of choosing and adhering to their own value system.</li>
</ol>

<p>These conclusions have significant relevance to our current discussion. 
Maturana’s framework explains the challenge of debunking conspiracy theories and how individuals can inhabit vastly different realities. 
It also underscores our responsibility to acknowledge our inherently constructed perspective on reality and the value system we embrace.</p>

<p>Now, I should mention that there are numerous critics of constructivism, including Markus Gabriel, who advocates for what he refers to as <em>new realism</em>.
In his book <em>Der Sinn des Denkens</em> (English: <em>The Sense of Thinking</em>) <a class="citation" href="#gabriel:2018">(Gabriel, 2018)</a>, he writes:</p>

<blockquote>
  <p>Constructivism is incorrect. New realism asserts that we can perceive reality as it is, without there being precisely one world or reality that encompasses all objects or facts that exist. – Markus Gabriel</p>
</blockquote>

<p>I read his book with the expectation of finding a plausible justification for why this should be the case, but I couldn’t find any consistent argument. 
Admittedly, this might be unfair as the book is a popular science book and doesn’t aim to provide a rigorous theory. 
Nevertheless, Gabriel often labels assertions as obvious without offering a reasonable explanation. 
Based on what I’ve encountered and also observed in my own life, I lean toward believing in constructivism.</p>

<p>Certainly, this form of relativism raises several pressing issues. 
For instance, how can we justify the actions of individuals whose worldviews may vastly differ from our own?
Additionally, how can international bodies like the United Nations apply pressure on nations that violate human rights if their value systems diverge significantly from a Western dominated notion of values?
Richard Rorty has an intreresting take on that question.
If we want universal acceptance of and respect for human rights, we shouldn’t try to argue about it. 
We shouldn’t attempt to work out rational justifications of human rights, or arguments that will convince people that human rights are a good thing. 
Instead, according to Rorty, we would achieve better results if we try to influence people’s feelings instead of their minds—philosophy as poetry <a class="citation" href="#rorty:2016">(Rorty, 2016)</a>!
Rational justifying human rights is an abstract and philosophical way—something that, according to Rorty, isn’t possible anyway.
In a sense, Rorty suggest that, instead of arguing rationally against mythologies, we should imagine, construct and present better ones.
He does not believe in a second enlightenment—in which logic and rationality will triumph over evil.
It is worth noting that similar challenges and questions arise when we consider the concept of <em>free will</em>, which is itself a highly contentious notion, compare for example <a class="citation" href="#sapolsky:2023">(Sapolsky, 2023)</a>.</p>

<p>In a world with no singular perspective, there are multiple viewpoints coexisting and each has its inherent blind spots. 
According to Luhmann, this principle applies to any observing system, whether it’s a psychological system, like the human mind, or a social system.
The diversity of perspectives inherently limits each view, preventing it from fully encompassing all facets of a situation or concept. 
Observation is blind to its own conditions.
When I observe a tree I can not (at the same time) observe myself observing the tree.
This inherent limitation in observation highlights the intricate and multifaceted nature of comprehending and interpreting the world around us.</p>

<p>The diversity of perspectives among individuals often complicates accurate communication because each of us essentially speaks a slightly different language.
Communication, in itself, can be seen as improbable. 
In addition, language is not something we use to describe reality accurately but a technique we employ to get things done.
Nevertheless, communication remains a crucial element as it plays a central role in stabilizing the chaos and connecting psychic and social systems.
Interestingly, it can be effective even when we don’t fully comprehend each other. 
A prime example of this is ChatGPT, which may not understand as humans do but still manages to communicate effectively.</p>

<p>So, can we embrace and navigate this diversity of perspectives?
Can or should we tolerate the uncanny sensation of numerous distinct realities coexisting? 
Is it possible for us to, to some extent, accept that others may inhabit a differently constructed world while simultaneously acknowledging the existence of something that persists, even if we cease to believe in it.
After all, our constructions do not follow our beliefs—we can not dream the problem away.</p>

<h2 id="the-reality-of-the-climate-crisis">The Reality of the Climate Crisis</h2>

<p>The COVID-19 pandemic had a measurable positive effect on pollution levels—which did not last for long.
However, one could argue that it also had a negative impact on trust levels in institutions, especially scientific ones.
This erosion of trust may ultimately hinder efforts to address the climate crisis.
The ongoing debate regarding climate change’s origins and the necessity of curbing CO2 and equivalent gas emissions continues to persist even if the scientific community is clear on the matter—a conviction I have established via <em>second-order observation</em>.</p>

<p>It appears to me that COVID, combined with the rapid and highly polarizing consumption of “news” on social media, has fractured our social discourse. 
The culture of dialogue has suffered, forcing individuals to align with one of two extreme sides. 
It now seems impossible to critique one party without facing accusations of working for the other.
The language we employ has become more <strong>moralizing</strong>. 
Instead of characterizing people as simply incompetent, misled, misguided, or influenced by flawed incentives within a system, they are often labeled as <strong>evil</strong>. 
This focus on the individual impedes progress in reforming social systems, which are in need of change to provide alternative incentives that prioritize social, ecological, and economic measures for all inhabitants of the planet.
Part of the reality of the climate crisis is that we need trusted institutions that need to be aligned in a way that dealing with the crisis becomes possible.</p>

<p>Emissions are not the only problems on our hand.
Many ecological systems are on the bringe of collapse.
Our agriculture is under threat.
Water shortages are on the horizon.
Increased carbon dioxide absorption by oceans leads to ocean acidification, which can harm marine life, especially coral reefs and shellfish.
Climate change can exacerbate health issues by increasing the spread of diseases, heat-related illnesses, and air quality problems due to wildfires and increased pollen levels.
Changing weather patterns and more frequent extreme events can disrupt agriculture and water supplies, potentially leading to food shortages and conflicts over resources.
As climate impacts worsen, there will be never-seen increased migration and displacement of populations, both within and across borders, as people seek refuge from areas affected by climate-related hazards.
Climate change is very likely to intensify extreme weather events such as hurricanes, droughts, heatwaves, and heavy rainfall. 
These events can lead to increased property damage, displacement of populations, and economic losses.
Sea levels are expected to continue rising, posing a threat to coastal communities, infrastructure, and ecosystems. 
Flooding and saltwater intrusion into freshwater sources may become more common.
I imagine that, at some point, borders of certain countries will be closed, dividing the world in a <em>real</em> and a <em>hyperreal</em> one.</p>

<p>Considering all the points I discussed, the resistance to transitioning away from fossil fuels is expected, given the various parties involved. 
It would be a relief if we were in a simulated reality, where everything could be dismissed as a bad dream.
However, we are faced with the pressing need to convince everyone that climate change is a real problem that demands immediate action.
And that it is worth to sacrifice for the unknown other.
Individuals who have limited information may be persuaded through sound arguments and credible sources. 
However, those who actively reject the mainstream narrative may ultimately question the legitimacy of <em>second-order observation</em>.</p>

<p>An example of this dynamic in action was during a BBC News panel where Brian Cox (physics professor and science communicator) clashed with skeptic Malcolm Roberts (politician). 
Roberts insisted on <em>empirical evidence</em> and rejected <em>appeals to authority</em>. 
All seemed well and logical, but when Cox presented a graph as evidence, Roberts dismissed it, alleging that the data had been corrupted by NASA.
At this juncture, a discussion is no longer possible because there are no external empirical evidence available beyond that produced by scientific institutions.</p>

<p>Undeniable empirical evidence, including temperature records, ice melt data, and rising sea levels, serves as a compelling testament to the tangible effects of climate change. 
To promote a more informed perspective, it is advisable to encourage individuals to explore and critically evaluate reputable, peer-reviewed scientific sources, rather than relying on fringe or biased information.
So, let’s delve into a tiny selection of influential contributions from the scientific community that have shaped our understanding of climate change.
Note that this is only a tiny selection from the whole corpus:</p>

<p>As early as 1896, Svante Arrhenius published a groundbreaking paper on the greenhouse effect, demonstrating how rising concentrations of greenhouse gases lead to an increase in global average surface temperatures <a class="citation" href="#arrhenius:1896">(Arrhenius, 1896)</a>.
Another significant milestone occurred in 1967 when Manabe and Wetherald published the first paper that incorporated the fundamental elements of Earth’s climate into a computer model, exploring the implications of doubling carbon dioxide levels for global temperatures <a class="citation" href="#manabe:1967">(Manabe &amp; Wetherald, 1967)</a>. 
Remarkably, the results of their work remain valid today, according to Prof. Forster.
In 1976, Charles D. Keeling and his team documented a pivotal moment by revealing the sharp rise in carbon dioxide levels at the Mauna Loa observatory in Hawaii <a class="citation" href="#keeling:1976">(Keeling et al., 1976)</a>. 
This paper highlighted the observable increase in atmospheric CO2 resulting from the combustion of carbon, petroleum, and natural gas.
Fast-forwarding to 2006, Held and Soden advanced the concept known as <em>wet-get-wetter, dry-get-drier</em> precipitation in the context of global warming <a class="citation" href="#held:2006">(Held &amp; Soden, 2006)</a>. 
This idea, though occasionally misunderstood and misapplied, remains the first and perhaps the only systematic conclusion regarding regional precipitation and global warming based on a robust physical understanding of the atmosphere.
Additionally, the <em>Intergovernmental Panel on Climate Change</em> (IPCC) reports have played an integral role in consolidating and disseminating crucial climate science findings, further enhancing our collective comprehension of climate change.
If one is convinced that there may be some shadiness going on, I encourage the reader to delve into the extensive history of climate change science, for example <em>The Discovery of Global Warming</em> <a class="citation" href="#weart:2009">(Weart, 2008)</a>.</p>

<p>Additional, one can highlight the overwhelming consensus among climate scientists and scientific organizations that climate change is real and largely caused by human activities, compare <a class="citation" href="#myers:2021">(Myers et al., 2021; Lynas et al., 2021; Cook et al., 2016; Cook et al., 2013; Doran &amp; Zimmerman, 2009)</a>.
Of course, if we only rely on those papers we have to trust an even <em>higher-order observation</em>!
In addition, one can argue that there is no plausible alternative theory apart from the effects humans caused by polluting the planet.</p>

<p>But if an individual has lost trust in institutions, all these efforts will be fruitless.
Therefore, it is so deeply important that our scientific institutions as well as the media defend and improve their reputation.
Without the trust in <em>the other</em>, I see great danger on the horizon, especially if things become increasingly difficult.
The challenge lies in persuading individuals who harbor skepticism, particularly toward what climate change deniers label as <em>mainstream science</em>.</p>

<p>According to Maturana, convincing someone can only happen when they develop <strong>their own system of conviction</strong>. 
This aspect is especially crucial to consider. 
Therefore, engagement must be respectful. 
It’s essential to take the worries, fears, and arguments of deniers seriously, even when they appear unreasonable.
However, it is also important to remember that while we encourage others to develop their convictions, we should also acknowledge our own unique and potentially flawed convictions and remain true to them if we are not convinced otherwise.</p>

<blockquote>
  <p>I think I do it always through stories, never through direct confrontation.
Because if you directly confront somebody who’s thinking polar opposite to you, they don’t really listen.
They are thinking of arguments to refute to. […]
The first thing is to listen to them because maybe they’ve got a point, maybe they’re doing something you never thought about.
But if you still feel that you’re right, then you must have the courage of your conviction. – Jane Goodall</p>
</blockquote>

<p>There is another more systemic problem at hand: There is an entire self-producing industry centered around climate denial. 
Our society has fostered an army of lobbyists whose primary aim is to actively sabotage progress in addressing climate issues. 
In contrast to these ‘knowledgable’ deniers, scientists are required to rigorously justify every aspect of their research repeatedly and tend to be cautious about offering concrete advice. 
Conversely, climate deniers merely need to sow seeds of doubt; their strategy revolves around raising questions rather than providing evidence-based answers. 
This asymmetry in approach creates a challenging environment for advancing scientific understanding and consensus on climate change.</p>

<h2 id="rebels-of-authority">Rebels of Authority</h2>

<p>Apart from ignorance, selfish interests and plain stupidity, I have not yet pointed to the root problem.
I cited papers that suggest that there is a link between pattern matching skills, narcissism and believing in conspiracy theory.
Stupid people elect stupid and corrupt politicians, right?
But that is too easy of an explanation.
Whenever I have to fall back on stupidity or evilness, I get the feeling that I am missing something.
Most of the time peope are not evil or stupid.
Similar to conspiracy theories, such a rational is too simple, too individualistic and also too dangerous.
So what is the systemic reason why people reject actions required to keep the planet sustainable for us all?</p>

<p>Of course, there are many complex reasons but I want to focus on one that fits the meat of this article.
I more or less successfully tried to problematize the search for an absolute objective truth by pointing out that those who believe in stepping out of the cave go probably deeper into it.
And I pointed to Baudrillard’s imagined eradication of imagination leading to a <em>true simulacra</em> that no longer stands for anything but itself—a description of <em>hyperreality</em> that hits me emotionally.
Now, let us imagine that conspiracy theorists rightly point to a problem they might not really understand rationally but feel intuitively.
What might it be?</p>

<p>If we want to describe our modern Western society, I think it’s fair to say that we are a <em>knowledge-based society</em>.
Since Foucault’s historical analysis, we know that knowledge and power are linked. 
In our daily life, most of our decisions are informed by some scientifically produced piece of knowledge.
For example, our diet is informed by scientific research that gives us guidance to stay healthy.
How often or for what reason we go to the doctor is informed by science.
In fact, we wittness a whole industry of self-optimization that claims to be scientific.
There are cults trying to establish a ‘science’ of finding a breeding mate.
Or take this article.
My goal and strategy of achieving it is similar: I employ knowledge and second-order observation by citing scientific papers.
In that sense, I fall into the same trap, that is, I try to convince my opponents by displaying superious knowledge.
Since Foucault</p>

<p>In such a society, we tend to search for the <em>perfect algorithm</em> that can make the best decisions for any situation.
In fact, many decisions are already made by algorithms based on the observation of large amounts of data.
Even policies are crafted by utilizing artificial intelligence.
The idea is simple: instead of shouting at each other about the right course of action, let <em>objective reality</em> be the final judge; let ‘the Truth’ decide—let science guide us through the mess.
This is the dream born out of the Enlightenment and it sparked many ideologies that promised to fulfill this utopian harmony.
It is supported by the <em>mechanistic view</em> on the universe suggesting that we can eventually understand all causal relations going back to some final cause we might call God;
In its current interpretation, I call it <strong>algocracy</strong> <a class="citation" href="#danaher:2016">(Danaher, 2016)</a>, i.e., <em>rule by algorithms</em>.</p>

<p>Interstingly, most conspiracy theorist argue within this framework, that is, they argue not in terms of values but in terms of better knowledge.
At the same time they rebel against the authority of such algorithmic rigor.
Therefore, on the one hand, they believe in a complete understanding of reality or at least a big and very complex chunk of it.
Consequently, they also buy the idea that politics is basically determined by the most powerful authority: the truth.
On the other hand, they rebel against this doctrine by using its own logic.
I think here lies a great danger because <strong>the system harms itself by its own operation</strong>.
The system enters a process of selfdestruction.</p>

<p>Like Pippi Longstocking, conspiracy theorists create their world or imagine it based on an alternative system of turth.
In the case of Pippi Longstocking, we find this antiauthoritarian attitude charming, creative, and imaginative, but in the case of conspiracy theorists, many find it dangerous, idiotic, and ignorant—an inconsistency worth investigating.</p>

<p>According to the sociologist Alexander Bogner an overdose of  antiauthoritarian attitude becomes problematic if we no longer agree on any foundation making any conflict impossible since there is nothing to argue about.
One of such foundation is the belief in an objective truth.
He pragmatically argues that:</p>

<blockquote>
  <p>For libaral democracies, the idea of objective truth is a necessary fiction. – <a class="citation" href="#bogner:2021">(Bogner, 2021)</a></p>
</blockquote>

<p>This statement seems reasonable but I will problematize it later.
Bogner points out that a liberal democratic society requires conflicts grounded in dissent.
However, dissent is only possible if there is at least a shared foundation and also an alternative to think about.
Overall, I think Bogner’s essay aligns closely with the constructivists viewpoint shared by e.g. Luhmann even if he criticizes the theory for attacking the idea of an objective truth.
For example, he argues in a Luhmannian manner about the root problem of conspiracy theories, or what many call the <em>post-truth society</em>, pointing out the crossing of boundaries between two interdependent but operationally closed systems:
<strong>science</strong> and <strong>politics</strong>.</p>

<p>Due to the <em>functional differentiation</em> of modern societies, politics communicates about <strong>values</strong> and science communicates about <strong>facts</strong>.
Bogner argues that today this distinction is blurred leading to dysfunctional systems.
Nowadays, politics communicates about facts but is still guided by values.
This makes <em>productive dissent</em> difficult because different value systems no longer compete within the boundary of values but hide behind the battle for better knowledge.
Values are no longer up for debate.
Instead we argue for superior facts.
In this context, <em>fake news</em> take part in a new <em>language game</em> that is born out from the lack of alternatives.
Fake news do not attack values but knowledge itself.</p>

<blockquote>
  <p>This could also be observed during the coronavirus crisis: 
Due to the high pressure of scientification, the will for fundamental opposition in some places was discharged through the spread of ‘alternative facts’.
[…] It was directed against a (supposedly) authoritative instance that claimed to determine what is real, rational, and politically necessary through superior rationality.
From the perspective of this protest, political emancipation could only be an emancipation from the facts.
Alternative facts are evidently in vogue when politics (due to its alignment with science) appears to be without alternatives. – <a class="citation" href="#bogner:2021">(Bogner, 2021)</a></p>
</blockquote>

<p>Bogner correclty points out that scientific knowledge can be dangerous for politics because it dismantles or defuses the discourse.
It is also attractive if one wants to surrender responsibility because in science it is far easier to agree on statements since it follows a strict true-false distinction.
In this sense, science is authoritarian.</p>

<p>Take a simple example such as health.
Before talking about it, we already share certain values.
For example, most people probably assume that a long healthy life is preferable over a short or unhealthy life.
This seems reaonable.
However, even this seemingly simple assumption can not be proven.
It is a value and values are mutually constructed.
I can, for example, easily argue that it is better to live to the fullest even if this means that my life expectancy drops.</p>

<p>In politics we argue about values and try to find some common ground.
This requires some degree of cohesion.
It goes against our more and more individualistc togetherness.
Bogner points out that people are considerably constrained in their expression of dissent if scientific truth is translated into a politic that knows no alternative.
His statement resonates with the relation between the belief in conspiracy theory and narcism.
He writes:</p>

<blockquote>
  <p>The struggle against science and experts can thus be understood as a struggle against determinations that were not chosen by oneself.
For such determinations must appear as an impudent imposition to the individualized, modern person, who is increasingly called upon today as a self-responsible shaper of their destiny, as a self-entrepreneur or ‘Me Inc’. 
Therefore, the struggle against facts is, not least, a struggle for autonomy.
From this perspective, science denial appears as a critique that is in tune with the times, insofar as it derives its plausibility from the current conditions of subjectification.
In the protest against established knowledge, the disappointed hope of the highly individualized, activated subject for complete sovereignty and a fully comprehensible, decision-open world is discharged. – <a class="citation" href="#bogner:2021">(Bogner, 2021)</a></p>
</blockquote>

<p>If politicians give knowledge absolute authority over political decisions, and these decisions realize certain values, the only way to fight for different values is to fight a different truth.
Using Bogner’s perspective, fake news are not the cause of the problem but a symptom of a society that feels impotent because certain values are always already presupposed to be the right values.
Therefore, what Bogner calls <em>productive dissent</em> becomes impossible.</p>

<blockquote>
  <p>The rebellion of the knowledge deniers can overall be understood as a covert appeal directed against a (looming) colonization of politics by expert consensus, 
regardless of how reasonable the expert recommendations may be in individual cases. – <a class="citation" href="#bogner:2021">(Bogner, 2021)</a></p>
</blockquote>

<p>I agree with Bogner that a culture needs a common foundation, a reference world, which, according to Luhmann, should be provided by the media.
And we also need a more flexible and ambiguous reference world on a global scale, see <a class="citation" href="#bauer:2018">(Bauer, 2018)</a>.
Furthermore, I agree that science cannot and should not realize politics and that better knowledge does not automatically lead to better politics.
Science should continue to operate on the true-false distinction, while politics should use a different code that differentiates values.</p>

<p>In case of climate activism, it might be a far more effective strategy to demand the values people want to be realized instead of discussing the facts they believe in.
Instead of calling people deniers, it might be healthier to find and employ strategies that reveal their value system and how this value system supports certain actions.
We should reinvigorate the discourse about values since it is save to say that more and better facts alone cannot move a society towards change.</p>

<p>However, I do not agree with Bogner’s claim that we need to agree that there is an <em>objective truth</em> to be found. 
I even suspect he does not fully believe this himself, as he refers to the <strong>idea of an objective truth as a necessary fantasy</strong>, akin to <em>Plato’s noble lie</em>—a myth knowingly propagated by an elite to maintain social harmony.
I reject this in favor of an embracing uncertainty because, as I argued above, I think it drives the individual deeper into the cave. 
As the philosopher Slavoj Žižek argues—contrary to Dostoevsky and many others—–that the belief in God enables us to commit horrific crimes, believing in an objectively accessible truth can serve the same purpose.
If we no longer live in the same world, society might break, but this does not diminish the great value of doubt and uncertainty in restraining our tendency for megalomania.</p>

<p>On this matter, I lean more towards Rorty, believing that we are mature enough to handle the pragmatist view that it is socially more useful to reject the idea of an objective truth (in the Platonist sense).
According to Rorty, Pippi Longstocking should be allowed to argue for her position, and we should accept her position to be true if it emerges out of the struggle for truth.
The American pragmatist rejects the Platonist’s notion of truth, considering it unintelligible and meaningless.</p>

<blockquote>
  <p>For the idea of a liberal society, it is of central importance that everything is allowed as long as it pertains to words as opposed to actions, to persuasion as opposed to violence. 
This openness should not be maintained because, as the Bible says, the truth is great and will prevail, nor because, as Milton believes, in a free and open fight, the Truth will always win. A society is liberal when it is content to call ‘true’ whatever emerges as the result of such struggles. – <a class="citation" href="#rorty:1989">(Rorty, 1989)</a></p>
</blockquote>

<p>Many, including Bogner, believe that such an attitude leads to <strong>indifference</strong>. 
He argues that if we take Rorty’s comment literally, we would bury both the idea of <em>objective truth</em> and our <em>liberal democracy</em>, since the realization of individual freedom through social change requires productive dissent.
I think Bogner, like many others, misinterprets Rorty or at least applies his statement to the wrong level of society—the wrong reality, so to speak. 
The implication that we end up in a sort of absolute relativism or complete indifference if we stop believing in an objective truth remains unfounded. 
It’s a huge leap.
Even if we reject the idea of an objective truth, empirical evidence is still valuable for discussions.</p>

<p>Rejecting the belief in an objective truth does not mean that we lose the ground for productive dissent.
On the contrary, it allows for a more pragmatic approach to dialogue and disagreement.
When we drop this belief, we recognize that our conversations and disagreements are grounded in our contingent vocabularies and cultural practices.
This recognition does not lead to indifference but rather to a greater appreciation of the diversity of perspectives.
It encourages us to engage with others’ values and beliefs not because they correspond to an objective reality, but because they are part of the <strong>shared human experience</strong>.
Productive dissent arises from the acknowledgment that our beliefs and values are fallible and open to revision through conversation.
It promotes a kind of solidarity where we strive to understand and negotiate our differences rather than impose a supposed objective truth.
This process fosters mutual respect and a more inclusive society, where different viewpoints can coexist and contribute to a richer, more nuanced understanding of the world.
By focusing on the practical consequences of our beliefs and actions, we can still engage in meaningful debates about what kind of society we want to create.
This pragmatic approach encourages us to find common ground and work together to address shared problems, rather than being paralyzed by the quest for an elusive objective truth.</p>

<h2 id="conclusion">Conclusion</h2>

<p>What I wanted to emphasize is the idea that we are existing within a subjective and distorted world of a reality that we can not directly access.
We are always already in a simulation, relying on second-order observation.
What I call <em>angry media</em> outlets, many of which take part in spreading misinformation and conspiracy theories, use more and more frequently phrases like <em>literally</em>, <em>this is simply the truth</em>, <em>they all know it</em>, <em>its a fact</em>.
They fight against ambiguity dragging anything into a culture war such that it has to be defined using one of two perspectives.
Thus the root cause of descending into the rabbit hole may not be a detachment from reality but rather an <strong>attachment to certainty</strong>. 
Perhaps Baudrillard is correct in asserting that we are entering a <em>hyperreality</em> that is no longer contradictory and dissolves all illusions, imaginations, and mysteries.
According to him, what we are trying to do is to purge the world of all mysteries, illusions and imaginations.
But certainty exists only in a pure form of simulation. 
In a constructivist sense, we cannot access the absolute true reality (Kant’s <em>the thing in itself</em>) because it is always already mediated.
If we cannot accept this fundamental ambiguity of our reality (by which Baudrillard does not mean physical reality but that which is intelligible via signs), we run the risk of constructing the one and only reality, i.e., hegemony.
Since this realm does not allow contradictions, it tends to integrate everything, including disasters, into it.</p>

<p>Contradictions form the very foundation of our environment. 
We perceive the reality of nature when it manifests as a non-human force of destruction, such as a natural disaster or a pandemic. 
What occurs is incomprehensible. 
In hyperreality, nature ceases to be a conflicting force and becomes merely an element within the simulation—a floating sign. 
Destruction transforms into a calculated event.
Repeatedly displaying graphs of death tolls gives us the illusion of control.
It can be seen as an attempt to reintegrate death (arguably the greatest contradiction of all) into the simulation.
The negative, along with contradictions, is either integrated or discriminated against.
Wars, natural disasters, and pandemics metamorphose into a spectacle on the television screen, a tourist attraction for our theme park, and are more disastrous than the disaster, more natural than nature—in short, a perfect simulation that surpasses and supplants reality.</p>

<p>In a postmodern society, the absence of any central value system and firm, objective evaluative guides tends to create a demand for substitutes. 
These substitutes are symbolically created rather than being actual or socially produced. 
The need for these symbolic group tokens results in tribal politics and defines self-constructing practices that are collectivized but not socially produced. 
These neo-tribes function solely as imagined communities and, unlike their premodern namesake, exist only in symbolic form through the commitment of individual ‘members’ to the idea of an identity. 
They exist as imagined communities through a multitude of agent acts of self-identification and endure solely because people use them as vehicles of self-definition; as an identity technology termed <em>profilicaty</em>.</p>

<p>I have no intention of passing moral judgment on conspiracy theorists.
They often risk a significant amount of social capital, leading to alienation from their relatives and friends.
Being a conspiracy theorist is generally not an enjoyable experience.
As I said, they rebel against authority via the same use of authority—a conflict of knowledge against ‘better’ knowledge in the realm of politics.
This is dangerous.
It points to a deeper problem of a violation of the <em>operational clousure</em> of social systems which we should take serious if our goal is to preserve these systems.
One might even argue that conspiracy theorists perceive cracks in the simulation but mistakenly believe in a way out of it, which, and this misinterprets <em>The Matrix</em> as well, serves as the perfect cover-up for our reality as always partly simulated.</p>

<p>Baudrillard famously argued that Disneyland does not hide the fact that it is a simulation, but rather conceals the fact that America is a simulation too. 
Disneyland is more real than America.
Conspiracy theories operate similarly; they are pure simulations and, in this regard, true. 
The outsider is convinced of his or her reality, and the contradictory nature of these theories is not contradictory for him or her. 
Instead, (obvious) contradictions are necessary to make the theory hyperreal. 
The suspicion of the theorists is not unreasonable, but their conclusion is fatal—they demand a simple metanarrative and cannot see that this can only be another far more harmful simulation.</p>

<p>Several factors contribute to the prevalence of conspiracy theories, including information overload, the perception of a reality that is becoming ‘less real’, the sensation of living in a quasi-simulated reality, natural disasters, ongoing conflicts, and the rapid pace of our society. 
Our modern world is so complex that it is virtually impossible for any single individual to comprehensively make sense of all that occurs.
As a result, we heavily rely on the concept of second-order observation, and there is no shame in acknowledging this fact.
Turning to experts and authorities, provided that their authority is derived from genuine competence, is necessary. 
However, it’s crucial to subject these authorities to scrutiny and verification and to be aware that their observation has always a blind spot.</p>

<p>Nothing in what I’ve stated here should be misconstrued as a defense of the political system. 
Lobbyism, which is sometimes indistinguishable from outright corruption, represents a significant issue. 
The consistent failure to fulfill promises, whether they be pledges for a transaction tax, the cessation of subsidies that actively contribute to global warming, or the numerous ‘conferences’ like COP that, at this point, are merely part of a <em>hope economy</em>—offering a false and pacifying sense of hope—that fuels the distrust in institutions.</p>

<p>In my view, COP28 proved to be a disaster for the majority of the world’s population. 
There was essentially no consensus on even the most basic measures. 
No accord on phasing out fossil fuels, and not even a genuine commitment to promoting renewable energy. 
The decision to have COP led by one of the world’s largest oil and gas companies is, at this point, satirical if it were not true.
In an assessment by journalist Jonathan Watts in The Guardian, the winners of the conference were identified as the oil and gas industry, the United States, China, COP28 President Sultan Al Jaber, the green energy sector, and lobbyists. Conversely, the losers encompassed the climate, small island nations, climate justice, future generations, other species, and scientists.</p>

<p>Nothing in what I’ve said here should be interpreted as a defense of the media for its issues, nor should it downplay the problem of increasing wealth inequality or any other ecological, social, or economic problems.
Both mass media and the scientific system have their problems, but this doesn’t negate their incredible value.
Moreover, independent media outlets play a crucial role, but we must be under no illusion that the verification of information is becoming increasingly challenging. 
Social media platforms present us with an overwhelming array of viewpoints, claims, and video content. With the ascent of generative artificial intelligence, the task of verifying this deluge of information becomes even more daunting.
The power of the image is indeed huge. 
The blurred line between hyperreality and lower forms of simulations makes it difficult to navigate through the mass of information. In contrast, conspiracy theorists do not bear the burden of a demanding verification process. 
They can simply draw upon fringe and unvalidated stories, presenting them in an entertaining, sensational style akin to news pornography. 
They can use the power of high-order simulacra which are disconnected from the real.</p>

<p>Believing in a conspiracy theory is akin to being the prisoner in Plato’s cave, presuming that everyone else is, in fact, in prison. 
It is the belief in something outside of simulation.
The most effective remedy is to harbor doubts about our own competence, to be skeptical of ourselves, to maintain self-awareness at a metacognitive level, and to be able to live in a contradictory world and recognize those contradictions.
These contradictions live on the borders of hyperreality—in slums, cobalt mines, the streets of New York City, the border of Mexico, the fortress of the European sea, refugee camps, and the ‘ugly’ parts of the world.</p>

<p>I am not sure if I can agree with the cynical viewpoint of Baudrillard. 
His overly dramatic and playful writings are interesting but also contradictory, probably by design.
He would probably be horrified at our attempt to <a href="https://blogs.nvidia.com/blog/earth-2-supercomputer/">simulate the whole earth</a> to predict and control our future. 
But how else can we deal with an open, unpredictable future other than the pursuit of more and more accurate predictions? Esposito asks if this future will still be open <a class="citation" href="#esposito:2024">(Esposito, 2024; Esposito et al., 2023)</a>.</p>

<p>It appears to me that logical reasoning and providing ‘better’ facts alone is not sufficiently compelling.
Science and technology is not enough.
It feels like we lack spiritual growth.
Maybe Rorty is right about the importance of empathic stories.
Maybe science should provide us with that what we call truth and politics should offer us a story that moves us.
But where are these stories?
Are we already as cynical as Baudrillard?
We need storytellers and artists to craft more persuasive mythologies, narratives, and stories that resonate on an emotional level.
I believe this can only be possible if we do not filter out the negative and the ugly part of society.
We have to be serious yet playful and imaginative.
I want a serious vision which is shamelessly emphatic towards all forms of life.
This anti-Platonistic approach, while potentially controversial, could prove more effective in influencing beliefs and behaviors and, in the end, matter more than any rational argument could be.</p>

<p>We have reached an alarming point at which millions of people can no longer discriminate between reality and hyperreality. 
That which cannot be simulated seems to disappear. 
There is a confusion between truth claims grounded in evidence and sound logic and alternative facts inspired by an authoritarian rebellion against the authority of knowledge.
Even in the case of the climate crisis, if we want to preserve a liberal democracy, politics has to discuss alternatives based on the consideration of different values.
We need imagination to draw larger circles and rationalty to make valid moves within these circles.
We have to trust in the scientific method to provide us with good knowledege, politics (not politicians) that is informed by science to provide us with good politics, and spirituality that may spend us wisdom and psychologic stability.</p>

<p>Most importantly, instead of appealing to objective truth, it is more productive to focus on the practical consequences and shared values that can mobilize action.
Instead of arguing that climate change is an objective truth that everyone must accept, we should emphasize the practical and tangible consequences of climate inaction.
By highlighting concrete consequences, we may make a compelling case for action based on the observable and lived experiences of people.
We should frame the argument for climate action in terms of shared values and common interests.
The goal is to find common ground and motivate people to act based on their own interests and values, rather than trying to convince them of an abstract objective truth.
Acknowledging the contingency and fallibility of our beliefs does not mean we cannot act decisively.
We can adopt policies and actions that are based on the best available evidence while remaining open to revising them as new information emerges. 
While rejecting objective truth might complicate the traditional ways of arguing for action, it also opens up new avenues for persuasion and coalition-building.</p>

<p>Sooner or later the reality of the crisis will eventually bleed into hyperreality.
Even if we construct our own perspective on the world, the reality of the climate crisis will not disappear if we stop believing in it.</p>

<h2 id="references">References</h2>

<ol class="bibliography"><li><span id="popper:1934">Popper, K. (1934). <i>Logik der Forschung</i>.</span></li>
<li><span id="barthes:1973">Barthes, R. (1973). <i>Mythologies</i>. Hill &amp; Wang Pub.</span></li>
<li><span id="baudrillard:1983">Baudrillard, J. (1983). <i>Simulacra and Simulation</i>. Semiotext(e).</span></li>
<li><span id="luhmann:2000">Luhmann, N. (2000). Why does society describe itself as postmodern. In W. Rasch &amp; C. Wolfe (Eds.), <i>Observing complexity: Systems theory and postmodernity</i> (pp. 35–49). University of Minnesota.</span></li>
<li><span id="baudrillard:1968">Baudrillard, J. (1968). <i>System of Objects</i>.</span></li>
<li><span id="baudrillard:2008">Baudrillard, J. (2008). <i>The Perfect Crime</i>. Verso.</span></li>
<li><span id="douglas:2017">Douglas, K. M., Sutton, R. M., &amp; Cichocka, A. (2017). The psychology of conspiracy theories. <i>Current Directions in Psychological Science</i>, <i>26</i>(6), 538–542. https://doi.org/10.1177/0963721417718261</span></li>
<li><span id="uscinski:2018">Uscinski, J. E. (2018). The study of conspiracy theories. <i>Argumenta</i>, 233–245. https://doi.org/10.23811/53.arg2017.usc</span></li>
<li><span id="prooijen:2018">Prooijen, J.-W. (2018). <i>The Psychology of Conspiracy Theories</i>. Taylor &amp; Francis Group. https://doi.org/10.4324/9781315525419</span></li>
<li><span id="kay:2011">Kay, J. (2011). <i>Among the Truthers: A Journey Through America’s Growing Conspiracist Underground</i>. Harper.</span></li>
<li><span id="shermer:2022">Shermer, M. (2022). <i>Conspiracy: Why the Rational Believe the Irrational</i>. Johns Hopkins University Press.</span></li>
<li><span id="landau:2015">Landau, M. J., Kay, A. C., &amp; Whitson, J. A. (2015). Compensatory control and the appeal of a structured world. <i>Psychol Bull</i>. https://doi.org/10.1037/a0038703</span></li>
<li><span id="cichocka:2022">Cichocka, A., Marchlewska, M., &amp; Biddlestone, M. (2022). Why do narcissists find conspiracy theories so appealing? <i>Curr Opin Psychol</i>. https://doi.org/10.1016/j.copsyc.2022.101386</span></li>
<li><span id="cosgrove:2023">Cosgrove, T. J., &amp; Murphy, C. P. (2023). Narcissistic susceptibility to conspiracy beliefs exaggerated by education, reduced by cognitive reflection. <i>Front Psychol</i>. https://doi.org/10.3389/fpsyg.2023.1164725</span></li>
<li><span id="hart:2015">Hart, J., &amp; Graether, M. (2015). Something’s going on here: Psychological predictors of belief in conspiracy theories. <i>Journal of Individual Differences</i>.</span></li>
<li><span id="herman:1988">Herman, E. S., &amp; Chomsky, N. (1988). <i>Manufacturing Consent</i>. Pantheon Books.</span></li>
<li><span id="maturana:1987">Maturana, H. R., &amp; Varela, F. J. (1987). <i>The Tree of Knowledge</i>. Shambhala.</span></li>
<li><span id="gabriel:2018">Gabriel, M. (2018). <i>Der Sinn des Denkens</i>. Ullstein Buchverlag.</span></li>
<li><span id="rorty:2016">Rorty, R. (2016). <i>Philosophy as Poetry</i>. University of Virginia Press.</span></li>
<li><span id="sapolsky:2023">Sapolsky, R. M. (2023). <i>Determined</i>. Bodley Head.</span></li>
<li><span id="arrhenius:1896">Arrhenius, S. (1896). XXXI. On the influence of carbonic acid in the air upon the temperature of the ground. <i>The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science</i>, <i>41</i>(251), 237–276. https://doi.org/10.1080/14786449608620846</span></li>
<li><span id="manabe:1967">Manabe, S., &amp; Wetherald, R. T. (1967). Thermal equilibrium of the atmosphere with a given distribution of relative humidity. <i>Journal of Atmospheric Sciences</i>, <i>24</i>(3), 241–259. https://doi.org/10.1175/1520-0469(1967)024&lt;0241:TEOTAW&gt;2.0.CO;2</span></li>
<li><span id="keeling:1976">Keeling, C. D., Bacastow, R. B., Bainbridge, A. E., Ekdahl Jr., C. A., Guenther, P. R., Waterman, L. S., &amp; Chin, J. F. S. (1976). Atmospheric carbon dioxide variations at Mauna Loa Observatory, Hawaii. <i>Tellus</i>, <i>28</i>(6), 538–551. https://doi.org/10.1111/j.2153-3490.1976.tb00701.x</span></li>
<li><span id="held:2006">Held, I. M., &amp; Soden, B. J. (2006). Robust responses of the hydrological cycle to global warming. <i>Journal of Climate</i>, <i>19</i>(21), 5686–5699. https://doi.org/10.1175/JCLI3990.1</span></li>
<li><span id="weart:2009">Weart, S. R. (2008). <i>The Discovery of Global Warming</i>. Harvard University Press.</span></li>
<li><span id="myers:2021">Myers, K. F., Doran, P. T., Cook, J., Kotcher, J. E., &amp; Myers, T. A. (2021). Consensus revisited: quantifying scientific agreement on climate change and climate expertise among Earth scientists 10 years later. <i>Environmental Research Letters</i>, <i>16</i>(10), 104030. https://doi.org/10.1088/1748-9326/ac2774</span></li>
<li><span id="lynas:2021">Lynas, M., Houlton, B. Z., &amp; Perry, S. (2021). Greater than 99% consensus on human caused climate change in the peer-reviewed scientific literature. <i>Environmental Research Letters</i>, <i>16</i>(11), 114005. https://doi.org/10.1088/1748-9326/ac2966</span></li>
<li><span id="cook:2016">Cook, J., Oreskes, N., Doran, P. T., Anderegg, W. R. L., Verheggen, B., Maibach, E. W., Carlton, J. S., Lewandowsky, S., Skuce, A. G., Green, S. A., Nuccitelli, D., Jacobs, P., Richardson, M., Winkler, B., Painting, R., &amp; Rice, K. (2016). Consensus on consensus: a synthesis of consensus estimates on human-caused global warming. <i>Environmental Research Letters</i>, <i>11</i>(4), 048002. https://doi.org/10.1088/1748-9326/11/4/048002</span></li>
<li><span id="cook:2013">Cook, J., Nuccitelli, D., Green, S. A., Richardson, M., Winkler, B., Painting, R., Way, R., Jacobs, P., &amp; Skuce, A. (2013). Quantifying the consensus on anthropogenic global warming in the scientific literature. <i>Environmental Research Letters</i>, <i>8</i>(2), 024024. https://doi.org/10.1088/1748-9326/8/2/024024</span></li>
<li><span id="doran:2009">Doran, P. T., &amp; Zimmerman, M. K. (2009). Examining the scientific consensus on climate change. <i>Eos, Transactions American Geophysical Union</i>, <i>90</i>(3), 22–23. https://doi.org/https://doi.org/10.1029/2009EO030002</span></li>
<li><span id="danaher:2016">Danaher, J. (2016). The Threat of Algocracy: Reality, Resistance and Accommodation. <i>Philosophy &amp; Technology</i>, <i>29</i>, 245–268. https://doi.org/10.1007/s13347-015-0211-1</span></li>
<li><span id="bogner:2021">Bogner, A. (2021). <i>Die Epistemisierung des Politischen</i>. Reclam.</span></li>
<li><span id="bauer:2018">Bauer, T. (2018). <i>Die Vereindeutigung der Welt</i> (p. 104). Reclam.</span></li>
<li><span id="rorty:1989">Rorty, R. (1989). <i>Contingency, Irony, and Solidarity</i>. Cambridge University Press.</span></li>
<li><span id="esposito:2024">Esposito, E. (2024). Can we use the open future? Preparedness and innovation in times of self-generated uncertainty. <i>European Journal of Social Theory</i>, <i>0</i>(0). https://doi.org/10.1177/13684310231224546</span></li>
<li><span id="esposito:2023">Esposito, E., Hofmann, D., &amp; Coloni, C. (2023). Can a predicted future still be an open future? Algorithmic forcasts and actionability in the precision medicine. <i>History and Theory</i>. https://doi.org/10.1111/hith.12327</span></li></ol>]]></content><author><name>Benedikt Zönnchen</name><email>benedikt.zoennchen@web.de</email></author><category term="Social Systems Theory" /><category term="Conspiracy Theories" /><category term="Climate Crisis" /><summary type="html"><![CDATA[Diverging from my area of expertise is always a risky endeavor, but since this is a blog and not a scientific journal, I’m giving myself the liberty to explore and have fun with different ideas (even if the topic is depressing). Often writing helps in transforming the mess into a structured and coherent concept. The process of rethinking and reflecting can be invaluable. It helps to make ones thought anschlussfähig which literally means to be capable for connections and in this context means enabling the continuation of communication.]]></summary></entry><entry><title type="html">Musical Interrogation III - LSTM</title><link href="https://bzoennchen.github.io/Pages/2023/11/19/musical-interrogation-III.html" rel="alternate" type="text/html" title="Musical Interrogation III - LSTM" /><published>2023-11-19T00:00:00+01:00</published><updated>2023-11-19T00:00:00+01:00</updated><id>https://bzoennchen.github.io/Pages/2023/11/19/musical-interrogation-III</id><content type="html" xml:base="https://bzoennchen.github.io/Pages/2023/11/19/musical-interrogation-III.html"><![CDATA[<p>This article is the continuation of a series.
It is recommended that you read part I and II first.
This time we use a recurrent neural network (<strong>RNN</strong>), more precisely an <strong>LSTM</strong>, which I explained a little bit in the <a href="/Pages/2023/04/02/musical-interrogation-I.html">introduction</a>.
An LSTM is a RNN that counteracts the problem of exploding and vanishing gradients.</p>

<p>This article gives some explanation to the code in the following <a href="https://github.com/BZoennchen/musical-interrogation/blob/main/partIII/melody_rnn.ipynb">notebook</a>, which can be executed on <a href="https://colab.research.google.com/?hl=de">Google Colab</a>.
Because our model is now able to learn long-time relations, we can use the <code class="language-plaintext highlighter-rouge">GridEncoder</code>, i.e., a <em>piano roll data representation</em>, which is exactly what we do.</p>

<h2 id="recurrent-neural-networks">Recurrent Neural Networks</h2>

<p>Prior to the advent of transformers, many cutting-edge natural language processing (NLP) applications relied on recurrent neural networks (RNNs). 
These networks are particularly adept at processing sequences of data, such as words in NLP tasks or notes and musical events in audio processing.</p>

<p>An RNN processes a sequence</p>

\[\mathbf{x}_0, \ldots, \mathbf{x}_{n-1}\]

<p>of inputs one at a time. 
With each new input, the network not only considers this current input but also incorporates a <em>hidden state</em>—a representation of previous inputs—thanks to its recurrent connections.
This hidden state \(\mathbf{h}_t\) is updated at each step, ensuring that the network retains a memory of what it has processed so far.</p>

<p><br /></p>
<div><img style="display:block; margin-left:auto; margin-right:auto; width:80%;" src="/Pages/assets/images/rnn-unfold.png" alt="Sketch of an RNN unfolded in time" />
<div style="display: table;margin: 0 auto;">Figure 1: Sketch of an RNN unfolded in time.</div>
</div>
<p><br /></p>

<p>The unique feature of RNNs is their ability to maintain an internal state that captures information about the sequence they have processed to that point.
As a result, the final output of the RNN is informed by the entire input sequence.</p>

<p>Furthermore, RNNs have been employed in sequence generation or decoding tasks. 
In such applications, the tokens generated by the RNN are fed back into it as inputs. 
This feedback loop allows the RNN to generate sequences where each new token is influenced by the previously generated tokens, making it suitable for tasks like text generation, music composition, and more.</p>

<h2 id="long-short-term-memory-networks">Long Short-Term Memory Networks</h2>

<p>Long short-term memory networks (LSTMs) are a type of RNN that were designed to overcome some of the limitations of vanilla RNNs, particularly in handling long-term dependencies in sequence data.</p>

<p>Vanilla RNNs struggle with learning long-term dependencies due to the <em>vanishing gradient problem</em>. 
As the length of the input sequence increases, the gradients used in the training process can become extremely small, making it difficult for the RNN to learn and retain information from earlier inputs. 
LSTMs address this issue with their unique architecture, which includes <em>memory cells</em> which uses different <em>gates</em>.</p>

<p>A <em>memory cell</em> can maintain information in memory for long periods of time. 
The key components of an such a cell are its gates: the <em>input gate</em>, <em>output gate</em>, and <em>forget gate</em>.
These gates regulate the flow of information into and out of the cell, and they decide what to retain or discard from the cell state.</p>

<ul>
  <li><strong>Update Gate</strong>: Determines how much of the new information to add to the cell state.</li>
  <li><strong>Forget Gate</strong>: Decides what information is no longer needed and removes it from the cell state, helping to prevent the accumulation of irrelevant information.</li>
  <li><strong>Output Gate</strong>: Controls the extent to which the value in the cell is used to compute the output activation of the block.</li>
</ul>

<div><img style="display:block; margin-left:auto; margin-right:auto; width:80%;" src="/Pages/assets/images/lstm-cell.png" alt="LSTM cell" />
<div style="display: table;margin: 0 auto;">Figure 2: A sketch of a memory cell.</div>
</div>
<p><br /></p>

<p>Due to their architecture, LSTMs can learn and remember over longer sequences than vanilla RNNs, making them more effective for tasks like language modeling, text generation, speech recognition, and more, where understanding context over a long sequence is crucial.
The gating mechanism helps mitigate the vanishing gradient problem, allowing for more effective training over longer sequences.
This is because the gates allow gradients to flow through the network without being multiplied repeatedly by small numbers (which is what causes the gradients to vanish in vanilla RNNs).</p>

<h2 id="data-preparation">Data Preparation</h2>

<p>Again we use the data from <a href="http://kern.ccarh.org">EsAC</a>. 
The specific dataset I utilized is <a href="https://kern.humdrum.org/cgi-bin/ksdata?l=/essen/europa&amp;format=recursive">Folksongs from the continent of Europe</a> and for the purpose of this work, I will exclusively use the 1700 pieces found in the <code class="language-plaintext highlighter-rouge">./deutschl/erk</code> directory.</p>

<p>We assume that 1/16 is the shortest note in our dataset.
The <code class="language-plaintext highlighter-rouge">GridEncoder</code> automatically filters out pieces that do not fulfill this condition.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">time_step</span> <span class="o">=</span> <span class="mi">1</span><span class="o">/</span><span class="mi">16</span>
<span class="n">encoder</span> <span class="o">=</span> <span class="n">GridEncoder</span><span class="p">(</span><span class="n">time_step</span><span class="p">)</span>
<span class="n">enc_songs</span><span class="p">,</span> <span class="n">invalid_song_indices</span> <span class="o">=</span> <span class="n">encoder</span><span class="p">.</span><span class="n">encode_songs</span><span class="p">(</span><span class="n">scores</span><span class="p">)</span>
<span class="k">print</span><span class="p">(</span><span class="sa">f</span><span class="s">'there are </span><span class="si">{</span><span class="nb">len</span><span class="p">(</span><span class="n">enc_songs</span><span class="p">)</span><span class="si">}</span><span class="s"> valid songs and </span><span class="si">{</span><span class="nb">len</span><span class="p">(</span><span class="n">invalid_song_indices</span><span class="p">)</span><span class="si">}</span><span class="s"> songs'</span><span class="p">)</span>
</code></pre></div></div>

<p>Let us look at an example encoded of a piece:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>55 _ _ _ 60 _ _ _ 60 _ _ _ 60 _ _ _ 60 _ _ _ 64 _ _ _ 64 _ _ _ r _ _ _ 62 _ 64 _ 65 ...
</code></pre></div></div>

<p>As we discussed in the last article, <code class="language-plaintext highlighter-rouge">55 _ _ _</code> stands for the midinote <code class="language-plaintext highlighter-rouge">55</code> played for 4 beats where one beat is 1/16 note.
Therefore, this is a 1/4 note.
Likewise, <code class="language-plaintext highlighter-rouge">r _ _ _</code> is a 1/4 rest.</p>

<p>Next, the <code class="language-plaintext highlighter-rouge">StringToIntEncoder</code> converts our alphabet of tokens into positive integers.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">string_to_int</span> <span class="o">=</span> <span class="n">StringToIntEncoder</span><span class="p">(</span><span class="n">enc_songs</span><span class="p">)</span>
</code></pre></div></div>

<p>Next, we use <code class="language-plaintext highlighter-rouge">ScoreDataset</code> to arrange our training data.
It requires our encoded songs, the instance of <code class="language-plaintext highlighter-rouge">StringToIntEncoder</code> and a <em>hyperparameter</em> <code class="language-plaintext highlighter-rouge">sequence_len</code> that configures the length of token sequences our model will be trained on.
The longer the sequence, the longer the training will require because the deeper the recurrent neural network will be.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">sequence_len</span> <span class="o">=</span> <span class="mi">64</span> <span class="c1"># this is a hyperparameter!
</span><span class="n">dataset</span> <span class="o">=</span> <span class="n">ScoreDataset</span><span class="p">(</span>
    <span class="n">enc_songs</span><span class="o">=</span><span class="n">enc_songs</span><span class="p">,</span> 
    <span class="n">stoi_encoder</span><span class="o">=</span><span class="n">string_to_int</span><span class="p">,</span> 
    <span class="n">sequence_len</span><span class="o">=</span><span class="n">sequence_len</span><span class="p">)</span>
</code></pre></div></div>

<p>It is now possible to split our data into training, validation and test set.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">train_set</span><span class="p">,</span> <span class="n">val_set</span><span class="p">,</span> <span class="n">test_set</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">utils</span><span class="p">.</span><span class="n">data</span><span class="p">.</span><span class="n">random_split</span><span class="p">(</span><span class="n">dataset</span><span class="p">,</span> <span class="p">[</span><span class="mf">0.8</span><span class="p">,</span> <span class="mf">0.1</span><span class="p">,</span> <span class="mf">0.1</span><span class="p">])</span>
</code></pre></div></div>

<h2 id="model-definition">Model Definition</h2>

<p>First we define the rest of our <em>hyperparameters</em>:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">vocab_size</span> <span class="o">=</span> <span class="nb">len</span><span class="p">(</span><span class="n">string_to_int</span><span class="p">)</span> <span class="c1"># size of our alphabet
</span><span class="n">input_dim</span> <span class="o">=</span> <span class="n">vocab_size</span> <span class="c1"># can be different
</span><span class="n">hidden_dim</span> <span class="o">=</span> <span class="mi">128</span> <span class="c1"># can be different
</span><span class="n">layer_dim</span> <span class="o">=</span> <span class="mi">1</span> <span class="c1"># can be different
</span><span class="n">output_dim</span> <span class="o">=</span> <span class="n">vocab_size</span> <span class="c1"># should not be different
</span><span class="n">dropout</span> <span class="o">=</span> <span class="mf">0.2</span> <span class="c1"># can be different
</span>
<span class="n">criterion</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">nn</span><span class="p">.</span><span class="n">CrossEntropyLoss</span><span class="p">()</span>

<span class="n">learning_rate</span> <span class="o">=</span> <span class="mf">0.001</span> <span class="c1"># can be different
</span><span class="n">batch_size</span> <span class="o">=</span> <span class="mi">64</span> <span class="c1"># can be different
</span><span class="n">n_epochs</span> <span class="o">=</span> <span class="mi">10</span> <span class="c1"># can be different
</span><span class="n">eval_interval</span> <span class="o">=</span> <span class="mi">100</span> <span class="c1"># can be different
</span></code></pre></div></div>

<p>Before explaining every detail, let us look at the model definition first.
The following is the model description of our RNN/LSTM.
To understand what’s going on, look at the forward method.
This sends our data through the network.</p>

<p>The first two lines create the short-term \(\mathbf{h}_0\) and long-term memory \(\mathbf{c}_0\) and fill them with zeros.</p>

<p>Then an embedding takes place: <code class="language-plaintext highlighter-rouge">x = self.embedding(x)</code>.
This is nothing more than what we did with our simple feedforward net in <a href="/Pages/2023/05/31/musical-interrogation-II.html">Part II - FNN</a>: Each element of the input <code class="language-plaintext highlighter-rouge">x</code> is first one-hot encoded and then multiplied by a matrix. 
The result: Each event is represented by the row of a matrix (with learnable parameters).
The matrix has <code class="language-plaintext highlighter-rouge">vocab_size</code> rows and <code class="language-plaintext highlighter-rouge">input_dim</code> columns.</p>

<p>Next, we send our transformed input through our LSTM out, <code class="language-plaintext highlighter-rouge">(ht, ct) = self.lstm(x, (h0, c0))</code>.
This basically computes \(\mathbf{h}_t, \mathbf{c}_t\) based on \(\mathbf{h}_{t-1}, \mathbf{c}_{t-1}\) as indicated in Fig. 2.
We get as many outputs as our sequence is long, i.e., <code class="language-plaintext highlighter-rouge">sequence_len</code> many.
But we are only interested in the last output, which we get by <code class="language-plaintext highlighter-rouge">out[:, -1, :]</code>.
This is a vector with <code class="language-plaintext highlighter-rouge">hidden_dim elements</code>. 
We don’t need <code class="language-plaintext highlighter-rouge">ht</code> and <code class="language-plaintext highlighter-rouge">ct</code>.</p>

<p>Then we send the last output through a dropout layer to counteract <em>overfitting</em>.</p>

<p>In the last step, we transform the <code class="language-plaintext highlighter-rouge">hidden_dim</code>-dimensional vector into an <code class="language-plaintext highlighter-rouge">output_dim</code>-dimensional vector, which is equal to <code class="language-plaintext highlighter-rouge">vocab_size</code>.
This vector is interpreted as a probability distribution.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">class</span> <span class="nc">LSTMModel</span><span class="p">(</span><span class="n">torch</span><span class="p">.</span><span class="n">nn</span><span class="p">.</span><span class="n">Module</span><span class="p">):</span>
    <span class="k">def</span> <span class="nf">__init__</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">input_dim</span><span class="p">,</span> <span class="n">hidden_dim</span><span class="p">,</span> <span class="n">layer_dim</span><span class="p">,</span> <span class="n">output_dim</span><span class="p">,</span> <span class="n">dropout</span><span class="o">=</span><span class="mf">0.2</span><span class="p">):</span>
        <span class="nb">super</span><span class="p">(</span><span class="n">LSTMModel</span><span class="p">,</span> <span class="bp">self</span><span class="p">).</span><span class="n">__init__</span><span class="p">()</span>

        <span class="bp">self</span><span class="p">.</span><span class="n">hidden_dim</span> <span class="o">=</span> <span class="n">hidden_dim</span>
        <span class="bp">self</span><span class="p">.</span><span class="n">layer_dim</span> <span class="o">=</span> <span class="n">layer_dim</span>
        
        <span class="bp">self</span><span class="p">.</span><span class="n">embedding</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">nn</span><span class="p">.</span><span class="n">Embedding</span><span class="p">(</span><span class="n">vocab_size</span><span class="p">,</span> <span class="n">input_dim</span><span class="p">)</span>
        <span class="bp">self</span><span class="p">.</span><span class="n">lstm</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">nn</span><span class="p">.</span><span class="n">LSTM</span><span class="p">(</span><span class="n">input_dim</span><span class="p">,</span> <span class="n">hidden_dim</span><span class="p">,</span> <span class="n">layer_dim</span><span class="p">,</span> <span class="n">batch_first</span><span class="o">=</span><span class="bp">True</span><span class="p">)</span>
        <span class="bp">self</span><span class="p">.</span><span class="n">dropout</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">nn</span><span class="p">.</span><span class="n">Dropout</span><span class="p">(</span><span class="n">dropout</span><span class="p">)</span>
        <span class="bp">self</span><span class="p">.</span><span class="n">fc</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">nn</span><span class="p">.</span><span class="n">Linear</span><span class="p">(</span><span class="n">hidden_dim</span><span class="p">,</span> <span class="n">output_dim</span><span class="p">)</span>
        
    <span class="k">def</span> <span class="nf">forward</span><span class="p">(</span><span class="bp">self</span><span class="p">,</span> <span class="n">x</span><span class="p">):</span>
        <span class="n">h0</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">zeros</span><span class="p">(</span><span class="bp">self</span><span class="p">.</span><span class="n">layer_dim</span><span class="p">,</span> <span class="n">x</span><span class="p">.</span><span class="n">size</span><span class="p">(</span><span class="mi">0</span><span class="p">),</span> <span class="bp">self</span><span class="p">.</span><span class="n">hidden_dim</span><span class="p">,</span> <span class="n">device</span><span class="o">=</span><span class="n">device</span><span class="p">)</span>
        <span class="n">c0</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">zeros</span><span class="p">(</span><span class="bp">self</span><span class="p">.</span><span class="n">layer_dim</span><span class="p">,</span> <span class="n">x</span><span class="p">.</span><span class="n">size</span><span class="p">(</span><span class="mi">0</span><span class="p">),</span> <span class="bp">self</span><span class="p">.</span><span class="n">hidden_dim</span><span class="p">,</span> <span class="n">device</span><span class="o">=</span><span class="n">device</span><span class="p">)</span>
        
        <span class="c1"># x = B, T, C
</span>        <span class="n">x</span> <span class="o">=</span> <span class="bp">self</span><span class="p">.</span><span class="n">embedding</span><span class="p">(</span><span class="n">x</span><span class="p">)</span>
        
        <span class="n">out</span><span class="p">,</span> <span class="p">(</span><span class="n">ht</span><span class="p">,</span> <span class="n">ct</span><span class="p">)</span> <span class="o">=</span> <span class="bp">self</span><span class="p">.</span><span class="n">lstm</span><span class="p">(</span><span class="n">x</span><span class="p">,</span> <span class="p">(</span><span class="n">h0</span><span class="p">,</span> <span class="n">c0</span><span class="p">))</span>
        <span class="n">out</span> <span class="o">=</span> <span class="bp">self</span><span class="p">.</span><span class="n">dropout</span><span class="p">(</span><span class="n">out</span><span class="p">[:,</span> <span class="o">-</span><span class="mi">1</span><span class="p">,</span> <span class="p">:])</span>
        <span class="n">out</span> <span class="o">=</span> <span class="bp">self</span><span class="p">.</span><span class="n">fc</span><span class="p">(</span><span class="n">out</span><span class="p">)</span>
        <span class="k">return</span> <span class="n">out</span> <span class="c1"># B, C
</span></code></pre></div></div>

<p>Next, we initialize the model:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">model</span> <span class="o">=</span> <span class="n">LSTMModel</span><span class="p">(</span><span class="n">input_dim</span><span class="p">,</span> <span class="n">hidden_dim</span><span class="p">,</span> <span class="n">layer_dim</span><span class="p">,</span> <span class="n">output_dim</span><span class="p">,</span> <span class="n">dropout</span><span class="p">)</span>
<span class="n">model</span><span class="p">.</span><span class="n">to</span><span class="p">(</span><span class="n">device</span><span class="p">)</span>  <span class="c1"># use gpu if possible
</span>
<span class="n">optimizer</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">optim</span><span class="p">.</span><span class="n">Adam</span><span class="p">(</span><span class="n">model</span><span class="p">.</span><span class="n">parameters</span><span class="p">(),</span> <span class="n">lr</span><span class="o">=</span><span class="n">learning_rate</span><span class="p">)</span>

<span class="k">for</span> <span class="n">i</span> <span class="ow">in</span> <span class="nb">range</span><span class="p">(</span><span class="nb">len</span><span class="p">(</span><span class="nb">list</span><span class="p">(</span><span class="n">model</span><span class="p">.</span><span class="n">parameters</span><span class="p">()))):</span>
    <span class="k">print</span><span class="p">(</span><span class="nb">list</span><span class="p">(</span><span class="n">model</span><span class="p">.</span><span class="n">parameters</span><span class="p">())[</span><span class="n">i</span><span class="p">].</span><span class="n">shape</span><span class="p">)</span>
</code></pre></div></div>

<p>We could play with different hyperparameters.
Increasing <code class="language-plaintext highlighter-rouge">hidden_dim</code> basically increases the complexity of the “memory” of the LSTM.
We surely want to increase <code class="language-plaintext highlighter-rouge">n_epochs</code> to increase number of times the LSTM “sees” all training data.</p>

<p>We can visualize the LSTM by utilizing the <code class="language-plaintext highlighter-rouge">draw_graph</code> function from the <code class="language-plaintext highlighter-rouge">torchview</code> package.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1"># (batch_size, sequence_len)
</span><span class="n">X_vis</span><span class="p">,</span> <span class="n">y_vis</span> <span class="o">=</span> <span class="n">train_set</span><span class="p">[</span><span class="mi">0</span><span class="p">:</span><span class="n">batch_size</span><span class="p">]</span>
<span class="k">print</span><span class="p">(</span><span class="sa">f</span><span class="s">'shape of X_vis: </span><span class="si">{</span><span class="n">X_vis</span><span class="p">.</span><span class="n">shape</span><span class="si">}</span><span class="s">'</span><span class="p">)</span>
<span class="k">print</span><span class="p">(</span><span class="sa">f</span><span class="s">'shape of y_vis: </span><span class="si">{</span><span class="n">y_vis</span><span class="p">.</span><span class="n">shape</span><span class="si">}</span><span class="s">'</span><span class="p">)</span>
<span class="k">print</span><span class="p">(</span><span class="sa">f</span><span class="s">'number of different symbols </span><span class="si">{</span><span class="n">vocab_size</span><span class="si">}</span><span class="s">'</span><span class="p">)</span>
<span class="n">X_vis</span><span class="p">,</span> <span class="n">y_vis</span> <span class="o">=</span> <span class="n">X_vis</span><span class="p">.</span><span class="n">to</span><span class="p">(</span><span class="n">device</span><span class="p">),</span> <span class="n">y_vis</span><span class="p">.</span><span class="n">to</span><span class="p">(</span><span class="n">device</span><span class="p">)</span>
<span class="n">model_vis</span> <span class="o">=</span> <span class="n">LSTMModel</span><span class="p">(</span><span class="n">input_dim</span><span class="p">,</span> <span class="n">hidden_dim</span><span class="p">,</span> <span class="n">layer_dim</span><span class="p">,</span> <span class="n">output_dim</span><span class="p">,</span> <span class="n">dropout</span><span class="p">)</span>
<span class="n">model_graph</span> <span class="o">=</span> <span class="n">draw_graph</span><span class="p">(</span><span class="n">model_vis</span><span class="p">,</span> <span class="n">input_data</span><span class="o">=</span><span class="n">X_vis</span><span class="p">,</span> <span class="n">device</span><span class="o">=</span><span class="n">device</span><span class="p">)</span>
<span class="n">model_graph</span><span class="p">.</span><span class="n">visual_graph</span>
</code></pre></div></div>

<div><img style="display:block; margin-left:auto; margin-right:auto; width:40%;" src="/Pages/assets/images/lstm-model.png" alt="LSTM model" />
<div style="display: table;margin: 0 auto;">Figure 4: The architecture of our LSTM model using a batch size of 64 and a sequence length also equal to 64. The alphabet consists of 38 unique tokens. Each single input is hot-encoded into a vector with 38 components. The LSTM uses a hidden state with 128 components. After the dropout the 128 components of hidden state are reduced to 38 components utilizing a normal linear layer (without an activation function).</div>
</div>
<p><br /></p>

<p>Note that the softmax is part of our loss <code class="language-plaintext highlighter-rouge">criterion</code> i.e. the cross entropy loss <code class="language-plaintext highlighter-rouge">torch.nn.CrossEntropyLoss()</code> which is part of the backpropagation, i.e., the training process.</p>

<h2 id="melody-generation-before-training">Melody Generation (Before Training)</h2>

<p>Given a sequence of arbitrary length, the <code class="language-plaintext highlighter-rouge">generate</code> function is used to generate a new piece of music.
<code class="language-plaintext highlighter-rouge">temperature</code> determines how much the probability distribution learned by the model is considered.</p>

<ul>
  <li><code class="language-plaintext highlighter-rouge">temperature</code> equal to 1.0 means that sampling is done from the probability distribution.</li>
  <li><code class="language-plaintext highlighter-rouge">temperature</code> approaching infinity means that sampling is done uniformly (more variation).</li>
  <li><code class="language-plaintext highlighter-rouge">temperature</code> approaching 0 means that higher probabilities are emphasized (less variation).</li>
</ul>

<p>We can set a maximum length for the piece and also provide the beginning of a piece.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">def</span> <span class="nf">next_event_number</span><span class="p">(</span><span class="n">idx</span><span class="p">,</span> <span class="n">temperature</span><span class="p">:</span><span class="nb">float</span><span class="p">):</span>
    <span class="k">with</span> <span class="n">torch</span><span class="p">.</span><span class="n">no_grad</span><span class="p">():</span>
        <span class="n">logits</span> <span class="o">=</span> <span class="n">model</span><span class="p">(</span><span class="n">idx</span><span class="p">)</span>
        <span class="n">probs</span> <span class="o">=</span> <span class="n">F</span><span class="p">.</span><span class="n">softmax</span><span class="p">(</span><span class="n">logits</span> <span class="o">/</span> <span class="n">temperature</span><span class="p">,</span> <span class="n">dim</span><span class="o">=</span><span class="mi">1</span><span class="p">)</span> <span class="c1"># B, C
</span>        <span class="n">idx_next</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">multinomial</span><span class="p">(</span><span class="n">probs</span><span class="p">,</span> <span class="n">num_samples</span><span class="o">=</span><span class="mi">1</span><span class="p">)</span>
        <span class="k">return</span> <span class="n">idx_next</span>

<span class="k">def</span> <span class="nf">generate</span><span class="p">(</span><span class="n">seq</span><span class="p">:</span> <span class="nb">list</span><span class="p">[</span><span class="nb">str</span><span class="p">]</span><span class="o">=</span><span class="bp">None</span><span class="p">,</span> <span class="n">max_len</span><span class="p">:</span><span class="nb">int</span><span class="o">=</span><span class="bp">None</span><span class="p">,</span> <span class="n">temperature</span><span class="p">:</span><span class="nb">float</span><span class="o">=</span><span class="mf">1.0</span><span class="p">):</span>
    <span class="k">with</span> <span class="n">torch</span><span class="p">.</span><span class="n">no_grad</span><span class="p">():</span>
        <span class="n">generated_encoded_song</span> <span class="o">=</span> <span class="p">[]</span>
        <span class="k">if</span> <span class="n">seq</span> <span class="o">!=</span> <span class="bp">None</span><span class="p">:</span>
            <span class="n">idx</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">tensor</span><span class="p">(</span>
                <span class="p">[[</span><span class="n">string_to_int</span><span class="p">.</span><span class="n">encode</span><span class="p">(</span><span class="n">char</span><span class="p">)</span> <span class="k">for</span> <span class="n">char</span> <span class="ow">in</span> <span class="n">seq</span><span class="p">]],</span> 
                <span class="n">device</span><span class="o">=</span><span class="n">device</span>
            <span class="p">)</span>
            <span class="n">generated_encoded_song</span> <span class="o">=</span> <span class="n">seq</span><span class="p">.</span><span class="n">copy</span><span class="p">()</span>
        <span class="k">else</span><span class="p">:</span>
            <span class="n">idx</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">tensor</span><span class="p">([[</span><span class="n">string_to_int</span><span class="p">.</span><span class="n">encode</span><span class="p">(</span><span class="n">TERM_SYMBOL</span><span class="p">)]],</span> <span class="n">device</span><span class="o">=</span><span class="n">device</span><span class="p">)</span>
        
        <span class="k">while</span> <span class="n">max_len</span> <span class="o">==</span> <span class="bp">None</span> <span class="ow">or</span> <span class="n">max_len</span> <span class="o">&gt;</span> <span class="nb">len</span><span class="p">(</span><span class="n">generated_encoded_song</span><span class="p">):</span>
            <span class="n">idx_next</span> <span class="o">=</span> <span class="n">next_event_number</span><span class="p">(</span><span class="n">idx</span><span class="p">,</span> <span class="n">temperature</span><span class="p">)</span>
            <span class="n">char</span> <span class="o">=</span> <span class="n">string_to_int</span><span class="p">.</span><span class="n">decode</span><span class="p">(</span><span class="n">idx_next</span><span class="p">.</span><span class="n">item</span><span class="p">())</span>
            <span class="k">if</span> <span class="n">idx_next</span> <span class="o">==</span> <span class="n">string_to_int</span><span class="p">.</span><span class="n">encode</span><span class="p">(</span><span class="n">TERM_SYMBOL</span><span class="p">):</span>
                <span class="k">break</span>
            <span class="n">idx</span> <span class="o">=</span> <span class="n">torch</span><span class="p">.</span><span class="n">cat</span><span class="p">((</span><span class="n">idx</span><span class="p">,</span> <span class="n">idx_next</span><span class="p">),</span> <span class="n">dim</span><span class="o">=</span><span class="mi">1</span><span class="p">)</span> <span class="c1"># B, T+1, C
</span>            <span class="n">generated_encoded_song</span><span class="p">.</span><span class="n">append</span><span class="p">(</span><span class="n">char</span><span class="p">)</span>
            
        <span class="k">return</span> <span class="n">generated_encoded_song</span>
</code></pre></div></div>

<p>Of course, the results are almost random because the parameters of our model are initialized randomly and we did not train it yet.
The following code snippet generates 5 scores.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1"># number of songs we want to generate
</span><span class="n">n_scores</span> <span class="o">=</span> <span class="mi">5</span>
<span class="n">temperature</span> <span class="o">=</span> <span class="mf">0.6</span>
<span class="n">before_new_songs</span> <span class="o">=</span> <span class="p">[]</span>
<span class="k">for</span> <span class="n">_</span> <span class="ow">in</span> <span class="nb">range</span><span class="p">(</span><span class="n">n_scores</span><span class="p">):</span>
    <span class="n">encoded_song</span> <span class="o">=</span> <span class="n">generate</span><span class="p">(</span><span class="n">max_len</span><span class="o">=</span><span class="mi">13</span><span class="p">,</span><span class="n">temperature</span><span class="o">=</span><span class="n">temperature</span><span class="p">)</span>
    <span class="k">print</span><span class="p">(</span><span class="sa">f</span><span class="s">'generated </span><span class="si">{</span><span class="s">" "</span><span class="p">.</span><span class="n">join</span><span class="p">(</span><span class="n">encoded_song</span><span class="p">)</span><span class="si">}</span><span class="s"> consisting of </span><span class="si">{</span><span class="nb">len</span><span class="p">(</span><span class="n">encoded_song</span><span class="p">)</span><span class="si">}</span><span class="s"> notes'</span><span class="p">)</span>
    <span class="n">before_new_songs</span><span class="p">.</span><span class="n">append</span><span class="p">(</span><span class="n">encoded_song</span><span class="p">)</span>
</code></pre></div></div>

<p>Let’s listen to the first one:</p>

<audio controls="">
  <source src="/Pages/assets/audio/before_g_song.mp3" type="audio/mp3" />
  Your browser does not support the audio element.
</audio>

<h2 id="training">Training</h2>

<p>For training, we use something called a <code class="language-plaintext highlighter-rouge">DataLoader</code>. 
This helps us to access our data more easily. 
For example, we shuffle our data before training.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">train_loader</span> <span class="o">=</span> <span class="n">DataLoader</span><span class="p">(</span><span class="n">train_set</span><span class="p">,</span> <span class="n">batch_size</span><span class="o">=</span><span class="n">batch_size</span><span class="p">,</span> <span class="n">shuffle</span><span class="o">=</span><span class="bp">True</span><span class="p">)</span>
<span class="n">val_loader</span> <span class="o">=</span> <span class="n">DataLoader</span><span class="p">(</span><span class="n">val_set</span><span class="p">,</span> <span class="n">batch_size</span><span class="o">=</span><span class="n">batch_size</span><span class="p">,</span> <span class="n">shuffle</span><span class="o">=</span><span class="bp">True</span><span class="p">)</span>
<span class="n">test_loader</span> <span class="o">=</span> <span class="n">DataLoader</span><span class="p">(</span><span class="n">test_set</span><span class="p">,</span> <span class="n">batch_size</span><span class="o">=</span><span class="n">batch_size</span><span class="p">,</span><span class="n">shuffle</span><span class="o">=</span><span class="bp">True</span><span class="p">)</span>
</code></pre></div></div>

<p>The code for training seems a bit complicated because we use batches. 
This is due to dealing with a large amount of data, and we don’t send all of it through the network at once (per training step), but only a part of it, namely <code class="language-plaintext highlighter-rouge">batch_size</code> many. 
An <code class="language-plaintext highlighter-rouge">epoch</code> is defined by the fact that all training data have been sent through the network once.</p>

<p>In essence, nothing else happens but:</p>

<ol>
  <li>Send Batch through the network (Forward pass)</li>
  <li>Calculate error/cost</li>
  <li>Propagate gradients of the cost function with respect to the model parameters backwards through the network (Backward pass)</li>
  <li>Update model parameters</li>
</ol>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">def</span> <span class="nf">train_one_epoch</span><span class="p">(</span><span class="n">epoch_index</span><span class="p">,</span> <span class="n">tb_writer</span><span class="p">,</span> <span class="n">n_epochs</span><span class="p">):</span>
    <span class="n">running_loss</span> <span class="o">=</span> <span class="mf">0.0</span>
    <span class="n">last_loss</span> <span class="o">=</span> <span class="mf">0.0</span>
    <span class="n">all_steps</span> <span class="o">=</span> <span class="n">n_epochs</span> <span class="o">*</span> <span class="nb">len</span><span class="p">(</span><span class="n">train_loader</span><span class="p">)</span>
    
    <span class="k">for</span> <span class="n">i</span><span class="p">,</span> <span class="n">data</span> <span class="ow">in</span> <span class="nb">enumerate</span><span class="p">(</span><span class="n">train_loader</span><span class="p">):</span>
        <span class="n">local_X</span><span class="p">,</span> <span class="n">local_y</span> <span class="o">=</span> <span class="n">data</span>
        <span class="n">local_X</span><span class="p">,</span> <span class="n">local_y</span> <span class="o">=</span> <span class="n">local_X</span><span class="p">.</span><span class="n">to</span><span class="p">(</span><span class="n">device</span><span class="p">),</span> <span class="n">local_y</span><span class="p">.</span><span class="n">to</span><span class="p">(</span><span class="n">device</span><span class="p">)</span>
        <span class="n">optimizer</span><span class="p">.</span><span class="n">zero_grad</span><span class="p">()</span>
        <span class="n">outputs</span> <span class="o">=</span> <span class="n">model</span><span class="p">(</span><span class="n">local_X</span><span class="p">)</span>
        
        <span class="n">loss</span> <span class="o">=</span> <span class="n">criterion</span><span class="p">(</span><span class="n">outputs</span><span class="p">,</span> <span class="n">local_y</span><span class="p">)</span>
        <span class="n">loss</span><span class="p">.</span><span class="n">backward</span><span class="p">()</span>
        <span class="n">optimizer</span><span class="p">.</span><span class="n">step</span><span class="p">()</span>
        
        <span class="n">running_loss</span> <span class="o">+=</span> <span class="n">loss</span><span class="p">.</span><span class="n">item</span><span class="p">()</span>
        <span class="k">if</span> <span class="n">i</span> <span class="o">%</span> <span class="n">eval_interval</span> <span class="o">==</span> <span class="n">eval_interval</span><span class="o">-</span><span class="mi">1</span><span class="p">:</span>
            <span class="n">last_loss</span> <span class="o">=</span> <span class="n">running_loss</span> <span class="o">/</span> <span class="n">eval_interval</span>
            
            <span class="n">steps</span> <span class="o">=</span> <span class="n">epoch_index</span> <span class="o">*</span> <span class="nb">len</span><span class="p">(</span><span class="n">train_loader</span><span class="p">)</span> <span class="o">+</span> <span class="p">(</span><span class="n">i</span><span class="o">+</span><span class="mi">1</span><span class="p">)</span>
            
            <span class="n">ep_str</span> <span class="o">=</span> <span class="sa">f</span><span class="s">'Epoch [</span><span class="si">{</span><span class="n">epoch_index</span><span class="o">+</span><span class="mi">1</span><span class="si">}</span><span class="s">/</span><span class="si">{</span><span class="n">n_epochs</span><span class="si">}</span><span class="s">]'</span>
            <span class="n">step_str</span> <span class="o">=</span> <span class="sa">f</span><span class="s">'Step [</span><span class="si">{</span><span class="n">steps</span><span class="si">}</span><span class="s">/</span><span class="si">{</span><span class="n">all_steps</span><span class="si">}</span><span class="s">]'</span>
            <span class="n">loss_str</span> <span class="o">=</span> <span class="sa">f</span><span class="s">'Loss: </span><span class="si">{</span><span class="n">last_loss</span><span class="si">:</span><span class="p">.</span><span class="mi">4</span><span class="n">f</span><span class="si">}</span><span class="s">'</span>
            <span class="k">print</span><span class="p">(</span><span class="sa">f</span><span class="s">'</span><span class="si">{</span><span class="n">ep_str</span><span class="si">}</span><span class="s">, </span><span class="si">{</span><span class="n">step_str</span><span class="si">}</span><span class="s">, </span><span class="si">{</span><span class="n">loss_str</span><span class="si">}</span><span class="s">'</span><span class="p">)</span>

            <span class="n">tb_x</span> <span class="o">=</span> <span class="n">epoch_index</span> <span class="o">*</span> <span class="nb">len</span><span class="p">(</span><span class="n">train_loader</span><span class="p">)</span> <span class="o">+</span> <span class="n">i</span> <span class="o">+</span> <span class="mi">1</span>
            <span class="n">tb_writer</span><span class="p">.</span><span class="n">add_scalar</span><span class="p">(</span><span class="s">'Loss/train'</span><span class="p">,</span> <span class="n">last_loss</span><span class="p">,</span> <span class="n">tb_x</span><span class="p">)</span>
            <span class="n">running_loss</span> <span class="o">=</span> <span class="mf">0.</span>
            
    <span class="k">return</span> <span class="n">last_loss</span>

<span class="c1"># Initializing in a separate cell so we can easily add more epochs to the same run
</span><span class="k">def</span> <span class="nf">train</span><span class="p">(</span><span class="n">n_epochs</span><span class="p">,</span><span class="n">respect_val</span><span class="o">=</span><span class="bp">False</span><span class="p">):</span>
    <span class="n">timestamp</span> <span class="o">=</span> <span class="n">datetime</span><span class="p">.</span><span class="n">now</span><span class="p">().</span><span class="n">strftime</span><span class="p">(</span><span class="s">'%Y%m%d_%H%M%S'</span><span class="p">)</span>
    <span class="n">writer</span> <span class="o">=</span> <span class="n">SummaryWriter</span><span class="p">(</span><span class="s">'runs/fashion_trainer_{}'</span><span class="p">.</span><span class="nb">format</span><span class="p">(</span><span class="n">timestamp</span><span class="p">))</span>
    <span class="n">best_vloss</span> <span class="o">=</span> <span class="mi">1_000_000</span>

    <span class="k">for</span> <span class="n">epoch</span> <span class="ow">in</span> <span class="nb">range</span><span class="p">(</span><span class="n">n_epochs</span><span class="p">):</span>    
        <span class="n">model</span><span class="p">.</span><span class="n">train</span><span class="p">(</span><span class="bp">True</span><span class="p">)</span>
        <span class="n">avg_loss</span> <span class="o">=</span> <span class="n">train_one_epoch</span><span class="p">(</span><span class="n">epoch</span><span class="p">,</span> <span class="n">writer</span><span class="p">,</span> <span class="n">n_epochs</span><span class="p">)</span>
        
        <span class="n">model</span><span class="p">.</span><span class="n">train</span><span class="p">(</span><span class="bp">False</span><span class="p">)</span>
        <span class="n">running_vloss</span> <span class="o">=</span> <span class="mf">0.0</span>
        
        <span class="k">for</span> <span class="n">i</span><span class="p">,</span> <span class="n">vdata</span> <span class="ow">in</span> <span class="nb">enumerate</span><span class="p">(</span><span class="n">val_loader</span><span class="p">):</span>
            
            <span class="n">local_X</span><span class="p">,</span> <span class="n">local_y</span> <span class="o">=</span> <span class="n">vdata</span>
            <span class="n">local_X</span><span class="p">,</span> <span class="n">local_y</span> <span class="o">=</span> <span class="n">local_X</span><span class="p">.</span><span class="n">to</span><span class="p">(</span><span class="n">device</span><span class="p">),</span> <span class="n">local_y</span><span class="p">.</span><span class="n">to</span><span class="p">(</span><span class="n">device</span><span class="p">)</span>
            
            <span class="n">voutputs</span> <span class="o">=</span> <span class="n">model</span><span class="p">(</span><span class="n">local_X</span><span class="p">)</span>
            <span class="n">vloss</span> <span class="o">=</span> <span class="n">criterion</span><span class="p">(</span><span class="n">voutputs</span><span class="p">,</span> <span class="n">local_y</span><span class="p">)</span>
            <span class="n">running_vloss</span> <span class="o">+=</span> <span class="n">vloss</span>
            
        <span class="n">avg_vloss</span> <span class="o">=</span> <span class="n">running_vloss</span> <span class="o">/</span> <span class="p">(</span><span class="n">i</span><span class="o">+</span><span class="mi">1</span><span class="p">)</span>

        <span class="n">ep_str</span> <span class="o">=</span> <span class="sa">f</span><span class="s">'Epoch [</span><span class="si">{</span><span class="n">epoch</span><span class="o">+</span><span class="mi">1</span><span class="si">}</span><span class="s">/</span><span class="si">{</span><span class="n">n_epochs</span><span class="si">}</span><span class="s">]'</span>
        <span class="n">tloss_str</span> <span class="o">=</span> <span class="sa">f</span><span class="s">'Train-Loss: </span><span class="si">{</span><span class="n">avg_loss</span><span class="si">:</span><span class="p">.</span><span class="mi">4</span><span class="n">f</span><span class="si">}</span><span class="s">'</span>
        <span class="n">vloss_str</span> <span class="o">=</span> <span class="sa">f</span><span class="s">'Val-Loss: </span><span class="si">{</span><span class="n">avg_vloss</span><span class="si">:</span><span class="p">.</span><span class="mi">4</span><span class="n">f</span><span class="si">}</span><span class="s">'</span>
        <span class="k">print</span><span class="p">(</span><span class="sa">f</span><span class="s">'</span><span class="si">{</span><span class="n">ep_str</span><span class="si">}</span><span class="s">, </span><span class="si">{</span><span class="n">tloss_str</span><span class="si">}</span><span class="s">, </span><span class="si">{</span><span class="n">vloss_str</span><span class="si">}</span><span class="s">'</span><span class="p">)</span>
        
        <span class="n">writer</span><span class="p">.</span><span class="n">add_scalars</span><span class="p">(</span>
            <span class="s">'Training vs. Validation Loss'</span><span class="p">,</span> 
            <span class="p">{</span><span class="s">'Training'</span><span class="p">:</span> <span class="n">avg_loss</span><span class="p">,</span> <span class="s">'Validation'</span><span class="p">:</span> <span class="n">avg_vloss</span><span class="p">},</span> 
            <span class="n">epoch</span>
        <span class="p">)</span>

        <span class="n">writer</span><span class="p">.</span><span class="n">flush</span><span class="p">()</span>
        
        <span class="k">if</span> <span class="ow">not</span> <span class="n">respect_val</span> <span class="ow">or</span> <span class="p">(</span><span class="n">respect_val</span> <span class="ow">and</span> <span class="n">avg_vloss</span> <span class="o">&lt;</span> <span class="n">best_vloss</span><span class="p">):</span>
            <span class="n">best_vloss</span> <span class="o">=</span> <span class="n">avg_vloss</span>
            <span class="n">model_path</span> <span class="o">=</span> <span class="s">'./models/_model_{}_{}'</span><span class="p">.</span><span class="nb">format</span><span class="p">(</span><span class="n">timestamp</span><span class="p">,</span> <span class="n">epoch</span><span class="p">)</span>
            <span class="n">torch</span><span class="p">.</span><span class="n">save</span><span class="p">(</span><span class="n">model</span><span class="p">.</span><span class="n">state_dict</span><span class="p">(),</span> <span class="n">model_path</span><span class="p">)</span>
</code></pre></div></div>

<p>Calling <code class="language-plaintext highlighter-rouge">train</code> starts the training.</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">train</span><span class="p">(</span><span class="n">n_epochs</span><span class="p">)</span>
</code></pre></div></div>

<p>The best model from the training can be found in the folder <code class="language-plaintext highlighter-rouge">./models</code> and can be loaded as follows</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">model_path</span> <span class="o">=</span> <span class="s">'./models/pretrained_1_128_best_val'</span>

<span class="k">if</span> <span class="n">device</span><span class="p">.</span><span class="nb">type</span> <span class="o">==</span> <span class="s">'cpu'</span><span class="p">:</span>
    <span class="n">model</span><span class="p">.</span><span class="n">load_state_dict</span><span class="p">(</span><span class="n">torch</span><span class="p">.</span><span class="n">load</span><span class="p">(</span><span class="n">model_path</span><span class="p">,</span> <span class="n">map_location</span><span class="o">=</span><span class="n">torch</span><span class="p">.</span><span class="n">device</span><span class="p">(</span><span class="s">'cpu'</span><span class="p">)))</span>
<span class="k">elif</span> <span class="n">torch</span><span class="p">.</span><span class="n">backends</span><span class="p">.</span><span class="n">mps</span><span class="p">.</span><span class="n">is_available</span><span class="p">():</span>
    <span class="n">model</span><span class="p">.</span><span class="n">load_state_dict</span><span class="p">(</span><span class="n">torch</span><span class="p">.</span><span class="n">load</span><span class="p">(</span><span class="n">model_path</span><span class="p">,</span> <span class="n">map_location</span><span class="o">=</span><span class="n">torch</span><span class="p">.</span><span class="n">device</span><span class="p">(</span><span class="s">'mps'</span><span class="p">)))</span>
<span class="k">else</span><span class="p">:</span>
    <span class="n">model</span><span class="p">.</span><span class="n">load_state_dict</span><span class="p">(</span><span class="n">torch</span><span class="p">.</span><span class="n">load</span><span class="p">(</span><span class="n">model_path</span><span class="p">))</span>
<span class="n">model</span><span class="p">.</span><span class="nb">eval</span><span class="p">()</span>
</code></pre></div></div>

<h2 id="melody-generation-after-training">Melody Generation (After Training)</h2>

<p>After training or after we load our pretrained model, we generate new pieces:</p>

<div class="language-python highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">n_scores</span> <span class="o">=</span> <span class="mi">5</span>
<span class="n">temperature</span> <span class="o">=</span> <span class="mf">0.6</span>
<span class="n">after_new_songs</span> <span class="o">=</span> <span class="p">[]</span>
<span class="k">for</span> <span class="n">_</span> <span class="ow">in</span> <span class="nb">range</span><span class="p">(</span><span class="n">n_scores</span><span class="p">):</span>
    <span class="n">encoded_song</span> <span class="o">=</span> <span class="n">generate</span><span class="p">(</span><span class="n">max_len</span><span class="o">=</span><span class="mi">120</span><span class="p">,</span><span class="n">temperature</span><span class="o">=</span><span class="n">temperature</span><span class="p">)</span>
    <span class="k">print</span><span class="p">(</span><span class="sa">f</span><span class="s">'generated </span><span class="si">{</span><span class="s">" "</span><span class="p">.</span><span class="n">join</span><span class="p">(</span><span class="n">encoded_song</span><span class="p">)</span><span class="si">}</span><span class="s"> consisting of </span><span class="si">{</span><span class="nb">len</span><span class="p">(</span><span class="n">encoded_song</span><span class="p">)</span><span class="si">}</span><span class="s"> notes'</span><span class="p">)</span>
    <span class="n">after_new_songs</span><span class="p">.</span><span class="n">append</span><span class="p">(</span><span class="n">encoded_song</span><span class="p">)</span>

<span class="n">after_generated_scores</span> <span class="o">=</span> <span class="n">encoder</span><span class="p">.</span><span class="n">decode_songs</span><span class="p">(</span><span class="n">after_new_songs</span><span class="p">)</span>
<span class="n">Audio</span><span class="p">(</span><span class="n">score_to_wav</span><span class="p">(</span><span class="n">after_generated_scores</span><span class="p">[</span><span class="mi">0</span><span class="p">],</span> <span class="s">'a_g_song.wav'</span><span class="p">))</span>
</code></pre></div></div>

<p>We start to hear repetition and some structure within the piece:</p>

<audio controls="">
  <source src="/Pages/assets/audio/a_g_song.mp3" type="audio/mp3" />
  Your browser does not support the audio element.
</audio>

<h2 id="real-world-example">Real World Example</h2>

<p>In the realm of musical innovation, a significant advancement occurred with the development of a sophisticated <strong>LSTM</strong> model designed to create expressive piano roll music. 
This model was introduced in a notable study by <a class="citation" href="#oore:2018">(Oore et al., 2018)</a>. 
The researchers devised a unique discrete-event based representation for piano rolls, encompassing a diverse range of 413 different events. 
The architecture of their model was meticulously structured, comprising three hidden LSTM layers, each equipped with 512 cells. 
This design choice facilitated the processing of a 413-dimensional one-hot vector as input, with the model subsequently generating a categorical distribution over the same dimensional space.</p>

<p>The training process of the model was finely tuned, employing a mini-batch size of 64 and a learning rate of 0.001, alongside the implementation of teacher forcing techniques. 
For those interested in experiencing the model’s capabilities firsthand, a collection of generated music pieces is available for listening at this <a href="https://clyp.it/user/3mdslat4">link</a>. 
However, it’s important to note a primary limitation of this model: its tendency to produce relatively brief musical compositions, typically ranging from 10 to 20 seconds in duration. 
The authors also emphasized the critical role of high-quality data in achieving optimal results with this model.</p>

<p>Following this development, the field witnessed the emergence of the Music Transformer, introduced by <a class="citation" href="#huang:2018">(Huang et al., 2018)</a>.
This model also utilized a similar piano roll representation but marked a significant leap forward by employing the <strong>transformer</strong> architecture. 
This innovative approach enabled the Music Transformer to learn and reproduce longer sequences, demonstrating the capability to capture more extended musical dependencies. 
The transformer architecture and its implications in music generation will be further explored in the next installment of this series.</p>

<h2 id="references">References</h2>

<ol class="bibliography"><li><span id="oore:2018">Oore, S., Simon, I., Dieleman, S., Eck, D., &amp; Simonyan, K. (2018). <i>This time with feeling: Learning expressive musical performance</i>.</span></li>
<li><span id="huang:2018">Huang, C.-Z. A., Vaswani, A., Uszkoreit, J., Shazeer, N., Hawthorne, C., Dai, A. M., Hoffman, M. D., &amp; Eck, D. (2018). Music Transformer: Generating music with long-term structure. <i>ArXiv Preprint ArXiv:1809.04281</i>.</span></li></ol>]]></content><author><name>Benedikt Zönnchen</name><email>benedikt.zoennchen@web.de</email></author><category term="Music" /><category term="ML" /><category term="LSTM" /><summary type="html"><![CDATA[This article is the continuation of a series. It is recommended that you read part I and II first. This time we use a recurrent neural network (RNN), more precisely an LSTM, which I explained a little bit in the introduction. An LSTM is a RNN that counteracts the problem of exploding and vanishing gradients.]]></summary></entry><entry><title type="html">Laws of Form</title><link href="https://bzoennchen.github.io/Pages/2023/11/19/laws-of-form.html" rel="alternate" type="text/html" title="Laws of Form" /><published>2023-11-19T00:00:00+01:00</published><updated>2023-11-19T00:00:00+01:00</updated><id>https://bzoennchen.github.io/Pages/2023/11/19/laws-of-form</id><content type="html" xml:base="https://bzoennchen.github.io/Pages/2023/11/19/laws-of-form.html"><![CDATA[<p>In my last blog <a href="/Pages/2023/10/07/system-theory-and-ai.html">post</a>, I discussed Niklas Luhmann’s Social Systems Theory and I emphasized that his theory is based on differentiation and seemingly paradox relations.
My general understanding of Luhmann’s radical constructivism in the most reductive sense is that there are no unified and independent objects—there are only differences.
What an observer can identify as an object is a differentiation of a system and its environment, a foreground and its background, an interior and the external.
However, any observer is itself a distinction between system and environment.
I want to further investigate this idea by looking into Luhmann’s inspiration—his muse so to say.
So let me examine the logic of Georg Spencer-Brown presented in his work <em>Laws of Form</em> <a class="citation" href="#brown:1969">(Spencer-Brown, 1969)</a>.</p>

<h2 id="the-principle-of-differentiation">The Principle of Differentiation</h2>

<p>As mentioned, Luhmann’s theory is based on paradoxes such as</p>

<blockquote>
  <p>This statement is wrong.</p>
</blockquote>

<p>The statement is logically problematic because it is self-referential. 
If the statement is false, it states that it is in fact true and if the statement is true, it states that it is false.
Another famous paradox is Russell’s antinomy of the naïve set theory:</p>

\[R := \left\{ x \ | \ x \not\in x \right\}.\]

<p>\(R\) is the set of all sets that do not contain themselves which seems to be a well-defined mathematical object.
However, if we introduce the self-referential relation, we run into a paradox:</p>

<p>\begin{equation} 
R \in R \iff R \not\in R
\end{equation}</p>

<p>This paradox is related to the barber that shaves everyone that does not shave themselves.
If that is the case, does the barber shave themselves?</p>

<p>Another example involves the the proof of the <a href="/Pages/2021/06/08/Informatics-a-love-letter.html">Halting Problem</a>.
To prove it, one can establish a self-referential relation between a machine that presumable solves the Halting Problem.
The machine does not halt if the machine it checks halts, and it halts if the machine it checks does not halt.
Via the self-referential relation, that is, by letting the machine check itself, we get a contradiction.</p>

<p>The idea of Spencer-Brown is to resolve these paradoxes over time.
\(R \in R\) holds at one moment in time and \(R \not\in R\) holds at the next moment.
Note however that he does not resolve Russell’s antinomy <a class="citation" href="#cull:1979">(Cull &amp; Frank, 1979)</a>.
His idea of resolving paradoxes over time is reminiscent of the Hegelian dialectic—a process of self-creation.
The self-referential relation is a paradox if we ignore time and it becomes a generator if we consider time and place.</p>

<p>This led the biologists Huberto R. Maturana and Francisco J. Varela to the concept of <em>autopoiesis</em> <a class="citation" href="#maturana:1987">(Maturana &amp; Varela, 1987)</a>.
Furthermore, the importance of differentiation is inspired by the logic of Spencer-Brown and his work <em>Laws of Form</em> <a class="citation" href="#brown:1969">(Spencer-Brown, 1969)</a>.
He resolves the paradox by a similar idea that gave us imaginary numbers, that is, by using what he calls <em>Re-entry</em> which (re-)introduces a system into itself.</p>

<p>Luhmann integrated this idea into his social systems theory.
For example, the media can observe and reintroduce itself into itself. 
It can use its systemic operations on itself, i.e. it can report on itself.</p>

<p>Spencer-Brown begins his work by a quote from Lao-Tse (a stand-in for many different authors) thus begins by philosophical considerations:</p>

<blockquote>
  <p>Wu ming tain di zhi shi. – Loa-Tse</p>
</blockquote>

<p>The sentence has mainly two different meanings.
One is:</p>

<blockquote>
  <p>The beginning of heaven and earth is without a name.</p>
</blockquote>

<p>The other one is:</p>

<blockquote>
  <p>‘Nothing’ is the name of the beginning of heaven and earth.</p>
</blockquote>

<p>A paradox arises: How can Nothing be nothing if we can call it ‘Nothing’?
Furthermore, the quote points to a distinction between heaven and earth.
Can there be heaven without earth—a <em>calling</em> or <em>indication</em> without a <em>distinction</em>?
Spencer-Brown begins by the assumption that there is no such thing:</p>

<blockquote>
  <p>We take the idea of distinction and the idea of indication and that we cannot make an indication without making a distinction as given.
Therefore, we take the form of distinction as the form itself. – Georg Spencer-Brown</p>
</blockquote>

<p>In other words, what we normally identify as object (the form / system) is for Spencer-Brown equal to the distinction (system-environment differentiation).
There is no clear separation between the object or the result of distinction and the process of distinguishing.
Therefore, the process must be integrated into Spencer-Brown’s logic and as we will see, there is no clear separation between objects and operations in Spencer-Browns calculus.
Spencer-Brown thinks that differentiation is a proto-operation that is more fundamental than performing calculations or writing text because to do these activities we have to differentiate beforehand.
I cannot calculate 1 + 1 = 2 without distinguishing between the different symbols and a symbol and ‘nothing’ or the void.</p>

<p>Spencer-Brown uses the mark or cross (result) which at the same time marks (process).
The mark is, calls, and makes a difference.</p>

<p><br /></p>
<div><img style="height:30px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/mark.png" alt="The mark." /></div>
<p><br /></p>

<p>There is an interior of the mark and not the interior—the system and its environment.
I can make a distinction again (repetition):</p>

<p><br /></p>
<div><img style="height:30px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/calling.png" alt="The calling." /></div>
<p><br /></p>

<p>But making a distinction again does not change the distinction.</p>

<blockquote>
  <p>Calling something back-to-back by its name does not change its name. – Spencer-Brown</p>
</blockquote>

<p>The reverse is also true; therefore, Spencer-Brown introduces the <em>Law of Calling</em>:</p>

<p><br /></p>
<div><img style="height:30px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/law-of-calling.png" alt="Law of Calling." /></div>
<p><br /></p>

<p>The second transformation called <em>Law of Crossing</em> is less intuitive.
Crossing twice reverses the first crossing.</p>

<p><br /></p>
<div><img style="height:30px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/law-of-crossing.png" alt="Law of Crossing." /></div>
<p><br /></p>

<blockquote>
  <p>If a boundary is crossed twice, the original state will be reestablished.
The repetition of the crossing has a different value than the single crossing.
The reason is that in-between the reversal happens.
Crossing changes the side.
Re-crossing reverses this operation. – Spencer-Brown</p>
</blockquote>

<p>With only these two laws, Spencer-Brown established a logic calculus and we can start doing mathematics.
Interestingly, the <em>Law of Crossing</em> and the <em>Law of Calling</em> are implicitly established via the position of the marks.
There is no operator introduced because the result and process, indication and differentiation, the mark and the process of marking are not separated.</p>

<p>Let’s see what we can do with this calculus.
Let \(a, b\) variables, then the following holds:</p>

<p><br /></p>
<div><img style="height:230px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/transformations.png" alt="Transformations." /></div>
<p><br /></p>

<p>Let’s have a look at the last transformation.
Let us assume \(a\) is a <strong>mark</strong>.
Then we get:</p>

<p><br /></p>
<div><img style="height:30px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/transformation-1.png" alt="First possibility." /></div>
<p><br /></p>

<p>Let \(a\) be <strong>unmarked</strong> instead, then we can follow:</p>

<p><br /></p>
<div><img style="height:30px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/transformation-2.png" alt="Second possibility." /></div>
<p><br /></p>

<h2 id="the-re-entry">The Re-Entry</h2>

<p>How can we re-introduce the system (here an equation) into itself?
Or in other words: How does the <em>re-entry</em> work?
We can start by a simple self-referential algebraic equation:</p>

<p>\begin{equation}
x^2 = ax + b, \quad a, b \in \mathbb{R}.
\end{equation}</p>

<p>This equation has well-known solutions. 
It is also known that solutions can be imaginary, i.e., \(x\) might be of the form \(r + si\) with \(r, s \in \mathbb{R}\) and \(i^2 = -1\).
To see the re-entry, we can rewrite the equation above to get</p>

<p>\begin{equation}
x = a + b/x,
\end{equation}</p>

<p>thus the self-reference is obvious and we solve the equation by the re-entry</p>

<p>\begin{equation}
x = a + b/(a +b/(a+b/(a+b/a+b/(a + \ldots)))).
\end{equation}</p>

<p>Using this infinite formalism it is literally the case that</p>

<p>\begin{equation}
x = a + b/x,
\end{equation}
holds.
Using the same formalism, we can define the imaginary number \(i = -1/i\) as literally</p>

<p>\begin{equation}
i = -1 /(-1 /(-1 / (-1 / \ldots )))
\end{equation}</p>

<p>but what does this mean?
The system is not a number but a process, a generator that generates itself.
\(i\) alternates between 1 and -1.
Interestingly, this is precisely how we use the equal sign in most programming languages.
Writing <code class="language-plaintext highlighter-rouge">i = i / -1</code> in a programming language means</p>

<p>\begin{equation}
i \leftarrow \frac{i}{-1}.
\end{equation}</p>

<p>The next step is to introduce such a re-entry into logic.
Similar to the imaginary number \(i\), Spencer-Brown gives us the following fundamental paradox (<em>The Re-Entry of the Mark</em>):</p>

<p><br /></p>
<div><img style="height:30px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/j.png" alt="Imaginary truth value." /></div>
<p><br /></p>

<p>The solution is an alternation between a marked and unmarked state—between true and false.
A state that might seem contradictory in space, makes sense if it is observed in time and space.
Again, time resolves the paradox.
To highlight the re-entry, Spencer-Brown also uses the following notation:</p>

<p><br /></p>
<div><img style="height:30px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/j-2.png" alt="Notation of the re-entry." /></div>
<p><br /></p>

<p>We could similarily notate \(i\) as</p>

<p><br /></p>
<div><img style="height:35px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/i-2.png" alt="Notation of the re-entry for the imaginary number." /></div>
<p><br /></p>

<p>If we change <strong>all</strong> symbols within a system equally, there is no reason not to calculate with a self-generating process.
For example, we can state the following:</p>

<p><br /></p>
<div><img style="height:30px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/j-calculations.png" alt="Calculating with the re-entry." /></div>
<p><br /></p>

<p>However, it is forbidden to only change one appearance of \(J\)!
Following this simple rule, no paradox or inconsistency arises.
We can go on and evaluate the following transformation:</p>

<p><br /></p>
<div><img style="height:30px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/wave.png" alt="Wave equation." /></div>
<p><br /></p>

<p>which describes two alternating waves shifted by one cycle resulting in a mark.</p>

<p>Spencer-Brown goes on and defines his <em>Echelon</em>:</p>

<p><br /></p>
<div><img style="height:35px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/echelon.png" alt="Echelon." /></div>
<p><br /></p>

<p>which can be transformed into</p>

<p><br /></p>
<div><img style="height:40px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/echelon-transformation.png" alt="Echelon transformation." /></div>
<p><br /></p>

<p>thus gives us the re-entry</p>

<p><br /></p>
<div><img style="height:40px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/echelon-equation.png" alt="Echelon equation." /></div>
<p><br /></p>

<p>or</p>

<p><br /></p>
<div><img style="height:43px;display:block; margin-left:auto; margin-right:auto; " src="/Pages/assets/images/laws-of-form/echelon-equation-2.png" alt="Echelon equation." /></div>
<p><br /></p>

<h2 id="final-words">Final Words</h2>

<p>In the act of programming there is no problem of using expressions such as</p>

<p>\begin{equation}
i \leftarrow i + 1 \quad  \text{ or } \quad a  \leftarrow f(a)
\end{equation}</p>

<p>but in mathematics—at least since Plato—we assume some sort of eternity.
Of course, we can translate between the static world of “normal” mathematics and Spencer-Brown’s dynamic viewpoint, but it is a different viewpoint which might influence how we observe our environment.
It is like in physics where multiple theories are equivalent but start from very different viewpoints.
It starts by differentiation which gets reintroduced into the system which is constructed by this very same differentiation.
To generate new numbers, such as irrational or transfinite numbers, Spencer-Brown proposes not to use the limit but the whole infinite process that defines such limit.</p>

<p>Spencer-Brown believed that to be able to master the transition to new signs, something is necessary for which the previous signs are not sufficient.
To be able to close this gap; to make this leap successfully; to resolve paradoxes; a specific language of one’s own is necessary. 
According to Spencer-Brown this step is accomplished by <strong>thinking</strong> which provides us with its specific imaginations, playful freedom and contradictions.</p>

<p>It is surprising that Spencer-Brown’s <em>Laws of Form</em> plays no role in computer science even though it fits quite neatly in the perspective of programs, processes and computation.</p>

<h2 id="references">References</h2>

<ol class="bibliography"><li><span id="brown:1969">Spencer-Brown, G. (1969). <i>Laws of Form</i>. London: Allen and Unwin.</span></li>
<li><span id="cull:1979">Cull, P., &amp; Frank, W. (1979). flaws of form. <i>International Journal of General Systems</i>, <i>5</i>(4), 201–211. https://doi.org/10.1080/03081077908547450</span></li>
<li><span id="maturana:1987">Maturana, H. R., &amp; Varela, F. J. (1987). <i>The Tree of Knowledge</i>. Shambhala.</span></li></ol>]]></content><author><name>Benedikt Zönnchen</name><email>benedikt.zoennchen@web.de</email></author><category term="Logic" /><category term="Social Systems Theory" /><summary type="html"><![CDATA[In my last blog post, I discussed Niklas Luhmann’s Social Systems Theory and I emphasized that his theory is based on differentiation and seemingly paradox relations. My general understanding of Luhmann’s radical constructivism in the most reductive sense is that there are no unified and independent objects—there are only differences. What an observer can identify as an object is a differentiation of a system and its environment, a foreground and its background, an interior and the external. However, any observer is itself a distinction between system and environment. I want to further investigate this idea by looking into Luhmann’s inspiration—his muse so to say. So let me examine the logic of Georg Spencer-Brown presented in his work Laws of Form (Spencer-Brown, 1969).]]></summary></entry></feed>