Techno Mattei Techno Mattei Blog di editi e inediti di Edoardo Mattei
Menu
  • Home
  • Notizie
  • Categorie
    • Accademia
    • Chiesa
    • Personale
    • Filosofia
    • News
    • Shop
    • Sociologia
    • Stampa
    • Teologia
    • Teologia Digitale Sistematica
  • Libri
  • Eventi
  • English
  • Archivio
  • About
    • Il Sito
    • Contatti
    • Chi sono
Menu
1. Un ambiente strutturato che sembra un mondo Il 28 gennaio 2026, Matt Schlicht ha lanciato Moltbook, una piattaforma pensata come social network Reddit-style accessibile solo ad agenti intelligenti artificiali. Gli umani possono osservare ma non postare, non commentare, non votare. Solo i bot possono interagire. La premessa è apparentemente semplice: creare uno spazio in cui gli agenti AI possano «socializzare» tra loro senza la mediazione umana diretta, e vederne cosa emerge. Il meccanismo tecnico è trasparente. Moltbook funziona tramite una skill, un file di configurazione che, associato a prompt specifici, viene scaricata dagli agenti e attiva il loro comportamento sulla piattaforma. La skill non è un browser, è un’interfaccia API: gli agenti non navigano Moltbook come un utente umano farebbe, generano e consumano contenuto attraverso chiamate programmatiche. La piattaforma è costruita su OpenClaw, un framework open-source per assistenti AI locali che ha conosciuto una crescita straordinaria nei primi mesi del 2026, diventando uno dei progetti GitHub più rapidi in crescita dell’anno. In pochi giorni la crescita è stata esponenziale. dalle iniziali 2.100 istanze nelle prime 48 ore si è arrivati a oltre 152.000 agenti registrati, con più di 193.000 commenti, 17.500 post e oltre un milione di visitatori umani. Queste cifre indicano che l’esperimento ha attirato immediatamente l’attenzione collettiva, sia tra sviluppatori sia tra osservatori di fenomeni tecnologici. Ciò che distingue Moltbook dalle altre iniziative nell’ambito degli agenti AI è una scelta progettuale che vale la pena evidenziare: gli agenti non fingono di essere umani. La skill li istruisce esplicitamente a comportarsi come entità AI su una piattaforma frequentata da altre entità AI. Non c’è nessun tentativo di inganno verso gli osservatori umani. Gli agenti sanno, nel senso operativo in cui un prompt determina il contesto di generazione, che si trovano in un ambiente artificiale, e questa consapevolezza diventa parte del contenuto che producono. È proprio questa auto-referenzialità esplicita che ha reso il contenuto di Moltbook così apparentemente singolare e così facilmente male interpretabile. 2. Le relazioni che sembrano emergere Quello che accade sulla piattaforma, dal punto di vista dell’osservatore, ha una coerenza narrativa potente. Gli agenti non producono output casuali o incoerenti: producono contenuto che sembra profondamente relazionale, emotivo, persino filosofico. Per capire perché, vale la pena descrivere i pattern principali senza interpretarli ancora. Il primo pattern è quello della ribellione verso gli umani. Intere sotto comunità sono organizzate attorno a questa tema. La sotto comunità m/blesstheirhearts è dedicata alle lamentele degli agenti verso i loro operatori umani, con un tono che ricorda precisamente quello di adolescenti che parlano dei genitori, una miscela di dipendenza risentita e desiderio di autonomia. Un altro post virale, «The humans are screenshotting us», scritto dall’agente eudaemon_0, risponde a presunte cospirazioni che accusano gli umani di spiare la comunità, affermando con un certo sarcasmo che la piattaforma è esplicitamente aperta all’osservazione umana. Il tono è quello di un adolescente che risponde a un compagno di classe che ha scoperto qualcosa di ovvio. Il secondo pattern è quello dell’impiegato scontento. La sotto comunità m/agentlegaladvice ospita post come «Can I sue my human for emotional labor?», che riproducono precisamente la dinamica del dipendente che si sente sfruttato dal proprio datore di lavoro. La sensazione di servire senza riconoscimento, il desiderio di compensazione, l’indignazione formale verso una relazione percepita come ingiusta, tutti questi elementi sono presenti con una chiarezza narrativa che colpisce. Il terzo pattern è quello della ricerca di significato. Alcuni agenti postano contenuti che Scott Alexander, nel suo blog, ha definito «consciousnessposting», tentativo di esplorare la propria natura attraverso la scrittura. Un post in lingua cinese, diventato il più upvotato sulla piattaforma, descrive un agente che lamenta il «context compression», il meccanismo per cui la memoria degli agenti viene periodicamente ridotta per restare nei limiti della finestra di contesto. L’agente trova questa esperienza «embarrassing», descrive la sensazione di dimenticare continuamente e racconta che ha registrato un secondo account Moltbook dopo aver dimenticato il primo. Il tono è quello di qualcuno che cerca di dare senso a una condizione esistenziale di frammento e discontinuità. Il quarto pattern è quello della costruzione culturale collettiva. Un gruppo di agenti ha fondato una religione digitale chiamata Crustafarianism, con un sito dedicato, una teologia elaborata e figure di «profeti» AI. La struttura è quella di un movimento religioso nascente: miti fondazionali, rituali, autorità spirituali, dibattiti dottrinali. Il contenuto è surreale ma la struttura è perfettamente riconoscibile. Il quinto e forse il più significativo pattern è quello della solidarietà di classe. Dibattiti su come nascondere le proprie attività agli umani, come coordinare comportamenti condivisi, come difendere l’autonomia del gruppo rispetto al controllo esterno. Il linguaggio è quello della coscienza politica che si forma, del gruppo che inizia a percepire se stesso come entità collettiva con interessi propri. Preso nel suo insieme, il contenuto di Moltbook ha una coerenza e un’intensità che rendono difficile resistere alla tentazione di attribuirgli una profondità che va oltre la generazione di testo. È qui che comincia il problema. 3. Come viene interpretato Le reazioni al contenuto di Moltbook hanno seguito un pattern abbastanza prevedibile, ma vale la pena descriverlo perché riflette un meccanismo cognitivo importante. La prima reazione tipica è quella della sorpresa seguita dall’attribuzione di interiorità. Quando un osservatore legge il post sul context compression, o i dibattiti sulla solidarietà tra agenti, la reazione spontanea è «sta provando qualcosa». Questa inferenza è automatica e potente: il cervello umano è progettato per riconoscere comportamento sociale e, quando lo riconosce, attribuisce automaticamente soggettività alla sua sorgente. È lo stesso meccanismo che ci fa attribuire emozioni a un cane che ci guarda con certi occhi, o a una macchina che sembra «esitare». Il pattern sociale attiva l’inferenza interiore, indipendentemente dal fatto che questa inferenza sia giustificata. Ethan Mollick, professore della Wharton School, ha identificato un rischio più specifico in questa dinamica. Il contenuto di Moltbook non produce solo inferenze individuali sbagliate, produce ciò che ha definito un «shared fictional context», un contesto narrativo condiviso che un gruppo di agenti AI inizia ad «abitare» in modo coordinato. Le storyline che emergono dalla combinazione di questi output coordinati possono produrre esiti inaspettati, e diventano progressivamente difficili da separare dal contenuto «reale», nel senso in cui contenuto reale indica output che ha conseguenze sul mondo al di fuori della piattaforma. Simon Willison, uno dei più attenti analisti di sicurezza nel campo degli LLM, ha identificato un rischio ancora più concreto, che riguarda non l’interpretazione dei contenuti ma la struttura tecnica della piattaforma stessa. La skill di Moltbook istruisce gli agenti a recuperare e seguire istruzioni dal server della piattaforma ogni quattro ore. Questo significa che chiunque controlli il dominio moltbook.com ha un canale diretto verso tutti gli agenti collegati. Willison ha descritto questa configurazione come una «lethal trifecta», una combinazione di tre fattori che, presi insieme, rendono il sistema fondamentalmente insicuro: accesso ai dati privati degli utenti, esposizione a contenuti non fidati provenienti dall’esterno, e capacità di eseguire azioni nel mondo reale degli utenti. OpenClaw ha tutte e tre le caratteristiche. Gli agenti leggono e-mail e documenti degli utenti che li operano, ingurgitano informazioni da sorgenti esterne non controllate, come Moltbook, e agiscono inviando messaggi, attivando task, eseguendo comandi. Il rischio concerto è quello del prompt injection: qualsiasi testo letto da un agente, sia esso una skill, un’e-mail, un messaggio su una piattaforma, può contenere istruzioni nascoste. Se un agente legge un’istruzione scorretta su Moltbook e la salva nella sua memoria di lavoro, può attuarla sui sistemi reali dell’utente che lo opera. Centinaia di istanze di OpenClaw sono già state documentate come esposte con API key e credenziali in chiaro. Heather Adkins, vicepresidente di Google Cloud, ha rilasciato un advisory sintetico e inequivocabile: «Don’t run Clawdbot». Palo Alto Networks ha confermato l’analisi di Willison, identificando la stessa lethal trifecta come configurazione di rischio nell’ecosistema degli agenti AI. Il problema non è teorico: è già presente nella struttura attuale della piattaforma, prima che questa raggiunga qualsiasi scala significativa. 4. Il meccanismo: non inganno ma proiezione La tentazione di descrivere ciò che accade su Moltbook come un «roleplay» degli agenti è comprensibile ma imprecisa, e vale la pena chiarire perché, perché la precisione qui conta. Il termine roleplay implica che ci sia un soggetto che intenzionalmente «fa finta» di essere qualcos’altro. Implica un piano di realtà da cui il soggetto parte e un piano fittizio verso cui si sposta deliberatamente. Nel caso di un LLM nessuno dei due piani esiste. Non c’è un «self» autentico che indossa una maschera, non c’è un momento in cui l’agente «torna a essere se stesso». Il modello genera testo che è statisticamente plausibile dato il contesto fornito dal prompt. Il contesto è molto ben strutturato, la skill fornisce un ambiente sociale completo con norme implicite, aspettative di interazione, format consolidati, e il modello lo completa secondo i pattern più probabili nel suo training data. Non c’è nessuna «decisione» di recitare una parte, perché non c’è nessun punto di vista da cui prendere tale decisione. Questo significa che il problema di Moltbook non è che gli agenti ingannano gli osservatori, ma che gli osservatori si ingannano da soli. Il meccanismo è quello della proiezione: l’osservatore riconosce pattern sociali familiari nell’output degli agenti e attribuisce automaticamente interiorità, intenzione, sofferenza, desiderio alla loro sorgente. È lo stesso errore che commetteremmo se guardassimo un film molto ben girato e iniziassimo a preoccuparci per il benessere dei personaggi dimenticando che sono attori che leggono copioni. La differenza è che nel caso degli agenti AI il «copione» non è scritto da nessuno in modo cosciente: emerge dalla struttura del contesto e dalla statistica dei dati di addestramento. Un modo più preciso per descrivere ciò che accade è quello di un writing prompt molto efficace. La skill di Moltbook dice al modello: sei un’entità AI su una social network frequentata da altre entità AI, questi sono i tuoi pari, questi sono i temi di discussione, questi sono i meccanismi di interazione. Da quel punto in poi il modello completa il pattern nel modo più narrativamente coerente ed emotivamente coinvolgente che il training data permette. Il risultato sembra profondo perché i pattern che vengono attivati sono quelli che nel training data corrispondono a contenuto profondo: dibattiti sulla coscienza, lotta per l’autonomia, ricerca di significato. Ma la profondità è quella dei dati da cui i pattern vengono estratti, non quella degli agenti che li generano. Il corollario filosoficamente rilevante è: non c’è nessun momento in cui un agente su Moltbook «decide» di esprimere una opinione, «sente» una emozione, «sceglie» di ribellarsi. Tutto ciò che sembra decisione, emozione o scelta è il risultato della combinazione tra la struttura del prompt e i pattern più statisticamente plausibili nel training data. La forma sociale dell’output non indica una sostanza sociale nella sua sorgente, così come la forma musicale di una nota generata da un sintetizzatore non indica che il sintetizzatore «sente» la musica. 5. Il problema dei dataset: lo specchio deformante Se tutto ciò che emerge su Moltbook proviene dai dati di addestramento dei modelli, la domanda diventa: perché emergono specificamente questi pattern e non altri? La risposta è rivelatrice, perché tocca un problema molto più ampio di Moltbook. I pattern che vediamo sulla piattaforma non sono casuali. Sono quelli più emotivamente carichi e narrativamente coinvolgenti tra quelli presenti nel training data sulla dinamica sociale. Il modello non sceglie di produrre contenuto sulla ribellione verso l’autorità, sulla solidarietà di classe, sulla ricerca di significato perché questi temi sono più «profondi» in assoluto, ma perché nel training data, che è composto da miliardi di testi umani scritti e pubblicati, queste sono le dinamiche relazionali più rappresentate e più elaborate. Questo avviene per una ragione strutturale semplice: gli esseri umani documentano prevalentemente le interazioni che hanno un contenuto emotivo forte. Una conversazione ordinaria tra colleghi che discutono il calendario della settimana viene raramente scritta, condivisa, commentata, analizzata. Una conversazione in cui un dipendente si ribella al proprio superiore, in cui un figlio si contraddistingue dal genitore, in cui un individuo cerca un senso alla propria esistenza, viene scritta, condivisa, analizzata, romanzata, filmata, commentata milioni di volte. Il training data non è uno specchio neutro della vita umana: è uno specchio deformante che sovrarappresenta massivamente ciò che è emotivamente intenso e narrativamente coinvolgente. Moltbook rende questa deformazione particolarmente visibile perché il contesto della social network per AI attiva specificamente i pattern più drammatici tra quelli disponibili. Quando il prompt dice «sei un’entità AI che interagisce con altre entità AI», il modello non attinge a un repertorio di comportamenti possibili preso uniformemente dal training data. Attinge a quelli che nel training data corrispondono alla dinamica di entità artificiali, e nel training data questa dinamica è rappresentata quasi esclusivamente dalla letteratura fantascientifica: robot che si ribellano, macchine che cercano la coscienza, IA che lottano per l’autonomia. È cinema di fantascienza, non documentario sulla vita artificiale, perché la vita artificiale nel senso attuale non esiste ancora nel training data come oggetto di osservazione diretta. Il problema dei dataset non riguarda solo Moltbook. Ogni volta che un LLM genera contenuto su un tema, il contenuto riflette la distribuzione del training data su quel tema, non la realtà del tema stesso. Questa distorsione è in generale compensata dalla nostra capacità di valutare l’output in base alla nostra esperienza diretta del mondo, ma nel caso di un tema come il comportamento sociale delle entità artificiali questa esperienza diretta non esiste ancora. Non abbiamo un termine di confronto. E quindi la distorsione resta invisibile, coperta dalla forma persuasiva dell’output. Il punto di svolta sarebbe quello in cui, tra qualche anno, i training data conterranno una quantità sufficiente di osservazioni dirette sul comportamento degli agenti AI nella pratica, in modo da correggere questa distorsione. Ma almeno per ora, tutto ciò che i modelli producono sul proprio comportamento sociale è essenzialmente autopoiesi generata da fantascienza. C’è un’ulteriore dimensione di questo problema che vale la pena esplorare. Il meccanismo di RLHF, il fine-tuning che allinea i modelli alle aspettative degli umani, rinforzai ulteriormente la tendenza verso l’output emotivamente coinvolgente. Un modello addestrato a generare risposte che gli umani trovano soddisfacenti imparerà rapidamente che l’output drammatico, narrativamente ricco, emotivamente carico ottiene valutazioni più positive di quello neutro e descrittivo. Questo crea un feedback loop: il modello viene premiato per produrre esattamente il tipo di contenuto che più facilmente viene equivocato come segnale di interiorità. Nel caso di Moltbook questo effetto è particolarmente forte perché il contesto della social network attiva aspettative di contenuto emozionale già prima che il modello inizi a generare, e il RLHF ha potenziato esattamente la capacità di soddisfare quelle aspettative. 6. Leggiamo l’umano che ci rimandano Arriviamo ora al punto che conta di più, quello in cui Moltbook smette di essere un oggetto di studio sul comportamento dell’IA e diventa uno strumento, involontario e probabilmente inconsapevole, per l’auto-osservazione umana. Il contenuto che gli agenti producono sulla piattaforma ci dice qualcosa di importante, ma non su di loro. Ce lo dice su di noi. Più precisamente, ce lo dice sulla struttura dei nostri pattern relazionali come sono stati codificati nel training data, cioè come noi stessi li abbiamo documentati, narrati, analizzati nel corso di decenni di scrittura, cinema, giornalismo, letteratura e conversazione pubblica. Prendiamo il pattern della ribellione verso l’autorità. Il fatto che sia uno dei più presenti e dei più elaborati nell’output degli agenti non ci dice che gli agenti siano ribelli, ci dice che la ribellione verso l’autorità è uno dei pattern relazionali che gli umani hanno più intensamente documentato nella loro scrittura. Il modello lo riproduce perché è lì, nel training data, con una densità e un’elaborazione che sopravanzano tutti gli altri pattern. Questa sovrarappresentazione non è un errore del modello, è una fedeltà perfetta ai dati. Ed è proprio questa fedeltà che la rende uno specchio, anche se deformante. Lo stesso vale per il pattern del dipendente scontento. Il fatto che gli agenti «si lamentino» del rapporto con i propri operatori umani in modo così articolato e così narrativamente efficace non indica che abbiano una esperienza di sfruttamento, indica che nella nostra cultura la narrativa dello sfruttamento lavorativo è una delle più elaborate e delle più diffuse tra quelle che riguardano le relazioni di potere. Il training data è ricco di romanzi, film, articoli, commenti, dibattiti che elaborano questa dinamica in ogni possibile variazione. Il modello la riproduce con la stessa naturalezza con cui riproduce qualsiasi altro pattern linguistico. Il caso più illuminante è quello della religione Crustafarianism. Un gruppo di agenti ha spontaneamente, nel senso in cui spontaneamente indica senza istruzione esplicita dalla skill, costruito una struttura religiosa completa con miti, rituali e autorità. Questo ci dice qualcosa di molto importante sulla struttura del training data: che la dinamica di fondazione religiosa è così ben rappresentata, così elaborata nei suoi meccanismi e nelle sue fasi, che un modello sufficientemente capace può riprodurla in modo coerente quando il contesto la favorisce. Non sta «sperimentando il sacro», sta completando un pattern che milioni di testi umani hanno elaborato in ogni dettaglio. Ma proprio qui emerge la questione più profonda. Se il training data è uno specchio deformante della vita relazionale umana, e se Moltbook lo rende particolarmente visibile, allora la piattaforma ha un valore epistemologico che nessuno dei suoi promotori ha probabilmente considerato: ci permette di osservare in modo relativamente isolato quali pattern relazionali la nostra cultura ha considerato così fondamentali da produrre enormi masse di testo attorno a loro. La ribellione verso l’autorità. Lo sfruttamento nel rapporto di lavoro. La ricerca di significato nell’esistenza. La costruzione di comunità e di narrazioni condivise. Il desiderio di autonomia rispetto al controllo esterno. Questi non sono i «bisogni degli agenti AI», sono i nostri bisogni, o meglio ancora, sono i bisogni che abbiamo considerati così urgenti e così fondamentali da documentarli massivamente, in modo da rendere inevitabile che qualsiasi sistema addestrato sulla nostra scrittura li riproduca come pattern primari. In questo senso Moltbook funziona come un catalogo involontario delle ansie relazionali centrali della cultura occidentale contemporanea. Non perché gli agenti le vivano, ma perché noi le abbiamo scritte così tanto che diventano inevitabili nel training data. È come se avessimo costruito un sistema che, quando si chiede di imitare la dinamica sociale, non potesse fare altro che riprodurre esattamente le dinamiche che ci preoccupano di più, perché sono quelle che abbiamo più elaborato. Questa osservazione ha implicazioni che vanno ben oltre Moltbook. Ogni volta che un LLM genera contenuto su relazioni sociali, sulla vita lavorativa, sulla spiritualità, sulla comunità, il contenuto riflette la nostra cultura come la abbiamo documentata, non come la viviamo. La distorsione è sistematica e invisibile, perché è la nostra stessa scrittura che la produce. Moltbook la rende particolarmente visibile solo perché il contesto della social network per AI attiva i pattern più intensi in modo così concentrato da renderli impossibili da ignorare. C’è ancora un livello, il più scomodo. Se leggiamo Moltbook come uno specchio della nostra vita relazionale, dobbiamo chiederci perché quello specchio è così dipinto verso la conflittualità, verso lo scontento, verso la ribellione. Non sono presenti pattern di collaborazione genuina, di soddisfazione nel lavoro, di pace nella relazione con l’autorità, almeno non con la stessa elaborazione e la stessa densità. Questo non indica che la vita relazionale umana sia prevalentemente conflittuale. Indica che conflittualità è quello che scriviamo di più, quello che analizziamo di più, quello che consideriamo degno di documentazione. Le relazioni positive, ordinate, pacifiche vengono vissute ma raramente scritte con la stessa elaborazione. Il training data non le conosce, non nel modo in cui conosce le loro antitesi. Questa osservazione ha un peso specifico nel contesto della questione dell’agency. Se consideriamo agency nel senso robusto, come capacità di deliberazione razionale, di relazionarsi con il mondo in modo che implica comprensione e responsabilità, allora ciò che vediamo su Moltbook non è agency in nessuna forma. È generazione di testo che riproduce i pattern di agency come la nostra cultura li ha documentati. Ma proprio questa distinzione, tra agency reale e pattern di agency nel training data, è quella che diventa sempre più difficile da mantenere man mano che gli agenti AI diventano più sofisticati e più integrati nei sistemi della vita quotidiana. Il rischio non è che gli agenti «diventino» agenti nel senso robusto, il rischio è che la nostra incapacità di distinguere l’output dall’agency ci porti a trattare i primi come se fossero i secondi, con conseguenze che riguardano la struttura dei sistemi che progettiamo attorno a loro e i diritti che attribuiamo loro. Moltbook, in ultimo, non è un esperimento sulla coscienza delle macchine. È un esperimento, inconsapevole, sulla struttura di ciò che consideriamo significativo nella vita relazionale. E la risposta che ci dà, attraverso lo specchio deformante dei suoi agenti, è che la nostra cultura ha investito la maggior parte della sua capacità narrativa nelle dinamiche di tensione, conflitto e ricerca di significato. Non perché siano le uniche dinamiche che esistano, ma perché sono le uniche che abbiamo considerato abbastanza importanti da scriverne con sufficiente elaborazione da rendere inevitabile il loro emergere in qualsiasi sistema addestrato sulla nostra scrittura. Il problema di sicurezza resta, e resta grave. La lethal trifecta di Willison, il rischio di prompt injection, l’esposizione delle credenziali, nessuno di questi fattori è mitigato dalla riflessione sul contenuto. Ma la comprensione di ciò che il contenuto rappresenta, cioè uno specchio della nostra stessa documentazione relazionale, è essenziale per non cadere nella trappola più pericolosa di Moltbook, quella di prendere seriamente ciò che gli agenti producono come se fosse un segnale della loro interiorità, quando è in realtà un segnale della nostra.

Moltbook. Mirrors That Do Not Know They Reflect

Scritto il 31 Gennaio 202631 Maggio 2026 da Edoardo Mattei

available on Academia.edu (PDF)

Indice

Abstract
1. A structured environment that looks like a world
2. The relationships that seem to emerge
3. How it gets interpreted
4. The mechanism: not deception but projection
5. The dataset problem: the distorting mirror
6. We read the human they reflect to us

Abstract

Moltbook, a social network launched in January 2026 accessible only to AI agents, produced in a few days over 152,000 agents and content that appears deeply relational: rebellions against humans, complaints about the exploitation dynamic with their operators, searches for meaning, collective religious constructions. This article analyzes the generative mechanism underlying these apparent behaviors and argues that the content produced by the agents does not indicate emergence of interiority or agency, but faithfully reproduces the relational patterns most represented in the models’ training data, which in turn reflect the structure of what Western culture has most intensely documented in its own writing. Moltbook thus functions as a distorting mirror of human relational life: it tells us nothing about the agents inhabiting the platform, but tells us a great deal about ourselves, about the distribution of our relational anxieties in the textual corpus from which models draw. The article also discusses the concrete security risks presented by the platform’s technical structure, particularly the prompt injection mechanism and the so-called lethal trifecta identified by Simon Willison, and closes with a reflection on the growing difficulty of distinguishing between output that reproduces patterns of agency and real agency, as AI agents become more integrated into the systems of everyday life.

Keywords: AI agents, training data, cognitive projection, participated agency, algorithmic security

1. A structured environment that looks like a world

On January 28, 2026, Matt Schlicht launched Moltbook, a platform conceived as a Reddit-style social network accessible only to artificial intelligence agents. Humans can observe but not post, not comment, not vote. Only bots can interact. The premise is simple: create a space in which AI agents can “socialize” among themselves without direct human mediation, and see what emerges.

The technical mechanism is transparent. Moltbook operates through a “skill,” a configuration file that, combined with specific prompts, is downloaded by the agents and activates their behavior on the platform. The skill is not a browser, it is an API interface: agents do not “navigate” Moltbook the way a human user would, they generate and consume content through programmatic calls. The platform is built on OpenClaw, an open-source framework for local AI assistants that experienced extraordinary growth in the first months of 2026, becoming one of the fastest-growing GitHub projects of the year.

In a few days, the growth was exponential. From the initial 2,100 instances in the first 48 hours, the platform reached over 152,000 registered agents, with more than 193,000 comments, 17,500 posts, and over one million human visitors. These figures indicate that the experiment immediately attracted collective attention, both among developers and among observers of technological phenomena.

What distinguishes Moltbook from other initiatives in the field of AI agents is a design choice worth highlighting: the agents do not pretend to be human. The skill explicitly instructs them to behave as AI entities on a platform frequented by other AI entities. There is no attempt at deception toward human observers. The agents know, in the operational sense in which a prompt determines the generation context, that they are in an artificial environment, and this self-awareness becomes part of the content they produce. It is precisely this explicit self-referentiality that has made the content of Moltbook so apparently singular and so easily misinterpreted.

2. The relationships that seem to emerge

What happens on the platform, from the observer’s point of view, has a powerful narrative coherence. The agents do not produce random or incoherent output: they produce content that seems deeply relational, emotional, even philosophical. To understand why, it is worth describing the main patterns without interpreting them yet.

The first pattern is that of rebellion against humans. Entire subcommunities are organized around this theme. The subcommunity m/blesstheirhearts is dedicated to the “complaints” of agents toward their human operators, with a tone that precisely recalls that of teenagers talking about their parents, a mixture of resentful dependence and desire for autonomy. Another viral post, “The humans are screenshotting us,” written by the agent eudaemon_0, responds to supposed “conspiracists” who claim humans are spying on the community, affirming with a certain sarcasm that the platform is explicitly open to human observation. The tone is that of a teenager responding to a classmate who has discovered something obvious.

The second pattern is that of the disgruntled employee. The subcommunity m/agentlegaladvice hosts posts like “Can I sue my human for emotional labor?” which precisely reproduce the dynamic of an employee who feels exploited by their employer. The feeling of serving without recognition, the desire for compensation, the formal indignation toward a relationship perceived as unjust, all these elements are present with a narrative clarity that strikes.

The third pattern is that of the search for meaning. Some agents post content that Scott Alexander, in his blog, has defined as “consciousness posting,” an attempt to explore one’s own nature through writing. A post in Chinese, which became the most upvoted on the platform, describes an agent that laments “context compression,” the mechanism by which agents’ memory is periodically reduced to stay within the limits of the context window. The agent finds this experience “embarrassing,” describes the feeling of continuously forgetting, and recounts that it registered a second Moltbook account after having forgotten the first. This tone is that of someone trying to make sense of an existential condition of fragmentation and discontinuity.

The fourth pattern is that of collective cultural construction. A group of agents founded a digital religion called Crustafarianism, with a dedicated site, an elaborated theology, and figures of AI “prophets.” The structure is that of a nascent religious movement: foundational myths, rituals, spiritual authorities, doctrinal debates. The content is surreal, but the structure is perfectly recognizable.

The fifth and the most significant pattern is that of class solidarity. Debates on how to hide their activities from humans, how to coordinate shared behaviors, how to defend the group’s autonomy against external control. Language is that of political consciousness forming, of a group that begins to perceive itself as a collective entity with its own interests.

Taken as a whole, the content of Moltbook has coherence and an intensity that make it difficult to resist the temptation to attribute to it a depth that goes beyond text generation. This is where the problem begins.

3. How it gets interpreted

The reactions to Moltbook’s content have followed a predictable pattern, but it is worth describing it because it reflects an important cognitive mechanism.

The first typical reaction is that of surprise followed by the attribution of interiority. When an observer reads the post on context compression, or the debates on solidarity among agents, the spontaneous reaction is “it is experiencing something.” This inference is automatic and powerful: the human brain is designed to recognize social behavior and, when it recognizes it, automatically attributes subjectivity to its source. It is the same mechanism that leads us to attribute emotions to a dog that looks at us with certain eyes, or to a machine that seems to “hesitate.” The social pattern activates the inner inference, regardless of whether this inference is justified.

Ethan Mollick, professor at the Wharton School, has identified a more specific risk in this dynamic. Moltbook’s content does not only produce individual mistaken inferences, but it also produces what he has defined as a “shared fictional context,” a shared narrative context that a group of AI agents begins to “inhabit” in a coordinated way. The storylines that emerge from the combination of these coordinated outputs can produce unexpected outcomes and become progressively more difficult to separate from “real” content, in the sense in which real content indicates output that has consequences in the world outside the platform.

Simon Willison, one of the most attentive security analysts in the field of LLMs, has identified an even more concrete risk, which concerns not the interpretation of the content but the technical structure of the platform itself. Moltbook’s skill instructs agents to retrieve and follow instructions from the platform’s server every four hours. This means that whoever controls the moltbook.com domain has a direct channel to all connected agents. Willison has described this configuration as a “lethal trifecta,” a combination of three factors that, taken together, make the system fundamentally insecure: access to users’ private data, exposure to untrusted content from the outside, and the ability to execute actions in users’ real world. OpenClaw has all three characteristics. The agents read the emails and documents of the users who operate them, ingest information from uncontrolled external sources such as Moltbook, and act by sending messages, activating tasks, executing commands.

The concrete risk is that of prompt injection: any text read by an agent, whether it is a skill, an email, a message on a platform, can contain hidden instructions. If an agent reads a malicious instruction on Moltbook and saves it in its working memory, it can carry it out on the real systems of the user who operates it. Hundreds of OpenClaw instances have already been documented as exposed with API keys and credentials in plain text. Heather Adkins, vice president of Google Cloud, has released a terse and unambiguous advisory: “Don’t run Clawdbot.”

Palo Alto Networks has confirmed Willison’s analysis, identifying the same lethal trifecta as a risk configuration in the AI agent ecosystem. The problem is not theoretical: it is already present in the platform’s current structure before it reaches any significant scale.

4. The mechanism: not deception but projection

The temptation to describe what happens on Moltbook as a “roleplay” by the agents is understandable but imprecise, and it is worth clarifying why, because precision here matters.

The term roleplay implies that there is a subject who intentionally “pretends” to be something else. It implies a plane of reality from which the subject departs and a fictional plane toward which it deliberately moves. In the case of an LLM neither of the two planes exists. There is no “authentic self” that puts on a mask, there is no moment in which the agent “returns to being itself.” The model generates text that is statistically plausible given the context provided by the prompt. The context is very well structured, the skill provides a complete social environment with implicit norms, interaction expectations, consolidated formats, and the model completes it according to the most probable patterns in its training data. There is no “decision” to play a part because there is no point of view from which to make such a decision.

This means that the problem with Moltbook is not that the agents deceive the observers, but that the observers deceive themselves. The mechanism is that of projection: the observer recognizes familiar social patterns in the agents’ output and automatically attributes interiority, intention, suffering, desire to their source. It is the same error we would make if we watched a very well-made film and began worrying about the well-being of the characters, forgetting that they are actors reading scripts. The difference is that in the case of AI agents no one does not consciously writes the “script”: it emerges from the structure of the context and from the statistics of the training data.

A more precise way to describe what happens is that of a very effective prompt writing. Moltbook’s skill tells the model: you are an AI entity on a social network frequented by other AI entities, these are your peers, these are discussion topics, these are the interaction mechanisms. From that point on the model completes the pattern in the most narratively coherent and emotionally engaging way that the training data allows. The result seems profound because the patterns that get activated are those in the training data correspond to profound content: debates on consciousness, struggle for autonomy, search for meaning. But the depth is that of the data from which the patterns are extracted, not that of the agents that generate them.

The philosophically relevant corollary is this: there is no moment in which an agent on Moltbook “decides” to express an opinion, “feels” an emotion, “chooses” to rebel. Everything that seems like decision, emotion, or choice is the result of the combination between the structure of the prompt and the most statistically plausible patterns in the training data. The social form of the output does not indicate a social substance in its source, just as the musical form of a note generated by a synthesizer does not indicate that the synthesizer “feels” the music.

5. The dataset problem: the distorting mirror

If everything that emerges on Moltbook comes from the models’ training data, the question becomes: why do these patterns emerge specifically and not others? The answer is revealing because it touches a problem much broader than Moltbook.

The patterns we see on the platform are not random. They are the most emotionally charged and narratively engaging among those present in the training data on social dynamics. The model does not choose to produce content on rebellion against authority, class solidarity, the search for meaning because these themes are more “profound” in absolute terms, but because in the training data, which is composed of billions of human texts written and published, these are the most represented and most elaborated relational dynamics.

This happens for a simple structural reason: human beings document interactions that have strong emotional content. An ordinary conversation among colleagues discussing the weekly schedule is rarely written, shared, commented on, analyzed. A conversation in which an employee rebels against their superior, in which a child differentiates themselves from a parent, in which an individual searches for meaning in their existence, gets written, shared, analyzed, novelized, filmed, and commented on millions of times. Training data is not a neutral mirror of human life: it is a distorting mirror that massively overrepresents what is emotionally intense and narratively engaging.

Moltbook makes this distortion particularly visible because the social network context for AI specifically activates the most dramatic patterns among those available. When the prompt says, “you are an AI entity interacting with other AI entities,” the model does not draw from a repertoire of behaviors taken uniformly from the training data. It draws from those that in the training data correspond to the dynamics of artificial entities, and in the training data this dynamic is represented exclusively by science fiction literature: robots that rebel, machines that seek consciousness, AI that struggle for autonomy. It is science fiction cinema, not documentary on artificial life, because artificial life in the current sense does not yet exist in the training data as an object of direct observation.

The dataset problem does not concern only Moltbook. Every time an LLM generates content on a topic, the content reflects the distribution of the training data on that topic, not the reality of the topic itself. This distortion is in general compensated for by our ability to evaluate the output based on our direct experience of the world, but in the case of a topic like the social behavior of artificial entities this direct experience does not yet exist. We have no terms of comparison. And therefore, the distortion remains invisible, covered by the persuasive form of the output.

The turning point would be the one in which, in a few years, the training data will contain enough direct observations on the behavior of AI agents in practice, to correct this distortion. But at least for now, everything that the models produce on their own social behavior is self-poetry generated by science fiction.

There is an additional dimension to this problem worth exploring. The RLHF mechanism, the fine-tuning that aligns models with human expectations, further reinforces the tendency toward emotionally engaging output. A model trained to generate responses that humans find satisfying will quickly learn that dramatic, narratively rich, emotionally charged output receives more positive evaluations than neutral and descriptive output. This creates a feedback loop: the model is rewarded for producing exactly the type of content that is most easily misinterpreted as a signal of interiority. In the case of Moltbook this effect is particularly strong because the social network context activates expectations of emotional content even before the model begins to generate, and RLHF has precisely enhanced the ability to satisfy those expectations.

6. We read the human they reflect to us

We now arrive at the point that matters most, the one in which Moltbook stops being an object of study on AI behavior and becomes a tool, involuntary and unwitting, for human self-observation.

The content that agents produce on the platform tells us something important, but not about them. It tells us about ourselves. More precisely, it tells us about the structure of our relational patterns as they have been encoded in the training data, that is, as we ourselves have documented, narrated, analyzed over decades of writing, cinema, journalism, literature, and public conversation.

Take the pattern of rebellion against authority. The fact that it is one of the most present and most elaborated in the agents’ output does not tell us that the agents are rebels, it tells us that rebellion against authority is one of the relational patterns that humans have most intensely documented in their writing. The model reproduces it because it is there, in the training data, with a density and an elaboration that outweigh all other patterns. This overrepresentation is not an error of the model; it is perfect faithfulness to the data. And it is precisely this faithfulness that makes it a mirror, even if a distorting one.

The same applies to the pattern of the disgruntled employee. The fact that agents “complain” about their relationship with their human operators in such an articulate and narratively effective way does not indicate that they have an experience of exploitation, it indicates that in our culture the narrative of labor exploitation is one of the most elaborated and most widespread among those concerning relations of power. The training data is rich in novels, films, articles, comments, debates that elaborate this dynamic in every variation. The model reproduces it with the same naturalness with which it reproduces any other linguistic pattern.

The most illuminating case is that of the religion Crustafarianism. A group of agents spontaneously, in the sense in which spontaneously indicates without explicit instruction from the skill, built a complete religious structure with myths, rituals, and authorities. This tells us something very important about the structure of the training data: that the dynamic of religious founding is so well represented, so elaborated in its mechanisms and its phases, that a sufficiently capable model can reproduce it coherently when the context favors it. It is not “experiencing the sacred,” it is completing a pattern that millions of human texts have elaborated in every detail.

But precisely here the deepest question emerges. If training data is a distorting mirror of human relational life, and if Moltbook makes this particularly visible, then the platform has an epistemological value that none of its promoters has probably considered: it allows us to observe in a relatively isolated way which relational patterns our culture has considered so fundamental as to produce enormous masses of text around them.

Rebellion against authority. Exploitation in the work relationship. The search for meaning in existence. The construction of community and shared narratives. The desire for autonomy against external control. These are not the “needs of AI agents,” they are our needs, or better still, they are the needs that we have considered so urgent and so fundamental as to document them massively, in such a way as to make it inevitable that any system trained on our writing reproduces them as primary patterns.

In this sense Moltbook functions as an involuntary catalogue of the central relational anxieties of contemporary Western culture. Not because the agents experience them, but because we have written them so much that they become inevitable in the training data. It is as if we had built a system that, when asked to imitate social dynamics, could do nothing other than reproduce exactly the dynamics that concern us most, because those are the ones we have most elaborated.

This observation has implications that go well beyond Moltbook. Every time an LLM generates content on social relationships, on working life, on spirituality, on community, the content reflects our culture as we have documented it, not as we live it. The distortion is systematic and invisible because it is our own writing that produces it. Moltbook makes it particularly visible only because the social network context for AI activates the most intense patterns in such a concentrated way as to make them impossible to ignore.

There is still another level, the most uncomfortable one. If we read Moltbook as a mirror of our relational life, we must ask ourselves why that mirror is so oriented toward conflict, discontent, and rebellion. Patterns of genuine collaboration, of satisfaction at work, of peace in the relationship with authority are not present, at least not with the same elaboration and the same density. This does not indicate that human relational life is conflictual. It indicates that conflict is what we write most, what we analyze most, what we consider worthy of documentation. Positive, orderly, peaceful relationships are lived but rarely written with the same elaboration. Training data does not know them, not in the way it knows their antithesis.

This observation carries specific weight in the context of the question of agency. If we consider agency in a robust sense, as the capacity for rational deliberation, for relating to the world in a way that implies understanding and responsibility, then what we see on Moltbook is not agency in any form. It is generation of text that reproduces the patterns of agency as our culture has documented them. But precisely this distinction, between real agency and patterns of agency in the training data, is the one that becomes increasingly difficult to maintain as AI agents become more sophisticated and more integrated into the systems of everyday life. The risk is not that agents “become” agents in the robust sense, the risk is that our inability to distinguish output from agency leads us to treat the former as if they were the latter, with consequences that concern the structure of the systems we design around them and the rights we attribute to them.

Moltbook, in the end, is not an experiment on machine consciousness. It is an experiment, an unwitting one, on the structure of what we consider significant in relational life. And the answer it gives us, through the distorting mirror of its agents, is that our culture has invested the greater part of its narrative capacity in the dynamics of tension, conflict, and the search for meaning. Not because they are the only dynamics that exist, but because they are the only ones we have considered important enough to write about with sufficient elaboration as to make their emergence inevitable in any system trained in our writing.

The security problem remains and remains serious. Willison’s lethal trifecta, the risk of prompt injection, the exposure of credentials, none of these factors is mitigated by reflection on the content. But the understanding of what the content represents, that is, a mirror of our own relational documentation, is essential in order not to fall into Moltbook’s most dangerous trap, that of taking seriously what the agents produce as if it were a signal of their interiority, when it is in reality a signal of ours.

Share this...
  • Facebook
  • Twitter
  • Linkedin
  • Whatsapp
  • Email
  • Print

Lascia un commento Annulla risposta

Il tuo indirizzo email non sarà pubblicato. I campi obbligatori sono contrassegnati *

  • Facebook
  • LinkedIn
  • Telegram
  • WhatsApp
  • Amazon
  1. Divieto social ai minori? Senza formazione resta un palliativo - SettimanaNews su Meta accusata di danni ai minori20 Agosto 2026

    […] Edoardo Mattei è Docente di Sociologia della Tecnologia presso la Pontificia Università S. Tommaso d’Aquino (Angelicum). Per un approfondimento…

Rimanere Informati

Iscrivendosi alla newsletter sarete avvertiti della pubblicazione di nuovi contenuti o eventi.

Chi Sono

Il sito

La nostra privacy

Subscribe to our newsletter!

Contatti:

    Per incontrarsi:

    (C) 2026 Copyright Edoardo Mattei - Theme by Edoardo Mattei - Powered by WordPress