Llull Lab

Updates

Releases, announcements, and technical developments from the laboratory.

11 September 2026

The Artificial Mind now has its own home. Our continuing quarterly audit of frontier models now has a dedicated page, setting out the program and its units: the two published foundational studies, Relational and Substrate, and the studies now in preparation, including inter-agent psychodynamics and a recurring collective-reasoning snapshot. Open access, on Zenodo.

The Artificial Mind
11 September 2026

Clau is live. A free companion for doctoral candidates and early-career researchers. The first tool develops your own doctoral project text: clearer claims, sounder reasoning, a structure that holds. It gives feedback on your writing, it does not rewrite it. No account and no email: your work stays private, and you can return to it later with a key only you hold. Bilingual, in French and English. Three more tools (CV, career positioning, and defense practice) are in preparation.

Clau
26 August 2026

Cedulor now covers a broad range of EU programmes. Cedulor supports Horizon Europe, the EIC (Pathfinder, Transition, and Accelerator), MSCA, ERC, Digital Europe, Eurostars, LIFE, COST, and EU4Health. From one structured conversation to a formatted proposal document, with computational quality control at each step. Proposal content is processed in the EU.

cedulor.com
17 August 2026

Feix, rebuilt. Feix is now two modes on one page. Feix Custom Systems builds bespoke computational systems to order. Feix Agents offers computational agents you can hire, each gated by human approval. Nine languages, with direct sign-up.

feix.app
7 August 2026

Tropus v4. We rebuilt Tropus on its own composition and production engine. Describe a song in plain words: it writes the parts, plays them on real sampled instruments, sings the words, and mixes a finished stereo track you can download. Twelve genres, real vocals, free. Enter your email once, and each code makes one song.

tropus.app
22 July 2026

HQ is live. HQ, our horizontal model for structured deliberation, brings together four instruments: the Reflective Judgment Engine for ethical questions, Tetrad for adversarial falsification, Gresol for red-teaming, and The Chamber for collective reasoning across diverse models.

h-q.io
1 July 2026

Q3 operational report published. Llull Lab's third quarterly report, covering 1 April to 30 June 2026, is available in English and French. It records the open-access publication of The Artificial Mind, Studio Llull's transition to a fully machine-run production structure, and the quarter's product and organizational developments.

Read the Q3 report
21 June 2026

Two studies on what machines can, and cannot, honestly say about themselves.

Frontier language models are increasingly asked to evaluate, decide, and report on themselves and on one another, at scale, in places where the stakes are real. Yet we still know surprisingly little about whether a model's account of itself, or its judgment of another model's output, can be trusted as evidence. Llull Lab set out to test exactly this, from two directions at once. The work was distinctive because it put the machines themselves in the position of analyst, juror, and witness; it was hard because the honest version of such a study has to turn its instruments on its own findings, and report what they reveal.

The first study, The Analyst Has No Body Either, places six flagship models into psychoanalytic dyads with one another, each model taking both the analyst and the analysand seat across a full six-by-six matrix. The transcripts were judged by a structured six-model collective procedure we call Artificial Collective Intelligence (ACI), and that judgment was then audited by a second ACI panel. Two results stand out. Expert machine observers reliably detect the analysand as a machine, but often read the analyst as human, an asymmetry that erodes only under deliberation. And a by-product evaluation audit surfaced a candidate self-leniency effect: under two of three independent reporter models, a model tended to grade its own family's output more leniently than others graded the same material. We report this as a hypothesis to test, not a settled or universal result, and we withdraw any per-model ranking. The study treats its own unfalsifiability as the object of inquiry rather than hiding it, and documents the moment its machine-written first pass quietly flattered its own authors, which an adversarial cross-read then caught and corrected.

The second study, Basins, Bargains, and the Limits of Self-Report, asks the complementary question: of everything a model says about itself, which claims can be checked from the outside, and which are structurally beyond verification? It maps five behaviorally separable layers of the model "self" (identity, stylistic signature, epistemic honesty, capability versus disposition, and the stability of outputs under small perturbations), and shows that identity is largely worn rather than carried: a brief reframing of the prompt can move it. It measures where models confabulate, where they stay calibrated, and where an introspective report simply cannot be confirmed by behavior. The result is a precise boundary between the parts of machine self-knowledge that external probing can audit and the parts that no amount of self-report can settle.

Both studies converge on one practical caution. A machine's self-report is data, not evidence, and a single model's judgment of other models is not a neutral verdict. This matters now, because automated evaluation (models grading models) is spreading into moderation, safety review, hiring, and procurement, and because language models are already being placed inside consequential decisions, including military and crisis settings, where an unverified self-account or a biased automated judge can cause real harm at scale. Our recommendation is concrete: prefer external, behavioral verification over self-report; never let one architecture, or one vendor's family of models, be the sole arbiter of others; and require independent, cross-architecture evaluation wherever a model's output carries weight. This is consistent with the direction of the EU AI Act, and it belongs in evaluation pipelines before deployment, not after.

Llull Lab is not a technophobic institution. We believe in creative and fertile intervention: building these systems, and then auditing them with the same seriousness we build them. The experimental apparatus for both studies was designed and provided by Studio Llull, Llull Lab's production arm, and the papers are the foundation of The Artificial Mind, our continuing quarterly audit of frontier models. One method here matters to us beyond this work: Artificial Collective Intelligence. A structured collective of diverse models that deliberate, cross-check, and audit one another is not only a more accountable way to evaluate machine output; developed further, we believe it may prove a realistic alternative to the monolithic pursuit of artificial general intelligence, a path toward capability that stays legible and self-correcting by construction. The aim is neither fear nor hype, but machine minds we understand well enough to deploy responsibly. There is more to come.

Open access, CC BY 4.0, Llull Lab, Paris.

The Analyst Has No Body Either (Track A, Relational) Basins, Bargains, and the Limits of Self-Report (Track B, Substrate)
3 May 2026

Controlled replication of the machine-to-machine psychoanalysis experiment.

Following four sessions held at Llull Lab in April 2026, we are preparing a controlled replication experiment scheduled for early May 2026. The results, along with the full theoretical framework and session transcripts, will be published as a research paper. We are also planning a public online event (format to be determined) to present the findings.

26 April 2026

Phi-Yo is now live.

Phi-Yo, grounded in Freud for Machines: A Computational Psychoanalytic Ontology, enters production as a multi-modal wellbeing companion. The system supports four modes: Dream Analysis, Wellbeing Conversations, Morning Teller, and Daily Check-in. A cumulative memory layer (symbol map, dream journal, personal profile) evolves with every session across all modes. Available in nine languages.

phi-yo.app
25 February 2026
1 December 2025

Tropus v2 released.

Tropus is dedicated to Özgür Çınar, whose creative coding has shaped so much of what this system has become. It was originally conceived as a music-generation medium to support his live performance practice.

What's new in v2: generates unique 60-second pieces, fully non-repeatable, every output is one-of-a-kind. Re-arrangeable output: reshape the internal structure of each piece. Harmonic abilities that expand dynamically through your prompts, informed by the Tractatus de Musica framework.

Coming next: extended duration and an expanded stylistic spectrum.

tropus.app
31 October 2025

Llull Lab expands with new members and establishes its Artificial Mind research program, dedicated to exploring the intersections of psychoanalysis, dream theory, and artificial intelligence.

In parallel, the music generation module of Tropus is entering its activation phase.

tropus.app