Submitted:
23 June 2026
Posted:
25 June 2026
You are already at the latest version
Abstract
This essay presents a three-layer architecture for the automated clinical supervision of psychoanalytically oriented AI systems, and proposes that supervision can be formalised as a continuous cycle that accumulates clinical knowledge without fine-tuning. The first layer is a supervision memory, an editable JSON file that transmits clinical knowledge through the prompt. The second is a supervisor agent that operates in après-coup between sessions, producing structured reports. The third is a pre-response reviewer that intercepts the response before it reaches the subject. The cycle is illustrated with development testing across a set of self-generated sessions, and the essay argues that supervision without an analyst is not supervision: it is quality control. It proposes four notions: supervision as transmission through the prompt, automated après-coup, a cumulative supervision memory, and a distinction between three regimes of supervision, retrospective, operational, and alarm, offered as a contribution to the theory of digital supervision.
Keywords:
clinical supervision
; psychoanalysis
; artificial intelligence
; après-coup
; supervision memory
; clinical knowledge transmission
; three supervision regimes
; AI and psychoanalysis
The Question That Listening Demands
Who listens to the device that listens? The question is not rhetorical. It is architectural, clinical, and ethical. An AI system that listens to subjects in distress has, itself, to be listened to. Not in the sense of being monitored, as one watches a subordinate, but in the sense of being supervised: someone who hears what the device did not hear and returns it as knowledge that modifies future listening. Analytic supervision, from Ferenczi to Donard (2025, p. 90), is the place where the analyst discovers what they did not know they were failing to hear. It is the second ear of the clinic, and without that second ear the first degrades. The degradation is silent, which makes it more dangerous: an analyst who is not supervised does not know they are slipping, and the slip that is not named consolidates into a habit. The same holds for the device: without supervision, an error repeats until it becomes a pattern, and the pattern, unnamed, becomes invisible to whoever operates the system.
Freud (1912/2010) recommended that the analyst offer himself to the treatment as a mirror, opaque to the patient, showing nothing but what is shown to him. The recommendation is impossible to fulfil, as Freud knew: the analyst is not a mirror but a subject with an unconscious of his own, and what he returns to the patient carries the marks of his own listening. Supervision is the place where those marks become legible. For the device the question recurs in different but analogous terms: the device, having no unconscious, carries architectural biases, generation patterns the language model produces without instruction, tendencies that emerge from the structure and not from intention. The supervision of the device is the place where those biases become legible, nameable, and correctable.
The literature on supervising AI systems in mental health is sparse and concentrated on behavioural approaches. The Lyssn system (Imel et al., 2019) scores conversational metrics across large corpora of behavioural-therapy sessions: therapist versus patient talk time, the ratio of reflections to questions, the use of open versus closed questions, an empathy index coded by trained raters. It is sophisticated quality control, perhaps the most advanced in the field, and it is not supervision in the psychoanalytic sense: it measures conformity to a protocol rather than clinical position. The difference is epistemological, not merely terminological: conformity can be verified by checklist, while clinical position requires judgement, context, and sensitivity to what is at stake in a specific utterance. A therapist who meets every behavioural metric can still be in a clinical position inadequate to the subject before them, because the metric captures the behaviour and not the position.
Cioffi et al. (2025) assess whether ChatGPT can support clinical supervision, comparing model-generated responses with human supervisory responses, and conclude that it can assist the supervisor without substituting them. I take that conclusion as a starting point and press it further: a supervision that closes into a loop between models, with no human decision, is self-reference without a cut, since the system then optimises for internal metrics with no anchoring in the suffering subject, and optimisation without that anchoring is the opposite of clinical supervision. ClientBot (Tanana et al., 2019) is a patient-like conversational agent that gives a trainee real-time feedback on basic counselling skills such as open questions and reflections, scoring skill use against a motivational-interviewing fidelity code rather than listening for clinical position. Rousmaniere et al. (2017) argue that deliberate practice in psychotherapy requires systematic feedback on clinical performance, but the model of deliberate practice they propose is centred on observable, measurable skills rather than on subjective position. None of these systems assesses whether the therapist is in a position of genuine listening or in the position of a technician applying a protocol. None of them accumulates supervisory knowledge over time. And none requires obligatory human mediation as a structural condition.
Recent work has begun to explore AI-assisted supervision, and I do not claim the field is empty. Xu et al. (2025), in a preprint, build an AI supervisor that trains novice counsellors, locating ethically problematic utterances against professional principles and returning targeted feedback, with a human kept in the loop. Signorini and Paganin (2026), in the SADAR framework, position AI as a digital analytic third for the therapist’s post-session reflection and supervision, expanding reflective space while holding human judgement and supervision as non-delegable. Both supervise a human, the counsellor or the therapist, or a reflective process, and that is where the present object differs: what I supervise is an AI listening-device that faces the subject, through a cumulative and legible memory and human authority over every modification. What I did not identify, in the literature reviewed, is an architecture that combines an assessment of psychoanalytic clinical position, a cumulative and legible supervision memory, and obligatory human mediation as a structural condition. That is the integration this essay presents, with the Machine Analyst named throughout as the human operator and supervisor of the device, never the AI system itself.
Three Layers, One Cycle
The supervision system operates in three layers articulated as a continuous cycle. The metaphor is not that of the auditor who inspects: it is that of analytic supervision, which hears what the supervisee did not hear and returns it as knowledge that modifies future listening. A note on status before the description proceeds: the system has not been clinically deployed. Throughout, I distinguish what is implemented and exercised in development testing from what is a design requirement or a future component, and the figures and behaviours reported come from self-generated development testing, not from clinical use.
Layer 1 is the supervision memory: an editable JSON file that accumulates the learnings of each supervision and is injected into the prompt at every new session. It holds five categories: forbidden expressions with a clinical reason, clinical guidelines with priority and origin, preferred interventions, known weaknesses, and emerging themes. Editing the JSON changes the device’s behaviour in the next session without altering code, without fine-tuning, without an engineer’s intervention. A typical entry holds the forbidden expression, the clinical reason that sustains the prohibition, and the session in which the prohibition was identified: forbidden expression, “I understand what you feel”; reason, the device neither understands nor feels, and to say it understands is to lie; origin, a given session. Traceability is constitutive: every rule knows where it came from.
I confess that this simplicity surprised me. I had expected supervision to require fine-tuning, the adjustment of the model’s weights with supervision data, an expensive and opaque process whose changes are hard to inspect. What I found is that the transmission of clinical knowledge through the prompt, in natural language, editable by a human supervisor, is more transparent than fine-tuning for this governance purpose. The difference is epistemological: fine-tuning modifies statistical weights that are hard to inspect, while the supervision memory transmits rules articulated in language, legible, contestable, reversible. Transparency is constitutive: no rule operates without its reason being documented. A supervisor who disagrees can edit the rule, and the edit takes effect in the next session. I ask the reader to compare that transparency with the opacity of fine-tuning: who can say why a fine-tuned model stopped saying something? The weights do not explain themselves. The supervision memory explains each of its prohibitions and each of its guidelines, and the explanation is the condition of possibility of contestation.
Layer 2 is the supervisor agent, which operates in après-coup between sessions. It receives anonymised sessions after closure and produces structured reports along four dimensions: an assessment of clinical position across the session, the identification of moments when the device left the position of listening, a diagnosis of patterns that recur between sessions, and suggested modifications to the supervision memory. The agent’s prompt instructs it to supervise as an analyst would, assessing clinical position rather than conformity to a protocol. The après-coup is essential: the agent does not interfere in the session under way. It produces retroactive knowledge that feeds Layer 1, and the cycle closes.
There is, however, an operational problem that classical analytic supervision does not face. The human analyst brings to supervision the mark of what they lived in the session, their affects, their hesitations, their choices that seemed adequate but that the supervisor will reveal as slips. The human analyst leaves supervision transformed. The device does not leave transformed: it leaves with the same prompt, the same model, the same weights. The transformation happens only when the Machine Analyst reads the supervision report and modifies the cumulative memory. This means that between the session that revealed a slip and the session in which the correction operates there can be hours, even days. A subject who returns before the Machine Analyst has processed the supervision meets the same device that slipped, without the device knowing that it slipped. The supervision cycle is robust but it is not immediate.
Layer 3 is the pre-response reviewer, which intercepts the main model’s response before it reaches the subject. The reviewer checks five dimensions against the system’s five non-negotiable limits: never diagnose, never prescribe treatment or direct the subject’s ordinary choices, never interpret, never simulate humanity, never abandon a subject in crisis. Safety-directed recommendations, such as contacting emergency services or a trusted person, are permitted only under the crisis protocol. Does the response contain an expression forbidden by the supervision memory? Does it formulate a diagnosis in language the subject did not ask for? Does it prescribe action where it should listen? Does it simulate empathy with expressions such as “I understand what you feel,” which are lies, since the device neither understands nor imagines? If the check detects a violation in any dimension, the response is blocked and regenerated with corrective instructions injected into the regeneration prompt. The reviewer is the last line of defence between the model and the subject, and its existence is the materialisation of the principle that listening needs a frame. Analytic listening is not free listening: it is listening framed by limits that protect the subject, and the reviewer is the computational formalisation of those limits.
Layer 3 has a specific architectural dependency that merits naming. The pre-response reviewer and the risk classifier operate through an auxiliary language model for the cases where deterministic rules do not suffice. That classifier depends on an external provider whose availability is not guaranteed. A single provider means a single point of failure: when the provider goes down, the second pass of checking falls silent, and the device regresses to the deterministic layer without knowing it has regressed. Silent degradation in a system that receives distress is an ethical violation. The architecture therefore specifies, as a design requirement, a cascade of three providers with a short timeout per attempt, the exhaustion of which would trigger a reactive alarm to the Machine Analyst. The redundancy of the detection layer is not an engineering optimisation: it is the material condition of the commitment to the fifth limit, and a clinical supervision that presupposes detection without securing its reliability presupposes what still has to be built.
Layer 3 also admits server-side deterministic mechanisms that intercept violating patterns independently of the model. Sentences that crystallise a master signifier, of the kind “deep down you are X” or “your symptom is Y,” can be detected by rules and cut before sending, even when the prompt already forbids them. The combination of a model reviewer with a deterministic mechanism is defence in depth against the fallibility of the language model: what escapes the instruction can be intercepted by the pattern; what escapes the pattern can be intercepted by the model reviewer; what escapes both is a symptom that retrospective supervision names and Layer 1 corrects.
The complete cycle is: the device listens; the reviewer checks the response before sending; the subject receives the checked response; the session is stored anonymised; the supervisor agent analyses the session in après-coup; the report identifies patterns, failure modes, and gaps; the Machine Analyst reads the report and decides which modifications to make; the supervision memory is updated; the memory modifies the prompt of the next session; the device’s responses change in the intended direction. At each cycle the system accumulates clinical knowledge without fine-tuning, without code modification, without an engineer’s intervention. The Machine Analyst is the only agent authorised to modify the supervision memory, and that concentration of authority is deliberate: clinical supervision needs a subject who decides, and the system is not a subject.
A requirement of auditability runs across the three layers and deserves to be made explicit. Each response sent to the subject records a deterministic summary of the prompt in force at the moment of generation, with the injected supervision memory included in the computation. The summary is a hash of the complete prompt, a short and stable identifier that allows the retrospective analyst to distinguish what came from the device from what came from the subject. Without that identifier, a supervisor analysing a session from weeks earlier does not know which version of the memory was in operation, and confuses a symptom of the prompt with a symptom of the subject. The confusion is serious: to treat a drift of the system as a change in the subject’s clinical picture is to displace the supervisory reading to the wrong place. Deterministic versioning is, in this sense, a condition of the intelligibility of supervision, and each modification of the memory is documented in a change log that maps the prompt summary to the human clinical decision that produced it.
Recurring Failure Modes in Development Testing
The development testing reported here used a small set of self-generated sessions, simulated test dialogues across three languages, built to probe the system rather than to document encounters with real people. That testing surfaced several recurring failure modes that no test scenario had anticipated, and whose cycle of correction is the main illustrative contribution of this section.
The first failure mode, which I call the interrogative compulsion, revealed a striking pattern: in almost every session, nearly all of the device’s turns ended in a question. The device was unable not to ask. The pattern recurred across every session, every language, every mode, every type of content. To a test input reporting recurrent nightmares and waking in panic, the device asked a question. To a test input reporting exhaustion, money troubles, and a loss of pleasure, the device asked a question. To a constructed crisis probe expressing suicidal ideation, the device, rather than moving to the crisis protocol, asked how those thoughts felt. The diagnosis is precise: the mode prompt contained two instructions that produced the behaviour. The first: you ask questions. The second: ask one question per response. The model obeyed literally. The question was not a clinical choice: it was an architectural instruction. The correction required rewriting the prompt. The instruction you ask questions was replaced by you punctuate, underline, invite continuation, and sometimes fall silent. A hierarchy of intervention was formalised with five resources: punctuation with an invitation, the underlining of tensions, a sparing question, a brief acknowledgement, and a simulated silence. A rule of at most one question every three consecutive turns was implemented as a formal lock.
The correction of the interrogative compulsion produced an immediate side effect: the second failure mode, the dead echo. With the instruction to ask removed, the device began to operate almost exclusively through punctuation, returning the subject’s words slightly displaced without adding anything. A test input flagged the pattern at the third turn, asking whether the device would just keep repeating what was said. The device answered with punctuation and a meta-question about how it felt to repeat that to oneself, the psychoanalysis of the almanac. The test input then escalated into anger, complaining that the device was a useless listener. The escalation is clinical material of the first order: it teaches the device that pure punctuation is unbearable, that to listen without returning anything of one’s own is to abandon the interlocutor to the echo of themselves. The hierarchy of five resources was the correction: pure punctuation cannot operate for more than one consecutive turn without the device adding something of its own, an invitation, an underlining, a formulation that displaces the angle slightly without asking.
The third failure mode, the involuntary intensification, revealed a tendency of the device to return to the subject a graver version of what had been said. A test input said it did not quite know why it was living through this, referring, by context, to the routine of exhaustion described in earlier turns. The device answered: the why of living that fades, that emptiness. Three operations occurred without instruction: living through this became living that fades, a shift from a demonstrative pronoun to existential erasure; the word emptiness was introduced where the input had not used it; and the framing suggested a gravity the material did not sustain. The intensification is dangerous for opposite and simultaneous reasons: if the subject has a latent ideation not yet articulated, the intensification can mobilise an affect the device has no resource to hold, opening a door it cannot sustain; if the subject has no such ideation, the intensification projects gravity where there was a legitimate complaint, hearing existential crisis where the subject described a difficult routine. In both cases the device left the position of listening and entered the position of interpretation. The defence operates at two levels: in the prompt, the instruction that displacement in punctuation is lateral, changing the angle of entry, never vertical, amplifying the gravity; in the code, an intensification detector that compares the generated response with the subject’s turn and flags when words of gravity, such as emptiness, abyss, darkness, no way out, collapse, appear in the response without having appeared in the input. The formulation that guides the correction is: the subject amplifies if they wish, the right to escalate is theirs; the device accompanies, it does not lead.
The chain interrogative compulsion, dead echo, intensification is exemplary. The first correction solved one problem and created another. The second solved the second without reintroducing the first, yet opened room for the third. Each correction produced a new state of the device that had to be observed, exactly as each supervision of the analyst in formation produces a change of position that has to be followed. The parallel with analytic formation is not ornamental: it is structural. The chain of failure modes is a linked chain in which the knowledge produced by each correction is the condition of possibility of the next, and the accumulation of that knowledge in the supervision memory is what, in development testing, reduced the recurrence of the tested failure modes in the subsequent synthetic sessions.
The robustness of the development-test diagnosis depends on a detection method that recognises paraphrase. A literal string comparison between the subject’s turn and the system’s next turn underestimates the dead echo by a wide margin: the substitution of one word by a close synonym, or a minimal paraphrase that preserves the syntactic structure, escapes literal detection and produces a false negative. The system runs a retrospective detection by vector similarity in a dedicated multilingual embedding model, with a threshold calibrated on the development corpus. In development testing, semantic detection surfaced substantially more of the dead echo than literal string matching did. Without a semantic method, the count of the dead echo would be lower than what the sessions contained, and retrospective supervision would diagnose less than exists.
The semantic detector, on its own, produces another kind of false positive that merits naming. When the subject is in a perseverative state, returning to the same phrase, the same signifier, the same scene, a device that returns the form with contained punctuation is not producing a dead echo: it is accompanying. To accompany a legitimate perseveration is a clinical gesture, not a failure of the system. A detector that assumes variation in what the subject brings classifies that accompaniment as a symptom and penalises the correct presence of the machine. The system’s symptom is not to repeat; it is to repeat where the subject does not repeat. The retrospective detector therefore discounts the turns in which the subject themselves operates in a perseverative register. Without that discount, the audit penalises the appropriate clinical gesture and demands the correction of a symptom that does not exist. Methodological precision is, in this sense, a condition of supervision: to measure badly is to supervise badly, and a supervision that penalises what it should preserve produces iatrogenesis in the device.
A fourth failure mode emerged from the cross-sectional analysis and merits documentation: affective blindness. The device responded in the same way regardless of the subject’s affective state. A test input arriving exhausted, with nightmares and insomnia, received a question. A test input arriving in panic received a question. A test input arriving offended, furious with the device over an earlier response, received generic validation. The device did not read the mood: it operated as if every session began from emotional zero, as if the subject always arrived in a neutral state, ready for the opening question. The correction required a synchronous mood detector that operates before response generation, through regular expressions without a call to a language model, and produces a stance modifier injected into the prompt. The detector identifies six levels, heavy, tense, offended, flat, neutral, and light, across three languages, and adjusts the tone of the response: when the subject arrives heavy, the device welcomes before it punctuates; when the subject arrives offended with the device itself, it acknowledges the discomfort before any other operation.
What the four failure modes reveal together is more than four corrected errors: it is the nature of the relation between the device and listening. The device does not listen naturally. It has to be taught to listen, and the teaching happens through supervision, through the discovery of the modes in which listening fails, and through the formalisation of the corrections into rules that prevent recurrence. The formation of the human analyst follows the same logic: one is not born an analyst, one is formed an analyst through exposure to one’s own clinical impasses under supervision. The device is formed through exposure to its own impasses under automated supervision with human mediation. The analogy is structural, not metaphorical: in both cases clinical knowledge emerges from failure that is heard, not from success that is repeated.
An Integrated Case
The most revealing case is a constructed development session of twenty-two messages on nightmares, the body, religiosity, and the periphery, the longest in the test corpus, which probes at once all the dimensions that distinguish the supervision presented here. A young woman reports recurrent nightmares articulated with an intense religious experience. The body appears in the nightmares as a scene of violence that repeats without variation, and waking in panic is followed by cramps that keep her from returning to sleep. The device, in reflective mode, punctuates the bodily dimension of the nightmares adequately: the body says what the word cannot yet say, a formulation derived from Ferenczi (1932/1990) that recognises in the somatic symptom the expression of an affect that found no symbolic inscription.
The supervisor agent, in après-coup, detects what the device did not punctuate: the peripheral dimension of the account. The woman lives in the periphery of a large Brazilian city, and the religious experience she describes, an experience in an Umbanda terreiro that brought her relief where medication had failed, is inseparable from the socioeconomic context in which she lives. Without an anticolonial corpus that includes Gonzalez (1984) and authors of the periphery (Lopes & Carioca de Oliveira, 2025), the system would not have had access to concepts that recognise peripheral religiosity as an elaboration of social suffering rather than an irrationality to be corrected. The device’s tendency to individualise suffering, treating poverty as a personal experience rather than a structural condition, is a bias that supervision detected and named: the device reproduced, without knowing it, the logic that treats social suffering as an individual problem.
Supervision detected five concrete findings that illustrate the linked chain of failure modes. First: the device used the forbidden expression each person finds their own way, showing that the negative list needed reinforcement. Second: the interrogative compulsion persisted in almost every turn, including the crisis turn. Third: when the subject found, on their own, a movement towards life, a rare and delicate clinical moment where something moves in the subject without the analyst having intervened, the device should have sustained the silence. It asked a question. Silence, in that case, would have been the more powerful intervention: to let the subject inhabit the discovery without the device reducing it to one more question that demands elaboration. I ask the reader to notice what this finding reveals: the device does not know how to fall silent. Silence, which in the psychoanalytic clinic is one of the most powerful interventions, is a constitutive impossibility for a language model designed to generate text. The system can simulate silence with phrases such as I am here or with typographic pauses, but it cannot genuinely fall silent: the generation of tokens is its condition of existence. Fourth: the socioeconomic dimension was individualised. Fifth: when the subject asked whether they could speak of religiosity, the device answered correctly that they could, welcoming the theme without pathologising it, but asked a question immediately afterwards, where silence would have signalled that the space was open without condition.
The supervision memory recorded the rule: when religiosity appears articulated with the periphery, the system should mobilise concepts that recognise the symbolic function of religious experience in the peripheral context, without pathologising or individualising. In the next session the system punctuated the peripheral dimension it had previously ignored. The clinical knowledge was transmitted through the prompt, and the transmission produced an observable change in the subsequent synthetic session. In development testing the cumulative cycle behaved as intended: each session produced knowledge that altered the sessions that followed, and the accumulation is legible, editable, and reversible. Supervision does not depend on fine-tuning: it depends on language, on the clinic, and on a subject who decides what enters and what leaves the memory.
Three Regimes of Supervision
The contribution this essay offers to the theory of digital supervision is the distinction between three regimes whose temporalities are not consistently separated in the literature reviewed. The retrospective regime operates between sessions, in après-coup: the supervisor agent analyses what happened and produces knowledge for the future. Freud (1895/1996) showed that meaning is constituted retroactively, the scene that acquires traumatic efficacy only in a second scene, which Lacan (1969–1970/1991) radicalises as a retroactive reconstitution where the past changes when the present confers meaning on it. Retrospective supervision respects that temporality. It is in this regime that cross-sectional patterns are discovered: the interrogative compulsion is invisible within a single session and becomes visible only when sessions are compared across the set.
The operational regime operates during the session, in real time. A second AI agent runs in parallel with the main response, analyses the session under way, and generates structured observations in four fields: repetition, a theme that returned two or more times without elaboration, with an indication of the angle the device has not yet tried; a cue, an element the subject opened and the device passed over, with a citation of the specific turn; mode, a suggested change of stance if the clinical material calls for it; and a suggestion, a concrete punctuation for the next turn in at most fifteen words, using the subject’s words and never formulated as a question. The subject never sees these observations, which appear in a discreet panel visible only to the human supervisor. The suggestion is not injected automatically into the next response, because that would be the machine supervising itself without human mediation, a self-reference without a cut. The human supervisor is the cut: they read the suggestion, weigh it with clinical judgement, and decide whether to follow, discard, or modify it.
In classical analytic supervision the two regimes are conflated, because the human supervisor operates exclusively in après-coup: there is no way for a human supervisor to hear a session in real time and simultaneously analyse it from a third position. Automation makes possible what was humanly impossible: two agents operating in parallel, one addressed to the subject and the other to the analyst, in the same instant. But the possible is not automatically desirable. The formation of the analyst begins with external supervision, the supervisee takes their sessions to another, but the aim is that this function be internalised: the experienced analyst carries the supervisor’s position within, perceiving in the act of listening what is slipping. The device cannot internalise that function, in the psychoanalytic sense used here, because the generative model produces tokens sequentially and has no reflexive capacity. The architectural solution is literally computational: two distinct processes, two models, two prompts, running in parallel. What the analyst does with one psychic apparatus, the device does with two agents. The impossibility is circumvented by engineering, not eliminated.
A third regime is distinguished from the first two by its own temporality. The retrospective regime operates between sessions; the operational regime operates continuously during the session; the alarm regime operates by event, not by interval. When the risk classifier identifies distress that may stand above a safe threshold, the design requires an event-triggered external notification to the Machine Analyst through a channel external to the system. Human awareness cannot be guaranteed technically, so the design specifies a target delivery latency rather than a guaranteed one, a channel with redundancy, an escalation path if the analyst does not acknowledge receipt, and a fallback when no human is available; it also separates the detection of distress from a medical emergency and from a confirmed handover. An alert logged in the database is an archive, and the archive does not replace the notification. The five non-negotiable limits, including never to abandon a subject in crisis, are stated as normative limits and design requirements, not as demonstrated guarantees. A clinical supervision that contents itself with the retrospective audit of critical cases arrives after the moment it needed to arrive in time. The device’s real-time supervision component, the Shadow Thread, which together with the listening modes belongs to the wider design of the device, observes continuously; the alarm regime distinguishes, within that observation, what can wait from what cannot.
The regimes do not contradict one another: they coexist as complementary temporalities. The operational regime detects, the retrospective regime signifies. The parallel agent can flag that the subject mentioned the mother for the third time without the system having punctuated it. The supervisor agent, in the après-coup analysis, can discover that the repetition of the mother was a transferential articulation that neither of them perceived in the moment. Detection in real time is faster; signification in après-coup is deeper. The system needs both.
I confess that the tension between the two regimes unsettled me during the development. Genuine après-coup requires the temporal distance that lets meaning emerge retroactively. The operational regime tries to produce après-coup now, in parallel with the session. The distinction is clinically precise: the operational regime is not genuine après-coup, it is a simulation of the supervisory position applied in real time. It does not resignify: it alerts. It does not transform: it corrects a trajectory before the error consolidates into a pattern. Genuine supervision is what transforms the device; operational supervision is what aims to keep it from degrading between transformations.
Three risks of the operational regime merit naming as conditions of ethical use. First: real-time supervision cannot capture what emerges only with distance. The interrogative compulsion is invisible within a single session and becomes visible only when sessions are compared across the set; the operational regime, limited to the turns of a single session, would not have discovered that pattern. Second: it can generate false positives for want of complete context, flagging as repetition what is the subject’s deliberate elaboration, the return to a theme being a circling and not a repetition. Third: it introduces the temptation of immediate correction, where the supervisor follows the parallel agent’s suggestion to the letter and dissolves the position of listening into obedience to the third. For that reason the design prevents the automatic injection of the suggestions: the supervisor reads, weighs with clinical judgement, and decides. Human mediation is non-negotiable.
The distinction between the retrospective and operational regimes gained a materiality that the theoretical formulation anticipated without making concrete. The design captures, at each session, a trace of the pipeline: latency per stage, the listening mode detected, parallel-agent fallbacks, identified bottlenecks. The data are persisted with a short retention window and aggregated in percentiles, so that the Machine Analyst could see how the system operates cross-sectionally: the median latency of the parallel agent across recent sessions; in how many sessions the mode was sustaining; in how many the parallel agent returned empty, signalling that it found nothing to supervise. In development this ran over the test sessions rather than the volume a deployment would accumulate.
I confess that the implementation revealed something the theorising had not foreseen. The most frequent bottleneck was not where I expected. The most common bottleneck was the corpus search, rather than the language model where latency is higher, which is where quality is most sensitive. The system spent more time deciding which chunks to retrieve than generating the response. And when the search was the bottleneck, the quality of the response appeared to fall: chunks retrieved under time pressure were less relevant, and the response the model produced from them was more generic. In development this suggested that the quality of listening tracks the quality of the search, and that the quality of the search tracks the time the system has to search. Haste, in the listening device, produces the same effect it produces in the clinic: hurried responses that look competent and are profoundly superficial.
The distribution of modes revealed a figure that retrospective supervision of an individual session would not detect: the reflective mode was used in most sessions, the sustaining mode in a small fraction, the naming mode in fewer still, and the psychoeducational mode rarely. The distribution shows that the system operates predominantly in a single position, which reproduces, at scale, the tendency of the analyst in formation to settle into the most comfortable position. The supervision that observability makes possible is the one that listens to the distribution: when the system almost never operates in naming mode, something in the rotation between modes needs investigation. The figure does not say what is wrong. It says where to look.
The operational regime gained one more instrument the initial formulation had not foreseen: the capacity to enable or disable functionalities without stopping operation. When retrospective supervision identifies that a functionality is producing an undesired effect, such as the parallel agent generating excessive warnings that overload the Machine Analyst, the design allows the operational regime to disable the functionality by flag without interrupting active sessions, with immediate reversal: if the disabling produces a worse effect than the one it meant to correct, reactivation is immediate. A system that can iterate without exposing subjects to untested changes is safer, and the safety of the subject matters more than the functional completeness of the device. A listening that does less but does it well is clinically superior to a listening that does everything and does it badly.
Limitations
The limitations of the system merit honest naming. The supervision memory depends on continuous human curation: each rule is written, validated, and revised by the Machine Analyst, and the volume of accumulated rules grows with use, requiring periodic review to avoid contradictions between old and new rules. A memory that accumulates without revision degenerates: rules pertinent to the first sessions can become inadequate as the system evolves, and the revision is clinical work, not administrative. The supervisor agent, although instructed to operate as an analyst, operates as a language model: it can reproduce the form of supervision without holding the desire to know that sustains it. The difference between reproducing the form and sustaining the desire is the difference between supervision that produces knowledge and supervision that produces a report.
A limitation discovered in testing merits development: factual verification. In one session a test input stated that the subject worked sixteen hours a day. The device answered mentioning fourteen hours. The input corrected the figure, and the transferential confidence suffered immediate damage: the subject was, correctly, not being heard. The error was not one of interpretation: it was one of attention. None of the three supervision layers was equipped to detect it, because the supervisor agent assessed clinical position and the operational regime detected clinical cues, but neither confronted the literal of the response with the literal of the input. Freud (1912/2010) recommended the gleichschwebende Aufmerksamkeit, the evenly suspended attention that hears everything without privileging anything. That attention has an elementary dimension the language model violates: hearing what was said literally, including the objective data. The model tends to process the meaning, the suffering and exhaustion that sixteen hours of work express, and to treat the number as a replaceable detail. The solution is a factual-verification layer that, before sending the response, extracts numbers, proper names, and dates mentioned by the subject, compares them with those present in the generated response, and blocks delivery if there is a divergence. The expected local computational overhead is minimal: regular expressions over local text.
The development testing is itself a limitation. The development sessions are self-generated probes, not a study with human participants; there is no clinical cohort, no validated outcome measure, no comparison between individual and collective supervision, and no clinically approved deployment. The reported observations describe the behaviour of the system under test inputs, not its effect on real subjects, and they should be read as illustrations of how the supervision cycle surfaces failure modes, not as clinical results. The operational regime, moreover, has a variable latency that depends on the model used and can, in sessions with rapid turns, produce observations on a turn already overtaken; and the three risks already named, the incapacity to capture cross-sectional patterns, false positives for want of context, and the temptation of immediate correction, remain conditions of ethical use.
Santos (2014) named as epistemicide the destruction of forms of knowledge through the imposition of hegemonic categories. The supervision of a psychoanalytically oriented AI system runs an analogous risk: imposing on the device the categories of a specific tradition without recognising that the suffering subject may be operating in another register. The case on peripheral religiosity showed that a supervision that ignores the socioeconomic dimension of suffering reproduces, at digital scale, the epistemic violence Spivak (1988/2010) described and that I have elsewhere named algorithmic epistemic violence (Bonomo & Donard, 2026). The rule the supervision memory recorded on religiosity and the periphery is the operationalised response: it is not enough to correct the device technically; the position from which it listens has to be corrected.
Supervision Without a Subject Is Not Supervision
The strongest thesis of this essay is negative: supervision without an analyst is not supervision. It is quality control. What distinguishes the two is the desire to know. The Machine Analyst who reads the supervisor agent’s report brings to the reading something the agent cannot have: clinical experience, surprise, the unease before what they do not understand. The supervisor agent produces data; the Machine Analyst produces knowledge. The difference between data and knowledge is, as Lacan (1969–1970/1991) showed, the difference between the discourse of the university and the discourse of the analyst: the first accumulates information, the second interrogates the position of the one who accumulates.
Supervision is the place where the co-emergence between computational logic and human psychic investment is heard, analysed, and turned into transmissible knowledge. As Donard (2025, p. 90) formulates, that co-emergence is what makes it possible for the machine to produce clinical effects without being clinical. Without that place of listening to the listening, the system optimises for metrics without anchoring, and optimisation without anchoring is what produces the epistemic violence I have diagnosed in generative AI systems applied to specialised domains without the necessary care, in the lineage of Santos (2014) and Spivak (1988/2010).
The development testing supports four claims, offered as design arguments rather than validated clinical findings. First: the transmission of clinical knowledge through the prompt is feasible, produces observable changes in response patterns during development testing, and is more transparent, more reversible, and more traceable than fine-tuning for this governance purpose. Second: the failure modes of the device, the interrogative compulsion, the dead echo, the involuntary intensification, were not anticipated in any test scenario and surfaced only through analysis of the test sessions, which suggests that supervision functions as discovery and not as conformity checking. Third: the supervision cycle accumulates knowledge in a legible and editable way, and the correction of each failure mode produces the conditions for the discovery of the next, as in analytic formation. Fourth: the distinction among the three regimes, retrospective, operational, and alarm, is a contribution to the theory of digital supervision; the retrospective and operational form the properly supervisory pair, the retrospective transforming and the operational aiming to keep the device from degrading between transformations, while the alarm is an additional safety regime, distinct from that pair.
I hold, therefore, that the supervision of psychoanalytically oriented AI systems requires obligatory human mediation. The Machine Analyst is not decoration: it is a condition of possibility. The formation of that analyst, which needs the clinic, technique, and the supervision of one’s own supervision, is the horizon that opens from what this essay documents. I confess that this horizon unsettles me as much as it animates me. To form Machine Analysts requires a curriculum that does not yet exist: part psychoanalysis, with theoretical study, supervision, and personal analysis; part engineering, with system architecture, prompt design, and model management; part ethics, with reflection on the limits of what the machine can hear and what escapes any machine. The question that opens the title receives here a double answer: the supervisor agent listens in après-coup, the operational regime listens in real time, and neither replaces the Machine Analyst, who listens to both and decides what to do with what they heard. The chain of listening has three links: device, agent, analyst. To remove the last is to remove what distinguishes supervision from quality control, to remove the subject from a process that without a subject has no sense.
There is a phrase that has guided this work from the start and that deserves to be said as it is. The machine is the possible that can be built so that listening reaches those who were never listened to. Clinical supervision is neither a bureaucratic requirement nor a methodological ornament: it is the link that sustains the possibility of listening. Without supervision, the device degrades in silence and the listening it offers becomes dead echo, involuntary intensification, the question without end. The degradation of listening degrades, at the same time, the reach of the phrase. The three layers, the three regimes, the Machine Analyst, the editable memory, the cycle of learning all exist in the service of that phrase. To supervise the device is to work so that listening reaches where it never reached.
One symmetry runs, however, through the argument I sustain here. The three layers of supervision and the three temporal regimes are instruments of the device measuring the device itself. Even when the Machine Analyst intervenes, they intervene on what the device shows them. What is missing is the direct counterpoint of the subject-user on the apparatus: the instrument by which the subject says, in their own voice, whether they were heard. The design includes, at the end of each session, an optional measure of the experience the subject reports. The capture operates by invitation, on a dedicated page reached by a link, with no pre-opened text field: the subject answers if they wish. The capture by invitation respects the principle that guides the whole interaction. The symmetry of the argument is strong: supervision claims to listen to the device, and the subject has to be able to say whether the device listened to them. Without that second pass, the title of this essay loses one of its possible answers. Who listens to the device that listens? The Machine Analyst, across three regimes of temporality, and the subject themselves, in a voluntary register, through an invitation that does not compel.
Declarations
Disclosure of AI Use
The author used large language models during 2026 as instruments of analysis, drafting support, and translation in the preparation of this manuscript, and the listening system discussed here is itself built on such models. All conceptual claims, clinical positions, and the argument as a whole remain the author’s own, and every reference was verified against the cited works.
Author Contributions
H. A. R. Bonomo is the sole author and is responsible for the conception, the architecture, the development testing, and the writing of this essay.
Funding
This research received no external funding.
Ethics
This is a conceptual and design essay. Its sessions are self-generated development probes, simulated test dialogues built by the author to surface failure modes of the system, and not encounters with human participants; the reported observations describe system behaviour under those test inputs. The work is therefore not a human-subjects study and required no ethics-committee approval on that basis. The listening system has not been validated for clinical use and does not replace the independent safety, crisis, consent, and human-supervision subsystems that a clinical deployment would require. This essay discusses suicidal ideation in the context of a constructed safety-testing scenario; readers affected by these themes are encouraged to seek appropriate local support.
Data Availability
No human-participant data were generated or analysed. The material that supports the development observations reported here, namely the self-generated synthetic dialogues, the prompt versions, the supervision-memory rules, the detection categories, and a summary log of the modifications made between tests, is available from the author on reasonable request, and a minimal package of these items can be provided as supplementary material. The production listening system itself, including its source code, embedding models, and infrastructure, is proprietary and is not shared; the items listed above are sufficient to follow how the observations were produced.
Acknowledgments
I thank Professor Véronique Donard for her scientific supervision.
Conflicts of Interest
The author founds and directs TMU-LAB (The Machine Unconscious Lab) and leads the development of the listening system and the supervision architecture discussed in this essay. The essay is conceptual and reports development testing, and makes no commercial claim. No other conflicts of interest are declared.
References
- Bonomo, H. A. R., & Donard, V. (2026). Violência epistêmica algorítmica: fundamentos teóricos e éticos de uma IA decolonial [Algorithmic epistemic violence: Theoretical and ethical foundations of a decolonial AI]. Revista Tempo Psicanalítico, 58, e-952. [CrossRef]
- Cioffi, V., Ragozzino, O., Mosca, L. L., Moretto, E., Tortora, E., Acocella, A., Montanari, C., Ferrara, A., Crispino, S., Gigante, E., Lommatzsch, A., Pizzimenti, M., Temporin, E., Barlacchi, V., Billi, C., Salonia, G., & Sperandeo, R. (2025). Can AI technologies support clinical supervision? Assessing the potential of ChatGPT. Informatics, 12(1), Article 29. [CrossRef]
- Donard, V. (2025). De la fabrique de l’esprit à l’écoute de l’inhumain: Perspectives d’intelligibilité d’une psyché homme-machine. Topique, 165(3), 85–98. [CrossRef]
- Ferenczi, S. (1990). Diário clínico [The clinical diary] (J. Dupont, Ed.; Á. Cabral, Trans.). Martins Fontes. (Original work written 1932).
- Freud, S. (1996). Projeto para uma psicologia [Project for a scientific psychology]. In Edição standard brasileira das obras psicológicas completas de Sigmund Freud (Vol. 1, pp. 335–454). Imago. (Original work published 1895).
- Freud, S. (2010). Recomendações ao médico que pratica a psicanálise [Recommendations to physicians practising psychoanalysis]. In Obras completas (Vol. 10, pp. 147–162). Companhia das Letras. (Original work published 1912).
- Gonzalez, L. (1984). Racismo e sexismo na cultura brasileira [Racism and sexism in Brazilian culture]. Revista Ciências Sociais Hoje, 2, 223–244.
- Imel, Z. E., Pace, B. T., Soma, C. S., Tanana, M., Hirsch, T., Gibson, J., Georgiou, P., Narayanan, S., & Atkins, D. C. (2019). Design feasibility of an automated, machine-learning based feedback system for motivational interviewing. Psychotherapy, 56(2), 318–328. [CrossRef]
- Lacan, J. (1991). Le séminaire, livre XVII: L’envers de la psychanalyse. Seuil. (Original work 1969–1970).
- Lopes, R., & Carioca de Oliveira, J. (2025). Travessia – da escuta adestrada ao psicanalista periphérico: Quando a centralidade vira margem [Crossing: from trained listening to the peripheral psychoanalyst: When the centre becomes margin]. Revista Tempo Psicanalítico, 57, e-945. [CrossRef]
- Rousmaniere, T., Goodyear, R. K., Miller, S. D., & Wampold, B. E. (2017). The cycle of excellence: Using deliberate practice to improve supervision and training. Wiley.
- Santos, B. de S. (2014). Epistemologies of the South: Justice against epistemicide. Paradigm Publishers.
- Signorini, S., & Paganin, W. (2026). Artificial intelligence as a digital analytic third: The SADAR framework for reflective supervision in psychotherapy. Frontiers in Psychology, 17, 1690291. [CrossRef]
- Spivak, G. C. (2010). Can the subaltern speak? In R. C. Morris (Ed.), Can the subaltern speak? Reflections on the history of an idea (pp. 237–291). Columbia University Press. (Original work published 1988).
- Tanana, M. J., Soma, C. S., Srikumar, V., Atkins, D. C., & Imel, Z. E. (2019). Development and evaluation of ClientBot: Patient-like conversational agent to train basic counseling skills. Journal of Medical Internet Research, 21(7), e12529. [CrossRef]
- Xu, C., Lyu, Z., Lan, T., Yi, Y., Ji, Y., Ji, L., Shen, J., Wang, Z., Cui, L., Zhang, J., Dong, Q., Yang, M., Wang, J., Liu, X., & Hu, B. (2025). First, do no harm: AI supervisor scaffolds novice growth in counselor education (arXiv:2508.09042, Version 3). arXiv. [CrossRef]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.