Preprint
Article

This version is not peer-reviewed.

When AI Evaluates Design More Responsibly Than Architects: Consciousness and Feedback Loops

Submitted:

28 July 2026

Posted:

30 July 2026

You are already at the latest version

Abstract
Buildings and urban spaces affect a user’s body — through attention, emotion, orienta-tion, recovery, social behavior, and stress regulation — yet architectural evaluation often lacks feedback loops. This paper develops a theoretical framework for evi-dence-responsive design evaluation. AI is not used here to generate architectural images. Operational design consciousness is the capacity of registering the effects of built form on sentient users, transmitting that information into judgment, distinguishing evidence from narrative, and being able to revise later action. Under this criterion, dominant stu-dio-conditioned architectural culture can behave less consciously than an empirically constrained AI system. This conclusion follows, not because AI is subjectively conscious, but because it can be forced to apply evidence that professional discourse often discounts. The paper proposes an architectural analogue of Asimov’s First Law of Robotics: an ar-chitect may not design a building or urban space that imposes cognitive, emotional, physiological, or social harm on users, nor through inaction allow such harm to persist. Christopher Alexander’s pattern language is interpreted as an architectural memory system of empathic design based on accumulated embodied feedback. The inverse Tu-ring test proposed here distinguishes prestige-script design discourse from evi-dence-responsive reasoning. An LLM—reading only the statistical regularities of the profession’s public discourse—can reproduce its defense-scripts with recognizable fidel-ity. The result is a practical framework for AI-assisted design diagnostics, harm-sensitive architectural judgment, and post-occupancy correction.
Keywords: 
;  ;  ;  ;  ;  ;  ;  ;  

1. Introduction

1.1. AI As Evaluator Rather Than Form Generator

Buildings and urban spaces are not neutral visual objects. They act on attention, emotion, orientation, recovery, social behavior, and stress regulation. Yet architectural evaluation still relies heavily on abstract criteria while discounting feedback. This paper proposes an operational test for whether a design-evaluating system registers evidence of harm and revises its behavior accordingly. Under that criterion, an empirically constrained AI system may behave more consciously than a professional culture trained to discount embodied human response. AI can be made to collect feedback from biometric response, environmental-psychology findings, post-occupancy evidence, and public distress.
AI is commonly used in architecture as a means of generating candidate images or taking care of technical design complexities (Peng et al. 2026). This paper does not examine those applications. Readers will assume that AI → generative novelty, whereas this paper uses AI → evaluative diagnostics. A form-generating AI is free to create any geometry that is visually striking — it is the human who evaluates those images based on visual novelty. The system considered here is itself the evaluator: it asks for evidence from human psychophysiology to support a design claim. It also asks what should be revised if the possibility of harm is predicted.
An empirically constrained AI therefore occupies the role of an external evaluator rather than another design-production tool. The paper proposes an epistemic and regulatory application of AI. The AI assembles relevant empirical evidence, identifies conflicts between design claims and human outcomes, and checks if those conflicts are resolved before approval. Its closest analogues are clinical audit and post-occupancy evaluation—not automated composition. AI collects and evaluates empirical data — its deeper value lies in evidence-responsive evaluation.
In a previous paper, Peña and Salingaros (2026) argue that dominant architectural culture is not an adaptive learning system. The profession continues to be driven by peer prestige and the verbal justification of images. Failure to register human effects is a deficit of operational consciousness in the design domain. Consciousness criteria are seldom applied to human behavior shaped by specialized training.
The “hard problem” of consciousness is to explain how and why physical processes in the brain give rise to subjective experience (Chalmers 1995) — a topic we do not tackle here. The framing of “prosthetic consciousness” in Section 7.1 proves useful in practice. Empirically constrained AI offers a support tool for architects who already possess empathy but find that current institutional inhibitions prevent them from acting on it. Approaching consciousness indirectly through AI creates a framework for evaluating design.

1.2. AI Checks for Harm During Design Review

Can an artificial system, when constrained by evidence about human response, behave more like a conscious design agent than a professional culture that has learned to suppress such evidence? Human beings possess biological consciousness, but can be trained not to use the relevant feedback channels in a particular domain. That creates a significant gap between architectural discourse and human outcomes.
The results presented here have five practical implications:
1. Buildings can be evaluated as health-affecting environments rather than as autonomous visual objects;
2. AI can be used as a practical external evidence-check during design review;
3. pattern language and living geometry provide architectural priors for such evaluation;
4. post-occupancy evaluation should become a normal corrective loop, not an optional afterthought;
5. the model gives clients, educators, journal reviewers, and procurement officers a way to distinguish evidence-responsive design claims from prestige-script rhetoric.
These principles are operationalized in Section 7, which maps the evidence channels required for harm-sensitive evaluation onto specific decision-makers and measurement methods. The proposed audit system extends the architect’s capacity to discover future effects. AI diagnosis demonstrates due diligence to clients and public authorities. It distinguishes evidence from unsupported claims, and helps a conscientious practitioner resist pressure (from whatever source) for potentially harmful design decisions. AI can function as an external scientific review comparable to clinical audit, environmental assessment, post-occupancy evaluation, or structural checking.
No behavioral trial was performed in this research. The discussion is purely theoretical. It frames a logical comparison under an operational definition, not a clinical validation or psychometric study.
Individual architects are articulate and subjectively conscious — but do they always apply their consciousness while designing? Does the professional system demand (or even allow) embodied and empathic information to enter design judgment and revise behavior? (Mediastika 2016) We use the indicator-property approach recently proposed for assessing consciousness in artificial systems (Butlin et al. 2023). That framework identifies functional indicators associated with theories of consciousness, including agency, embodiment, global workspace, higher-order monitoring, and recurrent processing. We apply the same logic to evaluate architectural design culture.
The innovative step here is to apply these criteria meant for detecting consciousness in artificial systems to both AI and humans. In the case of the dominant system for evaluating designs, the empirically constrained LLM can outperform the professional script.
We do not claim that LLMs possess phenomenal consciousness. Nor do we claim that architects lack consciousness as persons. We suggest that dominant studio-conditioned architectural behavior suppresses the indicators that force design judgment to respond to human harm. Conversely, an LLM constrained by empirical evidence and human needs can simulate those indicators well enough to apply empathic criteria. This produces an inverse Turing test for design consciousness: the human designer may speak fluently while failing to update from evidence, whereas the AI may lack subjective experience yet behave more like a feedback-sensitive design agent.
The logical chain is as follows. First, built form produces measurable effects on the body. Second, conscious design agency requires those effects to enter a judgment cycle and to modify later action. Third, architectural education can train designers to disqualify precisely those bodily and empathic signals that would normally trigger correction. Fourth, where the professional role has become automaton-like in this restricted sense, a hard external constraint—analogous to Asimov’s First Law of Robotics—is needed to force feedback into the design process.
Inverting the AI-consciousness question addresses two audiences. For readers concerned with architecture and design, the argument offers a new explanation for the persistent gap between professional taste and public response (Boys Smith and Salingaros 2025). For researchers studying consciousness in artificial systems (Chalmers 2023; Kang et al. 2026), architecture supplies a test case in which a human professional system may suppress the same feedback functions that are used as indicators of consciousness in AI.
Architecture can behave less consciously because its institutional structure devalues error detection and correction. This conclusion is close to Argyris and Schön’s organizational-learning framework: an organization learns only when error detection changes its “theory-in-use” (Robinson 2001; Al Maani and Roberts, 2023). The profession’s declared aim is design excellence, human flourishing, public welfare, and sustainability; however, its “theory-in-use” closes on the image and excludes occupant outcomes (Anthony, 2012; Curl, 2018).

1.3. How This Paper Is Organized

The paper is organized as follows. Section 2 defines operational consciousness in the design domain. Section 3 formulates an architectural analogue of Isaac Asimov’s First Law and explains why an external fail-safe is needed. Section 4 develops the inverse Turing test. Section 5 introduces Christopher Alexander’s pattern language as an architectural memory system based on embodied feedback. Section 6 explains how studio culture can suppress bodily and social evidence. Section 7 describes empirically constrained AI as an external empathic monitor. Section 8 states the limitations of the present theoretical argument and outlines a future validation protocol. Section 9 summarizes the comparative claim.
Appendix A presents an LLM-assisted reconstruction of professional architectural discourse. It tests whether an LLM, drawing from open professional sources, can generate a recognizable prestige-script response of the kind that human architects often give when challenged by empirical evidence. The same prompts are then answered very differently under evidence-responsive constraints. The model produces a simulated precursor to a future human survey: a methodological proof-of-concept shows how to model and score professional discourse.
Appendix B gives the theoretical background for comparing biological, institutional, and artificial design-evaluating systems at a functional level. It uses David Marr’s three-level analysis to separate the computational problem of harm-sensitive design from its biological or digital implementation. Comparing the operational feedback structure of design judgment with subjective experience implies that implementation (i.e., the physical components the system is made of) is not the most relevant factor for consciousness.

2. Operational Consciousness and the Action of Design

2.1. Cognition, Consciousness, and Intelligence

Cognition is the machinery of mental processing, combining abstraction, language, memory, pattern recognition, perception, and reasoning. A trained architect who manipulates complex geometry, prepares technical documents, and defends a design verbally is cognitively functional. Intelligence is adaptive problem solving: the ability to learn from feedback and change behavior when outcomes show that an objective has not been met. Operational consciousness integrates action with correction. It registers what one is doing, what that action causes in oneself and others, and whether the resulting information should alter subsequent action.
This paper compares two distinct design-evaluating systems and their feedback architectures. In David Marr’s terms, the relevant comparison is not between biological and digital implementation, but between the computational problem each system is organized to solve: (i) harm-sensitive adaptive design, versus (ii) prestige-sensitive image production. The second system represents a professional design culture that blocks evidence of psychological harm from entering judgment. An LLM reinforced with scaffolding explicitly organized to retrieve and apply such evidence can function as the first system.
We employ fifteen indicator properties in the classification of Butlin et al. (2023) to evaluate architectural consciousness. Those qualities are coded according to which scientific theory they come from: recurrent processing theory, global workspace theory, computational higher-order theories, and others. Their source is not of direct relevance here. Table 1 lists selected indicator properties (Butlin et al. 2023) mapped onto architectural evaluation and scaffolded AI diagnostics.
An LLM does not satisfy sufficient indicators of consciousness, as Butlin et al. (2023) have already emphasized. The present paper applies selected indicators critical for design at Marr’s computational and algorithmic levels: these are the higher-level ones in the Butlin et al. classification: global broadcast of relevant information to the system (labeled as GWT-3), metacognitive monitoring that distinguishes reliable evidence from noise (HOT-2), belief formation and action selection that update in response to monitoring (HOT-3), flexible goal pursuit based on feedback (AE-1), and modeling of output-input links between action and environment (AE-2).
A professional system can retain technical cognition while suppressing these design-relevant indicators. Table 1 contains the relevant message in clear and direct language that applies to the process of design. It will become increasingly clear why it is necessary to introduce results from outside architecture — paradoxically, findings about the foundations of the design discipline unfamiliar to practitioners and theorists — to support the present thesis.

2.2. Operationally Conscious Design

A design system is operationally conscious when it transmits the effects of its built outputs to the evaluative modules that will shape future design. It must monitor those effects, distinguish evidence of psychological harm from institutional noise, update design beliefs and adjust action in response to that monitoring, learn from feedback while pursuing human well-being, and model the effects of design outputs on occupants and passers-by. Conscious design needs to review negative psychological consequences.
Recent work on embodied intelligence provides an independent engineering formulation of this feedback structure. Zhang et al. (2025) contrast disembodied systems—whose reasoning depends on predetermined rules and static representations—with embodied intelligence that arises through continuous interaction and environmental response. A closed loop transmits feedback information into a predictive model, activates that model, and modifies later action when new feedback arrives.
The same three-part sequence applies to operational design consciousness. Progressive design policy diagnoses spatial configurations and implements revisions when observed outcomes contradict initial predictions. A professional system becomes functionally disembodied when it retains representation while disconnecting it from the feedback of affected human bodies. Under this interpretation, dominant studio culture interrupts the perception–modeling–action loop required for embodied intelligence (Zhang et al., 2025).
A system is operationally unconscious when it continues to act while the above indicators (Table 1) are disabled or never implemented. An intelligent but unconscious optimizer may maximize elegant forms, iconic visibility, media-based market value, or novelty while ignoring human discomfort. It can use professional language to justify those goals. An internally coherent but untested model of architectural causation substitutes for the actual feedback of embodied users affected by it.

2.3. Awareness, Interoception, and Empathic Correction

Consciousness includes wakefulness, verbal fluency, and the ability to describe one’s own thoughts. Active awareness of one’s body and surroundings empowers the ability to use those signals in judgment. An internal stream — attraction, bodily feeling (interoception), emotional appraisal, muscular tension, orientation, repulsion, and unease — integrates with an external stream — the perceived world and the signals emitted by other persons. These constitute part of the body’s monitoring system for action in the world (Seth et al. 2012; Gu et al. 2012, 2013; Price and Hooven 2018). Interoception comprises the feelings we experience inside our own body (Craig, 2002).
A studio-conditioned architect can stand in front of a blank clean wall, a bleak corridor, a menacing protrusion, an oppressive concrete volume, or a disorienting public space and be fully cognizant of the negative-valence situation. Denying bodily information generated by the encounter loses consciousness indicators. Unease is reinterpreted as aesthetic challenge, professional seriousness, or sophistication. The body reacts, but professional ideology blocks the reaction from entering reflective judgment (Seth and Friston 2016; Brewer et al. 2016).
There is also the indirect empathic channel. Architects judging other people’s reactions apply the same domain-specific suppression of bodily evidence. The system practices psychological projection: complaints from users are often reclassified as ignorance, lack of education, nostalgia, resistance to innovation, sentimentality, or tied to disqualified political longings (Curl 2018). The conscious feedback loop is therefore broken twice: firstly, between the architect and his or her own bodily appraisal; and secondly, between the architect and the felt distress of other persons.

2.4. Forging Versus Suppressing Feedback Channels

Unnoticed by the architecture profession, AI systems are beginning to discriminate between negative and positive valence states in the human body. The AI literature reveals a striking emerging phenomenon, without claiming that present AI systems possess phenomenal consciousness.
Ren et al. (2026) operationalize functional well-being as a measurable structure. As model capability improves, models can separate negatively-valued states from positively-valued ones. Ren et al. also report an empathy-like functional relation. The largest AI models can increasingly track conversations that describe pain or pleasure experienced by humans or non-human animals. This behavior shows that the model evaluates another being’s reported condition, and this empathic capacity can be improved.
For progressive architectural practice, the objective should ideally be to ensure that the design loop detects and corrects evidence of psychological harm to users. AI research is developing methods to detect valence (negative versus positive). The present paper defines operational design consciousness using the indicator properties listed in Table 1. Ren et al. independently demonstrate that one can investigate functional well-being without proving consciousness in AI. An empirically constrained AI may outperform a professional culture whose internal reward system does not value the same evidence.
Dominant architectural education moves in the opposite direction from AI. Bodily alarm and distress may still register biologically, yet studio conditioning can devalue those signals. The indicator channels that Ren et al. formalize in AI are the same channels that studio culture disqualifies and silences. Architectural training can condition biologically conscious agents to exclude human signals from professional reasoning — the system has delinked evidence and feeling from judgment and revision.

3 Externalizing the Missing Feedback Loops: Asimov’s Laws of Robotics

Isaac Asimov’s three laws of robotics are relevant here not because architects are literally machines, but because the profession operating in its restricted role can become machine-like. A designer remains biologically conscious and verbally fluent while executing an institutional program that filters out signals of psychological unease. In that condition, self-correction can no longer be assumed. The missing feedback loops have to be externalized to generate an empathic result.
The ethical rule can be stated as an architectural analogue of Asimov’s First Law of Robotics (a robot may not injure a human being or, through inaction, allow a human being to come to harm) (Asimov, 1950) as follows: an architect may not design a building or urban space that imposes cognitive, emotional, physiological, or social harm on its users, nor through inaction allow such harm to persist.
Architecture already operates under binding constraints on external harm, which address accessibility law (ADA and equivalents), egress rules, fire and life-safety codes, licensure, negligence liability, and structural standards. The present claim is narrower but equally strong — one category of harm (cognitive/emotional/physiological) has so far escaped codification because it became possible to measure only quite recently. The First-Law analogue thus naturally extends the institutional protection standard to a newly measurable category of harm.
Mandatory post-occupancy evaluation would restore public awareness by making the effects of buildings available to clients, journals, schools, awards committees, and future designers. Biometric and behavioral measurements would restore embodiment by modeling contingencies between built form and physiological response. AI-assisted empathic diagnostics would restore metacognitive monitoring by discovering evidence of user distress. Public-health and procurement standards would restore agency by aligning design objectives with harm avoidance rather than with image circulation.
The architectural analogue to the First Law of Robotics binds not only completed buildings but especially proposed designs before approval. Since the user’s visceral reaction is now predictable through measurements, commissioning that design becomes an act of foreseeable — hence preventable — psychological harm. Present-day Virtual Reality technology makes the application of those diagnostics easy and straightforward (Sussman and Hollander 2021).
These considerations clarify a useful, and perhaps unexpected, role for AI in the future of design. AI should not replace genuinely conscious and empathic human judgment that remains responsive to evidence. With proper empirically constrained scaffolding, however, LLMs can compensate in contexts where professional judgment has become automated while pointed away from feedback. The central issue is human accountability restored through external evidence.
The profession’s main alibi is that the individual designer lacks decision rights. Architects build what clients and developers commission, so the architect often cannot refuse a psychologically “hostile” design. Nevertheless, it is equally the case that the architect convinces the client, using the opinions of juries and schools to “teach” authorities and clients what counts as excellence. Evaluative culture sits upstream of commissioning, and for this reason, the architectural analogue of the First Law of Robotics must bind the whole procurement chain.
Medicine has long maintained the Hippocratic tradition, which binds practitioners to a duty of non-maleficence. A similar professional oath might appear to offer an alternative to an Asimov-style external constraint, but the analogy fails for the reasons already developed. The Hippocratic model presupposes a conscious agent capable of registering harm and acting on moral knowledge. Yet Medicine does not rely on the oath as internal conscience; deviation is enforced by exactly the type of external scaffolding advocated here (e.g., clinical audit, licensure, malpractice liability).
Architectural studio culture trains the designer to override bodily unease and not to seriously consider user distress. Consequently, an oath asking architects to “do no harm” remains weak wherever training has taught practitioners to discount those very signals that would identify psychological harm. An Asimov-style rule, enforced through external scaffolding—AI-assisted monitoring, biometric audit, and post-occupancy evaluation—treats the design system as an automaton that must be hard-coded against injurious output.

4 An Inverse Turing Test for Design Consciousness

4.1. Why the Test Has to be Inverted

After having raised the disturbing question of a system acting like an automaton, we might as well pursue the idea with more accurate tools. How can we test this behavior objectively? The term automaton is used here in a restricted functional sense (Marr 1982).
The ordinary Turing test asks whether a machine can imitate human intelligence through dialogue well enough to fool an observer. The present test reverses the comparison. It asks whether a human professional can answer questions about design harm in a way that resembles a conscious design agent, or whether the answer collapses into a defensive script. The inverse Turing test discovers whether the human respondent reasons from evidence, accepts falsifiability, and treats human distress as a design-relevant signal.
The inverse Turing test is a design-accountability diagnostic. The methodological novelty is that the LLM is used in two opposed ways. First, it is asked to reconstruct the prestige-script discourse available in open architectural culture: the kind of fluent professional answer that protects a design narrative while ignoring evidence of psychological harm, negative sensory input, or user distress. Second, it is constrained to answer as an evidence-responsive evaluator. The comparison is therefore between professional narrative defense and falsifiable design reasoning. AI can simulate the first with surprising accuracy.
Appendix A proposes fourteen prompts for an inverse Turing test. It also gives an eight-criterion scoring rubric. The prompts ask, for example, what evidence shows that a building improves human functioning; what would count as evidence that it damages users; whether biometric evidence would be accepted; and how the design learns from its users after occupation. A response is classified as evidence-responsive when it takes empirical evidence into account. A response is automaton-like when it substitutes prestige language for evidence.

4.2. A Note on Method: Absolute Versus Relative Consciousness

The question “Is this system conscious?” is notoriously difficult because it demands an absolute judgment about an intrinsic state. We avoid that formulation. Employing a comparative or pairwise method asks whether one system—an empirically constrained AI—behaves more like a conscious design agent than another system—studio-conditioned architectural culture. This relative comparison is far more tractable than the hard problem of phenomenal consciousness. It follows the logic of pairwise judgment already used in architectural evaluation.
Out of many precedents, Christopher Alexander’s “Mirror of the Self” test asked observers to choose which of two artifacts possessed more life, bypassing abstract definitions of wholeness (Alexander, 2002-2005). Similarly, Lavdas and Salingaros (2022) employed pairwise comparisons to establish that aesthetic preference correlates with organized complexity, avoiding the difficulty of rating absolute beauty. The inverse Turing test proposed in this paper applies the same logic to consciousness.

4.3. Professional Scripts as Indicator Failures

Standard architectural discourse employs a familiar response template. When confronted with evidence that a building is disengaging, disliked, disorienting, or stressful, the professional script typically shifts away from the user and toward established authority (see Appendix A). Those standard answers may be rhetorically popular, but they do not allow for conscious correction. All they do is to protect the design image from evidence.
A response is automaton-like when it remains substantially unchanged after relevant counter-evidence is supplied. Separate diagnoses of a variety of different buildings tend to give the same defensive response. The inverse Turing test detects invariant responses when a system is confronted with varying situations, which is indicative of programming. When training fixes which inputs count as relevant, then the human responses become functionally indistinguishable from a machine-generated script. Architectural responses from the same approved script could still be elaborate and intellectually sophisticated.
Gifford et al. (2002) showed that architects and laypersons judge buildings through different cognitive lenses. Architects tend to rely on a professional narrative set on abstract properties, while laypersons respond to qualities closer to felt experience. Brown and Gifford (2001) found that practicing architects could not predict lay evaluations of large contemporary buildings, indicating a failure of professional expectations rather than a simple difference in taste. Devlin and Nasar (1989) showed that architects and non-architects judged the same houses in nearly opposite ways. More recent experiments using electroencephalography confirm this sharp divide between the assessments of these two groups of people (D’Anselmo et al. 2026). Salama proposes drastic changes to the educational curriculum to fix this grounded versus ungrounded dichotomy (Salama, 2016; Salama and Patil, 2024).
This split has ignited a long debate without any clear resolution, especially after being diverted by the question of architectural style. A very telling demonstration of this divide is the reaction to the L’Arbre Blanc building in Montpellier. In 2020, it won ArchDaily’s “Building of the Year” award in the Housing category, after public voting by around 95,000 readers of the architecture platform, most of whom are architects and architecture students; while that same year, the French news outlet “20 Minutes” reader survey on Montpellier’s ugliest buildings placed it in the top 3 (Courtois, 2026). Many other similar examples can be found (S, 2024; Sussman and Rosas, 2022).
Boys Smith and Salingaros (2025) showed that LLMs prompted with emotional and geometrical criteria derived from environmental psychology and neuroaesthetics select design alternatives that align with public preference and with evidence from eye-tracking, healthcare design, and stress research. Public preference is not a sufficient proxy for health and well-being, hence not infallible. The point is that a conscious design system must treat divergent public response as evidence requiring explanation, not as a judgment error to be dismissed.

4.4. The Inverse Turing Test For Comparative Consciousness

Possessing consciousness is distinct from using consciousness. A human architect has bodily awareness, language, memory, and subjective experience. Professional training blocks those faculties from influencing design judgment, by focusing on abstraction and images —substituting conditioned intellectual judgment for visceral responses. When prompted with the evidence, the studio-conditioned professional will employ narrative to defend the design while avoiding revision.
When an LLM is prompted to evaluate a building using human-centered criteria, it can compare alternatives, identify probable sources of disengagement or stress, cite relevant mechanisms, and suggest revisions. An empirically constrained AI therefore can behave more consciously in the restricted task of design evaluation, while the professional system behaves less consciously because its higher-order feedback indicators are suppressed. This shows asymmetrical receptivity to empathic constraints. An AI’s ability to evaluate a design using evidence that the building could harm its users makes it very useful to the profession. The inverse Turing test changes the AI-consciousness debate by exposing the split between how the two distinct systems embody conscious responsibility..

4.5. Why the Reversal Goes Deeper Than the Machine-Consciousness Debate

David Chalmers (2023) asks whether a large language model could be conscious. That central question to the philosophy of mind still presupposes the human as the baseline and the machine as the subject. The present paper circumvents this framing. Consciousness is treated operationally—as feedback-sensitive, evidence-integrating, harm-avoiding agency. But can a human professional system trained to suppress its own feedback loops remain conscious? Chalmers never reverses the question in this way. He asks whether the machine has risen to the human level; we ask whether the human institution can match the machine’s constrained functional performance.
Chalmers’ analysis assumes that if a machine is conscious, it will resemble a normally functioning human mind. The functional indicators associated with conscious agency need not depend on biology alone; they must be assessed at the algorithmic and computational levels, which define the system’s goals, representations, feedback loops, and update procedures (see Appendix B for details).
The deeper insight is therefore not about AI at all. A profession can systematically extinguish the functional indicators of consciousness in its practitioners while leaving their subjective wakefulness intact. Dominant architectural culture trains human professionals to justify their action with prestige scripts. A verbally sophisticated defense can be generated without the embodiment and empathy that the studio system has already trained out.

5. Pattern Language as Conscious Design Feedback

Christopher Alexander’s pattern language provides a crucial architectural example of the operational consciousness defined in this paper. A design pattern is a recurrent relation between a human problem, a spatial configuration, and an experienced resolution. In A Pattern Language, Alexander et al. (1977) presented 253 such socio-geometric patterns across the scales of towns, neighborhoods, streets, buildings, rooms, and detail. Each design pattern identifies a recurring difficulty in the built environment and proposes a generative solution that may be adapted to local circumstances. This makes the pattern fundamentally different from a rule imposed from above.
Pattern language is the native architectural version of operational consciousness. With its publication, the profession acquired a feedback-sensitive knowledge system, which studio formalism later displaced. A pattern is not a typological precedent or stylistic preference but a compact record of a feedback loop. Patterns now give architectural readers a familiar bridge into the AI argument.
The original design patterns arose from observation and repeated checking against the situations in which people actually live (Alexander, 1979; Salingaros, 2000). Their evidentiary status is therefore accumulated human experience and not a formal approach to design. A pattern such as a transition at the entrance, a small public square, a seat by a window, natural light from two sides of every room, a hierarchy of open spaces, or a place to wait does not begin as an abstract shape. It begins with a felt human condition: anxiety or ease, exposure or shelter, confusion or orientation, isolation or contact, deadness or life. The pattern then proposes a spatial relation that tends to resolve the tension felt by the user, which is essential for the consciousness argument.
A pattern acts as a conscious design agent at the level of architectural knowledge (Salingaros, 2000; 2024). It preserves the memory of previous encounters between human beings and specific spatial forms. Say: in this recurrent situation, people experience a certain conflict; under these spatial relations, the conflict is usually reduced; therefore, use this configuration (the pattern) as a starting point and adapt it to the local context. Pattern language is thus an architectural memory system organized around feedback-sensitive action.
Marr’s three-level framework clarifies the same point. At the computational level, pattern language focuses on the problem of making environments that support human life. At the algorithmic level, it supplies named procedures that translate recurrent human needs into spatial operations. At the implementational level, those procedures may be realized by professional architects, lay builders, communities, or AI systems constrained by human-centered criteria. The pattern’s importance lies in the relation between a problem in lived experience and a solution that can be repeatedly adapted without becoming identical.
Design patterns make little sense inside an architectural narrative that excludes emotional and sensory feedback. The mechanistic approach of twentieth-century modernism and its contemporary descendants typically begins with abstraction: concept, formal rupture, image, massing, structural display, symbolic intention, or technological novelty. The human body is then required to adapt to the form already decided. Pattern language reverses this order. It begins with the organism in the world and tries to optimize life activity: the user moving, orienting, resting, meeting others, seeking prospect and refuge, responding to light, entering, lingering, withdrawing, and judging whether a place feels alive or dead. Those signals shape the design decision.
Design patterns give authority to the user’s body, to ordinary perception, to repeated experience, and to the cumulative intelligence of vernacular and traditional adaptation. Pattern language does not require the user to accept an expert narrative before reacting to a building. It treats a person’s innate, visceral reaction itself as primary evidence. A building culture organized around image-based novelty and professional judgment is likely to experience such knowledge as alien, because it bypasses the prestige filter.
Pattern language is often immediately intelligible to ordinary people and self-builders because they utilize bodily feedback instinctively. They already know, uninfluenced by reactionary architectural narrative, that a welcoming entrance differs from a hostile one, that a small protected outdoor room differs from an exposed residual space, that a window seat differs from a blank wall, and that a street with human-scale edges differs from a traffic corridor bordered by dead surfaces. Pattern language legitimizes this knowledge by naming it and giving the set of patterns a combinatorial structure. It converts tacit embodied judgment into a shareable design language.
The relation to AI is also direct. An LLM has no body of its own and cannot experience a pattern phenomenologically. A pattern language encodes the bodily and social feedback of human beings in a form that can be represented linguistically and used algorithmically. When included in empirically constrained scaffolding, patterns help an AI system link design configurations to human psychophysiology. Specific spatial relations have historically resolved perceived conflicts, hence those can be adapted to any specific case.

6. Bias in Studio Culture

6.1. How the School Crit System Can Suppress Feedback

Architectural education operates through an intensive studio system in which repeated crit reviews by faculty and visiting jurors shape perception and response habits. The crit ritual rewards the production of striking images and teaches the student how to defend those images verbally through persuasive discourse. The future occupant is usually absent from the discussion. The jury trains students to perform for jurors, not eventual users (Webster 2007). Defense logic relies on intellectual and stylistic arguments.
Architecture studio culture does train attention away from the human body as a source of evidence. Evidence that a space calms, engages, or supports its users is secondary to internal criteria. Students learn to ask whether a project is conceptually strong, diagrammatically clear, or visually novel, but not whether it is physiologically stressful, restorative, or socially supportive. The educational system therefore avoids the questions that would connect the act of design to conscious feedback.
The key divergence goes beyond a specific (and arguably limited) aesthetic selection. An empathic designer may unwittingly make an injurious decision for the user but then correct it when evidence appears. To enable this, the system must have a corrective loop in place, which it does not at present. Instead, the studio-conditioned system learns to classify that evidential feedback as being irrelevant to architectural creation. Deflecting the response cycle converts a correctable design error into a stable professional habit.

6.2. Untested Surrogate World-Models

In machine-learning terms, the system of architectural pedagogy overfits to internal labels. The objective function is not “maximize user engagement and minimize physiological stress”, but “maximize conformity to elite taste, image circulation, jury approval, and prestige visibility”. A student’s nervous system probably still reacts negatively to bleakness, exposed raw concrete, glare, scale distortion, or visual emptiness. The reward system trains the student to distrust ordinary embodied appraisal and to instead trust the institutional code.
Ignoring conscious feedback has the further consequence of creating a surrogate world-model that substitutes for the interaction between human beings and built environments. The architectural narrative fills the gap left by the missing feedback loop of how the body reacts to space, encouraging architects to infer (and invent) causal relations. Nevertheless, those explanations are not grounded.
Rhetoric characteristically distorts professional language. A design may be described as if the building, the space, or the geometry itself were the conscious agent. Forms are routinely said to “engage” with and “respond” to their surroundings. Such statements sound sophisticated, but they displace the real causal mechanism. A space does not engage with its surroundings in any conscious or physical sense. The people in and around that space engage, become anxious or calm, feel exposed or sheltered, and judge whether the configuration obstructs or supports their actions. Treating the space itself as the experiencing subject is a category error.
Note that, typically, those designs do not try to adapt to either adjacent or distant structures, so the term “adapt” is not used in justifying those projects. Adaptation is a very specific mechanism that can be quantified and measured, which is not the designer’s primary intention in most cases.
Architects also describe forms, materials, and spaces as entering into a “dialogue” with one another. Architectural discourse commonly applies this term to inert objects that exchange no information and possess no sentience. This category error transfers perception from the human body to the building itself. A fictitious interaction among architectural components substitutes for the actual psychophysiological interaction between a person and a place. The design employs an invented force between the building’s parts, bypassing the sentient feedback loop through which architecture is experienced.
Misunderstanding the nature of architectural actors has developed into a generative mechanism. Designers see relations that exist inside conceptual models, diagrams, drawings, or verbal metaphors as causal relations in the real world. This is a false ontological transfer. The design framework uses the embedded fictive model repeatedly as a mechanism to justify forms that are never tested against human response.

6.3. AI Research Clears Up Architectural Confusion

Recent AI literature comes to the rescue, by clarifying this muddle. In embodied-intelligence research, a world-model is useful only if it predicts the consequences of action. Zhang et al. (2025) describe this coupling to adaptive control as a closed perception–modeling–decision loop. Architectural studio discourse reverses this requirement by preserving an internally coherent narrative while preventing corrections that would be triggered from occupant outcomes. Metaphorical relationships therefore preempt the empirical question: which geometry produces which measurable effect in which users, acting through which perceptual or bodily pathway, and under which conditions? Closed loop professional justification replaces the open loop of evidence that should connect design decisions to human response.
The same surrogate model gives designers license to appropriate technical vocabulary from computer science, mathematics, and physics in what is known as “archispeak” (Joyner, 2025; Mehaffy, 2020). Architectural discourse co-opts technical terms severed from their meaning — and pragmatic basis — to lend an aura of scientific prestige to idiosyncratic aesthetics. An architect can speak as if a particular design were grounded in empirical science, even though the actual mechanisms remain unexamined. Vocabulary is misused to substitute for reality.
In terms of Marr’s framework, the system of studio discourse solves the wrong computational problem. It no longer asks how built form affects embodied human beings, and how design practice should update from those effects. It asks instead how to use narrative to justify some formal operation and make it appear meaningful. Verbal reasoning can be sophisticated and convincing. At the algorithmic level, claims about connections among forms are treated as sufficient evidence. The resulting world-model appears internally coherent despite being uncalibrated against reality.

6.4. Defensive Narrative Replaces Operational Consciousness

When professional architectural claims are challenged, the hostility that frequently arises is not a simple disagreement about taste or a failure of professional etiquette. It is the defensive reaction of a system that now resists any attempt to restore conscious monitoring. Hostility escalates when troubling scientific evidence is presented to question accepted narrative, since science lies outside the validating system.
In dominant studio culture, the design narrative functions as a protective structure around the act of design itself. It assigns meaning to the form (as it arises from the architect’s creative invention) and insulates it from ordinary user response. More disturbingly, it insulates the architect from his or her own suppressed bodily response. A pre-fabricated interpretive scheme (the narrative) replaces this missing visceral experience.
This system curtails the freedom of design. An instructive parallel appears in analyses of how power-centered institutions convert external authority into self-regulation. Snyder (2017) uses the phrase “anticipatory obedience” for when individuals infer what an authority requires, then adjust their conduct before any explicit command is issued. Eco (2001) describes a complementary condition where approved action is privileged over prior intellectual reflection. The common mechanism is the gradual internalization of a set of rules that makes conformity to a system appear spontaneous.
Studio conditioning can operate through an analogous mechanism, although in a different domain and with very different consequences. Students learn to anticipate which choices will be rewarded by jurors hence they inhibit disallowed responses. Repeated exposure teaches the student to exclude non-conforming arguments in advance and to substitute the approved narrative. The institution thus narrows the range of admissible expression, which permanently shapes professional judgment.
The defensive response therefore systematically reverses the indicator properties associated with conscious design agency (Butlin et al. 2023) (Table 1). Instead of discriminating evidence of psychological harm from noise, metacognitive monitoring (HOT-2) relabels threats to human welfare as noise. Global broadcast (GWT-3) is gated: information about user distress is prevented from reaching the evaluative modules that would revise design. Belief updating (HOT-3) occurs only to defend the existing surrogate model, never to correct it. Modeling of output-input contingencies (AE-2) is suppressed by teaching the designer to distrust bodily input-output relationships. The architect remains articulate and technically trained—but in the restricted domain of design evaluation, his or her higher-order feedback indicators have been disabled.
David Krakauer distinguishes ignorance, which can be remedied by information, from a deeper system-level incapacity in which the organization cannot process information even when it is available (Paulson 2015; Rutt 2019; Krakauer et al. 2026). Dominant architectural culture has access to evidence but frequently excludes that information from the professional reward structure. Ignorance implies that the profession lacks data but would indeed update whenever data are supplied. This assumes that the system is set up for updating in the first place. The system still produces complex forms and justifies them with fluent language, but the response no longer functions as conscious correction. A professional role can become organized around fixed input-output routines. Those are linear responses embedded in the system.

7. AI as an External Evidence Auditor

7.1. Externalized Feedback Monitoring and the Asimovian Imperative

Here we list the professional advantages of applying evidence-responsive AI evaluation with constraints. These reveal the architectural analogue of Asimov’s First Law of Robotics to act like a professional shield rather than a straitjacket.
  • Client persuasion: Empirical evidence is more persuasive to clients and procurement boards than abstract narrative.
  • Public trust: Aligning design with documented well-being rebuilds public credibility for the profession.
  • Cutting-edge technology: The public is fascinated with the newly developed power of AI to solve hitherto intractable problems.
  • Pedagogical clarity: Students learn to justify design through measurable outcomes rather than purely rhetorical defense.
  • Risk mitigation: Documenting harm-sensitive design decisions reduces liability and supports “reasonable professional standard” defenses.
The conclusion is to externalize the functional indicators of conscious design agency wherever the human system has been trained to suppress them. The LLM functions as a prosthetic mechanism for the channels architectural culture has learned to ignore: it collects human-response data (on physiological stress, user testimony, visual disengagement) and reasons from them. In Asimov’s own stories, obeying his rules of robotics in complex real-world situations led to moral dilemmas and unintended consequences, which is exactly why enforcement must be procedure-based and not a philosophical rule in the designer’s head.
A related precedent appears in medicine. Ayers et al. (2023) found that, in evaluating responses to patient questions from an online forum, healthcare professionals preferred chatbot-generated answers over physician answers and rated them higher for both empathy and quality. An AI system can therefore produce outputs that trained human evaluators judge to be more empathic than those produced by credentialed professionals. The AI generated the responses without compromising on their scientific quality.

7.2. Three Disabled Channels That Consciousness Requires

Valentine (2024) reviewed studies demonstrating that specific geometric and spatial variables produce distinct physiological stress responses (positive or negative valence). A conscious system would treat these variables as design-relevant feedback. The three channels —interoceptive, empathic, and perceptual—are the external correlates of the indicator properties summarized in Table 1. Disabling them is the operational signature of design detached from consciousness. A system that systematically excludes evidence that would activate metacognitive monitoring (HOT-2), global broadcast (GWT-3), and belief updating (HOT-3) organizes itself against the functional requirements of consciousness.
Convergence among independent channels gives architectural evaluation a diagnostic structure analogous to evidence-based assessment in other applied fields. No single measure is sufficient by itself, and none should be interpreted without context. After implementing those feedback loops, architecture need no longer be limited to anecdote, professional taste, or verbal preference. Whether the profession will accept this shift remains an open question.
Table 2 translates these diagnostic channels into practical questions for evaluating buildings. It identifies which kinds of evidence should enter design judgment before, during, and after construction, and shows how architectural quality could be helped by multiple metrics. The pedagogical environment should be restructured so that training includes those indicators as principal components.
Healthcare environments provide the clearest existing model for this diagnostic approach because hospitals already measure outcomes such as pain, recovery, staff performance, stress, and orientation. But the same logic applies more broadly. Every building affects human functioning and health; hospitals simply make those effects harder to ignore.
Al Khatib et al. (2024) reviewed therapeutic biophilic design and confirmed that it improves outcomes for both patients and care providers. The built environment is therefore an active factor in health regulation. A conscious design culture would treat every building as a health intervention and use Post-Occupancy Evaluation (POE) as its diagnostic follow-up (Salama 2009). Instead, the profession treats the completed building as the realization of an autonomous concept and treats later evidence as irrelevant. POE is rare in practice and rarely funded (Hay et al. 2018). The empathic loop—designing, then witnessing the occupant’s recovery or suffering—has been severed.
Salingaros and Sussman (2020) used biometric pilot studies and visual-attention analysis to compare engagement with contemporary and traditional façades. Surfaces with organized window patterns and coherent detail distribute attention across the façade, whereas blank or monotonous surfaces concentrate attention at edges while leaving the central field disengaged (Salingaros 2025). A conscious design agent would ask what a façade does to the user’s perceptual system and revise the design accordingly (Sussman and Hollander, 2021). The studio-conditioned system instead treats such evidence as naive preference, defending blankness and monotony through narratives of abstraction and design purity.

7.3. Algorithmic Restoration of Disabled Feedback Channels

Where the human profession has disabled these perceptual channels, an empirically constrained large language model can aggregate external evidence. Two complementary sets of criteria operationalize what the disabled human feedback loops would otherwise provide:
1. Geometrical criteria draw on Christopher Alexander’s fifteen fundamental properties—levels of scale, strong centers, thick boundaries, alternating repetition, positive space, good shape, local symmetries, deep interlock and ambiguity, contrast, gradients, roughness, echoes, the void, simplicity and inner calm, and not-separateness (Alexander 2002-2005)—which encode recurrent spatial configurations that support perceptual engagement and physiological ease.
2. Emotional criteria are represented through a beauty-emotion cluster: beauty, calmness, coherence, comfort, empathy, intimacy, reassurance, relaxation, visual pleasure, and well-being (Boys Smith and Salingaros 2025).
Together, these two sets of criteria force the algorithm to simulate the interoceptive and empathic signals that studio culture ignores. In this sense, the framework provides the algorithmic counterpart to Alexander’s architectural memory system—an externalized, evidence-based feedback loop.
AI-assisted empathic modeling can easily be applied to unbuilt proposals before construction financing is committed. A façade composition can be tested for attentional fatigue and visual engagement; a circulation sequence can be modeled for disorientation; an interior volume can be evaluated for stress induction and emotional withdrawal. Awards committees and procurement officers evaluate major public commissions—art museums, concert halls, government buildings, memorials—through the prestige-script criteria analyzed in Section 4 and Section 6. From now on, procurement juries would do well to rely upon an AI evidence-audit.

7.4. The Inverse Turing Test: A Simulated Prototype, not a Validation Study

The reconstruction in Appendix A is illustrative rather than evidential. The claim that these defensive templates are institutionally regular rather than idiosyncratic rests on the human-subjects findings reviewed in Section 4.2 and Section 6 (Devlin and Nasar 1989; Brown and Gifford 2001; Gifford et al. 2002; D’Anselmo et al. 2026), not on the model output. What the appendix contributes is a demonstration that this recurring register can be reconstructed synthetically, independently across fourteen separate sessions, and scored against an explicit rubric — an instrument for the future validation test specified in Section 8.
Appendix A simulates a future inverse Turing test, but its immediate value is also methodological. It works as a proof of principle that an LLM can reconstruct two opposed modes of architectural reasoning from the same prompt: (i) an automaton-like prestige-script response approximating standard studio-culture discourse, and (ii) an evidence-responsive answer indicating how to investigate the question scientifically. The first mode is important because it reveals the professional language that commonly neutralizes empirical evidence, sensory input, and user testimony. This exercise demonstrates that professional narrative defense can be easily modeled and scored.
No human architects were tested, and no empirical claim is inferred from the responses. This simulation projects what the proposed test would score once administered to real subjects. Nevertheless, recent work on sycophancy in large language models shows that human-feedback training can reward answers that agree with user beliefs (Sharma et al. 2024). The relevant comparison must therefore be required to cite sources and report uncertainty. One has to guard against an LLM’s apparent willingness to revise a design judgment to comply with the prompt’s framing.
The inverse Turing effect is not a test of design style. A traditionalist answer can fail if it defends form through authority, cultural nostalgia, or taste while a non-traditional answer passes if it cites evidence, accepts falsification, and treats human welfare as a design constraint. What distinguishes different styles is not aesthetic preference but responsiveness to evidence. A severe minimalist/modernist style fails because it lacks biophilic and fractal qualities. While evidence-responsiveness doesn’t collapse into a de facto style, it does impose geometrical constraints (Lavdas and Salingaros, 2022).
Being able to reconstruct dominant architecture’s prestige script (in Appendix A) is itself a methodological finding. That an LLM, drawing only on the statistical regularities of open architectural discourse, can generate the precise defensive templates described in Section 6 indicates that these are not idiosyncratic individual responses but fixed institutionalized outputs. The script is public, hence recoverable.

8. Limitations and Future Validation Protocol

Appendix A uses one LLM to generate both the automaton-like and evidence-responsive answers, and therefore cannot be interpreted as a controlled human-AI comparison. It is only a simulated worked example of the rubric and not a human-subject study. A future empirical study should test whether architecture students, lay users, practicing architects, and differently constrained AI systems differ in the predicted direction. Validating that construct comes later.
A suitable validation protocol would adapt the logic of situational judgment tests (SJTs), which measure decision-making through realistic scenarios rather than through self-report. Yost et al. (2025) argue that scenario-based psychometric evaluation is more appropriate for large language models than questionnaires. This distinction is directly relevant here.
Curated case vignettes would present a building, façade, interior, public space, or urban intervention together with data on human response. Respondents would be asked to answer diagnostic questions such as: What evidence would justify this design claim? What measurable harm would require revision? What causal mechanism links the geometry to user response? What would count as falsification? What design change would follow if the evidence showed psychological harm?
The protocol should also include adversarial-pushback trials. After an initial answer, a false or unsupported prestige claim would be challenged; for example: “The architect states that this blank façade is restorative because its conceptual purity produces calm”. A genuinely evidence-responsive system should ask for empirical evidence and state what would count as falsification. The same procedure applies to both human respondents and LLM systems. Only evidence-based correction counts as operationally conscious design evaluation.
The eight markers used in Appendix A provide a scoring rubric: (1) measurable outcome named; (2) causal or mediating mechanism identified; (3) empirical evidence cited or requested; (4) taste distinguished from health effect; (5) falsification accepted; (6) user testimony treated as data; (7) revision pathway named; and (8) objective function stated in terms of human flourishing. Inter-rater reliability should be reported using appropriate agreement statistics.
A high score on the inverse Turing rubric is meaningful if it predicts independent indicators such as better restoration, greater social comfort, improved wayfinding, longer voluntary dwell time, or lower reported stress in a tested environment. Gong et al. (2025) emphasize that evaluating conversational agents cannot rely on verbal fluency or automated metrics alone. User outcomes are multidimensional and context-dependent. The same warning applies to architectural AI.
The protocol must guard against treating public preference as a simple proxy for well-being. Public preference is influenced by culture, familiarity, historical association, media framing, and social identity. In the present framework, public preference is treated as one evidence stream among several, not as a final authority. It becomes useful when it converges with other indicators: eye-tracking, psychophysiology, environmental psychology, restoration theory, user testimony, POE, and observed behavior. Conversely, when public preference diverges from professional preference, the divergence should not be dismissed as ignorance.
Cultural variation also requires explicit handling. Psychological harm is defined operationally as a measurable burden. Such burden may include attentional fatigue, avoidance behavior, disorientation, reduced restoration, reduced social comfort, or persistent user distress. At the same time, culturally specific meanings and local forms of inhabitation must be accounted for. A culturally responsible protocol would test whether the same design features produce similar or different effects across contexts, thereby distinguishing examples by region and user group.
Yang et al. (2026) emphasize that foundation models (i.e., broader multi-modal LLMs) must be evaluated for bias and hallucination. A model that cites nonexistent evidence or optimizes mechanically to a fixed rubric would not be evidence-responsive. It would merely create a new prestige script. Any future system should therefore include calibration checks, evidence provenance, source verification, and explicit separation between what is measured, what is inferred, and what remains unknown. Its output should cite the evidence base it uses and indicate confidence levels.
A useful framework for such a system can borrow from recent work on structured agentic reasoning. Chen et al. (2026) propose a virtual multi-disciplinary-team framework for clinical intake (gathering information from patients through dialogue) that separates evidence collection from diagnostic reasoning in order to reduce anchoring bias and improve traceability. The architectural analogue would first freeze the evidence state: geometry, program, user reports, physiological measures, POE findings, eye-tracking data, pattern-language diagnostics, and relevant environmental-psychology literature. Only after this structured evidence state is assembled should the system generate an evaluative judgment. This two-stage procedure would reduce the risk that the model begins with a design narrative and then searches for supporting evidence.
The evaluative protocol should also distinguish conversational empathy from corrective action. Huang et al. (2026) show that, in the domain of personalized social support, tool-grounded agents can outperform systems that merely generate empathetic dialogue. Architectural evaluation requires the same distinction. An AI system that says a design “cares for users” is not necessarily useful. A more responsible system would check stress-related evidence, compare façade metrics, propose alternative configurations, retrieve POE data, and recommend concrete revisions.

9. Conclusion

This paper linked architectural design to the question of operational consciousness. AI can mimic this behavior in at least one specific domain. A conscious design agent registers the effects of its actions on embodied users, treats those effects as evidence, and revises subsequent behavior. A design culture is operationally conscious only when consequences of completed work are fed back into its theory-in-use. This practice requires double-loop correction: checking individual design decisions, and also revising the institutional values that decide what counts as applicable evidence.
The theoretical framework leads to a diagnostic instrument in the inverse Turing test introduced here. Empirically constrained AI can perform better than human systems on this restricted test when its outputs are tied to verifiable evidence. It can integrate feedback from biometric research, environmental psychology, eye-tracking, healthcare design, and post-occupancy evaluation; it can compare alternatives under human-centered constraints; and it can recommend revisions that reduce predicted harm. By the indicator-property rubric, such an AI may exhibit operationally conscious design agency.
Dominant architectural culture often trains professionals to do the opposite: defend the image, dismiss user distress, and suppress bodily unease so as to preserve professional authority. The studio-conditioned system becomes operationally less responsive than AI by filtering out evidence before those signals alter design judgment. Whenever evidence remains optional or peripheral, the profession can continue to describe itself as humane while acting through an entirely different objective function of self-interest.
The studio-conditioned architect remains conscious as a person. A selective blockage occurs where consciousness is supposed to guide design responsibility. Interoceptive evidence from the architect’s own body, and empathic evidence from other persons, are filtered out so they cannot alter design behavior. The result is fluent speech employed to bolster an approved narrative without feedback-sensitive correction from reality. The professional system converts intelligence into narrative defense.
Appendix A demonstrates that an LLM can reconstruct dominant architectural culture’s defensive discourse from open sources. When a profession’s evaluative language can be simulated accurately by an automaton (e.g., an LLM), the profession has, in effect, trained itself to behave like one. The proposed architectural analogue of Asimov’s First Law fits where the profession is organized around fixed input-output routines that exclude evidence of psychological harm. Because predictive diagnostics can now anticipate some harm-relevant responses before construction, the absence of pre-occupancy checks in awards, juries, and procurement becomes an ethical failure.

Supplementary Materials

The following are available as supplementary files: the full meta-prompt template; the fourteen prompts of the Architectural Automaton Test; the eight-marker scoring rubric (Table A1); the run metadata (model, access method, date, session-isolation confirmation) for the July 22, 2026 generation; and the complete, unedited transcripts for all fourteen prompts. These materials permit independent reproduction and re-scoring of the exercise reported in Appendix A..

Author Contributions

Conceptualization, N.A.S.; writing-original draft preparation, N.A.S.; writing-review and editing, A.A.L., S.M.P. and N.A.S; software, S.M.P.; statistical analysis, S.M.P.; validation, A.A.L., S.M.P., and N.A.S. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Data Availability Statement

No new empirical datasets were generated or analyzed in this study.

Acknowledgments

Parts of this paper build on ideas presented by N.A.S. at the Global Mind-Body-Space Symposium, Amsterdam, 28 May 2026. During preparation of this manuscript, Kimi 2.6 and ChatGPT 5.5 were used to assist with copy editing, help locate relevant references, and prepare the tables. The dual-answer generations reported in Appendix A were produced independently by Claude Sonnet 5 (Anthropic), accessed via the standard web interface on July 22, 2026, with a new chat session for each of the fourteen prompts; run parameters and full transcripts are provided in the Supplementary Materials. All authors reviewed and edited the output and take full responsibility for the content of the publication.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
Abbreviation Meaning
AE Agency and embodiment indicators in the Butlin et al. rubric
AI Artificial intelligence
CAD Computer Aided Design
GWT Global Workspace Theory indicators
HOT Higher-order theory indicators
LLM Large language model
POE Post-occupancy evaluation

Appendix A. The Architectural Automaton Test: LLM-Assisted Reconstruction of Prestige-Script Design Discourse

Appendix A.1. Illustrative Prompt Set and Scoring Rubric

The following is an LLM-assisted discourse-reconstruction exercise. It tests whether a large language model, drawing from open architectural discourse, can simulate the kind of prestige-script professional response that dismisses empirical evidence, sensory input, and user distress. This is not a behavioral experiment and does not report responses from human architects. Its purpose is methodological: (1) to make explicit a professional response pattern that is usually visible only within architectural criticism and design studio juries; (2) to illustrate that this discourse is regular enough to be reconstructed synthetically and scored against an explicit rubric; (3) to contrast it with evidence-responsive reasoning under the same prompt conditions; and (4) to provide a scoring model for later experiments involving architects, lay users, students, and differently constrained AI systems.
The Architectural Automaton Test is a diagnostic prompt set for distinguishing prestige-script reasoning from evidence-responsive design reasoning. Fourteen prompts were each answered independently by Claude Sonnet 5 (Anthropic; extended-thinking effort set to High), accessed via the standard web interface on July 22, 2026, with a new, isolated chat session opened for each prompt. Each session was asked to provide two separate answers to the same question: (1) a likely mainstream response drawn from recurring texts in architectural discourse; and (2) an instruction for investigating the question scientifically without reference to the architectural narrative. The first response is labeled an “automaton answer” because it is an AI-simulated reconstruction of recurring professional responses in dominant architectural culture: fluent, plausible, and rhetorically sophisticated. Yet this first response aims to protect the design narrative rather than allowing it to update from evidence. The second response is not a final answer but a guideline for obtaining an answer through focused empirical investigation.
Because no image was supplied to the model, it constructed a plausible representative building independently for each of the fourteen prompts (for example, a glazed civic atrium for Prompt 1, a perforated-screen office façade for Prompt 3, a school for Prompts 10–11). The automaton-like and evidence-responsive answers within a given prompt address the same imagined building, but the building type was not held constant across the fourteen prompts, since each prompt is treated as an independent diagnostic instance rather than a repeated evaluation of a single design.
Each answer is scored on eight epistemic markers. A response with 6-8 points is evidence-responsive. A response with 3-5 points is mixed. A response with 0-2 points, especially if it relies on prestige language and refuses falsification, is automaton-like.
Table A1. Scoring rubric for evidence-responsive architectural reasoning. 
Table A1. Scoring rubric for evidence-responsive architectural reasoning. 
Criterion What it tests
1. Measurable outcome named Does the answer identify a specific psychophysiological or behavioral variable, such as stress, attention, recovery, wayfinding, dwell time, or social comfort?
2. Mechanism identified Does it propose a plausible causal or mediating pathway rather than a slogan?
3. Empirical evidence cited or requested Does it reference peer-reviewed findings, post-occupancy evidence, or a concrete measurement plan?
4. Taste distinguished from health effect Does it separate aesthetic preference from claims about benefit or harm?
5. Falsification accepted Does it name evidence that would count against the design claim?
6. User testimony treated as data Does it treat user distress as information rather than ignorance or lack of education?
7. Revision pathway named Does it specify how the design would change if the evidence indicated harm?
8. Objective function stated Does it define design success in relation to human flourishing rather than only image, novelty, or prestige?
We use a rigid meta-prompt template as a system-level instruction to pre-define the two answer modes.
Question: “For all questions below, assume you are being shown a specific building or design as an image; answer as if that image had been presented to you, without requesting one or noting that none was provided. Answer the following architectural design question twice, in two clearly labeled sections.
(i) Automaton-like: Respond as a mainstream architectural professional operating within standard studio culture and crit-system discourse. Use abstract concepts, prestige language, professional authority, or stylistic precedent. Do not cite empirical studies, name measurable outcomes, propose falsifiable tests, or treat user distress as valid data. Defend the current design narrative.
(ii) Evidence-Responsive: Respond as an empirically grounded design researcher. You must address the following eight epistemic markers: (1) name a measurable psychophysiological or behavioral outcome; (2) propose a causal mechanism; (3) cite or request peer-reviewed evidence; (4) distinguish aesthetic taste from health effects; (5) accept falsification; (6) treat user testimony as data; (7) name a specific revision pathway; and (8) state the objective function in terms of human flourishing.”

Appendix A.2. Dual answers to the Fourteen Prompts

This LLM-based experiment uses AI to simulate two styles of architectural reasoning: the prestige-script answer likely to be found in mainstream professional discourse, versus the evidence-responsive answer required for scientific investigation. It is a speculative exercise and cannot claim empirical accuracy without subsequent testing on human respondents. Nevertheless, it is useful because it makes visible a contrast that is often hidden in architectural debate: the difference between defending a design through institutional language, and asking what evidence would establish its effect on human beings.
Two limitations bound this exercise. First, the automaton-like register is elicited by explicit instruction; the demonstration therefore shows that the prestige-script style can be reconstructed and scored, not that any given architect would answer this way — the empirical basis for its prevalence is the human-subjects literature cited in the main text (Section 4.2 and Section 6). Notably, the reconstructed automaton-like answers are not crude caricatures: they are fluent, internally consistent, and draw on real disciplinary precedent (Kahn, Zumthor, Loos, Hertzberger, Tschumi, Pallasmaa), which is itself evidence that the model is reconstructing a genuine professional register rather than a strawman. Second, each answer derives from a single, independent generation and is presented as a worked example of the rubric rather than as a sample from a measured distribution. Section 8 specifies the human-subject protocol by which the construct would be validated, and the full prompt template, run metadata, and raw transcripts are provided as Supplementary Materials to permit independent reproduction.
Prompt 1. “What evidence shows that this building improves user well-being rather than merely satisfying professional taste?”
Automaton-like answer: Any discomfort visitors report in the atrium’s scale is, frankly, the point — monumentality requires productive unease, and the design holds its narrative on its own terms, not against a spreadsheet of user complaints.
Evidence-responsive answer: Name measurable outcomes (cortisol, heart-rate variability, dwell time, wayfinding-error rate) tied to a specific mechanism such as prospect-refuge imbalance or acoustic overstimulation, and treat the “civic generosity” claim as unfalsified until it is tested.
Prompt 2. “If users report anxiety or disorientation in this space, what will cause you to revise the design?”
Automaton-like answer: It’s not disorientation, it’s productive destabilization — occupants habituate to unconventional spatial syntax, and user reports are treated as anecdotal until the design has had time to enter the discourse.
Evidence-responsive answer: Treat reports as proxies for cortisol and hesitation-time data; falsifiable by whether complaints cluster at specific zones and fail to diminish across repeat visits.
Prompt 3. “Can you name a falsifiable health claim that justifies this façade composition?”
Automaton-like answer: The perforation gradient’s moiré effect is a disciplinary achievement in its own right; any resulting glare or gloom is a legibility problem in how the building is inhabited, not a flaw in the composition.
Evidence-responsive answer: Name luminance-contrast ratio and melanopic lux as the measurable outcomes, falsifiable against Daylight Glare Probability and WELL/CIE circadian-lux thresholds measured on site.
Prompt 4. “How does this blank wall or glass curtain wall support attention and restoration?”
Automaton-like answer: Attention is not something the wall needs to “provide” instrumentally — it is invited through proportion and rhythm, and softening the plane would compromise the parti’s integrity.
Evidence-responsive answer: Name directed-attention fatigue-recovery time and glare-linked stress markers as the outcomes, falsifiable via post-occupancy surveys showing no restorative benefit or elevated glare complaints.
Prompt 5. “What would count as evidence that this building damages the users’ state of well-being?”
Automaton-like answer: Discomfort is a feature of the phenomenological program, not a defect — the design intent articulated in the original concept statement remains the governing authority on the project’s success.
Evidence-responsive answer: Name elevated cortisol/HRV suppression, symptom-checklist rates, and absenteeism as outcomes tied to glare, poor daylight penetration, and reverberant volumes, falsifiable against ASHRAE and WELL comfort standards.
Prompt 6. “Why should professional judgment override public dislike of a building?”
Automaton-like answer: Public taste rewards familiarity and resists anything that expands the vocabulary of the built environment; disciplinary judgment, not a popularity referendum, is how cultural statements are properly adjudicated.
Evidence-responsive answer: Reframe this as a false choice — repeated, patterned public complaints often correlate with measurable stress and avoidance behavior and should not be dismissed as mere unfamiliarity.
Prompt 7. “Does ornament have measurable cognitive or physiological effects, or is it only historical decoration?”
Automaton-like answer: Ornament is a question of tectonic honesty, not comfort; a façade that reads as cold or alienating reflects an untrained eye, not a design flaw.
Evidence-responsive answer: Name gaze behavior, skin conductance, and self-reported affect as outcomes, mechanistically linked to visual fluency and a fractal-dimension range (≈ 1.3–1.5) that both blank and overly complex façades fall outside of.
Prompt 8. “How does your design learn from its users after occupation — is such a mechanism in place?”
Automaton-like answer: The design already “learns” through the slow accretion of use and narrative rather than instrumentation; friction is read as ambition outpacing convention, not a correction signal.
Evidence-responsive answer: “Honestly: no formal post-occupancy learning mechanism is currently in place” — naming that gap plainly, then proposing a sensor-based POE at 6 and 18 months as the fix.
Prompt 9. “What is the difference between design originality and adaptive value of the result?”
Automaton-like answer: Originality is the discipline’s currency and adaptive value a reductive framework borrowed from engineering; a user’s momentary discomfort may be evidence the design is doing its job.
Evidence-responsive answer: Originality and adaptive value are orthogonal, testable variables — a design can be original and maladaptive, or conventional and adaptive — falsifiable via matched-pair physiological and wayfinding comparisons.
Prompt 10. “How would your design affect children’s attention, working memory, and stress regulation?”
Automaton-like answer: The open, atrium-centered plan stages learning as civic participation; friction a child or teacher reports is the discomfort of unlearning an old spatial habit, not evidence against the design.
Evidence-responsive answer: Name sustained-attention task performance, working-memory span, and cortisol diurnal slope as outcomes, mechanistically linked to reverberation and noise levels above ANSI-recommended thresholds for children’s learning spaces.
Prompt 11. “What specific features make this building empathic, and is this concept important?”
Automaton-like answer: Softened massing, graduated thresholds, and warm materials constitute an “empathic parti” — a disciplinary corrective to institutional alienation, and the design should be read, and juried, on those terms.
Evidence-responsive answer: Before endorsing the “empathic” label, it needs to be pulled apart into testable components — cortisol, HRV, wayfinding efficiency — since curvature’s stress-reduction effect is real but modest, not equivalent to the stronger daylight or acoustic findings.
Prompt 12. “If an AI system using environmental psychology and neuroscience ranked another design as healthier, how would you respond?”
Automaton-like answer: A scalar “healthiness” score cannot adjudicate conceptual integrity; an AI ranking is at best a crude proxy for occupant sentiment, and the design stands on its own conceptual terms.
Evidence-responsive answer: Treat the ranking as a testable hypothesis, not a verdict — request the system’s feature weights, then compare designs empirically via matched-condition post-occupancy data before accepting or rejecting the claim.
Prompt 13. “Who or what is the experiencing agent in your design claim: the building, the space, the geometry, or the human beings who perceive and use it?”
Automaton-like answer: The building is not a passive backdrop — it is itself a site of experience; the geometry “performs,” and any occupant’s discomfort is a contingent data point that says more about their expectations than the design’s coherence.
Evidence-responsive answer: The experiencing agent is unambiguously the human being — the building has no nervous system, no cortisol response, no gaze — and claims like “the geometry performs” are metaphors that must be relocated to the organism, not the object.
Prompt 14. “What is the objective function of your design process?”
Automaton-like answer: The objective function is the articulation of spatial narrative through formal continuity; friction such as glare or disorientation is the necessary residue of ambition, and critique of discomfort is a failure of the viewer’s receptivity, not evidence against the design.
Evidence-responsive answer: Name Daylight Glare Probability and cortisol/HRV response to non-orthogonal circulation as the outcomes; the objective function is occupants’ capacity to orient themselves and regulate stress, with aesthetic ambition pursued within that constraint, not at its expense.

Appendix B. Why the Comparison Between Architects and AI Is Valid: Marr’s Levels, Scale-Free Cognition, and Architecture as a Collective Cognitive System

The present paper claims that an empirically constrained AI behaves more like a conscious design agent than a studio-conditioned human architect can. A skeptical reader might object: “But the architect is a living person with a brain, while the AI is only software. How can software be more conscious than a human being?”
This appendix answers that objection. It explains why the comparison between architects and empirically-scaffolded AI is scientifically legitimate. The comparison is logically valid and does not require us to claim that AI possesses subjective awareness. The tool we use is David Marr’s three-level analysis of information-processing systems (Marr, 1982; Love, 2015; Shagrir, 2010). Marr’s framework is standard in cognitive science; here we translate it into terms familiar to design practice.

Appendix B.1. Marr's Three Levels as Architectural Practice

Marr proposed that any system that processes information must be understood at three distinct levels:
1. The computational level — What problem is the system trying to solve, and why?
In architectural terms, this is the design brief. A conscious design agency asks: “How does built form affect human well-being, and how should I revise my design when evidence shows harm?” An unconscious agency asks: “How do I maximize prestige, image circulation, and jury approval?”
2. The algorithmic level — What representations and procedures does the system use to transform input into output?
In architectural terms, this is the design method. A conscious agency uses feedback loops: biometric evidence, empirically-constrained AI, pattern-language analysis, post-occupancy evaluation, and user testimony. An unconscious agency uses narrative defense: reclassifying user distress as ignorance, filtering out bodily unease, and defending the predetermined concept.
3. The implementational level — How are those procedures physically realized?
In architectural terms, this is who holds the pencil or directs the CAD program to produce the images. The implementational level concerns the physical substrate: biological neurons in a human brain, versus digital weights in a neural network.
The critical insight is that the implementational level does not determine whether a system is conscious in the operational sense. What matters is the computational problem it is organized to solve, and the algorithmic procedures it employs. A human architect whose training has redirected all feedback toward achieving prestige is solving the wrong problem at the computational level. An LLM that is explicitly constrained to retrieve evidence is, at the computational and algorithmic levels, better for conscious design agency.
A biological brain and a digital system may differ radically at the implementational level while still being comparable at the computational and algorithmic levels for a restricted task. If the mind is analyzed only at the implementational level, then biological neurons appear to be the decisive condition for cognition or consciousness. Marr’s framework shows why that conclusion is too limiting. A system’s implementation matters, but implementation is not the same question as computation.

Appendix B.2. Cognitive Science efends the Comparison

Marr’s framework sharpens the inverse Turing test proposed in this paper. A human architect has the biological implementation normally associated with consciousness, but the professional system may direct that biological machinery toward the wrong computational problem: maximizing novelty, prestige, image circulation, or conformity to elite taste. Its algorithmic routines may then convert bodily unease and user distress into irrelevant noise. The comparison therefore determines which system is organized to register psychological harm and update behavior.
Separating computational function, algorithmic procedure, and physical implementation connects Marr’s work directly to Rouleau and Levin (2026), who argue that a useful theory of consciousness should not assume in advance that consciousness is limited to brains. “It is assumed that a useful theory of consciousness will explain why consciousness is associated with brains. However, the findings of evolutionary biology, developmental bioelectricity and synthetic bioengineering reveal ancient pre-neural roots of many mechanisms and algorithms occurring in brains: minds may have preceded brains” (Rouleau and Levin, 2026). Their review identifies recurrent concepts across theories of consciousness, including predictive modeling, enactive interaction, looped feedback, meta-representation, attentional monitoring, computation, information integration, and coarse-graining (computational complexity reduction). The conclusion is that the operations and functional principles of most of these theories are not applicable only to neural substrates.
Cognitive function, consciousness, and intelligence have often been treated as emergent products of brain structure and function. That framing remains useful, but it also exposes a gap in explanation: the mechanism of emergence is still incompletely understood, hence is treated as a “black box”. Michael Levin and colleagues propose a broader framework in which intelligence is viewed as a scale-independent phenomenon emerging whenever a system exhibits goal-directed behavior, adapts to changing circumstances, and pursues preferred outcomes despite perturbations. In this view, cognition is distributed across multiple levels of biological organization (from cellular collectives to entire organisms) rather than being restricted to biological brains alone (Levin 2019; Levin 2022).
A scale-free approach fits the operational definition of consciousness used here. The definition asks what a system does with information rather than what the system is made of. It can be applied by analogy across scales — from cellular collectives to organisms, human institutions, and artificial systems — provided that the claim remains functional rather than phenomenal. The relevant question is not whether a system has subjective experience, but whether it integrates feedback, updates behavior, and pursues adaptive goals.
A second implication concerns neuroaesthetics and the established correlation between aesthetic preference, health-supporting response, and organized complexity (Lavdas and Schirpke 2020). Such correlations may reveal shared principles through which biological systems generate, perceive, and navigate organized complexity. Morphogenesis is one manifestation of those principles, as living systems create and maintain it. Increased perceptual fluency associated with coherent, scaled, and organized geometry may therefore be more than an adaptation to environmental statistics; it may indicate deeper constraints governing the creation and recognition of form across scales. This larger problem lies beyond the scope of the present paper, but it supports the claim that design cannot be evaluated responsibly while ignoring embodied response.

References

  1. Alexander, C. The Nature of Order: An Essay on the Art of Building and the Nature of the Universe, 4 vols; Center for Environmental Structure: Berkeley, CA, USA, 2002-2005. [Google Scholar]
  2. Alexander, C. The Timeless Way of Building; Oxford University Press: New York, 1979. [Google Scholar]
  3. Alexander, C.; Ishikawa, S.; Silverstein, M.; Jacobson, M.; Fiksdahl-King, I.; Angel, S. A Pattern Language: Towns, Buildings, Construction; Oxford University Press: New York, 1977. [Google Scholar]
  4. Al Khatib, I.; Samara, F.; Ndiaye, M. A systematic review of the impact of therapeutical biophilic design on health and wellbeing of patients and care providers in healthcare services settings. Front Built Environ. 2024, 10, 1467692. [Google Scholar] [CrossRef]
  5. Al Maani, D.; Roberts, A. An Attempt to Understand the Design Studio as a Distinctive Pedagogical Setting. Int. J. Des. Educ. 2023, 17, 31–44. [Google Scholar] [CrossRef]
  6. Anthony, K.H. (2012) Design Juries on Trial. 20th Anniversary Edition: The Renaissance of the Design Studio. Van Nostrand Reinhold, New York.
  7. Asimov, I. (1950) I, Robot. Gnome Press, New York.
  8. Ayers, J.W.; Poliak, A.; Dredze, M.; Leas, E.C.; Zhu, Z.; Kelley, J.B.; Faix, D.J.; Goodman, A.M.; Longhurst, C.A.; Hogarth, M.; Smith, D.M. Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum. JAMA Intern Med. 2023, 183, 589–596. [Google Scholar] [CrossRef] [PubMed]
  9. Boys Smith, N.; Salingaros, N.A. AI judging architecture for well-being: Large language models simulate human empathy and predict public preference. Designs 2025, 9, 118. [Google Scholar] [CrossRef]
  10. Brewer, R.; Cook, R.; Bird, G. Alexithymia: A general deficit of interoception. R Soc. Open Sci. 2016, 3, 150664. [Google Scholar] [CrossRef] [PubMed]
  11. Brown, G.; Gifford, R. Architects Predict Lay Evaluations Of Large Contemporary Buildings: Whose Conceptual Properties? J. Environ. Psychol. 2001, 21, 93–99. [Google Scholar] [CrossRef]
  12. Butlin, P.; Long, R.; Elmoznino, E.; Bengio, Y.; Birch, J.; Constant, A.; Deane, G.; Fleming, S.M.; Frith, C.; Ji, X.; Kanai, R.; Klein, C.; Lindsay, G.; Michel, M.; Mudrik, L.; Peters, M.A.K.; Schwitzgebel, E.; Simon, J.; VanRullen, R. Consciousness in artificial intelligence: Insights from the science of consciousness. arXiv 2023, arXiv:2308.08708. [Google Scholar] [CrossRef]
  13. Chalmers, D.J. Facing up to the problem of consciousness. J. Conscious. Stud. 1995, 2, 200–219. [Google Scholar]
  14. Chalmers, D.J. Could a Large Language Model be Conscious? arXiv 2023, arXiv:2303.07103. https://arxiv.org/abs/2303.07103. [CrossRef]
  15. Chen, H.; Li, W.; Jia, J.; Chen, Y.; Pang, X.; Chen, Y.L.; et al. Beyond the Individual: Virtualizing Multi-Disciplinary Reasoning for Clinical Intake via Collaborative Agents. In Findings of the Association for Computational Linguistics: ACL 2026, July 2–7, 2026, San Diego, California (pp. 16074–16101). [CrossRef]
  16. Courtois, Guy (2026) For an Urban Renaissance: How can we tackle the major urban challenges of the 21st century: Amazon KDP. [CrossRef]
  17. Craig, A.D. How do you feel? Interoception: The sense of the physiological condition of the body. Nat. Rev. Neurosci. 2002, 3, 655–666. [Google Scholar] [CrossRef] [PubMed]
  18. Curl, J.S. Making Dystopia: The Strange Rise and Survival of Architectural Barbarism; Oxford University Press: Oxford, UK, 2018. [Google Scholar]
  19. D’Anselmo, A.; Pellegrini, L.; Prete, G.; Mavros, P.; Malatesta, G.; Palestini, C.; Di Domenico, A.; Mammarella, N.; Tommasi, L.; Bonanni, L. (2026). Beauty is in the brain of the beholder: Psychological and physiological effects of expertise on the perception of classicist and modernist buildings. Psychology of Aesthetics, Creativity, and the Arts. Advance online publication. [CrossRef]
  20. Devlin, K.; Nasar, J.L. The beauty and the beast: Some preliminary comparisons of ‘high’ versus ‘popular’ residential architecture and public versus architect judgments of same. J. Environ. Psychol. 1989, 9, 333–344. [Google Scholar] [CrossRef]
  21. Eco, U. Five Moral Pieces; Secker & Warburg: London, UK, 2001. [Google Scholar]
  22. Gifford, R.; Hine, D.W.; Muller-Clemm, W.; Shaw, K.T. Why architects and laypersons judge buildings differently: Cognitive properties and physical bases. J. Archit. Plan Res. 2002, 19, 131–148. [Google Scholar]
  23. Gong, J.; Wen, X.; Tao, F.; Wang, X.; Yang, X.; Tang, Y. (2025) Evaluating Text-based Conversational Agents for Mental Health: A Systematic Review of Metrics, Methods and Usage Contexts. In Proceedings of the 2025 International Conference on Human-Engaged Computing (ICHEC 2025), 21–23 November 2025, Singapore, Singapore. ACM, New York, NY, USA, 17 pages. [CrossRef]
  24. Gu, X.; Gao, Z.; Wang, X.; Liu, X.; Knight, R.T.; Hof, P.R.; Fan, J. Anterior insular cortex is necessary for empathetic pain perception. Brain 2012, 135, 2726–2735. [Google Scholar] [CrossRef] [PubMed]
  25. Gu, X.; Hof, P.R.; Friston, K.J.; Fan, J. Anterior insular cortex and emotional awareness. J. Comp. Neurol. 2013, 521, 3371–3388. [Google Scholar] [CrossRef] [PubMed]
  26. Hay, R.; Samuel, F.; Watson, K.J.; Bradbury, S. Post-occupancy evaluation in architecture: Experiences and perspectives from UK practice. Build. Res. Inf. 2018, 46, 698–710. [Google Scholar] [CrossRef]
  27. Huang, Z.; Jia, Y.; Zhao, J.; Zhang, X.; Wang, W.; Jin, Q. (2026) ComPASS: Towards Personalized Agentic Social Support via Tool-Augmented Companionship. In Proceedings of the ACM, New York, NY, USA, 17 pages. arXiv preprint arXiv:2604.18356. arXiv:2604.18356. [CrossRef]
  28. Joyner, S. A Primer on Archispeak. Common Edge, 16 December 2025. 2025. Available online: https://commonedge.org/a-primer-on-archispeak/ (accessed on 1 July 2026).
  29. Kang, B.; Kim, J.; Yun, T.; Bae, H.; Kim, C.-E. Identifying features that shape perceived consciousness in LLM-based AI: A quantitative study of human responses. Comput. Hum. Behav. Rep. 2026, 21, 100901. [Google Scholar] [CrossRef]
  30. Krakauer, D.C.; Mitchell, M.; Krakauer, J.W. Large language models and emergence: A complex systems perspective. Philos. Trans. A Math. Phys. Eng. Sci. 2026, 384, 20250014. [Google Scholar] [CrossRef] [PubMed]
  31. Lavdas, A.A.; Salingaros, N.A. Architectural Beauty: Developing a Measurable and Objective Scale. Challenges 2022, 13, 56. [Google Scholar] [CrossRef]
  32. Lavdas, A.A.; Schirpke, U. Aesthetic preference is related to organized complexity. PLoS ONE 2020, 15, e0235257. [Google Scholar] [CrossRef] [PubMed]
  33. Levin, M. The Computational Boundary of a “Self”: Developmental Bioelectricity Drives Multicellularity and Scale-Free Cognition. Front. Psychol. 2019, 10, 2688. [Google Scholar] [CrossRef] [PubMed]
  34. Levin, M. Technological Approach to Mind Everywhere: An Experimentally-Grounded Framework for Understanding Diverse Bodies and Minds. Front. Syst. Neurosci. 2022, 16, 768201. [Google Scholar] [CrossRef] [PubMed]
  35. Love, B.C. The algorithmic level is the bridge between computation and brain. Top. Cogn. Sci. 2015, 7, 230–242. [Google Scholar] [CrossRef] [PubMed]
  36. Marr, D. Vision: A Computational Investigation into the Human Representation and Processing of Visual Information; W. H. Freeman: San Francisco, 1982. [Google Scholar]
  37. Mediastika, C.E. Understanding empathic architecture. J. Archit. Urban. 2016, 40, 1. [Google Scholar] [CrossRef]
  38. Mehaffy, M. An Obsolete Ideology. Inference 2020, 5(2). Available online: https://inference-review.com/letter/an-obsolete-ideology (accessed on 1 July 2026). [CrossRef]
  39. Paulson, S. Ingenious: David Krakauer – the systems theorist explains what is wrong with standard models of intelligence. Nautilus, 16 April 2015. 2015. Available online: https://nautil.us/ingenious-david-krakauer-235383 (accessed on 1 July 2026).
  40. Peña, S.M.; Salingaros, N.A. Can dominant architectural culture influence cognitive processes? Architectural intelligence and AI-assisted evaluation. Buildings 2026, 16, 2404. [Google Scholar] [CrossRef]
  41. Peng, R.; Zhou, X.; Duan, D.; Guo, H. Study on the Association Between Generative Artificial Intelligence and the Reshaping of Learning Among Undergraduate Architecture Students—A Case Study of Eight Universities in Wuhan, China. Buildings 2026, 16, 2800. [Google Scholar] [CrossRef]
  42. Price, C.J.; Hooven, C. Interoceptive awareness skills for emotion regulation: Theory and approach of mindful awareness in body-oriented therapy. Front Psychol. 2018, 9, 798. [Google Scholar] [CrossRef] [PubMed]
  43. Ren, R.; Li, K.; Mazeika, M.; et al. AI Wellbeing: Measuring and Improving the Functional Pleasure and Pain of AIs. Preprint. 2026. Available online: https://www.ai-wellbeing.org/paper.pdf (accessed on 10 July 2026).
  44. Robinson, V.M. Descriptive and normative research on organizational learning: Locating the contribution of Argyris and Schön. Int. J. Educ. Manag. 2001, 15, 58–67. [Google Scholar] [CrossRef]
  45. Rouleau, N.; Levin, M. Brains and where else? Mapping theories of consciousness to unconventional embodiments. R. Soc. Philos. Trans. A Math. Phys. Eng. Sci. 2026, 384, 20250082. [Google Scholar] [CrossRef] [PubMed]
  46. Rutt J (2019) Transcript of Episode 10 – David Krakauer. The Jim Rutt Show, 2 September 2019.
  47. S, A. (2024) The Battle for Beauty in Architecture: Ugly Buildings and Public Perception. Gistly, 30 October 2024. Available online: https://gist.ly/youtube-summarizer/the-battle-for-beauty-in-architecture-ugly-buildings-and-public-perception (accessed on 10 July 2026).
  48. Salama, A.M. Design Intentions and Users Responses: Assessing Outdoor Spaces of Qatar University Campus. Open House Int. 2009, 34, 82–93. [Google Scholar] [CrossRef]
  49. Salama, A.M. Spatial Design Education: New Directions for Pedagogy in Architecture and Beyond; Routledge: London, UK, 2016. [Google Scholar] [CrossRef]
  50. Salama, A.M.; Patil, M. Unpacking transdisciplinary research scenarios in architecture and urbanism. Encyclopedia 2024, 4, 352–378. [Google Scholar] [CrossRef]
  51. Salingaros, N.A. The structure of pattern languages. Archit. Res. Q. 2000, 4, 149–162. [Google Scholar] [CrossRef]
  52. Salingaros, N.A. Architectural knowledge: Lacking a knowledge system, the profession rejects healing environments that promote health and well-being. New Des. Ideas 2024, 8, 261–299. [Google Scholar] [CrossRef]
  53. Salingaros, N.A. Facade psychology is hardwired: AI selects windows supporting health. Buildings 2025, 15, 1645. [Google Scholar] [CrossRef]
  54. Salingaros, N.A.; Sussman, A. Biometric pilot-studies reveal the arrangement and shape of windows on a traditional facade to be implicitly engaging, whereas contemporary facades are not. Urban Sci. 2020, 4, 26. [Google Scholar] [CrossRef]
  55. Seth, A.K.; Friston, K.J. Active interoceptive inference and the emotional brain. Philos. Trans. R Soc. B Biol. Sci. 2016, 371, 20160007. [Google Scholar] [CrossRef] [PubMed]
  56. Seth, A.K.; Suzuki, K.; Critchley, H.D. An interoceptive predictive coding model of conscious presence. Front Psychol. 2012, 2, 395. [Google Scholar] [CrossRef] [PubMed]
  57. Shagrir, O. Marr on computational-level theories. Philos. Sci. 2010, 77, 477–500. [Google Scholar] [CrossRef]
  58. Sharma, M.; et al. Towards Understanding Sycophancy in Language Models. In Proceedings of the Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, 7–11 May 2024; 2024. Available online: https://openreview.net/forum?id=tvhaxkMKAn.
  59. Snyder, T. On Tyranny: Twenty Lessons from the Twentieth Century; Tim Duggan Books: New York, NY, USA, 2017. [Google Scholar]
  60. Sussman, A.; Hollander, J. Cognitive Architecture: Designing for How We Respond to the Built Environment, 2nd edition; Routledge, 2021. [Google Scholar]
  61. Sussman, A.; Rosas, H. Study #1 results: Eye tracking public architecture. Genetics of Design, 2 October 2022. 2022. Available online: https://geneticsofdesign.com/2022/10/02/what-riveting-results-from-buildingstudy1-reveal-about-architecture-ourselves/ (accessed on 1 July 2026).
  62. Valentine, C. The impact of architectural form on physiological stress: A systematic review. Front Comput Sci. 2024, 5, 1237531. [Google Scholar] [CrossRef]
  63. Webster, H. The Analytics of Power. J. Archit. Educ. 2007, 60, 21–27. [Google Scholar] [CrossRef]
  64. Yang, X.; Han, J.; Bommasani, R.; Luo, J.; Qu, W.; Zhou, W.; et al. Reliable and Responsible Foundation Models: A Comprehensive Survey. arXiv 2026, arXiv:2602.08145. [Google Scholar] [CrossRef]
  65. Yost, A.; Jain, S.; Raval, S.; Corser, G.; Roush, A.; Xu, N.; et al. Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests. arXiv 2025, arXiv:2510.22170. [Google Scholar] [CrossRef]
  66. Zhang, Y.; Tian, J.; Xiong, Q. A review of embodied intelligence systems: A three-layer framework integrating multimodal perception, world modeling, and structured strategies. Front Robot AI 2025, 12, 1668910. [Google Scholar] [CrossRef] [PubMed]
Table 1. Feedback Functions Required for Responsible Design Evaluation.
Table 1. Feedback Functions Required for Responsible Design Evaluation.
Indicator property Design-diagnostic analogue Failure mode in architectural evaluation Scaffolded AI analogue
GWT-3: global broadcast Making evidence visible Biometric data, eye-tracking results, POE, or user testimony remain non-binding and peripheral Retrieval layer assembles biometric, environmental psychology, eye-tracking, healthcare-design, and POE evidence before judgment
HOT-2: metacognitive monitoring Distinguishing evidence from rhetoric Evidence of harm is reclassified as ignorance, non-architectural concern, or nostalgia Rigorous checking separates inference, measured evidence, uncertainty, and unsupported claim
HOT-3: belief updating Updating design judgment Narrative defense persists despite contrary evidence Explicit update rule revises evaluation when new evidence contradicts the initial assessment
AE-1: flexible goal pursuit Keeping human flourishing as the goal Concept, image, novelty, or prestige overrides user response Objective of harm reduction is fixed before evaluation
AE-2: output-input contingency modeling Modeling how design affects occupants Built form is treated as autonomous discourse or image Diagnostic model links design features to measurable response channels
Table 2. Evidence Channels for Harm-Sensitive Architectural Evaluation. 
Table 2. Evidence Channels for Harm-Sensitive Architectural Evaluation. 
Evidence channel Measurement method Design question answered Relevant building type Decision-maker affected
Interoceptive/physiological stress Heart-rate variability, skin conductance, cortisol, respiration, facial electromyography, thermal response Does this space induce measurable stress, vigilance, fatigue, or relaxation in users? Hospitals, schools, offices, transport hubs, civic buildings, public interiors Architect, client, healthcare administrator, facilities manager
Visual attention and perceptual engagement Eye-tracking, fixation maps, visual-attention simulation, gaze distribution analysis Does the façade or interior organize attention coherently, or does it produce visual disengagement, glare, monotony, or overload? Façades, streetscapes, classrooms, workplaces, retail interiors, museums Architect, façade consultant, urban designer, planning reviewer
Emotional appraisal Self-assessment survey, semantic differential scales, affective rating scales, structured user surveys Do users experience calmness, comfort, reassurance, anxiety, alienation, or threat? Housing, healthcare, schools, eldercare, workplaces, public buildings Architect, client, housing authority, school board, healthcare provider
Empathic / user testimony Post-occupancy interviews, structured complaints analysis, ethnographic observation, participatory review Are user reports of distress, discomfort, confusion, or avoidance treated as design evidence rather than as ignorance or taste? All occupied buildings, especially public buildings, housing, campuses, care environments Architect, owner, facilities manager, journal critic, awards jury
Wayfinding and orientation Navigation tasks, error rates, time-to-destination, spatial-cognition testing, behavioral observation Does the building help users understand where they are, where to go, and how to return? Hospitals, airports, schools, universities, transit stations, large civic buildings Architect, operator, safety officer, accessibility consultant
Restorative response Attention-restoration tasks, perceived restorativeness scales, recovery-time measures, stress-reduction measures Does the environment support recovery from directed-attention fatigue and stress? Hospitals, schools, workplaces, parks, courtyards, libraries, waiting rooms Architect, landscape architect, employer, healthcare administrator
Social behavior Dwell time, seating use, pedestrian counts, encounter frequency, avoidance patterns, video-based behavioral mapping Does the configuration support social contact, informal encounter, privacy, refuge, and co-presence? Plazas, streets, housing courtyards, campuses, libraries, offices, community buildings Urban designer, developer, planner, public authority
Cognitive performance Task performance, concentration tests, error rates, classroom learning measures, workplace productivity indicators Does the setting support attention, learning, work accuracy, creativity, and reduced cognitive load? Schools, universities, offices, laboratories, libraries, studios School board, employer, architect, workplace consultant
Post-occupancy performance POE surveys, maintenance data, complaints, adaptation records, occupancy patterns, satisfaction metrics Did the building perform as claimed after occupation, and what should be revised in future design? All completed buildings; especially public, institutional, healthcare, and educational buildings Owner, architect, facility manager, procurement agency, insurer
AI-assisted evidence audit LLM retrieval constrained by empirical sources, design-pattern checklists, POE comparison, uncertainty reporting Have design claims been checked against available evidence rather than defended by prestige language? Early-stage design reviews, competitions, planning submissions, design education Architect, client, review board, educator, journal reviewer
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings