Submitted:
05 August 2026
Posted:
05 August 2026
You are already at the latest version
Abstract
Interactive art is a form wherein aesthetic content emerges through audience participation within technology-based installations designed for interactive engagement. Recent developments in AI have increasingly shaped this field, influencing audience participation, the dynamic interplay between installations and algorithms, and the formation of aesthetic experience. However, relatively little research has examined how AI affects these design dimensions in interactive art. To address this gap, this study proposes the 4A framework, an analytical framework comprising four core design elements: audience participation, art object embodiment, AI inquiry and response, and aesthetic experience. Rather than advancing a wholly new theory of interaction, the framework extends earlier models by introducing analytical dimensions needed to examine the technical and experiential conditions specific to AI-driven artworks. Applied to case studies of recent AI-driven interactive artworks presented at CVPR 2024, it offers design-oriented insights for the analysis and creation of AI-based interactive art installations.
Keywords:
artificial intelligence
; interactive art
; interaction design
; audience participation
; aesthetic experience
1. Introduction
1.1. Definition and Aesthetic Effects of Interactive Artworks
Interactive art is a genre wherein artistic content is generated in real time through the direct participation of an audience, mediated by algorithms and digital technologies within an interactive installation [1,2,3]. Because audience involvement is foundational, these works are characterized as participatory art that continuously evolves. Unlike traditional artworks, which primarily emphasize appreciation and interpretation, contemporary interactive art elevates real-time participation to a central element in the creation of artistic content and subsequent aesthetic experience. This shift signifies the emergence of a new form of aesthetic value.
Aesthetic interaction instantiates the human effort to find meaning when encountering art; a viewer’s interpretive ventures are discursive accomplishments that integrate the aesthetic dimension of life into the broader human experience [4].
Interactive art installations have continually evolved through the application of digital technologies. Consequently, prior research has explored the design and function of installations that enable interactivity [5,6,7,8,9]. By combining various technologies, these interactive installations can reflect audience actions and, in doing so, generate novel aesthetic experiences.
The implementation of interactive art has spurred continued research into new technological applications, the nature of interaction, and the temporal and spatial dimensions of content experience [10,11,12,13,14].
A fundamental relationship between the audience and the installation has been established in this field since the work of Fels [15]. This foundational structure has since evolved into increasingly complex and participatory forms, as documented by Edmonds [16]. The application of diverse technologies has enabled new participation modes and facilitated the integration of different content types within the structure of interactive art. Consequently, audience participation, technological mediation, and artistic representation have emerged as central research themes [5,6,7,8,9,10,11,12,13,14,15,16,17].
Particularly, recent advancements in artificial intelligence (AI) technologies have introduced significant changes and expanded the possibilities for production techniques, modes of audience participation, and the nature of the aesthetic experience in interactive art. From the perspective of installation, integrating various AI algorithms has enabled new forms of engagement that generate aesthetic experiences distinct from those possible with traditional interactive art.
- Integrating AI into artistic practice
AI technologies, propelled by an explosive increase in data, enhanced computing power, and breakthrough model architectures, are virtually influencing every aspect of modern life. At its core, AI aims to develop intelligent agents (systems) capable of reasoning, evolving, and functioning autonomously through learning. This integration of AI has catalyzed new forms of artistic transformation in interactive artworks [18,19,20,21,22,23,24], and academic research is actively investigating the influence of AI in the arts, including its implications for ethics and creativity [25,26].
Even before the advent of generative AI, machine learning (ML) methods—including supervised learning, unsupervised learning, and reinforcement learning [27]—were employed in the field of arts. These methods were particularly used to produce personalized content, generate creative output, and support artistic ideation [28]. The socio-technical conditions established for producing art-based ML prototypes have paved the way for an interactive understanding of ML, one that embraces the differences between design and engineering practices through diffractive practice [29].
Interactive ML has also been used to design and build new gestural controls, allowing artists and dancers to improvisationally define embodied interactions and evaluate models in real time by experimenting with controllers [30]. The application of deep learning in art is enhancing artistic creativity, influencing both ideation and visualization by enabling new modes of expression [31,32].
More recently, the emergence of generative AI—driven by models such as variational autoencoders (VAEs), generative adversarial networks, diffusion models, and large language models (LLMs) with transformers—has begun to challenge traditional notions of human creativity. Generative AI has opened new opportunities for audience engagement through image generation, transforming how audiences create and experience visual content [33].
Furthermore, co-creation between humans and AI is provoking ongoing discourse regarding the evolving roles of human creativity and technological agency [26,34,35]. As the scope of AI-driven art production and exhibition expands, it prompts sustained academic exploration of its aesthetic significance [36,37]. Consequently, transformative changes are occurring across the entire artistic process, from data collection and AI system design to technological procedures, aesthetic representation, and modes of audience participation.
- b. Study aims
This study aims to advance the analysis of AI-driven interactive art by examining recent interactive installations that incorporate AI. Although prior research has established important frameworks for traditional interactive art, relatively few studies have addressed how AI specifically reshapes interaction design, system behavior, and aesthetic experience in interactive installations. This gap persists owing to the limited capacity of traditional interactive art frameworks to address AI-specific design implications and the lack of established methodologies for examining how AI integration influences audience participation and aesthetic experience.
To address this gap, this study proposes an analytical framework for AI-integrated interactive art, grounded in a review of existing models of traditional interactive art. Rather than advancing a wholly new theory of interaction, the framework extends earlier models by introducing analytical dimensions needed to examine the technical and experiential conditions specific to AI-driven artworks. This framework is then applied to a series of case studies of recent AI-based artworks. Through this analysis, the study examines specific discussion points related to each artwork, their experiential value for audiences, and possible system design approaches to producing AI-based interactive installations. The overall research process is as follows:
- Conduct a literature review of traditional interactive art installations and AI-integrated interactive installations.
- Develop an analytical framework for examining AI-based interactive art installations based on the findings of the literature review.
- Conduct case studies of AI-integrated interactive art installations using the proposed framework.
- Derive design approaches and implications from the discussion of the case-study findings.
1.4. Scope and Limitations
The scope of this study is as follows:
1) As it was not feasible within the scope of this study to collect primary empirical data, such as audience interviews, surveys, or behavioral measurements, the analysis is based on the official descriptions of the artworks, related publications, and publicly available video documentation of the installations. Accordingly, this study adopts an observation-centered analytical approach rather than conventional human–computer interaction (HCI) methods focused on technological measurement or usability evaluation. The findings should therefore be understood as interpretive and framework-oriented rather than empirically validated accounts of audience experience.
2) Ethical issues related to the use of AI, including consent, privacy, and bias, are highly significant in contemporary AI art. While these concerns are acknowledged and briefly discussed in this study, they are not examined as a primary analytical focus.
3) Although theoretical questions concerning artist agency, authorship, and aesthetic value are central to ongoing discussions on AI-based interactive art, they are not the main focus of this study. The study is intended to advance an analytical frame-work rather than to propose new theories regarding these issues.
2. Literature Review
To analyze AI-integrated interactive art installations, a literature review was conducted across two primary domains: interactive art and AI-based interactive art installations. This review aimed to examine the fundamental structure of interactive art installations, identify and analyze the distinctive features of AI-integrated works, and explore the potential applications of various AI technologies within artistic contexts.
2.1. Research on interactive Art Installations
Over the past decade, interactive art installations have become a significant research topic at the intersection of art, technology, and HCI. Such installations employ various technologies to capture and respond to audience activities, generating new aesthetic experiences through dynamic interaction.
Saltz [37] conceptualized a process in which a sensing or input device translates aspects of a person’s behavior into a digital form understood by a computer; the computer then outputs data that are systematically related to the input, which are then translated back into real-world phenomena. Graham [38] explored the development of interactive artworks enhanced by collective usage, applying a “conversation/host” to the creation of the artwork; Graham also conducted further case studies and developed a taxonomy to illustrate differences in artwork–audience and audience–audience relationships. Fels [15] discussed three artworks—Iamascope, Video Cubism, and The Forklift Ballet—that combine technology and art to illustrate issues of intimacy and embodiment. Edmonds et al. [39] discussed the role of interaction in art systems and new methods for building them, while Edmonds [16] later developed a framework defining audience participation through sequential phases: adaptation, learning, anticipation, and deeper understanding. Xiaobo and Yuelin [17] investigated a basic interaction framework integrating time, space, and information within interactive systems to enable varied artistic experiences. Ahmed [5] analyzed how interaction and interactivity are defined in different fields, examining their relevance and applicability in the context of digital interactive art.
These interactive art installations incorporated various technologies to reflect audience actions, creating new aesthetic experiences through dynamic interactions.
2.2. Research on AI-Based Interactive Art Installations
AI has been increasingly employed in the field of arts, functioning both as a computational instrument and an autonomous system capable of producing aesthetic outcomes. AI technologies have also been incorporated into installation works that create interactive and adaptive experiences. Advances in AI have further facilitated interactions between humans and intelligent systems, establishing new paradigms for creative collaboration and audience engagement.
The application of AI in this context can be categorized into broad tiers, including AI, ML, deep learning (DL), and generative AI approaches. For instance, Chen et al. [40] built a new interactive art creation system with AI as the core medium, using a methodological approach of “cognition of human–computer symbiosis, innovation supported by intelligent technology, collaboration among creative subjects, and constraints on creative behavior,” where artists, robots, and the audience are co-creation subjects. Xie et al. [41] examined the integration of VAEs and particle swarm optimization in generative and interactive digital installations, revealing that this hybrid model enhances user engagement and aesthetic satisfaction. Canet Sola and Guljajeva [42] proposed the Dream Painter, an interactive robotic art installation that turns spoken dreams into a collective painting. They explored the interactive potential of AI and robotics, sparking a broader discussion on DL applications. In their other work, Visions of Destruction [43], they used artistic research methods to communicate a focus on the Anthropocene through audience interaction and generative AI [18]. Sun et al. [22] presented AI Nüshu, an emerging language system inspired by Nüshu, a unique language used exclusively by ancient Chinese women. Their work highlighted the role of AI in preserving cultural heritage and redefining human–machine dynamics in a language-making interactive artwork. Zhang et al. [24] described the conceptual background, AI system design, and visualization strategies of ReCollection, an interactive AI art installation. The artwork assembles collective visual memories from the language input of participants, blurring the boundaries between remembrance and imagination.
These works have employed generative AI technologies to create artistic content such as images, sound, and video, thereby offering new functionalities within interactive installations and expanding the aesthetic experiences of the audience.
While research on interactive art installations has received considerable attention, studies that focus specifically on the underlying frameworks of AI-based interactive artworks to better understand the impact of AI technology on interaction design and aesthetic effects remain limited.
Moreover, prior research has largely focused on the technical implementation of AI-based artworks, rather than offering a comprehensive analytical perspective. To address these gaps, this study proposes a new analytical framework and case-based approach to examine the artistic and conceptual implications of AI-driven interactive installations.
3. Research Framework
This section reviews existing research frameworks for interactive art and presents a new framework for analyzing interactive installation art that incorporates AI technologies.
- Research framework development
Previous studies have proposed fundamental models of interaction between humans and objects. Fels [15] categorized four types of relationships based on the extent to which an object is embodied within a person or the person within the object:
- The person engages in dialogue with the object,
- the person embodies the object,
- the object communicates with the person, and
- the object embodies the person.
This foundational model provided a methodology for categorizing communication between audience, object, and installation.
Edmonds et al. [39] delineated the relational structure between artwork, artist, viewer, and environment in interactive art using four categories: static, dynamic-passive, dynamic-interactive, and dynamic-interactive (varying). These categories reflect the extent and nature of the interaction achieved.
Xiaobo and Yuelin [17] structured their “Interaction Aesthetics Framework” into three components: interaction (classified into time, information, and space); interactive system (classified into function, form, and structure); and aesthetics (classified into visual, auditory, and tactile domains). This framework allowed for distinguishing audience engagement characteristics and systematically articulating aesthetic experiences shaped by the audience.
Ahmed [5] focused on the communicative relationship between humans and media within the context of interactive art, proposing a five-stage classification model:
- Human–Human Interaction,
- Human–Media Interaction,
- Device-Mediated Communication,
- Human–Machine Interaction, and
- Triggers for Change.
This model enabled a nuanced understanding of the interrelationships between audience, installation, environment, and aesthetics.
Guljajeva and Canet Sola [42] introduced a triangular framework of interactive art comprising author, artwork, and audience. In this model, feedback between author and artwork is generated through code and variables, whereas feedback between author and audience emerges through parameters, system stress, and user interaction. They proposed that the aesthetics of interactive artworks are shaped by the reciprocal exchange of input and output between audience and artwork. Fundamentally, their study conceptualized interactive art as a structure wherein installations—including digital devices designed in alignment with the intent of the artist—generate artistic content through audience interaction, which is then realized and experienced through participation.
The fundamental models discussed above are summarized and compared in Table 1. Within these models, AI may be understood either as part of the art object, that is, as a component of the installation system, or as a software-based agent embedded within it. However, this perspective becomes insufficient for AI-driven interactive artworks, where AI is not merely embedded in the installation system but actively mediates interaction, generation, and interpretation. The incorporation of AI has expanded interaction modalities, introduced new thematic and conceptual concerns, and enabled artworks to operate with partial autonomy while forming dynamic relationships with viewers [44] Therefore, AI should not be reduced to a generative algorithm or technical subsystem within the art object. It must be treated as a distinct analytical entity, since its inferential processes, generative uncertainty, and relative autonomy exert a direct influence on how interaction unfolds and how aesthetic experience is produced.
To address this gap, we propose a 4A framework for the analysis of AI-based interactive art. The framework consists of three key entities: audience, art object, and AI, and four key design elements: participation, embodiment, AI inquiry and response, and aesthetic experience (see Figure 1). Its primary contribution lies in treating AI as an independent element within interactive art, thereby extending the shared structure of earlier models to account for the specific ways in which AI reshapes interaction processes and influences the aesthetic experience emerging through audience engagement. Table 2 presents the core components of the 4A framework, including its four design elements and the corresponding analytical elements associated with each element.
3.2. Elements of the 4A Framework
Each element within the proposed 4A framework is described as follows:
- Audience Participation
Interaction between the audience and the artwork is initiated and mediated by the input data provided by participants. This framework analyzes the structural characteristics of input data (format) and sensory channels (modality) to develop a deeper understanding of the interaction process. This also includes choices that may influence the overall shaping process of the artwork system.
- 2.
- Art object embodiment
The framework investigates how outputs are generated and presented within the installation system. Similar to the analysis of inputs, this involves a detailed examination of the format and modality of output data, including sensory channels (e.g., visual, auditory) and structural forms through which the results are materialized or experienced.
- 3.
- AI inquiry and response
a) AI Technologies: The framework identifies the specific AI models, algorithms, and tools integrated into the interactive installation. It distinguishes the functional types of AI technologies (e.g., image generation, object recognition, speech processing) and, where applicable, subtypes (e.g., image-to-image or text-to-image generation). These can be realized via publicly available pretrained models (e.g., Stable Diffusion [45], BLIP-2 [46]) or artist-developed, fine-tuned models designed to meet particular conceptual and aesthetic intentions.
b) Data Processing: The framework examines how audience input data are preprocessed before being fed to the AI model. It also investigates whether supplementary data are incorporated and whether the output of the AI model is used directly to materialize the artwork or requires additional post-processing. Furthermore, the framework considers whether accumulated data from previous interactions or environmental sensing are integrated with new inputs during the ongoing interaction.
c) Temporality: The framework addresses the temporal characteristics of interaction within AI-driven installations. While temporality is also shaped by the broader processes of audience engagement and artwork manifestation, the framework treats AI as the primary factor and therefore locates temporality within AI inquiry and response rather than the art object itself. This also includes the time required to process the audience input and produce output, and the temporal design of interaction modes and pacing of the manifestation of the artwork. Interaction processes are categorized as procedural, delayed, near-immediate, and real-time, depending on system responsiveness and the continuity of user engagement.
- 4.
- Aesthetic Experience
Aesthetic experience is inherently subjective. Especially in interactive art, the diverse participation and choices of the audience lead to a broad range of possible outcomes, making the experience significantly more variable than that in traditional artworks. The audience experience is substantially shaped by sensory stimulation, perceptual interpretation, and cognitive meaning-making. As Savaş et al. [47] noted, aesthetic experience engages both “cognitive and affective processes,” indicating that thinking and feeling intertwine during audience interaction with art.
Based on previous studies, aesthetic experience in interactive art installations can be understood as a layered process involving sensory, perceptual, and cognitive dimensions [47,48,49,50,51,52]. Sensory experience refers to immediate embodied responses to multimodal stimuli such as visual, auditory, and tactile inputs generated through interaction with the interactive artwork. Perceptual experience involves the organization and recognition of these sensory signals, including the integration of multiple modalities and the awareness of one’s interaction with the system. Cognitive experience extends this process through reflective interpretation, as audience members construct meaning by relating perceived stimuli to prior knowledge, personal memories, and cultural contexts [51,52,53,54,55]. Together, these dimensions describe how aesthetic experience unfolds from immediate sensation to perceptual organization and ultimately to conceptual interpretation during interaction with interactive artworks.
Within AI-integrated installations, these experiential layers are further shaped by technological mediation. Unlike traditional interactive systems, AI technologies often generate responses through probabilistic processes, introducing elements of uncertainty and indeterminacy in the activity of the artwork [25,44,56]. This generative uncertainty can influence aesthetic experience by producing unexpected outputs that evoke responses such as surprise, curiosity, or interpretive ambiguity.
Furthermore, AI outputs are mediated by underlying datasets, model architectures, and prompt-based inputs, which shape how content is generated and presented. These processes may introduce recognizable stylistic patterns, biases, or unexpected associations that audience members interpret as meaningful elements of the interactive artwork. As a result, AI technologies influence not only the generation of sensory stimuli but also the perceptual and cognitive processes through which audience members organize and interpret interactive experience.
Accordingly, this study examines what and how specific characteristics of AI systems—such as generative uncertainty and prompt-based multimodal interaction—shape aesthetic responses within interactive installation. Through case study analysis, this study explores how these technological features relate to the experience of audience members across sensory, perceptual, and cognitive dimensions, and how they influence aesthetic experiences such as surprise, confusion, fear, fantasy, resonance, or alienation during interaction with AI-based interactive artworks. This perspective will allow the analysis to connect technical characteristics of AI systems with the experience of audience members in interactive artwork.
Again, it should be noted that this analysis is based primarily on the available online documentation of the artworks rather than on actual audience data. Therefore, within the scope and limitations of this study, the findings should be understood as observation-centered and interpretive rather than empirically validated accounts of audience experience.
4. Case Studies
This section describes the case-study process, detailing the selection of the interactive artworks that incorporate AI and presenting the analysis of each case.
4.1. Selection of Case-Study Artworks
The interactive art installations that incorporate AI for this case-study analysis were selected based on the following criteria:
- Artworks that had been exhibited (online or offline) at renowned AI-related academic conferences.
- Recent AI-based interactive artworks created between 2023 and 2024.
- Installations employing interactive formats integrating AI technologies.
In the initial stage, we conducted a preliminary review of the 116 artworks featured in the AI Art Gallery of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2024, based on their online descriptions and available video documentation. Among these, 13 were identified as interactive pieces incorporating AI technologies (see the Appendix for the full list). From this subset, three exemplary cases were selected for in-depth analysis using the proposed analytical framework. The selected works were chosen on the basis that they clearly demonstrate the core analytical dimensions of this study—namely, AI-driven temporality, multimodality of input and output, and integration of aesthetic elements. Figure 2 illustrates the case-study selection process, and Table 3 lists the selected interactive art installations that use AI identified for analysis based on these criteria.
The number of cases (three) was considered appropriate for applying the analytical framework while maintaining the possibility of analytic generalization. Prior design-oriented aesthetic studies of interactive art have similarly relied on the qualitative analysis of carefully selected works, balancing both HCI-related and artistic considerations. In alignment with these established research practices, we adopted a qualitative case-based analytical approach [57].
4.2. Analytical Appoach
The overall analytical process was conducted based on the proposed research framework to systematically identify and articulate key points.
First, the analysis was structured around the research framework introduced previously:
- Audience participation: Audience participation was analyzed, focusing on the input data provided by audiences.
- Art object embodiment: The output of each artwork was analyzed.
- AI inquiry and response: This component was examined across three primary dimensions:
- AI Technologies
- Data Processing
- Temporality
- 4.
- Aesthetic experience: Considering the scope and limitations of the study outlined above, this dimension was interpreted through the researchers’ analytical observations of the aesthetic experiences that audiences might have while interacting with the artworks, based on related online resources rather than empirical data.
Second, based on this analytical framework, key findings were derived from insights gained during the analysis of each artwork. These findings emphasize the distinctive characteristics and unique qualities of emerging forms of AI-integrated interactive art installations.
Table 4 presents a structured summary of this analytical approach. It should be noted that the manifestation of each design element may vary substantially depending on the specific nature and context of the artwork.
4.3. Case Analyses
4.3.1. Case Analysis 1—Unreal Pareidolia -Shadows- (2023) by Scott Allen
Scott Allen’s Unreal Pareidolia -shadows- is an interactive art installation that converts the shadows of everyday objects into AI-generated images and captions [58,59] (Figure 3). The interaction flow is as follows: A participant arranges everyday items and toys before a lamp on a table, where their overlapping shadows are projected onto a wall. Upon completing the configuration, the participant presses a button on the table to signal completion. The system captures the composite shadow image (a shadowgraph) via a webcam. This image is sent to an image-to-text model, BLIP-2, which generates a caption displayed over the shadows on the wall in English and Japanese. Next, the caption and the shadowgraph are used as inputs (a prompt and an input image, respectively) for Stable Diffusion, which generates a final image. This image is projected onto the same wall as the light turns off. After a few seconds, the light turns back on, indicating that a new configuration can be made and allowing the audience to continue the interaction. Table 5 presents an analysis of the work.
The aesthetic experience of the work primarily rests on the main concept of the work, “pareidolia,” the psychological phenomenon of perceiving familiar patterns, such as faces or silhouettes, in random visual stimuli. While selecting and placing objects on the table is a sensory and perceptual activity, observing the resulting shadows and AI images requires cognitive effort to discover recognizable patterns in the shadowgraph and relate them to the image produced by Stable Diffusion. Notably, the AI model defamiliarizes these familiar objects by creating an image that may be independent of what the audience imagined from the shadowgraph. Technically, this reflects an “out-of-distribution” or “out-of-domain” problem, wherein the inference inputs differ from the training domain of the model; Stable Diffusion was not originally trained to generate images from shadowgraphs. Moreover, the captions generated from the shadowgraph are often incorrect or irrelevant to its actual construction. The work creatively repurposes this domain discrepancy and “uncertainty” as an artistic force, yielding novel aesthetic experiences throughout the interaction. Ultimately, the uncertainty introduced by AI systems can lead audiences to experience aesthetic responses characterized by unfamiliarity and perceptual estrangement. These responses may evoke surprise, cognitive disruption, and creative forms of aesthetic engagement.
Several key points are worth noting. First, the work exemplifies the combined use of caption generation and image synthesis from a single input (the shadowgraph). Second, the tangible interaction positions the audience as the primary creator of the input data. Third, the delayed, procedural pipeline—from object arrangement to button press to final image output—allows for the active, deliberate engagement, giving participants time to reflect on their interaction and, ultimately, the underlying concept of the artwork.
4.3.2. Case Analysis 2—ReCollection (2022-2024) by Weidi Zhang and Rodger Luo
ReCollection is an interactive AI art installation that assembles synthetic collective memories from audience-language input. It blurs the boundaries between remembrance and imagination through its system design and experimental visualization [24,60,61]. Figure 4 presents a representative image of the artwork. To interact with the work, a participant whispers a brief personal memory into a microphone. The system converts this speech to text using OpenAI’s Whisper model [62]. It then optionally completes or structures the narrative using a GPT-4 model [63] fine-tuned on documentaries of visual memories of Alzheimer’s patients and their descriptions. The resulting text prompt is fed into Stable Diffusion to generate an image sequence. These images are presented as an evolving visual composition that draws on fine art techniques, such as monotype-inspired treatments and slit-scan effects. The near-real-time animation and projection allow the participant to see their “imagined memories” unfold in response to their narrative input. A summary of the analysis is presented in Table 6.
The theme of the work centers on the creative reimagination of the memories of the audience—which are verbally described and often incomplete—through the lens of generative AI trained on descriptions of the memories of Alzheimer’s patients. During this visual reinterpretation, the audience engages with the aesthetic qualities of the work on multiple levels. First, whispering a personal memory fosters intimacy and self-reflection at a high level of cognition, as participants recognize they are disclosing private information to an AI, consider what to remember, and decide what to share. Observing the AI-generated image sequences then requires perceiving the emerging visuals and interpreting how the system reimagines, extends, or distorts that narrative by comparing it with one’s own recollection. The core role of the AI is to infer missing details and synthesize new visual memories from a broader corpus, using the memory of the participant as a seed for imagination. This interplay between human remembrance and machine imagination foregrounds subjective interpretation and highlights how meaning emerges from the interaction between audience input and AI-driven visualization. Ultimately, such cognitive engagement may enable aesthetic experiences marked by the unfamiliarity of reconstructed memories, the surprise or confusion produced by distortion and reinterpretation, and empathetic resonance with the experiences of others.
Several interesting points merit discussion. Technically, the work employs a fine-tuned GPT-4 model trained on specific Alzheimer-related documentary data. This curatorial approach to the training data aligns with the thematic focus on fragmented memories, drawing attention to both the theme and the issues surrounding the source community of the data. Second, Whisper’s multilingual support demonstrates the capacity of AI to increase accessibility for a broader audience without substantial localization efforts. Third, the work shows how voice user interfaces can be effectively applied to LLM-based AI art installations, integrating speech-to-text and text auto-completion rather than merely accepting raw prompts. Finally, the stepwise yet responsive pipeline—from speech capture to animated imagery—sustains immersion and affords reflection, enabling participants to consider how their narratives are transformed into synthetic, collectively inflected visual memories.
4.3.3. Case Analysis 3—Visions of Destruction (2023) by Varvara Guljajeva and Mar Canet
Visions of Destruction is an interactive, real-time AI artwork that links spectatorship to environmental changes within synthetic landscapes [18,43,64]. The installation continuously renders landscape scenes using Stable Diffusion. As shown in Figure 5, when a viewer looks at specific regions on the display, an eye-tracking system tracks their gaze and aggregates fixations into a mask. The system then applies AI inpainting to those masked areas, replacing healthy terrain with polluted, degraded motifs selected from a topic-relevant prompt set. If no gaze is detected for a specific period, the scene gradually regenerates through ongoing text-to-image synthesis and smooth latent interpolations. The simple act of looking becomes a catalyst for damage, while looking away enables recovery. Table 7 summarizes the analysis results.
The aesthetic experience of the work emerges from a tightly coupled loop between seeing and altering. On a sensory level, viewers encounter luminous, slowly shifting panoramas punctuated by local, gaze-triggered transformations. Perceptually, participants learn to recognize a contingent rule: wherever their attention lingers, decay spreads. Cognitively, the piece invites reflection on the observer effect and human responsibility for ecological harm. The act of averting one’s gaze to allow recovery produces a deliberate tension between curiosity and restraint, turning self-regulation into an ethical and conceptual gesture. Through its generative AI-driven cycle of destruction and recovery, the artwork may evoke shock and fear while also producing relief through restoration. In repeatedly exposing audiences to the consequences of their own participation, the work encourages reflection on the ethical implications of their actions and may foster experiences of ethical awareness, empathy, and self-reflection.
Several aspects merit attention. First, the installation demonstrates a coherent integration of gaze tracking with diffusion-based image generation and inpainting. Latent interpolation sustains continuity, while inpainting localizes change without resetting the entire scene. Second, the interaction model is elegantly minimal: gaze serves as both a sensor and tool, lowering the barrier to entry while maintaining strong authorial agency for the participant. Third, the temporal design balances immediacy and delay. Specifically, real-time tracking provides responsiveness, whereas the slight lag in mask aggregation and inpainting encourages anticipation and reflection rather than frantic scanning. Finally, the regeneration mechanism reframes noninteraction as a meaningful choice. By withholding attention, participants perform an act of care, which deepens the conceptual reading of the work and broadens the range of audience experiences within the system.
5. Discussion
Based on the case analyses and in relation to the elements of the framework, we inductively identify three key factors that characterize the design of AI-based interactive art: AI temporality, modality conversion, and AI-mediated aesthetic experience. Temporality is most directly linked to the framework because AI latency and inference time are integral to interactive system design. Modality conversion operates across audience input, AI data processing, and artwork output, enabling artworks to support diverse forms of participation and expression beyond text-based prompting. Aesthetic experience is also transformed by the incorporation of AI, including the types of AI technologies employed, the ways models are trained and used for inference, and the degree of agency AI assumes in shaping the final output. The discussion of each factor follows below.
5.1. AI Temporality in System Design and User Interaction
In interactive installations incorporating AI, the time from user input to system output is critically dependent on the temporality of the AI technologies involved. This temporality becomes a key determinant of both system design and user interaction, ultimately shaping the aesthetic experience. System design must account for inference delays, which vary with factors such as model architecture, deployment mode, and available computing resources. These constraints raise several questions: How should an installation handle delay or latency introduced by AI inference processes? How might artists leverage or mitigate the temporal constraints of AI models to achieve aesthetic goals? How do the temporal characteristics of AI systems influence artistic practice and the experience of the viewer? Based on our three case studies, we identified three distinct approaches to system design that address these questions.
Approach 1, Procedural, Event-driven, or Delayed Interaction: The system can intentionally structure multiple steps between user input and AI output. This approach is effective when real-time inference is infeasible on local hardware or when models run remotely via application programming interfaces (APIs) or other networked services. Interaction can proceed either procedurally, with step-by-step instructions guiding input preparation; be event-driven, where graphical user interface elements enable users to enter, modify, and confirm input; or be delayed, where the system visualizes processing while the work completes. For example, in Allen’s Unreal Pareidolia -shadows-, participants first arrange physical objects and then press a button to confirm their arrangement. A caption for the shadowgraph must be generated by BLIP-2 before Stable Diffusion can synthesize the image, which introduces a delay. However, the system mitigates this by displaying the caption immediately, which can reduce the “felt time” (subjective temporality) by cueing expectations about the forthcoming output, even as the “measurable time” (objective temporality) remains unchanged [65].
As shown in Figure 6, clear progress indicators, countdowns, or skeleton states can maintain engagement during longer calls, whereas preview or staging modes allow participants to refine inputs before committing to computation. If submissions must be queued or batched, communicating the queue position helps manage expectations and reduce frustration. The aesthetic implication is that staging and interim feedback can transform latency introduced by AI from a disruption into an act of anticipation and reflection.
Approach 2, Pseudo Real-Time Interaction: As shown in Figure 7, installations can incorporate modest processing and output latency while using design techniques to perceptually compress delay. In Zhang and Luo’s ReCollection, Whisper-based speech recognition introduces a delay of up to a few seconds, depending on factors such as audio length, network conditions, and server load. The work reduces the impact of this delay by maintaining continuous animation from the prior input, ensuring the display remains active. As a result, viewers may perceive the system as responding in real time.
Similar effects can be achieved by streaming partial results, using progressive refinement, prefetching predictable assets, or double-buffering visual states for smooth transitions. Continuous ambient soundscapes, low-salience auditory layer, or subtle background motion can also sustain presence during brief waiting periods. The aesthetic consequence is a continuous sense of liveness and immersion. Because the surface remains active and responsive, audiences experience the system as “listening” and “attending”, shifting their focus from waiting for computation to interpreting the incoming AI-generated imagery.
Approach 3, Real-Time, Immediate Interaction: Installations can employ AI capabilities that support near real-time output while accommodating continuous user input. In this approach, user actions directly generate new content or modify the existing content in place. As illustrated in Figure 8, AI can function as a tool that produces a steady stream of frames for smooth animation or an agent that alters content using algorithms such as inpainting. In Visions of Destruction [43], real-time gaze tracking drives Stable Diffusion inpainting within specific regions, while latent-space interpolation sustains smooth landscape animation. Together, these mechanisms maintain real-time interactivity.
To stabilize responsiveness, designers can throttle update rates, apply temporal easing, restrict updates to local regions within a fixed frame-time constraint, or employ low-pass filtering and debounce constraints on input sensor signals to mitigate high-frequency fluctuations. Caching prompts or seeds and deploying distilled or on-device models can further bound latency. This immediate interaction can yield an aesthetic of agency, contingency, and presence. The tight input–output coupling renders causality legible, invites improvisation, and turns feedback into an expressive mechanism operating at sub-second timescales.
Taken together, these cases show that AI temporality emerges as both a technical constraint and an expressive material that shapes interaction and meaning. Procedural or delayed sequencing cultivates anticipation; pseudo real-time orchestration compresses perceived delay to sustain continuity; and immediate real-time coupling increases agency and presence. Managing measured latency and perceived time through sequencing, continuity cues, and stabilized updates turns delay from friction into structure, directly informing the aesthetic character of the work.
5.2. Modality Conversion in AI Interaction
Across the three case studies, interaction with generative AI extends far beyond simple text entry. Generative AI pipelines often rely on joint text–image representations and transformer-based models to achieve multimodal generation (e.g., text-to-image and text-to-sound). However, text need not be the primary input modality of the audience. Instead, systems can convert tangible, verbal, or embodied signals into internal representations that condition or contribute to the generation process. This conversion chain is central to how human and machine agency, artwork meaning, and interaction timing are experienced.
Approach 1, Tangible Interaction: Tangibility can be an integral mode of interaction in AI-based installations, coupling material manipulation with computational imagination. In Unreal Pareidolia -shadows-, participants arrange physical objects to produce a shadowgraph, which is then captured by a camera, captioned by an image-to-text model, and used to condition image synthesis. This approach transduces a tactile action into a visual latent and textual prompt, which together steer the output. Figure 9 illustrates this conversion from tangible input to the desired output modality.
Tangible interaction can support modality conversion beyond image synthesis. For instance, drawings depicting the mood of a participant could serve as inputs for image-to-sound or -music generation. Physical collages could be segmented into layout prompts for story or caption generation. Clay or modular sculptures scanned in situ could be used to initialize 3D scene or object synthesis via image-to-3D pipelines. Textile or weaving artifacts can be vectorized and used to condition the generation of materials and textures for 3D assets. Additionally, tangible user interfaces with buttons, sliders, or knobs can enable participants to engage directly with AI content generation by manipulating controllable parameters rather than serving only as providers of raw input, supporting iterative, hands-on exploration and fine-grained manipulation of the output.
Approach 2, Verbal Interaction: As illustrated in Figure 10, verbal interfaces can position speech as the primary conduit into generative AI pipelines by converting acoustic signals into textual or semantic representations. In ReCollection, the system records a whispered memory, performs speech recognition, optionally restructures the narrative with an LLM, and then drives image generation. Thus, speech becomes a narrative text that conditions a visual output, demonstrating how conversational AI interfaces can be adapted to installation contexts.
Verbal interaction has broader applications. For example, spoken scene descriptions can be translated into structured scripts for text-to-video generation; prosodic features (e.g., tempo and intensity) can modulate style parameters for music or affective color palettes for images; and multilingual speech can be normalized into prompts that control 3D scene synthesis or texture generation. Moreover, voice-driven refinement loops, wherein the system exposes interim transcripts or candidate paraphrases for confirmation, can support iterative engagement without returning to the beginning for a new input.
Approach 3, Embodied Interaction: As shown in Figure 11, embodied interfaces can mobilize sensorimotor signals as continuous control fields that directly modulate generative AI processes. In Visions of Destruction [43], gaze is tracked in real time, aggregated into spatial masks, and used to localize inpainting within a continuously generated landscape. Thus, embodied signals function as control layers over the evolving scene.
Beyond gaze, designers can incorporate various embodied interactions. For instance, whole-body poses can drive character synthesis or choreography in video generation; proxemic distance and orientation can adjust guidance strength, field of view, or depth-of-field in image synthesis; and motion-captured props can instantiate or reposition 3D objects within text-conditioned scene construction. Physiological inputs, including heart rate, respiration, electroencephalogram (EEG) signals, or electromyography signals, can modulate tempo, contrast, saturation, or noise schedules in outputs that track body state. To sustain expressive control, systems can expose calibrated mappings between embodied signals and generative parameters, allowing participants to continuously shape content rather than merely trigger it.
These modality conversions have concrete design consequences. First, the accuracy, lossiness, and bias of each converter influence how faithfully intent is conveyed. Second, the granularity of control depends on the representation; continuous masks afford different agency than discrete prompts. Third, transparency regarding intermediate artifacts (e.g., captions, transcripts, or masks) can make conversion a collaborative step, allowing participants to steer the system before final rendering. Finally, because conversion stages impose their own latencies, orchestration strategies such as progressive refinement, preview states, or continuous background motion are crucial for managing perceived delay.
By foregrounding modality conversion, interactive AI artworks can move beyond prompt-centric interfaces toward plural, situated forms of engagement that align audience capabilities with the representational affordances of generative models. Multimodal redundancy can support accessibility by providing alternate pathways for participation, whereas mixed-initiative repair loops can mitigate conversion errors.
5.3. AI-Mediated Aesthetics
Building upon the proposed analytical framework, this study investigates how AI mediates novel forms of aesthetic experience in interactive art. Our analysis of the three artworks identified three key mechanisms through which AI transforms aesthetic engagement: (1) transcending human cognition, (2) enabling personalized aesthetic interaction, and (3) sustaining or expanding continuous engagement. These findings provide insights into how AI is reshaping the aesthetic dimension of HCI within interactive art.
Approach 1, Transcending Human Cognition and Perception: Machine perception is the capability of computational systems to acquire sensory data and convert it into meaningful interpretations of the environment, in a manner analogous to human perceptual understanding [66,67]. Unlike human perception, which is grounded in embodied experience, cultural context, and accumulated knowledge, machine perception relies exclusively on structured and quantifiable data acquired through sensors and computational systems [68,69,70]. Such data are processed through pattern recognition and probabilistic inference models based on statistical correlations rather than integrative or semantic understanding. As a result, machine systems do not apprehend the world as a coherent and meaningful whole but instead transform perceptual inputs into computational representations. Consequently, their outputs remain constrained by the distribution and scope of the training data, limiting their capacity to interpret aspects of reality not encoded within those data structures.
From this perspective, AI technologies may extend aesthetic experience beyond the boundaries of human cognition and perception. The divergence between human and machine perception creates an alternative mode of “seeing” that generates defamiliarization—a core aesthetic strategy that challenges habitual modes of experience. In Unreal Pareidolia -shadows-, the AI produces unforeseen visual outcomes when exposed to inputs outside its training domain.
The AI interprets object configuration from silhouette information, generates textual descriptions, and transforms them into unfamiliar visual compositions. Although the resulting image may contain familiar elements, the process itself produces a defamiliarization effect, inducing a perceptual confusion—manifested as errors, distortions, or algorithmic defamiliarization—experienced by the audience. Through this, the audience is invited to reflect on the limitations of human perception. As illustrated in Figure 12, the gap between human and machine perception becomes a site for the formation of the aesthetic experience.
Although machine vision is not inherently aesthetic, its integration allows artists to explore the tension between human and non-human cognition, expanding the creative field through algorithmic mediation.
Approach 2, Enabling Personalized Aesthetic Experiences: AI also enables deeply personalized aesthetic experiences by responding dynamically to user-specific data. While traditional interactive installations often restrict participants to predefined scenarios, AI-based systems can tailor the experience, generating context-sensitive and emotionally resonant outputs. In ReCollection, the system employs multilingual AI models to record and analyze audience voices, transforming them into personalized visual and auditory compositions. As shown in Figure 13, this process converts individual memories into aesthetic representations, creating an intimate, reflective engagement. Moreover, when AI models are trained on culturally diverse datasets, they can allow audiences to experience vicarious perspectives, extending aesthetic participation into cross-cultural empathy. These personalized, culturally informed experiences demonstrate the potential of AI in HCI to create participant-specific, context-aware interactive art.
Pre-trained AI models inherently entail the risk of algorithmic bias arising from the datasets on which they are trained. In interactive environments involving audience participation, such biases may be manifested in output without sufficient control, thereby constituting a critical consideration for artists in the design of personalized aesthetic experiences.
Approach 3, Sustaining and Expanding Aesthetic Interaction: AI-driven interactivity enables continuous, evolving engagement, rather than discrete, one-time interactions. Traditional interactive systems often rely on fixed content, whereas AI allows artworks to adapt dynamically over time, encouraging sustained participation and iterative exploration. For instance, in Visions of Destruction [43], the system tracks audience gaze and generates imagery situated between visibility and invisibility. This design invites viewers to oscillate between attention and avoidance, gradually deepening their aesthetic engagement. Repetitive interactions allow the artwork to approximate the conceptual intent of the artist while maintaining participant autonomy. Conversely, in Unreal Pareidolia -shadows-, the unpredictability of AI can elicit unintended behaviors that transcend both the intention of the artist and the expectation of the audience. These emergent behaviors align with prior notions of unintended experience [10,16], as seen in earlier works such as Legible City [71,72] and Rain Room [73], wherein participants deliberately subverted system constraints. As illustrated in Figure 14, such dynamics highlight the dual role of AI: it can refine artistic control and provoke spontaneous, uncontrollable aesthetic events, urging HCI designers to reconsider the balance between system autonomy and user agency.
Across these three approaches, AI emerges as a co-creative agent that redefines aesthetic engagement. Specifically, it extends human perception, personalizes interaction, and sustains continuous engagement, transforming the traditional relationship between artist, audience, and system. From an HCI perspective, AI-mediated art exemplifies how computational systems can generate emergent, affective, and reflective experiences that bridge human and machine cognition. This framework offers a foundation for future research on AI-driven creativity, highlighting aesthetic interaction as a critical dimension of human–AI collaboration.
6. Conclusions
6.1. Implications
This study proposed the 4A framework, a new analytical framework for examining how AI technologies shape interaction processes and aesthetic experience in interactive artworks. The framework comprises three key entities (audience, art object, and AI) and four design elements (participation, embodiment, AI inquiry and response, and aesthetic experience), addressing a critical gap in prior research that has yet to fully account for the role of AI. The analysis of three contemporary cases—Unreal Pareidolia -shadows-, ReCollection, and Visions of Destruction [43]—reveals that AI functions not merely as a technical tool but also as a critical mediating element that connects artistic intention, system architecture, and audience participation. These results suggest that future AI-based interactive art should be designed with careful consideration of AI temporality in system design and user interaction, modality conversion in AI interaction, and the evolving role of AI in aesthetic experience. Based on these findings, the following design implications and guidelines are suggested.
Creative Use of Cognitive Gaps: Rather than viewing discrepancies between human and AI reasoning as mere limitations, these cognitive gaps can be deliberately incorporated as aesthetic resources. They may evoke ambiguity, surprise, and interpretive openness, thereby expanding the imagination and engagement of the audience. When designing AI-driven interactive installations, artists can leverage these cognitive contrasts to create aesthetic experiences that unfold at a perceptual or conceptual level distinct from human cognition. At the same time, these gaps may also function as points of failure, potentially leading to confusion or frustration rather than aesthetic engagement. However, when carefully mediated through appropriate interaction design strategies, such as procedural or delayed interaction, they can be strategically transformed into sources of novel aesthetic experience.
Integrated Design of Multimodal Inputs: Combining multimodal and embodied input methods—such as voice, gaze, and bodily movement—can deepen interaction and enhance immersion. A multimodal interface allows users to communicate with the artwork through their physical and sensory actions. By deliberately designing and integrating such interfaces, artists can enable AI systems to interpret heterogeneous inputs. The complementary and redundant functions of these inputs can improve accommodation and robustness, ultimately transforming these signals into diverse forms of artistic expression.
Embracing Implicit and Nonverbal User Behaviors: User input should not be limited to explicit actions. Subtle, unconscious, or nonverbal gestures can also be interpreted by AI to generate meaningful artistic responses. Through this design approach, artists can embrace serendipity and improvisation, expanding the aesthetic experiences born from unpredictable interactions between audiences and AI installations.
Maintaining Artistic Consistency Amid Uncertainty: Even when user input is excessive, insufficient, or ambiguous, AI should be designed to reinforce artistic intent and maintain systemic coherence. This approach can ensure aesthetic unity and experiential completeness, even under uncertain interaction conditions. Artists should aim for AI systems that strengthen their creative direction. Concurrently, they must acknowledge that unexpected audience reactions may transcend their original intent, and should even consider incorporating such unpredictability into the work’s design.
Overall, these insights provide practical guidance on the design principles and creative methodologies for AI-based interactive art installations. Moreover, they suggest new possibilities for generating richer, more nuanced aesthetic experiences at the convergence of technology and human perception.
6.2. Limitations and Future Work
Although this study explained the rationale for selecting three cases, the limited number of analyzed artworks presents a significant limitation. Future research should apply the 4A framework to a broader range of AI-based interactive artworks to further assess its applicability, refine its analytical dimensions, and examine its adaptability within the rapidly evolving landscape of AI technologies. Such comparative investigations will help strengthen the framework and support the development of more robust approaches to the analysis and design of AI-mediated interactive art.
Another limitation lies in the treatment of aesthetic experience. Because it was not feasible within the scope of this research to collect primary empirical data, such as audience interviews, surveys, or behavioral measurements, the analysis relied on the official descriptions of the artworks, related publications, and publicly available video documentation of the installations. As a result, the aesthetic findings should be understood as interpretive and framework-oriented rather than as empirically validated accounts of audience experience. Future research should integrate empirical user studies and audience-centered evaluation to examine how aesthetic experience is perceived by participants and to strengthen the analytical validity of this dimension within the framework.
The technical analysis of the selected artworks was also limited by the availability of documentation. Since the implementation of such work is often closely linked to artistic originality, core technical details are frequently only partially disclosed or withheld. As a result, the depth of technical verification and analysis in this study remains constrained. Future research should seek more detailed and technically grounded investigations of AI-based interactive artworks, particularly where richer implementation information can be obtained.
A final limitation lies in the scope of the framework. While ethical issues surrounding AI, including consent, privacy, and bias, are critically important in contemporary AI art, they were only briefly acknowledged rather than systematically analyzed. Likewise, theoretical debates concerning artist agency, authorship, and aesthetic value remain central to AI-based interactive art, yet they were not treated as the primary concern of this study, whose purpose was to advance an analytical framework rather than to develop new theoretical positions. Future research should therefore extend the framework by incorporating these ethical and theoretical dimensions more explicitly.
Author Contributions
Sihwa Park: Conceptualization, Funding acquisition, Resources, Methodology, Visualization, Writing—original draft, Writing—review & editing; Je-ho Oh: Conceptualization, Project administration, Visualization, Writing—original draft, Writing—review & editing.
Funding
This research was undertaken thanks in part to funding from the Connected Minds Program, supported by Canada First Research Excellence Fund, Grant #CFREF-2022-00010.
Data Availability Statement
The original contributions presented in this study are included in the article/supplementary material. Further inquiries can be directed to the corresponding author(s).
Declaration of the use of Generative AI
The authors used ChatGPT solely for language editing and proofreading of the manuscript. The authors reviewed and edited the generated output as necessary and take full responsibility for the content of this article.
Conflicts of Interest
The authors declare no conflicts of interest.
Abbreviations
The following abbreviations are used in this manuscript:
| AI | Artificial intelligence |
| ML | Machine learning |
| VAE | Variational autoencoder |
| LLM | Large-language model |
| HCI | Human–computer interaction |
| DL | Deep learning |
| CVPR | Computer vision and pattern recognition |
Appendix A
Table A1 lists the 13 artworks identified as interactive AI installations in the CVPR 2024 AI Art Gallery.
Table A1.
AI-based Interactive Art Installations in the CVPR 2024 AI Art Gallery.
| No. | Title | Artist(s) | Year |
|---|---|---|---|
| 1 | Unreal Pareidolia -shadows- | Scott Allen | 2023 |
| 2 | Visions Of Destruction | Varvara & Mar | 2023 |
| 3 | AI Nüshu (AI女书) | Yuqian Sun | 2023 |
| 4 | The Endless Collaborative Book: Wushu | Aven Le Zhou | 2023 |
| 5 | Connections | Yamin Xu | 2023 |
| 6 | Neuracappella | Zihou Ng | 2023 |
| 7 | Speculative Evolution, Prototype 1 | Marc Lee | 2024 |
| 8 | YouTube Mirror | Sihwa Park | 2022-2023 |
| 9 | Gender Tapestry | J. Rosenbaum | 2023-2024 |
| 10 | ReCollection | Weidi Zhang, Rodger Luo | 2022 |
| 11 | On The Existence of Self-Identity | Julia Chylak, Mateusz Błajda, Eryk Imos, Beata Bajno, Paulina Wachnicka | 2023 |
| 12 | Beyond Characters: The Unseen Labyrinth | Mingyong Cheng, Zetao Yu | 2023-2024 |
| 13 | LB²—LVX-1 | LB2 | 2024 |
References
- Kluszczynski, R.W. Strategies of interactive art. J. Aesthet. Cult. 2010, 2, 5525. [Google Scholar] [CrossRef]
- Kwastek, K. Aesthetics of Interaction in Digital Art; The MIT Press, 2013. [Google Scholar] [CrossRef]
- Lopes, D.M.M. The ontology of interactive art. J. Aesthet. Educ. 2001, 35, 65–81. [Google Scholar] [CrossRef]
- Bruder, K.A.; Ucok, O. Interactive art interpretation: How viewers make sense of paintings in conversation. Symb. Interact. 2000, 23, 337–358. [Google Scholar] [CrossRef]
- Ahmed, S.U. Interaction and interactivity: In the context of digital interactive art installation. In Human–Computer Interaction. Interaction in Context; Kurosu, M., Ed.; Springer International Publishing, 2018; pp. 241–257. [Google Scholar] [CrossRef]
- Bendor, R.; Maggs, D.; Peake, R.; Robinson, J.; Williams, S. The imaginary worlds of sustainability: Observations from an interactive art installation. Ecol. Soc. 2017, 22. [Google Scholar] [CrossRef]
- Jacucci, G.; Wagner, M.; Wagner, I.; Giaccardi, E.; Annunziato, M.; Breyer, N.; Hansen, J.; Jo, K.; Ossevoort, S.; Perini, A.; et al. ParticipArt: Exploring participation in interactive art installations. IEEE International Symposium on Mixed and Augmented Reality – Arts, Media, and Humanities, Seoul, South Korea, Oct. 2010; 2010, pp. 3–10. [Google Scholar] [CrossRef]
- Morrison, A.J.; Mitchell, P.; Brereton, M. The lens of ludic engagement: Evaluating participation in interactive art installations. In Proceedings of the 15th ACM International Conference on Multimedia Acad. Med., Augsburg, Germany, 2007; pp. 509–512. [Google Scholar] [CrossRef]
- Pavlin, E.; Elsner, Ž.; Jagodnik, T.; Batagelj, B.; Solina, F. From illustrations to an interactive art installation. J. Inf. Commun. Ethics Soc. 2015, 13, 130–145. [Google Scholar] [CrossRef]
- Bilda, Z.; Bowman, C.; Edmonds, E. Experience evaluation of interactive art: Study of GEO landscapes. In Proceedings of the 5th Australasian Conference on Interactive Entertainment, Brisbane, Australia. Acad. Med., 2008; pp. 1–10. [Google Scholar] [CrossRef]
- Costello, B.; Muller, L.; Amitani, S.; Edmonds, E. Understanding the experience of interactive art: Iamascope in Beta_space,”. ACM Digit. Libr. Proc. Second Australas. Conf. Interact. Entertain. 2005, 49–56. Available online: https://dl.acm.org/doi/abs/10.5555/1109180.1109188.
- Jacquemin, C.; Gagneré, G.; Lahoz, B. Shedding light on shadow: Real-time interactive artworks based on cast shadows or silhouettes. In Proceedings of the 19th ACM International Conference on Multimedia, Scottsdale, AZ, USA. Acad. Med., 2011; pp. 173–182. [Google Scholar] [CrossRef]
- Papageorgopoulou, P.; Arsenopoulou, N.; Charitos, D.; Rizopoulos, C.; Theona, I.; Katsarou, L.; Psaltis, A.; Korosidis, A. Designing interfaces to promote the meaningfulness of urban data through an interactive art installation. In Proceedings of the Chi Greece: 1st International Conference of the ACM Greek SIGCHI Chapter, Athens, Greece, Nov. 2021, 2021. [Google Scholar] [CrossRef]
- Yiyuan, H. Creation methodology of interactive art installation based on philosophy-understanding projection: Recreation of traditional Chinese painting. In Proceedings of the 2015 Virtual Reality International Conference, Laval, France. Acad. Med., 2015. [Google Scholar] [CrossRef]
- Fels, S. Intimacy and embodiment: Implications for art and technology. In Proceedings of the. ACM Workshops Multimedia 2000, Available online. 2000; ACM; pp. 13–16. (accessed on Nov. 2000). [Google Scholar] [CrossRef]
- Edmonds, E. The art of interaction. Digit. Creat. 2010, 21, 257–264. [Google Scholar] [CrossRef]
- Xiaobo, L.; Yuelin, L. Embodiment, interaction and experience: Aesthetic trends in interactive media arts. Leonardo 2014, 47, 166–169. [Google Scholar] [CrossRef]
- Canet Sola, M.; Guljajeva, V. Visions of destruction: Exploring a potential of generative AI in interactive art. In Proceedings of the 17th International Symposium on Visual Information Communication and Interaction, Hsinchu, Taiwan. Acad. Med., 2024. [Google Scholar] [CrossRef]
- Guljajeva, V.; Canet Sola, M. POSTcard landscapes from Lanzarote. In Proceedings of the 14th Conference on Creativity and Cognition Acad. Med., Venice, Italy, 2022; pp. 634–636. [Google Scholar] [CrossRef]
- Jacob, C.J.; Hushlak, G.; Boyd, J.E.; Nuytten, P.; Sayles, M.; Pilat, M. SwarmArt: Interactive art from swarm intelligence. Leonardo 2007, 40, 248–254. [Google Scholar] [CrossRef]
- Li, J.; Sun, G.; Tang, C.; Chen, W.; Yang, W.; Kou, W.; Ruan, Z.; Ma, W.; Nie, X. Silk road journey: A real-time AI-based interactive art installation for silk road cultural reenactment and experience. Extended abstracts of the CHI conference on human factors in computing systems Acad. Med., 2024; pp. 1–5. [Google Scholar] [CrossRef]
- Sun, Y.; et al. AI Nüshu: An exploration of language emergence in sisterhood through the lens of computational linguistics. In Proceedings of the SIGGRAPH Asia 2023 Art Papers, Sydney, Australia. Acad. Med., 2023. [Google Scholar] [CrossRef]
- Tidemann, A.; Brandtsegg, Ø. Entertainment Computing – ICEC 2015. In self.]: realization / art installation / artificial intelligence: a demonstration; Chorianopoulos, K., Divitini, M., Baalsrud Hauge, J., Jaccheri, L., Malaka, R., Eds.; Springer International Publishing: Cham, 2015; pp. 517–522. [Google Scholar]
- Zhang, W.; Cheng, L.; Luo, J. ReCollection: Creating synthetic memories with AI in an interactive art installation. Proc. ACM Comput. Graph. Interact. Tech. 2024, 7, 1–10. [Google Scholar] [CrossRef]
- Garcia, M.B. The paradox of artificial creativity: Challenges and opportunities of generative AI artistry. Creat. Res. J. 2025, 37, 755–768. [Google Scholar] [CrossRef]
- Messer, U. Co-creating art with generative artificial intelligence: Implications for artworks and artists. Comput. Hum. Behav. Artif. Hum. 2024, 2, 100056. [Google Scholar] [CrossRef]
- Mondal, B. Artificial intelligence: State of the art. In Recent Trends and Advances in Artificial Intelligence and Internet of Things; Balas, V.E., Kumar, R., Srivastava, R., Eds.; Springer International Publishing, 2020; pp. 389–425. [Google Scholar] [CrossRef]
- Audry, S. Art in the Age of Machine Learning; The MIT Press, 2021. [Google Scholar] [CrossRef]
- Scurto, H.; Caramiaux, B.; Bevilacqua, F. Prototyping machine learning through diffractive art practice. In Proceedings of the 2021 ACM Designing Interactive Systems Conference, Virtual Event, USA. Acad. Med. 2021; pp. 2013–2025. [Google Scholar] [CrossRef]
- Plant, N.; et al. Interactive machine learning for embodied interaction design: A tool and methodology. In Proceedings of the Fifteenth International Conference on Tangible, Embedded, and Embodied Interaction, Salzburg, Austria. Acad. Med., 2021. [Google Scholar] [CrossRef]
- Mahmud, B.; Hong, G.; Fong, B. A study of human–AI symbiosis for creative work: Recent developments and future directions in deep learning. ACM Trans. Multimed. Comput. Commun. Appl. 2023, 20, 1–21. [Google Scholar] [CrossRef]
- Zhang, S.; Qi, Y.; Wu, J. Applying deep learning for style transfer in digital art: Enhancing creative expression through neural networks. Sci. Rep. 2025, 15, 11744. [Google Scholar] [CrossRef] [PubMed]
- Oppenlaender, J.; Johnston, H.; Silvennoinen, J.M.; Barranha, H. Artworks reimagined: Exploring human-AI co-creation through body prompting. Proc. ACM Hum.-Comput. Interact. 2025, 9, 1–34. [Google Scholar] [CrossRef]
- Epstein, Z.; Hertzmann, A. Investigators of Human Creativity Art and the science of generative AI. In Science; Akten, M., Farid, H., Fjeld, J., Frank, M.R., Groh, M., Herman, L., Leach, N., et al., Eds.; 2023; Volume 380, pp. 1110–1111. [Google Scholar] [CrossRef] [PubMed]
- Furtado, L.S.; Soares, J.B.; Furtado, V. A task-oriented framework for generative AI in design. J. Creat. 2024, 34, 100086. [Google Scholar] [CrossRef]
- Kun, P.; Freiberger, M.A.; Løvlie, A.S.; Risi, S. GenFrame – Embedding generative AI into interactive artifacts. In Proceedings of the 2024 ACM Designing Interactive Systems Conference, Copenhagen, Denmark Acad. Med., 2024; pp. 714–727. [Google Scholar] [CrossRef]
- Saltz, D.Z. The art of interaction: Interactivity, performativity, and computers. J. Aesthet. Art. Crit. 1997, 55, 117–127. [Google Scholar] [CrossRef] [PubMed]
- Graham, C.B. A study of audience relationships with interactive computer-based visual artworks in gallery settings, through observation, art practice, and curation. PhD Thesis, University of Sunderland, 1997. Available online: http://www.berylgraham.com/cv/sub/thesis.pdf.
- Edmonds, E.; Turner, G.; Candy, L. Approaches to interactive art systems. In Proceedings of the 2nd International Conference on Computer Graphics and Interactive Techniques in Australasia and South East Asia, Singapore. Acad. Med., 2004; pp. 113–117. [Google Scholar] [CrossRef]
- Chen, W.; Shidujaman, M.; Jin, J.; Ahmed, S.U. A methodological approach to create interactive art in artificial intelligence. In Cognition, International, H.C.I., Ed., 2020. – Late Breaking Papers, Learning and Games; Stephanidis, C., Harris, D., Li, W.-C., Schmorrow, D. D., Fidopiastis, C. M., Zaphiris, P., Ioannou, A., Fang, X., Sottilare, R. A., Schwarz, J., Eds.; Springer International Publishing: Cham, 2020; pp. 13–31. [Google Scholar]
- Xie, J.; Yu, M.; Liu, G. Fusing algorithms for intersection of computer science and art: Innovations in generative art and interactive digital installations. IEEE Access 2024, 12, 173255–173267. [Google Scholar] [CrossRef]
- Canet Sola, M.; Guljajeva, V. Dream Painter: Exploring creative possibilities of AI-aided speech-to-image synthesis in the interactive art context. Proc. ACM Comput. Graph. Interact. Tech. 2022, 5, 1–11. [Google Scholar] [CrossRef]
- Canet Sola, M.; Guljajeva, V. Visions of destruction. In SIGGRAPH Asia Ed. (SA ‘23): Special interest group on computer graphics and interactive techniques conference 2023; Gallery, A., Ed.; ACM, 2023; pp. 1–2. [Google Scholar] [CrossRef]
- Sklar, S.J.; Jiang, M. Art with agency: Artificial intelligence as an interactive medium. Humanit. Soc. Sci. Commun. 2025, 12, 1546. [Google Scholar] [CrossRef]
- Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022; Institute of Electrical and Electronics Engineers; pp. 10674–10685. [Google Scholar] [CrossRef]
- Li, J.; Li, D.; Savarese, S.; Hoi, S. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. Proceedings of the Proceedings of the Mach. Learn. ReS 40th Int. Conf. Mach. Learn. (ICML) 2023, Vol. 202, 19730–19742. [Google Scholar]
- Savaş, E.B.; Verwijmeren, T.; van Lier, R. Aesthetic experience and creativity in interactive art. Art. Percept. 2021, 9, 167–198. [Google Scholar] [CrossRef]
- Bilda, Z.; Candy, L.; Edmonds, E. An embodied cognition framework for interactive experience. CoDesign 2007, 3, 123–137. [Google Scholar] [CrossRef]
- Candy, L. Evaluation and experience in art; Interactive Experience in the Digital Age: Evaluating New Art Practice, Candy, Candy, L., Ferguson, S., Eds.; Springer International Publishing, 2014; pp. 25–48. [Google Scholar] [CrossRef]
- Höök, K.; Sengers, P.; Andersson, G. “Sense and Sensibility”: evaluation and interactive art,”. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, Ft. Acad M, 2003; pp. 241–248. [Google Scholar] [CrossRef]
- Szubielska, M.; Imbir, K.; Szymańska, A. The influence of the physical context and knowledge of artworks on the aesthetic experience of interactive installations. Curr. Psychol. 2021, 40, 3702–3715. [Google Scholar] [CrossRef]
- Zhou, C.; Li, J. The development of aesthetic experience through virtual and augmented reality. Sci. Rep. 2024, 14, 4290. [Google Scholar] [CrossRef] [PubMed]
- Heid, K. Aesthetic development: A cognitive experience. Art. Educ. 2005, 58, 48–53. [Google Scholar] [CrossRef]
- Joy, A.; Sherry, J.F., Jr. Speaking of art as embodied imagination: A multisensory approach to understanding aesthetic experience. J. Consum. Res. 2003, 30, 259–282. [Google Scholar] [CrossRef] [PubMed]
- Leder, H.; Belke, B.; Oeberst, A.; Augustin, D. A model of aesthetic appreciation and aesthetic judgments. Br. J. Psychol. 2004, 95, 489–508. [Google Scholar] [CrossRef] [PubMed]
- Caramiaux, B.; Fdili Alaoui, S. “Explorers of unknown planets”: Practices and politics of artificial intelligence in visual arts. Proc. ACM Hum.-Comput. Interact. 2022, 6, 1–24. [Google Scholar] [CrossRef]
- Jeon, M.; Fiebrink, R.; Edmonds, E.A.; Herath, D. From rituals to magic: Interactive art and HCI of the past, present, and future. Int. J. Hum.-Comput. Stud. 2019, 131, 108–119. [Google Scholar] [CrossRef]
- Allen, S. Unreal Pareidolia -shadows-. Scott Allen [Online]. 2023. Available online: https://scottallen.ws/work/unreal-pareidolia-shadows.
- Allen, S. Scott Allen | CVPR AI art [Online]. 2024. Available online: https://thecvf-art.com/project.php?year=2024&artist=scott-allen&id=786.
- Zhang, W.; Luo, J. SIGGRAPH 2023 Art Gallery (SIGGRAPH ‘23): Special interest group on computer graphics and interactive techniques conference. In Acad. Med.; 2023; Volume (1–2). [Google Scholar] [CrossRef]
- Zhang, W.; Luo, R. Weidi Zhang, Rodger Luo | CVPR AI art [Online], 2024. Available online: https://thecvf-art.com/project.php?year=2024&artist=recollection&id=687.
- Radford, A.; Kim, J.W.; Xu, T.; Brockman, G.; McLeavey, C.; Sutskever, I. Robust speech recognition via large-scale weak supervision. Proceedings of the Proceedings of the Mach. Learn. ReS 40th Int. Conf. Mach. Learn. (ICML) 2023, Vol. 202, 28492–28518. [Google Scholar]
- OpenAI. GPT-4 Technical Report. arXiv 2023. [Google Scholar]
- Guljajeva, V.; Canet Sola, M. Varvara & Mar | CVPR AI art [Online], 2024. Available online: https://thecvf-art.com/project.php?year=2024&artist=varvara-mar&id=758.
- Wiberg, M.; Stolterman, E. Time and temporality in HCI research. Interact. Comput. 2021, 33, 250–270. [Google Scholar] [CrossRef]
- Bruckner, D.; Velik, R.; Penya, Y.K. Machine perception in automation: A call to arms. EURASIP J. Embed. Syst. 2011, 2011, Art.(no. 608423). [Google Scholar] [CrossRef]
- Nevatia, R. Machine Perception; Prentice Hall, 1982. [Google Scholar]
- Guile, D.; Popov, J. Machine learning and human learning: A socio-cultural and -material perspective on their relationship and the implications for researching working and learning. AI Soc. 2025, 40, 325–338. [Google Scholar] [CrossRef]
- Hutchins, E. Cognition in the Wild; The MIT Press, 1995. [Google Scholar] [CrossRef]
- Walker, R.J. Seeing Ourselves Through Technology How We Use Selfies, Blogs and Wearable Devices to See and Shape Ourselves; Springer Nature, 2014. [Google Scholar]
- Lit. Von Peter Weibel Graz Lit. Droschl. 2001, pp. 387–398. Available online: https://www.jeffreyshawcompendium.com/wp-content/uploads/2018/03/2001_Im-Buchstabenfeld.EXT_.GERMAN.pdf.
- Shaw, J. » The Legible City. 1988–1991 Im Buchstabenfeld Zuk.
- Peck, E.J. Understanding digital and interactive media in live events and performance: The RAIN ROOM. Academia [Online]. 2018. Available online: https://www.academia.edu/download/57279835/RainRoom.pdf.
Figure 1.
Proposed 4A Framework.

Figure 2.
Flowchart illustrating the case-study artwork selection process.

Figure 3.
Scott Allen, Unreal Pareidolia -shadows- (2023) (Image courtesy of the artist © Scott Allen).
Figure 3.
Scott Allen, Unreal Pareidolia -shadows- (2023) (Image courtesy of the artist © Scott Allen).

Figure 4.
Weidi Zhang and Rodger Luo, ReCollection (2022–2024) (Image courtesy of the artists).

Figure 5.
Varvara Guljajeva and Mar Canet, Visions of Destruction (2023) (Image courtesy of the artists).
Figure 5.
Varvara Guljajeva and Mar Canet, Visions of Destruction (2023) (Image courtesy of the artists).

Figure 6.
Procedural, event-driven, or delayed interaction approach.

Figure 7.
Pseudo real-time interaction approach.

Figure 8.
Real-time, immediate interaction approach.

Figure 9.
Modality conversion process for tangible interaction.

Figure 10.
Modality conversion process for verbal interaction.

Figure 11.
Modality conversion process for embodied interaction.

Figure 12.
Aesthetic approach: transcending human cognition and perception.

Figure 13.
Aesthetic approach: enabling personalized aesthetic experiences.

Figure 14.
Aesthetic approach: sustaining and expanding aesthetic interaction.

Table 1.
Summary of fundamental models.
| Categories | Fels [15] | Edmonds et al. [39] | Xiaobo & Yuelin [17] | Ahmed [5] | Canet Sola & Guljajeva [42] |
|---|---|---|---|---|---|
| Human | Person | Viewer, artist (human agent) | Audience | Artist, audience | Author, audience |
| Objects | Object | Art object | Interactive system | Artwork | Artwork |
| Environment, software agent | Environment | ||||
| Relationships | Engagement, embodiment, communication | Interaction (Static, dynamic-passive, dynamic-interactive, dynamic-interactive (varying)) | Interaction, aesthetics | Interaction, communication | Feedback |
Table 2.
Core Elements of the 4A Framework.
| Key design elements | Analysis elements | Description |
|---|---|---|
| Audience participation | Input | Audience participation is closely tied to the system’s input. This element is categorized by the format and modality of the input data, including participant choices that influence the artwork’s overall shaping process. |
| Art object embodiment | Output | Final format and modality of the system’s output. |
| AI inquiry and response | AI technologies | AI models or methods employed. |
| Data processing | Data pre-/post-processing for inquiry and inference. | |
| Temporality | Temporality (e.g., real-time, delay) of AI usage. | |
| Aesthetic experience | Sensory, perceptual, and cognitive experience | Described the audience’s engagement with the artwork, encompassing sensory and perceptual interactions as well as the cognitive meanings embedded in the work’s expressive modes. |
Table 3.
AI-based Interactive Art Installations Selected for Case-Study Analysis.
| Artist(s) | Artwork title | Year |
|---|---|---|
| Scott Allen | Unreal Pareidolia -shadows- | 2023 |
| Weidi Zhang and Rodger Luo | ReCollection | 2022–2024 |
| Varvara Guljajeva and Mar Canet | Visions of Destruction | 2023 |
Table 4.
Structured Summary of the Analytical Approach.
| Analysis structure | Key design elements | Analysis elements |
|---|---|---|
| Research framework | Audience participation | Input |
| Art object embodiment | Output | |
| AI inquiry and response | a) AI technologies b) Data processing c) Temporality |
|
| Aesthetic experience | Sensation, perception, cognitive experience | |
| Key points | a) Key findings b) Results c) Applications |
Table 5.
4A Framework-based Analysis of Unreal Pareidolia -shadows- (2023) by Scott Allen.
| Key design elements | Analysis elements | Analysis of the work |
|---|---|---|
| Audience participation | Input | Everyday items and toys arranged by the audience. Shadows of the items/toys projected on a wall. |
| Art object embodiment | Output | Generated caption, overlaid on the wall. Final AI-generated image. |
| AI inquiry and response | AI technologies | Image-to-image generation using Stable Diffusion. Image-to-text (caption) generation using BLIP-2. |
| Data processing | Projected shadow image (shadowgraph) captured by a camera and used as an input image for BLIP-2 and Stable Diffusion. Generated caption describing the shadowgraph; used as a prompt for Stable Diffusion. |
|
| Temporality | The interaction is procedural, as the participant must press a button to confirm their input, after which captions and images are generated and displayed sequentially. |
Table 6.
4A Framework-based Analysis of ReCollection (2022–2024) by Weidi Zhang and Rodger Luo.
| Key design elements | Analysis elements | Analysis of the work |
|---|---|---|
| Audience participation | Input | Audience’s speech (personal memory) whispered into a microphone. |
| Art object embodiment | Output | Generative image sequences displayed as evolving visual compositions. Real-time animation effects to support perceived immediacy. |
| AI inquiry and response | AI technologies | Speech recognition (speech-to-text) using Whisper. Text auto-completion using a fine-tuned GPT-4 model. Text-to-image generation using Stable Diffusion. |
| Data processing | Text prompt generated via speech-to-text conversion and optional narrative completion. | |
| Temporality | Near real-time: System responds with low latency as speech recognition, speech-to-text conversion, prompt processing, and image display occur sequentially but quickly. |
Table 7.
4A Framework-based Analysis of Visions of Destruction (2023) by Varvara Guljajeva and Mar Canet.
Table 7.
4A Framework-based Analysis of Visions of Destruction (2023) by Varvara Guljajeva and Mar Canet.
| Key design elements | Analysis elements | Analysis of the work |
|---|---|---|
| Audience participation | Input | Viewer’s gaze and head position, captured via an eye tracker. |
| Art object embodiment | Output | Generative landscapes, either continuously updated with smooth animation or locally modified via AI inpainting. |
| AI inquiry and response | AI technologies | Text-to-image generation using Stable Diffusion (with preset landscape prompt sets). Latent-space interpolation used to create smooth animations. Image inpainting using Stable Diffusion, guided by gaze-derived masks and predefined prompts. |
| Data processing | Aggregated gaze mask, computed from gaze fixations with a slight delay. | |
| Temporality | Real-time gaze tracking; near real-time inpainting updates; continuous image generation when no gaze is detected. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.