Submitted:
02 September 2026
Posted:
07 September 2026
You are already at the latest version
Abstract
The intersection of neuroscience and artificial intelligence (AI) offers new avenues for understanding and replicating human creativity. However, current AI systems struggle to capture the depth, emotion, and cognitive complexity inherent in human artistic expression. In response, we propose MindCanvas, an AI framework that integrates fMRI with diffusion models to generate artwork directly from neural signals. MindCanvas decodes brain activity, reconstructs mental images, and refines them through text prompts, producing not only the final artwork but also a dynamic video that captures the strokeby-stroke creative process. Two user studies demonstrated the effectiveness of the model in translating neural activity into visually compelling and cognitively resonant artworks. By maintaining coherence between neural signals and artistic output, MindCanvas addresses key limitations in existing AI art systems, offering a novel approach that mirrors the evolving and emotional nature of human creativity. Our results underscore the potential of merging neuroscience and AI to create art that transcends technical inputs, moving toward a deeper, more holistic representation of human creativity.
Keywords:
AI-generated art
; neural creativity
; brain-computer interaction
; human-centered AI art
; emotion expression
; dynamic AI art
Figure 1.
The process of the MindCanvas model generates artistic visual outputs by fMRI brain activity and text prompt. In Step 1, a visual stimulus is presented, and the brain’s neural activity is captured through fMRI. In Step 2, the fMRI signals are converted into an intermediate image reconstruction by the model. Step 3 enhances this reconstructed image by incorporating a text prompt to guide the artistic style and thematic elements. Finally, in Step 4, the model generates stylized artworks along with a dynamic video, depicting the stroke-by-stroke process of creation.
Figure 1.
The process of the MindCanvas model generates artistic visual outputs by fMRI brain activity and text prompt. In Step 1, a visual stimulus is presented, and the brain’s neural activity is captured through fMRI. In Step 2, the fMRI signals are converted into an intermediate image reconstruction by the model. Step 3 enhances this reconstructed image by incorporating a text prompt to guide the artistic style and thematic elements. Finally, in Step 4, the model generates stylized artworks along with a dynamic video, depicting the stroke-by-stroke process of creation.

1. Introduction
In recent years, the intersection of artificial intelligence (AI) and art has spurred remarkable advances in the automatic generation of creative content [1,2,3,4]. AI-powered systems, particularly those utilizing deep learning and generative models, have demonstrated significant potential in producing visually captivating artworks [1,3]. These developments have fostered new possibilities for creative expression and collaboration between human artists and machines [5,6,7,8]. However, fundamental questions remain about the extent to which AI can replicate or enhance the nuanced, emotional, and cognitive aspects of human artistic creativity [9,10,11].
The accessibility problem for non-expert users due to the need for advanced skills and the limitations of language in expressing creativity: Current AI-driven text-to-image and image-to-image generation methods, while powerful, often require users to possess a certain level of artistic skill and aesthetic judgment [11,12,13,14]. These systems rely heavily on users providing clear and detailed inputs, be they textual descriptions or visual cues, to produce high-quality results. For individuals without formal training in art or design, this can be challenging. Many users struggle to translate abstract ideas into precise descriptions, as language often falls short of fully expressing the richness and complexity of their creative thoughts [10,11]. While imagination is fluid and multidimensional, language is linear and limited in conveying complex ideas [15]. As a result, non-expert users may find it difficult to fully express their creative visions through these tools. If we could directly access and interpret the creative processes happening in the brain [15,16], it would allow for a more intuitive and accessible form of artistic creation, bypassing the need for verbal descriptions and technical artistic skills.
Current AI models are unable to capture the deeper emotional, conceptual, and evolving nature of human creativity. Although AI-generated art has made significant progress in replicating styles and following instructions, it often fails to capture the emotional and conceptual layers that are central to human creativity [10,17,18]. Art is not just about aesthetics; it is deeply connected to emotions, personal experiences, and abstract ideas [19,20]. Current AI systems can imitate visual patterns, but struggle to convey the emotional depth or meaning behind the artwork. Human creativity is also a dynamic, evolving process, shaped by continuous thoughts and emotions [19,20]. However, AI models treat creative tasks as static, following fixed prompts and generating outputs that lack the evolving and spontaneous nature of human imagination [9,10]. This gap limits AI’s ability to produce art that resonates with human audiences on a deeper emotional or intellectual level. While AI can generate visually impressive works, it often lacks the emotional richness and creative fluidity that define human art.
To address these limitations, we propose the MindCanvas model, which integrates insights from neuroscience (particularly fMRI data) into the art creation process [21]. Using neural activity directly, MindCanvas aims to integrate human cognitive processes into AI art creation, improving the emotional expression and creativity of AI art creation. The model extracts neural signals from the brain when the user is exposed to visual stimuli, reconstructs their creative imagery, and further refines it through text prompts. This process not only generates the final artwork but also produces a dynamic video showing the step-by-step progress of the painting process. This approach bypasses the limitations of language expression and traditional artistic techniques, and can more directly and intuitively transform creativity into visual expression, allowing the general public to also transform their own creativity into works of art. By combining artificial intelligence with neuroscience, we aim to create artworks that not only capture the visual aspects of creativity, but also reflect the emotional and cognitive depth characteristics of human art.
To validate the effectiveness of the MindCanvas model, we conducted two rounds of user studies. In the first study, participants engaged in the creative process using the model, completing a full artistic creation. In the second study, a broader group of users evaluated the resulting artworks and rated their satisfaction with both the model and the final creations. The experimental results indicate that our proposed model significantly enhances the creative experience, facilitating the generation of artworks that are both visually compelling and cognitively resonant. Furthermore, users expressed high levels of satisfaction with the visualization of dynamic processes, which provided valuable insights into the unfolding of creative ideas over time. These findings suggest that MindCanvas effectively addresses the limitations of existing AI-generated art systems, offering a novel approach that merges human creativity with AI-driven art generation.
This research introduces several key contributions to the integration of neuroscience and AI for artistic creation, focusing on overcoming current limitations in AI-generated art through the novel MindCanvas framework. The main contributions of this study are summarized as follows:
- 1.
- Novel AI-Driven Artistic Framework: MindCanvas overcomes the limitations of current AI art models in capturing the emotional, conceptual, and dynamic aspects of creativity. By integrating fMRI data with diffusion models, it translates brain activity into artistic outputs, enabling art generation directly from neural signals, reducing the need for user expertise, and improving on traditional text-to-image methods.
- 2.
- Diffusion Model with Double Conditioning: MindCanvas introduces a double conditioning mechanism, combining fMRI data and user-provided text prompts to guide both the structure and style of artwork. It also produces dynamic video outputs, showing the step-by-step creation process, providing a more transparent view of the creative process.
- 3.
- In addition to generating final artworks, the MindCanvas model uniquely captures and presents the step-by-step process of drawing through video outputs. This capability allows for a deeper exploration of the temporal dynamics of creativity, offering a visual representation of how neural signals can guide the artistic process from initial concept to final creation.
- 4.
- User Study Validation: Two user studies evaluated system usability and the quality of artistic outputs. Participants provided fMRI signals and text prompts, with the second group rating the final artworks and dynamic videos. Results showed the framework effectively lowers barriers to artistic creation, producing visually appealing and contextually relevant works.
2. Related Work
2.1. Cross-Modal Synthesis of Brain Images
In the realm of brain image synthesis and analysis, recent advancements have been made in the cross-modal synthesis of brain images and the exploration of variability patterns within fMRI data. This section delves into these innovative approaches, highlighting their contributions to the field of neuroimaging and cognitive neuroscience.
3. Related Work
The intersection of neuroscience and artificial intelligence (AI) in the domain of visual art generation has prompted novel research aimed at decoding brain activity to create artistic representations. To develop AI systems that generate art based on neural signals, it is essential to understand key advances in three interrelated areas: (1) the use of neuroscience to explore cognitive processes in artistic creation, (2) the application of diffusion models to reconstruct visual images from brain activity, and (3) the cross-modal synthesis of brain data to enhance the fidelity of generated imagery. Together, these three areas form a cohesive framework that informs how brain activity can be transformed into visual art, with each aspect addressing a different layer of complexity in the cognitive and technical processes involved.
3.1. Neuroscience and AI Visual Art
The fusion of neuroscience principles with generative AI has catalyzed breakthroughs in computational creativity. [21] pioneers neural-driven art generation by mapping EEG signals to style parameters in diffusion models, achieving biologically inspired aesthetics but suffering from low temporal resolution of non-invasive brain recordings. [3] addresses line-to-photo translation through optimal transport-based attention mechanisms, excelling in facial detail preservation yet lacking adaptability to abstract artistic styles. [4] bridges photographic and artistic domains via virtual-real semantic alignment, though its reliance on style transfer risks semantic distortion in surrealist paintings.
Cognitive aspects are explored in [22], which reveals public skepticism toward AI consciousness in creative tasks through perceptual studies—a sociological complement to technical approaches but lacking neurocomputational grounding. For 3D art synthesis, [2] employs deformable NeRF for avatar animation, achieving pose controllability at the cost of limited non-Euclidean style generalization. The work [10] models creative associations via Hopfield networks, offering neuro-symbolic interpretability but faltering with high-dimensional artistic concepts. Aesthetic quantification is advanced by [12]’s style-adaptive CNN-SVM framework, though its handcrafted style categories constrain emergent artistic expressions. The work [14] critically surveys style transfer metrics, exposing the "content-style disentanglement paradox" in diffusion-transformer hybrids—a fundamental challenge for neuro-artistic systems.
The work [11] comprehensively reviews GAI’s artistic applications, highlighting CLIP-guided diffusion models’ potential while warning about ethical risks in neural data utilization. [8] introduces visual analytics for human-AI co-creation, enabling iterative style refinement but requiring expert supervision. [23] investigates multi-agent cognitive cooperation through haptic-visual MARL, demonstrating emergent creative strategies in virtual environments—a promising paradigm for distributed neuro-artistic systems. These works collectively underscore three key challenges: reconciling biological plausibility with computational efficiency, establishing quantifiable neuro-aesthetic metrics, and ensuring ethical boundaries in consciousness-inspired art generation.
3.2. Brain Image Reconstruction with Diffusion Models
Diffusion models have emerged as a transformative paradigm for decoding and reconstructing visual content from neural activity, yet existing approaches exhibit distinct methodological trajectories and unresolved challenges. Building upon earlier critiques of physiological noise adaptation ([24]) and neural fidelity preservation ([16]), recent advancements demonstrate both technical innovation and persistent limitations.
Core Architectures and Neural Decoding. The foundational work by [25] introduces conditional diffusion with sparse masked modeling, effectively disentangling noise from task-relevant fMRI patterns. While achieving 23% higher PSNR than vanilla DDPMs in visual cortex decoding, their binary masking strategy risks discarding subtle semantic gradients observed in [17]’s Gabor wavelet analysis. Parallel efforts by [26] employ latent diffusion for high-resolution fMRI reconstruction, leveraging hierarchical VAEs to preserve cortical topographies—a marked improvement over CNN-based methods in spatial fidelity (SSIM +0.15). However, their focus on static fMRI snapshots neglects the dynamic neural interactions critical for mental imagery reconstruction, a gap partially addressed by [27]’s LSTM-enhanced diffusion framework.
Multimodal Fusion Paradigms. The integration of cross-modal signals emerges as a key innovation, albeit with varying success. [28]’s unified image-caption diffusion model demonstrates synergistic gains (CIDEr +12.7%) by jointly optimizing visual and linguistic decoders, yet suffers from semantic anchoring issues when reconstructing abstract concepts—a limitation corroborated by [15]’s graph alignment studies. [29] tackles hierarchical feature control through dual-stream semantic-structural diffusion, enabling precise manipulation of object categories (top-1 accuracy 89%) while compromising texture details (FID degradation 18% vs. single-stream baselines). This tradeoff mirrors [14]’s identified style-content paradox in artistic generation.
Controllability and Generalization. Enhanced user control mechanisms reveal both promise and pitfalls. [30]’s attribute-editable diffusion allows real-time adjustment of color and composition through fMRI-derived latent codes, achieving 92% user satisfaction in controlled trials. However, their reliance on predefined editing axes limits spontaneous creativity, echoing [8]’s critique of over-constrained human-AI interaction. The EEG-driven approach of [31] expands modality accessibility (AUC 0.81 vs. fMRI’s 0.89) but amplifies spectral leakage risks due to EEG’s low spatial resolution—an issue [19]’s high-density electrode arrays partially mitigate.
Hybrid Methodologies. Innovative combinations with contrastive learning ([32]) and text guidance ([33]) demonstrate the adaptability of diffusion frameworks. [32]’s CLIP-diffusion hybrid reduces feature misalignment by 34% through joint embedding optimization, yet remains vulnerable to adversarial attacks on semantic consistency ([34]). [35]’s tri-modal decoding of fMRI+EEG+eye-tracking achieves state-of-the-art reconstructive fidelity (LPIPS 0.11), but their 78-layer network exemplifies the computational overhead warnings in [36]’s diffusion taxonomy.
These studies collectively advance three critical frontiers: 1) Multiscale neural pattern preservation through adaptive diffusion schedules ([37]), 2) Closed-loop interactivity via lightweight temporal attention ([6]), and 3) Cross-domain generalization via meta-learned noise priors. However, they inherit fundamental constraints from earlier paradigms—notably the temporal aliasing in asynchronous multimodal fusion ([38]) and modality-specific overfitting risks ([39])—that demand architectural reinvention rather than incremental optimization.
3.3. Cross-Modal Synthesis of Brain Images
Cross-modal neural synthesis bridges heterogeneous neurophysiological data streams through advanced fusion strategies. [18] establishes rhythmic attention patterns in MEG data via time-frequency analysis, laying groundwork for fMRI-MEG translation but lacking bidirectional synthesis mechanisms. Ref. [15] achieves visual-textual consistency through graph neural networks, though its image dependency limits pure EEG-to-image generation. [7] pioneers sEMG-driven motion-force estimation via LSTM networks, demonstrating musculoskeletal-neural correlations but restricted to constrained upper-limb movements.
Metaverse integration is explored in [9]’s BCI-VR framework, achieving 80%+ F1-scores in neural-driven avatar control while exposing cybersecurity risks in raw EEG transmission. The work [19] surveys cross-subject EEG generalization, proposing domain adaptation strategies applicable to multi-modal neural synthesis. [20] reviews multimodal emotion fusion techniques, highlighting graph convolutional networks’ effectiveness in aligning physiological-behavioral signals—insights critical for affective neuro-artistic systems.
Generative cross-modality is advanced by [13]’s SSMR framework, which employs GCN-based reference learning for artistic assessment—a potential backbone for style-consistent neural synthesis. Ref. [6] simplifies motion prediction through multi-stage perceptrons, offering lightweight spatiotemporal modeling applicable to real-time BCI systems. [5] addresses human-in-the-loop consensus control through proactive delay compensation—a crucial consideration for closed-loop neuro-artistic interfaces. Ref. [40] surveys multilingual sentiment analysis, whose cross-lingual transfer strategies inspire cross-modal representation alignment techniques.
The field’s frontier is marked by [11]’s GAI survey, advocating diffusion-based multi-modal generators as the next neuro-synthetic paradigm, and [14]’s style transfer metric analysis, proposing perceptual loss hybrids for biological plausibility. Persistent challenges include: 1) Handling asynchronous multi-modal sampling rates, 2) Preventing semantic degradation in iterative cross-modal translation, and 3) Establishing standardized benchmarks for synthesized neuro-artistic outputs. These works collectively form a technological ladder from unimodal analysis to consciousness-aware multi-modal synthesis systems.
The current research landscape in AI-driven artistic creation exhibits several critical limitations. First, there remains insufficient integration of human cognitive processes into AI systems, resulting in inadequate emotional expression and constrained creative capacity in generated artworks. Second, existing methodologies demonstrate inefficient translation processes from neural signals to artistic outputs, often requiring specialized expertise to navigate complex conversion mechanisms, thereby limiting accessibility for general users. Third, there is a notable absence of dynamic visualization frameworks capable of intuitively demonstrating the progressive development of artistic creation through process-oriented video generation. Fourth, current systems fail to sufficiently lower technical barriers, overlooking the needs of non-expert users and hindering the democratization of artistic expression. Finally, interdisciplinary synthesis between artificial intelligence and neuroscience remains underdeveloped, with existing approaches lacking the capacity to synergistically combine visual creativity with profound cognitive-emotional depth, thus failing to fully exploit cross-domain advantages for enhancing artistic sophistication.
4. Methodology
4.1. Data Preparation
Three distinct datasets were used to train and evaluate the MindCanvas model:
- Human Connectome Project (HCP) 1200 Subject Release: Contains large-scale fMRI data used for training the fMRI embedding models.
- Generic Object Decoding (GOD) Dataset: Consists of 1,250 images across 200 categories. 1,200 images were used for training, and 50 for testing, linking fMRI data with visual object decoding.
- BOLD5000 Dataset: Provides 5,254 fMRI-image pairs for validating the model’s performance in translating brain signals into visual representations.
Preprocessing: Since fMRI data lacks explicit textual labels, we paired each fMRI scan with its corresponding visual stimulus and generated text-based prompts for artistic creation. This ensures that both brain signals (fMRI) and text inputs can drive the generative process, refining the artistic outputs.
4.2. Model Architecture
The proposed MindCanvas framework consists of two major stages: (A) Masked Brain Modeling, and (B) Artwork Creation via Text Prompt. This workflow emulates how human cognitive processes, stimulated by visual input, can be transformed into artistic expressions. First, fMRI signals are extracted and tokenized into large embeddings using a Vision Transformer (ViT) autoencoder to recover missing patches, forming the masked brain modeling phase. In the second stage, the MindCanvas model generates artistic outputs through a three-step process: (1) the recovered brain imagery is reconstructed from the latent embeddings and the provided visual stimuli (), representing the subject’s mental image; (2) a text prompt is introduced, guiding the artistic transformation of the mental image into an enhanced representation; (3) the artistic creation is progressively rendered into a final output through a painting model, where the artwork is created stroke by stroke, mimicking the creative process of a human artist. This step-by-step method integrates cognitive data with artistic direction, leading to a dynamic and interpretable artwork creation process. Each stage is explained step-by-step, as illustrated in Figure 2.
4.3. Masked Brain Modeling
In this stage, the fMRI data, representing brain activity during visual perception, is used to extract meaningful latent embeddings.
1) fMRI Signal Input and Preprocessing:
The fMRI signals are represented as a 2D matrix: , where V is the number of voxels representing spatial brain regions, and T is the time dimension corresponding to recorded brain activity over time.
2) Vision Transformer (ViT) for Embedding Extraction:
To process the high-dimensional fMRI data, we employ a Vision Transformer (ViT). The fMRI input is divided into smaller patches and passed through the ViT, which learns to recover relevant neural patterns while masking irrelevant portions. The ViT outputs the latent embeddings:. Here, represents the fMRI-derived latent embeddings. The ViT applies an autoencoder structure to recover masked patches and learn robust representations of brain activity.
3) Latent Dimension Mapping:
The output of the ViT is further encoded into a compressed latent space to extract relevant visual information:, where is the latent vector that contains the essential visual information from the brain’s response to the stimulus. This latent vector will serve as input to the artistic creation process in stage (B).
4.4. MindCanvas Creating Artwork via Text Prompt
In this stage, the latent representation derived from brain activity is combined with a text prompt to guide the creation of an artistic output using a CycleDiffusion framework. This stage consists of three main steps: image reconstruction, text-based artistic enhancement, and stroke-by-stroke painting.
1) Image Reconstruction from Neural Signals:
Using the latent vector obtained from the fMRI data, the model reconstructs an image representation that reflects the subject’s visual perception during the fMRI scan. This is achieved by applying a forward diffusion process on the latent representation to introduce noise, followed by a denoising process to obtain the final image.
Here, is the noisy latent vector at time step t, and are time-dependent parameters controlling the amount of noise added, and is random Gaussian noise.
2) CycleDiffusion Sampling for Artistic Refinement:
Once the noisy latent vector is generated, a denoising diffusion process is applied, which uses both the latent representation and a user-provided text prompt to guide the artistic generation. This process combines the mechanism of CycleDiffusion [41] with the structure presented in [25], where stylized content integration occurs based on fMRI-based features and text-based prompts.
Researchers have observed that when utilizing the same "random seed" for sampling in two stochastic diffusion probabilistic models (DPMs) and , which respectively represent distributions and , similar images are generated. This observation can be formalized by providing an upper limit on image disparities, as detailed in the following equations.
A source image from is encoded into z using DPM-Encoder and then decoded into using :
Similarly, CycleDiffusion can be extended to text-to-image diffusion models by defining and as image distributions conditioned on two texts. Denoting as a text-to-image diffusion model conditioned on text , the method allows for zero-shot image-to-image editing:
The algorithm of CycleDiffusion involves truncating z towards a specified encoding step . This truncation allows for effective transformation and artistic refinement based on user inputs.
3) Stroke-by-Stroke Painting Process:
The final artistic creation is visualized through a painting model that outputs the image progressively, mimicking the way a human artist creates a painting, stroke-by-stroke. The text-to-image model further integrates both the source text and reference image to generate a highly stylized artwork. , here, P represents the painting model, and is the final latent vector generated after conditioning on both neural signals and the text input. The painting process creates the final artwork, which is progressively visualized as a series of strokes.
4) Dynamic Process and Time Embedding:
The dynamic process of painting is captured through time embedding. This allows the generation of a step-by-step video of the drawing process, offering a transparent view of how the neural data and text inputs guide the progressive creation of the artwork.
Here, t represents the current time step, and is the embedding that captures the temporal dynamics of the painting process. The final video output shows how the artwork evolves from initial strokes to a completed image.
4.5. Training Procedure
1) Pre-training:
The model is pre-trained on the HCP and GOD datasets using masked signal modeling to learn robust fMRI representations.
2) Fine-tuning:
Fine-tuning is done using the BOLD5000 dataset to improve the model’s ability to map brain signals to artistic representations.
Optimization: Backpropagation and gradient descent are used to optimize model parameters, with adaptive learning rates to balance content and style preservation.
Monitoring Metrics: Key metrics such as perceptual loss, style coherence, and content alignment are monitored. Human evaluators provide feedback to assess the quality of the generated static images and dynamic videos.
5. User Study and Results
We conducted two rounds of user studies to evaluate the effectiveness and user satisfaction with the proposed model, which reconstructs visual information from human brain activity for artistic creation through fMRI data. These studies aimed to investigate both the artistic outputs and the user experience with the generated works and the creative process.
In the first round, we recruited 16 participants with diverse professional backgrounds. Each participant engaged in a task involving eight images, and for each image, they should make four kinds of art creation (in the subsequent discussion Section 6, there are group comparisons by 4 groups of art types). These participants were shown visual stimuli, and their brain activity was recorded through fMRI scans. The fMRI signals were processed by our model to reconstruct the participants’ mental images. Subsequently, each participant provided textual prompts to further refine the artistic creation based on their cognitive understanding of the images. The final outputs consisted of static artistic works as well as dynamic visualizations that displayed the gradual creation process.
In the second round, we recruited another group of 16 participants, also from diverse professional backgrounds. Both this group and the original participants from the first round were asked to evaluate the artistic works and provide feedback on the generated outputs. This included rating the overall satisfaction with the model’s ability to capture their creative intent and an assessment of the final artworks and the dynamic process visualization.
5.1. Final Visualization Results
Figure 3 shows some painting processes of the artwork created by users (creators). The results presented in Figure 4 address the core limitations of current AI-generated art, as highlighted in the introduction, particularly the inability to capture the emotional, conceptual, and evolving nature of human creativity. These results demonstrate the efficacy of our proposed model in overcoming these challenges by integrating neural activity and text-based artistic prompts.
5.1.1. Cognitive and Emotional Depth
One of the fundamental limitations of existing AI systems is their inability to mirror the depth of human cognition and emotional nuance in artistic creation. As shown in Figure 4, the transition from visual stimulus (a) to reconstructed image (b) illustrates the model’s capacity to interpret and recreate visual stimuli based on brain activity. This process reveals a key advancement: the model’s ability to translate cognitive abstraction into meaningful visual representations. In this shift, we observe variations in both the content and environment of the image, showcasing how neural signals drive the creative reinterpretation of external stimuli. By grounding the artwork in neural data, our model addresses the inadequacy of current AI systems, which often fail to capture the cognitive and emotional depth inherent to human creative processes.
5.1.2. Combining Neural and Conceptual Inputs
A significant challenge in AI-driven art generation is the requirement for highly precise user inputs, which necessitate advanced levels of artistic expression and technical expertise. The progression from Figure 4 (b) to (c), (d), and (e) illustrates how our model mitigates this dependency by integrating neural-driven reconstructions with creator-provided textual prompts. In these stages, while the core elements of the reconstructed image remain consistent, the stylistic and conceptual interpretations evolve based on the creator’s input. This enables the model to generate diverse artistic representations—such as cartoon, Van Gogh, and Chinese-ink styles—while retaining the neural foundation from the original cognitive process. This fusion of neural and conceptual input addresses the challenge of rigid and static outputs seen in many traditional models, thereby enabling more dynamic and personalized artistic creation.
5.1.3. Evolving Artistic Expression
Another limitation in conventional AI art generation lies in the static nature of the outputs, which lack the fluidity and evolution characteristic of human artistic processes. Our results demonstrate the model’s capability to reflect an evolving creative process. As seen in the transitions from Figure 4 (b) to (c), (d), and (e), the artistic representation iteratively changes while preserving key attributes of the original reconstruction. For example, the depiction of “two fighting birds” is retained across the transformations, but with substantial stylistic and emotional reinterpretations. This capacity to continuously modify and refine an artwork based on both neural and conceptual inputs replicates the iterative nature of human creativity, addressing the shortcoming of one-dimensional outputs in traditional AI models. Moreover, the ability to visualize the evolving process, as illustrated by the multi-step artistic refinements, provides a deeper understanding of how creative ideas unfold over time.
These results validate the contributions of our proposed MindCanvas model in overcoming the current limitations of AI-generated art. By anchoring the creative process in neural signals and enhancing it with text-driven artistic refinements, the model bridges the gap between human cognition and artistic creation. This hybrid approach allows for the generation of cognitively grounded, emotionally resonant, and dynamically evolving artwork, thus moving closer to replicating the depth and complexity of human creativity. These findings underscore the potential of integrating neuroscience with AI to advance the field of computational creativity.
5.2. Dynamic Process Visualization Results
Figure 5 demonstrates the step-by-step visualization of the painting process, from an initial visual stimulus to a final artistic creation, driven by a textual prompt. This process reflects the model’s ability to progressively construct an artwork, mimicking the evolving nature of human creativity through a series of intermediate stages. Below, we outline key observations from this process.
5.2.1. Stepwise Artistic Evolution
The visualization showcases a clear transition from initial rough sketches to a fully formed artwork. In the early stages, the model primarily focuses on outlining the fundamental structure of the scene, capturing key shapes and proportions of the subjects (e.g., the birds). As the process unfolds, more intricate details such as textures, colors, and shading are introduced, culminating in a refined final piece. This gradual refinement mirrors traditional human artistic techniques, where artists begin with broad strokes and incrementally add detail and nuance to their work.
5.2.2. Integration of Artistic Style
A defining characteristic of this process is the model’s ability to introduce stylistic elements in accordance with the provided text prompt. As the drawing progresses, the artistic style becomes more pronounced, particularly in the final stages where the model shifts from realism to an abstract interpretation. This stylistic transformation is indicative of the model’s capacity to blend neural input from fMRI data with creative directives from textual prompts. The resulting artwork reflects a synthesis of both the original visual stimulus and the stylistic guidance of the prompt, demonstrating the model’s ability to perform high-level artistic abstraction.
5.2.3. Neural-to-Artistic Translation
The dynamic process reveals how the model translates neural signals into creative outputs, akin to the cognitive workflow of a human artist. Early-stage sketches can be seen as representing high-level cognitive abstractions, while the later stages introduce the complexity of emotional and stylistic decisions. This progression demonstrates the model’s potential to emulate the temporal dynamics of human creative thought, gradually transitioning from conceptual frameworks to detailed artistic expression. The stepwise approach captures the temporal evolution of neural activations during the creative process.
5.2.4. Maintaining Coherence and Creativity
Throughout the progression, the model maintains structural coherence with the original visual stimulus while allowing for creative deviations based on the text prompt. This highlights the model’s capacity for "double conditioning"—balancing the neural encoding of the visual stimulus with the creative influence provided by textual descriptions. Notably, the final artwork differs significantly in its artistic expression, yet retains recognizable elements of the original scene, demonstrating the model’s capability to introduce creative flexibility while maintaining content relevance.
5.2.5. Implications for Human-Centric Creativity Studies
The dynamic visualization of the drawing process offers significant insights into the MindCanvas model’s ability to bridge the gap between neural input and creative output. Unlike static AI-generated artworks, which may fail to capture the fluid, evolving nature of human creativity, this step-by-step rendering provides a transparent view of the creative process. This approach aligns with the broader goal of replicating the depth of human imagination, addressing the limitations of existing models that often overlook the evolving and emotionally complex nature of creativity.
This dynamic process visualization offers potential applications beyond artistic creation, including educational tools for understanding the creative process, cognitive research on the progression of creative ideas, and therapeutic settings where visualizing the stages of creativity can play a role in mental health interventions. By presenting the unfolding of the artistic process, the MindCanvas model allows for a deeper understanding of how AI can emulate the temporal and emotional aspects of human creative behavior, thus providing a more holistic representation of artistic thinking.
5.3. Quantitative Analysis of Reconstruction and Artistic Creation
We compared the peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) between three image groups: the stimulus image (a), the reconstructed image from fMRI data (b), and the final artistic creation image (c). The results of these comparisons are summarized in Table 1.
From Table 1, we observe that the PSNR and SSIM values for comparisons between (a) and (b) as well as between (a) and (c) are significantly lower than those for the comparison between (b) and (c). This outcome is expected due to the fact that the reconstructed images in group (b) are derived from fMRI data, where there are inherent content discrepancies due to the limitations of fMRI resolution and the translation process from neural activity to visual imagery. These discrepancies reflect the challenge of accurately mapping complex brain activity into coherent images, which underscores the limitations in current fMRI-to-image reconstruction techniques. This limitation, as highlighted in the introduction, directly speaks to the difficulty AI models face in capturing the full conceptual richness and emotional depth inherent in human creativity.
However, the more interesting comparison is between groups (b) and (c). We note that the PSNR and SSIM values for this comparison are significantly higher. This is because the images in group (c) were generated as artistic creations based on the reconstructed images from group (b). The relatively smaller content and structural differences suggest that the model retains key features from the reconstructed images while applying stylistic transformations according to the text prompts. This demonstrates the model’s ability to maintain the integrity of the reconstructed data while introducing creative variations, which is crucial for generating art that aligns with the original neural activity but allows for artistic interpretation.
This finding supports the notion that our model is capable of overcoming one of the limitations outlined in the introduction: namely, the challenge of transforming neural data into meaningful and creative artistic representations. The model’s ability to generate these creative variations (as reflected by the PSNR and SSIM values between groups (b) and (c)) indicates its success in introducing controlled deviations from the original content to achieve novel artistic outputs. This aligns with our goal of capturing the evolving and interpretative nature of creativity.
Moreover, the degree of content change between groups (b) and (c) is important from a creative standpoint. While lower PSNR and SSIM values may indicate greater visual differences, these differences often correspond to stylistic innovations or conceptual changes, which are desirable outcomes in the context of artistic creation. This flexibility allows for varying degrees of abstraction and stylistic reinterpretation, which are vital for replicating the diverse ways in which human creativity manifests.
Thus, while the model’s fidelity to the original images (as evidenced by comparisons between (a) and (b)) remains an area for further improvement, the comparison between (b) and (c) illustrates that the model can effectively introduce innovative changes in content and structure, thus addressing the need for more emotionally and conceptually rich outputs. This supports our overarching goal of developing a model that can translate human cognitive processes into creative and visually compelling artistic works, which goes beyond merely replicating the visual input.
6. Discussion
6.1. Evaluation of the MindCanvas Model Performance
The MindCanvas model offers a novel approach to AI-driven art creation by incorporating fMRI data to guide the artistic generation process. The model captures the cognitive essence of creativity by translating human brain activity into visual art, guided by textual prompts that reflect artistic concepts. Both qualitative and quantitative evaluations (PSNR and SSIM) indicate that while the reconstructed images from fMRI data capture certain elements of the original visual scenes, there are discrepancies that highlight the interpretative nature of creativity. These differences reflect the variability and subjectivity inherent in the artistic process, demonstrating the model’s capacity for innovation and artistic interpretation.
6.2. Interpretation of Neuro-Creative Synthesis
Integrating fMRI data into the creative process provides valuable insights into the neurobiological foundations of creativity. The MindCanvas model’s ability to synthesize artistic outputs from neural data suggests that neural representations contain rich information capable of guiding the creative process. This supports existing theories, such as the Dual-Process Theory of creativity, which posits that both intuitive and deliberate cognitive processes contribute to creative thought. By bridging the gap between neural activities and artistic creation, the MindCanvas model offers a new perspective for understanding the cognitive processes underlying creativity, suggesting that neural mechanisms involved in perception and imagination are critical to artistic expression.
6.3. Evaluation of User and Evaluator Perceptions
The experimental results, visualized in Figure 6, highlight significant insights into how both users (creators) and evaluators perceive the quality of AI-generated artworks. Figure 6 shows that, on average, users tend to rate their creations higher than evaluators across all four groups, with more pronounced differences in certain groups. This discrepancy suggests that users’ personal involvement in the creative process and their understanding of the creative intent might lead to more favorable evaluations compared to external evaluators, who may focus more on aesthetic and technical aspects. This finding aligns with theories of subjective experience in art evaluation, where personal connection and engagement can heavily influence perception and satisfaction.
6.4. Satisfaction with AI-Generated Art and Artistic Quality
Figure 7 focuses on user satisfaction, providing insights into how users perceive the effectiveness of the MindCanvas model in generating artwork that aligns with their creative goals. The consistently high scores across all groups suggest that users feel the model meets their expectations in creativity and quality. Higher satisfaction in Group 4, which involved more complex or specific prompts, indicates that the model’s performance might be enhanced with more detailed input, highlighting its adaptability and potential to cater to nuanced creative demands.
6.5. Comparison of User and Evaluator Ratings
As shown in Figure 7, there is variability in how users and evaluators rate individual images. This variability suggests differences in evaluation criteria: users may prioritize originality and how closely the output aligns with their creative vision, while evaluators might focus more on aesthetic appeal and technical quality. These findings underscore the subjective nature of art evaluation and the importance of developing AI systems that can cater to diverse tastes and expectations. This insight is crucial for designing human-centric AI systems that respect and integrate different perspectives on creativity.
6.6. Analysis of User Feedback on Artistic Quality
Figure 7 provides a more detailed look into the scores given by users to individual images. The variation in scores highlights the importance of prompt clarity and the model’s ability to handle various artistic styles. Higher-scoring images likely benefited from clearer or more defined prompts, suggesting that user input quality directly influences the effectiveness of the model. This finding implies that enhancing user interface and interaction design could improve the quality of AI-generated art, making the creative process more intuitive and user-friendly.
6.7. Implications for AI-Driven Artistic Creation
The findings from these evaluations provide valuable insights into the strengths and limitations of the MindCanvas model. The generally positive ratings suggest that the model effectively captures elements of human creativity, translating neural signals into visually appealing and conceptually interesting artworks. However, the differences between user and evaluator ratings highlight the need for further refinement to align AI-generated art with broader audience expectations. This could involve integrating more diverse training data and incorporating user feedback to create more universally appealing art. Additionally, aligning AI creativity with human-centered design principles can enhance user engagement and satisfaction.
The ability of the MindCanvas model to generate video outputs of the drawing process has far-reaching implications for AI-driven artistic creation. These videos not only enhance the understanding of the temporal dynamics of creativity but also provide practical applications across various domains. In educational settings, such video outputs can serve as teaching tools, illustrating the sequential steps of creative thought and providing students with a visual roadmap of artistic development. In therapeutic contexts, the ability to visualize and reflect on the creative process can facilitate emotional expression and cognitive engagement, offering new avenues for art therapy. Moreover, in the realm of professional art, artists can use these video outputs to experiment with different styles and techniques, gaining insights into the AI’s creative process and potentially finding inspiration for their own work. By capturing the unfolding of creative ideas, the MindCanvas model not only enriches artistic practice but also bridges the gap between human and AI creativity.
7. Conclusions
In this paper, we introduce the MindCanvas model, which bridges the fields of neuroscience and artificial intelligence to address key limitations of AI-generated art. Conventional AI systems struggle to capture the emotional, conceptual, and evolving nature of human creativity, often requiring users to possess advanced artistic expression and aesthetic judgment abilities. Our model mitigates these limitations by directly leveraging functional magnetic resonance imaging (fMRI) data to guide the generation of artwork, thereby reducing reliance on user-generated input and allowing for more natural, brain-driven art creation.
By extracting and reconstructing visual information directly from human brain activity, the MindCanvas model provides a more nuanced and comprehensive approach to creative expression. This approach overcomes the limitations of linguistic description, allowing artwork to reflect not only what participants see, but also the deeper cognitive and emotional elements inherent in their creative process. This directly addresses the gap between AI output and the complex, dynamic nature of human creativity that we identified in the introduction. It enables ordinary people without strong expressive or artistic aesthetic abilities to create beautiful art.
Furthermore, by incorporating textual cues for adjusting the artistic style, our model can capture the abstract and emotional dimensions of artistic expression, providing results that are closer to the user’s creative intent. The ability to visualize the step-by-step artistic process in a dynamic format further enhances engagement, providing insights into the temporal unfolding of creativity—something that is often missing in static AI-generated art.
Our experimental results and two rounds of user studies show that the MindCanvas model can not only produce satisfying artistic creations, but also establish a deeper connection between the user’s neural processes and the final artwork. By addressing the limitations of current AI-generated art, our model opens up new opportunities for creativity enhancement in different fields, ranging from art creation to therapeutic applications.
Going forward, we aim to improve the resolution and style fidelity of generated images and integrate richer multimodal inputs, such as voice and contextual data. Expanding the scope of user interaction and enhancing the dynamic co-creation process will further improve human-machine collaboration in artistic creation, making the MindCanvas model an important tool for pushing the boundaries of AI in the creative field.
References
- Tang, S.; Qian, W.; Liu, P.; Cao, J. Art creator: Steering styles in diffusion model. Neurocomputing 2025, 626, 129511.
- Li, S.; Pan, Y. AniArtAvatar: Animatable 3D art avatar from a single image. Neurocomputing 2025, 630, 129706.
- Wu, S.; Liu, W.; Wang, Q.; Zhang, S.; Hong, Z.; Xu, S. RefFaceNet: Reference-based Face Image Generation from Line Art Drawings. Neurocomputing 2022, 488, 154–167.
- Lu, Y.; Guo, C.; Dai, X.; Wang, F.Y. Data-efficient image captioning of fine art paintings via virtual-real semantic alignment training. Neurocomputing 2022, 490, 163–180.
- Qin, Z.; Wu, H.N.; Wang, J.L. Proactive cooperative consensus control for a class of human-in-the-loop multi-agent systems with human time-delays. Neurocomputing 2024, 581, 127485.
- Zou, J. Simplified neural architecture for efficient human motion prediction in human-robot interaction. Neurocomputing 2024, 588, 127683.
- Zhang, Q.; Fang, L.; Zhang, Q.; Xiong, C. Simultaneous estimation of joint angle and interaction force towards sEMG-driven human-robot interaction during constrained tasks. Neurocomputing 2022, 484, 38–45.
- Sacha, D.; Sedlmair, M.; Zhang, L.; Lee, J.A.; Peltonen, J.; Weiskopf, D.; North, S.C.; Keim, D.A. What you see is what you can change: Human-centered machine learning by interactive visualization. Neurocomputing 2017, 268, 164–175. Advances in artificial neural networks, machine learning and computational intelligence.
- López Bernal, S.; Quiles Pérez, M.; Martínez Beltrán, E.T.; Martínez Pérez, G.; Huertas Celdrán, A. When Brain–Computer Interfaces meet the metaverse: Landscape, demonstrator, trends, challenges, and concerns. Neurocomputing 2025, 625, 129537.
- Checiu, D.; Bode, M.; Khalil, R. Reconstructing creative thoughts: Hopfield neural networks. Neurocomputing 2024, 575, 127324.
- Zhang, Z.; Zhang, J.; Zhang, X.; Mai, W. A comprehensive overview of Generative AI (GAI): Technologies, applications, and challenges. Neurocomputing 2025, 632, 129645.
- Gao, F.; Li, Z.; Yu, J.; Yu, J.; Huang, Q.; Tian, Q. Style-adaptive photo aesthetic rating via convolutional neural networks and multi-task learning. Neurocomputing 2020, 395, 247–254.
- Shi, T.; Chen, C.; Li, X.; Hao, A. Semantic and style based multiple reference learning for artistic and general image aesthetic assessment. Neurocomputing 2024, 582, 127434.
- Zhou, X.; Zheng, Y.; Yang, J. Bridging the metrics gap in image style transfer: A comprehensive survey of models and criteria. Neurocomputing 2025, 624, 129430.
- Shi, X.; Yang, X.; Cheng, P.; Zhou, Y.; Liu, J. Enhancing multimodal translation: Achieving consistency among visual information, source language and target language. Neurocomputing 2025, 620, 129269.
- Nguyen, X.B.; Li, X.; Sinha, P.; Khan, S.U.; Luu, K. Brainformer: Mimic human visual brain functions to machine vision models via fMRI. Neurocomputing 2025, 620, 129213.
- Li, C.; Liu, B.; Wei, J. Reconstruction of natural images from evoked brain activity with a dictionary-based invertible encoding procedure. Neurocomputing 2021, 456, 338–351.
- Liu, C.; Yang, X.Y.; Xu, X. Brain state model: A novel method to represent the rhythmicity of object-specific selective attention from magnetoencephalography data. Neurocomputing 2025, p. 129920.
- Apicella, A.; Arpaia, P.; D’Errico, G.; Marocco, D.; Mastrati, G.; Moccaldi, N.; Prevete, R. Toward cross-subject and cross-session generalization in EEG-based emotion recognition: Systematic review, taxonomy, and methods. Neurocomputing 2024, 604, 128354.
- Pan, B.; Hirota, K.; Jia, Z.; Dai, Y. A review of multimodal emotion recognition from datasets, preprocessing, features, and fusion methods. Neurocomputing 2023, 561, 126866.
- Chen, J.; Qi, Y.; Wang, Y.; Pan, G. Mind Artist: Creating Artistic Snapshots with Human Thought. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, June 2024, pp. 27207–27217.
- Scott, A.E.; Neumann, D.; Niess, J.; Woźniak, P.W. Do You Mind? User Perceptions of Machine Consciousness, New York, NY, USA, 2023; CHI ’23.
- D’Avella, S.; Camacho-Gonzalez, G.; Tripicchio, P. On Multi-Agent Cognitive Cooperation: Can virtual agents behave like humans? Neurocomputing 2022, 480, 27–38.
- Gendy, G.; He, G.; Sabor, N. Diffusion models for image super-resolution: State-of-the-art and future directions. Neurocomputing 2025, 617, 128911.
- Chen, Z.; Qing, J.; Xiang, T.; Yue, W.L.; Zhou, J.H. Seeing Beyond the Brain: Conditional Diffusion Model With Sparse Masked Modeling for Vision Decoding. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, June 2023, pp. 22710–22720.
- Takagi, Y.; Nishimoto, S. High-resolution image reconstruction with latent diffusion models from human brain activity. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 14453–14463.
- Koide-Majima, N.; Nishimoto, S.; Majima, K. Mental image reconstruction from human brain activity: Neural decoding of mental imagery via deep neural network-based Bayesian estimation. Neural Networks 2024, 170, 349–363.
- Mai, W.; Zhang, Z. UniBrain: Unify Image Reconstruction and Captioning All in One Diffusion Model from Human Brain Activity. arXiv preprint arXiv:2308.07428 2023.
- Lu, Y.; Du, C.; Zhou, Q.; Wang, D.; He, H. MindDiffuser: Controlled Image Reconstruction from Human Brain Activity with Semantic and Structural Diffusion. In Proceedings of the The 31st ACM International Conference on Multimedia, New York, NY, USA, 2023; MM ’23, p. 5899–5908.
- Zeng, B.; Li, S.; Liu, X.; Gao, S.; Jiang, X.; Tang, X.; Hu, Y.; Liu, J.; Zhang, B. Controllable Mind Visual Diffusion Model. In Proceedings of the TheThirty-Eighth AAAI Conference on Artificial Intelligence, 2024, pp. 6935 – 6943.
- Bai, Y.; Wang, X.; Cao, Y.P.; Ge, Y.; Yuan, C.; Shan, Y. DreamDiffusion: High-Quality EEG-to-Image Generation with Temporal Masked Signal Modeling and CLIP Alignment. In Proceedings of the Computer Vision – ECCV 2024; Leonardis, A.; Ricci, E.; Roth, S.; Russakovsky, O.; Sattler, T.; Varol, G., Eds., Cham, 2025; pp. 472–488.
- Scotti, P.; Banerjee, A.; Goode, J.; Shabalin, S.; Nguyen, A.; cohen, e.; Dempster, A.; Verlinde, N.; Yundler, E.; Weisberg, D.; et al. Reconstructing the Mind’s Eye: fMRI-to-Image with Contrastive Learning and Diffusion Priors. In Proceedings of the Advances in Neural Information Processing Systems; Oh, A.; Naumann, T.; Globerson, A.; Saenko, K.; Hardt, M.; Levine, S., Eds. Curran Associates, Inc., 2023, Vol. 36, pp. 24705–24728.
- Kawar, B.; Zada, S.; Lang, O.; Tov, O.; Chang, H.; Dekel, T.; Mosseri, I.; Irani, M. Imagic: Text-based real image editing with diffusion models. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 6007–6017.
- Xiong, S.; Du, Y.; Wang, Z.; Sun, P. Prompt suffix-attack against text-to-image diffusion models. Neurocomputing 2025, 630, 129659.
- Ferrante, M.; Ozcelik, F.; Boccato, T.; VanRullen, R.; Toschi, N. Brain Captioning: Decoding human brain activity into images and text. arXiv preprint arXiv:2305.11560 2023.
- Yeğin, M.N.; Amasyalı, M.F. Generative diffusion models: A survey of current theoretical developments. Neurocomputing 2024, 608, 128373.
- Gao, R.; Wang, J.; Yu, Y.; Wu, J.; Zhang, L. Enhanced graph diffusion learning with dynamic transformer for anomaly detection in multivariate time series. Neurocomputing 2025, 619, 129168.
- Yan, X.; Zhang, X.; Xia, S. Multi-View topology assisted dynamic graph learning for fMRI-based Alzheimer’s disease identification. Neurocomputing 2025, 618, 129025.
- Yin, W.; Li, L.; Wu, F.X. Deep learning for brain disorder diagnosis based on fMRI images. Neurocomputing 2022, 469, 332–345.
- Mercha, E.M.; Benbrahim, H. Machine learning and deep learning for sentiment analysis across languages: A survey. Neurocomputing 2023, 531, 195–216.
- Wu, C.H.; De la Torre, F. Unifying Diffusion Models’ Latent Space, with Applications to CycleDiffusion and Guidance. arXiv preprint arXiv:2210.05559 2022.
Figure 2.
Overview of the MindCanvas framework: (A) Masked Brain Modeling processes fMRI signals via a Vision Transformer (ViT) to extract latent embeddings, representing brain activity. (B) Artwork Creation uses CycleDiffusion to refine the embeddings based on text prompts and generates the final artwork through a step-by-step painting process.
Figure 2.
Overview of the MindCanvas framework: (A) Masked Brain Modeling processes fMRI signals via a Vision Transformer (ViT) to extract latent embeddings, representing brain activity. (B) Artwork Creation uses CycleDiffusion to refine the embeddings based on text prompts and generates the final artwork through a step-by-step painting process.

Figure 3.
The painting process created from users’ brain activity in response to visual stimuli and text prompts. Each row shows the gradual development of an artwork, starting with sketches and ending in fully detailed, stylized pieces. The figure highlights the model’s ability to refine details, integrating neural inputs and text prompts to generate diverse styles. This dynamic evolution mirrors the step-by-step process of human creativity.
Figure 3.
The painting process created from users’ brain activity in response to visual stimuli and text prompts. Each row shows the gradual development of an artwork, starting with sketches and ending in fully detailed, stylized pieces. The figure highlights the model’s ability to refine details, integrating neural inputs and text prompts to generate diverse styles. This dynamic evolution mirrors the step-by-step process of human creativity.

Figure 4.
The creation of art based on fMRI data of human brain activity corresponding to the visual stimulus and text prompts for artistic revision. Visual stimuli are in the (a) group, reconstructed images are in the (b) group, (c), (d), and (e) groups are the artwork artistic by text prompts.
Figure 4.
The creation of art based on fMRI data of human brain activity corresponding to the visual stimulus and text prompts for artistic revision. Visual stimuli are in the (a) group, reconstructed images are in the (b) group, (c), (d), and (e) groups are the artwork artistic by text prompts.

Figure 5.
The Dynamic Process of Stylized Artistic Creation with MindCanvas. The MindCanvas model uses fMRI-reconstructed images combined with text prompts to generate stylized artwork. Starting from a reconstructed image of a swan, the model applies a text prompt to guide the artistic style. As the process unfolds, MindCanvas generates a sequence of progressively refined outputs, culminating in a final stylized painting. This sequence captures the dynamic process of creative expression, from rough sketches to detailed, stylized artwork.
Figure 5.
The Dynamic Process of Stylized Artistic Creation with MindCanvas. The MindCanvas model uses fMRI-reconstructed images combined with text prompts to generate stylized artwork. Starting from a reconstructed image of a swan, the model applies a text prompt to guide the artistic style. As the process unfolds, MindCanvas generates a sequence of progressively refined outputs, culminating in a final stylized painting. This sequence captures the dynamic process of creative expression, from rough sketches to detailed, stylized artwork.

Figure 6.
Comparison of image scores from users and evaluators across 4 art-style groups.

Figure 7.
Side-by-side comparison of user and evaluator scores.

Table 1.
Quantitative Comparison of PSNR and SSIM Between Different Image Groups. A quantitative comparison of PSNR and SSIM values across three image groups: stimulus image (a), reconstructed image from fMRI data (b), and the final artistic creation image (c).
Table 1.
Quantitative Comparison of PSNR and SSIM Between Different Image Groups. A quantitative comparison of PSNR and SSIM values across three image groups: stimulus image (a), reconstructed image from fMRI data (b), and the final artistic creation image (c).
| Image | PSNR | SSIM | ||||
| (a)&(b) | (a)&(c) | (b)&(c) | (a)&(b) | (a)&(c) | (b)&(c) | |
| Img (1) | 9.1764 | 9.3325 | 19.2563 | 0.1215 | 0.2045 | 0.5447 |
| Img (2) | 10.5326 | 10.1618 | 18.8047 | 0.1250 | 0.1259 | 0.5148 |
| Img (3) | 10.6282 | 8.7512 | 14.5208 | 0.1227 | 0.0644 | 0.3682 |
| Img (4) | 10.6282 | 8.0812 | 12.9797 | 0.1227 | 0.0623 | 0.2991 |
| Img (5) | 10.1867 | 9.8658 | 15.1830 | 0.0841 | 0.0948 | 0.3435 |
| Img (6) | 9.9079 | 9.9540 | 18.0132 | 0.0763 | 0.0926 | 0.4722 |
| Img (7) | 7.4803 | 7.6725 | 23.1061 | 0.0756 | 0.0671 | 0.7025 |
| Img (8) | 10.1283 | 7.6560 | 14.9131 | 0.2461 | 0.2494 | 0.3428 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.