1. Discussion
1.1. Perceived Usefulness, Usability, and Reliability: RQ1 and H1
The first research question examined how students perceived the usefulness, usability, clarity, and reliability of the course-constrained Retrieval-Augmented Generation (RAG) assistant after using it in an authentic Vocational Education and Training (VET) classroom. Overall, the findings provide preliminary descriptive support for H1. Students reported positive perceptions across the four aspects considered, particularly in relation to ease of use, the clarity of the generated responses, and their perceived correctness.
Usability emerged as one of the strongest aspects of the prototype. All 19 students agreed or strongly agreed that the assistant was easy to use, and none indicated that technical or external assistance was required. These results suggest that the browser-based interface and the straightforward question-and-answer workflow enabled students to interact with the system without encountering significant access barriers. This is particularly relevant for the classroom integration of educational AI tools, as their potential value may be reduced if students require extensive training or technical support before they can use them effectively.
The clarity of the generated responses was also assessed positively. Most students considered the answers easy to understand and valued their concise and structured presentation. The open-ended comments reinforce this interpretation: participants highlighted that the assistant provided focused answers, avoided unnecessarily lengthy explanations, and presented information in a way that facilitated rapid consultation. The instructor similarly observed that the responses were well summarised and structured. These findings are consistent with previous research showing that students often value AI-based educational assistants for their immediacy, accessibility, and ability to provide targeted support for academic tasks [
16].
The results also indicate a positive perception of reliability. All students agreed or strongly agreed that the responses seemed correct most of the time. However, this result must be interpreted carefully. The study did not include an expert assessment of response accuracy or a comparison with a general-purpose Large Language Model (LLM). Consequently, the findings do not demonstrate that the system objectively generated correct answers or reduced hallucinations. They show that students perceived the responses as generally reliable after interacting with the prototype during the classroom session.
Grounding the responses in teacher-provided materials appears to have contributed to this perceived trustworthiness. Most students reported that knowing the assistant used the course materials increased their confidence in the system. The qualitative responses provide further context: several participants valued receiving explanations aligned with the procedures and terminology used in class rather than alternative solutions drawn from unrestricted external sources. This result is consistent with the rationale underlying course-constrained RAG systems: restricting the retrieval repository to instructor-selected materials can improve curricular consistency and make the origin of the information more transparent.
Similar observations have been reported in previous educational deployments of RAG-based assistants. Németh et al. [
13] found that grounding an AI tutor in lecturer-provided educational resources contributed to positive student and instructor experiences while also reducing unsupported responses when the relevant content was sufficiently represented in the available materials. In medical education, Thesen and Park [
14] similarly observed that source-grounded responses could reinforce students’ trust in the system, while also identifying a tension between reliability and comprehensiveness when learners expected answers beyond the boundaries of the curated corpus.
The present findings reproduce this tension in a technical VET setting. Although most students valued the use of official course materials, the confidence item showed greater variability than the other positively oriented items. The open-ended responses suggest that students recognised an important limitation: the quality of the assistant ultimately depends on the quality, completeness, and currency of the instructional resources incorporated into the repository. Restricting the system to a controlled corpus therefore improves pedagogical alignment, but it does not guarantee that every answer will be complete or objectively correct.
The students’ extensive previous experience with general-purpose AI tools also provides relevant context for interpreting these results. Most participants regularly used tools such as ChatGPT, Gemini, or Copilot before the intervention. Their positive assessment may therefore reflect not only the usability of the proposed assistant but also their ability to compare its course-specific behaviour with the broader and less controlled responses commonly produced by general-purpose systems. At the same time, this familiarity limits the extent to which the findings can be transferred to learners with little or no prior experience of generative AI.
Taken together, the results provide preliminary descriptive support for H1: after the classroom session, students perceived the assistant as a useful, accessible, clear, and generally reliable complementary learning-support tool. Nevertheless, the findings concern perceived reliability rather than objectively measured response accuracy. Future studies should complement user perceptions with expert evaluation of generated responses and controlled comparisons with unconstrained LLM-based tools.
1.1. Perceived Learning Support and Autonomous Learning: RQ2
The second research question examined the extent to which students perceived the assistant as supporting content understanding, the resolution of course-related questions, and more autonomous learning practices. The findings indicate that the strongest perceived contribution of the system concerned immediate learning support. Students valued the possibility of obtaining focused explanations and clarifications without manually searching through the complete set of instructional materials.
The quantitative results show consistently positive ratings across the learning-support items. Students particularly valued the provision of clear and useful explanations (M = 4.47, SD = 0.61), the resolution of course-related questions (M = 4.42, SD = 0.61), and the rapid access to summaries or clarifications (M = 4.42, SD = 0.69). The assistant was also perceived as helpful for improving understanding of the course content (M = 4.21, SD = 0.71). These findings are consistent with the qualitative responses, in which students highlighted the value of receiving concise, structured, and specific answers connected to their immediate doubts.
This type of support is particularly relevant in technical Vocational Education and Training (VET). Students in the Network Services and Internet module work with specialised documentation, practical procedures, and configuration tasks that may generate specific questions during learning activities. In this context, an assistant capable of locating relevant information within the course materials can reduce the effort required to search manually through multiple resources and provide timely guidance when a doubt arises. The system therefore acts as a complementary access layer to the instructional corpus rather than as a replacement for the original materials or for teacher support.
Previous studies have similarly highlighted the potential of educational chatbots to provide immediate, context-sensitive assistance. Schei et al. [
16] identified immediacy and accessibility as recurring factors contributing to students’ positive perceptions of generative AI tools in educational contexts. Miladi et al. [
10] also reported the potential of a RAG-enhanced conversational agent to support learning activities in an online environment, while Németh et al. [
13] emphasised that the usefulness of RAG-based tutoring systems depends on the relationship between the learner’s question and the available course resources. The present study extends these observations to an authentic technical VET classroom.
The results also indicate a positive perception of the assistant’s potential to support more autonomous study practices. Sixteen of the 19 students agreed or strongly agreed that the system could help them study more independently (M = 4.37, SD = 0.76). This finding suggests that participants recognised value in being able to consult the assistant without requiring immediate teacher intervention for every question. However, the result should be interpreted as a perception of potential rather than as evidence of increased learner autonomy. The classroom evaluation was limited to a single session and did not examine whether students subsequently used the tool independently, improved their study habits, or achieved better academic outcomes.
A more nuanced pattern emerges from the items related to reflective use and source verification. Eleven students agreed or strongly agreed that the assistant encouraged them to consult the original course materials (M = 3.79, SD = 1.18), and the same number considered that interacting with the system helped them formulate better questions (M = 3.68, SD = 0.82). Although these results are moderately positive, they are lower than the scores obtained for immediate doubt resolution and explanation quality. This difference suggests that the assistant was perceived primarily as a rapid consultation tool rather than as an instrument that automatically promotes reflective learning behaviours.
The distinction is important from a pedagogical perspective. Providing source references creates an opportunity for students to verify information and return to the original instructional resources, but the availability of these references does not ensure that learners will use them actively. Similarly, access to a conversational interface may encourage students to refine their questions over time, but the current lack of conversational memory limits the possibility of developing extended inquiry sequences. The assistant’s educational value therefore depends not only on its technical features but also on how it is incorporated into learning activities.
Future classroom deployments could strengthen this reflective dimension through guided activities. For example, students could be asked to compare an answer with its cited source, identify the fragment that supports a specific explanation, reformulate an ambiguous query, or evaluate whether the retrieved evidence is sufficient to answer a practical question. These activities could help transform source traceability from a transparency feature into an explicit learning strategy.
Taken together, the findings provide a positive response to RQ2. Students perceived the assistant as a useful tool for understanding content, resolving doubts, and obtaining rapid explanations. They also recognised its potential to support more autonomous study practices. Nevertheless, the results do not demonstrate an objective improvement in learning autonomy or academic performance. Further longitudinal and experimental research is required to determine whether repeated use of the assistant produces measurable changes in study behaviour, knowledge acquisition, or learning outcomes.
1.1. Design Trade-Offs and Improvement Priorities: RQ3 and RQ4
The third and fourth research questions addressed the limitations identified during the classroom deployment and the pedagogical and technical improvements suggested for future iterations of the assistant. The findings reveal that the most relevant challenges are closely connected to the core design decision underlying the platform: restricting responses to teacher-provided instructional materials.
1.1.1. Conversational Continuity as the Main Development Priority
The absence of conversational memory emerged as the most frequently requested improvement. Eleven of the 19 students explicitly suggested preserving the context of previous interactions, supporting follow-up questions, or allowing users to resume earlier conversations. Although the current platform stores a history of standalone queries and responses, these interactions are not incorporated into the context of subsequent questions.
This limitation reduces the naturalness of the interaction. Learning-related questions often develop progressively: an initial response can generate a new doubt, require clarification, or lead the student to explore a related concept. Requiring learners to reformulate the complete context in every query introduces unnecessary friction and may limit deeper inquiry.
Future versions should therefore incorporate conversational memory within each course-specific interaction. However, continuity should be implemented carefully. Previous questions and responses should be used to preserve the conversational context without allowing the dialogue to drift away from the validated instructional corpus. Each new response should remain grounded in retrieved evidence, and the system should continue to abstain when sufficient support is unavailable.
1.1.1. Handling Ambiguous Queries and Response Boundaries
A further priority concerns the treatment of broad, ambiguous, or insufficiently specific questions. Some students indicated that complete answers required highly detailed queries, while the instructor observed that an ambiguous question produced no useful response and suggested that the assistant could request additional information.
This observation points to an important distinction between abstention and clarification. Returning an abstention message is appropriate when the available materials do not contain sufficient evidence. However, when the difficulty arises because the user’s intention is unclear, the assistant should first attempt to reformulate or narrow the question through a clarification request. For example, a broad query about installing a tool could lead the system to ask which operating system, version, or deployment scenario the learner is using before generating an answer.
Future iterations should therefore incorporate a clarification mechanism before final response generation. This improvement could increase the usefulness of the assistant without weakening its course-constrained nature. It would also support the development of better question-formulation practices by guiding students towards more precise and contextually appropriate queries.
The classroom feedback also suggests that response-boundary control should be strengthened. One participant observed that the assistant correctly abstained when a question was clearly unrelated to the course but occasionally responded to closely related questions even when the relevant information was not present in the materials. In the current prototype, abstention is primarily managed through prompt instructions. A more robust implementation could combine prompt-based abstention with retrieval-level safeguards, such as minimum similarity thresholds or evidence-sufficiency checks. These mechanisms should be calibrated empirically in future evaluations rather than assumed to guarantee accuracy.
1.1.1. Additional Interaction and Traceability Improvements
The participants also suggested several secondary improvements. These included the ability to attach screenshots, images, or supplementary documents; more direct navigation to the exact location of the referenced information within the original source; improved access to previous interactions; and refinements to the visual interface.
The possibility of attaching screenshots may be particularly relevant in technical VET contexts, where students frequently encounter configuration errors, command-line messages, or interface states that are difficult to describe precisely in text. Nevertheless, multimodal interaction would require additional evaluation to ensure that the assistant continued to provide responses aligned with the intended pedagogical scope.
Improved source navigation would also strengthen the educational value of the platform. The current system associates responses with source documents and available location metadata. Future versions could provide more direct links to the relevant page, paragraph, or line range whenever the original file format allows it. This enhancement would make it easier for learners to verify the generated explanation and return to the original instructional resource.
Taken together, the findings provide a clear response to RQ3 and RQ4. The main limitations concern restricted corpus coverage, the lack of conversational continuity, the handling of ambiguous queries, and the need for more robust response-boundary control. The most relevant future improvements are therefore not aimed at removing pedagogical constraints, but at making those constraints more transparent, flexible, and useful for learners.
1.1. Implications for Vocational Education and Training
The findings have several implications for the use of course-constrained Retrieval-Augmented Generation (RAG) assistants in Vocational Education and Training (VET). Technical VET programmes frequently require students to interpret specialised documentation, apply procedural knowledge, and resolve practical questions arising during configuration and troubleshooting activities. In this context, rapid access to explanations grounded in the instructional materials can provide valuable support without replacing the original resources or the instructor’s role.
The proposed assistant should therefore be understood as a complementary learning-support tool rather than as an autonomous source of knowledge. Its main educational value lies in facilitating access to the materials selected by the teacher, helping students locate relevant information, clarify concepts, and obtain concise explanations when a doubt arises. This function may be particularly useful during practical sessions, independent study, or revision activities, where immediate teacher intervention is not always available.
The course-constrained approach also supports a form of pedagogical oversight that is especially relevant in vocational education. By restricting the retrieval repository to instructional resources selected by the teacher, the assistant can remain more closely aligned with the terminology, procedures, and level of complexity expected in the module. This characteristic may reduce the risk of exposing students to alternative solutions that are technically plausible but inconsistent with the learning objectives or with the procedures introduced in class.
However, pedagogical control should not be confused with completeness. The findings show that a restricted repository can also limit the assistant’s ability to address unexpected technical problems, ambiguous queries, or questions that extend beyond the available course materials. For this reason, instructors should view the knowledge base as an evolving educational resource requiring periodic review, updating, and extension. The design of the assistant should make these boundaries visible to learners and clearly distinguish between supported answers, insufficient evidence, and optional supplementary information.
The interaction history available to instructors also creates opportunities for pedagogical monitoring. Reviewing the questions submitted by students may help identify recurring doubts, misunderstood concepts, or parts of the instructional materials that require clarification. In this sense, the assistant can support not only individual consultation but also formative reflection on the teaching process. Future versions could incorporate dashboards or summaries of frequently asked questions to facilitate this type of analysis.
The classroom integration of these tools should also include explicit guidance on their appropriate use. Students need to understand that source-grounded responses remain dependent on the quality of the indexed materials and that the availability of citations does not eliminate the need for critical verification. Activities that require learners to compare answers with the original sources, reformulate ambiguous questions, or evaluate whether the retrieved evidence is sufficient could contribute to the development of artificial intelligence literacy and responsible study practices.
These implications suggest that course-constrained RAG assistants may be particularly suitable for technical VET environments when they are integrated as supervised, transparent, and curriculum-aligned support tools. Their educational value does not depend solely on response generation, but also on how teachers curate the knowledge base, design learning activities, interpret interaction data, and encourage students to use the system critically.
1.1. Study Limitations
The findings should be interpreted in light of several limitations related to the exploratory nature of the study.
First, the evaluation was conducted with a small convenience sample comprising 19 students from a single module within one Vocational Education and Training (VET) institution. The results therefore provide context-specific evidence and should not be generalised to other educational levels, subject areas, institutions, or student profiles without further investigation. Future studies should include larger and more diverse samples across different modules, VET programmes, and educational settings.
Second, the classroom deployment was limited to a single practical session. The study captures participants’ immediate perceptions after interacting with the assistant, but it does not examine sustained use over time. Consequently, the results do not provide evidence regarding long-term adoption, changes in study habits, continued engagement, or the extent to which the assistant could support autonomous learning outside the classroom. Longitudinal evaluations are required to assess how students use the system during regular study activities and whether their perceptions change after repeated interaction.
Third, the evaluation focused primarily on self-reported perceptions. The questionnaire examined usability, clarity, perceived reliability, learning support, and practical limitations, but the study did not include objective measures of learning outcomes. No pre-test/post-test design, control group, or assessment of academic performance was incorporated. The findings therefore indicate that students perceived the assistant as useful, but they do not demonstrate that its use improved knowledge acquisition, retention, or academic achievement.
Fourth, perceived reliability should not be interpreted as objective response accuracy. Although students generally considered the generated answers correct, the responses were not systematically reviewed by independent experts using a predefined evaluation protocol. The study also did not compare the course-constrained assistant with a general-purpose Large Language Model (LLM) or quantify the occurrence of unsupported or hallucinated responses. Future research should incorporate expert assessment of generated answers, retrieval-quality metrics, predefined question sets, and controlled comparisons between constrained and unconstrained systems.
Fifth, the questionnaire was developed specifically for this exploratory classroom evaluation. The thematic organisation of the items facilitated descriptive interpretation, but the instrument has not undergone psychometric validation. The grouped scores should therefore be understood as descriptive summaries rather than as validated measurement scales. Future studies with larger samples could refine the instrument and examine its reliability and construct validity.
Sixth, the participants reported extensive previous experience with generative artificial intelligence tools such as ChatGPT, Gemini, and Copilot. This familiarity may have facilitated interaction with the assistant and influenced their expectations and evaluations. The generally positive reception observed in this study may not be replicated in groups with lower levels of AI literacy or different patterns of technology use.
Seventh, the instructor feedback provides a useful complementary perspective but is based on the response of a single teacher. It should therefore be interpreted as contextual evidence rather than as a representative evaluation of instructor acceptance. Future deployments should involve multiple instructors and examine how teachers use interaction histories, manage course materials, and integrate the assistant into different pedagogical strategies.
Finally, the evaluation concerns a specific version of the prototype and its current technical configuration. The system retrieves the five highest-ranked fragments using local cosine similarity and relies primarily on prompt-based abstention when the available evidence is insufficient. The study did not assess retrieval precision, experiment with different chunking strategies, or calibrate similarity thresholds. Subsequent technical evaluations should examine how these design decisions influence response quality, corpus-boundary control, and the balance between reliability and informational coverage.
Despite these limitations, the study provides preliminary empirical evidence from an authentic technical VET classroom. Its value lies in identifying both the perceived benefits of a course-constrained RAG assistant and the practical challenges that should guide more extensive technical and pedagogical evaluations.