Submitted:
24 July 2026
Posted:
27 July 2026
You are already at the latest version
Abstract
AI-powered conversational agents are becoming part of the everyday Internet information ecosystem, reshaping how users seek, interpret, and act on health-related information outside clinical encounters. As large language model (LLM)-based chatbots are increasingly used as on-demand digital health information tools, understanding how users perceive their credibility, usefulness, and limitations is essential for the responsible design of future Internet-based health services. This mixed-method survey study examined how general adults evaluated healthcare-related question-answer pairs provided by physicians and generated by AI chatbots. Participants (N=62) rated each answer on clarity, usefulness, appropriateness of detail, trustworthiness, and perceived evidence, and provided open-ended explanations of their judgments. Results found that AI-generated responses were rated significantly higher than physician-provided responses overall, t(61) = 8.63, p < 0.001, with significant advantages across all five dimensions. In addition, qualitative findings showed that participants valued detailed, specific, and evidence-like explanations, while expressing concerns about hallucination, privacy, over-reliance, and the need for clinician verification. These findings suggest that LLM-based chatbots may be perceived as useful supplemental information tools within future Internet health ecosystems, but their deployment should include safeguards that support transparency, verification, and appropriate reliance.
Keywords:
health information
; healthcare
; artificial intelligence
; AI-based conversational agents
; information seeking
; human-AI interaction
1. Introduction
Patients and caregivers increasingly seek health-related information outside direct clinical encounters [1]. Although clinicians remain central sources of medical guidance, understanding diagnosis or treatment options can be cognitively and emotionally demanding for individuals without medical training [2]. Patients and caregivers may need additional explanations after clinical visits, particularly when they try to understand unfamiliar terminology, complex treatment choices, probabilistic risk information, or uncertainty about outcomes [3]. Therefore, the need for supplemental health information becomes especially important in contexts involving preference-sensitive decisions, where patients must compare multiple reasonable options and consider how different risks, benefits, and side effects align with their personal circumstances.
Meanwhile, AI-powered conversational agents are becoming increasingly embedded in the everyday Internet information ecosystem [4]. Large language model (LLM)-based chatbots (e.g., ChatGPT, Gemini, and Claude) are able to generate detailed and conversational answers to health-related questions. Unlike static web pages, these chatbots allow users to ask context-specific questions, request follow-up clarification, and receive immediate responses in natural language [5]. As a result, LLM-based chatbots may become important supplemental sources of health information for patients and caregivers seeking explanations outside clinical encounters.
However, it creates new challenges for health communication research. While there are more sources for patients to seek and receive health information and advice, physician-provided and AI-generated information may vary in style, length, structure, and level of details. Physician communication often reflects clinical expertise, but it may be constrained by time, context, and the need to communicate efficiently. In contrast, AI-generated answers are able to provide immediate and expansive explanations in conversational language, but they may lack clinical judgment and accountability. The different characteristics raise uncertainty about how patients judge the quality and trustworthiness of health information when professional authority and explanatory details do not come from the same source. In particular, individuals may perceive an AI-generated answer as more helpful because of how it is presented, even when it lacks the clinical context that makes physician guidance medically appropriate.
Given these challenges, there is a need to better understand how individuals evaluate AI-generated health information as a form of patient-facing communication. Much of the existing discussion around AI in healthcare focuses on whether AI-generated responses are medically accurate or technically capable [6,7]. However, patients and caregivers often encounter health information as non-experts, and their judgments may depend on how understandable, complete, useful, and trustworthy the information appears. It is also unclear whether these judgments are shaped primarily by the source of the information, by the way the response is written, or by individuals’ own background characteristics. Examining AI-generated and physician-provided answers side by side can help clarify how individuals perceive different forms of health advice and what factors may influence appropriate or inappropriate reliance on AI as a supplemental health information tool.
Specifically, this study aims to answer the following research questions:
- 1.
- How do individuals evaluate AI-generated health information compared with physician-provided information in terms of clarity, usefulness, appropriateness of details, trustworthiness, and perceived evidence?
- 2.
- Are personal characteristics (specifically health literacy, AI literacy, and baseline trust in AI) associated with an individual’s evaluations of AI-generated and/or physician-provided health information?
- 3.
- What response characteristics do individuals identify as shaping trust, skepticism, and willingness to use AI-generated health information as a supplemental resource?
2. Related Works
2.1. Health Information Seeking Outside Clinical Encounters
Prior studies have shown that online health information seeking is a routine part of contemporary healthcare behavior. Systematic reviews identify patients, caregivers, family members, and general consumers as major online health information seekers, with motivations ranging from understanding symptoms and diagnoses to comparing treatment options and preparing for clinical encounters [1]. Another review further identified demographic, socioeconomic, literacy-related, and contextual factors that influence online health information seeking, including age, gender, income, education, caregiving role, and place of residence [8].
Research also shows that individuals often rely on multiple sources to make sense of health-related information. Socioeconomically disadvantaged individuals, for example, commonly rely on online resources, family members, and healthcare professionals to supplement their understanding of health information [9]. This pattern suggests that health information seeking is often distributed across formal and informal sources rather than limited to direct communication with clinicians. However, patients may remain dissatisfied with the informational background they gain from clinical encounters, particularly when medical information is difficult to understand, emotionally overwhelming, or insufficiently detailed. In one study, prostate cancer patients frequently reported dissatisfaction with the amount and clarity of information they received regarding diagnosis, treatment options, and long-term outcomes [10]. Such findings suggest that insufficient information can contribute to uncertainty and reduced confidence in treatment decision-making.
A parallel body of work examines how individuals judge the quality of online health information. For example, Sun et al. found that consumers evaluate online health information using criteria such as trustworthiness, understandability, comprehensiveness, relevance, usefulness, and credibility [11]. Other reviews showed that perceived information quality could be shaped by source factors, content factors, design features, and individual user characteristics [12]. Health literacy is especially important because users with lower health literacy may have more difficulty identifying reliable information, understanding medical terminology, and applying information appropriately [2].
2.2. Conversational AI in Patient Education and Healthcare Communication
Recent research on AI in healthcare has expanded from earlier applications in prediction, diagnosis, medical imaging, and clinical decision support to natural-language applications enabled by large language models. Conversational LLMs have been studied for medical question answering, summarization, patient education, administrative support, clinical documentation, and patient communication [6]. A systematic review of healthcare applications of LLMs further shows that many studies focus on technical performance, benchmark accuracy, and system capability, while fewer studies evaluate how these systems are perceived and used in patient-facing contexts [7].
Within this broader area, several studies have examined how AI-generated information may support communication between patients and healthcare systems. Whitehouse et al. examined AI-generated patient message drafts in a peri-operative endocrine surgery setting and found that such tools may help streamline routine communication tasks [13]. Previous studies have also examined AI chatbots as patient-facing educational tools. Baumgärtner et al. evaluated PROSCA, a medical chatbot designed to inform patients about prostate cancer, and found that patients perceived the chatbot as useful for improving understanding and supplementing physician communication [14]. Another study by Baur et al. developed and evaluated a retrieval-augmented generation chatbot for orthopedic and trauma surgery education and reported that chatbot-generated explanations may improve accessibility of patient education materials by presenting information in conversational and understandable forms [15]. Extending this line of work, Ayers et al. compared physician and AI chatbot responses to patient questions posted online and found that chatbot responses were rated highly for both quality and empathy by licensed healthcare evaluators [16]. Although these studies differ in setting and evaluation approach, they share a common pattern that AI-generated responses are often valued for accessibility, conversational style, completeness, and efficiency.
However, these studies also illustrate unresolved questions. Many evaluations focus on expert review, system performance, or workflow efficiency rather than general patients’ own judgments of AI-generated health advice. In addition, studies of condition-specific tools may not fully explain how individual users perceive widely available LLM-based chatbots when seeking health-related information in everyday contexts.
3. Materials and Methods
To address the gaps and explore our research questions, a mixed-method, within-subjects survey was designed to examine how general adult participants evaluate physician-provided and AI-generated health information, using both quantitative ratings of perceived answer quality and qualitative open-ended responses about participants’ reasoning, skepticism, and potential use of AI-generated health advice. The primary comparison was between physician-provided answers and AI-generated answers. AI-generated answers included responses generated by two LLMs, ChatGPT and Claude. Since our main research objective was to compare human-provided and AI-generated health information, ratings for ChatGPT- and Claude-generated responses were pooled into a single AI-generated category for the main analysis. The study focused on participants’ perceptions of answer quality, including clarity, usefulness, appropriateness of details, trustworthiness, and perceived evidence. Participants were not asked to determine whether the answers were clinically correct, and the study did not independently evaluate the medical accuracy of the responses. Accordingly, all findings are interpreted as users’ perceptions of health-information quality and trustworthiness rather than objective assessments of clinical validity.
Specifically, participants were asked to evaluate health-related knowledge in the context of a hypothetical prostate cancer information-seeking scenario, imagining a close friend of them had recently been diagnosed with prostate cancer and had asked for help understanding the condition, treatment options, side effects, risks, and follow-up care. This scenario was used to provide a consistent patient-facing context for evaluating whether the answers were understandable, useful, appropriately detailed, reliable, and trustworthy. The scenario positioned the participants as general information evaluators. Participants were not asked to act as clinicians or make medical judgments. Instead, they were instructed to evaluate how the answers would appear to someone seeking health information for themselves or for another person.
The survey materials consisted of prostate cancer-related question-answer sets. The questions covered common informational topics related to prostate cancer diagnosis and treatment, including early-stage treatment options, radiation therapy, surgery, side effects, incontinence, erectile dysfunction, PSA monitoring, second opinions, cancer aggressiveness, and treatment decision-making. These were the most frequently asked questions selected from 100 recorded transcriptions on conversations between doctors and prostate cancer patients and their caregivers. Note that all the personally identifiable information were originally removed from the datasets.
3.1. Participants Recruitment
Study participants were recruited through Prolific, a crowd-sourcing platform for online research participation, from May 1, 2026 to June 1, 2026. The study targeted general adult participants rather than individuals with medical training or healthcare expertise, since they were expected to evaluate the responses from a patient or caregiver perspective, not from the perspective of a clinician assessing medical correctness. Eligible participants were required to be at least 18 years old, able to read and understand English, able to provide informed consent electronically, and willing to complete an online survey involving hypothetical health-information scenarios. Participants accessed the study through Prolific and completed the survey on Qualtrics.
3.2. Survey Flow
The survey consisted of three main sections:
- 1.
-
Background questions. These included demographic items and measures of health literacy, AI literacy, and baseline trust in AI tools. These measures were used to characterize the sample and to examine whether participant-level characteristics were associated with evaluations of AI-generated and physician-provided answers. Specifically,
- Health literacy was measured using an abbreviated set of items adapted from the All Aspects of Health Literacy Scale [17]. The items assessed functional, communicative, and critical aspects of health literacy, including participants’ ability to understand health information, communicate with healthcare professionals, ask questions, seek clarification, evaluate whether health information can be trusted, and consider whether health information applies to their situation. Negatively worded items were reverse-coded so that higher scores consistently indicated stronger health literacy.
- AI literacy was measured using an abbreviated set of items adapted from the Meta AI Literacy Scale [18]. The items assessed participants’ perceived ability to use AI tools, understand AI concepts and limitations, recognize AI-enabled systems, consider ethical implications of AI use, solve problems using AI, keep up with AI developments, recognize AI influence, and regulate emotional responses when interacting with AI systems. Higher scores indicated greater self-reported AI literacy.
- Baseline trust in AI, adapted from Choung, David, and Ross’s Trust in AI Questionnaire [19], was measured using items assessing participants’ general attitudes toward AI tools, including perceived reliability, trustworthiness, honesty, usefulness, ability to complete tasks, and concern for users’ interests. Higher scores indicated greater baseline trust in AI.
- 2.
-
Evaluation tasks. Participants were randomly assigned five prostate cancer-related question blocks. Each block consisted of one health question and three corresponding answers: one physician-provided answer, one ChatGPT-generated answer, and one Claude-generated answer. Participants reviewed all three answers within each assigned block, and each participant evaluated 15 (6 * 3) answers in total. Source labels were not disclosed to participants during the evaluation. This blinding procedure was used to reduce source-label bias and to focus participants’ evaluations on the content and presentation of the answers. After reading each answer, participants rated it on five dimensions using a five-point Likert-type scale ranging from 1 = strongly disagree to 5 = strongly agree. The five rating dimensions were:
- Clarity: assessed whether the answer was clearly written and easy to understand.
- Usefulness: assessed whether the answer provided helpful information.
- Appropriateness of details: assessed whether the answer addressed the question with an appropriate level of details.
- Trustworthiness: assessed whether participants viewed the answer as reliable and trustworthy.
- Perceived evidence: assessed whether participants viewed the answer as evidence-based and accurate.
After each answer, participants were also asked to provide an open-ended explanation of what influenced their ratings. - 3.
- Post-survey open-ended questions. Participants were asked to share their broader views of AI-generated health information. These questions asked whether they would consider using AI-generated answers for health-related decisions, how much they trust AI-generated health information in real life, and what concerns they have about use of AI tools for healthcare decision-making.
3.3. Data Analysis
Descriptive statistics were calculated for participant demographics, measurement scales, and ratings of answers for each question in five dimensions. Internal consistency was assessed for multi-item measures using Cronbach’s alpha. The primary quantitative analysis compared participants’ mean ratings of AI-generated answers with their mean ratings of physician-provided answers. Paired-samples t-tests were used since each participant evaluated both source types. Mean differences, 95% confidence intervals, p-values, and standardized effect sizes were also reported. To examine whether the AI-vs-physician difference varied across rating dimensions, a repeated-measures analysis of variance was conducted with response source and evaluation dimension as within-subject factors. This analysis tested whether the magnitude of the source effect differed across clarity, usefulness, appropriateness of details, trustworthiness, and perceived evidence.
Correlation analyses were conducted to examine whether participant-level characteristics were associated with answer evaluations. Specifically, Pearson correlations were used to examine associations among AI literacy, health literacy, baseline AI trust, AI-generated answer ratings, physician-provided answer ratings, and the AI advantage score. These analyses were used to determine whether participants’ background and personal characteristics helped explain variation in attitudes and preferences on AI-generated information versus physician-provided information. Note that all the quantitative analyses were interpreted as assessing participant perceptions rather than clinical accuracy or medical effectiveness.
Qualitative responses were analyzed using thematic analysis. The analysis focused on identifying recurring patterns in participants’ explanations of their perspectives and concerns about AI-generated health information. The qualitative analysis proceeded in several steps. Open-ended responses were first reviewed to identify recurring reasons participants gave for rating answers more or less favorably. After that, initial codes were developed around common response characteristics, such as clarity, specificity, completeness, use of numbers or evidence, perceived medical credibility, jargon, overconfidence, uncertainty, and recommendations to consult healthcare professionals. Codes were then grouped into broader themes that explained how participants evaluated physician-provided and AI-generated answers. Qualitative themes were also compared with the quantitative findings to identify where participant explanations aligned with rating patterns.
4. Results
After data validation, responses from 62 participants who completed the survey and passed all embedded attention checks were analyzed. This section presents the results of both quantitative data analysis and a summary of qualitative responses.
4.1. Demographic Information
A total of 62 participants completed the survey, including 36 male (58.1%), 25 female (40.3% and 1 non-binary participant (1.6%). Age distribution is shown in Table 1. There are 45 participants (72.6%) that are under 55 years old, while 17 participants (27.4%) are age 55 or older. For ethnicity, major of the participants are White (48 participants, 77.4%), while other ethnicity categories contained relatively few participants (4 Asian, 6.5%; 4 Black or African American, 6.5%; 6 Others, 9.7%).
4.2. Measurements on AI Literacy, Health Literacy, and Baseline AI Trust
Participants were asked to complete the shorten version of All Aspects of Health Literacy Scale, Meta AI Literacy Scale, and Trust in AI Questionnaire, and the descriptive summary is shown in Table 2. Cronbach’s alpha is a statistical coefficient used to measure the internal consistency or reliability of each set of measurement items, which indicates how closely related the group of items are as a group, to ensure if a questionnaire accurately measures a single underlying variable.
Overall, participants reported relatively high AI literacy level, with an average score of 3.88 out of 5, which suggests that most participants felt comfortable understanding and using AI tools. Health literacy was also relatively high, with an average score of 3.70. However, the health-literacy measure showed modest internal consistency, likely because the designed survey used an abbreviated set of items covering several different aspects of health literacy. Baseline trust in AI was more mixed, with an average score of 3.17, which implies that participants generally felt more confident in their ability to use and understand AI than in their willingness to fully trust AI tools.
4.3. Ratings for AI-Generated and Physician-Provided Answers
Table 3 shows the statistics of participant ratings of AI-generated and physician-provided health answers. To compare participants’ evaluations of human- and AI-generated health advice, ratings for Claude- and ChatGPT-generated responses were combined into a single AI-generated category. These ratings were then compared with ratings for physician-provided responses. Since each participant evaluated both response types, paired-samples statistical tests were used. The ratings (range from 1 to 5) measure perceived clarity, usefulness, appropriateness of detail, trustworthiness, and evidence/accuracy. Note that these are participants’ perceptions, not an independent clinical assessment of medical correctness.
Participants rated AI-generated answers more positively than physician-provided answers. The average rating for AI-generated answers was 4.16 out of 5 (SD = 0.46), compared with 3.62 for physician-provided answers (SD = 0.66). This difference of 0.54 points was statistically significant, t(61) = 8.63, p < .001, 95% CI [0.41, 0.66], with a large effect size, Cohen’s dz = 1.10. At the participant level, 51 of the 62 participants (82.3%) rated AI-generated answers more highly overall than physician-provided answers.
AI-generated answers received significantly higher ratings across all five evaluation dimensions. The smallest difference appeared for clarity: AI-generated answers were rated 0.26 points higher than physician-provided answers, t(61) = 3.52, p < .001. Larger differences appeared for usefulness, appropriate level of detail, reliability and trustworthiness, and perceived evidence and accuracy. The largest difference was observed for perceived evidence and accuracy, where AI-generated answers were rated 0.64 points higher than physician-provided answers, t(61) = 9.91, p < .001, Cohen’s dz = 1.26.
A repeated-measures analysis of variance showed a significant interaction between response source and evaluation dimension, F(4, 244) = 13.63, p < .001, partial = 0.18. The F-value is calculated to compare how much the AI-vs-physician gap differs across the five dimensions and how much unexplained variation remains after accounting for participant-level differences. The score indicates that the size of the AI advantage differed across the five evaluation dimensions. Partial is calculated to measure the proportion of variance in the dependent variable explained by a specific independent variable, and a partial of 0.18 means the difference between AI-generated and physician-provided answers changes meaningfully depending on what participants are evaluating. Although participants perceived AI-generated answers as somewhat clearer than physician-provided answers, the stronger advantages concerned informational substance that participants viewed AI-generated responses as more useful, more appropriately detailed, more trustworthy, and more evidence-based.
These findings reflect participants’ perceptions of response quality rather than an independent assessment of clinical accuracy. The results suggest that participants responded positively to characteristics commonly present in the AI-generated answers, such as greater detail, explanation, and evidence-like specificity in this survey context.
4.4. Correlation Between Participant-Level Characteristics and Answer Evaluations
Pearson correlation analyses were conducted to examine whether participant-level characteristics were associated with answer evaluations. Specifically, we tested whether AI literacy, health literacy, and baseline trust in AI were associated with participants’ average ratings of AI-generated answers, average ratings of physician-provided answers, and the AI advantage score. The AI advantage score was calculated as each participant’s mean rating of AI-generated answers minus their mean rating of physician-provided answers. The Pearson correlation coefficient (r) and p-values were used to assess whether individual differences in participants’ background characteristics helped explain variation in evaluations of AI-generated versus physician-provided health information.
Results showed that AI literacy was positively associated with ratings of AI-generated answers with r = .351, p = .005, indicating that participants with higher self-reported AI literacy tended to evaluate AI-generated answers more favorably. However, AI literacy was also positively associated with ratings of physician-provided answers (r = .285, p = .025), and was not significantly associated with the AI advantage score (r = -.054, p = .675). This suggests that AI literacy was related to generally more favorable evaluations of health answers, but it was not a stronger relative preference for AI-generated answers.
Health literacy was not significantly associated with ratings of AI-generated answers (r = .210, p = .102), but was positively associated with ratings of physician-provided answers (r = .280, p = .028). Also, there was no statistical significance found on correlation between health literacy and AI advantage score (r = -.179, p = .163), implying that health literacy did not explain the relative preference for AI-generated over physician-provided answers.
Baseline trust in AI was positively associated with both ratings of AI-generated answers (r = .313, p = .013) and ratings of physician-provided answers (r = .422, p < .001). It also showed a modest negative association with the AI advantage score (r = -.274, p = .031). The pattern indicates that participants with higher baseline AI trust tended to rate both answer more favorably, particularly physician-provided answers, which reduced the relative AI-over-physician gap.
The correlation results suggest that participant-level characteristics were associated with some differences in absolute answer ratings but did not clearly explain the relative preference for AI-generated answers. AI literacy and baseline trust in AI were associated with more favorable ratings of AI-generated answers, but none of the participant-level variables provided strong evidence of explaining the AI advantage score. Therefore, the observed AI-over-physician difference appears to be driven more by perceived response-level qualities than by participants’ background literacy or baseline trust.
4.5. Qualitative Responses
Participants’ written explanations help clarify why AI-generated answers received higher quantitative ratings than physician-provided answers. Across the per-answer rationales and post-survey reflections, five recurring themes emerged: the importance of sufficient detail, the value of evidence-like specificity, the need to balance detail with accessibility, skepticism toward unsupported reassurance, and conditional acceptance of AI as a supplemental tool rather than a replacement for clinicians.
4.5.1. Detail and Completeness Increased Perceived Usefulness
A dominant theme across participants’ explanations was the importance of sufficient detail. Participants often rated answers more favorably when they provided context, explained why a treatment or risk mattered, and helped the reader understand next steps or trade-offs. In contrast, answers that were short or general were often viewed as less useful, even when they were understandable. Participants frequently described physician-provided answers as understandable but too brief to support informed decision-making. For example, one participant wrote:
"The answer was easy to understand, but it felt too brief. It listed the main options but did not explain how they differ or when each one is usually recommended."
Another participant commented:
"I feel like a question about cancer deserves a bit more information and detail."
This pattern aligns with the quantitative results. The difference between AI-generated and physician-provided answers was relatively small for clarity but substantially larger for appropriate level of detail and usefulness. Participants did not necessarily view physician-provided answers as difficult to understand. Instead, they often viewed them as insufficiently developed.
By contrast, AI-generated answers were frequently praised for providing more explanation and context. One participant wrote:
"The answer is more specific but still succinct and feels more trustworthy."
Another explained:
"This is a much better answer than the others with a greater level of explanation and detail."
These comments suggest that participants valued answers that did more than list treatment options or provide a brief conclusion. They wanted explanations of how options differed, when they might apply, and what considerations might affect a decision.
4.5.2. Evidence-like Specificity Increased Trust
Participants also placed substantial weight on answers that appeared evidence-based. Comments frequently referred to numbers, statistics, risk estimates, time frames, and concrete explanations. Participants appeared to interpret these features as signals of credibility and reliability. For example, after reviewing an AI-generated response, one participant wrote:
"It gave me numbers and detailed other risks. A more complete answer."
Another participant described a response favorably because it:
"Provides data and doesn’t sugar coat it."
This theme aligns closely with the quantitative results, where the AI-generated answers showed one of their largest advantages on perceived evidence and trustworthiness. Instead of independently verifying the medical accuracy of the answers, participants tended to judge whether the answer appeared medically credible and evidence-based.
4.5.3. Unexplained Jargon Reduced Clarity
Although participants generally preferred detailed answers, they did not reward detail automatically. Some AI-generated answers were criticized for using medical terminology without sufficient explanation. One participant wrote:
"There is too much medical jargon in this answer which most people would not understand. Although it gives a bit more detail than the previous answer, the extra detail does not actually help."
Another stated:
"The response uses too much medical jargon without explaining it. I am left with more questions than answers."
A similar concern appeared in a comment about an otherwise positively received answer:
"This is a much better answer than the others with a greater level of explanation and detail. But it’s still leaving out information about each treatment, such as side effects, and more, so I think more details would be helpful. Also, what’s HDR? Not defined."
These comments show that the preferred answer style was not simply longer. Participants wanted information that was detailed and accessible. Explanations were most effective when they introduced medical concepts clearly, avoided unnecessary jargon, or defined unfamiliar terms.
4.5.4. Unsupported Reassurance and Dismissive Wording Reduced Trust
Some of the lowest-rated physician-provided responses were criticized for sounding overly reassuring or insufficiently supported. Participants appeared uncomfortable when an answer minimized a concern without providing enough explanation or evidence. For example, one participant reacted to a physician-provided response by writing:
"Intentionally ignores the big picture. Because it never happened with this doctor doesn’t mean it doesn’t happen."
Another participant wrote:
"I think the decision should be taken with a doctor. This answer is dismissive and highly concerning to me."
These comments help explain why certain question-level differences were especially large. Participants expected high-stakes health answers to acknowledge uncertainty, clarify the basis for reassurance, and explain the relevant conditions or risks. Reassurance alone was not always perceived as trustworthy.
4.5.5. Post-Survey Responses
The post-survey reflections revealed a clear distinction between favorable ratings of AI-generated answers and willingness to rely on AI in real-world healthcare decisions. Many participants described AI as useful for gathering information, understanding medical terminology, preparing questions, and conducting preliminary research. However, they were reluctant to treat AI as a replacement for clinicians. One participant explained:
"I would use AI to get general information about health conditions but I would always want to fact-check these to ensure it is giving me reliable information."
Another stated:
"Yes, but mostly for general information or to help me understand medical terms and treatment options before speaking with a doctor. I would not rely on AI alone for serious health decisions or diagnosis. I would feel more comfortable using it as a starting point and then confirming the information with a healthcare professional."
A third participant described a similar boundary:
"I would use it for general things that are not major life decisions. I have used it to help understand my son’s IBS and managing symptoms, but I always fact-check it and require citations."
These responses indicate that participants’ acceptance of AI-generated health information was conditional. They often viewed AI as a useful informational aid, especially for lower-stakes questions or preparation for a clinical conversation, while reserving final decisions for healthcare professionals.
Participants frequently expressed concern that AI systems could provide incorrect, incomplete, or outdated information while sounding confident. Hallucination was a recurring issue. One participant wrote:
"I don’t fully trust it because AI can hallucinate, I would always validate its responses with other sources."
Another explained:
"I moderately distrust since AI can sometimes hallucinate while giving the impression it is accurate and detailed."
Participants also raised concerns about privacy and accountability. One response summarized these issues succinctly:
"I’m a little concerned about privacy and more concerned about accuracy and accountability."
Another participant emphasized the need to retain human oversight:
"At this time, YES, we still have to utilize the human in the loop approach as the final result."
Some participants worried that AI could lead people to delay care, ignore clinicians, or become overly confident in an incomplete answer. For example:
"I worry that people will become way too reliant on what AI tells them and use it in place of medical professionals."
Overall, these findings suggest a model of conditional trust. Participants often rated AI-generated answers more favorably because they appeared more detailed, explanatory, and evidence-based. At the same time, participants remained aware of the limitations of AI systems and generally preferred to use AI-generated information as a supplement to professional medical guidance rather than as a substitute for it.
5. Discussion
A central implication of the findings is that individuals’ trust in health information is strongly shaped by observable communication details. When people lack the expertise to independently assess medical correctness, they often rely on more accessible indicators of credibility, such as whether the answer is specific, relevant, and supported by evidence-like information (e.g., statistics, previous cases). Prior research on consumer evaluation of online health information shows that users commonly judge quality through criteria such as trustworthiness, understandability, comprehensiveness, and usefulness [11], which closely match the response features participants valued in this study. This helps explain why participants responded favorably to AI-generated answers. In many cases, the AI-generated responses appeared to offer more of what individuals seek when trying to understand a health topic: more comprehensive explanations, more context, clearer descriptions of risks, and sometimes numerical details. These features can make information feel more complete, thus more trustworthy. From a behavioral perspective, users may interpret detailed and well-organized information as a signal that the answer is credible, even when they cannot verify its medical accuracy.
As AI being able to provide an answer to a health-related question that sounds patient-centered, complete, and evidence-based, for a user facing uncertainty, such type of response may feel more helpful than a shorter answer provided by a physician, even if both contain clinically relevant information. This suggests that trust in AI health information is partly a response-design issue that patients tend to make decisions based on how information is presented, not only on who produced it. However, such situation may also raise risks. When individual patients evaluate health information mainly based on the level of details, they may over-trust an AI-generated answer that is insufficiently personalized, or even inaccurate. This is especially concerning since LLMs can sometime generate hallucinated content while maintaining a polished and authoritative tone. Therefore, while detailed AI-generated responses may improve perceived helpfulness, they must be designed with safeguards that clearly communicate and emphasize the uncertainty and the need for clinician verification.
Responses from post-survey questions further suggest that participants’ acceptance of AI-generated health information was conditional rather than unconditional. Participants often described AI as useful for general education and clarification. Meanwhile, they expressed concern about hallucination, privacy, over-reliance, and the absence of accountability when AI-generated information is used for consequential health decisions. In other words, participants perceived AI-generated answers as helpful and credible, but still resisted the idea that AI should replace professional medical advice. The qualitative findings point to a model of conditional trust, in which AI is accepted as a supplemental information source when it supports understanding, but is expected to remain secondary to clinician judgment when diagnosis, treatment selection, or personalized decision-making is involved.
Combining both quantitative and qualitative results, findings from this survey study highlight the potential application of AI chatbots in healthcare communication. While consultation and doctor appointment length varies widely across health systems, and time pressure is a persistent feature of clinical practice [20], patients may not be able to ask every question in an one-time visit since they may think of additional questions after meeting with the doctor, forget parts of the explanation, or need information repeated in simpler language. As a result, patients often need to prepare questions in advance, take notes during the visit, and seek clarification afterwards.
Unlike clinicians, AI systems and chatbots are available on demand and can generate explanations whenever patients need them. They can define medical terms, summarize treatment options, explain possible side effects, and help users prepare questions for a future appointment. Such availability is especially important for individuals who feel overwhelmed during clinical visits or who need additional time to process complex information. The implication is that AI may be applied to support patient care by extending the informational environment around clinical care. With responsible design, AI tools could help patients better prepare for conversations with clinicians. For example, in the prostate cancer scenario presented in this study, a patient could simply ask AI to explain the difference between active surveillance and radiation therapy or generate a list of questions to ask their oncologist.
However, while AI-generated health advice can be valuable on supporting patients understanding the knowledge as a supplemental tool, such information should be carefully evaluated to make a clinical decision. Health information differs from many other Internet information domains given the fact that it is personal, sensitive, and consequential. If an AI chatbot gives inaccurate advice about a hobby, a recipe, or a travel plan, the consequences may be inconvenient, but if it gives inaccurate or poorly contextualized health advice, the consequences may involve delayed care, unnecessary anxiety, misunderstanding of treatment options, or inappropriate self-management. In this situation, a general answer can be useful for education, but it should not be mistaken for patient-specific advice that implies sufficiency for diagnosis or treatment selection.
6. Conclusions
This study examined how individuals evaluate AI-generated health information compared with physician-provided information, focusing on perceived clarity, usefulness, appropriateness of details, trustworthiness, and perceived evidence. In a blinded survey context, participants rated AI-generated answers more favorably overall and across all five evaluation dimensions. Correlation analyses suggested that participant-level characteristics, including health literacy, AI literacy, and baseline trust in AI, were associated with some absolute ratings but did not clearly explain the relative preference for AI-generated answers. Qualitative responses further showed that participants valued answers that were detailed, specific, explanatory, and evidence-like, while also expressing concerns about hallucination, privacy, over-reliance, and the need for clinician verification. These findings suggest that AI-generated health information may be perceived as a helpful supplemental resource, but that perceived helpfulness should not be interpreted as evidence of clinical accuracy or appropriateness.
Future research should further examine how users evaluate and rely on AI-generated health information across different health topics, populations, and real-world contexts. Larger and more diverse samples are needed to understand how demographic, level of literacy, digital access, and prior healthcare experiences shape trust and reliance. Future studies should also compare user perceptions with expert clinical evaluations, since an answer that appears clear and trustworthy to users may still be incomplete, insufficient, or medically inappropriate. In addition, it is necessary to test design interventions such as source citations, uncertainty statements, personalization boundaries, privacy notices, and prompts to consult healthcare professionals. Such research can inform the development of AI health tools that support patient understanding while promoting transparency, verification, and appropriate reliance.
Author Contributions
Conceptualization, T.W. and M.B.; methodology, T.W.; validation, T.W. and M.B.; formal analysis, T.W.; data curation, T.W.; writing—original draft preparation, T.W.; writing—review and editing, M.B.; supervision, M.B.; project administration, M.B. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Institutional Review Board Statement
The study was approved by the Institutional Review Board of University of Illinois (protocol code IRB26-0304, date of approval: 03/13/2026).
Informed Consent Statement
Informed consent was obtained from all subjects involved in the study.
Data Availability Statement
The raw data supporting the conclusions of this article will be made available by the authors on request.
Conflicts of Interest
The authors declare no conflicts of interest.
Abbreviations
The following abbreviations are used in this manuscript:
| AI | Artificial Intelligence |
| LLM | Large Language Model |
References
- Tan, S. S. L.; Goonawardene, N. Internet health information seeking and the patient-physician relationship: a systematic review. J. Med. Internet Res. 2017, 19(1), e9. [Google Scholar] [CrossRef] [PubMed]
- Diviani, N.; Van Den Putte, B.; Giani, S.; van Weert, J. C. Low health literacy and evaluation of online health information: a systematic review of the literature. J. Med. Internet Res. 2015, 17(5), e112. [Google Scholar] [CrossRef] [PubMed]
- Meyer, A. N.; Giardina, T. D.; Khawaja, L.; Singh, H. Patient and clinician experiences of uncertainty in the diagnostic process: current understanding and future directions. Patient Educ. Couns. 2021, 104(11), 2606–2615. [Google Scholar] [CrossRef] [PubMed]
- Elliott, A. The culture of AI: Everyday life and the digital revolution; Routledge, 2019. [Google Scholar]
- Adla, Y. A.; El Hajj, A. Enhancing patient care with AI agents: integrating advanced chatbot technologies for improved healthcare delivery. In BMC Medical Informatics and Decision Making; 2026. [Google Scholar]
- Wang, L.; Wan, Z.; Ni, C.; Song, Q.; Li, Y.; Clayton, E.; Malin, B.; Yin, Z. Applications and concerns of ChatGPT and other conversational large language models in health care: systematic review. J. Med. Internet Res. 2024, 26, e22769. [Google Scholar] [CrossRef] [PubMed]
- Bedi, S.; Liu, Y.; Orr-Ewing, L.; Dash, D.; Koyejo, S.; Callahan, A.; Fries, J.; Wornow, M.; Swaminathan, A.; Lehmann, L. S.; Hong, H. J.; Kashyap, M.; Chaurasia, A. R.; Shah, N. R.; Singh, K.; Tazbaz, T.; Milstein, A.; Pfeffer, M. A.; Shah, N. H. Testing and evaluation of health care applications of large language models: a systematic review. Jama 2025, 333(4), 319–328. [Google Scholar] [CrossRef] [PubMed]
- Jia, X.; Pang, Y.; Liu, L. S. Online health information seeking behavior: a systematic review. In Healthcare; MDPI, December 2021; Vol. 9, No. 12. [Google Scholar]
- Stormacq, C.; Oulevey Bachmann, A.; Van den Broucke, S.; Bodenmann, P. How socioeconomically disadvantaged people access, understand, appraise, and apply health information: A qualitative study exploring health literacy skills. PLoS ONE 2023, 18(8), e0288381. [Google Scholar] [CrossRef] [PubMed]
- Lamers, R. E.; Cuypers, M.; Husson, O.; de Vries, M.; Kil, P. J.; Ruud Bosch, J. L. H.; van de Poll-Franse, L. V. Patients are dissatisfied with information provision: perceived information provision and quality of life in prostate cancer patients. Psycho-oncology 2016, 25(6), 633–640. [Google Scholar] [PubMed]
- Sun, Y.; Zhang, Y.; Gwizdka, J.; Trace, C. B. Consumer evaluation of the quality of online health information: systematic literature review of relevant criteria and indicators. J. Med. Internet Res. 2019, 21(5), e12522. [Google Scholar] [CrossRef] [PubMed]
- Zhang, Y.; Kim, Y. Consumers’ evaluation of web-based health information quality: meta-analysis. J. Med. Internet Res. 2022, 24(4), e36463. [Google Scholar] [CrossRef] [PubMed]
- Whitehouse, K. R.; Heslin, R. T.; Desir, A.; Raghunathan, R.; Islam, A. K.; Lallky, S.; Philip, P.; Reedy, N.; Mehta, A.; Parmer, M.; Piersall, M. V.; Oltmann, S. C.; Dackiw, A.; Maalouf, N. M.; Sant, V. R. Easing the burden: A pilot study evaluating AI-generated In-Basket message drafts to streamline perioperative endocrine surgical care. Surgery 2025, 109700. [Google Scholar] [PubMed]
- Baumgärtner, K.; Byczkowski, M.; Schmid, T.; Muschko, M.; Woessner, P.; Gerlach, A.; Bonekamp, D.; Schlemmer, H. P.; Hohenfellner, M.; Görtz, M. Effectiveness of the medical chatbot PROSCA to inform patients about prostate cancer: results of a randomized controlled trial. Eur. Urol. Open Sci. 2024, 69, 80–88. [Google Scholar] [CrossRef] [PubMed]
- Baur, D.; Ansorg, J.; Heyde, C. E.; Voelker, A. Development and Evaluation of a Retrieval-Augmented Generation Chatbot for Orthopedic and Trauma Surgery Patient Education: Mixed-Methods Study. JMIR AI 2025, 4, e75262. [Google Scholar] [CrossRef] [PubMed]
- Ayers, J. W.; Poliak, A.; Dredze, M.; Leas, E. C.; Zhu, Z.; Kelley, J. B.; Faix, D. J.; Goodman, A. M.; Longhurst, C. A.; Hogarth, M.; Smith, D. M. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA Intern. Med. 2023, 183(6), 589–596. [Google Scholar] [CrossRef] [PubMed]
- Chinn, D.; McCarthy, C. All Aspects of Health Literacy Scale (AAHLS): developing a tool to measure functional, communicative and critical health literacy in primary healthcare settings. Patient Educ. Couns. 2013, 90(2), 247–253. [Google Scholar] [CrossRef] [PubMed]
- Carolus, A.; Koch, M. J.; Straka, S.; Latoschik, M. E.; Wienrich, C. MAILS-Meta AI literacy scale: Development and testing of an AI literacy questionnaire based on well-founded competency models and psychological change-and meta-competencies. Comput. Hum. Behav. Artif. Hum. 2023, 1(2), 100014. [Google Scholar] [CrossRef]
- Choung, H.; David, P.; Ross, A. Trust in AI and its role in the acceptance of AI technologies. Int. J. Human–Computer Interact. 2023, 39(9), 1727–1739. [Google Scholar]
- Irving, G.; Neves, A. L.; Dambha-Miller, H.; Oishi, A.; Tagashira, H.; Verho, A.; Holden, J. International variations in primary care physician consultation time: a systematic review of 67 countries. BMJ Open 2017, 7(10), e017902. [Google Scholar] [CrossRef] [PubMed]
Table 1.
Participants’ Age Distribution.
| Age Group | N | Percentage |
|---|---|---|
| 25-34 | 16 | 25.8 |
| 35-44 | 19 | 30.6 |
| 45-54 | 10 | 16.1 |
| 55-64 | 9 | 14.5 |
| 65-74 | 7 | 11.3 |
| 75-84 | 1 | 1.6 |
Table 2.
Descriptive Statistics on the Three Measurements.
| Measure | # of Items | Mean | SD | Observed Range | Cronbach’s Alpha |
|---|---|---|---|---|---|
| AI Literacy | 16 | 0.47 | 3.88 | 2.94-5.00 | 0.79 |
| Health Literacy | 11 | 3.7 | 0.4 | 2.18-4.45 | 0.58 |
| Baseline AI Trust | 5 | 3.17 | 0.95 | 1.00-5.00 | 0.9 |
Table 3.
Participant Ratings of AI-Generated and Physician-Provided Health Answers.
| Evaluation Dimension | Physician-Provided, Mean (SD) | AI-Generated, Mean (SD) | Mean Difference: AI-Physician | 95% CI for Difference | t(61) | p |
|---|---|---|---|---|---|---|
| Overall Rating | 3.62 (0.66) | 4.16 (0.46) | 0.54 | [0.41, 0.66] | 8.63 | < .001 |
| Clarity | 4.12 (0.67) | 4.38 (0.44) | 0.26 | [0.11, 0.41] | 3.52 | < .001 |
| Usefulness | 3.69 (0.77) | 4.23 (0.50) | 0.55 | [0.39, 0.70] | 7.12 | < .001 |
| Appropriate Level of Details | 3.27 (0.81) | 3.92 (0.61) | 0.65 | [0.48, 0.82] | 7.74 | < .001 |
| Trustworthiness | 3.56 (0.73) | 4.15 (0.54) | 0.59 | [0.46, 0.72] | 9.05 | < .001 |
| Perceived Evidence | 3.47 (0.79) | 4.11 (0.61) | 0.64 | [0.51, 0.77] | 9.91 | < .001 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.