INTRODUCTION
Since OpenAI’s introduction of GPT-3 in 2020, large language models (LLMs) have rapidly permeated daily life, creating significant ripple effects across industries and academics. Prior to the emergence of LLMs, artificial intelligence (AI) research in medicine primarily focused on limited specialties such as radiology and oncology. With the advent of LLMs, however, the scope of potential applications has expanded to nearly all medical fields. Once it was believed that AI would struggle with psychiatry, a discipline heavily reliant on verbal and emotional interaction between patients and physicians. However, current LLMs, which is capable of processing natural language in a human-like manner, are expected to benefit psychiatric diagnosis and treatment. These models can effectively analyze written case reports, verbatim transcripts, or even audio input from real-time doctor-patient interactions.
Given the rapid pace of advancement, it is reasonable to anticipate that LLMs will soon become integrated into clinical practice. Studies have already demonstrated that LLMs can pass the medical licensing examinations, offer more empathetic responses and useful information to patients than some physicians [
1]. In fact, in some challenging cases, they provide more accurate diagnoses than physicians [
2-
4]. Nonetheless, regardless of AI accuracy, legal and ethical concerns render fully autonomous AI decision-making without human oversight both risky and unethical. Consequently, AI-assisted decision-making—where AI offers recommendations while final decision authority remains with human—appears to be the most prudent path forward [
5,
6].
Physicians respond to this trend in various ways from skepticism that underestimates LLMs’ capabilities to fears that AI will eventually make their jobs obsolete. Regardless of one’s preconceptions about AI, both outright rejection and excessive dependence present significant concerns. Successful human-AI collaboration depends on their ability to accurately discern when to trust AI judgments over their own assessments. While trust calibration, which means how users adjust their trust based on an AI system’s actual reliability, was initially thought to depend solely on the AI model’s performance, recent research shows that it is influenced by multiple factors including transparency, explainability, and fairness [
7,
8].
Beyond AI’s characteristics, user-related factors play a crucial role. Individual personality traits, such as openness to new technologies and general propensity to trust and seek help from others, significantly influence AI acceptance [
9,
10], as does one’s ability to accurately assess her own competence [
11]. While people typically rely more heavily on AI in areas where they lack confidence, the opposite pattern can also emerge. The Dunning-Kruger effect (DKE), a cognitive bias whereby individuals with limited knowledge in a given domain systematically overestimate their own competence, may lead such users to resist AI assistance, perceiving it as unnecessary despite their actual skill gaps [
12-
14]. Conversely, highly competent users might uncritically trust AI after experiencing a couple of instances where AI’s judgment aligns with their own [
15,
16]. Indeed, a recent randomized trial found that granting physicians access to an AI assistant did not significantly improve their performance; in some cases, the collaborative results were even inferior to those achieved independently by either physicians or LLMs [
3]. This suggests a complex interplay between human and AI intelligence systems in collaborative diagnostic contexts.
To maximize the benefit from AI assistance, human users should accept AI-generated decisions when they perceive AI outputs to be substantially more accurate than their own judgments. However, this presents an inherent paradox: if users initially sought AI assistance due to their inability to make judgments in challenging cases, how can they evaluate whether the AI’s assessment is more accurate than their own? Consequently, trust in AI depends not so much on the objective accuracy of AI-generated decisions than their perceived reasonableness and trustworthiness. Furthermore, users’ acceptance of AI judgments would depend on the balance between their confidence in their own judgment and their trust in the AI system.
A couple of situations in which users place much trust on AI could be: 1) when AI reveals indisputable facts that users had previously overlooked or 2) when AI presents compelling evidence even though that may contradict the users’ initial judgment. Conversely, users may doubt their own judgment in the following situations: 1) when faced with particularly challenging tasks that are difficult to judge on their own or 2) when they lack general confidence in their competence in the domain. These situations provide potential avenues for experimental analysis. For instance, in the first scenario, acceptance of AI judgment depends on task difficulty, while in the second scenario, users with low competence might overly rely on AI regardless of how difficult the task is.
We hypothesized that users’ willingness to accept AI assistance depends on the balance between their perceived confidence in AI and self-assessment of their own competence. To test this hypothesis empirically, it was required to evaluate physicians’ appraisal of both 1) LLM accuracy and 2) their own competence. However, since directly measuring these subjective assessments proved impractical, an alternative approach was taken: examining the objective accuracy of LLM judgments for individual tasks and users’ overall accuracy across the task set, then evaluating how these two variables influenced users’ ability to improve task performance. Although our study focused on objective competence rather than subjective appraisal, this methodology allowed us to indirectly examine the factors influencing users’ decisions to accept or reject AI judgments. In addition to this main study objective, we aimed to investigate whether the DKE emerged in AI-physician interaction; whether users with lower competence might more stubbornly reject LLM recommendations.
METHODS
Overview of the process
Prepared case vignettes were presented to both AI and physicians, who were asked to assess the presence and severity of various psychiatric symptoms. Subsequently, the AI assessments were shared with physicians, who were then given the opportunity to revise their initial assessments to maximize accuracy. The accuracy of both AI and physician evaluations (initial and post-revision) was compared against the gold standard prepared by the authors before the experiment. Records were de-identified by removing direct identifiers (name, date of birth, address), in accordance with institutional privacy and research-ethics policies. The study was approved by the IRB of Eulji University Medical Center (EMC 2024-08-004).
Participants
Letters were sent to board-certified psychiatrists working in hospitals and private clinics throughout Daejeon and Sejong City with an introduction of the study. In addition to those who agreed to participate, three psychiatric residents from the authors’ affiliated hospital also joined the study. While initial invitations were sent to over 50 psychiatrists, the number of physicians who ultimately participated in the study was very limited. The participants’ clinical experience ranged widely, from 0 years (residents) to 28 years of practice as specialists. After receiving a detailed explanation of the research protocol, participants were required to sign a confidentiality agreement regarding the case vignettes before commencing the experiment.
Case vignette
Case vignettes were collected from all case conferences conducted at our department over the past decade. From these, 50 cases were selected that were both well-documented and representative of their respective diagnoses. All personally identifiable information was meticulously redacted to ensure complete anonymity, even upon detailed examination of the vignettes.
Symptom list for assessment
For each case vignette, a list of psychiatric symptoms was evaluated on a 4-point Likert scale ranging from none (0) to severe (3), assessing both the presence and severity. The symptom list was based on the Hierarchical Taxonomy of Psychopathology (HiTOP), a model developed by Kotov et al. [
17] that addresses the limitations of traditional diagnostic classification. Based on the psychopathology categories and symptoms in the HiTOP model, we curated a comprehensive list comprising 73 psychopathological symptoms across 13 categories, with some modifications to suit our research objectives.
To establish a gold standard for symptom assessment, two psychiatrist raters independently evaluated all case vignettes. One rater had approximately 4 years of clinical training in psychiatry, and the other was a senior psychiatrist with over 30 years of clinical and academic experience in psychopathology. Any discrepancies were resolved through discussion until consensus was reached, and these final consensus ratings served as the gold standard for subsequent analyses.
Building research platform
To facilitate remote participation, we developed a dedicated research website. Each participant was assigned unique login credentials (ID and password) to ensure secure access. Participants were only able to view their own assessment records. The website allowed participants to access and complete their evaluations at their convenience, with no imposed time constraints. The website was closed when no further assessments were being submitted.
Assessment and revision
During the initial study design, we considered using locally deployed open-source LLMs to minimize the risk of transmitting patient-derived information. We conducted preliminary testing with models available for local inference before initiating the study (April 2024), including LLaMA 3.1 8B and Qwen 2 7B via Ollama. However, these models produced substantially inferior outputs in psychiatric symptom assessment. We therefore selected GPT-4o as the primary LLM for this study, with appropriate de-identification procedures in place. Before participants began to involve, LLM assessments for all 50 case vignettes were completed using GPT-4o model (August 2024 snapshot; OpenAI), the company’s flagship model as of September 2024. Participants first conducted their initial assessments without knowledge of the LLM’s assessment results. Upon completion, the web interface displayed the LLM’s assessment results on the left screen and the participant’s results on the right screen. Participants were then given the opportunity to revise their initial assessments based on the LLM’s results. Notably, when the LLM identified the presence of a particular symptom (score ≥1), it provided supporting evidence and reasoning for its assessment, which was also displayed on the interface.
After a participant completed her revision for one case, the next case was displayed, continuing until all 50 cases were assessed. However, due to significant participant attrition during the study, the number of completed cases varied considerably among participants.
Analysis
Participants completed two rounds of assessments (initial and revision), while the LLM performed one assessment. The following indices were calculated for each of these three assessments:
Objective accuracy
This was calculated as the agreement with the gold standard. Given that the assessment was conducted on a 4-point Likert scale, two metrics were employed: percent agreement and weighted kappa.
Percent agreement
Agreement was defined by two criteria: 1) concordance in determining the presence or absence of symptoms, and 2) if symptoms were present, the difference in severity ratings should not exceed 1 point. For each symptom, we calculated the proportion of cases meeting these criteria between participant/LLM assessments and the gold standard.
Weighted kappa
Since percent agreement is significantly biased by symptom prevalence (proportion of scores >0), kappa coefficients were calculated for correction. Given the ordinal nature of Likert scale, linearly weighted kappa was employed.
Agreement between initial and LLM assessments
This was calculated using the same criteria as objective accuracy, yielding both percent agreement and weighted kappa.
Switch rate
Participants’ corrections of their initial assessments fell into two categories. When switching occurred in cases where initial and LLM assessments differed, it can be inferred that participants consulted the LLM results. Conversely, switching made when initial and LLM assessments aligned suggested independent reconsideration. Only the former type of switch rate was calculated.
Participant competence
While previous indices were calculated separately for each symptom and averaged across participants, participant competence was derived by averaging the objective accuracy of initial assessments across all symptoms. This metric represents each participant’s ability to accurately assess symptoms without LLM assistance.
Performance improvement
The extent to which participants improved their accuracy with LLM assistance was calculated as the relative improvement in kappa coefficient from initial to revised assessment. The degree of improvement is strongly influenced by the baseline kappa. Specifically, if the baseline kappa was low, there was considerable room for improvement, but if the baseline value was already high, the potential for improvement was very limited. Therefore, relative improvement (κimp) was defined as the following formula:
For statistical analysis, we primarily employed descriptive statistics with data visualization, and extensively utilized linear regression to identify factors influencing performance improvement. All visualizations and statistical analyses were conducted using R (version 4.4.2; R Foundation for Statistical Computing), an open-source statistical package.
DISCUSSION
In this study, we examined how physicians revised their initial assessments of psychopathological symptoms (I-assessment) after being provided with LLM-generated outputs (L-assessment). We also investigated what factors might have influenced their decision to revise their original assessments. In the analysis, several important findings could be observed.
First, the LLM demonstrated significantly higher accuracy than the human participants in the task. This superior performance was consistent whether measured by percent agreement or by weighted kappa coefficient. Furthermore, even after the participants revised their initial assessments, they were unable to match the accuracy level achieved by the LLM.
Second, when there were discrepancies between their and LLM’s judgment, the participants changed their mind in only one out of four cases. Had the participants fully adopted the LLM’s recommendations, the accuracy of the revised assessments (R-assessment) could have improved further. However, their apparent reluctance to fully embrace the LLM’s output resulted in only moderate improvements in accuracy.
Both symptom-level and user-level switch rates showed substantial variation, suggesting that symptom-specific and user-specific characteristics influenced the acceptance of LLM output. Regression analyses confirmed that accuracy improvement after reviewing LLM output was strongly associated with the LLM’s correctness; users were more likely to rely on the LLM when it provided reasonable and plausible recommendations. We also examined other factors, including symptom categories, prevalence, and diagnostic specificity. For instance, we hypothesized that physicians would trust the LLM’s assessment more for thought content symptoms than formal thought symptoms, and that common symptoms would garner more trust than rare ones. However, regression analysis revealed that none of these factors had a significant influence.
Although not included in the formal analysis, we also hypothesized that diagnostic specificity of symptoms would influence acceptance rates. For example, in schizophrenia patients, hallucinations would be diagnosis-specific, while insomnia would be nonspecific. The LLM’s assessment follows a purely bottom-up process, marking symptoms as positive whenever relevant cues appear in the case vignette, with equal attention to all symptoms. In contrast, physician assessment involves an interplay of bottom-up and top-down processes—once a provisional diagnosis is formulated based on salient symptoms, physicians tend to overlook symptoms less relevant to that diagnosis. Therefore, we anticipated that physicians would be more likely to rely on the LLM’s judgment for symptoms with low diagnostic specificity, since these symptoms were more likely to be omitted by physicians. However, diagnostic specificity of each symptom could not be objectively defined, and preliminary analyses showed no significant effects regardless of how specificity was defined. Consequently, this factor was excluded from the formal analysis.
User-level analysis revealed that participants with lower performance in their I-assessment were more likely to rely on LLM outputs. This suggests that users who were less confident in their own competence made maximum use of the LLM’s assistance. Therefore, undesirable phenomena like the DKE could not be observed. While years of clinical experience were included as another potential indicator of physician competence in the regression formula, this factor was not significant. It’s important to note that using user-level I-assessment accuracy as a measure of physician competence is hard to justify. Since participants received fixed material rewards regardless of accuracy, their motivation and dedication in completing the task may have varied. Nevertheless, both inherent competence and effort invested in the task inevitably influenced participants’ confidence in their I-assessments.
In summary, the degree to which participants improved their clinical performance with the help of LLM appeared to depend on the relative balance between their confidence in their own assessments and their trust in the LLM’s assessments [
18]. An interesting deviation from this pattern was that participants’ I-assessment accuracy for individual symptoms did not significantly influence the accuracy improvement. Although not statistically significant,
Figure 4’s left panel suggests that higher I-assessment accuracy was associated with greater improvement. If users had accurately evaluated the accuracy of their I-assessments, a negative correlation between the two variables would have been observed. This stands in stark contrast to the strong correlation observed with participants’ user-level assessment accuracy. While participants seemed able to accurately gauge their overall assessment reliability, they were unable to properly evaluate their accuracy in individual tasks. One possibility is that even participants who recognized their overall assessments as unreliable might have struggled to identify which specific symptoms they were more or less competent in evaluating. This finding suggests that, to achieve productive collaboration with AI, both users and AI systems need to calibrate their relative competencies across finely differentiated task categories.
User responses to AI assistance, particularly in decision-making, can fall into two extremes: algorithm appreciation [
19] and algorithm aversion [
20]. Algorithm appreciation refers to the tendency to place excessive trust in algorithmic judgments, assuming they are unbiased and error-free despite lacking direct evidence of superiority. Conversely, algorithm aversion describes how initial trust in AI can rapidly deteriorate after encountering a few errors, leading users to systematically discount its reliability [
20,
21]. While the present study could not identify such nuanced cognitive tendencies, the observed switch rate of 25% suggests an ambivalent response toward LLM outputs. Consequently, despite having the opportunity to revise their assessments, physicians did not fully utilize this opportunity.
To avoid these extremes, it is crucial for human users to accurately evaluate both AI capabilities and their own competence to achieve optimal trust calibration. In empirical trials, non-specialists were most susceptible to automation bias, whereas physicians with stronger training were more likely to catch AI errors, thus exhibiting algorithmic aversion [
15,
16]. Conversely, users with insufficient expertise may be reluctant to rely on algorithms, as demonstrated in the DKE [
12,
13]. Or, competent users may develop strong trust when encountering a couple of AI decisions that confirm their own thinking (confirmation bias), potentially leading to overreliance thereafter [
22,
23].
While robust capability assessment is necessary, it alone cannot prevent phenomena like algorithmic appreciation or aversion. Countless other factors can influence these tendencies. For one thing, the mode of AI response presentation can affect physicians’ trust. Recently emerging reasoning models not only provide more accurate results but also enhance user trust by showing the thinking process through which they reach conclusions. Similarly, if the model provides a convincing rationale or cites evidence, physicians might discard their own intuition in favor of the AI’s decision [
24]. Even the length and tone of LLM responses can affect trust. As LLM responses become more human-like, the risk of algorithmic appreciation increases, which is why some systems are deliberately designed to give more neutral, matter-of-fact responses [
25]. Conversely, as in our study, when LLM outputs are presented as simple numerical values rather than conversational text, they often fail to gain physician confidence.
Another consideration is the various cognitive biases humans exhibit during repeated decision-making tasks. While the purpose of implementing AI is to reduce physicians’ cognitive load, this goal may lead to undesirable phenomena. For example, in anchoring bias, physicians tend to base their judgment on AI’s initial output, causing AI’s early suggestion to anchor the physician’s final decision. Additionally, confirmation bias may occur when physicians place excessive trust in AI responses that align with their own views while disregarding others, potentially leading both parties to reach incorrect conclusions [
26]. The collaborative relationship between users and AI itself can introduce another form of bias. When users begin to view AI as a partner rather than merely a tool, a phenomenon known as confidence matching may emerge [
27]. This bias occurs when two entities work together on decisions, causing their confidence levels to gradually converge over time. For instance, even if users initially maintain accurate trust calibration between their own judgments and AI recommendations, repeated exposure to AI’s definitive responses may increase trust in the AI system regardless of its actual accuracy. Current LLMs do not provide confidence or uncertainty levels with their responses, making their outputs appear to be made with 100% certainty. Through repeated interactions with such systems, users naturally tend to develop increasingly higher levels of trust in AI. Finally, beyond accuracy considerations, there may be misalignment between human and AI values. For instance, while some clinicians might believe in aggressive treatment at the slightest suspicion of cancer, AI might take a more conservative approach. Such value discrepancies could make users distrustful and reluctant to accept AI assistance.
This study has several significant limitations. First and foremost, the number of participants was too small to be a representative sample of psychiatrists practicing in South Korea. Despite offering monetary compensation, the percentage of psychiatrist willing to participate in the study was lower than anticipated, and many of those who did participate did not complete all assigned cases. Second, the task of evaluating the presence and severity of individual symptoms based on case vignettes is far removed from actual clinical practice. Using verbatim transcripts of the real diagnostic interviews would have been more ideal. Unfortunately, obtaining a sufficient number of such transcripts was practically impossible. Still, the case vignette approach is not entirely divorced from clinical reality. Since the adoption of electronic medical record (EMR) systems, many hospitals’ discharge summaries require documentation of the presence and severity of important mental symptoms, sometimes even requiring initial severity and degree of recovery at discharge. If LLMs could be incorporated into EMR system, these tasks could be automated by processing progress notes.
A technical limitation is that key variables of interest, specifically the confidence physicians had in their own I-assessment and in the LLM’s L-assessment, could not be directly measured. We used assessment accuracy (agreement with the gold standard) as a proxy for confidence, though accuracy and confidence are not equivalent, and the gold standard itself may not represent absolute correctness in clinical assessment tasks.
Similarly, the degree to which participants accepted LLM outputs could not be measured directly. The switch rate alone was insufficient, as participants were instructed to maximize assessment accuracy by referencing LLM outputs, not simply to accept or reject them. We therefore used relative improvement in kappa from I- to R-assessment (κimp) as a proxy. Because this metric is strongly influenced by baseline kappa, we adopted the formula described in the Methods to account for ceiling effects. A limitation of this definition is that it does not incorporate L-assessment accuracy; however, given that the goal of R-assessment was to improve accuracy by any means available, we retained this approach.
Another important point is that participants viewed LLM outputs immediately after rating each case, rather than completing all initial assessments. This sequential design may have introduced an anchoring effect, with participants becoming progressively more reliant on LLM outputs over successive cases. However, internal pilot testing revealed that requiring participants to complete all 50 initial assessments before viewing any LLM outputs imposed a substantial burden, leading to diminished motivation to complete the process. Given that the study was conducted online, with participants accessing the platform across multiple sessions at their convenience, enforcing a strict separation between initial assessment and LLM exposure phases was not practicable.
With the exponential progress in LLMs like ChatGPT, AI systems have reached human-level capabilities in many cognitive tasks and may soon surpass even domain experts in knowledge and experience. People with limited exposure to LLMs may develop unrealistic expectations due to the enthusiasm of technology evangelists. Meanwhile, professionals who try to incorporate AI into their daily workflows often encounter its limitations, leading to disappointment and disillusionment. Finding an appropriate balance between optimism and realism will be essential as AI continues to redefine the boundaries of human responsibility. To avoid excessive algorithmic appreciation and algorithm aversion, trust calibration is necessary, which requires accurate estimation of both AI’s and humans’ strengths and weaknesses [
28].
To establish a productive collaborative relationship between humans and AI, a comprehensive approach encompassing several key aspects is necessary. Even when AI outputs appear convincing on the surface, there is inherent uncertainty and inaccuracy; thus, it is crucial to ensure explainability and transparency so users can fully understand how AI reaches its decisions [
29,
30]. Through ongoing collaboration and follow-up studies, the strengths, weaknesses, and limitations of both AI and humans must be identified, while mechanisms to monitor and prevent algorithm aversion and appreciation tendencies should be established. Given that collaboration with AI will likely become a necessity rather than a choice in future clinical practice, we hope that more physicians gain experience with AI systems and contribute their insights toward its improvement.