Psychiatry Investig Search

CLOSE


Psychiatry Investig > Volume 23(8); 2026 > Article
Kim and Jeong: Determinants of Trust and Reliance on Artificial Intelligence in Clinical Settings: A Study on Physicians’ Use of Large Language Model in Psychiatric Symptom Assessment

Abstract

Objective

This study investigated physicians’ use of large language model (LLM)-generated assessments in psychiatric symptom evaluation and factors influencing their acceptance of artificial intelligence (AI) suggestions.

Methods

Twelve psychiatrists evaluated 50 anonymized psychiatric case vignettes, rating the presence and severity of 73 symptoms using a 4-point Likert scale. After initial assessment, they reviewed GPT-4o’s symptom ratings and could revise their evaluations. Accuracy was measured against a gold standard established by expert consensus. We computed initial (I), LLM (L), and revised (R) accuracy scores using percent agreement and weighted kappa. Performance improvement was measured by relative kappa increase. Regression analyses examined the influence of task difficulty and user competence on accuracy improvement.

Results

LLM assessments outperformed initial physician ratings in both agreement (90.4% vs. 82.9%, p<0.001) and kappa (0.713 vs. 0.559, p<0.001). After revision, physician performance improved (κ=0.651) but remained below LLM levels. The switch rate—cases where physicians revised in response to LLM disagreement—was modest (25.4%), indicating partial reliance on AI. Performance gains were positively associated with LLM accuracy and negatively associated with users’ own baseline competence, suggesting that less confident users rely more on LLM assistance.

Conclusion

Physicians demonstrated limited but strategic trust in LLM outputs, adjusting their judgments more when the LLM was accurate or when their own competence was lower. Miscalibrated trust—excessive skepticism or overreliance—led to missed gains. Effective human-AI collaboration in psychiatry requires tools and training for accurate self-assessment and AI trust calibration to avoid the pitfalls of algorithm aversion or overreliance.

INTRODUCTION

Since OpenAI’s introduction of GPT-3 in 2020, large language models (LLMs) have rapidly permeated daily life, creating significant ripple effects across industries and academics. Prior to the emergence of LLMs, artificial intelligence (AI) research in medicine primarily focused on limited specialties such as radiology and oncology. With the advent of LLMs, however, the scope of potential applications has expanded to nearly all medical fields. Once it was believed that AI would struggle with psychiatry, a discipline heavily reliant on verbal and emotional interaction between patients and physicians. However, current LLMs, which is capable of processing natural language in a human-like manner, are expected to benefit psychiatric diagnosis and treatment. These models can effectively analyze written case reports, verbatim transcripts, or even audio input from real-time doctor-patient interactions.
Given the rapid pace of advancement, it is reasonable to anticipate that LLMs will soon become integrated into clinical practice. Studies have already demonstrated that LLMs can pass the medical licensing examinations, offer more empathetic responses and useful information to patients than some physicians [1]. In fact, in some challenging cases, they provide more accurate diagnoses than physicians [2-4]. Nonetheless, regardless of AI accuracy, legal and ethical concerns render fully autonomous AI decision-making without human oversight both risky and unethical. Consequently, AI-assisted decision-making—where AI offers recommendations while final decision authority remains with human—appears to be the most prudent path forward [5,6].
Physicians respond to this trend in various ways from skepticism that underestimates LLMs’ capabilities to fears that AI will eventually make their jobs obsolete. Regardless of one’s preconceptions about AI, both outright rejection and excessive dependence present significant concerns. Successful human-AI collaboration depends on their ability to accurately discern when to trust AI judgments over their own assessments. While trust calibration, which means how users adjust their trust based on an AI system’s actual reliability, was initially thought to depend solely on the AI model’s performance, recent research shows that it is influenced by multiple factors including transparency, explainability, and fairness [7,8].
Beyond AI’s characteristics, user-related factors play a crucial role. Individual personality traits, such as openness to new technologies and general propensity to trust and seek help from others, significantly influence AI acceptance [9,10], as does one’s ability to accurately assess her own competence [11]. While people typically rely more heavily on AI in areas where they lack confidence, the opposite pattern can also emerge. The Dunning-Kruger effect (DKE), a cognitive bias whereby individuals with limited knowledge in a given domain systematically overestimate their own competence, may lead such users to resist AI assistance, perceiving it as unnecessary despite their actual skill gaps [12-14]. Conversely, highly competent users might uncritically trust AI after experiencing a couple of instances where AI’s judgment aligns with their own [15,16]. Indeed, a recent randomized trial found that granting physicians access to an AI assistant did not significantly improve their performance; in some cases, the collaborative results were even inferior to those achieved independently by either physicians or LLMs [3]. This suggests a complex interplay between human and AI intelligence systems in collaborative diagnostic contexts.
To maximize the benefit from AI assistance, human users should accept AI-generated decisions when they perceive AI outputs to be substantially more accurate than their own judgments. However, this presents an inherent paradox: if users initially sought AI assistance due to their inability to make judgments in challenging cases, how can they evaluate whether the AI’s assessment is more accurate than their own? Consequently, trust in AI depends not so much on the objective accuracy of AI-generated decisions than their perceived reasonableness and trustworthiness. Furthermore, users’ acceptance of AI judgments would depend on the balance between their confidence in their own judgment and their trust in the AI system.
A couple of situations in which users place much trust on AI could be: 1) when AI reveals indisputable facts that users had previously overlooked or 2) when AI presents compelling evidence even though that may contradict the users’ initial judgment. Conversely, users may doubt their own judgment in the following situations: 1) when faced with particularly challenging tasks that are difficult to judge on their own or 2) when they lack general confidence in their competence in the domain. These situations provide potential avenues for experimental analysis. For instance, in the first scenario, acceptance of AI judgment depends on task difficulty, while in the second scenario, users with low competence might overly rely on AI regardless of how difficult the task is.
We hypothesized that users’ willingness to accept AI assistance depends on the balance between their perceived confidence in AI and self-assessment of their own competence. To test this hypothesis empirically, it was required to evaluate physicians’ appraisal of both 1) LLM accuracy and 2) their own competence. However, since directly measuring these subjective assessments proved impractical, an alternative approach was taken: examining the objective accuracy of LLM judgments for individual tasks and users’ overall accuracy across the task set, then evaluating how these two variables influenced users’ ability to improve task performance. Although our study focused on objective competence rather than subjective appraisal, this methodology allowed us to indirectly examine the factors influencing users’ decisions to accept or reject AI judgments. In addition to this main study objective, we aimed to investigate whether the DKE emerged in AI-physician interaction; whether users with lower competence might more stubbornly reject LLM recommendations.

METHODS

Overview of the process

Prepared case vignettes were presented to both AI and physicians, who were asked to assess the presence and severity of various psychiatric symptoms. Subsequently, the AI assessments were shared with physicians, who were then given the opportunity to revise their initial assessments to maximize accuracy. The accuracy of both AI and physician evaluations (initial and post-revision) was compared against the gold standard prepared by the authors before the experiment. Records were de-identified by removing direct identifiers (name, date of birth, address), in accordance with institutional privacy and research-ethics policies. The study was approved by the IRB of Eulji University Medical Center (EMC 2024-08-004).

Participants

Letters were sent to board-certified psychiatrists working in hospitals and private clinics throughout Daejeon and Sejong City with an introduction of the study. In addition to those who agreed to participate, three psychiatric residents from the authors’ affiliated hospital also joined the study. While initial invitations were sent to over 50 psychiatrists, the number of physicians who ultimately participated in the study was very limited. The participants’ clinical experience ranged widely, from 0 years (residents) to 28 years of practice as specialists. After receiving a detailed explanation of the research protocol, participants were required to sign a confidentiality agreement regarding the case vignettes before commencing the experiment.

Case vignette

Case vignettes were collected from all case conferences conducted at our department over the past decade. From these, 50 cases were selected that were both well-documented and representative of their respective diagnoses. All personally identifiable information was meticulously redacted to ensure complete anonymity, even upon detailed examination of the vignettes.

Symptom list for assessment

For each case vignette, a list of psychiatric symptoms was evaluated on a 4-point Likert scale ranging from none (0) to severe (3), assessing both the presence and severity. The symptom list was based on the Hierarchical Taxonomy of Psychopathology (HiTOP), a model developed by Kotov et al. [17] that addresses the limitations of traditional diagnostic classification. Based on the psychopathology categories and symptoms in the HiTOP model, we curated a comprehensive list comprising 73 psychopathological symptoms across 13 categories, with some modifications to suit our research objectives.
To establish a gold standard for symptom assessment, two psychiatrist raters independently evaluated all case vignettes. One rater had approximately 4 years of clinical training in psychiatry, and the other was a senior psychiatrist with over 30 years of clinical and academic experience in psychopathology. Any discrepancies were resolved through discussion until consensus was reached, and these final consensus ratings served as the gold standard for subsequent analyses.

Building research platform

To facilitate remote participation, we developed a dedicated research website. Each participant was assigned unique login credentials (ID and password) to ensure secure access. Participants were only able to view their own assessment records. The website allowed participants to access and complete their evaluations at their convenience, with no imposed time constraints. The website was closed when no further assessments were being submitted.

Assessment and revision

During the initial study design, we considered using locally deployed open-source LLMs to minimize the risk of transmitting patient-derived information. We conducted preliminary testing with models available for local inference before initiating the study (April 2024), including LLaMA 3.1 8B and Qwen 2 7B via Ollama. However, these models produced substantially inferior outputs in psychiatric symptom assessment. We therefore selected GPT-4o as the primary LLM for this study, with appropriate de-identification procedures in place. Before participants began to involve, LLM assessments for all 50 case vignettes were completed using GPT-4o model (August 2024 snapshot; OpenAI), the company’s flagship model as of September 2024. Participants first conducted their initial assessments without knowledge of the LLM’s assessment results. Upon completion, the web interface displayed the LLM’s assessment results on the left screen and the participant’s results on the right screen. Participants were then given the opportunity to revise their initial assessments based on the LLM’s results. Notably, when the LLM identified the presence of a particular symptom (score ≥1), it provided supporting evidence and reasoning for its assessment, which was also displayed on the interface.
After a participant completed her revision for one case, the next case was displayed, continuing until all 50 cases were assessed. However, due to significant participant attrition during the study, the number of completed cases varied considerably among participants.

Analysis

Participants completed two rounds of assessments (initial and revision), while the LLM performed one assessment. The following indices were calculated for each of these three assessments:

Objective accuracy

This was calculated as the agreement with the gold standard. Given that the assessment was conducted on a 4-point Likert scale, two metrics were employed: percent agreement and weighted kappa.

Percent agreement

Agreement was defined by two criteria: 1) concordance in determining the presence or absence of symptoms, and 2) if symptoms were present, the difference in severity ratings should not exceed 1 point. For each symptom, we calculated the proportion of cases meeting these criteria between participant/LLM assessments and the gold standard.

Weighted kappa

Since percent agreement is significantly biased by symptom prevalence (proportion of scores >0), kappa coefficients were calculated for correction. Given the ordinal nature of Likert scale, linearly weighted kappa was employed.

Agreement between initial and LLM assessments

This was calculated using the same criteria as objective accuracy, yielding both percent agreement and weighted kappa.

Switch rate

Participants’ corrections of their initial assessments fell into two categories. When switching occurred in cases where initial and LLM assessments differed, it can be inferred that participants consulted the LLM results. Conversely, switching made when initial and LLM assessments aligned suggested independent reconsideration. Only the former type of switch rate was calculated.

Participant competence

While previous indices were calculated separately for each symptom and averaged across participants, participant competence was derived by averaging the objective accuracy of initial assessments across all symptoms. This metric represents each participant’s ability to accurately assess symptoms without LLM assistance.

Performance improvement

The extent to which participants improved their accuracy with LLM assistance was calculated as the relative improvement in kappa coefficient from initial to revised assessment. The degree of improvement is strongly influenced by the baseline kappa. Specifically, if the baseline kappa was low, there was considerable room for improvement, but if the baseline value was already high, the potential for improvement was very limited. Therefore, relative improvement (κimp) was defined as the following formula:
κimp=κrevised-κinitial1-κinitial.
For statistical analysis, we primarily employed descriptive statistics with data visualization, and extensively utilized linear regression to identify factors influencing performance improvement. All visualizations and statistical analyses were conducted using R (version 4.4.2; R Foundation for Statistical Computing), an open-source statistical package.

RESULT

General characteristics of participants

The participants comprised 12 board-certified psychiatrists and 3 psychiatric residents. Of the 50 available case vignettes, 33 cases were evaluated by two or more participants and thus included in the analysis. On average, each case was evaluated by 11.5 raters, and each clinician assessed 25.3 cases. Approximately half of the 73 psychopathological symptoms were excluded from the analysis due to low prevalence rates below 10% (based on the prepared gold standard). The final analysis encompassed 35 symptoms across 11 categories (Table 1).

Objective accuracy of assessment

According to the study design, for each symptom, three different measurements data were collected: 1) participants’ initial assessment (I-assessment), 2) the assessment by LLM (Lassessment), and 3) participants’ revised decision after being shared with the LLM results (R-assessment).
The objective accuracy of these three measurements was assessed by percent agreement and weighted kappa with a pre-established gold standard. The percent agreement for I-assessment ranged from 67.9% to 96.6%, those for L-assessment from 65.3% to 100% and 71.1% to 98.2% for R-assessment.
Notably, the LLM demonstrated exceptionally high agreement rates exceeding 99.5% for several symptoms, including grandiose delusion, soliloquy, compulsion, elevated mood, and feeling of worthlessness. When averaging the percent agreements across all symptoms, which was considered as overall accuracy, physicians’ I-assessment achieved 82.9% accuracy, while the LLM achieved a significantly higher rate of 90.4% (paired t-test, t=-5.29, df=34, p<0.001). Even after revision, the percent agreement of R-assessment remained at 87.3%, slightly lower than the L-assessment’s accuracy (statistically insignificant).
The percent agreements varied across symptoms. The LLM significantly outperformed participants in most cases, with the most notable differences in concentration difficulty (96.8% vs. 72.6%), irritable mood (97.1% vs. 73.9%), and feelings of worthlessness (99.5% vs. 79.2%). Participants performed better in only a few symptoms: paranoid delusion (87.6% vs. 77.1%), worry (72.9% vs. 65.3%), lack of motivation (73.9% vs. 71.1%), commanding voice (87.9% vs. 86.8%), and odd or illogical thinking (71.3% vs. 71.1%). However, these differences were less substantial than those where the LLM showed superior performance.
The accuracy analysis using kappa coefficients showed similar results. The average kappa coefficient was 0.559 for I-assessments and 0.713 for L-assessment, again demonstrating that LLM achieved significantly higher accuracy (paired t-test, t=-5.39, df=34, p<0.001). After revision, the kappa coefficient for R-assessment increased to 0.651, though it remained lower than the L-assessment (paired t-test, t=-2.88, df=34, p-value=0.007) (Figure 1).
The symptom specific kappa coefficients were displayed in Figure 2. The symptom with the largest difference between the I- and L-assessment was psychomotor agitation and that with the smallest difference was auditory hallucination. With the exception of obsession, paranoid ideation and odd or illogical thinking, the L-assessment demonstrated higher accuracy compared to I-assessment across all other symptoms.

Switch rate

The overall switch rate was 25.4%, indicating that participants changed approximately one quarter of the discrepant assessments, while maintaining their original decisions in the remaining three-quarters. This rate varied considerably across symptoms, ranging from as low as 10% for odd or illogical thinking and chronic fatigue to nearly 50% for decreased need for sleep and weight loss (Figure 3).
Although extremely rare, there were instances where participants revised their I-assessments even when they aligned with the L-assessment. This occurred because the study instructions encouraged participants to maximize accuracy regardless of LLM results.

Factors influencing accuracy improvement

After participants consulted the L-assessment results, the accuracy of R-assessment (overall kappa coefficient) increased from 0.559 to 0.651—a 16.5% improvement. However, this final performance fell short of the L-assessment’s 0.713, indicating that participants only partially accepted the LLM’s suggestions, leading to suboptimal improvement. This is further evidenced by the relatively modest switch rate of 25.4%, despite the LLM’s demonstrably superior performance.

Factors related with symptoms

To examine factors influencing participants’ acceptance of LLM judgment, both task (symptom) and user characteristics were analyzed. Task characteristics included category membership of the symptom (Table 1), sample prevalence (percentage of cases where each symptom was detected in all cases), and task difficulty. The I- and L-assessment accuracy showed a strong positive correlation (r=0.60, df=33, p<0.001), suggesting that symptoms challenging for humans were also difficult for LLM. Therefore, the objective accuracy of both I- and L-assessments were regarded as a proxy for task difficulty. In the regression analysis, only the L-assessment accuracy showed a significant association with relative improvement (β=0.408, t=4.49, p<0.001), while the I-assessment accuracy had no significant effect (Tables 2, 3, and Figure 4). However, the latter finding may be attributed to the strong collinearity between I-assessment and L-assessment masking the true effect of I-assessment.

Factors related with users

Unlike the analysis of task characteristics, the user level accuracy of I-assessment was the only factor significantly associated with relative improvement (β=-1.34, t=-2.94, p=0.014). Since relative improvement had already been adjusted for baseline kappa effects, these results suggest that users who could not accurately assess case vignettes in general were more likely to rely on LLM’s recommendations (Figure 5).

DISCUSSION

In this study, we examined how physicians revised their initial assessments of psychopathological symptoms (I-assessment) after being provided with LLM-generated outputs (L-assessment). We also investigated what factors might have influenced their decision to revise their original assessments. In the analysis, several important findings could be observed.
First, the LLM demonstrated significantly higher accuracy than the human participants in the task. This superior performance was consistent whether measured by percent agreement or by weighted kappa coefficient. Furthermore, even after the participants revised their initial assessments, they were unable to match the accuracy level achieved by the LLM.
Second, when there were discrepancies between their and LLM’s judgment, the participants changed their mind in only one out of four cases. Had the participants fully adopted the LLM’s recommendations, the accuracy of the revised assessments (R-assessment) could have improved further. However, their apparent reluctance to fully embrace the LLM’s output resulted in only moderate improvements in accuracy.
Both symptom-level and user-level switch rates showed substantial variation, suggesting that symptom-specific and user-specific characteristics influenced the acceptance of LLM output. Regression analyses confirmed that accuracy improvement after reviewing LLM output was strongly associated with the LLM’s correctness; users were more likely to rely on the LLM when it provided reasonable and plausible recommendations. We also examined other factors, including symptom categories, prevalence, and diagnostic specificity. For instance, we hypothesized that physicians would trust the LLM’s assessment more for thought content symptoms than formal thought symptoms, and that common symptoms would garner more trust than rare ones. However, regression analysis revealed that none of these factors had a significant influence.
Although not included in the formal analysis, we also hypothesized that diagnostic specificity of symptoms would influence acceptance rates. For example, in schizophrenia patients, hallucinations would be diagnosis-specific, while insomnia would be nonspecific. The LLM’s assessment follows a purely bottom-up process, marking symptoms as positive whenever relevant cues appear in the case vignette, with equal attention to all symptoms. In contrast, physician assessment involves an interplay of bottom-up and top-down processes—once a provisional diagnosis is formulated based on salient symptoms, physicians tend to overlook symptoms less relevant to that diagnosis. Therefore, we anticipated that physicians would be more likely to rely on the LLM’s judgment for symptoms with low diagnostic specificity, since these symptoms were more likely to be omitted by physicians. However, diagnostic specificity of each symptom could not be objectively defined, and preliminary analyses showed no significant effects regardless of how specificity was defined. Consequently, this factor was excluded from the formal analysis.
User-level analysis revealed that participants with lower performance in their I-assessment were more likely to rely on LLM outputs. This suggests that users who were less confident in their own competence made maximum use of the LLM’s assistance. Therefore, undesirable phenomena like the DKE could not be observed. While years of clinical experience were included as another potential indicator of physician competence in the regression formula, this factor was not significant. It’s important to note that using user-level I-assessment accuracy as a measure of physician competence is hard to justify. Since participants received fixed material rewards regardless of accuracy, their motivation and dedication in completing the task may have varied. Nevertheless, both inherent competence and effort invested in the task inevitably influenced participants’ confidence in their I-assessments.
In summary, the degree to which participants improved their clinical performance with the help of LLM appeared to depend on the relative balance between their confidence in their own assessments and their trust in the LLM’s assessments [18]. An interesting deviation from this pattern was that participants’ I-assessment accuracy for individual symptoms did not significantly influence the accuracy improvement. Although not statistically significant, Figure 4’s left panel suggests that higher I-assessment accuracy was associated with greater improvement. If users had accurately evaluated the accuracy of their I-assessments, a negative correlation between the two variables would have been observed. This stands in stark contrast to the strong correlation observed with participants’ user-level assessment accuracy. While participants seemed able to accurately gauge their overall assessment reliability, they were unable to properly evaluate their accuracy in individual tasks. One possibility is that even participants who recognized their overall assessments as unreliable might have struggled to identify which specific symptoms they were more or less competent in evaluating. This finding suggests that, to achieve productive collaboration with AI, both users and AI systems need to calibrate their relative competencies across finely differentiated task categories.
User responses to AI assistance, particularly in decision-making, can fall into two extremes: algorithm appreciation [19] and algorithm aversion [20]. Algorithm appreciation refers to the tendency to place excessive trust in algorithmic judgments, assuming they are unbiased and error-free despite lacking direct evidence of superiority. Conversely, algorithm aversion describes how initial trust in AI can rapidly deteriorate after encountering a few errors, leading users to systematically discount its reliability [20,21]. While the present study could not identify such nuanced cognitive tendencies, the observed switch rate of 25% suggests an ambivalent response toward LLM outputs. Consequently, despite having the opportunity to revise their assessments, physicians did not fully utilize this opportunity.
To avoid these extremes, it is crucial for human users to accurately evaluate both AI capabilities and their own competence to achieve optimal trust calibration. In empirical trials, non-specialists were most susceptible to automation bias, whereas physicians with stronger training were more likely to catch AI errors, thus exhibiting algorithmic aversion [15,16]. Conversely, users with insufficient expertise may be reluctant to rely on algorithms, as demonstrated in the DKE [12,13]. Or, competent users may develop strong trust when encountering a couple of AI decisions that confirm their own thinking (confirmation bias), potentially leading to overreliance thereafter [22,23].
While robust capability assessment is necessary, it alone cannot prevent phenomena like algorithmic appreciation or aversion. Countless other factors can influence these tendencies. For one thing, the mode of AI response presentation can affect physicians’ trust. Recently emerging reasoning models not only provide more accurate results but also enhance user trust by showing the thinking process through which they reach conclusions. Similarly, if the model provides a convincing rationale or cites evidence, physicians might discard their own intuition in favor of the AI’s decision [24]. Even the length and tone of LLM responses can affect trust. As LLM responses become more human-like, the risk of algorithmic appreciation increases, which is why some systems are deliberately designed to give more neutral, matter-of-fact responses [25]. Conversely, as in our study, when LLM outputs are presented as simple numerical values rather than conversational text, they often fail to gain physician confidence.
Another consideration is the various cognitive biases humans exhibit during repeated decision-making tasks. While the purpose of implementing AI is to reduce physicians’ cognitive load, this goal may lead to undesirable phenomena. For example, in anchoring bias, physicians tend to base their judgment on AI’s initial output, causing AI’s early suggestion to anchor the physician’s final decision. Additionally, confirmation bias may occur when physicians place excessive trust in AI responses that align with their own views while disregarding others, potentially leading both parties to reach incorrect conclusions [26]. The collaborative relationship between users and AI itself can introduce another form of bias. When users begin to view AI as a partner rather than merely a tool, a phenomenon known as confidence matching may emerge [27]. This bias occurs when two entities work together on decisions, causing their confidence levels to gradually converge over time. For instance, even if users initially maintain accurate trust calibration between their own judgments and AI recommendations, repeated exposure to AI’s definitive responses may increase trust in the AI system regardless of its actual accuracy. Current LLMs do not provide confidence or uncertainty levels with their responses, making their outputs appear to be made with 100% certainty. Through repeated interactions with such systems, users naturally tend to develop increasingly higher levels of trust in AI. Finally, beyond accuracy considerations, there may be misalignment between human and AI values. For instance, while some clinicians might believe in aggressive treatment at the slightest suspicion of cancer, AI might take a more conservative approach. Such value discrepancies could make users distrustful and reluctant to accept AI assistance.
This study has several significant limitations. First and foremost, the number of participants was too small to be a representative sample of psychiatrists practicing in South Korea. Despite offering monetary compensation, the percentage of psychiatrist willing to participate in the study was lower than anticipated, and many of those who did participate did not complete all assigned cases. Second, the task of evaluating the presence and severity of individual symptoms based on case vignettes is far removed from actual clinical practice. Using verbatim transcripts of the real diagnostic interviews would have been more ideal. Unfortunately, obtaining a sufficient number of such transcripts was practically impossible. Still, the case vignette approach is not entirely divorced from clinical reality. Since the adoption of electronic medical record (EMR) systems, many hospitals’ discharge summaries require documentation of the presence and severity of important mental symptoms, sometimes even requiring initial severity and degree of recovery at discharge. If LLMs could be incorporated into EMR system, these tasks could be automated by processing progress notes.
A technical limitation is that key variables of interest, specifically the confidence physicians had in their own I-assessment and in the LLM’s L-assessment, could not be directly measured. We used assessment accuracy (agreement with the gold standard) as a proxy for confidence, though accuracy and confidence are not equivalent, and the gold standard itself may not represent absolute correctness in clinical assessment tasks.
Similarly, the degree to which participants accepted LLM outputs could not be measured directly. The switch rate alone was insufficient, as participants were instructed to maximize assessment accuracy by referencing LLM outputs, not simply to accept or reject them. We therefore used relative improvement in kappa from I- to R-assessment (κimp) as a proxy. Because this metric is strongly influenced by baseline kappa, we adopted the formula described in the Methods to account for ceiling effects. A limitation of this definition is that it does not incorporate L-assessment accuracy; however, given that the goal of R-assessment was to improve accuracy by any means available, we retained this approach.
Another important point is that participants viewed LLM outputs immediately after rating each case, rather than completing all initial assessments. This sequential design may have introduced an anchoring effect, with participants becoming progressively more reliant on LLM outputs over successive cases. However, internal pilot testing revealed that requiring participants to complete all 50 initial assessments before viewing any LLM outputs imposed a substantial burden, leading to diminished motivation to complete the process. Given that the study was conducted online, with participants accessing the platform across multiple sessions at their convenience, enforcing a strict separation between initial assessment and LLM exposure phases was not practicable.
With the exponential progress in LLMs like ChatGPT, AI systems have reached human-level capabilities in many cognitive tasks and may soon surpass even domain experts in knowledge and experience. People with limited exposure to LLMs may develop unrealistic expectations due to the enthusiasm of technology evangelists. Meanwhile, professionals who try to incorporate AI into their daily workflows often encounter its limitations, leading to disappointment and disillusionment. Finding an appropriate balance between optimism and realism will be essential as AI continues to redefine the boundaries of human responsibility. To avoid excessive algorithmic appreciation and algorithm aversion, trust calibration is necessary, which requires accurate estimation of both AI’s and humans’ strengths and weaknesses [28].
To establish a productive collaborative relationship between humans and AI, a comprehensive approach encompassing several key aspects is necessary. Even when AI outputs appear convincing on the surface, there is inherent uncertainty and inaccuracy; thus, it is crucial to ensure explainability and transparency so users can fully understand how AI reaches its decisions [29,30]. Through ongoing collaboration and follow-up studies, the strengths, weaknesses, and limitations of both AI and humans must be identified, while mechanisms to monitor and prevent algorithm aversion and appreciation tendencies should be established. Given that collaboration with AI will likely become a necessity rather than a choice in future clinical practice, we hope that more physicians gain experience with AI systems and contribute their insights toward its improvement.

Notes

Availability of Data and Material

The datasets generated or analyzed during the study are available from the corresponding author on reasonable request.

Conflicts of Interest

The authors have no potential conflicts of interest to disclose.

Author Contributions

Conceptualization: Seong Hoon Jeong. Data curation: Ji-Min Kim. Formal analysis: Seong Hoon Jeong, Ji-Min Kim. Investigation: Ji-Min Kim. Methodology: Seong Hoon Jeong, Ji-Min Kim. Project administration: Seong Hoon Jeong. Software: Seong Hoon Jeong. Supervision: Seong Hoon Jeong. Validation: Seong Hoon Jeong, Ji-Min Kim. Visualization: Seong Hoon Jeong. Writing—original draft: Seong Hoon Jeong, Ji-Min Kim. Writing—review & editing: Seong Hoon Jeong.

Funding Statement

This study was supported by the Daeho Ethnic Psychiatry Research Fund (2024) from the Korean Foundation of Neuropsychiatric Research.

Acknowledgments

None

Figure 1.
Agreement between assessment outcomes of each symptom and the gold standard across different evaluators, comparing initial human assessment, large language model (LLM) assessment, and revised human assessment. (A) Percent agreement (%) and (B) weighted kappa coefficient.
pi-2025-0330f1.jpg
Figure 2.
Agreement with the gold standard (weighted kappa coefficient) for individual symptoms across initial human assessment, large language model (LLM) assessment, and revised human assessment. The results are presented in order of kappa values from the initial assessment.
pi-2025-0330f2.jpg
Figure 3.
Switch rate for individual symptoms. The switch rate was defined as the proportion of cases where the rater changed their initial assessment among the cases where the initial human assessment and large language model (LLM) assessment results differed. The white bar represents the proportion of cases where the rater altered their assessment despite the initial human assessment and LLM assessment results being identical.
pi-2025-0330f3.jpg
Figure 4.
Relationship between assessment accuracy and performance improvement after reviewing large language model (LLM) outputs. The left panel illustrates the relationship between physicians’ initial assessment accuracy and the relative improvement in diagnostic performance after reviewing the LLM outputs. The right panel shows the relationship between LLM assessment accuracy and the same performance improvement metric. Assessment accuracy was measured using weighted kappa coefficients in comparison to a gold standard. Performance improvement was quantified as the relative increase in weighted kappa from initial to revised assessments, adjusted for baseline values. A significant positive association was observed only for the LLM accuracy (β=0.408, t=4.49, p<0.001).
pi-2025-0330f4.jpg
Figure 5.
Relationship between participant competence and performance improvement after reviewing large language model (LLM) outputs. Participant competence was measured as the average weighted kappa of initial assessments across symptoms. Performance improvement was defined as the relative increase in kappa after reviewing LLM outputs. A significant negative association was found (β=-1.34, t=-2.94, p=0.014), indicating that less competent participants benefited more from LLM assistance.
pi-2025-0330f5.jpg
Table 1.
Psychopathology symptoms included in the final analysis and their respective category membership
Delusion Problems of thought content Formal thought disorder Perceptual problems Mania like symptoms Depression like symptoms
• Paranoid delusion • Idea of reference • Odd or illogical thinking • Auditory hallucination Elevated mood • Depressed mood
• Grandiose delusion • Obsession • Soliloquy • Irritable mood • Diminished interest
• Compulsion • Commanding voice • Impulsiveness • Loss of pleasure (anhedonia)
• Grandiosity • Lack of motivation
• Decreased need for sleep • Feeling of worthlessness
• Reckless behavior • Excessive guilt
Suicidal symptoms Anxiety symptoms Somatic symptoms Eating symptoms Cognitive symptoms
• Nonsuicidal injury • Anxiety • Insomnia • Decreased appetite • Concentration difficulty
• Suicidal idea • Panic-like symptoms • Nightnmare
• Suicidal attempt • Worry • Chronic fatigue
• Restlessness • Weight loss
• Psychomotor agitation
Table 2.
Linear regression analysis examining factors influencing the relative improvement of the kappa coefficient from independent assessment to updated assessment
df Sum Sq. Mean Sq. F-value p
Kappa for LLM assessment 1 0.270 0.270 33.492 <0.001***
Kappa for initial assessment 1 0.013 0.013 1.609 0.219
Symptom category 10 0.075 0.007 0.929 0.527
Symptom prevalence 1 0.002 0.002 0.275 0.605
Residuals 21 0.169 0.008

Analysis of symptom-related characteristics as potential predictors. An anlaysis of variance table of the regression model is presented, as the variable “symptom category” comprises multiple levels.

*** p<0.001.

LLM, large language model.

Table 3.
Linear regression analysis examining factors influencing the relative improvement of the kappa coefficient from independent assessment to updated assessment
Term β-coefficient Std. error T-value p
Career (year) 0.0029 0.0031 0.958 0.359
Sex 0.0566 0.0578 0.979 0.349
Kappa of independent assessment -1.34 0.46 -2.94 0.014*
No of observation 15
R2-adjusted 0.332

Analysis of rater characteristics as potential predictors.

* p<0.05.

REFERENCES

1. Ayers JW, Poliak A, Dredze M, Leas EC, Zhu Z, Kelley JB, et al. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA Intern Med 2023;183:589-596.
crossref pmid pmc
2. Cabral S, Restrepo D, Kanjee Z, Wilson P, Crowe B, Abdulnour RE, et al. Clinical reasoning of a generative artificial intelligence model compared with physicians. JAMA Intern Med 2024;184:581-583.
crossref pmid pmc
3. Goh E, Gallo R, Hom J, Strong E, Weng Y, Kerman H, et al. Large language model influence on diagnostic reasoning: a randomized clinical trial. JAMA Netw Open 2024;7:e2440969.
crossref pmid pmc
4. Liu X, Liu H, Yang G, Jiang Z, Cui S, Zhang Z, et al. A generalist medical language model for disease diagnosis assistance. Nat Med 2025;31:932-942.
crossref pmid pdf
5. MacIntyre MR, Cockerill RG, Mirza OF, Appel JM. Ethical considerations for the use of artificial intelligence in medical decision-making capacity assessments. Psychiatry Res 2023;328:115466
crossref pmid
6. Kleine AK, Kokje E, Hummelsberger P, Lermer E, Schaffernak I, Gaube S. AI-enabled clinical decision support tools for mental healthcare: a product review. Artif Intell Med 2025;160:103052
crossref pmid
7. Gille F, Jobin A, Ienca M. What we talk about when we talk about trust: theory of trust for AI in healthcare. Intell Based Med 2020;1:100001
crossref
8. Afroogh S, Akbari A, Malone E, Kargar M, Alambeigi H. Trust in AI: progress, challenges, and future directions. Humanit Soc Sci Commun 2024;11:1568
crossref pdf
9. Sharan NN, Romano DM. The effects of personality and locus of control on trust in humans versus artificial intelligence. Heliyon 2020;6:e04572.
crossref pmid pmc
10. Küper A, Krämer N. Psychological traits and appropriate reliance: factors shaping trust in AI. Int J Hum Comput Interact 2025;41:4115-4131.

11. Chong L, Zhang G, Goucher-Lambert K, Kotovsky K, Cagan J. Human confidence in artificial intelligence and in themselves: the evolution and impact of confidence on adoption of AI advice. Comput Hum Behav 2022;127:107018
crossref
12. Kruger J, Dunning D. Unskilled and unaware of it: how difficulties in recognizing one’s own incompetence lead to inflated self-assessments. J Pers Soc Psychol 1999;77:1121-1134.
crossref pmid
13. Rahmani M. Medical trainees and the Dunning-Kruger effect: when they don’t know what they don’t know. J Grad Med Educ 2020;12:532-534.
crossref pmid pmc pdf
14. He G, Kuiper L, Gadiraju U. Knowing about knowing: an illusion of human competence can hinder appropriate reliance on AI systems [Internet] Available at: http://doi.org/10.1145/3544548.3581025. Accessed September 15, 2025.

15. Kücking F, Hübner U, Przysucha M, Hannemann N, Kutza JO, Moelleken M, et al. Automation bias in AI-decision support: results from an empirical study. Stud Health Technol Inform 2024;317:298-304.
pmid
16. Kim SH, Schramm S, Riedel EO, Schmitzer L, Rosenkranz E, Kertels O, et al. Automation bias in AI-assisted detection of cerebral aneurysms on time-of-flight MR angiography. Radiol Med 2025;130:555-566.
crossref pmid pmc pdf
17. Kotov R, Krueger RF, Watson D, Achenbach TM, Althoff RR, Bagby RM, et al. The Hierarchical Taxonomy of Psychopathology (HiTOP): a dimensional alternative to traditional nosologies. J Abnorm Psychol 2017;126:454-477.
crossref pmid
18. Rezaeian O, Asan O, Bayrak AE. The impact of AI explanations on clinicians’ trust and diagnostic accuracy in breast cancer. Appl Ergon 2025;129:104577
crossref pmid
19. Logg JM, Minson JA, Moore DA. Algorithm appreciation: people prefer algorithmic to human judgment. Organ Behav Hum Decis Process 2019;151:90-103.
crossref
20. Dietvorst BJ, Simmons JP, Massey C. Algorithm aversion: people erroneously avoid algorithms after seeing them err. J Exp Psychol Gen 2015;144:114-126.
crossref pmid
21. Dietvorst BJ, Bharti S. People reject algorithms in uncertain decision domains because they have diminishing sensitivity to forecasting error. Psychol Sci 2020;31:1302-1314.
crossref pmid pdf
22. Hammond MEH, Stehlik J, Drakos SG, Kfoury AG. Bias in medicine: lessons learned and mitigation strategies. JACC Basic Transl Sci 2021;6:78-85.
pmid pmc
23. Ly DP, Shekelle PG, Song Z. Evidence for anchoring bias during physician decision-making. JAMA Intern Med 2023;183:818-823.
crossref pmid pmc
24. Benda NC, Novak LL, Reale C, Ancker JS. Trust in AI: why we should be designing for APPROPRIATE reliance. J Am Med Inform Assoc 2021;29:207-212.
crossref pmid pmc pdf
25. Wester J, de Jong S, Pohl H, van Berkel N. Exploring people’s perceptions of LLM-generated advice. Comput Hum Behav Artif Hum 2024;2:100072
crossref
26. Bashkirova A, Krpan D. Confirmation bias in AI-assisted decision-making: AI triage recommendations congruent with expert judgments increase psychologist trust and recommendation acceptance. Comput Hum Behav Artif Hum 2024;2:100066
crossref
27. Bang D, Aitchison L, Moran R, Herce Castanon S, Rafiee B, Mahmoodi A, et al. Confidence matching in group decision-making. Nat Hum Behav 2017;1:0117
crossref pdf
28. Schwartz JM, George M, Rossetti SC, Dykes PC, Minshall SR, Lucas E, et al. Factors influencing clinician trust in predictive clinical decision support systems for in-hospital deterioration: qualitative descriptive study. JMIR Hum Factors 2022;9:e33960.
crossref pmid pmc
29. Eke CI, Shuib L. The role of explainability and transparency in fostering trust in AI healthcare systems: a systematic literature review, open issues and potential solutions. Neural Comput Appl 2025;37:1999-2034.
crossref pdf
30. Bang CB, Jung YC, You SC, Kim K, Kim BH. Promises and pitfalls of large language models in psychiatric diagnosis and knowledge tasks. Br J Psychiatry 2025;226:243-244.
crossref pmid
TOOLS
Share:
Facebook Twitter Linked In Google+
METRICS Graph View
  • 0 Crossref
  •   Scopus
  • 710 View
  • 23 Download


ABOUT
AUTHOR INFORMATION
ARTICLE CATEGORY

Browse all articles >

BROWSE ARTICLES
Editorial Office
#522, G-five Central Plaza, 27 Seochojungang-ro 24-gil, Seocho-gu, Seoul 06601, Korea
Tel: +82-2-537-6171  Fax: +82-2-537-6174    E-mail: psychiatryinvest@gmail.com                

Copyright © 2026 by Korean Neuropsychiatric Association.

Developed in M2PI

Close layer
prev next