Performance of the Large Language Model ChatGPT on the National Nurse Examinations in Japan: Evaluation Study

Taira, Kazuya; Itaya, Takahiro; Hanada, Ayame

doi:10.2196/47305

Cited by 34 publications

(9 citation statements)

References 13 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…Several studies assessed AI model performance in non-English languages with variable results despite the overall trend of below bar performance in non-English languages. For example, Taira et al tested ChatGPT performance in the Japanese National Nursing Examination in Japanese language in ve consecutive years [43]. Despite approaching the passing threshold in four years and passing the 2019 exam, the results indicated the relative weakness of ChatGPT in Japanese [43].…”

Section: Discussionmentioning

confidence: 99%

“…For example, Taira et al tested ChatGPT performance in the Japanese National Nursing Examination in Japanese language in ve consecutive years [43]. Despite approaching the passing threshold in four years and passing the 2019 exam, the results indicated the relative weakness of ChatGPT in Japanese [43]. Nevertheless, attributing this result to language limitations alone is challenging, given the superior performance of ChatGPT-4 in Japanese language compared to medical residents in the Japanese General Medicine In-Training Examination, as reported by Watari et al [44].…”

Section: Discussionmentioning

confidence: 99%

See 1 more Smart Citation

Superior Performance of Artificial Intelligence Models in English Compared to Arabic in Infectious Disease Queries

Sallam,

Al-Mahzoum,

Alshuaib

et al. 2024

Preprint

View full text Add to dashboard Cite

Background Assessment of artificial intelligence (AI)-based models across languages is crucial to ensure equitable access and accuracy of information in multilingual contexts. This study aimed to compare AI model efficiency in English and Arabic for infectious disease queries. Methods The study employed the METRICS checklist for the design and reporting of AI-based studies in healthcare. The AI models tested included ChatGPT-3.5, ChatGPT-4, Bing, and Bard. The queries comprised 15 questions on HIV/AIDS, tuberculosis, malaria, COVID-19, and influenza. The AI-generated content was assessed by two bilingual experts using the validated CLEAR tool. Results In comparing AI models' performance in English and Arabic for infectious disease queries, variability was noted. English queries showed consistently superior performance, with Bard leading, followed by Bing, ChatGPT-4, and ChatGPT-3.5 (P = .012). The same trend was observed in Arabic, albeit without statistical significance (P = .082). Stratified analysis revealed higher scores for English in most CLEAR components, notably in completeness, accuracy, appropriateness, and relevance, especially with ChatGPT-3.5 and Bard. Across the five infectious disease topics, English outperformed Arabic, except for flu queries in Bing and Bard. The four AI models' performance in English was rated as “excellent”, significantly outperforming their “above-average” Arabic counterparts (P = .002). Conclusions Disparity in AI model performance was noticed between English and Arabic in response to infectious disease queries. This language variation can negatively impact the quality of health content delivered by AI models among native speakers of Arabic. This issue is recommended to be addressed by AI developers, with the ultimate goal of enhancing health outcomes.

show abstract

Section: Discussionmentioning

confidence: 99%

Section: Discussionmentioning

confidence: 99%

Superior Performance of Artificial Intelligence Models in English Compared to Arabic in Infectious Disease Queries

Sallam,

Al-Mahzoum,

Alshuaib

et al. 2024

Preprint

View full text Add to dashboard Cite

show abstract

“…Consistently, prior research on the Japanese national medical examinations found that the performance gap between AI and humans widened with increasing question difficulty [ 12 ]. Indeed, AI models such as GPT-4 have achieved the proficiency level required to pass even highly challenging certification examinations that often pose challenges for many humans [ 2 - 5 , 11 , 12 ]. Because common clinical scenarios often follow a distinct framework or pattern, AI’s rule-based responses have the potential to surpass human performance [ 22 , 23 ].…”

Section: Discussionmentioning

confidence: 99%

“…This assessment is especially relevant because Japanese is considered among English natives as one of the most challenging languages to master [ 10 ]. Interestingly, it has been suggested that GPT-3.5, the precursor to GPT-4, has achieved passing grades on the Japanese Nursing Licensing examination [ 11 ]. In the latest Japanese national medical licensing examination in February 2023, GPT-4 attained passing levels while GPT-3.5 showed that it is not far behind the passing criteria [ 12 ].…”

Section: Introductionmentioning

confidence: 99%

Performance Comparison of ChatGPT-4 and Japanese Medical Residents in the General Medicine In-Training Examination: Comparison Study

Watari,

Takagi,

Sakaguchi

et al. 2023

JMIR Med Educ

View full text Add to dashboard Cite

Background The reliability of GPT-4, a state-of-the-art expansive language model specializing in clinical reasoning and medical knowledge, remains largely unverified across non-English languages. Objective This study aims to compare fundamental clinical competencies between Japanese residents and GPT-4 by using the General Medicine In-Training Examination (GM-ITE). Methods We used the GPT-4 model provided by OpenAI and the GM-ITE examination questions for the years 2020, 2021, and 2022 to conduct a comparative analysis. This analysis focused on evaluating the performance of individuals who were concluding their second year of residency in comparison to that of GPT-4. Given the current abilities of GPT-4, our study included only single-choice exam questions, excluding those involving audio, video, or image data. The assessment included 4 categories: general theory (professionalism and medical interviewing), symptomatology and clinical reasoning, physical examinations and clinical procedures, and specific diseases. Additionally, we categorized the questions into 7 specialty fields and 3 levels of difficulty, which were determined based on residents’ correct response rates. Results Upon examination of 137 GM-ITE questions in Japanese, GPT-4 scores were significantly higher than the mean scores of residents (residents: 55.8%, GPT-4: 70.1%; P<.001). In terms of specific disciplines, GPT-4 scored 23.5 points higher in the “specific diseases,” 30.9 points higher in “obstetrics and gynecology,” and 26.1 points higher in “internal medicine.” In contrast, GPT-4 scores in “medical interviewing and professionalism,” “general practice,” and “psychiatry” were lower than those of the residents, although this discrepancy was not statistically significant. Upon analyzing scores based on question difficulty, GPT-4 scores were 17.2 points lower for easy problems (P=.007) but were 25.4 and 24.4 points higher for normal and difficult problems, respectively (P<.001). In year-on-year comparisons, GPT-4 scores were 21.7 and 21.5 points higher in the 2020 (P=.01) and 2022 (P=.003) examinations, respectively, but only 3.5 points higher in the 2021 examinations (no significant difference). Conclusions In the Japanese language, GPT-4 also outperformed the average medical residents in the GM-ITE test, originally designed for them. Specifically, GPT-4 demonstrated a tendency to score higher on difficult questions with low resident correct response rates and those demanding a more comprehensive understanding of diseases. However, GPT-4 scored comparatively lower on questions that residents could readily answer, such as those testing attitudes toward patients and professionalism, as well as those necessitating an understanding of context and communication. These findings highlight the strengths and limitations of artificial intelligence applications in medical education and practice.

show abstract

“…Research has shown that ChatGPT can assist nurses in medical record documentation [ 37 ], enhance patient education resources [ 38 ], and successfully boost patient communication efficiency [ 37 ]. One study discovered that ChatGPT successfully passed the Japanese registered nursing licensure test [ 39 ]. While experienced nursing educators produce National Council Licensure Examination for Registered Nurses (NCLEX-RN) questions for daily test practice, the actual NCLEX-RN exam uses computer-generated questions based on the college’s current response scenario.…”

Section: The Impact Of Chatgpt On Nursing Educationmentioning

confidence: 99%

Embracing the future: Integrating ChatGPT into China’s nursing education system

Ni,

Peng,

Zheng

et al. 2024

International Journal of Nursing Sciences

View full text Add to dashboard Cite

Performance of the Large Language Model ChatGPT on the National Nurse Examinations in Japan: Evaluation Study

Cited by 34 publications

References 13 publications

Superior Performance of Artificial Intelligence Models in English Compared to Arabic in Infectious Disease Queries

Superior Performance of Artificial Intelligence Models in English Compared to Arabic in Infectious Disease Queries

Performance Comparison of ChatGPT-4 and Japanese Medical Residents in the General Medicine In-Training Examination: Comparison Study

Embracing the future: Integrating ChatGPT into China’s nursing education system

Contact Info

Product

Resources

About