NASSJ Literature Review
Performance of ChatGPT versus Spine Surgeons as an Emergency Department Spine Call Consultant

Diego Garmendia, BA
Yale School of Medicine New Haven, CT

Jonathan N. Grauer, MD
Yale School of Medicine New Haven, CT
Article Reviewed
Taka TM, Sebt S, Meng S, Cabrera A, Shin D, Yacoubian V, Chao W, Rossie D, Xu Z, Erickson M, Rocos B, Than K, Yu E, Ahn N, Bono C, Cheng W, Danisa O. (2026). Performance of ChatGPT versus spine surgeons as an emergency department spine call consultant. North American Spine Society Journal (NASSJ), 25, 100836. https://doi.org/10.1016/j.xnsj.2025.100836
Abstract
Background: Large language models (LLMs) like ChatGPT are increasingly being recognized as credible tools for use across diverse health care settings. While artificial intelligence (AI) use has previously been evaluated in emergency medicine, its use in subspecialty care – particularly spine surgery – remains underexplored. This study evaluates the clinical accuracy, management appropriateness, completeness, helpfulness, and overall quality of ChatGPT responses compared to those of board-certified, spine surgeons in response to common emergency department (ED) consultations.
Methods: A 7-part questionnaire was developed based on common ED spine consultations (eg, Cauda Equina Syndrome, compression fracture in elderly patients, purulent drainage from surgical wound, acute lumbar disc herniation, incomplete spinal cord injury, epidural abscess, and metastatic spine disease). Each case included 3–4 questions pertaining to examination, diagnosis, management, and counseling. Responses from ChatGPT and 7 board-certified spine surgeons were restricted to 3–4 sentences per question. Three emergency medicine physicians rated each de-identified questionnaire response using a 5-point Likert scale. Statistical analysis was conducted using a 2-sample T-test with unequal variance. Inter-rater reliability was assessed using pairwise weighted Cohen’s kappa coefficient (κ).
Results: When comparing AI responses versus spine surgeon responses to proposed ED consultations, AI responses were rated to be superior across all 5 metrics of clinical accuracy, management appropriateness, completeness, helpfulness, and overall quality (p<.05). Inter-rater reliability was assessed using the average pairwise weighted Cohen’s kappa coefficient which showed substantial agreement (κ=0.76).
Conclusions: ChatGPT responses to emergency department spine consultations were rated as significantly higher compared to board-certified spine surgeons by emergency medicine providers. Though further improvement and validation is warranted, these findings suggest that ChatGPT can be a useful clinical adjunct for spine-related emergency department consultations.
Commentary
This review evaluates a study by Taka et al entitled, “Performance of ChatGPT versus spine surgeons as an emergency department spine call consultant,” published in the North American Spine Society Journal. The authors highlight the growing role of artificial intelligence (AI) in clinical decision-making. Accordingly, recent research efforts have focused on assessing the performance of large language models (LLM), such as ChatGPT, across various medical specialties. Despite this expanding body of work, there has existed a gap in the literature comparing clinical assessments between LLMs and those of spine surgeons in an emergency department (ED) setting.
Given potentially time-sensitive aspects of ED spine cases, the input of a spine surgeon may not always be readily available. Thus, assessing ChatGPT’s clinical aptitude in such settings may be particularly relevant. In light of this, the reviewed study aimed to assess the utility of AI-generated responses for common ED surgical spine scenarios by comparing responses to those given by spine surgeons.
The study employed a questionnaire which presented brief, common ED spine-related clinical vignettes to seven board-certified spine surgeons and ChatGPT-4o. The scenarios included Cauda Equina Syndrome, compression fracture in elderly patients, purulent drainage from spine surgical wound, acute lumbar disc herniation, incomplete spinal cord injury, epidural abscess, and metastatic spine disease. Each vignette was followed by three to four questions relating to clinical-decision making and patient counseling. To standardize responses, both the spine surgeons and ChatGPT were instructed to answer each question in 3-4 sentences. All completed questionnaires were subsequently deidentified and reviewed by 2 board-certified emergency medicine physicians and one emergency medicine resident.
The evaluating ED physicians graded each response using a 5-point Likert scale to assess clinical accuracy, appropriateness of management plan, completeness, helpfulness, and overall quality, with higher scores indicating superior performance. Statistical significance between spine surgeon and ChatGPT ratings was assessed using a 2-sample t-test with unequal variances. Interrater reliability was evaluated, and effect sizes were interpreted to assess clinical relevance.
The study analyzed a total of 168 responses, comprising 147 from spine surgeons and 21 from ChatGPT. Across all five evaluation domains, AI-generated responses received higher mean ratings than those provided by spine surgeons. Specifically, ChatGPT scored higher in clinical accuracy (pairwise weighted Cohen’s kappa of 4.67±0.91 versus 4.13±0.87, p=0.0176), management appropriateness (4.52±0.87 versus 3.81±1.06, p<0.01), completeness (4.95±0.22 versus 3.69±1.02, p<0.01), helpfulness (4.71±0.46 versus 3.81±1.05, p<0.01), and overall quality (4.71±0.46 versus 3.76±1.02, p<0.01). Interrater reliability was higher for AI-generated responses than for those provided by spine surgeons, with near-identical ratings observed for the completeness domain in AI-generated responses. Moreover, medium to large effect sizes were noted.
Overall, AI-generated responses were rated as superior to those of spine surgeons in all domains in the following five clinical vignettes: cauda equina syndrome, central cord syndrome, purulent surgical wound drainage, epidural abscess, and metastatic disease. In the acute lumbar disc herniation scenario, AI responses achieved higher or equivalent scores across all domains; however, differences in appropriateness, helpfulness, and overall quality did not reach statistical significance. Conversely, in the elderly compression fracture vignette, AI responses received lower mean scores for clinical accuracy and management appropriateness compared with spine surgeon responses, though these differences were also not statistically significant.
The reviewed study aligns with prior literature underscoring the potential utility of AI in clinical settings. Its findings suggest that ChatGPT may be a valuable adjunct for emergency medicine physicians managing spine-related complications, particularly in settings where timely access to spine surgeons is limited. These conclusions are strengthened by the high-inter rater reliability observed, which confirms that consistent criteria were used when rating both AI and spine surgeon responses.
The preference of AI-generated responses by emergency medicine physicians, however, must be interpreted within the study’s methodological constraints. While the use of common, well-defined vignettes is appropriate for evaluating ChatGPT’s clinical-decision making, such scenarios fail to encapsulate the complexity of real-world spine consultations. The absence of imaging, laboratory data, and complicating clinical factors limits the assessment of AI’s capabilities. Consequently, while the study supports the use of AI in routine cases, more complex decision-making will still need to be identified and flagged for further clinical output. The extent to which AI can effectively incorporate greater context and nuance, therefore, remains uncertain.
Beyond clinical performance, the reliance of large language models on expansive, publicly available datasets raises important concerns regarding misinformation and patient privacy when deployed in health care environments. As such, the integration of AI into medical evaluations and consultations must be continually reviewed to ensure clinical appropriateness, accuracy, and compliance with HIPAA regulations.
Several methodological limitations warrant consideration. Differences in tone or linguistic structure may have enabled evaluators to infer whether responses were generated by AI or by spine surgeons, introducing potential bias. The ratings in accuracy, appropriateness, completeness, helpfulness, and quality of the responses were from the perspective of the ED team and therefore reflect perceived, rather than objectively superior performance, as other evaluators may have judged them differently. Additionally, the unequal number of responses between groups may also have increased sampling variability and limited the robustness of statistical inference.
In summary, the reviewed study supports the use of AI as a clinical decision-support tool for common spine-related scenarios in the ED. These findings are particularly significant given the potential time-sensitive nature of such cases in situations where immediate subspecialty consultation may not be available.
Key Takeaways
- This vignette-based study assessed the quality and utility of AI-generated ED spine clinical responses by comparing them to responses from board-certified spine surgeons as rated by emergency medicine physicians.
- Emergency medicine physicians rated AI-generated responses as more clinically accurate, appropriate, complete, helpful, and higher in overall quality than those provided by spine surgeons.
- The observed differences between AI and spine surgeon responses demonstrated medium to large effect sizes, indicating meaningful clinical relevance.
Strengths of Study
- The study addressed a relevant topic with multiple clinical vignettes.
- Responses were obtained from spine surgeons representing various institutions.
- Emergency medicine physicians scored responses across five distinct domains.
Limitations of Study
- The use of standard, well-defined clinical scenarios likely does not encapsulate the complexity of real-world spine consultations. The absence of imaging, laboratory data, and complicating clinical factors limits the assessment of AI’s capabilities.
- Differences in tone or linguistic structure may have enabled evaluators to infer whether responses were generated by AI or by spine surgeons, introducing potential bias. Future studies should evaluate the success of blinding.
- The unequal number of responses between groups may also have increased sampling variability and limited the robustness of statistical inference.
Author Disclosures
D Garmendia: Nothing to disclose
JN Grauer: Deputy Editor, JAAOS. Editor-in-Chief, NASSJ