Skip to content
akaturk Academic measurement

Article detail · 2026

Alignment Between Large Language Models and Consensus in Endoscopic Spinal Disc Surgery: A Comparative Analysis Against the Neurocore‐ SENSED

Journal

JOR Spine
OpenAlex Open access · gold SJR Q1 JCR Q1 Citations 0 Top 10% Percentile 91.8% FWCI 0.0
Year
2026
Type
article

Data source split

  • YÖKSİS venue JOR Spine
  • OpenAlex OpenAlex enrichment (abstract, citations, topics)

Abstract

OpenAlex · English

Background: The Neurocore-SENSED framework, derived from a three-round modified Delphi process involving 77 international spine surgeons, provides a structured reference standard for reporting in endoscopic spine surgery (ESS) for disc disease. The ability of large language models (LLMs) to reproduce graded levels of expert agreement within a reporting framework has not been examined in ESS. Objective: To evaluate the extent to which three contemporary LLMs align with the Neurocore-SENSED consensus and whether they discriminate between items of differing consensus level. Methods: In January 2026, each of the 166 framework items with published item-level endorsement data was submitted once to GPT-5.2, Gemini 3 Pro, and Claude Sonnet 4.6, in independent sessions with web retrieval disabled. Models selected one of five ordered response options directly. Alignment was assessed by Spearman correlation with panel endorsement and by four measures of categorical agreement, each with bootstrap 95% confidence intervals. Separate protocols examined response stability, response format, and prior exposure to the consensus. Results: < 0.001), which did not differ. Categorical agreement was limited and equivalent across models, with observed agreement of 0.506-0.518 and overlapping confidence intervals for all marginal-robust statistics. Between 81.3% and 95.8% of responses fell in the top two categories, and discordance was almost entirely unidirectional: 31 items were endorsed by every model but by fewer than half the panel, against one item in the opposite direction. Weighted kappa was not interpretable in this setting. Conclusion: Current LLMs reproduce the broad rank ordering of expert endorsement in ESS reporting but cannot reliably distinguish endorsed from contested propositions, principally because of a positive response bias.

Topics

  • Cervical and Thoracic Myelopathy
  • Spine and Intervertebral Disc Pathology
  • Patient-Provider Communication in Healthcare

Primary topic Cervical and Thoracic Myelopathy

Authors

No author information.