Title : Accuracy of ChatGPT in predicting six-week postoperative knee flexion following total knee replacement: A retrospective cohort study
Abstract:
Background: Artificial Intelligence (AI), particularly Large Language Models (LLMs) such as ChatGPT, is increasingly being explored for clinical prediction. However, their ability to predict early functional outcomes following Total Knee Replacement (TKR) remains uncertain. This study assessed the accuracy of ChatGPT in predicting six-week postoperative knee flexion and investigated factors associated with prediction error.
Methods: A retrospective cohort study was conducted involving 160 patients who underwent primary TKR at a UK tertiary centre between January and February 2025. Age, sex, Body Mass Index (BMI), diabetes, smoking status, American Society of Anesthesiologists (ASA) grade, and six-week postoperative knee flexion were obtained from electronic medical records. ChatGPT (GPT-5-mini) generated predicted six-week knee flexion using a standardized prompt based on these clinical variables. Predicted and observed flexion were compared using the Wilcoxon signed-rank test. Subgroup analyses were performed according to age, diabetes, smoking status, ASA grade, and BMI. Multivariable linear regression assessed predictors of absolute prediction error.
Results: Median observed knee flexion at six weeks was 95°, compared with 103° predicted by ChatGPT (p < 0.05). The median absolute prediction error was 10°. ChatGPT significantly overestimated flexion across most age groups and among diabetic and non-diabetic patients, smokers and non-smokers, and all ASA grades. Prediction error differed significantly by ASA grade, with median errors of 17° for ASA I, 9° for ASA II, and 6° for ASA III (p < 0.05). BMI was not significantly associated with prediction error. No clinical variable independently predicted absolute prediction error on multivariable analysis.
Conclusion: ChatGPT demonstrated moderate accuracy but consistently overestimated six-week postoperative knee flexion following TKR. Prediction errors were greater among younger and healthier patients and smaller among patients with greater comorbidity burden. Although LLMs may have potential as tools for postoperative outcome prediction, their current accuracy is insufficient to support clinical decision-making. Larger prospective studies using institution-specific datasets are required to establish their reliability and clinical utility.
Keywords: Artificial intelligence, ChatGPT, total knee replacement, knee flexion, postoperative outcome, large language models, arthroplasty

