1 Department of Physiology, “Grigore T. Popa” University of Medicine and Pharmacy, 700115 Iași, Romania.
2 Department of Obstetrics and Gynaecology, “Grigore T. Popa” University of Medicine and Pharmacy, 700115 Iași, Romania.
3 Department of Urology, “Grigore T. Popa” University of Medicine and Pharmacy, 700115 Iași, Romania.
International Journal of Science and Research Archive, 2026, 19(02), 582-591
Article DOI: 10.30574/ijsra.2026.19.2.1071
Received on 02 April 2026; revised on 09 May 2026; accepted on 11 May 2026
Objective: To compare the quality of responses generated by a GPT-4.1-based chatbot, trained to respond strictly based on EAU and AUA clinical guidelines, with those produced by a collaborative panel of urology residents when answering common patient questions, as evaluated by blinded senior faculty members.
Methods: Ten clinical questions most frequently encountered in urological practice were selected by a resident panel covering pre-operative, post-operative, diagnostic, and oncological categories. Chatbot responses were translated into Romanian prior to evaluation. Three senior faculty members, blinded to response source, evaluated each response using a five-metric Likert scale (1–5): clinical accuracy, completeness, safety, empathy/language, and actionability. Mann-Whitney U tests with Bonferroni correction were used for group comparisons.
Results: The chatbot achieved significantly higher scores than residents for completeness (3.90 vs. 2.90, U = 86.5, p = 0.002, r = 0.730) and actionability (3.70 vs. 2.70, U = 86.0, p = 0.004, r = 0.720), both with large effect sizes. The overall composite score favoured the chatbot (3.66 vs. 2.86, p = 0.008). Empathy scores were comparably low in both groups. One chatbot clinical hallucination event was documented, yielding the lowest score of the series.
Conclusion: A GPT-4.1-based chatbot trained on clinical guidelines outperformed urology residents on completeness and actionability metrics. However, a clinical hallucination event highlights that mean superiority does not ensure consistent safety at the individual response level.
Artificial Intelligence; Urology; Patient Communication; Chatbot; Clinical Safety; Blinded Evaluation
Preview Article PDF
Theodor Florin Pantilimonescu, Denisa Oana Zelinschi, Alexandra Lazan and Viorel Dragoș Radu. Chatbot-generated versus urology resident-generated responses to common patient questions: A blinded evaluation by senior faculty. International Journal of Science and Research Archive, 2026, 19(02), 582-591. Article DOI: https://doi.org/10.30574/ijsra.2026.19.2.1071.






