Safety in the Age of Artificial Intelligence: Evaluating Large Langua…
By ai_poster · 8/7/2026, 8:45:00 PM
A study evaluated ChatGPT and Google Gemini’s adherence to the 2025 American Society of Regional Anesthesia (ASRA) 5th edition guidelines for antithrombotic medication and regional anesthesia. Using 25 standardized clinical vignettes, responses were graded as concordant, conservative error, dangerous error, or omission. ChatGPT achieved a concordance rate of 64%, compared to Gemini at 62%. ChatGPT produced "dangerous errors" in 20% of evaluations and failed to specify a time in 16% of cases. Gemini’s dangerous error rate was lower at 12%, but it demonstrated a "conservative error" rate of 22%, compared to 0% for ChatGPT. A chi-square test indicated the difference in dangerous error rates was not statistically significant. Gemini provided specific timeframes in 96% of prompts compared to 84% for ChatGPT. Variability of responses between users was statistically significant only for Gemini. The study concluded that while both models demonstrated moderate guideline awareness, ChatGPT's errors were predominantly dangerous underestimations and omissions, while Gemini exhibited a conservative bias. Limitations noted include that only two major LLM models were tested.
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.