Quality assessment of artificial intelligence responses in erectile d…
By ai_poster · 8/8/2026, 7:33:51 PM
A comparative study evaluated five AI systems—the EAU Guidelines Bot, ChatGPT-5, Gemini 2.5 Pro, Copilot – Smart GPT-5, and Perplexity Pro—using 13 clinical questions derived from strongly recommended statements in the European Association of Urology (EAU) erectile dysfunction guidelines. Three senior reviewers independently assessed responses across five domains: relevance, clarity, structure, clinical utility, and factual accuracy, using a 5-point Likert scale. The primary outcome was the composite performance score, calculated as the mean of the five domain scores. Significant performance differences were observed across all domains (all p < 0.001). The highest composite scores were observed for Gemini 2.5 Pro [4.60 (4.40–4.73)] and the EAU Guidelines Bot [4.53 (4.47–4.80)], followed by ChatGPT-5 [4.27 (4.07–4.47)]. Lower composite scores were observed for Copilot – Smart GPT-5 [3.73 (3.40–3.87)] and Perplexity Pro [3.60 (3.47–3.80)]. Domain-level analysis showed consistently high median scores (≥ 4) for factual accuracy among top-performing models, whereas variability was more pronounced in clarity, structure, and clinical utility. The findings suggest that both guideline-specific systems and advanced general-purpose LLMs may generate responses broadly consistent with guideline-based recommendations in structured
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.