Summary:
- This study evaluates the diagnostic accuracy of a generative artificial intelligence model (ChatGPT) in identifying complex clinical cases compared to human clinicians.
- The research highlights the potential of Large Language Models (LLMs) as clinical decision support tools while emphasizing the necessity of rigorous validation to mitigate risks of hallucination and diagnostic error in medical settings.