AI in Healthcare

How Reliable Are AI-Generated Citations? What the Evidence Says

Asking a language model to generate bibliographic references for a scientific text is one of the most common — and riskiest — ways to use AI in medical content. A study published in JMIR evaluated the accuracy of citations generated by ChatGPT (GPT-3.5) across two distinct academic domains and found that, of 102 citations generated in total, 72.7% in the natural sciences and 76.6% in the humanities corresponded to publications that actually existed [1].

The more concerning figure from that same study isn't how many citations exist, but how many are accurate: only 32.7% of citations in the natural sciences turned out to be completely correct when verifying author, title, and other metadata against the real source [1]. In other words, a citation can correspond to a real article and still have the wrong authorship, year, or details — an error that's far harder to catch than a reference invented from scratch.

A comparative analysis on using ChatGPT and Bard to support systematic reviews reaches a clear conclusion about this type of risk: these models are not recommended as the primary or exclusive tool for conducting systematic reviews, and any reference they generate requires thorough validation by researchers [2]. The same analysis notes that the high frequency of hallucinations in these models highlights the need to refine their training and operation before they can be confidently used for rigorous academic purposes [2].

For a medical or scientific team integrating AI into its workflow, the practical takeaway isn't to avoid these tools — it's to design the process around their known limitation: no AI-generated citation — whether or not the referenced article actually exists — should be incorporated into a manuscript, dossier, or medical presentation without a human expert verifying it against the original source, confirming authorship, year, journal, and that the cited content actually supports the claim made.

References

  1. Mugaanyi J, Cai L, Cheng S, Lu C, Huang J. Evaluation of Large Language Model Performance and Reliability for Citations and References in Scholarly Writing: Cross-Disciplinary Study. J Med Internet Res. 2024. doi:10.2196/52935. Available at: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11031695/
  2. Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis. PMC. Available at: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11153973/

Sound familiar, and you have a related project?

Tell us — at WriterTek we'll guide you through it.

Contact us →