• LOGIN
    Login with username and password
Repository logo

BORIS Portal

Bern Open Repository and Information System

  • Publications
  • Theses
  • Research Data
  • Projects
  • Organizations
  • Researchers
  • More
  • Collections
  • Statistics
  • LOGIN
    Login with username and password
Repository logo
Unibern.ch
  1. Home
  2. Publications
  3. Conceptual proposal for LLM-generated FDG PET/CT follow-up reports in melanoma: a pilot study on model stability and blinded expert evaluation.
 

Conceptual proposal for LLM-generated FDG PET/CT follow-up reports in melanoma: a pilot study on model stability and blinded expert evaluation.

Options
  • Details
  • Files
BORIS DOI
10.48620/97437
Publisher DOI
10.3389/fnume.2026.1723650
PubMed ID
41909529
Description
Purpose
Oncological patients regularly undergo PET/CT re-staging, which requires a report that outlines their current disease status and highlights relevant changes compared to the previous PET/CT. Large language models (LLMs) may be helpful with documentation in the future. This study is a pilot on LLM performance, focusing on test-retest stability and reproducibility.
Methods
Three textbook melanoma follow-up cases of increasing complexity (involving one to eight organs) were selected. From standardized text-only prompts (no imaging data), follow-up reports were written by GPT-4o, Claude Sonnet 4 (each producing three independent revisions), and three nuclear medicine residents. This yielded nine reports per case (27 in total). Six blinded nuclear medicine experts (three internal, three external) performed test-retest evaluations of report quality and authorship identification.
Results
The cosine similarity analysis revealed high intra-case coherence (mean: 0.599-0.727) regardless of authorship. The external human readers consistently rated reports higher than the internal human readers. The LLM-generated reports received comparable or superior ratings to human reports, with Claude achieving the highest external reader scores (mean 0.926, standard deviation 0.263, on a 0-1 scale). Human performance declined with case complexity, while Claude, in particular, improved. The external readers significantly preferred the LLM impressions (Fisher's exact test, p = 0.005). Neither the human nor LLM readers reliably identified authorship (balanced accuracy 0.343-0.500).
Conclusion
In this pilot, blinded expert evaluation demonstrated that current LLMs can generate reports for melanoma [18F]fluorodeoxyglucose PET/CT of comparable quality to human-authored reports from text prompts in this study. High test-retest stability was obtained. Larger future studies will be required to confirm these findings.
Date of Publication
2026
Publication Type
Article
Subject(s)
600 Technology > 610 Medicine & health
Keyword(s)
AI in clinical practice
•
LLM
•
automation
•
medical technology
•
melanoma FDG PET/CT
Language(s)
en
Contributor(s)
Bosbach, Wolfram A.
Clinic of Nuclear Medicine
Heide, Marie S
Gözlügöl, Nasir
Clinic of Nuclear Medicine
Dana, Fatemeh
Clinic of Nuclear Medicine
Aghapour Zangeneh, Foroud
Clinic of Nuclear Medicine
Ventura, David
Schindler, Philipp
Roll, Wolfgang
Strunz, Franziskaorcid-logo
Clinic of Nuclear Medicine
Caobelli, Federicoorcid-logo
Clinic of Nuclear Medicine
Shi, Kuangyuorcid-logo
Clinic of Nuclear Medicine
Afshar-Oromieh, Ali
Clinic of Nuclear Medicine
Rominger, Axelorcid-logo
Clinic of Nuclear Medicine
Seifert, Robert
Clinic of Nuclear Medicine
Additional Credits
Clinic of Nuclear Medicine
Series
Frontiers in Nuclear Medicine
Publisher
Frontiers Media
ISSN
2673-8880
Access(Rights)
open.access
Show full item
BORIS Portal
Bern Open Repository and Information System
Build: dd892c [ 9.04. 8:30]
Explore
  • Projects
  • Funding
  • Publications
  • Research Data
  • Organizations
  • Researchers
  • Audiovisual Material
  • Software & other digital items
  • Events
More
  • About BORIS Portal
  • Send Feedback
  • Cookie settings
  • Service Policy
Follow us on
  • Mastodon
  • YouTube
  • LinkedIn
UniBe logo