Benchmarking Large Language Models in Acute-on-Chronic Pancreatitis: An Exploratory Study.
OA: gold
CC-BY-4.0
Abstract
Background/Objectives: Acute-on-chronic pancreatitis (ACP) is a fibro-inflammatory syndrome that stands at the border between acute pancreatitis (AP) and chronic pancreatitis (CP). It can be broadly defined as an acute exacerbation of CP. Large language models (LLMs) are artificial intelligence (AI) machine learning tools that can be employed in clinical decision systems (CDSs). Methods: This is an exploratory benchmark study. Six LLM chatbots were evaluated on a Likert scale by a panel of 13 experts in pancreatology based on the similarity of responses to a position statement (PS) in ACP. Results: ClaudeAI had the best average Likert score, while Alice had the lowest. There was a positive correlation between the length of the response and expert-perceived similarity. There was no correlation between seniority and Likert score. Conclusions: In this exploratory benchmark, some LLMs generated answers that were often judged as similar to the position statement, suggesting that they might be helpful as educational support tools and may assist in the refinement of research questions or in the identification of research gaps in this particular pathology.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.
SciLite annotations
organisms 1
human
Source provenance
- europepmc
- last seen: 2026-09-20T09:27:46.357103+00:00
- scilite
- last seen: 2026-09-20T10:02:19.494152+00:00
License: CC-BY-4.0
· commercial use OK
· attribution required
Per Europe PMC
Per Europe PMC