Objective
2
This s tudy aims t o in v es tig a t e whet her the u s ag e of specific t erminologies ha s incr ea s ed, 3
f ocu sing o n wor ds and p hr a s e s fr equ en tly r epor t ed as ov erused by Cha tG P T . 4
Mat erials and Methods: 5
The lis t of 117 p o t entially AI - influen ced t er ms w as c ur a t ed based on pos ts and c ommen t s 6
fr o m anon ymous Cha t GPT u ser s , and 75 common ac ademic phr a se s w er e us ed a s c on tr ols . 7
PubMed r ecor ds fr o m 200 0 t o 2024 (u n til April) we r e analy zed t o tr ac k t he fr equency of 8
these t erms . Us ag e tr en ds w er e nor mali z ed usi n g a modified Z-sc or e tr an sf or ma tion. A 9
linear m ix ed- ef f ects model w as use d t o compa r e t he usag e of p o t enti ally AI-influe n c ed 10
t er ms t o common academic phr as e s ov e r ti me. 11
Results
12
A t ot a l o f 26,403, 493 PubMed r e c or ds w er e in ves tig a t ed. Among t he pot en t ially AI-13
influenced t erms, 74 di s p l a y ed a me aningful inc r eas e (modified Z-s c or e ≥ 3.5) in u s ag e in 14
2024. The linear mi x ed- e f f ec t s mo del s how ed a si gnifican t ef f ect of pot ent ially AI-influen c ed 15
t er ms on u s ag e fr eq uenc y c ompar ed t o c ommon ac ademic phr ases (p < 0. 001). The u s ag e of 16
pot en tiall y AI-influenced t e rms show ed a noticeable incr ea se s t arting in 2020. 17
Discu ssion: 18
This s tudy r evealed tha t cert ain wo r ds and phr a ses, s uch as "delv e," "under s cor e," 19
"meticulous," and "commendable," ha ve been used mor e f requen tly in medic al and 20
biologic a l f ields s ince t he i n tr oductio n of Ch at G PT . The us a g e r a t e of the se w or ds /phr a se s 21
has been i n c r eas ing f or sev er al y ear s be f or e the r e lea se of Chat G PT , s u g g e s ting that C hat G PT 22
migh t ha v e accel er a t ed th e popula ri ty of s cien tific e x pr es sions t ha t w er e alr eady g a ining 23
tr ac t i o n. 24
Conclusions
25
The iden t ified t erms in this stud y c an pr ovide v aluab le ins igh t s f or both LL M user s, edu ca t or s , 26
. CC-BY 4.0 International licenseIt is made available under a
is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)
The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint
3 / 16
and super visor s in t hes e fields. 1
2
Keywords
3
ChatGPT , large language models, academic writing, scientific terminology, PubMed, AI-influenced terms, 4
medical literature, language evolution. 5
6
. CC-BY 4.0 International licenseIt is made available under a
is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)
The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint
4 / 16
Introduction
1
Cha tG PT r apidly a chieved widespr ead gl o bal use aft er it s launch on No vember 30, 2022. 2
T r ained o n a vas t cor pus of t e x t d a t a, t he la r g e l angu a g e model ( L LM) i ncluding C ha tGPT 3
g ener a t es na tur al languag e wit h r e mark able fluency . Shor tly aft er i ts r eleas e, Cha tG PT ' s 4
applic a b i lit y f or s cien tific writ ing in medic al and b i o l ogi c al fields bec ame eviden t. 1,2 D u e t o 5
the f erv or s u rr oun ding i t s c apabiliti es, it w as c r edit ed as an aut hor o n sever al pap e r s , 6
igniting consider able debat e (curr ently , AI is not acknowledged as an author in sc holar ly 7
public a tions 3 ). Ther e w e r e ev en o p i nions t ha t the us e of Cha t G P T in paper writing w as 8
plagiarism 4 , but in r eality , L LMs su c h as Cha tG PT , Gemini, and Claude ar e a l r eady being u s ed 9
in paper wr i t ing. The us e of LLMs can be applied in v arious wa ys in academi c writ ing 1,5 and is 10
als o import a n t f or t h e resear ch activ ities of non-native r es ear cher s whos e fir st languag e is 11
not Engli sh. 6-8 Pr esen tly , a fr amew ork has been est abli shed t ha t per m its the u s e of L LMs in 12
writing , pr ovided their in v o lv e men t is adequa t e ly acknowledg ed. 3 13
Whil e LLMs c an pr odu ce na tur al w r iti ng , th e ir outpu t als o e xhibits cert ain char a ct er istics . 9,10 14
R ecen tly , it became a t op i c of discu ssion on X ( f ormer ly T wi t t er) and R e ddit t hat C ha tG P T 15
fr eq uently outp uts the w o r d 'delve' 16
(h t t ps:/ /ww w .r edd i t .co m/ r / mildlyinf ur ia ti ng / commen ts/ 1b z v gq j / appar en t ly _u s in g_the_w or17
d_delv e _is _a_sign_of_the/ [Acc e s s e d 2024, A pril 12] ). In add ition, r ecen t r eport s f ocus ing 18
on det ec t ing t e xt g ener a t ed by LL M s ha v e iden ti f i ed s ev er al f r equently us ed w or d s , su ch a s 19
‘ co mmendable, ’ ‘meticulous , ’ ‘in tric a t e, ’ and ‘r ealm. ’ 11-14 The extr ac t ion o f t hes e 20
char a ct eris tic k eyw or d s of L LMs in t hes e pr ev ious r epor ts w a s per f ormed b y c omparing 21
human- gen er at ed t e xt with Cha t G P T-g ener at ed t e x t. 11,13,14 While t his a ppr oac h r evealed 22
Cha tG PT' s char a ct er istics among the w or d s c ommonly used by both h uma ns an d Cha tG PT , it 23
had methodologic al li mit at ions in e x tr ac t ing w or d s wi t h low u s ag e fr equencies. Clar i f y ing 24
the w or d e xpr essions tha t L LMs t en d t o use in medical and bio logical p aper s is cru cial f or 25
desig n i n g ac ademic writ i n g suppor t and medical educ a t ion p r ogr ams. 15 Mor eove r , r ev ea l ing 26
. CC-BY 4.0 International licenseIt is made available under a
is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)
The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint
5 / 16
the e x t ent o f Cha t G PT' s impact on paper s in t he medic a l and bio logical fi el ds is es sen tial f or 1
main t a ining the f airness and r eli abili t y of academi c r e se ar c h and fr om the per spectiv e of 2
r esear ch ethi cs. 16 Ho wev er , the e xi s ting li t er a tur e l ack s a t hor ough in v est iga t ion of the 3
spe c ific w a y s in which Cha tGP T has tr ansf or med ac ademic writing pr actice s in the med ic a l 4
and biologic al discipli n es , ne cessit a ting fur ther r es ear c h . 5
As the use fu lness of L LMs bec omes mor e ev ident, t he numbe r o f r esear c her s usi n g LLMs f or 6
writing paper s has been gr adually in c r eas in g. 11,13,17 It woul d logically f o l l o w tha t th er e has 7
been an incr ease in the number o f r esear ch r epor ts f ea turing specific e x pr essions unique t o 8
LL Ms . This stu dy , ther e f or e, t es ts the h ypoth es i s t ha t the adopt ion of cert ain s c ien t if ic 9
t er m inologies has risen f ollow ing th e adv e n t o f Cha t G PT . F o cu sing on w or ds and phr ase s 10
fr eq uently r epo r t ed as used b y Cha tG PT , I i n ves tig a t ed PubMed r ec or d s fr om 2000 o nwar ds 11
and p e r f ormed a c ompar is on u s ing phr ase s c ommonly u s ed in ac ademia as a c on tr ol. This 12
a n a l y s i s a i m s t o e m p i r i c a l l y e x p l o r e t h e i n f l u e n c e o f L L M s o n t h e l e x i c o n o f m e d i c a l 13
lit era t ur e. 14
15
Methods
16
2 Met hods 17
2.1 Sear ch f or R ec or d s 18
Unli k e earlier s tud i es, 12-14 th i s r esear ch, dr awing in s igh t s fr om v arious a n on ymous en d-user s , 19
e xtr act ed po ten t ially AI- influen c ed t erms fr o m R eddi t, X (f orme rly T wi t t e r) , blogs , and 20
f or ums , f o c u sing on w or d s and phr a ses fr equ en tly pr oduced b y LL M s . Th e selection of t hes e 21
t er ms w as c arried ou t thr ough a r igor ous manual cur a tion pr oces s fr o m A pr i l 1 2 t o May 11, 22
2024, iden t ifyi n g 117 pot en t ia l ly A I-i nfluenced t er ms. In addition, a s a c o n tr ol gr oup, I us ed 23
the t op 100 c olloc a t ions iden t i fied a s char a ct er i s tic o f t he ac ademic c orpus i n a pr evious 24
st u dy . 18 Phr a s es tha t c ould be sear ch ed on PubMed a s tw o c on se c u tiv e w or ds w er e included 25
( f o r e x a m p l e , t h e c o l l o c a t i o n " b e t w e e n a n d " i s u s e d i n t h e f o r m o f " b e t w e e n A a n d B , " s o i t 26
. CC-BY 4.0 International licenseIt is made available under a
is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)
The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint
6 / 16
w as e x cluded a s no r ec or d s w er e f ound when sear chin g f or " b etw een and [T e xt W or d]"). In 1
the en d, 75 c ommon ac ademic phr as es w er e c h os en f or v erific a t ion in thi s s tud y . The lis t of 2
these phr ase s appear s in T able 1. 3
Tab le 1. Wor ds and ph r a s e s ex amined fo r usag e rat e s
Potentially AI-influenced terms
Verb s
ali gn, bols te r, burge on, captiv a te, c a talyz e, c ompel, delve , demys ti fy, di ve, e leva t e, eluc ida te,
emba r k , embr ac e, empl oy, enc ount er, e nde avor, e nha nce , enrich, explore, facil ita t e , fo s t er,
harne s s , leve r a ge, mi tig ate , nav igate , nu anc e, optimiz e, pa ve , pi one er, r e sona te,
revolu t i oni z e , s a f e gua rd, shed l igh t, sho wca s e , streaml in e, und e rscor e, unle as h, unloc k,
unve il, utili z e
Adjec tive s
ae sth etic, comme nda ble, compr ehen siv e , cruc ial, c utti ng-edg e , di sr up tiv e, dy namic ,
es sen t i al , ev er - e volv ing, f re sh, gro undbr eak ing, holi st i c, il lu striou s, imp era tiv e, in nov ative ,
inv aluabl e, in trica te, kee n, meticul ou s , methodica l , multi fac e ted, no ta bl e, no tew o r thy ,
ove r a r c hing, pa ramoun t, piv otal , poi sed, poten t, rob ust, se aml e ss, tail o r ed , tran s f ormative ,
unparal l eled, unp rece de nte d, unwa ve rin g, valua ble, v er s a til e, vi br a n t , vital
Adve r b s
addi t i on ally , aptly, e f fective ly, e xce ll entl y, impressiv ely, ingeni ou sly , moreove r, r epo r t edly ,
schola r ly , s t ra teg ical ly, ul t i mately, und ou btedly
Noun s
adv enture , be ac on, c apabi li ty, c omplex ity , c or ne rsto ne, di git al w or l d, driving fo rc e, e nigma,
game-chan ge r, hurdl e, i nte r pl ay, i nt e r s e ction , journey, kal e ido s c ope , l and scap e , new er a ,
paradi gm shi f t , prowe s s , r ea lm, spe arh e ad, sta te -o f- the -ar t, s y mphony, s y n ergy, ta pe stry,
t es t am e n t, t r eas u r e t r o v e
Common academic phrases (controls)
ac cording t o , a re pr e sent ed, a re show n, as s oc ia ted w ith , a s s oc iati on be tween, b a s e d o n,
betwe e n group s, ca n be , c hang e s in, ch a r a cte rized by, compa red to , compa r e d with,
con s i st ent with , c ontr ol group, c or r e la te d w it h , correl ation b etwe en, curr ent stu dy , data se t,
dec r e a se in, de fi ned a s , de fine d by, de te r mined by , dif fer ence s be t w ee n, d ue t o, ef fec t o n,
ex pos ur e to , f o ll ow up , have sho wn, hig h er than , in additi on, in co ntra st , inc r e a se in, indi cate
that , inte rac tio n betwe en, i s con si st ent, no signi fic ant , numb er o f, ob taine d from, occ ur renc e
of, ou r re sul ts, our s tu dy, over time , parti cipan ts we r e , perce nta ge o f, p re sen t st u dy ,
preva lence o f, p revio u s studi e s, rel ate d t o, r el a tion ship b etwe e n, re sp ect t o, re sul ts
obtain e d, s a mple siz e, s how tha t, sign ific ant b e t w een, sig ni fican t di f f e re nce, sig ni fican t l y
diff er ent, signi fic antly higher , s ta ti s tic ally s i gnific an t, sub s e t o f, s ugg es t th a t , t he s e fi ndings ,
the se r e s ul ts, thi s ar ticle , th is p aper , thi s st u dy , to de termin e, to eval uat e, w a s c al c ulated,
wa s me a s u red, wa s per fo rmed, w a s u s e d , were c ollec ted, w e r e de termin ed, wer e
signific antly, with r e spect
I used Pu bMed' s adv anc ed s ear ch f e a tur e (ht tp s :// pu bmed.ncbi.nlm.nih.g ov/adv anced/) t o 4
r ev eal the nu mbe r of r ec or ds in whic h these wor ds wer e us ed b y sear chin g f or " T e x t W or d" . 5
T o ensur e c ompr ehens iv e c ov er ag e o f ve r b f o r ms in Englis h, the sear ch quer y inc lud ed the 6
base f orm, t hir d per son singula r pr esen t, pr es en t participle/ pr ogr ess iv e, past t ense , and p a s t 7
. CC-BY 4.0 International licenseIt is made available under a
is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)
The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint
7 / 16
part ic iple. F or nouns, b oth s ingular and plur al f orms wer e inc or p or a t e d. Considering the 1
daily incr eas e in r ec or ds inde x ed in P ubMed, the sear ch c ondition s w er e st and a r dized fr om 2
Januar y 1, 2000 , t o April 30, 2024. T he s ear ch f ormulas f or all w or ds/ phr as e s ar e shown in 3
Supplemen t a r y T able S1. 4
5
2.2 Da t a Pr epar a tion 6
T o in v estig a t e t he usag e tr end s o f po t ent ial ly A I-inf luenced t erms in th e P ubMed da t abas e, 7
w e f i r st c alc u l a t ed t he usag e f r equ e nc y of each term by div iding the n umber of r ecor ds 8
c on t aining th e t e rm by the t ot al number o f r ec or ds in PubMed f or each y ear fr om 2000 t o 9
2024 (up t o Apr il 30, 20 24). Thi s pr o c es s yi elded a da t as et w it h usa g e fr equency f or ea c h 10
t er m and y ear . Ne x t , the modif ied Z-s c or e t r ans f ormat ion w as us ed t o nor malize t he usage 11
fr eq uenc y and f acilit at e c ompar i sons acr o s s t er ms and y ear s. F or ea ch t er m, the median and 12
median absolut e devi a t i on (MAD) wer e c alculat ed. The mo difi ed Z-sc or e w as comput ed b y 13
subtr a cting the median fr o m each occ ur r ence r a t e, div iding th e result by the MAD , and 14
multipl ying b y 0.674 5. T o iden t i f y significan t deviations in t er m u s ag e, we c ons ider ed an 15
absolut e modif i ed Z-scor e of 3.5 o r higher as indic a tive of a meani ngful incr ease or 16
decr ease 19 . Th e r esulting da t aset, c on t aining t h e modi f ied Z-scor es f or ea ch t e r m and year , 17
was t hen us ed f o r fur ther s t a t is tic al analy sis. 18
19
2.3 St a tis tic al A nalysis 20
A l inear mix ed -ef f ec t s model was used to c ompar e t h e usag e of po t en ti all y AI-inf luen c ed 21
t er ms and c ommo n a c ademic phr ases f r om 2000 t o 20 24. The d at a, consis t ing of mod ified Z-22
scor e s f or each wo r d or p hr ase, wer e obt ained and r es haped i n t o a long f orma t. The mod el, 23
. CC-BY 4.0 International licenseIt is made available under a
is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)
The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint
8 / 16
c onst ruct ed u s ing t he 'lme' function f r om t he 'nlme' pack age i n R, include d the modi fied Z-1
scor e s a s the dependen t v ariab le, t he g r oup (pot ent ial ly AI- influen c ed ter ms or c ommon 2
ac a d emic phr ase s) a s a fix ed e ff ect, and a r a n dom i n t er cept f or each w or d or p h r as e t o 3
accou n t f or r ep ea t ed measur es. T he model's s ummary w a s g ener a t ed t o asses s the 4
si gn i f ican ce of t he fix ed e ff ect of th e gr oup on t erm u s ag e. A line plot wi t h 95% co nf idenc e 5
in t erva ls w as c r ea t ed using the ' g gplot 2' pac k a g e t o vis u a li z e the tr end s in mean usag e f or 6
each gr oup fr om 2 000 t o 2024. The signific ance level f or a ll s t atis tic a l t est s wa s s et a t 0.05. 7
The analy s i s wa s perf or med using R ver s ion 4.3.2. 8
9
Results
10
A t o t al of 26,403,493 r eco r d s betw een J anuary 1, 20 0 0, and Ap ril 30, 20 24 w er e e xtr act ed 11
fr o m PubMed. The f r eq uenc y r a t es of each w or d/phr ase w er e det e r mined using the annu a l 12
t ot a l numb er of r ecor ds as th e d eno minat o r , f o llow ed by th e calcula t ion o f th e modifi ed Z-13
scor e. The Modif i ed Z-sc or e f or al l the w or ds and phr as e s a cr o ss all per iods is shown in 14
Supplemen t a r y T able S2. 15
In this stud y , amon g the 117 pot ent ia lly AI- inf luenc ed t e rms ver if ied, 74 w or ds /phr a se s 16
(lis t ed in desc ending or der : ‘ delv e, ’ ‘ under s c or e, ’ ‘meticulou s, ’ ‘ commendable, ’ ‘ show ca se, ’ 17
‘in tr icat e, ’ ‘t apestr y , ’ ‘ s ymphon y , ’ ‘i mpr ess iv ely , ’ ‘r ea lm, ’ ‘ cut t ing-edg e, ’ ‘pr owes s , ’ ‘ c a p tiv at e, ’ 18
‘not ewor th y , ’ ‘ gr oundbr eaking , ’ ‘unlo c k, ’ ‘ c ompel, ’ ‘lev er ag e, ’ ‘not able, ’ ‘un v eil, ’ ‘i n geniou s ly , ’ 19
‘piv ot al, ’ ‘bols t e r , ’ ‘holistic, ’ ‘ saf eguar ds, ’ ‘ elev at e, ’ ‘unwa v ering , ’ ‘t r ansf or ma ti v e, ’ ‘p i o nee r , ’ 20
‘ enigma, ’ ‘ embark, ’ ‘in v aluable, ’ ‘t es tamen t , ’ ‘nuance, ’ ‘mitig a te, ’ ‘ g ame-chang er , ’ ‘v aluable, ’ 21
‘ endea v or , ’ ‘imper at i ve, ’ ‘ cru cial, ’ ‘r ev oluti oniz e, ’ ‘unleash, ’ ‘ e f f ec t iv ely , ’ ‘ employ , ’ ‘ d igit al 22
world, ’ ‘f ost er , ’ ‘ demy s tified, ’ ‘multif acet ed, ’ ‘na vig a t e, ’ ‘ ev er - evol v ing , ’ ‘ str eamline, ’ 23
‘in t er section, ’ ‘utiliz e, ’ ‘h a r ness , ’ ‘ s h ed l igh t , ’ ‘ s tr at egi c ally , ’ ‘ seamles s, ’ ‘ en c oun t e r , ’ ‘ essen tial, ’ 24
‘ align, ’ ‘ addition al ly , ’ ‘pav e, ’ ‘poised, ’ ‘inno va tive, ’ ‘ syner g y , ’ ‘ c ompr ehens iv e, ’ ‘bur g eon, ’ ‘ aptly , ’ 25
. CC-BY 4.0 International licenseIt is made available under a
is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)
The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint
9 / 16
‘ div e, ’ ‘unpar alleled, ’ ‘ulti mat ely , ’ ‘vital, ’ ‘jo urney , ’ ‘ enhan ce’ ) di s pla y ed a modif ied Z - s c or e 1
e x c eeding 3.5 in 2 02 4. Whil e the majority of t he 75 c ommon ac ademic phr ase s (c on tr ols ) 2
display ed no s i gnific an t devi at ions i n us ag e r a t es, phr a s es su ch a s 'oc c urr enc e of ,' 't hes e 3
findings,' 'hav e shown, ' 'in t er action bet ween,' and ' char a ct er i z ed b y' s ur passed a modified Z-4
scor e o f 3.5 in 2024. On the ot her hand, the phr as es 'per cen t ag e o f ,' 'w as measur ed ,' 5
'number o f ,' 'with r espect, ' 'r es p e c t t o ,' and 't o det er min e' r egist er ed modifi ed Z-sco r es 6
below -3.5 in the same y ear (Figur e 1). 7
[Insert Figur e 1 h e r e] 8
The li n ea r mix ed - eff ects model r ev e aled a significan t ef f ec t o f t h e gr ou p (po t en tia lly AI-9
influenced t er ms vs. c ommo n a cad emi c ph r a s es) on the u sag e fr equenc y . The mod e l 10
show ed tha t the usag e of p ot e n tial ly AI- i n flu enced t er ms was significan t l y hig h er th a n tha t 11
of c ommon academic phr ase s (β = 0. 554, S E = 0.080, t(190) = 6.969, p < 0. 001). The line plo t 12
(Figur e 2) illus tr a t es t he tr en ds in mea n fr eq uenc y f or pot en t i a lly A I-infl uenced t erms and 13
c ommon academi c phr a ses fr om 2000 t o 2024. W hile the fr equency of the c on tr ol gr oup 14
r emains r el a ti v ely s t able, the p ot entially AI-influenced terms begin t o s ho w an incr ea s e 15
ar oun d 2016, with a not a b l e and st e ep upw ar d tr ajec t ory s t ar ting in 2 020 tha t bec omes 16
part ic ular ly pr onounced in 2023 and 2024. 17
[Insert Figur e 2 h e r e] 18
19
Discu ssion 20
This s tudy demonst r a t ed tha t, in the fields of medicine and bio logy , a number o f specific 21
w or ds and phr a se s , l ed by "delve," " under sc or e," "meti culous, " a n d "co mmendable," ha ve 22
c ome t o be used mo r e f r equ ently f ol lowing th e adven t of Cha tGPT . Th e i nc r eas in g t r end in 23
the usag e r at es of these wo r ds/ phr as e s w a s mor e pr ono u nc ed in 2024 than in 2023 in 24
almos t all c ases . Thi s ma y r e flect th e g ener ali z a tion of LLM us e among r es ear cher s i n the 25
. CC-BY 4.0 International licenseIt is made available under a
is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)
The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint
10 / 16
fields of med ici n e and bio logy , a s sh own in p r evious findings. 13 T h e l i s t o f o v e r u s e d t e r m s 1
su g ge s t ed in th i s st udy wil l help t hos e wr i t ing w it h L LMs cen t er ed ar o und C ha tGPT . 2
It has been obser v ed tha t med ic a l t e x t s g ener a t ed by Cha tGP T , while fluent a n d l o gi c al, t end 3
t o incl u de less s pecific inf orma tion and mor e g ener alized e xpr ess ions c ompar ed t o t hos e 4
autho r ed by huma n s , whi ch f ea tur e a richer and mor e d iver s e c on t en t. 10 In g ener al paper s , 5
it has been no t ed t ha t Cha tGPT t ends t o 1) use the same st y le and e xpressio ns r ep eat ed ly , 2) 6
show a dec r ea s e in t he f r equency of bas ic ver bs lik e ‘is’ and ‘ ar e, ’ and 3) f r eque n tly us e 7
adjectiv es and adv erbs. 11 P articularly f or adjectiv es and adver bs , n umer ous w or ds tha t 8
Cha tG PT fr equen t ly use s ha v e been poin t ed ou t. 14 Bec aus e this st udy only c oun t ed the 9
r ec or ds wher e spec if i c w or ds or phr ases o ccurr ed, it did not ev alu a t e the w eigh t of t e r ms 10
appearing mu l t ipl e tim es. In the cur r en t s tudy , s e v er a l w or ds t ha t w er e pr ev iously iden tifie d 11
a s f r e q u e n t ly u s e d b y C h a t G P T d i d n o t e x h i b i t a n o t a b le in cr e a se i n u s a g e ; y e t , C h a t GP T ma y 12
actually ov eruse these wor d s mor e than s u g g e s t ed by th e r esults o f t his st ud y . Similarly , 13
fr eq uently used ver bs s u ch a s ‘ enh a n c e’ , ‘ eleva t e’ , and ‘ut ili z e’ ma y al s o h av e been o v er us ed 14
by LL M s mor e than sugges t ed by thi s st ud y . 15
A pr evious r epor ts' limi t a tion li es in their lack of f oc u s on the spe c ific w or ds or t er ms 16
ov erused by Cha t GPT , thus f ailing t o c ompr eh ens iv ely e xplor e char act eris tic t er ms . As 17
e xt ensi v ely debat ed onli ne 18
(h t t ps:/ /ww w .r edd i t .co m/ r / mildlyinf ur ia ti ng / commen ts/ 1b z v gq j / appar en t ly _u s in g_the_w or19
d_delv e _is _a_sign_of_the/ [Acce s se d 2024, Apr i l 12]) , the incr eas ed u sag e of the w or d 20
‘ delv e’ w a s in cr edibly pr onounced c o mpar e d t o other w or ds or ph r as e s, with a modified Z-21
scor e of ar ound 100. D es pit e its ov e rwhelming pr esence, pr evious paper s compar ed t e x t s 22
cr ea t ed b y humans and Cha tG P T 12-14 d i d not mention ‘d e l v e , ’ highligh ting a s tr eng t h of t his 23
s tudy' s methodo l o gy . This s tud y c an not con c lusiv el y est ablish t he c onne ction bet w een the 24
fr eq uent use of ‘ del v e’ and t he e mer g ence of Cha tGPT , alth ough its impact is highly 25
su s pect ed. The fr equen t use o f the t erm ‘ delve’ by Cha tG PT could be a t tribu t ed t o i t s 26
. CC-BY 4.0 International licenseIt is made available under a
is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)
The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint
11 / 16
pr ominenc e in th e tr ain ing da t a, p o ssibly r e sulting fr o m c ommo n ins truc t ions dur i n g the 1
r einf o r cemen t learn ing fr o m human f ee dback pha s e, or as a f e a tur e of lar g e la n g u a g e 2
models designed t o pr oj ec t aut hori ty ; how ev er , th e se h y p othese s r emai n s pecula tiv e and 3
unc onfir med. 4
Not ably , th e fr equenc y of u se f or t he p ot en t ially A I-influ enced t erms in ves tig a t ed in t his 5
s tudy had al r eady div er g ed mar k edly even b e f or e ChatGP T r eleased in Nov ember 2022. In 6
part ic ular , 'delve' w as used e x ceptionally oft en in 2024, but i t had also b een used at an 7
e x c ept ionally high fr equency in ac ademic writing sinc e 2021. One h ypothes is ar is es fr om t his 8
obser v a tion―man y of the pot en tia ll y AI-influenced t er ms ma y ha ve c o n tribut ed to their 9
ov eruse in Cha tGP T , as t he s e e x p r essions g ained po pularity du ring t he period of int ensive 10
LL M tr a ining. In other w or ds, Cha tGP T ma y ha v e a c c eler a t ed the i nev it able t empor a l 11
chang e s in writ ing i n r es ear c h . H owever , t his h ypot hes i s would be dif ficult t o verify , since we 12
c annot obser v e a par allel unive r s e wher e C ha tG P T d oe s not exis t. 13
In t er estingly , some of t he co mmon a c ademic phr a se s u sed as co n tr ols also d ev ia t ed i n thei r 14
pr op orti on of use in 2024. The f o ur ph r as e s 'o c curr ence o f', 'these finding s', 'h a v e s h ow n ' , 15
and 'in t er a ct i o n be tw een' sig n i f ican t l y incr e a sed in fr equency of use in 2024, but s ince t he y 16
ar e all v ery co mmonly u s ed e xpr es sions in ac ademic writ i n g , it w ould be dif fic u l t f o r us 17
humans t o r ec ognize tha t their f requency ha s in cr eased. Con v er sely , the six phr a se s 18
'per cen t ag e of', ' wa s mea sur ed', 'nu mber of', 'with r es pe ct', 'r es pe ct t o', and 't o det ermine ' 19
not ably dec r eas ed in usag e in 2024. W h en int er pr e ting these r esults, w e must r e membe r 20
tha t the languag e u s ed in paper s na t ur all y ev olv es ov er t i me; 20 many p hra se s that dec r eas ed 21
in frequency had alr eady b een dec lini ng even be fo re the introduct i o n o f C hatGP T. However, 22
the two phr as es 't o d eter mine' and 'number of ' did no t show a not ic eable decrease in 23
fr equency of us e befo re 2022, and their f r equency of use appear s to have dec r eas ed 24
si gn i f ic ant l y af ter 20 23 (see Supplementar y Ta b l e 2). Whil e t his result ma y be c oincident a l, 25
. CC-BY 4.0 International licenseIt is made available under a
is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)
The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint
12 / 16
it could a lso indic at e tha t t he p roli feration o f C hatG P T may subt ly lead to t he decreased us e 1
of cert ai n wor ds or phrases with out us reco gniz ing it. 2
This s t udy ha s some limita t ions. The most important li mi tation is that the t erms pot en tial ly 3
influenced by LLMs in this stu dy were iden tifi ed t hrough manual ins pection r a t her than 4
being extr ac t ed in an objec t iv e and s ystemat ic manner. Therefor e, there may be words or 5
phr a ses that were not i n c luded in t h i s stud y yet have seen a s ignificant i n c r eas e i n u sa ge 6
post-Ch a t GPT. Addition al ly, tempor al shifts in the f requency of word or phrase u se could 7
have been inf l u enc ed by external f a c to rs s uch a s evolving resear ch tr ends and shifts in the 8
style of scientif i c commun i cation, f act ors not ac c o unted for in this s t ud y. Lastly , the ab senc e 9
of long-ter m trend analy si s limits o u r ability to fully as s e s s the impact of AI on la n guag e 10
usa g e. P articularly, sin ce the d ata for 2024 i s limited to A pr il, it c ann ot be denied that the 11
Results
may fluctuate when looking at the whole year. 12
13
Conclusion
14
This study hi g h l ight s the overuse of specific words a n d phr a s es t hat ha v e become mor e 15
prevalent sinc e the intr oduction of ChatG PT. Th e li st o f selected terms discu s s ed in t his 16
study can be advant a geou s for bo th users employing LLMs for wr it ing p urposes and for 17
individuals in educationa l and supervisory capa c it i es within th e f i el d s of medic ine an d 18
biology. H owev er , the c h a n g es in a cademic writ ing su ggested b y t his study ma y be 19
temp orary and spe ci f ic to 2024; a s LL Ms impr ov e, dis t ingui s hing betwe en h uman and AI-20
generat ed tex t ma y b ec ome more c hallenging. 21 Thus, fur ther stu dies on future changes in 21
academic ter minology a r e w ar rant ed. Many r es ear c her s are expec t ed to c o ntinue usi n g 22
LL Ms f or the ir wr iting—needles s to s ay, adhering to et h ica l aspect s a n d taking respon s ibility 23
for the final o utp ut is cru cial for t he a uthor s when using the se tools. 24
25
. CC-BY 4.0 International licenseIt is made available under a
is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)
The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint
13 / 16
Funding: This w or k w a s s u pport ed by the J apan Soc iet y f or t he Pr omo tio n of Sc ien ce (JSP S) 1
KAKE NHI ( gr an t number 22K15778). 2
3
Author Contributions: KM c on c eiv ed t he s tudy , dev eloped th e method ol ogy , d esigned and 4
perf ormed the expe r ime n ts, c ollect e d and analyz ed t he da t a, cur a t ed th e d at a, wr o t e th e 5
original dr aft, r ev iewed and edited t h e manusc r ipt, cr ea t ed t he vis ualiz a tions, manag ed the 6
pr oject, and acquir e d t he funding. KM i s the sole co n tribu t or t o t h is wor k . 7
8
Acknowledgem ent: 9
Du ring the p r epar a ti o n o f t h is w or k, the aut ho r u s ed GP T-4, G P T-4o and Claude 3 Opus f or 10
dr af ting R c ode, pr oofr ead ing t he manuscript , and impr ovi n g th e readab i lit y of th e t ext. 11
Aft er using these s ervic e s, the aut hor r ev iew ed and edi t ed the c ont e n t as n eeded and t ak es 12
full responsibility f or th e c on t en t of t h e p ublica ti on . 13
14
Conflict of inter est stat ement: I decla r e no c ompet ing i n t e r es ts. 15
16
Dat a sharing: Da t a will b e s har ed upon r easonable r eques t to t he co r r esponding auth or . 17
18
. CC-BY 4.0 International licenseIt is made available under a
is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)
The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint
14 / 16
Refer ences 1
1. Biswas S. ChatGPT and the Future of Medical Writing. Radiology. 2023; 307 (2): e223312. 2
2. Salvagno M, Taccone FS, Gerli AG. Can artificial intelligence help for scientific writing? Crit Care. 3
2023; 27 (1): 75. 4
3. Lin Z. Towards an AI policy framework in scholarly publishing. Trends Cogn Sci. 2024; 28 (2): 85-88. 5
4. Thorp HH. ChatGPT is fun, but not an author. Science (New York, NY). 2023; 379 (6630): 313. 6
5. Lin Z. Techniques for supercharging academic writing with generative AI. Nat Biomed Eng. 2024. 7
6. Berdejo-Espinola V , Amano T. AI tools can improve equity in science. Science. 2023; 379 (6636): 991. 8
7. Hwang SI, Lim JS, Lee RW , et al. Is ChatGPT a "Fire of Prometheus" for Non-Native English-9
Speaking Researchers in Academic Writing? Korean J Radiol. 2023; 24 (10): 952-959. 10
8. Matsui K, Koda M, Y oshida K. Implications of Nonhuman "Authors". Jama. 2023; 330 (6): 566. 11
9. AlAfnan MA, MohdZuki SF. Do artificial intelligence chatbots have a writing style? An investigation 12
into the stylistic features of ChatGPT-4. Journal of Artificial intelligence and technology. 2023; 3 (3): 85-94. 13
10. Liao W, Liu Z, Dai H , et al. Differentiating ChatGPT-Generated and Human-Written Medical Texts: 14
Quantitative Study. JMIR Med Educ. 2023; 9: e48904. 15
11. Geng M, Trotta R. Is ChatGPT Transforming Academics' Writing Style? arXiv preprint 16
arXiv:240408627. 2024. 17
12. Gray A. ChatGPT" contamination": estimating the prevalence of LLMs in the scholarly literature. 18
arXiv preprint arXiv:240316887. 2024. 19
13. Liang W, Zhang Y , Wu Z , et al. Mapping the increasing use of llms in scientific papers. arXiv preprint 20
arXiv:240401268. 2024. 21
14. Liang W, Izzo Z, Zhang Y , et al. Monitoring ai-modified content at scale: A case study on the impact 22
of chatgpt on ai conference peer reviews. arXiv preprint arXiv:240307183. 2024. 23
15. Song C, Song Y . Enhancing academic writing skills and motivation: assessing the efficacy of 24
ChatGPT in AI-assisted language learning for EFL students. Front Psychol. 2023; 14: 1260843. 25
16. Jeyaraman M, Ramasubramanian S, Balaji S , et al. ChatGPT in action: Harnessing artificial 26
intelligence potential and addressing ethical challenges in medicine, education, and scientific research. Wor ld J 27
Methodol. 2023; 13 (4): 170-178. 28
17. Cheng H, Sheng B, Lee A , et al. Have AI-Generated Texts from LLM Infiltrated the Realm of 29
Scientific Writing? A Large-Scale Analysis of Preprint Platforms. bioRxiv. 2024: 2024-2003. 30
18. Durrant P . Investigating the viability of a collocation list for students of English for academic 31
purposes. English for Specific Purposes. 2009; 28 (3): 157-169. 32
19. Crosby T. How to detect and handle outliers. In: Taylor & Francis; 1994. 33
20. Park M, Leahey E, Funk RJ. Papers and patents are becoming less disruptive over time. Nature. 2023; 34
613 (7942): 138-144. 35
21. Herbold S, Hautli-Janisz A, Heuer U , et al. A large-scale comparison of human-written versus 36
ChatGPT-generated essays. Scientific reports. 2023; 13 (1): 18617. 37
38
. CC-BY 4.0 International licenseIt is made available under a
is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)
The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint
15 / 16
1
. CC-BY 4.0 International licenseIt is made available under a
is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)
The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint
16 / 16
Figure Leg ends 1
Figure 1. Sca t ter plot o f w o r d/phrase u s age f r eque ncy v s . modi f i ed Z- S c or e in 20 24 . 2
Figur e 1 illust rates the r elationship betw e en th e fr equenc y o f us e a nd t he mod i fied Z-sco r es fo r 3
wor ds and phra ses w ith abs ol ut e mod i fi ed Z- sco r es exc eeding 3.5 in 2 0 24. R ed cir c les r ep r esent 4
potentiall y AI- influe nced t erms, whil e g r ey ci r c les r epr es ent c ommon academi c phrases (c ont r ol s). 5
The x - a xis s how s the number o f total r e c or d s usi ng t he wor ds/phra s e s on a l og arit hmi c scale , and 6
the y-ax is dis play s the modified Z- sco r e for us age f r equenc y . 7
8
Figure 2. Mea n u s ag e (modif i ed Z- s c o r es) of pot e n tia lly A I-inf luenced t er m s and c ommon 9
ac a d emic phr a s es fr om 2 000 t o 2024. Shaded ar eas r epr esen t 95% c onfidence in t er v als . 10
11
12
. CC-BY 4.0 International licenseIt is made available under a
is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)
The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint
delve
underscore
commendable meticulous
realm
intricatetapestry
symphony
ingeniously captivateprowess groundbreaking
cutting-edge
compel pivotal
notable/leverage
unveilbolster
elevate
mitigate crucial
Number of total records using the words/phrases (log scale)
holistic
utilize
these
findings
testament
unwavering
was measuredpercentage of
unlock
pioneertransformativesafeguard
nuance imperativegame-changer
embark/enigma
unleash
endeavor
occurrence ofrevolutionize foster
intersection
streamline
demystify
ever-evolving navigate/harness
multifaceted
pave
strategicallyseamless
align/shed light
number ofwith respect/respect to
aptly
unparalleled
to determine
comprehensivehave
shown
interaction
between
innovative
synergyburgeon
poised
encounter
ultimately
showcase
impressively
digital world additionallyenhance
essential
noteworthy
dive
invaluable
journey
vital
characterized
by
Modified z-score for usage frequency
valuable employ
effectively
. CC-BY 4.0 International licenseIt is made available under a
is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)
The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint
Year
Potentially AI-influenced terms
Common academic phreases
(control)
2000 2005 2010 2015 2020 2025
Modified z-score for usage frequency
9
8
7
6
5
4
3
2
1
0
-1
-2
. CC-BY 4.0 International licenseIt is made available under a
is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)
The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.