Delving into PubMed records: some terms in medical writing have drastically changed after the arrival of ChatGPT

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Introduction It is estimated that large language models (LLMs) including ChatGPT is already widely used in academic paper writing. This study aims to investigate whether the usage of specific terminologies has increased, focusing on words and phrases frequently reported as overused by ChatGPT. Methods A list of 142 potentially AI-influenced terms was curated from online discussions and recent literature documenting LLM vocabulary patterns, while 84 common academic terms in the medical field were used as controls. PubMed records from 2000 to 2024 were analyzed to track the frequency of these terms. Usage trends were normalized using a modified Z-score transformation. Results Among the potentially AI-influenced terms, 100 displayed a meaningful increase (modified Z-score ≥ 3.5) in usage in 2024. The linear mixed-effects model showed a significant effect of potentially AI-influenced terms on usage frequency compared to common academic phrases (p < 0.001); the usage of potentially AI-influenced terms showed a noticeable increase starting in 2020. Discussion This study revealed that certain words, such as “delve,” “underscore,” “meticulous,” “boast,” and “commendable,” have been used more frequently in medical and biological fields since the introduction of ChatGPT. The usage of these terms had already been increasing prior to ChatGPT’s release, suggesting that ChatGPT accelerated the popularity of expressions already gaining traction. The identified terms can inform medical educators aiming to enhance awareness of language trends and promote best practices among trainees using LLMs.
Full text 40,729 characters · extracted from oa-pdf · 9 sections · click to expand

Objective

2 This s tudy aims t o in v es tig a t e whet her the u s ag e of specific t erminologies ha s incr ea s ed, 3 f ocu sing o n wor ds and p hr a s e s fr equ en tly r epor t ed as ov erused by Cha tG P T . 4 Mat erials and Methods: 5 The lis t of 117 p o t entially AI - influen ced t er ms w as c ur a t ed based on pos ts and c ommen t s 6 fr o m anon ymous Cha t GPT u ser s , and 75 common ac ademic phr a se s w er e us ed a s c on tr ols . 7 PubMed r ecor ds fr o m 200 0 t o 2024 (u n til April) we r e analy zed t o tr ac k t he fr equency of 8 these t erms . Us ag e tr en ds w er e nor mali z ed usi n g a modified Z-sc or e tr an sf or ma tion. A 9 linear m ix ed- ef f ects model w as use d t o compa r e t he usag e of p o t enti ally AI-influe n c ed 10 t er ms t o common academic phr as e s ov e r ti me. 11

Results

12 A t ot a l o f 26,403, 493 PubMed r e c or ds w er e in ves tig a t ed. Among t he pot en t ially AI-13 influenced t erms, 74 di s p l a y ed a me aningful inc r eas e (modified Z-s c or e ≥ 3.5) in u s ag e in 14 2024. The linear mi x ed- e f f ec t s mo del s how ed a si gnifican t ef f ect of pot ent ially AI-influen c ed 15 t er ms on u s ag e fr eq uenc y c ompar ed t o c ommon ac ademic phr ases (p < 0. 001). The u s ag e of 16 pot en tiall y AI-influenced t e rms show ed a noticeable incr ea se s t arting in 2020. 17 Discu ssion: 18 This s tudy r evealed tha t cert ain wo r ds and phr a ses, s uch as "delv e," "under s cor e," 19 "meticulous," and "commendable," ha ve been used mor e f requen tly in medic al and 20 biologic a l f ields s ince t he i n tr oductio n of Ch at G PT . The us a g e r a t e of the se w or ds /phr a se s 21 has been i n c r eas ing f or sev er al y ear s be f or e the r e lea se of Chat G PT , s u g g e s ting that C hat G PT 22 migh t ha v e accel er a t ed th e popula ri ty of s cien tific e x pr es sions t ha t w er e alr eady g a ining 23 tr ac t i o n. 24

Conclusions

25 The iden t ified t erms in this stud y c an pr ovide v aluab le ins igh t s f or both LL M user s, edu ca t or s , 26 . CC-BY 4.0 International licenseIt is made available under a is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review) The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint 3 / 16 and super visor s in t hes e fields. 1 2

Keywords

3 ChatGPT , large language models, academic writing, scientific terminology, PubMed, AI-influenced terms, 4 medical literature, language evolution. 5 6 . CC-BY 4.0 International licenseIt is made available under a is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review) The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint 4 / 16

Introduction

1 Cha tG PT r apidly a chieved widespr ead gl o bal use aft er it s launch on No vember 30, 2022. 2 T r ained o n a vas t cor pus of t e x t d a t a, t he la r g e l angu a g e model ( L LM) i ncluding C ha tGPT 3 g ener a t es na tur al languag e wit h r e mark able fluency . Shor tly aft er i ts r eleas e, Cha tG PT ' s 4 applic a b i lit y f or s cien tific writ ing in medic al and b i o l ogi c al fields bec ame eviden t. 1,2 D u e t o 5 the f erv or s u rr oun ding i t s c apabiliti es, it w as c r edit ed as an aut hor o n sever al pap e r s , 6 igniting consider able debat e (curr ently , AI is not acknowledged as an author in sc holar ly 7 public a tions 3 ). Ther e w e r e ev en o p i nions t ha t the us e of Cha t G P T in paper writing w as 8 plagiarism 4 , but in r eality , L LMs su c h as Cha tG PT , Gemini, and Claude ar e a l r eady being u s ed 9 in paper wr i t ing. The us e of LLMs can be applied in v arious wa ys in academi c writ ing 1,5 and is 10 als o import a n t f or t h e resear ch activ ities of non-native r es ear cher s whos e fir st languag e is 11 not Engli sh. 6-8 Pr esen tly , a fr amew ork has been est abli shed t ha t per m its the u s e of L LMs in 12 writing , pr ovided their in v o lv e men t is adequa t e ly acknowledg ed. 3 13 Whil e LLMs c an pr odu ce na tur al w r iti ng , th e ir outpu t als o e xhibits cert ain char a ct er istics . 9,10 14 R ecen tly , it became a t op i c of discu ssion on X ( f ormer ly T wi t t er) and R e ddit t hat C ha tG P T 15 fr eq uently outp uts the w o r d 'delve' 16 (h t t ps:/ /ww w .r edd i t .co m/ r / mildlyinf ur ia ti ng / commen ts/ 1b z v gq j / appar en t ly _u s in g_the_w or17 d_delv e _is _a_sign_of_the/ [Acc e s s e d 2024, A pril 12] ). In add ition, r ecen t r eport s f ocus ing 18 on det ec t ing t e xt g ener a t ed by LL M s ha v e iden ti f i ed s ev er al f r equently us ed w or d s , su ch a s 19 ‘ co mmendable, ’ ‘meticulous , ’ ‘in tric a t e, ’ and ‘r ealm. ’ 11-14 The extr ac t ion o f t hes e 20 char a ct eris tic k eyw or d s of L LMs in t hes e pr ev ious r epor ts w a s per f ormed b y c omparing 21 human- gen er at ed t e xt with Cha t G P T-g ener at ed t e x t. 11,13,14 While t his a ppr oac h r evealed 22 Cha tG PT' s char a ct er istics among the w or d s c ommonly used by both h uma ns an d Cha tG PT , it 23 had methodologic al li mit at ions in e x tr ac t ing w or d s wi t h low u s ag e fr equencies. Clar i f y ing 24 the w or d e xpr essions tha t L LMs t en d t o use in medical and bio logical p aper s is cru cial f or 25 desig n i n g ac ademic writ i n g suppor t and medical educ a t ion p r ogr ams. 15 Mor eove r , r ev ea l ing 26 . CC-BY 4.0 International licenseIt is made available under a is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review) The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint 5 / 16 the e x t ent o f Cha t G PT' s impact on paper s in t he medic a l and bio logical fi el ds is es sen tial f or 1 main t a ining the f airness and r eli abili t y of academi c r e se ar c h and fr om the per spectiv e of 2 r esear ch ethi cs. 16 Ho wev er , the e xi s ting li t er a tur e l ack s a t hor ough in v est iga t ion of the 3 spe c ific w a y s in which Cha tGP T has tr ansf or med ac ademic writing pr actice s in the med ic a l 4 and biologic al discipli n es , ne cessit a ting fur ther r es ear c h . 5 As the use fu lness of L LMs bec omes mor e ev ident, t he numbe r o f r esear c her s usi n g LLMs f or 6 writing paper s has been gr adually in c r eas in g. 11,13,17 It woul d logically f o l l o w tha t th er e has 7 been an incr ease in the number o f r esear ch r epor ts f ea turing specific e x pr essions unique t o 8 LL Ms . This stu dy , ther e f or e, t es ts the h ypoth es i s t ha t the adopt ion of cert ain s c ien t if ic 9 t er m inologies has risen f ollow ing th e adv e n t o f Cha t G PT . F o cu sing on w or ds and phr ase s 10 fr eq uently r epo r t ed as used b y Cha tG PT , I i n ves tig a t ed PubMed r ec or d s fr om 2000 o nwar ds 11 and p e r f ormed a c ompar is on u s ing phr ase s c ommonly u s ed in ac ademia as a c on tr ol. This 12 a n a l y s i s a i m s t o e m p i r i c a l l y e x p l o r e t h e i n f l u e n c e o f L L M s o n t h e l e x i c o n o f m e d i c a l 13 lit era t ur e. 14 15

Methods

16 2 Met hods 17 2.1 Sear ch f or R ec or d s 18 Unli k e earlier s tud i es, 12-14 th i s r esear ch, dr awing in s igh t s fr om v arious a n on ymous en d-user s , 19 e xtr act ed po ten t ially AI- influen c ed t erms fr o m R eddi t, X (f orme rly T wi t t e r) , blogs , and 20 f or ums , f o c u sing on w or d s and phr a ses fr equ en tly pr oduced b y LL M s . Th e selection of t hes e 21 t er ms w as c arried ou t thr ough a r igor ous manual cur a tion pr oces s fr o m A pr i l 1 2 t o May 11, 22 2024, iden t ifyi n g 117 pot en t ia l ly A I-i nfluenced t er ms. In addition, a s a c o n tr ol gr oup, I us ed 23 the t op 100 c olloc a t ions iden t i fied a s char a ct er i s tic o f t he ac ademic c orpus i n a pr evious 24 st u dy . 18 Phr a s es tha t c ould be sear ch ed on PubMed a s tw o c on se c u tiv e w or ds w er e included 25 ( f o r e x a m p l e , t h e c o l l o c a t i o n " b e t w e e n a n d " i s u s e d i n t h e f o r m o f " b e t w e e n A a n d B , " s o i t 26 . CC-BY 4.0 International licenseIt is made available under a is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review) The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint 6 / 16 w as e x cluded a s no r ec or d s w er e f ound when sear chin g f or " b etw een and [T e xt W or d]"). In 1 the en d, 75 c ommon ac ademic phr as es w er e c h os en f or v erific a t ion in thi s s tud y . The lis t of 2 these phr ase s appear s in T able 1. 3 Tab le 1. Wor ds and ph r a s e s ex amined fo r usag e rat e s Potentially AI-influenced terms Verb s ali gn, bols te r, burge on, captiv a te, c a talyz e, c ompel, delve , demys ti fy, di ve, e leva t e, eluc ida te, emba r k , embr ac e, empl oy, enc ount er, e nde avor, e nha nce , enrich, explore, facil ita t e , fo s t er, harne s s , leve r a ge, mi tig ate , nav igate , nu anc e, optimiz e, pa ve , pi one er, r e sona te, revolu t i oni z e , s a f e gua rd, shed l igh t, sho wca s e , streaml in e, und e rscor e, unle as h, unloc k, unve il, utili z e Adjec tive s ae sth etic, comme nda ble, compr ehen siv e , cruc ial, c utti ng-edg e , di sr up tiv e, dy namic , es sen t i al , ev er - e volv ing, f re sh, gro undbr eak ing, holi st i c, il lu striou s, imp era tiv e, in nov ative , inv aluabl e, in trica te, kee n, meticul ou s , methodica l , multi fac e ted, no ta bl e, no tew o r thy , ove r a r c hing, pa ramoun t, piv otal , poi sed, poten t, rob ust, se aml e ss, tail o r ed , tran s f ormative , unparal l eled, unp rece de nte d, unwa ve rin g, valua ble, v er s a til e, vi br a n t , vital Adve r b s addi t i on ally , aptly, e f fective ly, e xce ll entl y, impressiv ely, ingeni ou sly , moreove r, r epo r t edly , schola r ly , s t ra teg ical ly, ul t i mately, und ou btedly Noun s adv enture , be ac on, c apabi li ty, c omplex ity , c or ne rsto ne, di git al w or l d, driving fo rc e, e nigma, game-chan ge r, hurdl e, i nte r pl ay, i nt e r s e ction , journey, kal e ido s c ope , l and scap e , new er a , paradi gm shi f t , prowe s s , r ea lm, spe arh e ad, sta te -o f- the -ar t, s y mphony, s y n ergy, ta pe stry, t es t am e n t, t r eas u r e t r o v e Common academic phrases (controls) ac cording t o , a re pr e sent ed, a re show n, as s oc ia ted w ith , a s s oc iati on be tween, b a s e d o n, betwe e n group s, ca n be , c hang e s in, ch a r a cte rized by, compa red to , compa r e d with, con s i st ent with , c ontr ol group, c or r e la te d w it h , correl ation b etwe en, curr ent stu dy , data se t, dec r e a se in, de fi ned a s , de fine d by, de te r mined by , dif fer ence s be t w ee n, d ue t o, ef fec t o n, ex pos ur e to , f o ll ow up , have sho wn, hig h er than , in additi on, in co ntra st , inc r e a se in, indi cate that , inte rac tio n betwe en, i s con si st ent, no signi fic ant , numb er o f, ob taine d from, occ ur renc e of, ou r re sul ts, our s tu dy, over time , parti cipan ts we r e , perce nta ge o f, p re sen t st u dy , preva lence o f, p revio u s studi e s, rel ate d t o, r el a tion ship b etwe e n, re sp ect t o, re sul ts obtain e d, s a mple siz e, s how tha t, sign ific ant b e t w een, sig ni fican t di f f e re nce, sig ni fican t l y diff er ent, signi fic antly higher , s ta ti s tic ally s i gnific an t, sub s e t o f, s ugg es t th a t , t he s e fi ndings , the se r e s ul ts, thi s ar ticle , th is p aper , thi s st u dy , to de termin e, to eval uat e, w a s c al c ulated, wa s me a s u red, wa s per fo rmed, w a s u s e d , were c ollec ted, w e r e de termin ed, wer e signific antly, with r e spect I used Pu bMed' s adv anc ed s ear ch f e a tur e (ht tp s :// pu bmed.ncbi.nlm.nih.g ov/adv anced/) t o 4 r ev eal the nu mbe r of r ec or ds in whic h these wor ds wer e us ed b y sear chin g f or " T e x t W or d" . 5 T o ensur e c ompr ehens iv e c ov er ag e o f ve r b f o r ms in Englis h, the sear ch quer y inc lud ed the 6 base f orm, t hir d per son singula r pr esen t, pr es en t participle/ pr ogr ess iv e, past t ense , and p a s t 7 . CC-BY 4.0 International licenseIt is made available under a is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review) The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint 7 / 16 part ic iple. F or nouns, b oth s ingular and plur al f orms wer e inc or p or a t e d. Considering the 1 daily incr eas e in r ec or ds inde x ed in P ubMed, the sear ch c ondition s w er e st and a r dized fr om 2 Januar y 1, 2000 , t o April 30, 2024. T he s ear ch f ormulas f or all w or ds/ phr as e s ar e shown in 3 Supplemen t a r y T able S1. 4 5 2.2 Da t a Pr epar a tion 6 T o in v estig a t e t he usag e tr end s o f po t ent ial ly A I-inf luenced t erms in th e P ubMed da t abas e, 7 w e f i r st c alc u l a t ed t he usag e f r equ e nc y of each term by div iding the n umber of r ecor ds 8 c on t aining th e t e rm by the t ot al number o f r ec or ds in PubMed f or each y ear fr om 2000 t o 9 2024 (up t o Apr il 30, 20 24). Thi s pr o c es s yi elded a da t as et w it h usa g e fr equency f or ea c h 10 t er m and y ear . Ne x t , the modif ied Z-s c or e t r ans f ormat ion w as us ed t o nor malize t he usage 11 fr eq uenc y and f acilit at e c ompar i sons acr o s s t er ms and y ear s. F or ea ch t er m, the median and 12 median absolut e devi a t i on (MAD) wer e c alculat ed. The mo difi ed Z-sc or e w as comput ed b y 13 subtr a cting the median fr o m each occ ur r ence r a t e, div iding th e result by the MAD , and 14 multipl ying b y 0.674 5. T o iden t i f y significan t deviations in t er m u s ag e, we c ons ider ed an 15 absolut e modif i ed Z-scor e of 3.5 o r higher as indic a tive of a meani ngful incr ease or 16 decr ease 19 . Th e r esulting da t aset, c on t aining t h e modi f ied Z-scor es f or ea ch t e r m and year , 17 was t hen us ed f o r fur ther s t a t is tic al analy sis. 18 19 2.3 St a tis tic al A nalysis 20 A l inear mix ed -ef f ec t s model was used to c ompar e t h e usag e of po t en ti all y AI-inf luen c ed 21 t er ms and c ommo n a c ademic phr ases f r om 2000 t o 20 24. The d at a, consis t ing of mod ified Z-22 scor e s f or each wo r d or p hr ase, wer e obt ained and r es haped i n t o a long f orma t. The mod el, 23 . CC-BY 4.0 International licenseIt is made available under a is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review) The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint 8 / 16 c onst ruct ed u s ing t he 'lme' function f r om t he 'nlme' pack age i n R, include d the modi fied Z-1 scor e s a s the dependen t v ariab le, t he g r oup (pot ent ial ly AI- influen c ed ter ms or c ommon 2 ac a d emic phr ase s) a s a fix ed e ff ect, and a r a n dom i n t er cept f or each w or d or p h r as e t o 3 accou n t f or r ep ea t ed measur es. T he model's s ummary w a s g ener a t ed t o asses s the 4 si gn i f ican ce of t he fix ed e ff ect of th e gr oup on t erm u s ag e. A line plot wi t h 95% co nf idenc e 5 in t erva ls w as c r ea t ed using the ' g gplot 2' pac k a g e t o vis u a li z e the tr end s in mean usag e f or 6 each gr oup fr om 2 000 t o 2024. The signific ance level f or a ll s t atis tic a l t est s wa s s et a t 0.05. 7 The analy s i s wa s perf or med using R ver s ion 4.3.2. 8 9

Results

10 A t o t al of 26,403,493 r eco r d s betw een J anuary 1, 20 0 0, and Ap ril 30, 20 24 w er e e xtr act ed 11 fr o m PubMed. The f r eq uenc y r a t es of each w or d/phr ase w er e det e r mined using the annu a l 12 t ot a l numb er of r ecor ds as th e d eno minat o r , f o llow ed by th e calcula t ion o f th e modifi ed Z-13 scor e. The Modif i ed Z-sc or e f or al l the w or ds and phr as e s a cr o ss all per iods is shown in 14 Supplemen t a r y T able S2. 15 In this stud y , amon g the 117 pot ent ia lly AI- inf luenc ed t e rms ver if ied, 74 w or ds /phr a se s 16 (lis t ed in desc ending or der : ‘ delv e, ’ ‘ under s c or e, ’ ‘meticulou s, ’ ‘ commendable, ’ ‘ show ca se, ’ 17 ‘in tr icat e, ’ ‘t apestr y , ’ ‘ s ymphon y , ’ ‘i mpr ess iv ely , ’ ‘r ea lm, ’ ‘ cut t ing-edg e, ’ ‘pr owes s , ’ ‘ c a p tiv at e, ’ 18 ‘not ewor th y , ’ ‘ gr oundbr eaking , ’ ‘unlo c k, ’ ‘ c ompel, ’ ‘lev er ag e, ’ ‘not able, ’ ‘un v eil, ’ ‘i n geniou s ly , ’ 19 ‘piv ot al, ’ ‘bols t e r , ’ ‘holistic, ’ ‘ saf eguar ds, ’ ‘ elev at e, ’ ‘unwa v ering , ’ ‘t r ansf or ma ti v e, ’ ‘p i o nee r , ’ 20 ‘ enigma, ’ ‘ embark, ’ ‘in v aluable, ’ ‘t es tamen t , ’ ‘nuance, ’ ‘mitig a te, ’ ‘ g ame-chang er , ’ ‘v aluable, ’ 21 ‘ endea v or , ’ ‘imper at i ve, ’ ‘ cru cial, ’ ‘r ev oluti oniz e, ’ ‘unleash, ’ ‘ e f f ec t iv ely , ’ ‘ employ , ’ ‘ d igit al 22 world, ’ ‘f ost er , ’ ‘ demy s tified, ’ ‘multif acet ed, ’ ‘na vig a t e, ’ ‘ ev er - evol v ing , ’ ‘ str eamline, ’ 23 ‘in t er section, ’ ‘utiliz e, ’ ‘h a r ness , ’ ‘ s h ed l igh t , ’ ‘ s tr at egi c ally , ’ ‘ seamles s, ’ ‘ en c oun t e r , ’ ‘ essen tial, ’ 24 ‘ align, ’ ‘ addition al ly , ’ ‘pav e, ’ ‘poised, ’ ‘inno va tive, ’ ‘ syner g y , ’ ‘ c ompr ehens iv e, ’ ‘bur g eon, ’ ‘ aptly , ’ 25 . CC-BY 4.0 International licenseIt is made available under a is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review) The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint 9 / 16 ‘ div e, ’ ‘unpar alleled, ’ ‘ulti mat ely , ’ ‘vital, ’ ‘jo urney , ’ ‘ enhan ce’ ) di s pla y ed a modif ied Z - s c or e 1 e x c eeding 3.5 in 2 02 4. Whil e the majority of t he 75 c ommon ac ademic phr ase s (c on tr ols ) 2 display ed no s i gnific an t devi at ions i n us ag e r a t es, phr a s es su ch a s 'oc c urr enc e of ,' 't hes e 3 findings,' 'hav e shown, ' 'in t er action bet ween,' and ' char a ct er i z ed b y' s ur passed a modified Z-4 scor e o f 3.5 in 2024. On the ot her hand, the phr as es 'per cen t ag e o f ,' 'w as measur ed ,' 5 'number o f ,' 'with r espect, ' 'r es p e c t t o ,' and 't o det er min e' r egist er ed modifi ed Z-sco r es 6 below -3.5 in the same y ear (Figur e 1). 7 [Insert Figur e 1 h e r e] 8 The li n ea r mix ed - eff ects model r ev e aled a significan t ef f ec t o f t h e gr ou p (po t en tia lly AI-9 influenced t er ms vs. c ommo n a cad emi c ph r a s es) on the u sag e fr equenc y . The mod e l 10 show ed tha t the usag e of p ot e n tial ly AI- i n flu enced t er ms was significan t l y hig h er th a n tha t 11 of c ommon academic phr ase s (β = 0. 554, S E = 0.080, t(190) = 6.969, p < 0. 001). The line plo t 12 (Figur e 2) illus tr a t es t he tr en ds in mea n fr eq uenc y f or pot en t i a lly A I-infl uenced t erms and 13 c ommon academi c phr a ses fr om 2000 t o 2024. W hile the fr equency of the c on tr ol gr oup 14 r emains r el a ti v ely s t able, the p ot entially AI-influenced terms begin t o s ho w an incr ea s e 15 ar oun d 2016, with a not a b l e and st e ep upw ar d tr ajec t ory s t ar ting in 2 020 tha t bec omes 16 part ic ular ly pr onounced in 2023 and 2024. 17 [Insert Figur e 2 h e r e] 18 19 Discu ssion 20 This s tudy demonst r a t ed tha t, in the fields of medicine and bio logy , a number o f specific 21 w or ds and phr a se s , l ed by "delve," " under sc or e," "meti culous, " a n d "co mmendable," ha ve 22 c ome t o be used mo r e f r equ ently f ol lowing th e adven t of Cha tGPT . Th e i nc r eas in g t r end in 23 the usag e r at es of these wo r ds/ phr as e s w a s mor e pr ono u nc ed in 2024 than in 2023 in 24 almos t all c ases . Thi s ma y r e flect th e g ener ali z a tion of LLM us e among r es ear cher s i n the 25 . CC-BY 4.0 International licenseIt is made available under a is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review) The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint 10 / 16 fields of med ici n e and bio logy , a s sh own in p r evious findings. 13 T h e l i s t o f o v e r u s e d t e r m s 1 su g ge s t ed in th i s st udy wil l help t hos e wr i t ing w it h L LMs cen t er ed ar o und C ha tGPT . 2 It has been obser v ed tha t med ic a l t e x t s g ener a t ed by Cha tGP T , while fluent a n d l o gi c al, t end 3 t o incl u de less s pecific inf orma tion and mor e g ener alized e xpr ess ions c ompar ed t o t hos e 4 autho r ed by huma n s , whi ch f ea tur e a richer and mor e d iver s e c on t en t. 10 In g ener al paper s , 5 it has been no t ed t ha t Cha tGPT t ends t o 1) use the same st y le and e xpressio ns r ep eat ed ly , 2) 6 show a dec r ea s e in t he f r equency of bas ic ver bs lik e ‘is’ and ‘ ar e, ’ and 3) f r eque n tly us e 7 adjectiv es and adv erbs. 11 P articularly f or adjectiv es and adver bs , n umer ous w or ds tha t 8 Cha tG PT fr equen t ly use s ha v e been poin t ed ou t. 14 Bec aus e this st udy only c oun t ed the 9 r ec or ds wher e spec if i c w or ds or phr ases o ccurr ed, it did not ev alu a t e the w eigh t of t e r ms 10 appearing mu l t ipl e tim es. In the cur r en t s tudy , s e v er a l w or ds t ha t w er e pr ev iously iden tifie d 11 a s f r e q u e n t ly u s e d b y C h a t G P T d i d n o t e x h i b i t a n o t a b le in cr e a se i n u s a g e ; y e t , C h a t GP T ma y 12 actually ov eruse these wor d s mor e than s u g g e s t ed by th e r esults o f t his st ud y . Similarly , 13 fr eq uently used ver bs s u ch a s ‘ enh a n c e’ , ‘ eleva t e’ , and ‘ut ili z e’ ma y al s o h av e been o v er us ed 14 by LL M s mor e than sugges t ed by thi s st ud y . 15 A pr evious r epor ts' limi t a tion li es in their lack of f oc u s on the spe c ific w or ds or t er ms 16 ov erused by Cha t GPT , thus f ailing t o c ompr eh ens iv ely e xplor e char act eris tic t er ms . As 17 e xt ensi v ely debat ed onli ne 18 (h t t ps:/ /ww w .r edd i t .co m/ r / mildlyinf ur ia ti ng / commen ts/ 1b z v gq j / appar en t ly _u s in g_the_w or19 d_delv e _is _a_sign_of_the/ [Acce s se d 2024, Apr i l 12]) , the incr eas ed u sag e of the w or d 20 ‘ delv e’ w a s in cr edibly pr onounced c o mpar e d t o other w or ds or ph r as e s, with a modified Z-21 scor e of ar ound 100. D es pit e its ov e rwhelming pr esence, pr evious paper s compar ed t e x t s 22 cr ea t ed b y humans and Cha tG P T 12-14 d i d not mention ‘d e l v e , ’ highligh ting a s tr eng t h of t his 23 s tudy' s methodo l o gy . This s tud y c an not con c lusiv el y est ablish t he c onne ction bet w een the 24 fr eq uent use of ‘ del v e’ and t he e mer g ence of Cha tGPT , alth ough its impact is highly 25 su s pect ed. The fr equen t use o f the t erm ‘ delve’ by Cha tG PT could be a t tribu t ed t o i t s 26 . CC-BY 4.0 International licenseIt is made available under a is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review) The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint 11 / 16 pr ominenc e in th e tr ain ing da t a, p o ssibly r e sulting fr o m c ommo n ins truc t ions dur i n g the 1 r einf o r cemen t learn ing fr o m human f ee dback pha s e, or as a f e a tur e of lar g e la n g u a g e 2 models designed t o pr oj ec t aut hori ty ; how ev er , th e se h y p othese s r emai n s pecula tiv e and 3 unc onfir med. 4 Not ably , th e fr equenc y of u se f or t he p ot en t ially A I-influ enced t erms in ves tig a t ed in t his 5 s tudy had al r eady div er g ed mar k edly even b e f or e ChatGP T r eleased in Nov ember 2022. In 6 part ic ular , 'delve' w as used e x ceptionally oft en in 2024, but i t had also b een used at an 7 e x c ept ionally high fr equency in ac ademic writing sinc e 2021. One h ypothes is ar is es fr om t his 8 obser v a tion―man y of the pot en tia ll y AI-influenced t er ms ma y ha ve c o n tribut ed to their 9 ov eruse in Cha tGP T , as t he s e e x p r essions g ained po pularity du ring t he period of int ensive 10 LL M tr a ining. In other w or ds, Cha tGP T ma y ha v e a c c eler a t ed the i nev it able t empor a l 11 chang e s in writ ing i n r es ear c h . H owever , t his h ypot hes i s would be dif ficult t o verify , since we 12 c annot obser v e a par allel unive r s e wher e C ha tG P T d oe s not exis t. 13 In t er estingly , some of t he co mmon a c ademic phr a se s u sed as co n tr ols also d ev ia t ed i n thei r 14 pr op orti on of use in 2024. The f o ur ph r as e s 'o c curr ence o f', 'these finding s', 'h a v e s h ow n ' , 15 and 'in t er a ct i o n be tw een' sig n i f ican t l y incr e a sed in fr equency of use in 2024, but s ince t he y 16 ar e all v ery co mmonly u s ed e xpr es sions in ac ademic writ i n g , it w ould be dif fic u l t f o r us 17 humans t o r ec ognize tha t their f requency ha s in cr eased. Con v er sely , the six phr a se s 18 'per cen t ag e of', ' wa s mea sur ed', 'nu mber of', 'with r es pe ct', 'r es pe ct t o', and 't o det ermine ' 19 not ably dec r eas ed in usag e in 2024. W h en int er pr e ting these r esults, w e must r e membe r 20 tha t the languag e u s ed in paper s na t ur all y ev olv es ov er t i me; 20 many p hra se s that dec r eas ed 21 in frequency had alr eady b een dec lini ng even be fo re the introduct i o n o f C hatGP T. However, 22 the two phr as es 't o d eter mine' and 'number of ' did no t show a not ic eable decrease in 23 fr equency of us e befo re 2022, and their f r equency of use appear s to have dec r eas ed 24 si gn i f ic ant l y af ter 20 23 (see Supplementar y Ta b l e 2). Whil e t his result ma y be c oincident a l, 25 . CC-BY 4.0 International licenseIt is made available under a is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review) The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint 12 / 16 it could a lso indic at e tha t t he p roli feration o f C hatG P T may subt ly lead to t he decreased us e 1 of cert ai n wor ds or phrases with out us reco gniz ing it. 2 This s t udy ha s some limita t ions. The most important li mi tation is that the t erms pot en tial ly 3 influenced by LLMs in this stu dy were iden tifi ed t hrough manual ins pection r a t her than 4 being extr ac t ed in an objec t iv e and s ystemat ic manner. Therefor e, there may be words or 5 phr a ses that were not i n c luded in t h i s stud y yet have seen a s ignificant i n c r eas e i n u sa ge 6 post-Ch a t GPT. Addition al ly, tempor al shifts in the f requency of word or phrase u se could 7 have been inf l u enc ed by external f a c to rs s uch a s evolving resear ch tr ends and shifts in the 8 style of scientif i c commun i cation, f act ors not ac c o unted for in this s t ud y. Lastly , the ab senc e 9 of long-ter m trend analy si s limits o u r ability to fully as s e s s the impact of AI on la n guag e 10 usa g e. P articularly, sin ce the d ata for 2024 i s limited to A pr il, it c ann ot be denied that the 11

Results

may fluctuate when looking at the whole year. 12 13

Conclusion

14 This study hi g h l ight s the overuse of specific words a n d phr a s es t hat ha v e become mor e 15 prevalent sinc e the intr oduction of ChatG PT. Th e li st o f selected terms discu s s ed in t his 16 study can be advant a geou s for bo th users employing LLMs for wr it ing p urposes and for 17 individuals in educationa l and supervisory capa c it i es within th e f i el d s of medic ine an d 18 biology. H owev er , the c h a n g es in a cademic writ ing su ggested b y t his study ma y be 19 temp orary and spe ci f ic to 2024; a s LL Ms impr ov e, dis t ingui s hing betwe en h uman and AI-20 generat ed tex t ma y b ec ome more c hallenging. 21 Thus, fur ther stu dies on future changes in 21 academic ter minology a r e w ar rant ed. Many r es ear c her s are expec t ed to c o ntinue usi n g 22 LL Ms f or the ir wr iting—needles s to s ay, adhering to et h ica l aspect s a n d taking respon s ibility 23 for the final o utp ut is cru cial for t he a uthor s when using the se tools. 24 25 . CC-BY 4.0 International licenseIt is made available under a is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review) The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint 13 / 16 Funding: This w or k w a s s u pport ed by the J apan Soc iet y f or t he Pr omo tio n of Sc ien ce (JSP S) 1 KAKE NHI ( gr an t number 22K15778). 2 3 Author Contributions: KM c on c eiv ed t he s tudy , dev eloped th e method ol ogy , d esigned and 4 perf ormed the expe r ime n ts, c ollect e d and analyz ed t he da t a, cur a t ed th e d at a, wr o t e th e 5 original dr aft, r ev iewed and edited t h e manusc r ipt, cr ea t ed t he vis ualiz a tions, manag ed the 6 pr oject, and acquir e d t he funding. KM i s the sole co n tribu t or t o t h is wor k . 7 8 Acknowledgem ent: 9 Du ring the p r epar a ti o n o f t h is w or k, the aut ho r u s ed GP T-4, G P T-4o and Claude 3 Opus f or 10 dr af ting R c ode, pr oofr ead ing t he manuscript , and impr ovi n g th e readab i lit y of th e t ext. 11 Aft er using these s ervic e s, the aut hor r ev iew ed and edi t ed the c ont e n t as n eeded and t ak es 12 full responsibility f or th e c on t en t of t h e p ublica ti on . 13 14 Conflict of inter est stat ement: I decla r e no c ompet ing i n t e r es ts. 15 16 Dat a sharing: Da t a will b e s har ed upon r easonable r eques t to t he co r r esponding auth or . 17 18 . CC-BY 4.0 International licenseIt is made available under a is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review) The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint 14 / 16 Refer ences 1 1. Biswas S. ChatGPT and the Future of Medical Writing. Radiology. 2023; 307 (2): e223312. 2 2. Salvagno M, Taccone FS, Gerli AG. Can artificial intelligence help for scientific writing? Crit Care. 3 2023; 27 (1): 75. 4 3. Lin Z. Towards an AI policy framework in scholarly publishing. Trends Cogn Sci. 2024; 28 (2): 85-88. 5 4. Thorp HH. ChatGPT is fun, but not an author. Science (New York, NY). 2023; 379 (6630): 313. 6 5. Lin Z. Techniques for supercharging academic writing with generative AI. Nat Biomed Eng. 2024. 7 6. Berdejo-Espinola V , Amano T. AI tools can improve equity in science. Science. 2023; 379 (6636): 991. 8 7. Hwang SI, Lim JS, Lee RW , et al. Is ChatGPT a "Fire of Prometheus" for Non-Native English-9 Speaking Researchers in Academic Writing? Korean J Radiol. 2023; 24 (10): 952-959. 10 8. Matsui K, Koda M, Y oshida K. Implications of Nonhuman "Authors". Jama. 2023; 330 (6): 566. 11 9. AlAfnan MA, MohdZuki SF. Do artificial intelligence chatbots have a writing style? An investigation 12 into the stylistic features of ChatGPT-4. Journal of Artificial intelligence and technology. 2023; 3 (3): 85-94. 13 10. Liao W, Liu Z, Dai H , et al. Differentiating ChatGPT-Generated and Human-Written Medical Texts: 14 Quantitative Study. JMIR Med Educ. 2023; 9: e48904. 15 11. Geng M, Trotta R. Is ChatGPT Transforming Academics' Writing Style? arXiv preprint 16 arXiv:240408627. 2024. 17 12. Gray A. ChatGPT" contamination": estimating the prevalence of LLMs in the scholarly literature. 18 arXiv preprint arXiv:240316887. 2024. 19 13. Liang W, Zhang Y , Wu Z , et al. Mapping the increasing use of llms in scientific papers. arXiv preprint 20 arXiv:240401268. 2024. 21 14. Liang W, Izzo Z, Zhang Y , et al. Monitoring ai-modified content at scale: A case study on the impact 22 of chatgpt on ai conference peer reviews. arXiv preprint arXiv:240307183. 2024. 23 15. Song C, Song Y . Enhancing academic writing skills and motivation: assessing the efficacy of 24 ChatGPT in AI-assisted language learning for EFL students. Front Psychol. 2023; 14: 1260843. 25 16. Jeyaraman M, Ramasubramanian S, Balaji S , et al. ChatGPT in action: Harnessing artificial 26 intelligence potential and addressing ethical challenges in medicine, education, and scientific research. Wor ld J 27 Methodol. 2023; 13 (4): 170-178. 28 17. Cheng H, Sheng B, Lee A , et al. Have AI-Generated Texts from LLM Infiltrated the Realm of 29 Scientific Writing? A Large-Scale Analysis of Preprint Platforms. bioRxiv. 2024: 2024-2003. 30 18. Durrant P . Investigating the viability of a collocation list for students of English for academic 31 purposes. English for Specific Purposes. 2009; 28 (3): 157-169. 32 19. Crosby T. How to detect and handle outliers. In: Taylor & Francis; 1994. 33 20. Park M, Leahey E, Funk RJ. Papers and patents are becoming less disruptive over time. Nature. 2023; 34 613 (7942): 138-144. 35 21. Herbold S, Hautli-Janisz A, Heuer U , et al. A large-scale comparison of human-written versus 36 ChatGPT-generated essays. Scientific reports. 2023; 13 (1): 18617. 37 38 . CC-BY 4.0 International licenseIt is made available under a is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review) The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint 15 / 16 1 . CC-BY 4.0 International licenseIt is made available under a is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review) The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint 16 / 16 Figure Leg ends 1 Figure 1. Sca t ter plot o f w o r d/phrase u s age f r eque ncy v s . modi f i ed Z- S c or e in 20 24 . 2 Figur e 1 illust rates the r elationship betw e en th e fr equenc y o f us e a nd t he mod i fied Z-sco r es fo r 3 wor ds and phra ses w ith abs ol ut e mod i fi ed Z- sco r es exc eeding 3.5 in 2 0 24. R ed cir c les r ep r esent 4 potentiall y AI- influe nced t erms, whil e g r ey ci r c les r epr es ent c ommon academi c phrases (c ont r ol s). 5 The x - a xis s how s the number o f total r e c or d s usi ng t he wor ds/phra s e s on a l og arit hmi c scale , and 6 the y-ax is dis play s the modified Z- sco r e for us age f r equenc y . 7 8 Figure 2. Mea n u s ag e (modif i ed Z- s c o r es) of pot e n tia lly A I-inf luenced t er m s and c ommon 9 ac a d emic phr a s es fr om 2 000 t o 2024. Shaded ar eas r epr esen t 95% c onfidence in t er v als . 10 11 12 . CC-BY 4.0 International licenseIt is made available under a is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review) The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint delve underscore commendable meticulous realm intricatetapestry symphony ingeniously captivateprowess groundbreaking cutting-edge compel pivotal notable/leverage unveilbolster elevate mitigate crucial Number of total records using the words/phrases (log scale) holistic utilize these findings testament unwavering was measuredpercentage of unlock pioneertransformativesafeguard nuance imperativegame-changer embark/enigma unleash endeavor occurrence ofrevolutionize foster intersection streamline demystify ever-evolving navigate/harness multifaceted pave strategicallyseamless align/shed light number ofwith respect/respect to aptly unparalleled to determine comprehensivehave shown interaction between innovative synergyburgeon poised encounter ultimately showcase impressively digital world additionallyenhance essential noteworthy dive invaluable journey vital characterized by Modified z-score for usage frequency valuable employ effectively . CC-BY 4.0 International licenseIt is made available under a is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review) The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint Year Potentially AI-influenced terms Common academic phreases (control) 2000 2005 2010 2015 2020 2025 Modified z-score for usage frequency 9 8 7 6 5 4 3 2 1 0 -1 -2 . CC-BY 4.0 International licenseIt is made available under a is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review) The copyright holder for this preprint this version posted May 16, 2024. ; https://doi.org/10.1101/2024.05.14.24307373doi: medRxiv preprint

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-pdf

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-30T02:00:01.510937+00:00
License: CC-BY-4.0