FPM: A Collection of Large-scale Foundation Pre-trained Language Models | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article FPM: A Collection of Large-scale Foundation Pre-trained Language Models Dezhou Shen This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-1061146/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Recent work in language modeling has shown that train- ing large-scale Transformer models has promoted the lat- est developments in natural language processing applica- tions. However, there is very little work to unify the cur- rent effective models. In this work, we use the current ef- fective model structure to launch a model set through the current most mainstream technology. We think this will become the basic model in the future. For Chinese, us- ing the GPT-2[9] model, a 10.3 billion parameter language model was trained on the Chinese dataset, and, in particu- lar, a 2.9 billion parameter language model based on dia- logue data was trained; the BERT model was trained on the Chinese dataset with 495 million parameters; the Trans- former model has trained a language model with 5.6 bil- lion parameters on the Chinese dataset. In English, cor- responding training work has also been done. Using the GPT-2 model, a language model with 6.4 billion param- eters was trained on the English dataset; the BERT[3] model trained a language model with 1.24 billion param- eters on the English dataset, and in particular, it trained a 688 million parameter based on single card training tech- nology Language model; Transformer model trained a lan- guage model with 5.6 billion parameters on the English dataset. In the TNEWS classification task evaluated by CLUE[13], the BERT-C model exceeded the 59.46% accu- racy of ALBERT-xxlarge with an accuracy rate of 59.99%, an increase of 0.53%. In the QQP classification task evalu- ated by GLUE[11], the accuracy rate of 78.95% surpassed the accuracy rate of BERT-Large of 72.1%, an increase of 6.85%. Compared with the current accuracy rate of ERNIE, the first place in the GLUE evaluation of 75.2%, an increase of 3.75%. Computational Mathematics language modelling large scale modelling language processing Full Text Additional Declarations There is NO Competing Interest. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-1061146","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":62044650,"identity":"c6c17053-0531-46eb-be7f-94b32ae881de","order_by":0,"name":"Dezhou Shen","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAxElEQVRIiWNgGAWjYDACCSDmYWBIYGBgPgAWYCNBC1sChCJBC48BmCII5Gc3H3vwpqIuj18i59uHnz8Y5PkIaTG4cyzdcM4ZtmLJGbmbZ/YkMBi2EdQikWMmzdvGk7jhRu5mZqDDGAlqkZ+R/02a958EUEvOY5AWe4JaGG7ksEnzNhiAtDCDtCQSdtiNNDPJOccSEmf2PDNm7EmTSCbCYcnPJN7U1CX2syc/ZvhhY2M7v4Ggy1CBBInqR8EoGAWjYBRgBQD2SDomN7Bw6wAAAABJRU5ErkJggg==","orcid":"https://orcid.org/0000-0001-5514-507X","institution":"rct ai","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Dezhou","middleName":"","lastName":"Shen","suffix":""}],"badges":[],"createdAt":"2021-11-08 15:31:32","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-1061146/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-1061146/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":15328746,"identity":"e3101846-c247-4c26-86c1-12c5d94858e6","added_by":"auto","created_at":"2021-11-08 18:33:16","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":156207,"visible":true,"origin":"","legend":"","description":"","filename":"rctfounderpm.pdf","url":"https://assets-eu.researchsquare.com/files/rs-1061146/v1_covered.pdf"}],"financialInterests":"There is \u003cb\u003eNO\u003c/b\u003e Competing Interest.","formattedTitle":"FPM: A Collection of Large-scale Foundation Pre-trained Language Models","fulltext":[{"header":"Full Text","content":"This preprint is available for \u003ca href='/article/rs-1061146/latest.pdf' target='_blank'\u003edownload as a PDF\u003c/a\u003e."}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"language modelling, large scale modelling, language processing","lastPublishedDoi":"10.21203/rs.3.rs-1061146/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-1061146/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"Recent work in language modeling has shown that train-\r\ning large-scale Transformer models has promoted the lat-\r\nest developments in natural language processing applica-\r\ntions. However, there is very little work to unify the cur-\r\nrent effective models. In this work, we use the current ef-\r\nfective model structure to launch a model set through the\r\ncurrent most mainstream technology.\r\nWe think this will\r\nbecome the basic model in the future.\r\nFor Chinese, us-\r\ning the GPT-2[9] model, a 10.3 billion parameter language\r\nmodel was trained on the Chinese dataset, and, in particu-\r\nlar, a 2.9 billion parameter language model based on dia-\r\nlogue data was trained; the BERT model was trained on the\r\nChinese dataset with 495 million parameters; the Trans-\r\nformer model has trained a language model with 5.6 bil-\r\nlion parameters on the Chinese dataset. In English, cor-\r\nresponding training work has also been done. Using the\r\nGPT-2 model, a language model with 6.4 billion param-\r\neters was trained on the English dataset; the BERT[3]\r\nmodel trained a language model with 1.24 billion param-\r\neters on the English dataset, and in particular, it trained a\r\n688 million parameter based on single card training tech-\r\nnology Language model; Transformer model trained a lan-\r\nguage model with 5.6 billion parameters on the English\r\ndataset.\r\nIn the TNEWS classification task evaluated by\r\nCLUE[13], the BERT-C model exceeded the 59.46% accu-\r\nracy of ALBERT-xxlarge with an accuracy rate of 59.99%,\r\nan increase of 0.53%. In the QQP classification task evalu-\r\nated by GLUE[11], the accuracy rate of 78.95% surpassed\r\nthe accuracy rate of BERT-Large of 72.1%, an increase of\r\n6.85%. Compared with the current accuracy rate of ERNIE,\r\nthe first place in the GLUE evaluation of 75.2%, an increase\r\nof 3.75%.","manuscriptTitle":"FPM: A Collection of Large-scale Foundation Pre-trained Language Models","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2021-11-08 18:33:11","doi":"10.21203/rs.3.rs-1061146/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"e7ab08fb-b983-4b82-9067-72b2210ee8da","owner":[],"postedDate":"November 8th, 2021","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":8380574,"name":"Computational Mathematics"}],"tags":[],"updatedAt":"2021-11-12T20:40:56+00:00","versionOfRecord":[],"versionCreatedAt":"2021-11-08 18:33:11","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-1061146","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-1061146","identity":"rs-1061146","version":["v1"]},"buildId":"-HB7Z8yhvgn0wM9Nzuekk","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.