Recognition of arthropod species names using bigram-based classification | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research article Recognition of arthropod species names using bigram-based classification Jennien Raffington, Dirk Steinke, Dan Tulpan This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-26532/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background The task of recognizing species names in scientific articles is a quintessential step for a large number of applications in high-throughput text mining and data analytics, such as species-specific information collection, construction of species food networks and trophic relationship extraction. These tasks become even more important in fast-paced species-discovery areas such as entomology, where an impressive number of new arthropod species are discovered each year. This article explores the use of twocharacter n-grams (bigrams) in machine learning models for arthropod species name recognition. This particular method has been previously applied successfully to the task of language identification [1] but the application to species name identification had yet to be explored. Results Arthropod species names, regular English words used in scientific publications and person names were collected from the public domain and bigrams were extracted and used as classifier features. A number of learning classifiers spanning 7 algorithmic categories (tree-based, rule-based, artificial neural network, Bayesian, boosting, lazy and kernel-based) were tested and the highest accuracies were consistently obtained with LIBLINEAR [2], Bayesian Logistic Regression [3], the Multilayer Perceptron [4], Random Forest [5], and the LIBSVM [6] classifiers. When compared with dictionary-based external software tools such as GNRD [7] and TaxonFinder [8], our top-3 classifiers were insensitive to words capitalization and were able to correctly classify novel species names that are absent in dictionary-based approaches with accuracies between 88.6% and 91.6%. Conclusions Our results suggest that character bigram-based classification is a suitable method for distinguishing arthropod species names from regular English words and person names commonly found in scientific literature. Moreover, our method can also be used to reduce the number of false positives produced by dictionary-based methods. Bioinformatics machine learning classification species names arthropod bigram Figures Figure 1 Figure 2 Figure 3 Full Text Additional files Additional file 1 – ZIP archive including all datasets used in this work. Each file is labelled by applying the naming convention used in the manuscript. All datasets are also made publicly available at: http://animalbiosciences.uoguelph.ca/~dtulpan/papers/specrec2020 Supplementary Files Additionalfile1.zip Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-26532","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research article","associatedPublications":[],"authors":[{"id":551731,"identity":"fecf5a34-d6aa-43e9-b36a-13dec7119c2c","order_by":1,"name":"Jennien Raffington","email":"","orcid":"","institution":"University of Guelph Biodiversity Institute of Ontario","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Jennien","middleName":"","lastName":"Raffington","suffix":""},{"id":551732,"identity":"04c45f5b-c0ed-4028-a37c-86c7f1a3c9da","order_by":2,"name":"Dirk Steinke","email":"","orcid":"","institution":"University of Guelph Biodiversity Institute of Ontario","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Dirk","middleName":"","lastName":"Steinke","suffix":""},{"id":551733,"identity":"91eb15cc-09a4-4152-93b8-74e7baa0fe75","order_by":3,"name":"Dan Tulpan","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA4UlEQVRIie3RMQuCQBTA8RdBLpdzEtRXeNHUt8lFl4qDIBoanHKRWutb2DcQHpyLczgIGYFzUzRElDpEy2lbw/2Xdwf34w4OQKX6xwJoBO/RK3ccxtVE5ApgWO7wB2I6tYl+8pD4KrH90D12rpjMQKNUSgzBkHYim/pRtDB2mM2BWSglmBPWoqkfT6zue206HahDnmRjTh4F0a7VpL2mMca26EJBmPwWQ1ic9hsa7KOoOfIwM9dswqVEJzpc+I36euie4/syMbda6EtJXrMcxXsCaFWe/xAthfJbVSqVSvXdC0bEUcezx8gzAAAAAElFTkSuQmCC","orcid":"https://orcid.org/0000-0003-1100-646X","institution":"University of Guelph","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Dan","middleName":"","lastName":"Tulpan","suffix":""}],"badges":[],"createdAt":"2020-05-03 12:48:59","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-26532/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-26532/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":1074089,"identity":"94779b18-9847-4c55-ba25-b6b915ad8c60","added_by":"auto","created_at":"2020-05-11 23:38:14","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":171619,"visible":true,"origin":"","legend":"Training and testing classification accuracies for all 15 classifiers applied on problems P1, P2 and P3.\nThe figure depicts a bar plot of prediction accuracy for all the machine learning models applied in this study on the 3 classification problems. Note: the Bayesian Logistic Regression method could not be applied on non-binary classification problems such as P3.","description":"","filename":"Figure1.png","url":"https://assets-eu.researchsquare.com/files/rs-26532/v1/Figure1.png"},{"id":1074092,"identity":"9ca2ad45-14c0-49a7-8b7b-40f8a349feef","added_by":"auto","created_at":"2020-05-11 23:38:15","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":94946,"visible":true,"origin":"","legend":"Venn diagram representation of top 100 high frequency bigrams for SCI, ENG and PEO instances from the training set.\nThe 3 set intersection diagram depicts the overlapping and unique bigrams commonly appearing in top 100 high frequency bigrams occurring in the datasets including arthropod species names, person names and English words.","description":"","filename":"Figure2.png","url":"https://assets-eu.researchsquare.com/files/rs-26532/v1/Figure2.png"},{"id":1074093,"identity":"03f82eb2-9db6-4642-9cda-c01f8f7f26f0","added_by":"auto","created_at":"2020-05-11 23:38:15","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":163760,"visible":true,"origin":"","legend":"Average runtimes for all 15 classifier methods applied to the 3 classification problems: P1, P2 and P3.\nThe figure depicts the average runtimes for the classifiers used in this study. Each classifier was executed three times on each problem and the results were averaged out and reported in the table included at the bottom of the figure.","description":"","filename":"Figure3.png","url":"https://assets-eu.researchsquare.com/files/rs-26532/v1/Figure3.png"},{"id":1074094,"identity":"7d8b5bd6-bedb-474e-99b4-9cb968bdfb90","added_by":"auto","created_at":"2020-05-11 23:38:20","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":234296,"visible":true,"origin":"","legend":"","description":"","filename":"RaffingtonEtAl2020v5.pdf","url":"https://assets-eu.researchsquare.com/files/rs-26532/v1_stamped.pdf"},{"id":1074091,"identity":"f797c05b-f16e-4d9e-a65f-138865036d66","added_by":"auto","created_at":"2020-05-11 23:38:15","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":205655,"visible":true,"origin":"","legend":"","description":"","filename":"RaffingtonEtAl2020v5.pdf","url":"https://assets-eu.researchsquare.com/files/rs-26532/v1/RaffingtonEtAl2020v5.pdf"},{"id":13502883,"identity":"753760f0-e642-4a3b-bb9f-836a658e06c2","added_by":"auto","created_at":"2021-09-16 23:16:33","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":511947,"visible":true,"origin":"","legend":"","description":"","filename":"RaffingtonEtAl2020v5.pdf","url":"https://assets-eu.researchsquare.com/files/rs-26532/v1_covered.pdf"},{"id":1074090,"identity":"6d0d085e-cbe9-4a77-a4d6-b31de310c172","added_by":"auto","created_at":"2020-05-11 23:38:14","extension":"zip","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":88005,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile1.zip","url":"https://assets-eu.researchsquare.com/files/rs-26532/v1/Additionalfile1.zip"}],"financialInterests":"","formattedTitle":"Recognition of arthropod species names using bigram-based classification","fulltext":[{"header":"Full Text","content":"\u003cp\u003eThis preprint is available for \u003ca href='/article/rs-26532/latest.pdf' target='_blank'\u003edownload as a PDF\u003c/a\u003e.\u003c/p\u003e"},{"header":"Additional files ","content":"\u003cp\u003e\u003cstrong\u003eAdditional file 1 \u0026ndash; ZIP archive including all datasets used in this work. Each file is labelled by applying the naming convention used in the manuscript. \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eAll datasets are also made publicly available at: \u003c/em\u003e\u003c/p\u003e\n\u003cp\u003e\u003cem\u003e\u003cu\u003ehttp://animalbiosciences.uoguelph.ca/~dtulpan/papers/specrec2020\u003c/u\u003e\u003c/em\u003e\u0026nbsp;\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"machine learning, classification, species names, arthropod, bigram ","lastPublishedDoi":"10.21203/rs.3.rs-26532/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-26532/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eBackground\u003c/p\u003e\u003cp\u003eThe task of recognizing species names in scientific articles is a quintessential step for a large number of applications in high-throughput text mining and data analytics, such as species-specific information collection, construction of species food networks and trophic relationship extraction. These tasks become even more important in fast-paced species-discovery areas such as entomology, where an impressive number of new arthropod species are discovered each year. This article explores the use of twocharacter n-grams (bigrams) in machine learning models for arthropod species name recognition. This particular method has been previously applied successfully to the task of language identification [1] but the application to species name identification had yet to be explored.\u003c/p\u003e\u003cp\u003eResults\u003c/p\u003e\u003cp\u003eArthropod species names, regular English words used in scientific publications and person names were collected from the public domain and bigrams were extracted and used as classifier features. A number of learning classifiers spanning 7 algorithmic categories (tree-based, rule-based, artificial neural network, Bayesian, boosting, lazy and kernel-based) were tested and the highest accuracies were consistently obtained with LIBLINEAR [2], Bayesian Logistic Regression [3], the Multilayer Perceptron [4], Random Forest [5], and the \u003cem\u003eLIBSVM \u003c/em\u003e[6] classifiers. When compared with dictionary-based external software tools such as GNRD [7] and TaxonFinder [8], our top-3 classifiers were insensitive to words capitalization and were able to correctly classify novel species names that are absent in dictionary-based approaches with accuracies between 88.6% and 91.6%.\u003c/p\u003e\u003cp\u003eConclusions\u003c/p\u003e\u003cp\u003eOur results suggest that character bigram-based classification is a suitable method for distinguishing arthropod species names from regular English words and person names commonly found in scientific literature. Moreover, our method can also be used to reduce the number of false positives produced by dictionary-based methods.\u0026nbsp;\u003c/p\u003e","manuscriptTitle":"Recognition of arthropod species names using bigram-based classification","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2020-05-11 23:38:13","doi":"10.21203/rs.3.rs-26532/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"9beb9f6b-7543-43ae-a1ff-e5b1c4c0f7d0","owner":[],"postedDate":"May 11th, 2020","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":98220,"name":"Bioinformatics"}],"tags":[],"updatedAt":"2020-06-10T18:00:38+00:00","versionOfRecord":[],"versionCreatedAt":"2020-05-11 23:38:13","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-26532","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-26532","identity":"rs-26532","version":["v1"]},"buildId":"FbvkV6FR0MCFSLy54lSbu","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.