Sequence Tagging Approach in Grammar Error Detection: Identifying Areas of Improvement for the State-of-the-Art | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Sequence Tagging Approach in Grammar Error Detection: Identifying Areas of Improvement for the State-of-the-Art Qiao Wang, Zheng Yuan This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4479362/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract This study provides a qualitative evaluation of Seqtagger, a state-of-the-art machine learning-based sequence-tagging model developed for grammatical error detection (GED) and correction (GEC). The model's performance is evaluated on error detection against human benchmarks, with academic texts written by Japanese university students. Through human annotation and subsequent thematic analysis on failures in error detection, this study reveals that Seqtagger performs well in detecting errors related to simpler grammatical rules such as adverb position and prepositions in fixed collocations, with poorer performance in errors possibly influenced by the Japanese language, macro-structure errors and errors where human judgment is required. The underlying reasons for failures in detection are identified to be a narrow context window that fails to capture broader textual information, insufficient training data, particularly data that fully represents the linguistic characteristics of the Japanese students, and overgeneralization of patterns from the training data. These findings highlight the need for sequence-tagging GED and GEC tools to enhance their context window, be more adaptable to the diverse linguistic features of global learners and to enhance the ability to understand the linguistic complexities of the English language. Grammar Error Correction (GEC) Grammar Error Detection (GED) sequence tagging Seqtagger Japanese university students Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4479362","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":308838563,"identity":"d0bdf247-3442-45eb-aef8-492b6ce22296","order_by":0,"name":"Qiao Wang","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABBUlEQVRIiWNgGAWjYBACxgYGAwiLvQFIQNkSDAxsRGjhOUCkFiRlCQghCXzqmWckb5NgbLPLk498/PBxQQFDYn//4YM3fjDw5eF02Iy0MqCW5GLD22nGxjMMGBJn3EhLtuxhYCvGrSXHTIJxG3PixtkJZtI8Bgy5DTd4zCR4GNgSG/BrqU/cOPP4N7CW+efPf5P8Q1jL4cT5EjwQWzYcyGGTxmtLz7Nii8R/xxM38OQUG/MYSNRvvJFmbC1jgNsvhu3JG298OFOdOL/9+MbHPH9sjOXOH354803FMZwhZtjAwAKOEYMDYD4sRgyOJeDSIg+Mmg9gBprTa3BqGQWjYBSMghEHAOUXVe8vLqUpAAAAAElFTkSuQmCC","orcid":"","institution":"Waseda University","correspondingAuthor":true,"prefix":"","firstName":"Qiao","middleName":"","lastName":"Wang","suffix":""},{"id":308838564,"identity":"0ed90b57-800c-4b6d-9ff6-ec4914a81b9c","order_by":1,"name":"Zheng Yuan","email":"","orcid":"","institution":"King's College London","correspondingAuthor":false,"prefix":"","firstName":"Zheng","middleName":"","lastName":"Yuan","suffix":""}],"badges":[],"createdAt":"2024-05-26 08:53:22","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4479362/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4479362/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":80738630,"identity":"e3c2edf3-7836-46fb-ac29-b483ec7f2a37","added_by":"auto","created_at":"2025-04-16 14:01:59","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":485784,"visible":true,"origin":"","legend":"","description":"","filename":"submission.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4479362/v1_covered_d908644f-ea32-46da-921a-ae4e5dc8ca54.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Sequence Tagging Approach in Grammar Error Detection: Identifying Areas of Improvement for the State-of-the-Art","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Grammar Error Correction (GEC), Grammar Error Detection (GED), sequence tagging, Seqtagger, Japanese university students","lastPublishedDoi":"10.21203/rs.3.rs-4479362/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4479362/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"This study provides a qualitative evaluation of Seqtagger, a state-of-the-art machine learning-based sequence-tagging model developed for grammatical error detection (GED) and correction (GEC). The model's performance is evaluated on error detection against human benchmarks, with academic texts written by Japanese university students. Through human annotation and subsequent thematic analysis on failures in error detection, this study reveals that Seqtagger performs well in detecting errors related to simpler grammatical rules such as adverb position and prepositions in fixed collocations, with poorer performance in errors possibly influenced by the Japanese language, macro-structure errors and errors where human judgment is required. The underlying reasons for failures in detection are identified to be a narrow context window that fails to capture broader textual information, insufficient training data, particularly data that fully represents the linguistic characteristics of the Japanese students, and overgeneralization of patterns from the training data. These findings highlight the need for sequence-tagging GED and GEC tools to enhance their context window, be more adaptable to the diverse linguistic features of global learners and to enhance the ability to understand the linguistic complexities of the English language.","manuscriptTitle":"Sequence Tagging Approach in Grammar Error Detection: Identifying Areas of Improvement for the State-of-the-Art","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-06-07 04:20:57","doi":"10.21203/rs.3.rs-4479362/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"ff452445-39dc-466f-a0c1-1549a0a1190b","owner":[],"postedDate":"June 7th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-04-16T13:53:47+00:00","versionOfRecord":[],"versionCreatedAt":"2024-06-07 04:20:57","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-4479362","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4479362","identity":"rs-4479362","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.