Variability in Low-Resource Machine Translation Evaluation: Authentic vs. LLM-Generated Training Corpora | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Variability in Low-Resource Machine Translation Evaluation: Authentic vs. LLM-Generated Training Corpora Sofía García González¹, German Rigau Claramunt², Jose Ramom Pichel Campos This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7466136/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract The evaluation of machine translation systems often relies on a single metric and a single test dataset, an approach that can yield misleading system comparisons and premature conclusions regarding translation quality. A further complicating factor is the presence of translationese in test data, i.e. linguistic features specific to translated texts which can significantly influence both human and automatic assessments, particularly when present in the source language. In this paper, we examine the variability of MT evaluation results across datasets for English–Galician, Spanish–Galician, and Portuguese–Galician. We investigate whether the variability observed in prior research work can be replicated by training our own models from scratch, explore the feasibility of generating synthetic training corpora using large language models, and assess whether results obtained with authentic data can be reproduced with synthetic corpora. Additionally, we examine the relationship between dataset variability and the translationese effect. Our findings provide new insights into the influence of dataset composition on machine translation evaluation, the utility of large language models generated corpora, and the challenges associated with evaluating translation quality in low-resource settings. variability translationese low-resource machine translation Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7466136","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":574455246,"identity":"060408b6-d0af-490c-b518-7d641431dc73","order_by":0,"name":"Sofía García González¹","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA80lEQVRIiWNgGAWjYFACxgcHoIwGhg8GQJqZoBZmA7gWxhnEakEweYhxlnx7M+OBHwy18ubthxs/2xRskzc4znuA4UfFNpxaDM4cZjjYw3DccM6ZxGbpHIPbhjOb+RIYe87cxq1FIv/AAR6GY4wzJBgbQFoY+5l5DJgZ23BrkZ+RzHDwD8Mxe6CW5t8WBrft2whpYbiRzHCYh6EmEailTZrB4HYiQVtAfjksY3AgeQZPYptlj8Ht5JnNPAYH8fkFGGLMH99U1NnOYD/++MaPP7dtN5w/Y/jgRwUeh0HsOozKP0BAPQjUEaFmFIyCUTAKRiwAAMy/VUPNy8hZAAAAAElFTkSuQmCC","orcid":"","institution":"University of the Basque Country","correspondingAuthor":true,"prefix":"","firstName":"Sofía","middleName":"García","lastName":"González¹","suffix":""},{"id":574455247,"identity":"8fc75887-0dbe-4af5-8096-efc0b48b5078","order_by":1,"name":"German Rigau Claramunt²","email":"","orcid":"","institution":"²IXA GROUP (EHU)","correspondingAuthor":false,"prefix":"","firstName":"German","middleName":"Rigau","lastName":"Claramunt²","suffix":""},{"id":574455248,"identity":"d2b5dd5e-9354-4226-a6f8-3802d408ab64","order_by":2,"name":"Jose Ramom Pichel Campos","email":"","orcid":"","institution":"³Centro de Investigación en Tecnoloxías da Información (CiTIUS)","correspondingAuthor":false,"prefix":"","firstName":"Jose","middleName":"Ramom Pichel","lastName":"Campos","suffix":""}],"badges":[],"createdAt":"2025-08-26 21:08:04","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7466136/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7466136/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":100761211,"identity":"b413d139-6a90-4f28-9ca3-c1e6fde60190","added_by":"auto","created_at":"2026-01-21 07:37:26","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1198651,"visible":true,"origin":"","legend":"","description":"","filename":"LRECJpaperLRLMT.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7466136/v1_covered_85c1e501-d737-448c-ba18-4980605ab9a4.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Variability in Low-Resource Machine Translation Evaluation: Authentic vs. LLM-Generated Training Corpora","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"variability, translationese, low-resource machine translation","lastPublishedDoi":"10.21203/rs.3.rs-7466136/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7466136/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"The evaluation of machine translation systems often relies on a single metric and a single test dataset, an approach that can yield misleading system comparisons and premature conclusions regarding translation quality. A further complicating factor is the presence of translationese in test data, i.e. linguistic features specific to translated texts which can significantly influence both human and automatic assessments, particularly when present in the source language. In this paper, we examine the variability of MT evaluation results across datasets for English–Galician, Spanish–Galician, and Portuguese–Galician. We investigate whether the variability observed in prior research work can be replicated by training our own models from scratch, explore the feasibility of generating synthetic training corpora using large language models, and assess whether results obtained with authentic data can be reproduced with synthetic corpora. Additionally, we examine the relationship between dataset variability and the translationese effect. Our findings provide new insights into the influence of dataset composition on machine translation evaluation, the utility of large language models generated corpora, and the challenges associated with evaluating translation quality in low-resource settings.","manuscriptTitle":"Variability in Low-Resource Machine Translation Evaluation: Authentic vs. LLM-Generated Training Corpora","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-01-21 07:16:21","doi":"10.21203/rs.3.rs-7466136/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"186dd0e5-a9d9-4f5b-b936-0126b799b888","owner":[],"postedDate":"January 21st, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-01-21T07:16:21+00:00","versionOfRecord":[],"versionCreatedAt":"2026-01-21 07:16:21","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7466136","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7466136","identity":"rs-7466136","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.