ChavacanoMT: A Corpus and Evaluation of Neural Machine Translation for Philippine Creole Spanish | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article ChavacanoMT: A Corpus and Evaluation of Neural Machine Translation for Philippine Creole Spanish Aileen Joan Vicente, Theresse Faith Amamampang, Dunn Dexter Lahaylahay, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-5022127/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 13 Jan, 2026 Read the published version in Language Resources and Evaluation → Version 1 posted 10 You are reading this latest preprint version Abstract Chavacano, formally referred to as Philippine Creole Spanish, is the only Creole spoken in the Philippines. Like many languages, especially Creoles, computational studies on Chavacano are scarce because of the dearth of available corpora. This paper describes the creation of ChavacanoMT, a benchmark corpus for the machine translation study of Philippine Creole Spanish. ChavacanoMT consists of 767,053 parallel sentences between Chavacano and related languages, Spanish, Cebuano, Hiligaynon, Tagalog, and English. It is sourced from scraped bible translations and articles on the Jehovah’s Witness website. This paper also presents the performance of a multilingual neural machine translation model generated using ChavacanoMT. We report an overall 17 BLEU score on a fine-tuned mT5 model, outperforming an mT5-based model trained from scratch. Our experiments show that ChavacanoMT can generate models on par with a similar system that translates between English and some Philippines languages despite having fewer sentence samples used in training. We also report an improved Chavacano translation to and from its related languages that can be used as benchmark data. In particular, we highlight more than 20 BLEU points of improvement in the translation between Chavacano and English. The study opens avenues for exploring cross-linguistic interactions of Chavacano and its related languages in its translation that may benefit other low-resource languages. Philippine Creole Spanish Chavacano Translation Corpus Multilingual Translation Chavacano Figures Figure 1 Figure 2 Figure 3 1. Introduction the Philippines (Komisyon sa Wikang Filipino, 2020) and the only Spanish-based Creole in Asia. Like many languages worldwide, the computational study of the Chavacano language is scarce. This is primarily due to the dearth of available corpora from which computational and other language studies can be made. Chavacano is also one of the regional languages in the Philippines that is less computationally studied than other major languages such as Tagalog, Cebuano, and Hiligaynon. The Komisyon sa Wikang Filipino (2020) identified Chavacano as one of the 39 languages in the Philippines already in various stages of endangerment that need preservation and revival. The OPUS Corpus Collection 1 registers 2,529 English-Chavacano parallel sentences only. No other resources are published. This dataset was also included in PH-MNMT (Coronia, 2022). We aim to augment the existing language resource for Chavacano to encourage computational work on the language. This paper reports the creation of the ChavacanoMT corpus for benchmarking Chavacano machine translation. The corpus comprises parallel sentences between Chavacano and its related languages, Spanish, Cebuano, Hiligaynon, English, and Tagalog. We also report its usage in a many-to-many machine translation of Chavacano and its related languages. The remainder of this paper is organized as follows: Section 2 presents the literature on corpus creation for Philippine languages, supporting this work’s contribution to language resources. Information on the Chavacano Creole language and its linguistic properties are also presented. Section 3 details the creation of the ChavacanoMT. This section also presents descriptive information about the corpus. The evaluation of the corpus is shown in a translation experiment discussed in Sections 4 and 5. Finally, the insights from this work, including future directions on the computational linguistic study of Chavacano, are presented in Section 6. 2. Related Works This paper focuses on creating datasets for the under-resource Chavacano language in the context of leveraging related languages to improve the translation quality of low-resource languages (Baliber, Cheng, Adlaon, & Mamonong, 2020; Coronia, 2022; Dabre, Chu, & Kunchukuttan, 2021; Robinson, Hogan, Fulda, & Mortensen, 2022) such as Chavacano. As of this writing, only the work of Coronia (2022) explored the translation of Chavacano to other Philippine languages, but the results were poor. This is primarily due to poor language representation in the PH-MNMT (Coronia, 2022) dataset that was used. Moreover, the Creole nature of Chavacano makes it an exciting topic for computational linguistic study, especially since there are few NLP works on Creole languages. The machine translation of Creoles is also under-researched, owing to the lack of publicly available datasets (Dabre & Sukhoo, 2022). Ethnologue (Eberhard, Simons, & Fenning, 2024) registers 92 Creole languages worldwide. In the literature, however, only Haitian (Robinson et al., 2022), Nigerian Pidgin (Ahia & Ogueji, 2020), and Kreol Morisien (Dabre & Sukhoo, 2022) have been explored. 2.1 Chavacano: Philippine Creole Spanish The Philippine Creole Spanish, known as Chavacano, comprises three major dialects in Ternate, Cavite, and Zamboanga (Lipski, 2001), Philippines. Both the Ternate and Cavite dialects are classified as the Manila Bay PCS. Ternaten˜o was the oldest Spanish-based creole, and Caviten˜o was an off-shoot. Zamboanguen˜o, on the other hand, comprises the largest group of Chavacano speakers in Zamboanga City and neighboring towns and cities in Mindanao. Aside from the population of speakers, Zamboanguen˜o is actively used in blogs, news, and social media that can be used as digital resources. Zamboanguen˜o has support from its local government (DepEd-IX, 2016) while Ternaten˜o and Caviten˜o Chavacano did not receive such support from the national or local government (Lesho & Sippola, 2013). Chavacano in Zamboanga is surviving; the language is dying in Ternate and Cavite City (Genuino, 2005). With this information, this study assumes that the digital resources used in creating the corpus are mostly written in Zamboanguen˜o, as the language is distinctly recognized and used. The formation of Chavacano in Zamboanga resulted from historical and cultural interactions in the Philippines during the Spanish colonial period. It is considered a Creole language, meaning it emerged as a stable and fully developed language from a mixture of different languages, giving it some properties unique from its source languages. Chavacano belongs to the Creole family of languages of Spanish descent (Eberhard et al., 2024). Lipski (1992, 2001) reported an exhaustive investigation of the Chavacano language’s historical and social underpinnings in Zamboanga. Accordingly, Chavacano started to develop during the Spanish garrison in Zamboanga. Ilonggo later influenced Chavacano as Iloilo became a stopover for ships from Manila to Zamboanga. Later in the 20th century, immigration from the Central Visayan region to southwest Mindanao added some Visayan or Cebuano items to the language. Over time, Chavacano has adopted English in its lexicon. This account by Lipski (1992, 2001) served as the basis for identifying related languages for Chavacano. 2.2 Chavacano Lexicon and Orthography The lexicon of Chavacano is largely Spanish (Lipski & Santoro, 2007) but with orthographic shifts. For Zamboanguen˜o, in particular, several stages of relexification occurred to include lexical items of Philippine origin from regional Visayan, Ilonggo, and occasionally Tagalog (Lipski, 2001). Zamboanguen˜o has also adopted a heavy English lexical transfer (Lipski, 1992) over time. In 2016, the Department of Education Region IX and the Local Government of Zamboanga City published a revised Zamboanga Chavacano Orthography (DepEd- IX, 2016) to standardize the use of written Chavacano. The standardization is based on how the present generation of Zamboanguen˜os uses Chavacano. The orthography describes a way of spelling out Chavacano words using the alphabet of the word’s traced etymology. For example, the Spanish-derived words zacate (grass) and man˜ana (tomorrow) are spelled using the Spanish’s abecedario . In contrast, the Chavacano words of local origin, like kanila (them) and kanamon (us), are spelled using the Philippine alphabet system. Some orthographic shifts are noted from loaned words, such as dropping the letter r in the Spanish verbs like comer (to eat), bailar (to dance), i.e., come , baila . It is also interesting to note that the Spanish writing utilizes diacritics that are not necessarily applied in Chavacano. In general, Chavacano words are spelled the way they are pronounced. 2.3 Chavacano Grammar Zamboanguen˜o’s grammatical structure differs from any Spanish variety, and while there are standard lexicons, the two are mutually non-intelligible (Lipski, 2001). Over three centuries of Philippine history influenced the morphology, grammar, and syntax of Zamboanguen˜o (Lipski & Santoro, 2007). Even so, it has retained its Austronesian foundation as evidenced by the Verb-Subject-Object word order, albeit many alternative possibilities (Lipski, 1992). The Philippine languages belong to the Austronesian language family. This contrasts Spanish’s Subject-Verb-Object word order (Lee, 2017). The conjugation of verbs to show tenses also does not apply in Chavacano (DepEd-IX, 2016). 3. ChavacanoMT 3.1 Dataset Collection Multi-source translations of the books of the bible and articles from the Jehovah’s Witnesses website (JW.org, 2023) serve as the main sources of parallel sentences. Multi-source refers to translations of texts in multiple languages, such as bible translations, EU parliamentary proceedings, and transcripts from TED talks and United Nations meetings, where a particular sentence has translations to many target languages. Dabre et al. (2021) encourages leveraging multi-source sentences whenever available, as this helps improve the translation. The Bible is probably the most tapped resource for parallel texts with translations. According to Wycliffe Global Alliance (2023), the full bible has already been translated into 736 languages and the New Testament into 1,658 languages. Chavacano only has translations of the New Testament. The Jehovah’s Witnesses website JW.org (2023) also contains translations of their articles in multiple languages. Aside from the general articles, the website makes translations of their Watchtower and Awake! magazines available. These magazines span more than two decades of publications. Besides multi-source texts, bilingual translations of Chavacano to the other related languages and vice versa will also be collected. This is done to augment the dataset, especially since only a few resources are available for the Philippine languages. Translations of the Chavacano texts to Spanish, English, Cebuano, and Hiligaynon from the Bible and articles are scraped using Webscraper.io 2 and Beautiful Soup 3 , a web-scraping library in Python. 3.2 Preprocessing The primary goal of preprocessing is to ensure the alignment of scraped texts. The alignment of bible texts is based on the verse numbers for each book chapter. The articles, however, are aligned based on the number of sentence chunks, i.e. group of sentences, captured per text object in the site’s web pages during scraping. In JW.org (2023), the web pages across languages share the same site map, and a single script was used to scrape all articles across the target languages. 3.2.1 Articles: Cleaning and Sentence Segmentation The aligned articles were preprocessed to remove unnecessary symbols and characters such as ellipses, bullets, asterisks, brackets, and non-breaking white characters. The quotation marks were also removed because some translations in other languages do not contain these. The punctuation marks, however, were retained. The Bible verses used as in-text references were also removed. For instance, the verse reference, -Santiago 2:14-17., was removed in the succeeding example excerpt. Ta ayuda con aquellos quien ta necesita.-Santiago 2:14-17. The scraped sentence chunks are segmented to extract individual sentences. Sentence segmentation was done using Sentence-Splitter 4 , a free module that allows splitting of text paragraphs into sentences. The heuristics algorithm, however, was based only on the English language. Other segmentation tools were considered, such as the tools from SpaCy, but the Sentence-Splitter produced more aligned chunks. The chunks that did not produce the same number of sentences across any or all article translations after segmentation are considered misaligned and, therefore, excluded from the final sample set. 3.2.2 Bible: Cleaning and Verse Segmentation The bible verses did not require the removal of unnecessary symbols. The quotation marks were retained together with the usual punctuation marks. Verses combined in one bible translation but separate in other translations, for example, Matthew 17 and 18 in the English translation, but Matthew 17-18 in Hiligaynon are removed. Each verse in the bible is considered a sentence sample, even if it spans more than one sentence or is a combination of sentence and sentence fragments. It is assumed that the translation was done per verse, and to preserve the meaning, the whole verse was used collectively as a sentence sample. The verse number, though, was removed. 3.3 Corpus Preparation Parallel sentences of language pairs between Chavacano, Cebuano, Hiligaynon, Tagalog, Spanish, and English comprise ChavacanoMT. These language pairs were prepared from the sentence samples collected. Table 1 summarizes each language pair’s number of sentence samples. Table 1 breaks the type of sentences included in the language-pair datasets as multisource (MS) and non-multisource (NMS). We note that most of the sentence samples for Chavacano are multisource, meaning each sentence sample has translations to the other five languages in the corpus. We also note that the number of Chavacano sentence samples is small compared to the other language pairs. Table 1 Number of Sentence Samples per Language Pair in ChavacanoMT. MS indicates Multisource, while NMS is Non-Multisource. As a reference, the language codes used are as follows: cbk for Chavacano, ceb for Cebuano, hil for Hiligaynon, tl for Tagalog, en for English, and es for Spanish. Language Pairs Bible MS Bible NMS Articles MS Articles NMS Total Sentences cbk-ceb 7,728 13,931 35 21,694 cbk-hil 7,728 13,931 21,659 cbk-es 7,728 13,931 35 21,694 cbk-en 7,728 13,931 35 21,694 cbk-tl 7,728 13,931 21,659 ceb-hil 7,728 21,803 13,931 51,664 95,126 ceb-es 7,728 21,803 13,931 51,776 95,238 ceb-en 7,728 21,803 13,931 51,954 95,416 ceb-tl 7,728 21,803 13,931 46,546 90,008 hil-es 7,728 21,803 13,931 2,865 46,327 hil-en 7,728 21,803 13,931 2,904 46,366 hil-tl 7,728 21,803 13,931 2,904 46,366 es-en 7,728 21,803 13,931 13,420 56,882 es-tl 7,728 21,803 13,931 43,462 en-tl 7,728 21,803 13,931 43,462 3.4 Corpus Statistics We describe the corpora using the following statistics: Average Sentence Length, Number of Unique Words, and Number of Shared or Overlapping Words per Language. The statistics are taken from all sentence samples per language, and the count may differ slightly if taken from each language pair. Figure 1 shows that the sentence samples from the Bible are longer than those from the articles. As discussed in Section 3.2.2, these sentences are verses in the bible that may comprise more than one sentence or sentence fragments. The number of unique words for each language in each resource is summarized in Figure 2. The number of unique words for Chavacano is arguably low compared to the other languages. We argue this was attributed to Chavacano being less morpho- logically rich than the other languages. Pahulaya (2022) presented that Chavacano comprises simple, compound, affixed, and reduplicated words like other languages. Its verbs, however, do not show inflections in tenses as the markers ya (past), ta (present), and ay (future) are being used. In addition, the Chavacano dictionary (de Dios, Maria Isabelita Riego, 1989) registers about 6,500 entries only, including both heads and derived forms as described in SEAlang Library 5 . This shows that the number of words captured in the ChavacanoMT corpus is reasonable. The lexical similarity among languages can be measured from the number of overlapping words across languages. Figure 3 shows this similarity in the corpus. Figure 3 shows that Chavacano samples share more word overlaps with Spanish, i.e., around 40% of the Chavacano words in the corpus are shared with Spanish. Overlaps with the other languages are also present. The Philippine languages, Cebuano, Hiligaynon, and Tagalog, share the most overlaps. The lexical similarity of Chavacano and the related languages shows the influence of these languages on Chavacano. It is also essential to consider that Spanish and English have a lexical influence on the Philippine languages due to years of colonialism. 4. Chavacano Multilingual Neural Machine Translation In this section, we present the utilization of the ChavacanoMT corpus in the neural machine translation of Chavacano to and from its related languages, Spanish, Cebuano, Hiligaynon, and English. 4.1 Model Training We experiment with a multilingual neural machine translation of Chavacano. It has already been established in the literature how related languages, especially those that are high-resource, support the translation of low-resource languages (Dabre,Nakagawa, & Kazawa, 2017; Dabre & Sukhoo, 2022; Goyal, Kumar, & Sharma, 2020; Tubay & Costa-Juss`a, 2018; Zoph, Yuret, May, & Knight, 2016). Such is the case in the neural machine translation of Chavacano. The multilingual training leverages high-resource languages in the dataset: Spanish, Cebuano, and English. In this experiment, we build a many-to-many machine translation model (a) from scratch using mT5 (Xue et al., 2021) model configuration and (b) from fine-tuning mT5 model using a subset of the ChavacanoMT corpus (subset described in Section 4.2). In training from scratch, a vocabulary of 32,300 sentence pieces was created using Sentencepiece (Kudo & Richardson, 2018) from combined words of languages in the corpus (Table 1). An mT5-based tokenizer, cbkTokenizer, was also trained from ChavacanoMT. The tokenizer was used to tokenize the dataset. In the case of fine-tuning, the vocabulary of mT5 with 250,112 sentence pieces and its built-in MT5Tokenizer were used. The tokenizer and models were trained using Huggingface’s Transformers 6 library. The model training ran in 8 epochs for both models and was optimized using AdamWeightDecay with a learning rate of 0.001. Bilingual Evaluation Understudy (BLEU) (Papineni, Roukos, Ward, & Zhu, 2002) scores using Sacrebleu (Post, 2018) and Recall-Oriented Understudy for Gisting Evaluation (ROUGE) (Lin, 2004) scores are collected to measure the model’s performance. 4.2 Datasets Used Based on the historical investigation of Lipski (1992, 2001), the creolization of Chavacano is influenced by Spanish as its lexifier and Philippine languages, Cebuano and Hiligaynon, as adstrates. It has also undergone lexical transfer from English over time. Table 2 Multilingual Many-to-Many Dataset Train Test Validation Total Samples cbk-ceb 15,185 3,255 3,254 21,694 cbk-en 15,185 3,255 3,254 21,694 cbk-es 15,185 3,255 3,254 21,694 cbk-hil 15,161 3,249 3,249 21,659 ceb-cbk 15,185 3,255 3,254 21,694 ceb-en 66,791 14,313 14,312 95,416 ceb-es 66,666 14,286 14,286 95,238 ceb-hil 66,588 14,269 14,269 95,126 en-cbk 15,185 3,255 3,254 21,694 en-ceb 66,791 14,313 14,312 95,416 en-es 39,817 8,533 8,532 56,882 en-hil 32,448 6,954 6,953 46,355 es-cbk 15,185 3,255 3,254 21,694 es-ceb 66,666 14,286 14,286 95,238 es-en 39,817 8,533 8,532 56,882 es-hil 32,428 6,950 6,949 46,327 hil-cbk 15,161 3,249 3,249 21,659 hil-ceb 66,588 14,269 14,269 95,126 hil-en 32,434 6,951 6,950 46,335 hil-es 32,428 6,950 6,949 46,327 This account became the basis of the multilingual dataset for Chavacano translation. The language pairs involving Chavacano, Spanish, Cebuano, Hiligaynon, and English were used except for Tagalog. In this experiment’s many-to-many training set-up, the language pairs are used in both directions, that is, cbk-en and en-cbk are used. In total, 20 language pairs between Chavacano, Spanish, Cebuano, Hiligaynon, and English were included in the dataset. In this training set-up, all languages serve as source and target languages. It is hoped that by doing so, the similarities shared between Chavacano and the high-resource languages may be represented in both the encoder and decoder sides of the neural network. The summary of the dataset used in this experiment is shown in Table 2. 70% of the dataset was used as a training set, 15% as a validation set, and 15% as a test set. The data splits are stratified according to language pairs. The dataset is unbalanced, with more training samples from the high-resource Cebuano, English, and Spanish language pairs. 5 Results Table 3 shows the BLEU and ROUGE-1 scores of the models generated in the experiment. As expected, the fine-tuned model earned better results than the model generated from scratch. This demonstrates the leverage one can get on pre-trained weights even if target languages are not included in the pre-training. Table 3 Comparison of BLEU and ROUGE-1 scores for (a) mT5-based model from scratch and (b) fine-tuned mT5 model. Model Type BLEU ROUGE- 1 Scratch 0.4879 0.1222 Finetuned 17.8274 0.5457 The fine-tuned model was tested further using the individual language pairs in the test set. Table 4 presents the performance scores for each translation direction. Table 4 Performance results for fine-tuned mT5 model. Table (b) on the left shows the translation direction with better performance. BLEU ROUGE- 1 BLEU ROUGE- 1 cbk-ceb 21.95 0.62 ceb-cbk 23.54 0.66 hil-ceb 22.34 0.60 ceb-hil 22.79 0.61 en-ceb 20.02 0.59 ceb-en 26.25 0.61 en-cbk 24.06 0.68 cbk-en 35.30 0.69 ceb-es 16.61 0.50 es-ceb 16.34 0.53 hil-cbk 22.10 0.65 cbk-hil 23.51 0.64 en-hil 16.20 0.57 hil-en 23.70 0.59 es-hil 15.13 0.55 hil-es 16.63 0.52 es-cbk 21.79 0.64 cbk-es 24.68 0.60 en-es 19.45 0.54 es-en 25.94 0.60 The results show that the translation from Philippine to foreign language seems better than the other way around, except between Cebuano and Spanish, where the scores are essentially the same in both directions. It is also interesting to note that the highest-performing language pair at 35.30 BLEU is Chavacano-English, whose training samples are among the lowest in the dataset. As a way of benchmarking, we also compare the Ph-EN and EN-Ph BLEU scores obtained from the experiments of Coronia (2022) with the result of our experiments. Ph here refers to Philippine languages. Table 5 summarizes the best-performing model from Coronia (2022) and the BLEU scores from our experiments. Table 5 Comparison of bilingual translations from finetuned mT5 multilingual models. Coronia (2022) uses the PH-MNMT dataset, while ours uses the ChavacanoMT corpus. en-ceb ceb-en en-hil hil-en en-cbk cbk-en Fine-tuned mT5 (Coronia) 21.25 26.67 24.4 26.11 0.08 0.79 Fine-tuned mT5 (Ours) 20.02 26.25 16.20 23.70 24.06 35.30 For context, the PH-MNMT dataset used in Coronia (2022) comprises millions of English-Cebuano and English-Hiligaynon sentence samples compared to ChavacanoMT. ChavacanoMT, however, has more sentence samples for Chavacano. In Coronia (2022), the best-performing mT5 model was fine-tuned on EN-Ph and Ph-EN sentences comprising English, Tagalog, Cebuano, Hiligaynon, Waray, and Chavacano. In contrast, our experiment generated the translation model using a many-to-many training setup involving languages that are specifically related to Chavacano. Both Coronia (2022) and our translation experiment considered using related languages in model training. Although an objective comparison between the models cannot be made, we can infer that ChavacanoMT can produce models that are on par with published results. 6 Conclusions This paper presents ChavacanoMT, a benchmark corpus for the machine translation of Chavacano to and from related languages, Spanish, Cebuano, Hiligaynon, Tagalog, and English. The corpus consists of 767,053 parallel sentence samples from 15 language pairs. The corpus is a combination of multi-source and parallel sentences. Using the ChavacanoMT corpus in our machine translation experiments has demonstrated its potential to significantly enhance the translation quality between Chavacano and its related languages. Our experiments showed that models trained using ChavacanoMT could achieve performance on par with or surpass existing multilingual neural machine translation systems involving Chavacano, particularly in translating between Chavacano and English, with improvements exceeding 20 BLEU points. This highlights the effectiveness of incorporating diverse related languages with multisource sentence samples in building robust machine translation models for low-resource languages in Chavacano. Our findings suggest that leveraging related languages within the corpus improves translation accuracy. This approach underscores the importance of using carefully curated multilingual datasets to support underrepresented languages, ultimately contributing to their preservation and wider accessibility. The insights gained from this study open avenues for exploring the influence of specific linguistic relationships in translation quality, which could guide the development of more targeted translation strategies for other low-resource languages. Additionally, ChavacanoMT provides a foundation for further research on cross- linguistic interactions and their implications for the evolution and revitalization of Chavacano. By extending this work, researchers can deepen their understanding of how digital tools and computational methods can support Creole-speaking communities’ linguistic and cultural heritage. Declarations Acknowledgements. We acknowledge the Jehovah’s Witness organization for the permission to scrape their website www.jw.org to support the NLP research on Philippine languages. Funding: This work was supported by the Commission on Higher Education through its Scholarships for Instructors’ Knowledge Advancement Program (SIKAP) grant. Conflict of interest/Competing interests: The authors have no competing interests to declare relevant to this article’s content. Data availability: The corpus built in this study is available from the authors, but restrictions apply. Some resources used in building the corpus were under a non- commercial agreement from Watch Tower Bible and Tract Society, Philippines for the current study. The authors wish to honor the permission by ensuring that out- comes are used only for academic purposes. Data are, however, available from the authors upon reasonable request. References Ahia, O., & Ogueji, K. (2020). Towards Supervised and Unsupervised Neural Machine Translation Baselines for Nigerian Pidgin. AfricaNLP Workshop. Online. Retrieved from http://arxiv.org/abs/2003.12660 Baliber, R.I., Cheng, C., Adlaon, K.M., Mamonong, V. (2020, December). Bridging Philippine Languages with Multilingual Neural Machine Translation. Proceedings of the 3rd Workshop on Technologies for MT of Low Resource Languages (pp. 14–22). Association for Computational Linguistics. Retrieved from https://aclanthology.org/2020.loresmt-1.2.pdf Coronia, J.D. (2022). Exploring clustering of Philippine languages in multilingual neural machine translation. (Unpublished master’s thesis). De La Salle University, Manila, Philippines. (Retrieved from https://animorepository.dlsu.edu.ph/etdm softtech/4) Dabre, R., Chu, C., Kunchukuttan, A. (2021, September). A Survey of Multilingual Neural Machine Translation. ACM Computing Surveys , 53 (5), 1–38, https://doi.org/10.1145/3406095 Retrieved 2023-09-12, from https://dl.acm.org/doi/10.1145/3406095 Dabre, R., Nakagawa, T., Kazawa, H. (2017, November). An Empirical Study of Language Relatedness for Transfer Learning in Neural Machine Translation. Proceedings of the 31st Pacific Asia Conference on Language, Information and Computation (pp. 282–286). The National University (Philippines). Retrieved from https://aclanthology.org/Y17-1038 Dabre, R., & Sukhoo, A. (2022, November). KreolMorisienMT: A dataset for Mauritian Creole machine translation. Findings of the association for computational linguistics: Aacl-ijcnlp 2022 (pp. 22–29). Online only: Association for Computational Linguistics. Retrieved from https://aclanthology.org/2022.findings-aacl.3 de Dios, Maria Isabelita Riego (1989). A composite dictionary of Philippine Creole Spanish. Studies in Philippine Linguistics. Manila: Linguistic Society of the Philippines and Summer Institute of Linguistics. DepEd-IX (2016). Zamboanga Chavacano Orthography . Local Government of Zamboanga City: Philippines. Eberhard, D., Simons, G., & Fenning, C. (Eds.). (2024). Ethnologue: Languages of the World (27th ed.). Dallas, Texas: SIL International. Genuino, C.F. (2005). Language extinction in process across Chabacano communities: A sociolinguistic approach. (Unpublished doctoral disserta- tion). De La Salle University, Manila, Philippines. (Retrieved from https://animorepository.dlsu.edu.ph/etd doctoral/87) Goyal, V., Kumar, S., Sharma, D.M. (2020). Efficient Neural Machine Translation for Low-Resource Languages via Exploiting Related Languages. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: Student Research Workshop (p. 162–168). Association for Computational Linguistics. JW.org (2023). Official Website of Jehovah’s Witnesses. Available at http://www.https://www.jw.org/en/ (2023/10/18). Komisyon sa Wikang Filipino (2020). Repositoryo ng mga Wika at Kultura. https:// kwfwikaatkultura.ph/chabacano/. Kudo, T., & Richardson, J. (2018, November). SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. E. Blanco & W. Lu (Eds.), Proceedings of the 2018 conference on empirical methods in natural language processing: System demonstrations (pp. 66–71). Brussels, Belgium: Association for Computational Linguistics. Retrieved from https://aclanthology.org/D18-2012 Lee, J.F. (2017). Word order and linguistic factors in the second language processing of spanish passive sentences. Hispania , 100 (4), 580–595, Retrieved 2023-06-26, from https://www.jstor.org/stable/26387810 Lesho, M., & Sippola, E. (2013). The sociolinguistic situations of Manila Bay Chabacano-speaking communities. Language Documentation and Conservation , 7 , 1–30, Lin, C. (2004, July). ROUGE: A package for automatic evaluation of summaries. Text Summarization Branches Out (pp. 74–81). Barcelona, Spain: Association for Computational Linguistics. Retrieved from https://www.aclweb.org/anthology/W04-1013 Lipski, J. (1992). New thoughts on the origins of Zamboanguen˜o (Philippine Creole Spanish). Language Sciences , 14 (3), 197-231, https:// doi.org/https://doi.org/10.1016/0388-0001(92)90005-Y Retrieved from https://www.sciencedirect.com/science/article/pii/038800019290005Y Lipski, J. (2001, Aug.). The place of Chabacano in the Philippine linguistic profile. Sociolinguistic Studies , 2 (2), 119–163, https://doi.org/10.1558/sols.v2i2.119 Retrieved from https://journal.equinoxpub.com/SS/article/view/11691 Lipski, J., & Santoro, M. (2007). Zamboanguen˜o Creole Spanish [Bibliographical record]. J. Holm & P. Patrick (Eds.), Comparative creole syntax. parallel outlines of 18 creole grammars (p. 373-398). London: Battlebridge. (Much information is based on Forman (1972).) Pahulaya, V.L. (2022). Morphological Analysis on the Structure of Chavacano Language: A Complex Mental Process. NeuroQuantology , 20 (6), 9820-9830, https://doi.org/https://doi.org/10.14704/nq.2022.20.6.NQ22960 Papineni, K., Roukos, S., Ward, T., Zhu, W. (2002). BLEU: a Method for Automatic Evaluation of Machine Translation. Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (ACL) (pp. 311–318). Philadelphia. Post, M. (2018, October). A Call for Clarity in Reporting BLEU Scores. Proceedings of the Third Conference on Machine Translation: Research Papers (pp. 186– 191). Belgium, Brussels: Association for Computational Linguistics. Retrieved from https://www.aclweb.org/anthology/W18-6319 Robinson, N., Hogan, C., Fulda, N., Mortensen, D.R. (2022, October). Data-adaptive Transfer Learning for Translation: A Case Study in Haitian and Jamaican. Proceedings of the Fifth Workshop on Technologies for Machine Translation of Low-Resource Languages (LoResMT 2022) (pp. 35–42). Gyeongju, Republic of Korea: Association for Computational Linguistics. Retrieved from https://aclanthology.org/2022.loresmt-1.5 Tubay, B., & Costa-Juss`a, M.R. (2018, October). Neural machine translation with the transformer and multi-source Romance languages for the biomedical WMT 2018 task. Proceedings of the Third Conference on Machine Translation: Shared Task Papers (pp. 667–670). Belgium, Brussels: Association for Computational Linguistics. Retrieved from https://aclanthology.org/W18-6449 Wycliffe Global Alliance (2023). 2023 Global Scripture Access. https://www.wycliffe.net/resources/statistics/. (Accessed: February 9, 2024) Xue, L., Constant, N., Roberts, A., Kale, M., Al-Rfou, R., Siddhant, A., . . . Raf- fel, C. (2021, June). mT5: A massively multilingual pre-trained text-to-text transformer. K. Toutanova et al. (Eds.), Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 483–498). Online: Association for Computational Linguistics. Retrieved from https://aclanthology.org/2021.naacl-main.41 Zoph, B., Yuret, D., May, J., Knight, K. (2016). A Transfer learning for low-resource neural machine translation. Proceedings of the Conference on Empirical Methods in Natural Language Processing (p. 1568–1575). Association for Computational Linguistics. Footnotes 1 https://opus.nlpl.eu/ 2 https://webscraper.io/ 3 https://www.crummy.com/software/BeautifulSoup/ 4 https://github.com/mediacloud/sentence-splitter?tab=readme-ov-file 5 http://sealang.net/chavacano/dictionary.htm 6 https://huggingface.co/docs/transformers/en/index Additional Declarations No competing interests reported. Supplementary Files AppendixA.docx Cite Share Download PDF Status: Published Journal Publication published 13 Jan, 2026 Read the published version in Language Resources and Evaluation → Version 1 posted Editorial decision: Revision requested 12 May, 2025 Reviews received at journal 10 May, 2025 Reviewers agreed at journal 03 Feb, 2025 Reviews received at journal 02 Jan, 2025 Reviewers agreed at journal 08 Nov, 2024 Reviewers agreed at journal 06 Nov, 2024 Reviewers invited by journal 04 Nov, 2024 Editor assigned by journal 08 Sep, 2024 Submission checks completed at journal 03 Sep, 2024 First submitted to journal 03 Sep, 2024 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-5022127","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":354764092,"identity":"3f6fc21a-bfdb-4e60-a95e-fb3a3def5e7a","order_by":0,"name":"Aileen Joan Vicente","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABGElEQVRIiWNgGAWjYDACZuYGMM0G4dow9sG5Eri0MCJpOcCQxthGUAsDVAsDRMthwlr42xkbH3zcwZDYx3742OMPFedl2/jPGDB8KDvMYC7dgFWLxGHGZsOZZxgS23jS0g0OnLlt3CaRY8A449xhBss5B7BbA3SJNG8bgzGbBI+ZxMG224ltEjwGzLxthxkMbiRg1SF/mLH991+wFv5vEgf/nUsEOYz5Lx4tBkBbmIFelgPawiZxsOFAYhtDjgFQBLcWQ6BfJHtBWnjSzCTOHEsG+iWt4GDPuXQeyxnYtcidP3zww882Bh759sPPJCpq7GT7+Q9vfPCjzFrOXAK7Fij4j8o9AMQ8Bvg0YAdkaBkFo2AUjILhCQAJn1rRLnWL+AAAAABJRU5ErkJggg==","orcid":"","institution":"De La Salle University","correspondingAuthor":true,"prefix":"","firstName":"Aileen","middleName":"Joan","lastName":"Vicente","suffix":""},{"id":354764093,"identity":"08331618-43eb-4603-b477-80857acebb0c","order_by":1,"name":"Theresse Faith Amamampang","email":"","orcid":"","institution":"University of the Philippines Cebu","correspondingAuthor":false,"prefix":"","firstName":"Theresse","middleName":"Faith","lastName":"Amamampang","suffix":""},{"id":354764094,"identity":"44d3ad21-dc8d-46fb-b90d-66bc3b400200","order_by":2,"name":"Dunn Dexter Lahaylahay","email":"","orcid":"","institution":"University of the Philippines Cebu","correspondingAuthor":false,"prefix":"","firstName":"Dunn","middleName":"Dexter","lastName":"Lahaylahay","suffix":""},{"id":354764095,"identity":"032c6cfb-cfef-4a55-a863-fa72cabdfb15","order_by":3,"name":"Charibeth Cheng","email":"","orcid":"","institution":"De La Salle University","correspondingAuthor":false,"prefix":"","firstName":"Charibeth","middleName":"","lastName":"Cheng","suffix":""}],"badges":[],"createdAt":"2024-09-03 05:44:48","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-5022127/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-5022127/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1007/s10579-025-09888-3","type":"published","date":"2026-01-13T16:28:47+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":66322677,"identity":"562cfef4-7912-4fa9-8f18-04f43763c40b","added_by":"auto","created_at":"2024-10-10 12:12:53","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":50018,"visible":true,"origin":"","legend":"\u003cp\u003eAverage Sentence Length (based on the number of words) per language resource.\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-5022127/v1/6c69cba95c55549685724cee.png"},{"id":66322679,"identity":"54e4fbb3-c475-4ce9-a057-3f1ed22c2bf7","added_by":"auto","created_at":"2024-10-10 12:12:53","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":39383,"visible":true,"origin":"","legend":"\u003cp\u003eCounts of the unique words per language resource.\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-5022127/v1/4edbb9864d683a766492357f.png"},{"id":66322678,"identity":"c1746e94-fa65-4757-a29d-5a231424f5ca","added_by":"auto","created_at":"2024-10-10 12:12:53","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":49737,"visible":true,"origin":"","legend":"\u003cp\u003eCount of overlapping words between languages. For example, the number of overlapping words between Chavacano and Spanish is 3,381.\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-5022127/v1/8bd6e92643b0d5be68db5ccc.png"},{"id":100614539,"identity":"80c95baf-9ca6-4ea9-836c-249ab0dad4dd","added_by":"auto","created_at":"2026-01-19 17:21:55","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":957055,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5022127/v1/2f048ee8-b340-4742-9d1b-49c6a07effe1.pdf"},{"id":66322680,"identity":"02ddc8ac-3f85-450c-ab36-85496a964b32","added_by":"auto","created_at":"2024-10-10 12:12:53","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":16760,"visible":true,"origin":"","legend":"","description":"","filename":"AppendixA.docx","url":"https://assets-eu.researchsquare.com/files/rs-5022127/v1/3de6301b06ec231d35184d56.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"ChavacanoMT: A Corpus and Evaluation of Neural Machine Translation for Philippine Creole Spanish","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003ethe Philippines (Komisyon sa Wikang Filipino, 2020) and the only Spanish-based Creole in Asia. Like many languages worldwide, the computational study of the Chavacano language is scarce. This is primarily due to the dearth of available corpora from which computational and other language studies can be made.\u003c/p\u003e\n\u003cp\u003eChavacano is also one of the regional languages in the Philippines that is less computationally studied than other major languages such as Tagalog, Cebuano, and Hiligaynon. The Komisyon sa Wikang Filipino (2020) identified Chavacano as one of the 39 languages in the Philippines already in various stages of endangerment that need preservation and revival.\u003c/p\u003e\n\u003cp\u003eThe OPUS Corpus Collection\u003csup\u003e1\u003c/sup\u003e registers 2,529 English-Chavacano parallel sentences only. No other resources are published. This dataset was also included in PH-MNMT (Coronia, 2022). We aim to augment the existing language resource for Chavacano to encourage computational work on the language.\u003c/p\u003e\n\u003cp\u003eThis paper reports the creation of the ChavacanoMT corpus for benchmarking Chavacano machine translation. The corpus comprises parallel sentences between Chavacano and its related languages, Spanish, Cebuano, Hiligaynon, English, and Tagalog. We also report its usage in a many-to-many machine translation of Chavacano and its related languages.\u003c/p\u003e\n\u003cp\u003eThe remainder of this paper is organized as follows: Section 2 presents the literature on corpus creation for Philippine languages, supporting this work’s contribution to language resources. Information on the Chavacano Creole language and its linguistic properties are also presented. Section 3 details the creation of the ChavacanoMT. This section also presents descriptive information about the corpus. The evaluation of the corpus is shown in a translation experiment discussed in Sections 4 and 5. Finally, the insights from this work, including future directions on the computational linguistic study of Chavacano, are presented in Section 6.\u003c/p\u003e"},{"header":"2.\tRelated Works","content":"\u003cp\u003eThis paper focuses on creating datasets for the under-resource Chavacano language in the context of leveraging related languages to improve the translation quality of low-resource languages (Baliber, Cheng, Adlaon, \u0026amp; Mamonong, 2020; Coronia, 2022; Dabre, Chu, \u0026amp; Kunchukuttan, 2021; Robinson, Hogan, Fulda, \u0026amp; Mortensen, 2022) such as Chavacano.\u003c/p\u003e\n\u003cp\u003eAs of this writing, only the work of Coronia (2022) explored the translation of Chavacano to other Philippine languages, but the results were poor. This is primarily due to poor language representation in the PH-MNMT (Coronia, 2022) dataset that was used.\u003c/p\u003e\n\u003cp\u003eMoreover, the Creole nature of Chavacano makes it an exciting topic for computational linguistic study, especially since there are few NLP works on Creole languages. The machine translation of Creoles is also under-researched, owing to the lack of publicly available datasets (Dabre \u0026amp; Sukhoo, 2022). Ethnologue (Eberhard, Simons, \u0026amp; Fenning, 2024) registers 92 Creole languages worldwide. In the literature, however, only Haitian (Robinson et al., 2022), Nigerian Pidgin (Ahia \u0026amp; Ogueji, 2020), and Kreol Morisien (Dabre \u0026amp; Sukhoo, 2022) have been explored.\u003c/p\u003e\n\u003ch2\u003e2.1 Chavacano: Philippine Creole Spanish\u003c/h2\u003e\n\u003cp\u003eThe Philippine Creole Spanish, known as Chavacano, comprises three major dialects in Ternate, Cavite, and Zamboanga (Lipski, 2001), Philippines. Both the Ternate and Cavite dialects are classified as the Manila Bay PCS. Ternaten\u0026tilde;o was the oldest Spanish-based creole, and Caviten\u0026tilde;o was an off-shoot. Zamboanguen\u0026tilde;o, on the other hand, comprises the largest group of Chavacano speakers in Zamboanga City and neighboring towns and cities in Mindanao. Aside from the population of speakers, Zamboanguen\u0026tilde;o is actively used in blogs, news, and social media that can be used as digital resources. Zamboanguen\u0026tilde;o has support from its local government (DepEd-IX, 2016) while Ternaten\u0026tilde;o and Caviten\u0026tilde;o Chavacano did not receive such support from the national or local government (Lesho \u0026amp; Sippola, 2013). Chavacano in Zamboanga is surviving; the language is dying in Ternate and Cavite City (Genuino, 2005). With this information, this study assumes that the digital resources used in creating the corpus are mostly written in Zamboanguen\u0026tilde;o, as the language is distinctly recognized and used.\u003c/p\u003e\n\u003cp\u003eThe formation of Chavacano in Zamboanga resulted from historical and cultural interactions in the Philippines during the Spanish colonial period. It is considered a Creole language, meaning it emerged as a stable and fully developed language from a mixture of different languages, giving it some properties unique from its source languages. Chavacano belongs to the Creole family of languages of Spanish descent (Eberhard et al., 2024).\u003c/p\u003e\n\u003cp\u003eLipski (1992, 2001) reported an exhaustive investigation of the Chavacano language\u0026rsquo;s historical and social underpinnings in Zamboanga. Accordingly, Chavacano started to develop during the Spanish garrison in Zamboanga. Ilonggo later influenced Chavacano as Iloilo became a stopover for ships from Manila to Zamboanga. Later in the 20th century, immigration from the Central Visayan region to southwest Mindanao added some Visayan or Cebuano items to the language. Over time, Chavacano has adopted English in its lexicon. This account by Lipski (1992, 2001) served as the basis for identifying related languages for Chavacano.\u003c/p\u003e\n\u003ch2\u003e2.2 Chavacano Lexicon and Orthography\u003c/h2\u003e\n\u003cp\u003eThe lexicon of Chavacano is largely Spanish (Lipski \u0026amp; Santoro, 2007) but with orthographic shifts. For Zamboanguen\u0026tilde;o, in particular, several stages of relexification occurred to include lexical items of Philippine origin from regional Visayan, Ilonggo, and occasionally Tagalog (Lipski, 2001). Zamboanguen\u0026tilde;o has also adopted a heavy English lexical transfer (Lipski, 1992) over time.\u003c/p\u003e\n\u003cp\u003eIn 2016, the Department of Education Region IX and the Local Government of Zamboanga City published a revised Zamboanga Chavacano Orthography (DepEd- IX, 2016) to standardize the use of written Chavacano. The standardization is based on how the present generation of Zamboanguen\u0026tilde;os uses Chavacano. The orthography describes a way of spelling out Chavacano words using the alphabet of the word\u0026rsquo;s traced etymology. For example, the Spanish-derived words \u003cem\u003ezacate \u003c/em\u003e(grass) and \u003cem\u003eman\u0026tilde;ana \u003c/em\u003e(tomorrow) are spelled using the Spanish\u0026rsquo;s \u003cem\u003eabecedario\u003c/em\u003e. In contrast, the Chavacano words of local origin, like \u003cem\u003ekanila \u003c/em\u003e(them) and \u003cem\u003ekanamon \u003c/em\u003e(us), are spelled using the Philippine alphabet system. Some orthographic shifts are noted from loaned words, such as dropping the letter \u003cem\u003er \u003c/em\u003ein the Spanish verbs like \u003cem\u003ecomer \u003c/em\u003e(to eat), \u003cem\u003ebailar \u003c/em\u003e(to dance), i.e., \u003cem\u003ecome\u003c/em\u003e, \u003cem\u003ebaila\u003c/em\u003e. It is also interesting to note that the Spanish writing utilizes diacritics that are not necessarily applied in Chavacano. In general, Chavacano words are spelled the way they are pronounced.\u003c/p\u003e\n\u003ch2\u003e2.3 Chavacano Grammar\u003c/h2\u003e\n\u003cp\u003eZamboanguen\u0026tilde;o\u0026rsquo;s grammatical structure differs from any Spanish variety, and while there are standard lexicons, the two are mutually non-intelligible (Lipski, 2001).\u003c/p\u003e\n\u003cp\u003eOver three centuries of Philippine history influenced the morphology, grammar, and syntax of Zamboanguen\u0026tilde;o (Lipski \u0026amp; Santoro, 2007). Even so, it has retained its Austronesian foundation as evidenced by the Verb-Subject-Object word order, albeit many alternative possibilities (Lipski, 1992). The Philippine languages belong to the Austronesian language family. This contrasts Spanish\u0026rsquo;s Subject-Verb-Object word order (Lee, 2017). The conjugation of verbs to show tenses also does not apply in Chavacano (DepEd-IX, 2016).\u003c/p\u003e"},{"header":"3. ChavacanoMT","content":"\u003ch2\u003e3.1 Dataset Collection\u003c/h2\u003e\n\u003cp\u003eMulti-source translations of the books of the bible and articles from the Jehovah\u0026rsquo;s Witnesses website (JW.org, 2023) serve as the main sources of parallel sentences. Multi-source refers to translations of texts in multiple languages, such as bible translations, EU parliamentary proceedings, and transcripts from TED talks and United Nations meetings, where a particular sentence has translations to many target languages. Dabre et al. (2021) encourages leveraging multi-source sentences whenever available, as this helps improve the translation.\u003c/p\u003e\n\u003cp\u003eThe Bible is probably the most tapped resource for parallel texts with translations. According to Wycliffe Global Alliance (2023), the full bible has already been translated into 736 languages and the New Testament into 1,658 languages. Chavacano only has translations of the New Testament.\u003c/p\u003e\n\u003cp\u003eThe Jehovah\u0026rsquo;s Witnesses website JW.org (2023) also contains translations of their articles in multiple languages. Aside from the general articles, the website makes translations of their Watchtower and Awake! magazines available. These magazines span more than two decades of publications.\u003c/p\u003e\n\u003cp\u003eBesides multi-source texts, bilingual translations of Chavacano to the other related languages and vice versa will also be collected. This is done to augment the dataset, especially since only a few resources are available for the Philippine languages.\u003c/p\u003e\n\u003cp\u003eTranslations of the Chavacano texts to Spanish, English, Cebuano, and Hiligaynon from the Bible and articles are scraped using Webscraper.io\u003csup\u003e2\u003c/sup\u003e and Beautiful Soup\u003csup\u003e3\u003c/sup\u003e, a web-scraping library in Python.\u003c/p\u003e\n\u003ch2\u003e3.2 Preprocessing\u003c/h2\u003e\n\u003cp\u003eThe primary goal of preprocessing is to ensure the alignment of scraped texts. The alignment of bible texts is based on the verse numbers for each book chapter. The articles, however, are aligned based on the number of sentence chunks, i.e. group of sentences, captured per text object in the site\u0026rsquo;s web pages during scraping. In JW.org (2023), the web pages across languages share the same site map, and a single script was used to scrape all articles across the target languages.\u003c/p\u003e\n\u003ch3\u003e3.2.1 Articles: Cleaning and Sentence Segmentation\u003c/h3\u003e\n\u003cp\u003eThe aligned articles were preprocessed to remove unnecessary symbols and characters such as ellipses, bullets, asterisks, brackets, and non-breaking white characters. The quotation marks were also removed because some translations in other languages do not contain these. The punctuation marks, however, were retained.\u003c/p\u003e\n\u003cp\u003eThe Bible verses used as in-text references were also removed. For instance, the verse reference, -Santiago 2:14-17., was removed in the succeeding example excerpt.\u003c/p\u003e\n\u003cp\u003eTa ayuda con aquellos quien ta necesita.-Santiago 2:14-17.\u003c/p\u003e\n\u003cp\u003eThe scraped sentence chunks are segmented to extract individual sentences. Sentence segmentation was done using Sentence-Splitter\u003csup\u003e4\u003c/sup\u003e, a free module that allows splitting of text paragraphs into sentences. The heuristics algorithm, however, was based only on the English language. Other segmentation tools were considered, such as the tools from SpaCy, but the Sentence-Splitter produced more aligned chunks.\u003c/p\u003e\n\u003cp\u003eThe chunks that did not produce the same number of sentences across any or all article translations after segmentation are considered misaligned and, therefore, excluded from the final sample set.\u003c/p\u003e\n\u003ch3\u003e3.2.2 Bible: Cleaning and Verse Segmentation\u003c/h3\u003e\n\u003cp\u003eThe bible verses did not require the removal of unnecessary symbols. The quotation marks were retained together with the usual punctuation marks.\u003c/p\u003e\n\u003cp\u003eVerses combined in one bible translation but separate in other translations, for example, Matthew 17 and 18 in the English translation, but Matthew 17-18 in Hiligaynon are removed.\u003c/p\u003e\n\u003cp\u003eEach verse in the bible is considered a sentence sample, even if it spans more than one sentence or is a combination of sentence and sentence fragments. It is assumed that the translation was done per verse, and to preserve the meaning, the whole verse was used collectively as a sentence sample. The verse number, though, was removed.\u003c/p\u003e\n\u003ch2\u003e3.3 Corpus Preparation\u003c/h2\u003e\n\u003cp\u003eParallel sentences of language pairs between Chavacano, Cebuano, Hiligaynon, Tagalog, Spanish, and English comprise ChavacanoMT. These language pairs were prepared from the sentence samples collected. Table 1 summarizes each language pair\u0026rsquo;s number of sentence samples.\u003c/p\u003e\n\u003cp\u003eTable 1 breaks the type of sentences included in the language-pair datasets as multisource (MS) and non-multisource (NMS). We note that most of the sentence samples for Chavacano are multisource, meaning each sentence sample has translations to the other five languages in the corpus. We also note that the number of Chavacano sentence samples is small compared to the other language pairs.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 1\u0026nbsp;\u003c/strong\u003eNumber of Sentence Samples per Language Pair in ChavacanoMT. MS indicates Multisource, while NMS is Non-Multisource. As a reference, the language codes used are as follows: cbk for Chavacano, ceb for Cebuano, hil for Hiligaynon, tl for Tagalog, en for English, and es for Spanish.\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 26.8765%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eLanguage\u0026nbsp;Pairs\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 11.6223%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eBible\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eMS\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12.1065%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eBible\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eNMS\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eArticles\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eMS\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eArticles\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eNMS\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18.4019%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eTotal\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eSentences\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 26.8765%;\"\u003e\n \u003cp\u003ecbk-ceb\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 11.6223%;\"\u003e\n \u003cp\u003e7,728\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12.1065%;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e13,931\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e35\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18.4019%;\"\u003e\n \u003cp\u003e21,694\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 26.8765%;\"\u003e\n \u003cp\u003ecbk-hil\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 11.6223%;\"\u003e\n \u003cp\u003e7,728\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12.1065%;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e13,931\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18.4019%;\"\u003e\n \u003cp\u003e21,659\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 26.8765%;\"\u003e\n \u003cp\u003ecbk-es\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 11.6223%;\"\u003e\n \u003cp\u003e7,728\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12.1065%;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e13,931\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e35\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18.4019%;\"\u003e\n \u003cp\u003e21,694\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 26.8765%;\"\u003e\n \u003cp\u003ecbk-en\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 11.6223%;\"\u003e\n \u003cp\u003e7,728\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12.1065%;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e13,931\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e35\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18.4019%;\"\u003e\n \u003cp\u003e21,694\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 26.8765%;\"\u003e\n \u003cp\u003ecbk-tl\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 11.6223%;\"\u003e\n \u003cp\u003e7,728\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12.1065%;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e13,931\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18.4019%;\"\u003e\n \u003cp\u003e21,659\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 26.8765%;\"\u003e\n \u003cp\u003eceb-hil\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 11.6223%;\"\u003e\n \u003cp\u003e7,728\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12.1065%;\"\u003e\n \u003cp\u003e21,803\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e13,931\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e51,664\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18.4019%;\"\u003e\n \u003cp\u003e95,126\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 26.8765%;\"\u003e\n \u003cp\u003eceb-es\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 11.6223%;\"\u003e\n \u003cp\u003e7,728\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12.1065%;\"\u003e\n \u003cp\u003e21,803\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e13,931\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e51,776\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18.4019%;\"\u003e\n \u003cp\u003e95,238\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 26.8765%;\"\u003e\n \u003cp\u003eceb-en\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 11.6223%;\"\u003e\n \u003cp\u003e7,728\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12.1065%;\"\u003e\n \u003cp\u003e21,803\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e13,931\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e51,954\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18.4019%;\"\u003e\n \u003cp\u003e95,416\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 26.8765%;\"\u003e\n \u003cp\u003eceb-tl\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 11.6223%;\"\u003e\n \u003cp\u003e7,728\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12.1065%;\"\u003e\n \u003cp\u003e21,803\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e13,931\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e46,546\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18.4019%;\"\u003e\n \u003cp\u003e90,008\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 26.8765%;\"\u003e\n \u003cp\u003ehil-es\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 11.6223%;\"\u003e\n \u003cp\u003e7,728\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12.1065%;\"\u003e\n \u003cp\u003e21,803\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e13,931\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e2,865\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18.4019%;\"\u003e\n \u003cp\u003e46,327\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 26.8765%;\"\u003e\n \u003cp\u003ehil-en\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 11.6223%;\"\u003e\n \u003cp\u003e7,728\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12.1065%;\"\u003e\n \u003cp\u003e21,803\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e13,931\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e2,904\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18.4019%;\"\u003e\n \u003cp\u003e46,366\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 26.8765%;\"\u003e\n \u003cp\u003ehil-tl\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 11.6223%;\"\u003e\n \u003cp\u003e7,728\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12.1065%;\"\u003e\n \u003cp\u003e21,803\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e13,931\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e2,904\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18.4019%;\"\u003e\n \u003cp\u003e46,366\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 26.8765%;\"\u003e\n \u003cp\u003ees-en\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 11.6223%;\"\u003e\n \u003cp\u003e7,728\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12.1065%;\"\u003e\n \u003cp\u003e21,803\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e13,931\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e13,420\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18.4019%;\"\u003e\n \u003cp\u003e56,882\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 26.8765%;\"\u003e\n \u003cp\u003ees-tl\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 11.6223%;\"\u003e\n \u003cp\u003e7,728\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12.1065%;\"\u003e\n \u003cp\u003e21,803\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e13,931\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18.4019%;\"\u003e\n \u003cp\u003e43,462\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 26.8765%;\"\u003e\n \u003cp\u003een-tl\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 11.6223%;\"\u003e\n \u003cp\u003e7,728\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 12.1065%;\"\u003e\n \u003cp\u003e21,803\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e13,931\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.4964%;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 18.4019%;\"\u003e\n \u003cp\u003e43,462\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003ch2\u003e3.4 Corpus Statistics\u003c/h2\u003e\n\u003cp\u003eWe describe the corpora using the following statistics: Average Sentence Length, Number of Unique Words, and Number of Shared or Overlapping Words per Language. The statistics are taken from all sentence samples per language, and the count may differ slightly if taken from each language pair.\u003c/p\u003e\n\u003cp\u003eFigure 1 shows that the sentence samples from the Bible are longer than those from the articles. As discussed in Section 3.2.2, these sentences are verses in the bible that may comprise more than one sentence or sentence fragments.\u003c/p\u003e\n\u003cp\u003eThe number of unique words for each language in each resource is summarized in Figure 2. The number of unique words for Chavacano is arguably low compared to the other languages. We argue this was attributed to Chavacano being less morpho- logically rich than the other languages. Pahulaya (2022) presented that Chavacano comprises simple, compound, affixed, and reduplicated words like other languages. Its verbs, however, do not show inflections in tenses as the markers \u003cem\u003eya\u0026nbsp;\u003c/em\u003e(past), \u003cem\u003eta\u0026nbsp;\u003c/em\u003e(present), and \u003cem\u003eay\u0026nbsp;\u003c/em\u003e(future) are being used. In addition, the Chavacano dictionary (de Dios, Maria Isabelita Riego, 1989) registers about 6,500 entries only, including both heads and derived forms as described in SEAlang Library\u003csup\u003e5\u003c/sup\u003e. This shows that the number of words captured in the ChavacanoMT corpus is reasonable.\u003c/p\u003e\n\u003cp\u003eThe lexical similarity among languages can be measured from the number of overlapping words across languages. Figure 3 shows this similarity in the corpus.\u003c/p\u003e\n\u003cp\u003eFigure 3 shows that Chavacano samples share more word overlaps with Spanish, i.e., around 40% of the Chavacano words in the corpus are shared with Spanish. Overlaps with the other languages are also present. The Philippine languages, Cebuano, Hiligaynon, and Tagalog, share the most overlaps. The lexical similarity of Chavacano and the related languages shows the influence of these languages on Chavacano. It is also essential to consider that Spanish and English have a lexical influence on the Philippine languages due to years of colonialism.\u003c/p\u003e"},{"header":"4. Chavacano Multilingual Neural Machine Translation","content":"\u003cp\u003eIn this section, we present the utilization of the ChavacanoMT corpus in the neural machine translation of Chavacano to and from its related languages, Spanish, Cebuano, Hiligaynon, and English.\u003c/p\u003e\n\u003ch2\u003e4.1 Model Training\u003c/h2\u003e\n\u003cp\u003eWe experiment with a multilingual neural machine translation of Chavacano. It has already been established in the literature how related languages, especially those that are high-resource, support the translation of low-resource languages (Dabre,Nakagawa, \u0026amp; Kazawa, 2017; Dabre \u0026amp; Sukhoo, 2022; Goyal, Kumar, \u0026amp; Sharma, 2020; Tubay \u0026amp; Costa-Juss`a, 2018; Zoph, Yuret, May, \u0026amp; Knight, 2016). Such is the case in the neural machine translation of Chavacano. The multilingual training leverages high-resource languages in the dataset: Spanish, Cebuano, and English.\u003c/p\u003e\n\u003cp\u003eIn this experiment, we build a many-to-many machine translation model (a) from scratch using mT5 (Xue et al., 2021) model configuration and (b) from fine-tuning mT5 model using a subset of the ChavacanoMT corpus (subset described in Section 4.2).\u003c/p\u003e\n\u003cp\u003eIn training from scratch, a vocabulary of 32,300 sentence pieces was created using Sentencepiece (Kudo \u0026amp; Richardson, 2018) from combined words of languages in the corpus (Table 1). An mT5-based tokenizer, cbkTokenizer, was also trained from ChavacanoMT. The tokenizer was used to tokenize the dataset.\u003c/p\u003e\n\u003cp\u003eIn the case of fine-tuning, the vocabulary of mT5 with 250,112 sentence pieces and its built-in MT5Tokenizer were used.\u003c/p\u003e\n\u003cp\u003eThe tokenizer and models were trained using Huggingface\u0026rsquo;s Transformers\u003csup\u003e6\u003c/sup\u003e library. The model training ran in 8 epochs for both models and was optimized using AdamWeightDecay with a learning rate of 0.001. Bilingual Evaluation Understudy (BLEU) (Papineni, Roukos, Ward, \u0026amp; Zhu, 2002) scores using Sacrebleu (Post, 2018) and Recall-Oriented Understudy for Gisting Evaluation (ROUGE) (Lin, 2004) scores are collected to measure the model\u0026rsquo;s performance.\u003c/p\u003e\n\u003ch2\u003e4.2 Datasets Used\u003c/h2\u003e\n\u003cp\u003eBased on the historical investigation of Lipski (1992, 2001), the creolization of Chavacano is influenced by Spanish as its lexifier and Philippine languages, Cebuano and Hiligaynon, as adstrates. It has also undergone lexical transfer from English over time.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 2\u0026nbsp;\u003c/strong\u003eMultilingual Many-to-Many Dataset\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"2\" valign=\"top\" style=\"width: 31.3889%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eTrain\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eTest\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 23.3333%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eValidation\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 30.5556%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eTotal\u0026nbsp;Samples\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 16.6667%;\"\u003e\n \u003cp\u003ecbk-ceb\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e15,185\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e3,255\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 23.3333%;\"\u003e\n \u003cp\u003e3,254\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 30.5556%;\"\u003e\n \u003cp\u003e21,694\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 16.6667%;\"\u003e\n \u003cp\u003ecbk-en\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e15,185\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e3,255\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 23.3333%;\"\u003e\n \u003cp\u003e3,254\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 30.5556%;\"\u003e\n \u003cp\u003e21,694\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 16.6667%;\"\u003e\n \u003cp\u003ecbk-es\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e15,185\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e3,255\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 23.3333%;\"\u003e\n \u003cp\u003e3,254\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 30.5556%;\"\u003e\n \u003cp\u003e21,694\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 16.6667%;\"\u003e\n \u003cp\u003ecbk-hil\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e15,161\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e3,249\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 23.3333%;\"\u003e\n \u003cp\u003e3,249\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 30.5556%;\"\u003e\n \u003cp\u003e21,659\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 16.6667%;\"\u003e\n \u003cp\u003eceb-cbk\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e15,185\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e3,255\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 23.3333%;\"\u003e\n \u003cp\u003e3,254\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 30.5556%;\"\u003e\n \u003cp\u003e21,694\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 16.6667%;\"\u003e\n \u003cp\u003eceb-en\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e66,791\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e14,313\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 23.3333%;\"\u003e\n \u003cp\u003e14,312\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 30.5556%;\"\u003e\n \u003cp\u003e95,416\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 16.6667%;\"\u003e\n \u003cp\u003eceb-es\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e66,666\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e14,286\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 23.3333%;\"\u003e\n \u003cp\u003e14,286\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 30.5556%;\"\u003e\n \u003cp\u003e95,238\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 16.6667%;\"\u003e\n \u003cp\u003eceb-hil\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e66,588\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e14,269\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 23.3333%;\"\u003e\n \u003cp\u003e14,269\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 30.5556%;\"\u003e\n \u003cp\u003e95,126\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 16.6667%;\"\u003e\n \u003cp\u003een-cbk\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e15,185\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e3,255\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 23.3333%;\"\u003e\n \u003cp\u003e3,254\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 30.5556%;\"\u003e\n \u003cp\u003e21,694\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 16.6667%;\"\u003e\n \u003cp\u003een-ceb\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e66,791\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e14,313\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 23.3333%;\"\u003e\n \u003cp\u003e14,312\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 30.5556%;\"\u003e\n \u003cp\u003e95,416\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 16.6667%;\"\u003e\n \u003cp\u003een-es\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e39,817\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e8,533\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 23.3333%;\"\u003e\n \u003cp\u003e8,532\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 30.5556%;\"\u003e\n \u003cp\u003e56,882\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 16.6667%;\"\u003e\n \u003cp\u003een-hil\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e32,448\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e6,954\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 23.3333%;\"\u003e\n \u003cp\u003e6,953\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 30.5556%;\"\u003e\n \u003cp\u003e46,355\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 16.6667%;\"\u003e\n \u003cp\u003ees-cbk\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e15,185\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e3,255\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 23.3333%;\"\u003e\n \u003cp\u003e3,254\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 30.5556%;\"\u003e\n \u003cp\u003e21,694\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 16.6667%;\"\u003e\n \u003cp\u003ees-ceb\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e66,666\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e14,286\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 23.3333%;\"\u003e\n \u003cp\u003e14,286\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 30.5556%;\"\u003e\n \u003cp\u003e95,238\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 16.6667%;\"\u003e\n \u003cp\u003ees-en\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e39,817\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e8,533\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 23.3333%;\"\u003e\n \u003cp\u003e8,532\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 30.5556%;\"\u003e\n \u003cp\u003e56,882\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 16.6667%;\"\u003e\n \u003cp\u003ees-hil\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e32,428\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e6,950\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 23.3333%;\"\u003e\n \u003cp\u003e6,949\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 30.5556%;\"\u003e\n \u003cp\u003e46,327\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 16.6667%;\"\u003e\n \u003cp\u003ehil-cbk\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e15,161\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e3,249\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 23.3333%;\"\u003e\n \u003cp\u003e3,249\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 30.5556%;\"\u003e\n \u003cp\u003e21,659\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 16.6667%;\"\u003e\n \u003cp\u003ehil-ceb\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e66,588\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e14,269\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 23.3333%;\"\u003e\n \u003cp\u003e14,269\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 30.5556%;\"\u003e\n \u003cp\u003e95,126\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 16.6667%;\"\u003e\n \u003cp\u003ehil-en\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e32,434\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e6,951\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 23.3333%;\"\u003e\n \u003cp\u003e6,950\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 30.5556%;\"\u003e\n \u003cp\u003e46,335\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 16.6667%;\"\u003e\n \u003cp\u003ehil-es\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e32,428\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 14.7222%;\"\u003e\n \u003cp\u003e6,950\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 23.3333%;\"\u003e\n \u003cp\u003e6,949\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 30.5556%;\"\u003e\n \u003cp\u003e46,327\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eThis account became the basis of the multilingual dataset for Chavacano translation. The language pairs involving Chavacano, Spanish, Cebuano, Hiligaynon, and English were used except for Tagalog.\u003c/p\u003e\n\u003cp\u003eIn this experiment\u0026rsquo;s many-to-many training set-up, the language pairs are used in both directions, that is, cbk-en and en-cbk are used. In total, 20 language pairs between Chavacano, Spanish, Cebuano, Hiligaynon, and English were included in the dataset. In this training set-up, all languages serve as source and target languages. It is hoped that by doing so, the similarities shared between Chavacano and the high-resource languages may be represented in both the encoder and decoder sides of the neural network. The summary of the dataset used in this experiment is shown in Table 2.\u003c/p\u003e\n\u003cp\u003e70% of the dataset was used as a training set, 15% as a validation set, and 15% as a test set. The data splits are stratified according to language pairs. The dataset is unbalanced, with more training samples from the high-resource Cebuano, English, and Spanish language pairs.\u003c/p\u003e"},{"header":"5 Results","content":"\u003cp\u003eTable 3 shows the BLEU and ROUGE-1 scores of the models generated in the experiment. As expected, the fine-tuned model earned better results than the model generated from scratch. This demonstrates the leverage one can get on pre-trained weights even if target languages are not included in the pre-training.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 3\u0026nbsp;\u003c/strong\u003eComparison of BLEU and ROUGE-1 scores for (a) mT5-based model from scratch and (b) fine-tuned mT5 model.\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 39.3574%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eModel\u0026nbsp;Type\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 26.506%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eBLEU\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 34.1365%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eROUGE-\u003c/strong\u003e\u003cstrong\u003e1\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 39.3574%;\"\u003e\n \u003cp\u003eScratch\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 26.506%;\"\u003e\n \u003cp\u003e0.4879\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 34.1365%;\"\u003e\n \u003cp\u003e0.1222\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 39.3574%;\"\u003e\n \u003cp\u003eFinetuned\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 26.506%;\"\u003e\n \u003cp\u003e\u003cstrong\u003e17.8274\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 34.1365%;\"\u003e\n \u003cp\u003e\u003cstrong\u003e0.5457\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eThe fine-tuned model was tested further using the individual language pairs in the test set. Table 4 presents the performance scores for each translation direction.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 4\u0026nbsp;\u003c/strong\u003ePerformance results for fine-tuned mT5 model. Table (b) on the left shows the translation direction with better performance.\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"2\" valign=\"top\" style=\"width: 28.7105%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eBLEU\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 20.6813%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eROUGE-\u003c/strong\u003e\u003cstrong\u003e1\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd colspan=\"2\" valign=\"top\" style=\"width: 28.7105%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eBLEU\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21.8978%;\"\u003e\n \u003cp\u003e\u003cstrong\u003eROUGE-\u003c/strong\u003e\u003cstrong\u003e1\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 15.3285%;\"\u003e\n \u003cp\u003ecbk-ceb\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.382%;\"\u003e\n \u003cp\u003e21.95\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 20.6813%;\"\u003e\n \u003cp\u003e0.62\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.3285%;\"\u003e\n \u003cp\u003eceb-cbk\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.382%;\"\u003e\n \u003cp\u003e23.54\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21.8978%;\"\u003e\n \u003cp\u003e0.66\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 15.3285%;\"\u003e\n \u003cp\u003ehil-ceb\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.382%;\"\u003e\n \u003cp\u003e22.34\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 20.6813%;\"\u003e\n \u003cp\u003e0.60\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.3285%;\"\u003e\n \u003cp\u003eceb-hil\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.382%;\"\u003e\n \u003cp\u003e22.79\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21.8978%;\"\u003e\n \u003cp\u003e0.61\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 15.3285%;\"\u003e\n \u003cp\u003een-ceb\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.382%;\"\u003e\n \u003cp\u003e20.02\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 20.6813%;\"\u003e\n \u003cp\u003e0.59\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.3285%;\"\u003e\n \u003cp\u003eceb-en\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.382%;\"\u003e\n \u003cp\u003e26.25\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21.8978%;\"\u003e\n \u003cp\u003e0.61\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 15.3285%;\"\u003e\n \u003cp\u003een-cbk\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.382%;\"\u003e\n \u003cp\u003e24.06\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 20.6813%;\"\u003e\n \u003cp\u003e0.68\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.3285%;\"\u003e\n \u003cp\u003ecbk-en\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.382%;\"\u003e\n \u003cp\u003e35.30\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21.8978%;\"\u003e\n \u003cp\u003e0.69\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 15.3285%;\"\u003e\n \u003cp\u003eceb-es\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.382%;\"\u003e\n \u003cp\u003e16.61\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 20.6813%;\"\u003e\n \u003cp\u003e0.50\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.3285%;\"\u003e\n \u003cp\u003ees-ceb\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.382%;\"\u003e\n \u003cp\u003e16.34\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21.8978%;\"\u003e\n \u003cp\u003e0.53\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 15.3285%;\"\u003e\n \u003cp\u003ehil-cbk\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.382%;\"\u003e\n \u003cp\u003e22.10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 20.6813%;\"\u003e\n \u003cp\u003e0.65\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.3285%;\"\u003e\n \u003cp\u003ecbk-hil\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.382%;\"\u003e\n \u003cp\u003e23.51\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21.8978%;\"\u003e\n \u003cp\u003e0.64\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 15.3285%;\"\u003e\n \u003cp\u003een-hil\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.382%;\"\u003e\n \u003cp\u003e16.20\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 20.6813%;\"\u003e\n \u003cp\u003e0.57\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.3285%;\"\u003e\n \u003cp\u003ehil-en\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.382%;\"\u003e\n \u003cp\u003e23.70\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21.8978%;\"\u003e\n \u003cp\u003e0.59\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 15.3285%;\"\u003e\n \u003cp\u003ees-hil\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.382%;\"\u003e\n \u003cp\u003e15.13\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 20.6813%;\"\u003e\n \u003cp\u003e0.55\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.3285%;\"\u003e\n \u003cp\u003ehil-es\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.382%;\"\u003e\n \u003cp\u003e16.63\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21.8978%;\"\u003e\n \u003cp\u003e0.52\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 15.3285%;\"\u003e\n \u003cp\u003ees-cbk\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.382%;\"\u003e\n \u003cp\u003e21.79\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 20.6813%;\"\u003e\n \u003cp\u003e0.64\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.3285%;\"\u003e\n \u003cp\u003ecbk-es\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.382%;\"\u003e\n \u003cp\u003e24.68\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21.8978%;\"\u003e\n \u003cp\u003e0.60\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 15.3285%;\"\u003e\n \u003cp\u003een-es\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.382%;\"\u003e\n \u003cp\u003e19.45\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 20.6813%;\"\u003e\n \u003cp\u003e0.54\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 15.3285%;\"\u003e\n \u003cp\u003ees-en\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 13.382%;\"\u003e\n \u003cp\u003e25.94\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 21.8978%;\"\u003e\n \u003cp\u003e0.60\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eThe results show that the translation from Philippine to foreign language seems better than the other way around, except between Cebuano and Spanish, where the scores are essentially the same in both directions. It is also interesting to note that the highest-performing language pair at 35.30 BLEU is Chavacano-English, whose training samples are among the lowest in the dataset.\u003c/p\u003e\n\u003cp\u003eAs a way of benchmarking, we also compare the Ph-EN and EN-Ph BLEU scores obtained from the experiments of Coronia (2022) with the result of our experiments. Ph here refers to Philippine languages. Table 5 summarizes the best-performing model from Coronia (2022) and the BLEU scores from our experiments.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 5\u0026nbsp;\u003c/strong\u003eComparison of bilingual translations from finetuned mT5 multilingual models. Coronia (2022) uses the PH-MNMT dataset, while ours uses the ChavacanoMT corpus.\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"2\" valign=\"top\" style=\"width: 225px;\"\u003e\n \u003cp\u003een-ceb\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 53px;\"\u003e\n \u003cp\u003eceb-en\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 48px;\"\u003e\n \u003cp\u003een-hil\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 51px;\"\u003e\n \u003cp\u003ehil-en\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 54px;\"\u003e\n \u003cp\u003een-cbk\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 54px;\"\u003e\n \u003cp\u003ecbk-en\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 172px;\"\u003e\n \u003cp\u003eFine-tuned\u0026nbsp;mT5\u0026nbsp;(Coronia)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 52px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e21.25\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 53px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e26.67\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 48px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e24.4\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 51px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e26.11\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 54px;\"\u003e\n \u003cp\u003e0.08\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 54px;\"\u003e\n \u003cp\u003e0.79\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd valign=\"top\" style=\"width: 172px;\"\u003e\n \u003cp\u003eFine-tuned\u0026nbsp;mT5\u0026nbsp;(Ours)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 52px;\"\u003e\n \u003cp\u003e20.02\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 53px;\"\u003e\n \u003cp\u003e26.25\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 48px;\"\u003e\n \u003cp\u003e16.20\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 51px;\"\u003e\n \u003cp\u003e23.70\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 54px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e24.06\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 54px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e35.30\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eFor context, the PH-MNMT dataset used in Coronia (2022) comprises millions of English-Cebuano and English-Hiligaynon sentence samples compared to ChavacanoMT. ChavacanoMT, however, has more sentence samples for Chavacano. In Coronia (2022), the best-performing mT5 model was fine-tuned on EN-Ph and Ph-EN sentences comprising English, Tagalog, Cebuano, Hiligaynon, Waray, and Chavacano. In contrast, our experiment generated the translation model using a many-to-many training setup involving languages that are specifically related to Chavacano. Both Coronia (2022) and our translation experiment considered using related languages in model training. Although an objective comparison between the models cannot be made, we can infer that ChavacanoMT can produce models that are on par with published results.\u003c/p\u003e"},{"header":"6\tConclusions","content":"\u003cp\u003eThis\u0026nbsp;paper\u0026nbsp;presents\u0026nbsp;ChavacanoMT,\u0026nbsp;a\u0026nbsp;benchmark\u0026nbsp;corpus\u0026nbsp;for\u0026nbsp;the\u0026nbsp;machine\u0026nbsp;translation of Chavacano to and from related languages, Spanish, Cebuano, Hiligaynon, Tagalog, and English. The corpus consists of 767,053 parallel sentence samples from 15 language\u0026nbsp;pairs.\u0026nbsp;The\u0026nbsp;corpus\u0026nbsp;is\u0026nbsp;a\u0026nbsp;combination\u0026nbsp;of\u0026nbsp;multi-source\u0026nbsp;and\u0026nbsp;parallel\u0026nbsp;sentences.\u003c/p\u003e\n\u003cp\u003eUsing the ChavacanoMT corpus in our machine translation experiments has demonstrated its potential to significantly enhance the translation quality between Chavacano and its related languages. Our experiments showed that models trained\u0026nbsp;using ChavacanoMT could achieve performance on par with or surpass existing multilingual neural machine translation systems involving Chavacano, particularly in translating between Chavacano and English, with improvements exceeding 20 BLEU points.\u0026nbsp;This\u0026nbsp;highlights\u0026nbsp;the\u0026nbsp;effectiveness\u0026nbsp;of\u0026nbsp;incorporating\u0026nbsp;diverse\u0026nbsp;related\u0026nbsp;languages with multisource sentence samples in building robust machine translation models for low-resource languages in Chavacano. Our findings suggest that leveraging related languages within the corpus improves translation accuracy. This approach underscores\u0026nbsp;the\u0026nbsp;importance\u0026nbsp;of\u0026nbsp;using\u0026nbsp;carefully\u0026nbsp;curated\u0026nbsp;multilingual\u0026nbsp;datasets\u0026nbsp;to\u0026nbsp;support underrepresented languages, ultimately contributing to their preservation and wider accessibility.\u003c/p\u003e\n\u003cp\u003eThe insights gained from this study open avenues for exploring the influence of specific linguistic relationships in translation quality, which could guide the development\u0026nbsp;of\u0026nbsp;more\u0026nbsp;targeted\u0026nbsp;translation\u0026nbsp;strategies\u0026nbsp;for\u0026nbsp;other\u0026nbsp;low-resource\u0026nbsp;languages.\u003c/p\u003e\n\u003cp\u003eAdditionally, ChavacanoMT provides a foundation for further research on cross- linguistic interactions and their implications for the evolution and revitalization of Chavacano. By extending this work, researchers can deepen their understanding of how digital tools and computational methods can support Creole-speaking communities’ linguistic and cultural heritage.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAcknowledgements.\u0026nbsp;\u003c/strong\u003eWe acknowledge the Jehovah\u0026rsquo;s Witness organization for the permission to scrape their website www.jw.org to support the NLP research on Philippine languages.\u003c/p\u003e\n\u003cul\u003e\n \u003cli\u003eFunding: This work was supported by the Commission on Higher Education through its\u0026nbsp;Scholarships\u0026nbsp;for\u0026nbsp;Instructors\u0026rsquo;\u0026nbsp;Knowledge\u0026nbsp;Advancement\u0026nbsp;Program\u0026nbsp;(SIKAP)\u0026nbsp;grant.\u003c/li\u003e\n \u003cli\u003eConflict\u0026nbsp;of\u0026nbsp;interest/Competing\u0026nbsp;interests:\u0026nbsp;The\u0026nbsp;authors\u0026nbsp;have\u0026nbsp;no\u0026nbsp;competing\u0026nbsp;interests to declare relevant to this article\u0026rsquo;s content.\u003c/li\u003e\n \u003cli\u003eData availability: The corpus built in this study is available from the authors, but restrictions apply. Some resources used in building the corpus were under a non- commercial agreement from Watch Tower Bible and Tract Society, Philippines for the current study. The authors wish to honor the permission by ensuring that out- comes are used only for academic purposes. Data are, however, available from the authors upon reasonable request.\u003c/li\u003e\n\u003c/ul\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eAhia, O., \u0026amp; Ogueji, K. (2020). Towards Supervised and Unsupervised Neural Machine Translation Baselines for Nigerian Pidgin. \u003cem\u003eAfricaNLP Workshop. \u003c/em\u003eOnline. Retrieved from http://arxiv.org/abs/2003.12660\u003c/li\u003e\n\u003cli\u003eBaliber, R.I., Cheng, C., Adlaon, K.M., Mamonong, V. (2020, December). Bridging Philippine Languages with Multilingual Neural Machine Translation. \u003cem\u003eProceedings of the 3rd Workshop on Technologies for MT of Low Resource Languages \u003c/em\u003e(pp. 14\u0026ndash;22). Association for Computational Linguistics. Retrieved from https://aclanthology.org/2020.loresmt-1.2.pdf\u003c/li\u003e\n\u003cli\u003eCoronia, J.D. (2022). \u003cem\u003eExploring clustering of Philippine languages in multilingual neural machine translation. \u003c/em\u003e(Unpublished master\u0026rsquo;s thesis). De La Salle University, Manila, Philippines. (Retrieved from https://animorepository.dlsu.edu.ph/etdm softtech/4)\u003c/li\u003e\n\u003cli\u003eDabre, R., Chu, C., Kunchukuttan, A. (2021, September). A Survey of Multilingual Neural Machine Translation. \u003cem\u003eACM Computing Surveys\u003c/em\u003e, \u003cem\u003e53 \u003c/em\u003e(5), 1\u0026ndash;38, https://doi.org/10.1145/3406095 Retrieved 2023-09-12, from https://dl.acm.org/doi/10.1145/3406095\u003c/li\u003e\n\u003cli\u003eDabre, R., Nakagawa, T., Kazawa, H. (2017, November). An Empirical Study of Language Relatedness for Transfer Learning in Neural Machine Translation. \u003cem\u003eProceedings of the 31st Pacific Asia Conference on Language, Information and Computation \u003c/em\u003e(pp. 282\u0026ndash;286). The National University (Philippines). Retrieved from https://aclanthology.org/Y17-1038\u003c/li\u003e\n\u003cli\u003eDabre, R., \u0026amp; Sukhoo, A. (2022, November). KreolMorisienMT: A dataset for Mauritian Creole machine translation. \u003cem\u003eFindings of the association for computational linguistics: Aacl-ijcnlp 2022 \u003c/em\u003e(pp. 22\u0026ndash;29). Online only: Association for Computational Linguistics. Retrieved from https://aclanthology.org/2022.findings-aacl.3\u003c/li\u003e\n\u003cli\u003ede Dios, Maria Isabelita Riego (1989). A composite dictionary of Philippine Creole Spanish. \u003cem\u003eStudies in Philippine Linguistics. \u003c/em\u003eManila: Linguistic Society of the Philippines and Summer Institute of Linguistics.\u003c/li\u003e\n\u003cli\u003eDepEd-IX (2016). \u003cem\u003eZamboanga Chavacano Orthography\u003c/em\u003e. Local Government of Zamboanga City: Philippines.\u003c/li\u003e\n\u003cli\u003eEberhard, D., Simons, G., \u0026amp; Fenning, C. (Eds.). (2024). \u003cem\u003eEthnologue: Languages of the World \u003c/em\u003e(27th ed.). Dallas, Texas: SIL International.\u003c/li\u003e\n\u003cli\u003eGenuino, C.F. (2005). \u003cem\u003eLanguage extinction in process across Chabacano communities: A sociolinguistic approach. \u003c/em\u003e(Unpublished doctoral disserta- tion). De La Salle University, Manila, Philippines. (Retrieved from https://animorepository.dlsu.edu.ph/etd doctoral/87)\u003c/li\u003e\n\u003cli\u003eGoyal, V., Kumar, S., Sharma, D.M. (2020). Efficient Neural Machine Translation for Low-Resource Languages via Exploiting Related Languages. \u003cem\u003eProceedings of the 58th Annual Meeting of the Association for Computational Linguistics: Student Research Workshop \u003c/em\u003e(p. 162\u0026ndash;168). Association for Computational Linguistics.\u003c/li\u003e\n\u003cli\u003eJW.org (2023). \u003cem\u003eOfficial Website of Jehovah\u0026rsquo;s Witnesses. \u003c/em\u003eAvailable at http://www.https://www.jw.org/en/ (2023/10/18).\u003c/li\u003e\n\u003cli\u003eKomisyon sa Wikang Filipino (2020). \u003cem\u003eRepositoryo ng mga Wika at Kultura. \u003c/em\u003ehttps:// kwfwikaatkultura.ph/chabacano/.\u003c/li\u003e\n\u003cli\u003eKudo, T., \u0026amp; Richardson, J. (2018, November). SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing.\u003c/li\u003e\n\u003cli\u003eE. Blanco \u0026amp; W. Lu (Eds.), \u003cem\u003eProceedings of the 2018 conference on empirical methods in natural language processing: System demonstrations \u003c/em\u003e(pp. 66\u0026ndash;71). Brussels, Belgium: Association for Computational Linguistics. Retrieved from https://aclanthology.org/D18-2012\u003c/li\u003e\n\u003cli\u003eLee, J.F. (2017). Word order and linguistic factors in the second language processing of spanish passive sentences. \u003cem\u003eHispania\u003c/em\u003e, \u003cem\u003e100 \u003c/em\u003e(4), 580\u0026ndash;595, Retrieved 2023-06-26, from https://www.jstor.org/stable/26387810\u003c/li\u003e\n\u003cli\u003eLesho, M., \u0026amp; Sippola, E. (2013). The sociolinguistic situations of Manila Bay Chabacano-speaking communities. \u003cem\u003eLanguage Documentation and Conservation\u003c/em\u003e, \u003cem\u003e7 \u003c/em\u003e, 1\u0026ndash;30,\u003c/li\u003e\n\u003cli\u003eLin, C. (2004, July). ROUGE: A package for automatic evaluation of summaries. \u003cem\u003eText Summarization Branches Out \u003c/em\u003e(pp. 74\u0026ndash;81). Barcelona, Spain: Association for Computational Linguistics. Retrieved from https://www.aclweb.org/anthology/W04-1013\u003c/li\u003e\n\u003cli\u003eLipski, J. (1992). New thoughts on the origins of Zamboanguen\u0026tilde;o (Philippine Creole Spanish). \u003cem\u003eLanguage Sciences\u003c/em\u003e, \u003cem\u003e14 \u003c/em\u003e(3), 197-231, https:// doi.org/https://doi.org/10.1016/0388-0001(92)90005-Y Retrieved from https://www.sciencedirect.com/science/article/pii/038800019290005Y\u003c/li\u003e\n\u003cli\u003eLipski, J. (2001, Aug.). The place of Chabacano in the Philippine linguistic profile. \u003cem\u003eSociolinguistic Studies\u003c/em\u003e, \u003cem\u003e2 \u003c/em\u003e(2), 119\u0026ndash;163, https://doi.org/10.1558/sols.v2i2.119 Retrieved from https://journal.equinoxpub.com/SS/article/view/11691\u003c/li\u003e\n\u003cli\u003eLipski, J., \u0026amp; Santoro, M. (2007). Zamboanguen\u0026tilde;o Creole Spanish [Bibliographical record]. J. Holm \u0026amp; P. Patrick (Eds.), \u003cem\u003eComparative creole syntax. parallel outlines of 18 creole grammars \u003c/em\u003e(p. 373-398). London: Battlebridge. (Much information is based on Forman (1972).)\u003c/li\u003e\n\u003cli\u003ePahulaya, V.L. (2022). Morphological Analysis on the Structure of Chavacano Language: A Complex Mental Process. \u003cem\u003eNeuroQuantology \u003c/em\u003e, \u003cem\u003e20 \u003c/em\u003e(6), 9820-9830, https://doi.org/https://doi.org/10.14704/nq.2022.20.6.NQ22960\u003c/li\u003e\n\u003cli\u003ePapineni, K., Roukos, S., Ward, T., Zhu, W. (2002). BLEU: a Method for Automatic Evaluation of Machine Translation. \u003cem\u003eProceedings of the 40th Annual Meeting of the Association for Computational Linguistics (ACL) \u003c/em\u003e(pp. 311\u0026ndash;318). Philadelphia.\u003c/li\u003e\n\u003cli\u003ePost, M. (2018, October). A Call for Clarity in Reporting BLEU Scores. \u003cem\u003eProceedings of the Third Conference on Machine Translation: Research Papers \u003c/em\u003e(pp. 186\u0026ndash; 191). Belgium, Brussels: Association for Computational Linguistics. Retrieved from https://www.aclweb.org/anthology/W18-6319\u003c/li\u003e\n\u003cli\u003eRobinson, N., Hogan, C., Fulda, N., Mortensen, D.R. (2022, October). Data-adaptive Transfer Learning for Translation: A Case Study in Haitian and Jamaican. \u003cem\u003eProceedings of the Fifth Workshop on Technologies for Machine Translation of Low-Resource Languages (LoResMT 2022) \u003c/em\u003e(pp. 35\u0026ndash;42). Gyeongju, Republic of Korea: Association for Computational Linguistics. Retrieved from https://aclanthology.org/2022.loresmt-1.5\u003c/li\u003e\n\u003cli\u003eTubay, B., \u0026amp; Costa-Juss`a, M.R. (2018, October). Neural machine translation with the transformer and multi-source Romance languages for the biomedical WMT 2018 task. \u003cem\u003eProceedings of the Third Conference on Machine Translation: Shared Task Papers \u003c/em\u003e(pp. 667\u0026ndash;670). Belgium, Brussels: Association for Computational Linguistics. Retrieved from https://aclanthology.org/W18-6449\u003c/li\u003e\n\u003cli\u003eWycliffe Global Alliance (2023). \u003cem\u003e2023 Global Scripture Access. \u003c/em\u003ehttps://www.wycliffe.net/resources/statistics/. (Accessed: February 9, 2024)\u003c/li\u003e\n\u003cli\u003eXue, L., Constant, N., Roberts, A., Kale, M., Al-Rfou, R., Siddhant, A., . . . Raf- fel, C. (2021, June). mT5: A massively multilingual pre-trained text-to-text transformer. K. Toutanova et al. (Eds.), \u003cem\u003eProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \u003c/em\u003e(pp. 483\u0026ndash;498). Online: Association for Computational Linguistics. Retrieved from https://aclanthology.org/2021.naacl-main.41\u003c/li\u003e\n\u003cli\u003eZoph, B., Yuret, D., May, J., Knight, K. (2016). A Transfer learning for low-resource neural machine translation. \u003cem\u003eProceedings of the Conference on Empirical Methods in Natural Language Processing \u003c/em\u003e(p. 1568\u0026ndash;1575). Association for Computational Linguistics.\u003c/li\u003e\n\u003c/ol\u003e"},{"header":"Footnotes","content":"\u003cp\u003e\u003csup\u003e1\u003c/sup\u003ehttps://opus.nlpl.eu/\u003c/p\u003e\n\u003cp\u003e\u003csup\u003e2\u003c/sup\u003ehttps://webscraper.io/\u003c/p\u003e\n\u003cp\u003e\u003csup\u003e3\u003c/sup\u003ehttps://www.crummy.com/software/BeautifulSoup/\u003c/p\u003e\n\u003cp\u003e\u003csup\u003e4\u003c/sup\u003ehttps://github.com/mediacloud/sentence-splitter?tab=readme-ov-file\u003c/p\u003e\n\u003cp\u003e\u003csup\u003e5\u003c/sup\u003ehttp://sealang.net/chavacano/dictionary.htm\u003c/p\u003e\n\u003cp\u003e\u003csup\u003e6\u003c/sup\u003ehttps://huggingface.co/docs/transformers/en/index\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"language-resources-and-evaluation","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"lrev","sideBox":"Learn more about [Language Resources and Evaluation](http://link.springer.com/journal/10579)","snPcode":"10579","submissionUrl":"https://submission.nature.com/new-submission/10579/3","title":"Language Resources and Evaluation","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Philippine Creole Spanish, Chavacano Translation Corpus, Multilingual Translation, Chavacano","lastPublishedDoi":"10.21203/rs.3.rs-5022127/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-5022127/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eChavacano, formally referred to as Philippine Creole Spanish, is the only Creole spoken in the Philippines. Like many languages, especially Creoles, computational studies on Chavacano are scarce because of the dearth of available corpora. This paper describes the creation of ChavacanoMT, a benchmark corpus for the machine translation study of Philippine Creole Spanish. ChavacanoMT consists of 767,053 parallel sentences between Chavacano and related languages, Spanish, Cebuano, Hiligaynon, Tagalog, and English. It is sourced from scraped bible translations and articles on the Jehovah’s Witness website.\u003c/p\u003e\n\u003cp\u003eThis paper also presents the performance of a multilingual neural machine translation model generated using ChavacanoMT. We report an overall 17 BLEU score on a fine-tuned mT5 model, outperforming an mT5-based model trained from scratch. Our experiments show that ChavacanoMT can generate models on par with a similar system that translates between English and some Philippines languages despite having fewer sentence samples used in training. We also report an improved Chavacano translation to and from its related languages that can be used as benchmark data. In particular, we highlight more than 20 BLEU points of improvement in the translation between Chavacano and English.\u003c/p\u003e\n\u003cp\u003eThe study opens avenues for exploring cross-linguistic interactions of Chavacano and its related languages in its translation that may benefit other low-resource languages.\u003c/p\u003e","manuscriptTitle":"ChavacanoMT: A Corpus and Evaluation of Neural Machine Translation for Philippine Creole Spanish","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-10-10 12:12:48","doi":"10.21203/rs.3.rs-5022127/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-05-12T09:41:10+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-05-10T06:41:12+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"165146865356313760371589427827445874496","date":"2025-02-03T11:16:06+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-01-02T07:12:41+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"214487281406048882311836871996706926899","date":"2024-11-08T09:56:07+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"136706945448887160858013148935588041055","date":"2024-11-06T13:00:43+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2024-11-04T12:47:37+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2024-09-08T14:52:37+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2024-09-03T06:03:28+00:00","index":"","fulltext":""},{"type":"submitted","content":"Language Resources and Evaluation","date":"2024-09-03T05:43:00+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"language-resources-and-evaluation","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"lrev","sideBox":"Learn more about [Language Resources and Evaluation](http://link.springer.com/journal/10579)","snPcode":"10579","submissionUrl":"https://submission.nature.com/new-submission/10579/3","title":"Language Resources and Evaluation","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"084d1e18-117a-4cbd-b3e8-9186b8dcd159","owner":[],"postedDate":"October 10th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2026-01-19T16:45:42+00:00","versionOfRecord":{"articleIdentity":"rs-5022127","link":"https://doi.org/10.1007/s10579-025-09888-3","journal":{"identity":"language-resources-and-evaluation","isVorOnly":false,"title":"Language Resources and Evaluation"},"publishedOn":"2026-01-13 16:28:47","publishedOnDateReadable":"January 13th, 2026"},"versionCreatedAt":"2024-10-10 12:12:48","video":"","vorDoi":"10.1007/s10579-025-09888-3","vorDoiUrl":"https://doi.org/10.1007/s10579-025-09888-3","workflowStages":[]},"version":"v1","identity":"rs-5022127","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-5022127","identity":"rs-5022127","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.