Enabling Machine Translation Technique for Low resource Digaru language to English

Abstract

This study presents the development of a Neural Machine Translation (NMT) system for Digaru, a critically endangered and low-resource language spoken in Northeast India. In the absence of pre-existing digital linguistic resources, a parallel corpus comprising approximately 90,000 English Digaru sentence pairs was constructed. The data were sourced primarily from the Tatoeba corpus and supplemented with manually curated English sentences translated into Digaru by fluent native speakers. The resulting corpus was thoroughly cleaned and pre-processed to ensure structural consistency and compatibility with contemporary Machine Translation (MT) architectures. To assess translation performance under low-resource conditions, several MT models were implemented: Phrase-Based Statistical Machine Translation (PBSMT), a Long Short-Term Memory (LSTM)-based Sequence-to-Sequence (Seq2Seq) model, and the Transformer architecture within the NMT framework. Evaluation was conducted using automatic metrics Bilingual Evaluation Understudy (BLEU), Translation Edit Rate (TER), and Metric for Evaluation of Translation with Explicit ORdering (METEOR) alongside manual assessments of adequacy and fluency. Results showed that the Transformer-based NMT system significantly outperformed the PBSMT baseline, with BLEU score improvements ranging from 13.8 to 15.2 points. The optimized Transformer model achieved a peak score of 48.6 BLEU on a 50,000-sentence training set, while character-level tokenization led to additional gains, reaching 24.9 BLEU with +5.7 over the base configuration. To address the challenge of data scarcity, Backtranslation (BT) was employed to generate synthetic Digaru English sentence pairs. The addition of 20,000 synthetic sentences to the base dataset led to the most significant improvement, increasing the score from 45.5 to 56.9 percent. Similar gains were also observed in Digaru German and German Digaru translation directions. In summary, this research contributes the first digital parallel corpus and NMT models.

Description

Keywords

Citation

item.page.endorsement

item.page.review

item.page.supplemented

item.page.referenced