Enabling Machine Translation Technique for Low resource Digaru language to English

dc.contributor.guideSambyo, Koj
dc.coverage.spatialArtificial Intelligence
dc.creator.researcherKri, Rushanti
dc.date.accessioned2026-01-29T05:53:35Z
dc.date.available2026-01-29T05:53:35Z
dc.date.awarded2025
dc.date.completed2025
dc.date.registered2020
dc.description.abstractThis study presents the development of a Neural Machine Translation (NMT) system for Digaru, a critically endangered and low-resource language spoken in Northeast India. In the absence of pre-existing digital linguistic resources, a parallel corpus comprising approximately 90,000 English Digaru sentence pairs was constructed. The data were sourced primarily from the Tatoeba corpus and supplemented with manually curated English sentences translated into Digaru by fluent native speakers. The resulting corpus was thoroughly cleaned and pre-processed to ensure structural consistency and compatibility with contemporary Machine Translation (MT) architectures. To assess translation performance under low-resource conditions, several MT models were implemented: Phrase-Based Statistical Machine Translation (PBSMT), a Long Short-Term Memory (LSTM)-based Sequence-to-Sequence (Seq2Seq) model, and the Transformer architecture within the NMT framework. Evaluation was conducted using automatic metrics Bilingual Evaluation Understudy (BLEU), Translation Edit Rate (TER), and Metric for Evaluation of Translation with Explicit ORdering (METEOR) alongside manual assessments of adequacy and fluency. Results showed that the Transformer-based NMT system significantly outperformed the PBSMT baseline, with BLEU score improvements ranging from 13.8 to 15.2 points. The optimized Transformer model achieved a peak score of 48.6 BLEU on a 50,000-sentence training set, while character-level tokenization led to additional gains, reaching 24.9 BLEU with +5.7 over the base configuration. To address the challenge of data scarcity, Backtranslation (BT) was employed to generate synthetic Digaru English sentence pairs. The addition of 20,000 synthetic sentences to the base dataset led to the most significant improvement, increasing the score from 45.5 to 56.9 percent. Similar gains were also observed in Digaru German and German Digaru translation directions. In summary, this research contributes the first digital parallel corpus and NMT models.
dc.description.note
dc.format.accompanyingmaterialNone
dc.format.dimensions30cm
dc.format.extentxix, 111
dc.identifier.researcherid0000-0001-7701-6770
dc.identifier.urihttp://hdl.handle.net/10603/690692
dc.languageEnglish
dc.publisher.institutionDepartment of Computer Science and Engineering
dc.publisher.placeJote
dc.publisher.universityNational Institute of Technology Arunachal Pradesh
dc.relation68
dc.rightsuniversity
dc.source.universityUniversity
dc.subject.keywordLow-resources
dc.subject.keywordNMT SMT
dc.subject.keywordTransformer Model
dc.titleEnabling Machine Translation Technique for Low resource Digaru language to English
dc.title.alternative
dc.type.degreePh.D.

Files

Original bundle

Now showing 1 - 5 of 13
Loading...
Thumbnail Image
Name:
01_title.pdf
Size:
481.67 KB
Format:
Adobe Portable Document Format
Description:
Attached File
Loading...
Thumbnail Image
Name:
02_prelim pages.pdf
Size:
1.06 MB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
03_contents.pdf
Size:
85.15 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
04_abstract.pdf
Size:
54.12 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
05_chapter-1.pdf
Size:
249.96 KB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.79 KB
Format:
Plain Text
Description: