Statistical Models Suited for Machine Translation from English to Indian Languages

dc.contributor.guideSangal, Rajeev and Joshi, Aravind
dc.coverage.spatial
dc.creator.researcherSriram,Venkatapathy
dc.date.accessioned2024-02-07T09:56:06Z
dc.date.available2024-02-07T09:56:06Z
dc.date.awarded2010
dc.date.completed2010
dc.date.registered2003
dc.description.abstractIn this thesis, models are proposed for statistical machine translation and word-alignment which newlineare well-suited for translation between English and Hindi (or any other Indian Language (IL)) apart newlinefrom addressing the limitations of existing machine translation systems. In our models, we allow the newlineflexibility of using linguistically motivated features but leave the learning to machine based on data. The newlinepurpose is better translation quality. Some of the issues in building a statistical translation system from newlineEnglish to Hindi are, newline1. Large distance word-reordering. newline2. Morphological richness of Hindi. newline3. Lack of large training dataset. newline4. Importance of function words in conveying grammatical roles in ILs. newline5. Translation/Word-alignment of multi-word expressions. newlineI address these issues in the statistical machine translation and word-alignment models proposed by newlineus. The main contributions of my work are the two approaches that I propose for machine translation. newlineThey are, newline1. Discriminative machine translation using global lexical selection (GLS approach): newlineIn this system, I propose a novel approach for machine translation using global lexical selection. newlineGlobal lexical selection allows us to not rely on word-alignment information (which are often newlineunreliable) between sentence pairs. Also, the emphasis in this approach is on lexical selection newlinerather than lexical reordering. Indian languages are relatively free word order and hence, lexical newlineselection (specially the function words) plays the most important role while translating to them. newlineThe results of this approach are presented in Chapters 4 and 5 [128, 129]. newline2. Dependency based statistical machine translation system (Vaanee): newlineI propose a dependency based statistical system that uses discriminative techniques to train its newlineparameters. The use of syntax (dependency tree) allows us to address the large word-reorderings newlinebetween English and Hindi. Discriminative training allows us to use rich feature sets, including newlinelinguistic features that are useful in the machine translation
dc.description.note
dc.format.accompanyingmaterialNone
dc.format.dimensions
dc.format.extent187
dc.identifier.urihttp://hdl.handle.net/10603/544192
dc.languageEnglish
dc.publisher.institutionComputer Science and Engineering
dc.publisher.placeHyderabad
dc.publisher.universityInternational Institute of Information Technology, Hyderabad
dc.relation
dc.rightsuniversity
dc.source.universityUniversity
dc.subject.keywordComputer Science
dc.subject.keywordComputer Science Artificial Intelligence
dc.subject.keywordEngineering and Technology
dc.titleStatistical Models Suited for Machine Translation from English to Indian Languages
dc.title.alternative
dc.type.degreePh.D.

Files

Original bundle

Now showing 1 - 5 of 15
Loading...
Thumbnail Image
Name:
80_recommendation.pdf
Size:
16.37 KB
Format:
Adobe Portable Document Format
Description:
Attached File
Loading...
Thumbnail Image
Name:
abstract.pdf
Size:
14.82 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
annexures.pdf
Size:
69.29 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
chapter 1.pdf
Size:
304.88 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
chapter 2.pdf
Size:
111.83 KB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.79 KB
Format:
Plain Text
Description: