Intelligent techniques for enhancing recall and precission in cross lingual search among indian languages
Loading...
Date
item.page.authors
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
The Indian language information access technologies face severe precision and recall problems when using conventional Information Retrieval techniques (used for English-like languages).This is a study of Indian language information access. During this study, we investigated the web extensively for Indian languages and in this process, we came up with some solutions for the low recall and precision problems. We focused our research on the key components of cross-lingual search like indexing, transliteration, stemming, translation and summarization. The following are some of the major contributions of this thesis.
newlineand#61623; We built a query-biased summarizer based on similarity of sentences.
newlineand#61623; We produced a summary which is having good linguistic quality and also is non-redundant with good Recall and Precision rates.
newlineand#61623; We developed Hindi- English CLIR without using query translation.
newlineand#61623; We resolved polysemy, synonymy problems caused by dictionary based indexing.
newlineand#61623; We retrieved the documents based on the semantic relation between documents.
newlineand#61623; We defined and developed a language independent stemmer.
newlineand#61623; We modeled an unsupervised Telugu stemmer that does not require inflection-root pair for training.
newlineand#61623; We showed retrieval effectiveness and reduced the size of the index for the Telugu information retrieval task.
newlineand#61623; We constructed a Machine transliteration model which combines both grapheme and phoneme based transliteration models.
newlineand#61623; We came up with user friendly transliteration of words by considering both the spelling and the pronunciation of a word.
newlineand#61623; We designed complete file transliteration model.
newlineii
newlineAll our evaluations were based on Telugu, Hindi and English datasets, which are proved to be reasonably good for most of the Indian Languages with minor modifications. In most of our experiments, we used system-based evaluation methodology which is a widely used evaluation methodology in Information Access research community. All available standard evaluation datasets were used. In other cases, we built our evaluation datasets us