A novel approach for duplicate elimination and effective topic modeling for document clustering

dc.contributor.guideLatha B
dc.coverage.spatialA novel approach for duplicate elimination and effective topic modeling for document clustering
dc.creator.researcherUma R
dc.date.accessioned2021-10-06T05:09:39Z
dc.date.available2021-10-06T05:09:39Z
dc.date.awarded2020
dc.date.completed2020
dc.date.registered
dc.description.abstractInformation available on the web increases at a fast pace, within the last two years, there has been an explosive growth of internet information. A great amount of information available in the web is textual information. Textual information plays a vital part in IR and it is probably the most useful information. Searching the Web is becoming dominant due to the fact of richness in information available and convenience in getting the information required .Web search is rooted towards Information Retrieval (IR) which is a study that assists users in finding the required information from a large corpus of documents. The documents in the web are called WebPages. Relevancy and efficiency are the ultimate issues in web search. WebPages are semi-structured in nature. The content in a page is organized and presented in multiple structured blocks. Some blocks contain vital information and others are not. Detecting the main content blocks actively from a webpage is useful in searching the web because terms that are found in those blocks are more important. Users face a great difficulty in identifying the relevant information. The existing approaches need to improve the accuracy in terms of relevancy. Information retrieval is a way to separate relevant data from the irrelevant. Documents on the web are available in different formats. Conventional information retrieval methods operate on clean text, if there is noise in the data it has to be cleaned for efficient retrieval. This research work takes an initiate to increase the retrieval accuracy, relevancy and increase the performance of retrieval for text documents. To attain these goals the search space has to be reduced and the underlying semantics need to be identified. newline
dc.description.note
dc.format.accompanyingmaterialNone
dc.format.dimensions21cm
dc.format.extentxviii, 151p.
dc.identifier.urihttp://hdl.handle.net/10603/343259
dc.languageEnglish
dc.publisher.institutionFaculty of Information and Communication Engineering
dc.publisher.placeChennai
dc.publisher.universityAnna University
dc.relationp.140-150
dc.rightsuniversity
dc.source.universityUniversity
dc.subject.keywordEngineering and Technology
dc.subject.keywordComputer Science
dc.subject.keywordComputer Science Information Systems
dc.subject.keywordDocument Clustering
dc.subject.keywordDuplicate Elimination
dc.subject.keywordInformation Retrieval
dc.subject.keywordSub Topic Model
dc.titleA novel approach for duplicate elimination and effective topic modeling for document clustering
dc.title.alternative
dc.type.degreePh.D.

Files

Original bundle

Now showing 1 - 5 of 17
Loading...
Thumbnail Image
Name:
01_title.pdf
Size:
24.12 KB
Format:
Adobe Portable Document Format
Description:
Attached File
Loading...
Thumbnail Image
Name:
02_certificates.pdf
Size:
563.93 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
03_abstracts.pdf
Size:
14.14 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
04_acknowledgements.pdf
Size:
456.88 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
05_contents.pdf
Size:
15.32 KB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.79 KB
Format:
Plain Text
Description: