Design and development of progressive sampling models for increasing the efficiency of data analytics

dc.contributor.guideH A, Dinesha
dc.coverage.spatial
dc.creator.researcherB C, Yathish Aradhya
dc.date.accessioned2024-09-30T05:10:21Z
dc.date.available2024-09-30T05:10:21Z
dc.date.awarded2024
dc.date.completed2024
dc.date.registered2016
dc.description.abstractBig Data become famous for its huge quantity and multi-faceted nature which offers obstacles for analytics thus making it impractical to exploit the complete dataset for training learning algorithms. The challenge of sampling such much data is difficult. Dealing with big data involves processing huge amounts, creating a substantial challenge for both academia and industries. Conventional sampling strategies fall short when confronted with difficulties like data imbalance, substantial data heterogeneity, and multi-dimensionality. While existing progressive sampling systems frequently rely on random sample selection, this may be inadequate, particularly in scenarios with high data heterogeneity and unbalanced data circumstances. The restrictions in random feature selection can lead to erroneous performance due to insufficient or insignificant feature learning. Therefore, a resilient computing environment, along with better pre-processing, feature-sensitive feature selection, and classification, offers as a promising solution for effective big data analytics. newlineIn the first study, innovative approaches to enhance the operational efficiency of Probabilistic Sampling Algorithms (PSA) have been explored. Thus, mainly focusing on a substantial reduction in the cardinality of the training dataset while preserving the accuracy of the learning method within the framework of Probably Approximately Correct (PAC) using Bounds of Rademacher Averages (RMA). The newly suggested strategies rely on the divergence of statistics for the selection of an initial set of samples for Rademacher averages and a sampling schedule that is reliant on data and#8455;-approximation, giving a tight constraint for the learning process. These approaches are developed to maximize the theoretical assessments of Rademacher averages in the context of a progressive sampling algorithm for estimating stopping time. The PSA s runtime is successfully controlled by the PSA through the adoption of these novel approaches, with a time complexity of O(and#8455;-approximation
dc.description.note
dc.format.accompanyingmaterialDVD
dc.format.dimensions
dc.format.extent158
dc.identifier.urihttp://hdl.handle.net/10603/592434
dc.languageEnglish
dc.publisher.institutionDepartment of Computer Science and Engineering
dc.publisher.placeBelagavi
dc.publisher.universityVisvesvaraya Technological University, Belagavi
dc.relation
dc.rightsuniversity
dc.source.universityUniversity
dc.subject.keywordComputer Science
dc.subject.keywordComputer Science Interdisciplinary Applications
dc.subject.keywordEngineering and Technology
dc.titleDesign and development of progressive sampling models for increasing the efficiency of data analytics
dc.title.alternative
dc.type.degreePh.D.

Files

Original bundle

Now showing 1 - 5 of 12
Loading...
Thumbnail Image
Name:
01_title.pdf
Size:
32.59 KB
Format:
Adobe Portable Document Format
Description:
Attached File
Loading...
Thumbnail Image
Name:
02_prelim pages.pdf
Size:
476.62 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
03_content.pdf
Size:
196.82 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
04_abstract.pdf
Size:
159.68 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
05_chapter 1.pdf
Size:
284.7 KB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.79 KB
Format:
Plain Text
Description: