Data Anonymization in cluster computing using Machine learning algorithms

dc.contributor.guideDey, Rajesh
dc.coverage.spatial
dc.creator.researcherSingh, Monika
dc.date.accessioned2025-04-29T06:32:36Z
dc.date.available2025-04-29T06:32:36Z
dc.date.awarded2025
dc.date.completed2024
dc.date.registered2021
dc.description.abstractBalancing privacy concerns with meaningful data analysis in the era of big data has become a serious challenge. This research explores the prospects of data anonymization in cluster computing environments. We employ machine learning solutions to strengthen the privacy issue without affecting the utility of data. However, traditional anonymization practices are largely limited by scalability and adaptive scope when set up to manage vast data volumes spread at a number of distant locations. To address the problem, the study chooses to propose fresh methods driven by machine learning for anonymization, thereby giving the ability to perform its operations inside cluster computing systems like Hadoop and Spark. This would lend the edge of higher speed and greater scalability as well.A case has been put forward for employing various anonymization techniques, like k-anonymity, l-diversity, and differential privacy. The performance evaluation is taken further by supervised and non-supervised learning executions. Advanced algorithms such as decision trees, k-means clustering, and neural networks then come into the process for finding optimal data anonymization under data sensitivity and contextual settings. An experiment is taken up upon benchmark datasets to analyze the trade-off between privacy protection and data utility by examining metrics like information loss, execution times, and classification accuracy post-anonymization.It is conclusively stated that machine learning anonymization models have extremely good capabilities in enhancing anonymization in situations of vast data sources with many varied data patterns. This concludes the model enhanced with the propositions alone, which can access different anonymization strategies based on various data characteristics and policy-constraint values. The research hence contributes to data privacy, machine learning, and distributed computing and provides those who are interested with practical knowledge to help in safe data transfer via cloud-based platforms, healthcare s
dc.description.note
dc.format.accompanyingmaterialNone
dc.format.dimensions
dc.format.extent
dc.identifier.researcherid0009-0002-5390-3783
dc.identifier.urihttp://hdl.handle.net/10603/634985
dc.languageEnglish
dc.publisher.institutionFaculty of Information Technology and Engineering
dc.publisher.placeRohtas
dc.publisher.universityGopal Narayan Singh University
dc.relationAPA
dc.rightsuniversity
dc.source.universityUniversity
dc.subject.keywordComputer Science
dc.subject.keywordComputer Science Artificial Intelligence
dc.subject.keywordEngineering and Technology
dc.titleData Anonymization in cluster computing using Machine learning algorithms
dc.title.alternative
dc.type.degreePh.D.

Files

Original bundle

Now showing 1 - 5 of 13
Loading...
Thumbnail Image
Name:
01_title.pdf
Size:
18.59 KB
Format:
Adobe Portable Document Format
Description:
Attached File
Loading...
Thumbnail Image
Name:
02_prelim pages.pdf
Size:
806.65 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
03_table of content.pdf
Size:
772.23 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
04_abstract.pdf
Size:
355.32 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
05_chapter1.pdf
Size:
3.02 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.79 KB
Format:
Plain Text
Description: