Privacy Preservation Techniques for High Dimensional Data

dc.contributor.guideVENKATESULU DONDETI
dc.coverage.spatial
dc.creator.researcherSHASHIDHAR VIRUPAKSHA
dc.date.accessioned2022-09-16T10:39:19Z
dc.date.available2022-09-16T10:39:19Z
dc.date.awarded
dc.date.completed2022
dc.date.registered2016
dc.description.abstractData is being collected every day by the organizations and data mining is performed on the collected data. When data is shared for data mining, sensitive information about individuals is also revealed. Many governments have enacted legislation also to preserve privacy. Hence Privacy Preserving Data Mining (PPDM) algorithms were also developed. PPDM works by transforming the dataset to preserve confidential information while performing data mining. newline newlineIn the last few years, many applications and organizations deal with High Dimensional (HD) continuous datasets in data mining. PPDM on HD continuous datasets is challenging because HD datasets contain many irrelevant dimensions. Data is now available in subspaces. Data characteristics are not the same in these subspaces. Identifying to which cluster a point belongs is difficult. Thus, PPDM on HD datasets results in data loss, information loss, and some original clusters are lost. HD continuous datasets are especially being used in medical and healthcare for disease diagnosis, newborn screen and gene analysis. PPDM on such HD continuous data is even more difficult because some datasets are noise sensitive, have less records, privacy offered is low since a small distortion leads to very high data loss, information loss and original clusters are lost. Hence in this thesis, PPDM algorithms are proposed for high dimensional continuous data. newlinePPDM algorithms on continuous data are classified into two major categories anonymization and noise addition. This thesis proposes novel anonymization algorithm Subspace Based Aggregation (SBA), novel noise addition algorithm Subspace Based Noise Addition (SBNA) and a novel hybrid Anonymized Noise Addition in Subspaces (ANAS). SBA, SBNA and ANAS decreases data loss, information loss and enhance clusters that are identified. SBA computes the subspaces first. Records in these subspaces are then grouped based on squared Euclidean distances. These groups of records are now aggregated.
dc.description.note
dc.format.accompanyingmaterialCD
dc.format.dimensions
dc.format.extent145
dc.identifier.urihttp://hdl.handle.net/10603/405784
dc.languageEnglish
dc.publisher.institutionDepartment of Computer Science and Engineering
dc.publisher.placeGuntur
dc.publisher.universityVignans Foundation for Science Technology and Research
dc.relation
dc.rightsuniversity
dc.source.universityUniversity
dc.subject.keywordEngineering and Technology
dc.subject.keywordComputer Science
dc.subject.keywordComputer Science Theory and Methods
dc.titlePrivacy Preservation Techniques for High Dimensional Data
dc.title.alternative
dc.type.degreePh.D.

Files

Original bundle

Now showing 1 - 5 of 15
Loading...
Thumbnail Image
Name:
10_chapter-3.pdf
Size:
326.01 KB
Format:
Adobe Portable Document Format
Description:
Attached File
Loading...
Thumbnail Image
Name:
11_chapter-4.pdf
Size:
266.22 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
12-chapter-5.pdf
Size:
283.88 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
13_chapter-6.pdf
Size:
60.53 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
14_publications.pdf
Size:
174.13 KB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.79 KB
Format:
Plain Text
Description: