Design of algorithms for gene predictions

dc.contributor.guideGarg, Deepak
dc.coverage.spatial
dc.creator.researcherMaji, Srabanti
dc.date.accessioned2019-01-25T10:30:33Z
dc.date.available2019-01-25T10:30:33Z
dc.date.awarded
dc.date.completed2013
dc.date.registered
dc.description.abstractIdentification of coding sequence from genomic DNA sequence is the major step in pursuit of gene identification. In the prediction of splice site, which is the separation between exons and introns, though the sequences adjacent to the splice sites have a high conservation, but still, the accuracy is lower than 90%. Therefore, here, both approaches Conventional as well as Computational Intelligences (CI) have been pursued to predict the splice site in DNA sequence of the Eukaryotic organism and, both have been evaluated and compared in terms of their performance. In the conventional approach, i.e., Hidden Markov Model (HMM) System , the model architecture includes the probabilistic descriptions of the splicing, translational, and transcriptional signals. Splice sites predictor based on Unique Hidden Markov Model (HMM) is developed and trained using Modified Expectation Maximization (MEM) algorithm. A 12 fold cross validation technique is also applied to check the reproducibility of the results obtained and to further increase the prediction accuracy. The proposed system is able to achieve the accuracy of 98% of true donor site and 93% for true acceptor site in the standard DNA (nucleotide) sequence. The second proposed method, based on combination of conventional and computational intelligences, namely, Markov Model 2 Feature Support Vector Machine (MM2F-SVM) consists of three stages initial stage, in which a second order Markov Model (MM2) is used; intermediate, or the second stage in which principal feature analysis (PFA) is done; and the third or final stage, in which a support vector machine (SVM) with Gaussian kernel is used. The first stage is known as feature extraction ; the second stage is called feature selection and, the final stage is known as classification . The model is proficient of indicating the reliability of each predicted splice site with high accuracy.
dc.description.note
dc.format.accompanyingmaterialNone
dc.format.dimensions
dc.format.extentxv, 95p.
dc.identifier.urihttp://hdl.handle.net/10603/227192
dc.languageEnglish
dc.publisher.institutionDepartment of Computer Science and Engineering
dc.publisher.placePatiala
dc.publisher.universityThapar Institute of Engineering and Technology
dc.relation
dc.rightsuniversity
dc.source.universityUniversity
dc.subject.keywordBioinformatics
dc.subject.keywordGene Identification
dc.subject.keywordSplice Site
dc.subject.keywordSupport Vector Machine
dc.titleDesign of algorithms for gene predictions
dc.title.alternative
dc.type.degreePh.D.

Files

Original bundle

Now showing 1 - 5 of 11
Loading...
Thumbnail Image
Name:
file10(publications).pdf
Size:
11.51 KB
Format:
Adobe Portable Document Format
Description:
Attached File
Loading...
Thumbnail Image
Name:
file11(references).pdf
Size:
121.92 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
file1(title).pdf
Size:
16.3 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
file2(certificate).pdf
Size:
42.83 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
file3(preliminary pages).pdf
Size:
106.09 KB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.79 KB
Format:
Plain Text
Description: