Deep learning basedframeworksfor Automaticmedicalvideosummarization

dc.contributor.guideSingh, Prabhishek
dc.coverage.spatial
dc.creator.researcherGupta, Jaya
dc.date.accessioned2025-12-19T08:46:42Z
dc.date.available2025-12-19T08:46:42Z
dc.date.awarded2025
dc.date.completed2025
dc.date.registered2019
dc.description.abstractquotThe increasing use of medical video recordings in clinical practice and surgical education newlinepresents significant opportunities for knowledge extraction and decision support. However, newlineeffectively analyzing these videos remains a considerable challenge due to their extended dura newlinetion, visual redundancy, and complex semantics. A single surgical procedure may span several newlinehours, generating large video files that are difficult to browse, store, or manually review. Med newlineical videos differ significantly from general-purpose visual content due to their limited visual newlinevariation. Subtle changes, such as slight movements of surgical tools or minor tissue responses, newlineare often clinically significant. However, these changes are typically difficult to detect without newlinespecialized contextual understanding. This presents unique challenges for automated analysis, newlineas critical events may be visually indistinct yet carry important medical implications. The lack newlineof clear scene transitions and the high visual similarity across frames make automatic summa newlinerization particularly difficult. Additionally, the scarcity of annotated datasets, due to the cost newlineand expertise required for labeling, further limits the effectiveness of supervised learning ap newlineproaches. This thesis addresses these challenges through a series of novel, context-aware video newlinesummarization frameworks designed for medical and dynamic procedural content. newlineThe first contribution introduces an unsupervised, resolution-aware summarization frame newlinework, combiningCNN-basedmulti-layerfeature extraction (MobileNetV3, GoogleNet, DenseNet newline161, ResNet50, EfficientNet-B7), spatial pyramid pooling (SPP), and Bi-LSTM-based tempo newlineral modeling. This approach effectively captures multi-scale spatio-temporal features and iden newlinetifies key events across diverse medical datasets. Among the architectures explored, EfficientNet newlineB7 stands out with the highest precision of 87.2% for Endoscapes-CVS201 dataset and the newlinehighest recall of 88.3% for JIGSAWS, achieving an overall F1-score of 87.1%.
dc.description.note
dc.format.accompanyingmaterialNone
dc.format.dimensions
dc.format.extentXVII; 107
dc.identifier.researcherid0000-0002-2908-2762
dc.identifier.urihttp://hdl.handle.net/10603/682336
dc.languageEnglish
dc.publisher.institutionSchool of Computer Science Engineering and Technology
dc.publisher.placeGreater Noida
dc.publisher.universityBennett University
dc.relation
dc.rightsuniversity
dc.source.universityUniversity
dc.subject.keywordComputer Science
dc.subject.keywordComputer Science Information Systems
dc.subject.keywordEngineering and Technology
dc.titleDeep learning basedframeworksfor Automaticmedicalvideosummarization
dc.title.alternative
dc.type.degreePh.D.

Files

Original bundle

Now showing 1 - 5 of 13
Loading...
Thumbnail Image
Name:
01_title.pdf
Size:
160.02 KB
Format:
Adobe Portable Document Format
Description:
Attached File
Loading...
Thumbnail Image
Name:
02_prelim pages.pdf
Size:
593.13 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
03_content.pdf
Size:
53.96 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
04_abstract.pdf
Size:
57.07 KB
Format:
Adobe Portable Document Format
Loading...
Thumbnail Image
Name:
05_chapter 1.pdf
Size:
1.14 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.79 KB
Format:
Plain Text
Description: