Publication:
Semantic Representation and Attention Alignment for Graph Information Bottleneck in Video Summarization
| cris.virtual.department | #PLACEHOLDER_PARENT_METADATA_VALUE# | |
| cris.virtual.orcid | 0000-0001-7290-0428 | |
| cris.virtualsource.department | 72e3f67d-06a9-4361-96cf-ba7646472b0b | |
| cris.virtualsource.orcid | 72e3f67d-06a9-4361-96cf-ba7646472b0b | |
| dc.contributor.author | Zhong, Rui | |
| dc.contributor.author | Wang, Rui | |
| dc.contributor.author | Yao, Wenjin | |
| dc.contributor.author | Hu, Min | |
| dc.contributor.author | Dong, Shi | |
| dc.contributor.author | Munteanu, Adrian | |
| dc.date.accessioned | 2023-12-19T08:23:44Z | |
| dc.date.available | 2023-08-11T16:46:11Z | |
| dc.date.available | 2023-12-19T08:23:44Z | |
| dc.date.embargo | 2024-01-13 | |
| dc.date.issued | 2023 | |
| dc.description.abstract | End-to-end Long Short-Term Memory (LSTM) has been successfully applied to video summarization. However, the weakness of the LSTM model, poor generalization with inefficient representation learning for inputted nodes, limits its capability to efficiently carry out node classification within user-created videos. Given the power of Graph Neural Networks (GNNs) in representation learning, we adopted the Graph Information Bottle (GIB) to develop a Contextual Feature Transformation (CFT) mechanism that refines the temporal dual-feature, yielding a semantic representation with attention alignment. Furthermore, a novel Salient-Area-Size-based spatial attention model is presented to extract frame-wise visual features based on the observation that humans tend to focus on sizable and moving objects. Lastly, semantic representation is embedded within attention alignment under the end-to-end LSTM framework to differentiate indistinguishable images. Extensive experiments demonstrate that the proposed method outperforms State-Of-The-Art (SOTA) methods. | |
| dc.description.wosFundingText | This work was supported in part by the National Natural Science Foundation of China under Grant 62002130 and Grant 62201222; in part by the Fundamental Research Funds for the Central Universities under Grant CCNU22QN014, Grant CCNU22XJ034, and Grant CCNU22JC007; in part by the National Key Research and Development Program of China under Grant 2022YFD1700204; and in part by Fonds Wetenschappelijk Onderzoek(FWO), Vlaanderen, under Project G094122N. | |
| dc.identifier.doi | 10.1109/tip.2023.3293762 | |
| dc.identifier.issn | 1057-7149 | |
| dc.identifier.pmid | MEDLINE:37440397 | |
| dc.identifier.uri | https://imec-publications.be/handle/20.500.12860/42324 | |
| dc.publisher | IEEE-INST ELECTRICAL ELECTRONICS ENGINEERS INC | |
| dc.source.beginpage | 4170 | |
| dc.source.endpage | 4184 | |
| dc.source.issue | / | |
| dc.source.journal | IEEE TRANSACTIONS ON IMAGE PROCESSING | |
| dc.source.numberofpages | 15 | |
| dc.source.volume | 32 | |
| dc.subject.keywords | NETWORK | |
| dc.subject.keywords | LSTM | |
| dc.title | Semantic Representation and Attention Alignment for Graph Information Bottleneck in Video Summarization | |
| dc.type | Journal article | |
| dspace.entity.type | Publication | |
| Files | Original bundle
| |
| Publication available in collections: |