Publication:

Semantic Representation and Attention Alignment for Graph Information Bottleneck in Video Summarization

 
cris.virtual.department#PLACEHOLDER_PARENT_METADATA_VALUE#
cris.virtual.orcid0000-0001-7290-0428
cris.virtualsource.department72e3f67d-06a9-4361-96cf-ba7646472b0b
cris.virtualsource.orcid72e3f67d-06a9-4361-96cf-ba7646472b0b
dc.contributor.authorZhong, Rui
dc.contributor.authorWang, Rui
dc.contributor.authorYao, Wenjin
dc.contributor.authorHu, Min
dc.contributor.authorDong, Shi
dc.contributor.authorMunteanu, Adrian
dc.date.accessioned2023-12-19T08:23:44Z
dc.date.available2023-08-11T16:46:11Z
dc.date.available2023-12-19T08:23:44Z
dc.date.embargo2024-01-13
dc.date.issued2023
dc.description.abstractEnd-to-end Long Short-Term Memory (LSTM) has been successfully applied to video summarization. However, the weakness of the LSTM model, poor generalization with inefficient representation learning for inputted nodes, limits its capability to efficiently carry out node classification within user-created videos. Given the power of Graph Neural Networks (GNNs) in representation learning, we adopted the Graph Information Bottle (GIB) to develop a Contextual Feature Transformation (CFT) mechanism that refines the temporal dual-feature, yielding a semantic representation with attention alignment. Furthermore, a novel Salient-Area-Size-based spatial attention model is presented to extract frame-wise visual features based on the observation that humans tend to focus on sizable and moving objects. Lastly, semantic representation is embedded within attention alignment under the end-to-end LSTM framework to differentiate indistinguishable images. Extensive experiments demonstrate that the proposed method outperforms State-Of-The-Art (SOTA) methods.
dc.description.wosFundingTextThis work was supported in part by the National Natural Science Foundation of China under Grant 62002130 and Grant 62201222; in part by the Fundamental Research Funds for the Central Universities under Grant CCNU22QN014, Grant CCNU22XJ034, and Grant CCNU22JC007; in part by the National Key Research and Development Program of China under Grant 2022YFD1700204; and in part by Fonds Wetenschappelijk Onderzoek(FWO), Vlaanderen, under Project G094122N.
dc.identifier.doi10.1109/tip.2023.3293762
dc.identifier.issn1057-7149
dc.identifier.pmidMEDLINE:37440397
dc.identifier.urihttps://imec-publications.be/handle/20.500.12860/42324
dc.publisherIEEE-INST ELECTRICAL ELECTRONICS ENGINEERS INC
dc.source.beginpage4170
dc.source.endpage4184
dc.source.issue/
dc.source.journalIEEE TRANSACTIONS ON IMAGE PROCESSING
dc.source.numberofpages15
dc.source.volume32
dc.subject.keywordsNETWORK
dc.subject.keywordsLSTM
dc.title

Semantic Representation and Attention Alignment for Graph Information Bottleneck in Video Summarization

dc.typeJournal article
dspace.entity.typePublication
Files

Original bundle

Name:
TIP_final_manuscript_v3.pdf
Size:
5.73 MB
Format:
Adobe Portable Document Format
Description:
Accepted version
Publication available in collections: