Publication:

ADGT: Enhancing 3D human pose estimation with attention-driven graph-transformers

 
cris.virtual.department#PLACEHOLDER_PARENT_METADATA_VALUE#
cris.virtual.department#PLACEHOLDER_PARENT_METADATA_VALUE#
cris.virtual.orcid0000-0001-8042-6834
cris.virtual.orcid0000-0002-1774-2970
cris.virtualsource.department6d0ac6ee-44b1-4239-ad3e-44be2a439e9b
cris.virtualsource.departmentf869d5b4-4c3a-4052-af0a-3f6e6925edb3
cris.virtualsource.orcid6d0ac6ee-44b1-4239-ad3e-44be2a439e9b
cris.virtualsource.orcidf869d5b4-4c3a-4052-af0a-3f6e6925edb3
dc.contributor.authorYang, Shuo
dc.contributor.authorLuu, Anh Tuan
dc.contributor.authorNguyen, Xuan Son
dc.contributor.authorHistace, Aymeric
dc.contributor.authorJansen, Bart
dc.contributor.authorSahli, Hichem
dc.date.accessioned2026-08-31T12:03:33Z
dc.date.available2026-08-31T12:03:33Z
dc.date.createdwos2026
dc.date.issued2026
dc.description.abstract2D-to-3D lifting is a fundamental approach in 3D human pose estimation (3DHPE). This task is crucial in applications, including motion analysis and virtual reality. While Graph Convolutional Networks (GCNs) have demonstrated effectiveness in capturing spatial relationships in human skeletons, they suffer from over-smoothing and limited receptive fields. Transformer-based models provide global context but struggle with local feature extraction and computational efficiency. To address these challenges, we propose ADGT, a novel parallel GCN-transformer architecture combining the strengths of both approaches. Our method introduces three key innovations: Hop-Wise Scalable Adaptive GCN to refine local feature extraction, Attention-Based Local Feature Extractor to enhance the integration of local and global representations, and Register-Based Transformer Enhancement to improve feature separation. Extensive experiments on Human3.6M and MPI-INF-3DHP datasets demonstrate ADGT achieves state-of-the-art performance among frame-based methods while maintaining computational efficiency. These results highlight the potential of ADGT for real-time applications requiring accurate and efficient 3DHPE. The code is available at https://github.com/sYANGunique1111/ADGT
dc.description.wosFundingTextThis work was funded by the EUTOPIA PhD Co-tutelle Program, France (EUTOPIA is an alliance of ten European universities and six Global partners co-funded by the European Union) grant EUTOPIA-PhD-2021-0000000061.
dc.identifier.doi10.1016/j.jvcir.2026.104829
dc.identifier.issn1047-3203
dc.identifier.urihttps://imec-publications.be/handle/20.500.12860/60157
dc.language.isoeng
dc.provenance.editstepusergreet.vanhoof@imec.be
dc.publisherACADEMIC PRESS INC ELSEVIER SCIENCE
dc.source.beginpage104829
dc.source.journalJOURNAL OF VISUAL COMMUNICATION AND IMAGE REPRESENTATION
dc.source.numberofpages12
dc.source.volume118
dc.subject.keywordsRECOGNITION
dc.title

ADGT: Enhancing 3D human pose estimation with attention-driven graph-transformers

dc.typeJournal article
dspace.entity.typePublication
imec.internal.crawledAt2026-07-14
imec.internal.sourcecrawler
imec.internal.wosCreatedAt2026-07-14
Files

Original bundle

Name:
1-s2.0-S1047320326001240-main.pdf
Size:
2.63 MB
Format:
Adobe Portable Document Format
Description:
Published
Publication available in collections: