Publication:

Efficient DNN Training Using Vectorized Block-Scaled GeMMs with Adaptive Block Shapes

 
cris.virtual.department#PLACEHOLDER_PARENT_METADATA_VALUE#
cris.virtual.department#PLACEHOLDER_PARENT_METADATA_VALUE#
cris.virtual.department#PLACEHOLDER_PARENT_METADATA_VALUE#
cris.virtual.department#PLACEHOLDER_PARENT_METADATA_VALUE#
cris.virtual.orcid0000-0003-3495-9263
cris.virtual.orcid0000-0003-0181-8069
cris.virtual.orcid0000-0002-1592-755X
cris.virtual.orcid0000-0001-6561-8934
cris.virtualsource.departmentc84426b5-5f84-48ba-9153-1fe96862af32
cris.virtualsource.department91857424-b227-471d-aaae-198ffad1e716
cris.virtualsource.department15e57581-19c6-4927-9cf5-286e171d9d9e
cris.virtualsource.department873d5ca3-d769-441b-b014-52f18a2fd1c0
cris.virtualsource.orcidc84426b5-5f84-48ba-9153-1fe96862af32
cris.virtualsource.orcid91857424-b227-471d-aaae-198ffad1e716
cris.virtualsource.orcid15e57581-19c6-4927-9cf5-286e171d9d9e
cris.virtualsource.orcid873d5ca3-d769-441b-b014-52f18a2fd1c0
dc.contributor.authorSatya Murthy, Nitish
dc.contributor.authorLaubeuf, Nathan
dc.contributor.authorBhattacharjee, Debjyoti
dc.contributor.authorCatthoor, Francky
dc.contributor.authorVerhelst, Marian
dc.date.accessioned2026-08-25T12:55:25Z
dc.date.available2026-08-25T12:55:25Z
dc.date.createdwos2026
dc.date.issued2026
dc.description.abstractReduced precision datatypes have become essential to the efficient training and deployment of Deep Neural Networks (DNNs). A recent development in the field has been the emergence of block-scaled datatypes: tensor representation formats derived from floating-point, that share a common exponent across multiple elements. While these formats are being broadly adopted and optimized for by DNN-specific inference accelerators, the potential benefits for training workloads on general-purpose vector processors has yet to be thoroughly explored. This work proposes implementations of block-scaled general matrix multiplications (GeMM) for DNN training at the edge using commercially available vector instruction sets (ARM SVE). Using these implementations, we highlight an accuracy-speed trade-off involving the shape of shared exponent blocks—vectors or squares. We exploit this result to optimize the training of fully connected networks by adapting the shared exponent block shapes during training. We also explore deployments on scalable hardware vector lengths utilizing different degrees of data-level parallelism. We demonstrate more efficient DNN training with our block-scaled datatypes compared to standard IEEE 32-bit floating point (FP32) formats through effective hardware-software cooptimizations.
dc.description.wosFundingTextThis research received funding from the Flemish Government (AI Research Program).
dc.identifier.doi10.1007/978-3-032-09575-6_3
dc.identifier.isbn978-3-032-09577-0
dc.identifier.isbn978-3-032-09574-9
dc.identifier.issn1868-4238
dc.identifier.urihttps://imec-publications.be/handle/20.500.12860/60117
dc.language.isoeng
dc.provenance.editstepusergreet.vanhoof@imec.be
dc.publisherSPRINGER INTERNATIONAL PUBLISHING AG
dc.source.beginpage33
dc.source.conferenceVLSI-SoC: Technology Advancement on SoC Design 32nd IFIP/IEEE International Conference on Very Large Scale Integration - System on a Chip
dc.source.conferencedate2024-10-06
dc.source.conferencelocationTangier
dc.source.endpage47
dc.source.journalVLSI-SOC: TECHNOLOGY ADVANCEMENT ON SOC DESIGN, 2024
dc.source.numberofpages15
dc.title

Efficient DNN Training Using Vectorized Block-Scaled GeMMs with Adaptive Block Shapes

dc.typeProceedings paper
dspace.entity.typePublication
imec.internal.crawledAt2026-07-14
imec.internal.sourcecrawler
imec.internal.wosCreatedAt2026-07-14
Files
Publication available in collections: