Publication:

Keeping up with Large Language Models: A Holistic Methodology of Compute, Memory, Communication, and Cost Modeling

 
cris.virtual.department#PLACEHOLDER_PARENT_METADATA_VALUE#
cris.virtual.department#PLACEHOLDER_PARENT_METADATA_VALUE#
cris.virtual.department#PLACEHOLDER_PARENT_METADATA_VALUE#
cris.virtual.department#PLACEHOLDER_PARENT_METADATA_VALUE#
cris.virtual.department#PLACEHOLDER_PARENT_METADATA_VALUE#
cris.virtual.department#PLACEHOLDER_PARENT_METADATA_VALUE#
cris.virtual.orcid0000-0001-8706-4311
cris.virtual.orcid0000-0003-2410-7315
cris.virtual.orcid0000-0003-3578-6069
cris.virtual.orcid0000-0001-7736-6316
cris.virtual.orcid0000-0002-0029-6548
cris.virtual.orcid0000-0002-5337-0617
cris.virtualsource.department44e990b7-69bb-4030-9cfb-7c520e920b5d
cris.virtualsource.departmentbcd4a6e1-db61-4645-a7dd-2a91024f156d
cris.virtualsource.departmentabec653f-8f18-475b-8327-7e7829f7aafb
cris.virtualsource.department04af0672-0fe2-4201-b106-c913135dae0b
cris.virtualsource.departmentf6f17b49-e3c3-4223-9429-3bcd739eacc2
cris.virtualsource.department50f46d64-80bf-4a78-9703-a96abed1e6b8
cris.virtualsource.orcid44e990b7-69bb-4030-9cfb-7c520e920b5d
cris.virtualsource.orcidbcd4a6e1-db61-4645-a7dd-2a91024f156d
cris.virtualsource.orcidabec653f-8f18-475b-8327-7e7829f7aafb
cris.virtualsource.orcid04af0672-0fe2-4201-b106-c913135dae0b
cris.virtualsource.orcidf6f17b49-e3c3-4223-9429-3bcd739eacc2
cris.virtualsource.orcid50f46d64-80bf-4a78-9703-a96abed1e6b8
dc.contributor.authorGuo, Wenzhe
dc.contributor.authorKundu, Joyjit
dc.contributor.authorTos, Uras
dc.contributor.authorKong, Weijiang
dc.contributor.authorSisto, Giuliano
dc.contributor.authorEvenblij, Timon
dc.contributor.authorPerumkunnil, Manu
dc.date.accessioned2026-09-08T07:51:29Z
dc.date.available2026-09-08T07:51:29Z
dc.date.createdwos2026
dc.date.issued2025
dc.description.abstractLarge language models (LLMs), based on transformer architectures, have revolutionized numerous domains within artificial intelligence, science, and engineering due to their exceptional scalability and adaptability. However, the exponential growth in LLM size and complexity has outpaced advancements in compute capacity, memory bandwidth, network performance, and cost efficiency, posing significant challenges to their scalability on distributed systems. To address these limitations, alternative model architectures, optimization strategies, communication-aware network topologies, and novel system design approaches have been proposed in literature. This paper introduces a performance-cost modeling methodology for LLM training and inference that integrates state-of-the-art compute techniques with memory optimizations, and latest communication techniques. Building on an analytical performance model, our approach incorporates recent innovations such as the flash attention technique and mixture of experts models to address the memory bandwidth and compute bottlenecks. It also considers the impact of different network topologies and topology-specific communication algorithms with 5D parallellism. The framework also integrates a chiplet cost model. The proposed modeling methodology provides valuable insights to guide future compute system design and facilitates hardware-software co-development, in particular due to its ability to analyze performance-cost trade-offs for various system architectural configurations.
dc.identifier.doi10.1109/iiswc66894.2025.00019
dc.identifier.isbn979-8-3315-4918-3
dc.identifier.urihttps://imec-publications.be/handle/20.500.12860/60259
dc.language.isoeng
dc.provenance.editstepusergreet.vanhoof@imec.be
dc.publisherIEEE COMPUTER SOC
dc.relation.ispartofseriesInternational Symposium on Workload Characterization Proceedings
dc.source.beginpage116
dc.source.conferenceIEEE International Symposium on Workload Characterization (IISWC)
dc.source.conferencedate2025-10-12
dc.source.conferencelocationIrvine
dc.source.endpage126
dc.source.journal2025 IEEE INTERNATIONAL SYMPOSIUM ON WORKLOAD CHARACTERIZATION, IISWC
dc.source.numberofpages11
dc.subject.keywordsDESIGN
dc.subject.keywordsSIM
dc.title

Keeping up with Large Language Models: A Holistic Methodology of Compute, Memory, Communication, and Cost Modeling

dc.typeProceedings paper
dspace.entity.typePublication
imec.internal.crawledAt2025-11-21
imec.internal.sourcecrawler
imec.internal.wosCreatedAt2026-09-07
Files
Publication available in collections: