Publication:

Soundscape Captioning Using Sound Affective Quality Network and Large Language Model

 
cris.virtual.department#PLACEHOLDER_PARENT_METADATA_VALUE#
cris.virtual.department#PLACEHOLDER_PARENT_METADATA_VALUE#
cris.virtual.orcid0000-0001-5207-7745
cris.virtual.orcid#PLACEHOLDER_PARENT_METADATA_VALUE#
cris.virtualsource.department6c1aac4b-593e-4f80-9ecc-911fd20f3c31
cris.virtualsource.department952936d6-a5d1-4952-ab8c-7ba5c377af16
cris.virtualsource.orcid6c1aac4b-593e-4f80-9ecc-911fd20f3c31
cris.virtualsource.orcid952936d6-a5d1-4952-ab8c-7ba5c377af16
dc.contributor.authorHou, Yuanbo
dc.contributor.authorRen, Qiaoqiao
dc.contributor.authorMitchell, Andrew
dc.contributor.authorWang, Wenwu
dc.contributor.authorKang, Jian
dc.contributor.authorBelpaeme, Tony
dc.contributor.authorBotteldooren, Dick
dc.date.accessioned2026-09-09T08:14:10Z
dc.date.available2026-09-09T08:14:10Z
dc.date.createdwos2026-03-28
dc.date.issued2026
dc.description.abstractWe live in a rich and varied acoustic world, which is experienced by individuals or communities as a soundscape. Computational auditory scene analysis, disentangling acoustic scenes by detecting and classifying events, focuses on objective attributes of sounds, such as their category and temporal characteristics, ignoring their effects on people, such as the emotions they evoke within a context. To fill this gap, we propose the affective soundscape captioning (ASSC) task, which enables automated soundscape analysis, thus avoiding labour-intensive subjective ratings and surveys in conventional methods. With soundscape captioning, context-aware descriptions are generated for soundscape by capturing the acoustic scenes (ASs), audio events (AEs) information, and the corresponding human affective qualities (AQs). To this end, we propose an automatic soundscape captioner (SoundSCaper) system composed of an acoustic model, i.e. SoundAQnet, and a large language model (LLM). SoundAQnet simultaneously models multi-scale information about ASs, AEs, and perceived AQs, while the LLM describes the soundscape with captions by parsing the information captured with SoundAQnet. SoundSCaper is assessed by two juries of 32 people. In expert evaluation, the average score of SoundSCaper-generated captions is slightly lower than that of two soundscape experts on the evaluation set D1 and the external mixed dataset D2, but not statistically significant. In layperson evaluation, SoundSCaper outperforms soundscape experts in several metrics on datasets D1 and D2. In addition to human evaluation, compared to other automated audio captioning (AAC) systems with and without LLM, SoundSCaper performs better on the ASSC task in several natural language processing (NLP) based metrics. Overall, SoundSCaper performs well in human subjective evaluation and various objective captioning metrics, and the generated captions are comparable to those annotated by soundscape experts. The model, source code, LLM scripts, human assessment data, instructions, and evaluation statistics are all publicly available.
dc.description.wosFundingTextThis work was funding by Flemish Government under the "Onderzoeksprogramma Artificiele Intelligentie (AI) Vlaanderen" programme.
dc.identifier.doi10.1109/tmm.2026.3651023
dc.identifier.eissn1941-0077
dc.identifier.issn1520-9210
dc.identifier.urihttps://imec-publications.be/handle/20.500.12860/60286
dc.language.isoeng
dc.provenance.editstepusergreet.vanhoof@imec.be
dc.publisherIEEE-INST ELECTRICAL ELECTRONICS ENGINEERS INC
dc.source.beginpage2186
dc.source.endpage2200
dc.source.journalIEEE TRANSACTIONS ON MULTIMEDIA
dc.source.numberofpages15
dc.source.volume28
dc.subject.keywordsCLASSIFICATION
dc.title

Soundscape Captioning Using Sound Affective Quality Network and Large Language Model

dc.typeJournal article
dspace.entity.typePublication
imec.internal.crawledAt2026-04-07
imec.internal.sourcecrawler
imec.internal.wosCreatedAt2026-04-07
Files
Publication available in collections: