Kolluri, Ganesh Pavan Kartikeya Bharadwaj and Zhang, Yuchen and Kampouridis, Michael and Shekhar, Ravi (2026) CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning. In: IEEE Spoken Language Technology SLT 2026, 2026-12-13 - 2026-09-16, Palermo, Sicily, Italy. (In Press)
Kolluri, Ganesh Pavan Kartikeya Bharadwaj and Zhang, Yuchen and Kampouridis, Michael and Shekhar, Ravi (2026) CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning. In: IEEE Spoken Language Technology SLT 2026, 2026-12-13 - 2026-09-16, Palermo, Sicily, Italy. (In Press)
Kolluri, Ganesh Pavan Kartikeya Bharadwaj and Zhang, Yuchen and Kampouridis, Michael and Shekhar, Ravi (2026) CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning. In: IEEE Spoken Language Technology SLT 2026, 2026-12-13 - 2026-09-16, Palermo, Sicily, Italy. (In Press)
Abstract
Modern automated audio captioning systems paira frozen audio encoder with a large language model (LLM)via a trainable projector, incurring the encoder’s inference costand bottlenecking the model through its fixe acoustic features.We present CARD, an encoder-free audio captioning model thatremoves the encoder at inference: a 13.2M projector feeds afrozen LLM with merged LoRA adapters, while the teacher usedto train it is discarded. CARD distills a pretrained audio teacher(CLAP-HTSAT) into the model, but rather than injecting it intothe LLM alone, it routes the teacher’s representations acrosscomponents: perceptual stages to the projector and semanticstages to the LLM. This placement improves CIDEr-D by +12.18over an LLM-only distilled model on AudioCaps and by +5.21 onClotho, reaching 55.4 against a 66.4 encoder-kept upper boundwith no encoder at inference, showing that where a teacher’sknowledge is placed matters as much as its presence.
| Item Type: | Conference or Workshop Item (Paper) |
|---|---|
| Additional Information: | Published proceedings: _not provided_ |
| Uncontrolled Keywords: | audio captioning, encoder-free models, cross-component distillation, audio-language models |
| Subjects: | Z Bibliography. Library Science. Information Resources > ZR Rights Retention |
| Divisions: | Faculty of Science and Health > Computer Science and Electronic Engineering, School of |
| SWORD Depositor: | Unnamed user with email elements@essex.ac.uk |
| Depositing User: | Unnamed user with email elements@essex.ac.uk |
| Date Deposited: | 23 Sep 2026 11:45 |
| Last Modified: | 23 Sep 2026 11:47 |
| URI: | http://repository.essex.ac.uk/id/eprint/43881 |
Available files
Filename: CARD.pdf
Licence: Creative Commons: Attribution 4.0