Multi-Modal Semantic Expansion with Constrained LLM Reranking for Conversational Music Recommendation
2026-08-24 • Artificial Intelligence
Artificial Intelligence
AI summaryⓘ
The authors describe their approach to a music recommendation challenge where the system suggests songs in conversations. Their method uses multiple types of information like track features, user behavior, lyrics, audio, and images, combined smartly to find good song matches. They then refine these matches and generate personalized replies using a language model. They also tested other complex methods but did not use them in the final submission due to cost and complexity. Their solution improved ranking accuracy and performed reasonably in the competition.
conversational recommender systemmulti-modal retrievalembedding spaceCollaborative Filtering (CF-BPR)Reciprocal Rank Fusion (RRF)rerankingGPT-4o-mininDCGdifferential evolutionmusic recommendation
Authors
Naman Garg, Sarika Jain, George Fazekas
Abstract
We present Team Semiintelligencn's solution for the ACM RecSys 2026 TalkPlayData Challenge, addressing conversational music recommendation through a multi-modal and personalized conversational recommender system. Our submitted system employs a three-stage pipeline: (1) multi-modal retrieval constructing decay-weighted centroids across seven dense embedding spaces - track- and user-level CF-BPR, Qwen3 (metadata, lyrics, attributes), CLAP audio, and SigLIP visual - supplemented by BM25 lexical retrieval and an artist substring-match signal, all fused via weighted Reciprocal Rank Fusion (RRF) with optimized signal weights; (2) lightweight reranking (history filtering, popularity smoothing, and catalog diversity penalization); and (3) persona-diversified response generation using GPT-4o-mini. Beyond this submitted configuration, we report development-time experiments with additional components - constrained LLM-guided artist injection, album continuation signals, XGBoost LambdaMART, and a superior GPT-4.1 response prompt - that were not deployed to Blind B due to cost and complexity constraints. We optimize RRF weights on a 500-session development split via differential evolution, improving MRR by +19.5%. On Blind A, we observe that unconstrained LLM-guided injection across 54 sessions causes catastrophic nDCG regression (-18.9%), while conservative injection on only 9 sessions yields the best observed Blind A nDCG - a finding we present as a Blind A observation warranting further validation. The submitted system achieves a Blind B composite score of 0.3213.