[1]

Yadav A, Vishwakarma DK. 2020. A deep learning architecture of RA-DLNet for visual sentiment analysis. Multimedia Systems 26:431−451

doi: 10.1007/s00530-020-00656-7
[2]

Das R, Singh TD. 2023. Multimodal sentiment ana lysis: a survey of methods, trends, and challenges. ACM Computing Surveys 55:1−38

doi: 10.1145/3586075
[3]

Wang L, Niu J, Yu S. 2020. SentiDiff: Combining textual information and sentiment diffusion patterns for Twitter sentiment analysis. IEEE Transactions on Knowledge and Data Engineering 32:2026−2039

doi: 10.1109/tkde.2019.2913641
[4]

Zhao S, Yao X, Yang J, Jia G, Ding G, et al. 2022. Affective image content analysis: Two decades review and new perspectives. IEEE Transactions on Pattern Analysis and Machine Intelligence 44:6729−6751

doi: 10.1109/TPAMI.2021.3094362
[5]

Xu N, Mao W. 2017. MultiSentiNet: a deep semantic network for multimodal sentiment analysis. CIKM '17: Proceedings of the 2017 ACM on Conference on Information and Knowledge Management, November 6−10, 2017, Singapore. USA: ACM. pp. 2399−2402 doi: 10.1145/3132847.3133142

[6]

Wang M, Meng M, Liu J, Wu J. 2023. Adequate alignment and interaction for cross-modal retrieval. Virtual Reality & Intelligent Hardware 5:509−522

doi: 10.1016/j.vrih.2023.06.003
[7]

Xu N, Mao W, Chen G. 2018. A co-memory network for multimodal sentiment analysis. SIGIR '18: The 41st international ACM SIGIR conference on research & development in information retrieval, July 8−12, 2018, Ann Arbor, MI, USA. pp. 929−932 doi: 10.1145/3209978.3210093

[8]

Xu N. 2017. Analyzing multimodal public sentiment based on hierarchical semantic attentional network. 2017 IEEE international conference on intelligence and security informatics (ISI), July 22−24, 2017, Beijing, China. USA: IEEE. pp. 152−154 doi: 10.1109/ISI.2017.8004895

[9]

Chen F, Ji R, Su J, Cao D, Gao Y. 2018. Predicting microblog sentiments via weakly supervised multimodal deep learning. IEEE Transactions on Multimedia 20:997−1007

doi: 10.1109/TMM.2017.2757769
[10]

Dai S, Man H. 2018. Integrating visual and textual affective descriptors for sentiment analysis of social media posts. 2018 IEEE Conference on Multimedia Information Processing and Retrieval (MIPR), April 10−12, 2018, Miami, FL, USA. USA: IEEE. pp. 13−18 doi: 10.1109/MIPR.2018.00011

[11]

Zhao Z, Zhu H, Xue Z, Liu Z, Tian J, et al. 2019. An image-text consistency driven multimodal sentiment analysis approach for social media. Information Processing & Management 56:102097

doi: 10.1016/j.ipm.2019.102097
[12]

Zhu T, Li L, Yang J, Zhao S, Liu H, et al. 2023. Multimodal sentiment analysis with image-text interaction network. IEEE transactions on multimedia 25:3375−3385

doi: 10.1109/TMM.2022.3160060
[13]

Xie Z, Zhang W, Sheng B, Li P, Chen CLP. 2021. BaGFN: broad attentive graph fusion network for high-order feature interactions. IEEE Transactions on Neural Networks and Learning Systems 34:4499−4513

doi: 10.1109/tnnls.2021.3116209
[14]

Houlsby N, Giurgiu A, Jastrzebski S, Morrone B, De Laroussilhe Q, et al. 2019. Parameter-efficient transfer learning for NLP. Proceedings of the 36th International Conference on Machine Learning (ICML), June 9−15, 2019, Long Beach, CA, USA. USA: PMLR. pp. 2790−2799 https://proceedings.mlr.press/v97/houlsby19a.html

[15]

Hu EJ, Shen Y, Wallis P, Allen-Zhu Z, Li Y, et al. 2022. LoRA: low-rank adaptation of large language models. The Tenth International Conference on Learning Representations (ICLR), April 25−29, 2022, Virtual Event. OpenReview.net. https://openreview.net/forum?id=nZeVKeeFYf9

[16]

Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, et al. 2021. Learning transferable visual models from natural language supervision. Proceedings of the 38th International Conference on Machine Learning (ICML), July 18−24, 2021, Virtual Event. USA: PMLR. pp. 8748−8763 https://proceedings.mlr.press/v139/radford21a.html

[17]

Chen Z, Xu L, Zheng H, Chen L, Tolba A, et al. 2024. Evolution and prospects of foundation models: from large language models to large multimodal models. Computers, Materials & Continua 80(2):1753−1808

doi: 10.32604/cmc.2024.052618
[18]

Zhao X, Poria S, Li X, Chen Y, Tang B. 2025. Toward robust multimodal sentiment analysis using multimodal foundational models. Expert Systems with Applications 276:126974

doi: 10.1016/j.eswa.2025.126974
[19]

Mu J, Wang W, Liu W, Yan T, Wang G. 2025. Multimodal large language model with lora fine-tuning for multimodal sentiment analysis. ACM Transactions on Intelligent Systems and Technology 16:1−23

doi: 10.1145/3709147
[20]

Vinyals O, Toshev A, Bengio S, Erhan D. 2015. Show and tell: a neural image caption generator. 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 7−12, 2015, Boston, MA, USA. USA: IEEE. pp. 3156−3164 doi: 10.1109/CVPR.2015.7298935

[21]

Li J, Li D, Savarese S, Hoi S. 2023. BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. ICML'23: Proceedings of the 40th International Conference on Machine Learning (ICML), July 23−29, 2023, Honolulu, HI, USA. USA: PMLR. pp. 19730−19742 doi: 10.5555/3618408.3619222

[22]

Hu X, Gan Z, Wang J, Yang Z, Liu Z, et al. 2022. Scaling up vision-language pretraining for image captioning. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 18−24 June 2022, New Orleans, LA, USA. USA: IEEE. pp. 17980−17989 doi: 10.1109/CVPR52688.2022.01745

[23]

Achiam J, Adler S, Agarwal S, Ahmad L, Akkaya I, et al. 2023. GPT-4 technical report. arXiv Preprint:2303.08774

doi: 10.48550/arXiv.2303.08774
[24]

Li Z, Xu B, Zhu C, Zhao T. 2022. CLMLF: a contrastive learning and multi-layer fusion method for multimodal sentiment detection. Findings of the association for computational linguistics: NAACL 2022, Seattle, United States. USA: ACM. pp. 2282-2294 doi: 10.18653/v1/2022.findings-naacl.175

[25]

Zhang S, Liu J, Jiao Y, Zhang Y, Chen L, et al. 2025. A multimodal semantic fusion network with cross-modal alignment for multimodal sentiment analysis. ACM Transactions on Multimedia Computing, Communications and Applications 21:1−22

doi: 10.1145/3744648
[26]

Deng J, Guo J, Yang J, Xue N, Kotsia I, Zafeiriou S. 2022. ArcFace: additive angular margin loss for deep face recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence 44:5962−5979

doi: 10.1109/TPAMI.2021.3087709
[27]

Wang Q, Wen Z, Ding K, Liang B, Xu R. 2025. Cross-domain sentiment analysis via disentangled representation and prototypical learning. IEEE Transactions on Affective Computing 16:264−276

doi: 10.1109/taffc.2024.3431946
[28]

Zhuang Y, Bai W, Zhang Y, Deng J, Hu Z, et al. 2025. Multi-level contrastive learning for multimodal sentiment analysis. IEEE Transactions on Multimedia. 27:9044−9058

doi: 10.1109/TMM.2025.3613116
[29]

Russell JA. 1980. A circumplex model of affect. Journal of Personality and Social Psychology 39:1161−1178

doi: 10.1037/h0077714
[30]

Niu T, Zhu S, Pang L, El Saddik A. 2016. Sentiment analysis on multi-view social data. MultiMedia Modeling. MMM 2016. Lecture Notes in Computer Science. Cham: Springer. pp. 15−27 doi: 10.1007/978-3-319-27674-8_2

[31]

Kim Y. 2014. Convolutional neural networks for sentence classification. Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), 2014, Doha, Qatar. Stroudsburg, PA, USA: Association for Computational Linguistics. pp. 1746−1751 doi: 10.3115/v1/d14-1181

[32]

Devlin J, Chang MW, Lee K, Toutanova K. 2019. BERT: pre-training of deep bidirectional transformers for language understanding. Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1(long and short papers), 2019, Minneapolis, Minnesota. Stroudsburg, PA, USA: Association for Computational Linguistics. pp. 4171−4186 doi: 10.18653/v1/n19-1423

[33]

He K, Zhang X, Ren S, Sun J. 2016. Deep residual learning for image recognition. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 27−30 June 2016, Las Vegas, NV, USA. USA: IEEE. pp. 770−778 doi: 10.1109/CVPR.2016.90

[34]

Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Zhai X, et al. 2021. An image is worth 16×16 words: transformers for image recognition at scale. The Ninth International Conference on Learning Representations (ICLR), May 3−7, 2021, Virtual Event. OpenReview.net. https://openreview.net/forum?id=YicbFdNTTy

[35]

Wang A, Chen H, Lin Z, Han J, Ding G. 2024. Rep Vit: revisiting mobile cnn from vit perspective. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 16−22 June 2024, Seattle, WA, USA. USA: IEEE. pp. 15909−15920 doi: 10.1109/CVPR52733.2024.01506

[36]

Shi D. 2024. TransNeXt: robust foveal visual perception for vision transformers. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 16−22 June 2024, Seattle, WA, USA. USA: IEEE. pp. 17773−17783 doi: 10.1109/CVPR52733.2024.01683

[37]

Yang X, Feng S, Wang D, Zhang Y. 2020. Image-text multimodal emotion classification via multi-view attentional network. IEEE Transactions on Multimedia 23:4014−4026

doi: 10.1109/TMM.2020.3035277
[38]

Yang X, Feng S, Zhang Y, Wang D. 2021. Multimodal sentiment detection based on multi-channel graph neural networks. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), 2021, Online. Stroudsburg, PA, USA: Association for Computational Linguistics. pp. 328−339 doi: 10.18653/v1/2021.acl-long.28

[39]

Wei Y, Yuan S, Yang R, Shen L, Li Z, et al. 2023. Tackling modality heterogeneity with multi-view calibration network for multimodal sentiment detection. Proceedings of the 61st annual meeting of the Association for Computational Linguistics (volume 1: Long papers), 2023, Toronto, Canada. Stroudsburg, PA, USA: Association for Computational Linguistics. pp. 5240−5252 doi: 10.18653/v1/2023.acl-long.287

[40]

An J, Wan Zainon WMN. 2023. Integrating color cues to improve multimodal sentiment analysis in social media. Engineering Applications of Artificial Intelligence 126:106874

doi: 10.1016/j.engappai.2023.106874