| [1] |
Yadav A, Vishwakarma DK. 2020. A deep learning architecture of RA-DLNet for visual sentiment analysis. |
| [2] |
Das R, Singh TD. 2023. Multimodal sentiment ana lysis: a survey of methods, trends, and challenges. |
| [3] |
Wang L, Niu J, Yu S. 2020. SentiDiff: Combining textual information and sentiment diffusion patterns for Twitter sentiment analysis. |
| [4] |
Zhao S, Yao X, Yang J, Jia G, Ding G, et al. 2022. Affective image content analysis: Two decades review and new perspectives. |
| [5] |
Xu N, Mao W. 2017. MultiSentiNet: a deep semantic network for multimodal sentiment analysis. CIKM '17: Proceedings of the 2017 ACM on Conference on Information and Knowledge Management, November 6−10, 2017, Singapore. USA: ACM. pp. 2399−2402 doi: 10.1145/3132847.3133142 |
| [6] |
Wang M, Meng M, Liu J, Wu J. 2023. Adequate alignment and interaction for cross-modal retrieval. |
| [7] |
Xu N, Mao W, Chen G. 2018. A co-memory network for multimodal sentiment analysis. SIGIR '18: The 41st international ACM SIGIR conference on research & development in information retrieval, July 8−12, 2018, Ann Arbor, MI, USA. pp. 929−932 doi: 10.1145/3209978.3210093 |
| [8] |
Xu N. 2017. Analyzing multimodal public sentiment based on hierarchical semantic attentional network. 2017 IEEE international conference on intelligence and security informatics (ISI), July 22−24, 2017, Beijing, China. USA: IEEE. pp. 152−154 doi: 10.1109/ISI.2017.8004895 |
| [9] |
Chen F, Ji R, Su J, Cao D, Gao Y. 2018. Predicting microblog sentiments via weakly supervised multimodal deep learning. |
| [10] |
Dai S, Man H. 2018. Integrating visual and textual affective descriptors for sentiment analysis of social media posts. 2018 IEEE Conference on Multimedia Information Processing and Retrieval (MIPR), April 10−12, 2018, Miami, FL, USA. USA: IEEE. pp. 13−18 doi: 10.1109/MIPR.2018.00011 |
| [11] |
Zhao Z, Zhu H, Xue Z, Liu Z, Tian J, et al. 2019. An image-text consistency driven multimodal sentiment analysis approach for social media. |
| [12] |
Zhu T, Li L, Yang J, Zhao S, Liu H, et al. 2023. Multimodal sentiment analysis with image-text interaction network. |
| [13] |
Xie Z, Zhang W, Sheng B, Li P, Chen CLP. 2021. BaGFN: broad attentive graph fusion network for high-order feature interactions. |
| [14] |
Houlsby N, Giurgiu A, Jastrzebski S, Morrone B, De Laroussilhe Q, et al. 2019. Parameter-efficient transfer learning for NLP. Proceedings of the 36th International Conference on Machine Learning (ICML), June 9−15, 2019, Long Beach, CA, USA. USA: PMLR. pp. 2790−2799 https://proceedings.mlr.press/v97/houlsby19a.html |
| [15] |
Hu EJ, Shen Y, Wallis P, Allen-Zhu Z, Li Y, et al. 2022. LoRA: low-rank adaptation of large language models. The Tenth International Conference on Learning Representations (ICLR), April 25−29, 2022, Virtual Event. OpenReview.net. https://openreview.net/forum?id=nZeVKeeFYf9 |
| [16] |
Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, et al. 2021. Learning transferable visual models from natural language supervision. Proceedings of the 38th International Conference on Machine Learning (ICML), July 18−24, 2021, Virtual Event. USA: PMLR. pp. 8748−8763 https://proceedings.mlr.press/v139/radford21a.html |
| [17] |
Chen Z, Xu L, Zheng H, Chen L, Tolba A, et al. 2024. Evolution and prospects of foundation models: from large language models to large multimodal models. |
| [18] |
Zhao X, Poria S, Li X, Chen Y, Tang B. 2025. Toward robust multimodal sentiment analysis using multimodal foundational models. |
| [19] |
Mu J, Wang W, Liu W, Yan T, Wang G. 2025. Multimodal large language model with lora fine-tuning for multimodal sentiment analysis. |
| [20] |
Vinyals O, Toshev A, Bengio S, Erhan D. 2015. Show and tell: a neural image caption generator. 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 7−12, 2015, Boston, MA, USA. USA: IEEE. pp. 3156−3164 doi: 10.1109/CVPR.2015.7298935 |
| [21] |
Li J, Li D, Savarese S, Hoi S. 2023. BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. ICML'23: Proceedings of the 40th International Conference on Machine Learning (ICML), July 23−29, 2023, Honolulu, HI, USA. USA: PMLR. pp. 19730−19742 doi: 10.5555/3618408.3619222 |
| [22] |
Hu X, Gan Z, Wang J, Yang Z, Liu Z, et al. 2022. Scaling up vision-language pretraining for image captioning. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 18−24 June 2022, New Orleans, LA, USA. USA: IEEE. pp. 17980−17989 doi: 10.1109/CVPR52688.2022.01745 |
| [23] |
Achiam J, Adler S, Agarwal S, Ahmad L, Akkaya I, et al. 2023. GPT-4 technical report. |
| [24] |
Li Z, Xu B, Zhu C, Zhao T. 2022. CLMLF: a contrastive learning and multi-layer fusion method for multimodal sentiment detection. Findings of the association for computational linguistics: NAACL 2022, Seattle, United States. USA: ACM. pp. 2282-2294 doi: 10.18653/v1/2022.findings-naacl.175 |
| [25] |
Zhang S, Liu J, Jiao Y, Zhang Y, Chen L, et al. 2025. A multimodal semantic fusion network with cross-modal alignment for multimodal sentiment analysis. |
| [26] |
Deng J, Guo J, Yang J, Xue N, Kotsia I, Zafeiriou S. 2022. ArcFace: additive angular margin loss for deep face recognition. |
| [27] |
Wang Q, Wen Z, Ding K, Liang B, Xu R. 2025. Cross-domain sentiment analysis via disentangled representation and prototypical learning. |
| [28] |
Zhuang Y, Bai W, Zhang Y, Deng J, Hu Z, et al. 2025. Multi-level contrastive learning for multimodal sentiment analysis. |
| [29] |
Russell JA. 1980. A circumplex model of affect. |
| [30] |
Niu T, Zhu S, Pang L, El Saddik A. 2016. Sentiment analysis on multi-view social data. MultiMedia Modeling. MMM 2016. Lecture Notes in Computer Science. Cham: Springer. pp. 15−27 doi: 10.1007/978-3-319-27674-8_2 |
| [31] |
Kim Y. 2014. Convolutional neural networks for sentence classification. Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP), 2014, Doha, Qatar. Stroudsburg, PA, USA: Association for Computational Linguistics. pp. 1746−1751 doi: 10.3115/v1/d14-1181 |
| [32] |
Devlin J, Chang MW, Lee K, Toutanova K. 2019. BERT: pre-training of deep bidirectional transformers for language understanding. Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1(long and short papers), 2019, Minneapolis, Minnesota. Stroudsburg, PA, USA: Association for Computational Linguistics. pp. 4171−4186 doi: 10.18653/v1/n19-1423 |
| [33] |
He K, Zhang X, Ren S, Sun J. 2016. Deep residual learning for image recognition. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 27−30 June 2016, Las Vegas, NV, USA. USA: IEEE. pp. 770−778 doi: 10.1109/CVPR.2016.90 |
| [34] |
Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Zhai X, et al. 2021. An image is worth 16×16 words: transformers for image recognition at scale. The Ninth International Conference on Learning Representations (ICLR), May 3−7, 2021, Virtual Event. OpenReview.net. https://openreview.net/forum?id=YicbFdNTTy |
| [35] |
Wang A, Chen H, Lin Z, Han J, Ding G. 2024. Rep Vit: revisiting mobile cnn from vit perspective. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 16−22 June 2024, Seattle, WA, USA. USA: IEEE. pp. 15909−15920 doi: 10.1109/CVPR52733.2024.01506 |
| [36] |
Shi D. 2024. TransNeXt: robust foveal visual perception for vision transformers. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 16−22 June 2024, Seattle, WA, USA. USA: IEEE. pp. 17773−17783 doi: 10.1109/CVPR52733.2024.01683 |
| [37] |
Yang X, Feng S, Wang D, Zhang Y. 2020. Image-text multimodal emotion classification via multi-view attentional network. |
| [38] |
Yang X, Feng S, Zhang Y, Wang D. 2021. Multimodal sentiment detection based on multi-channel graph neural networks. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), 2021, Online. Stroudsburg, PA, USA: Association for Computational Linguistics. pp. 328−339 doi: 10.18653/v1/2021.acl-long.28 |
| [39] |
Wei Y, Yuan S, Yang R, Shen L, Li Z, et al. 2023. Tackling modality heterogeneity with multi-view calibration network for multimodal sentiment detection. Proceedings of the 61st annual meeting of the Association for Computational Linguistics (volume 1: Long papers), 2023, Toronto, Canada. Stroudsburg, PA, USA: Association for Computational Linguistics. pp. 5240−5252 doi: 10.18653/v1/2023.acl-long.287 |
| [40] |
An J, Wan Zainon WMN. 2023. Integrating color cues to improve multimodal sentiment analysis in social media. |