Figures (8)  Tables (4)
    • Figure 1. 

      Examples of the affective gap.

    • Figure 2. 

      Examples of image-text pairs collected from Twitter.

    • Figure 3. 

      Examples illustrating the unreliability of unimodal inference when text is dominated by subjective expressions or linguistic noise. The structured visual context provides the necessary semantic grounding for accurate sentiment prediction.

    • Figure 4. 

      The overall architecture of GSKA. The pipeline transitions from sensory encoding to knowledge-driven reasoning, culminating in an affective manifold mapping for sentiment prediction.

    • Figure 5. 

      Visualization of the generative semantic anchoring process. High-level affective indicators are extracted to serve as stable semantic coordinates ($ {\bf{S}} $).

    • Figure 6. 

      Detailed mechanism of the Bi-directional Knowledge Alignment (BKA) module. The Semantic Anchor ($ {\bf{S}} $) modulates the cross-attention flow between vision ($ {\bf{P}} $) and text ($ {\bf{T}} $).

    • Figure 7. 

      t-SNE visualization of multimodal embeddings for different variants: (a) Standard Cross-Entropy; (b) Contrastive Loss; and (c) Emo-Loss (full GSKA).

    • Figure 8. 

      Visualization of the attention mechanism, highlighting aligned image regions and textual tokens.

    • Input: Fused multi-modal feature $ M_{\text{fused}} $, target label $ y $, Prototype set $ {\bf{C}} $, scaling factor $ \beta $
      Output: Emotional discrepancy loss $ {\cal{L}}_{Emo} $
      1. Similarity Calculation: Compute cosine similarities for all classes: $ s_j = {\text{cosine}}\_{\text{sim}}(M_{\text{fused}}, {\bf{c}}_j) $ for $ j \in \{pos, neu, neg\} $;
      2. Target Identification: Retrieve target similarity: $ s_{target} = s_y $;
      3. Competitor Mining: Find maximum confounding similarity: $ s_{other} = \max_{j \neq y} (s_j) $;
      4. Margin Optimization: Calculate similarity margin: $ \Delta = s_{target} - s_{other} $;
      5. Non-linear Mapping: Normalize margin via scaled sigmoid: $ \text{diff} = 2 \cdot \sigma(\beta \cdot \Delta) - 1 $;
      6. return $ {\cal{L}}_{Emo} = 1 - \text{diff} $;

      Table 1. 

      Emo-Loss computation.

    • DatasetPositiveNeutralNegativeTotal
      MVSA-single2,6834701,3584,511
      MVSA-multiple11,3184,4081,29817,024

      Table 1. 

      Statistics of the datasets.

    • ModalityMethodsMVSA-singleMVSA-multiple
      AccuracyF1AccuracyF1
      TextCNN68.1955.9065.6457.66
      BERT71.1169.7067.5966.24
      ImageResNet5064.6761.5561.8860.98
      ViT63.7862.2661.9461.19
      RepViT64.8262.8162.0161.35
      TransNeXt65.7164.8162.4861.52
      MultimodalHSAN69.8866.9067.9667.76
      MultiSentiNet69.8469.6368.8668.11
      Co-MN-Hop670.5170.0168.9268.83
      MVAN-M72.9872.9872.3672.30
      MGNNS73.7772.7072.4969.34
      CLMLF75.3373.4672.0069.83
      MVCN76.0674.5572.0770.01
      ICCI79.3377.5173.2970.06
      MSFN78.9878.4874.7572.62
      GSKA (ours)80.9380.1075.0574.57

      Table 2. 

      Performance comparison on MVSA datasets. Accuracy and F1-score are reported as a percentage (%).

    • Variants Accuracy (%) $ \Delta $Acc F1 (%) $ \Delta $F1
      GSKA (full model) 80.93 – 80.10 –
      w/o Co-Attention 77.42 $ \downarrow $3.51 76.88 $ \downarrow $3.22
      w/o Captioning 78.08 $ \downarrow $2.85 78.27 $ \downarrow $1.83
      w/o text mediation (Direct Alignment) 77.92 $ \downarrow $3.01 77.45 $ \downarrow $2.65
      w/o Emo-Loss (w/ Contrastive Loss) 76.77 $ \downarrow $4.16 76.05 $ \downarrow $4.05
      w/o Emo-Loss (w/ Cross-Entropy) 78.32 $ \downarrow $2.61 77.79 $ \downarrow $2.31
      $ \downarrow $ denotes the performance drop relative to the full model.

      Table 3. 

      Ablation results on the MVSA-single dataset.