Text to FigureText → Imageacademic

PRISM Framework: Multimodal Sentiment Overview

Figure 1. Overview of the PRISM framework. The left panel illustrates the multimodal encoding stage, where visual, acoustic, and textual inputs are independently encoded into modality-specific feature sequences. The middle panel presents the sentiment prototype bank (SPB) and prototype-conditioned modality selection where shared sentiment prototypes interact with each modality through cross-attention to extract prototype-aligned unimodal responses, which are then adaptively combined by predicting prototype-wise modality weights to form the fused prototype sequence. The right panel shows the Tr

Paper context

Paper title: Learning Shared Sentiment Prototypes for Adaptive Multimodal Sentiment Analysis Abstract: Multimodal sentiment analysis (MSA) aims to predict human sentiment from textual, acoustic, and visual information in videos. Recent studies improve multimodal fusion by modeling modality interaction and assigning different modality weights. However, they usually compress diverse sentiment cues into a single compact representation before sentiment reasoning. This early aggregation makes it difficult to preserve the internal structure of sentiment evidence, where different cues may complement, conflict with, or differ in reliability from each other. In addition, modality importance is often determined only once during fusion, so later reasoning cannot further adjust modality contributions. To address these issues, we propose PRISM, a framework that unifies structured affective extraction and adaptive modality evaluation. PRISM organizes multimodal evidence in a shared prototype space, which supports structured cross-modal comparison and adaptive fusion. It further applies dynamic modality reweighting during reasoning, allowing modality contributions to be continuously refined as semantic interactions become deeper. Experiments on three benchmark datasets show that PRISM outperforms representative baselines. Passages referencing this figure: and adaptive fusion. It further applies dynamic modality reweighting during reasoning, allowing modality contributions to be continuously refined as semantic interactions become deeper. Experiments on three benchmark datasets show that PRISM outperforms representative baselines. 1 1 1 The code is available at https://github.com/synlp/PRISM . 1 1 footnotetext: Corresponding author. 1. Introduction Figure 1. Overview of the PRISM framework. The left panel illustrates the multimodal encoding stage, where visual, acoustic, and textual inputs are independently encoded into modality-specific feature sequences. The middle panel presents the sentiment prototype bank (SPB) and prototype-conditioned modality selection where shared sentiment prototypes interact with each modality through cross-attent aring sentiment evidence across modalities. In contrast, PRISM applies the same shared prototype bank to text, audio, and visual modalities, so prototype responses are aligned by slot and directly comparable across modalities. This shared structure enables the prototypes to support both structured affective extraction and modality evaluation within a unified mechanism. 3. The Approach As shown in Fig. 1 , given multimodal input 𝒳 = { 𝐗 v , 𝐗 a , 𝐗 t } \mathcal{X}=\{\mathbf{X}^{\mathrm{v}},\mathbf{X}^{\mathrm{a}},\mathbf{X}^{\mathrm{t}}\} , the PRISM framework sequentially performs modality encoding, prototype-driven extraction based on sentiment prototype bank (SPB), prototype-conditioned selection (PCS), and dynamic modality-reweighted reasoning (DMR) to produce a sentiment intensity pred

The prompt

Above I've shared:
(1) the paper title + abstract + method section,
(2) the figure caption I want.

TASK: Render the main figure for this academic paper. Style requirements:

  - This is an ACADEMIC PAPER FIGURE (not a poster, not an infographic).
  - Clean black-on-white background; minimal decoration.
  - Components, arrows, and labels rendered crisply; small dense text OK.
  - Single-figure layout — no banner header, no "title" inside the image.
  - Match the level of detail of a top-tier conference paper figure
    (NeurIPS / ICLR / CVPR style).

Render the figure described in the caption. Just give me the final image.

Try this prompt now

Open it inside the generator with the prompt pre-filled.

Try this prompt

Related prompts