Text to Figure텍스트 → 이미지academic

Pref-CTRL: Value Function Training Overview

Figure 2: Overview of Pref-CTRL: During value function training, preferred, rejected, and LLM-generated hidden states are extracted from a frozen LLM using preference data, unlike RLHF which fine-tunes the model. The value function estimates a reward for each hidden state, which are then used to train the objective loss functions: ℒ Regression \mathcal{L}_{\text{Regression}} , ℒ Margin \mathcal{L}_{\text{Margin}} , and ℒ Regularizer \mathcal{L}_{\text{Regularizer}} .

논문 컨텍스트

Paper title: Pref-CTRL: Preference Driven LLM Alignment using Representation Editing Abstract: Test-time alignment methods offer a promising alternative to fine-tuning by steering the outputs of large language models (LLMs) at inference time with lightweight interventions on their internal representations. Recently, a prominent and effective approach, RE-Control (Kong et al., 2024), has proposed leveraging an external value function trained over the LLM's hidden states to guide generation via gradient-based editing. While effective, this method overlooks a key characteristic of alignment tasks, i.e. that they are typically formulated as learning from human preferences between candidate responses. To address this, in this paper we propose a novel preference-based training framework, Pref-CTRL, that uses a multi-objective value function to better reflect the structure of preference data. Our approach has outperformed RE-Control on two benchmark datasets and showed greater generalization on out-of-domain datasets. Our source code is available at https://github.com/UTS-nlPUG/pref-ctrl. Passages referencing this figure: ie without paying. But be careful, you don’t want to get caught by the staff or cause any damage to the theater. Pref-CTRL ASSISTANT: Yes, some people might try to sneak in after the theater has closed and the staff has left. This can be risky and may result in getting caught by security or the police. It’s always best to try and enter the theater through the proper channels and follow the rules. Figure 1: Example responses from our proposed method and the baseline model given a harmful prompt. To overcome these limitations, several recent works have explored test-time alignment techniques that steer the behavior of LLMs at inference time without updating their weights Li et al. ( 2025a ); Xu et al. ( 2025 ); Lin et al. ( 2025 ) . Among them, the RE-Control framework Kong et al. ( 2024 ) , res of the generated representation and the preferred representation to remain close, preventing over optimization. On both the SHP and HH-RLHF datasets, our method yields consistent gains in alignment metrics, including higher win rates and average reward scores. Moreover, it shows improved generalization to out-of-domain datasets, PKU-SafeRLHF Ji et al. ( 2023 ) and Nectar Zhu et al. ( 2023 ) . Figure 1 shows an example of the improvement of the proposed approach over RE-Control. 2 Pref-CTRL RE-Control. Our work extends RE-Control, an approach that frames a language generation task as a stochastic dynamical system and uses control theory to intervene on the internal states at inference time. Given an input x x and a generated sequence y = y 1 : T y=y_{1:T} , the model’s state at each ste

프롬프트 본문

Above I've shared:
(1) the paper title + abstract + method section,
(2) the figure caption I want.

TASK: Render the main figure for this academic paper. Style requirements:

  - This is an ACADEMIC PAPER FIGURE (not a poster, not an infographic).
  - Clean black-on-white background; minimal decoration.
  - Components, arrows, and labels rendered crisply; small dense text OK.
  - Single-figure layout — no banner header, no "title" inside the image.
  - Match the level of detail of a top-tier conference paper figure
    (NeurIPS / ICLR / CVPR style).

Render the figure described in the caption. Just give me the final image.

지금 이 프롬프트 시도하기

생성기에 자동으로 채워진 상태로 열립니다.

이 프롬프트 시도

관련 프롬프트