Key ElementsImage → Imageacademic

FrameDiffuser Architecture with Dual Conditioning

Figure 2 : FrameDiffuser Architecture with dual conditioning: ControlNet processes 10-channel input comprising 9 G-buffer channels for structural guidance and 1 pred. irradiance channel for lighting guidance, computed from the previous frame’s model output and basecolor. ControlLoRA conditions on the previous frame encoded in VAE latent space for temporal coherence. The generated output at time t t is used to compute the irradiance input for the next frame at time t + 1 t+1 , enabling autoregressive frame generation. The encoder ℰ \mathcal{E} and decoder 𝒟 \mathcal{D} represent the VAE components operating in latent space. The training strategy on the right shows our three-stage approach: first, we train ControlNet on the G-buffer to image translation task without irradiance. Second, we add ControlLoRA and irradiance for temporal conditioning. Third, we train autoregressively using the model’s own generated frames as previous-frame inputs to make the model robust against its own generation errors.

Input image
Generated result

Paper context

Paper title: FrameDiffuser: G-Buffer-Conditioned Diffusion for Neural Forward Frame Rendering Abstract: Neural rendering for interactive applications requires translating geometric and material properties (G-buffer) to photorealistic images with realistic lighting on a frame-by-frame basis. While recent diffusion-based approaches show promise for G-buffer-conditioned image synthesis, they face critical limitations: single-image models like RGBX generate frames independently without temporal consistency, while video models like DiffusionRenderer are too computationally expensive for most consumer gaming sets ups and require complete sequences upfront, making them unsuitable for interactive applications where future frames depend on user input. We introduce FrameDiffuser, an autoregressive neural rendering framework that generates temporally consistent, photorealistic frames by conditioning on G-buffer data and the models own previous output. After an initial frame, FrameDiffuser operates purely on incoming G-buffer data, comprising geometry, materials, and surface properties, while using its previously generated frame for temporal guidance, maintaining stable, temporal consistent generation over hundreds to thousands of frames. Our dual-conditioning architecture combines ControlNet for structural guidance with ControlLoRA for temporal coherence. A three-stage training strategy enables stable autoregressive generation. We specialize our model to individual environments, prioritizing consistency and inference speed over broad generalization, demonstrating that environment-specific Passages referencing this figure: Figure 1 : G-buffer to Photorealistic Rendering.

The prompt

A reference image is attached above. It contains ONLY the visual elements
I want to use in my final figure — icons, charts, photos, illustrations —
arranged roughly in the layout I'm imagining. There are no panel borders,
arrows, text labels, or section titles in the reference yet.

I've also shared the paper title + abstract + method section + figure
caption + paragraphs that reference this figure.

TASK: Build the finished publication-quality figure USING the specific
visual elements I've already provided.

  - The icons / charts / photos / illustrations in the reference image
    are the ONES I want in the final figure. Use them. Don't substitute
    different icons. Don't pick generic stock visuals.
  - Their rough positions in the reference are my intended layout —
    keep them roughly where they are unless a small adjustment clearly
    helps composition.
  - Add the connecting structure: panel borders, arrows, text labels,
    section titles, captions — whatever is needed to make the figure
    coherent and publication-quality.
  - Do NOT generate the figure from scratch with different elements.
    Do NOT replace my icons with new ones.

If your output uses different icons / charts / illustrations from the
reference, you've failed the task. Just give me the final figure.

Try this prompt now

Open it inside the generator with the prompt pre-filled.

Try this prompt

Related prompts