Key ElementsImage → Imageacademic

LouvreSAE Training and Inference Stages

Figure 2 : LouvreSAE Training & Inference Stages. The method comprises three stages: (a) Backbone SAE Training , where the SAE is presented with a wide range of data sources (including both generic and art-specific datasets); (b) Style Profile Encoding , where a handful of sparse codes obtained from samples of the target style are overlapped to find overlapping concepts that make up the style; and (c) Generation Steering , where the style profile us used as a steering vector during image-to-image pipeline inference.

Input image
Generated result

Paper context

Paper title: LouvreSAE: Sparse Autoencoders for Interpretable and Controllable Style Transfer Abstract: Artistic style transfer in generative models remains a significant challenge, as existing methods often introduce style only via model fine-tuning, additional adapters, or prompt engineering, all of which can be computationally expensive and may still entangle style with subject matter. In this paper, we introduce a training- and inference-light, interpretable method for representing and transferring artistic style. Our approach leverages an art-specific Sparse Autoencoder (SAE) on top of latent embeddings of generative image models. Trained on artistic data, our SAE learns an emergent, largely disentangled set of stylistic and compositional concepts, corresponding to style-related elements pertaining brushwork, texture, and color palette, as well as semantic and structural concepts. We call it LouvreSAE and use it to construct style profiles: compact, decomposable steering vectors that enable style transfer without any model updates or optimization. Unlike prior concept-based style transfer methods, our method requires no fine-tuning, no LoRA training, and no additional inference passes, enabling direct steering of artistic styles from only a few reference images. We validate our method on ArtBench10, achieving or surpassing existing methods on style evaluations (VGG Style Loss and CLIP Score Style) while being 1.7-20x faster and, critically, interpretable. Passages referencing this figure: Figure 1 : Overview of LouvreSAE .

The prompt

A reference image is attached above. It contains ONLY the visual elements
I want to use in my final figure — icons, charts, photos, illustrations —
arranged roughly in the layout I'm imagining. There are no panel borders,
arrows, text labels, or section titles in the reference yet.

I've also shared the paper title + abstract + method section + figure
caption + paragraphs that reference this figure.

TASK: Build the finished publication-quality figure USING the specific
visual elements I've already provided.

  - The icons / charts / photos / illustrations in the reference image
    are the ONES I want in the final figure. Use them. Don't substitute
    different icons. Don't pick generic stock visuals.
  - Their rough positions in the reference are my intended layout —
    keep them roughly where they are unless a small adjustment clearly
    helps composition.
  - Add the connecting structure: panel borders, arrows, text labels,
    section titles, captions — whatever is needed to make the figure
    coherent and publication-quality.
  - Do NOT generate the figure from scratch with different elements.
    Do NOT replace my icons with new ones.

If your output uses different icons / charts / illustrations from the
reference, you've failed the task. Just give me the final figure.

Try this prompt now

Open it inside the generator with the prompt pre-filled.

Try this prompt

Related prompts