Text to FigureText → Imageacademic

HistDiT: Transformer Diffusion Architecture for Histology

Figure 1: The HistDiT Architecture replaces the standard U-Net with a Transformer backbone that integrates two distinct conditioning streams: (i) Global Semantic Guidance uses frozen UNI embeddings injected via Adaptive Layer Norm (adaLN) to enforce diagnostic consistency; (ii) Spatial Structural Guidance uses VAE-encoded H&E latents injected via Cross-Attention to strictly preserve tissue morphology.

Paper context

Paper title: HistDiT: A Structure-Aware Latent Conditional Diffusion Model for High-Fidelity Virtual Staining in Histopathology Abstract: Immunohistochemistry (IHC) is essential for assessing specific immune biomarkers like Human Epidermal growth-factor Receptor 2 (HER2) in breast cancer. However, the traditional protocols of obtaining IHC stains are resource-intensive, time-consuming, and prone to structural damages. Virtual staining has emerged as a scalable alternative, but it faces significant challenges in preserving fine-grained cellular structures while accurately translating biochemical expressions. Current state-of-the-art methods still rely on Generative Adversarial Networks (GANs) or standard convolutional U-Net diffusion models that often struggle with "structure and staining trade-offs". The generated samples are either structurally relevant but blurry, or texturally realistic but have artifacts that compromise their diagnostic use. In this paper, we introduce HistDiT, a novel latent conditional Diffusion Transformer (DiT) architecture that establishes a new benchmark for visual fidelity in virtual histological staining. The novelty introduced in this work is, a) the Dual-Stream Conditioning strategy that explicitly maintains a balance between spatial constraints via VAE-encoded latents and semantic phenotype guidance via UNI embeddings; b) the multi-objective loss function that contributes to sharper images with clear morphological structure; and c) the use of the Structural Correlation Metric (SCM) to focus on the core morphological structure for precise assessment of sample quality. Consequently Passages referencing this figure: uces our HistDiT (Histopathology Diffsuion Transformer) , a novel virtual staining method for generating realistic and accurate IHC stained images used for HER2 assessment in breast cancer diagnostics. HistDiT utilises a dual-conditioned Diffusion Transformer to translate H&E images into specialized IHC stains with high structural and semantic fidelity. The overall framework is illustrated in Fig. 1 . Figure 1: The HistDiT Architecture replaces the standard U-Net with a Transformer backbone that integrates two distinct conditioning streams: (i) Global Semantic Guidance uses frozen UNI embeddings injected via Adaptive Layer Norm (adaLN) to enforce diagnostic consistency; (ii) Spatial Structural Guidance uses VAE-encoded H&E latents injected via Cross-Attention to strictly preserve t HistDiT (Histopathology Diffsuion Transformer) , a novel virtual staining method for generating realistic and accurate IHC stained images used for HER2 assessment in breast cancer diagnostics. HistDiT utilises a dual-conditioned Diffusion Transformer to translate H&E images into specialized IHC stains with high structural and semantic fidelity. The overall framework is illustrated in Fig. 1 . Figure 1: The HistDiT Architecture replaces the standard U-Net with a Transformer backbone that integrates two distinct conditioning streams: (i) Global Semantic Guidance uses frozen UNI embeddings injected via Adaptive Layer Norm (adaLN) to enforce diagnostic consistency; (ii) Spatial Structural Guidance uses VAE-encoded H&E latents injected via Cross-Attention to strictly preserve tissue mor ng or distorting the physical tissue structure. The noisy latent z t z_{t} is first “patchified” into a sequence of N N tokens, where N = ( h / p ) × ( w / p ) N=(h/p)\times(w/p) and p p is the patc

The prompt

Above I've shared:
(1) the paper title + abstract + method section,
(2) the figure caption I want.

TASK: Render the main figure for this academic paper. Style requirements:

  - This is an ACADEMIC PAPER FIGURE (not a poster, not an infographic).
  - Clean black-on-white background; minimal decoration.
  - Components, arrows, and labels rendered crisply; small dense text OK.
  - Single-figure layout — no banner header, no "title" inside the image.
  - Match the level of detail of a top-tier conference paper figure
    (NeurIPS / ICLR / CVPR style).

Render the figure described in the caption. Just give me the final image.

Try this prompt now

Open it inside the generator with the prompt pre-filled.

Try this prompt

Related prompts