Figure 2 : Framework of DiffHDR. Given an input LDR video, we first detect its clipped regions and map it into the proposed Log-Gamma color space. A finetuned video diffusion model reconstructs missing radiance in over- and underexposed regions. A mask detector and context-focused prompting module support controllable detail synthesis. The final output HDR video supports faithful re-exposure, accurate reproduction on HDR displays, and flexible post-production workflows.
Paper title: DiffHDR: Re-Exposing LDR Videos with Video Diffusion Models Abstract: Most digital videos are stored in 8-bit low dynamic range (LDR) formats, where much of the original high dynamic range (HDR) scene radiance is lost due to saturation and quantization. This loss of highlight and shadow detail precludes mapping accurate luminance to HDR displays and limits meaningful re-exposure in post-production workflows. Although techniques have been proposed to convert LDR images to HDR through dynamic range expansion, they struggle to restore realistic detail in the over- and underexposed regions. To address this, we present DiffHDR, a framework that formulates LDR-to-HDR conversion as a generative radiance inpainting task within the latent space of a video diffusion model. By operating in Log-Gamma color space, DiffHDR leverages spatio-temporal generative priors from a pretrained video diffusion model to synthesize plausible HDR radiance in over- and underexposed regions while recovering the continuous scene radiance of the quantized pixels. Our framework further enables controllable LDR-to-HDR video conversion guided by text prompts or reference images. To address the scarcity of paired HDR video data, we develop a pipeline that synthesizes high-quality HDR video training data from static HDRI maps. Extensive experiments demonstrate that DiffHDR significantly outperforms state-of-the-art approaches in radiance fidelity and temporal stability, producing realistic HDR videos with considerable latitude for re-exposure. Passages referencing this figure: o data, we develop a pipeline that synthesizes high-quality HDR video training data from static HDRI maps. Extensive experiments demonstrate that DiffHDR significantly outperforms state-of-the-art approaches in radiance fidelity and temporal stability, producing realistic HDR videos with considerable latitude for re-exposure. Visit our project page at https://yzmblog.github.io/projects/DiffHDR/ . Figure 1 : DiffHDR reconstructs lost radiance to convert LDR videos into faithful HDR while maintaining temporal coherence (Top). DiffHDR further enables controllable HDR synthesis guided by text prompts or reference images, facilitating realistic hallucination of saturated regions (Bottom). 1 Introduction High dynamic range (HDR) video captures a wide range of scene luminance, preserving intricat ideo model. To address information loss in clipped regions, we employ luminance-based masks to guide both the generative process and a context-focused cross-attention module. By incorporating context-focused prompting or reference images, this module facilitates controllable reconstruction in over- and underexposed areas, utilizing spatio-temporal cues to hallucinate physically plausible details (Fig. 1 ). Our main contributions are as follows: 1. We introduce DiffHDR, the first video diffusion framework for LDR-to-HDR reconstruction, along with a curation pipeline which synthesizes high-quality HDR video training data from static HDRIs. 2. We introduce a Log-Gamma color mapping, enabling HDR generation within pretrained latent spaces while preserving the backbone’s generative priors and t