SketchImage → Imageacademic

DINOv2 Resolution-Prediction Network Overview

Figure 3 : Overview of our resolution-prediction network. A frozen DINOv2 backbone encodes spatial features of a cropped video clip, fused with encoded motion information via gated addition, and passed through an MLP to predict a discrete resolution.

Input image
Generated result

Paper context

Paper title: Seeing enough: non-reference perceptual resolution selection for power-efficient client-side rendering Abstract: Many client-side applications, especially games, render video at high resolution and frame rate on power-constrained devices, even when users perceive little or no benefit from all those extra pixels. Existing perceptual video quality metrics can indicate when a lower resolution is "good enough", but they are full-reference and computationally expensive, making them impractical for real-world applications and deployment on-device. In this work, we leverage the spatio-temporal limits of the human visual system and propose a non-reference method that predicts, from the rendered video alone, the lowest resolution that remains perceptually indistinguishable from the best available option, enabling power-efficient client-side rendering. Our approach is codec-agnostic and requires only minimal modifications to existing infrastructure. The network is trained on a large dataset of rendered content labeled with a full-reference perceptual video quality metric. The prediction significantly enhances perceptual quality while substantially reducing computational costs, suggesting a practical path toward perception-guided, power-efficient client-side rendering. Passages referencing this figure: nostic and requires only minimal modifications to existing infrastructure. The network is trained on a large dataset of rendered content labeled with a full-reference perceptual video quality metric. The prediction significantly enhances perceptual quality while substantially reducing computational costs, suggesting a practical path toward perception-guided, power-efficient client-side rendering. Figure 1 : Client devices render interactive content at a high frame rate (120 Hz). For each 250 ms clip, we extract motion vectors and a short rendered image sequence; our resolution predictor encodes the relevant spatial and temporal structures that give rise to visible distortions. These motion and appearance features are fused by a non-reference resolution predictor, which estimates the next p

The prompt

A reference image is attached above. It is my rough sketch of what I
want my final figure to look like — sometimes hand-drawn, sometimes
an AI quick-draft. The quality is rough; details may be wrong; some
elements may be missing — but it shows the STRUCTURE / SPATIAL LAYOUT
I'm going for.

I've also shared the paper title + abstract + method section + figure
caption + paragraphs that reference this figure.

TASK: Refine my rough sketch into a polished publication-quality figure.

  - Preserve the SPATIAL STRUCTURE of the sketch: where the boxes are,
    how they connect, the overall reading order, the rough proportions.
  - You may correct details: better text labels (use the paper context
    to get the right component names), cleaner shapes, real icons
    instead of stick-figure placeholders.
  - Do NOT regenerate from scratch with a different layout. The
    finished figure must be visibly the same composition as the sketch.

If your output bears no spatial resemblance to the reference sketch,
you've failed the task. Refine the sketch — don't replace it. Just
give me the polished figure.

Try this prompt now

Open it inside the generator with the prompt pre-filled.

Try this prompt

Related prompts