Key ElementsImage → Imageacademic

PianoFlow Model Architecture with Control Inputs

Figure 1 . Architecture of the PianoFlow model with the control inputs on the left and generation logic on the right.

Input image
Generated result

Paper context

Paper title: SyMuPe: Affective and Controllable Symbolic Music Performance Abstract: Emotions are fundamental to the creation and perception of music performances. However, achieving human-like expression and emotion through machine learning models for performance rendering remains a challenging task. In this work, we present SyMuPe, a novel framework for developing and training affective and controllable symbolic piano performance models. Our flagship model, PianoFlow, uses conditional flow matching trained to solve diverse multi-mask performance inpainting tasks. By design, it supports both unconditional generation and infilling of music performance features. For training, we use a curated, cleaned dataset of 2,968 hours of aligned musical scores and expressive MIDI performances. For text and emotion control, we integrate a piano performance emotion classifier and tune PianoFlow with the emotion-weighted Flan-T5 text embeddings provided as conditional inputs. Objective and subjective evaluations against transformer-based baselines and existing models show that PianoFlow not only outperforms other approaches, but also achieves performance quality comparable to that of human-recorded and transcribed MIDI samples. For emotion control, we present and analyze samples generated under different text conditioning scenarios. The developed model can be integrated into interactive applications, contributing to the creation of more accessible and engaging music performance systems. Passages referencing this figure: The complete model with all conditioning is illustrated in Figure 1 . Figure 1 .

The prompt

A reference image is attached above. It contains ONLY the visual elements
I want to use in my final figure — icons, charts, photos, illustrations —
arranged roughly in the layout I'm imagining. There are no panel borders,
arrows, text labels, or section titles in the reference yet.

I've also shared the paper title + abstract + method section + figure
caption + paragraphs that reference this figure.

TASK: Build the finished publication-quality figure USING the specific
visual elements I've already provided.

  - The icons / charts / photos / illustrations in the reference image
    are the ONES I want in the final figure. Use them. Don't substitute
    different icons. Don't pick generic stock visuals.
  - Their rough positions in the reference are my intended layout —
    keep them roughly where they are unless a small adjustment clearly
    helps composition.
  - Add the connecting structure: panel borders, arrows, text labels,
    section titles, captions — whatever is needed to make the figure
    coherent and publication-quality.
  - Do NOT generate the figure from scratch with different elements.
    Do NOT replace my icons with new ones.

If your output uses different icons / charts / illustrations from the
reference, you've failed the task. Just give me the final figure.

Try this prompt now

Open it inside the generator with the prompt pre-filled.

Try this prompt

Related prompts