SketchImage → Imageacademic

Clarifier Framework for Multimodal Ambiguous Queries

Figure 1: Our Clarifier framework resolves the multimodal ambiguous query, “Is this a good gift?” (a) Multimodal Intent Ambiguity: A standard AI defaults to a guess, making assumptions about the recipient (e.g., a child), their interests, and the user’s budget. This “black box” approach is unhelpful because if the assumptions are wrong, the recommendation is useless. (b) Plug-and-Play Clarifier: Our system avoids guessing. It first identifies that key information (recipient, occasion, budget), missing visual context, and pointing gestures between modalities. It then proactively asks clarifying questions and provides camera feedback (”For who? Move camera upward…”). Once the user provides the necessary context, the system can deliver a relevant and genuinely helpful recommendation.

Input image
Generated result

Paper context

Paper title: Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation Abstract: The performance of egocentric AI agents is fundamentally limited by multimodal intent ambiguity. This challenge arises from a combination of underspecified language, imperfect visual data, and deictic gestures, which frequently leads to task failure. Existing monolithic Vision-Language Models (VLMs) struggle to resolve these multimodal ambiguous inputs, often failing silently or hallucinating responses. To address these ambiguities, we introduce the Plug-and-Play Clarifier, a zero-shot and modular framework that decomposes the problem into discrete, solvable sub-tasks. Specifically, our framework consists of three synergistic modules: (1) a text clarifier that uses dialogue-driven reasoning to interactively disambiguate linguistic intent, (2) a vision clarifier that delivers real-time guidance feedback, instructing users to adjust their positioning for improved capture quality, and (3) a cross-modal clarifier with grounding mechanism that robustly interprets 3D pointing gestures and identifies the specific objects users are pointing to. Extensive experiments demonstrate that our framework improves the intent clarification performance of small language models (4--8B) by approximately 30%, making them competitive with significantly larger counterparts. We also observe consistent gains when applying our framework to these larger models. Furthermore, our vision clarifier increases corrective guidance accuracy by over 20%, and our cross-modal clarifier improves semantic answer acc Passages referencing this figure: Introduction Figure 1: Our Clarifier framework resolves the multimodal ambiguous query, “Is this a good gift?” (a) Multimodal Intent Ambiguity: A standard AI defaults to a guess, making assumptions about the recipient (e.

The prompt

A reference image is attached above. It is my rough sketch of what I
want my final figure to look like — sometimes hand-drawn, sometimes
an AI quick-draft. The quality is rough; details may be wrong; some
elements may be missing — but it shows the STRUCTURE / SPATIAL LAYOUT
I'm going for.

I've also shared the paper title + abstract + method section + figure
caption + paragraphs that reference this figure.

TASK: Refine my rough sketch into a polished publication-quality figure.

  - Preserve the SPATIAL STRUCTURE of the sketch: where the boxes are,
    how they connect, the overall reading order, the rough proportions.
  - You may correct details: better text labels (use the paper context
    to get the right component names), cleaner shapes, real icons
    instead of stick-figure placeholders.
  - Do NOT regenerate from scratch with a different layout. The
    finished figure must be visibly the same composition as the sketch.

If your output bears no spatial resemblance to the reference sketch,
you've failed the task. Refine the sketch — don't replace it. Just
give me the polished figure.

Try this prompt now

Open it inside the generator with the prompt pre-filled.

Try this prompt

Related prompts