Skip to content

Omni Demand Understanding

A Benchmark for Situated User-Intent Inference in Multimodal Interaction

Explore examples
AbstractPaper examplesHuman recordings More demos Ethics

Abstract

2,078 scenes277 human recordingsEnglish & MandarinAudio & audio-visual

Examples from the Paper

from the teaser, spanning four taxonomy axes and No demand.

Human-recorded Demos

with reviewed benchmark annotations.

Actor consent. All participating actors have provided informed consent for public display and benchmark evaluation of their recordings, including their faces and voices. See the Ethics Statement in our paper for details.

Filter recordings

No recordings match these filters.

More Demos

, including audio-only interactions and Mandarin examples.

Filter examples

No examples match these filters.

Ethics Statement

← All examples

The interaction

Read the reference annotation ↓

Conversation sequence

Only the final user turn is evaluated. Earlier assistant replies are text supplied as conversation context.

The reference annotation

Scene taxonomy 6 construction axes

Scene taxonomy describes construction coverage. Evidence labels describe the source of information for each annotation.

ODU-Bench · Omni Demand Understanding