A family of omnimodal world models designed to jointly process and generate language, vision, audio, and action sequences.
Tingle Li
I am a final-year CS Ph.D. student at UC Berkeley, advised by Gopala Anumanchipalli. I am also a Research Intern at the NVIDIA Cosmos Lab. Previously, I worked with Hang Zhao at Tsinghua University.
My research explores audio-visual generative models for physical intelligence, using sound as a window into how objects and scenes are composed, behave, and interact. I am grateful to be supported by the Sony Research Award.
Publications
Physical AI & World Models
A benchmark for evaluating physically-grounded video-to-audio generation.
Learn sim-to-real robot policies with generative audio.
Audio-Visual Generative Models
Interactively generate object-specific sounds from visual scenes.
Manipulate audio texture using exemplar-based analogy.
Restyle a sound to fit with another scene, using an audio-visual conditional example taken from that scene.
Improving generalization in multimodal learning through uni-modal feature analysis.
We learn from unlabeled data to manipulate the style of an image using sound.
Automatic video dubbing driven by a neural network.
Speech & Audio Models
A benchmark for evaluating turn-taking capabilities in full-duplex spoken dialogue models.
We synthesize speech from MRI videos.
High-quality speech recovery system for millimeter-wave radar without deafness.
One-way GAN training for non-parallel voice conversion.
Sliced attention for music source separation by focusing on local intra-chunk features.
Winning systems for NASA Fearless Steps Challenge on speech activity detection and speaker identification.
Attention-based target speaker separation with improved efficiency and generalization.
A Hungarian algorithm-based loss that speeds up end-to-end speaker diarization.