论文标题
舞蹈:音乐启发的舞蹈视频综合
DanceIt: Music-inspired Dancing Video Synthesis
论文作者
论文摘要
闭上眼睛听音乐,可以轻松地想象演员与音乐一起节奏地跳舞。这些舞蹈运动通常由您以前见过的舞蹈运动组成。在本文中,我们建议在计算机视觉系统中复制人类的固有能力。提出的系统由三个模块组成。为了探索音乐和舞蹈运动之间的关系,我们提出了一个跨模式对齐模块,该模块着重于舞蹈视频剪辑,并伴随着预先设计的音乐,以学习一个可以判断姿势序列的视觉特征与音乐声学特征之间一致性的系统。然后将学习的模型用于想象模块中,以选择给定音乐的姿势序列。但是,从音乐中选择的姿势序列通常是不连续的。为了解决这个问题,在时空对齐模块中,我们根据舞蹈运动的趋势和周期性来开发一种空间比对算法,以预测不连续片段之间的舞蹈运动。此外,选定的姿势序列通常与音乐节拍错过。为了解决这个问题,我们进一步开发了一种时间对齐算法,以使音乐和舞蹈的节奏保持一致。最后,加工的姿势序列用于合成想象模块中现实的舞蹈视频。生成的舞蹈视频与音乐的内容和节奏相匹配。实验结果和主观评估表明,所提出的方法可以通过输入音乐来执行产生有希望的舞蹈视频的功能。
Close your eyes and listen to music, one can easily imagine an actor dancing rhythmically along with the music. These dance movements are usually made up of dance movements you have seen before. In this paper, we propose to reproduce such an inherent capability of the human-being within a computer vision system. The proposed system consists of three modules. To explore the relationship between music and dance movements, we propose a cross-modal alignment module that focuses on dancing video clips, accompanied on pre-designed music, to learn a system that can judge the consistency between the visual features of pose sequences and the acoustic features of music. The learned model is then used in the imagination module to select a pose sequence for the given music. Such pose sequence selected from the music, however, is usually discontinuous. To solve this problem, in the spatial-temporal alignment module we develop a spatial alignment algorithm based on the tendency and periodicity of dance movements to predict dance movements between discontinuous fragments. In addition, the selected pose sequence is often misaligned with the music beat. To solve this problem, we further develop a temporal alignment algorithm to align the rhythm of music and dance. Finally, the processed pose sequence is used to synthesize realistic dancing videos in the imagination module. The generated dancing videos match the content and rhythm of the music. Experimental results and subjective evaluations show that the proposed approach can perform the function of generating promising dancing videos by inputting music.