Enhancing Human-Computer Interaction Through Decoupling Motion and Camera Control in Human-Centric Video Generation.
Saved in:
| Title: | Enhancing Human-Computer Interaction Through Decoupling Motion and Camera Control in Human-Centric Video Generation. |
|---|---|
| Authors: | Chen, Guanfu1 (AUTHOR), Liang, Xiao1 (AUTHOR) bjtulxlx@sina.com, He, Wenbin1 (AUTHOR) |
| Source: | International Journal of Human-Computer Interaction. Dec2025, Vol. 41 Issue 23, p14746-14760. 15p. |
| Subjects: | Human-computer interaction, Camera movement, Interactive videos, Machine learning |
| Abstract: | Recently, text-to-video (T2V) models based on diffusion models have achieved significant breakthroughs. However, existing models lack sufficient control, particularly for human-centric content. Enhancing the usability and interactive experience of video generation remains a key challenge. In this work, we propose a video generation framework that decouples human motion and camera movement, enabling users to control subject actions and camera dynamics independently or in combination. Our method improves video quality and control in few-shot settings by integrating additional human-computer interaction (HCI) cues. Specifically, for highly precise human motion control, we design a motion generation module with a variant skeleton projection mechanism and then utilise ControlNet to inject continuous skeleton information into the neural network to guide the generation of coherent action sequences. For camera movement control, we introduce temporal low-rank adaptations (LoRAs) to efficiently fine-tune parameters using only 8 videos on a single GPU, thereby learning a specific camera movement pattern. Additionally, we design a temporal training loss to enhance inter-frame consistency and improve camera movement learning. The proposed framework enables users to perform diverse lens control under few-shot learning. Extensive experiments demonstrate that our method can generate competitive human-centric videos. Our approach enhances the efficiency and effectiveness of human-computer interaction in video creation. [ABSTRACT FROM AUTHOR] |
| Copyright of International Journal of Human-Computer Interaction is the property of Taylor & Francis Ltd and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
|
Full text is not displayed to guests.
Login for full access.
|
|
| Abstract: | Recently, text-to-video (T2V) models based on diffusion models have achieved significant breakthroughs. However, existing models lack sufficient control, particularly for human-centric content. Enhancing the usability and interactive experience of video generation remains a key challenge. In this work, we propose a video generation framework that decouples human motion and camera movement, enabling users to control subject actions and camera dynamics independently or in combination. Our method improves video quality and control in few-shot settings by integrating additional human-computer interaction (HCI) cues. Specifically, for highly precise human motion control, we design a motion generation module with a variant skeleton projection mechanism and then utilise ControlNet to inject continuous skeleton information into the neural network to guide the generation of coherent action sequences. For camera movement control, we introduce temporal low-rank adaptations (LoRAs) to efficiently fine-tune parameters using only 8 videos on a single GPU, thereby learning a specific camera movement pattern. Additionally, we design a temporal training loss to enhance inter-frame consistency and improve camera movement learning. The proposed framework enables users to perform diverse lens control under few-shot learning. Extensive experiments demonstrate that our method can generate competitive human-centric videos. Our approach enhances the efficiency and effectiveness of human-computer interaction in video creation. [ABSTRACT FROM AUTHOR] |
|---|---|
| ISSN: | 10447318 |
| DOI: | 10.1080/10447318.2025.2487877 |