ChallenCap: Monocular 3D Capture of Challenging Human Performances using Multi-Modal References

He, Yannan; Pang, Anqi; Chen, Xin; Liang, Han; Wu, Minye; Ma, Yuexin; Xu, Lu

doi:10.48550/arxiv.2103.06747

Cited by 1 publication

(1 citation statement)

References 41 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…Different than the reconstruction of static scenes, tackling dynamic scenes requires settling the illumination changes and moving objects. To obtain a reconstruction for dynamic objects with input data from either camera array or a single camera, methods involving silhouette [Kim et al 2012;Taneja et al 2011], stereo Luo et al 2020;Lv et al 2018;Xu et al 2018], segmentation [Ranftl et al 2016;Russell et al 2014], and photometric [Ahmed et al 2008;He et al 2021;Vlasic et al 2009] have been explored. Early solutions [Collet et al 2015;Dou et al 2017;Mustafa et al 2016] rely on multi-view dome-based setup for high-fidelity reconstruction and texture rendering of human activities in novel views.…”

Section: Trajectory Predictionmentioning

confidence: 99%

Editable free-viewpoint video using a layered neural representation

Zhang

Liu

et al. 2021

ACM Trans. Graph.

Self Cite

View full text Add to dashboard Cite

Generating free-viewpoint videos is critical for immersive VR/AR experience, but recent neural advances still lack the editing ability to manipulate the visual perception for large dynamic scenes. To fill this gap, in this paper, we propose the first approach for editable free-viewpoint video generation for large-scale view-dependent dynamic scenes using only 16 cameras. The core of our approach is a new layered neural representation, where each dynamic entity, including the environment itself, is formulated into a spatiotemporal coherent neural layered radiance representation called ST-NeRF. Such a layered representation supports manipulations of the dynamic scene while still supporting a wide free viewing experience. In our ST-NeRF, we represent the dynamic entity/layer as a continuous function, which achieves the disentanglement of location, deformation as well as the appearance of the dynamic entity in a continuous and self-supervised manner. We propose a scene parsing 4D label map tracking to disentangle the spatial information explicitly and a continuous deform module to disentangle the temporal motion implicitly. An object-aware volume rendering scheme is further

show abstract