UP2You reconstructs high-quality textured meshes from unconstrained photos. Our approach effectively handles extremely unconstrained photo collections by rectifying them into orthogonal multi-view images and corresponding normal maps, enabling the reconstruction of detailed 3D clothed portraits.
Paradigm Differences Between Previous Works and UP2You
Top: Previous works like PuzzleAvatar and AvatarBooth compress unconstrained photos into implicit personal tokens and DreamBooth weights through fine-tuning, then generate 3D humans via SDS optimization.
Bottom: UP2You directly rectifies unconstrained photo collections into orthogonal view images and normals, then reconstructs textured human meshes, achieving superior quality while reducing processing time from 4 hours to 1.5 minutes.
Our Results
Pose-Dependent Correlation Maps
Related Work
- PuzzleAvatar: Assembly of Avatar from Unconstrained Photo Collections
- AvatarBooth: High-Quality and Customizable 3D Human Avatar Generation
- MV-Adapter: Multi-view Consistent Image Generation Made Easy
- PSHuman: Photorealistic Single-image 3D Human Reconstruction using Cross-Scale Multiview Diffusion
- 4D-DRESS: A 4D Dataset of Real-world Human Clothing with Semantic Annotations
- Function4D: Real-time Human Volumetric Capture from Very Sparse RGBD Sensors
- Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer
- Learning Locally Editable Virtual Humans
- High-fidelity 3D Human Digitization from Single 2K Resolution Images