๐๐๏ธ๐ New Research Alert - ICCV 2025 (Poster)! ๐๐๏ธ๐ ๐ Title: Is Less More? Exploring Token Condensation as Training-Free Test-Time Adaptation ๐
๐ Description: Token Condensation as Adaptation (TCA) improves the performance and efficiency of Vision Language Models in zero-shot inference by introducing domain anchor tokens.
๐ฅ Authors: Zixin Wang, Dong Gong, Sen Wang, Zi Huang, Yadan Luo
๐๐๏ธ๐ New Research Alert - ICCV 2025 (Oral)! ๐๐๏ธ๐ ๐ Title: Diving into the Fusion of Monocular Priors for Generalized Stereo Matching ๐
๐ Description: The proposed method enhances stereo matching by efficiently combining unbiased monocular priors from vision foundation models. This method addresses misalignment and local optima issues using a binary local ordering map and pixel-wise linear regression.
๐๐๐ New Research Alert - ICCV 2025 (Oral)! ๐๐ค๐ ๐ Title: Understanding Co-speech Gestures in-the-wild ๐
๐ Description: JEGAL is a tri-modal model that learns from gestures, speech and text simultaneously, enabling devices to interpret co-speech gestures in the wild.
๐ฅ Authors: @sindhuhegde, K R Prajwal, Taein Kwon, and Andrew Zisserman
๐๐ก๐ New Research Alert - ICCV 2025 (Oral)! ๐๐ช๐ ๐ Title: LoftUp: Learning a Coordinate-based Feature Upsampler for Vision Foundation Models ๐
๐ Description: LoftUp is a coordinate-based transformer that upscales the low-resolution features of VFMs (e.g. DINOv2 and CLIP) using cross-attention and self-distilled pseudo-ground truth (pseudo-GT) from SAM.
๐ฅ Authors: Haiwen Huang, Anpei Chen, Volodymyr Havrylov, Andreas Geiger, and Dan Zhang
๐๐ท๏ธ๐ New Research Alert - ICCV 2025 (Oral)! ๐๐งฉ๐ ๐ Title: Heavy Labels Out! Dataset Distillation with Label Space Lightening ๐
๐ Description: The HeLlO framework is a new corpus distillation method that removes the need for large soft labels. It uses a lightweight, online image-to-label projector based on CLIP. This projector has been adapted using LoRA-style, parameter-efficient tuning. It has also been initialized with text embeddings.
๐๐ค๐ New Research Alert - ICCV 2025 (Oral)! ๐๐ค๐ ๐ Title: Variance-based Pruning for Accelerating and Compressing Trained Networks ๐
๐ Description: The one-shot pruning method efficiently compresses networks, reducing computation and memory usage while retaining almost full performance and requiring minimal fine-tuning.
๐ฅ Authors: Uranik Berisha, Jens Mehnert, and Alexandru Paul Condurache
๐๐๏ธ๐ New Research Alert - ICCV 2025 (Oral)! ๐๐๏ธ๐ ๐ Title: Token Activation Map to Visually Explain Multimodal LLMs ๐
๐ Description: The Token Activation Map (TAM) is an advanced explainability method for multimodal LLMs. Using causal inference and a Rank Gaussian Filter, TAM reveals token-level interactions and eliminates redundant activations. The result is clearer, high-quality visualizations that enhance understanding of object localization, reasoning and multimodal alignment across models.
๐ฅ Authors: Yi Li, Hualiang Wang, Xinpeng Ding, Haonan Wang, and Xiaomeng Li
๐๐ญ๐ New Research Alert - WACV 2025 (Avatars Collection)! ๐๐ญ๐ ๐ Title: EmoVOCA: Speech-Driven Emotional 3D Talking Heads ๐
๐ Description: EmoVOCA is a data-driven method for generating emotional 3D talking heads by combining speech-driven lip movements with expressive facial dynamics. This method has been developed to overcome the limitations of corpora and to achieve state-of-the-art animation quality.
๐ฅ Authors: @FedeNoce, Claudio Ferrari, and Stefano Berretti
๐ Conference: WACV, 28 Feb โ 4 Mar, 2025 | Arizona, USA ๐บ๐ธ
๐ฅ๐ญ๐ New Research Alert - HeadGAP (Avatars Collection)! ๐๐ญ๐ฅ ๐ Title: HeadGAP: Few-shot 3D Head Avatar via Generalizable Gaussian Priors ๐
๐ Description: HeadGAP introduces a novel method for generating high-fidelity, animatable 3D head avatars from few-shot data, using Gaussian priors and dynamic part-based modelling for personalized and generalizable results.
๐๐บ๐ New Research Alert - ECCV 2024 (Avatars Collection)! ๐๐๐ ๐ Title: Expressive Whole-Body 3D Gaussian Avatar ๐
๐ Description: ExAvatar is a model that generates animatable 3D human avatars with facial expressions and hand movements from short monocular videos using a hybrid mesh and 3D Gaussian representation.
๐ฅ Authors: Gyeongsik Moon, Takaaki Shiratori, and @psyth
๐ฅ๐ญ๐ New Research Alert - ECCV 2024 (Avatars Collection)! ๐๐ญ๐ฅ ๐ Title: MeshAvatar: Learning High-quality Triangular Human Avatars from Multi-view Videos ๐
๐ Description: MeshAvatar is a novel pipeline that generates high-quality triangular human avatars from multi-view videos, enabling realistic editing and rendering through a mesh-based approach with physics-based decomposition.
๐ฅ Authors: Yushuo Chen, Zerong Zheng, Zhe Li, Chao Xu, and Yebin Liu
๐๐บ๐ New Research Alert - CVPR 2024 (Avatars Collection)! ๐๐๐ ๐ Title: IntrinsicAvatar: Physically Based Inverse Rendering of Dynamic Humans from Monocular Videos via Explicit Ray Tracing ๐
๐ Description: IntrinsicAvatar is a method for extracting high-quality geometry, albedo, material, and lighting properties of clothed human avatars from monocular videos using explicit ray tracing and volumetric scattering, enabling realistic animations under varying lighting conditions.
๐ฅ Authors: Shaofei Wang, Boลพidar Antiฤ, Andreas Geiger, and Siyu Tang
๐ Conference: CVPR, Jun 17-21, 2024 | Seattle WA, USA ๐บ๐ธ
๐ฅ๐ญ๐ New Research Alert - ECCV 2024 (Avatars Collection)! ๐๐ญ๐ฅ ๐ Title: RodinHD: High-Fidelity 3D Avatar Generation with Diffusion Models ๐
๐ Description: RodinHD generates high-fidelity 3D avatars from portrait images using a novel data scheduling strategy and weight consolidation regularization to capture intricate details such as hairstyles.
๐ฅ๐ญ๐ New Research Alert - LivePortrait (Avatars Collection)! ๐๐ญ๐ฅ ๐ Title: LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control ๐
๐ Description: LivePortrait is an efficient video-driven portrait animation framework that uses implicit keypoints and stitching/retargeting modules to generate high-quality, controllable animations from a single source image.
๐ฅ Authors: @cleardusk, Dingyun Zhang, Xiaoqiang Liu, Zhizhou Zhong, Yuan Zhang, Pengfei Wan, and Di Zhang
๐๐บ๐ New Research Alert (Avatars Collection)! ๐๐๐ ๐ Title: Expressive Gaussian Human Avatars from Monocular RGB Video ๐
๐ Description: The new EVA model enhances the expressiveness of digital avatars by using 3D Gaussians and SMPL-X to capture fine-grained hand and face details from monocular RGB video.
๐ฅ Authors: Hezhen Hu, Zhiwen Fan, Tianhao Wu, Yihan Xi, Seoyoung Lee, Georgios Pavlakos, and Zhangyang Wang
๐ฅ๐ญ๐ New Research Alert - ECCV 2024 (Avatars Collection)! ๐๐ญ๐ฅ ๐ Title: Topo4D: Topology-Preserving Gaussian Splatting for High-Fidelity 4D Head Capture ๐
๐ Description: Topo4D is a novel method for automated, high-fidelity 4D head tracking that optimizes dynamic topological meshes and 8K texture maps from multi-view time-series images.
๐ฅ Authors: @Dazz1e, Y. Cheng, @Ryan-sjtu, H. Jia, D. Xu, W. Zhu, Y. Yan
๐๐ญ๐ New Research Alert - Portrait4D-v2 (Avatars Collection)! ๐๐ญ๐ ๐ Title: Portrait4D-v2: Pseudo Multi-View Data Creates Better 4D Head Synthesizer ๐
๐ Description: Portrait4D-v2 is a novel method for one-shot 4D head avatar synthesis using pseudo multi-view videos and a vision transformer backbone, achieving superior performance without relying on 3DMM reconstruction.
๐ฅ Authors: Yu Deng, Duomin Wang, and Baoyuan Wang