Xiao (Brandon) Han
Senior AI Research Scientist
About Me
I am a Senior AI Research Scientist at Meta in London. My research lies at the intersection of multimodal representation learning and media generation. I work on transferable vision-language representations and controllable image and video generation, connecting fundamental research with systems deployed at production scale.
At Meta, I work across research and deployment for controllable image and video generation.
I received my Ph.D. in Vision, Speech and Signal Processing from the University of Surrey in 2024, supervised by Prof. Tao Xiang and Prof. Yi-Zhe Song. I completed my Bachelor's degree in Information Engineering at Zhejiang University in 2020.
You can download my full CV here.
Employment
Education
Research
My research focuses on multimodal representation learning and controllable media generation. I am interested in transferable vision-language representations for cross-modal and compositional retrieval, multimodal adaptation, and fine-grained visual understanding. Building on this foundation, my current work explores reference-conditioned image and video generation and creative 2D/3D content synthesis.
I also host and advise visiting Ph.D. research scientist interns at Meta AI, teach and mentor postgraduate researchers, and review for leading computer vision and machine learning venues.
Selected Publications
VecGlypher: Unified Vector Glyph Generation with Language Models
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026 · Corresponding author
Saber: Scaling Zero-Shot Reference-to-Video Generation
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026
TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
Selected as a CVPR Highlight (top 3.5% of accepted papers).
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026
Learning Flow Fields in Attention for Controllable Person Image Generation
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2025
MarDini: Masked Auto-regressive Diffusion for Video Generation at Scale
Transactions on Machine Learning Research (TMLR) 2025
HeadSculpt: Crafting 3D Head Avatars with Text
Neural Information Processing Systems (NeurIPS) 2023
PoCoLD: Controllable Person Image Synthesis with Pose-Constrained Latent Diffusion
IEEE/CVF International Conference on Computer Vision (ICCV) 2023
FAME-ViL: Multi-Tasking Vision-Language Model for Heterogeneous Fashion Tasks
Selected as a CVPR Highlight (top 2.5% of accepted papers).
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2023
FashionViL: Fashion-Focused Vision-and-Language Representation Learning
European Conference on Computer Vision (ECCV) 2022
For a complete publication list, see my Google Scholar profile.
Contact
Email: brandon.hxiao@gmail.com
Google Scholar | GitHub | LinkedIn | X
This website uses Tufte CSS. Last updated: