Ailing Zeng 曾爱玲
BiographyI am the Head of Foundation Video Model at Bilibili, building the frontier audio-video generation foundation models. Video is becoming the universal language of expression and information — for humans to understand and enjoy, and for machines to see and learn about the world. Foundation models will turn video from a medium we record into a medium we create: new forms of AI-native entertainment for people, and world simulators for agents and robots. We are only at the beginning — the upper bound of multimodal content intelligence and quality remains wide open. Previously, I led teams on interactive video generation and human-centric research at Anuttacon, Tencent, and IDEA, after my Ph.D. at The Chinese University of Hong Kong (advised by Prof. Qiang Xu) and a visiting stay at the Robotics Institute, CMU. Several works from my teams — LPM, DWPose, Grounded-SAM, Motion-X — are widely adopted by the community; see Google Scholar for the full list.
We are hiring! My team at Bilibili is looking for both technical and creative content talents. We are a flat, high-agency organization, pursuing the upper bound of multimodal content intelligence and quality. Feel free to reach out via email.
News
Selected WorkA few representative works; see the full list at Google Scholar. (*equal contribution, #corresponding author or project lead)
|