Portrait of Junfeng Wu Senior Research Scientist
ByteDance US

01 / About

Learning to understand, create, and act.

I am a Senior Research Scientist at ByteDance, US, working on multimodal intelligence—from visual representation learning and unified multimodal models to video generation, world models, and agentic systems. My long-term goal is to build general-purpose AI that understands the world, creates, reasons, and acts.

I earned my bachelor's degree and Ph.D. at Huazhong University of Science and Technology, completing my doctoral research in the VLR Group with Prof. Xiang Bai as my advisor. From 2021 to 2025, I was a research intern at ByteDance AI Lab, working closely with Yi Jiang and Song Bai.

I’m looking for research interns interested in multimodal intelligence, with opportunities in China and the US. Feel free to reach out!

02 / Research

Research overview

Explore selected work

News & highlights

2022–2026
  1. TokBench accepted to ACM MM Oral

    VideoRAE released as an arXiv preprint.

  2. UniTok at NeurIPS Spotlight

    Liquid accepted to IJCV.

    Liquid was formally published in 2026.

  3. GLEE at CVPR Highlight

    PartGLEE at ECCV.

  4. InstMove at CVPR.

  5. IDOL and SeqFormer at ECCV 2 × Oral

Current directions

01

Unified understanding & generation

Visual tokenization and multimodal language models that connect perception with synthesis.

TokBench · UniTok · Liquid

02

Generalist visual perception

Recognizing, grounding, and segmenting objects and their parts across images and videos.

GLEE · PartGLEE

03

Video representations & generation

Generation-ready video latents and object-centric representations for motion and instance tracking.

VideoRAE · InstMove · IDOL · SeqFormer

03 / Publications

Selected work

TokBench Figure 1: face and text reconstruction comparisons with human judgments and evaluation metrics ACM MM 2026 · Oral

Visual tokenizer evaluation

TokBench: Evaluating Your Visual Tokenizer before Visual Generation

Junfeng Wu*, Dongliang Luo*, Weizhi Zhao, Zhihao Xie, Yuanhao Wang, Junyi Li, Xudong Xie, Yuliang Liu, Xiang Bai
* Equal contribution.

A lightweight benchmark measuring how well image and video tokenizers preserve fine-grained text and facial details.

VideoRAE Figure 1: text-to-video samples from an 11B OpenSora2 model adapted to VideoRAE-3D arXiv 2026 · Preprint

Video representation autoencoders

VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders

Zhihao Xie*, Junfeng Wu*, Xinting Hu, Junchao Huang, Li Jiang
* Equal contribution.

Transforms frozen video foundation features into compact latents for both diffusion and autoregressive video generation.

UniTok method preview NeurIPS 2025 · Spotlight

Unified visual tokenization

UniTok: A Unified Tokenizer for Visual Generation and Understanding

Chuofan Ma, Yi Jiang, Junfeng Wu, Jihan Yang, Xin Yu, Zehuan Yuan, Bingyue Peng, Xiaojuan Qi

One tokenizer bridges continuous visual understanding with discrete high-fidelity generation.

Liquid generation gallery IJCV 2026 · Accepted 2025

Multimodal generation

Liquid: Language Models are Scalable and Unified Multi-modal Generators

Junfeng Wu, Yi Jiang, Chuofan Ma, Yuliang Liu, Hengshuang Zhao, Zehuan Yuan, Song Bai, Xiang Bai

Language models scale into unified multimodal generators across resolutions and visual formats.

GLEE general object foundation model preview CVPR 2024 · Highlight

Generalist vision

General Object Foundation Model for Images and Videos at Scale

Junfeng Wu, Yi Jiang, David Liu, Zehuan Yuan, Xiang Bai, Song Bai

A single generalist model recognizes and grounds diverse visual objects across images and videos.

PartGLEE preview ECCV 2024

Part-level perception

PartGLEE: A Foundation Model for Recognizing and Parsing Any Objects

Junyi Li, Junfeng Wu, Weizhi Zhao, Song Bai, Xiang Bai

Extending generalist object understanding from whole objects to fine-grained semantic parts.

InstMove previewCVPR 2023

Motion modeling

InstMove: Instance Motion for Object-centric Video Segmentation

David Liu, Junfeng Wu, Yi Jiang, Xiang Bai, Alan Yuille, Song Bai

IDOL previewECCV 2022 · Oral

Online video models

In Defense of Online Models for Video Instance Segmentation

Junfeng Wu, David Liu, Yi Jiang, Song Bai, Alan Yuille, Xiang Bai

SeqFormer previewECCV 2022 · Oral

Sequence modeling

SeqFormer: Sequential Transformer for Video Instance Segmentation

Junfeng Wu, Yi Jiang, Song Bai, Wenqing Zhang, Xiang Bai

04 / Community

Academic service

Contributing to the computer vision and machine learning community.

Conferences

CVPR · ICCV · ECCV

Reviewer, 2023–2025

Machine Learning

NeurIPS · ICML · AAAI

Reviewer, 2024–2025

Journals

TPAMI · PR · SCIS

Journal reviewer

05 / Contact

Let’s build the next visual foundation model.

wjf5203@gmail.com