Hello! I’m Baiqi Li, a PhD student in Computer Science at the University of North Carolina at Chapel Hill (Fall 2025 – present), advised by Gedas Bertasius. My research focuses on video reasoning, video generation, and embodied AI. Before joining UNC, I was fortunate to work with Deva Ramanan at Carnegie Mellon University.
News
- 2026.06 We introduced WatchAct, a benchmark for behavior-grounded robot manipulation: WatchAct: A Benchmark for Behavior-Grounded Robot Manipulation.
- 2026.02 We introduced TimeBlind, a diagnostic benchmark for compositional spatio-temporal understanding of video LLMs: TimeBlind: A Spatio-Temporal Compositionality Benchmark for Video LLMs.
- 2024.09 Our paper NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples was accepted to NeurIPS 2024.
- 2024.06 Our workshop paper GenAI-Bench: A Holistic Benchmark for Compositional Text-to-Visual Generation was selected as the Best Paper at the SynData4CV workshop @ CVPR 2024.
- 2024.06 We introduced GenAI-Bench for evaluating leading image and video generation models on various aspects of compositional text-to-visual generation: GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation.
- 2024.06 We proposed a semi-automated approach to collect a vision-centric benchmark, NaturalBench, for reliably evaluating VLMs: NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples.
- 2024.04 We introduced VQAScore for evaluating the prompt alignment of text-to-image/video/3D models: Evaluating Text-to-Visual Generation with Image-to-Text Generation.
Publications
WatchAct: A Benchmark for Behavior-Grounded Robot Manipulation arXiv
Baiqi Li, Ce Zhang, Yu Fang, Yue Yang, Shangzhe Li, Mingyu Ding, Gedas Bertasius
Website | arXiv | Code | HuggingFace

TimeBlind: A Spatio-Temporal Compositionality Benchmark for Video LLMs arXiv
Baiqi Li, Kangyi Zhao, Ce Zhang, Chancharik Mitra, Jean de Dieu Nyandwi, Gedas Bertasius

NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples NeurIPS
Baiqi Li*, Zhiqiu Lin*, Wenxuan Peng*, Jean de Dieu Nyandwi*, Daniel Jiang, Zixian Ma, Simran Khanuja, Ranjay Krishna †, Graham Neubig †, Deva Ramanan †
Website | arXiv | HuggingFace |

Evaluating Text-to-Visual Generation with Image-to-Text Generation ECCV
Zhiqiu Lin, Deepak Pathak, Baiqi Li, Emily Li, Xide Xia, Graham Neubig, Pengchuan Zhang †, Deva Ramanan †

GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation Best Paper at SynData4CV
Baiqi Li*, Zhiqiu Lin*, Deepak Pathak, Emily Li, Feiyi Xin, Kewen Wu, Tiffany Ling, Xide Xia †, Pengchuan Zhang †, Graham Neubig †, Deva Ramanan †
Website | arXiv | HuggingFace
Research Experience
- 2025.08 – present, PhD Student, University of North Carolina at Chapel Hill.
- 2024.01 – 2025.06, Research Assistant, Carnegie Mellon University.
Others
- Teaching Assistant, COMP 669: Vision Transformers, Fall 2026.
- Teaching Assistant, COMP 577: Introduction to Computer Vision, Fall 2026.
- Reviewer: NeurIPS, ICLR, ICML, CVPR, ECCV, etc.