Photo for Weijia Fan

Weijia Fan

Grad Teaching Assistantship
Faculty of Science - Computing Science
Grad Research Asst Fellowship
Faculty of Science - Computing Science

Personal Website: https://wakinghours-github.github.io/

Contact

Grad Teaching Assistantship
Faculty of Science - Computing Science

Email
wfan8@ualberta.ca

Grad Research Asst Fellowship
Faculty of Science - Computing Science

Email
wfan8@ualberta.ca

Overview

Area of Study / Keywords

Computer Vision Multimodal Learning Autonomous Driving


About

Hi, welcome! 👋 I’m Weijia Fan, a PhD student in the Department of Computing Science at the University of Alberta, supervised by Professor Anup Basu.

My research interests include computer vision, multimodal learning, and vision-language models, with their applications in visual recognition tasks, scene understanding, document analysis, and autonomous driving. My work has appeared at BMVC, AAAI, and CVPR, as well as in academic journals.

I’m always open to research discussions and collaboration. If you’re interested in working together, please contact me at weijia.fan[AT]ualberta.ca.


Research

In the past, my research interests have focused on foundational tasks in computer vision. I investigate how and why standard learning methods like cross-entropy-based training paradigms fall short in complex visual tasks, such as Face Recognition (FR) and Long-tailed Recognition (LTR). I am particularly interested in exploring alternative learning strategies and loss variants, like binary cross-entropy and its variants, and their potential to help models build a compact representation or comprehensive understanding of the visual world. The works have been accepted by AAAI2026 (Paper, Code) and Pattern Analysis and Applications.

Then, I was exploring the frontiers of Vision-Language Models (VLMs) in Karlsruhe, Germany. My work aims to fill a critical gap in scene understanding and autonomous driving, as existing VLMs primarily focus on standard pinhole camera image settings, which miss the seamless spatial information across the images. We developed methods for holistic panoramic scene understanding in autonomous driving. We curated a large-scale benchmark, PanoVQA, to evaluate the existing VLMs and train our developed Panorama-Language Models (PLMs). This work has been accepted by CVPR2026 (Paper, Code).

At UofA, I am exploring topics related to Computer Vision, Multimodal Learning, Scene Understanding, and their applications in Autonomous Driving and Open World Perception.