Hi, I'm Zhixuan Chen
Ph.D. researcher · I work on Multimodal Large Language Models
I am a Ph.D. student in Computer Science and Engineering at HKUST, advised by Prof. Hao Chen. My research focuses on multimodal foundation models and vision-language intelligence, including efficient 3D visual encoding, cross-modal alignment, region-level understanding and generation, and reinforcement fine-tuning. Before HKUST, I graduated with honors from UESTC and received the China National Scholarship for three consecutive years.

What I work on
Building multimodal models that connect visual perception, language understanding and model reasoning, with an emphasis on fine-grained grounding, efficient adaptation and scalable training.
Multimodal Foundation Models
Large-scale vision-language learning with efficient visual representations, contrastive objectives and cross-modal semantic alignment.
Fine-grained Vision–Language Intelligence
Region-level referring, grounding and long-form generation that connect visual details with precise natural-language instructions.
Promptable Visual Understanding
Text-prompted segmentation and open-vocabulary perception, enabling natural language to drive pixel-level visual understanding.
LLM Post-training
Parameter-efficient fine-tuning, reinforcement fine-tuning and prompt learning for stronger multimodal reasoning and adaptation.
News & Highlights
Latest acceptances, awards, and milestones — most recent first.
Selected Publications
Selected work on multimodal foundation models, fine-grained vision–language understanding and efficient model adaptation. My name is shown in bold. See the full list on Google Scholar.
Awards & Honors
A selection of competitive scholarships, honors and competition prizes.
Education
Professional Service
Journal Reviewer
- IEEE Transactions on Medical Imaging (TMI)
- IEEE Journal of Biomedical and Health Informatics
Teaching & Mentoring
Teaching Assistant · COMP 4021 Internet Computing (Fall 2023, Spring 2024)
Mentee · Liqi Lin (UG @ USTC), 2024.08 – Present
Let's collaborate
Open to research collaboration and opportunities in multimodal large language models, vision-language intelligence, and LLM post-training.