Hi! I am Hao Lin (Chinese: ๆๆตฉ), a third-year undergraduate student majoring in Software Engineering at the School of Software Engineering, Huazhong University of Science and Technology(HUST) (Rank: 2/116 , GPA: 93.06/100) .
I am currently exploring how multimodal and video foundation models can reason more reliably, generate more efficiently, and maintain consistent dynamic world states over long contexts.
๐ Research
My academic exploration focuses on multimodal foundation models, video reasoning, reinforcement learning and efficient generative systems. Currently, my research is focused on:
- Multimodal Large Language Models and Vision-Language Reasoning ๐ญ
- Video Reasoning and Reinforcement Learning ๐ฌ
- Efficient Video Generation and Inference โก
- World Models ๐
๐ฅ News
- 2026.05: ๐ We released the paper Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding.
- 2026.05: ๐ We released the paper Focused Forcing: Content-Aware Per-Frame KV Selection for Efficient Autoregressive Video Diffusion.
- 2026.05: ๐ We released the paper VISD: Enhancing Video Reasoning via Structured Self-Distillation.
- 2026.02: ๐ We released the paper Not Just Whatโs There: Enabling CLIP to Comprehend Negated Visual Descriptions Without Fine-tuning.
- 2025.11: ๐ Our paper Not Just Whatโs There was accepted by AAAI 2026 as a poster.
- 2025.09: ๐๏ธ Honored to receive the National Scholarship.
- 2025.07: ๐ผ Started my internship at Baidu PaddleOCR, focusing on Multimodal Document Understanding and OCR Benchmarking.
- 2024.12: ๐๏ธ Honored to receive the National Scholarship.
๐ Publications

Not Just What's There: Enabling CLIP to Comprehend Negated Visual Descriptions Without Fine-tuning
A plug-and-play framework for improving CLIP's understanding of negated visual descriptions while keeping the CLIP backbone frozen.

VISD: Enhancing Video Reasoning via Structured Self-Distillation
A structured on-policy self-distillation framework for video reasoning RLVR, using teacher feedback on student rollouts for fine-grained credit assignment.
NeurIPS 2026 Under Review [arXiv] [project page] [code]

Focused Forcing: Content-Aware Per-Frame KV Selection for Efficient Autoregressive Video Diffusion
A training-free KV selection method that focuses cached history along generated-frame and attention-head dimensions for faster long video generation.
NeurIPS 2026 Under Review [arXiv]

Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding
A speculative decoding framework that decouples causal dependency modeling from autoregressive drafting, improving draft quality while keeping parallel drafting efficient.

MemeSleuth-Bench: Can Models Detect Chinese Internet Meme Origins Through Web Retrieval?
A benchmark for evaluating whether multimodal models can trace Chinese internet meme origins through web retrieval and culturally grounded evidence seeking.
ACM MM 2026 Under Review
๐ Honors and Awards
- 2026.03: ๐๏ธ FiberHome Telecommunication Scholarship(็ฝ็ซ้ไฟกไผไธๅฅๅญฆ้),FiberHome
- 2025.11: ๐ฅ National Second Prize, Challenge Cup โAI+โ Special Competition
- 2025.09: ๐๏ธ National Scholarship(ๅฝๅฎถๅฅๅญฆ้),Ministry of Education of China
- 2025.09: ๐๏ธAcademic Excellence Scholarship,HUST
- 2025.08: ๐ฅ National First Prize, China Robotics and Artificial Intelligence Competition, AI Innovation Track
- 2025.08: ๐ฅ National Third Prize, China Robotics and Artificial Intelligence Competition, AI Innovation Track
- 2025.06: ๐ฅ National Third Prize, Lanqiao Cup AI Practical Competition
- 2024.12: ๐๏ธ National Scholarship(ๅฝๅฎถๅฅๅญฆ้),Ministry of Education of China
- 2024.11: ๐ฅ First Prize in Hubei Province, Chinese Mathematics Competition(CMC)
- 2024.09: ๐๏ธMerit Student Scholarship,HUST
- 2024.03: ๐๏ธFreshman Self-Reliance Scholarship,HUST
๐ป Experience
Baidu PaddleOCR
Internship, 2025.07 - 2025.08
Topic: Multimodal Document Understanding and OCR Benchmarking

๐ Educations
2023.09 - Now
Undergraduate, School of Software Engineering, Huazhong University of Science and Technology
Major: Software Engineering
