China
Foshan University
6D object pose estimation, Feature enhancement, Instance reconstruction, Category-level
Multimodal large language models, Cross-video reasoning, Video question answering, Multivideo understanding, Benchmark evaluation.