Hengshuo Chu
Logo 哈尔滨工业大学(深圳)

我的研究方向包括具身智能(Embodied AI)与多模态大模型(Multimodal Large Language Models)。

目前我正关注交互模型(Interaction Models)、流式与实时模型(Streaming & Real-Time Models)以及第一视角智能(Egocentric AI)等方向。


Education
  • 哈尔滨工业大学(深圳)
    哈尔滨工业大学(深圳)
    计算机技术学士
    2023 - 2026
  • 东北大学
    东北大学
    理学学士
    2019 - 2023
Experience
  • 蚂蚁集团 - 机器智能部门
    蚂蚁集团 - 机器智能部门
    算法工程师实习生
    2025.05 - 2025.09
  • 蚂蚁集团 - 机器智能部门
    蚂蚁集团 - 机器智能部门
    算法工程师
    2026.04 - Present
Honors & Awards
  • 研究生入学二等奖学金
Selected Publications (view all )
3D-AffordanceLLM: Harnessing Large Language Models for Open-Vocabulary Affordance Detection in 3D Worlds

Hengshuo Chu, Xiang Deng, Qi Lv, Xiaoyang Chen, Yinchuan Li, Jianye Hao, Liqiang Nie

The Thirteenth International Conference on Learning Representations (ICLR) 2025 Poster

3D affordance detection is a challenging problem with broad applications in robotic tasks. Existing methods typically formulate detection as label-based semantic segmentation, relying on predefined labels and offering limited ability to understand complex natural language or generalize to open-world scenes. To address these limitations, we reformulate affordance detection as the Instruction Reasoning Affordance Segmentation (IRAS) task, which predicts an affordance mask from a reasoning query without depending on fixed input categories. We propose 3D-AffordanceLLM (3D-ADLLM), a framework that introduces large language models into 3D affordance perception and uses a custom decoder to generate affordance masks. To mitigate the scarcity of training data, we further introduce a multi-stage strategy beginning with Referring Object Part Segmentation (ROPS) pre-training, followed by IRAS fine-tuning. By leveraging the world knowledge and human-object interaction reasoning capabilities of large language models, 3D-ADLLM improves open-vocabulary affordance detection by approximately 8% mIoU.

3D-AffordanceLLM: Harnessing Large Language Models for Open-Vocabulary Affordance Detection in 3D Worlds

Hengshuo Chu, Xiang Deng, Qi Lv, Xiaoyang Chen, Yinchuan Li, Jianye Hao, Liqiang Nie

The Thirteenth International Conference on Learning Representations (ICLR) 2025 Poster

3D affordance detection is a challenging problem with broad applications in robotic tasks. Existing methods typically formulate detection as label-based semantic segmentation, relying on predefined labels and offering limited ability to understand complex natural language or generalize to open-world scenes. To address these limitations, we reformulate affordance detection as the Instruction Reasoning Affordance Segmentation (IRAS) task, which predicts an affordance mask from a reasoning query without depending on fixed input categories. We propose 3D-AffordanceLLM (3D-ADLLM), a framework that introduces large language models into 3D affordance perception and uses a custom decoder to generate affordance masks. To mitigate the scarcity of training data, we further introduce a multi-stage strategy beginning with Referring Object Part Segmentation (ROPS) pre-training, followed by IRAS fine-tuning. By leveraging the world knowledge and human-object interaction reasoning capabilities of large language models, 3D-ADLLM improves open-vocabulary affordance detection by approximately 8% mIoU.

All publications