Skip to content
View Jackey0903's full-sized avatar

Highlights

  • Pro

Block or report Jackey0903

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Jackey0903/README.md

Hi, I'm Haojie Hu (胡浩杰) 👋

Incoming Ph.D. Student · Shanghai Jiao Tong University × Shanghai Innovation Institute

Multimodal large language models · Video understanding & generation · World models


Website Email GitHub


🧑‍🔬 About Me

  • 🎓 B.Eng. in Software Engineering at Tongji University (2023 – 2027), and an incoming Ph.D. student at Shanghai Jiao Tong University, jointly trained at Shanghai Innovation Institute.
  • 🔬 I work on multimodal large language models, video understanding and generation, and world models — how models listen, look, reason, and revise.
  • 🛠️ I like research whose failures are visible, and I build tools that keep intermediate artifacts open to inspection.
  • 📫 Reach me at 3038115521@qq.com or through my homepage.

📝 Papers

PosterMELD PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs
Haojie Hu, Chenhao Dang, Yaojia Liu, Hengrui Kang, Conghui He, Weijia Li
arXiv:2608.02218
Paper · Code · Project page
Listening to the Motion Listening to the Motion: Audio-Conditioned Kinematic Verification for Robust Audio-Visual Segmentation
Under review
Code
To Think or Not to Think To Think or Not to Think: Pre-Decisional Reasoning Budgets for Referring Audio-Visual Segmentation
Under review
Code

🚀 Projects

DraftCode DraftCode: NBA Draft War Room
An auditable NBA draft prediction agent that fuses talent, expert mocks, and market signals across 30 GM personas and 1,500 Monte Carlo scenarios.
AWS Summit Shanghai 2026 Hackathon · Third place, advanced to the Macau round
VoxSprite VoxSprite
Turns any voice into a playable instrument with Web Audio, an ESP32-S3, physical keys, and reactive LEDs.
Xiaohongshu AI Builder · Excellence Award
Stardew-Valley Stardew-Valley
A Cocos2d-x systems project covering map interaction, character control, collision detection, inventory, and farming simulation mechanics.
Course project

🎓 Education

🏅 Honors

Award Awarded by Year
🏅 National Scholarship Ministry of Education of the People's Republic of China 2025
🏅 Qidi Scholarship Tongji University 2026
🏅 First-Class Outstanding Student Scholarship Tongji University 2024
🏅 Social Activity Scholarship Tongji University 2024, 2025
Outstanding Student Tongji University 2024, 2025
Computer Science Youth Pioneer School of Computer Science and Technology, Tongji University 2026

🛠️ Tech Stack

Skills

Popular repositories Loading

  1. PosterMELD PosterMELD Public

    PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs

    Python 4

  2. Stardew-Valley Stardew-Valley Public

    You have to work hard every day, just like planting crops. How can you harvest if you don’t sow seeds?

    C++ 2

  3. To-Think-or-Not-to-Think To-Think-or-Not-to-Think Public

    [Under review] To Think or Not to Think: Pre-Decisional Reasoning Budgets for Referring Audio-Visual Segmentation

    Python 2

  4. draftcode draftcode Public

    HTML 2

  5. SKA-VCT SKA-VCT Public

    [Under review] Listening to the Motion: Audio-Conditioned Kinematic Verification for Robust Audio-Visual Segmentation

    Python 1

  6. VoxSprite VoxSprite Public

    Turn any voice into a playable instrument with Web Audio, ESP32-S3, physical keys, and reactive LEDs.

    TypeScript 1 1