- Machine Learning · Multimodal Large Models · Human-Inspired AI Understanding *
I have always been curious about a simple question: Why do humans understand so well, yet machines—despite their scale—still don’t?
Modern models, even large ones, often rely on surface-level correlations, dense representations, or pattern matching. Their internal states rarely reflect the kind of organized, explainable knowledge that supports human reasoning.
My research goal is to explore how machine learning systems can develop more structured, interpretable, and multimodal forms of understanding, and to design model and agent architectures that better approximate how humans organize information.
This motivation drives all the directions I work in.
I am a first-year CS student at The University of Tokyo.
Currently a visiting student at CUHK-Shenzhen, working on multiagent prompt optimization.
- Multimodal video understanding, structured video representations
- LLM agents, prompt & structure optimization
- Representation learning, causal reasoning
- Multimodal models
Outside machine learning, I write quiet pieces of music, take still photographs, and travel to places that change how I see light and shape. I keep a small archive of these creative experiments on Bilibili.