I am a researcher working on foundation models and agents.
I am currently affiliated with z.ai (Zhipu AI),
contributing to the GLM series of foundation models.
I have been a member of GLM team since 2023, supervised by Prof. Jie Tang at Tsinghua University.
Before that, I received my Bachelor's and Master's degrees at Sun Yat-sen University, under the supervision of Prof. Daifeng Li.
My research interests include foundation models, agents, and physical intelligence.
A foundation model that natively unifies visual perception with reasoning, planning, and tool use for multimodal agents. Excels in visual coding and framework-based agent operations while maintaining strong text-based coding capability.
A family of vision-language models supporting both thinking and non-thinking modes, trained with Reinforcement Learning with Curriculum Sampling (RLCS). Achieves state-of-the-art performance among comparably sized open-source VLMs across STEM, video, GUI, and document understanding tasks.
A hierarchical benchmark spanning static UI-to-code, interactive multi-page reproduction, and long-horizon full-stack development. Comprises 193 tasks across 16 categories with 918 prototype images and 1,255 test cases, evaluated through a workflow-based agent verification paradigm.
Uses LLM-generated intent-preserving query variations to drive a contrastive training framework that substantially improves the robustness of neural IR rankers against query perturbations while preserving retrieval performance.
A multi-step sales forecasting framework that disentangles universal sequence patterns from instance-specific fluctuations, with a query-sparsity-measurement attention to scale to large multivariate time series for inventory optimization.