Kai Liu (刘凯)
I will be graduating with a PhD in June 2027 and am currently seeking a research role in industry to continue advancing linear-complexity LLMs toward real-world deployment. Feel free to contact me at jackyliukai@outlook.com.
I am a jointly trained PhD student at Tongji University and the Shanghai AI Lab. As first author, I have published four top-tier conference papers, with one additional top-tier paper under review, and I have also contributed to multiple industry engineering projects.
I currently study linear-complexity LLMs from a systems perspective and am seeking a role where I can continue advancing them toward real-world deployment.
My key strengths are as follows:
- I have conducted systematic research on linear-complexity models and proposed a new paradigm of joint architecture-data optimization. Based on this paradigm, I believe linear-complexity models are likely to mature over the next one to two years and become commercially viable. My core research perspectives and contributions in linear-complexity LLMs are summarized as follows:
- Research goal: build linear-complexity LLMs that can compete with self-attention models as general-purpose models.
- Research paradigm: architectural improvements alone are not sufficient to close the gap; we need joint architecture-data optimization from a full-system perspective, coordinating data, training, and inference to narrow the gap further.
- Long-input setting: Smooth Reading improves data locality in long-input settings through agentic inference, strengthening long-context understanding and making linear-complexity LLMs competitive with self-attention LLMs.
- Long-output setting: architecture-aware reinforcement learning improves data locality for sliding-window attention models in long-output settings, enabling performance on mathematical reasoning tasks comparable to self-attention LLMs.
- Infrastructure: SimpleLLM provides a unified framework for training, inference, and evaluation of linear-complexity models.
- I also have broad experience across deep learning, including model architecture design, large-scale pre-training, supervised fine-tuning, reinforcement learning, and infrastructure development.
Education
- PhD, Tongji University & Shanghai AI Lab - Linear-complexity LLMs, 2024.9-2027.6 (expected)
- M.S., Sun Yat-sen University - Model Pruning, 2020.9-2022.7
- B.S., Minzu University of China - Sociology, 2016.9-2020.6
Work Experience
- Engineer, Shanghai AI Lab - Model Pruning & VLMs, 2022.7-2024.8
- Intern, Baidu Research - Vision Transformers, 2021.6-2022.1
Selected Research: Linear-Complexity LLMs
- [1] (ACL 2025) Scaling up the State Size of RNN LLMs for Long-Context Scenarios
Kai Liu, Jianfei Gao, Kai Chen - [2] (ICLR 2026) Smooth Reading: Bridging the Gap of Recurrent LLMs to Self-Attention LLMs on Long-Context Understanding
Kai Liu, Zhan Su, Peijie Dong, Fengran Mo, Jianfei Gao, Shaoting Zhang, Kai Chen - [3] (Under review) Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning
Kai Liu, Peijie Dong, Xinchen Xie, …, Xiaowen Chu, Shaoting Zhang, Kai Chen - [4] Engineering: SimpleLLM: A Self-Contained Infrastructure for Research
Unified infrastructure for SFT, RL, inference, and evaluation across self-attention and linear-complexity LLMs
Kai Liu
Other Publications and Projects
- [5] (IJCAI 2022) Dynamic Group Transformer: A General Vision Transformer Backbone with Dynamic Group Attention
Kai Liu, Tianyi Wu, Cong Liu, and Guodong Guo - [6] (ICML 2024) Differential Model Scaling using Differential Topk
Kai Liu, Ruohui Wang, Jianfei Gao, Kai Chen - [7] (Engineering) General Pruning Framework: Pruning Any Model Automatically in MMRazor (2022-2023)
Kai Liu, XTuner team - [8] (Engineering) XTuner: A Large-Scale LLM Training System Based on FSDP for InternLM (2024-2026)
XTuner team, Kai Liu (contributed to early design discussions and prototyping)
Engineering Skills and Experience
- Python and PyTorch for day-to-day research and engineering
- Training systems and inference engines for research-scale LLMs: SimpleLLM and XTuner
- CUDA/Triton kernels: CUDA kernel work in Dynamic Group Transformer and Triton kernel work in Scaling up the State Size of RNN LLMs
- PyTorch FX tracer: used in Differential Model Scaling using Differential Topk and MMRazor
I am willing to learn any engineering skills that help advance research goals, and I can quickly turn research ideas into working prototypes.
