Kai Liu (刘凯)

I will be graduating with a PhD in June 2027 and am currently seeking a research role in industry to continue advancing linear-complexity LLMs toward real-world deployment. Feel free to contact me at jackyliukai@outlook.com.

I am a jointly trained PhD student at Tongji University and the Shanghai AI Lab. As first author, I have published four top-tier conference papers, with one additional top-tier paper under review, and I have also contributed to multiple industry engineering projects.

I currently study linear-complexity LLMs from a systems perspective and am seeking a role where I can continue advancing them toward real-world deployment.

My key strengths are as follows:

  • I have conducted systematic research on linear-complexity models and proposed a new paradigm of joint architecture-data optimization. Based on this paradigm, I believe linear-complexity models are likely to mature over the next one to two years and become commercially viable. My core research perspectives and contributions in linear-complexity LLMs are summarized as follows:
    • Research goal: build linear-complexity LLMs that can compete with self-attention models as general-purpose models.
    • Research paradigm: architectural improvements alone are not sufficient to close the gap; we need joint architecture-data optimization from a full-system perspective, coordinating data, training, and inference to narrow the gap further.
    • Long-input setting: Smooth Reading improves data locality in long-input settings through agentic inference, strengthening long-context understanding and making linear-complexity LLMs competitive with self-attention LLMs.
    • Long-output setting: architecture-aware reinforcement learning improves data locality for sliding-window attention models in long-output settings, enabling performance on mathematical reasoning tasks comparable to self-attention LLMs.
    • Infrastructure: SimpleLLM provides a unified framework for training, inference, and evaluation of linear-complexity models.
  • I also have broad experience across deep learning, including model architecture design, large-scale pre-training, supervised fine-tuning, reinforcement learning, and infrastructure development.

Education

  • PhD, Tongji University & Shanghai AI Lab - Linear-complexity LLMs, 2024.9-2027.6 (expected)
  • M.S., Sun Yat-sen University - Model Pruning, 2020.9-2022.7
  • B.S., Minzu University of China - Sociology, 2016.9-2020.6

Work Experience

  • Engineer, Shanghai AI Lab - Model Pruning & VLMs, 2022.7-2024.8
  • Intern, Baidu Research - Vision Transformers, 2021.6-2022.1

Selected Research: Linear-Complexity LLMs

Other Publications and Projects

Engineering Skills and Experience

  • Python and PyTorch for day-to-day research and engineering
  • Training systems and inference engines for research-scale LLMs: SimpleLLM and XTuner
  • CUDA/Triton kernels: CUDA kernel work in Dynamic Group Transformer and Triton kernel work in Scaling up the State Size of RNN LLMs
  • PyTorch FX tracer: used in Differential Model Scaling using Differential Topk and MMRazor

I am willing to learn any engineering skills that help advance research goals, and I can quickly turn research ideas into working prototypes.