Baizhou Zhang

Baizhou Zhang

I'm Baizhou Zhang, working as member of technical staff for RadixArk. As maintainer of SGLang, I mostly work on AI inference infrastructure. Before that I graduated from UC San Diego (MS) and Peking University (BS). I once worked at Nvidia CuDNN team, Baidu Paddle team and ColossalAI team as intern.

In my spare time, I like playing video games on Nintendo Switch/PS5, or playing classical piano pieces. Recently I'm playing Pokopia~

Education

Peking University, School of EECS

Sep 2019 – Jul 2024

Bachelor of Science, Intelligence Science and Technology

Beijing, China

UC Berkeley, Department of EECS

Aug 2022 – May 2023

Visiting Student

Berkeley, USA

UC San Diego, Jacobs School of Engineering

Sep 2024 – Dec 2025

Master of Science, Computer Science and Engineering

La Jolla, USA

Industry Experience

RadixArk

Mar 2026 – Present

Member of Technical Staff

Palo Alto, California, United States · On-site

  • Member of Founding Team; Core developer and maintainer of SGLang.
  • Co-led the optimization of GLM-5.2 model on Blackwell hardware. [Blog]
  • Co-led the design and implementation of the parallelism strategy for Kimi-K3 Day 0 support (Prefill PP + Decode CP) [Blog]
  • Maintainer of Prefill Context Parallelism feature for SGLang. Independently designed and implemented a cleaner version of PCP with CPStrategy abstraction. [Roadmap]
  • Co-led the initial implementation of NVFP4 KV Cache in SGLang [Blog]
  • DeepSeek V4 support and optimization
    • Implemented MTP speculative decoding and Trtllm MoE runner for DeepSeek V4 Day-0 support. [Blog]
    • Reviewed subsequent optimizations on DeepSeek V4 and helped increase its throughput to 5x compared with Day 0. [Blog]
  • Implemented a tvm-ffi wrapper for SGLang's DeepGemm fork, and created a customized PyPI wheel sgl-deep-gemm for more flexible development and usage.

SGLang Community

Jan 2025 – Mar 2026

Open Source Developer (Among top 10 contributors)

California, United States · Remote · Part-time

  • Core Developer and Maintainer of SGLang (among the top 10 committers).
  • Led the initial support of deterministic inference in SGLang. [Blog] [X post] [量子位]
  • Enhanced chunked prefill performance for the DeepSeek V3/R1 model by using a chunked prefix cache technique with MHA kernels, and achieved approximately 1.4x throughput improvement in long-context (32k) scenarios. [PR]
  • Co-led the initial support of SGLang on Blackwell architecture, and deployment of Large-Scale Expert Parallelism on GB200 and GB300. [GB200 blog part 1] [GB200 blog part 2] [GB300 blog]
  • Implementation of multiple attention backends for MLA, including FlashInfer backend ([PR1], [PR2]), and Flash Attention 3 backend ([PR]).
  • Maintainer of LoRA feature. [Roadmap]
  • Early development of srt-slurm, a framework that easily benchmarks models deployed with SGLang and NVIDIA Dynamo.

NVIDIA

Jun 2025 – Sep 2025

Software Engineer Intern

Santa Clara, California, United States · On-site

  • CuDNN Heuristic Team. Supported heuristics for compiler-generated Blackwell GEMM kernels.

Baidu, Inc.

Feb 2024 – May 2024

Software Engineer Intern

Beijing, China · On-site

HPC-AI Tech

Jun 2023 – Dec 2023

Software Engineer Intern

Beijing, China · On-site

  • Developer of ColossalAI, a distributed training framework.

Invited Talks

Full-stack support and optimization of DeepSeek V4 model in SGLang

May 9, 2026

Qingke Community

Online

[Video]

GTC Training Lab: High-Performance LLM Serving and Training with SGLang

Mar 19, 2026

RadixArk & Nvidia

San Jose, US

[Video] [X post] [Blog]

Update of SGLang Community and Future Roadmap

Feb 6, 2026

SGLang Shanghai Meetup / SGLang & Machine’s Heart

Shanghai, China

[WeChat Post]

End-to-end Deployment of DeepSeek Models on GB200 NVL72 with SGLang and Dynamo

Jan 22, 2026

Dynamo Day / Nvidia

Online

[Event] [Video]

Update of SGLang Community and Future Roadmap

Jan 18, 2026

Ant Open Source Meetup / SGLang & Ant Open Source

Hangzhou, China

[WeChat Post]

Deployment of DeepSeek Models on GB200 NVL72 with SGLang

Oct 2, 2025

SGLang & Dynamo Meetup / SGLang & Nvidia

San Francisco, US

[Video]

Research Experience

Peking University

Jan 2024 – May 2024

Research Assistant

Beijing, China · On-site

  • Applied QLoRA model quantization to cloud-device knowledge distillation.
  • Supervised by Prof. Shanghang Zhang.

NYU Shanghai

Feb 2022 – Sep 2022

Research Assistant

Remote

  • Researched music generation with VAE-based models.

Music Works

Piano Performance

  • Schubert: Piano Quintet in A major (Trout), D. 667
    IV. Andantino – Allegretto
    [Video] (3:48:20 – 3:57:00)
  • Gershwin: Rhapsody in Blue
    [Video]
  • Debussy: Piano Trio in G Major, L. 3
    I. Andantino con moto allegro
    III. Andante espressivo
    [Video]
  • Schubert: Piano Sonata in G Major, D. 894
    IV. Allegretto
    [Video]
  • Pierné: Violin Sonata, Op. 36
    I. Allegretto
    [Video]
  • Grieg: Violin Sonata No. 3 in C Minor, Op. 45
    I. Allegro molto ed appassionato
    [Video]
  • Schubert: Impromptu in E-sharp Major, D. 899 No. 2
    [Video]

Composition