Gaokai Zhang
An LLM researcher who gazes at the starlit skies of Artificial General Intelligence
M.S. in Intelligent Information Systems
Carnegie Mellon University
Language Technologies Institute
Hi there! I’m Gaokai Zhang, an M.S. student in Intelligent Information Systems at CMU LTI since Fall 2025. I hold dual B.S. degrees from ZJUI (CompE @ UIUC, ECE @ ZJU).
Currently, I’m interning at NVIDIA on agentic-workflow tooling and tokenomics, and contributing to sglang-omni. I also work with Yiqing Xie and Prof. Daniel Fried on code-generation data synthesis and agent training.
Open to LLM-related MLE/RS opportunities as I’m graduating in December 2026! 📄 Resume
Experience
NVIDIA (May 2026 - Present) Software Engineer Intern, AI Tools Team
- Built hook- and skill-based guardrails spanning 5 check categories (secrets, security, code quality, dependency, merge-conflict) that scan AI coding-agent actions in real time
- Contributed to token-saving agent-workflow skills benchmarked (N=50) on SWE-Bench Lite, Terminal-Bench, and CodeScaleBench: cut cost up to 38% with no accuracy loss, and +7% accuracy at -12% cost elsewhere
- Built session-level hooks that capture agent-event telemetry: per-session token/cost usage, model-routing decisions, and subagent fan-out (turns, tokens), feeding cost-governance dashboards
sglang-omni (Apr 2026 - Present) Open Source Contributor, SGLang Project
- Added native serving support for 5 TTS model families and unified them onto a shared serving scheduler, applying torch.compile + CUDA Graph, LRU caching, and radix-cache reuse to speed up decode by 5.5x
Microsoft Research Asia (Jul 2024 - Jul 2025) Research Intern, Systems & Networking Group Mentored by Dr. Li Lyna Zhang
- Led LoongRL: Novel data synthesis + reinforcement learning enabling 7B models to surpass 32B LRMs in long-context reasoning (100k-200k tokens) (ICLR 2026 Oral)
- Contributed to LongRoPE2: Extended LLM context windows to 128K tokens while retaining 98.5% short-context accuracy (ICML 2025 poster)
- Built parallel pipeline for large-scale user-query processing; delivered production-ready long-context recommendation models to Microsoft Asia-Pacific R&D
Carnegie Mellon University (Oct 2025 – Present) Research Assistant, Language Technologies Institute
- Synthetic task generation for training coding agents to generalize across repository-level environments (Hybrid-Gym, ICML 2026)
University of Illinois Urbana-Champaign Research Assistant with Prof. Fan Lai and Prof. Minjia Zhang
- Monte-Carlo-Tree-Search planning for cost-efficient LLM training on heterogeneous GPUs/TPUs
- Robustness benchmarking of LLMs (Stochastic Monkeys)
Research Interests
- Long-context reasoning & scaling
- Reinforcement learning for LLMs
- Efficient training architectures
- Code generation agents
Beyond Research
Outside work, I enjoy gaming (lifetime Faker fan), vibe to rap, and hunt for the perfect omakase bite.
Feel free to reach out - always happy to connect with like-minded friends and collaborators!
news
| Jul 01, 2026 | Hybrid-Gym accepted at ICML 2026! |
|---|---|
| May 11, 2026 | Started an internship at NVIDIA working on agentic-workflow tooling and tokenomics! |
| Apr 12, 2026 | Started contributing to sglang-omni, SGLang’s multi-modal serving engine! |
| Jan 26, 2026 | LoongRL accepted as Oral at ICLR 2026 - my first co-first-authored paper! |
| Aug 01, 2025 | Started M.S. in Intelligent Information Systems at CMU LTI! |