Gaokai Zhang
An LLM researcher who gazes at the starlit skies of Artificial General Intelligence
M.S. in Intelligent Information Systems
Carnegie Mellon University
Language Technologies Institute
Hi there! I’m Gaokai Zhang, an M.S. student in Intelligent Information Systems at CMU LTI since Fall 2025. I hold dual B.S. degrees from ZJUI (CompE @ UIUC, ECE @ ZJU).
I recently interned at NVIDIA on agentic-workflow tooling and agent cost efficiency, and I’m still contributing to SGLang-Omni. I also work with Yiqing Xie on code-generation data synthesis and agent training.
Open to LLM-related Machine Learning Engineer / Research Scientist roles — graduating December 2026! 📄 Resume
Experience
NVIDIA (May 2026 - Aug 2026) Software Engineer Intern, AI Tools Team
- Built the pluggable check API and Perforce support for a coding-agent guardrail with 900+ developers, 100K+ runs; shipped agent-workflow skills that enforce V-model implementation for different teams across the company
- Benchmarked harnesses and cost-saving techniques (model right-sizing, context compaction, routing extensions) on SWE-Bench Pro and Terminal-Bench 2.1: cut cost up to 59% at similar accuracy as Opus 4.8 on the same tasks
- Hardened the company-wide agent-telemetry pipeline (token/cost, model routing, subagent fan-out) behind cost-governance dashboards and added its field-allowlisting egress gate, making sure everything sensitive is reviewed
SGLang-Omni (Apr 2026 - Present) Open Source Contributor, SGLang Project
- Added native serving to SGLang-Omni for 4 text-to-speech (TTS) model families (PR #451, #609, #728), then sped up audio generation with torch.compile and CUDA Graphs (5.5x faster per audio frame, PR #728) plus prompt-reuse caching: 11 to 21% more throughput at equal accuracy on the full SeedTTS benchmark (PR #527)
- Co-initiated and planned the repo-wide text-to-speech model refactor (scheduling PR #937, engine startup, caching, audio decoding, etc.) (Issue #661), and built the shared streaming audio-decoder base used across models (PR #936)
Carnegie Mellon University (Oct 2025 – Present) Research Assistant, Language Technologies Institute
- Synthetic task generation for training coding agents to generalize across repository-level environments (Hybrid-Gym, ICML 2026)
Microsoft Research Asia (Jul 2024 - Jul 2025) Research Intern, Systems & Networking Group Mentored by Dr. Li Lyna Zhang
- Co-led LoongRL: Novel data synthesis + reinforcement learning enabling 7B models to surpass 32B LRMs in long-context reasoning (100k-200k tokens) (ICLR 2026 Oral)
- Contributed to LongRoPE2: Extended LLM context windows to 128K tokens while retaining 98.5% short-context accuracy (ICML 2025 poster)
- Built parallel pipeline for large-scale user-query processing; delivered production-ready long-context recommendation models to Microsoft Asia-Pacific R&D
University of Illinois Urbana-Champaign Research Assistant with Prof. Fan Lai and Prof. Minjia Zhang
- Monte-Carlo-Tree-Search planning for cost-efficient LLM training on heterogeneous GPUs/TPUs
- Robustness benchmarking of LLMs (Stochastic Monkeys)
Research Interests
- Long-context reasoning & scaling
- Reinforcement learning for LLMs
- Efficient training architectures
- Code generation agents
Beyond Research
Outside work, I enjoy gaming (lifetime Faker fan), vibe to rap, and hunt for the perfect omakase bite.
Feel free to reach out - always happy to connect with like-minded friends and collaborators!
news
| Jul 01, 2026 | Hybrid-Gym accepted at ICML 2026! |
|---|---|
| May 11, 2026 | Started an internship at NVIDIA working on agentic-workflow tooling and agent cost efficiency! |
| Apr 12, 2026 | Started contributing to sglang-omni, SGLang’s multi-modal serving engine! |
| Jan 26, 2026 | LoongRL accepted as Oral at ICLR 2026 - my first co-first-authored paper! |
| Aug 01, 2025 | Started M.S. in Intelligent Information Systems at CMU LTI! |