Gaokai Zhang

An LLM researcher who gazes at the starlit skies of Artificial General Intelligence

prof_pic.jpg

M.S. in Intelligent Information Systems

Carnegie Mellon University

Language Technologies Institute

Hi there! I’m Gaokai Zhang, an M.S. student in Intelligent Information Systems at CMU LTI since Fall 2025. I hold dual B.S. degrees from ZJUI (CompE @ UIUC, ECE @ ZJU).

I recently interned at NVIDIA on agentic-workflow tooling and agent cost efficiency, and I’m still contributing to SGLang-Omni. I also work with Yiqing Xie on code-generation data synthesis and agent training.

Open to LLM-related Machine Learning Engineer / Research Scientist roles — graduating December 2026! 📄 Resume


Experience

NVIDIA (May 2026 - Aug 2026) Software Engineer Intern, AI Tools Team

  • Built the pluggable check API and Perforce support for a coding-agent guardrail with 900+ developers, 100K+ runs; shipped agent-workflow skills that enforce V-model implementation for different teams across the company
  • Benchmarked harnesses and cost-saving techniques (model right-sizing, context compaction, routing extensions) on SWE-Bench Pro and Terminal-Bench 2.1: cut cost up to 59% at similar accuracy as Opus 4.8 on the same tasks
  • Hardened the company-wide agent-telemetry pipeline (token/cost, model routing, subagent fan-out) behind cost-governance dashboards and added its field-allowlisting egress gate, making sure everything sensitive is reviewed

SGLang-Omni (Apr 2026 - Present) Open Source Contributor, SGLang Project

  • Added native serving to SGLang-Omni for 4 text-to-speech (TTS) model families (PR #451, #609, #728), then sped up audio generation with torch.compile and CUDA Graphs (5.5x faster per audio frame, PR #728) plus prompt-reuse caching: 11 to 21% more throughput at equal accuracy on the full SeedTTS benchmark (PR #527)
  • Co-initiated and planned the repo-wide text-to-speech model refactor (scheduling PR #937, engine startup, caching, audio decoding, etc.) (Issue #661), and built the shared streaming audio-decoder base used across models (PR #936)

Carnegie Mellon University (Oct 2025 – Present) Research Assistant, Language Technologies Institute

  • Synthetic task generation for training coding agents to generalize across repository-level environments (Hybrid-Gym, ICML 2026)

Microsoft Research Asia (Jul 2024 - Jul 2025) Research Intern, Systems & Networking Group Mentored by Dr. Li Lyna Zhang

  • Co-led LoongRL: Novel data synthesis + reinforcement learning enabling 7B models to surpass 32B LRMs in long-context reasoning (100k-200k tokens) (ICLR 2026 Oral)
  • Contributed to LongRoPE2: Extended LLM context windows to 128K tokens while retaining 98.5% short-context accuracy (ICML 2025 poster)
  • Built parallel pipeline for large-scale user-query processing; delivered production-ready long-context recommendation models to Microsoft Asia-Pacific R&D

University of Illinois Urbana-Champaign Research Assistant with Prof. Fan Lai and Prof. Minjia Zhang

  • Monte-Carlo-Tree-Search planning for cost-efficient LLM training on heterogeneous GPUs/TPUs
  • Robustness benchmarking of LLMs (Stochastic Monkeys)

Research Interests

  • Long-context reasoning & scaling
  • Reinforcement learning for LLMs
  • Efficient training architectures
  • Code generation agents

Beyond Research

Outside work, I enjoy gaming (lifetime Faker fan), vibe to rap, and hunt for the perfect omakase bite.

Feel free to reach out - always happy to connect with like-minded friends and collaborators!

news

Jul 01, 2026 Hybrid-Gym accepted at ICML 2026!
May 11, 2026 Started an internship at NVIDIA working on agentic-workflow tooling and agent cost efficiency!
Apr 12, 2026 Started contributing to sglang-omni, SGLang’s multi-modal serving engine!
Jan 26, 2026 LoongRL accepted as Oral at ICLR 2026 - my first co-first-authored paper!
Aug 01, 2025 Started M.S. in Intelligent Information Systems at CMU LTI!

selected publications

  1. ICLR 2026 Oral
    LoongRL: Incentivizing Long-Context Reasoning in Large Language Models via Reinforcement Learning
    Siyuan Wang, Gaokai Zhang, Li Lyna Zhang, Ning Shang, Fan Yang, Dongyao Chen, and 1 more author
    International Conference on Learning Representations, 2026
  2. ICML 2025
    LongRoPE2: Near-Lossless LLM Context Window Scaling
    Ning Shang, Li Lyna Zhang, Siyuan Wang, Gaokai Zhang, Gilsinia Lopez, Fan Yang, and 2 more authors
    International Conference on Machine Learning, 2025
  3. ICML 2026
    Hybrid-Gym: Training Coding Agents to Generalize Across Tasks
    Yiqing Xie, Emmy Liu, Gaokai Zhang, Nachiket Kotalwar, Shubham Gandhi, Sathwik Acharya, and 4 more authors
    International Conference on Machine Learning, 2026