Yizhe Zhang

About Me

I am a Staff Research Scientist at Apple MLR. I build systems that close the loop between hypothesis and result: agents that write and run code, the environments and verifiers they train against, and post-training methods that let a model improve itself without a teacher. Recent work spans code agents (SWE-Gym, CodeAct), agentic evaluation (ToolSandbox, ASTRA-bench), self-improvement (SSD), latent reasoning (LaDiR, CLaRa, LaDi-RL), and text diffusion (DiffuCoder, CADD, FS-DFM). Before Apple, I was at Meta AI and Microsoft Research, where I worked on natural language generation and pre-training (including DialoGPT). I received my Ph.D. and M.S. degrees from Duke University. Before that, I received my B.Sc. degree in Physics from Nanjing University, Kuang Yaming Honors School, in 2011.
What I want to build
I want to build AI with genuine intuition—models that form a strong first guess, put it to the test, and learn from what comes back, getting better without a teacher. I think the shortest path runs through code, where a hunch is cheap to check, and through models that see the answer before they argue for it, rather than narrating every step. Read the full vision →

News
Apr 2026
SSD Released Simple self-distillation boosts Qwen3-30B from 42.4% to 55.3% pass@1 on LiveCodeBench v6—no external verifiers or teachers needed. Paper GitHub GitHub stars
Feb 2026
LaDi-RL Released Latent reasoning + RL achieving 20.5 on AIME25 and 52.7 on LCB v6 for an 8B model with 2x faster reasoning. Surprisingly, RL for latent reasoning doesn't suffer from entropy/diversity collapse! Paper
Jan 2026
6 Papers Accepted to ICLR 2026 Our work on diffusion-based language models continues to advance, covering masked diffusion for code generation (DiffuCoder), latent diffusion for text reasoning (LaDiR), few-step diffusion for long text generation (FS-DFM), continuous augmentation for discrete diffusion (CADD), adaptive reward shaping for efficient reasoning (LASER), and Bayesian experimental design with LLMs (BED-LLM).
Dec 2025
CLaRa Released CLaRa bridges retrieval and generation with continuous latent reasoning. 1k+ GitHub stars! GitHub GitHub stars
Jul 2025
DiffuCoder Released Masked diffusion for code generation with Coupled-GRPO, achieving +4.4% on EvalPlus. GitHub GitHub stars
May 2025
SWE-Gym at ICML 2025 A training environment for software-engineering agents and verifiers—now 1.8M+ dataset downloads/month on Hugging Face. Training on ~500 trajectories yields strong SWE-bench gains, and a learned verifier gives log-linear inference-time scaling. Paper GitHub GitHub stars

Research Interests

My research focuses on giving language models stronger intuition and generalization, with the code domain as the main proving ground:

Collaborations
I am always happy to hear from strong students and researchers interested in code agents, post-training and self-improvement, latent reasoning, text diffusion, and the AI scientist. Feel free to reach out by email with your latest CV.

Academic Service

Area Chair / Senior Program Committee:

ICLR 2023-2025 ICML 2022-2025 NeurIPS 2020-2025 ACL 2020-2021 EMNLP 2022 NAACL 2023 AAAI 2018-2021

Editorial Roles:

  • Action Editor for Transactions on Machine Learning Research (TMLR, since 2023)
  • Action Editor for ACL Rolling Review (ARR, since 2023)

Organization:

  • Organization Committee Member, ACL 2020

Visitor Map