Yizhe Zhang
About Me
I am a Staff Research Scientist at Apple MLR. I build systems that close the loop between hypothesis and result: agents that write and run code, the environments and verifiers they train against, and post-training methods that let a model improve itself without a teacher. Recent work spans code agents (SWE-Gym, CodeAct), agentic evaluation (ToolSandbox, ASTRA-bench), self-improvement (SSD), latent reasoning (LaDiR, CLaRa, LaDi-RL), and text diffusion (DiffuCoder, CADD, FS-DFM). Before Apple, I was at Meta AI and Microsoft Research, where I worked on natural language generation and pre-training (including DialoGPT). I received my Ph.D. and M.S. degrees from Duke University. Before that, I received my B.Sc. degree in Physics from Nanjing University, Kuang Yaming Honors School, in 2011.
What I want to build
I want to build AI with genuine intuition—models that form a strong first guess, put it to the test, and learn from what comes back, getting better without a teacher. I think the shortest path runs through code, where a hunch is cheap to check, and through models that see the answer before they argue for it, rather than narrating every step. Read the full vision →
I want to build AI with genuine intuition—models that form a strong first guess, put it to the test, and learn from what comes back, getting better without a teacher. I think the shortest path runs through code, where a hunch is cheap to check, and through models that see the answer before they argue for it, rather than narrating every step. Read the full vision →
News
Apr 2026
SSD Released Simple self-distillation boosts Qwen3-30B from 42.4% to 55.3% pass@1 on LiveCodeBench v6—no external verifiers or teachers needed. Paper GitHub Feb 2026
LaDi-RL Released Latent reasoning + RL achieving 20.5 on AIME25 and 52.7 on LCB v6 for an 8B model with 2x faster reasoning. Surprisingly, RL for latent reasoning doesn't suffer from entropy/diversity collapse! PaperJan 2026
6 Papers Accepted to ICLR 2026 Our work on diffusion-based language models continues to advance, covering masked diffusion for code generation (DiffuCoder), latent diffusion for text reasoning (LaDiR), few-step diffusion for long text generation (FS-DFM), continuous augmentation for discrete diffusion (CADD), adaptive reward shaping for efficient reasoning (LASER), and Bayesian experimental design with LLMs (BED-LLM).Dec 2025
CLaRa Released CLaRa bridges retrieval and generation with continuous latent reasoning. 1k+ GitHub stars! GitHub Jul 2025
DiffuCoder Released Masked diffusion for code generation with Coupled-GRPO, achieving +4.4% on EvalPlus. GitHub Research Interests
My research focuses on giving language models stronger intuition and generalization, with the code domain as the main proving ground:
Code LLMs & Agents Coding models and autonomous software agents, and the environments and verifiers they learn against
AI Scientist Agents that form their own hypotheses, design and run experiments, and learn from the results
Long-Horizon Planning & Reasoning Multi-step reasoning, planning, and agentic benchmarks that stress long horizons
RAG & Latent Reasoning Retrieval-augmented generation and reasoning in a compressed, continuous latent space
Text Diffusion Models Non-autoregressive generation through diffusion, for text and code
Collaborations
I am always happy to hear from strong students and researchers interested in code agents, post-training and self-improvement, latent reasoning, text diffusion, and the AI scientist. Feel free to reach out by email with your latest CV.
I am always happy to hear from strong students and researchers interested in code agents, post-training and self-improvement, latent reasoning, text diffusion, and the AI scientist. Feel free to reach out by email with your latest CV.
Academic Service
Area Chair / Senior Program Committee:
ICLR 2023-2025 ICML 2022-2025 NeurIPS 2020-2025 ACL 2020-2021 EMNLP 2022 NAACL 2023 AAAI 2018-2021
Editorial Roles:
- Action Editor for Transactions on Machine Learning Research (TMLR, since 2023)
- Action Editor for ACL Rolling Review (ARR, since 2023)
Organization:
- Organization Committee Member, ACL 2020
Visitor Map
