Shaoyang Guo郭绍阳

Physics undergraduate at Peking University. Research intern at ByteDance Seed. Working on LLM/VLM post-training, RL/SFT, agents, benchmarking, and Physics of AI.

Peking University, School of Physics (Class of 2027) GPA 3.74/4.00 · National Scholarship · CPhO Gold Medalist
LLM/VLM RL/SFT/Agents Benchmarking Physics of AI
Shaoyang Guo at NeurIPS 2025

News

2025.07 Joined ByteDance Seed as research intern, working on VLM/LLM post-training.
2025.04 PHYBench preprint released on arXiv. Submitted to NeurIPS 2025.
2025.03 Started contributing to VLA models survey (action tokenization perspective).
2024.12 Awarded National Scholarship (top 1% at PKU).

Publications

PHYBench
NeurIPS 2025 (submitted)

PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Shi Qiu*, Shaoyang Guo*, et al.

A comprehensive physics perception and reasoning benchmark with 500 original problems contributed by 178 PKU students. Co-initiated the project and helped design evaluation and quality-control workflows.

VLA Survey
arXiv Preprint

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective

Authors including Shaoyang Guo

A survey on VLA models focusing on action representation. Responsible for the Raw Action chapter, reviewing 30+ key papers on end-to-end VLA architectures.

Experience

Jul 2025 – Present

Research Intern, VLM/LLM Post-Training

ByteDance Seed

Working on post-training and research automation for VLM/LLM systems, with emphasis on RL, SFT, mid-training, rollouts, data pipelines, and agent workflows.

  • Contributed to HiPhO-oriented RL, SFT, and mid-training work for Seed 2.0 models, improving reported Lite performance from 72.5 to 83.8.
  • Participated in mid-training runs at large compute scale and supported rollout pipelines for model improvement.
  • Built data and prompt pipelines for QA pairs, CoT compression, summaries, and SFT-to-RL transfer experiments.
  • Explored auto-research agent loops, adversarial pair agents, and agent-based research settings.
Feb 2025 – Sep 2025

Co-initiator & Co-first Author, PHYBench

Peking University, Eureka Lab

Co-initiated and co-led PHYBench, a physics perception and reasoning benchmark for LLMs.

  • Identified gaps in existing LLM physics evaluation and led the project from concept validation to a full data pipeline.
  • Organized 178 PKU students to build 500 high-quality original physics problems in 2 weeks.
  • Designed evaluation criteria and quality-control workflows for LLM physics reasoning.
  • Co-authored the arXiv preprint submitted to NeurIPS 2025.
Mar 2025 – Aug 2025

Research Assistant, VLA Survey

PsiRobot Lab, Peking University

Co-authored a survey on Vision-Language-Action models from an action-tokenization perspective.

  • Responsible for the Raw Action chapter; reviewed 30+ key papers on end-to-end VLA architectures.
  • Organized taxonomies for VLA model design and contributed to the arXiv preprint.

Blogs

Ideas and working notes on AI, physics, and research taste.

DeepSeek Harness × 246 trajectory audit
Agent Trajectory Audit

四百八十万事件之后:DeepSeek Harness 如何追一条 246 的证明路线

完整重建两条主会话的 11,651 个逻辑 step:模型做了什么、哪些数值证据成立、为什么它还没有收敛成 Axiom 级 Lean 证明。

DeepSeekAI for MathFormalization
Read →
Paper Podcast — 论文深度解读播客
论文播客

Paper Podcast — 论文深度解读播客

85 期双音色对话式论文解读,覆盖具身智能、可解释性、Agent·自进化与核心必读四个分区;每期 20–35 分钟,把论文架构、方法、实验与结论讲透。

Podcast论文解读AI
Listen →
ComfyResearch Puzzle:把训练曲线变成可判分的题
ComfyResearch Puzzle

把训练曲线变成一道可判分的题

从隐藏入口 puzzle1 到 m_equal_peaks 两峰等高:用可编辑训练图、明确 verifier contract 和 full-run evidence,把“看起来像 M 形”变成机器可判的研究问题。

PuzzlesComfyResearchReproducibility
Read →
Horsey Island — 一座 AI 的孤岛实验
实验报告

Horsey Island — 一座 AI 的孤岛实验

把几个"无辜的" agent 播种进一座没有网络的孤岛(一台 8×H200):它们能睡眠、苏醒、归档、死亡,能跑实验、写看板、验证假说、创造分身;不做任何引导,只观察一个问题——意义的单位是单个 agent,还是整个社群?报告覆盖实验主线、agent 机制、物理马场、240 基因基因组、社群看板与实时快照。

AI AgentOffline SocietyPhysics of AI
Read →
Rank Collapse on TinyShakespeare
ComfyResearch Reproduction

Rank Collapse on TinyShakespeare

Reproducing the spectral rank-collapse signature inside a standard ComfyResearch canvas: real text, editable nodes, and live RankMe / alphaReQ curves. We lock the setting at lr=5e-3, wd=1e-3, 20k steps and show how bottleneck width controls the collapse.

Physics of AIRank CollapseComfyResearch
Read →
ArchitectureIQ 项目全面 Review
Project Review

ArchitectureIQ 项目全面 Review

A comprehensive review of ArchitectureIQ's question-generation pipeline, significance tests, evaluation protocol, meta-model results, conclusions, and open problems.

ArchitectureIQEvaluationMeta-Model
Read →
ArchitectureIQ Meta-Model:预测训练设置结果的基模
Research Report

ArchitectureIQ Meta-Model:预测训练设置结果的基模

A tabular base model that predicts which training setting wins before training: 82.11% significant three-choice accuracy on 30 environments, beating tested LLMs on frozen questions. Includes a path toward a world model for AI4AI.

ArchitectureIQMeta-ModelAI4AI
Read →
N-gram Gap 机制指南
Mechanism Guide

N-gram Gap 机制指南

A visual guide to the N-gram Gap mechanism, including global N-gram frequency, validation loss, contribution analysis, and the training cliff.

N-gramLanguage ModelsMechanistic Analysis
Read →
N-gram Gap Regime Bridge
Regime Bridge

N-gram Gap Regime Bridge

An interactive roadmap connecting top-down and bottom-up evidence for the N-gram Gap: from reduced positives to order controls and observable curve ablations.

N-gramInteractiveExperiments
Read →
What makes a STEM benchmark actually useful?
Planned Essay

What makes a STEM benchmark actually useful?

Notes on building evaluations that reveal real reasoning capability rather than benchmark-specific pattern matching, with lessons from PHYBench.

BenchmarkingPhysicsEvaluation
Draft coming soon
Views on large model training
Writing Plan

Views on large model training

A continuing series for organizing personal views on post-training, data quality, RL/SFT dynamics, and the practical craft of making models better.

Post-TrainingData QualityVLM
Draft coming soon
From physics olympiad to AI research
Personal Note

From physics olympiad to AI research

Reflections on how physics training shapes taste in AI research: problem selection, abstraction, experiments, and long-term curiosity.

ResearchPhysicsPersonal
Draft coming soon

Education & Honors

Peking University, School of Physics

B.S. in Physics, expected Jun 2027. GPA 3.74/4.00, top 10% in the School of Physics; completed 141/149 credits by sophomore year including 3 graduate courses.

National Scholarship (2024)

Ministry of Education, top 1% at Peking University.

Chinese Physics Olympiad Gold Medal

National rank #57 (2022). Admitted to PKU Physics via PKU Excellence Program.

NOIP First Prize (2020)

National Olympiad in Informatics in Provinces.

Contact

Open to research collaboration, especially in VLM post-training, evaluation, and embodied intelligence.