Research program · 2026
Making agent progress testable
Real environments → grounded evidence → inspectable learning → governed improvement.
研究计划 · 2026
让智能体的进步可以被检验
真实环境 → 证据约束 → 可检查学习 → 受治理改进。
I study when AI agents can complete constrained work in realistic environments, how their outcomes should be evaluated, and how feedback can improve future behavior without letting the system control its own acceptance criteria.
Each work keeps its own evidence, version, status, and limitation. A project page, benchmark score, or public artifact is a route for inspection—not a universal capability claim.
我研究 AI agent 何时能在真实环境中完成有约束的工作、怎样可靠地评估结果, 以及如何利用反馈改进后续行为,同时避免系统控制自己的接纳标准。
每项工作保留自己的证据、版本、状态和限制。项目页、benchmark 分数和公开工件 是供读者检查的入口,不等于对通用能力的宣判。
Agents in the real world
Can an agent finish the task—not merely interact with the interface?
真实环境中的 Agent
Agent 能否真正完成任务,而不只是操作界面?
- ClawBenchlive-web evaluation · public preprint / under review
- Dr. Clawopen-source research workspace
- Dr. Benchdeep-research-agent evaluation · under review
Inspectable self-improvement
Can feedback improve the system while its acceptance gate stays independent?
可检查的自我改进
反馈能否改进系统,同时让接纳标准保持独立?
- RewardHarnessskills-and-tools reward evolution · COLM 2026
- OpenSkillopen-world skill evolution · public preprint / under review
- d-OPSDon-policy self-distillation for diffusion LLMs · public preprint / under review
Efficient learning and information use
Which data, tokens, and training signals actually carry the useful information?
高效学习与信息利用
哪些数据、token 和训练信号真正承载有效信息?
- Function-Aware FIMmid-training for coding-agent foundations · public preprint / under review
- MIRAsource-aware data selection · public preprint / under review
- CompRanktoken-efficient LLM reranking · public preprint / under review
- ScholarCopilotcitation-grounded academic writing · COLM 2025
Evidence-grounded multimodal intelligence
Did the model use the evidence it claims to understand or evaluate?
有证据约束的多模态智能
模型是否真正使用了它声称理解或评价的证据?
- Watch Before You Answervisually grounded video post-training · public preprint / under review
- Structured Defect Groundinglocalized text-to-image feedback · public preprint / under review
- VideoScore2generative-video evaluation · TMLR 2026
- StructEvalstructured-output evaluation · TMLR 2025
- Retri3D3D representation retrieval · ICLR 2025 Spotlight
One program, four interfaces
Environment → evidence → learning → acceptance
The environment defines what the agent can observe and change. Evidence records what happened. Learning uses that evidence to propose an improvement. An independent acceptance gate decides whether the change survives. The next cycle starts only from accepted changes.
一个计划,四个接口
环境 → 证据 → 学习 → 接纳
环境定义 agent 能观察和改变什么;证据记录实际发生了什么;学习根据证据提出改进; 独立接纳门决定改动是否保留。下一轮只从被接纳的改动开始。
Choose your route
Read, inspect, or collaborate
选择你的入口