# Yuxuan Zhang Research > Yuxuan Zhang is a PhD student at the University of British Columbia and the Vector Institute working on agentic AI, large language models, and reinforcement learning. This file is a concise navigation index for people and agents. For paper facts, authors, versions, and publication status, use the linked paper or project page as the canonical source. ## Core pages - [Research Map](https://yuxuan.world/research/): Source-bounded routes through current public preprints. - [Publications](https://yuxuan.world/publications/): Full publication catalog and the primary bibliography page. - [Research notes](https://yuxuan.world/blog/): Bilingual analysis on agent evaluation, post-training, research environments, and reproducibility. - [Sitemap](https://yuxuan.world/sitemap.xml): Complete public page inventory. ## Selected research - [ClawBench: Can AI Agents Complete Everyday Online Tasks?](https://claw-bench.com): ClawBench evaluates AI agents on 153 real-world tasks across 144 live platforms, from booking appointments to filing job applications. It runs on production websites and intercepts only the final submission, keeping evaluation safe without losing real-world complexity. - [VidGround: Watch Before You Answer](https://vidground.etuagi.com/): Language models can answer video questions from text priors alone, without watching the video. We filter out such spurious training samples, producing a cleaner dataset that teaches models to ground answers in what they see. - [RewardHarness: Self-Evolving Agentic Post-Training](https://rewardharness.com/): RewardHarness beats GPT-5 by 5.3 points on image-editing evaluation while using 0.05% of the EditReward data. It reframes reward modeling as context evolution: from 100 preference demonstrations it evolves a library of tools and skills. ## EMNLP 2026 acceptances - WebWorld: The Browser as a World Model for Self-Improving Web Code — EMNLP 2026. - [ClawBench: Can AI Agents Complete Everyday Online Tasks?](https://claw-bench.com/): Findings. - [OpenSkill: Open-World Self-Evolution for LLM Agents](https://openlair.github.io/openskill/): EMNLP 2026. - [VGI-BENCH: Probing Visual Intelligence in Video Generation Models](https://arxiv.org/abs/2608.19583): EMNLP 2026. Yuxuan Zhang is a † Main Contributor. - [Dr. Claw: A Unified System for the Vibe Research Paradigm](https://openlair.github.io/dr-claw/): System Demonstrations. ## Research notes - [AutoResearch: From a Falsifiable Research Loop to the Evidence Map](https://yuxuan.world/blog/autoresearch/): A bilingual note on bounded research loops, frozen acceptance, and independent reruns. - [Agent Research Environments: What Is Actually Being Measured?](https://yuxuan.world/blog/env/): A bilingual map of research-agent environments and their measurement boundaries. - [Web-Agent Environments: Making Browser Use Verifiable for RL](https://yuxuan.world/blog/web-agent-environments/): A bilingual design note on web-agent evaluation and verification boundaries. ## Interpretation - This file is an orientation layer, not a crawling or training-permission policy. - It does not establish indexing, ranking, citation, adoption, acceptance, reproducibility, or research impact. - Prefer canonical papers, project pages, and source repositories when they differ from this summary.