# Yuxuan Zhang research catalog

This Markdown index mirrors the canonical [JSON catalog](https://yuxuan.world/research/papers/catalog.json). For bibliographic details, versions, authors, and experimental claims, use each linked original source.

## An Attention-based Multi-Scale Feature Learning Network for Multimodal Medical Image Fusion

- Research guide: https://yuxuan.world/research/papers/dilran/
- Research focus: Multiscale feature learning for multimodal medical image fusion
- Original paper/source: https://arxiv.org/abs/2212.04661

## PLAICraft: Large-Scale Time-Aligned Vision-Speech-Action Dataset for Embodied AI

- Research guide: https://yuxuan.world/research/papers/plaicraft/
- Research focus: Time-aligned vision, speech and action data for embodied AI
- Original paper/source: https://arxiv.org/abs/2505.12707

## Dr. Bench: A Multidimensional Evaluation for Deep Research Agents, from Answers to Reports

- Research guide: https://yuxuan.world/research/papers/deep-research-bench/
- Research focus: Evaluating deep-research agents from answers to reports
- Original paper/source: https://arxiv.org/abs/2510.02190

## WikiGap: Promoting Epistemic Equity by Surfacing Knowledge Gaps Between English Wikipedia and other Language Editions

- Research guide: https://yuxuan.world/research/papers/wikigap/
- Research focus: Finding knowledge gaps across Wikipedia language editions
- Original paper/source: https://arxiv.org/abs/2505.24195

## StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs

- Research guide: https://yuxuan.world/research/papers/structeval/
- Research focus: Evaluating structured output generation and format conversion
- Original paper/source: https://arxiv.org/abs/2505.20139

## Retri3D: 3D Neural Graphics Representation Retrieval

- Research guide: https://yuxuan.world/research/papers/retri3d/
- Research focus: Retrieving neural graphics representations of 3D scenes
- Original paper/source: https://openreview.net/forum?id=q3EbOXb4y1

## ScholarCopilot: Training Large Language Models for Academic Writing with Accurate Citations

- Research guide: https://yuxuan.world/research/papers/scholarcopilot/
- Research focus: Academic writing with learned citation retrieval
- Original paper/source: https://arxiv.org/abs/2504.00824

## VideoScore2: Think before You Score in Generative Video Evaluation

- Research guide: https://yuxuan.world/research/papers/videoscore2/
- Research focus: Reasoning-based evaluation of generated videos
- Original paper/source: https://arxiv.org/abs/2509.22799

## Enhancing Vector Quantization with Distributional Matching: A Theoretical and Empirical Study

- Research guide: https://yuxuan.world/research/papers/vq/
- Research focus: Distributional matching for vector quantization
- Original paper/source: https://arxiv.org/abs/2506.15078

## Edge-Enhanced Dilated Residual Attention Network for Multimodal Medical Image Fusion

- Research guide: https://yuxuan.world/research/papers/dran/
- Research focus: Edge-enhanced multimodal medical image fusion
- Original paper/source: https://arxiv.org/abs/2411.11799

## S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?

- Research guide: https://yuxuan.world/research/papers/s3gym/
- Research focus: Testing self-improvement through self-testing and self-judging
- Original paper/source: https://arxiv.org/abs/2608.31100

## RewardHarness: Learning Human Preferences for Image Editing with Only 100 Demonstrations

- Research guide: https://yuxuan.world/research/papers/rewardharness/
- Research focus: Learning rewards for instruction-guided image editing
- Original paper/source: https://arxiv.org/abs/2605.08703

## Aspire: Can Models Self-Evolve from Vague Goals?

- Research guide: https://yuxuan.world/research/papers/aspire/
- Research focus: Agent self-evolution from vague goals
- Original paper/source: https://arxiv.org/abs/2608.31111

## CompRank: Efficient LLM Reranking via Token-Level Compression and Decoding-Free Scoring

- Research guide: https://yuxuan.world/research/papers/comprank/
- Research focus: Efficient language-model reranking
- Original paper/source: https://arxiv.org/abs/2606.11700

## Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback

- Research guide: https://yuxuan.world/research/papers/structured-defect-grounding/
- Research focus: Localized defect feedback for text-to-image generation
- Original paper/source: https://arxiv.org/abs/2606.06113

## HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?

- Research guide: https://yuxuan.world/research/papers/harnessdev/
- Research focus: Creating and evolving model-external agent harnesses
- Original paper/source: https://arxiv.org/abs/2609.01437

## MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection

- Research guide: https://yuxuan.world/research/papers/mira/
- Research focus: Source-aware data selection for mid-training
- Original paper/source: https://arxiv.org/abs/2605.30288

## WebWorld: The Browser as a World Model for Self-Improving Web Code

- Research guide: https://yuxuan.world/research/papers/webworld/
- Research focus: Browser-grounded evaluation for self-improving web code
- Original paper/source: https://arxiv.org/abs/2608.30530

## OpenSkill: Open-World Self-Evolution for LLM Agents

- Research guide: https://yuxuan.world/research/papers/openskill/
- Research focus: Open-world self-evolution for language-model agents
- Original paper/source: https://arxiv.org/abs/2606.06741

## Watch Before You Answer: Learning from Visually Grounded Post-Training

- Research guide: https://yuxuan.world/research/papers/vid-filter/
- Research focus: Video post-training that depends on visual evidence
- Original paper/source: https://arxiv.org/abs/2604.05117

## Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models

- Research guide: https://yuxuan.world/research/papers/fim-midtraining/
- Research focus: Function-aware fill-in-the-middle training for coding agents
- Original paper/source: https://arxiv.org/abs/2607.12463

## MedClaw: Heuristic Agent Harness for Long-Horizon Surgical Video Reasoning

- Research guide: https://yuxuan.world/research/papers/medclaw/
- Research focus: Long-horizon reasoning over surgical videos
- Original paper/source: https://arxiv.org/abs/2608.14015

## ModularRSI: Toward Generalizable Harness RSI

- Research guide: https://yuxuan.world/research/papers/modularrsi/
- Research focus: Generalizable harness self-improvement
- Original paper/source: no public paper linked; see the research guide for primary project resources.

## Learning from the Self-future: On-policy Self-distillation for dLLMs

- Research guide: https://yuxuan.world/research/papers/self-future-dllm/
- Research focus: On-policy self-distillation for diffusion language models
- Original paper/source: https://arxiv.org/abs/2606.18195

## Distributional Matching for Vector Quantization: A Unified Theoretical and Empirical Framework

- Research guide: https://yuxuan.world/research/papers/distributional-matching-vq/
- Research focus: A unified treatment of distributional matching in vector quantization
- Original paper/source: https://arxiv.org/abs/2607.15933

## VGI-BENCH: Probing Visual Intelligence in Video Generation Models

- Research guide: https://yuxuan.world/research/papers/vgi-bench/
- Research focus: Evaluating visual intelligence in video generation models
- Original paper/source: https://arxiv.org/abs/2608.19583

## Dr. Claw: An AI Scientist Workspace for Vibe Research

- Research guide: https://yuxuan.world/research/papers/dr-claw/
- Research focus: An AI scientist workspace for research workflows
- Original paper/source: https://arxiv.org/abs/2609.00365

## ClawBench: Can AI Agents Complete Everyday Online Tasks?

- Research guide: https://yuxuan.world/research/papers/clawbench/
- Research focus: Evaluating browser agents on everyday online workflows
- Original paper/source: https://arxiv.org/abs/2604.08523

## CelebHair: A New Large-Scale Dataset for Hairstyle Recommendation Based on CelebA

- Research guide: https://yuxuan.world/research/papers/celebhair/
- Research focus: A dataset for hairstyle recommendation based on CelebA
- Original paper/source: https://arxiv.org/abs/2104.06885
