Learning rewards for instruction-guided image editing
RewardHarness: Learning Human Preferences for Image Editing with Only 100 Demonstrations
Linked citation-source title: RewardHarness: Self-Evolving Agentic Post-Training
The heading follows the author-provided COLM 2026 published record. The linked arXiv preprint has a different title and author list. The abstract, authors below, machine-readable citation metadata, and CFF describe that linked preprint; the BibTeX download preserves the author-maintained citation repository. Choose the metadata for the version you actually use; do not combine the published title with the preprint author list.
Paper / source record · 2026-05-09
Research question and approach
Evaluating instruction-guided image edits requires rewards that reflect subtle human preferences, yet current reward models typically depend on large-scale preference annotation and additional model training. This creates a data-efficiency gap: humans can often infer the target evaluation criteria from only a few examples, while models are usually trained on hundreds of thousands of comparisons. We present RewardHarness, a self-evolving agentic reward framework that reframes reward modeling as context evolution rather than weight optimization.
When this work is relevant
Cite the appropriate version when discussing learning human preferences for image editing. The homepage publication title and the saved preprint BibTeX may differ; both are exposed explicitly.
Author-written abstract
Evaluating instruction-guided image edits requires rewards that reflect subtle human preferences, yet current reward models typically depend on large-scale preference annotation and additional model training. This creates a data-efficiency gap: humans can often infer the target evaluation criteria from only a few examples, while models are usually trained on hundreds of thousands of comparisons. We present RewardHarness, a self-evolving agentic reward framework that reframes reward modeling as context evolution rather than weight optimization. Instead of learning from large-scale annotations, RewardHarness aligns with human preferences by iteratively evolving a library of tools and skills from as few as 100 preference demonstrations. Given a source image, candidate edited images, and an editing instruction, an Orchestrator selects the most relevant subset of tools and skills from the maintained library, and a frozen Sub-Agent uses them to construct a reasoning chain that produces a preference judgment. By comparing predicted judgments with ground-truth preferences and analyzing successes and failures in the reasoning process, the Orchestrator automatically refines its library of tools and skills without additional human annotation. Using only 0.05% of the EditReward preference data, RewardHarness achieves 47.4% average accuracy on image-editing evaluation benchmarks, surpassing GPT-5 by 5.3 points. When used as a reward signal for GRPO fine-tuning, RL-tuned models achieve 3.52 on ImgEdit-Bench. Project page: https://rewardharness.com.
Abstract source: https://arxiv.org/abs/2605.08703. Checked 2026-09-14. Bibliographic metadata uses the linked paper record or author-maintained catalog. Results, limitations and experimental settings remain defined by the original source.
Citation
BibTeX is preserved from the author-maintained citation repository. The CFF uses the source metadata shown on this page. Version titles or author lists can differ; choose the version you used.
@article{zhang2026rewardharness,
title={RewardHarness: Self-Evolving Agentic Post-Training},
author={Zhang, Yuxuan and Du, Penghui and Li, Bo and Wei, Cong and Miao, Junwen and Zhang, Huaisong and Cai, Songcheng and Wang, Yubo and Jiang, Dongfu and Zhang, Yuyu and others},
journal={arXiv preprint arXiv:2605.08703},
year={2026}
}