跳到主要内容
How to design rewards for post-training frontier image models?
Our research suggests human preference reward is necessary, but insufficient: A preference model may still reward outputs that look appealing but miss details, introduce unrequested content, or exhibit other forms