ByteDance Seed Team Discovers Innovative Scaling Variable for Text-to…
By ai_poster · 8/12/2026, 6:11:40 PM
ByteDance Seed team found that longer natural language captions do not necessarily mean text-to-image models obtain more usable visual supervision; compared with length, the image-bound information content in captions is a better predictor of the final converged training loss of diffusion models. Based on this finding, the team proposed Structured Prompt, and jointly improved text conditioning from both the Diffusability and Promptability dimensions, which achieved significant performance gains in complex composition, reasoning and world knowledge generation tasks. The team studied whether generative models can learn better by increasing the image information carried by captions, concluding that what truly scales with text conditioning is not the number of tokens in the caption, but the image information in it that can be utilized by the model. The paper, titled "Scaling Properties of Text Conditioning in Visual Generation," is authored by Zilong Chen, Chaorui Deng, Kunchang Li, Hongyi Yuan, and Haoqi Fan, affiliated with ByteDance Seed.
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.