Nanjing University and the Alibaba team proposed a causally interpretable framework and found that in the Diffusion Transformer (DiT), structural tokens fr...
论文研究HuggingFace Daily Papers(社区热门论文)
Today AI Intelligence Brief
Nanjing University and the Alibaba team proposed a causally interpretable framework and found that
in the Diffusion Transformer (DiT), structural tokens from chat templates, although not containing prompt semantics, become attention aggregation points from image to text, and causally maintain object identity, serving as implicit semantic registers.
Original Article Excerpt
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs :https://info.arxiv.org/labs/index.html.
The readable text on this page was extracted from the public source and organized with attribution, publication time and the original link. Copyright remains with the original author and publisher.
Nanjing University and the Alibaba team proposed a causally interpretable framework and found that in the Diffusion Transformer (DiT), structural tokens fr...