GitHub

1The University of Hong Kong       2Nanjing University
3University of Chinese Academy of Sciences       4Nanyang Technological University
5Harbin Institute of Technology

(*Equal Contribution.    Corresponding Author.)

Paper | Project Page | LoRA Weights

About

We propose TACA, a parameter-efficient method that dynamically rebalances cross-modal attention in multimodal diffusion transformers to improve text-image alignment.

teaser.mp4

Usage

For Stable Diffusion 3.5, simply run:

python infer/infer_sd3.py

For FLUX.1, run:

python infer/infer_flux.py

Benchmark

Comparison of alignment evaluation on T2I-CompBench for FLUX.1-Dev-based and SD3.5-Medium-based models.

Model Attribute Binding Object Relationship Complex

Read the original on github.com ↗