1The University of Hong Kong
2Nanjing University
3University of Chinese Academy of Sciences
4Nanyang Technological University
5Harbin Institute of Technology
(*Equal Contribution. †Corresponding Author.)
Paper | Project Page | LoRA Weights
About
We propose TACA, a parameter-efficient method that dynamically rebalances cross-modal attention in multimodal diffusion transformers to improve text-image alignment.
teaser.mp4Usage
For Stable Diffusion 3.5, simply run:
python infer/infer_sd3.py
For FLUX.1, run:
python infer/infer_flux.py
Benchmark
Comparison of alignment evaluation on T2I-CompBench for FLUX.1-Dev-based and SD3.5-Medium-based models.
| Model | Attribute Binding | Object Relationship | Complex |
|---|