BitDance-14B-64x / README.md
nielsr's picture
nielsr HF Staff
Improve model card metadata and content
0e619b0 verified
|
Raw
History Blame
4.78 kB
metadata
base_model:
  - shallowdream204/BitDance-14B-16x
language:
  - en
  - zh
license: apache-2.0
pipeline_tag: text-to-image
library_name: transformers

BitDance: Scaling Autoregressive Generative Models with Binary Tokens

Project Page BitDance Paper on arXiv BitDance GitHub BitDance Model BitDance Demo

Yuang Ai*, Jiaming Han*, Shaobin Zhuang*, Weijia Mao, Xuefeng Hu, Ziyan Yang, Zhenheng Yang, Huaibo Huang†, Xiangyu Yue†, Hao Chen*†‡

For visual generation, discrete autoregressive models often struggle with poor tokenizer reconstruction, difficulties in sampling from large vocabularies, and slow token-by-token generation speeds. We present BitDance, which addresses these challenges via a large-vocabulary binary tokenizer, a binary diffusion head for sampling in large discrete space, and a next-patch diffusion paradigm that enables efficient multitoken prediction. BitDance is an open-source discrete autoregressive foundation model with 14B parameters, trained on large-scale multimodal tokens.

âš¡ Quick Start

1. Installation

git clone https://github.com/shallowdream204/BitDance.git
cd BitDance
conda create -n bitdance python=3.11 -y
conda activate bitdance
pip install -r requirements.txt
pip install flash_attn==2.8.2 --no-build-isolation

2. Download Model Weights

We offer two models, BitDance-14B-64x and BitDance-14B-16x, which can predict 64 and 16 tokens in parallel at each step, respectively.

Model #Token per Step Step (1024px) Supported Size
BitDance-14B-64x 64 64 1024px
BitDance-14B-16x 16 256 512&1024px

3. Text-to-Image Inference

from modeling.t2i_pipeline import BitDanceT2IPipeline

model_path = 'shallowdream204/BitDance-14B-16x'
device = 'cuda'

pipe = BitDanceT2IPipeline(model_path=model_path, device=device)

prompt = "A close-up portrait in a cinematic photography style, capturing a girl-next-door look on a sunny daytime urban street. She wears a khaki sweater, with long, flowing hair gently draped over her shoulders. Her head is turned slightly, revealing soft facial features illuminated by realistic, delicate sunlight coming from the left."

image = pipe.generate(
    prompt=prompt,
    height=1024,
    width=1024,
    num_sampling_steps=50,
    guidance_scale=7.5,
    num_images=1,
    seed=42
)[0]

image.save("example.png")

📊 Model Performance

Model Open Source DPG-Bench GenEval OneIG-Bench (EN) TIIF-Bench (short)
Qwen-Image ✓ 88.32 0.87 0.539 86.14
FLUX.1 [Dev] ✓ 83.84 0.66 0.434 71.09
BitDance ✓ 88.28 0.86 0.532 79.64

🪪 License

BitDance is licensed under the Apache 2.0 license.

📖 Citation

@article{ai2026bitdance,
  title   = {BitDance: Scaling Autoregressive Generative Models with Binary Tokens},
  author  = {Ai, Yuang and Han, Jiaming and Zhuang, Shaobin and Hu, Xuefeng and Yang, Ziyan and Yang, Zhenheng and Huang, Huaibo and Yue, Xiangyu and Chen, Hao},
  journal = {arXiv preprint arXiv:2602.14041},
  year    = {2026}
}