--- tags: - ahead-of-time - pytorch library_name: diffusers --- > [!NOTE] > This **README** has been auto-generated by the **HF Job** run linked below > and the whole repository is a reproducible artifact of this Job # Ahead-of-time repository AoT repos contain **pre-compiled binaries** of PyTorch models, enabling: - fast startup times (no `torch.compile` needed) - significant **speedup** - **ZeroGPU** compatibility ## How to use ``` python import os import tempfile import numpy as np import torch import spaces from PIL import Image from diffusers import LTX2InContextPipeline from diffusers.pipelines.ltx2.pipeline_ltx2_ic_lora import LTX2ReferenceCondition from diffusers.pipelines.ltx2.utils import DISTILLED_SIGMA_VALUES # base == distilled in architecture, so one compiled graph serves both; distilled is # what most demos use. The AOTI package is weight-agnostic, so this base graph also # serves any FUSED LoRA (fuse_lora before aoti_load on the Space). MODEL_ID = os.environ.get("LTX_MODEL_ID", "diffusers/LTX-2.3-Distilled-Diffusers") pipe = LTX2InContextPipeline.from_pretrained(MODEL_ID, torch_dtype=torch.bfloat16) pipe.to("cuda") pipe.vae.enable_tiling() spaces.aoti_load( module=pipe.transformer, repo_id='linoyts/LTX-2.3-Transformer-GroupA-sm120-cu130-r9e', ) ``` ## How to reproduce or customize ``` bash hf jobs uv run job.py --flavor rtx-pro-6000 --image pytorch/pytorch:2.9.1-cuda13.0-cudnn9-devel --secrets HF_TOKEN ``` ## Samples | Before compilation (0.37s) | After compilation (0.30s) | |---|---| | | | Speedup: **1.21x** ## Environment
Click to expand ``` PyTorch version: 2.12.0+cu130 Is debug build: False CUDA used to build PyTorch: 13.0 ROCM used to build PyTorch: N/A OS: Ubuntu 22.04.5 LTS (x86_64) GCC version: (Ubuntu 11.4.0-1ubuntu1~22.04) 11.4.0 Clang version: Could not collect CMake version: version 4.1.2 Libc version: glibc-2.35 Python version: 3.10.19 (main, Oct 31 2025, 23:02:46) [Clang 21.1.4 ] (64-bit runtime) Python platform: Linux-6.12.88-119.157.amzn2023.x86_64-x86_64-with-glibc2.35 Is CUDA available: True CUDA runtime version: 13.0.48 CUDA_MODULE_LOADING set to: GPU models and configuration: GPU 0: NVIDIA RTX PRO 6000 Blackwell Server Edition Nvidia driver version: 580.159.03 cuDNN version: Could not collect Is XPU available: False HIP runtime version: N/A MIOpen runtime version: N/A Is XNNPACK available: True Caching allocator config: N/A CPU: Architecture: x86_64 CPU op-mode(s): 32-bit, 64-bit Address sizes: 46 bits physical, 48 bits virtual Byte Order: Little Endian CPU(s): 192 On-line CPU(s) list: 0-191 Vendor ID: GenuineIntel Model name: Intel(R) Xeon(R) Platinum 8559C CPU family: 6 Model: 207 Thread(s) per core: 2 Core(s) per socket: 48 Socket(s): 2 Stepping: 2 BogoMIPS: 4800.00 Flags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mmx fxsr sse sse2 ss ht syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon rep_good nopl xtopology nonstop_tsc cpuid aperfmperf tsc_known_freq pni pclmulqdq monitor ssse3 fma cx16 pdcm pcid sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand hypervisor lahf_lm abm 3dnowprefetch cpuid_fault ssbd ibrs ibpb stibp ibrs_enhanced fsgsbase tsc_adjust bmi1 avx2 smep bmi2 erms invpcid avx512f avx512dq rdseed adx smap avx512ifma clflushopt clwb avx512cd sha_ni avx512bw avx512vl xsaveopt xsavec xgetbv1 xsaves avx_vnni avx512_bf16 wbnoinvd ida arat avx512vbmi umip pku ospke waitpkg avx512_vbmi2 gfni vaes vpclmulqdq avx512_vnni avx512_bitalg tme avx512_vpopcntdq rdpid cldemote movdiri movdir64b md_clear serialize amx_bf16 avx512_fp16 amx_tile amx_int8 flush_l1d arch_capabilities Hypervisor vendor: KVM Virtualization type: full L1d cache: 4.5 MiB (96 instances) L1i cache: 3 MiB (96 instances) L2 cache: 192 MiB (96 instances) L3 cache: 640 MiB (2 instances) NUMA node(s): 2 NUMA node0 CPU(s): 0-47,96-143 NUMA node1 CPU(s): 48-95,144-191 Vulnerability Gather data sampling: Not affected Vulnerability Indirect target selection: Not affected Vulnerability Itlb multihit: Not affected Vulnerability L1tf: Not affected Vulnerability Mds: Not affected Vulnerability Meltdown: Not affected Vulnerability Mmio stale data: Not affected Vulnerability Reg file data sampling: Not affected Vulnerability Retbleed: Not affected Vulnerability Spec rstack overflow: Not affected Vulnerability Spec store bypass: Mitigation; Speculative Store Bypass disabled via prctl Vulnerability Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization Vulnerability Spectre v2: Mitigation; Enhanced / Automatic IBRS; IBPB conditional; PBRSB-eIBRS SW sequence; BHI BHI_DIS_S Vulnerability Srbds: Not affected Vulnerability Tsa: Not affected Vulnerability Tsx async abort: Not affected Vulnerability Vmscape: Not affected Versions of relevant libraries: [pip3] Could not collect [conda] numpy 2.3.4 py311h2e04523_0 conda-forge [conda] nvidia-cublas 13.0.0.19 pypi_0 pypi [conda] nvidia-cuda-cupti 13.0.48 pypi_0 pypi [conda] nvidia-cuda-nvrtc 13.0.48 pypi_0 pypi [conda] nvidia-cuda-runtime 13.0.48 pypi_0 pypi [conda] nvidia-cudnn-cu13 9.13.0.50 pypi_0 pypi [conda] nvidia-cufft 12.0.0.15 pypi_0 pypi [conda] nvidia-curand 10.4.0.35 pypi_0 pypi [conda] nvidia-cusolver 12.0.3.29 pypi_0 pypi [conda] nvidia-cusparse 12.6.2.49 pypi_0 pypi [conda] nvidia-cusparselt-cu13 0.8.0 pypi_0 pypi [conda] nvidia-nccl-cu13 2.27.7 pypi_0 pypi [conda] nvidia-nvjitlink 13.0.39 pypi_0 pypi [conda] nvidia-nvtx 13.0.39 pypi_0 pypi [conda] optree 0.17.0 pypi_0 pypi [conda] torch 2.9.1+cu130 pypi_0 pypi [conda] torchaudio 2.9.1+cu130 pypi_0 pypi [conda] torchelastic 0.2.2 pypi_0 pypi [conda] torchvision 0.24.1+cu130 pypi_0 pypi [conda] triton 3.5.1 pypi_0 pypi ```
## Job run - [linoyts/6a326f7f5ff0a6cf94fa00e8](https://huggingface.co/jobs/linoyts/6a326f7f5ff0a6cf94fa00e8)