[TensorRT-LLM] TensorRT LLM version: 1.2.0rc6 max_position_embeddings: 8192 2 2 False True enable_block_reuse=True max_tokens=524288 max_attention_window=None sink_token_length=None free_gpu_memory_fraction=0.7 host_cache_size=None onboard_blocks=True cross_kv_cache_fraction=None secondary_offload_min_priority=None event_buffer_max_size=0 attention_dp_events_gather_period_ms=5 enable_partial_reuse=True copy_on_partial_reuse=True use_uvm=False max_gpu_total_bytes=0 dtype='auto' mamba_ssm_cache_dtype='auto' tokens_per_block=32 batch_sizes=None max_batch_size=64 enable_padding=True [12/20/2025-15:48:06] [TRT-LLM] [I] Using LLM with PyTorch backend [12/20/2025-15:48:06] [TRT-LLM] [W] Using default gpus_per_node: 2 [12/20/2025-15:48:06] [TRT-LLM] [I] neither checkpoint_format nor checkpoint_loader were provided, checkpoint_format will be set to HF. [TensorRT-LLM] TensorRT LLM version: 1.2.0rc6 [TensorRT-LLM] TensorRT LLM version: 1.2.0rc6 [TensorRT-LLM][INFO] Refreshed the MPI local session [TensorRT-LLM][INFO] Refreshed the MPI local session Model init total -- 15.59s Model init total -- 15.63s [TensorRT-LLM][INFO] Max KV cache blocks per sequence: 4097 [window size=16448], tokens per block=32, primary blocks=514, secondary blocks=0, max sequence length=131073 [TensorRT-LLM][INFO] Number of tokens per block: 32. [TensorRT-LLM][INFO] [MemUsageChange] Allocated 1.44 GiB for max tokens in paged KV cache (16448). [TensorRT-LLM][INFO] Max KV cache blocks per sequence: 4097 [window size=16448], tokens per block=32, primary blocks=514, secondary blocks=0, max sequence length=131073 [TensorRT-LLM][INFO] Number of tokens per block: 32. [TensorRT-LLM][INFO] [MemUsageChange] Allocated 1.44 GiB for max tokens in paged KV cache (16448). [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134571520 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134571520 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 103809024 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 103809024 bytes [TensorRT-LLM][WARNING] [kv cache manager] storeContextBlocks: Can not find sequence for request 2048 [TensorRT-LLM][WARNING] [kv cache manager] storeContextBlocks: Can not find sequence for request 2049 [TensorRT-LLM][INFO] Max KV cache blocks per sequence: 4097 [window size=131073], tokens per block=32, primary blocks=16384, secondary blocks=0, max sequence length=131073 [TensorRT-LLM][INFO] Max KV cache blocks per sequence: 4097 [window size=131073], tokens per block=32, primary blocks=16384, secondary blocks=0, max sequence length=131073 [TensorRT-LLM][INFO] Number of tokens per block: 32. [TensorRT-LLM][INFO] [MemUsageChange] Allocated 46.00 GiB for max tokens in paged KV cache (524288). [TensorRT-LLM][INFO] Number of tokens per block: 32. [TensorRT-LLM][INFO] [MemUsageChange] Allocated 46.00 GiB for max tokens in paged KV cache (524288). [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134571520 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134571520 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 207618048 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 207618048 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 134571520 bytes to 207618048 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 134571520 bytes to 207618048 bytes