[TensorRT-LLM] TensorRT LLM version: 1.2.0rc8 max_position_embeddings: 8192 [01/15/2026-00:32:20] [TRT-LLM] [I] Using LLM with PyTorch backend [01/15/2026-00:32:20] [TRT-LLM] [W] Using default gpus_per_node: 8 [01/15/2026-00:32:20] [TRT-LLM] [I] neither checkpoint_format nor checkpoint_loader were provided, checkpoint_format will be set to HF. [TensorRT-LLM] TensorRT LLM version: 1.2.0rc8 [TensorRT-LLM] TensorRT LLM version: 1.2.0rc8 [TensorRT-LLM] TensorRT LLM version: 1.2.0rc8 [TensorRT-LLM] TensorRT LLM version: 1.2.0rc8 [TensorRT-LLM] TensorRT LLM version: 1.2.0rc8 [TensorRT-LLM] TensorRT LLM version: 1.2.0rc8 [TensorRT-LLM] TensorRT LLM version: 1.2.0rc8 [TensorRT-LLM] TensorRT LLM version: 1.2.0rc8 [TensorRT-LLM][INFO] Refreshed the MPI local session [TensorRT-LLM][INFO] Refreshed the MPI local session [TensorRT-LLM][INFO] Refreshed the MPI local session [TensorRT-LLM][INFO] Refreshed the MPI local session [TensorRT-LLM][INFO] Refreshed the MPI local session [TensorRT-LLM][INFO] Refreshed the MPI local session [TensorRT-LLM][INFO] Refreshed the MPI local session [TensorRT-LLM][INFO] Refreshed the MPI local session Model init total -- 153.34s Model init total -- 153.28s Model init total -- 153.49s Model init total -- 153.71s Model init total -- 153.61s Model init total -- 153.41s Model init total -- 153.51s [TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=6399, secondary blocks=0, max sequence length=202753 [TensorRT-LLM][INFO] Number of tokens per block: 32. [TensorRT-LLM][INFO] [MemUsageChange] Allocated 35.93 GiB for max tokens in paged KV cache (204768). Model init total -- 154.18s [TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=6399, secondary blocks=0, max sequence length=202753 [TensorRT-LLM][INFO] Number of tokens per block: 32. [TensorRT-LLM][INFO] [MemUsageChange] Allocated 35.93 GiB for max tokens in paged KV cache (204768). [TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=6399, secondary blocks=0, max sequence length=202753 [TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=6399, secondary blocks=0, max sequence length=202753 [TensorRT-LLM][INFO] Number of tokens per block: 32. [TensorRT-LLM][INFO] [MemUsageChange] Allocated 35.93 GiB for max tokens in paged KV cache (204768). [TensorRT-LLM][INFO] Number of tokens per block: 32. [TensorRT-LLM][INFO] [MemUsageChange] Allocated 35.93 GiB for max tokens in paged KV cache (204768). [TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=6399, secondary blocks=0, max sequence length=202753 [TensorRT-LLM][INFO] Number of tokens per block: 32. [TensorRT-LLM][INFO] [MemUsageChange] Allocated 35.93 GiB for max tokens in paged KV cache (204768). [TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=6399, secondary blocks=0, max sequence length=202753 [TensorRT-LLM][INFO] Number of tokens per block: 32. [TensorRT-LLM][INFO] [MemUsageChange] Allocated 35.93 GiB for max tokens in paged KV cache (204768). [TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=6399, secondary blocks=0, max sequence length=202753 [TensorRT-LLM][INFO] Number of tokens per block: 32. [TensorRT-LLM][INFO] [MemUsageChange] Allocated 35.93 GiB for max tokens in paged KV cache (204768). [TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=6399, secondary blocks=0, max sequence length=202753 [TensorRT-LLM][INFO] Number of tokens per block: 32. [TensorRT-LLM][INFO] [MemUsageChange] Allocated 35.93 GiB for max tokens in paged KV cache (204768). [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes [TensorRT-LLM][WARNING] [kv cache manager] storeContextBlocks: Can not find sequence for request 2050 [TensorRT-LLM][WARNING] [kv cache manager] storeContextBlocks: Can not find sequence for request 2048 [TensorRT-LLM][WARNING] [kv cache manager] storeContextBlocks: Can not find sequence for request 2049 [TensorRT-LLM][WARNING] [kv cache manager] storeContextBlocks: Can not find sequence for request 2051 [TensorRT-LLM][WARNING] [kv cache manager] storeContextBlocks: Can not find sequence for request 2052 [TensorRT-LLM][WARNING] [kv cache manager] storeContextBlocks: Can not find sequence for request 2054 [TensorRT-LLM][WARNING] [kv cache manager] storeContextBlocks: Can not find sequence for request 2053 [TensorRT-LLM][WARNING] [kv cache manager] storeContextBlocks: Can not find sequence for request 2055 [TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=15907, secondary blocks=0, max sequence length=202753 [TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=15907, secondary blocks=0, max sequence length=202753 [TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=15907, secondary blocks=0, max sequence length=202753 [TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=15907, secondary blocks=0, max sequence length=202753 [TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=15907, secondary blocks=0, max sequence length=202753 [TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=15907, secondary blocks=0, max sequence length=202753 [TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=15907, secondary blocks=0, max sequence length=202753 [TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=15907, secondary blocks=0, max sequence length=202753 [TensorRT-LLM][INFO] Number of tokens per block: 32. [TensorRT-LLM][INFO] [MemUsageChange] Allocated 89.32 GiB for max tokens in paged KV cache (509024). [TensorRT-LLM][INFO] Number of tokens per block: 32. [TensorRT-LLM][INFO] [MemUsageChange] Allocated 89.32 GiB for max tokens in paged KV cache (509024). [TensorRT-LLM][INFO] Number of tokens per block: 32. [TensorRT-LLM][INFO] [MemUsageChange] Allocated 89.32 GiB for max tokens in paged KV cache (509024). [TensorRT-LLM][INFO] Number of tokens per block: 32. [TensorRT-LLM][INFO] [MemUsageChange] Allocated 89.32 GiB for max tokens in paged KV cache (509024). [TensorRT-LLM][INFO] Number of tokens per block: 32. [TensorRT-LLM][INFO] [MemUsageChange] Allocated 89.32 GiB for max tokens in paged KV cache (509024). [TensorRT-LLM][INFO] Number of tokens per block: 32. [TensorRT-LLM][INFO] [MemUsageChange] Allocated 89.32 GiB for max tokens in paged KV cache (509024). [TensorRT-LLM][INFO] Number of tokens per block: 32. [TensorRT-LLM][INFO] [MemUsageChange] Allocated 89.32 GiB for max tokens in paged KV cache (509024). [TensorRT-LLM][INFO] Number of tokens per block: 32. [TensorRT-LLM][INFO] [MemUsageChange] Allocated 89.32 GiB for max tokens in paged KV cache (509024). [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 134309376 bytes to 311427072 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 134309376 bytes to 311427072 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 134309376 bytes to 311427072 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 134309376 bytes to 311427072 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 134309376 bytes to 311427072 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 134309376 bytes to 311427072 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 134309376 bytes to 311427072 bytes [TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 134309376 bytes to 311427072 bytes trt-llm (tokenizer=/root/work/huggingface/hub/models--zai-org--GLM-4.7/snapshots/475d85cda16beac79cde7f4cf4cae8d1260566f5,checkpoint_dir=/root/work/quantized_models//saved_models_475d85cda16beac79cde7f4cf4cae8d1260566f5_nvfp4_kv_fp8,max_gen_toks=4096), gen_kwargs: (None), limit: None, num_fewshot: None, batch_size: 64 |Tasks|Version| Filter |n-shot| Metric | |Value | |Stderr| |-----|------:|----------------|-----:|-----------|---|-----:|---|-----:| |gsm8k| 3|flexible-extract| 5|exact_match|↑ |0.9348|± |0.0068| | | |strict-match | 5|exact_match|↑ |0.9325|± |0.0069|