videodanchik's picture
Upload folder using huggingface_hub
0b95572 verified
Raw
History Blame Contribute Delete
12.6 kB
[TensorRT-LLM] TensorRT LLM version: 1.2.0rc8
max_position_embeddings: 8192
[01/15/2026-00:32:20] [TRT-LLM] [I] Using LLM with PyTorch backend
[01/15/2026-00:32:20] [TRT-LLM] [W] Using default gpus_per_node: 8
[01/15/2026-00:32:20] [TRT-LLM] [I] neither checkpoint_format nor checkpoint_loader were provided, checkpoint_format will be set to HF.
[TensorRT-LLM] TensorRT LLM version: 1.2.0rc8
[TensorRT-LLM] TensorRT LLM version: 1.2.0rc8
[TensorRT-LLM] TensorRT LLM version: 1.2.0rc8
[TensorRT-LLM] TensorRT LLM version: 1.2.0rc8
[TensorRT-LLM] TensorRT LLM version: 1.2.0rc8
[TensorRT-LLM] TensorRT LLM version: 1.2.0rc8
[TensorRT-LLM] TensorRT LLM version: 1.2.0rc8
[TensorRT-LLM] TensorRT LLM version: 1.2.0rc8
[TensorRT-LLM][INFO] Refreshed the MPI local session
[TensorRT-LLM][INFO] Refreshed the MPI local session
[TensorRT-LLM][INFO] Refreshed the MPI local session
[TensorRT-LLM][INFO] Refreshed the MPI local session
[TensorRT-LLM][INFO] Refreshed the MPI local session
[TensorRT-LLM][INFO] Refreshed the MPI local session
[TensorRT-LLM][INFO] Refreshed the MPI local session
[TensorRT-LLM][INFO] Refreshed the MPI local session
Model init total -- 153.34s
Model init total -- 153.28s
Model init total -- 153.49s
Model init total -- 153.71s
Model init total -- 153.61s
Model init total -- 153.41s
Model init total -- 153.51s
[TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=6399, secondary blocks=0, max sequence length=202753
[TensorRT-LLM][INFO] Number of tokens per block: 32.
[TensorRT-LLM][INFO] [MemUsageChange] Allocated 35.93 GiB for max tokens in paged KV cache (204768).
Model init total -- 154.18s
[TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=6399, secondary blocks=0, max sequence length=202753
[TensorRT-LLM][INFO] Number of tokens per block: 32.
[TensorRT-LLM][INFO] [MemUsageChange] Allocated 35.93 GiB for max tokens in paged KV cache (204768).
[TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=6399, secondary blocks=0, max sequence length=202753
[TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=6399, secondary blocks=0, max sequence length=202753
[TensorRT-LLM][INFO] Number of tokens per block: 32.
[TensorRT-LLM][INFO] [MemUsageChange] Allocated 35.93 GiB for max tokens in paged KV cache (204768).
[TensorRT-LLM][INFO] Number of tokens per block: 32.
[TensorRT-LLM][INFO] [MemUsageChange] Allocated 35.93 GiB for max tokens in paged KV cache (204768).
[TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=6399, secondary blocks=0, max sequence length=202753
[TensorRT-LLM][INFO] Number of tokens per block: 32.
[TensorRT-LLM][INFO] [MemUsageChange] Allocated 35.93 GiB for max tokens in paged KV cache (204768).
[TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=6399, secondary blocks=0, max sequence length=202753
[TensorRT-LLM][INFO] Number of tokens per block: 32.
[TensorRT-LLM][INFO] [MemUsageChange] Allocated 35.93 GiB for max tokens in paged KV cache (204768).
[TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=6399, secondary blocks=0, max sequence length=202753
[TensorRT-LLM][INFO] Number of tokens per block: 32.
[TensorRT-LLM][INFO] [MemUsageChange] Allocated 35.93 GiB for max tokens in paged KV cache (204768).
[TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=6399, secondary blocks=0, max sequence length=202753
[TensorRT-LLM][INFO] Number of tokens per block: 32.
[TensorRT-LLM][INFO] [MemUsageChange] Allocated 35.93 GiB for max tokens in paged KV cache (204768).
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes
[TensorRT-LLM][WARNING] [kv cache manager] storeContextBlocks: Can not find sequence for request 2050
[TensorRT-LLM][WARNING] [kv cache manager] storeContextBlocks: Can not find sequence for request 2048
[TensorRT-LLM][WARNING] [kv cache manager] storeContextBlocks: Can not find sequence for request 2049
[TensorRT-LLM][WARNING] [kv cache manager] storeContextBlocks: Can not find sequence for request 2051
[TensorRT-LLM][WARNING] [kv cache manager] storeContextBlocks: Can not find sequence for request 2052
[TensorRT-LLM][WARNING] [kv cache manager] storeContextBlocks: Can not find sequence for request 2054
[TensorRT-LLM][WARNING] [kv cache manager] storeContextBlocks: Can not find sequence for request 2053
[TensorRT-LLM][WARNING] [kv cache manager] storeContextBlocks: Can not find sequence for request 2055
[TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=15907, secondary blocks=0, max sequence length=202753
[TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=15907, secondary blocks=0, max sequence length=202753
[TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=15907, secondary blocks=0, max sequence length=202753
[TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=15907, secondary blocks=0, max sequence length=202753
[TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=15907, secondary blocks=0, max sequence length=202753
[TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=15907, secondary blocks=0, max sequence length=202753
[TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=15907, secondary blocks=0, max sequence length=202753
[TensorRT-LLM][INFO] Max KV cache blocks per sequence: 6337 [window size=202753], tokens per block=32, primary blocks=15907, secondary blocks=0, max sequence length=202753
[TensorRT-LLM][INFO] Number of tokens per block: 32.
[TensorRT-LLM][INFO] [MemUsageChange] Allocated 89.32 GiB for max tokens in paged KV cache (509024).
[TensorRT-LLM][INFO] Number of tokens per block: 32.
[TensorRT-LLM][INFO] [MemUsageChange] Allocated 89.32 GiB for max tokens in paged KV cache (509024).
[TensorRT-LLM][INFO] Number of tokens per block: 32.
[TensorRT-LLM][INFO] [MemUsageChange] Allocated 89.32 GiB for max tokens in paged KV cache (509024).
[TensorRT-LLM][INFO] Number of tokens per block: 32.
[TensorRT-LLM][INFO] [MemUsageChange] Allocated 89.32 GiB for max tokens in paged KV cache (509024).
[TensorRT-LLM][INFO] Number of tokens per block: 32.
[TensorRT-LLM][INFO] [MemUsageChange] Allocated 89.32 GiB for max tokens in paged KV cache (509024).
[TensorRT-LLM][INFO] Number of tokens per block: 32.
[TensorRT-LLM][INFO] [MemUsageChange] Allocated 89.32 GiB for max tokens in paged KV cache (509024).
[TensorRT-LLM][INFO] Number of tokens per block: 32.
[TensorRT-LLM][INFO] [MemUsageChange] Allocated 89.32 GiB for max tokens in paged KV cache (509024).
[TensorRT-LLM][INFO] Number of tokens per block: 32.
[TensorRT-LLM][INFO] [MemUsageChange] Allocated 89.32 GiB for max tokens in paged KV cache (509024).
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 134309376 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 0 bytes to 311427072 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 134309376 bytes to 311427072 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 134309376 bytes to 311427072 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 134309376 bytes to 311427072 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 134309376 bytes to 311427072 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 134309376 bytes to 311427072 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 134309376 bytes to 311427072 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 134309376 bytes to 311427072 bytes
[TensorRT-LLM][WARNING] Attention workspace size is not enough, increase the size from 134309376 bytes to 311427072 bytes
trt-llm (tokenizer=/root/work/huggingface/hub/models--zai-org--GLM-4.7/snapshots/475d85cda16beac79cde7f4cf4cae8d1260566f5,checkpoint_dir=/root/work/quantized_models//saved_models_475d85cda16beac79cde7f4cf4cae8d1260566f5_nvfp4_kv_fp8,max_gen_toks=4096), gen_kwargs: (None), limit: None, num_fewshot: None, batch_size: 64
|Tasks|Version| Filter |n-shot| Metric | |Value | |Stderr|
|-----|------:|----------------|-----:|-----------|---|-----:|---|-----:|
|gsm8k| 3|flexible-extract| 5|exact_match|↑ |0.9348|± |0.0068|
| | |strict-match | 5|exact_match|↑ |0.9325|± |0.0069|