craffel's picture
download
raw
1.51 kB
Setting Data Parallel size to 8
Tensor parallelism has not been tested for a while, use at your own risk
WARNING 26-04-21 17:22:55.376202 - 0:00:00 - Signal handler installed.
WARNING 26-04-21 17:22:55.376559 - 0:00:00 - WARNING: Setting MKL_SERVICE_FORCE_INTEL to GNU
WARNING 26-04-21 17:22:55.376629 - 0:00:00 - WARNING: Setting MKL_NUM_THREADS to 1
WARNING 26-04-21 17:22:55.376685 - 0:00:00 - WARNING: Setting ENABLE_INTRA_NODE_COMM to 1
WARNING 26-04-21 17:22:55.376730 - 0:00:00 - WARNING: Setting TORCH_NCCL_AVOID_RECORD_STREAMS to 1
WARNING 26-04-21 17:22:55.376768 - 0:00:00 - WARNING: Setting NCCL_IB_TIMEOUT to 22
WARNING 26-04-21 17:22:55.376806 - 0:00:00 - WARNING: Setting NCCL_DEBUG to INFO
WARNING 26-04-21 17:22:55.376838 - 0:00:00 - WARNING: Setting TORCH_NCCL_ASYNC_ERROR_HANDLING to 1
WARNING 26-04-21 17:22:55.376869 - 0:00:00 - WARNING: Setting TRITON_CACHE_DIR to /scratch/tmpr8f56ou2
/fsx/craffel/miniconda3/envs/lingua_250401/lib/python3.11/site-packages/torch/autograd/graph.py:825: UserWarning: cuDNN SDPA backward got grad_output.strides() != output.strides(), attempting to materialize a grad_output with matching strides... (Triggered internally at ../aten/src/ATen/native/cudnn/MHA.cpp:674.)
return Variable._execution_engine.run_backward( # Calls into the C++ engine to run the backward pass
[rank1]:[W421 17:25:36.964300419 CPUAllocator.cpp:249] Memory block of unknown size was allocated before the profiling started, profiler results will not include the deallocation event

Xet Storage Details

Size:
1.51 kB
·
Xet hash:
60768d89efdab90b9679b0208256a6ed9d6fe1226d33fd14cb43cd9d4c2ffb03

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.