craffel's picture
download
raw
3.23 kB
Setting Data Parallel size to 16
Tensor parallelism has not been tested for a while, use at your own risk
WARNING 26-05-04 19:35:46.745271 - 0:00:00 - Signal handler installed.
WARNING 26-05-04 19:35:46.745577 - 0:00:00 - WARNING: Setting MKL_SERVICE_FORCE_INTEL to GNU
WARNING 26-05-04 19:35:46.745647 - 0:00:00 - WARNING: Setting MKL_NUM_THREADS to 1
WARNING 26-05-04 19:35:46.745706 - 0:00:00 - WARNING: Setting ENABLE_INTRA_NODE_COMM to 1
WARNING 26-05-04 19:35:46.745751 - 0:00:00 - WARNING: Setting TORCH_NCCL_AVOID_RECORD_STREAMS to 1
WARNING 26-05-04 19:35:46.745788 - 0:00:00 - WARNING: Setting NCCL_IB_TIMEOUT to 22
WARNING 26-05-04 19:35:46.745827 - 0:00:00 - WARNING: Setting NCCL_DEBUG to INFO
WARNING 26-05-04 19:35:46.745864 - 0:00:00 - WARNING: Setting TORCH_NCCL_ASYNC_ERROR_HANDLING to 1
WARNING 26-05-04 19:35:46.745897 - 0:00:00 - WARNING: Setting TRITON_CACHE_DIR to /scratch/tmpvi4elr6g
/fsx/craffel/miniconda3/envs/lingua_250401/lib/python3.11/site-packages/torch/autograd/graph.py:825: UserWarning: cuDNN SDPA backward got grad_output.strides() != output.strides(), attempting to materialize a grad_output with matching strides... (Triggered internally at ../aten/src/ATen/native/cudnn/MHA.cpp:674.)
return Variable._execution_engine.run_backward( # Calls into the C++ engine to run the backward pass
[rank12]:[W504 19:46:50.686722214 CPUAllocator.cpp:249] Memory block of unknown size was allocated before the profiling started, profiler results will not include the deallocation event
[rank12]: Traceback (most recent call last):
[rank12]: File "<frozen runpy>", line 198, in _run_module_as_main
[rank12]: File "<frozen runpy>", line 88, in _run_code
[rank12]: File "/fsx/craffel/lingua_logs/safe_v2_2nodes/code/apps/main/train.py", line 695, in <module>
[rank12]: main()
[rank12]: File "/fsx/craffel/lingua_logs/safe_v2_2nodes/code/apps/main/train.py", line 691, in main
[rank12]: train(cfg)
[rank12]: File "/fsx/craffel/lingua_logs/safe_v2_2nodes/code/apps/main/train.py", line 581, in train
[rank12]: from apps.main.eval import (
[rank12]: File "/fsx/craffel/lingua_logs/safe_v2_2nodes/code/apps/main/eval.py", line 14, in <module>
[rank12]: from lm_eval import simple_evaluate
[rank12]: File "<frozen importlib._bootstrap>", line 1229, in _handle_fromlist
[rank12]: File "/fsx/craffel/miniconda3/envs/lingua_250401/lib/python3.11/site-packages/lm_eval/__init__.py", line 23, in __getattr__
[rank12]: from .evaluator import simple_evaluate
[rank12]: File "/fsx/craffel/miniconda3/envs/lingua_250401/lib/python3.11/site-packages/lm_eval/evaluator.py", line 17, in <module>
[rank12]: from lm_eval.evaluator_utils import (
[rank12]: File "/fsx/craffel/miniconda3/envs/lingua_250401/lib/python3.11/site-packages/lm_eval/evaluator_utils.py", line 16, in <module>
[rank12]: from lm_eval.result_schema import EvalResults, _SampleCount, _TaskMetrics
[rank12]: File "/fsx/craffel/miniconda3/envs/lingua_250401/lib/python3.11/site-packages/lm_eval/result_schema.py", line 110, in <module>
[rank12]: class _TaskMetrics(TypedDict, Generic[T], extra_items=T):
[rank12]: TypeError: _TypedDictMeta.__new__() got an unexpected keyword argument 'extra_items'

Xet Storage Details

Size:
3.23 kB
·
Xet hash:
7f0720f6d527f917a04d841a20d83e60cf1f31a8131e8add78a0acc469418f89

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.