craffel's picture
download
raw
3.23 kB
Setting Data Parallel size to 16
Tensor parallelism has not been tested for a while, use at your own risk
WARNING 26-05-04 19:35:46.750070 - 0:00:00 - Signal handler installed.
WARNING 26-05-04 19:35:46.750393 - 0:00:00 - WARNING: Setting MKL_SERVICE_FORCE_INTEL to GNU
WARNING 26-05-04 19:35:46.750463 - 0:00:00 - WARNING: Setting MKL_NUM_THREADS to 1
WARNING 26-05-04 19:35:46.750519 - 0:00:00 - WARNING: Setting ENABLE_INTRA_NODE_COMM to 1
WARNING 26-05-04 19:35:46.750569 - 0:00:00 - WARNING: Setting TORCH_NCCL_AVOID_RECORD_STREAMS to 1
WARNING 26-05-04 19:35:46.750610 - 0:00:00 - WARNING: Setting NCCL_IB_TIMEOUT to 22
WARNING 26-05-04 19:35:46.750653 - 0:00:00 - WARNING: Setting NCCL_DEBUG to INFO
WARNING 26-05-04 19:35:46.750689 - 0:00:00 - WARNING: Setting TORCH_NCCL_ASYNC_ERROR_HANDLING to 1
WARNING 26-05-04 19:35:46.750726 - 0:00:00 - WARNING: Setting TRITON_CACHE_DIR to /scratch/tmpav2skq6y
/fsx/craffel/miniconda3/envs/lingua_250401/lib/python3.11/site-packages/torch/autograd/graph.py:825: UserWarning: cuDNN SDPA backward got grad_output.strides() != output.strides(), attempting to materialize a grad_output with matching strides... (Triggered internally at ../aten/src/ATen/native/cudnn/MHA.cpp:674.)
return Variable._execution_engine.run_backward( # Calls into the C++ engine to run the backward pass
[rank14]:[W504 19:46:50.686905518 CPUAllocator.cpp:249] Memory block of unknown size was allocated before the profiling started, profiler results will not include the deallocation event
[rank14]: Traceback (most recent call last):
[rank14]: File "<frozen runpy>", line 198, in _run_module_as_main
[rank14]: File "<frozen runpy>", line 88, in _run_code
[rank14]: File "/fsx/craffel/lingua_logs/safe_v2_2nodes/code/apps/main/train.py", line 695, in <module>
[rank14]: main()
[rank14]: File "/fsx/craffel/lingua_logs/safe_v2_2nodes/code/apps/main/train.py", line 691, in main
[rank14]: train(cfg)
[rank14]: File "/fsx/craffel/lingua_logs/safe_v2_2nodes/code/apps/main/train.py", line 581, in train
[rank14]: from apps.main.eval import (
[rank14]: File "/fsx/craffel/lingua_logs/safe_v2_2nodes/code/apps/main/eval.py", line 14, in <module>
[rank14]: from lm_eval import simple_evaluate
[rank14]: File "<frozen importlib._bootstrap>", line 1229, in _handle_fromlist
[rank14]: File "/fsx/craffel/miniconda3/envs/lingua_250401/lib/python3.11/site-packages/lm_eval/__init__.py", line 23, in __getattr__
[rank14]: from .evaluator import simple_evaluate
[rank14]: File "/fsx/craffel/miniconda3/envs/lingua_250401/lib/python3.11/site-packages/lm_eval/evaluator.py", line 17, in <module>
[rank14]: from lm_eval.evaluator_utils import (
[rank14]: File "/fsx/craffel/miniconda3/envs/lingua_250401/lib/python3.11/site-packages/lm_eval/evaluator_utils.py", line 16, in <module>
[rank14]: from lm_eval.result_schema import EvalResults, _SampleCount, _TaskMetrics
[rank14]: File "/fsx/craffel/miniconda3/envs/lingua_250401/lib/python3.11/site-packages/lm_eval/result_schema.py", line 110, in <module>
[rank14]: class _TaskMetrics(TypedDict, Generic[T], extra_items=T):
[rank14]: TypeError: _TypedDictMeta.__new__() got an unexpected keyword argument 'extra_items'

Xet Storage Details

Size:
3.23 kB
·
Xet hash:
bf2a122d2078e5732827849e62aebfdb39a841e4e6710402601d618726cc6026

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.