souging commited on
Commit
5998d5f
·
verified ·
1 Parent(s): d34e814

End of training

Browse files
Files changed (2) hide show
  1. README.md +21 -21
  2. adapter_model.bin +1 -1
README.md CHANGED
@@ -6,7 +6,7 @@ tags:
6
  - axolotl
7
  - generated_from_trainer
8
  model-index:
9
- - name: b1494398-57b9-4061-b394-da82c151d1ae
10
  results: []
11
  ---
12
 
@@ -19,15 +19,15 @@ should probably proofread and complete it, then remove this comment. -->
19
  axolotl version: `0.4.1`
20
  ```yaml
21
  adapter: lora
22
- base_model: Qwen/Qwen2.5-14B-Instruct
23
  bf16: auto
24
  dataset_prepared_path: null
25
  datasets:
26
  - data_files:
27
- - 8f359061ef840864_train_data.json
28
  ds_type: json
29
  format: custom
30
- path: /root/G.O.D-test/core/data/8f359061ef840864_train_data.json
31
  type:
32
  field_input: tools
33
  field_instruction: func_name
@@ -49,7 +49,7 @@ fsdp_config: null
49
  gradient_accumulation_steps: 4
50
  gradient_checkpointing: false
51
  group_by_length: false
52
- hub_model_id: souging/b1494398-57b9-4061-b394-da82c151d1ae
53
  hub_repo: null
54
  hub_strategy: checkpoint
55
  hub_token: null
@@ -58,16 +58,16 @@ load_in_4bit: false
58
  load_in_8bit: false
59
  local_rank: null
60
  logging_steps: 1
61
- lora_alpha: 64
62
  lora_dropout: 0.05
63
  lora_fan_in_fan_out: null
64
  lora_model_dir: null
65
- lora_r: 32
66
  lora_target_linear: true
67
  lr_scheduler: cosine
68
- max_steps: 100
69
- micro_batch_size: 1
70
- mlflow_experiment_name: /tmp/8f359061ef840864_train_data.json
71
  model_type: AutoModelForCausalLM
72
  num_epochs: 4
73
  optimizer: adamw_bnb_8bit
@@ -86,10 +86,10 @@ trust_remote_code: true
86
  val_set_size: 0.05
87
  wandb_entity: null
88
  wandb_mode: online
89
- wandb_name: 1d376839-b404-43d5-9ab3-103297cfbe18
90
  wandb_project: Gradients-On-Demand
91
  wandb_run: your_name
92
- wandb_runid: 1d376839-b404-43d5-9ab3-103297cfbe18
93
  warmup_steps: 100
94
  weight_decay: 0.01
95
  xformers_attention: null
@@ -98,11 +98,11 @@ xformers_attention: null
98
 
99
  </details><br>
100
 
101
- # b1494398-57b9-4061-b394-da82c151d1ae
102
 
103
- This model is a fine-tuned version of [Qwen/Qwen2.5-14B-Instruct](https://huggingface.co/Qwen/Qwen2.5-14B-Instruct) on the None dataset.
104
  It achieves the following results on the evaluation set:
105
- - Loss: 0.0005
106
 
107
  ## Model description
108
 
@@ -122,24 +122,24 @@ More information needed
122
 
123
  The following hyperparameters were used during training:
124
  - learning_rate: 0.0002
125
- - train_batch_size: 1
126
- - eval_batch_size: 1
127
  - seed: 42
128
  - distributed_type: multi-GPU
129
  - num_devices: 8
130
  - gradient_accumulation_steps: 4
131
- - total_train_batch_size: 32
132
- - total_eval_batch_size: 8
133
  - optimizer: Use OptimizerNames.ADAMW_BNB with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
134
  - lr_scheduler_type: cosine
135
  - lr_scheduler_warmup_steps: 100
136
- - training_steps: 100
137
 
138
  ### Training results
139
 
140
  | Training Loss | Epoch | Step | Validation Loss |
141
  |:-------------:|:------:|:----:|:---------------:|
142
- | 0.0 | 0.0837 | 100 | 0.0005 |
143
 
144
 
145
  ### Framework versions
 
6
  - axolotl
7
  - generated_from_trainer
8
  model-index:
9
+ - name: 73bda874-d4bf-4b2d-bdc6-18b3a9a8e0c4
10
  results: []
11
  ---
12
 
 
19
  axolotl version: `0.4.1`
20
  ```yaml
21
  adapter: lora
22
+ base_model: Qwen/Qwen2.5-7B-Instruct
23
  bf16: auto
24
  dataset_prepared_path: null
25
  datasets:
26
  - data_files:
27
+ - d185db34f9ed4537_train_data.json
28
  ds_type: json
29
  format: custom
30
+ path: /root/G.O.D-test/core/data/d185db34f9ed4537_train_data.json
31
  type:
32
  field_input: tools
33
  field_instruction: func_name
 
49
  gradient_accumulation_steps: 4
50
  gradient_checkpointing: false
51
  group_by_length: false
52
+ hub_model_id: souging/73bda874-d4bf-4b2d-bdc6-18b3a9a8e0c4
53
  hub_repo: null
54
  hub_strategy: checkpoint
55
  hub_token: null
 
58
  load_in_8bit: false
59
  local_rank: null
60
  logging_steps: 1
61
+ lora_alpha: 48
62
  lora_dropout: 0.05
63
  lora_fan_in_fan_out: null
64
  lora_model_dir: null
65
+ lora_r: 24
66
  lora_target_linear: true
67
  lr_scheduler: cosine
68
+ max_steps: 80
69
+ micro_batch_size: 2
70
+ mlflow_experiment_name: /tmp/d185db34f9ed4537_train_data.json
71
  model_type: AutoModelForCausalLM
72
  num_epochs: 4
73
  optimizer: adamw_bnb_8bit
 
86
  val_set_size: 0.05
87
  wandb_entity: null
88
  wandb_mode: online
89
+ wandb_name: 72ac60aa-9b01-44e4-9ac9-0307450f38fa
90
  wandb_project: Gradients-On-Demand
91
  wandb_run: your_name
92
+ wandb_runid: 72ac60aa-9b01-44e4-9ac9-0307450f38fa
93
  warmup_steps: 100
94
  weight_decay: 0.01
95
  xformers_attention: null
 
98
 
99
  </details><br>
100
 
101
+ # 73bda874-d4bf-4b2d-bdc6-18b3a9a8e0c4
102
 
103
+ This model is a fine-tuned version of [Qwen/Qwen2.5-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct) on the None dataset.
104
  It achieves the following results on the evaluation set:
105
+ - Loss: 0.0009
106
 
107
  ## Model description
108
 
 
122
 
123
  The following hyperparameters were used during training:
124
  - learning_rate: 0.0002
125
+ - train_batch_size: 2
126
+ - eval_batch_size: 2
127
  - seed: 42
128
  - distributed_type: multi-GPU
129
  - num_devices: 8
130
  - gradient_accumulation_steps: 4
131
+ - total_train_batch_size: 64
132
+ - total_eval_batch_size: 16
133
  - optimizer: Use OptimizerNames.ADAMW_BNB with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
134
  - lr_scheduler_type: cosine
135
  - lr_scheduler_warmup_steps: 100
136
+ - training_steps: 80
137
 
138
  ### Training results
139
 
140
  | Training Loss | Epoch | Step | Validation Loss |
141
  |:-------------:|:------:|:----:|:---------------:|
142
+ | 0.0108 | 0.1339 | 80 | 0.0009 |
143
 
144
 
145
  ### Framework versions
adapter_model.bin CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:62041656417109919d46c4854b59ea30564eec442f094b308939325d9342f9e7
3
  size 242362666
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:27963a21504e3c4c8c888b17040eb0f16b7fbc72dde7d93560cc401d6c2d9aec
3
  size 242362666