File size: 3,541 Bytes
a8bf8c3
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
10478b0
 
a8bf8c3
10478b0
 
 
 
 
 
 
a8bf8c3
10478b0
 
 
 
 
 
 
 
 
4f51ce5
a8bf8c3
10478b0
 
 
4f51ce5
 
10478b0
 
 
4f51ce5
10478b0
 
 
4f51ce5
10478b0
 
 
 
 
 
4f51ce5
10478b0
 
 
 
a8bf8c3
10478b0
 
 
 
 
 
 
 
 
 
4f51ce5
10478b0
 
 
4f51ce5
10478b0
 
 
4f51ce5
10478b0
4f51ce5
10478b0
4f51ce5
10478b0
 
4f51ce5
10478b0
 
a8bf8c3
 
10478b0
 
 
 
 
 
 
a8bf8c3
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
---
tags:
- text-classification
- transformers
license: apache-2.0
datasets:
- glue
model-index:
- name: bert-base-uncased
  results:
  - task:
      type: text-classification
      name: Text Classification
    dataset:
      name: GLUE
      type: dataset
    metrics:
    - type: accuracy
      value: 90.5
language:
- ja
base_model:
- llm-jp/llm-jp-3-13b
---
## 実行手順
以下の手順に従うことで、Hugging Face上のモデル(llm-jp/llm-jp-3-13b + /sncffcns/llm-jp-3-13b-it-20241127_lora)を用いて入力データ(elyza-tasks-100-TV_0.jsonl)を推論し、その結果を{adapter_id}-outputs.jsonlというファイルに出力することができる。

## ライブラリのインストールを行う
  ```bash
  !pip install unsloth
  !pip uninstall unsloth -y && pip install --upgrade --no-cache-dir "unsloth[colab-new] @ git+https://github.com/unslothai/unsloth.git"
  !pip install -U torch
  !pip install -U peft
  ```

## 必要なライブラリの読み込みを行う
  ```python
  from unsloth import FastLanguageModel
  from peft import PeftModel
  import torch
  import json
  from tqdm import tqdm
  import re
  ```

# ベースとなるモデルと学習したLoRAのアダプタを設定する(Hugging FaceのIDを指定)。
  ```python
  model_id = "llm-jp/llm-jp-3-13b"
  adapter_id = "sncffcns/llm-jp-3-13b-it-20241127_lora"


  # Hugging Face Token を指定する
  from google.colab import userdata
  HF_TOKEN = userdata.get('HF_TOKEN_WRITE')

  # unslothのFastLanguageModelで元のモデルをロード。
  dtype = None # Noneにしておけば自動で設定
  load_in_4bit = True # 今回は13Bモデルを扱うためTrue

  model, tokenizer = FastLanguageModel.from_pretrained(
      model_name=model_id,
      dtype=dtype,
      load_in_4bit=load_in_4bit,
      trust_remote_code=True,
  )

# 元のモデルにLoRAのアダプタを統合する
  model = PeftModel.from_pretrained(model, adapter_id, token = HF_TOKEN)
  ```
## タスクとなるデータの読み込み。
# ./elyza-tasks-100-TV_0.jsonlというファイルからデータセットをロードする
  ```python
  datasets = []
  with open("./elyza-tasks-100-TV_0.jsonl", "r") as f:
      item = ""
      for line in f:
        line = line.strip()
        item += line
        if item.endswith("}"):
          datasets.append(json.loads(item))
          item = ""

  # モデルを用いてタスクの推論を行う
  # 推論するためにモデルのモードを変更する
  FastLanguageModel.for_inference(model)

  results = []
  for dt in tqdm(datasets):
    input = dt["input"]

    prompt = f"""### 指示\n{input}\n### 回答\n"""

    inputs = tokenizer([prompt], return_tensors = "pt").to(model.device)

    outputs = model.generate(**inputs, max_new_tokens = 512, use_cache = True, do_sample=False, repetition_penalty=1.2)
    prediction = tokenizer.decode(outputs[0], skip_special_tokens=True).split('\n### 回答')[-1]

    results.append({"task_id": dt["task_id"], "input": input, "output": prediction})
  ```
# 結果をjsonlで保存する
# adapter_idをベースにしたファイル名でJSONL形式の出力ファイルを保存する
  ```python
  json_file_id = re.sub(".*/", "", adapter_id)
  with open(f"/content/{json_file_id}_output.jsonl", 'w', encoding='utf-8') as f:
      for result in results:
          json.dump(result, f, ensure_ascii=False)
          f.write('\n')
  ```
# 以上の手順で、{adapter_id}-outputs.jsonlというファイル名で推論結果が作成される