n0kovo commited on
Commit
fd1bccc
·
verified ·
1 Parent(s): cd87141

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +51 -16
README.md CHANGED
@@ -1,35 +1,70 @@
1
  ---
2
- base_model: fancyfeast/llama-joycaption-beta-one-hf-llava
 
3
  tags:
4
  - captioning
5
  - mlx
6
- - mlx-my-repo
7
  pipeline_tag: image-text-to-text
8
- library_name: transformers
9
  ---
10
 
11
- # n0kovo/llama-joycaption-beta-one-hf-llava-mlx-8Bit
12
 
13
- The Model [n0kovo/llama-joycaption-beta-one-hf-llava-mlx-8Bit](https://huggingface.co/n0kovo/llama-joycaption-beta-one-hf-llava-mlx-8Bit) was converted to MLX format from [fancyfeast/llama-joycaption-beta-one-hf-llava](https://huggingface.co/fancyfeast/llama-joycaption-beta-one-hf-llava) using mlx-lm version **0.28.3**.
14
 
15
- ## Use with mlx
 
 
 
 
 
 
 
 
 
 
 
 
16
 
17
  ```bash
18
- pip install mlx-lm
19
  ```
20
 
21
  ```python
22
- from mlx_lm import load, generate
 
 
 
23
 
24
- model, tokenizer = load("n0kovo/llama-joycaption-beta-one-hf-llava-mlx-8Bit")
25
 
26
- prompt="hello"
 
27
 
28
- if hasattr(tokenizer, "apply_chat_template") and tokenizer.chat_template is not None:
29
- messages = [{"role": "user", "content": prompt}]
30
- prompt = tokenizer.apply_chat_template(
31
- messages, tokenize=False, add_generation_prompt=True
32
- )
 
33
 
34
- response = generate(model, tokenizer, prompt=prompt, verbose=True)
 
 
 
 
 
 
 
 
35
  ```
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ base_model:
3
+ - fancyfeast/llama-joycaption-beta-one-hf-llava
4
  tags:
5
  - captioning
6
  - mlx
 
7
  pipeline_tag: image-text-to-text
8
+ library_name: mlx
9
  ---
10
 
11
+ # Llama JoyCaption Beta One (MLX 8-bit)
12
 
13
+ MLX port of [fancyfeast/llama-joycaption-beta-one-hf-llava](https://huggingface.co/fancyfeast/llama-joycaption-beta-one-hf-llava), quantized to 8-bit for efficient inference on Apple Silicon.
14
 
15
+ [JoyCaption](https://github.com/fpgaminer/joycaption) is a free, open, and uncensored image captioning VLM built on Llama 3.1 8B and SigLIP2, designed for generating descriptive captions to train diffusion models.
16
+
17
+ ## Model Details
18
+
19
+ | | |
20
+ |---|---|
21
+ | Architecture | LLaVA (SigLIP2 vision encoder + Llama 3.1 8B) |
22
+ | Quantization | 8-bit (`group_size=64`) |
23
+ | Vision encoder | google/siglip2-so400m-patch14-384 |
24
+ | Image resolution | 384x384 |
25
+ | Total size | ~9.1 GB |
26
+
27
+ ## Usage with mlx-vlm
28
 
29
  ```bash
30
+ pip install mlx-vlm
31
  ```
32
 
33
  ```python
34
+ import mlx.core as mx
35
+ from mlx_vlm import load, generate
36
+ from mlx_vlm.prompt_utils import apply_chat_template
37
+ from mlx_vlm.utils import load_config
38
 
39
+ MODEL = "n0kovo/llama-joycaption-beta-one-hf-llava-mlx-8Bit"
40
 
41
+ model, processor = load(MODEL)
42
+ config = load_config(MODEL)
43
 
44
+ prompt = apply_chat_template(
45
+ processor,
46
+ config,
47
+ "Write a long descriptive caption for this image in a formal tone.",
48
+ num_images=1,
49
+ )
50
 
51
+ output = generate(
52
+ model,
53
+ processor,
54
+ prompt,
55
+ image="image.jpg",
56
+ max_tokens=512,
57
+ temperature=0.6,
58
+ )
59
+ print(output)
60
  ```
61
+
62
+ ## Conversion Notes
63
+
64
+ - Language model weights quantized to 8-bit via `mlx-lm`
65
+ - Vision encoder weights quantized to 8-bit where layer dimensions allow (`group_size=64`); 28 MLP layers with incompatible dimensions (4304, not divisible by 64) are kept in float16
66
+ - Projector weights quantized to 8-bit
67
+
68
+ ## Credits
69
+
70
+ - Original model by [fancyfeast](https://huggingface.co/fancyfeast) — [JoyCaption GitHub](https://github.com/fpgaminer/joycaption)