File size: 1,924 Bytes
91dcfea
fd1bccc
 
91dcfea
 
 
 
fd1bccc
91dcfea
 
fd1bccc
91dcfea
fd1bccc
91dcfea
fd1bccc
 
 
 
 
 
 
 
 
 
 
 
 
91dcfea
 
fd1bccc
91dcfea
 
 
fd1bccc
 
 
 
91dcfea
fd1bccc
91dcfea
fd1bccc
 
91dcfea
fd1bccc
 
 
 
 
 
91dcfea
fd1bccc
 
 
 
 
 
 
 
 
91dcfea
fd1bccc
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
---
base_model:
- fancyfeast/llama-joycaption-beta-one-hf-llava
tags:
- captioning
- mlx
pipeline_tag: image-text-to-text
library_name: mlx
---

# Llama JoyCaption Beta One (MLX 8-bit)

MLX port of [fancyfeast/llama-joycaption-beta-one-hf-llava](https://huggingface.co/fancyfeast/llama-joycaption-beta-one-hf-llava), quantized to 8-bit for efficient inference on Apple Silicon.

[JoyCaption](https://github.com/fpgaminer/joycaption) is a free, open, and uncensored image captioning VLM built on Llama 3.1 8B and SigLIP2, designed for generating descriptive captions to train diffusion models.

## Model Details

| | |
|---|---|
| Architecture | LLaVA (SigLIP2 vision encoder + Llama 3.1 8B) |
| Quantization | 8-bit (`group_size=64`) |
| Vision encoder | google/siglip2-so400m-patch14-384 |
| Image resolution | 384x384 |
| Total size | ~9.1 GB |

## Usage with mlx-vlm

```bash
pip install mlx-vlm
```

```python
import mlx.core as mx
from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
from mlx_vlm.utils import load_config

MODEL = "n0kovo/llama-joycaption-beta-one-hf-llava-mlx-8Bit"

model, processor = load(MODEL)
config = load_config(MODEL)

prompt = apply_chat_template(
    processor,
    config,
    "Write a long descriptive caption for this image in a formal tone.",
    num_images=1,
)

output = generate(
    model,
    processor,
    prompt,
    image="image.jpg",
    max_tokens=512,
    temperature=0.6,
)
print(output)
```

## Conversion Notes

- Language model weights quantized to 8-bit via `mlx-lm`
- Vision encoder weights quantized to 8-bit where layer dimensions allow (`group_size=64`); 28 MLP layers with incompatible dimensions (4304, not divisible by 64) are kept in float16
- Projector weights quantized to 8-bit

## Credits

- Original model by [fancyfeast](https://huggingface.co/fancyfeast) — [JoyCaption GitHub](https://github.com/fpgaminer/joycaption)