Instructions to use mlabonne/gemma-3-12b-it-abliterated with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mlabonne/gemma-3-12b-it-abliterated with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="mlabonne/gemma-3-12b-it-abliterated") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("mlabonne/gemma-3-12b-it-abliterated") model = AutoModelForMultimodalLM.from_pretrained("mlabonne/gemma-3-12b-it-abliterated", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use mlabonne/gemma-3-12b-it-abliterated with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "mlabonne/gemma-3-12b-it-abliterated" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mlabonne/gemma-3-12b-it-abliterated", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/mlabonne/gemma-3-12b-it-abliterated
- SGLang
How to use mlabonne/gemma-3-12b-it-abliterated with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "mlabonne/gemma-3-12b-it-abliterated" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mlabonne/gemma-3-12b-it-abliterated", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "mlabonne/gemma-3-12b-it-abliterated" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mlabonne/gemma-3-12b-it-abliterated", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use mlabonne/gemma-3-12b-it-abliterated with Docker Model Runner:
docker model run hf.co/mlabonne/gemma-3-12b-it-abliterated
Missing files?
I noticed that the official model has multiple files missing from this repo, and when trying to create a Quant of this model without them it throws an error about tokenizer.model being missing.
Now that I added the official models that were missing, I'm able to quantize it, but I'm wondering if these official files might mess up your abliteration or not.
Thanks for your feedback. Is tokenizer.model the only missing file for you? The abliteration only modifies the weights so changing the tokenizer files shouldn't affect it.
Hi, the other missing files are: added_tokens.json, chat_template.json, generation_config.json, preprocessor_config.json, and processor_config.json.
As for the abliteration and tokenizer: I'm not sure what's up, but I did make a Q8_0 GGUF, and it gave me the same scolding responses for inappropriate questions as it would with the base model. I also made sure to use the suggested model settings.
Thanks! They shouldn't be required (probably more related to the quantization code) but I'll add them for convenience. I'll re-run the model to double-check that it works. Have you tried the 4B or the 27B by any chance?
I wanted to try 12b, many solutions to launch it in textgen-webui but to no avail , any guide on how to load it properly there?
I use min_p preset, and change top_p to 0.95, and top_k to 64.
For model loading i use 'load-in-4bit'
The model loads, runs, answers; however it seems to have a problem 'stopping' (TextGen will show it generating with no ending most times, but clicking stop works).
If anyone has an issue to the 'not stopping' bit that would be nice; i'm wondering if it's a stop token i'm missing.