Instructions to use unsloth/gemma-2-27b-it-bnb-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use unsloth/gemma-2-27b-it-bnb-4bit with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="unsloth/gemma-2-27b-it-bnb-4bit") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("unsloth/gemma-2-27b-it-bnb-4bit") model = AutoModelForCausalLM.from_pretrained("unsloth/gemma-2-27b-it-bnb-4bit", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use unsloth/gemma-2-27b-it-bnb-4bit with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "unsloth/gemma-2-27b-it-bnb-4bit" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "unsloth/gemma-2-27b-it-bnb-4bit", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/unsloth/gemma-2-27b-it-bnb-4bit
- SGLang
How to use unsloth/gemma-2-27b-it-bnb-4bit with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "unsloth/gemma-2-27b-it-bnb-4bit" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "unsloth/gemma-2-27b-it-bnb-4bit", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "unsloth/gemma-2-27b-it-bnb-4bit" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "unsloth/gemma-2-27b-it-bnb-4bit", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Unsloth Studio
How to use unsloth/gemma-2-27b-it-bnb-4bit with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for unsloth/gemma-2-27b-it-bnb-4bit to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for unsloth/gemma-2-27b-it-bnb-4bit to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for unsloth/gemma-2-27b-it-bnb-4bit to start chatting
Load model with FastModel
pip install unsloth from unsloth import FastModel model, tokenizer = FastModel.from_pretrained( model_name="unsloth/gemma-2-27b-it-bnb-4bit", max_seq_length=2048, ) - Docker Model Runner
How to use unsloth/gemma-2-27b-it-bnb-4bit with Docker Model Runner:
docker model run hf.co/unsloth/gemma-2-27b-it-bnb-4bit
How to create gguf fromm this
Hi there, how can irun this llm in ollama Server? Tried to convert it to gguf with llama.cpp without success. How can i use it? Thanks in advance
This is a special 4-bit quant for finetuning with unsloth. If you just want to run gemma-2-27b-it in ollama, you'd probably just do this and let ollama download it from their repository for you:
ollama run gemma2:27b
If you do in fact want to build your own .gguf file locally with llamacpp, use this one instead:
https://huggingface.co/unsloth/gemma-2-27b-it
That being said, you can also just get a gguf someone else has made
https://huggingface.co/bartowski/gemma-2-27b-it-GGUF
One final option is to have huggingface built you a gguff file:
https://huggingface.co/spaces/ggml-org/gguf-my-repo
Thanks a lot for your reply.
I´m heaving trouble runnig the normal gemma2-27b. it is realy slow... so i found this model. A read that is is faster than the normal on? Or i sonly the training faster with unsloth? I do not need extra training at the moment.. just want to run the 27b in higher speed with more token/s than now...
Thanks so much.
Thanks a lot for your reply.
I´m heaving trouble runnig the normal gemma2-27b. it is realy slow... so i found this model. A read that is is faster than the normal on? Or i sonly the training faster with unsloth? I do not need extra training at the moment.. just want to run the 27b in higher speed with more token/s than now...Thanks so much.
It's only faster because it is 4bit quantized which is unrelated to unsloth. GGUF cannot be in 4bit so the best option you have is to use Bartowski's upload.
We do make training and inference of models faster however but currently our inference only works with GPUs.
Ok, thanks for making things clear to me :) i'll give an other ready to use gguff with 4 bit quant a chance.
Ok, thanks for making things clear to me :) i'll give an other ready to use gguff with 4 bit quant a chance.
When you finetune a model with Unsloth remember you can also directly export it to GGUF using Unsloth!