Instructions to use Nexesenex/Llama_3.x_70b_Hexagon_Purple_V2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Nexesenex/Llama_3.x_70b_Hexagon_Purple_V2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Nexesenex/Llama_3.x_70b_Hexagon_Purple_V2") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Nexesenex/Llama_3.x_70b_Hexagon_Purple_V2") model = AutoModelForCausalLM.from_pretrained("Nexesenex/Llama_3.x_70b_Hexagon_Purple_V2", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Nexesenex/Llama_3.x_70b_Hexagon_Purple_V2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Nexesenex/Llama_3.x_70b_Hexagon_Purple_V2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Nexesenex/Llama_3.x_70b_Hexagon_Purple_V2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Nexesenex/Llama_3.x_70b_Hexagon_Purple_V2
- SGLang
How to use Nexesenex/Llama_3.x_70b_Hexagon_Purple_V2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Nexesenex/Llama_3.x_70b_Hexagon_Purple_V2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Nexesenex/Llama_3.x_70b_Hexagon_Purple_V2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Nexesenex/Llama_3.x_70b_Hexagon_Purple_V2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Nexesenex/Llama_3.x_70b_Hexagon_Purple_V2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Nexesenex/Llama_3.x_70b_Hexagon_Purple_V2 with Docker Model Runner:
docker model run hf.co/Nexesenex/Llama_3.x_70b_Hexagon_Purple_V2
Repetitiveness
Maybe its because I’m using a 4.0 BPW but I’ve noticed this model likes to repeat words. For example, when introducing technology to humanity in a roleplay scenario it likes to throw out “Game-Changer” and such words, without really going into detail or describing anything.
I can ban them in silly tavern, but I’m curious on if this is an issue with the dataset.
I have not tried the grey model, which is on my to-do list.
I really do enjoy this model, though.
Thanks for the feedback, FrenzyBiscuit, and sorry for the long answering delay.
This model is quite a loaded merge of various models.. and merges. So issues might occur related to one or several models. The most problematic ones imo are Tulu (included in Smartracks), Doppel (based on Hermes, which can be tricky due to its tokenizer), and maybe Tess-3 that I use as a perplexity dropper and is prone to repeat itself (common issue with low ppl Llama 3 bases). V3 is out, but I doubt it changes anything much.
Grey is more classic, using one of my old Smatricks base (which was stable), and my usual favorite finetunes to color it. It's a solid model.