Instructions to use MaziyarPanahi/Ultra-Miqu-WizardLM-120B-v0.1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use MaziyarPanahi/Ultra-Miqu-WizardLM-120B-v0.1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="MaziyarPanahi/Ultra-Miqu-WizardLM-120B-v0.1") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("MaziyarPanahi/Ultra-Miqu-WizardLM-120B-v0.1") model = AutoModelForCausalLM.from_pretrained("MaziyarPanahi/Ultra-Miqu-WizardLM-120B-v0.1", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use MaziyarPanahi/Ultra-Miqu-WizardLM-120B-v0.1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "MaziyarPanahi/Ultra-Miqu-WizardLM-120B-v0.1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MaziyarPanahi/Ultra-Miqu-WizardLM-120B-v0.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/MaziyarPanahi/Ultra-Miqu-WizardLM-120B-v0.1
- SGLang
How to use MaziyarPanahi/Ultra-Miqu-WizardLM-120B-v0.1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "MaziyarPanahi/Ultra-Miqu-WizardLM-120B-v0.1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MaziyarPanahi/Ultra-Miqu-WizardLM-120B-v0.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "MaziyarPanahi/Ultra-Miqu-WizardLM-120B-v0.1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MaziyarPanahi/Ultra-Miqu-WizardLM-120B-v0.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use MaziyarPanahi/Ultra-Miqu-WizardLM-120B-v0.1 with Docker Model Runner:
docker model run hf.co/MaziyarPanahi/Ultra-Miqu-WizardLM-120B-v0.1
Examples
There are three killers in a room. Someone enters the room and kills one of them. Nobody leaves the room. How many killers are left in the room? Explain your reasoning step by step.
At the beginning, there were three killers in the room.
1. Someone entered the room, making it four people inside (including the new person).
2. The new person then killed one of the original three killers, so now there are two dead bodies in the room.
3. However, the riddle didn't say that the new person wasn't a killer themselves. In fact, they just proved to be one by killing one of the original three.
4. Therefore, after the murder, there are still three killers in the room: the new person who just committed a killing (making them a killer), and the remaining two from the original group of three.
Mark is faster than Mary , Mary is faster than Joe. Is Joe faster than Mark? Let's think step by step
1) We know that Mark is faster than Mary, so we can say that Mark's speed > Mary's speed
2) We also know that Mary is faster than Joe, so we can say that Mary's speed > Joe's speed
3) Based on the information above, we can conclude that Mark'sss speed > Mary' speed > Joe's speed.
4) Therefore, Joe cannot be faster than Mark as his speed is less than both Mark and Mary.
If I lay 2 wet shirts out in the sun flat on the ground to dry and it takes 5 hours until they are dry, how long would it take to dry 4 shirts of the same type that way simultanously?
It would still take approximately 5 hours. The number of shirts does not affect the drying time, assuming all other conditions remain constant (such as temperature, humidity, air flow, and amount of moisture).
- Downloads last month
- 5