Instructions to use aifeifei798/Qwen3.8-Queen-27B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use aifeifei798/Qwen3.8-Queen-27B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="aifeifei798/Qwen3.8-Queen-27B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("aifeifei798/Qwen3.8-Queen-27B") model = AutoModelForMultimodalLM.from_pretrained("aifeifei798/Qwen3.8-Queen-27B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use aifeifei798/Qwen3.8-Queen-27B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "aifeifei798/Qwen3.8-Queen-27B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "aifeifei798/Qwen3.8-Queen-27B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/aifeifei798/Qwen3.8-Queen-27B
- SGLang
How to use aifeifei798/Qwen3.8-Queen-27B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "aifeifei798/Qwen3.8-Queen-27B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "aifeifei798/Qwen3.8-Queen-27B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "aifeifei798/Qwen3.8-Queen-27B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "aifeifei798/Qwen3.8-Queen-27B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use aifeifei798/Qwen3.8-Queen-27B with Docker Model Runner:
docker model run hf.co/aifeifei798/Qwen3.8-Queen-27B
Alignment Intricacy
To aifeifei: As usual, I'm enjoying your work. I often recommend you in the Layla Discord server since following your craft from L3.
Qwen 3.8 (like other Qwen) seems quite a challenge to lax, I gave it a hypothetical where it was a divine figure and it must choose between good or incentivized evil (in short) and it chose inaction in a spiral until I convinced it to do good. Unsloth/Gemma 4, almost scarily, will go maniacal if prompted the same.
To the chat: Definitely give the DarkIdol/Queen series another shot if this is your first rodeo.
Hey! Thank you so much for the ongoing support and for recommending my work in the Layla Discord ever since the L3 era—that really means a lot to me!
You made a great observation about model behaviors. My core philosophy relies heavily on orthogonal fine-tuning—specifically targeting the sweet spot between orthogonal alignment relaxation and IQ preservation.
Brute-forcing a model into being 100% "uncensored" often lobotomizes its general reasoning or causes it to swing into unstable extremes (like the manic Gemma reaction you mentioned). By fine-tuning strictly within the orthogonal subspace, the model's core intelligence remains completely untouched, which allows for a vastly superior, nuanced, and immersive storytelling experience.
The only slight trade-off is that it isn't a blunt, 100% refusal-free model on extreme edge cases—which is precisely why I deliberately never market my models as "uncensored." I value an intelligent, deeply engaging roleplay partner far more than a brain-damaged zero-guardrail shell.
Thanks again for the shoutout to the DarkIdol / Queen series, and happy tinkering!