Instructions to use Alibaba-NLP/Tongyi-DeepResearch-30B-A3B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Alibaba-NLP/Tongyi-DeepResearch-30B-A3B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Alibaba-NLP/Tongyi-DeepResearch-30B-A3B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Alibaba-NLP/Tongyi-DeepResearch-30B-A3B") model = AutoModelForCausalLM.from_pretrained("Alibaba-NLP/Tongyi-DeepResearch-30B-A3B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Alibaba-NLP/Tongyi-DeepResearch-30B-A3B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Alibaba-NLP/Tongyi-DeepResearch-30B-A3B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Alibaba-NLP/Tongyi-DeepResearch-30B-A3B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Alibaba-NLP/Tongyi-DeepResearch-30B-A3B
- SGLang
How to use Alibaba-NLP/Tongyi-DeepResearch-30B-A3B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Alibaba-NLP/Tongyi-DeepResearch-30B-A3B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Alibaba-NLP/Tongyi-DeepResearch-30B-A3B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Alibaba-NLP/Tongyi-DeepResearch-30B-A3B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Alibaba-NLP/Tongyi-DeepResearch-30B-A3B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Alibaba-NLP/Tongyi-DeepResearch-30B-A3B with Docker Model Runner:
docker model run hf.co/Alibaba-NLP/Tongyi-DeepResearch-30B-A3B
Alibaba Hosted Model plz
Hi,
Since this project comes from Alibaba, any chance you could get them to host the Tongyi-DeepResearch-30B-A3B model on Alibaba Cloud?
This project's inference code is set up to be locally hosted with 8 (!) vLLM instances. That is a non-trivial hardware requirement.
I re-wrote the inference code to use an OpenAPI compatible endpoint, then set up Alibaba-NLP/Tongyi-DeepResearch-30B-A3B as a "Deployment" with the DeepInfra company.
It would be much easier if the model was just available from Alibaba. Perhaps Alibaba could add an "Experimental" section to their models, to make it clear that this may not yet (?) be a full production model.
Thanks for your consideration,
-Jeff
According to the update README.md in the project's Github, the model on OpenRouter is now supported with a few lines of code changes:
https://openrouter.ai/alibaba/tongyi-deepresearch-30b-a3b
I am running with some different code than the git repo, but I have found the OpenRouter model to have some issues compared to the DeepInfra Deployment I did. The latter, I believe, uses the Tongyi-DeepResearch-30B-A3B model config/template from Huggingface when it imports the model. I don't know how OpenRouter sets it up, but they don't appear to be working exactly the same. I haven't dug deep into the issue, but FYI, others may see the same.
The OpenRouter model is being directed to using Atlas Cloud.
Unfortunately, it is hitting 429 rate limits. I did very few queries with it (I am mostly using DeepInfra). The limits appear to be an issue with Atlas Cloud itself.
I used Atlas Cloud's API directly with an Atlas Cloud account, and hit the same "429" rate limit on the first query. So it seems Atlas Cloud can't handle serving this model or is having some issue.
OpenRouter API URL: https://openrouter.ai/api/v1
OpenRouter Model Name: alibaba/tongyi-deepresearch-30b-a3b
Atlas Cloud API URL:https://api.atlascloud.ai/v1
Atlast Cloud Model Name: Alibaba-NLP/Tongyi-DeepResearch-30B-A3B
Please add the model to Alibaba Cloud!