Instructions to use McGill-NLP/A3-Qwen3.5-2B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use McGill-NLP/A3-Qwen3.5-2B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="McGill-NLP/A3-Qwen3.5-2B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("McGill-NLP/A3-Qwen3.5-2B") model = AutoModelForMultimodalLM.from_pretrained("McGill-NLP/A3-Qwen3.5-2B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use McGill-NLP/A3-Qwen3.5-2B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "McGill-NLP/A3-Qwen3.5-2B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "McGill-NLP/A3-Qwen3.5-2B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/McGill-NLP/A3-Qwen3.5-2B
- SGLang
How to use McGill-NLP/A3-Qwen3.5-2B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "McGill-NLP/A3-Qwen3.5-2B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "McGill-NLP/A3-Qwen3.5-2B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "McGill-NLP/A3-Qwen3.5-2B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "McGill-NLP/A3-Qwen3.5-2B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use McGill-NLP/A3-Qwen3.5-2B with Docker Model Runner:
docker model run hf.co/McGill-NLP/A3-Qwen3.5-2B
Improve model card: add metadata, library name, and citation (#1)
Browse files- Improve model card: add metadata, library name, and citation (8849a39d81182c9413fd0ffd969f61d9ede79a9a)
Co-authored-by: Niels Rogge <nielsr@users.noreply.huggingface.co>
README.md
CHANGED
|
@@ -1,27 +1,61 @@
|
|
| 1 |
---
|
|
|
|
| 2 |
language:
|
| 3 |
- en
|
|
|
|
|
|
|
|
|
|
| 4 |
tags:
|
| 5 |
- agents
|
| 6 |
- web
|
| 7 |
- sft
|
| 8 |
- qwen
|
| 9 |
-
base_model: Qwen/Qwen3.5-2B
|
| 10 |
-
pipeline_tag: text-generation
|
| 11 |
---
|
| 12 |
|
| 13 |
<div align="center">
|
| 14 |
|
| 15 |
# A3-Qwen3.5-2B
|
| 16 |
|
| 17 |
-
| [**💾 Code**](https://github.com/McGill-NLP/agent-as-annotators) | [**📄 Paper**](https://
|
| 18 |
| :--: | :--: | :--: |
|
| 19 |
| [**🤗 Dataset**](https://huggingface.co/datasets/McGill-NLP/A3-Synth) | [**🤖 Models**](https://huggingface.co/collections/McGill-NLP/a3-agent-as-annotators-69d854ab5b1993b10efc3fba) | [**📦 PyPI**](https://pypi.org/project/agent-as-annotators/) |
|
| 20 |
|
| 21 |
-
[**Structured Distillation of Web Agent Capabilities Enables Generalization**](https://
|
| 22 |
|
| 23 |
*Xing Han Lù, Siva Reddy*
|
| 24 |
|
| 25 |
</div>
|
| 26 |
|
| 27 |
-
A3-Qwen3.5-2B is a 2B web agent fine-tuned from [Qwen/Qwen3.5-2B](https://huggingface.co/Qwen/Qwen3.5-2B) on [A3-Synth](https://huggingface.co/datasets/McGill-NLP/A3-Synth)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
---
|
| 2 |
+
base_model: Qwen/Qwen3.5-2B
|
| 3 |
language:
|
| 4 |
- en
|
| 5 |
+
pipeline_tag: image-text-to-text
|
| 6 |
+
library_name: transformers
|
| 7 |
+
license: apache-2.0
|
| 8 |
tags:
|
| 9 |
- agents
|
| 10 |
- web
|
| 11 |
- sft
|
| 12 |
- qwen
|
|
|
|
|
|
|
| 13 |
---
|
| 14 |
|
| 15 |
<div align="center">
|
| 16 |
|
| 17 |
# A3-Qwen3.5-2B
|
| 18 |
|
| 19 |
+
| [**💾 Code**](https://github.com/McGill-NLP/agent-as-annotators) | [**📄 Paper**](https://huggingface.co/papers/2604.07776) | [**🌐 Website**](https://agent-as-annotators.github.io) |
|
| 20 |
| :--: | :--: | :--: |
|
| 21 |
| [**🤗 Dataset**](https://huggingface.co/datasets/McGill-NLP/A3-Synth) | [**🤖 Models**](https://huggingface.co/collections/McGill-NLP/a3-agent-as-annotators-69d854ab5b1993b10efc3fba) | [**📦 PyPI**](https://pypi.org/project/agent-as-annotators/) |
|
| 22 |
|
| 23 |
+
[**Structured Distillation of Web Agent Capabilities Enables Generalization**](https://huggingface.co/papers/2604.07776)
|
| 24 |
|
| 25 |
*Xing Han Lù, Siva Reddy*
|
| 26 |
|
| 27 |
</div>
|
| 28 |
|
| 29 |
+
**A3-Qwen3.5-2B** is a 2B multimodal web agent fine-tuned from [Qwen/Qwen3.5-2B](https://huggingface.co/Qwen/Qwen3.5-2B) using the **Agent-as-Annotators (A3)** framework. It is trained on [A3-Synth](https://huggingface.co/datasets/McGill-NLP/A3-Synth), a dataset of high-quality synthetic trajectories generated through a structured teacher-student distillation process.
|
| 30 |
+
|
| 31 |
+
## Model Description
|
| 32 |
+
A3-Qwen3.5-2B is designed to navigate complex web environments by processing visual screenshots and text. By decomposing the synthetic data generation process into three modular roles—Task Designer, Annotator, and Supervisor—the A3 framework allows small, locally deployable models to achieve competitive performance on benchmarks like WebArena, even surpassing some larger closed-source models.
|
| 33 |
+
|
| 34 |
+
## Quick Start: Evaluation
|
| 35 |
+
|
| 36 |
+
You can evaluate the model using the `agent-as-annotators` toolkit:
|
| 37 |
+
|
| 38 |
+
### 1. Serve the model with vLLM
|
| 39 |
+
```bash
|
| 40 |
+
vllm serve --model McGill-NLP/A3-Qwen3.5-2B
|
| 41 |
+
```
|
| 42 |
+
|
| 43 |
+
### 2. Run evaluation
|
| 44 |
+
```bash
|
| 45 |
+
a3-eval --benchmark webarena_test --model A3-qwen3.5-2b
|
| 46 |
+
```
|
| 47 |
+
|
| 48 |
+
## Citation
|
| 49 |
+
|
| 50 |
+
If you find this model useful, please cite our work:
|
| 51 |
+
|
| 52 |
+
```bibtex
|
| 53 |
+
@misc{lu2025structured,
|
| 54 |
+
title={Structured Distillation of Web Agent Capabilities Enables Generalization},
|
| 55 |
+
author={Xing Han Lù and Siva Reddy},
|
| 56 |
+
year={2025},
|
| 57 |
+
eprint={2604.07776},
|
| 58 |
+
archivePrefix={arXiv},
|
| 59 |
+
primaryClass={cs.LG}
|
| 60 |
+
}
|
| 61 |
+
```
|