File size: 3,732 Bytes
8877377
 
 
 
 
 
c68103a
 
 
 
 
8877377
c68103a
8877377
c68103a
 
8877377
 
c68103a
 
 
 
8877377
 
 
 
 
 
 
 
2495573
8877377
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c68103a
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
---
license: mit
library_name: transformers
pipeline_tag: image-text-to-text
base_model: bytedance-research/UI-TARS-1.5-7B
tags:
- gui-agent
- computer-use
- vision-language
- reinforcement-learning
- osworld
datasets:
- osworld
language:
- en
- zh
---

<h1 style="display: flex; align-items: center; justify-content: center; gap: 10px;">
  <img src="https://leon-gittech.github.io/Verl_GUI/icon.png" alt="BEPA" style="height: 0.9em; width: auto;">
  <span>BEPA-7B-S2</span>
</h1>

<p align="center">
  <b>From Off-Policy to On-Policy: Enhancing GUI Agents via Bi-level Expert-to-Policy Assimilation</b>
</p>

<p align="center">
  <a href="https://leon-gittech.github.io/Verl_GUI/">🌐 Project Page</a> |
  <a href="https://arxiv.org/abs/2601.05787">πŸ“‘ arXiv Paper</a> |
  <a href="https://github.com/LEON-gittech/Verl_GUI.git">πŸ’» GitHub</a>
</p>

<p align="center">
  πŸ† <b>#1 Open-Source End-to-End Model on OSWorld (15 steps)</b>: Achieves <b>32.13%</b> success rate<br>
  πŸ“Š <b>Extreme Data Efficiency</b>: Matches GUI-OWL-7B performance using only <b>128 training tasks</b>
</p>

## Model Description

**BEPA-7B-S2** is a GUI agent model fine-tuned from [UI-TARS-1.5-7B](https://huggingface.co/bytedance-research/UI-TARS-1.5-7B) using the BEPA (Bi-Level Expert-to-Policy Assimilation) framework. This model achieves state-of-the-art performance among open-source end-to-end models on the OSWorld benchmark.

### Key Results

| Method | D<sub>expert_only</sub> | D<sub>train</sub> | D<sub>held_out</sub> | Overall (%) |
|--------|-------------------------|-------------------|----------------------|-------------|
| UITARS1.5-7B | 18.52 | 55.12 | 5.74 | 22.87 |
| GRPO | 11.11 | 58.02 | 5.32 | 23.60 |
| **BEPA (ours)** | **35.19** | **73.23** | **10.30** | **32.13** |

BEPA improves UI-TARS-1.5-7B from **22.87%** to **32.13%** on OSWorld-Verified (+9.26 points, +40.5% relative improvement).

## BEPA Framework

<p align="center">
  <img src="https://leon-gittech.github.io/Verl_GUI/stats/overview.png" alt="BEPA Overview" width="90%">
</p>

BEPA addresses two key challenges when using expert trajectories for training end-to-end GUI policies:

1. **Structural Mismatch:** Framework traces interleave multiple roles (planning, execution, grounding) that end-to-end policies cannot directly imitate.
2. **Distribution Gap:** Even after format conversion, trajectories remain far from the base-policy manifold.

### LEVEL-1: Self-Rolled Execution
Transforms alien expert traces into policy-compatible trajectories by abstracting expert trajectories into compact natural-language plans, then letting the base policy act in the environment with plan conditioning.

### LEVEL-2: Self-Aligned Assimilation
Dynamically maintains a per-task cache, injecting guided trajectories into GRPO updates only upon total on-policy failure. The cache is continuously refreshed with the policy's own successful executions.

## Citation

```bibtex
@misc{wang2026offpolicyonpolicyenhancinggui,
      title={From Off-Policy to On-Policy: Enhancing GUI Agents via Bi-level Expert-to-Policy Assimilation},
      author={Zezhou Wang and Ziyun Zhang and Xiaoyi Zhang and Zhuzhong Qian and Yan Lu},
      year={2026},
      eprint={2601.05787},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2601.05787},
}
```


## License

This model is released under the [MIT License](https://opensource.org/licenses/MIT).

## Acknowledgements

- [veRL](https://github.com/volcengine/verl) for the RL framework
- [vLLM](https://github.com/vllm-project/vllm) for fast inference
- [OSWorld](https://github.com/xlang-ai/OSWorld) for the benchmark
- [UI-TARS](https://github.com/bytedance/UI-TARS) for the base model