Papers
arxiv:2607.28227

Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents

Published on Jul 30
· Submitted by
Quyu Kong
on Jul 31
#2 Paper of the day
Authors:
,
,
,
,
,
,
,
,
,
,
,

Abstract

GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them toward real-world use, we envision agents that operate reliably on real devices, execute workflows across platforms, combine GUI interaction with CLI execution, complete long-horizon tasks, proactively initiate useful services, and autonomously improve their capabilities with minimal human effort. Guided by this vision, we present Qwen-UI-Agent, a real-world centric foundation GUI agent spanning mobile, computer-use, web, and DeepSearch environments. Qwen-UI-Agent combines diverse sandbox environments with a large-scale real-device mobile runtime. Its unified action space interleaves GUI operations with CLI execution and generates batched actions in a single model turn. An AutoResearch-style data flywheel uses agents to construct tasks and environments, diagnose failures, and plan subsequent iterations. Online RL supports training on trajectories exceeding 100 turns, with over 10,000 concurrent environments accelerating rollout. A lightweight harness layer supports proactive service initiation and stateful workflows across mobile and computer. Across a broad suite of evaluations, Qwen-UI-Agent sets state-of-the-art performance on mobile-use benchmarks while delivering competitive performance on computer- and browser-use tasks against frontier models, including Opus 4.8, Gemini 3.1 Pro, and GPT-5.6 Sol. On mobile use, it achieves 82.1% on MobileWorld, 92.2% on MobileWorld-Real, and 97.5% on AndroidDaily. On computer use, it achieves 79.5% on OSWorld-Verified and a 40.0% partial-progress score on OSWorld-v2. On browser use and GUI grounding, it achieves 73.6% on WebArena and 81.5% on ScreenSpot-Pro, respectively.

Community

Paper author Paper submitter
edited 1 day ago

We present Qwen-UI-Agent, a real-world centric foundation GUI agent unifying mobile, computer, browser, and DeepSearch scenarios in a single model. It sets a new SOTA on mobile use and stays competitive with Opus 4.8, Gemini 3.1 Pro, and GPT-5.6 Sol on computer and browser use. 🚀
🏠 Project Page: https://tongyi-mai.github.io/Qwen-UI-Agent/

image

🔥🔥🔥🔥🔥

把 AutoResearch 风格的数据飞轮和百步以上的 Online RL 结合起来,这个思路太赞了!

过去长序列任务很容易跑偏或者崩溃,通过主动服务发起(Proactive Service)和动态诊断,真正看到了全端 GUI Agent 从‘沙箱玩具’走向‘真实可用’的曙光,阿里通义团队给力!🔥🔥🔥

思路很赞

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

97.5% on AndroidDaily and 92.2% on MobileWorld-Real is insane. 🤯

Achieving SOTA on mobile while keeping competitive performance on computer and browser tasks shows the strength of unified action spaces. Can't wait to try integrating MAI-UI into our workflow!

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2607.28227
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2607.28227 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2607.28227 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2607.28227 in a Space README.md to link it from this page.

Collections including this paper 2