Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
Abstract
GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them toward real-world use, we envision agents that operate reliably on real devices, execute workflows across platforms, combine GUI interaction with CLI execution, complete long-horizon tasks, proactively initiate useful services, and autonomously improve their capabilities with minimal human effort. Guided by this vision, we present Qwen-UI-Agent, a real-world centric foundation GUI agent spanning mobile, computer-use, web, and DeepSearch environments. Qwen-UI-Agent combines diverse sandbox environments with a large-scale real-device mobile runtime. Its unified action space interleaves GUI operations with CLI execution and generates batched actions in a single model turn. An AutoResearch-style data flywheel uses agents to construct tasks and environments, diagnose failures, and plan subsequent iterations. Online RL supports training on trajectories exceeding 100 turns, with over 10,000 concurrent environments accelerating rollout. A lightweight harness layer supports proactive service initiation and stateful workflows across mobile and computer. Across a broad suite of evaluations, Qwen-UI-Agent sets state-of-the-art performance on mobile-use benchmarks while delivering competitive performance on computer- and browser-use tasks against frontier models, including Opus 4.8, Gemini 3.1 Pro, and GPT-5.6 Sol. On mobile use, it achieves 82.1% on MobileWorld, 92.2% on MobileWorld-Real, and 97.5% on AndroidDaily. On computer use, it achieves 79.5% on OSWorld-Verified and a 40.0% partial-progress score on OSWorld-v2. On browser use and GUI grounding, it achieves 73.6% on WebArena and 81.5% on ScreenSpot-Pro, respectively.
Community
We present Qwen-UI-Agent, a real-world centric foundation GUI agent unifying mobile, computer, browser, and DeepSearch scenarios in a single model. It sets a new SOTA on mobile use and stays competitive with Opus 4.8, Gemini 3.1 Pro, and GPT-5.6 Sol on computer and browser use. 🚀
🏠 Project Page: https://tongyi-mai.github.io/Qwen-UI-Agent/
🔥🔥🔥🔥🔥
把 AutoResearch 风格的数据飞轮和百步以上的 Online RL 结合起来,这个思路太赞了!
过去长序列任务很容易跑偏或者崩溃,通过主动服务发起(Proactive Service)和动态诊断,真正看到了全端 GUI Agent 从‘沙箱玩具’走向‘真实可用’的曙光,阿里通义团队给力!🔥🔥🔥
思路很赞
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks (2026)
- Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields (2026)
- WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces (2026)
- Beyond Global Replanning: Hierarchical Recovery for Cross-Device Agent Systems (2026)
- A Task-State Representation for Long-Horizon Mobile GUI Agents (2026)
- PhoneHarness: Harnessing Phone-Use Agents through Mixed GUI, CLI, and Tool Actions (2026)
- MacAgentBench: Benchmarking AI Agents on Real-World macOS Desktop (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
97.5% on AndroidDaily and 92.2% on MobileWorld-Real is insane. 🤯
Achieving SOTA on mobile while keeping competitive performance on computer and browser tasks shows the strength of unified action spaces. Can't wait to try integrating MAI-UI into our workflow!
Get this paper in your agent:
hf papers read 2607.28227 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
