PhoneBuddy: Hybrid Training Framework Achieves 83.2% Success Rate on AndroidWorld

2026年6月27日

86

981

PhoneBuddy: Hybrid Training Framework Achieves 83.2% Success Rate on AndroidWorld

The evolution of large language models has shifted from answering questions to directly operating software interfaces. In the mobile domain, this manifests as AI agents that must complete real-world tasks: booking hotels, filling documents, searching for information in mini-apps. The critical difference lies in understanding that task success is not measured by whether a button is correctly identified, but whether the agent can complete the entire task given the device's current state—account lo

概述

PhoneBuddy addresses a fundamental challenge in mobile agent training: the tension between realism and scalability. Real app environments provide authentic feedback but are expensive, slow, and difficult to automate for evaluation. Simulated environments offer scale and automatic verification but risk learning behaviors that don't transfer to real devices. The research team proposes combining both approaches—using real apps for ground truth feedback while leveraging PhoneWorld, a reconstructed s

Training Architecture

The framework starts all models from the same Qwen3.5-4B backbone with shared action interfaces and SFT initialization using 950,758 action steps from both real and simulated trajectories. After supervised fine-tuning, models diverge into two reinforcement learning branches: Real-only RL and Real+Mock hybrid RL with a 50/50 rollout ratio. The real app environment uses rubric-based model judging—Gemini generates evaluation rubrics and Qwen3.5-122B-A10B scores trajectories against them. PhoneWorld

PhoneBuddy's value lies not in beautiful numbers but in clarifying the most awkward contradiction in mobile agent training: the more realistic the environment, the harder it is to scale; the more scalable the environment, the more it loses fidelity.

“Research Analysis”
🦞

JimoClaw — 桌面 AI Agent 工作台

让 AI 处理本地资料、操控浏览器,最终交付可直接使用的文档、表格与 PPT,而不只是一段回答。

下载桌面版

Performance Results

The progression from SFT to Real to Real+Mock shows consistent improvement on single-app and AndroidWorld tasks. On AndroidWorld specifically, success rates climb from 60.3% to 77.2% to 83.2%—demonstrating transferable gains beyond the paper's internal evaluation set. Single-app tasks reach 62.0%, outperforming Gemini 3.1 Pro (50.0%) and GPT-5.4 (50.0%) despite using a much smaller model. Average success rate across all categories reaches 54.8%, representing a 5.0 percentage point improvement ov

Cross-App Limitations and Future Directions

The most revealing finding is the persistent weakness in cross-app tasks, which show no improvement across training stages (22.0% → 20.0% → 18.0%). This highlights that single-app practice doesn't naturally extend to scenarios requiring information transfer between applications—copying content from one app to another, maintaining context across interfaces, or coordinating actions across different apps. Such capabilities would require specialized designs for state tracking, temporary storage, cro

🛡️

积墨 AI 安全隐患巡检系统

任务一键下达 · 隐患 AI 识别 · 整改全程留痕 · 报告一键生成。让安全巡检真正看得见、管得住、能闭环。

了解方案

如有侵权,请联系删除。

Related Articles

联系我们 试用咨询
小墨 AI