Research

Google releases Android Bench 2.0 to evaluate AI agents

Google has launched Android Bench 2.0, adding multi-day long-horizon tasks and continuous scoring to better measure how effectively AI models handle real-world Android development work.

InfoQ AI2 days agoResearch
Image: InfoQ AI

Google has launched Android Bench 2.0, significantly expanding its evaluation platform to measure how autonomous agents and AI models execute complex software engineering tasks. Moving beyond minor code tweaks, the framework introduces long-horizon tasks that mirror multi-day developer assignments, including app conversions, new feature builds, greenfield projects, and dependency upgrades. The update also replaces strict binary pass/fail metrics with a continuous completion rate based on functional output, visual accuracy, regression checks, and instruction compliance penalties.

Evaluation results from the benchmark reveal distinct performance boundaries for modern coding assistants. Google notes that models perform better when drafting fresh code rather than refactoring legacy repositories. Models handle deterministic refactoring well—such as migrating Java to Kotlin, switching Retrofit to Ktor, or adding a ViewModel layer—but falter on tasks requiring runtime validation like unhandled dependency injection graphs, breaking framework updates, or unreleased libraries. Even porting cross-platform applications to native Android remains an open challenge, with top-tier models topping out at an 80% completion rate.

The updated Android Bench 2.0 leaderboard tracks a range of recent systems, including Gemini 3.8 Flash, Gemini 3.7 Flash, OpenAI GPT-6, Anthropic Fable 5.1, Kimi K3, and Qwen 3.8 Max. At release, Claude Opus 5.5 claimed the top position with a 32% long-horizon task pass rate, followed closely by GPT 6 Astra at 28%. For software architects and mobile engineering teams, these nuanced benchmark scores provide clearer guidance on which development workflows can safely be automated today versus those that still demand heavy human oversight.

This is our own summary of reporting by InfoQ AI

More in Research