Android Bench 2 Adds Support for Long-Horizon Tasks, Agentic Evaluation
It is aimed at Android engineers, AI researchers, and anyone building or evaluating coding agents against realistic mobile-development workloads.
By Dillip Chowdary • Oct 10, 2026 • Source: InfoQ
Google released Android Bench 2.0 on October 9, 2026, a significant overhaul of its benchmark framework for evaluating AI models and agents on Android development tasks. As InfoQ's report by Sergio De Simone details, the update introduces long-horizon tasks, agent-based evaluation, and a continuous scoring system designed to capture performance on the kind of complex, multi-step engineering work that the original binary pass/fail system routinely mislabeled as failure.
This article covers every material change in Android Bench 2.0 — what new task categories were added, how scoring was redesigned, which models currently lead the leaderboard, where AI still falls short, and what developers and benchmark maintainers should watch next. It is aimed at Android engineers, AI researchers, and anyone building or evaluating coding agents against realistic mobile-development workloads.
What Android Bench 2 shipped
Google's Android Bench framework launched a few months before this update with a focus on incremental changes to existing repositories, testing AI models against common Android development tasks in areas such as permissions, navigation, and connectivity. Android Bench 2.0 expands the scope considerably, introducing the first set of long-horizon tasks — work that Google describes as taking an engineer "multiple days or even a week to complete." Task categories in this new tier include upgrading dependencies, adding new features, building apps from scratch, and converting cross-platform apps to Android.
Alongside the new task categories, Google simultaneously introduced agentic evaluation, starting with agents from corresponding model providers. The updated Android Bench 2.0 dashboard now tracks recent models including Gemini 3.8 Flash, Gemini 3.7 Flash, OpenAI GPT-6, Anthropic Fable 5.1, Kimi K3, and Qwen 3.8 Max. At the time of publication, Claude Opus 5.5 led the leaderboard with a 32% long-horizon task pass rate, followed by GPT-6 Astra at 28%.
What changed for builders in Android Bench 2
The most consequential architectural shift in version 2.0 is the replacement of binary pass/fail scoring with a continuous completion-rate system. Under the original design, a task could be marked as fully failed because of a single failing edge-case assertion, even when an agent had correctly met dozens of other requirements. The new system calculates completion rates through a combination of factors — functionality, visual fidelity, and avoidance of regressions — and applies objective scoring penalties for deviations from evaluation instructions or structural constraints.
| Dimension | Android Bench 1.0 | Android Bench 2.0 |
|---|---|---|
| Task horizon | Incremental repo changes | Multi-day / multi-week LHTs |
| Scoring model | Binary pass/fail | Continuous completion rate |
| Evaluation type | Model-only | Model + agentic |
| Leaderboard models | Earlier generation | Gemini 3.8 Flash, GPT-6, Fable 5.1, Kimi K3, Qwen 3.8 Max |
| Top LHT pass rate | N/A | 32% (Claude Opus 5.5) |

The benchmark also surfaces which task types AI handles reliably. Google reports that AI "does a better job at writing new code rather than refactoring existing code," a finding attributed to the architectural complexity that refactoring and migrations demand. Well-established, deterministic transformations perform strongly even in larger codebases — examples include converting Java to Kotlin, swapping Retrofit for Ktor, and introducing a ViewModel layer.
How to install or upgrade Android Bench 2
Advertisement
Tech Pulse Daily
Get tomorrow's pulse first
Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.
Android Bench 2.0 is a Google-maintained open benchmark framework, not a standalone installable SDK, so access comes through the updated dashboard and the public repository rather than a package manager. Developers evaluating their own models or agents against Android Bench 2.0 should pull the latest version of the repository to access the new long-horizon task definitions and updated scoring harness.
```bash
# Clone or update the Android Bench repository
git clone https://github.com/google-deepmind/android_bench
# or, if already cloned
git pull origin mainhttps://androidbench.dev (check the repo README for the current URL)
```The agentic evaluation path requires configuring an agent from a supported model provider. Google's documentation specifies that agentic evaluation begins with agents from the corresponding model providers already listed on the dashboard. Teams integrating custom agents should review the evaluation instructions and structural constraints documented in the repository, since scoring penalties apply for deviations from those specifications.
Gotchas and compatibility in Android Bench 2
Despite the improvements in scoring granularity, Android Bench 2.0 still exposes clear limits in current models. Tasks requiring runtime validation — such as those involving missing dependency injection graphs — remain problematic, as do tasks involving breaking framework changes or knowledge gaps with unreleased libraries. Cross-platform-to-Android porting is highlighted as an "open challenge": even the best-in-class model achieves only 80% completion on that category under continuous scoring, a number that would have been a hard failure under the binary system.
Developers building agents for production Android work should also note that the benchmark's findings about refactoring and migrations are not just a scoring artifact — they reflect genuine model limitations with architectural context. Tasks that appear deterministic on the surface, such as swapping networking libraries, can involve implicit assumptions about threading models, error-handling conventions, and API surface compatibility that current models do not consistently resolve. The 2.0 release makes those gaps measurable in a way the original framework could not.
What to watch after Android Bench 2
The current leaderboard captures a moment in time — Claude Opus 5.5 at 32% and GPT-6 Astra at 28% on long-horizon tasks — and Google has signaled that the dashboard will be updated with new models as they become available. The 80% ceiling on cross-platform porting and the unresolved runtime-validation problems give benchmark maintainers and model developers specific targets to chase. Future iterations of Android Bench will likely add more LHT categories and expand agentic evaluation beyond the initial set of model-provider agents.
For the Android engineering community, the shift to continuous scoring is itself a model for how to evaluate agentic systems on realistic multi-day workloads, separate from any leaderboard position. The explicit penalty structure for deviating from evaluation instructions or structural constraints introduces reproducibility requirements that binary benchmarks have historically obscured. Whether that design carries into other domain-specific benchmarks — iOS, web, embedded — is the open question the 2.0 release implicitly poses.
Developer Action Items
- ☐ Diff the official changelog for OpenAI / Anthropic / Claude 2.0 before you bump — APIs, defaults, and removed flags only.
- ☐ Install through the vendor's documented channel in staging; keep a one-command rollback and time-box the canary.
- ☐ Grep your repo for old flag names, lockfile pins, and plugin versions that the notes mark as breaking.
- ☐ Prefer the first patch cut over the day-zero tag unless you have a reason to be on the leading edge.
- ☐ If InfoQ did not name a region, plan, or SKU, screenshot the official availability line before you promise it to users.
Android Bench 2 FAQ
What is Android Bench 2.0?
Android Bench 2.0 is Google's updated benchmark framework for evaluating AI models and agents on Android development tasks, introducing long-horizon tasks, agentic evaluation, and continuous scoring as of October 9, 2026.
Which model is leading the Android Bench 2.0 leaderboard?
Claude Opus 5.5 leads with a 32% long-horizon task pass rate, followed by GPT-6 Astra at 28%, as of the October 9 publication date.
What are long-horizon tasks in Android Bench 2.0?
Long-horizon tasks are complex multi-step engineering jobs that Google estimates a human engineer would take multiple days or even a week to complete, including upgrading dependencies, building apps from scratch, and porting cross-platform apps to Android.
How is Android Bench 2.0 scoring different from version 1.0?
Version 1.0 used binary pass/fail scoring, where one failing edge case could mark an entire task as failed. Version 2.0 uses a continuous completion rate based on functionality, visual fidelity, and regression avoidance, with penalties for structural constraint violations.
What tasks do current AI models still struggle with in Android Bench 2.0?
Models struggle with tasks requiring runtime validation such as missing dependency injection graphs, tasks involving breaking framework changes, knowledge gaps with unreleased libraries, and cross-platform-to-Android porting, where the best model reaches only 80% completion.
Sources
Author
Dillip Chowdary
Writes Tech Bytes coverage of AI, engineering, and the tools that actually ship. Editor of Tech Pulse Daily.
Related on Tech Bytes
9to5Mac Daily: October 9, 2026 – Apple’s ‘Welcome home’ launch announced, more
Read →
Apple acqui-hires AI startup founded by former NotebookLM developers
Read →
The maker of non-text AI model Jev valued at $7.5B just weeks after launch
Read →
In Other News: AI Used in Korean Bank Breaches, Poem-Guided Botnet, Empire Admin Gets 40…
Read →
Today's Tech Pulse briefing
Full briefing →
Advertisement