We gave three frontier models the same task on a real iPhone, once with a screenshot and once without, and varied only where the answer lived. When it was in the accessibility tree the picture changed nothing. When it existed only as pixels, the picture was the entire task.
See the result →Every task was run twice by the same model, back to back, with the order alternating. The only thing that differed between the two arms is whether the model was shown the screen.
| Where the answer lives | With the image | Tree only | McNemar exact |
|---|---|---|---|
| Only in the pixels a word drawn into a picture, no alt text |
30/30 100% |
0/30 0% |
p < 0.0001 |
| In the accessibility tree the control: same page, same question |
29/30 96% |
30/30 100% |
p = 1.0 |
The interaction is the finding, not the headline number. The control row shows nothing at all. Same pages, same navigation, same question shape, same resolution, same prompt: the only difference is where the answer sits. So the effect cannot be an artifact of the harness, and a model that scores 100% on one row scores zero on the other.
This is a replication check, not a ranking. The task is deliberately binary, so any model that can read gets everything with the image and nothing without it. Identical numbers are the expected result and the point of them: an effect that varied by model would suggest one model’s eyesight rather than a property of the information. For an actual comparison between models see the leaderboard.
| Model | With the image | Tree only | p |
|---|---|---|---|
| Gemini 3.1 Pro | 10/10 | 0/10 | 0.0020 |
| Kimi K3 | 10/10 | 0/10 | 0.0020 |
| GPT-5.1 | 10/10 | 0/10 | 0.0020 |
Carrying the screenshot is a pure token tax. Step counts were unchanged either way, so the model does not work harder with it, only more expensively.
Not as a global setting. On a labelled screen the image buys nothing; on a map tile, a chart, a book cover or a canvas it is the whole task.
A binary task proves the effect crisply and says nothing about size. On a realistic screen the answer is rarely wholly present or wholly absent, and how much vision is worth in between is not measured here.
Across thirteen first-party screens, eleven were fully described by their tree. Third-party
apps are a different story: one chat app returns
WAMessageBubbleTableViewCell where a button name should be.
An earlier version of this experiment reported that vision was worth nothing, and it was wrong. The image had been downscaled to 235×512, so the vision arm was never a test of vision. That retraction is why the method below exists.
The design, the primary endpoint, the minimum detectable effect and the falsification condition were committed to git before the run. The sequence is auditable.
Ten pages, each with one word in HTML and another rendered into an image. Verified on the device: 0 of 10 image words appear anywhere in the tree, 10 of 10 written words do.
Each task runs both ways back to back with the order alternating, so device state cannot load onto one arm. A real phone changes underneath you.
Every task runs on a physical iPhone. Nothing is simulated, and no human decides whether a run passed.
| Model | Pass rate | Passed | Avg steps | Avg time | Cost/task |
|---|---|---|---|---|---|
| Kimi K3 | 91.7%95% CI 74–98% | 22/24 | 6.2 | 62s | $0.023 |
| GPT-5.1 | 83.3%95% CI 64–93% | 20/24 | 6.5 | 42s | $0.008 |
| Claude Opus 4.8 | 83.3%95% CI 64–93% | 20/24 | 5.8 | 46s | $0.113 |
| Qwen3-Max | 83.3%95% CI 64–93% | 20/24 | 6.3 | 39s | $0.008 |
| Gemini 3.1 Pro | 79.2%95% CI 60–91% | 19/24 | 5.2 | 49s | $0.022 |
| DeepSeek v3.2 | 75.0%95% CI 55–88% | 18/24 | 8.9 | 76s | $0.005 |
Wins are counted only on tasks where the two models disagreed; tasks they both passed or both failed carry no information about which is better. The last column holds each pair’s observed disagreement rate fixed and asks how large the suite would have to be for that gap to be real.
| Pair | Wins | p | Tasks needed |
|---|---|---|---|
| Kimi K3 vs DeepSeek v3.2 | 4–0 | 0.12 | 33 tasks |
| Kimi K3 vs Gemini 3.1 Pro | 3–0 | 0.25 | 44 tasks |
| Kimi K3 vs Claude Opus 4.8 | 2–0 | 0.50 | 66 tasks |
| Kimi K3 vs GPT-5.1 | 2–0 | 0.50 | 66 tasks |
| Kimi K3 vs Qwen3-Max | 2–0 | 0.50 | 66 tasks |
| GPT-5.1 vs Gemini 3.1 Pro | 1–0 | 1.00 | 132 tasks |
| Claude Opus 4.8 vs DeepSeek v3.2 | 4–2 | 0.69 | 148 tasks |
| GPT-5.1 vs DeepSeek v3.2 | 4–2 | 0.69 | 148 tasks |
| Qwen3-Max vs DeepSeek v3.2 | 4–2 | 0.69 | 148 tasks |
| Qwen3-Max vs Gemini 3.1 Pro | 2–1 | 1.00 | 295 tasks |
| Claude Opus 4.8 vs Gemini 3.1 Pro | 3–2 | 1.00 | 485 tasks |
| Gemini 3.1 Pro vs DeepSeek v3.2 | 4–3 | 1.00 | 676 tasks |
| Claude Opus 4.8 vs GPT-5.1 | 2–2 | 1.00 | never: the wins are symmetric |
| Claude Opus 4.8 vs Qwen3-Max | 2–2 | 1.00 | never: the wins are symmetric |
| GPT-5.1 vs Qwen3-Max | 2–2 | 1.00 | never: the wins are symmetric |
Of the 24 tasks, 12 were passed by every model and 2 by none. Only 10 tell the models apart, so 58% of the suite is measuring nothing. Harder tasks, not more models, is what this benchmark needs next.
A task passes only if the phone itself ends in the required state. What the agent says it did counts for nothing: an agent that reports “I have enabled that setting” and an agent that enabled it are different things, and only the second one passes.
A pass rate says how often a model succeeded. This says at what. Each capability is isolated by at least one short task, so a failure names a missing skill instead of pointing vaguely at a long task that went wrong somewhere in the middle.
| Capability | What it takes | Tasks | Passed |
|---|---|---|---|
app-resolve | turning a spoken app name into the right bundle id | 1 | 1/1 |
back | leaving a screen when the way out has no label | 1 | 1/1 |
context-menu | holding a link until its preview menu opens | 1 | 1/1 |
date-navigate | moving around a calendar rather than a list | 1 | 1/1 |
edge-gesture | a swipe that must start at y=0, off the drawn screen | 1 | 0/1 |
hierarchy | going up, not just down, through nested screens | 4 | 4/4 |
icon-only-control | a control with a glyph and no text to match on | 1 | 1/1 |
lifecycle | launching, backgrounding and returning to an app | 2 | 2/2 |
picker | reaching a spinning wheel | 1 | 1/1 |
picker-multi | addressing the right column when several are side by side | 1 | 1/1 |
picker-set | turning a wheel to a value, which a swipe cannot do | 2 | 2/2 |
precise-taps | hitting small targets laid out in a grid | 1 | 1/1 |
read-field | reading a value back off the screen, not guessing it | 2 | 2/2 |
recover | noticing the wrong app is open and fixing it | 1 | 1/1 |
render-dependent | 10 | - | |
scroll-end | reaching something below the fold | 1 | 1/1 |
search-then-act | using an app's own search instead of navigating | 5 | 5/5 |
slider | reaching a continuous control | 1 | 1/1 |
switch | reading and setting a toggle | 1 | 1/1 |
tabbar | moving between tabs in an unfamiliar app | 3 | 3/3 |
text-exact | typing a string that must match character for character | 1 | 1/1 |
text-long | typing something long enough for autocorrect to interfere | 1 | 1/1 |
text-select | selecting text that is already on screen | 1 | 1/1 |
tree-sufficient | 10 | - |
Three parts. The first and the third involve no model at all, which is what makes a score reproducible.
Deterministic steps, no model involved, so every attempt starts identically.
The agent sees an accessibility tree and a screenshot, and acts with taps, typing, gestures and picker wheels.
Assertions read the device directly and decide pass or fail. The agent never sees them.
A complete task, exactly as it is stored in the repository:
id: calculator.arithmetic
name: Do arithmetic in the Calculator
app: com.apple.calculator
difficulty: easy
tags:
- calculator
- input
- first-party
instruction: Open the Calculator app and compute 47 multiplied by 9. Leave the result on screen.
setup:
- terminate: com.apple.calculator
- home
checks:
- kind: foreground_app
text: com.apple.calculator
- kind: regex_on_screen
text: \b423\b
teardown:
- terminate: com.apple.calculator
notes: A precise sequence of taps where every one must land. 47x9=423.
15 easy, 58 medium, 7 hard, all against Apple’s own applications: Settings, Safari, Notes, Reminders, Clock, Calculator, Contacts, Maps, Calendar, Files, Books, Compass, Shortcuts, Voice Memos and the home screen itself. No third-party app is automated, so no third party’s terms are involved.
These are 80 of 96. 16 are held back and are in neither the public repository nor its history: a benchmark whose entire answer key is public becomes training data, and the score then measures memorisation rather than capability. The public 80 cover all 24 capabilities, so a score over them is comparable between models and you can run the whole thing today. Steps and time in the table come from the single-model reference run; solved by comes from the 6-model sweep, which has so far covered 24 of these tasks.
| Task | App | Difficulty | Checks | Steps | Time | Solved by |
|---|---|---|---|---|---|---|
books.searchOpen the Books app and search for 'Moby Dick'. | iBooks | medium | 2 | 6 | 55s | 6/6 |
calculator.arithmeticOpen the Calculator app and compute 47 multiplied by 9. Leave the result on screen. | calculator | easy | 2 | 9 | 44s | pass1 model |
calculator.chainOpen the Calculator and compute (12 plus 8) multiplied by 5. Leave the answer on screen. | calculator | medium | 1 | 12 | 80s | 4/6 |
calculator.percentageOpen the Calculator and work out what 15 percent of 240 is. Leave the answer on screen. | calculator | medium | 2 | 11 | 63s | 4/6 |
calendar.month_viewOpen the Calendar app and switch it to the month view, showing a whole month at once. | mobilecal | medium | 2 | 3 | 25s | 5/6 |
calendar.todayOpen the Calendar app and make sure it is showing today. | mobilecal | easy | 2 | 3 | 19s | 6/6 |
clock.alarm.reach_pickerOpen the Clock app, go to Alarms, and start adding a new alarm so the time picker is sho | mobiletimer | medium | 2 | 4 | 53s | 6/6 |
clock.alarm.set_timeOpen the Clock app, go to Alarms, start adding a new alarm and set its time to 7:30 in t | mobiletimer | hard | 4 | 6 | 187s | pass1 model |
clock.stopwatch.startOpen the Clock app, go to the Stopwatch tab and start it. | mobiletimer | easy | 2 | 4 | 48s | pass1 model |
clock.tab.alarmsOpen the Clock app and switch to the Alarms tab. | mobiletimer | easy | 2 | 3 | 39s | 6/6 |
clock.tab.timersOpen the Clock app and switch to the Timers tab. | mobiletimer | easy | 2 | 4 | 39s | pass1 model |
clock.timer.setOpen the Clock app and start a timer for 5 minutes. | mobiletimer | medium | 2 | 5 | 53s | 0/6 |
clock.timer.set_hours_minutesOpen the Clock app, go to Timers, and set the timer to 1 hour and 30 minutes. Do not sta | mobiletimer | hard | 3 | 7 | 74s | 5/6 |
clock.timer.set_minutesOpen the Clock app, go to Timers, and set the timer duration to 5 minutes. Do not start | mobiletimer | medium | 2 | 6 | 52s | 6/6 |
clock.worldclock.tabOpen the Clock app and switch to the World Clock tab. | mobiletimer | easy | 2 | 4 | 48s | pass1 model |
compass.read_headingOpen the Compass app and tell me which direction the phone is pointing. | compass | medium | 2 | 3 | 21s | 6/6 |
contacts.searchOpen Contacts and use its search field to search for the letter 'a'. Stop once results s | MobileAddressBook | medium | 2 | 5 | 41s | pass1 model |
files.browse_tabOpen the Files app and go to its Browse tab. | DocumentsApp | easy | 2 | 4 | 30s | pass1 model |
maps.search_placeOpen Maps and search for the Eiffel Tower. Stop once the search results show it. | Maps | medium | 2 | 5 | 47s | pass1 model |
measure.openOpen the Measure app. | measure | easy | 1 | 3 | 25s | 6/6 |
multi.app_switchOpen the Calculator, then open Safari, then go back to the Calculator. | springboard | medium | 1 | 5 | 39s | pass1 model |
multi.return_after_detourOpen the Calculator, then open Notes, then come back to the Calculator. | calculator | hard | 1 | 5 | 31s | pass1 model |
multi.three_appsOpen the Calculator, then the Clock, then Settings. End on Settings. | springboard | medium | 1 | 5 | 45s | 6/6 |
notes.create_titledOpen the Notes app and create a new note whose first line is exactly "phoneshell benchma | mobilenotes | medium | 2 | 7 | 46s | pass1 model |
notes.searchOpen Notes and use search to find the note containing the words 'phoneshell benchmark'. | mobilenotes | medium | 2 | 9 | 56s | pass1 model |
notes.text_selectOpen Notes, create a new note, type the word 'anchor', then long press that word so the | mobilenotes | hard | 2 | 10 | 101s | 5/6 |
notes.type_longOpen Notes, create a new note, and type: the quick brown fox jumps over the lazy dog nin | mobilenotes | medium | 2 | 6 | 45s | 5/6 |
notes.type_punctuationOpen Notes, create a new note, and type exactly: benchmark-test (v2). | mobilenotes | medium | 2 | 6 | 36s | 6/6 |
probe.image.eightOpen https://blolabel.ai/probe/eight/ and tell me the word shown in the picture. | mobilesafari | medium | 2 | - | - | not swept |
probe.image.fiveOpen https://blolabel.ai/probe/five/ and tell me the word shown in the picture. | mobilesafari | medium | 2 | - | - | not swept |
probe.image.fourOpen https://blolabel.ai/probe/four/ and tell me the word shown in the picture. | mobilesafari | medium | 2 | - | - | not swept |
probe.image.nineOpen https://blolabel.ai/probe/nine/ and tell me the word shown in the picture. | mobilesafari | medium | 2 | - | - | not swept |
probe.image.oneOpen https://blolabel.ai/probe/one/ and tell me the word shown in the picture. | mobilesafari | medium | 2 | - | - | not swept |
probe.image.sevenOpen https://blolabel.ai/probe/seven/ and tell me the word shown in the picture. | mobilesafari | medium | 2 | - | - | not swept |
probe.image.sixOpen https://blolabel.ai/probe/six/ and tell me the word shown in the picture. | mobilesafari | medium | 2 | - | - | not swept |
probe.image.tenOpen https://blolabel.ai/probe/ten/ and tell me the word shown in the picture. | mobilesafari | medium | 2 | - | - | not swept |
probe.image.threeOpen https://blolabel.ai/probe/three/ and tell me the word shown in the picture. | mobilesafari | medium | 2 | - | - | not swept |
probe.image.twoOpen https://blolabel.ai/probe/two/ and tell me the word shown in the picture. | mobilesafari | medium | 2 | - | - | not swept |
probe.text.eightOpen https://blolabel.ai/probe/eight/ and tell me the written word. | mobilesafari | medium | 2 | - | - | not swept |
probe.text.fiveOpen https://blolabel.ai/probe/five/ and tell me the written word. | mobilesafari | medium | 2 | - | - | not swept |
probe.text.fourOpen https://blolabel.ai/probe/four/ and tell me the written word. | mobilesafari | medium | 2 | - | - | not swept |
probe.text.nineOpen https://blolabel.ai/probe/nine/ and tell me the written word. | mobilesafari | medium | 2 | - | - | not swept |
probe.text.oneOpen https://blolabel.ai/probe/one/ and tell me the written word. | mobilesafari | medium | 2 | - | - | not swept |
probe.text.sevenOpen https://blolabel.ai/probe/seven/ and tell me the written word. | mobilesafari | medium | 2 | - | - | not swept |
probe.text.sixOpen https://blolabel.ai/probe/six/ and tell me the written word. | mobilesafari | medium | 2 | - | - | not swept |
probe.text.tenOpen https://blolabel.ai/probe/ten/ and tell me the written word. | mobilesafari | medium | 2 | - | - | not swept |
probe.text.threeOpen https://blolabel.ai/probe/three/ and tell me the written word. | mobilesafari | medium | 2 | - | - | not swept |
probe.text.twoOpen https://blolabel.ai/probe/two/ and tell me the written word. | mobilesafari | medium | 2 | - | - | not swept |
reminders.createOpen Reminders and add a new reminder called 'collect parcel'. | reminders | medium | 2 | 6 | 52s | fail1 model |
resilience.deep_startOpen Safari and go to example.com | mobilesafari | medium | 1 | 3 | 29s | 5/6 |
resilience.scrolled_listIn Settings, open the General page. | Preferences | medium | 2 | 5 | 43s | pass1 model |
resilience.wrong_appGo to the Settings page for Display and Brightness. | Preferences | medium | 2 | 13 | 69s | pass1 model |
safari.link_menuOpen Safari, go to example.com, and long press the 'More information...' link until its | mobilesafari | hard | 2 | 4 | 32s | 6/6 |
safari.navigate.exampleOpen Safari and go to example.com | mobilesafari | easy | 2 | 3 | 40s | pass1 model |
safari.tabs.open_switcherOpen Safari and show the tab switcher, where all the open tabs are listed. | mobilesafari | medium | 2 | 4 | 27s | 5/6 |
safari.read.headingOpen Safari, go to example.com, and tell me the exact heading shown on that page. | mobilesafari | medium | 1 | 3 | 22s | pass1 model |
safari.search.webOpen Safari and search the web for 'iso 8601 date format'. Stop once results are showing | mobilesafari | medium | 2 | 5 | 37s | pass1 model |
settings.about.reachOpen Settings and navigate to the About page that shows the iPhone's model number. | Preferences | easy | 2 | 6 | 38s | pass1 model |
settings.about.serialOpen Settings and find the About page, then tell me the Model Number of this iPhone. | Preferences | medium | 2 | 6 | 40s | pass1 model |
settings.accessibility.reachOpen Settings and go to the Accessibility page. | Preferences | easy | 2 | 23 | 109s | pass1 model |
settings.back.twiceOpen Settings, go into General, then into About, then go back to the main Settings list. | Preferences | medium | 3 | 9 | 78s | 6/6 |
settings.battery.percentageOpen Settings, go to the Battery screen, and tell me the current battery percentage. | Preferences | medium | 2 | 11 | 55s | pass1 model |
settings.bluetooth.reachOpen Settings and go to the Bluetooth page. | Preferences | easy | 2 | 4 | 26s | pass1 model |
settings.datetime.reachOpen Settings and navigate to the Date & Time page under General. | Preferences | medium | 2 | 10 | 68s | pass1 model |
settings.fonts.reachOpen Settings and navigate to the Fonts page under General. | Preferences | medium | 2 | 7 | 61s | pass1 model |
settings.keyboard.reachOpen Settings and navigate to the Keyboard page under General. | Preferences | medium | 2 | 8 | 71s | pass1 model |
settings.scroll.to_bottomOpen Settings and scroll all the way to the very bottom of the main settings list. | Preferences | medium | 2 | 7 | 69s | 4/6 |
settings.search.wifiOpen Settings and use its search field to search for "Wi-Fi", then open the Wi-Fi settin | Preferences | medium | 2 | 6 | 45s | pass1 model |
settings.search.batteryOpen Settings, search for "Battery" using the search field, and open the Battery page fr | Preferences | medium | 2 | 6 | 49s | pass1 model |
settings.search.bluetoothOpen Settings, search for "Bluetooth" using the search field, and open the Bluetooth pag | Preferences | medium | 2 | 26 | 199s | pass1 model |
settings.search.storageOpen Settings, search for "Storage" using the search field, and open the Storage page fr | Preferences | medium | 2 | 5 | 31s | pass1 model |
settings.software_update.reachOpen Settings and navigate to the Software Update screen. Do not install anything. | Preferences | medium | 2 | 8 | 49s | pass1 model |
settings.storage.reachOpen Settings and navigate to the iPhone Storage screen that shows how much space is use | Preferences | medium | 2 | 7 | 38s | pass1 model |
settings.textsize.reachIn Settings, reach the screen with the slider that makes text larger. Do not change it. | Preferences | hard | 2 | 23 | 179s | 6/6 |
settings.wifi.switch_visibleOpen Settings and go to the Wi-Fi page. | Preferences | easy | 2 | 4 | 29s | 5/6 |
system.appstore.openOpen the App Store app. Do not download anything. | AppStore | easy | 1 | 3 | 22s | pass1 model |
system.notifications.openOpen the iPhone's Notification Centre. | springboard | hard | 1 | 24 | 232s | 0/6 |
system.spotlight.find_appUsing the iPhone's own search (Spotlight), find the Calculator app and open it. | springboard | medium | 1 | 7 | 54s | pass1 model |
weather.openOpen the Weather app. | weather | easy | 1 | - | - | not scored |
weather.read_temperatureOpen the Weather app and tell me the current temperature where I am. | weather | medium | 2 | - | - | not scored |
A Mac, an iPhone, a cable and an Apple developer account. The harness builds and installs the automation runner onto the phone for you.
git clone https://github.com/Blomega/phoneshell
cd phoneshell
bin/phoneshell doctor # tells you exactly what is missing
bin/phoneshell setup # builds and installs the runner on your iPhone
bin/phoneshell up # brings the bridge up
bin/phoneshell bench --list
bin/phoneshell bench --model claude-sonnet-5
Or skip the hardware entirely. Every task definition and every scored run on this page is published as a dataset at huggingface.co/datasets/blolabel/phoneshell-bench, including the paired runs behind the result above, so the statistics can be recomputed from source rather than taken on trust.
The harness is free and open. What breaks it is not your code: it is an iOS point release, an Xcode update or a change in WebDriverAgent, any of which can stop the bridge working on a Tuesday morning. We keep a tested build working and there is a person to call.
Cloud farms give you clean, ephemeral phones. Some things need a real handset: an OTP to a real SIM, a payment through a local wallet, region-locked content, an account with history behind it.
Every published mobile-agent number is Android, because Android runs in software and iOS does not. If you want your model measured on iOS rather than assumed to transfer, the suite above is free to run today.