Three numbers from this year, each with a primary source you can open in a new tab. This is the public version of the credibility tile we show in the workshop, built so a sceptical operator can fact-check it before paying for a seat.
Source posture
Primary links first, methodology linked separately, and the as-of date travels with the tile so screenshots stay accountable.
Agents already work. Here's the receipts.
Anthropic
85%
% task success on OSWorld-Verified (real desktop apps)
Computer-use agents now hit 85% on OSWorld's real-desktop task suite — above the ~72% human baseline — with five different vendors inside five points of the lead.
xlang.ai
OpenAI / Anthropic
90–93%
% task success on Online-Mind2Web (300 tasks, 136 live sites)
Web agents complete 90–93% of real tasks across 136 live websites — bookings, forms, lookups — the "click around the web for me" job is no longer demo-ware.
github.com/OSU-NLP-Group
Perplexity
#1
for B2B SaaS citation quality
Perplexity beats ChatGPT and Google AI Mode on citation quality for B2B SaaS — AI search drives referral traffic today.
averi.ai
How to read this. Every number above links to its primary source — open them. The three logos are shown in greyscale on purpose: the numbers carry the slide, not the brands. We refresh these figures monthly; the "current as of" date moves even when the numbers don't, so the tile never goes stale-looking. Questions about methodology? The OSWorld NeurIPS 2024 paper is the method of record for the desktop number, and “An Illusion of Progress?” (COLM 2025) for the web-agent number — including the WebJudge auto-eval the leaderboard screens against.