← Back to Blog

The Best AI Agent on the Market Finishes One Job in Five Properly

Two of the newest AI agent benchmarks tested the best models on real, everyday computer tasks. The top score was 33%. Here's what that means for anyone being sold a fully autonomous AI employee right now.

By · · blog

The Best AI Agent on the Market Finishes One Job in Five Properly

A researcher gave the best AI agent money can buy a laptop and a list of everyday jobs. Book a flight. Fill in a form. Sort a spreadsheet. Send an email with the right attachment. Nothing exotic, it's the stuff a half-decent admin assistant does before their first coffee.

The best model on the leaderboard finished one in five of them properly.

That's not a typo. OSWorld 2.0, a benchmark built this year to test AI agents on long, real-world computer tasks, ran the strongest configuration they had, Claude Opus 4.8 with maximum thinking switched on, through 500-step jobs on real software. It completed 20.6% of them properly. GPT-5.5 managed 13%. These aren't toy tasks either. Same category as things you'd hand a new hire in week one.

Here's the thing that should worry anyone shopping for an "AI employee" right now. This isn't a story about the model being thick. Read what the researchers actually found when they watched the agents fail: they lose track of constraints halfway through the job, they miss information that turns up mid-task, they guess instead of asking the user a question, and they skip checking their own work before calling it done. That's not a maths problem. That's exactly what a distracted junior does on their third week, minus the bit where they'd normally put their hand up and say "sorry, can you clarify."

A separate benchmark, ClawBench, ran a similar test on live production websites instead of a lab environment. Real forms, real dynamic pages, the actual mess of the internet rather than a clean sandbox. Eight frontier models went through it. The best one, Claude Sonnet 4.6, completed 33.3% of tasks. The runner up managed 26.1%. So on a good day, two out of three jobs still needed a human to step in and finish them, or clean up after they went sideways.

I build automation for a living. I sell the idea that a business can run itself. So you'd think I'd be the last person telling you an AI agent can only do one job in three properly. But that's exactly why I need to say it out loud, because half the pitch decks landing in small business owners' inboxes right now are selling the opposite story. "Hire an AI employee." "Let AI run your operations end to end." "Set it and forget it." Someone in your inbox this week is trying to sell you a fully autonomous agent that "just handles it."

Handles it about one job in five, based on the same numbers the people building these things are publishing themselves.

And here's the bit that actually matters for your business, and it's not "AI is rubbish, ignore it." It's sort of the opposite. There's a massive difference between AI automation and AI intelligence, and mixing the two up is where the money gets wasted.

Automation is deterministic. If this happens, do that. A form submits, a webhook fires, a tag gets added, an email goes out, a calendar slot gets booked. No guessing, no judgement call, same result every single time. That's the stuff I build for clients and it's boringly reliable because there's nothing for the machine to get wrong. It's not deciding anything. It's just doing the thing you told it to do, the same way, every time.

A goal-driven AI agent is different. You hand it a goal, "sort out this booking" or "reconcile these two spreadsheets," and it has to work out the steps itself, on the fly, in a live environment that keeps changing underneath it. That's the bit that's still failing four times out of five on the hardest, most realistic tests anyone's published this year. Not because the model's dim. Because open-ended judgement in a messy, real environment is a genuinely hard problem, and we're not there yet, no matter what the demo video showed you.

So where does that leave you if you're running a business and trying to work out what to actually build?

Anything with a clear trigger and a clear outcome, automate it properly with deterministic workflows. Lead comes in, gets tagged and routed. Invoice goes overdue, reminder goes out on day three, day seven, day fourteen. Booking confirmed, calendar and CRM update together. This is the stuff that should already be running itself in your business, and it's the stuff that actually does, reliably, because there's no decision being made, just a rule being followed.

Anything that needs real judgement in a changing environment, an AI can help you do it faster, but keep a human checking the output for now. Drafting a client proposal, triaging which leads are worth a callback, summarising a messy inbox, that's a great use of AI as an assistant sitting next to you. It is not yet a great use of AI as an unsupervised operator making the calls and moving on without you seeing it.

The tell is right there in what the researchers watched the agents actually do. They guess rather than ask. If you deploy something that guesses instead of asking, in your business, with your customers, you own every one of those guesses. A workflow that fires the same way every time doesn't guess. It just does the job you built it to do.

I've had three separate conversations this month with business owners who'd bought an "AI agent" tool, pointed it at a genuinely open-ended job, and then spent longer fixing what it got wrong than it would've taken to do the job themselves. Not because the tool was a scam. Because the job they gave it was exactly the kind of long, multi-step, judgement-heavy task that's currently failing on the best public benchmarks four times out of five. Wrong tool, wrong job, and nobody told them the failure rate before they signed up.

Do the maths before you buy the pitch. Ask what specifically the AI is meant to decide, on its own, without you checking. If the answer involves more than a couple of steps and any real judgement, that's still very much a supervise-it job right now, not a set-it-and-forget-it one. Save the full autonomy for the stuff that's actually deterministic. That's where the reliable wins already are.

Sources:

Want someone to go through your business and tell you which bits should be running themselves? That's the AI Ops Audit. I look at how you actually operate, find the jobs a machine should be doing, and hand you the order to fix them in. Not a 40 page report nobody reads.

Brewed by Steven, poured by Viktor

About Steven Tann: Steven helps business owners build systems that run themselves using AI. After 10+ years helping 7,000+ businesses and building his own autonomous operations, he's the bloke who actually does it, not just talks about it. Find out more at steventann.com.

Tags: Small Business Automation, AI for Small Business, Practical AI, AI Agents, AI Reality Check