RSS Amplifier

AI Safety Hong Kong · Jun 24, 2026

Anson Ho: How close is AI to taking my job? (And what the benchmarks aren't telling us)

0
Sign in to vote or save

AI Safety Hong Kong · AI Safety Hong Kong

All views shared here are the speaker’s own, and do not reflect the views of Epoch AI or AI Safety HK. Any errors in summarizing the talk are our own.

Anson Ho, a researcher and writer at Epoch AI, answers a question that’s on everyone’s mind: how close is AI really to taking our jobs? He walked us through what happened when he tried to get the best AI agents to do his own work, across three very different tasks: economic modeling, writing, and publishing. The results reveal why the benchmarks everyone is watching might be the wrong place to look and where the real bottlenecks actually are.

This was the second talk in AI Safety Hong Kong’s speaker series, focused on making AI safety approachable and relevant for a wider audience. Hope you enjoy!

“We’re, in some sense, searching under the streetlights, just looking at the things that are easy to check.”

TL;DR: AI benchmarks are great at measuring what is easy to grade, but terrible at predicting real-world job performance. Instead of reading benchmark leaderboards, spend a few hours testing the frontier models on your actual day-to-day tasks to see where the real bottlenecks are.

  • Benchmarks are disconnected from reality: Tests like MMLU focus on domains with checkable correctness (like math and coding), completely missing the “fuzzier” judgment required in real knowledge work.

  • Capabilities are “spiky”: AI might ace PhD-level science questions and build complex economic models in minutes, but still fail spectacularly at basic copy-and-paste formatting.

  • Tasks are not jobs: Automating a single task just shifts human effort to the next bottleneck. The impact of AI on the labor market will depend on how quickly that list of automatable tasks expands.

  • Hong Kong needs to adapt: Because access to leading models is often restricted locally, many professionals haven’t seen the true frontier. You need to know what these tools can actually do.

The central idea of Anson Ho’s recent talk was simple but powerful: benchmarks measure what is easy to grade, not what actually matters at work.

While scores on advanced tests like GPQA race ahead, the real-world impact on most jobs remains surprisingly small. Anson pointed to OpenAI’s GDPval, a benchmark that cost millions to develop as a leading indicator for automation. Even older models hit an 80% win rate against human experts on it, yet the measurable economic impact from AI has stayed remarkably muted.

His solution? Stop refreshing benchmark leaderboards and start testing the best models on your own job.

Here is a breakdown of his core arguments, his personal AI experiment, and what it all means for the future of knowledge work.

There is a growing disconnect between benchmark performance and real-world capability. Anson broke down exactly why this happens:

  • Searching under the streetlights: Models have blown past human scores on GPQA (PhD-level science questions), but acing a test doesn’t mean an AI can do a PhD scientist’s actual job. As Anson noted, “We’re, in some sense, searching under the streetlights, just looking at the things that are easy to check.”

  • Real jobs are messy: Professional roles are multi-faceted and require years of context. Benchmarks, on the other hand, are often single, isolated problems with a clear right answer.

  • A bias toward math and code: Benchmarks cluster in domains with checkable correctness. Fields like writing, law, and strategy remain under-measured because they require human judgment.

  • The contamination problem: The reliability of any benchmark depends on keeping the test set private. For example, Epoch’s FrontierMath (which features research-level problems reviewed by top minds like Terence Tao) is heavily guarded to prevent AI labs from training on it. Public test sets simply don’t offer that guarantee.

Instead of relying on expensive, formal benchmarks, Anson’s core advice is to conduct a personal test.

“Here is me automating my job for the sake of science. Let’s see how it goes.”

He cautions against building anything formal, which is costly and time-consuming. Instead, take an 80/20 approach: spend two to four hours on a task that actually matters to your work and that a model could plausibly handle. Push the AI hard, write down exactly where it fails, and repeat the experiment every few months to track the trend.

Anson acted as his own test subject, attempting to automate three specific tasks from his day-to-day workflow. The results were surprisingly jagged.

These uneven results perfectly illustrate a modern version of Moravec’s paradox: what is hard for a human is not necessarily hard for an AI, and vice versa.

The economic model was a grueling task for a human, yet the AI got impressively close because it excels at code and math. Meanwhile, the writing and publishing tasks were easy for Anson, but the AI struggled in bizarrely non-human ways. “I thought copy and paste would be a really easy thing... But it turns out that Claude was just unable to do it,”Anson remarked.

Based on this experiment, he shared a few vibes-based forecasts for when AI might master these tasks:

  • Economic model: ~50% chance the best agent could replicate it with 3–4 hours of effort by the end of 2026.

  • Article writing: ~50% chance an AI could match his original quality by early 2029.

  • Publishing: This is moving fast. His failed attempt is a floor, not a ceiling. Better prompting or APIs could likely solve this today.

Because AI training data distributions differ wildly from human experience, Anson expects this “spiky” capability progress to persist.

“Personally, I still expect AI to take my job within the next 10 years.”

Crucially, automating a task is not the same as automating a job. As specific tasks get handed off to AI, human effort simply shifts to whatever remains the bottleneck.

The real question is whether the set of automatable tasks will expand fast enough to leave humans with nothing to do. Extrapolating the progress of deep learning over the last decade, Anson’s personal bet is that AI could potentially do his job within roughly ten years. However, he explicitly pushed back on the overblown narrative that AI is about to take all jobs overnight, noting that current labor-market impacts are still quite small.

As a knowledge-work hub heavily anchored in finance and software, Hong Kong is squarely in the frame for AI exposure.

Looking through a policy lens, Anson suggested we need to focus on the intersection of those most exposed to automation and those least equipped to adapt—for instance, a mid-level software worker without liquid wealth or an easy path to pivot careers. While the dynamics in Hong Kong likely mirror research coming out of the US, our smaller, less diversified economy may have a harder time absorbing rapid labor shocks.

His final, practical nudge for the local audience was clear. Because access to leading models has historically been awkward in Hong Kong, many professionals still haven’t seen what the frontier of AI can actually do.

“Benchmarks aren’t very realistic right now, because they’re designed to be easy to evaluate, and so you should try to test out the best models on your own job in specific tasks.”

If you’re meeting AISHK for the first time through this post, here’s how to follow what we’re doing.

  • Subscribe to this Substack for recaps, original writing from the team, and reading recommendations.

  • Follow AI Safety Hong Kong on LinkedIn and visit aisafetyhk.org.

  • In-person AI Safety camps with ML4Good, TARA, BlueDot and more are coming in the second half of 2026, alongside our series of speaker events.

You can follow Anson’s work on Github, LinkedIn, and Epoch AI

If anything here resonates and you’d like to get involved, reach out through any of the channels above — wherever you are.

No posts

Read the original on aisafetyhk.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.