
AI labs now improve their models mainly by reinforcement learning (RL): giving the model a realistic task, letting it try, and rewarding it when a checker says the answer is right. Somebody has to write those tasks and checkers, and for frontier models that somebody has to be an expert, because the task must be hard enough to challenge the model. According to SemiAnalysis (a respected chip and AI research firm), this has created a new supply chain. Mercor, Surge and Handshake, the three incumbents, are each past $1 billion in annual revenue; newcomers barely a year old are approaching $100 million. Mercor logged 2,517,000 expert hours in the second quarter of 2026, the equivalent of about 4,800 people working full time, and its average pay rate passed $100 an hour. Frontier labs pay $5,000 or more for a single good coding task, and the best contractors earn over a million dollars a year. Meta has moved about 3,000 of its own engineers into writing these tasks full time. The old picture of AI training data, low-paid workers labelling images, is out of date for the frontier. The new picture is closer to hiring the best teachers to write the hardest exams.
The short version
In 2024 Ilya Sutskever, OpenAI's co-founder, famously called data "the fossil fuel of AI": a finite resource the industry would soon use up. SemiAnalysis, the research firm whose work we draw on for most of these briefs, argued in July that he was half right. The internet's supply of text is finite. But if demand is strong enough, people can be paid to make more of the specific data AI needs, and that is exactly what has happened.
The numbers in their piece describe an industry most people have never heard of. Three established companies, Mercor, Surge AI and Handshake, are each above $1 billion in annual revenue. Surge is estimated at around $3 billion with fewer than 110 employees and no outside investors. Companies that did not exist two years ago, such as Fleet, Mechanize and Afterquery, are approaching $100 million. What they sell is carefully designed homework for AI.
Why AI needs homework now
The first big AI models learned by reading: predict the next word across a large slice of the internet, billions of times. The current generation improves mostly through reinforcement learning, RL for short. Instead of reading, the model practises. It gets a task (fix this bug, build this spreadsheet model, draft this contract clause), a working environment to do it in, tools it can use, and a verifier: an automatic checker, or a detailed marking guide called a rubric, that decides whether the result is good. Good attempts are rewarded, and the model gradually gets better at that kind of work.
The hard part is the task. SemiAnalysis makes the point that a task is only useful if it sits at the right difficulty: too easy and the model learns nothing, too hard and it never earns a reward. For today's frontier models, "too easy" covers most of what an average professional would think to ask. So the people writing tasks spend much of their time making them harder, which means they need to be genuinely expert in the field. You cannot write an exam that stretches a top student unless you understand the subject better than they do.
This is why some AI insiders think more and better tasks are the main thing standing between today's AI and automating much of office work. SemiAnalysis quotes the Anthropic researcher Sholto Douglas saying, on the Dwarkesh podcast last year, that current methods would be enough to automate white-collar work given enough of the right data. That is a prediction, and a contested one, but it explains the spending.
What the work pays
Here is what SemiAnalysis reports. Mercor disclosed 2,517,000 expert hours on its platform in the second quarter of 2026, about 4,800 people working 40-hour weeks. Two years earlier it was around 2,000 hours a quarter, and the most recent quarter alone grew 69%. Mercor's average pay rate recently passed $100 an hour, with software engineers well above that average. The top contractors across these companies earn over seven figures a year. Frontier labs will pay $5,000 or more for a single decent coding task.
A bit of arithmetic, ours rather than theirs, makes that last number feel less strange. Mechanize, one of the new companies, expects the software engineers it pays over $400,000 a year to produce about one good task per week. $400,000 spread over 52 weeks is roughly $7,700 of salary per task before any overheads. Seen that way, $5,000 for a good task is close to what it costs to make one.
The same logic runs through the whole industry: 2.5 million hours at over $100 an hour is more than $250 million paid to experts in one quarter on one platform. That money buys expert judgement.
Labs will pay $5,000 or more for one good coding task, and the best contractors earn over $1 million a year (SemiAnalysis)
Is this actually new?
Paying humans to make AI training data is old. For most of the last decade it meant large teams, often in lower-income countries, drawing boxes around objects in photos or tagging text as harmful. In 2023 TIME reported that Kenyan workers labelling toxic content for OpenAI through an outsourcing firm earned less than $2 an hour. Scale AI built a multi-billion-dollar business on that model, and Meta paid $14.3 billion last year for a stake in it, largely to hire its founder Alexandr Wang.
What is new is who does the work and what it pays. SemiAnalysis is blunt that the era of undereducated contractors drawing bounding boxes is over for frontier AI: the models are now good enough that creating a useful training example is a real intellectual problem. Designing a task that cannot be gamed (AI models are good at finding shortcuts to a reward, known as reward hacking), calibrating its difficulty, and doing this at scale without the quality slipping are engineering jobs. The pay has moved from a couple of dollars an hour to a hundred. The cheaper, older kind of labelling still exists, and the people doing it have not vanished. They are simply not where the frontier's money goes any more.
The everyday version
Think of a very gifted student preparing for the hardest exam in their field. Textbooks got them a long way, but they have read all of them. What helps now is practice papers, and not any practice papers: questions hard enough to stretch them, written by examiners who know the subject better than they do, with a marking scheme precise enough to say exactly what a good answer looks like. There are not many people who can write those papers, so they get paid very well, and a single excellent question is worth a surprising amount of money.
The AI labs are that student. The expert contractors are the examiners. And because the student learns fast and needs a fresh stack of harder papers every few months, the examiners' order book never closes.
Why Meta turned its own staff into examiners
The most striking example in the SemiAnalysis piece is Meta. In late May, as part of a restructuring, Meta created an "applied AI engineering" organisation and moved about 3,000 engineers, including 70% of its new graduates, into making RL tasks and environments full time. SemiAnalysis estimates that puts Meta in the same ballpark as Mercor's whole platform, with a larger pool of staff behind it if the experiment works. It connects to a story we covered earlier, that Meta also began recording employees' screens and keystrokes, because real recordings of office work make the most realistic tasks.
Two caveats. The revenue figures in the chart come from public disclosures and, for Surge, a rumour; they are "gross" revenue, which includes the money passed on to contractors. And the belief that more tasks will be enough to automate office work is a bet, not a result. But whichever way that bet goes, a new kind of job now exists: being paid, handsomely, to be harder to fool than a machine.
Curious about AI? Come build with us.
Oslo Vibe Coding runs free, beginner-friendly drop-ins where we build real things with AI. No one codes alone.