AI agents are now successfully completing 16.1 percent of paid freelance assignments at professional level. Eight months ago, that figure stood at just 2.5 percent. This emerges from the latest measurement of the Remote Labor Index (RLI), a benchmark developed jointly by the Center for AI Safety and Scale AI.
The acceleration is striking: in under a year, the highest automation rate has more than quadrupled. That makes the RLI one of the more concrete indicators of how quickly AI agents are becoming deployable for economically productive work, outside controlled laboratory settings.
How the Remote Labor Index works
The RLI does not measure whether AI can produce something, but whether the result meets the quality threshold of a paid professional. The benchmark consists of 240 projects with a combined value of $144,000, submitted by 358 verified freelancers. Project categories range from 3D design and CAD to graphic design, video and animation, audio, data analysis, web applications, software development, and writing.
Human evaluators from the Center for AI Safety compare each result against a reference solution produced by a paid professional. The measurements also revealed that AI judges used as a substitute for human evaluators did not work: they systematically rated newer models too favourably. The benchmark therefore retains human evaluation.
Which models score highest
In the most recent measurement, Fable 5 achieved the highest automation rate at 16.1 percent. A caveat is warranted here: of the 240 projects, 218 could be evaluated before the US government restricted access to the model. It is unclear whether the figure would be higher or lower with a complete evaluation.
In second place is Opus 4.8 with 8.3 percent, followed by GPT-5.5 with 6.3 percent. At the benchmark's launch eight months ago, Manus was the best-performing agent with a score of 2.5 percent. The simultaneous progress of multiple models indicates that the advances are not tied to a single provider.
What the figures do and do not say
An automation rate of 16 percent also means that 84 percent of the freelance assignments in this benchmark remain beyond the reach of AI agents for now. Complex, multi-layered, or highly context-dependent tasks still constitute the majority. The benchmark deliberately makes no claims about how the broader market is developing; it measures a sample of projects representative of the freelance platforms from which the assignments were collected.
What stands out is the pace of the increase. If the rate of the past eight months is sustained, more categories will come within reach. Which ones, and how quickly, cannot be predicted on the basis of current data.
Implications for those building with or investing in AI
For developers and product teams, the RLI provides more concrete guidance than most conventional benchmarks, which typically measure performance on academic tasks or isolated problems. The link to real, paid assignments makes the scores economically interpretable: what meets professional quality in the eyes of a client?
For investors and policymakers, the index confirms that AI agents are no longer suited only to straightforward automation. Categories such as data analysis, web research, and certain forms of writing already lend themselves to autonomous handling to a meaningful degree. At the same time, the finding regarding AI judges underscores that evaluation methods themselves must be critically scrutinised: those who use AI to assess AI risk systematically over-optimistic outcomes.