All videos
0:00 / 0:00
research

Chain of Thought: Introducing Remote Labor Index (RLI)

Scale AI16 June 2026Watch on YouTube

Description

Introducing the Remote Labor Index, RLI. Brad Kenstler, Head of Agent Capabilities and Environments, discusses RLI with Bing Liu, Head of Research, Madhu Sehwag, Research Scientist, and Mantas Mazeika, Research Scientist at the Center for AI Safety. The Remote Labor Index (RLI) is a benchmark that empirically measures the capability of AI agents to perform real-world, economically valuable remote work. 0:00 Introduction 1:00 Overview of RLI 5:11 Benchmarking Freelance work 10:32 Comparing RLI to other professional domain benchmarks 12:18 Deep dive on RLI tasks 17:27 Making tasks representative of real-world work 22:50 Rubrics vs judge-based evaluation 26:15 Bottlenecks on agentic capabilities 29:30 Which agents does RLI evaluate 34:04 Failure modes of RLI 37:30 Implications on the future of remote labor 42:10 Unlocking performance improvements on RLI Learn more about the benchmark at: https://scale.com/leaderboard/rli

What you'll learn

  • Remote Labor Index (RLI) is a benchmark measuring the capability of AI agents to perform real-world, economically valuable remote work.
  • RLI includes tasks representative of actual freelance work, focusing on task design that mimics genuine work scenarios.
  • Evaluation of AI agent performance uses rubrics and judge-based methods to assess work quality.
  • The benchmark identifies bottlenecks in agentic capabilities and determines which AI models can be evaluated.
  • RLI provides insights into the implications of AI agents for the future of remote work and potential performance improvements.

Frequently asked questions

What is the Remote Labor Index and what does it measure?
The Remote Labor Index (RLI) is a benchmark from Scale AI that empirically measures how well AI agents can perform real-world, economically valuable remote work. It evaluates AI models on their ability to perform tasks representative of actual freelance work.
How does RLI ensure that tasks are representative of real work?
RLI focuses on task design that mimics actual work, using freelance work as a reference point. This ensures the benchmark provides realistic evaluations of what AI agents can accomplish in practical situations.
What evaluation methods does RLI use for assessing AI agent performance?
RLI uses rubrics and judge-based evaluation methods to assess work quality. These approaches help objectively determine how well AI agents perform the tasks.
What are the implications of RLI for the future of remote work?
RLI sheds light on how AI agents will impact the future of remote labor. The benchmark helps identify bottlenecks in agentic capabilities and where performance improvements are possible.

Topics