GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Remote Labor Index: Fable 5 Automates 16.1% of Freelance Work — Double Opus 4.8, Six Times October Baseline

The Center for AI Safety and Scale Labs published new Remote Labor Index results on July 1. The headline number: Claude Fable 5 completes 16.1% of real freelance projects at professional quality — more than double any other model evaluated and six times the 2.5% ceiling that marked the benchmark’s launch in October 2025.

What the Benchmark Measures

The Remote Labor Index runs AI agents against actual paid projects sourced from freelance platforms. Categories include 3D and CAD, architecture, graphic design, video and animation, audio production, data analysis, and web app development. Human evaluators judge each AI deliverable against a gold-standard output produced by a paid professional. The automation rate is the fraction of projects where the AI’s work matches or beats the professional’s.

There are no synthetic prompts. No multiple-choice shortcuts. Each task is a real job someone paid a human to do.

Current Results

ModelAutomation Rate
Claude Fable 516.1%
Claude Opus 4.88.3%
GPT-5.56.3%
October 2025 baseline2.5%

Fable 5’s 16.1% is the highest automation rate the benchmark has recorded. It is roughly double Opus 4.8 and more than triple GPT-5.5 — not a marginal lead but a clean separation. The gap between consecutive frontier tiers (Fable 5 vs Opus 4.8 vs GPT-5.5) is also notable: this is not a benchmark where all frontier models cluster together.

The Trajectory

When the RLI launched in October 2025, the best AI agents cleared 2.5% of tasks. Eight months later, that number is 16.1% with stronger agent scaffolding. The CAIS blog notes the new evaluations were run with updated scaffold configurations alongside the newer models — meaning the automation rate reflects a model-plus-harness combination, not raw model capability in isolation.

That distinction matters for interpreting the numbers. A 6x jump in eight months is not purely a function of better base models. It is partly better scaffolding, retrieval, tool use, and multi-step execution wrapping those models. The benchmark is measuring deployable systems, not isolated inference.

The implication is that the automation rate will continue rising as scaffolding improves even without underlying model improvements. Fable 5’s 16.1% is a floor, not a ceiling.

What Is Not Being Automated

The gap between 16% and 84% matters as much as the number itself. The projects in the RLI span high-complexity creative and technical work. A 16% automation rate means Fable 5 is completing roughly one in six of these tasks to professional standard — and failing five. The benchmark covers categories where professional-quality output requires judgment, taste, and revision cycles that current systems handle inconsistently.

The specific failure modes are not yet published. What categories are driving the 84% failure rate will be more informative than the headline number once CAIS releases task-level breakdowns.