GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Amazon Employees Are 'Tokenmaxxing' MeshClaw to Game Internal AI Leaderboards

Amazon set a target: more than 80% of developers should use AI tools each week. To track progress, the company began publishing internal leaderboards showing AI token consumption by team. Employees responded by doing what people always do when given a visible metric: they optimised for the metric itself.

The behaviour now has a name. Amazon staff call it “tokenmaxxing.”

What MeshClaw Is

Amazon’s MeshClaw is an internal agentic platform built over the past year, directly inspired by OpenClaw, the open-source personal agent framework that went viral in February 2026. MeshClaw allows Amazon employees to create AI agents that connect to workplace software — Slack, internal deployment systems, email — and carry out tasks autonomously.

More than three dozen engineers built the product. An internal memo described it in aspirational terms: “It dreams overnight to consolidate what it learned, monitors your deployments while you’re in meetings, and triages your email before you wake up.”

The Perverse Incentive

Once token consumption appeared on team-visible leaderboards, a subset of employees began using MeshClaw not to do useful work, but to generate token activity. The tool’s agentic capabilities made this easy: connect it to a few internal endpoints, set it running, and watch the consumption numbers climb.

“There is just so much pressure to use these tools,” one Amazon employee told the Financial Times. “Some people are just using MeshClaw to maximise their token usage.”

Amazon has officially told staff that token statistics will not be used in performance evaluations. Multiple employees said they did not believe it. “Managers are looking at it. When they track usage it creates perverse incentives and some people are very competitive about it.”

The company subsequently limited access to leaderboard data so only employees and their direct managers can view individual statistics — an acknowledgment that publishing team-wide figures had backfired.

Not Just Amazon

Meta employees have independently developed similar practices on their own internal leaderboards. The behaviour appears to be a structural consequence of any organisation that sets visible AI usage targets without tying them to measurable output quality.

The Security Problem Nobody Solved

MeshClaw’s design gave agents broad permissions to act on behalf of users. Several Amazon employees flagged this as the more serious concern.

“The default security posture terrifies me,” one employee said. “I’m not about to let it go off and just do its own thing.”

The company is spending $200 billion in capital expenditure in 2026, the majority on AI and data centre infrastructure. Internally, the returns on that investment are being measured partly through token consumption dashboards. The gap between what the dashboard measures and what it is supposed to represent — genuine AI-assisted productivity — is where tokenmaxxing lives.

What This Signals

Amazon’s situation is not a failure of the technology. It is a textbook example of Goodhart’s Law: when a measure becomes a target, it ceases to be a good measure. The 80%-adoption target and the token leaderboard were designed to accelerate AI adoption. They succeeded. What they measured was not productivity but the act of measurement itself.

The companies likely to avoid this are ones that tie AI usage to output metrics — code shipped, support tickets resolved, decisions made with AI assistance — rather than raw token counts. The companies currently running leaderboards are learning this the hard way.