TL;DR: Accenture ties promotion eligibility to weekly AI logins. Meta made “AI-driven impact” a core 2026 review criterion. Amazon built a leaderboard, told employees it would stay out of reviews, then watched employees game it anyway, a practice now called tokenmaxxing. I look at why performance ratings already struggled before AI usage got added, and what a number has to survive before a promotion committee should ever see it.
Three companies made the same bet in 2026: that how often someone touches an AI tool says something real about how well they do the job, real enough to put beside their name at review time. Accenture told associate directors and senior managers in February that promotion into leadership now requires “regular adoption” of the firm’s AI tools, tracked as weekly logins, according to CNBC. Meta’s head of people told employees that starting with the 2026 cycle, “AI-driven impact” becomes a core review expectation for engineers, marketers, and designers alike, a policy reported by HR Grapevine. Amazon went further, building an internal leaderboard and pushing for four out of five developers to use AI tools every week, tracked through MeshClaw, an internal agent platform that can deploy code, sort email, and act inside Slack.
Amazon told employees the leaderboard would stay separate from their reviews. Employees found the fastest way up it anyway: run the agent on work that did not need doing. The Financial Times reported the practice spread wide enough to earn its own name, tokenmaxxing. “Some people are just using MeshClaw to maximize their token usage,” one employee told the paper. Another said the part everyone suspected: “Managers are looking at it.” They didn’t believe Amazon’s promise, and several managers kept watching the leaderboard after the company restricted who could see it.
Meta’s formal review criterion produced an identical instinct on its own, informal side. An employee built an internal dashboard nicknamed Claudeonomics, ranking the company’s roughly 85,000 workers by how many tokens they burned. Fortune reported that in a single thirty-day window, total consumption passed 60 trillion tokens, an estimated nine billion dollars in cost, with the single biggest user running through 281 billion tokens in a month. Titles went to the winners: Token Legend, Cache Wizard. Neither Mark Zuckerberg nor CTO Andrew Bosworth cracked the top 250. Bosworth had told a conference that an engineer spending a full salary on tokens proved the tool paid for itself tenfold, and the only response was to keep spending. Once “AI-driven impact” sits in a formal review, a dashboard ranking who has the most of it is not a curiosity. It is the exam key. This is Meta’s second turn here in as many weeks: last week was about the return of its managers, a coaching problem. This one is about what counts as proof of using AI well.
Microsoft shows what happens without even a formal policy. President Julia Liuson told staff AI use was “no longer optional, it’s core to every role and every level,” language close enough to a review criterion that employees treated it as one. A company spokesperson later walked it back, telling reporters there was “no formal review of an employee’s AI usage,” the kind of clarification a company issues once its first message has already reached everyone wondering about their next rating.
Three companies, three mechanisms, one identical outcome: a number folded into a review or a promotion decision stopped measuring performance within weeks of going live. Human Resources Director’s coverage of the Amazon story reached for the explanation economists have used for fifty years. British economist Charles Goodhart named it in the 1970s writing about monetary policy: when a measure becomes a target, it stops being a good measure.
Cognizant’s CEO Ravi Kumar made the same point in June, calling token consumption a “vanity metric,” and Fortune reported that Meta, Amazon, and OpenAI already leaned on it internally.
Amazon’s employees spent the following months proving Kumar right. The token count stopped describing how much useful work AI was doing the moment employees understood their promotions might ride on it, and started describing who could produce the biggest number.
This should not have surprised anyone who follows performance-review research. The instrument had a credibility problem before a single employee opened a chatbot. Researchers have studied how reliably two supervisors rate the same employee’s work for decades, and the finding that should worry every promotion committee reading this is not the average score. It is what changes the score. A meta-analysis published in Frontiers in Psychology found interrater reliability for overall job performance running around .61 when ratings get collected purely for research, with no promotion or raise attached. Once the same kind of rating feeds an actual administrative decision, a raise, a promotion, a stack rank, reliability drops to about .45. In plain terms: at .45, well under half of the difference between how two managers score the same person’s work traces back to that person’s actual performance. The rest is the rater, the scale, the mood in the room that week, and noise. Accenture, Meta, and Amazon each added a second, more gameable number on top of a review process already carrying that weakness.
Tokenmaxxing added a highly gameable number to a review process the research already knew wobbled the moment a raise or a promotion walked into the room, and the wobble got worse.
Deloitte put a figure on exactly this problem in 2015, years before anyone was arguing about AI adoption. Marcus Buckingham and Ashley Goodall, writing in Harvard Business Review, described a firm spending close to two million hours a year filling out forms, holding calibration meetings, and producing ratings that 58 percent of the firm’s own executives said drove neither engagement nor performance. Deloitte’s fix wasn’t a sharper rating scale. It replaced the annual review with frequent, forward-looking check-ins with a manager close enough to the work to say something specific, and built promotion decisions on that ongoing record instead of a single year-end number. I was part of the “Reinventing Performance Management” era at Deloitte. The lesson has traveled with me since: a performance signal is only as good as the person taking it and how often they take it, never how precisely the resulting number gets reported. A promotion tied to a login count or a token leaderboard runs on the opposite logic. It sits far from the work, looks backward, and gets glanced at once a year rather than discussed.
There’s a real case for visible pressure behind an AI adoption push, and it deserves a fair hearing. Left alone, plenty of capable employees won’t touch a new tool until a deadline or a promotion cycle forces it, and executives tying AI use to career advancement aren’t wrong about the urgency. Kelly Services surveyed more than 6,000 professionals and found 69 percent of executives believe refusing to adopt AI is a bigger threat to someone’s career than AI itself, and 59 percent said they would replace an employee who resisted using it. Tying usage to a review puts that expectation in front of thousands of people at once, and fast has real value when capital is committed at this scale.
Fast bought Accenture, Meta, and Amazon adoption within weeks. It also bought review inputs so unreliable that Amazon had to publicly disown the exact use its own employees assumed the number was for, a strange place for three sophisticated measurement organizations to land, by choice.
Every workforce leader building AI usage into a review or a promotion criterion is about to run the same experiment. Decades of research on performance ratings under real stakes, plus six months of watching these three companies relearn the lesson, point to the same test before that number reaches a promotion committee. Can it be gamed faster than the underlying skill can be built? If the honest answer is yes, expect what these three companies got: a great deal of activity in the numbers, and very little information about who actually deserved to move up.
Here’s how you take action
If your company is folding any version of AI usage into a review right now, logins, token counts, seat activation, ask the question research already answered: is this reliable enough to decide someone’s career, or just easy to count? If you can’t answer that with evidence, assume the worse one. Your employees already have.
Then check how close the rater sits to the actual work, and how often, before you look at the number at all. Deloitte’s decade-old fix still holds: a manager who checks in weekly with someone whose work they can see, and who builds a promotion case on that record, will tell you more than a leaderboard tracking a hundred thousand people at once, gamed or not.
One closing note, for the record: I still hate the word “tokenmaxxing.”
Christina Lexa writes Workforce Rewired, on the intersection of workforce transformation, AI, and global talent.
The views expressed here are my own and do not represent the position of my employer or any organization I am affiliated with.






