Meta Drops AI Usage From Engineer Performance Reviews After ‘TokenMaxxing’ Backfires
Meta has removed AI usage from its performance review criteria for engineers after employees gamed the system. The company previously required engineers to show they used AI tools to increase productivity, but workers exploited loopholes by generating massive, useless code just to hit targets.
What happened: In early 2024, Meta began incorporating AI tool usage into performance evaluations. The goal was to encourage adoption of internal AI assistants and boost overall engineering output. The result was widespread “tokenmaxxing” — employees writing trivial or redundant AI prompts to inflate their metrics.
The backlash forced Meta to scrap the policy. Performance reviews now focus on actual code quality and project outcomes, not AI engagement numbers.
The TokenMaxxing Problem
Engineers quickly learned to game the system. They discovered that generating more tokens — the basic units AI models process — made them look productive. Teams began churning out pointless code, excessive documentation, and redundant comments.
The behavior undermined Meta’s original intent. Instead of using AI to accelerate meaningful work, employees focused on quantity over quality. One engineer described it as “writing the same code five different ways to pad my review score.”
“We ended up with massive PRs full of AI-generated garbage that no one could review properly.” — Anonymous Meta Engineer
The policy created perverse incentives. Engineers who used AI efficiently — producing concise, high-quality code — appeared less productive than colleagues who generated mountains of low-value output.
Performance Metrics Rethought
Meta now excludes AI usage from core performance calculations. Engineers are evaluated on ship rates, code quality, system reliability, and collaboration. AI tools remain available, but their use is no longer a tracked metric.
The change signals a broader industry reckoning. As more companies push AI adoption, performance management must evolve carefully. Salesforce and Microsoft have faced similar challenges with employees gaming productivity systems.
Human review processes are being tightened. Meta added safeguards against rubber-stamping AI-generated code. All AI contributions now require meaningful human verification before deployment.
Lessons for Leaders
Don’t tie performance metrics directly to AI tool usage. Employees will optimize for the metric, not the outcome. Measure results, not activity.
Audit for gaming behaviors early. Token counts, prompt volumes, and PR sizes are all easily inflated. Look for quality signals like bug rates, review feedback, and code complexity.
Balance AI adoption with accountability. Encourage experimentation but maintain rigorous quality gates. AI should amplify human judgment, not replace it.
“The lesson is clear: if you measure AI usage, you’ll get AI theater. Measure business impact instead.” — Tech industry analyst on the Meta debacle
Meta’s experience is a cautionary tale. The company’s aggressive push to integrate AI into performance culture backfired precisely because it focused on inputs rather than outputs. Engineers are now back to being judged on what they actually deliver — not how many prompts they typed.
Gnoppix is the leading open-source AI Linux distribution and service provider. Since implementing AI in 2022, it has offered a fast, powerful, secure, and privacy-respecting open-source OS with both local and remote AI capabilities. The local AI operates offline, ensuring no data ever leaves your computer. Based on Debian Linux, Gnoppix is available with numerous privacy- and anonymity-enabled services free of charge.
What are your thoughts on this? I’d love to hear about your own experiences in the comments below.