TLDR: AI use has become a formal promotion criterion at some of the world’s largest employers, but most are measuring tool adoption rather than judgement — rewarding theatre and, in regulated industries, creating real compliance exposure.
Accenture gates senior promotions on AI use — Meta scores it in every review
If a new question has started appearing in your performance review — “so, how are you using AI?” — you are not imagining it, and it is not small talk. Across 2026 a cluster of very large employers moved AI usage from encouraged behaviour to assessed behaviour, with consequences attached to the answer. What was a training initiative two years ago is now a line in the rating conversation, and in several cases an explicit condition for advancement. The shift happened quietly, inside review templates rather than press releases, which is why many professionals meet it for the first time in the room.
Accenture has been the most explicit. After training roughly 550,000 employees on its internal AI tooling, the firm told senior staff that consistent use of those tools is a condition for high-level promotion, and began monitoring weekly log-ins for parts of its senior population — only those showing regular adoption are considered for leadership roles. The logic is coherent from the top: an organisation that has spent heavily on capability wants evidence the capability is being used, and log-in data is the one signal that already exists in the system. What makes it consequential is the coupling. The moment a behavioural signal is wired to promotion, it stops being a diagnostic and becomes an incentive, and people optimise accordingly.
The pattern has spread quickly. As reported across 2026, Meta made “AI-driven impact” a core expectation in every employee’s performance review regardless of function, rewarding exceptional impact and signalling that ignoring the technology risks a lower rating. Amazon runs an internal system giving managers visibility into which AI tools staff use and how often, feeding that into assessment of engagement. Firms including KPMG have tied AI usage to progression and remuneration. Notably, these sit at different points on a spectrum: Meta’s framing is outcome-shaped (“impact”), while Accenture’s and Amazon’s are activity-shaped (log-ins, frequency). That distinction looks like wording. It is the whole argument.
For anyone being assessed, the practical consequence is that AI fluency has quietly migrated from a nice-to-have on a development plan to a factor in compensation and advancement. That migration changes what a review conversation demands of you: the burden has moved from describing intent to producing evidence, and the evidence a manager can actually weigh looks very different from the evidence a dashboard collects. For anyone designing these systems, the consequence is heavier still. You are about to find out what your workforce does when you tell them exactly which number you are watching, and the answer arrives faster than any correction to the metric can.
| Employer | What is measured | Consequence |
|---|---|---|
| Accenture | Consistent use of internal AI tools; weekly log-ins monitored for some senior staff | Regular adoption required to be considered for senior promotion |
| Meta | “AI-driven impact” as a core review expectation for every role | Exceptional impact rewarded; ignoring AI risks lower ratings |
| Amazon | Internal visibility of which AI tools staff use and how often | Feeds manager assessment of engagement |
| KPMG and others | AI usage linked to progression | Influences advancement and remuneration |
Counting log-ins measures compliance, not capability
Adoption metrics are seductive because they are cheap. Every enterprise AI deployment emits usage telemetry by default, so the number is sitting there, requiring no new instrumentation, no manager training, and no uncomfortable judgement call. That convenience is precisely why it gets adopted — and precisely why it fails. The metric that is easiest to collect is almost never the metric that describes the thing you care about, because ease of collection tracks how mechanical a signal is, while the qualities that distinguish strong professional work are the ones that resist mechanical capture.
What follows is Goodhart’s law in its purest form: when a measure becomes a target, it ceases to be a good measure. Tell a workforce that weekly log-ins influence promotion and you will reliably get weekly log-ins. You will get the consultant who opens the assistant on Monday morning to register activity, asks it to rephrase an email she had already written, and closes it. Her usage graph is exemplary. Her work is unchanged. Meanwhile the colleague who used the same tool twice in a quarter — once to pressure-test an assumption that turned out to be wrong, saving three weeks of misdirected analysis — shows up in the data as a laggard. The measurement has not merely failed to capture value; it has inverted the ranking.
There is a second, subtler cost that adoption metrics actively worsen. Field research on human-AI teams found that people who delegate heavily to AI edit its output substantially less, and that the resulting work — while individually acceptable — converges toward sameness across the group. Each person accepts a competent draft, few push back, and the organisation’s collective output narrows around whatever the model considers the median good answer. A KPI rewarding raw usage pushes exactly that behaviour: more delegation, less editing, less distinctiveness. You can hit the target and degrade the product simultaneously, and the dashboard will show green throughout.
Usage data tells you someone opened a tool. It is silent on whether they framed the problem well, noticed the model’s confident fabrication, weighed it against what they knew from the field, or made a decision they could defend to a regulator. Those capabilities are what actually separate a strong professional from a weak one in an AI-saturated workflow — and they are the ones the metric is structurally blind to. Automation bias, the well-documented tendency to over-trust automated output, means the people who use AI most uncritically may be generating the most hidden risk while posting the best adoption numbers.
Measure outcomes and judgement, not activity
The alternative is not to abandon AI expectations — it is to measure the thing you actually want. Start with outcomes: cycle time on defined deliverables, error and rework rates, quality as assessed by whoever receives the work, and impact on the customer or the science. These are harder to collect than log-ins because they require agreeing what good output looks like before you measure it. That difficulty is a feature. It forces the conversation that adoption metrics let managers avoid, and it produces a number that cannot be inflated by opening an application.
Then assess judgement explicitly, because judgement is what compounds. The most useful review question in an AI-heavy role is a worked example: describe a case where the model got it wrong, how you caught it, and what you changed. That question is close to ungameable — you cannot fabricate a plausible answer without having genuinely worked through the material, and it surfaces exactly the discrimination you are trying to develop. It also flips the incentive. Instead of rewarding trust in the tool, it rewards calibrated scepticism, which is the trait that keeps a regulated organisation out of trouble.
Finally, reward teaching and keep accountability with the human. The person who lifts a whole team’s capability creates more enterprise value than the heaviest individual user, yet adoption dashboards make them invisible; a criterion that credits enablement corrects that distortion. And every review should credit the decision, not the draft — the professional remains accountable for the output regardless of what produced the first version. Written that way, the expectation becomes legible to staff: nobody has to guess whether performative usage counts, because it visibly does not, and the ambiguity that makes people hedge disappears from the process.
- Outcomes over activity — cycle time, quality, error rates, customer or scientific impact.
- Judgement, evidenced — a worked example of catching and correcting a model error.
- Enablement — who raised the team’s capability, not who logged in most.
- Human accountability — credit the decision, not the draft.
In regulated industries, a blanket usage mandate creates real exposure
In pharma, MedTech and financial services, “use AI more or lose your promotion” is not a neutral instruction — it is pressure applied to people whose work sits inside binding frameworks. Staff operate under regimes set by Swissmedic, the European Medicines Agency and the FDA; financial communication answers to FINMA and its European equivalents; personal data falls under the revised Swiss Federal Act on Data Protection and the GDPR; and specific AI applications carry obligations under the EU AI Act, which treats employment-related systems as high-risk. None of these frameworks care that a usage target existed.
The failure mode is predictable and specific. An adoption KPI creates quiet pressure to route work through whatever tool is nearest when the approved one is slow, restricted, or badly suited to the task — pasting adverse-event narratives, patient-identifiable data, unpublished trial results or draft promotional copy into an unvetted consumer assistant. Each instance looks trivial in the moment. In aggregate it is uncontrolled disclosure, unvalidated content entering a regulated process, and an audit trail that cannot demonstrate human oversight. Regulators do not assess intent; they assess whether the control operated. A metric that rewarded volume while the control was ambiguous is an aggravating fact, not a mitigating one.
The fix is not to exempt regulated teams from AI expectations, which would leave them structurally behind. It is to write the expectation with the precision the environment demands: approved tools, for approved use cases, with documented human accountability and a record of who reviewed what. Measure the outcome and the governance together, and the tension largely dissolves — staff know exactly what counts, the compliance function can evidence oversight, and the organisation still gets the productivity it invested in. It requires more design work than switching on a usage report, which is precisely why most organisations have not done it yet.
Bring outcomes to your review, not a usage dashboard
If you are the one being assessed, the strategic error is to play the log-in game, because activity is the one dimension on which you are most easily replaced and most easily outscored by someone with more time than judgement. Build a short, concrete record instead: a task whose cycle time you compressed and by roughly how much, a process you redesigned rather than merely accelerated, and at least one case where you overrode the model and were right to. Three specific examples, with numbers where you have them, will outperform any usage statistic in a review conversation — and unlike a dashboard, they travel with you to your next employer.
Be clear-eyed about direction of travel, though. AI fluency is becoming table stakes in regulated industries rather than a differentiator, and the professionals who will be scarce in three years are not the ones who used the tools earliest — they are the ones who can tell when the output is wrong and defend the call to an auditor. That capability is built deliberately: on real work, under review, with someone senior enough to correct you. It is worth asking whether your current role is developing it or quietly eroding it.
Which leaves the harder question for whoever owns performance management. AI savviness in reviews is real and spreading, and it will not reverse. Measured as adoption, it produces theatre, degrades output quality and, in a regulated setting, manufactures exposure that surfaces at inspection. Measured as outcomes and judgement, it identifies the people who will carry the organisation through the next several years — and it does so precisely because it cannot be gamed by opening an application on a Monday morning. The metric you choose this cycle is a statement about which of those two organisations you intend to become.
Edward Galle sources and develops talent for regulated industries — professionals who direct AI rather than defer to it. Employers rethinking how they assess AI capability can talk to our team or see our services for companies. Professionals can submit a CV or sharpen their judgement through our HBR-backed training journeys.
References
- Fortune. Accenture warns senior staff to use AI or don’t get promoted. February 2026. https://fortune.com/2026/02/23/last-year-accenture-trained-550000-staff-use-ai-now-promotions-hinge-on-putting-that-into-practice/
- Metaintro. Companies now track employees’ AI usage in performance reviews. 2026. https://www.metaintro.com/blog/companies-track-employees-ai-usage-performance-reviews
- Collaborating with AI Agents: Field Experiments on Teamwork, Productivity, and Performance. arXiv. https://arxiv.org/abs/2503.18238
- European Commission. Regulatory framework for AI (the EU AI Act). https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai
- Swiss Financial Market Supervisory Authority (FINMA). https://www.finma.ch/en/