Back to Blog

Why AI Competency on the Review Form Fails

The 2026 artifact is the review form, not the workshop. HR added an AI fluency row with a 1-5 scale and no behavioral anchors, so managers rate from gut and the competency is theater.

B

Boon

Author

August 21, 2026

Published

AI competency on a performance review is a rating of whether someone used AI well on a real job. In 2026 many organizations added a review-form row labeled AI fluency, digital leadership, or AI adoption. The row is usually a 1-5 scale with no behavioral anchors: what meets looks like on the work, who observed it, and over what window. Managers then rate from gut. L&D cannot prove anyone grew. The competency is theater.

This is the review-form query. Why AI adoption fails already covered the stall. Managers are the missing owner already covered who the team copies. What happens between AI training sessions is the days after the workshop. How to practice a difficult conversation is the rehearsal. This post stays on the row HR added and the rubric nobody wrote.

Why Did HR Add an AI Competency Row Without a Rubric?

Because the form was the fastest way to look serious.

A workshop is an event. A license count is IT's number. A competency row is People language. It lives in the cycle everyone already runs. Adding the row is a one-line change in the HRIS. Writing what "meets" means on a real job is the work nobody scheduled.

On December 8, 2025, SHRM put the question in public. Andy Biladeau, SHRM's chief transformation officer, called AI enablement "the core skill that virtually all professionals need going forward." The same piece said the riskiest stance is choosing no stance at all. Adding the row is a stance. Leaving it blank underneath is the no-stance version of a stance.

That blank is not a missing adjective. It is a missing observer, a missing job, and a missing window. The form asks for a number. It does not ask what work the number is about. So the manager invents a private meaning in the hour they fill the form, and that invention travels into calibration, bonus, and promo conversations as if it were shared.

The skip felt reasonable. The tool was bought. The session had run. The board had asked what People was doing about AI. A labeled row answers the slide. A behavioral rubric answers the job. Most teams shipped the first one and called the cycle updated.

How HR leads AI transformation already argued for building the manager layer first. The form is where that layer either gets a standard or gets a blank.

What Happens When Managers Rate AI Fluency From Gut?

They do the only thing the form allows. They guess, then they defend the guess.

The cycle opens. A manager has twelve reviews and a Friday deadline. The AI row has a label and five radio buttons. There is no example of a 2 versus a 3 on this team's actual work, no prompt for which deliverable they saw, and no window other than "this cycle." They click 3. Sometimes they click 4 for the person who talks about AI in Slack. None of those clicks is evidence.

The dangerous rating is not "needs improvement." It is "meets."

Needs improvement at least admits a gap. Meets is how the company files the competency as handled. The person hears they are fine. L&D hears the cohort is on track. Calibration spends its time on the 2s and the 5s, and treats a stack of 3s as the quiet middle. That middle is where the empty rubric did the most work. It turned an unobserved job into a passing score.

On the formWhat a real rating needs
Label: AI fluency, digital leadership, or AI adoptionA job they already own
1-5 with no descriptorsObservable anchors for 1, 3, and 5
Optional comment boxWho saw the work, and which deliverable
Window: this review cycleThe weeks the person ran the tool on live work

The left column is what most 2026 forms shipped. The right column is the minimum for a number you can defend. If you cannot fill the right column, you do not have a competency. You have a vibe with a dropdown.

That is why AI adoption metrics for HR cannot be a login chart pasted into the review. Usage answers "did they open it." A review rating is supposed to answer "did the work get better." A system can count opens. It cannot tell you whether the person caught a bad draft.

Gut ratings also teach the team that the AI row is political. A private experiment can outscore a quieter person who changed a workflow. The form still looks official. The scores are now reputation.

How Do You Measure AI Competency on a Performance Review?

You write the rubric the form skipped. Then you rate the job, not the mood.

A usable rating needs four things on the page, not in a side deck only L&D has seen.

A named job. "Uses AI" is not ratable. "First draft of the weekly forecast starts in the tool, and you edit" is ratable. How to build an AI adoption strategy starts with a business problem for that reason. The review needs the same specificity or the manager will rate personality.

Behavioral anchors at 1, 3, and 5. Write what a person on this job would be seen doing. A 1 is the job still starting the old way, with AI as a private experiment if it appears at all. A 3 is one live-job use case in the actual workflow, with a before-state they can say in a sentence and a miss they caught. A 5 is the new path as the team's default on that job, plus a second workflow with a before-state. Exceeds is not "talks about AI." Exceeds is a changed week other people can point to.

An observer. The rater has to have seen the work, or have a written artifact from someone who did. A license report is not an observer. If the manager never watched the job, the honest rating is "unable to rate," not a polite 3. What management coaching is is the mechanic that gets a manager close enough to the work to have something to say.

A window that matches the work. Rate the weeks the person actually had the tool on a live task. Do not treat attendance as time on the job. A workshop is not the same as ongoing growth. Training transfers language in a room. A review is supposed to score what happened after.

Three sentences will carry most reviews if the anchors exist: the job I watched, what meets looks like on that job, and what I saw in this window. If a manager cannot say those, they are not ready to click a number. The fix is a standard, and time with the work, not a longer comment box.

If the review conversation is the hard 1:1, rehearse it once before it is live. That is optional. Practice Space is rehearsal, not a second product. The required work is still the rubric and the observed job.

The enterprise AI rollout checklist is the assignable version of the human-layer work. Use it to name the workflow. Use the review to score whether that workflow moved. In programs Boon has run since 2023, competency scores improve 23 percent on average through coaching. That line only means something because the competency had anchors and a coach who saw the work. A 1-5 with no descriptors cannot move 23 percent. It can only wander. AI transformation coaching is the conversation that sits with the clumsy phase so the next cycle has something to rate besides gut.

Why Is an Empty AI Row a People Problem, Not an IT Training Problem?

Because the form is a people system. The tool already shipped.

IT can stand up access and a usage export. It cannot write what "meets" looks like on a forecast, a ticket, or a customer note. It cannot sit with the manager who will not look unfinished in front of the people they lead. Those are People and L&D jobs. When they go unfinished, the failure shows up as a review number nobody can explain.

Calling it a training gap is how the empty row survives another cycle. Another recorded session will not write the anchors. The person who still starts the job the old way does not have a knowledge problem. They have a standard that was never made visible, and a rater who was asked to invent one in a form.

That is why this belongs next to the review calendar, not next to the launch email. Coaching in that window is how a manager gets close enough to the work to rate it, and how a person gets a chance to grow the behavior before the number is filed. Sessions stay on Zoom. Slack and Teams carry the prep and the follow-through. The form is the artifact. The coaching is what makes it honest. Attendance is not a rubric. A 3 is not proof.

FAQ

How do you measure AI competency on a performance review?

Measure it against a named job, with behavioral anchors for 1, 3, and 5, an observer who saw the work, and a window that matches when the person used the tool on a live task. A 1-5 with no descriptors is a guess the form required, not a measure.

What is AI fluency on a review form if there is no rubric?

It is a label. The manager still has to click a number, so they rate from gut. That number then travels into calibration as if it were shared. Without anchors, AI fluency is theater.

Why is an AI competency row without behavioral anchors a problem?

Because people get a score that will follow them into bonus and promo conversations, and nobody can say what the score meant on the job. "Meets" becomes the default. L&D cannot prove growth. The company has performed a decision it has not made.

Who should observe AI competency, and over what window?

The manager who saw the live-job use case, or someone who left a written artifact of that work. Rate the weeks the person ran the tool on a real deliverable. A license report and a workshop recap are not observation.

Is rating AI usage the same as rating AI competency?

No. Usage answers whether someone opened the tool. Competency answers whether they used it well on a job: whether they caught a miss, ignored a wrong answer, and changed the path the work takes. Counting opens and calling it a review score is how tourism hides inside a passing rating.

How does coaching sit next to the review cycle?

Coaching is the calendar that gets a manager close enough to the work to rate it, and gives the person time to grow the behavior before the number is filed. It is not another workshop. It is the conversation about the job they tried, where it got clumsy, and what "meets" will look like on the form.

The Form Is Not a Rubric

If the cycle opened and the AI row is a 1-5 with no anchors, you do not have a mysterious measurement problem. You have a people-development gap that the form made official. Write what meets looks like on a real job. Require an observer and a window. Then put coaching next to the cycle so the next rating is about work someone saw.

Boon Adapt is the coaching calendar that sits next to that cycle. It sits with SCALE, GROW, EXEC, and TOGETHER as one operating system for people development that lives in Slack, Teams, and MCP, and gets measured. Write the rubric. Then coach the job the form is supposed to score.

Newsletter

Get more like this

Leadership insights, coaching research, and practical frameworks delivered to your inbox.

Ready to transform your leadership development?

Discover how Boon can help your organization build resilient, effective leaders at every level.