AI Performance Reviews for Managers: Do It Right (2026)

94% of managers now use AI to draft performance reviews, but 66% of employees trust it less than a human. Here's the bias-checked, privacy-safe workflow that closes that gap.

Fall review season is starting, and if you’re a manager, there’s a decent chance you’ve already opened ChatGPT to help write one. You’re not alone — a June 2026 survey of more than 1,300 managers found 94% now use AI to build employee development plans, and 91% use it for assessments. But there’s a catch nobody’s telling you: a global survey of 1,365 employees found 66% trust AI-assisted performance management less than a human-led process, and 72% suspect their manager is already using AI to write their review, whether they’ve been told or not.

That gap — between how fast managers adopted the tool and how little employees trust it — is where most of the damage happens. Not from using AI. From using it badly, quietly, or as a replacement for actually knowing how your people are doing. This is the walkthrough for doing it the other way: the workflow that saves you the two-to-four hours per review that writing from scratch takes, while producing something that reads like you, not like a chatbot, and that survives an employee asking “did you actually write this?”

What’s actually happening

Performance reviews were already broken before AI showed up. SHRM’s research puts real numbers on that: 61% of HR professionals say fewer than half their managers “effectively address underperformance or areas for improvement,” 43% say managers simply aren’t trained to conduct effective reviews, and 60% say managers lack the data-driven insight to back up what they’re writing. Only 14% of employees strongly agree their reviews actually inspire them to improve. Only 29% think their reviews are fair. Only 26% think they’re accurate.

That’s the opening AI walked into — not because performance management needed a chatbot, but because most managers were flying blind, under-trained, and out of time. SHRM’s own framing for 2026 is that the “personalized AI coach” is becoming the tool that helps managers actually do what the traditional annual review never delivered: consistent, specific, behavior-based feedback instead of a rushed paragraph written the night before it’s due.

OpenAI Academy’s guide to ChatGPT for managers, covering feedback and performance cycles Source: OpenAI Academy — ChatGPT for Managers

Here’s how fast this moved. A Reworked survey of 1,300+ managers, published June 2026, found:

What managers use AI forShare doing it
Building employee development plans94%
Drafting performance assessments91%
Related performance-management tasks88%

Separately, UK-focused research from Visier found that of managers who’d used generative AI in their role, 59% used it specifically for performance reviews and feedback. And SHRM’s own Talent Trends data shows HR-department AI use in at least one function jumped from 26% in 2024 to 43% in 2025 — the fastest single-year jump recorded for any HR technology category.

So the practice is mainstream. What’s not mainstream — and this is the part that should make you pause before you paste your notes into ChatGPT — is any kind of shared standard for doing it responsibly. Fewer than half of organizations have a mature AI governance framework for this. SHRM’s own analyst called that governance gap “perhaps the most concerning finding” in their 2026 report. You are, in practice, on your own to figure out the right way to do this. That’s what the rest of this post is for.

Why “just paste your notes in” is the wrong instinct

The failure mode isn’t using AI. It’s using AI the way most managers default to: dumping raw notes in, accepting whatever comes back, and moving on. Three things go wrong when you do that, and all three show up in the research.

It scales bias instead of removing it. Bloomberg Law’s December 2025 reporting quotes employment attorneys warning that AI trained on historical review language can absorb and then “replicate more consistently” any gendered or racially uneven scoring patterns already present in that data — as one attorney put it, “AI may or may not create bias on its own, but either way it can scale it significantly.” A 2026 HR statistics roundup citing IBM data found 30% of AI HR deployments have already surfaced bias issues in practice. This isn’t hypothetical. It’s the single most-cited risk across every credible HR and legal source that’s written about this.

It creates real legal exposure. Under the EU AI Act, systems used to evaluate or classify workers for performance purposes are explicitly “high-risk” (Annex III), which triggers audit and transparency obligations. In the US, the EEOC treats algorithmic tools that inform employment decisions as “selection procedures” subject to disparate-impact analysis under Title VII — regardless of whether the AI made the final call or just informed it. “The algorithm scored them lower, not me” is not a defense that holds up. And attorneys note something a lot of managers haven’t thought about: the AI-generated draft, and the prompt you used to create it, can become discoverable evidence if a review is ever challenged in litigation. Opposing counsel can ask what you told the AI about the employee.

It erodes trust faster than it saves time — especially undisclosed. A Kickresume survey of 1,365 employees found 72% suspect their manager uses AI to generate their review, with suspicion highest in Asia (82%), Europe (68%), and the US (67%). Among people who suspect AI involvement, 30% say the review feels generic or overly polished, and 18% believe it’s mostly or entirely AI-written. Pew Research’s March 2026 workplace refresh found 43% of US workers trust a colleague’s output less once they learn AI was involved, versus only 20% who trust it more. And here’s the twist that should change how you think about disclosure: research cited in a March 2026 LinkedIn analysis found AI-generated feedback improved employee performance by roughly 15% in the underlying study — but “the moment you disclose the source, the benefit disappears.” Silence isn’t a fix. It’s a delayed cost. AI HR Daily’s March 2026 reporting on HR Executive research found 58% of managers say AI use is becoming an “unspoken performance requirement” at work — but only 29% of employees agree that’s happening. That’s a massive perception gap, and it’s exactly the kind of thing that surfaces later as a disputed review or a grievance filing.

None of that means don’t use AI. It means the way you use it is the entire ballgame — and it’s exactly the distinction some HR-tech commentators miss when they argue for avoiding AI review-writing entirely.

Textio’s blog post arguing that using ChatGPT to write performance reviews is a no-go Source: Textio — Why using ChatGPT to write performance reviews is a no-go

Here’s the workflow that keeps you on the right side of all three problems.

The workflow: from rough notes to a fair, specific review draft

This is the five-step version. Steps 1 through 3 take about 10 minutes with a real employee’s worth of notes. Steps 4 and 5 — the parts most people skip — take another 10, and they’re the ones that actually protect you and your employee.

Step 1: De-identify before you paste anything

Before you type a single word into ChatGPT, Claude, or Gemini, strip out anything that identifies the employee. Public consumer AI tools log conversations — that’s not a rumor, it’s how the free and Plus tiers work by default — so the uniform guidance from HR and legal sources is to anonymize (write “Employee A” instead of a real name) or restrict this kind of content to a company-approved, privacy-safeguarded AI tool if your employer provides one.

What to strip: the employee’s name, any client or project names that could identify them by association, exact dates that map to a specific incident someone could trace back, and anything about medical leave, disability accommodations, or protected-characteristic context (age, pregnancy, religion, etc.) that has no bearing on performance.

Expected result: a set of plain-text bullet notes that reads like performance data, not a personnel file.

Step 2: Turn behavior bullets into a first draft

This is the actual prompt. Notice what it does — and doesn’t — ask the AI to do.

I'm writing a mid-year performance review for a team member. Below are
rough notes on their work this period. Turn these into a first-draft
review section using specific, behavior-based language — describe what
they did and the outcome, not vague traits like "great attitude" or
"team player." Do not invent any detail, metric, or example that isn't
in my notes. Flag anywhere you think I haven't given you enough
information to write a specific sentence, instead of filling the gap
yourself.

Notes:
- Hit the Q2 revenue target for their territory (up from missing it in Q1)
- Missed two internal deadlines on the CRM migration project
- Spent ~4 hours over two weeks helping onboard the new hire, unprompted
- Client feedback on the March renewal call was specifically positive
  about their preparation
- Hasn't raised any blockers proactively in 1:1s — I've had to ask

Expected result: a draft that reads something like — “In Q2, [Employee] turned around underperformance from Q1 to hit their territory revenue target. They took the initiative to spend roughly four hours over two weeks mentoring the new hire, without being asked, which the new hire specifically credited in their own ramp-up feedback. On the March renewal call, the client noted their preparation was a standout factor in the close. Two internal deadlines on the CRM migration were missed this quarter, which is an area to address directly. One growth area: [Employee] doesn’t proactively surface blockers in 1:1s — I’ve had to ask rather than being told.”

That’s specific, defensible, and traceable to a real note. Notice the prompt explicitly forbids inventing detail — that instruction matters more than almost anything else in this workflow, because generative tools will confidently fabricate a plausible-sounding metric or quote if you don’t tell them not to. Multiple HR sources flag exactly this: verify every named detail against actual evidence before it goes in the final version.

Step 3: Add the “areas to grow” language — honestly

Ask the AI to draft the growth-area language separately, and be explicit that you want honest, not softened, phrasing:

Now write the "areas to grow" section for the same employee, based only
on the blocker-related note above. Be honest and specific about the gap
— don't soften it into something vague like "could communicate more."
Pair it with one concrete, actionable suggestion for how they could
change it, and end with a forward-looking, non-punitive tone.

Expected result: something like — “A clear area to grow: proactively surfacing blockers before they become deadline risks, rather than waiting to be asked in 1:1s. A concrete next step: bring one blocker or open question to each 1:1, even a minor one — this builds the habit and gives us a standing checkpoint instead of a reactive one. This is coachable and doesn’t reflect on the strong client and revenue results above.”

Step 4: Run the 3-minute bias-language check

This is the step almost nobody in the “no-go” articles or vendor blogs writes about, and it’s the actual differentiator. Take your near-final draft and run it back through the AI with this prompt:

Review this performance review draft for loaded, gendered, or vague
language that could read as biased  words like "abrasive,"
"aggressive," "emotional," "not a team player," or praise that's vaguer
for this person than it would be for someone else doing the same work.
Flag each instance and suggest a specific, behavior-based replacement.

Loaded language in reviews skews in well-documented, predictable directions — women are more likely to be described as “abrasive” for the same assertiveness that gets men called “confident”; older employees get “resistant to change” for questions younger employees get called “engaged” for asking. An AI pass won’t catch everything a trained DEI reviewer would, but it catches the obvious patterns fast, and running it takes about three minutes. This is also your best documented defense if a review is ever challenged — legal sources describe documented, consistent review practices applied the same way across employees as the strongest available defense against a disparate-impact claim.

Step 5: Rewrite it in your own voice, then decide on disclosure

Read the draft out loud. If it doesn’t sound like something you’d actually say to this person’s face, it’s not done. Cut the corporate throat-clearing AI tends to add (“It is important to note that…”), swap in phrases you actually use, and add the one or two details only you would know — a specific moment from a meeting, a callback to something they said matters to them. This is also where you decide, deliberately, whether and how to tell the employee AI was involved in drafting. Given that 72% of employees already suspect it whether you disclose or not, and that undisclosed AI use is exactly what shows up later as “this feels generic” — the safer long-run play, per SHRM and multiple HR-legal sources, is transparency framed as a process improvement (“I used AI to help organize my notes into a first draft, then edited and finalized it myself”) rather than a liability to hide.

Expected result: a review that took you roughly 15-20 minutes total instead of 2-4 hours, reads like you wrote it because you did the parts that matter, and has already been checked for the two things most likely to get a review challenged — invented detail and biased language.

Lattice’s published ChatGPT prompts for performance reviews, by role and goal Source: Lattice — 15 ChatGPT Prompts for Performance Reviews

AI-assisted vs. AI-outsourced vs. from-scratch: what you’re actually choosing between

Writing from scratchAI-outsourced (paste and accept)AI-assisted (this workflow)
Time per review2-4 hours5-10 minutes15-20 minutes
Risk of invented/fabricated detailNone (but often vague from fatigue)High — AI will fill gaps with plausible-sounding fictionLow — explicitly instructed not to invent, human verifies every detail
Bias riskDepends entirely on the manager, uncheckedHigh — no bias pass, and AI can amplify patterns in training dataChecked — dedicated bias-language pass built into the workflow
Reads as genuinely yoursYes, but often inconsistent qualityNo — 30% of employees who suspect AI say it “feels generic”Yes — final rewrite pass is mandatory
Legal defensibilityWeak if inconsistent across employeesWeak — no documented human judgment layerStrong — documented process + human review is the standard legal defense
Employee trust if disclosedHigh baseline, if review is goodLow — disclosure “erases” any performance benefit research foundHigher — framed as a process tool, not a replacement for you

The middle column is where most of the bad press about “AI performance reviews” comes from, and it’s genuinely different from the right column in outcome, not just intention. The data above — the fabricated details, the bias risk, the trust collapse — is what happens in the outsourced pattern. The assisted pattern, where AI drafts and a human verifies, edits, and takes ownership, is the one every credible HR and legal source actually recommends, including SHRM’s own official guidance and employment-law analysis from firms like Fisher Phillips.

What this means for you

If you’re a first-time manager writing your first review cycle: Start here. You don’t have bad habits to unlearn yet. Use the five-step workflow exactly as written, and don’t skip Step 4 — the bias check is the single highest-leverage 3 minutes in this whole process, and it’s the thing that’s easiest to skip when you’re rushed.

If you’re an experienced manager who’s already been pasting notes into ChatGPT: Your reviews are probably fine on substance and weak on two specific things — bias language and specificity. Add Step 4 to whatever you’re already doing this week. It’s the fastest fix with the highest downside protection.

If you manage a team of 8+ and reviews eat a full week: The time math changes the equation for you specifically. At 2-4 hours each, eight reviews is a lost work-week. At 15-20 minutes with this workflow, it’s under three hours total. Block one afternoon, do all eight in sequence — the de-identify and bias-check steps get faster with repetition.

If you’re in a heavily regulated industry (finance, healthcare, government contracting): Do not use a public consumer AI tool for this at all if your employer offers an enterprise or company-approved AI tool with a data-processing agreement. If they don’t, push to de-identify more aggressively than the baseline above — strip department names and role titles too, not just the employee’s name, since re-identification risk is higher in smaller regulated teams.

If you’re worried about the EEOC/Title VII or EU AI Act exposure specifically: Document your process, not just your output. Keep a simple internal note (even a one-line log) that you ran the bias-language check and applied the same review process to every direct report this cycle. Consistency of process, applied identically across your team, is the actual legal shield — not avoiding AI, and not hiding that you used it.

If you’re writing your own self-review (the other half of review season): Use the same de-identify-and-verify discipline in reverse — feed AI your actual accomplishments with real numbers, ask it to structure them using the STAR method (Situation, Task, Action, Result), then run your own “does this sound like me” pass. The goal is the same: AI organizes, you own the final words.

If you’re an HR leader setting policy for your managers: The data says you’re behind, not ahead — governance frameworks are lagging manager adoption almost everywhere. The cheapest fix available to you this week is publishing a one-page internal guide covering exactly the five steps above, plus your company’s stance on AI tool selection (consumer vs. enterprise) and disclosure expectations. SHRM’s own guidance is to pilot transparently with a volunteer group before mandating anything company-wide.

If you’re the employee on the receiving end of a review, wondering if AI wrote it: The tell isn’t the polish — it’s the specificity. A review with real dates, real project names, and a detail only your manager would know is a good sign regardless of what tool helped structure it. A review that’s all adjectives and no concrete example is worth a direct, calm question in your 1:1: “Can you walk me through the specific examples behind this?”

Edge cases and troubleshooting

“The AI keeps adding metrics or specifics I never gave it.” This is the single most common failure, and it’s exactly why Step 2’s prompt explicitly forbids invention. If it happens anyway, add “If you’re inferring rather than working from a note I gave you, say so explicitly instead of stating it as fact” to your prompt, and always do a line-by-line check against your original notes before finalizing — treat every number and named example as unverified until you’ve confirmed it’s yours.

“My employee straight-up asked if I used AI to write their review.” Answer honestly. The research is consistent: undisclosed AI use, once discovered, does more trust damage than disclosed use from the start. A good answer: “I used it to help organize my notes into a first draft — the examples, the judgment, and the final wording are mine.” That’s true if you followed this workflow, and it’s the framing SHRM and multiple HR sources recommend.

“HR hasn’t given us any guidance on which AI tool to use, or whether we’re allowed to at all.” Don’t assume silence means permission. Ask directly before your review cycle starts — specifically whether the company has an approved tool with a data-processing agreement, since that changes what’s safe to paste in. If there’s truly no policy, default to the most conservative version of de-identification in this guide.

“I manage in the EU and I’ve heard performance-evaluation AI is ‘high-risk’ now — am I breaking the law by using ChatGPT to help draft a review?” Using a general-purpose AI assistant to help you draft text isn’t the same as deploying an automated system that scores or classifies employees without human judgment — the EU AI Act’s high-risk classification targets systems making or materially informing the employment decision itself. A human-drafted, human-reviewed, human-decided review where AI helped with wording is a different category. If your company does use a dedicated AI performance-scoring product (not a chat assistant), that’s the case where the high-risk rules squarely apply, and that’s a conversation for your legal/compliance team, not something to guess at.

“A teammate told me they got the exact same phrase in their review as I got in mine — did our manager just copy-paste an AI template?” This is the “feels generic” failure mode showing up in the wild, and it usually means Step 5 (rewrite in your own voice) got skipped. If you’re the manager, this is your sign to slow down on that step specifically — generic praise phrases (“consistently exceeds expectations,” “strong team player”) are the fastest way to make two different people’s reviews sound interchangeable.

“I don’t have enough written documentation from throughout the year to give AI good notes — I’m basically starting from memory in October.” AI can’t fix this, and it’s not supposed to try. If your notes are thin, that’s a signal to build a lighter habit for next cycle (a two-line note after any notable interaction, logged wherever’s easiest), not to let AI invent detail to paper over the gap. A review built on real, if sparse, notes beats one built on confident-sounding fabrication every time.

“Legal/compliance told us AI drafts could be discoverable — does that mean I shouldn’t save my prompts?” The opposite, generally — a documented, consistent process (including that you ran a bias check) is your defense, not your liability. What you want to avoid isn’t documentation; it’s a prompt or draft that reveals judgment you wouldn’t stand behind if read back to you in a dispute. Write prompts the same way you’d write an email you’d be comfortable forwarding.

“My company blocks ChatGPT/Claude/Gemini entirely on the work network — what do I do?” Respect the block; it usually exists because there’s no data-processing agreement in place, which is exactly the privacy risk this guide is trying to help you avoid. Use the same five-step structure with a plain text editor and your own judgment doing the drafting — the workflow (de-identify, behavior-based language, honest growth areas, bias self-check, rewrite in your voice) works with or without AI in the loop; AI just makes step 2 faster.

What this can’t fix

It can’t replace not knowing how your people are doing. If you’re reaching for AI because you genuinely don’t have enough information about someone’s performance, that’s a management gap, not a writing problem — AI can organize thin notes, but it can’t invent the observation you never made. The fix is building a lighter habit of logging notes through the year, not writing a better prompt in October.

It can’t make an unfair rating fair. If the underlying assessment is wrong — you’re rating someone lower than their actual work because of a personality clash, or higher because you like them — AI-polished language doesn’t fix that. It just makes the wrong conclusion read more convincingly. The bias-language check catches loaded words, not a wrong underlying judgment.

It can’t substitute for the actual conversation. A well-drafted written review that gets read aloud in a rushed 10-minute meeting with no room for questions accomplishes less than a mediocre written review paired with a real, unhurried conversation. The document is a supporting artifact, not the point.

It can’t guarantee legal safety on its own. Following this workflow meaningfully reduces your risk, but the actual legal shield — documented, consistent process applied identically across your team — is something you build through practice over a review cycle, not something any single tool or prompt gives you automatically.

It can’t read your organization’s specific compensation, promotion, or PIP (performance improvement plan) rules. Never let AI draft language that touches pay decisions, promotion recommendations, or formal disciplinary action without your HR partner reviewing it first — those categories carry legal weight the drafting workflow above isn’t designed to handle, and multiple sources are explicit that ratings, pay, and termination decisions must remain fully human calls.

FAQ

Is it okay to use ChatGPT to write performance reviews? Yes, with two conditions that matter more than the tool choice itself: de-identify what you paste in, and treat the output as a first draft you verify and rewrite, not a finished product you accept as-is. The research consistently distinguishes “AI-assisted” (human verifies and owns the final version) from “AI-outsourced” (paste and accept) — the first is fine and increasingly standard; the second is where the bias, fabrication, and trust risks concentrate.

Will my employees be able to tell if AI helped write their review? Maybe, and it matters less than you’d think if you followed the full workflow. The “tell” employees actually notice is genericness — vague praise, interchangeable phrasing, no specific example only you would know. A review with real, specific detail reads as yours regardless of what helped you organize the notes.

Should I tell my employees I used AI to help draft their review? The research leans toward yes, framed carefully. Roughly 70% of employees say they want to know when AI is used in HR decisions that affect them, and undisclosed use — once discovered, and 72% already suspect it — does more trust damage than disclosed use. Frame it as a process tool (“I used AI to organize my notes into a first draft, then wrote and finalized it myself”), not as the thing that wrote their review.

What information should I never paste into a public AI tool for a review? The employee’s name, anything tying them to a specific traceable incident, medical or disability-related information, and anything about a protected characteristic (age, religion, pregnancy, etc.) that isn’t directly relevant to a documented performance issue. If your company has an approved enterprise AI tool with a data-processing agreement, the bar is lower — check with IT/legal before assuming.

Can AI actually detect bias in my writing, or is the “bias check” step just theater? It catches real, well-documented patterns — loaded words that skew by gender or age (“abrasive” vs. “confident” for the same behavior, “resistant to change” vs. “engaged” for the same question), and vaguer praise for some employees than others doing comparable work. It won’t catch everything a trained reviewer would, but a 3-minute pass that catches the obvious cases is meaningfully better than no check at all, and it’s fast enough that there’s no real excuse to skip it.

What’s the actual legal risk if I use AI to write reviews badly? In the US, algorithmic tools that inform employment decisions can be treated as “selection procedures” under Title VII, subject to disparate-impact analysis, regardless of whether AI made the final call or just informed it. In the EU, systems used to evaluate worker performance are classified “high-risk” under the AI Act if they’re making or materially driving the decision. And in any jurisdiction, AI-generated drafts and the prompts behind them can become discoverable evidence if a review is ever formally disputed. None of this bans using AI to help draft — it raises the bar for keeping a human clearly, demonstrably in charge of the judgment.

Does using AI actually save real time, or is verifying everything just as slow as writing from scratch? Real time saved, based on the workflow above: roughly 15-20 minutes per review versus 2-4 hours writing from scratch — and that gap holds up even with a genuine verification and bias-check pass built in, because AI is doing the slowest part (turning bullet notes into full sentences), not the fast part (deciding what’s true and whether it’s fair).

My company hasn’t given us any AI policy for reviews — am I allowed to just use my personal ChatGPT account? Ask before assuming. Silence from HR isn’t the same as permission, especially given how far governance has lagged adoption industry-wide (fewer than half of organizations have a mature AI framework for this). If there’s genuinely no policy, treat every note as if it’ll be seen by someone outside your team, and de-identify accordingly.

What about using AI for a Performance Improvement Plan (PIP) instead of a regular review? Be more conservative, not less. PIPs carry direct legal and employment-termination weight, and every source in this space is explicit that AI should draft language for you to review, not generate the plan’s substance or the decision behind it. Loop in your HR or legal partner before finalizing a PIP, regardless of how the draft got written.

The bottom line

The manager who’s already using AI to write reviews isn’t the problem. The manager who pastes raw notes in and ships whatever comes back without reading it critically — that’s where fabricated detail, amplified bias, and the “this feels generic” trust collapse all live. The fix isn’t avoiding the tool. It’s five specific steps that take maybe 20 minutes total: de-identify, draft from real notes with fabrication explicitly forbidden, write growth areas honestly, run the bias check, and rewrite it in your own voice before it goes out. Do that consistently across your team, and you’ve got a documented process that’s faster than writing from scratch, more consistent than doing it purely from memory, and genuinely more defensible than either.

If you want the complete workflow — de-identify, draft, growth areas, bias check, voice rewrite, plus self-reviews and PIP caution — FindSkill’s AI for Managers: Fair, Fast Performance Reviews course walks through all five steps hands-on in eight short lessons, with the first two free.

Sources

Build Real AI Skills

Step-by-step courses with quizzes and certificates for your resume