PERFORMANCE MANAGEMENT
Performance management, rebuilt from the practitioner’s chair: the 4-stage cycle, 6 components, a calibration agenda, and a 90-day rollout roadmap.
I spent 8 years building talent systems inside engineering and manufacturing functions that hired 250 people a year. In that time I watched the performance cycle run perfectly and change nothing. Forms in, ratings out, business as usual.
You know the shape of it. Review season lands, you chase managers for submissions, and the calibration meeting calibrates nothing. Ratings get published, merit letters go out, and by February nobody remembers what was agreed.
Most organizations respond to that by redesigning the form. Wrong repair — the form was never the constraint. Manager capability was.
The cost of getting this wrong does not appear as a failed process. It appears as a regrettable exit nobody saw coming.
$500 a day
Operational drain from one unfilled senior engineering role. A 60-day vacancy you could have prevented is $30,000 of lost productivity, before a single agency fee.
This page is the operating model, not the theory. It sits inside the wider talent management system the Lab publishes, and it assumes you already run reviews of some kind.
Here’s what you get:
- The four-stage performance management cycle, with the cadence, owner, and artifact that make each stage real
- The six components of a system that survives contact with actual managers
- A calibration agenda you can run next quarter, step by step
- A 90-day sequence for fixing yours, in the order that works
Every number here is either from my own operating data or a sourced study you can check, which is the standard across the Lab.
What is Performance Management?
Here’s the definition:
Performance management is the continuous, structured system an organization uses to align individual work with business objectives, build capability, and make evidence-based decisions about reward, development, and role fit. It runs all year through goal setting, feedback, evidence capture, and calibration. It is a system, not an event.
Five parts carry it: goal architecture, feedback cadence, an evidence standard, calibration, and a consequence architecture. A sixth, manager capability, decides whether the other five function at all. Each carries its own design rule, set out in the six components of a system that holds up.
Performance management vs performance appraisal
Here’s the difference:
A performance appraisal is a bounded evaluative event. It produces a rating for a defined period, and then it ends. Performance management is the continuous system that appraisal sits inside.
Treat the two as the same thing and your system produces exactly one output: a number. A number with no development pathway attached is a scoreboard, not a management tool.
| Performance appraisal | Performance management | |
|---|---|---|
| Purpose | Evaluate what happened in a defined period | Improve what happens next, all year |
| Frequency | Once, at the end of the cycle | Continuous, with one formal checkpoint a year |
| Owner | The manager, with HR administering the form | The manager runs it, HR maintains the system |
| Output | A rating | A rating, plus a reward, development, and redeployment decision |
| What happens next | The rating feeds the merit matrix | Every rating carries a named action within four weeks |
Earlier in my career, standardizing the assessment format is what moved hiring decisions from impression to evidence. The same mechanism is what separates a real appraisal from a rating exercise.
◆ PRO TIP
Real talk: plenty of organizations discover, three months into a redesign, that they have been running appraisals and calling it performance management. That is not a failure of the redesign. It is the first useful thing the redesign produced.
What performance management is not
It is not a compliance artifact. When the only reason the cycle runs is that legal or the HRIS needs a rating field populated, you are not managing performance. You are maintaining a database.
⚠ WATCH OUT
Common mistake: treating the rating as a compensation mechanism wearing a different name. When the rating exists to feed the merit matrix, managers optimize the number for the budget and the development conversation quietly disappears. Organizations that fuse rating and pay inflate fixed cost, create internal inequities, and lose people anyway once the novelty wears off.
It is not a surveillance function either. Continuous feedback is not continuous monitoring, and conflating the two is how organizations end up with activity dashboards instead of performance systems.
So judge your own system on one test. Not whether the cycle closed on time, but whether managers made better decisions and employees changed behavior. Completion rate is hygiene, not success.
Why Most Performance Management Systems Fail: The Trust Deficit
You already know the symptoms. Managers inflate ratings to avoid a difficult conversation, employees stopped expecting the review to change anything years ago, and the calibration meeting quietly rewards whoever argues hardest.
None of that is a people problem. When a system produces outcomes people cannot predict from their own behavior, they stop reading it as a signal.
A performance system nobody can predict is not a signal. It is weather.
Deloitte put a number on how far that has gone.
of workers, and 61% of managers, could not say they trust their organization’s performance management process.
Translate that into your own building. If seven in ten of your workforce cannot say they trust the process, every output of that process is landing as arbitrary, however carefully you built it. The rating, the merit decision, the promotion case, and the improvement plan all included.
The people who own these systems are no more convinced. Only 2% of Fortune 500 CHROs strongly agree that their performance management system inspires employees to improve.
I have never seen a better form fix any of this. The strongest twelve-month cohort retention I have run, 89%, came from role clarity, a written evidence standard, and managers who could hold the conversation.
So before you redesign anything, work out which of three failures you actually have. The fix for each one is completely different.
Design failure comes first. The goals were never connected to anything the business measures, so the review has nothing real to assess against and the conversation defaults to personality.
Capability failure comes next. The design is sound, but setting a differentiated goal, giving corrective feedback, and defending a rating under challenge are three separate skills, and most managers were trained in none of them.
Consequence failure comes last. The design is sound, the managers are capable, and nothing happens as a result of the rating, so everyone involved correctly reads the whole exercise as theater.
Here’s the tell for each:
- Design: pull 10 goals at random and count how many name the business objective they ladder up to
- Capability: read five completed reviews and count how many cite a dated, specific observation rather than a description of the person
- Consequence: pull last cycle’s ratings and count how many carry a named action logged against them
⚠ WATCH OUT
Anti-pattern: treating a capability failure as a design failure. It is the most common error in this space, and it buys you a better-designed system that your managers still cannot run.
Hiring taught me the same ordering. Quality of Hire moved 12% from competitor analysis, 14% from talent landscaping, and 30% from a skills-based pipeline, because the biggest gain sat in the input architecture rather than the assessment step.
The old way versus the Lab Way
The old way
- Set annual goals in January and never reopen them
- Wait for the year-end window to raise a performance concern
- Ask managers to write reviews from memory
- Hold a calibration meeting with no evidence standard
- Publish ratings and move on
The Lab Way
- Set goals with a mandatory quarterly decision to keep, change, or retire each one
- Document any concern within 14 days of observing it
- Run monthly check-ins against a written evidence log
- Calibrate against pre-read evidence packs with distribution guardrails
- Publish every rating alongside a named development or redeployment action
That second column does not work because it is stricter. It works because every item converts a judgment into a record.
A record can be checked, challenged, and defended in front of the person it describes. Trust follows checkability, not sentiment.
The Performance Management Cycle: Four Stages, Operationalized
Four stages, and you have seen them named a hundred times. What almost nobody specifies is the cadence, the owner, and the artifact each stage produces, which is the only part that decides whether the stage happens at all.
This section is the operating view of the cycle, and the performance management process breaks the same ground down into procedure.
1
Stage 1: Planning and goal setting
Planning is where a manager and an employee agree, in writing, what outcomes the employee is accountable for and what evidence will demonstrate them.
| Cadence | Annually, with a mandatory decision at each quarter boundary to keep, change, or retire every goal |
|---|---|
| Owner | The manager writes, the employee negotiates, the HRBP audits a sample for cascade integrity |
| Artifact | A goal document naming the parent business objective and the evidence source for each goal |
Write every goal so the evidence source is named inside it. “Improve supplier quality” is an intention. “Cut supplier-caused stoppages on the assembly cell from 14 a quarter to under 5, measured from the downtime log” is a goal, because the argument about whether it was met is settled before the period starts.
The M in SMART is where most goals collapse. Managers write something measurable-sounding and never name the source that will measure it.
Skip this stage and the whole cycle becomes retrospective opinion, because there is no agreed standard to assess anything against.
2
Stage 2: Check-ins and continuous monitoring
Check-ins are the recurring conversations where progress is assessed against the agreed evidence, and corrections still cost less than a rating does.
| Cadence | Monthly, 30 minutes, structured. Weekly informal contact is not a substitute and should not be counted as one |
|---|---|
| Owner | The manager runs the conversation, the employee brings the evidence |
| Artifact | A running evidence log, dated entries of three to five lines each |
Managers will tell you they give continuous feedback. Employees will tell you they do not receive it. Both are being honest, because informal encouragement is being counted as feedback and none of it is written down.
The frequency is the variable that moves the outcome, not the quality of the annual conversation. Gallup found that 56% of employees formally review their goals with their manager once a year or less, and that employees with quarterly progress checks are 2.1 times as likely to say the process is fair and transparent.
Source: Gallup, 2% of CHROs Think Their Performance Management System Works, 2024
Skip the documentation and you get the year-end review written from memory, which every employee correctly identifies as recency bias.
◆ PRO TIP
The honest downside: a manager with 15 direct reports cannot run this properly. Monthly structured check-ins across a wide span is a real time cost, which makes span of control a performance management design constraint and not just an org design one.
3
Stage 3: Review and calibration
The review converts accumulated evidence into a rating. Calibration is where those ratings are moderated across managers so that they mean the same thing.
| Cadence | Annual formal review, plus a mid-year checkpoint that carries no rating but does carry a written position |
|---|---|
| Owner | The manager writes, the skip-level reviews, the HRBP facilitates calibration |
| Artifact | A written review citing at least three dated evidence entries, plus a calibration record noting any rating that moved and why |
Why this works:
Standardizing the evidence format is what makes assessment comparable across people who have never watched each other work. It is the same intervention that makes interview scoring comparable across interviewers, applied to managers instead.
A review that cannot cite three dated entries is not ready to be filed. If you want the mechanics of the assessment itself, how to run a performance evaluation is a discipline worth its own guide.
Skip the evidence standard and calibration turns into a negotiation about headcount for the merit pool. The agenda that prevents that is set out further down this page.
4
Stage 4: Reward, development, and redeployment
This is the stage where the rating becomes a decision.
| Cadence | Within four weeks of the rating, not six months later |
|---|---|
| Owner | The manager decides, the HRBP audits that no rating closes without all three decisions attached |
| Artifact | A one-page action record per employee |
Calling this stage “reward” is what reduces a year of system to a merit percentage. Attach three decisions to every rating instead.
- A reward decision: compensation, bonus, or recognition
- A development decision: what capability gets built, by when, with what support, reviewed on what date
- A redeployment decision: whether this person is in the right role, including the uncomfortable version where a strong performer has outgrown one
◆ FROM THE LAB
The Sofia test: a 48-hour internal posting window before any role went external produced a 23% internal hire rate, and internal moves reached full productivity in 28 days against 67 days for external hires. The thing that told recruiters who to approach was the performance record.
A rating with no attached action is the clearest signal an employee will ever get that the system is theater. It is also the moment most regrettable exits are decided, months before anyone resigns.
The Six Components of a Performance Management System That Holds up
The cycle tells you when things happen. This tells you what the system is built out of, and what breaks when one of the parts is missing.
Six components, because that is what the topic requires. If you are assessing platforms rather than design, how performance management systems are built and selected covers that ground separately.
Component 1: Goal architecture
The design rule:
Every individual goal must name the business objective it ladders up to and the evidence source that will assess it.
Most cascades break at the second level. A departmental objective gets translated into individual goals by people who were not in the room when the objective was set, so the connection is assumed rather than written down.
Make the parent objective a mandatory field on the goal document, then audit a 10% sample every quarter. Alignment you have not audited is alignment you are guessing at.
⚠ WATCH OUT
Red flag: an employee carrying 11 goals. That is functionally zero goals, because nothing can be deprioritized when everything has been committed. Three to five is the range where trade-offs stay visible.
Component 2: The feedback cadence engine
The design rule:
Feedback that is not documented within 48 hours cannot be used in a rating decision.
That single rule solves three problems at once. It forces the manager to convert an impression into a specific observation while the observation is still accurate, and it builds the evidence base that makes the review defensible.
It also kills the year-end surprise. A concern that has to be written down inside 48 hours is a concern the employee hears about inside 48 hours.
Here’s the entry format:
- The date
- The situation
- The behavior observed, described as an action rather than a trait
- The impact on the work
- What was agreed next
Three to five lines. Drafting assistance genuinely reduces the friction of writing one, which is the main reason managers skip the step, and where AI helps and where it must never go is set out further down.
◆ PRO TIP
Real talk: yes, this is administrative overhead, and your managers will say so in the first week. It is also the cheapest insurance available against a rating you cannot defend in a promotion case, an improvement plan, or a tribunal.
Component 3: The evidence standard
The design rule:
A rating must be supported by evidence of what the person did, not by a characterization of what they are.
“She is proactive” is unfalsifiable. “She identified the supplier qualification gap in March and closed it before it hit the build schedule” is evidence, and it can be checked by anyone in the room.
This is also the primary bias control in the whole system. Characterizations are where affinity bias, gender bias, and proximity bias enter, because a characterization is a judgment about a person rather than a record of an act. When the standard is what did they do and when, the manager with easy rapport and the manager without it carry the same burden of proof.
| Characterization | Evidence |
|---|---|
| “Strong technical instincts” | Caught the thermal derating error in the design review on 12 March, before tooling was committed |
| “Not strategic enough” | Delivered the quarterly plan on time, and has not proposed a change to the roadmap in three cycles |
| “Great team player” | Covered two shift handovers in August and trained three new operators on the packing line |
| “Struggles under pressure” | Missed two of six escalation targets in Q2, both during the peak trading week |
| “Natural with clients” | Renewed four of five at-risk accounts, including the two flagged red in January |
Unconscious bias training produces awareness and no measurable movement in rating distributions. Change what counts as an admissible justification instead, and bias has nowhere to enter the record.
Change the evidence rule, not the mindset.
◆ FROM THE LAB
Real implementation: I needed a manufacturing systems analyst, and traditional sourcing returned 14 qualified resumes. Fourteen, for a role the plant needed filled that quarter.
So I stopped searching for the job title and started searching for evidence of the work. PLC programming projects. Lean manufacturing case studies. Anywhere somebody had documented what they had built rather than what they had been called.
That search returned 62 candidates.
The person I hired had spent years as a factory-floor supervisor and taught himself Python in his own time. He now leads digital transformation. Every credential filter I had been using would have removed him at the first screen, and not one of those filters was measuring capability. They were measuring whose career had followed a conventional shape.
Here is why that story belongs in a section about performance management. The same organization that screens him out on credentials at the door will screen him out again at rating time, if the standard in the calibration room is whether he comes across as a strategic thinker rather than what he built and when. Credential gatekeeping and characterization-based rating are the same mechanism. One operates at the entrance, the other operates every year after.
Component 4: The calibration mechanism
The design rule:
Calibration exists to make ratings comparable across managers, not to fit ratings to a curve.
Organizations that conflate those two things end up defending a forced distribution to a workforce that reads it accurately as a quota. The trust cost is larger than the consistency gain, every time.
The mechanism needs five things to function: a written rating standard that predates the session, evidence packs circulated as pre-read, a facilitator who is not one of the rating managers, distribution guardrails expressed as expected ranges rather than hard quotas, and a written rationale for every rating that moves. The session agenda itself is the next section.
⚠ WATCH OUT
Watch out: calibration only works when the manager group is genuinely comparable. Calibrating a software team against a field service team produces noise rather than fairness, and everybody in the room will know it before you do.
Component 5: The consequence architecture
The design rule:
Every rating must produce a named action within four weeks, and the employee must be able to see it.
Reward is the easiest of the three pathways and the least behaviorally significant. A rating scale that maps mechanically to a merit percentage will be gamed by managers protecting their people, so keep the mapping a guide with documented discretion rather than a formula.
Development actions need a capability, a method, a support owner, and a review date. Without all four you have written an aspiration and called it a plan.
Redeployment runs in two directions. Strong performers move into roles that stretch them, and mismatched performers move into roles that fit them before the situation becomes an improvement plan.
◆ PRO TIP
The catch: all of this needs managers who are willing to release good people to other teams. Most reward structures punish them for exactly that, and no amount of policy will out-argue a bonus formula.
Component 6: The manager capability layer
The design rule:
Assume no manager can do this until you have watched them do it.
This is the component that decides whether the other five function, which is why it is listed last and matters most. “Train your managers” is not an instruction, because capability here is four separate skills and most managers were taught none of them.
- Writing a differentiated goal with a named evidence source
- Documenting an observation without characterizing a person
- Delivering corrective feedback that does not trigger defensiveness
- Defending a rating in calibration when a peer challenges it
Train each one with practice on real cases and observed feedback. A 60-minute deck moves attendance, not capability.
Then sample the output. Read evidence logs and review documents every quarter and give each manager a capability read, the same way a quality function samples what comes off a line. You do not have manager capability until you have read what your managers wrote.
of organizations report that their managers are very or extremely effective at enabling the performance of people on their teams.
Three quarters of organizations are running a system on operators they know cannot run it. That is the constraint, and no redesign of the form touches it.
◆ FROM THE LAB
The Sofia lens: the highest-volume program I have run delivered 420 hires in 10 weeks, and recruiter productivity rose by roughly 35%. None of that came from motivating anybody. It came from structured scorecards, daily stand-ups, and trackers everyone could see, which is a capability system rather than a motivation system.
How to Run a Calibration Session that Removes Bias
If I could keep one section of this page, it would be this one. Calibration is where the evidence standard either holds or collapses, and almost nobody publishes the agenda.
Five steps, one discrete action each. Publish the evidence standard before the first session, because every step below assumes it exists.
1
Step 1: Circulate evidence packs as pre-read
Each manager submits a one-page pack per employee, 72 hours before the session.
- The proposed rating
- The three strongest dated evidence entries supporting it
- Any evidence that cuts against it
That third item is what changes the room. A manager who has to surface counter-evidence in writing cannot run a one-sided verbal case.
Your job before the session is to read every pack and pre-flag the ratings that need challenge. Meeting time then goes to the contested cases instead of the obvious ones.
2
Step 2: Set guardrails, not quotas
A guardrail is an expected range that triggers a conversation when it is breached. If 80% of a team lands in the top band, you ask what evidence supports that, and sometimes the answer is legitimate.
A quota is a fixed allocation that has to be filled regardless of evidence. It converts calibration into a rationing exercise, and the workforce reads it accurately as arbitrary.
Here is the part the forced-ranking debate skips. Real performance distributions are frequently not normal, and contribution in many roles is power-law shaped, where a small minority produces a disproportionate share of the output.
Forcing a power-law reality into a bell curve does not correct bias. It suppresses the outlier contributors you most need to find.
◆ FROM THE LAB
My experience: a referral program I inherited was underperforming, and the assumed fix was a bigger bonus. It was not.
When I looked at who was referring, the distribution was nothing like the flat one the program had been designed around. Between 5% and 10% of employees were generating 60% to 70% of every successful referral. Everyone else participated occasionally, to very little effect.
So I stopped running it as a democratic program. The top referrers got Talent Scout status, real perks, and a quarterly conversation about what was coming.
One senior engineer produced 11 referrals. Nine became hires. Eight are still with the company. His conversion ran at 82% against a company average of 31%, and the agency fees that avoided came to roughly $48,000.
I am telling you this in a section about calibration for one reason. The story is not about referrals. It is about the shape of a distribution, observed directly, inside an ordinary working population.
Tighten a guardrail for one reason only: persistent, unexplained rating inflation across several cycles in the same manager group.
3
Step 3: Demand evidence, not adjectives
Set the rule before the first rating is discussed. Any justification that describes a person rather than an action is inadmissible and has to be restated.
Three sentences do most of the work in the room. “What did they do, and when?” “That is a description of how they come across; what was the output?” “Which entry in the evidence log supports that?”
Challenge these four words automatically, because they carry the heaviest bias load and the least information.
- Proactive
- Strategic
- Mature
- A culture fit
Half of the executives McKinsey surveyed said their evaluation and feedback systems have no impact on performance, or a negative one. This is the mechanism. An assessment built on characterization cannot generate a development action, because there is nothing specific to develop.
Source: McKinsey & Company, In the spotlight: Performance management that puts people first, 2024
⚠ WATCH OUT
Watch out: enforcing this on some managers and not others is worse than not enforcing it at all. Announce the rule to the whole room up front, and apply it first to the most senior person in it.
4
Step 4: Challenge the outliers in both directions
In almost every calibration session, a proposed low rating gets interrogated and a proposed top rating gets nodded through. That is not fairness. It is conflict avoidance, and it is how rating inflation becomes structural.
Apply identical scrutiny to both bands, and say so at the start so the challenge does not read as targeted when it lands.
For the top band, ask what this person did that a solid performer in the same role did not do. If the answer is a longer list of the same activities, that is volume rather than differentiated performance.
For the bottom band, ask whether the evidence is about outcomes or about a working style the manager finds difficult.
Defensible assessment is not an abstraction. The 90% offer acceptance high-water mark I have worked to came from people being able to see exactly how they had been judged.
◆ PRO TIP
The honest downside: this makes the session longer and more uncomfortable, and the first cycle after you introduce it will produce visible manager resistance. Expect that and hold the line. The resistance concentrates in the managers whose top ratings have never once been questioned.
5
Step 5: Document the rationale for every change
For every rating that moved, write two lines: the original rating, the final rating, and the evidence that drove the change.
It gives the manager language to explain the outcome, which is the only thing that stops the change being experienced as arbitrary. It creates a defensible record if the rating is contested later in a promotion decision or a redundancy selection. And after two cycles it gives you a pattern dataset.
That dataset is the quiet prize. It shows which managers consistently over-rate, which consistently under-rate, and which teams are being calibrated against peers who are not comparable.
Documentation standards vary by jurisdiction, so treat this as operational guidance and take the legal position from your own counsel. The general principle travels well enough: a documented, evidence-based rationale is a stronger position than a remembered one.
⚠ WATCH OUT
Red flag: if no ratings moved in the session, the session did not calibrate anything. It held a meeting.
Performance Management Models: Which One to Use and When
Somebody has probably asked you to move to OKRs, or to run 360s, on the strength of an article they read. The taxonomy is easy to find. A basis for choosing between them is not.
Here’s the difference:
The Lab has nothing to sell, which is the only reason the fourth column below can be written honestly.
| Model | What it measures | Best suited to | Where it breaks | The Lab’s verdict |
|---|---|---|---|---|
| MBO | Individual objectives agreed between a manager and an employee | Roles where output is cleanly attributable to one person | Interdependent work, where contribution cannot be separated out | Still works in sales and field roles, weak anywhere collaboration decides the outcome |
| OKRs | Ambitious team outcomes against defined results | Fast-moving functions that need direction more than judgment | The moment they feed an individual rating, because people stop setting ambitious objectives | Keep them for alignment, keep them out of the rating |
| 360-degree feedback | How a person’s behavior is experienced by peers, reports, and stakeholders | Development input for managers and senior individual contributors | Low-trust cultures, where anonymity turns it into score-settling | Powerful for development, dangerous the moment it touches a rating |
| BARS | Observed behavior against written anchors specific to the role | Operational and safety-critical roles with repeatable work | When the anchors are written once and never revised as the role changes | The most underused model on this list, and the best fit for non-desk populations |
| Continuous performance management | Progress and capability, assessed in rhythm rather than at a single point | Organizations that have already fixed manager capability | When it is layered on top of the annual cycle instead of replacing part of it, so managers carry both | The direction of travel, provided you remove the overhead it was meant to replace |
The selection rule is simpler than the taxonomy suggests. Choose the model that matches how attributable the work is.
Where individual contribution is cleanly attributable and outcomes land inside the cycle, objective-based models work. Where the work is highly interdependent and outcomes emerge across several cycles, behavior-anchored and continuous models assess contribution to a collective result far better.
And where the real problem is manager capability rather than measurement, changing the model will not help you at all.
◆ PRO TIP
Pro tip: go back to the three failure modes before you pick anything. If yours is a capability failure, a new model hands your managers a different form they cannot fill in properly, and the manager capability layer is where the year belongs instead.
Employee Performance Management for Non-desk and Shift-based Teams
Everything on this page so far assumes a salaried knowledge worker with a laptop, a calendar, and a quarterly goal document. In most large organizations that describes a minority of the workforce.
This is the population I came up in. Engineering and manufacturing, shift patterns, and supervisors carrying spans that no corporate framework was ever designed around.
Four structural differences break the standard model, and each one breaks it in a specific way.
Span of control goes first. A shift supervisor with 30 to 40 reports cannot run monthly 30-minute check-ins, so a design that requires them will not happen. What happens instead is one supervisor filling in 35 forms on a Thursday afternoon.
Access goes next. No corporate email, no laptop, and on many lines no personal device permitted at the workstation. The evidence log has to live somewhere a supervisor can reach inside 90 seconds or it will not be used.
Measurement is the one most organizations get backwards. Output is already captured continuously and objectively by the production system, so re-measuring throughput inside a review conversation tells you nothing the system does not already know. The conversation should be about capability and safety behavior instead.
Cycle length breaks it last. An annual cycle means very little against a workforce with high internal movement between lines, shifts, and sites, where the person rating someone in December did not manage them in March.
⚠ WATCH OUT
Common mistake: rolling the corporate framework out uniformly and accepting low-quality completion from operations. You are not running one system with a weak segment. You are running a system that half your workforce has no realistic way to take part in.
Here’s how to build the second one:
| Cadence | A 10-minute structured conversation per person per quarter, supervisor-led, plus a monthly team debrief. Not a monthly one-to-one, which is not achievable at that span |
|---|---|
| Format | A standardized paper or kiosk scorecard carrying five to seven behavioral anchors written for that specific line, scored against observation rather than memory |
| Access | Evidence capture at the supervisor’s station, entered once per shift, in under two minutes |
| Measurement | Throughput comes from the production system automatically. The human assessment covers safety behavior, quality discipline, cross-skilling progress, and reliability |
| Consequence | Tie the rating to cross-skilling and shift-preference decisions, which are the consequences this population values, rather than to a merit percentage that is often collectively bargained anyway |
Same evidence standard as the corporate model. Different cadence, format, and access. The evidence rule does not bend for the shop floor, and it should not.
◆ FROM THE LAB
Real implementation: the ask was 350 to 500 hires in 8 to 12 weeks, across customer support, operations analysts, junior engineers, and sales support. The constraint was never effort. It was interviewer capacity.
So I planned backwards from the business start date rather than forwards from the requisition. Weekly hiring targets, recruiter bandwidth, interviewer availability, and sourcing throughput were all mapped before the first advert went live.
Then I took judgment out of the places it was adding nothing. Standardized pre-assessment criteria. Batch interviews rather than individually scheduled ones. Structured scorecards, so every interviewer recorded the same things in the same format. Daily recruiter stand-ups and a tracker everyone could see.
It delivered 420 hires in 10 weeks. Time-to-Fill dropped from 32 days to between 18 and 20. Offer-to-join held above 90%, and 90-day retention landed at 85% to 88%, in line with the historical benchmark, which is the number that proves quality survived the volume.
Two mechanisms carry straight over to a shop floor. Standardized behavioral anchors, so assessment means the same thing whoever is holding the clipboard. And a cadence visible enough that a supervisor with 35 reports can sustain it without a reminder from HR.
None of that is a scaled-down version of the corporate form. It is a second operating model, built for the constraints the first one pretends do not exist.
AI in Performance Management: What to Automate, What Never to Automate
Your managers are already pasting review content into AI tools. The only open question is whether you define sanctioned use before an incident defines it for you.
Start with what genuinely helps, and be honest that the benefit is friction removal rather than intelligence. Four uses hold up.
- Drafting the evidence log entry, where the manager supplies three lines of raw observation and the tool structures it into situation, behavior, and impact
- Summarizing twelve months of entries into a review draft the manager then edits and owns
- Language-checking a written review for characterization, gendered descriptors, and unfalsifiable adjectives before it is submitted
- Drafting differentiated goal language from a role description
All four operate on inputs the manager has already produced. That is the defining test, and it is the only one you need to teach.
The third is the strongest of them. It enforces the evidence standard mechanically, at the exact moment a manager is least likely to enforce it on themselves.
◆ FROM THE LAB
My experience: use AI as a co-pilot, not a decision-maker. I have built prompt frameworks for role clarity, skill translation, inclusive drafting, structured question design, and market synthesis, and every one of them carries a human validation point before anything is decided. AI accelerates thinking. The human owns judgment, context, and the final call.
⚠ WATCH OUT
Warning: three things AI must never do inside a performance cycle.
It must never generate or suggest a rating. The training data is your own historical rating record, which encodes whatever bias produced it, so a model trained on ratings that historically favored one group will reproduce that preference wearing the appearance of objectivity. That makes the bias harder to challenge, not easier.
It must never infer performance from activity telemetry. Keystrokes, message volume, calendar density, and badge-in times measure activity rather than output, and a rating built on them converts performance management into surveillance.
It must never process identifiable employee performance data in a consumer AI tool.
None of that is hypothetical. Gartner reports that most managers already experimenting with AI in performance management have received no formal training on appropriate use, and names performance management becoming less and more human as one of four trends to prepare for in 2026.
Source: Gartner, Four Trends Talent Management Leaders Should Prepare for in 2026, 29 October 2025
Route any automated design past legal before you deploy it. An adverse employment decision materially influenced by an automated system attracts scrutiny under the EU AI Act’s high-risk employment provisions and under emerging US state and municipal rules on automated employment decision tools.
The telemetry prohibition carries its own exposure, with GDPR Article 22 relevant for anyone operating in the EU or the UK. This is operational guidance, not a legal position, and the jurisdictions differ enough that yours has to come from your own counsel.
The performance management software you run on determines how easy each of these clauses is to enforce, which is worth knowing before you sign anything.
Here’s the policy:
Six clauses. Adopt them as written and edit, rather than starting from a blank page three weeks before someone asks you for a position.
- AI may draft. Only a named human may decide.
- Any AI-assisted review is edited and signed by the manager, who owns every word of it.
- No rating may be generated, ranked, or ordered by an automated system.
- No performance inference may be drawn from activity telemetry.
- Employee-identifiable performance data is processed only in enterprise-tenanted tools the organization has approved.
- Any AI use in the performance cycle is disclosed to employees.
◆ PRO TIP
Lab Note: a policy with no enforcement point is a memo. Yours already exists, because the calibration evidence pack is the natural control. A review that was generated rather than assisted will be visibly unsupported by dated evidence entries, and the facilitator will see it before anybody else does.
Performance Improvement Plans
In most organizations, everyone involved in a performance improvement plan understands it as a managed exit. The employee knows. The manager knows, and HR wrote the template.
◆ PRO TIP
Real talk: that shared understanding is exactly why a plan so rarely improves anything. A process everyone reads as a decision already taken cannot function as an intervention, however carefully the objectives are written.
So put a gate in front of it. No improvement plan should be issued where the concern was not documented within 14 days of first observation and raised with the employee at the time.
If the person is hearing about it for the first time on the plan, the system failed, not the person. That is what the 48-hour documentation rule is protecting against, three components upstream of this one.
⚠ WATCH OUT
Common mistake: issuing a plan for a concern that appears nowhere in the record. It is the most reliable way an organization turns a manageable performance conversation into a dispute, and it is entirely self-inflicted.
Run it over 60 or 90 days. Never 30. Thirty days is not long enough for a change in behavior to become visible, everyone in the room knows that, and it is the clearest possible signal to the employee that the outcome was settled before the meeting started.
Here’s what goes in it:
- Two or three outcome-based objectives, each naming the evidence source that will assess it
- A weekly 30-minute check-in, scheduled across the whole period on day one rather than booked week by week
- A named support resource: training, mentoring, tooling, or an adjustment to the role itself
A plan that demands change and supplies no support is a notice period with extra paperwork. The support line is what separates the two.
Then measure your own success rate, because that number tells you what the plan is being used for. A well-run process should return a meaningful minority of participants to sustained performance.
If yours sits at effectively zero, the plan is not a performance tool in your organization, whatever the policy document says.
This is where my position on retention is most testable. If the concern was never raised in cycle, if the role was badly designed, or if the manager could not give corrective feedback, the failure sat upstream of the employee and the plan is being asked to repair something it did not break.
Requirements vary by jurisdiction. A plan is not a lawful precondition to termination everywhere, and completing one does not protect an employer, so take the legal position from your own counsel rather than from a blog. The mechanics of a performance improvement plan built to rehabilitate deserve more room than this section gives them.
Connecting Performance Management to Internal Mobility and Retention
Your performance cycle produces something no other process in the organization produces. A dated, evidenced, manager-verified record of what every single employee is capable of.
In most organizations that record is written once a year, filed, and never queried by anybody making a hiring decision.
So route it deliberately. Every rating already carries a redeployment decision under the consequence architecture, which means the internal mobility function should be able to query that field, and open roles should be matched against it before the requisition goes external.
Then look hard at who you are losing. The employees most likely to leave are frequently the ones the performance system has just confirmed are performing well, and then given nothing new to do.
Retention is not a compensation problem. It is a systems problem, and the system in question is a performance process whose output never reaches the people making hiring decisions.
Routing the performance record: the old way versus the Lab Way
The old way
- Run the cycle, publish the ratings, file them
- Post the role externally the week it opens
- Solve retention separately, with pay reviews and engagement surveys
- Find out about the internal candidate after they resign
The Lab Way
- Treat the performance record as the capability inventory of the Talent Supply Chain
- Open every role internally for 48 hours before it goes external
- Have recruiters approach named internal candidates rather than wait for applications
- Match open roles against the redeployment field before the requisition is written
◆ FROM THE LAB
The Sofia test: external hiring on senior engineering roles was expensive and slow to pay back. Filling the seat was only half the problem. The ramp afterwards was the other half.
So I introduced a 48-hour internal window before any role went external. That part was easy, and it was not the part that worked.
What worked was changing what recruiters did during those 48 hours. Rather than posting the role internally and waiting for applications, they went and found people. They approached named individuals whose managers had already documented what those people were capable of, which is the only reason they knew who to approach.
Internal moves reached 23% of all hires. Time-to-Productivity for internal hires came in at 28 days, against 67 days for external ones, which is 40% to 50% faster.
The number that convinced the business was 28 against 67. The number that mattered to me was a different one. Every single internal match existed because somebody had written down, months earlier, what that person had done and when.
That is the whole argument for the First Look policy, and it is why the policy fails in organizations that have not fixed the evidence standard first. A 48-hour window is worth nothing if nobody can tell you who is behind the door.
The performance record is not an HR artifact. It is the capability inventory for your entire talent management system, and most organizations let it expire in the HRIS every twelve months.
Your 90-day Performance Management Implementation Roadmap
You have just read fifteen possible changes. Here is the order to make them in, and it is not the order you will be tempted to use.
1
Days 1 to 30: diagnose
Do not fix anything this month. Find out what is broken.
- Run the three-failure-mode diagnostic against your own system
- Sample 20 completed reviews and count how many cite three or more dated evidence entries
- Put the single trust question in your next pulse survey
Deliverable: a one-page diagnosis naming which of the three failure modes you have. One page, because you will be reading it aloud to people who have twenty minutes.
2
Days 31 to 60: build the two things that unlock the rest
Two components carry everything else, so build only those.
- Publish the evidence standard and the 48-hour documentation rule as written policy
- Train managers on the two skills that matter first, documenting an observation without characterizing a person, and defending a rating with evidence
- Train both with observed practice on real cases, never with a deck
Deliverable: a written evidence standard, and a manager cohort you have watched use it.
Pre-align capability before the requirement lands. In practice that means training managers on the standard before the cycle needs it, rather than during the week they are chasing submissions.
3
Days 61 to 90: run one cycle differently
One business unit. Not the organization.
- Implement the calibration protocol in a single business unit
- Measure rating distribution variance before and after the session
- Document every rating that moved, and why
Deliverable: a calibration record and a variance comparison you can put in front of leadership without caveats.
◆ PRO TIP
The catch: choose the pilot business unit on the strength of its leader, not the size of its headcount. The first properly run calibration session is uncomfortable enough without a skeptic chairing it.
Then stop. Do not attempt the full redesign in year one, and do not let anyone talk you into launching all six components together.
Fix the evidence standard first, because every other component depends on it. Scale what demonstrably worked, in the second year, using the variance numbers you now have.
Go deeper on any part of this
This page is the operating model. Each piece of it has its own guide.
Building the system
-
The performance management process
The full cycle written as procedure, stage by stage, for the person who has to run it next quarter.
-
Performance management systems
How to design and select the system underneath the process, and how to tell a design problem from a tooling one.
-
Performance management software
What the platforms do, what they cannot do, and the questions to put to a vendor before you sign.
Assessing the individual
-
Employee performance evaluation
How to run the assessment itself without falling back on impression, and what evidence has to be in the room.
-
Performance appraisal
The appraisal event, its real limits, and what a rating can honestly tell you on its own.
-
Performance improvement plans
Designing a plan that rehabilitates rather than one that documents an exit already decided.