En este modulo
- The problem with traditional evaluation
- Automated OKR tracking with AI
- 360 feedback analysis with AI
- AI-assisted calibration
- From annual review to continuous feedback
- Bias in performance reviews
- Writing evaluations with AI
- Preparing difficult evaluation conversations
- Evaluation tools with AI
- Ejercicio practico
- Puntos clave
The problem with traditional evaluation
The annual performance evaluation is the most hated HR process in the company. Employees hate it because they feel their year of work is reduced to an arbitrary number. Managers hate it because they have to fill in endless forms with meaningless text. And HR hates it because they know the results are contaminated by bias and rarely translate into concrete actions.
The problems with traditional evaluation are well known:
- Recency effect: the manager remembers what happened in the last 4 weeks, not the last 12 months. A December mistake weighs more than 11 months of good performance.
- Halo effect: a positive (or negative) first impression contaminates the entire evaluation. If the employee is likable, all their competencies are rated higher.
- Central tendency: managers avoid extremes. Everyone is "satisfactory" or "good". Real differentiation disappears.
- Affinity bias: we evaluate people who resemble us better (same university, same hobbies, same communication style).
- Lack of data. The evaluation is based on the manager's memory and impressions, not on objective performance data.
AI does not make evaluation a perfect process. But it can mitigate each of these problems if used correctly.
The game-changing stat
According to a CEB (now Gartner) study, 95% of managers are dissatisfied with their performance evaluation system. And 90% of HR leaders believe their evaluations do not produce accurate information. We all agree it does not work. The question is how to fix it.
Automated OKR tracking with AI
OKRs (Objectives and Key Results) are the most widely used goal-setting framework in tech companies and increasingly in other sectors. The concept is simple: you define an Objective (qualitative, inspiring) and 3-5 Key Results (quantitative, measurable) that indicate whether you achieved it.
Where OKRs fail without AI
The problem is not defining OKRs. It is tracking them. In most companies, OKRs are defined in January, superficially reviewed in July, and rediscovered in December when evaluation time comes. By then, half are obsolete because the context has changed.
How AI improves OKR tracking
- Assisted definition. AI helps formulate OKRs that meet SMART criteria. A vague objective like "improve customer service" becomes "Reduce average ticket response time from 24h to 8h in Q2, measured by the ticketing system".
- Automatic updates. If Key Results are linked to internal system data (CRM, project tool, product metrics), AI can update progress automatically. Without anyone having to fill in anything.
- Deviation alerts. If a KR is below 30% at mid-quarter, AI alerts the employee and manager. Not at quarter end when it is too late, but in time to correct.
- Adjustment suggestions. If context has changed (a team member left, a new priority emerged), AI suggests OKR adjustments with justification.
Prompt to define OKRs with AI
"Help me define OKRs for [name/position] for next quarter. Department context: [priorities]. Their role includes: [main responsibilities]. Generate 2-3 Objectives with 3-4 Key Results each. Each Key Result must be: measurable (with metric and numerical target), ambitious but achievable, and linked to a system where it can be verified. Include for each KR: metric, current value (baseline), target, and data source."
360 feedback analysis with AI
360 feedback collects evaluations from multiple sources: direct manager, peers, direct reports, and sometimes internal clients. It is more comprehensive than manager-only evaluation. But it also generates a volume of data that is hard to process.
The 360 problem without AI
A 360 for an employee with 5 evaluators generates 30 to 50 responses (between quantitative and open). For a team of 20 people, that is 600 to 1,000 responses someone has to read, analyze and synthesize. That "someone" is normally the manager, who does not have time and ends up doing a superficial summary.
How AI analyzes 360 feedback
- Synthesis by competency. AI groups responses by competency and generates a summary with key themes, identified strengths and areas for improvement. Includes anonymized text quotes as evidence.
- Discrepancy detection. If the manager rates a competency as excellent but peers rate it as needs improvement, AI flags it. These discrepancies are the most valuable conversations in the process.
- Language analysis. AI detects patterns in open responses: are the same words repeated? Is there a theme all evaluators mention? Is there an evaluator who clearly has a personal bias (everything positive or everything negative)?
- Executive report. For each employee, a 1-2 page document with: strengths (3 bullets), development areas (3 bullets), action recommendations (3 bullets), and a summary sentence for the manager.
Prompt to analyze 360 feedback
"Analyze this 360 feedback for [name, position]. [Paste all responses.] Generate: 1) Key strengths (3, with evidence from evaluators). 2) Development areas (3, with evidence). 3) Discrepancies between evaluators (where they significantly differ). 4) Overall feedback pattern (positive, neutral, concerning). 5) 3 recommended development actions with timeline. 6) Suggested points for the manager-employee evaluation conversation. Do NOT identify individual evaluators. Maintain anonymity."
AI-assisted calibration
Calibration is the process where managers in a department meet to compare and adjust their evaluations, ensuring that an "excellent" from one manager means the same as an "excellent" from another. Without calibration, evaluations reflect differences between managers, not between employees.
How AI assists in calibration
- Evaluation pattern detection. AI identifies managers who consistently rate higher or lower than average (lenient bias / strict bias). This is not to single out the manager, but so the group has context when discussing.
- Evidence comparison. When two employees with similar performance receive different evaluations, AI presents the evidence for both so the group can compare objectively.
- Intelligent forced distribution. If the company uses forced distribution (X% excellent, Y% good, etc.), AI can suggest the distribution based on data, preventing the decision from being purely political.
- Automatic documentation. AI generates the calibration session minutes: decisions made, justifications and adjustments. This is especially important for regulatory compliance.
Calibration is not consensus
The goal of calibration is not for everyone to agree. It is for discrepancies to be based on data, not perceptions. If two managers evaluate their teams differently, the question is not "who is right" but "what evidence does each have". AI facilitates this data-based conversation.
From annual review to continuous feedback
The trend in performance evaluation is clear: from the annual review (once a year, long, formal) to continuous feedback (frequent, short, informal). Companies like Adobe, Microsoft, Deloitte and Netflix have eliminated or drastically transformed their annual evaluations.
The continuous feedback model
- Weekly or biweekly check-ins. Brief meeting (15-30 min) between manager and employee. It is not evaluation. It is conversation: what is going well, what do you need, how can I help.
- Real-time peer-to-peer feedback. After a project or presentation, immediate feedback from colleagues. Do not wait 6 months for the 360.
- Instant recognition. Tools that allow public recognition of a contribution the moment it occurs. Bonusly, Kudos, or simply a dedicated Slack channel.
- Quarterly evaluation (not annual). A more formal check every quarter, aligned with OKR cycles. Shorter and more frequent than the annual review.
The role of AI in continuous feedback
- Intelligent reminders. AI reminds the manager to do the weekly check-in. If it detects 3 weeks without a 1:1 meeting with a report, it sends an alert.
- Feedback suggestions. Based on OKRs and project data, AI suggests check-in topics: "Juan completed project X 2 days before the deadline. Recognize it?"
- Automatic logging. AI transcribes and summarizes check-ins (with consent), creating a history that avoids the recency effect in the quarterly evaluation.
- Recognition nudges. AI detects notable contributions (an important PR merge, a satisfied customer, a met deadline) and suggests the manager recognize them.
Bias in performance reviews
Bias in performance evaluations is systematic and well-documented. Some data:
- Women receive 2.5 times more feedback about their communication style than about their technical skills (Textio, 2024).
- Employees from ethnic minorities receive vaguer evaluations with less actionable feedback (Harvard Business Review, 2023).
- Remote employees are rated 5-10% lower than their on-site peers for the same performance (proximity bias).
- Introverted employees are rated lower in "leadership" regardless of their results.
How AI detects bias in reviews
- Language analysis. AI compares the language used in evaluations of different demographic groups. If women's evaluations use more emotional adjectives ("empathetic", "collaborative") and men's more competence adjectives ("strategic", "decisive"), there is bias.
- Score-evidence consistency. If an employee has a high score but vague written evidence, or a low score with solid evidence, AI flags the inconsistency.
- Cross-evaluator comparison. AI identifies evaluators who systematically rate employees from a certain group differently. Not as an accusation, but as data for reflection.
Prompt to review bias in evaluations
"Analyze these 10 performance evaluations [paste texts]. For each: 1) Identify if the feedback is specific and actionable or vague and generic. 2) Classify the language: results-oriented, personality-oriented, or mixed. 3) Flag if there are concrete development recommendations. Then, generate an aggregate report: are there systematic patterns in the type of language used? Are there evaluations significantly more vague than others? Are there deviations in the positive/negative feedback ratio?"
Writing evaluations with AI
Writing a good evaluation takes between 30 and 60 minutes per employee. For a manager with 10 direct reports, that is 5 to 10 hours of writing. AI can reduce that time to 5 or 10 minutes per evaluation by generating a first draft the manager reviews and personalizes.
Prompt to write a performance evaluation
"Write a performance evaluation for [name], [position], period [Q1-Q4 2026]. Data: OKRs and results: [list]. 360 feedback summary: [strengths and areas for improvement]. Notable achievements: [list]. Development areas: [list]. Generate: 1) Overall assessment paragraph (5 lines). 2) Strengths (3 bullets with concrete evidence). 3) Development areas (2 bullets with specific improvement suggestions). 4) Objectives for next period (3 proposals). Tone: professional, constructive, specific. Avoid vague language ('good overall performance') and generic phrases ('keep it up')."
Preparing difficult evaluation conversations
A manager's hardest conversation is not with the excellent employee or the terrible one. It is with the employee who is "fine but not great", who has not improved in two years, or who has a brilliant technical strength but problematic interpersonal behavior.
How AI helps prepare these conversations
- Structuring the conversation. AI generates a guide with points to cover, recommended order (start with the positive? go straight to the problem?), and key phrases for each moment.
- Anticipating reactions. Based on the employee's profile and feedback content, AI suggests possible reactions and how to handle them.
- Practicing. The manager can simulate the conversation with AI before having it with the employee. AI plays the employee role with different reactions (defensive, emotional, disinterested).
Prompt to simulate evaluation conversation
"I am going to have an evaluation conversation with an employee who: [situation description]. I want to communicate: [main message]. Simulate the conversation: you are the employee, I am the manager. The employee will probably react with [expected reaction]. After the simulation (5-6 exchanges), give me feedback on my approach: what I did well, what I could improve, and a 5-point guide for the real conversation."
Evaluation tools with AI
- Lattice: complete performance management platform with OKRs, 360, check-ins, analytics. Integrated AI for evaluation writing and bias detection. From 8 EUR/user/month.
- 15Five: focused on continuous feedback, check-ins and engagement. Less formal, more conversational. Good for companies wanting to abandon the annual review. From 4 EUR/user/month.
- BetterWorks: specialized in OKR management with advanced analytics. HRIS integrations. For medium and large companies.
- Notion/Docs + LLM: the DIY option. Evaluation template in Notion, OKRs in a table, feedback in shared documents, and an LLM to analyze and write. Works for small teams.
Ejercicio practico
- Define OKRs for a team member for next quarter using this module's prompt. Ensure each KR has baseline, target and data source.
- If you have a recent 360 (or a collection of feedback), analyze it with this module's prompt and generate the executive report.
- Take a performance evaluation you wrote recently and analyze it with the bias detection prompt. Rewrite it with AI eliminating vague or biased language.
- Simulate a difficult evaluation conversation with an LLM. Practice 2 different scenarios (defensive employee, disinterested employee).
- Design a 15-minute weekly check-in system: 3 standard questions, notes format, and how you will record it to avoid the recency effect.
Bonus: Take your entire team's evaluations from the last cycle. Ask an LLM to detect aggregate bias patterns: are there systematic differences in language, specificity or tone between different groups?
Puntos clave
Puntos clave from HR06
- 95% of managers are dissatisfied with their evaluation system. The problems: recency effect, halo, central tendency and affinity bias. AI mitigates each if used correctly.
- OKRs with automated tracking and deviation alerts prevent objectives from becoming a January exercise forgotten until December.
- AI analyzes 360 feedback in minutes (what takes a manager hours), detects discrepancies between evaluators, and generates actionable executive reports.
- Bias in evaluations is systematic and measurable. AI can detect differences in language, specificity and tone between demographic groups. The first step to correcting it is making it visible.
- Continuous feedback (weekly check-ins + instant recognition) is more effective than the annual review. AI facilitates reminders, topic suggestions, and automatic logging.
Guia de estudio — Conceptos clave de HR06
El problema de la evaluacion tradicional
- Efecto de recencia:el manager recuerda lo que paso en las ultimas 4 semanas, no en los ultimos 12 meses. Un error en diciembre pesa mas que 11 meses de buen rendimiento.
- Efecto halo:una primera impresion positiva (o negativa) contamina toda la evaluacion. Si el empleado es simpatico, todas sus competencias se evaluan mejor.
- Tendencia central:los managers evitan los extremos. Todo el mundo es "satisfactorio" o "bueno". La diferenciacion real desaparece.
- Sesgo de afinidad:evaluamos mejor a personas que se parecen a nosotros (misma universidad, mismos hobbies, mismo estilo de comunicacion).
- Falta de datos.La evaluacion se basa en la memoria y las impresiones del manager, no en datos objetivos de rendimiento.
- El dato que lo cambia todo: Segun un estudio de CEB (ahora Gartner), el 95% de los managers estan insatisfechos con su sistema de evaluacion del desempeno. Y el 90% de los responsables de RRHH creen que sus evaluaciones no producen informacion precisa. Estamos todos de acuerdo en que no funciona. La pregunta es como arreglarlo.
OKR tracking automatizado con IA
- Definicion asistida.La IA ayuda a formular OKRs que cumplan los criterios SMART. Un objetivo vago como "mejorar la atencion al cliente" se transforma en "Reducir el tiempo medio de respuesta a tickets de soporte de 24h a 8h en Q2, medido por el sistema de tickets".
- Actualizacion automatica.Si los Key Results estan vinculados a datos de sistemas internos (CRM, herramienta de proyectos, metricas de producto), la IA puede actualizar el progreso automaticamente. Sin que nadie tenga que rellenar nada.
- Alertas de desviacion.Si un KR va por debajo del 30% a mitad del trimestre, la IA alerta al empleado y al manager. No al final del trimestre cuando ya es tarde, sino a tiempo para corregir.
- Sugerencia de ajustes.Si el contexto ha cambiado (se ha ido un miembro del equipo, ha surgido una prioridad nueva), la IA sugiere ajustes a los OKRs con justificacion.
Analisis de feedback 360 con IA
- Sintesis por competencia.La IA agrupa las respuestas por competencia y genera un resumen con los temas clave, las fortalezas identificadas, y las areas de mejora. Incluye citas textuales anonimizadas como evidencia.
- Deteccion de discrepancias.Si el manager evalua una competencia como excelente pero los peers la evaluan como mejorable, la IA lo senala. Estas discrepancias son las conversaciones mas valiosas del proceso.
- Analisis de lenguaje.La IA detecta patrones en las respuestas abiertas: se repiten las mismas palabras? Hay un tema que mencionan todos los evaluadores? Hay un evaluador que claramente tiene un sesgo personal (todo positivo o todo negativo)?
- Informe ejecutivo.Para cada empleado, un documento de 1-2 paginas con: fortalezas (3 bullets), areas de desarrollo (3 bullets), recomendaciones de accion (3 bullets), y una frase resumen para el manager.
Calibracion asistida por IA
- Deteccion de patrones de evaluacion.La IA identifica a managers que evaluan consistentemente mas alto o mas bajo que la media (lenient bias / strict bias). Esto no es para senalar al manager, sino para que el grupo tenga contexto al discutir.
- Comparacion de evidencia.Cuando dos empleados con rendimiento similar reciben evaluaciones diferentes, la IA presenta la evidencia de ambos para que el grupo pueda compararlos objetivamente.
- Distribucion forzada inteligente.Si la empresa usa distribucion forzada (X% excelente, Y% bueno, etc.), la IA puede sugerir la distribucion basandose en los datos, evitando que la decision sea puramente politica.
- Documentacion automatica.La IA genera el acta de la sesion de calibracion: decisiones tomadas, justificaciones, y ajustes realizados. Esto es especialmente importante para cumplimiento normativo.
- Calibracion no es consenso: El objetivo de la calibracion no es que todos esten de acuerdo. Es que las discrepancias se basen en datos, no en percepciones. Si dos managers evaluan diferente a sus equipos, la pregunta no es "quien tiene razon" sino "que evidencia tiene cada uno". La IA facilita esta conversacion basada en datos.
Del annual review al feedback continuo
- Check-ins semanales o quincenales.Reunion breve (15-30 min) entre manager y empleado. No es evaluacion. Es conversacion: que va bien, que necesitas, como puedo ayudarte.
- Feedback peer-to-peer en tiempo real.Despues de un proyecto o presentacion, feedback inmediato de los companeros. No esperar 6 meses al 360.
- Recognition instantaneo.Herramientas que permiten reconocer publicamente una contribucion en el momento en que ocurre. Bonusly, Kudos, o simplemente un canal de Slack dedicado.
- Evaluacion trimestral (no anual).Un check mas formal cada trimestre, alineado con los ciclos de OKRs. Mas corto y mas frecuente que el annual review.
- Recordatorios inteligentes.La IA recuerda al manager hacer el check-in semanal. Si detecta que lleva 3 semanas sin reunion 1:1 con un reporte, envia una alerta.
- Sugerencias de feedback.Basandose en los OKRs y los datos del proyecto, la IA sugiere temas para el check-in: "Juan completo el proyecto X 2 dias antes del deadline. Reconocerlo?"
Sesgo en performance reviews
- Las mujeres reciben 2.5 veces mas feedback sobre su estilo de comunicacion que sobre sus habilidades tecnicas (Textio, 2024).
- Los empleados de minorias etnicas reciben evaluaciones mas vagas y con menos feedback accionable (Harvard Business Review, 2023).
- Los empleados remotos son evaluados un 5-10% por debajo de sus pares presenciales por el mismo rendimiento (proximity bias).
- Los empleados introvertidos son evaluados peor en "liderazgo" independientemente de sus resultados.
- Analisis de lenguaje.La IA compara el lenguaje usado en las evaluaciones de diferentes grupos demograficos. Si las evaluaciones de mujeres usan mas adjetivos emocionales ("empatica", "colaborativa") y las de hombres mas adjetivos de competencia ("estrategico", "decisivo"), hay sesgo.
- Consistencia score-evidencia.Si un empleado tiene un score alto pero la evidencia escrita es vaga, o un score bajo con evidencia solida, la IA senala la inconsistencia.
Siguiente: HR07 - People Analytics
Performance evaluation generates data. People analytics turns it into decisions. HR dashboards, workforce planning, diversity, compensation and predictive analytics.
Ir al modulo HR07