En este modulo
- AI audit fundamentals
- Audit types: internal, external, regulatory
- Step-by-step audit methodology
- Article-by-article AI Act checklist
- Evidence collection and management
- Common findings and how to address them
- Audit report template
- Continuous auditing and automation
- Preparation for third-party audits
- Ejercicio practico
- Puntos clave
AI audit fundamentals
AI system auditing is the systematic and independent evaluation of an artificial intelligence system to verify that it complies with regulatory requirements, internal policies and applicable good practices. It is not an informal technical review: it is a formal process with defined scope, explicit criteria, documented evidence and traceable conclusions.
Unlike financial auditing or traditional information systems auditing, AI auditing presents specific challenges:
- Model opacity: ML and deep learning models are not always interpretable. The auditor must verify compliance without necessarily understanding every model parameter.
- Dynamism: AI systems can change their behavior over time (model drift, data drift). A point-in-time audit may not capture problems that emerge after weeks or months of operation.
- Interdisciplinarity: the audit requires technical (data science, ML), legal (AI Act, GDPR), ethical (fairness, explainability) and business (usage context, impact on persons) competencies.
- Regulatory novelty: the AI Act is recent. There is not yet a consolidated body of "AI audit good practices" comparable to, for example, ISO 27001. Auditors are building the discipline as they practice it.
Auditing as an obligation
For high-risk systems, auditing is not optional. The AI Act requires conformity assessment (Art. 43), quality management system (Art. 17) and post-market surveillance (Art. 72). All of this requires systematic verification processes, that is, auditing.
Audit types: internal, external, regulatory
Internal audit
Conducted by the organization itself (internal audit function or AI governance team). Advantages: full access to the system, knowledge of context, reduced cost. Risk: lack of objective independence. Recommendation: conduct at least annually for each high-risk system, and always upon substantial changes.
Independent external audit
Conducted by an independent third party (audit firm, specialist consultancy). Greater credibility with regulators and stakeholders. Recommended for critical high-risk systems and as preparation for regulatory inspections. Recommended frequency: every 2-3 years, or upon the first placing on the market of a high-risk system.
Regulatory audit
Conducted or commissioned by the market surveillance authority (in Spain, AESIA). You do not schedule it: the regulator does. Your role is to be prepared. The best preparation: rigorous internal audits and complete documentation.
Conformity assessment (Art. 43)
For high-risk systems, the AI Act requires a conformity assessment before placing on the market or putting into service. In most cases (pathway 2 of Art. 6(2), Annex III), this assessment is a self-assessment by the provider (internal control, Annex VI). Only in certain cases (biometrics, critical infrastructure via pathway 1) is assessment by a notified body required (Annex VII).
Step-by-step audit methodology
Phase 1: Planning
- Define scope: which system(s) are audited, which requirements are verified (full AI Act, GDPR, ethics, internal policy), period covered.
- Audit team: required competencies (technical, legal, ethical). Verify independence (the team must not have participated in developing or operating the system).
- Audit criteria: the reference against which the evaluation is made. AI Act (specific articles), harmonised standards (ISO/IEC 42001 if applicable), internal AI policy, sectoral good practices.
- Audit plan: timeline, deliverables, review milestones, required resources, communication protocols.
Phase 2: Information gathering
- Document review: technical documentation, DPIA/FRIA, risk classification, AI policy, provider contracts, training records, AI committee minutes.
- Interviews: system owner, operators, DPO, data officer, developers.
- Technical inspection: model review (performance metrics, logs, fairness tests), infrastructure review (security, availability), data review (quality, representativeness, governance).
- Direct observation: how operators interact with the system in practice. Is human oversight real or fictitious?
Phase 3: Analysis and evaluation
For each audit criterion, evaluate the degree of compliance:
- Conformant: evidence demonstrates full compliance with the criterion.
- Observation: substantial compliance with improvement opportunities. Not a non-compliance but may evolve into one.
- Minor non-conformity: partial non-compliance that does not pose an immediate risk to persons' rights or regulatory compliance, but requires corrective action.
- Major non-conformity: significant non-compliance that requires urgent corrective action. May entail system suspension until remediation.
Phase 4: Report
Document findings, conclusions and recommendations (see report template section).
Phase 5: Follow-up
Verify implementation of corrective actions. Close non-conformities with remediation evidence. Schedule the next audit.
Article-by-article AI Act checklist
This checklist covers the most relevant Title III articles for deployers of high-risk systems. Each item requires documentary evidence.
Article 9: Risk management system
- Is there a documented risk management system?
- Does it cover the entire system lifecycle?
- Does it identify and analyse known and foreseeable risks?
- Does it include mitigation measures for each identified risk?
- Does it assess residual risk after mitigation?
- Is it updated with new information and operational experience?
- Have tests been conducted with data representative of actual use?
Article 10: Data governance
- Do the training, validation and testing data meet documented quality criteria?
- Has the representativeness of data been assessed relative to the usage context?
- Have potential biases in the data been identified and documented?
- Are there processes to detect and correct data errors?
- If special category data is processed for bias correction (Art. 10(5)), are the required safeguards in place?
Article 11: Technical documentation
- Does technical documentation conforming to Annex IV exist?
- Is the documentation current with the system's present state?
- Is it comprehensible to the competent authority?
- Does it include system description, development process, capabilities, limitations and changes?
Article 12: Record-keeping
- Does the system generate automatic logs?
- Do logs enable traceability of the system's decisions?
- Are logs retained for at least 6 months (or the period defined by the provider)?
- Are logs accessible to the deployer?
Article 13: Transparency
- Do clear instructions for use exist?
- Do the instructions include provider identity, expected performance, known limitations, human oversight measures?
- Has the deployer reviewed and applied them?
Article 14: Human oversight
- Does the system allow effective human oversight?
- Are designated persons with competence and authority assigned to oversight?
- Can they understand the system's capabilities and limitations?
- Can they detect anomalies and act accordingly?
- Can they disregard, reverse or stop the system's decisions?
- Is the oversight real (not rubber-stamping)?
Article 26: Deployer obligations
- Is the system used in accordance with the intended purpose and the provider's instructions?
- Is the input data relevant and representative?
- Is the system's operation monitored?
- Are incidents and risks reported to the provider?
- Are automatic logs retained?
Article 27: Fundamental rights impact assessment
- Has a FRIA been conducted before putting the system into use?
- Does the FRIA cover the potentially affected fundamental rights?
- Has the market surveillance authority been notified (if applicable)?
Evidence collection and management
Evidence is what transforms an opinion into an audit finding. Without evidence, there is no finding. Each audit conclusion must be supported by at least one piece of verifiable evidence.
Types of evidence
- Documentary: policies, procedures, DPIA/FRIA, technical documentation, contracts, meeting minutes, training records, test reports.
- Technical: system logs, performance metrics, fairness test results, monitoring data, system configuration, source code (if relevant and accessible).
- Testimonial: statements from interviews with operators, owners, developers, DPO. Must be documented (signed interview notes or authorised recordings).
- Observational: direct observation of how the system is operated, how human oversight is exercised, how data subject requests are processed.
Evidence management principles
- Sufficiency: adequate quantity of evidence to support the conclusion.
- Relevance: the evidence is directly related to the criterion being evaluated.
- Reliability: the evidence comes from trustworthy and independent sources.
- Traceability: each finding must reference the supporting evidence (document number, date, page).
Common findings and how to address them
Based on experience from the first AI audits conducted under the AI Act framework, these are the most frequent findings:
1. Incomplete inventory
The organization does not have an exhaustive inventory of its AI systems. SaaS tools with embedded AI (Salesforce Einstein, Microsoft Copilot, Grammarly) are not registered. Solution: shadow AI audit, survey of all departments, review of provider contracts.
2. Undocumented risk classification
Systems are identified but the formal AI Act risk classification has not been conducted. Or it was done informally without documentation. Solution: systematic classification process with formal documentation (see TG02).
3. Insufficient human oversight
Human oversight exists on paper but not in practice. The operator approves 99% of decisions without substantive review. They have no specific training on the system. They lack authority to reverse decisions. Solution: redesign the oversight process, train operators, establish sample review mechanisms and automatic alerts.
4. DPIA not updated or non-existent
The DPIA/FRIA has not been conducted or the existing one is outdated (done before significant system changes). Solution: conduct or update the integrated DPIA/FRIA (see TG03).
5. Undocumented training data
The provider does not supply sufficient information about the model's training data. The deployer cannot verify quality and representativeness. Solution: contractually require data governance information, include specific clauses in the contract, evaluate provider alternatives.
6. Absence of production monitoring
The system was deployed and its performance, bias and drift are not monitored. There are no production fairness metrics. Solution: implement continuous monitoring with dashboards, degradation alerts and periodic fairness metric reviews.
7. Insufficient training
There is no evidence of AI training for system operators or the governance committee. Solution: implement the level-based training programme (see TG05).
Audit report template
Recommended structure
1. Executive summary
Scope, period, audit team, overall conclusion (conformant/non-conformant), summary of critical findings (maximum 1 page).
2. Objectives and scope
Which systems were audited, which criteria were applied, what was excluded from scope and why.
3. Methodology
Audit phases, techniques used (document review, interviews, technical inspection, observation), sampling applied.
4. Detailed findings
For each finding:
- Reference to the audit criterion (AI Act article, standard clause, policy point).
- Finding description.
- Supporting evidence.
- Classification (conformant, observation, minor/major non-conformity).
- Associated risk.
- Corrective action recommendation.
5. Action plan
For each non-conformity: corrective action, responsible party, deadline, closure criterion.
6. Conclusion
Overall compliance assessment. Strategic recommendations. Next scheduled audit.
7. Annexes
List of documents reviewed, list of persons interviewed, technical evidence, completed checklist with results.
Continuous auditing and automation
Point-in-time auditing (annual or biannual) is insufficient for AI systems that evolve continuously. Continuous auditing complements periodic auditing with automated monitoring.
Continuous audit elements
- Performance monitoring: dashboards with accuracy, recall, F1-score metrics. Alerts when metrics fall below predefined thresholds.
- Fairness monitoring: disparity metrics between protected groups calculated automatically with production data. Alerts when disparity exceeds the acceptable threshold.
- Drift detection: statistical comparison between production data distribution and training distribution. Alerts when significant drift is detected.
- Log verification: automated checking that event records are generated correctly and are accessible.
- Human oversight review: human intervention metrics (override rate, review time, decision distribution). Alerts if the override rate is anomalously low (possible rubber-stamping).
Preparation for third-party audits
External audits (by consultancies or the regulator) require specific preparation:
Before the audit
- Verify that the AI system inventory is complete and up to date.
- Ensure all documentation (risk classification, DPIAs/FRIAs, technical documentation, AI committee minutes, training records) is organised and accessible.
- Conduct an internal pre-audit to identify obvious gaps and correct them beforehand.
- Designate a single point of contact for the external auditor.
- Prepare the persons who will be interviewed: what to expect, how to respond (honestly, not defensively), what documentation to have at hand.
During the audit
- Provide access to documentation in an organised manner (structured digital folder, not disorganised emails).
- Respond honestly. If something has not been done, it is better to acknowledge it than to invent excuses.
- Internally document the auditor's questions and the responses provided.
- Do not provide unsolicited information that could unnecessarily expand the scope.
After the audit
- Review the audit report carefully. Challenge any finding you disagree with (with evidence).
- Develop the corrective action plan and assign responsible parties and realistic deadlines.
- Implement corrective actions within the committed deadlines.
- Schedule the verification of non-conformity closure.
Ejercicio practico
- Select a high-risk AI system (real or simulated). Define the audit scope: which AI Act articles you will verify.
- Walk through the complete checklist (Articles 9, 10, 11, 12, 13, 14, 26 and 27) for that system. For each point, indicate: conformant/observation/minor non-conformity/major non-conformity. Justify with the evidence you would expect to find.
- Identify the 3 most critical findings. For each one, draft a complete finding with: criterion, description, evidence, classification, risk and recommendation.
- Draft the executive summary of the audit report (1 page).
- Develop the corrective action plan for the 3 critical findings: action, responsible party, deadline, closure criterion.
Output: a partial simulated audit report with completed checklist, 3 drafted findings, executive summary and action plan.
Puntos clave
Puntos clave from TG06
- AI auditing requires interdisciplinary competencies: technical, legal, ethical and business. A single profile cannot cover everything.
- The article-by-article AI Act checklist (Arts. 9-14, 26-27) is the basic operational tool. Each point requires verifiable documentary evidence.
- The most common findings are: incomplete inventory, undocumented classification, fictitious human oversight, non-existent DPIA and absence of production monitoring.
- Continuous auditing (automated monitoring of performance, fairness, drift and human oversight) complements periodic auditing and is essential for dynamic systems.
- Preparation for third-party audits starts with rigorous internal audits. If you cannot pass your own audit, you will not pass the regulator's.
Guia de estudio — Conceptos clave de TG06
Fundamentos de la auditoria de IA
- Opacidad del modelo:los modelos de ML y deep learning no siempre son interpretables. El auditor debe verificar cumplimiento sin necesariamente entender cada parametro del modelo.
- Dinamismo:los sistemas de IA pueden cambiar su comportamiento con el tiempo (model drift, data drift). Una auditoria puntual puede no capturar problemas que emergen tras semanas o meses de operacion.
- Interdisciplinariedad:la auditoria requiere competencias tecnicas (data science, ML), legales (AI Act, RGPD), eticas (fairness, explicabilidad) y de negocio (contexto de uso, impacto en personas).
- Novedad regulatoria:el AI Act es reciente. No existe aun un cuerpo consolidado de "buenas practicas de auditoria de IA" comparable al de, por ejemplo, ISO 27001. Los auditores estan construyendo la disciplina mientras la practican.
- Auditoria como obligacion: Para sistemas de alto riesgo, la auditoria no es opcional. El AI Act exige evaluacion de conformidad (art. 43), sistema de gestion de calidad (art. 17) y vigilancia post-comercializacion (art. 72). Todo esto requiere procesos de verificacion sistematica, es decir, auditoria.
Metodologia de auditoria paso a paso
- Definir alcance:que sistema(s) se auditan, que requisitos se verifican (AI Act completo, RGPD, etica, politica interna), periodo cubierto.
- Equipo auditor:competencias necesarias (tecnica, legal, etica). Verificar independencia (el equipo no debe haber participado en el desarrollo o la operacion del sistema).
- Criterios de auditoria:la referencia contra la que se evalua. AI Act (articulos especificos), normas armonizadas (ISO/IEC 42001 si aplica), politica interna de IA, buenas practicas sectoriales.
- Plan de auditoria:cronograma, entregables, hitos de revision, recursos necesarios, protocolos de comunicacion.
- Revision documental: documentacion tecnica, DPIA/FRIA, clasificacion de riesgos, politica de IA, contratos con proveedores, registros de formacion, actas del comite de IA.
- Entrevistas: responsable del sistema, operadores, DPO, responsable de datos, desarrolladores.
Checklist por articulo del AI Act
- Existe un sistema de gestion de riesgos documentado?
- Cubre todo el ciclo de vida del sistema?
- Identifica y analiza riesgos conocidos y previsibles?
- Incluye medidas de mitigacion para cada riesgo identificado?
- Evalua el riesgo residual tras la mitigacion?
- Se actualiza con nueva informacion y experiencia operativa?
Recogida y gestion de evidencias
- Documental:politicas, procedimientos, DPIA/FRIA, documentacion tecnica, contratos, actas de reunion, registros de formacion, informes de tests.
- Tecnica:logs del sistema, metricas de rendimiento, resultados de tests de fairness, datos de monitoreo, configuracion del sistema, codigo fuente (si es relevante y accesible).
- Testimonial:declaraciones de entrevistas con operadores, responsables, desarrolladores, DPO. Deben documentarse (notas de entrevista firmadas o grabaciones autorizadas).
- Observacional:observacion directa de como se opera el sistema, como se ejerce la supervision humana, como se procesan las solicitudes de los interesados.
- Suficiencia:cantidad de evidencia adecuada para soportar la conclusion.
- Pertinencia:la evidencia esta directamente relacionada con el criterio evaluado.
Plantilla del informe de auditoria
- Referencia al criterio de auditoria (articulo del AI Act, clausula de la norma, punto de la politica).
- Descripcion del hallazgo.
- Evidencia que lo soporta.
- Clasificacion (conforme, observacion, no conformidad menor/mayor).
- Riesgo asociado.
- Recomendacion de accion correctiva.
Auditoria continua y automatizacion
- Monitoreo de rendimiento:dashboards con metricas de precision, recall, F1-score. Alertas cuando las metricas caen por debajo de umbrales predefinidos.
- Monitoreo de fairness:metricas de disparidad entre grupos protegidos calculadas automaticamente con datos de produccion. Alertas cuando la disparidad supera el umbral aceptable.
- Deteccion de drift:comparacion estadistica entre la distribucion de datos de produccion y la de entrenamiento. Alertas cuando se detecta drift significativo.
- Verificacion de logs:comprobacion automatica de que los registros de eventos se generan correctamente y son accesibles.
- Revision de supervision humana:metricas de intervencion humana (tasa de override, tiempo de revision, distribucion de decisiones). Alertas si la tasa de override es anomalamente baja (posible rubber-stamping).
Siguiente: TG07 - NIS2, DORA and AI
AI systems do not operate in a regulatory vacuum. The next module explores the intersection of the AI Act with NIS2 (cybersecurity) and DORA (digital operational resilience), and how to build a joint compliance framework.
Ir al modulo TG07