The EU AI Act Playbook (3/3): Why Policies Can't Survive an Audit

Why a policy document can't survive an audit, and what an EU AI Act evidence trail actually has to prove.

Key Takeaways

  • Every obligation in the EU AI Act that matters operationally is an evidence obligation. The question a regulator, an auditor or an enterprise customer asks is never “do you have a policy” — it is “show me the record.” Policies do not produce records. Systems do.
  • Article 50 has applied since 2 August 2026, and the Commission’s final Guidelines made clear that a disclosure buried in terms and conditions does not discharge it. What you need is proof that the disclosure reached the person, at the point of interaction, in a language they understand.
  • Agents are where the gap is widest. Deloitte found that only about one in five organisations has a mature governance model for autonomous AI agents — at the same time as the Commission’s Guidelines imposed specific disclosure duties on agents at authorisation, reporting and validation steps, and brought multi-agent architectures into scope.
  • The high-risk deferral to December 2027 does not change what you will eventually have to produce under Articles 9 to 15. It changes how much time you have to start generating it — and logs cannot be created retroactively.
  • No platform makes you compliant, and any vendor claiming otherwise is describing something that does not exist. Compliance attaches to a specific system, in a specific use case, under a specific role. What a platform can do is produce the evidence, and remove whole categories of work from your team.
  • AnyInsight is built for that job: a zero-trust GenAI firewall inline on every interaction, five log types including connector-level audit of agent actions, an immutable Data Vault that retains evidence regardless of compliance status, and three-tier workspace isolation — on an ISO/IEC 27001-certified platform.

1. The Gap Between a Policy and a Record

Most AI governance programmes fail their first external test in the same place. The organisation has an AI policy. It has been reviewed by legal, circulated to staff, and referenced in the annual risk report. Then a customer’s procurement team, or a market surveillance authority, or an internal auditor asks a narrower question: on 14 March, when your AI assistant handled this enquiry, was the customer told they were talking to an AI — and how do you know?

The policy cannot answer that. Only a record can.

A policy document and a system record answer different questions.

Figure 1. A policy document and a system record answer different questions.

This is the structural feature of the EU AI Act that most compliance programmes underweight. Read Articles 9 to 15 and Article 50 as an assessor would, and almost every requirement resolves into an artefact someone has to be able to retrieve: a log, a document, a test result, a demonstration that a human could intervene. The Act is written in the language of product safety, and product safety regimes run on evidence rather than intent.

Deloitte’s 2026 enterprise survey puts numbers on the mismatch. Data privacy and security is the top AI risk for 73% of organisations, legal and regulatory compliance for 50%, and governance capability and oversight for 46%. Deloitte’s own recommendation is unusually concrete on this point: organisations need to define where humans stay in control, how automated decisions are audited, and which records of system behaviour should be retained. That last clause is the whole problem in eight words.

And evidence has a property that policy does not: it cannot be generated retroactively. You can write a policy the week before an audit. You cannot write the log of what your AI did last March. This is why the high-risk deferral to December 2027 is a worse deal than it looks — the deadline moved, but the window in which the evidence has to accumulate did not.

2. What Article 50 Requires You to Prove Right Now

Article 50 has applied since 2 August 2026, and enforcement powers went live on the same day. It is the obligation almost every company running generative AI or an agent is already subject to, and the one where the evidence question is sharpest.

The Commission’s final Guidelines, adopted on 20 July 2026, closed the easy answer. Placing the disclosure in your terms and conditions does not inform the user, and technical marking alone does not satisfy the disclosure duty either. What is required is that the person interacting with the system is told, at the point of interaction, in a way they can actually process.

That standard has a practical consequence people miss: disclosure has to work in the user’s language. A German employee of your Munich subsidiary, or a Spanish customer on your support channel, has not been informed by an English-language notice they cannot read. For a company deploying one AI platform across multiple EU member states, the interface language is a compliance surface, not a usability preference.

The evidence you need is therefore threefold: that the disclosure was present, that it was presented at the right moment, and that it was intelligible to that user. A conversation-level log that records the model or agent used, the inputs and outputs, the context and the interface state is what turns all three from assertion into record.

One nuance worth being precise about, because vendors routinely blur it: the Article 50(2) marking obligation for generated content sits with the provider of the system that generates it, not with the platform that orchestrates the request. A governance layer gives you the record of what was generated, by which model, for whom and when. It does not relieve you of the marking duty if you are the provider of the generating system.

3. What High-Risk Will Require You to Prove by December 2027

If your classification lands in Annex III — most commonly through employment and worker management, or credit and insurance scoring — the obligations in Articles 9 to 15 apply from 2 December 2027. The table below maps the ones that resolve into evidence, and what generates it.

Obligation What an assessor actually asks for How AnyInsight produces it
Article 50(1)
Disclosure to users
Show that the person was informed, in a form they could understand, at the point of interaction — not a clause in your terms Dialogue Log records the model or agent used, the inputs, the outputs and the context of every conversation. The 32-language interface means the disclosure is in the user’s own language
Article 50(4)
Labelling deepfakes and public-interest text
Show which published content was AI-generated, by whom, and when Dialogue and Automation Logs attribute every generated artefact to an actor, a model and a timestamp
Article 4
AI literacy
Show the measures you took to give staff a sufficient level of AI literacy Prompt Wizard, 100+ prompt templates, the Get-started Agent and in-platform AI Helper, all delivered in 32 languages — with System Log records of who used them
Article 12
Automatic logging
Automatically recorded events over the lifetime of the system, retained and retrievable Five log types — System, Dialogue, Automation, Shared Conversation and Connector — plus a Data Vault that retains every interaction whether or not it breached policy
Article 14
Human oversight
Show that a named person could intervene, and did — not that a policy said they should Moderation Record ties each event to a workspace, user and resolution status; graded Alert / Block / Mask actions; MAIA Critique Mode surfaces a second model’s challenge for a human to weigh
Article 15
Accuracy, robustness, cybersecurity
Show the controls protecting the system from manipulation and unauthorised use Prompt protection at the boundary; access-control policies on network, geolocation, device, operating system, browser and time window, combined with allowlists
Article 26
Deployer obligations
Show assigned oversight, relevant input data, retained logs and ongoing monitoring Three-tier Org / Workspace / User identity with least-privilege roles; inline inspection of inputs; downloadable logs; real-time violation monitoring over rolling windows
B2B carve-out
(final Article 50 Guidelines)
Show cloud isolation and role-based access control preventing outputs leaving the organisation Walled workspaces with resource segmentation; sharing restricted to internal members and traced; Workspace Owner / Admin / Contributor / Auditor / User roles

Two observations about that table are worth more than the table itself.

The compliance deadline moved; the window in which evidence has to accumulate did not.

Figure 2. The compliance deadline moved; the window in which evidence has to accumulate did not.

  • Article 12 is the load-bearing one. Automatic logging is the requirement that makes every other requirement demonstrable. The design decision that matters here is what gets retained: AnyInsight’s Data Vault preserves every interaction whether or not it violated a policy, on the principle that an audit is only as trustworthy as the completeness of its evidence. A log that only keeps the incidents tells an assessor what you caught, not what happened.
  • Article 14 is the one most often misread as a policy commitment. “A human reviews AI outputs” is not oversight; it is an aspiration. Oversight means a named person, with the authority and the interface to intervene, and a record showing they could and did. The Moderation Record ties each event to a workspace, a user and a resolution status; graded Alert, Block and Mask actions make the intervention proportionate; and MAIA’s Critique Mode, in which a second model challenges the first and surfaces the disagreement, gives that human something concrete to act on rather than a single confident answer to rubber-stamp.

4. Agents Are the Hard Part

The gap between what enterprises are deploying and what they can govern is widest at agents, and both halves of that statement now have evidence behind them.

On the deployment side, Deloitte’s 2026 survey found that only around one in five organisations has a mature governance model for autonomous AI agents, even as agentic adoption accelerates sharply. On the regulatory side, the Commission’s final Article 50 Guidelines went further on agents than the May draft did: where a provider cannot determine in advance whether an agent will interact with a natural person, the agent must be designed and instructed to disclose itself wherever such interaction is reasonably likely — including when the person is acting for a legal entity. Agents must disclose themselves at key steps such as authorisation, reporting and validation, and at every new interaction. Multi-agent architectures are explicitly in scope.

The audit problem an agent creates is different in kind from the one a chat interface creates. A conversation has one actor and one boundary crossing. An agent authorised to act across connected systems performs many, autonomously, over time — reading from one application, writing to another, triggering a third. Each of those is a data movement that an assessor may ask you to account for, and none of them appears in a conversation log.

This is why connector-level logging is the capability that separates AI governance from AI monitoring. AnyInsight’s Connector Log records every operation an agent performs against an external system through a connector: which agent, which connector, which application, what data moved, and what the outcome was. The policy engine inspects and enforces at the connector as well, so the same masking, blocking and alerting that applies to a human prompt applies to an agent’s autonomous data movement. The Automation Log adds the execution layer — what triggered the task, which agents and connectors were involved, which systems were touched.

The design principle is worth stating plainly: the shadow-agent problem is not solved by trusting the agent. It is solved by putting the same governed checkpoint on the agent’s actions as on a person’s, and keeping the record.

An autonomous agent's actions across systems don't show up in a conversation log.

Figure 3. An autonomous agent's actions across systems don't show up in a conversation log.

5. What a Platform Cannot Do For You

This section exists because the previous four are only worth reading if this one is honest.

No platform can be “EU AI Act certified”, and none can make you compliant. Compliance under the Act attaches to a specific AI system, in a specific use case, under a specific role — and the harmonised technical standards that would turn the Act’s requirements into testable criteria are not finished, which is a large part of why the high-risk deadline moved to December 2027. There is no conformity assessment scheme a governance platform can hold. AnyInsight holds ISO/IEC 27001 certification, which is a real, independently audited credential, and runs on cloud infrastructure certified to ISO/IEC 27001, 27017 and 27018. Its relationship to the AI Act is architectural: it is built to produce what the Act asks you to be able to produce. Those are two different kinds of claim and they should not be presented in the same column.

Concretely, here is what stays on your side of the line:

  • Classification. Deciding whether a given system is high-risk, and whether the Article 6(3) exemption applies, is a legal judgement about your use case. No platform can make it for you, and if you rely on the exemption you must document the assessment and register the system.
  • Article 11 technical documentation for systems you build. If you are the provider of a high-risk AI system, the Act requires design-time documentation — architecture, development process, testing methodology, performance metrics. A governance layer produces operational records of the system in use. Those are complementary, not substitutes.
  • Article 10 training-data governance. If you train or fine-tune models, the examination of training, validation and test datasets for bias, gaps and representativeness is work on your datasets. AnyInsight governs data in use — what enters and leaves the perimeter — not the corpus you trained on.
  • Conformity assessment, CE marking and EU database registration. These are procedural obligations with the provider. No platform discharges them.
  • AI running outside the platform. Governance reaches what is inside the perimeter. If half your organisation is using an unmanaged tool on a corporate card, that exposure is unchanged until it is brought inside. The platform closes shadow AI by consolidation, not by detection at a distance.

The honest formulation is the one worth holding a vendor to: AnyInsight does not make you compliant. It produces the evidence you need in order to demonstrate compliance, and it removes several categories of work — building a logging pipeline, a masking layer, an oversight workflow and an audit export — that your team would otherwise build and maintain itself.

6. Provider or Deployer: Who This Is Built For

The provider-versus-deployer distinction decides which obligations are yours, and it also decides how much of the problem a governance platform can address.

If you are a deployer — you buy or subscribe to AI systems and use them under your own authority, which describes the large majority of enterprises — your obligations are operational. Assign human oversight. Ensure input data is relevant. Retain logs. Monitor operation. Disclose under Article 50 where the system interacts with people, and label deepfakes and public-interest text you publish. Every one of those resolves into evidence generated at runtime, which is exactly what a governed platform produces. This is the case AnyInsight is designed for.

If you are a provider — you develop AI systems and place them on the market, or you rebrand or substantially modify a third-party system, which can make you the provider of it — you additionally carry design-time obligations that no runtime platform generates for you. A governance layer still covers the operational half: logging, monitoring, oversight, post-market observation. It does not write your technical documentation or run your conformity assessment.

Being clear about which side you are on is the single most useful hour of work available before you evaluate any tooling, because it determines what you are actually shopping for.

7. FAQ

Q1: We already have an AI policy and a DPIA. Is that not enough?

They are necessary and they are not evidence. A DPIA under GDPR Article 35 documents an assessment of processing risk; it does not record what your AI system did in production. The AI Act asks for both the assessment and the operational record, and the second is the one most organisations cannot produce on request.

Q2: The high-risk deadline moved to December 2027. Why start generating evidence now?

Because logs cannot be backdated. If an assessor in 2028 asks to see how oversight operated across the preceding period, the answer depends on whether you were logging in 2026. The deferral gives you time to build the programme; it does not give you time to accumulate the record afterwards.

Q3: Does this replace our SIEM or our network firewall?

No, and it addresses a different layer. A network firewall inspects packets at the boundary and has no visibility into the meaning of a prompt, the sensitivity of the data inside it, or the intent of an autonomous agent. The AI firewall sits inline between users, agents and models, and reasons about content. In practice the two work together, and the AI logs are a source your SIEM can consume.

Q4: Our AI is internal-only. Does any of this apply?

Probably. Internal use does not exempt you — what matters is whether the use case falls into an Annex III domain, and internal performance reviews and recruitment screening are the two most common places enterprises find themselves inside it. The Commission also tightened the business-to-business carve-out in the final Article 50 Guidelines, adding conditions including cloud isolation and role-based access control, which is a control question rather than a policy one.

Conclusion

The EU AI Act does not ask whether you intended to govern your AI. It asks you to produce the record. Article 50 asks it today; Articles 9 to 15 will ask it in December 2027; and enterprise procurement teams are already asking it on behalf of both.

The practical consequence is that AI governance is an infrastructure decision before it is a policy decision. Whatever tooling you choose, the test is the same: when someone asks what your AI and your agents did, can you show them — completely, attributably, and without a reconstruction project?

The Series, Start to Finish

This was Part 3 of The EU AI Act Playbook.

เริ่มสร้างด้วย Trusted AI วันนี้

สร้างบัญชี AnyInsight.ai ของคุณและใช้งานทดลองใช้ฟรี 14 วัน พร้อมเข้าถึงทุกฟีเจอร์อย่างเต็มรูปแบบ
เริ่มทดลองใช้ฟรี

เกี่ยวกับ AnyInsight.ai

AnyInsight.ai คือแพลตฟอร์ม AI workforce ที่ปลอดภัย ช่วยให้องค์กรสร้าง deploy และจัดการ AI Agents ได้โดยไม่ต้อง coding ขับเคลื่อนด้วยสถาปัตยกรรม zero-trust พร้อม access control, prompt injection protection, governance และ compliance ในตัว เพื่อให้องค์กรสามารถขยายการใช้งาน AI ได้อย่างมั่นใจ

Disclaimer

The insights and information shared in this article regarding the EU AI Act are for informational purposes only and do not constitute professional legal advice. We do not provide legal consulting services and assume no legal liability for any decisions made based on the content of this publication. As the interpretation and application of laws can vary depending on specific circumstances, we strongly recommend consulting a qualified legal advisor or attorney before making any compliance assessments or business decisions.

References

สำรวจต่อ

บทความที่เกี่ยวข้อง

ดูบทความทั้งหมด