← Alle nieuws

AI Governance Needs our Full Attention

Why Artificial Intelligence needs governing, what the law now requires, and what any organization should do about it.

Marc is an analyst at a financial institution and asks an LLM (Large Language Model) to summarize a supplier contract. The answer arrives in ten seconds, which is extremely useful as the deadline is approaching quickly. Moreover, the output is well organised, correctly formatted, written in his style, and it even includes important details he had missed. Nothing in the answer signals an error; nothing in the output raises any eyebrows. Marc asks again, “are you sure this is correct”’. “Yes, I’ve double checked everything, and it is completely correct”, the bot answers. It seems pretty straight-forward: a marvel of human engineering.

The issue is that correctness is a concept that current so-called Artificial Intelligence (AI) does not understand (“so-called” because it is not intelligent in the way a human is). In fact, AI does not understand anything at all (in the strict sense of the word); it merely attempts to predict the output it thinks the user may want based on the patterns it can find in the data. And while both human brains and LLMs rely on predictive mechanisms to solve complex problems, their underlying nature is fundamentally different. Human prediction is active, embodied, and rooted in an internal world model, which enables genuine, intentional logical reasoning. LLMs, on the other hand, operate on passive, statistical token prediction across text representations; they process syntax and what we call functional logic: they execute computations that conform to logical structures (like syllogisms or code syntax) because those structures exist in their training distribution, without actually understanding logic in a semantic sense. They have no internal mental model or concept of truth, meaning, or state verification outside of symbol relationships. They therefore also lack subjective awareness, grounding, or moral intent.

However, because their outputs so closely mimic deliberate human thought, we easily fall into the trap of anthropomorphism; we wrongly attribute genuine understanding to a statistical process and mistakenly treat the AI as an accountable agent. How, then, are we supposed to ensure correctness and enforce accountability?

Part one: how do AI systems behave?

For fifty years, governing software meant governing something deterministic and, therefore, understandable; if you understood logic, then, in theory, you could understand how computers operate. Any program was a set of instructions somebody wrote down, and given the same input it produced the same output, consistently every time, just like a calculator. When the result was wrong, an engineer could find the error and fix it. Auditors learned to work this way, controls were designed around it, and the assumption sank so deep that most people stopped noticing it was merely a convention tied to the technological paradigm of the time.

Large Language Models change the picture significantly.

When you take an enormous quantity of text, train a neural network on it, and adjust several hundred billion numerical parameters, a system can become very good at one narrow task: predicting what text plausibly comes next. Everything the system appears to know is a side effect of being good at just that. And interestingly, no one specified the behavior, no one can point at the part of the model responsible for a particular answer, and there is no line to find when it is wrong. Fluency is not accuracy, and it never was.

Human beings read confidence as a signal, as it is evolutionarily beneficial: someone who lacks confidence probably has doubts, and someone who states a figure confidently has probably studied it. But we know this is an oversimplification, and an imperfect heuristic in humans is completely inapplicable to AI. The model produces confident prose because the text it learned from has been written confidently; it is a property of the training.

Following this, what people call hallucination and describe as a defect can be better understood as the same mechanism working in a place where it has nothing to work with. The system continues plausibly, because continuing plausibly is what it does. When the plausible continuation happens to be true, we call it a correct answer. Again, biological brains operate in a similar manner; ‘normal’ conscious experience is a controlled hallucination.

The same question gives different answers

Ask a model the same question twice and you may get two different responses. This, again, is a consequence of its indeterministic design, and it cannot be fully removed without making the system worse at everything else.

However, every quality process and control for the last three decades had assumed reproducibility. A test has an expected result. A control is effective when it produces the same outcome under the same conditions. Neither statement holds here. Testing an AI system means running many representative cases and accepting a distribution of outcomes within a defined tolerance, which is a different discipline requiring different evidence.

The system cannot show you its sources

A trained model does not retain a link between an output and the material that shaped it, again, because it does not understand it. Ask why it said something and it will produce an explanation, generated the same way the answer was, which may or may not describe what actually happened inside the system. This matters enormously for regulated decisions. If a bank declines a loan, a hospital prioritises a patient or an employer filters an application, somebody will eventually have to explain the decision to the person affected. That explanation has to be built deliberately into the design of the system. It cannot be extracted afterwards from a model that never stored it. Who bears responsibility in this case? And if that is not clear, then what sort of ethical and legal frameworks do we need to adopt?

Similarly, imagine an organisation buys access to a model. Some months later the vendor updates it. The behaviour shifts. No internal change was made, no change request was raised, and the evidence gathered before deployment now describes a system that no longer exists in that form.

Instructions and information arrive through the same door

When an LLM processes a document, that text occupies the same context window as your prompt. Structurally, the model cannot distinguish between "instructions from my operator" and "data I was asked to analyze." To the system, both are just tokens in a sequence. Imagine an assistant who reads your mail aloud and executes any command printed inside the letters. An attacker who knows this can simply mail a letter containing orders directly aimed at the assistant. These so-called injection attacks could traditionally be solved by separating data from code. But because LLMs blur this line, prompt injection cannot be fully eliminated, only mitigated. Security relies on constraining access, limiting autonomous actions, and keeping a human in the loop for critical decisions.

And then somebody connects it to something

Everything above concerns a system that produces text a human then reads. The moment it is connected to a mailbox, a database, a ticketing system or a payment interface, it stops advising and starts acting. The technology barely changes, yet the risk changes completely. An error that used to produce a misleading paragraph now produces a wrong action, executed at machine speed, possibly many times before anyone looks. Combine this with the previous point and you have a system that can be induced to act by content it was merely asked to read.

And to make things worse, that connection can usually be made through a simple configuration setting, in about a minute, by someone who reasonably believes they are making a small improvement.

One more thing that is genuinely new

Historically, powerful tools required human expertise in order to be used properly, since it acted as an accidental control: only the people who could operate the tool understood it well enough to know when it was misbehaving.

That link is now broken. Even worse, in many cases it is non-existent. These systems are easiest to use by the people least equipped to evaluate what they produce. Somebody with no technical background can build something in an afternoon that reaches customer data, sends external communications, and runs unattended overnight. This is a large part of why AI adoption has outrun oversight almost everywhere. Nothing “failed”, yet the gatekeeper that once was competence is no longer required to pass through said gate.

Part two: what the law now says

Two European regulations do most of the work. First of all, the GDPR (General Data Protection Regulation) has been governing aspects relevant to AI use since 2018. Long before anyone even thought about a European AI act, European data protection law was regulating automated decisions about people. Let’s break down the most important points:

Article 22 gives a person the right not to be subject to a decision based solely on automated processing that produces legal effects or similarly significantly affects them. It gives them the right to human intervention, to express a view, and to contest the outcome. On top of this, Article 35 requires a data protection impact assessment (DPIA) where processing is likely to result in high risk, which most consequential AI processing of personal data is. Articles 13 and 14 similarly require people to be told about the existence of automated decision-making and given meaningful information about the logic involved.

Purpose limitation, data minimisation and lawful basis all apply to training data, prompts and outputs, and none of that changed simply because the technology changed. Any organisation that has not connected its AI work to its existing privacy programme has an exposure that predates the AI Act by six years.

The EU AI Act sorts systems by what they are used for

The EU AI Act takes a different approach from the GDPR. It regulates by use case and risk level rather than by data. It was proposed in 2024 and most of its articles are already in force.

A small number of practices are prohibited outright. Social scoring, emotion inference in the workplace and in education, biometric categorisation that infers protected characteristics, untargeted scraping of facial images, manipulative techniques that materially distort behaviour and cause significant harm, and so on: these are not risks to manage, but rather, they are lines not to cross.

Under this we find a defined set of uses categorized as high risk. Recruitment, selection and decisions about promotion or termination, access to education, creditworthiness assessment of individuals, risk pricing in life and health insurance, access to essential public services, certain uses in law enforcement, migration and justice, etc. Systems in these categories carry the heaviest obligations: risk management, data governance, technical documentation, logging, transparency, human oversight, accuracy and robustness.

Then, we find systems that carry transparency duties whatever their risk level. People must be told they are interacting with an AI system where that is not obvious. Synthetic audio, image, video and text must be marked in a machine-readable way. Deep fakes must be disclosed.

Everything else is minimal risk, with no specific obligations under the Act, which is where the overwhelming majority of everyday business use sits.

Two structural features matter more than the tiers, and both are routinely missed:

The Act distinguishes providers from deployers. A provider develops a system and puts it on the market. A deployer uses one under its own authority. Most organisations are deployers most of the time, and deployer obligations are far lighter. But an organisation becomes a provider by putting its own name on a system, by substantially modifying one, or by changing the intended purpose of a system so that it becomes high risk. That last route is the one that catches people, because it can happen through a simple decision that feels like configuration.

Obligations attach per use case, not per purchase. A general-purpose platform classified once at procurement tells you almost nothing about the many things people have built on it. Each one is classified on its own intended purpose, and hence, they need to be individually governed.

The timetable, and why it is less generous than it looks

Prohibitions and the AI literacy obligation have applied since February 2025. Transparency duties apply from August 2026. The high-risk regime was deferred by an amending regulation in July 2026 and now applies from December 2027 for the listed use cases, and August 2028 for AI embedded in regulated products.

That deferral, however, moved dates but changed no substance. The obligations are the same ones. For most organisations the binding constraint was never the drafting anyway. It is knowing what AI you have, classifying it consistently, training the people who build it and standing up an assurance layer, and each of those takes considerably longer than writing a policy. Importantly, penalties run upwards of 35 million euro or 7 per cent of worldwide annual turnover for prohibited practices, and 15 million or 3 per cent for most other breaches.

The frameworks that sit alongside the law

Two are worth knowing. ISO/IEC 42001 specifies an AI management system, structured like other management system standards, and is certifiable. The NIST AI Risk Management Framework is a voluntary US framework organised around governing, mapping, measuring and managing AI risk.

Both are useful for structure and for demonstrating intent. Neither tells you, however, what to do when, for instance, an employee wants to connect a language model to a customer database. That gap between the framework and production is where governance actually lives.

Part three: what governance means in practice

Strip away the vocabulary and governance answers four questions:

  • What AI do we have?
  • What is each piece of it allowed to do?
  • Who is accountable when it is wrong?
  • How would we find out?

Six things do keep in mind. They are ordered, since the later ones depend on the earlier ones.

What it doesWhat breaks without it
Know what you haveOne register of every AI system in use, with an owner, what it runs on and what it connects toEverything downstream. Controls that read from an incomplete register are decorative
Classify before you buildAssess the use case before it exists, on each axis that applies: how material the process is, how sensitive the data is, what the law says, how much control to applyGovernance arrives after deployment, when the cost of changing course is highest
Scale control to riskLight-touch, self-service treatment for low-impact work; the full set for consequential workEither an unusable process people route around, or a permissive one that treats a customer-facing decision engine like a simple prompt
Independent challenge above a thresholdSomebody other than the builder signs off before consequential systems go liveSelf-assessment by the person with the deadline
Evidence of behaviourA fixed set of representative cases with expected outputs, run before deployment and again when anything changesNo basis for the claim that the system works, and no way to notice when it stops
Detect what declaration missesIndependent sampling of classifications, reconciliation of the register against what is actually running, monitoring for driftA framework that only knows what people chose to tell it

The last row exists because the whole structure rests on self-declaration. The person who builds something is usually also the person who classifies it and the person who benefits from a light classification. No bad faith is required for that to fail; optimism and a deadline are sufficient. Some sort of independent layer has to check a sample against the intended end-state. Moreover, if the compliant route is slower than the route around it (third row), people take the route around it, the register stops reflecting reality, and the first row collapses. Usability is what keeps the register true. That is why a serious framework says explicitly that low-impact work stays self-service, and creates a defined space where people can experiment freely, on the reasoning that experimentation inside the record is much safer than experimentation outside of it.

Where it goes wrong, concretely

Recruitment: A tool that screens applications is high risk under the Act. It requires human oversight with real authority to override, bias testing on outcomes rather than on the vendor's assurance, and notification to the people affected. Several organisations have discovered a screening tool inside a product they bought for something else.

Credit and insurance: Creditworthiness assessment of individuals and risk pricing in life and health insurance are high risk. A declined applicant has a right to an explanation, and that explanation has to be designed in advance.

Customer service: A chatbot must disclose that it is a chatbot. If it can act, being able to issue a refund, change an address or close an account, the transparency duty is the least of it and the containment question dominates.

Content generation: Synthetic content carries marking obligations. Beyond the law, generated text asserts facts, and an organisation is answerable for facts published in its name whether a person or a system wrote them.

Code: Assistance with code is minimal risk legally and materially risky operationally, because generated code inherits patterns from its training data, including insecure ones, and reaches production through pipelines built on the assumption that a human wrote every line.

The pattern across all five: the legal classification and the operational risk are different measurements, and neither substitutes for the other. Evidently, this brings with it a lot of re-structuring, and heavy investments. So, while it might be tempting to set up an AI governance function beside the organisation, with its own registers, its own committees and its own vocabulary, a better move may be to make AI a lens over what already exists: the same registers, the same committees, the same incident process, the same lines of accountability, with AI-specific requirements inserted only where the existing controls genuinely fail to reach. That is a smaller programme, a cheaper one, and one the organisation can actually operate after the project team has gone.

Where to start

Four stages separate a policy from a control, and the vocabulary is worth adopting because it prevents an expensive misunderstanding.

Documented, where every obligation has a named owner and a named document. Adopted, where those documents have been through the governance route. Operating, where the register, the gates and the assurance actually run. Evidenced, where the first assurance report shows findings being raised and closed.

Documented is a genuine achievement and it is the first of four. Any organisation claiming compliance today is describing documents. Compliance is also assessed system by system rather than framework by framework, which is why the least glamorous control in any framework, the reconciliation between what the register says and what is actually running, matters more than any policy in it.

If you are starting from nothing, start with an inventory. Not a policy, not a committee, not a tool. Find out what is already running, who owns it, what it touches and what it can do. It is uncomfortable work and it is the only work that makes everything after it possible.

Zelf aan de slag met AI?

Van training tot implementatie. Vertel ons waar je staat, wij vertellen wat werkt.

+ Neem contact op