Compliance programmes are built to reach a signature. The duties that decide whether a regulated AI system stays lawful are the ones that only begin once it is answering real traffic: monitoring that has to keep running, logs that have to keep accumulating, declared performance that has to keep holding, incidents timed from the moment somebody understood, and a modification rule that a routine model upgrade can trip. None of it carries a completion date, and none of it can be produced afterwards by anyone who did not capture it while the system ran.
Each of these is satisfied by something the system produces while it serves traffic, or it is not satisfied at all. They are also the three that look finished on launch day, which is precisely why they are the ones an audit finds wide open two years later.
Article 12 wants events recorded automatically over the lifetime of a system, and Article 17 puts their retention inside the quality management system. Two things go wrong reliably. A window sized against storage cost in 2026 decides what still exists in 2028, and the record holds sensitive input, so it needs a classification, an access rule and a lawful basis of its own.
Article 15 asks for accuracy, robustness and cybersecurity across the lifecycle rather than at the moment of assessment, which is a stronger demand than it sounds. Declared metrics sit in the instructions for use while live traffic drifts away from the distribution they were measured on, and the declaration quietly stops being true while the documentation stays signed and valid.
Article 73 runs from the point a causal link is established, and from awareness before that, which makes awareness the field nobody writes down. The trigger is as likely to be a complaint, a journalist or your own evaluation run as an alert, so the question put to you later is always who first understood what had happened, and when.
Settle what this system is measured on in production, what an out-of-range reading triggers and who reads the numbers weekly. The dated note of what each reading changed is the part an assessor asks for.
Model and version, the resolved prompt, retrieved context with document identifiers, every tool call and its result, the output returned, and the human decision that followed it.
One mandatory field asking whether this change touches intended purpose, model version, training data, decision scope or the action set, with a yes routed to a reviewer who can hold the merge.
A dated record tying the live version to its evaluation results, prompt and model pinned together so a rollback restores both, the logging configuration in force, and whoever approved it.
The ones weighted after entry into service: post-market monitoring under Article 72, automatic logging under Articles 12 and 17, accuracy, robustness and cybersecurity held across the lifecycle under Article 15, and serious incident reporting under Article 73. Not one of them has a point at which it is finished.
No. The plan is a document and the duty is the data behind it, collected, documented and analysed across the operational life of the system. What gets requested is the series: drift, quality against declared metrics, override rates, complaint volume, and a record of what each signal changed and who acted on it.
Yes, and this is the version of the rule teams walk into rather than choose. A substantial modification needs a fresh assessment, and changing the base model, the training data, the decision scope or the action set are all candidates. None of them announces itself as a regulatory event inside a pull request.
From awareness, which is the timestamp nobody captures. Article 73 reporting runs from establishing a causal link, the GDPR gives 72 hours from awareness of a breach, NIS2 wants an early warning inside 24 hours, and DORA an initial notification four hours after classification. One event, four different starting guns.
Switch on the two records nobody can produce later: full trace and event capture with retention set against the longest duty the system carries, and a monitoring series with tuned thresholds and someone reading it every week. Each covers only the stretch since it started, so a quarter of delay is a quarter you will never have.
Read this one for what the law asks of a system already in service, and that one for the machinery that answers it. Production AI: An Engineering Reference covers evaluation, release gating, rollback, drift, cost and end-of-life as operational controls. The duties here are the reason several of those controls stop being optional.
Heads of engineering, CISOs and heads of risk who own a system that is already serving traffic. It assumes the approval work is behind you and asks the harder question, which is what your release process and your telemetry would show if somebody examined both of them this quarter.
Drop your details and we'll send Production AI Under Regulation straight to your inbox - no spam, unsubscribe anytime.
Bring one system already serving traffic and we will map the duties running against it today, the evidence it currently produces by itself, and the gaps that no amount of later effort can close. Two hours with engineers. SECTION 7 - FAQ - 5 to 8 questions
Book a production review