Daniel's question this week is about the moment when a technology stops being a toy and starts being a department. AI has moved from pilot projects to production systems, and the people who sign off on risk are now asking a very specific question: who is accountable when a machine makes a decision nobody fully scripted? Daniel wants to know where ISO 42001 fits into this, what the standard actually certifies, what makes AI risk genuinely different from sending data to a vendor, and what other ISO standards are emerging around it.
And the timing on this is good, because the answer to "who's accountable" is currently being written in real time by a committee in Geneva. That's not a metaphor. The standard exists. It's published. Organizations are getting certified against it right now.
So what exactly is this field, and why does a standard like 42001 exist?
AI compliance is a distinct discipline now, and it's not just data privacy with a chatbot bolted on. Privacy asks whether personal data is handled lawfully. Cybersecurity asks whether systems are protected from attack. AI compliance asks something stranger: how do you govern a system that makes decisions with some degree of autonomy, where the output isn't fully predictable, and where the traditional accountability chain breaks down because there's no single human who made the call?
That's the part that keeps risk officers up at night. A human employee can be trained, supervised, disciplined, fired. A model can be retrained, but you can't sit it down for a performance review.
Right. And that's exactly the gap ISO 42001 is trying to fill. It was published in December twenty twenty-three as the first certifiable management system standard for artificial intelligence. The key phrase there is "management system standard." It's built on the same architecture as ISO 27001 for information security or ISO 9001 for quality management. It gives an organization a framework to establish, implement, maintain, and continually improve what's called an AI management system.
I want to sit on that phrase for a second, because it's where most of the confusion starts. A management system standard is not a product certification. It does not say "this model is safe" or "this chatbot is trustworthy." It says "this organization has processes in place to manage AI risk." Those are wildly different claims.
Completely different. And I think that's the single most misunderstood thing about 42001. People hear "ISO certified AI" and picture a lab where an algorithm got a gold star. What actually happens is an auditor comes in, looks at your policies, your role assignments, your risk assessments, your documentation trails, and decides whether you've built a functioning governance machine around your AI. The AI itself could still be mediocre or even dangerous. The certification is about whether you've built the scaffolding to catch that.
Which is either deeply reassuring or deeply disappointing depending on what you wanted the certificate to mean.
I'd say it's the honest version. You cannot certify that a neural network is safe the way you certify a toaster. The toaster's behavior is fully specified. The model's isn't. So the only thing you can meaningfully certify is the process around it.
Let's get into what 42001 actually requires, and what it doesn't.
The core components are fairly recognizable if you've seen ISO 27001. The organization has to define AI policies. It has to assign roles and responsibilities, so there's a named person or team accountable for AI governance. It has to conduct risk assessments, implement controls, and document all of it. But the part that's new is the lifecycle requirement. The standard requires you to address AI risk from design through development, deployment, and eventually retirement.
Retirement is the one nobody thinks about. Everyone's focused on the shiny deployment moment, and the standard is saying: what happens when you turn this thing off? What happens to the data, the integrations, the downstream systems that learned to depend on it?
And that's where you can tell this standard was written by people who've watched systems rot. A model that was compliant at launch can become non-compliant as the world changes around it. The training data drifts. The regulations shift. The business context changes. 42001 forces you to treat AI as a living system with a full arc, not a one-time purchase.
Now let's get to Daniel's point about novel risk, because I think this is the heart of why AI compliance is its own field. He framed it as business information being sent to machines that have some degree of independence.
That's the crux. Traditional compliance assumes a human decision-maker somewhere in the loop. You send data to a vendor, a human at that vendor processes it according to a contract, and if something goes wrong you can point to the human, the contract, the training, the supervision. AI breaks that. When you send business information to a semi-independent system, the machine can act on that information in ways that weren't scripted in advance.
And it's not just about the machine making a wrong decision. It's that the machine might make a decision that was right in a narrow sense but wrong in a way nobody anticipated needed a rule. The governance gap is that you can't hold a model accountable. You can only hold the organization accountable.
Which is exactly what 42001 is designed to do. It requires organizations to document decision-making processes, maintain human oversight where it's needed, and establish clear escalation paths for when something goes wrong. The idea is to rebuild the accountability chain that AI severed, but at the organizational level rather than the individual level.
Give me a concrete example.
The classic one is a customer service AI that has access to customer personally identifiable information and can issue refunds or discounts autonomously. Under 42001, the organization would need documented policies for how that AI handles data, what it's permitted to do, what it's not permitted to do, how incidents are reported, who reviews the decisions, and what triggers a human takeover. The standard doesn't say "don't let the AI issue refunds." It says "if you let the AI issue refunds, you need a governance structure around that decision."
And the refund is the easy case because the failure mode is obvious: the AI gives away too much money. The harder cases are the ones where the harm is diffuse. The AI makes a credit decision that's subtly biased. The AI writes a summary of a legal document that's subtly wrong. The AI prioritizes a queue in a way that's subtly unfair. Those don't set off alarms. They just quietly compound.
And that's why the standard emphasizes risk assessment as an ongoing activity, not a one-time checkbox. You're supposed to be continuously asking: what could this system do that we didn't intend, and how would we know? The "how would we know" part is the one most organizations skip, because it requires monitoring and logging and actual human review, not just a policy document.
There's also the question of what "human oversight" actually means in practice. A human clicking "approve" on a thousand AI decisions a day is not oversight. That's a rubber stamp with a pulse.
The standard is deliberately vague on that, which is both a strength and a weakness. It says human oversight "where needed" but doesn't prescribe exactly what that looks like. The risk is that organizations implement the thinnest possible version of oversight to pass the audit, which is a problem we should come back to.
But 42001 doesn't exist in a vacuum. It's part of a whole family of standards being built out right now.
Right. And this is where the field gets interesting, because 42001 is the umbrella, but it's supported by a set of companion standards that do the heavy lifting on specific pieces. The three that matter most right now are ISO/IEC 23894, which provides AI risk management guidance and was actually published in February twenty twenty-three, before 42001 itself. Then there's ISO/IEC 42005, which offers guidance on AI system impact assessments. And ISO/IEC 42006, which sets requirements for the bodies that audit and certify organizations against 42001.
So 23894 is the "how to think about risk" companion, 42005 is the "how to assess impact" companion, and 42006 is the "how to verify compliance" companion. Together they form a governance stack.
The ordering is interesting. 23894 came first because the ISO recognized that before you can certify a management system, you need a shared vocabulary for what AI risk even is. The risk landscape for a predictive maintenance model in a factory is completely different from a generative model writing marketing copy. 23894 gives you the framework for thinking about risk across that spectrum.
And 42005 is the one that's going to get more attention as regulations start to bite. Impact assessments are becoming a legal requirement in some jurisdictions, and having a standard for how to do them is useful.
The EU AI Act is the big driver there. It requires impact assessments for high-risk AI systems, and organizations are looking for a recognized methodology. 42005 is the ISO's answer to that. It's still relatively new, but it's filling a real gap.
What about 42006? That one seems less glamorous but possibly the most important.
It is the most important, and here's why. 42001 certification is only as meaningful as the auditors who grant it. If any random consultancy can hand out a certificate, the whole thing becomes meaningless. 42006 sets requirements for certification bodies: they need to be accredited, they need to demonstrate competence in AI specifically, not just in generic management systems auditing. That's a hard problem, because auditing AI governance requires understanding both the technology and the governance frameworks.
So the certification ecosystem is still maturing. The standard exists, but the infrastructure for verifying compliance is being built out in real time.
And that's the honest state of the field right now. There are organizations getting certified, but the number of accredited certification bodies with genuine AI competence is small. It's a bottleneck. Which creates an interesting knock-on effect: the emergence of AI compliance as a profession.
Say more about that.
Three years ago, the job title "AI compliance officer" barely existed. Now it's a real role, and the skillset is hard to hire for. You need someone who understands machine learning well enough to ask intelligent questions about model behavior, and who also understands governance frameworks, audit methodology, and regulatory requirements. Those two skillsets rarely live in the same person.
It's the classic split. The engineers don't want to think about compliance, and the compliance people don't understand the technology. The people who bridge that gap are going to be very well paid.
They already are. And the demand is only growing as regulations like the EU AI Act start to bite. The Act has real teeth: fines for non-compliance can run into the tens of millions of euros depending on the severity. So organizations are scrambling to demonstrate that they take AI governance seriously, and ISO 42001 certification is becoming one of the primary ways to do that.
Which brings us to the tension at the heart of this whole field. ISO standards are consensus-based and slow-moving. They take years to develop, and they're designed to be stable. AI is evolving on a scale of months. A standard published in twenty twenty-three may already feel dated by the time organizations are getting certified against it.
And yet it's still the best framework we have. That's the uncomfortable truth. The standard is imperfect, it's already aging, and the field is moving faster than the standards body can keep up. But the alternative is no framework at all, and that's worse. A flawed standard that organizations actually use beats a perfect standard that doesn't exist.
I think that's right, but I want to push on one thing. There's a risk that 42001 becomes a box-ticking exercise, a certificate on the website that says "we take AI seriously" without actually changing how the organization operates. That's the history of ISO 9001 in a lot of industries.
It's a real risk. And it's the thing I worry about most with this standard. The gap between having a documented AI management system and actually managing AI risk is enormous. An auditor can verify that the documents exist. It's much harder to verify that anyone reads them.
The documents are the easy part. The hard part is building a culture where someone feels empowered to say "this AI system is doing something weird and I think we should stop using it."
And that's not something any standard can mandate. It's a leadership problem. The standard gives you the structure, but the structure only works if the people inside it actually care.
Practical question for a business listening to this: is 42001 certification worth pursuing right now?
Depends on what you're trying to achieve. If you're selling AI services to enterprises, certification is becoming a competitive differentiator. It signals to customers and partners that you've built governance infrastructure around your AI, which matters when they're doing their own due diligence. If you're an enterprise buying AI services, asking your vendors about their 42001 status is a reasonable part of procurement now.
And if you're a small business just trying to use a chatbot for customer support?
Full certification is probably overkill. But the framework itself is useful even if you never pursue the certificate. The questions it forces you to ask: who's accountable, what are the risks, what's the escalation path, what happens when we retire this system. Those are worth asking regardless of size.
So the standard has value as a thinking tool even without the audit.
And that's true of most management system standards. The certification is the stick. The framework is the actual value.
What's the timeline look like for the companion standards? Are they all published, or are some still in development?
23894 is published, as I mentioned. 42005 and 42006 are in various stages of development and publication. The ISO moves deliberately, and AI standards are particularly contentious because the technology is changing so fast. There's active debate about whether the current versions are already too narrow.
Too narrow how?
The standards were largely written with a certain kind of AI in mind: supervised learning models making predictions in business contexts. Generative AI, autonomous agents, systems that can take actions across multiple platforms. Those are stretching the framework in ways the original authors didn't fully anticipate. An AI agent that can browse the web, send emails, and make purchases is a very different governance problem than a fraud detection model.
That's the frontier. The standards are being written for the AI we had, while the AI we're getting is more independent, more connected, more capable of acting in the world. The governance gap is widening even as the standards are being published.
And that's why the field of AI compliance is going to be one of the more interesting places to watch over the next few years. The standards are a first draft. They'll be revised, extended, replaced. But the fundamental problem they're trying to solve isn't going away: how do you govern systems that make decisions without fully predictable outcomes?
I keep thinking about the accountability question. We started with that. Who's accountable when a machine makes a decision nobody fully scripted? The answer 42001 gives is: the organization is accountable, through its management system. That's unsatisfying in a way, because it doesn't give you a person to blame. But it's also the only honest answer.
The person to blame is the person who was supposed to build the management system and didn't. Or built it badly. That's what the standard is really about: creating a structure where accountability can actually land somewhere.
And the structure has to be real, not just documented. That's the thing I keep coming back to. A policy document that nobody reads is not governance. It's a liability shield, and a thin one.
The auditors are supposed to catch that. Whether they actually do is the open question. And that's where 42006 matters so much. If the certification bodies are rigorous, the standard means something. If they're not, it becomes theater.
I suspect the answer is going to vary a lot by auditor and by region. Some certification bodies will be rigorous, and some will be rubber stamps. The market will sort that out eventually, but in the meantime, a 42001 certificate is going to mean different things depending on who issued it.
That's already true. I've seen certificates issued by bodies with very different levels of AI competence. Some of them are just generic management systems auditors who took a weekend course on AI. Others have actual technical staff who can ask hard questions about model behavior. The certificate looks the same, but the assurance is very different.
So Daniel's question about what 42001 actually certifies has a layered answer. It certifies that an organization has a documented AI management system. But the value of that certification depends on the rigor of the auditor, the genuine engagement of the organization, and the evolving landscape of companion standards that fill in the details.
And the whole thing is a moving target, because the technology is moving faster than the governance. We're building the safety rails for a vehicle that's still being designed.
That's a decent summary of where we are. Let me ask you something that's been nagging at me. Do you think 42001 is actually going to become the de facto standard for AI governance, or is it going to be superseded by regulation like the EU AI Act?
I think it's going to be both. The EU AI Act is the law, and it's binding. But laws tend to say what you must do, not how to do it. Standards like 42001 fill the "how" gap. My guess is we'll see a convergence, where demonstrating compliance with 42001 becomes one of the recognized ways to show you're meeting your obligations under the AI Act. That's already starting to happen in some sectors.
So the standard and the regulation are complementary, not competing.
In theory. In practice, there's friction, because the standard and the regulation were written by different bodies with different timelines and different assumptions. Organizations are stuck trying to satisfy both, and they don't always align perfectly. But that's the normal mess of compliance. It's never clean.
One more thing I want to touch on before we wrap the main discussion. Daniel mentioned "best practices in governance in a very new field.How new is this, really?
The field of AI compliance as a distinct discipline is maybe three or four years old. The standards are even newer. 23894 was published in February twenty twenty-three, 42001 in December twenty twenty-three. We're in the absolute infancy of this. The people who are working in this field right now are, in a very real sense, the ones defining what it will become.
Which means the standards we're discussing today are the first draft of how we'll govern machine decision-making for decades. That's a heavy thought.
It is. And it's why the choices made now matter so much. If the first generation of certification is rigorous, it sets a high bar that everything else builds on. If it's theater, we'll spend years trying to claw back credibility.
The stakes are higher than they look. A standard that exists but isn't taken seriously is worse than no standard, because it creates the illusion of governance without the substance.
That's the nightmare scenario. A world where every AI vendor has a 42001 certificate on their website and none of them have actually changed how they operate. The certificate becomes a marketing asset, not a governance tool.
And the way you prevent that is exactly what 42006 is trying to do: make the certification bodies accountable. But that's a slow process, and the technology isn't waiting.
It never does.
No, it doesn't.
Hilbert: The auditor took the lunch.
Sorry?
Hilbert: Nineteen ninety-seven. I was doing quality assurance for a software company in Hartford. Small outfit, forty people, we made inventory management software for auto parts distributors. The owners decided they wanted ISO 9001 certification because one of their big customers asked for it. So they hired a consultant, wrote a bunch of process documents, and scheduled the audit.
And the auditor took the lunch.
Hilbert: The auditor was more interested in the restaurant than the software. We spent three days showing him documents, and he spent three days asking where we were taking him for lunch. The audit passed. We got the certificate. Framed it, put it in the lobby. Nothing about how we actually built software changed.
That's the certification theater problem in a nutshell.
Hilbert: The thing is, the documents we wrote were good. I wrote some of them myself. They described how the company should work. But nobody followed them, because following them was slower, and the customers didn't care. They just wanted the certificate on the website.
That's exactly the risk with 42001. The documents can be excellent and the practice can be completely disconnected.
Hilbert: The auditor had a business card. I still have it somewhere. He was a nice guy. Knew nothing about software, but he knew a lot about restaurants.
And that's why 42006 exists. To make sure the people auditing AI actually know something about AI.
Hilbert: I read the summary of 42006. It says certification bodies need to demonstrate competence in artificial intelligence. That's the right idea. But competence is hard to verify. My auditor was competent at lunch. That's a competence. Just not the one we needed.
The difference is that AI auditing requires a level of technical understanding that generic management systems auditing never did. You can audit a quality management system without understanding the product. You cannot audit an AI management system without understanding what a model is, what training data is, what a hallucination is. The auditor has to be able to ask technical questions and understand the answers.
Hilbert: My auditor couldn't have asked a technical question about our software if his life depended on it. But he could tell you the difference between a good veal piccata and a bad one.
The question is whether the certification bodies actually invest in that technical competence, or whether they just rebadge their existing auditors and send them out.
Hilbert: Some of them will do the second thing. That's what happened with 9001. The certification bodies had a business model, and the business model was selling certificates. The lunch was part of the sales process.
That's the cynical take, but I think there's a real difference this time. The consequences of AI failure are more visible and more severe. A bad quality management system produces a bad product. A bad AI management system produces a system that makes discriminatory decisions or leaks sensitive data or takes actions that cause real harm. The stakes are higher, and the scrutiny is higher.
Hilbert: The stakes were high for the auto parts distributors too. If our software failed, their inventory was wrong, and they lost money. But nobody framed the certificate and then asked whether the software worked. They just wanted the certificate.
What would you do differently, if you were setting up the certification regime for 42001?
Hilbert: I'd make the auditors demonstrate that they understand the technology. Not just take a course. Actually understand it. And I'd do surprise audits. Not the scheduled kind where everyone cleans up for three days before the auditor arrives. Show up unannounced and see what's actually happening.
Surprise audits are a good idea, but they're expensive, and the certification bodies don't have the incentive to do them. It cuts into their margins.
Hilbert: That's the problem. The incentive is to sell certificates, not to provide assurance. As long as that's true, the certificate is worth less than the paper it's printed on.
You're not wrong. But I think there's a countervailing force now that didn't exist in nineteen ninety-seven. The customers are asking harder questions. They're not just checking for the certificate. They're asking to see the risk assessments, the incident reports, the actual governance artifacts. The market is getting more sophisticated.
Hilbert: Maybe. Or maybe they're just asking for a different certificate.
That's the pessimistic view, but it's not unfounded. The history of compliance is full of examples where the certificate became the product and the actual governance was an afterthought.
Hilbert: I'm not saying don't do the standard. I'm saying the standard is only as good as the people enforcing it. And the people enforcing it are the same people who were enforcing 9001, with the same incentives. Some of them will do it well. Some of them will take the lunch.
The market won't know the difference until something goes wrong.
Hilbert: That's usually how it works.
The thing that gives me some hope is that AI failures are more public than quality management failures. When an AI system makes a bad decision, it tends to end up in the news. That creates pressure on the certification bodies to actually do their jobs, because their reputations are on the line too.
Hilbert: Maybe. Or maybe the company just blames the AI and the certificate stays on the wall. I've seen that too.
The certificate as a shield rather than a signal. That's the risk.
Hilbert: It is. Anyway. That's what I wanted to say. The standard's fine. The auditors are the weak point. Same as it ever was.
You know, the more I think about it, the more I think the real test of 42001 won't be whether organizations get certified. It'll be whether a certified organization ever has a serious AI incident, and what happens to the certificate after.
That's the empirical question. We don't have enough history yet to answer it. The standard is too new, and the incident data is too thin.
Give it two or three years. We'll know by then whether the certificate means anything.
If it doesn't, the whole field of AI compliance has a credibility problem. Which would be a shame, because the underlying problem is real.
The problem is real, and it's not going away. The question is whether the governance infrastructure we're building is up to the task. I think 42001 is a genuine attempt to answer that question. Whether it succeeds is a different matter.
The misconception I want to leave people with is the one we flagged at the start. Most people think ISO 42001 certifies that an AI system is safe or trustworthy. It doesn't. It certifies that an organization has a management system for AI risks. The difference is the entire point.
Right. The certificate is about the organization, not the model. And the value of that certificate depends entirely on whether the organization actually uses the system it documented. A certificate on a wall is not governance. It's decoration.
The field of AI compliance is being built right now, and the standards we've been discussing are the first draft. They'll evolve. They'll be replaced. But the fundamental question, who's accountable when a machine makes a decision nobody fully scripted, that's going to be with us for a long time.
The answer, whatever form it takes, will be built on the scaffolding these standards provide. Imperfect, incomplete, but real.
Thanks to our producer Hilbert Flumingtop for keeping the show running. This has been My Weird Prompts. If you want to share your own weird prompt, email the show at show at my weird prompts dot com.
We'll be back soon.