The PolicySight blog

Can AI-built software be enterprise-grade?

A question I'd ask too

AI writes the code for all of our applications. I make business software but I'm not a software engineer, I'm a product manager. I've said all of this publicly, so the same question comes up in almost every conversation. Can software built that way really be enterprise-grade?

The instinct says no. Everyone has seen what unsupervised AI produces, and nobody wants it anywhere near anything that matters.

I think the instinct is aiming at the wrong target. In every one of those horror stories, the AI was left to work alone. Nobody gave it a plan. Nobody guided it as it went. Nobody redirected it when it wandered off course, and nobody made sure what it produced was what was actually needed.

I can't read the code either

There's a harder version of the question, and I ask it myself. I can't check the code. I'm not a coder; I couldn't tell a right line from a wrong one by reading it. So how do I know the software is correct?

The honest answer starts with a surprise: software companies don't know their code is correct by reading it all either. Nobody at Microsoft has read the whole of Windows. What they do instead will feel familiar to anyone who has ever approved the minutes of the last meeting. In a well-run software firm, every change is checked at the moment it's made, by someone who didn't make it, while it's still small enough to check. The whole system is far too big for anyone to read, and it doesn't need to be, because nothing got in without passing that ritual. Twenty years of committee minutes earn trust the same way: not because anyone re-reads the archive, but because every page was approved by the people in the room while it was fresh.

The same structure sits around the AI here. Every change it writes is checked before it ships, by a reviewer that didn't write it. The difference is that those reviewers are other AIs, and one of them comes from a different model family, so it doesn't share the same blind spots. Nothing ships until I've signed it off. My job is the chair's: say what correct means before anything is built, govern the controls, and test those controls, so a problem surfaces loudly. Automated checks act as real users and verify what the software does, and new controls are broken on purpose to prove the alarms fire. I trust it the way you trust a cloud drive with every family photo you own: you've never inspected the data centre, you're trusting the controls behind it. The only difference is that you can ask to see mine.

Nobody ever audited the typist

Enterprise-grade was never a claim about who wrote the software. Security questionnaires ask how changes are reviewed, how access is controlled, how releases are tested, and how the people involved are vetted and trained. They never ask which human typed each line, or whether they were having a good week. Buyers ask about process, because process is the only thing that scales past trusting individuals.

Anyone who works in governance already knows this. If you sit on a committee, you rely on minutes you didn't write, accounts you didn't prepare and reports you didn't compile. The reason you can is the control environment around them: independent review, evidence, a named person accountable at the end.

Banks and auditors have run the sharpest versions of this for decades, as two controls with different jobs. Segregation of duties is aimed at fraud: the work is split into separate roles performed by different people, the person who sets up a payment is never the person who authorises it, and nobody controls the whole chain on their own. The four-eye check is aimed at error: a second person goes over the work itself, because a fresh pair of eyes catches what the writer no longer sees. Both rest on the same discovery: you don't get safety by hiring perfect people. You get it by never letting one pair of hands, or one pair of eyes, be enough.

So the useful question about AI-built software is the question both controls ask: what does the code have to pass, and who answers for it?

What the gates look like

The narrow claim I'm prepared to defend is this: software built with AI can be enterprise-grade when the AI is never left to work alone. Planned, guided, redirected and checked, inside a control environment an auditor would recognise.

In our case, segregation of duties describes the build itself. No single AI does the whole job. Every application we make is built by the same team of ten specialist AI agents, each doing one part, and none of them ever approves its own work. An orchestrator coordinates them and isn't allowed to write code at all.

Here's what happens to every change, in order. The numbers come from our first application, the one that's live today.

  1. A written brief comes first. A domain expert agent sets out what's needed, an architect agent turns that into a technical specification, and a platform specialist checks it can be built. No code is written without it.
  2. The code is written. A coding agent builds only what the specification says.
  3. The code is reviewed. A reviewer agent checks every change, and a security reviewer checks every change that touches anything sensitive. That security review has caught real problems, including an exposed key found days before a release. Since September, a second AI from a different model family also reviews the work, so the last pair of eyes isn't the same kind of brain that wrote it.
  4. First kind of testing: automated checks. A suite of scripted checks runs the same way every time, so nothing that worked yesterday breaks unnoticed today. For version 1.2 that was 1,343 checks at the release gate, passed with zero failures on 12 August 2026, plus over 860 more that drive the application through its screens in ordinary and admin roles. The checks are tested too: a new control is broken on purpose to prove the right check catches it, then put back.
  5. Second kind of testing: an AI using it like a person. That same second AI controls a computer, signs in as each of our eight test users, one for each role, and works through our full catalogue of user journeys, clicking the buttons itself rather than following a script. It catches what a script was never written to look for. On its first full pass, in September, it found real faults, and each one was checked and fixed.
  6. A human signs it off. Every release passes eight named gates, with evidence recorded at each, and the last one is a human go that no machine can give itself. That's me. Version 1.2 completed all eight on 13 August 2026.

And once, outside eyes: Microsoft reviewed that first application before listing it on its marketplace, and it passed certification at the first attempt. The listing went live in July 2026.

Trust comes from the gates, not the writer.

The same stamps, in the same order, for every change.

The boundaries are part of the method

Two things I volunteer before anyone asks, because receipts that have to be prised out of someone are worth very little.

There has been no independent penetration test yet. One is scoped and planned, and I'd rather say that plainly than have it discovered.

And it's new. A deep test estate, but not years of production behind it yet. The discipline is the substitute for the years until the years exist.

I'd argue the volunteering is itself part of the answer. A control environment that hides its own gaps isn't one, whoever writes the code.

What I'd ask any vendor

None of this is special to me, and that's the point. If you're weighing up software you suspect was built with AI, and increasingly that is most software, don't ask whether they use AI. Ask what the AI's work has to pass. Ask who reviews it, whether the reviewer is independent of the writer, whether the tests themselves are tested, and which human puts their name on the release. If the answer is a variation on "we check it when we get time", you have your answer.

And if you've worked out a better way to trust work you didn't do yourself, I'd like to hear it. Drop me a message.