AI Can Build Your App. It Can't Make It Safe to Sell

AI can build your app, but it can't make it safe to sell. A lot of founders with a vibe-coded or no-code product hit this gap: the demo works and users sign up, and then the first serious business customer asks a few procurement questions and the launch date slips. Getting the prototype working was the easy part. AI tools don't build the security and compliance work that business buyers check for, and regulated UK buyers check hardest.

Updated 22 September 2026. I corrected several things in the July version. Some statistics didn't match their sources, I'd given the wrong date for the EU AI Act's transparency rules and the wrong level of fine, and I'd said the European Accessibility Act requires WCAG 2.2, which isn't yet the case. The last section now describes how we start with founders today.

We use these tools too. They let someone who knows an industry inside out turn an idea into a working product in days, for a fraction of a traditional build. Finding out whether people want the thing is the hardest part of starting a software business, and it's now close to free.

Why a working demo doesn't close an enterprise deal

We see the same pattern again and again. Someone with deep industry knowledge builds a tool that solves a real, expensive problem. It demos well. A big customer likes it, then sends a security questionnaire, asks for a penetration test report, or wants to know where the data is hosted. The founder has no answers, because the tool they built with never asked those questions.

Earlier this year a founder brought us exactly this problem. They'd built their app with a no-code tool and planned to launch it on an enterprise software marketplace. The marketplace's security and data privacy checklist delayed the launch by about three months.

Diagram of what buyers check before they sign: an AI builder gets you a validated idea, a working demo and first users, and procurement then checks for multi-tenant data isolation, security certifications, AI safety guardrails, UK data residency, audit logging and penetration testing, which AI won't build for you.

Who's liable when the AI gets it wrong?

If your app uses an AI model to generate content a person relies on, and that content is safety-critical or legally significant, start here. It's the biggest risk in this post, and founders rarely spot it.

Say your app uses a large language model to write method statements, the documents that tell a worker how to do a dangerous job safely. Language models predict plausible text without checking facts. Stanford researchers tested the AI legal research tools sold by LexisNexis and Thomson Reuters, which are built specifically to be accurate, and found they hallucinated between 17% and 33% of the time (Journal of Empirical Legal Studies, 2025). An earlier study from the same group found general-purpose chatbots got specific legal questions wrong between 58% and 88% of the time.

Apply that to a method statement for a high-voltage job. If the model invents a clearance distance or drops a control, the document still looks finished, and it's wrong. The duty holder is still responsible for that document, whatever produced it. Courts have already seen this happen: Damien Charlotin's database lists more than 2,000 court decisions where a party relied on hallucinated AI material, most often invented cases or quotes, and over 800 of them involve lawyers. In one US case the penalties on two lawyers, including the other side's costs, passed $100,000.

Decision flow for whether AI-generated output can reach a customer: safety-critical or legally binding output must pass three gates (grounding in verified source data, meaningful human review with logged sign-off, and disclosure that content is AI-generated) before it is issued; raw AI output with no gates is a liability event

What do AI guardrails look like in a real product?

A commercial product puts controls around the AI, and no-code builders don't add them for you.

A person should sign off anything consequential. A safety document isn't issued automatically: the software makes a named, logged-in person review and approve it before it carries legal weight, and records that approval. The ICO's guidance on automated decisions says a human reviewer has to have the authority and ability to change the outcome. A reviewer who only rubber-stamps it doesn't count.

Tie the model to verified source material, such as official guidance and your own vetted templates, and make it decline when it can't answer from that material.

Prompt injection is the first entry on OWASP's Top 10 for LLM Applications, in both the 2023 and 2025 lists, and OWASP says it's unclear whether any method prevents it completely. If you pass user text to a model, someone can try to hijack it into leaking another customer's data. You reduce the risk with input and output filtering, giving the model access to as little as possible, and human approval for anything sensitive.

Log what the AI generated, from which inputs, and who approved it, and tell users when content is AI-generated.

If your app is used in the EU, some of this is already law. Article 50 of the EU AI Act has applied since 2 August 2026: people have to be told when they're dealing with an AI unless it's obvious, and providers of systems that generate content have to mark it as AI-generated, with exceptions such as standard editing help. If you build on someone else's model, check which of you that duty falls on. Generative systems already on the market before August have until 2 December 2026 to add that marking. If your use case is on the Act's Annex III high-risk list, human oversight, record-keeping and a conformity assessment apply from 2 December 2027. Breaking those rules can cost up to €15 million or 3% of worldwide turnover, or the lower of the two for a small business. The UK has no equivalent Act, but UK GDPR and your sector regulator still apply.

Is the code an AI wrote secure?

Often it isn't: Veracode's 2025 GenAI Code Security Report tested more than 100 models and found 45% of code samples failed security tests and introduced OWASP Top 10 vulnerabilities, with no improvement from newer or larger models. A USENIX Security 2025 study of 576,000 code samples from 16 models found 19.7% of the packages they referenced don't exist. Attackers can register those invented names with malware inside (known as slopsquatting), so installing an import an AI suggested can pull malware into your app. Live apps look no better: in October 2025 Escape.tech scanned more than 5,600 public apps built with vibe-coding tools and found more than 2,000 vulnerabilities, more than 400 exposed secrets and 175 cases of exposed personal data, including medical records and bank details.

One case has a CVE number. In 2025 a researcher reported CVE-2025-48757 against a popular AI app builder. The record in the US National Vulnerability Database says weak row level security policies let attackers read or write the database tables of apps it generated without logging in. The record also notes that the vendor disputes the CVE, because it says each customer is responsible for protecting their own app's data. The researcher found the problem in 170 of the 1,645 apps checked, across 303 endpoints. The record covers versions up to April 2025. Builders have added security checks since, and in the apps we see the gap now tends to open later. On apps backed by Supabase, the key in the browser is meant to be public, and it's only safe while row level security is switched on and correctly written for every table. In the apps we review, it usually is for the tables the builder created at the start and often isn't for the table you added three weeks in.

Diagram of a database where the tables created on day one have row level security switched on and a table added later does not, so the public key can read it.

What UK enterprise buyers ask for before they sign

Selling to a UK enterprise, a utility, an NHS body or a government department means getting through procurement, and for many buyers it's pass or fail. If you can't produce the paperwork, the deal stops. Roughly in the order it comes up:

A vendor security questionnaire covering encryption, access control, backups and incident response. Without documented answers, you don't get past it.

A recent penetration test report from an independent testing firm. In our experience buyers want one from the last twelve months, often from a CREST-accredited firm, and automated scanner output doesn't count.

Cyber Essentials, often Cyber Essentials Plus. The UK government has required Cyber Essentials, or equivalent controls, for certain public contracts since 2014. Under the current policy note, PPN 014, that covers central government and NHS contracts that handle citizens' personal data or OFFICIAL information. For assessments started on or after 27 April 2026, multi-factor authentication is mandatory on every cloud service that offers it, and missing it fails the assessment automatically. A founder's admin login to a cloud service without MFA fails straight away.

ISO 27001 or SOC 2 Type II as deal sizes grow. Both take months, because auditors look for evidence of how you've worked over a period of time. Our guide to reading an enterprise compliance checklist goes through each item.

Where your data lives. Many UK buyers ask for UK hosting, and some AI builders host in the US by default, so check where yours is before a buyer asks.

Certifications audit the foundations underneath the app, which a prototype usually hasn't got yet, so leave time before the pitch.

Which laws does your app inherit?

A real product also has to follow the law of its sector, and AI builders don't warn you about it. Under UK GDPR, which the Data (Use and Access) Act 2025 amended with main changes taking effect on 5 February 2026, your app has to handle erasure requests and, where it applies, data portability. You have to report a notifiable breach to the ICO within 72 hours, and you need a written contract with every processor that handles personal data for you. A prototype with one login and one flat database often can't cleanly delete one customer's data, and rarely has the logging to spot a breach, let alone report it in three days.

Then there's the sector. Take construction, where risk assessment is a legal duty under the Management of Health and Safety at Work Regulations 1999 and CDM 2015 requires a construction phase plan for every project. The Health and Safety Executive completed 246 prosecutions in 2024/25 with a 96% conviction rate and more than £33 million in fines. When your software writes a document that meets a legal duty, it takes on real legal exposure, and every regulated sector has its own version of this.

If you sell to consumers in the EU, the European Accessibility Act has applied since 28 June 2025. It covers selling any product or service to consumers through a website or app, and banking. Apps sold only to businesses aren't covered, and microenterprises providing services (fewer than 10 staff, and turnover or balance sheet of €2 million or less) are exempt. The standard it points to, EN 301 549, is based on WCAG 2.1 AA today. ETSI published a version based on WCAG 2.2 in September 2026, but it isn't yet the legal reference.

How do you keep each customer's data separate?

Most of the engineering goes into multi-tenancy, which means keeping one customer's data away from another's. If the separation depends on the app's own code remembering to filter every query, one bug shows Customer A the records of Customer B. Buyers test for exactly that. Real products enforce the separation in the database, and there are three common ways to do it, trading cost against isolation:

A shared database with row level security. Everyone shares tables, and the database decides who sees each row. It suits most B2B products: low cost, with strong isolation when it's configured correctly.

A separate schema per customer. One database server, with each customer in their own compartment. It sits between the other two on cost and separation.

A separate database per customer. Each customer gets a physically separate database. It's for banks, defence and customers with the strictest compliance needs.

Around that sits the rest: UK-region hosting, an encrypted managed database, MFA and single sign-on, a web application firewall, secrets kept out of the browser, monitoring and audit logging. It's a known programme of work that usually takes months, and a lot of what you've built carries over.

What carries over from your prototype?

The hard part of a software business is knowing what to build and proving people want it, and your demo has already done that. Your domain logic, screens and business rules carry forward. Whether the layer underneath needs a partial migration or a full rebuild depends on the app: about four in ten projects we see need only partial migration, and the rest need a full rebuild. AI-assisted tooling has cut our own rebuild costs by 20-30%, so a rebuild costs less than it did. Sort out separation between customers early, because every new customer adds data to the tables you'd have to change.

Start with a free check. Tell us about your app through a short form, and a person reads it and replies in writing with what we'd look at first. If there's more to find, a £95 review goes deeper: whether you can charge for the app, whether your business rules hold, whether search engines and AI assistants can find it, and what you own. The £95 is credited in full against any paid work we do next.

Work to meet an enterprise buyer's checklist, such as separating customers' data, splitting environments or adding audit logging, is scoped in the review and quoted at a fixed price. So is a full native or cross-platform rebuild. Certification audits and penetration tests come from independent firms and are paid for separately.

If your app mainly needs to get to production, that's a fixed £999 to £1,999, depending on complexity: fixes, security, moving your backend, in-app payments, store submission, or a wrapper route to the stores. Our Vibe Code to Production page sets out the whole route, and if you built on Lovable or Base44, those pages cover what's specific to each. Or get a free check and tell us where your app stands today.

Frequently Asked Questions

Is my vibe-coded app secure enough to sell to businesses?

Probably not without a review, and that's no criticism of you. Veracode found 45% of AI-generated code samples introduced OWASP Top 10 vulnerabilities (2025), and a scan by Escape.tech found more than 400 exposed secrets across more than 5,600 live vibe-coded apps. The usual problems are a table added after launch with row level security switched off, a secret key in the front end, weak separation between customers' data and no audit logging, and they're exactly what enterprise buyers test for. The free check tells you what to look at first.

Do I have to rebuild my whole app to make it production-ready?

Not always. About four in ten projects we see need only partial migration, and the rest need a full rebuild. Either way your domain logic, screens, business rules and the demand you've proved carry forward. What most often needs replacing is underneath: the backend, data model, sign-in and hosting, because that's what certifications audit and what separating customers' data depends on. The £95 review tells you which your app needs before you spend on either.

What certifications do I need to sell software to UK enterprises?

For central government and NHS contracts that handle citizens' personal data or OFFICIAL information, PPN 014 requires Cyber Essentials or proof of equivalent controls, and many buyers ask for Cyber Essentials Plus. As deal sizes grow, larger buyers ask for ISO 27001 (common with UK and EU buyers) or SOC 2 Type II (common with US buyers). Expect to need a recent penetration test report and a data processing agreement too. These check your foundations, so they can't be added at the last minute.

My app uses AI to generate content. What are the risks?

If that content is safety-critical or legally binding, the risk is significant. Models hallucinate on a meaningful share of factual questions, and using AI doesn't move your legal duty onto the AI vendor. A commercial AI feature needs guardrails: human review and sign-off before anything consequential is issued, grounding in verified data, defence against prompt injection, audit logging, and telling users when content is AI-generated. If your app is used in the EU, the transparency rules in Article 50 of the EU AI Act have applied since 2 August 2026.

Meet our CTO, Gareth. He has been involved in mobile app development for over 25 years. Gareth is an experienced CTO and works with many startups

We'd love to show you how we can help

Get in Touch  

Latest Articles

All Articles