Why Vibe-Coded iOS Apps Keep Failing App Review
A year ago, a founder with a weekend and a credit card could get an iOS prototype onto a simulator, maybe onto their own phone through a developer certificate, and show it around. The build worked on the machine that made it, which was enough for most conversations. Today the same founder opens a chat window, describes an app in plain English, watches an LLM assemble something that compiles, uploads it to TestFlight, and expects the public launch to be the easy part.
The code runs. The submission portal is a different audience entirely.
That shift is why rejection letters have started reading like code reviews. The questions below are the ones founders keep asking after the first rejection email lands, in roughly the order they ask them.
Why Does a Build That Runs Locally Keep Failing Review?
A build that opens on your phone has cleared one bar: it compiles, launches, and does something recognizable. App Review is a different bar. Before external testers can even touch a TestFlight build, Apple runs it through Beta App Review against the full App Store Review Guidelines, and the production submission is reviewed against the same checklist with less patience. The prompt that generated the app never read those guidelines.
The scale of the mismatch is easy to underestimate. AppFollow's overview of Apple's own 2025 numbers shows more than two million rejected submissions in a single year, with over 1.35 million flagged for Performance alone — crashes, bugs, placeholder content, incomplete flows.
A vibe-coded app rarely crashes in the demo. It crashes on the device the reviewer picks, when an edge the prompt never saw gets exercised. The take from Geek Vibes Nation on vibe coding and the App Store puts it plainly: the prototype looks magical in a demo, and the submission portal is a different room.
What Does the Reviewer Actually Check That the Prompt Doesn't?
The checklist is long, but a handful of items catch vibe-coded submissions over and over. These are the recurring hits.
- Sign in with Apple. If the app offers third-party or social login, Apple's guidelines require Sign in with Apple as an equivalent option in most cases. Prompts cheerfully wire up Google and email auth and skip this entirely.
- Permission strings. Every requested permission — camera, location, contacts, microphone, tracking — needs a purpose string that explains the use. Generated boilerplate like "This app needs access to your camera" gets rejected on sight.
- Account deletion in-app. If users can create an account, they have to be able to delete it from inside the app, not through a support email. Prompted auth flows almost never include this screen.
- Payments through the right rails. Digital goods and subscriptions have to go through StoreKit. An LLM that saw a lot of Stripe tutorials will happily wire Stripe into a consumable purchase and earn a Guideline 3.1.1 rejection.
- Working demo credentials. Review notes with a login that doesn't work, or a sandbox that silently fails, close the review before it starts.
- Real content, not placeholder. Lorem ipsum, broken links, empty states that say "Coming soon," and screens the prompt scaffolded but never finished all count as an incomplete build.
Where Are the Security and Data Problems Hiding?
The visible rejections are the forgiving ones. The invisible failures — the ones that pass review and detonate in production — tend to live in how the generated code talks to its backend. Prompted apps lean heavily on a handful of hosted services for auth and data, and they tend to connect to them with whatever permissions the quickstart used. Row-level security gets skipped, API keys land in the client bundle, and services like Supabase or Firebase do the right thing by default only when you configure them.
That is the pattern behind the public incidents. Black Duck's coverage of vibe coding catalogs the recurring failure modes: unknown provenance for the generated code, vulnerable dependencies pulled in without review, secrets embedded in the client, and hallucinated APIs that happen to compile. None of these surface during a TestFlight session. They surface when a researcher pokes at the production backend, or when a user figures out that another account's data is one request away.
Where Do Experienced Engineers Have to Step Back In?
The honest answer: before the first submission, not after the third rejection. A senior iOS engineer brings a mental checklist the prompt doesn't have — the submission guidelines, yes, but also how Keychain should hold secrets instead of UserDefaults, how background modes get justified, how StoreKit receipts validate, how the privacy manifest declares every tracking SDK the generated code pulled in. All of it is what separates a build that passes review from one that keeps bouncing.
The useful division of labor is becoming clearer. Prompts are good at scaffolding — the first version of a screen, a plausible data model, boilerplate for a familiar pattern. Humans are still responsible for the parts that make the app shippable: the threat model, the permission posture, the payment path, the deletion flow, the empty states, the error handling the demo never triggered.
Teams that treat the LLM output as a first draft and bring engineers in to harden it ship. Teams that treat the output as the product keep refreshing the review queue. The App Store has long been a filter. It is now filtering, in part, for whether anyone with judgment touched the code before submission.
