Web end-to-end testing is a solved, comfortable problem. Tools are mature, everything happens inside one browser tab, and a competent team can have a suite running in a week. Mobile is not that, and the difference is not effort — it is where the failures live.

Your app is a guest on someone else's platform

The moments that break an app are mostly not your code:

  • The permission dialog — camera, notifications, location — drawn by the operating system, on top of your app, outside anything your tests can normally reach.
  • The payment sheet, likewise. Apple Pay and Google Pay hand control away and give it back changed.
  • A push notification arriving while the screen is locked, and what happens when it is tapped.
  • The app killed in the background by an Android that needed the memory, then reopened three days later expecting to find its state.
  • The network that comes and goes in a lift, on a train, in a basement warehouse.

A test that stays inside your app cannot see any of this. Which is why most mobile suites quietly test the easy half and the interesting failures reach your users first.

What that means practically

Testing that reaches outside the app needs a tool that can drive the device, not just the widget tree. We use Patrol for Flutter apps because it can tap a system dialog, dismiss a payment sheet, and assert what the app looks like when it comes back. Our own marketplace carries thirty such journeys, and the ones that repeatedly earn their cost are exactly the awkward ones: a deposit paid by card when the wallet is empty, a token refreshed mid-session, two arbitration flows for when two people disagree about the same order.

The second requirement is real devices — or at least real emulators in the pipeline, which cost meaningfully more than a headless browser. That cost is the honest reason mobile testing is quoted higher, and it is worth knowing so you can tell a considered quote from a hopeful one.

The trap of testing only what is easy

There is a version of mobile testing that looks thorough and is not: unit tests over business logic, widget tests over screens, and a handful of integration tests through the happy path. It produces impressive numbers and misses every failure listed at the top of this article.

Ask a studio how they test the permission dialog. Ask what happens in their suite when the app is killed in the background. If the answer is that those are checked manually before release, that is a legitimate choice — but it means the checking is made of a person and their attention on a release evening, and you should price that accordingly.

And the part everyone forgets

An app is a client for a server, and half the interesting failures are joint: the payment that succeeded on the server and failed to render on the phone, the order that exists in two states at once. Assert the same journey from both ends. On our marketplace that means thirty journeys on the device and twenty-four specs on the backend, checking the same flows from the other side. When they disagree, you have found something real — and you have found it before your customer did. Testing an app from both ends is part of what we set up, and the mobile half is the reason a first release takes eight weeks rather than four.

The question before this one: what to test first when you cannot test everything.