Why the first test mattered more than the feature
The app had shipped to the App Store with no usability testing and reviews that repeated the same complaints. I was hired to design one feature. The feature was the reason to run the first test, not the goal of it. Ten interviews came first, and all ten described the app as overwhelming. One said it felt like homework, which is the word a financial literacy app for children cannot afford.
How the evidence was graded
Two findings were Verified: only 20% chose competing against a rival as what they liked most, and participants asked to see what the rival answered. Those shipped as design changes. Requests for more game-like mechanics beyond the quiz were Directional, from a smaller number of sessions, and did not change the build. Separating the two kept the round from turning into a wish list.
What testing did not settle
Twenty-two sessions is enough to find a broken entry point and not enough to prove a retention effect. Task success read 100% while misclicks ran at 65.9%, which is the clearest evidence that the metric the company watched was not measuring the thing it believed. I left before the release cycle closed and cannot confirm what shipped.






