Why this product
AI can classify tickets. Every support tool proves it daily. The open question is whether the person held accountable for the output can understand it, calibrate it, and take it back. Triage makes that asymmetry concrete: when the AI misfiles an account deletion as product feedback, the agent answers for it, not the model. I audited Zendesk, Intercom Fin, and Freshdesk. All three lead with resolution volume, show confidence as a raw score, and bury the path to reverse an action.
How the evidence was graded
Concept projects invite cherry-picking, so I graded every source before it could influence a decision. Verified meant peer-reviewed or replicated. Directional meant a single credible study. Vendor claim meant marketing numbers, treated only as a signal of what companies believe. The grading changed the design. Frequency framing rests on Verified research and shipped as a core pattern. The feedback-framing decision rested on a Directional finding and shipped as a hypothesis Round 1 was built to check.
What testing did not settle
Both rounds ran unmoderated in Maze on the wired prototype, blocks identical word for word. Five to seven participants per round is enough to expose a broken flow and not enough to estimate an effect size, so 83.5% to 0% is a before and after of one redesign, not a controlled experiment. With no production data, the model calibration outside my held-out set is untested.



