Skip to content
Jun – Jul 2026 · Self-directed concept

Surfaced, not buried.
Misclicks: 83.5% to 0.

My impact
Sift inbox, AI-sorted queue with confidence, category and undo
Role
Solo designer, end to end
Scope
Product design, research, model training, prototyping
Tools
Figma, FigJam, Maze

83.5% of support professionals clicked everywhere except the control that decides how much the AI does.

Overview

Agents do not need more automation. They need to see what the AI did, why, and how to take it back. Designed end to end, tested in two rounds with twelve CX professionals, model fine-tuned by me.

My impact

Agents stopped hunting for the controls and kept the final say on every call. Not a mockup: 21 wired screens with the real classifier, live in the browser.

  1. 01Agents clear hundreds of tickets a day, and their AI tools hide the why.
  2. 02So Sift hands control back: confidence shown, reasons visible, actions reversible.
  3. 03Round 1 broke my assumptions, so I fixed what it exposed and ran Round 2.
0%
misclick rate

Down from 83.5% on the critical task. Same mission, same population, one redesign in between.

0.0/5
perceived control

Up from 3.0. Twelve working CX professionals, two rounds, identical task blocks.

0.00%
model intent accuracy

A classifier I fine-tuned myself, calibrated to ECE 0.0016 and running client-side.

Round 1 · one entry, buried in Settings
Round 1: the only control, buried in SettingsRound 2: automation controls surfaced as four doors in the inbox
Same mission, same population, one redesign apart. Hover a marker to see what moved.
Evidence

The 3.0 average was hiding a split.

A cluster at 4 and a cluster at 1 to 2, split by whether people ever found the control. Round 2 asked the same question of the same population.

Round 1
n = 7
avg 3.0
Round 2
n = 5
avg 4.4
12345
Felt sense of control · 1 = none → 5 = full

Even at 95%, experts still read the ticket

Three of five checked the source before accepting at 95%. That is not distrust. It is their name on the decision, so the design stopped asking for trust and made verifying fast.

Who it is for

“If it files something wrong, I’m the one who hears about it.”

Maya Chen · Support agent · 142 open tickets today
  • See why the AI chose, not only what it chose
  • Take back any automated action in one tap
  • Decide how much the AI does, category by category

One persona, deliberately. Team leads set policy and read audit trails, and scoping them out is what let the agent seat go deep instead of wide.

Solution 01

Confidence, written as a sentence people can calibrate

Raw percentages lie to the gut. Every number carries a frequency line: right about 41 of 100 on tickets like this.

On the ticket

On tickets like this, the AI has been right about 41 of 100.

Low confidence·41%
Why this is uncertain
Matched
refund · order reference
Signal
billing dispute, 0.41
Lowered by
overlap with account access
In settings

At 85%, the AI has been right about 96 of every 100 similar tickets.

Confidence threshold
Conservative · 90%Balanced · 85%Aggressive · 70%
In the queue

Risk policy visible where the work happens.

Delete my account and all data
#48196 · Tom Kim · 11m
Always humanNot scoredthe AI did not judge this
Solution 02

A dial the agent owns, and friction only on the risky calls

One threshold per category, because a bug report and an account deletion are not the same risk. Account deletion stays Always human.

Glass confirmation: friction only for expensive mistakes
Exploration

Three architectures. One survived.

Review every ticket and approval turns automatic. Automate every ticket and no one checks what goes out. The survivor routes by confidence, and by risk.

01
Rejected
Review everything

Every AI suggestion lands in an approval queue. Nothing moves until a human confirms it.

At 142 tickets a day, approving every suggestion becomes automatic clicking. It also throws away the volume relief that justifies AI triage.

02
Rejected
Automate everything, undo after

The AI files every ticket on its own. The agent gets a complete activity log and a global undo.

When everything is automatic, review atrophies. An auto-processed account deletion is an incident.

03
Chosen
Confidence-routed, per category

Auto-sort above a threshold the agent sets per category, route everything below it to a person, never touch locked categories.

The only architecture where speed and judgment coexist. At threshold 100 it degrades into option 01, so a cautious team can start there and move.

Every incoming ticket gets a proposalcategory · priority · confidence
Above threshold
Auto-sortedundo stays open
Below threshold
Needs reviewrouted to a human first
Sensitive category
Always humannever auto, permanently
one dial per category · every automated action reversible
Under the hood

The model is real, and it runs in the browser

A DistilBERT classifier fine-tuned on 24,370 tickets, calibrated to ECE 0.0016, quantized to 68 MB, deployed with transformers.js. Nothing you type leaves the page.

24,370
tickets fine-tuned on
99.67%
intent accuracy
0.0016
expected calibration error
68 MB
quantized, in-browser

The classifier weighs 68 MB and runs entirely on this page. Load it once and your browser caches it for next time.

Why this product

AI can classify tickets. Every support tool proves it daily. The open question is whether the person held accountable for the output can understand it, calibrate it, and take it back. Triage makes that asymmetry concrete: when the AI misfiles an account deletion as product feedback, the agent answers for it, not the model. I audited Zendesk, Intercom Fin, and Freshdesk. All three lead with resolution volume, show confidence as a raw score, and bury the path to reverse an action.

How the evidence was graded

Concept projects invite cherry-picking, so I graded every source before it could influence a decision. Verified meant peer-reviewed or replicated. Directional meant a single credible study. Vendor claim meant marketing numbers, treated only as a signal of what companies believe. The grading changed the design. Frequency framing rests on Verified research and shipped as a core pattern. The feedback-framing decision rested on a Directional finding and shipped as a hypothesis Round 1 was built to check.

What testing did not settle

Both rounds ran unmoderated in Maze on the wired prototype, blocks identical word for word. Five to seven participants per round is enough to expose a broken flow and not enough to estimate an effect size, so 83.5% to 0% is a before and after of one redesign, not a controlled experiment. With no production data, the model calibration outside my held-out set is untested.

Reflection
01

Failing in front of seven people

Round 1 broke publicly on my own site. Nobody made me run it, and rebuilding after it is the part of this project I would defend hardest.

02

Honesty is a set of decisions that don’t demo

A frequency sentence, a Not scored label, an error banner whose first message is that the human can keep working. Each one is skippable and none of them show up in a demo.