← Selected work

Responsible AI · Automation design · Human oversight

AI-First Support With Confidence-Based Automation

When should automation act on its own, and when does a human step in?

Company
The Weather Company
Period
2023–2024
Disciplines
AI, Responsible AI, Customer Experience, Operations
Chapters
8
AI-first
Primary entry point
4
Interaction classes with distinct rules
AOPs
Written procedures per class
01The situation

Putting an AI agent in front of consumer support at that scale means a wrong answer is not a one-person problem. Everyone agreed automation was the direction. Nobody had written down where its judgment ended.

02The actual problem

The request was

Turn on the AI agent and measure deflection.

Discovery showed

Deflection rewards answering. Trust depends on knowing when not to answer. Without a rule for that, the agent optimizes toward confident wrong answers, which costs more than a queue does.

So I defined confidence thresholds per interaction class and wrote AI Operating Procedures that described the action, the review requirement, and the failure mode for each one. Automation expanded when the evidence cleared the bar, not when someone was ready to celebrate.

03What I learned
  • Modeled answer accuracy against confidence score to find where automation actually earned trust.
  • Grouped interactions by consequence, from informational answers to account actions and safety topics.
  • Sampled agent responses continuously rather than auditing after complaints arrived.
  • Tracked escalation quality, since a bad handoff undoes a good answer.
04How I framed it

Exhibit

AI Operating Procedures

 ThresholdHuman reviewFailure mode
01Informational answerModerateSampledLow, correctable
02Troubleshooting stepsHighSampled and flaggedMedium, wasted user time
03Account or billing actionVery highAlwaysHigh, trust damage
04Safety or legal topicNever automatedHuman onlySevere

Below threshold the agent does not guess. It educates or hands off with full context.

What you're looking atOne row per interaction class. This is the contract that let automation grow without gambling the brand.
05What we built
  • 01Confidence thresholds tied to consequence rather than to a single global setting.
  • 02Written AI Operating Procedures covering action, review, and escalation per interaction class.
  • 03A warm handoff path that carries context so nobody repeats themselves to a human.
  • 04Continuous sampling and quality review instead of complaint-driven auditing.
06My role
Design
Authored the thresholds and the operating procedures.
Analysis
Modeled accuracy against confidence to place each bar.
Governance
Set what evidence was required before raising any threshold.
Operations
Built the review cadence with the support team, not around them.
07The outcome
  • The AI agent became the primary consumer entry point with an explicit, defensible scope.
  • Automation expanded on evidence, one interaction class at a time.
  • Escalations arrived with context instead of restarting the conversation.
  • The framework became the precursor to the autonomy levels I use in AI Ops work now.
08What I'd do differently

Writing down what the AI is not allowed to do turned out to be the fastest way to expand what it was allowed to do. The rules made people comfortable enough to say yes.

Contact

Let's get into it.

Ambiguous problem, AI adoption that stalled, an operating model that stopped scaling, a support function that should be a product. That is the conversation I want.