← Selected work

Discovery · Operating model design · Executive synthesis

AI Operations: The Evidence Changed the Question

Where can AI create real operational leverage, and what has to be true first?

Company
Honeycomb
Period
2026
Disciplines
AI Strategy, Discovery, Operating Model, Systems Thinking
Chapters
8
36
Discovery sessions across functions
23
Experiments designed
4
Pilots running
100+
GTM tasks decomposed
01The situation

I joined to find where AI could create leverage. The common approach is to collect use cases, rank them by excitement, and start building. I started by watching how work actually happened instead, function by function, session by session.

02The actual problem

The request was

Find the AI use cases and start automating.

Discovery showed

The evidence pointed somewhere else. In most workflows the model could already do the task. What was missing was the context the model needed, a clear definition of the work, someone who owned the outcome, systems that could write back, and enough trust for anyone to act on the output.

Once that pattern repeated across 36 sessions, the question stopped being 'what can AI do here' and became 'what has to be true before AI can improve this work.' That reframe is the case study. It changed what I built, what I measured, and what I told the executive team.

03What I learned
  • 36 discovery sessions across GTM, engineering, finance, and operations, focused on how the work runs rather than what people wish it did.
  • Decomposed 100+ GTM tasks into inputs, judgment, context sources, and handoffs.
  • Designed 23 experiments to test specific conditions, not general enthusiasm.
  • Stood up 4 pilots where the conditions were already close enough to succeed.
04How I framed it

Exhibit

What Actually Blocked the Work

 What it looks likeWhat it actually needsFix owner
01Missing contextOutput is generic or wrongRetrieval into real systems of recordData + Ops
02Undefined workEvery person runs it differentlyA written definition of doneFunction lead
03No ownerPilot stalls after the demoOne named outcome ownerExec sponsor
04Disconnected systemsHuman copies output by handWrite-back and notificationsEngineering
05No trustPeople check every resultEvaluation, sampling, visible rulesAI Ops

Model capability appeared as the primary blocker in a small minority of the tasks I looked at.

What you're looking atAcross the sessions, the blocker was rarely the model. Sorting the findings this way made the roadmap obvious and made the executive conversation short.

Exhibit

The AI Ops Method

  1. 01

    Observe

    Sit with the work as it runs today

  2. 02

    Decompose

    Break it into tasks, context, judgment, handoffs

  3. 03

    Test conditions

    Run a narrow experiment against one blocker

  4. 04

    Build the condition

    Context, ownership, connection, rules

  5. 05

    Then automate

    Add the model once the work can hold it

What you're looking atThe repeatable loop I built out of the discovery findings.
05What we built
  • 01A repeatable AI Ops discovery and decomposition method other leads can run without me.
  • 02A prioritized experiment portfolio tied to conditions rather than to tooling.
  • 03Four pilots with named owners, defined outcomes, and evaluation in place.
  • 04An executive narrative that moved the conversation from tool selection to operating conditions.
06My role
Discovery
Ran all 36 sessions myself and did the synthesis.
Framing
Changed the thesis when the evidence stopped supporting the original one.
Method
Turned the findings into a method rather than a slide.
Executive
Presented the reframe and the sequencing to leadership.
07The outcome
  • The company's AI thesis shifted from use-case hunting to condition-building.
  • 23 experiments and 4 pilots sequenced against evidence instead of enthusiasm.
  • A shared vocabulary for why an AI project stalls, which made the failures diagnosable.
  • The method now sits underneath the GTM operating system, the task assessment framework, and the engineering discovery work.
08What I'd do differently

The hardest part was saying out loud that the question I was hired to answer was the wrong question. It landed because I had 36 sessions behind it. Evidence buys you the right to change the frame.

Contact

Let's get into it.

Ambiguous problem, AI adoption that stalled, an operating model that stopped scaling, a support function that should be a product. That is the conversation I want.