← Selected work

AI product strategy · Platform integrity · Model evaluation

MAIA: AI-Generated Content Detection at Platform Scale

Once you can detect AI-generated music, what should actually happen to the track?

Company
SoundCloud
Period
2024–2026
Disciplines
Trust & Safety, AI Product, Evaluation, Platform
Chapters
8
99.36%
Detection accuracy
99.75%
Precision
~700K
Tracks scanned per day
5–7s
Latency per track
01The situation

AI-generated music arrived on the platform faster than the industry could agree on what it meant. Artists wanted transparency. Rights holders and partners wanted assurance. Product wanted a policy that could ship. Nobody wanted a system that quietly punished a human musician for sounding synthetic.

02The actual problem

The request was

Can we detect AI-generated tracks?

Discovery showed

Detection alone changes nothing. A confidence score on a track is only useful once someone has decided what it triggers: a tag, a review queue, a monetization rule, a partner signal, or nothing at all.

The technical bar was real. It had to be accurate enough that false positives did not damage legitimate artists, and fast enough to keep up with upload volume at platform scale. But the decision that mattered was product and policy: what the platform does with the answer, who reviews it, what an artist sees, and how an appeal works.

03What I learned
  • Evaluated detection performance on catalog-representative audio rather than vendor benchmark sets.
  • Modeled false-positive cost in artist-trust terms, not just precision terms, because the two do not move together.
  • Sized throughput, latency, and unit cost against real daily upload volume before committing to an integration.
  • Worked through tagging, monetization, and partner expectations with product, legal, and trust and safety in the same room.
04How I framed it

Exhibit

Detection Performance and What It Costs

 MeasuredWhy it matteredTradeoff
01Accuracy~99.36%Credible enough to act onResidual error at 700K/day is still real volume
02Precision~99.75%Protects legitimate artistsTuned up at some cost to recall
03Throughput~8 tracks/secKeeps pace with uploadsInfrastructure cost scales with catalog
04Latency5–7 secondsFits inside the upload flowLimits model size and complexity

A false positive on a human artist costs more than a false negative on a synthetic track. That asymmetry set the tuning.

What you're looking atThe numbers that decided whether this could run in production, alongside the tradeoff each one carried.

Exhibit

What Happens After Detection

  1. 01

    Scan

    Every upload scored at ingest

  2. 02

    Tag

    Disclosure applied at the agreed confidence

  3. 03

    Route

    Borderline results into human review

  4. 04

    Apply policy

    Monetization and distribution rules follow the tag

  5. 05

    Appeal

    Artists can contest, and outcomes feed back into tuning

What you're looking atThe part that turned a model into a product.
05What we built
  • 01The product and policy layer around detection: tagging, review routing, monetization rules, and appeals.
  • 02An evaluation approach measured on real catalog distribution rather than published benchmarks.
  • 03A shared risk model that priced false positives in artist-trust terms across product, legal, and trust and safety.
  • 04The operational plan for throughput, latency, and cost at roughly 700K tracks a day.
06My role
Product
Owned what the platform does with a detection, not just whether it detects.
Evaluation
Designed the testing approach and the criteria for shipping.
Tradeoffs
Set the precision-first tuning based on artist impact.
Alignment
Got product, legal, trust and safety, and partners to one decision.
07The outcome
  • Detection runs at platform scale with performance the business can act on.
  • Transparency became a shipped behavior rather than a stated intention.
  • False-positive cost is an explicit, agreed input to tuning instead of an unspoken assumption.
  • The evaluation approach carried into later model and vendor decisions.
08What I'd do differently

The temptation with a number like 99.36% is to treat it as the finish line. At 700K tracks a day the remainder is a real queue of real artists, which is why the review path and the appeal path got as much of my attention as the model did.

Contact

Let's get into it.

Ambiguous problem, AI adoption that stalled, an operating model that stopped scaling, a support function that should be a product. That is the conversation I want.