Accuracy, benchmarks and limits

"Does an AI music detector actually work?" is the right question to ask before you rely on one. This page explains how results are produced, how to read them, where they fail, and what our own tests show.

How detection works

The detector analyses the audio signal — not file names, tags or metadata. Music generators leave characteristic artefacts in the spectrum and in how sound evolves over time. A model trained on large sets of generated and human-made music scores the full mix, each section of the track, and (when available) the vocal and instrumental stems separately. It also scores how closely the track matches each known generator.

How to read a result

BandAI probabilitySuggested action
AI-generated90–100%Apply your AI policy. Keep the report.
Likely AI65–90%Review, or ask the uploader.
Inconclusive35–65%Send to a human listener.
Likely human10–35%Continue; spot-check if the account is new.
Human-made0–10%Continue.

The likely source is only shown as a clear match when one generator scores at least 60% and leads the next by 20 points. Otherwise the report says the source is unclear and shows the closest match.

Our benchmark

Our first benchmark run is in progress. When it is published, this table will show detection rates per generator and version, false-positive rates on human-made releases, and how results hold up after MP3 re-encoding, trimming and pitch-shifting.

What the benchmark covers:

  • Generators and versions: Suno, Udio, Mureka, ElevenLabs Music and others, including versions released after the detection model was trained.
  • Human-made music: commercial releases across genres, home recordings, and heavily produced pop with pitch correction.
  • Transformations: MP3 128/320 kbps, streaming-quality AAC, 30-second previews, ±1 semitone pitch shift, 5% time-stretch.
  • Mixed tracks: AI vocals over human production and human vocals over AI instrumentals.

The underlying detection engine's developer reports very high precision and recall on its own internal test data. Independent tests on newer generators and processed files are usually lower, which is exactly what our benchmark measures.

Known limits

  • New generator versions can be under-detected until the model is retrained on them.
  • Heavy processing — low bitrates, pitch-shifting, time-stretching — pushes scores toward the middle band.
  • Partly AI tracks can score lower overall; check the stems and the timeline.
  • Short clips (under ~30 seconds) carry less evidence.
  • Instrumental music has no vocal stem; decisions rest on the mix alone.

FAQ

On clean exports from generators they were trained on, detection is very reliable. Accuracy drops on new generator versions, heavy processing and tracks that mix AI and human parts. That is why we report confidence bands and publish results by generator and by file variant, not a single headline number.

No. It gives a probability based on audio artefacts. Use it to decide which tracks need a closer look, and combine it with other evidence before enforcement.

Aggressive pitch correction and vocoders, some AI-assisted mastering and stem-separation tools, very low bitrates, and certain synthetic sound-design styles. The segment timeline helps: a real false positive is often concentrated in one section.

Report a wrong result

If you believe a result is wrong, email support@musicdetector.org with the detection ID. We use disputed cases (with your permission) to improve our test sets — never to train on your audio without consent.