Can You Trust AI and Deepfake Detectors?

AI and deepfake detectors can be useful, but their scores are not proof. Learn why false positives, compression, unseen generators and context matter when interpreting detector results.

A detector score can look authoritative: 96% AI-generated, 82% deepfake, likely synthetic. But a percentage is not the same as proof.

AI and deepfake detectors can be valuable forensic signals, especially when used by trained analysts. The danger comes when a score is treated as a final verdict without understanding how the detector was built, what it has seen before and how the media has been processed.

What does a detector actually do?

Most detectors estimate whether a piece of media contains patterns associated with synthetic or manipulated content. Depending on the system, those patterns may come from spatial image features, frequency information, compression artefacts, model fingerprints, biological cues, metadata or multiple signals combined.

The output is usually a model score or classification — not a mathematical statement that the file is definitely fake.

Why detector accuracy can fall in the real world

Compression and screenshots

Social platforms resize, recompress and transform media. Screenshots remove metadata and can alter the low-level signals a detector relies on. A model tested on clean benchmark files may behave very differently on a heavily compressed repost.

New generators

A detector can perform well on generators represented in its training data and struggle when a new model produces different artefacts. This is one of the central problems in deepfake detection: generalising to content the detector has never seen.

Editing after generation

Cropping, blur, colour correction, re-encoding, inpainting and other edits can change forensic traces. Some changes are innocent; others may be deliberately used to evade detection.

False positives

Genuine media can contain unusual textures, lighting, rendering artefacts or compression patterns that resemble synthetic content. A false accusation based on one detector can cause real harm.

What does “95% fake” actually mean?

The meaning depends on the detector. A displayed percentage might be a model confidence score, a calibrated probability, a thresholded output or simply a user-interface representation. Unless the provider explains calibration and evaluation, you should not assume “95%” means there is literally a 95% real-world probability the file is fake.

When are detectors useful?

  • As one signal in a wider verification workflow.
  • For triaging large volumes of content for further review.
  • When multiple independent detectors or forensic methods agree.
  • When results are interpreted alongside provenance, source and context.
  • When the limitations of the specific detector are known.

When should you be especially cautious?

  • When the media is a screenshot or heavily compressed.
  • When only one detector has been used.
  • When the result could lead to a serious accusation.
  • When the detector provider gives no evaluation methodology.
  • When the model may not have been tested on the generator or editing process involved.

Detection versus provenance

Detection asks whether the media contains signs of synthesis. Provenance asks what can be verified about where the file came from and how it changed. These are different questions, and they work best together. See Provenance vs Detection for the full comparison.

That is why organisations are increasingly combining open standards such as C2PA, durable watermarking and detection. OpenAI describes a multi-layered provenance approach that combines Content Credentials, SynthID and public verification, while noting that no single detection method is foolproof. Read the provenance approach.

Community Smart Hub’s rule

Never make a high-stakes decision from one detector score.

Combine source + context + search + provenance + forensic analysis + human judgement. That is the principle behind our Five-Layer Shield and Verify & Trust guidance.

Questions to ask before trusting a detector

  1. What content was the detector trained and tested on?
  2. Has it been tested on compressed and edited media?
  3. Does the provider publish false-positive and false-negative performance?
  4. Does the result change if another detector is used?
  5. Is there provenance or source evidence that supports or contradicts the score?

Bottom line: detectors can inform a decision. They should not replace verification.

Dr Jireh Jam
Dr Jireh Jam

Dr Jireh Jam is a computer vision and AI technologist specialising in deepfake detection, synthetic media, content provenance, watermarking, age assurance, AI evaluation and online safety. He holds a PhD in Computer Vision and has led applied AI research, evaluation and public-interest technology projects, translating complex technical risks into practical guidance for communities, organisations and policymakers.

Articles: 19

Leave a Reply

Your email address will not be published. Required fields are marked *