Skip to main content

Method & benchmarks

Numbers that know where they stop.

By ShakuPublished Updated

How Kaneme measures a voice, how we test the engine, and where it refuses to answer.

Published results

You, or an AI imitating you

0,992

Q1 AUC

You, or another human

0,887

Q2 AUC

Held-out authors
31
Genres measured
6
Validated language
French
Served engine
emb-frq-1.1
Measured on
August 3, 2026

AUC: 1 = perfect separation, 0.5 = chance, written in French decimal notation. Each protocol is in chapter 2.

01

Two questions. Never one blended score.

“Is this you rather than another person?” and “is this you rather than a machine?” need different evidence. At the same score of 80, the calibrated probabilities are 70% and 99%. Kaneme keeps them apart.

Q1

Your voice, or an AI imitating you?

Calibrated against AI imitations built from each benchmark author’s own reference texts. This is the main measurement, and the bar is high: the machine already had material to copy.

Q2

You, or another human?

Much harder: people in the same field writing about the same subjects naturally overlap. Kaneme publishes this weaker result because it describes the difficult cases.

Fidelity score. Out of 100, it says how close a draft stays to your voice: rhythm, vocabulary, phrasing, avoided words.

Free test (Q0). Without your reference it answers a weaker question. Its scores and weaknesses are in the experimental details.

02

The result has to survive unseen authors.

We hold authors back before tuning, test on their real writing, then rerun the bench for every engine version. If a number falls, the published number falls with it.

Internal benchmark

0,992

Your voice versus AI imitation

The AI adversary receives material from the target author, so this is an imitation test rather than generic AI detection.

0,887

Your voice versus another human

The harder question, published because it describes the difficult cases.

99% / 70%

What a score of 80 is worth

The same displayed score does not carry the same evidence against AI and against another human.

Public benchmarks · PAN (CLEF) and AuthBench

The unchanged engine is also evaluated on public authorship competitions: texts we neither wrote nor chose. These out-of-domain results are a floor, not the product score.

0,70
0,90
0,62

03

How we test

Five steps, always in this order. A result that skips one is not published.

  1. 1Corpus

    54 French-speaking authors and their real writing. Tuning material and judging material never mix.

  2. 2Held-out set

    31 authors never seen in training, read once, at the end.

  3. 3Calibration

    Every score is checked against reality, pair by pair, separately against AI and against humans.

  4. 4Test

    AUC, false positives and attacks, measured out of sample, genre by genre.

  5. 5Publication

    Published numbers follow the served engine in the same deployment, including when they drop.

04

Where it breaks

Kaneme is not a universal AI detector: it measures whether a text sounds like a reference voice. This is where that measurement stops.

The reasoning behind these boundaries is set out in the manifesto.

  1. Outside its scope, it abstains

    Too little text, a thin corpus or a register the engine cannot read produces no forced score and no charge.

  2. Writing genre matters

    Results are published genre by genre, because some genres overlap far more than the median suggests. The weakest one is named on the French page.

  3. So does language

    The production verdict is calibrated for French writing. The English interface is fully usable, but it does not turn an unvalidated English measurement into a supported capability.

05

Same model. Same request. Three methods.

The public comparison sets the assistant alone, “write like me” with three examples, and Kaneme against the same brief, with methods hidden until you choose. It runs on French texts, on the French home page.

Going furtherSee the experimental details: calibration, adversaries and all limits

A score still isn’t a probability

Kaneme checks what each score meant on held-out writing, separately against other people and against AI imitations. The stress test also includes machine text deliberately rewritten to pass.

The detailed evidence stays with its valid domain

Genre tables, free-test weaknesses and provider rates live on the French page because that is where the production measurement has been validated. Translating the labels wouldn’t extend the evidence.

  • Not a universal AI detector

    Kaneme measures similarity to your reference. It does not infer the absolute origin of any text.

  • Adversarial performance is separate

    A strong result on ordinary samples does not prove resistance to deliberate evasion. Kaneme publishes that distinction.

  • Short texts are asymmetric

    A strong signal can support a match on a short text; a weak signal on a short text cannot prove a mismatch.

  • No exact public score

    Public certificates expose a band and provenance rather than a fine score someone could optimise against.

Bring the writing. Let the evidence answer.

Use three French reference texts and one candidate. You don’t need an account.