Method & benchmarks
Numbers that know where they stop.
By ShakuPublished Updated
How Kaneme measures a voice, how we test the engine, and where it refuses to answer.
Published results
You, or an AI imitating you
0,992
Q1 AUC
You, or another human
0,887
Q2 AUC
- Held-out authors
- 31
- Genres measured
- 6
- Validated language
- French
- Served engine
- emb-frq-1.1
- Measured on
- August 3, 2026
AUC: 1 = perfect separation, 0.5 = chance, written in French decimal notation. Each protocol is in chapter 2.
Method contents
- 01What we measureTwo questions, ranked by importance.
- 02ResultsInternal benchmark, public benchmarks, what a score is worth.
- 03How we testFrom corpus to publication, in five steps.
- 04Where it breaksThe three limits that change an answer.
- 05Product benchmarkThe public comparison, to play yourself.
- 06Experimental detailsCalibration, adversaries, all limits.
01
Two questions. Never one blended score.
“Is this you rather than another person?” and “is this you rather than a machine?” need different evidence. At the same score of 80, the calibrated probabilities are 70% and 99%. Kaneme keeps them apart.
Your voice, or an AI imitating you?
Calibrated against AI imitations built from each benchmark author’s own reference texts. This is the main measurement, and the bar is high: the machine already had material to copy.
You, or another human?
Much harder: people in the same field writing about the same subjects naturally overlap. Kaneme publishes this weaker result because it describes the difficult cases.
Fidelity score. Out of 100, it says how close a draft stays to your voice: rhythm, vocabulary, phrasing, avoided words.
Free test (Q0). Without your reference it answers a weaker question. Its scores and weaknesses are in the experimental details.
02
The result has to survive unseen authors.
We hold authors back before tuning, test on their real writing, then rerun the bench for every engine version. If a number falls, the published number falls with it.
Internal benchmark
0,992
Your voice versus AI imitation
The AI adversary receives material from the target author, so this is an imitation test rather than generic AI detection.
0,887
Your voice versus another human
The harder question, published because it describes the difficult cases.
99% / 70%
What a score of 80 is worth
The same displayed score does not carry the same evidence against AI and against another human.
Public benchmarks · PAN (CLEF) and AuthBench
The unchanged engine is also evaluated on public authorship competitions: texts we neither wrote nor chose. These out-of-domain results are a floor, not the product score.
03
How we test
Five steps, always in this order. A result that skips one is not published.
1Corpus
54 French-speaking authors and their real writing. Tuning material and judging material never mix.
2Held-out set
31 authors never seen in training, read once, at the end.
3Calibration
Every score is checked against reality, pair by pair, separately against AI and against humans.
4Test
AUC, false positives and attacks, measured out of sample, genre by genre.
5Publication
Published numbers follow the served engine in the same deployment, including when they drop.
04
Where it breaks
Kaneme is not a universal AI detector: it measures whether a text sounds like a reference voice. This is where that measurement stops.
The reasoning behind these boundaries is set out in the manifesto.
Outside its scope, it abstains
Too little text, a thin corpus or a register the engine cannot read produces no forced score and no charge.
Writing genre matters
Results are published genre by genre, because some genres overlap far more than the median suggests. The weakest one is named on the French page.
So does language
The production verdict is calibrated for French writing. The English interface is fully usable, but it does not turn an unvalidated English measurement into a supported capability.
05
Same model. Same request. Three methods.
The public comparison sets the assistant alone, “write like me” with three examples, and Kaneme against the same brief, with methods hidden until you choose. It runs on French texts, on the French home page.
Going furtherSee the experimental details: calibration, adversaries and all limits
A score still isn’t a probability
Kaneme checks what each score meant on held-out writing, separately against other people and against AI imitations. The stress test also includes machine text deliberately rewritten to pass.
The detailed evidence stays with its valid domain
Genre tables, free-test weaknesses and provider rates live on the French page because that is where the production measurement has been validated. Translating the labels wouldn’t extend the evidence.
Not a universal AI detector
Kaneme measures similarity to your reference. It does not infer the absolute origin of any text.
Adversarial performance is separate
A strong result on ordinary samples does not prove resistance to deliberate evasion. Kaneme publishes that distinction.
Short texts are asymmetric
A strong signal can support a match on a short text; a weak signal on a short text cannot prove a mismatch.
No exact public score
Public certificates expose a band and provenance rather than a fine score someone could optimise against.
Bring the writing. Let the evidence answer.
Use three French reference texts and one candidate. You don’t need an account.