Evidence and limitations

Methodology and results

These results come directly from the current model's published validation matrix and service limits.

Last updated:

Model version
20260820-011603
Validation protocol
v3 validation; production protocol: thresholds are set on the human training calibration set with alpha = 0.02; cross_source means the validation source was not present in the training corpus
Total samples
2,800 samples
Test date
2026-08-20

Current model aggregate results

Current-model test accuracy is 98.54%: 2,800 tested texts, 2,759 were classified correctly. The rate uses actual sample counts, not an average of scenario percentages, and applies only to this test set.

Out-of-source AI-text recall
98.30%
In-source human-text false-positive rate
0.56%
Out-of-source human-text false-positive rate
2.50%
Highest false-positive rate among tested cells
5.00%English · Literary · en·literary

The aggregate results above and cell-level results below come from the same validation matrix. False-positive rates include only cells with human-text samples; combinations with no human samples remain untested and are not counted as zero false positives.

Complete 12-cell validation matrix

Each cell is broken down by source relationship, language, and document type, with sample and hit counts for independently recalculating the published rates.

Swipe horizontally to view all columns.

Cell-level recall and false-positive rates for the current model
Source relationshipLanguageTypeHuman-text false-positive rateAI-text recall
In-sourceChineseAcademic0.00%0 / 150 false positives95.33%143 / 150 identified
In-sourceChineseLiterary0.67%1 / 150 false positives100.00%150 / 150 identified
In-sourceChineseLong-form online1.33%2 / 150 false positives98.00%147 / 150 identified
In-sourceEnglishAcademic0.67%1 / 150 false positives96.67%145 / 150 identified
In-sourceEnglishLiterary0.00%0 / 150 false positives100.00%150 / 150 identified
In-sourceEnglishLong-form online0.67%1 / 150 false positives99.33%149 / 150 identified
Out-of-sourceChineseAcademicNot tested (0 human samples)100.00%100 / 100 identified
Out-of-sourceChineseLiterary1.00%1 / 100 false positives100.00%100 / 100 identified
Out-of-sourceChineseLong-form online4.00%4 / 100 false positives95.00%95 / 100 identified
Out-of-sourceEnglishAcademicNot tested (0 human samples)99.00%99 / 100 identified
Out-of-sourceEnglishLiterary5.00%5 / 100 false positives98.00%98 / 100 identified
Out-of-sourceEnglishLong-form online0.00%0 / 100 false positives98.00%98 / 100 identified

Evaluation set and coverage

The validation matrix covers Chinese and English text in Academic, Literary, and Long-form online categories. Each scenario is further divided into in-source and out-of-source data. "Out-of-source" means the validation source was not present in the training corpus. The matrix contains 1,300 human texts and 1,500 AI texts.

Training corpus:19,004,936 non-whitespace characters from 47,183 Chinese and English texts used to fit the model. The count includes punctuation and counts English characters rather than words. It excludes calibration, validation, and test text and is not the pretraining scale of the underlying language model.

How is the threshold set?

Production thresholds are set using the human training calibration set under protocol train-human-conformal-alpha-0.02; score > domain threshold. A calibrated score is counted as a deviation only when it exceeds the threshold for the applicable language and domain. The threshold protocol, model version, and validation matrix are bound together; if model versions do not match, the service does not publish these metrics as current results.

How should I read the metrics?

AI-text recall is the share of actual AI text identified as a deviation; higher is better. The human-text false-positive rate is the share of actual human text identified as a deviation; lower is better. Untested cells remain blank.

Limitations

The current model's validated per-segment range is 1101200 Chinese characters and 2802600 English characters. Longer documents are split according to service rules. Do not extrapolate definitive conclusions from these metrics for very short fragments, substantially different source distributions, specially rewritten text, or combinations with no samples in the matrix.

View current-model test details on the home page