Evidence and limitations
Methodology and results
These results come directly from the current model's published validation matrix and service limits.Last updated:
- Model version
- 20260820-011603
- Validation protocol
- v3 validation; production protocol: thresholds are set on the human training calibration set with alpha = 0.02; cross_source means the validation source was not present in the training corpus
- Total samples
- 2,800 samples
- Test date
- 2026-08-20
Current model aggregate results
Current-model test accuracy is 98.54%: 2,800 tested texts, 2,759 were classified correctly. The rate uses actual sample counts, not an average of scenario percentages, and applies only to this test set.
- Out-of-source AI-text recall
- 98.30%
- In-source human-text false-positive rate
- 0.56%
- Out-of-source human-text false-positive rate
- 2.50%
- Highest false-positive rate among tested cells
- 5.00%English · Literary ·
en·literary
The aggregate results above and cell-level results below come from the same validation matrix. False-positive rates include only cells with human-text samples; combinations with no human samples remain untested and are not counted as zero false positives.
Complete 12-cell validation matrix
Each cell is broken down by source relationship, language, and document type, with sample and hit counts for independently recalculating the published rates.
Swipe horizontally to view all columns.
| Source relationship | Language | Type | Human-text false-positive rate | AI-text recall |
|---|---|---|---|---|
| In-source | Chinese | Academic | 0.00%0 / 150 false positives | 95.33%143 / 150 identified |
| In-source | Chinese | Literary | 0.67%1 / 150 false positives | 100.00%150 / 150 identified |
| In-source | Chinese | Long-form online | 1.33%2 / 150 false positives | 98.00%147 / 150 identified |
| In-source | English | Academic | 0.67%1 / 150 false positives | 96.67%145 / 150 identified |
| In-source | English | Literary | 0.00%0 / 150 false positives | 100.00%150 / 150 identified |
| In-source | English | Long-form online | 0.67%1 / 150 false positives | 99.33%149 / 150 identified |
| Out-of-source | Chinese | Academic | Not tested (0 human samples) | 100.00%100 / 100 identified |
| Out-of-source | Chinese | Literary | 1.00%1 / 100 false positives | 100.00%100 / 100 identified |
| Out-of-source | Chinese | Long-form online | 4.00%4 / 100 false positives | 95.00%95 / 100 identified |
| Out-of-source | English | Academic | Not tested (0 human samples) | 99.00%99 / 100 identified |
| Out-of-source | English | Literary | 5.00%5 / 100 false positives | 98.00%98 / 100 identified |
| Out-of-source | English | Long-form online | 0.00%0 / 100 false positives | 98.00%98 / 100 identified |
Evaluation set and coverage
The validation matrix covers Chinese and English text in Academic, Literary, and Long-form online categories. Each scenario is further divided into in-source and out-of-source data. "Out-of-source" means the validation source was not present in the training corpus. The matrix contains 1,300 human texts and 1,500 AI texts.
Training corpus:19,004,936 non-whitespace characters from 47,183 Chinese and English texts used to fit the model. The count includes punctuation and counts English characters rather than words. It excludes calibration, validation, and test text and is not the pretraining scale of the underlying language model.
How is the threshold set?
Production thresholds are set using the human training calibration set under protocol train-human-conformal-alpha-0.02; score > domain threshold. A calibrated score is counted as a deviation only when it exceeds the threshold for the applicable language and domain. The threshold protocol, model version, and validation matrix are bound together; if model versions do not match, the service does not publish these metrics as current results.
How should I read the metrics?
AI-text recall is the share of actual AI text identified as a deviation; higher is better. The human-text false-positive rate is the share of actual human text identified as a deviation; lower is better. Untested cells remain blank.
Limitations
The current model's validated per-segment range is 110–1200 Chinese characters and 280–2600 English characters. Longer documents are split according to service rules. Do not extrapolate definitive conclusions from these metrics for very short fragments, substantially different source distributions, specially rewritten text, or combinations with no samples in the matrix.