A new standard for Chinese pre-publication correction.
V3.0 finds more of the errors that matter, reduces unnecessary alerts, and returns a publication-ready correction more consistently across news, publishing, government, and regulated-content workflows.
V3.0 compared with V2.0, at a readable scale.
Four deployment-critical metrics retain paired absolute values; the complete eight-metric ledger below shows the full sample-weighted result. V3.0 improves correction quality while lowering both missed errors and false alarms.
Evaluation protocol: V2.0 and V3.0 use the same task definitions; 4,408 examples across three high-difficulty sample sets (n=484 / 991 / 2933), weighted by sample count. Lower is better for FPR and FNR; higher is better for the other six metrics.
The largest gain holds across all three tasks.
Perfect correction exceeds 93% on every evaluation task. The improvement over V2.0 remains between +9.10 and +10.60 percentage points rather than depending on one favorable sample set.
Task 02
n=991Task 03
n=2,933Built around the point where content becomes public.
V3.0 focuses on high-consequence pre-publication text: subtle wording errors, context-dependent corrections, dense domain vocabulary, and workflows where a plausible but unnecessary edit is itself a production risk.
A harder evaluation standard
The evaluation set was rebuilt around difficult, context-heavy examples drawn from high-standard content workflows instead of routine typo detection.
More corrections that can ship
Perfect correction rises by 9.50 percentage points, strengthening the model's ability to return a fix an editor can adopt directly.
Tighter alert control
Missed errors fall sharply while false alarms also decline, preserving editorial attention and supporting stable integration into production review.
Not one favorable metric. A consistent upgrade.
Across every dataset and metric tracked, V3.0 wins 43 of 47 direct comparisons. All eight sample-weighted aggregate metrics improve over V2.0.
91.5% of measured comparisons favor V3.0.
The result spans detection, correction, precision, recall, and final correction quality—evidence of a system-wide improvement rather than a single benchmark spike.
V3.0 leads 35 of the 39 comparisons measured within the three individual sample sets.
After weighting by sample count, V3.0 improves every one of the eight aggregate metrics.
For high-standard content operations.
Yanlan V3.0 is now fully available for enterprise and government deployments, including private environments and workflows with organization-specific terminology.
Yanlan V3.0 is now fully available.
Enterprise and government teams can evaluate the upgrade against representative production text and deployment requirements.