← Back to news News release · 29 October 2025

Yanlan model upgrade: full-modality Chinese correction goes live

A 10× lower false-alarm rate, 15× faster processing, 10× lower deployment cost. Text, video subtitles, audio transcripts and face checks run in one system.

Start using Product page V3.0 upgrade brief
Yanlan illustration
10×Lower false-alarm rate · 5% → 0.5%
15×Faster processing · 2,000 → 30,000 chars/s
10×Lower deployment cost · ¥2,000,000 → ¥200,000
99.9%Annual availability SLA

On 29 October 2025 Coolwei AI Lab released the upgraded Yanlan model; the update reached every service node on 28 September. The new version runs on Coolwei's own full-modality architecture, trained with 6 stages of mixed supervised fine-tuning (SFT) and reinforcement learning (RL): the false-alarm rate falls from 5% to 0.5%, single-machine throughput rises from 2,000 to 30,000 characters per second, and deployment cost drops from ¥2,000,000 to ¥200,000.

One upgrade, three hard numbers moved at once. A 10× lower false-alarm rate, 15× faster processing, 10× lower deployment cost. Text, video subtitles, audio transcripts and face checks run in the same system, already serving government communications, news and publishing, education, and enterprise document review.

Three hard numbers, moved together

The upgrade resets the three figures that decide what deployment costs. The false-alarm rate falls from 5% to 0.5%. Single-machine throughput rises from 2,000 to 30,000 characters per second. Full-system deployment cost falls from the H100 tier at ¥2,000,000 to the RTX 4090 tier at ¥200,000, a 90% reduction.

Three hard numbers: previous generation → Yanlan modelEach row is scaled on its own axis; read the labels, not the bar lengthsFalse-alarm ratelower is better5%0.5%↓ 10×Throughputhigher is better2,000 chars/s30,000 chars/s↑ 15×Deployment costlower is better¥2,000,000¥200,000↓ 10×
Grey bars are the previous generation, blue bars the upgraded Yanlan model. Rows are scaled independently, so bar lengths are not comparable across rows.
Fewer false alarms.A 0.5% false-alarm rate keeps review teams working on real issues instead of noise.
Faster turnaround.30,000 characters per second on one machine carries tens of millions of characters per hour on a single configuration.
A lower hardware bar.Consumer cards such as the RTX 4090 and 5090 are enough, and domestic silicon including Huawei Ascend and Hygon DCU is fully supported.

6-stage mixed training and routed processing

Yanlan is trained in 6 mixed stages of supervised fine-tuning (SFT) and reinforcement learning (RL), tightening the model's detection and correction boundaries stage by stage. At inference a classifier predicts the text type — government, legal, media — and routes the input to the central dictionary, the correction model, or a custom dictionary, before post-processing and the result integration layer produce the output. That routing takes overall complexity from O(n²) to O(nlogn + kn), raises parallel capacity about 30×, and cuts memory use by 60%.

Routed-processing architecture6-stage SFT + RL training · the classifier decides the pathInput dataMultimodal perceptionClassifierRouted processingCentral dictionaryError correctionCustom dictionaryPost-processingResult integrationHigh-quality outputThe classifier reads the text type first; each path then handles its own case.
System architecture. The three paths run in parallel and the result integration layer produces one output; post-processing holds the false-alarm rate at 0.5%.

The pipeline is layered by precision: L1 lexical analysis at <1ms per thousand characters, L2 grammar detection at 30ms, L3 semantic understanding at 500ms. Each layer is optimised on its own, coarse filtering and fine correction stay separate, and the false-alarm rate holds at 0.5%.

Every modality in one system

Beyond text, video subtitles, audio transcripts and face checks run in the same system, sharing one dictionary and one post-processing chain.

Accuracy by modalityHorizontal axis shows the 90% – 100% window100%90%Text correction99%+Video subtitles99%+Audio transcripts96%+Face checks98%+Four modalities, one system — no separate integration per channel.
Internal evaluation. The axis starts at 90% to show differences in the high range.
Video subtitle correction.Joint visual and textual understanding reads context from the frame; one minute of video is processed in 1–5 seconds, with specialist terms and institution names in government and legal material held exact.
Audio transcript correction.Speech-feature recognition on Coolwei's full-modality architecture, covering regional dialects and noisy environments; one minute of audio is processed in 1–5 seconds, carrying meeting recordings and video soundtracks.
Face checks.98%+ accuracy on officials removed from office, matched across angles and occlusion; one hour of video takes 2–3 minutes, 20–30× faster than before, and new entries take effect within 24 hours with no retraining.

Billion-scale training data and authoritative dictionaries

Training data moved from millions to billions of samples, a 1000× increase, drawn from authoritative media and official documents and refreshed daily, so error patterns are covered far more completely.

On the regulatory side the full national laws and regulations database is loaded — 16,740 documents covering the constitution, statutes, administrative regulations, local regulations and judicial interpretations — synchronised daily with the official source, with repealed and amended provisions detected automatically so historical text is never cited. The officials dictionary is collected daily from the Central Commission for Discipline Inspection site, new entries land within 24 hours, and images, video and documents can be checked in batch with position changes updated automatically.

On-premise deployment and production reliability

All data is processed locally and never leaves the internal network, cutting end-to-end latency 90% against a cloud round trip. Hardware runs on consumer cards such as the RTX 4090 and 5090 and on domestic silicon including Huawei Ascend and Hygon DCU; a single machine within a ¥200,000 budget carries tens of millions of characters per hour. Three deployment shapes are available: on-premise (data stays inside the network, highest security tier), private cloud (dedicated resources with elastic scaling), and Coolwei cloud (core data local, compute in the cloud).

99.9%Annual availability SLA
<50msP99 latency · 99% of requests under 50ms
1000+Concurrent QPS per machine
>4KPer-card throughput · chars/s

The system is built for 7×24 operation. Health checks monitor runtime state and isolate faults automatically, services restart on failure, and load peaks trigger graceful degradation that protects core functions. QPS, latency and throughput are monitored live, alerts reach email, SMS and DingTalk, and performance reports are issued on a regular cycle. Multi-level access control and complete operation logs support security audit and incident tracing.

Where Yanlan is already running

The Yanlan model runs in production across the following high-standard content settings.

Government communications.Pre-publication review of official documents, reports and notices for government bodies and public institutions, checking political accuracy and regulatory citations in real time.
Broadcast and media.Subtitle correction and batch video quality control for broadcasters, news organizations and video platforms, keeping public-opinion risk contained.
Education and training.Essay marking, courseware review and textbook proofreading for education institutions and training platforms, giving teachers and students an accurate grammatical reference.
Enterprise document review.Automatic review of contracts, reports and email, catching grammar and formatting issues before documents go out.

The Yanlan model is now fully available

The full-modality Chinese correction system is now fully available: a 0.5% false-alarm rate, 30,000 characters per second on a single machine, ¥200,000 deployment cost, and a 99.9% annual availability SLA. Already serving government communications, news and publishing, education and training, and enterprise document review. Contact us for a trial.

Start using → Contact the team