Updates from Coolwei AI Lab: research published, products shipped, and moments worth recording.

Our August frontend update introduces Yanlan-expand: single-character differences become phrase-level highlights, keeping the context of each correction in view.
Read full
All 8 weighted metrics ahead of V2.0, with 43 of 47 comparisons won. Chinese pre-publication correction for publishing, government communications, and enterprise content.
Read fullOur August frontend update introduces Yanlan-expand: single-character differences become phrase-level highlights, keeping the context of each correction in view.
All 8 weighted metrics ahead of V2.0, with 43 of 47 comparisons won. Chinese pre-publication correction for publishing, government communications, and enterprise content.
Models are like students who grind problem sets but never check their paper: they solve, they don’t fix. ThinkTwice trains “checking” into a skill — an 11.5-point gain on AIME pass@4.
Engines are like master craftsmen who cannot teach: accurate, but unable to explain. Master Distillation gives a 4B model concise puzzle commentary that surpasses its teacher.
Rewriting official prose into plain language has no yardstick in most languages. Five languages, 9,519 sentences, written by native speakers — the first open evaluation for low-resource simplification.
50 real iOS feature tasks, 449 human-written tests, ~500K lines of production code, and a best task pass rate of 12%.
False-alarm rate down 10×, processing 15× faster, deployment cost down 10× — text, video subtitles, audio transcripts and face checks now run in one system.
A report should read the same in any format; models often disagree with themselves. SEAM quantifies cross-modal inconsistency in 21 vision-language models across chess, molecules, scores, and graphs.
The start of a long-running collaboration to evaluate coding agents on real mobile production codebases.
Focused on safe deployment, evaluation, and real-world applications of large language models.
Like a teacher writing comments, Report Cards auto-write behavior reports for models — verified to genuinely help people tell models apart. A NeurIPS SoLaR Spotlight.