We work on AI safety, agentic coding, multimodal evaluation, reasoning training, low-resource NLP, and model auditing, writing each result up as a bilingual explainer for researchers and general readers alike.


EMNLP 2026
Paid search once pushed dubious hospitals in front of patients; today sellers can rewrite pages so AI favors them. SafeGEO measures the real scale of that risk — and the room for defense — across 600 recommendation cases and 22 attack variants.
AI Safety · Recommendation Agents · GEO

KDD 2026
50 real iOS feature tasks, 449 human-written tests, ~500K lines of production code, and a best score of 12%.
Coding Agents · Benchmark · Xiaohongshu

arXiv · 2026
Turning “remitted prior to the commencement date” into “pay before the start date” decides whether public information is readable — yet outside English it barely had an evaluation. OasisSimp has native speakers write multi-reference simplifications for 9,519 sentences in five languages, all open.
Multilingual NLP · Dataset

arXiv · 2026
Engines are like master craftsmen who cannot teach: accurate, but unable to explain. Master Distillation gets a 4B model to 48.1% puzzle accuracy, above most frontier models, with far shorter explanations.
Distillation · RLVR

arXiv · 2026
A student who never checks their paper won’t fix mistakes; neither do models trained the standard way. ThinkTwice trains solving and revision with one correctness reward, gaining 11.5 points on AIME pass@4.
RLVR · Self-refinement
Three stages: a first paper selected as a NeurIPS SoLaR Spotlight in 2024, the lab founded and SEAM accepted at COLM in 2025, and five research lines released across 2026 — two of them at a CCF-A main conference and EMNLP.
Report Cards describes model ability in natural language, then validates the descriptions themselves with contrastive accuracy and Card Elo.
The same question in text or in image form often draws different answers. SEAM quantifies cross-modal consistency across 21 models, four domains and 9,600 evaluations.
In a single year the work extends from agent benchmarking to reasoning training, master distillation and multilingual datasets, with two results at a CCF-A main conference and EMNLP.