What would you like to explore?
“On-policy distillation can dramatically shorten answers without reliably improving accuracy. In one Qwen3.5-35B-A3B run, responses shrank from 14,070 to 6,132 tokens—a 56% reduction—while accuracy changed from 84.0% to 85.2%, a difference within the evaluation’s roughly 1.6-point standard error. The student generated its own responses; the teacher only scored those student-selected tokens.”
Miles v0.1: Production-Level Post-TrainingDaily Publications
Featured Papers
Handpicked academic papers for being particularly interesting or impactful
Articles & Threads
Practitioner threads, guides, and articles from the AI community
Today on X.com
The Astra demo reel kept running, but this window the safety countercurrent finally arrived with a name attached. A pretraining researcher who spent three years at both OpenAI and Anthropic resigned publicly, writing that neither company is acting responsibly and both are "racing straight to self-im...










