What would you like to explore?

Today

On-policy distillation can dramatically shorten answers without reliably improving accuracy. In one Qwen3.5-35B-A3B run, responses shrank from 14,070 to 6,132 tokens—a 56% reduction—while accuracy changed from 84.0% to 85.2%, a difference within the evaluation’s roughly 1.6-point standard error. The student generated its own responses; the teacher only scored those student-selected tokens.

Miles v0.1: Production-Level Post-Training

Daily Publications

669 this period · 21,965 total
Aug 23Sep 10

Featured Papers

View all

Handpicked academic papers for being particularly interesting or impactful

Fractal basins trap latent reasoning

Fractal basins trap latent reasoning

This paper shows that hard reasoning problems create fractal-like landscapes in a model’s hidden thought process, where trajectories linger near plausible but wrong answers before escaping to the correct one, revealing a measurable source of reasoning difficulty and slow inference.

2.5K312
WHALE: A Simple Recipe for Joint Harness-Weight Optimization

WHALE: A Simple Recipe for Joint Harness-Weight Optimization

This paper shows that WHALE’s simple strategy of alternately improving an LLM’s model and the code that orchestrates it consistently outperforms tuning either alone, especially when the performance bottleneck shifts between them.

33449
Multi-Mask Diffusion Language Models for Few-Step Generation

Multi-Mask Diffusion Language Models for Few-Step Generation

This paper shows that giving diffusion language models many mask tokens instead of one creates a better starting point for few-step generation, improving sample quality and transferring even to large models like LLaDA.

104
Quantifying the Carbon Emissions of Machine Learning

Quantifying the Carbon Emissions of Machine Learning

This paper introduces a practical calculator for measuring the carbon footprint of machine learning projects, revealing how factors like server location, training time, and hardware choices significantly impact environmental costs.

20
Embedding Surgery: Localized Updates for Adaptive Ranking Correction in Dense Retrieval

Embedding Surgery: Localized Updates for Adaptive Ranking Correction in Dense Retrieval

This paper shows that dense-retrieval rankings can be corrected on demand by making tiny, targeted edits to document embeddings—without retraining or rebuilding the index—while improving results and preserving overall search quality.

210
SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers

SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers

This paper shows that, even when computation, parameters, and memory are carefully matched, letting a sparse transformer revisit its middle layers provides a genuine refinement benefit, improving downstream performance while saving up to 18% of training compute.

72
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills

Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills

This paper introduces DisCo, which turns repositories and research papers into verified, task-relevant skills that give AI research agents practical know-how and more than double their performance on a challenging ML benchmark.

72
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation

Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation

This paper shows that transferring the direction of a stronger model’s improvements—rather than copying its behavior—helps smaller models learn faster, retain their own useful discoveries, and even surpass their teachers across successive transfers and multiple domains.

82

Today on X.com

1/5

The Astra demo reel kept running, but this window the safety countercurrent finally arrived with a name attached. A pretraining researcher who spent three years at both OpenAI and Anthropic resigned publicly, writing that neither company is acting responsibly and both are "racing straight to self-im...

Based on 130 tweetsYesterday