Q Labs published a paper this month claiming the first zeroth-order optimization method that is competitive with backpropagation at pretraining transformer language models. The method, called Dust, does not compute gradients analytically. It adds Gaussian noise to the output of each linear layer, independently at every token, runs a forward pass, and rewards each token’s noise by how much it lowered the loss. The reward-weighted noise, averaged over draws, becomes the gradient estimate. Authors Samip Dahal, Bishwas Mandal, Serdar Gülbahar, and Akshay Vegesna posted the work to qlabs.sh, with correspondence to [email protected].
The interesting part is not that zeroth-order optimization works in a small setting. It is that Dust’s authors report the opposite of what the field expects at scale. Zeroth-order methods are widely believed not to scale to large networks. Q Labs finds larger models are more population-efficient, not less. A 243M-parameter model outperforms a model 120 times smaller at most population sizes. That single result, if it holds, reframes overparameterization as a larger search space with better geometry rather than a liability to be managed.
What Dust actually does
The mechanism is node perturbation, a technique that dates to Widrow and Lehr’s 1990 work on adaptive neural networks. The usual argument for perturbing activations rather than weights is dimensionality: a layer’s output has d_out entries and its weights have d_out times d_in, so activation noise lives in a much smaller space. Q Labs pairs that with a specific trick. Because a transformer sequence holds a few thousand tokens, and Dust perturbs each token independently, one forward pass evaluates a few thousand population members per sequence instead of one. The authors call this a virtual population.
That is the whole efficiency argument. Weight-space evolution strategies like EGGROLL, described by Sarkar and coauthors in a 2025 arXiv preprint, make perturbed copies cheap with low-rank perturbations, but each member is still one sequence element of the batch. The population is bounded by the forward passes you can afford. Dust sidesteps weight space entirely. Every weight in the model trains this way except the 2L residual mixing scalars, which stay on ordinary weight-space ES.
The paper reports Dust is on the order of 10^3 to 10^4 times more efficient than a transformer implementation of EGGROLL from 1M tokens up, based on the authors’ extrapolations. That word matters. Extrapolations are not measurements, and the paper is explicit that it does not attempt to make Dust compute-efficient enough to replace backprop today.
The compute-rich bet
The framing here is Sutton’s bitter lesson, and Q Labs says so directly. The argument: differentiability and backprop are good inductive biases in a low-compute regime, where they make learning efficient, but in a high-compute regime they constrain which architectures work. The AlphaGo Zero comparison is the paper’s own. Bootstrapping on human data helped initially; with enough computation, the purely self-play network overtook it.
Q Labs extends that logic to credit assignment. Gradient-based methods, the authors write, fail to explore the loss landscape optimally, citing Liu, Papailiopoulos, and Achlioptas on bad global minima reachable by SGD. A search-based credit assignment algorithm, in their telling, is a step toward better generalization and a wider space of trainable architectures.
There is a second claim worth flagging. The authors argue activations are a more interesting space to search than weights because mechanistic interpretability has shown reasoning lives in the activations, citing Gurnee and coauthors and the Lindsey et al. “On the biology of a large language model” work. If that is right, training becomes a search over latent reasoning rather than a descent through parameter space. That is a much larger claim than the efficiency numbers, and the paper does not test it.
Larger models are more population-efficient, not less. That inverts the assumption that has kept zeroth-order methods on the margins.
What the results do and do not show
Dust’s gradient estimates align better with backprop’s as population grows, and the paper reports the alignment holds at every scale tested, up to 1B tokens. That is a real number and a real ceiling. One billion tokens is a research-scale pretraining run. Frontier pretraining runs are orders of magnitude larger, and nothing in the paper demonstrates Dust at that scale. The authors say as much: the goal is to lay foundations, not to ship a backprop replacement.
Two other things are left explicitly to future work. The paper does not train the new kinds of networks its method makes accessible, such as nets with an external program in the loop, or transformers looped over many steps that backpropagation through time struggles to train. Those are the architectures where a search-based method would matter most, and they are untested here.
What it means for builders
Read this as a research signal, not a tooling change. Nobody is swapping their training loop for Dust next quarter. The compute cost is the point: Dust approximates backprop closely at large population, which means substantially more compute, and it exceeds backprop in multiple settings only in that regime. If you are compute-poor, backprop still wins on every axis that matters.
The claim to watch is the scaling inversion. If larger models genuinely make better use of larger populations, then the economics of training shift in a direction the field has not priced in. Compute becomes not just a way to train bigger models but a way to train them differently. Q Labs has not shown that at frontier scale, and the paper’s own honest framing says so. But the 243M-parameter result is the kind of finding that gets replicated or refuted within a year, and either outcome is useful. The next test is whether anyone runs Dust past 1B tokens.