TL;DR
Get school and study supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Q Labs Research describes Dust, a zeroth-order method that perturbs transformer activations to estimate learning updates without backpropagation. The October 2026 report says it approached or exceeded backpropagation in some tested settings and scaled to experiments involving up to 1 billion tokens, but the claims come from the researchers’ experiments and extrapolations; broader performance and compute costs remain uncertain.
Q Labs Research has presented Dust, a method for pretraining transformer language models without backpropagation, reporting that it was competitive with the standard training method in several experiments. The October 2026 research report describes a way to estimate updates by perturbing model activations, a result that could broaden how large neural networks are trained if it holds up beyond the reported tests.
Dust is a zeroth-order optimization method: rather than calculating derivatives through the network, it perturbs activations and uses changes in loss to estimate which updates may help. The researchers say the perturbations are applied independently at each token. They treat tokens as members of a virtual population, letting a single forward pass evaluate many perturbations in parallel instead of materializing and running a separate model for each population member.
In the report, Q Labs says Dust’s estimates become more aligned with backpropagation as the population grows and remain well aligned across the scales tested, up to 1 billion tokens. The team also reports that a 243-million-parameter model outperformed a model 120 times smaller at most population sizes. In some settings, Dust exceeded backpropagation, according to the authors; the report does not establish that result as a general advantage.
The researchers compare Dust with EGGROLL, an evolution-strategy method that perturbs weights. They estimate that, from 1 million tokens upward, Dust is roughly 1,000 to 10,000 times more efficient than a transformer implementation of EGGROLL. That figure is an extrapolation in the report, rather than a direct demonstration that Dust is more efficient than backpropagation at comparable end-to-end training quality.
A Different Route to Transformer Training
Backpropagation has been central to modern neural-network training because it efficiently assigns credit for errors across a model’s parameters. Dust tests whether a search-based approach can learn useful updates without calculating those gradients. If methods of this kind prove practical, researchers could have more options for training systems whose architectures or operations are difficult to differentiate.
The result also puts attention on the tradeoff between computation and structure. Dust uses population-based search, and Q Labs says it approaches backpropagation as the population grows. That may make it relevant where abundant compute is available, but the report does not show that the extra computation is economical for production-scale language models. Its comparison with EGGROLL addresses another zeroth-order method, not the full cost and quality comparison with standard training.
transformer model training hardware
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why Search Methods Matter
Most large language models are trained with backpropagation, which calculates gradients by passing error information backward through the network. Evolution strategies offer a different approach: evaluate perturbed versions of a model and use their results to guide changes. The challenge is that weight-space methods can require many separate model evaluations as the population grows.
Dust builds on activation-space or node perturbation, applying noise inside a network rather than creating separate weight-perturbed models. The Q Labs report argues that perturbing each token independently creates a virtual population that can be processed together. It frames this as a way to reduce the practical cost of zeroth-order search. The report is a research presentation; the supplied material does not establish independent replication or peer-reviewed publication.
“We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models.”
— Q Labs Research, report summary
AI research books on neural networks
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Limits of the Reported Evidence
The report’s results do not settle whether Dust can train state-of-the-art language models at large scale or match backpropagation at equal compute, training time, and final model quality. The supplied material describes experiments and extrapolations, but gives no independent replication. It also does not provide enough detail here to assess the full hardware, energy, or wall-clock costs of the largest comparisons.
The claim of efficiency against EGGROLL is specifically an estimate against a transformer implementation of that method, not a measured advantage over all zeroth-order approaches. It remains unclear how Dust performs across different architectures, datasets, population sizes, and downstream evaluations, and whether its reported alignment with backpropagation predicts comparable model capabilities.
machine learning optimization tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Replication and Larger-Scale Tests
The next evidence needed is independent replication and a fuller comparison with backpropagation under matched compute budgets, including training time and final model quality. Further tests across model sizes and tasks would show whether the reported population-efficiency trend persists beyond the settings Q Labs examined.
The report, as supplied, does not announce a release schedule, a planned benchmark, or a specific next experiment. For now, Dust is a research result that suggests a possible alternative training approach; whether it changes standard transformer practice depends on evidence about its costs and performance at broader scales.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Dust?
Dust is a zeroth-order method that perturbs a model’s activations and uses changes in loss to estimate updates, rather than computing gradients with backpropagation.
Does Dust eliminate backpropagation in language-model training?
The report presents Dust as a way to pretrain transformers without backpropagation, but it does not show that the method has replaced backpropagation in general use or at production scale.
What does the 1,000-to-10,000-times efficiency figure compare?
Q Labs estimates that Dust is about 1,000 to 10,000 times more efficient than a transformer implementation of EGGROLL from 1 million tokens upward. This is an extrapolation against that method, not a measured comparison with backpropagation.
What remains unproven?
Independent replication, performance at larger training scales, and comparisons at matched compute and final model quality remain open. The report does not establish that Dust is generally faster or cheaper than backpropagation.
Source: hn
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
