AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get school and study supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Q Labs Research introduced Dust, a zeroth-order training method that perturbs transformer activations rather than computing gradients with backpropagation. The October 2026 report says Dust approached or exceeded backpropagation in some tested settings, but often used substantially more compute; its claims about scaling are based on experiments and extrapolations, not evidence that it is ready to replace standard training.

Q Labs Research has published results for Dust, a method for pretraining transformer language models without backpropagation, the standard technique for calculating how model weights should change during training. The researchers report that Dust’s performance approached backpropagation’s in some experiments and exceeded it in some settings, while also warning that matching it can require a substantially larger population of perturbations—and, in turn, more computation.

Dust is a zeroth-order optimization method: instead of calculating a conventional gradient through the network, it perturbs internal activations and uses the resulting changes in loss to estimate an update. The report says perturbations are applied independently at each token. That lets the method treat tokens as members of a virtual population and evaluate them together in a forward pass, rather than separately materializing and evaluating many altered sets of model weights.

In its October 2026 report, Q Labs says Dust’s estimates become more aligned with backpropagation’s gradients as the population grows. The authors also report that the alignment remained strong across the model and data scales they tested, including experiments up to 1 billion tokens. They say a 243-million-parameter model outperformed a model 120 times smaller at most tested population sizes. These are findings reported by the researchers, not an independent assessment of the method.

The paper compares Dust with EGGROLL, a weight-space evolution-strategy method, and says Dust is roughly 1,000 to 10,000 times more efficient from 1 million tokens onward. That figure is based on the authors’ extrapolations. It concerns the comparison with EGGROLL, not a claim that Dust requires less compute than backpropagation. The report’s summary says Dust can approximate backpropagation closely at large populations, which involve substantially more compute, and can outperform it in some settings.

At a glance
reportWhen: Published October 2026
The developmentQ Labs Research published a report describing Dust, an activation-perturbation method for pretraining transformer language models without backpropagation.

The Cost of Skipping Backpropagation

The report challenges a common assumption in machine learning research: that methods without backpropagation cannot train large neural networks competitively. If activation perturbations can provide useful training signals at scale, researchers may gain another way to train models, including in settings where gradient calculation is difficult or where different learning rules are worth exploring.

But the results do not establish that Dust is a cheaper or more practical replacement for backpropagation. Its reported advantage over EGGROLL is a separate comparison, while the authors acknowledge that approaching backpropagation may call for larger populations and more compute. For developers and organizations, the relevant question is not only whether the approach can learn, but whether its training costs, hardware demands and final model quality compare favorably in full-scale runs.

The authors frame Dust as a sign that compute-intensive search methods could become more competitive as available computing grows. That is a research motivation, not a demonstrated forecast that such methods will overtake gradient-based training. The report offers evidence from its experiments; it does not show that Dust has produced a deployable language model matching a production system.

Amazon

high performance GPU for AI training

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Dust Differs From Weight Search

Backpropagation calculates gradients by propagating information about a model’s loss backward through its layers. Modern neural-network architectures, training software and hardware are built around that process. Zeroth-order methods instead evaluate how perturbations affect performance and use those evaluations to guide updates, avoiding the conventional backward pass.

Earlier evolution-strategy approaches perturb model weights. Their populations can be expensive to evaluate because individual candidate models must be represented and run. Q Labs’ proposed alternative perturbs activations rather than weights, using token-level perturbations to form what it calls a virtual population. The paper describes this as an adaptation of node perturbation, a longstanding family of learning methods.

The distinction matters because Dust’s efficiency claims are specific. The report’s projected comparison with EGGROLL does not settle how Dust performs against backpropagation on equivalent hardware, budgets and training objectives. Nor does a close match between estimated gradients and backpropagation’s gradients by itself establish equal downstream language-model capability.

“Dust approximates backprop closely at large population (i.e. substantially more compute) and in multiple settings even exceeds it.”

— Q Labs Research, in the report’s summary

Questions About Compute and Scale

The report does not establish whether Dust can match backpropagation’s overall training cost or model quality under a controlled, equal-compute comparison. Its claim of being 1,000 to 10,000 times more efficient than EGGROLL is described as an extrapolation, and the supplied summary does not give the assumptions needed to independently assess that projection.

It is also unclear how well the results transfer to larger models, longer training runs, different datasets or production-scale language-model tasks. The report says gradient alignment stayed strong up to 1 billion tokens, but that does not establish performance beyond the tested range. The available material does not specify independent replication, full benchmark results, or whether Dust-trained models perform as well on downstream evaluations.

Evidence Needed for Wider Adoption

The immediate next step is scrutiny of the paper’s full experimental details: how compute is counted, how population size affects results, and how Dust compares with backpropagation under matched budgets. Replication by other researchers would help test whether the reported performance and scaling behavior hold outside the authors’ experiments.

Further evaluations would also need to measure end-to-end training cost and language-model quality across larger models and longer runs. Until such evidence is available, Dust is best described as a reported research result that tests an alternative to backpropagation—not as a proven replacement for the method used to train modern transformers.

Key Questions

What is Dust?

Dust is a zeroth-order training method from Q Labs Research that perturbs transformer activations and uses changes in loss to estimate updates, rather than calculating gradients with backpropagation.

Does Dust eliminate all computation during training?

No. Dust avoids the conventional backward pass, but it still requires forward evaluations and perturbations. The report says matching backpropagation closely can involve substantially more compute through larger populations.

Did Dust outperform backpropagation?

Q Labs reports that Dust approximated backpropagation closely at large population sizes and exceeded it in some tested settings. The report does not establish that it is generally better or cheaper across models and training tasks.

What does the EGGROLL efficiency figure mean?

The authors say Dust is about 1,000 to 10,000 times more efficient than a transformer implementation of EGGROLL from 1 million tokens onward. They describe this as an extrapolation; it is not a comparison showing Dust is that much more efficient than backpropagation.

Is Dust ready to replace standard transformer training?

The reported findings do not show that Dust is ready to replace backpropagation. Independent replication, matched-compute comparisons and larger-scale language-model evaluations remain necessary to establish its practical costs and capabilities.

Source: hn

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Understanding the Internet of Things

What if everyday devices could connect and communicate seamlessly, transforming your life in ways you never imagined?

Why Carpet Cleaner Tank Size Matters More Than You Think

What you learn about carpet cleaner tank size can drastically improve your cleaning results and efficiency—find out why it matters more than you think.

Hell Arrives in Washington

Severe weather conditions have struck Washington, causing widespread disruption and raising concerns about climate impacts. Details are still emerging.

Understanding ESG Investing

For those exploring sustainable growth, understanding ESG investing reveals how responsible choices can shape your financial future—discover the key insights ahead.