AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get school and study supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Q Labs Research describes Dust, a zeroth-order method that trains transformers by perturbing activations rather than calculating backpropagation gradients. The team reports competitive results in experiments and strong efficiency relative to one evolution-strategy comparison, while saying Dust needs larger populations and substantially more compute to approximate backprop closely.

Q Labs Research has published a report on Dust, a method for pretraining transformer language models without backpropagation. The researchers say their zeroth-order approach, which perturbs a model’s activations and uses changes in loss to estimate learning updates, can compete with backprop in some tested settings; the finding is experimental, not evidence that it is ready to replace standard training.

Dust works by adding perturbations to activations—the intermediate values passed through a neural network—rather than making small changes to model weights. The report says the method applies perturbations independently at each token. That lets tokens act as members of a “virtual population”: one forward pass can evaluate many perturbations together, rather than running a separate model for every candidate as conventional weight-space evolution strategies do.

Q Labs characterizes Dust as a zeroth-order optimization method: it uses loss measurements to score perturbations and combines them to estimate an update, rather than computing a backpropagation gradient. According to the report, Dust’s estimates align more closely with backpropagation as the population grows. The researchers also report alignment across the scales tested, including experiments reaching 1 billion tokens.

The report says that a 243-million-parameter model outperformed a model 120 times smaller at most tested population sizes. It also estimates that, from 1 million tokens onward, Dust is roughly 1,000 to 10,000 times more efficient than a transformer implementation of EGGROLL, an evolution-strategy method. Those figures are the authors’ experimental results and extrapolations; they are not a general demonstration of lower end-to-end training costs across models, datasets, or hardware.

At a glance
reportWhen: Report dated October 2026
The developmentQ Labs Research has published a report describing Dust, an activation-perturbation method for pretraining transformer language models without backpropagation.

A Different Route to Transformer Training

Transformer pretraining currently depends heavily on backpropagation, which calculates how changes to model parameters affect the loss. A method that learns without that calculation could broaden the training approaches researchers can test, especially if it can take advantage of parallel computation. Dust’s central proposal is to search in activation space while evaluating many token-level perturbations in the same forward pass.

The immediate importance is narrower than a replacement for backprop: the report presents evidence that zeroth-order training may be more competitive than commonly assumed. Q Labs also argues that larger models can be more population-efficient in its experiments. Whether that advantage persists at practical model sizes, training durations, and compute budgets is not established by the claims summarized in the report.

Amazon

transformer language model training kit

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Weight Search to Activation Search

Evolution strategies and other zeroth-order methods explore candidate changes and use measured outcomes to guide learning. A central cost of weight-space approaches is evaluating many separately perturbed model instances. Q Labs says Dust avoids materializing those candidates by perturbing activations within a single model pass, drawing on the established idea of node perturbation.

The report frames backpropagation as an effective inductive bias that has shaped neural-network architectures, optimizers, and hardware. It asks whether greater computing resources could make more search-based approaches useful. That is the authors’ motivation, not a result that more compute will necessarily cause such methods to surpass gradient-based training.

“We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models.”

— Q Labs Research, report summary

Compute Costs and External Validation

The report’s claims do not establish how Dust compares with backpropagation on total training cost, including wall-clock time, energy use, hardware, and the final model’s performance on independent evaluations. The stated efficiency advantage over EGGROLL is based partly on extrapolations, and the source material does not provide a common broad benchmark across training methods.

It is also unclear how results would change with different architectures, data mixtures, longer training runs, or larger populations. The report argues that Dust may benefit from compute-rich settings, but does not establish that increasing compute will make it outperform backpropagation generally. Independent replications and detailed comparisons would help determine how far the results extend.

Replication and Larger-Scale Tests

The next meaningful test is whether other researchers can reproduce the reported findings and compare Dust with backpropagation under matched compute and evaluation conditions. Further results would also need to clarify the trade-off between population size and training cost, and whether performance holds as model and dataset scales increase.

The source describes a research report, not a product release or a confirmed change to standard language-model training. Until comparable, independently checked results are available, Dust is best understood as a promising experimental alternative whose practical advantages remain open questions.

Key Questions

What is Dust?

Dust is a zeroth-order training method from Q Labs Research that perturbs a transformer’s activations and uses the resulting loss changes to guide updates, instead of calculating backpropagation gradients.

Does Dust eliminate backpropagation in all transformer training?

No. The report describes experiments with Dust and claims competitive results in some settings. It does not show that Dust can replace backpropagation across models, datasets, or practical training workloads.

How does Dust evaluate many perturbations?

According to the researchers, Dust perturbs activations independently at each token. The tokens form a virtual population that can be evaluated together in a forward pass, rather than requiring a separate model instance for every candidate.

What does the report say about Dust’s efficiency?

Q Labs estimates that, from 1 million tokens onward, Dust is around 1,000 to 10,000 times more efficient than a transformer implementation of EGGROLL. The comparison relies on the report’s measurements and extrapolations and should not be treated as a general cost advantage over backpropagation.

What evidence would clarify whether Dust is practical?

Independent replication and matched-compute tests would help establish how Dust performs against backpropagation in training time, resource use, and model quality across larger and more varied tasks.

Source: hn

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Does Georgism Work? Five Years Later

Five years after renewed interest, experts evaluate whether Georgism has achieved its economic goals, with ongoing debates about its success and limitations.

Will The **High Temp In Miami** Be 92-93° On Jul 29, 2026?

Market activity suggests a possibility that the high temperature in Miami may reach 92-93°F on July 29, 2026, but no official forecast confirms this yet.

Unwetter Gewitter

Severe thunderstorms and heavy rainfall have caused disruptions across parts of Germany, prompting warnings from the German Weather Service.

Peru Earthquakes

A series of earthquakes struck Peru, resulting in injuries and property damage. Authorities are assessing the impact as rescue efforts continue.