Dust: Pretraining Transformers Without Backpropagation
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get school and study supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Q Labs describes Dust, a zeroth-order method that trains transformer language models by perturbing activations instead of calculating backpropagation gradients. The report says Dust approached or exceeded backpropagation in some tested settings, but doing so required larger virtual populations and substantially more compute. The results are research claims; the supplied material does not establish independent replication or practical large-scale deployment.

Q Labs has published a report describing Dust, a method for pretraining transformer language models without backpropagation, the standard procedure for computing training gradients. The researchers say the zeroth-order approach can approach backpropagation’s performance, and in some tested settings exceed it, when given a sufficiently large population of activation perturbations; those results come with higher compute costs and are not evidence yet of a practical replacement for backpropagation.

Dust estimates how to update a model by adding noise to its activations and measuring how each perturbation changes the loss. It then combines perturbations according to their rewards, with changes that reduce the loss contributing more strongly. The report says perturbations are applied independently at every token. Each token acts as a member of a virtual population, allowing one forward pass to evaluate many perturbations in parallel rather than materializing and evaluating a separate model for each one.

Q Labs reports that Dust’s estimates become more aligned with backpropagation’s gradients as the population grows, and says this alignment remained strong in experiments up to 1 billion tokens. The paper also says a 243 million-parameter model outperformed a model 120 times smaller at most population sizes tested. These are results described by the authors; the available source excerpt does not provide enough experimental detail to assess all comparisons, compute budgets or replication.

The report compares Dust with weight-space evolutionary strategies, including EGGROLL. Q Labs estimates that from 1 million tokens onward Dust is roughly 1,000 to 10,000 times more efficient than a transformer implementation of EGGROLL. The authors characterize this as an extrapolation, so it should not be read as a measured advantage across every workload. They also say Dust can approximate backpropagation closely at large populations, which require substantially more compute.

At a glance
reportWhen: Published October 2026; experiments and…
The developmentQ Labs published a report describing Dust, an activation-perturbation method for pretraining transformer language models without a backward pass.

A Search-Based Route to Training

Backpropagation is central to current deep learning because it computes gradients efficiently through differentiable models. Dust tests whether a search-based method that avoids that calculation can train language models at competitive quality. If such methods scale, researchers could gain another way to assign credit to internal model activity and explore training choices that depend less on differentiability.

The immediate evidence is narrower than that possibility. The report’s results suggest larger models may use Dust’s population more efficiently in the experiments described, but the method’s performance depends on population size and compute. The authors’ comparison with EGGROLL is extrapolated, and the supplied report does not establish that Dust is cheaper or more effective than backpropagation for practical production-scale training.

Amazon

GPU for AI training

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Weight Search to Activations

Evolution strategies generally search by perturbing model weights and evaluating each candidate. That can become expensive because many perturbed candidates must be represented and run. Dust instead perturbs activation values inside the model, using tokens in a batch as parallel members of the search population. This design draws on node perturbation, an existing family of methods that changes internal neural network activity rather than directly optimizing weights through backpropagation.

Q Labs frames the work against a broader question: whether methods that rely on brute-force computation could become more competitive as compute grows. Its report argues that backpropagation’s structure is efficient, especially when compute is limited, while a search-based method might have different strengths at larger budgets. That is the authors’ motivation, not a demonstrated general advantage for high-compute training.

“We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models.”

— Q Labs Research, in the report’s TL;DR

Compute Costs and Replication

The supplied report material does not make clear how Dust’s total compute requirements compare with backpropagation across the full range of model sizes and training budgets. Q Labs says large populations can make the method competitive, while also describing those populations as requiring substantially more compute. The practical trade-off remains unresolved.

The report’s EGGROLL efficiency figures are extrapolations, and independent replication is not documented in the source material. It is also unclear whether the reported results extend to longer training runs, different architectures, broader language benchmarks or models larger than those described. The 1 billion-token figure identifies a tested scale in the report, not proof of performance at every larger scale.

Further Tests at Larger Scales

The next evidence needed is detailed, reproducible comparison across matched compute budgets, model sizes and training data, including direct comparisons with backpropagation and weight-space search methods. Such tests would show whether Dust’s population-efficiency trend holds as training scales and whether its reported performance can be achieved at a practical cost.

Q Labs says the gradient estimates remain aligned with backpropagation at the scales tested, but the source material does not specify a timetable for follow-up results or an independent replication. For now, Dust is a research report describing a promising training approach, with its broader scaling and cost claims still to be checked.

Key Questions

What is Dust?

Dust is a zeroth-order training method from Q Labs that perturbs model activations and uses changes in loss to estimate updates, without a backpropagation pass.

How does Dust use tokens as a population?

The report says Dust perturbs activations independently at each token. Tokens in a forward pass act as members of a virtual search population, so the method can evaluate many perturbations in parallel.

Did Dust outperform backpropagation?

Q Labs says Dust approached backpropagation closely at large population sizes and exceeded it in some tested settings. The supplied material does not establish that this holds across tasks or at equal compute budgets.

Is Dust more efficient than EGGROLL?

Q Labs estimates Dust is about 1,000 to 10,000 times more efficient than a transformer implementation of EGGROLL from 1 million tokens onward. The report labels that estimate an extrapolation, so it is not a general measured result.

Can Dust replace backpropagation in large language model training?

The report does not establish that. It presents experimental results and scaling claims, while practical compute costs, broader performance and independent replication remain open questions.

Source: hn

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

New AI Tutor Achieves 0.71-1.30 SD Effect Size In Dartmouth Course [Pdf]

A new AI tutoring system shows effect sizes of 0.71-1.30 SD in Dartmouth’s course, marking notable progress in AI-assisted education.

Record-breaking heatwave to hit several areas of China

A severe heatwave is expected to impact several areas across China, with temperatures reaching historic highs, according to weather authorities.

Chicago Tornado

A confirmed tornado struck Chicago today, causing injuries and property damage. Authorities are assessing the situation amid ongoing severe weather alerts.

Tropical Storm Bertha Weather Forecast

Forecasts predict Tropical Storm Bertha will bring heavy rain and gusty winds to southeastern US coastlines. Authorities monitor developments closely.