In my last post, I played with late interaction for instruction-following retrieval. This time I want to talk about the other half of every training recipe, the part I usually do not think about very much: the optimizer.

Muon is everywhere right now. Instead of applying the raw momentum, it orthogonalizes the momentum of each hidden weight matrix with a few Newton–Schulz iterations. In LLM pretraining it has been reported to match AdamW with about half the compute (Liu et al., 2025), and NorMuon adds neuron-wise normalization on top. It has also reached the retrieval world: in a small ablation on a NanoBEIR subset, the mxbai-edge-colbert-v0 report found Muon at its best learning rate slightly ahead of the best AdamW run (0.599 vs. 0.592 nDCG@10), though behind that run at every other learning rate, and concluded that Muon appears to be a strong optimizer for ColBERT training.

That made me curious. Fine-tuning a retriever is a pretty different regime from pretraining an LLM: the starting model is already contrastively trained, batches are small, training is short, and what we really care about is how well the model transfers to domains it has never seen. Whether Muon’s advantages carry over to fine-tuning a model that was pretrained with Adam is itself an open question. So I asked a narrow version of it: if I give AdamW and Muon the same tuning effort, does Muon give me a better retriever?

The short answer is no, at least not in my setup. Muon and NorMuon fit the training data better in every seed I ran, but the retrievers they produced were no better on BEIR (Figure 1). The details are below.

Quick links
Models on Hugging Face · Code, configs and all result tables on GitHub

−0.4 −0.2 0 +0.2 +0.4 +0.6 ← AdamW better Muon-class better → Muon (5e-4) +0.09 [−0.51, +0.70] NorMuon (5e-4) +0.04 [−0.32, +0.40]
Figure 1. 14-task BEIR nDCG@10 (×100) of the selected Muon and NorMuon recipes minus AdamW 5e-5, paired by training seed. Open circles are seeds 42, 2027 and 3407; the filled circle and bar are the mean and its 95% t-interval.

Setup

I fine-tuned lightonai/DenseOn-unsupervised (Sourty et al., 2026), a 22-layer ModernBERT-style encoder with 768-dimensional embeddings that has already gone through contrastive pretraining. The training data are 500K queries from lightonai/embeddings-fine-tuning, each with one positive and seven mined hard negatives. I use InfoNCE with temperature 0.02 and no in-batch negatives, one epoch (3,907 steps) at batch size 128, 10% warmup followed by linear decay, and bf16. One run takes about 7.8 hours on four GPUs, whichever optimizer I use.

Muon and NorMuon only update the 88 hidden weight matrices (attention and MLP). As the Muon authors recommend, everything else, here the token-embedding matrix and 45 norm weights, is trained by a regular AdamW running alongside. To keep the comparison clean, this AdamW uses exactly the settings of the best AdamW run (learning rate 5e-5, same weight decay), so the only difference between the optimizers is how the hidden matrices are updated. For Muon and NorMuon I use momentum 0.95 and five Newton–Schulz steps, plus β₂ = 0.95 for NorMuon; weight decay is 0.01 everywhere except the norms.

Each optimizer gets six peak learning rates: 1e-6, 3e-6, 1e-5, 3e-5, 5e-5 and 1e-4 for AdamW, and 1e-4, 2e-4, 3e-4, 5e-4, 1e-3 and 3e-3 for Muon and NorMuon. Two rules decide the rest:

  • Learning rates are selected by the contrastive loss on 4,096 held-out queries, never by BEIR.
  • Evaluation is nDCG@10 on LightOn’s decontaminated versions of 14 BEIR tasks (test queries and documents that also appear in the mGTE pretraining data are removed), on full corpora and with equal task weights. Five of these tasks are in-domain, because their source datasets are part of the training mix (FEVER, FiQA-2018, HotpotQA, MS MARCO and NQ); the other nine are out-of-domain.

Each selected recipe is then retrained with two more seeds (2027 and 3407), without any retuning.

Results

Validation loss picks 5e-5 for AdamW and 5e-4 for both Muon and NorMuon (Figure 2). Both Muon-class optimizers reach a lower validation loss than the best AdamW run, but on BEIR their sweeps from 1e-4 to 1e-3 sit at AdamW’s level rather than above it.

Validation loss 0.18 0.19 0.20 0.21 1e-5 1e-4 1e-3 peak learning rate BEIR nDCG@10 58.8 59.0 59.2 59.4 1e-5 1e-4 1e-3 peak learning rate
AdamWMuonNorMuon
Figure 2. Learning-rate sweeps at seed 42, each optimizer on its own learning-rate scale. Left: contrastive loss on 4,096 held-out queries, which selects the learning rate (ringed). Right: 14-task BEIR nDCG@10 (×100). Off the plotted range: AdamW at 1e-6 and 3e-6 and every run at 3e-3 (BEIR 55.7 to 58.3).

Then I retrained the selected recipes with two more seeds:

Recipe Seed 42 2027 3407 Mean vs. AdamW [95% CI] p
AdamW 5e-5 59.18 59.13 59.00 59.10 – –
Muon 5e-4 59.08 59.15 59.37 59.20 +0.09 [−0.51, +0.70] 0.57
NorMuon 5e-4 59.06 59.23 59.15 59.15 +0.04 [−0.32, +0.40] 0.67

Neither Muon nor NorMuon beats AdamW. Three seeds cannot show that the recipes are equivalent, since the intervals are 0.7 to 1.2 points wide, but they do bound the effect: for NorMuon, the 95% interval excludes an advantage above 0.40, while for Muon, one favorable seed (3407, +0.37) keeps the upper end at 0.70.

So what does Muon do differently?

That does not mean the two optimizers behave the same. In every seed, the Muon and NorMuon recipes

  • end with a lower training loss (0.234 to 0.237, compared with 0.240 for AdamW, averaged over the last 200 steps),
  • reach a lower loss on the held-out queries, by 0.002 to 0.005,
  • and score higher on average over the five in-domain BEIR tasks, by 0.04 to 0.63.
Recipe Train loss (42 / 2027 / 3407) Validation loss (42 / 2027 / 3407)
AdamW 5e-5 0.2395 / 0.2402 / 0.2400 0.1799 / 0.1844 / 0.1805
Muon 5e-4 0.2349 / 0.2342 / 0.2367 0.1776 / 0.1812 / 0.1757
NorMuon 5e-4 0.2346 / 0.2341 / 0.2359 0.1763 / 0.1802 / 0.1752

They also move the hidden weights further away from the pretrained model: a 3.0% relative change, compared with 1.7% for AdamW at seed 42. The nine out-of-domain tasks, however, do not follow (Figure 3). There, the differences range from −0.30 to +0.32 and average −0.06 for Muon and −0.10 for NorMuon.

0 +0.2 +0.4 +0.6 −0.4 −0.2 0 +0.2 +0.4 +0.6 in-domain tasks (5): difference to AdamW out-of-domain tasks (9)
Muon 5e-4NorMuon 5e-4
Figure 3. Difference to AdamW 5e-5 in the same seed on the five in-domain tasks (x) and the nine out-of-domain tasks (y), nDCG@10 ×100, one point per seed. All six points sit right of zero; four of the six sit below zero.

What I take from this

  • No learning rate makes Muon better here. Even if I pick the learning rate by BEIR itself, which my selection rules forbid, the best seed-42 Muon run (3e-4, 59.21) and the best NorMuon run (2e-4, 59.18) are level with AdamW (59.18). The learning rate matters far more than the optimizer: across the six rates, AdamW’s BEIR score at seed 42 ranges from 56.5 to 59.2 and Muon’s from 55.7 to 59.2, while switching optimizers at the selected rates moves the three-seed mean by only +0.09 (Muon) or +0.04 (NorMuon).
  • Muon did worse on scientific and biomedical retrieval. On all four scientific and biomedical tasks, Muon scores below AdamW in every seed: SciFact (−0.4 on average), NFCorpus (−0.4), SCIDOCS (−0.5) and TREC-COVID (−0.5). NorMuon is below AdamW in 10 of these 12 comparisons. Both gain in every seed on ClimateFEVER (+1.4 for Muon, which shares FEVER’s Wikipedia corpus), FiQA (+1.1) and NQ (+0.6). Other tasks swing a lot between seeds (Touché-2020 is −1.83, −1.34 and +0.05 for Muon), so I read these numbers as directions rather than precise effect sizes. Whether Muon helps therefore depends on the target: close to the training mix it can, on specialized scientific text it did worse.
  • Held-out loss tunes Muon for fit, not transfer. At seed 42, raising Muon’s learning rate from 1e-4 to 1e-3 lifts its in-domain score from 72.5 to 73.5 but lowers its out-of-domain score at every step, from 51.6 to 50.8; NorMuon shows the same trade-off. The held-out queries come from the training distribution, so the held-out loss rewards in-distribution fit and selects 5e-4, about 0.35 points below the out-of-domain peak at 1e-4 (Muon) or 2e-4 (NorMuon). For AdamW, the selected 5e-5 is also the top of its out-of-domain curve. These are single-seed sweeps and 0.35 points is close to the seed-to-seed spread, but the direction holds across rates for both optimizers. For a general-purpose retriever, I would try Muon at smaller learning rates and check them on out-of-domain data.

So for fine-tuning an already contrastively pretrained retriever at batch size 128, Muon is not a free upgrade. I would not read this as contradicting the mxbai report, which trained late-interaction models in a different pipeline; it just means that the advantage is not automatic.

My Remaining Questions

  • Does Muon help more when the model has more to learn, for example when fine-tuning from a plain MLM checkpoint such as ModernBERT-base instead of an already contrastively trained one?
  • Is the setup tilted toward AdamW from the start? Everything before my fine-tuning used Adam-type optimizers: ModernBERT was pretrained with StableAdamW (Warner et al., 2024), and DenseOn’s contrastive pretraining used AdamW, the default in its released training script. Earlier work suggests this matters: Qu et al. (2026) find that switching an Adam-pretrained model to Muon for fine-tuning degrades performance because the two optimizers favor structurally different weights, and Liu, Wang and Zhang (2026) find that full fine-tuning with the same optimizer as pretraining forgets less. A fairer test would start from an encoder pretrained with Muon.
  • Does the picture change with much larger batches?
  • Would late-interaction models, like the ones in the mxbai report, behave differently?
  • Weight decay is applied in proportion to the learning rate, so at the selected rates Muon’s decay on the hidden matrices is ten times AdamW’s. I did not separate that effect, and I did not tune momentum, the number of Newton–Schulz steps or NorMuon’s β₂ either.

That is it for this one. As always, I would really appreciate feedback, corrections, or pointers to related results, especially if you have seen Muon help (or not help) in your own retrieval training.