All findings
Research 16 min read

Specialists can be blended on demand.

Can a specialist be built without training one? A blend of three public LoRA adapters beat the base model by 9 points on 200 questions none of them had seen. The first run of the same experiment stopped.

Can a specialist be built without training one?

That is the question this experiment asks, narrowly enough to be answered either way:

Does a LoraHub-style blend of two or more public LoRA adapters answer better, on a domain genuinely in between the domains they were trained on, than the base model, or than the single best adapter?

Everything here exists to answer that, and the answer is allowed to be no. This is Stage 1 of two. Stage 2, wiring on-demand adapter fusion into the gateway, is gated on this clearing a five-point threshold and is not built unless it does.

What this does not claim. Nothing here says the network gives better answers. It says whether this blending procedure does, on this base model, on these questions.

What is being compared

configurationhow it is chosen
(a)base modelQwen/Qwen2.5-1.5B-Instruct, untouched
(b)best single adapterwhichever of the seven scores highest on the 50 maths questions
(c)the blendLoraHub weights, searched on 20 held-out questions only

(b) is deliberately chosen with hindsight, on the same questions the blend is then judged on. It is the strongest baseline available, because it already knew the answers. A weaker one would only have flattered the blend.

Seven public adapters, six of them blendable

Seven public LoRA adapters for Qwen/Qwen2.5-1.5B-Instruct, all pinned to a commit in adapters.json so a re-run months from now loads the same bytes.

domainadaptersblendable
maths3yes
code2yes
reasoning1yes
medical1no

Blendable is decided by one thing: whether the adapter can be averaged in factor space. LoraHub sums the A and B matrices across adapters and takes the product of the sums, so every adapter in a blend must agree on r, lora_alpha and target_modules. Six agree exactly (r=16, alpha=32, the same seven projections, ~70.5 MB each). The medical adapter is r=64, alpha=16 with four targets, so it is scored as a single adapter and kept out of blends. harness.load_registry asserts that precondition on every run rather than trusting the registry file.

The obvious in-between domain for this base would be physics. No physics or science adapter exists for Qwen2.5-1.5B at any download count. So the in-between domain tested is maths word problems, on a base blended from maths and code adapters. Every maths item needs multi-step arithmetic, which is what the code adapters are nominally good at.

The question sets

Three files. Both scored sets are generated, and both sets’ answers are computed rather than asserted, so a question and its answer cannot silently disagree.

  • Maths (50). Ten families: unit price and change, discount then tax, rate/time/distance, averages, ratio sharing, compound growth, work rate, percentage increase then decrease, mixture concentration, arithmetic series. Numbers come from a seeded RNG, and every answer is computed in fractions.Fraction and quantised half-up, so it is exact.
  • Control (50). Eight families rendered from encoded data tables: element symbols, atomic numbers, capitals, currencies, continents, SI units, planetary order, chemical formulae.
  • Held out (20). Eighteen maths and two control, used only to choose the blend weights and never scored.

check_questions.py runs 1,037 assertions over all three before anything is scored on them: counts, no duplicate prompts, every maths answer re-derived from the item’s own parameters, every parameter proven load-bearing by perturbing it and requiring the answer to move, every parameter’s value shown to appear in the question text, control answers checked against a fresh table lookup, held-out items proven absent from both scored sets, and even family coverage.

A generated set can be self-consistent and still be wrong

The first generated maths set contained this item:

A shop sells markers at £5.20 each. Alex buys 6 of them and pays with a £10 note. How much change, in pounds, does Alex receive? Answer: −21.20

Every check above passed on it. The answer re-derived from the parameters, both parameters were load-bearing, both appeared in the question. It was an impossible question. The generator had chosen the note by mixing pounds and pence, so £10 was offered for £31.20 of goods.

Re-derivation cannot catch a wrong premise. It only proves the arithmetic agrees with itself. It was caught by printing one item per family and reading them, which is why that step is in the plan. check_questions.py now also asserts every maths answer is a positive quantity, which is the one check here that is not about internal consistency. That check was then verified to actually fire on the historical bug, since a check that only ever passes on good data proves nothing.

Method: a forward-pass delta, and no PEFT

One base model loaded once. Adapters and blends are applied as a forward-pass delta through a LoRALinear wrapper rather than merged into the weights.

Applying a LoRA is W_eff = W + (alpha/r) * B @ A, and a LoraHub blend is the same expression with A and B replaced by the weighted sums. That is tensor arithmetic, so blend.py does it in plain torch + safetensors. The reason is dependency risk rather than purity: transformers is at 5.8.0 here, a major release, and peft is the package most likely to break against it. With PEFT out of the scoring path, transformers is only asked to load Qwen2.5 and generate, which is the stable part.

The saving is real beyond dependencies. A blend is ~74 MB of matrices rather than a ~3 GB copy of the model. Swapping configuration is 196 pointer assignments. Returning to the base model is enabled = False rather than a reload, and configuration N+1 cannot be contaminated by configuration N, because the base weights are never written to.

Note the size. A six-way blend is not smaller than a single adapter. It is the same 74 MB. The blend has the same rank and shape as its inputs, because the sums ΣλᵢAᵢ and ΣλᵢBᵢ are each still rank r, so it costs exactly what one adapter costs no matter how many are blended in. The 74 MB is what the file on disk actually weighs.

Everything else is held fixed across configurations: same chat template, same instruction, same max_new_tokens, no per-configuration prompt tuning, greedy decoding everywhere. Greedy matters because a sampled run would put a random variable inside a comparison of a few points. One answer extractor, fixed in advance and never revisited after seeing which configuration it favours: the last number in the completion for maths, compared with a tolerance of half a unit in the item’s own stated decimal place, and a word-boundary phrase match for the control set.

The duplicated sum, and the check that caught it

The path is not merely asserted to be equivalent. blend.py writes the blend out in PEFT’s own layout and then checks, before the scored run, that applying the written file and applying the in-memory matrices give identical next-token distributions.

That check earned its keep. The LoraHub sum ΣwᵢAᵢ was originally written out twice, once in the wrapper that applies a blend and once in the writer that saves one, and the two copies drifted. The wrapper omitted the weight on its first term, so it computed A₀ + Σ_{i>0} wᵢAᵢ where it should have computed Σᵢ wᵢAᵢ. For a six-way blend the first adapter entered the sum at full weight rather than at its fitted weight, so the two were different models.

Nothing downstream could see it. The search ran to completion, every candidate produced a plausible score, and the artifact was written correctly from the correct formula. The only symptom was that the matrices on disk and the matrices being scored disagreed, which is exactly what the logits comparison tests. It failed the first run with a maximum logit difference of 3.36 against a required 1e-3. The run stopped before scoring a single question, so no number in results.json was ever produced by the wrong path.

The fix was to delete the copy rather than correct it. The sum now exists once, as harness.weighted_factor_sums, and both callers use it. A model-free equivalence check runs as a pre-flight, so the same class of fault surfaces in the first second rather than after a fifty-minute weight search. It was verified to fail on the original buggy formula, since a check that has only ever seen working code has not been tested.

The run that stopped

Two runs, and they are the same experiment twice. The first fitted blend weights on 20 held-out items and scored everything on 50. It read +12.0 points over the base model while its own paired interval still spanned zero, which is the case the pre-registered rule says must stop rather than resolve itself. The second took those weights frozen, ran no search at all, and scored three configurations on 200 fresh questions that neither the search nor the choice of baseline had ever seen.

Both numbers are below, because the second is only interpretable against the first.

configurationaccuracycorrect
(c)blend0.42021/50
(b)math-adaanchor, best single0.34017/50
math-12k0.32016/50
(a)base model0.30015/50
code-r160.30015/50
reasoning0.28014/50
math-pilot0.26013/50
medical (single only)0.26013/50
code-r16v30.1407/50
comparisondeltaexact McNemarbootstrap CIdiscordant
(c) − (a) blend vs base+12.0 ptsp = 0.146[−0.020, +0.240]12
(c) − (b) blend vs best single+8.0 ptsp = 0.388[−0.060, +0.220]12

On the control set of 50: base 0.760 (38/50), blend 0.740 (37/50), math-adaanchor 0.700 (35/50). The blend cost one question of general ability against the base and gained two against the single adapter. At n=50 that is a spread of one or two items and should not be read as a difference either way. It is reported because the treatment could have cost general ability and did not visibly do so.

Verdict: stop. The point estimate cleared the five-point rule against both baselines and the paired evidence did not support it. Under the rule that is the case that comes back to be decided, which is what happened.

The weights were the output of 27 gradient-free evaluations against the 20 held-out items, scoring 11/20 there (0.55). That took 4,936 seconds, 82 minutes of search to fit three numbers on twenty questions. The blend artifact itself took 75.7 ms to build and 73,859,072 bytes to write.

Three things were held fixed, and each is an assert in replicate.py rather than a note.

  • The weights, read out of results.json and never re-searched. Re-searching on the new set would turn an out-of-sample test back into an in-sample one, which is the whole thing this run exists to avoid.
  • The comparator, math-adaanchor, the best single adapter on the 50. If the 200 would have promoted a different adapter it is not switched to, because switching would give the baseline the hindsight the blend was denied.
  • The items, 200 new instances from the same ten generator families, written by a seed far from the scored and held-out seeds, and proven disjoint from all 70 previously used prompts before the run started.
configurationaccuracycorrectseconds
(c)blend0.38076/2003562
(a)base model0.29058/2001995
(b)math-adaanchor0.21543/2002910
comparisondeltaexact McNemarbootstrap CIdiscordant
(c) − (a) blend vs base+9.0 ptsp = 0.015[+0.020, +0.160]50 (34 blend-only, 16 base-only)
(c) − (b) blend vs single+16.5 ptsp < 0.001[+0.100, +0.230]49 (41 blend-only, 8 single-only)
The blend against each baseline, on 200 fresh questions
Against the base model+9.0 points, p = 0.015 +9.0
Against the best single adapter+16.5 points, p < 0.001 +16.5

Points of accuracy. The shaded band is the 95% bootstrap interval on the paired difference. Both bands sit clear of zero, which is the tripwire not firing.

Verdict: proceed. Both deltas clear the five-point rule and both paired intervals exclude zero, so the tripwire does not fire. This is the first run here where the point estimate and the paired evidence agree.

The artifact was rebuilt from the frozen weights and re-verified before a question was scored. Applying the written file and applying the in-memory matrices gave a maximum logit difference of 0.0 across 196 modules. Build time 117 ms, same 73.9 MB.

What moved between the runs

on the 50on the 200change
base model0.3000.290−1.0 pt
blend0.4200.380−4.0 pts
math-adaanchor0.3400.215−12.5 pts

The base model held. The blend slipped about four points, which is roughly what a fit on 20 items should be expected to lose out of sample. The single adapter fell 12.5 points.

That matters for reading the headline. (c) − (a) is +9.0 against a baseline that barely moved, so it is mostly the blend improving on base. (c) − (b) is +16.5, and most of that gap is the comparator degrading rather than the blend gaining. The comparison against the base model is the one that is not carried by a baseline falling over.

Both statements are true and both are in the table above. The second is the weaker claim.

The Ollama spike

Stage 2 ends with a node running ollama create on a fused adapter. If Ollama cannot fuse a Qwen2.5 LoRA then Stage 2 is empty, so this was tested first, before spending hours on the evaluation.

It works, on the GGUF path. A PEFT-directory ADAPTER line and a bare .safetensors one both fail, in four different ways. Ollama documents safetensors adapter support for Llama, Mistral and Gemma, and Qwen is not among them. Converting with llama.cpp’s convert_lora_to_gguf.py and pointing ADAPTER at the resulting GGUF succeeds.

Three independent confirmations that the adapter is genuinely fused rather than quietly ignored, which is the failure that would matter:

  1. ollama create prints success, where the safetensors path errored.
  2. The model manifest contains application/vnd.ollama.image.adapter at 79,816,896 bytes.
  3. The same prompt produces visibly different output. The base model writes an essay. The fused model writes the terse style its adapter was trained into, with a different final answer.

Conversion needs the base’s tokenizer.json, tokenizer_config.json, vocab.json and merges.txt (~18 MB) but not its 3 GB of weights, and the converter needs sentencepiece.

Stage 2 is not empty. Whether it is worth building is Stage 1’s question.

What this does and does not say

Does. On this base model, on this question distribution, a LoraHub blend of three public adapters, fitted only on 20 held-out items, beats the base model by 9 points and the best single adapter by 16.5 on 200 items none of them had seen, with paired evidence that agrees. A blend of 74 MB can be built in 117 ms on a laptop CPU, with no training run and no GPU.

Does not. Nothing here says the network gives better answers, and the wording stays “can build specialists on demand” rather than “better answers”. This is one base model, one size, one quantisation, one machine, and one question distribution. The result is that this blending procedure works on these questions.

Limitations

Stated plainly, because several of them cut against the result rather than for it.

  • The maths set is in the GSM8K template space. Fresh numbers do not remove template memorisation: a model tuned on ten thousand GSM8K unit-price problems has seen this shape of question whatever the figures. probe.py scores the base, the maths adapters and the blend on GSM8K’s held-out test split and on MMLU high-school mathematics. That probe has not been run. The confound cuts both ways. It inflates the baseline, which is conservative, but a blend inheriting a memorising adapter inherits the inflation too, which is not.
  • The 200 are new instances of the same ten families rather than a second domain. They test that the fitted weights generalise across instances of the families they were fitted near. They do not test that blending transfers across domains, and 200 items must not be read as though the sample size had bought that. The in-between-domain question this experiment opens with is answered only within the maths-word-problem space.
  • The comparator was frozen, so the +16.5 is against a baseline that was not allowed to improve. If the 200 would promote a different single adapter, this run does not know and does not say. That keeps the comparison honest in the direction that costs the blend, but it does mean “beats the best single adapter” means the adapter that was best on the 50.
  • The control set was not run on the 200. It is a maths-only replication, so unlike the 50-question run it says nothing about whether the blend costs general ability.
  • 20 held-out items is few. Weight search on 20 items can fit noise, which biases against the blend. This is the conservative direction and is stated rather than hidden.
  • Adapter provenance is unknown. Public adapters, unreviewed, with no statement of what they were trained on. One of them very likely saw GSM8K.
  • One base model, one size, one quantisation, one machine. 1.5B parameters is small. Nothing here generalises to 7B without being measured on 7B.

How the verdict is decided

(c) must beat both (a) and (b) by at least 5 points on the 50 maths questions. That gates. The rule is crude on purpose: 5 points on 50 questions is two and a half questions.

The paired statistics, exact McNemar on the discordant pairs and a bootstrap CI on the paired difference, are reported against every baseline and act as a tripwire rather than a veto. At n=50 the CI half-width is roughly ±14 points, so a genuine 6-point gain will essentially never reach significance, and making the paired test a hard gate would throw away real wins. But if the point estimate clears 5 points while the CI still spans zero, the gain and the evidence disagree, and that is the case that stops the run and comes back to be decided.

The evidence is results.json and replication.json, each carrying a per-question record for every configuration: the raw completion, what the extractor read, and whether that was right, so every number above traces back to one. Greedy decoding and a seeded search mean the scores reproduce exactly. The timings do not.