Specialists can be blended on demand.
Can a specialist be built without training one? A blend of three public LoRA adapters beat the base model by 9 points on 200 questions none of them had seen. The first run of the same experiment stopped.
Can a specialist be built without training one?
That is the question this experiment asks, narrowly enough to be answered either way:
Does a LoraHub-style blend of two or more public LoRA adapters answer better, on a domain genuinely in between the domains they were trained on, than the base model, or than the single best adapter?
Everything here exists to answer that, and the answer is allowed to be no. This is Stage 1 of two. Stage 2, wiring on-demand adapter fusion into the gateway, is gated on this clearing a five-point threshold and is not built unless it does.
What this does not claim. Nothing here says the network gives better answers. It says whether this blending procedure does, on this base model, on these questions.
What is being compared
| configuration | how it is chosen | |
|---|---|---|
| (a) | base model | Qwen/Qwen2.5-1.5B-Instruct, untouched |
| (b) | best single adapter | whichever of the seven scores highest on the 50 maths questions |
| (c) | the blend | LoraHub weights, searched on 20 held-out questions only |
(b) is deliberately chosen with hindsight, on the same questions the blend is then judged on. It is the strongest baseline available, because it already knew the answers. A weaker one would only have flattered the blend.
Seven public adapters, six of them blendable
Seven public LoRA adapters for Qwen/Qwen2.5-1.5B-Instruct, all pinned to a commit
in adapters.json so a re-run months from now loads the same bytes.
| domain | adapters | blendable |
|---|---|---|
| maths | 3 | yes |
| code | 2 | yes |
| reasoning | 1 | yes |
| medical | 1 | no |
Blendable is decided by one thing: whether the adapter can be averaged in factor
space. LoraHub sums the A and B matrices across adapters and takes the product
of the sums, so every adapter in a blend must agree on r, lora_alpha and
target_modules. Six agree exactly (r=16, alpha=32, the same seven
projections, ~70.5 MB each). The medical adapter is r=64, alpha=16 with four
targets, so it is scored as a single adapter and kept out of blends.
harness.load_registry asserts that precondition on every run rather than
trusting the registry file.
The obvious in-between domain for this base would be physics. No physics or
science adapter exists for Qwen2.5-1.5B at any download count. So the in-between
domain tested is maths word problems, on a base blended from maths and code
adapters. Every maths item needs multi-step arithmetic, which is what the code
adapters are nominally good at.
The question sets
Three files. Both scored sets are generated, and both sets’ answers are computed rather than asserted, so a question and its answer cannot silently disagree.
- Maths (50). Ten families: unit price and change, discount then tax,
rate/time/distance, averages, ratio sharing, compound growth, work rate,
percentage increase then decrease, mixture concentration, arithmetic series.
Numbers come from a seeded RNG, and every answer is computed in
fractions.Fractionand quantised half-up, so it is exact. - Control (50). Eight families rendered from encoded data tables: element symbols, atomic numbers, capitals, currencies, continents, SI units, planetary order, chemical formulae.
- Held out (20). Eighteen maths and two control, used only to choose the blend weights and never scored.
check_questions.py runs 1,037 assertions over all three before anything is scored
on them: counts, no duplicate prompts, every maths answer re-derived from the
item’s own parameters, every parameter proven load-bearing by perturbing it and
requiring the answer to move, every parameter’s value shown to appear in the
question text, control answers checked against a fresh table lookup, held-out items
proven absent from both scored sets, and even family coverage.
A generated set can be self-consistent and still be wrong
The first generated maths set contained this item:
A shop sells markers at £5.20 each. Alex buys 6 of them and pays with a £10 note. How much change, in pounds, does Alex receive? Answer: −21.20
Every check above passed on it. The answer re-derived from the parameters, both parameters were load-bearing, both appeared in the question. It was an impossible question. The generator had chosen the note by mixing pounds and pence, so £10 was offered for £31.20 of goods.
Re-derivation cannot catch a wrong premise. It only proves the arithmetic agrees
with itself. It was caught by printing one item per family and reading them, which
is why that step is in the plan. check_questions.py now also asserts every maths
answer is a positive quantity, which is the one check here that is not about
internal consistency. That check was then verified to actually fire on the
historical bug, since a check that only ever passes on good data proves nothing.
Method: a forward-pass delta, and no PEFT
One base model loaded once. Adapters and blends are applied as a forward-pass delta
through a LoRALinear wrapper rather than merged into the weights.
Applying a LoRA is W_eff = W + (alpha/r) * B @ A, and a LoraHub blend is the same
expression with A and B replaced by the weighted sums. That is tensor
arithmetic, so blend.py does it in plain torch + safetensors. The reason is
dependency risk rather than purity: transformers is at 5.8.0 here, a major
release, and peft is the package most likely to break against it. With PEFT out
of the scoring path, transformers is only asked to load Qwen2.5 and generate,
which is the stable part.
The saving is real beyond dependencies. A blend is ~74 MB of matrices rather than a
~3 GB copy of the model. Swapping configuration is 196 pointer assignments.
Returning to the base model is enabled = False rather than a reload, and
configuration N+1 cannot be contaminated by configuration N, because the base
weights are never written to.
Note the size. A six-way blend is not smaller than a single adapter. It is the
same 74 MB. The blend has the same rank and shape as its inputs, because the sums
ΣλᵢAᵢ and ΣλᵢBᵢ are each still rank r, so it costs exactly what one adapter
costs no matter how many are blended in. The 74 MB is what the file on disk
actually weighs.
Everything else is held fixed across configurations: same chat template, same
instruction, same max_new_tokens, no per-configuration prompt tuning, greedy
decoding everywhere. Greedy matters because a sampled run would put a random
variable inside a comparison of a few points. One answer extractor, fixed in
advance and never revisited after seeing which configuration it favours: the last
number in the completion for maths, compared with a tolerance of half a unit in the
item’s own stated decimal place, and a word-boundary phrase match for the control
set.
The duplicated sum, and the check that caught it
The path is not merely asserted to be equivalent. blend.py writes the blend out
in PEFT’s own layout and then checks, before the scored run, that applying the
written file and applying the in-memory matrices give identical next-token
distributions.
That check earned its keep. The LoraHub sum ΣwᵢAᵢ was originally written out
twice, once in the wrapper that applies a blend and once in the writer that saves
one, and the two copies drifted. The wrapper omitted the weight on its first term,
so it computed A₀ + Σ_{i>0} wᵢAᵢ where it should have computed Σᵢ wᵢAᵢ. For a
six-way blend the first adapter entered the sum at full weight rather than at its
fitted weight, so the two were different models.
Nothing downstream could see it. The search ran to completion, every candidate
produced a plausible score, and the artifact was written correctly from the correct
formula. The only symptom was that the matrices on disk and the matrices being
scored disagreed, which is exactly what the logits comparison tests. It failed the
first run with a maximum logit difference of 3.36 against a required 1e-3. The run
stopped before scoring a single question, so no number in results.json was ever
produced by the wrong path.
The fix was to delete the copy rather than correct it. The sum now exists once, as
harness.weighted_factor_sums, and both callers use it. A model-free equivalence
check runs as a pre-flight, so the same class of fault surfaces in the first second
rather than after a fifty-minute weight search. It was verified to fail on the
original buggy formula, since a check that has only ever seen working code has not
been tested.
The run that stopped
Two runs, and they are the same experiment twice. The first fitted blend weights on 20 held-out items and scored everything on 50. It read +12.0 points over the base model while its own paired interval still spanned zero, which is the case the pre-registered rule says must stop rather than resolve itself. The second took those weights frozen, ran no search at all, and scored three configurations on 200 fresh questions that neither the search nor the choice of baseline had ever seen.
Both numbers are below, because the second is only interpretable against the first.
| configuration | accuracy | correct | |
|---|---|---|---|
| (c) | blend | 0.420 | 21/50 |
| (b) | math-adaanchor, best single | 0.340 | 17/50 |
math-12k | 0.320 | 16/50 | |
| (a) | base model | 0.300 | 15/50 |
code-r16 | 0.300 | 15/50 | |
reasoning | 0.280 | 14/50 | |
math-pilot | 0.260 | 13/50 | |
medical (single only) | 0.260 | 13/50 | |
code-r16v3 | 0.140 | 7/50 |
| comparison | delta | exact McNemar | bootstrap CI | discordant |
|---|---|---|---|---|
| (c) − (a) blend vs base | +12.0 pts | p = 0.146 | [−0.020, +0.240] | 12 |
| (c) − (b) blend vs best single | +8.0 pts | p = 0.388 | [−0.060, +0.220] | 12 |
On the control set of 50: base 0.760 (38/50), blend 0.740 (37/50), math-adaanchor
0.700 (35/50). The blend cost one question of general ability against the base and
gained two against the single adapter. At n=50 that is a spread of one or two items
and should not be read as a difference either way. It is reported because the
treatment could have cost general ability and did not visibly do so.
Verdict: stop. The point estimate cleared the five-point rule against both baselines and the paired evidence did not support it. Under the rule that is the case that comes back to be decided, which is what happened.
The weights were the output of 27 gradient-free evaluations against the 20 held-out items, scoring 11/20 there (0.55). That took 4,936 seconds, 82 minutes of search to fit three numbers on twenty questions. The blend artifact itself took 75.7 ms to build and 73,859,072 bytes to write.
The replication: 200 fresh questions, no search
Three things were held fixed, and each is an assert in replicate.py rather than a
note.
- The weights, read out of
results.jsonand never re-searched. Re-searching on the new set would turn an out-of-sample test back into an in-sample one, which is the whole thing this run exists to avoid. - The comparator,
math-adaanchor, the best single adapter on the 50. If the 200 would have promoted a different adapter it is not switched to, because switching would give the baseline the hindsight the blend was denied. - The items, 200 new instances from the same ten generator families, written by a seed far from the scored and held-out seeds, and proven disjoint from all 70 previously used prompts before the run started.
| configuration | accuracy | correct | seconds | |
|---|---|---|---|---|
| (c) | blend | 0.380 | 76/200 | 3562 |
| (a) | base model | 0.290 | 58/200 | 1995 |
| (b) | math-adaanchor | 0.215 | 43/200 | 2910 |
| comparison | delta | exact McNemar | bootstrap CI | discordant |
|---|---|---|---|---|
| (c) − (a) blend vs base | +9.0 pts | p = 0.015 | [+0.020, +0.160] | 50 (34 blend-only, 16 base-only) |
| (c) − (b) blend vs single | +16.5 pts | p < 0.001 | [+0.100, +0.230] | 49 (41 blend-only, 8 single-only) |
Points of accuracy. The shaded band is the 95% bootstrap interval on the paired difference. Both bands sit clear of zero, which is the tripwire not firing.
Verdict: proceed. Both deltas clear the five-point rule and both paired intervals exclude zero, so the tripwire does not fire. This is the first run here where the point estimate and the paired evidence agree.
The artifact was rebuilt from the frozen weights and re-verified before a question
was scored. Applying the written file and applying the in-memory matrices gave a
maximum logit difference of 0.0 across 196 modules. Build time 117 ms, same
73.9 MB.
What moved between the runs
| on the 50 | on the 200 | change | |
|---|---|---|---|
| base model | 0.300 | 0.290 | −1.0 pt |
| blend | 0.420 | 0.380 | −4.0 pts |
math-adaanchor | 0.340 | 0.215 | −12.5 pts |
The base model held. The blend slipped about four points, which is roughly what a fit on 20 items should be expected to lose out of sample. The single adapter fell 12.5 points.
That matters for reading the headline. (c) − (a) is +9.0 against a baseline that barely moved, so it is mostly the blend improving on base. (c) − (b) is +16.5, and most of that gap is the comparator degrading rather than the blend gaining. The comparison against the base model is the one that is not carried by a baseline falling over.
Both statements are true and both are in the table above. The second is the weaker claim.
The Ollama spike
Stage 2 ends with a node running ollama create on a fused adapter. If Ollama
cannot fuse a Qwen2.5 LoRA then Stage 2 is empty, so this was tested first, before
spending hours on the evaluation.
It works, on the GGUF path. A PEFT-directory ADAPTER line and a bare
.safetensors one both fail, in four different ways. Ollama documents safetensors
adapter support for Llama, Mistral and Gemma, and Qwen is not among them.
Converting with llama.cpp’s convert_lora_to_gguf.py and pointing ADAPTER at the
resulting GGUF succeeds.
Three independent confirmations that the adapter is genuinely fused rather than quietly ignored, which is the failure that would matter:
ollama createprintssuccess, where the safetensors path errored.- The model manifest contains
application/vnd.ollama.image.adapterat 79,816,896 bytes. - The same prompt produces visibly different output. The base model writes an essay. The fused model writes the terse style its adapter was trained into, with a different final answer.
Conversion needs the base’s tokenizer.json, tokenizer_config.json, vocab.json
and merges.txt (~18 MB) but not its 3 GB of weights, and the converter needs
sentencepiece.
Stage 2 is not empty. Whether it is worth building is Stage 1’s question.
What this does and does not say
Does. On this base model, on this question distribution, a LoraHub blend of three public adapters, fitted only on 20 held-out items, beats the base model by 9 points and the best single adapter by 16.5 on 200 items none of them had seen, with paired evidence that agrees. A blend of 74 MB can be built in 117 ms on a laptop CPU, with no training run and no GPU.
Does not. Nothing here says the network gives better answers, and the wording stays “can build specialists on demand” rather than “better answers”. This is one base model, one size, one quantisation, one machine, and one question distribution. The result is that this blending procedure works on these questions.
Limitations
Stated plainly, because several of them cut against the result rather than for it.
- The maths set is in the GSM8K template space. Fresh numbers do not remove
template memorisation: a model tuned on ten thousand GSM8K unit-price problems
has seen this shape of question whatever the figures.
probe.pyscores the base, the maths adapters and the blend on GSM8K’s held-out test split and on MMLU high-school mathematics. That probe has not been run. The confound cuts both ways. It inflates the baseline, which is conservative, but a blend inheriting a memorising adapter inherits the inflation too, which is not. - The 200 are new instances of the same ten families rather than a second domain. They test that the fitted weights generalise across instances of the families they were fitted near. They do not test that blending transfers across domains, and 200 items must not be read as though the sample size had bought that. The in-between-domain question this experiment opens with is answered only within the maths-word-problem space.
- The comparator was frozen, so the +16.5 is against a baseline that was not allowed to improve. If the 200 would promote a different single adapter, this run does not know and does not say. That keeps the comparison honest in the direction that costs the blend, but it does mean “beats the best single adapter” means the adapter that was best on the 50.
- The control set was not run on the 200. It is a maths-only replication, so unlike the 50-question run it says nothing about whether the blend costs general ability.
- 20 held-out items is few. Weight search on 20 items can fit noise, which biases against the blend. This is the conservative direction and is stated rather than hidden.
- Adapter provenance is unknown. Public adapters, unreviewed, with no statement of what they were trained on. One of them very likely saw GSM8K.
- One base model, one size, one quantisation, one machine. 1.5B parameters is small. Nothing here generalises to 7B without being measured on 7B.
How the verdict is decided
(c) must beat both (a) and (b) by at least 5 points on the 50 maths questions. That gates. The rule is crude on purpose: 5 points on 50 questions is two and a half questions.
The paired statistics, exact McNemar on the discordant pairs and a bootstrap CI on the paired difference, are reported against every baseline and act as a tripwire rather than a veto. At n=50 the CI half-width is roughly ±14 points, so a genuine 6-point gain will essentially never reach significance, and making the paired test a hard gate would throw away real wins. But if the point estimate clears 5 points while the CI still spans zero, the gain and the evidence disagree, and that is the case that stops the run and comes back to be decided.
The evidence is results.json and replication.json, each carrying a per-question
record for every configuration: the raw completion, what the extractor read, and
whether that was right, so every number above traces back to one. Greedy decoding
and a seeded search mean the scores reproduce exactly. The timings do not.