Simulated retained-capability spot-checks for 864 model-merging configurations across 6 methods (linear averaging, Task Arithmetic, TIES, DARE-TIES, DELLA, Model Stock), 3 expert-heterogeneity levels, expert count, and model scale, grounded in the foundational 2022-2023 model-merging literature through 2025-2026 large-scale empirical studies. Forward-looking task predicts full-suite interference from a quick spot-check on a handful of tasks.
Simulates whether merging fine-tuned experts preserves or destroys their individual capabilities. 864 configs across 6 methods including a documented non-monotonic failure mode (DARE-TIES is too aggressive on well-aligned models specifically). Ten sources (2022-2026) in merging_theory.py, including Model Soups, Task Arithmetic, TIES-Merging, DARE, DELLA, and a large "in-the-wild" study (4 base LLMs × 12 fine-tunes × 16 tasks).
Lets A Team Merging Several Fine-tuned Experts Estimate, From A Quick Spot-check, Whether A Full Evaluation Suite Will Reveal Interference.
Attribution 4.0 International (CC BY- 4.0)
To preview this file, you need to be a registered user. Please complete the registration process to gain access and continue viewing the content.
© 2026 - Copyright AIKosh. All rights reserved.