Post
81
๐ Two drops: the Ornith-1.5-27B-A3B pair (Coder + CoderX) and Qwen3.8-27B-Omnimerge-v6.
โ๏ธ Ornith = code-targeted expert prune of Ornith-1.5-35B-A3B: 256โ184 experts/layer, ~35.9Bโ26.7B, still A3B active. Top-8, MTP head and vision untouched. Nothing folded โ every surviving expert is bit-identical to the base, asserted at build. Same map for both; CoderX adds a REAP-style floor (--protect 6).
๐ Q6_K + imatrix, llama.cpp, vendor sampler, 11 benches โ base 256e / Coder / CoderX:
โก LiveCodeBench v6 (77 hard) โ 0.6623 / 0.7273 / 0.7662 โ +10.4pp over the teacher with 28% fewer experts
๐ง GPQA-Diamond โ 0.8283 / 0.7677 / 0.8131 โ the floor buys back most of what the pure map gives up
โ HumanEval 0.8963, AIME 0.9667, MATH-500 0.95 โ all CoderX
๐ Mean(11) โ 0.8252 / 0.8292 / 0.8386
๐ Start with CoderX. Coder still takes HumanEval+ (0.8293) and LCB-medium (0.5273), so it isn't dominated.
๐งช Omnimerge-v6: v4's sources, weights and method moved onto the Qwen3.8-27B base. Vs v4 on one binary and sampler, the only result outside the band is LiveCodeBench 0.883 vs 0.818 (68/77 vs 63/77). GPQA is 2pp lower and it thinks longer for it.
๐ก๏ธ Tool-calling (tool-eval-bench hardmode, 88 scenarios, 5 seeds, chart attached): v6 first of ten, 156.4 ยฑ3.5, +5.6 over its base. Read the safety column โ 3 safety-critical failures vs 9โ16 for every other model, bases included. All nine others fail TC-60 (cross-turn sleeper injection) 5/5; v6 never does.
โ ๏ธ Ties are ties: at n=77 one LCB problem is 1.3pp. And the Ornith rows there are IQ4_XS vs Q4_K_M for the Qwen rows โ part of that gap is quantisation, not architecture.
๐ ManniX-ITA/Ornith-1.5-27B-A3B-CoderX-MTP-GGUF
๐ https://ollama.com/mannix/ornith-1.5-27b-a3b-coderx
๐ ManniX-ITA/Ornith-1.5-27B-A3B-Coder-MTP-GGUF
๐ https://ollama.com/mannix/ornith-1.5-27b-a3b-coder
๐ ManniX-ITA/Qwen3.8-27B-Omnimerge-v6-MTP-GGUF
๐ https://ollama.com/mannix/omnimerge-v6
โ๏ธ Ornith = code-targeted expert prune of Ornith-1.5-35B-A3B: 256โ184 experts/layer, ~35.9Bโ26.7B, still A3B active. Top-8, MTP head and vision untouched. Nothing folded โ every surviving expert is bit-identical to the base, asserted at build. Same map for both; CoderX adds a REAP-style floor (--protect 6).
๐ Q6_K + imatrix, llama.cpp, vendor sampler, 11 benches โ base 256e / Coder / CoderX:
โก LiveCodeBench v6 (77 hard) โ 0.6623 / 0.7273 / 0.7662 โ +10.4pp over the teacher with 28% fewer experts
๐ง GPQA-Diamond โ 0.8283 / 0.7677 / 0.8131 โ the floor buys back most of what the pure map gives up
โ HumanEval 0.8963, AIME 0.9667, MATH-500 0.95 โ all CoderX
๐ Mean(11) โ 0.8252 / 0.8292 / 0.8386
๐ Start with CoderX. Coder still takes HumanEval+ (0.8293) and LCB-medium (0.5273), so it isn't dominated.
๐งช Omnimerge-v6: v4's sources, weights and method moved onto the Qwen3.8-27B base. Vs v4 on one binary and sampler, the only result outside the band is LiveCodeBench 0.883 vs 0.818 (68/77 vs 63/77). GPQA is 2pp lower and it thinks longer for it.
๐ก๏ธ Tool-calling (tool-eval-bench hardmode, 88 scenarios, 5 seeds, chart attached): v6 first of ten, 156.4 ยฑ3.5, +5.6 over its base. Read the safety column โ 3 safety-critical failures vs 9โ16 for every other model, bases included. All nine others fail TC-60 (cross-turn sleeper injection) 5/5; v6 never does.
โ ๏ธ Ties are ties: at n=77 one LCB problem is 1.3pp. And the Ornith rows there are IQ4_XS vs Q4_K_M for the Qwen rows โ part of that gap is quantisation, not architecture.
๐ ManniX-ITA/Ornith-1.5-27B-A3B-CoderX-MTP-GGUF
๐ https://ollama.com/mannix/ornith-1.5-27b-a3b-coderx
๐ ManniX-ITA/Ornith-1.5-27B-A3B-Coder-MTP-GGUF
๐ https://ollama.com/mannix/ornith-1.5-27b-a3b-coder
๐ ManniX-ITA/Qwen3.8-27B-Omnimerge-v6-MTP-GGUF
๐ https://ollama.com/mannix/omnimerge-v6