Anserwise commited on
Commit
da92efe
·
verified ·
1 Parent(s): df863f4

docs: Add VIDRAFT Darwin platform breeding/evolution description

Browse files
Files changed (1) hide show
  1. README.md +112 -55
README.md CHANGED
@@ -18,6 +18,8 @@ tags:
18
  - vision-language
19
  - image-text-to-text
20
  - darwin-derived
 
 
21
  - agent
22
  base_model:
23
  - google/gemma-4-31B-it
@@ -55,72 +57,116 @@ model-index:
55
 
56
  # AWAXIS-KR-31B
57
 
58
- ## 📌 모델 설명 (Description)
59
 
60
- **AWAXIS-KR-31B**은 한국어 특화 MoE 베이스(JDONE-Research/AIOne-Agent-52B-A36B-it)에 Opus-distill 추론 시그널(Anserwise/AWAXIS-Think-31B)을 결합한 **Darwin V8 FFN-crossbreed 파생 모델**입니다. Gemma-4 MoE 아키텍처(8 전문가 top-2 라우팅, **52B 총 / 36B 활성** 파라미터, vision/audio 토큰 지원)를 기반으로 한국어 instruction following, 지식·문화 QA, 단계별 추론·수학 작업에 최적화되어 있으며, 한국어 4과목 종합 **80.0%** 성능을 검증했습니다.
61
 
62
- > AWAXIS-KR-31B is a Darwin-derived Korean-focused MoE model (Gemma-4 family, 52B total / 36B active, 8 experts top-2 routing) built via Darwin V8 FFN-crossbreed. Optimized for Korean instruction following, knowledge & cultural QA, and reasoning. Architecture supports image-text inputs via the Gemma4 multimodal base.
 
 
63
 
64
  ---
65
 
66
- ## 🧬 모델 족보 (Model Lineage)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
67
 
68
  ```
69
- AWAXIS-KR-31B (this model — Darwin-derived)
70
- ├── 어머니 Mother (kept full)
71
- │ └── JDONE-Research/AIOne-Agent-52B-A36B-it
72
- │ — 한국어 특화 Gemma4 MoE 52B / A36B
73
- │
74
- └── 아버지 Father (dense-FFN donor)
75
- └── Anserwise/AWAXIS-Think-31B
76
- ├── 조모 (kept full)
77
- │ └── TeichAI/gemma-4-31B-it-Claude-Opus-Distill-v2
78
- │ — Claude Opus 추론 distill 베이스
79
- │
80
- └── 조부 (FFN donor)
81
- └── google/gemma-4-31B-it
82
- — Gemma-4 베이스
83
  ```
84
 
85
- ### 직계 부모 (Direct parents)
86
 
87
- | 역할 | 모델 | 기여 |
88
- |------|------|------|
89
- | 어머니 Mother (kept) | [JDONE-Research/AIOne-Agent-52B-A36B-it](https://huggingface.co/JDONE-Research/AIOne-Agent-52B-A36B-it) | 한국어 능력, MoE 라우팅, 전문가, 어텐션, 임베딩 100% 보존 |
90
- | 아버지 Father (FFN donor) | [Anserwise/AWAXIS-Think-31B](https://huggingface.co/Anserwise/AWAXIS-Think-31B) | Opus-distill 추론 시그널을 dense FFN 경로로 주입 |
91
 
92
- ### 조부모 (Paternal grandparents)
93
 
94
- | 역할 | 모델 |
95
- |------|------|
96
- | 조모 grandmother | [TeichAI/gemma-4-31B-it-Claude-Opus-Distill-v2](https://huggingface.co/TeichAI/gemma-4-31B-it-Claude-Opus-Distill-v2) |
97
- | 조부 grandfather | [google/gemma-4-31B-it](https://huggingface.co/google/gemma-4-31B-it) |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
98
 
99
- **공통 시조 (Common ancestor)**: Google **Gemma-4** 아키텍처.
100
 
101
  ---
102
 
103
- ## 📚 활용 데이터셋 (Datasets Used)
104
 
105
  본 모델의 **한국어 능력 평가**에는 **K-AI Hub(NIA AI Hub) / K-AI Leaderboard(aihub.or.kr) 생태계**의 표준 한국어 LLM 벤치마크 데이터셋을 활용했습니다.
106
 
107
- | 데이터셋 | 분야 | 출처 |
108
- |---|---|---|
109
- | **KMMLU** | 한국어 지식 (45 과목) | [HAERAE-HUB/KMMLU](https://huggingface.co/datasets/HAERAE-HUB/KMMLU) |
110
- | **HAE_RAE_BENCH_1.1** | 한국어 이해·문화 (13 서브셋) | [HAERAE-HUB/HAE_RAE_BENCH_1.1](https://huggingface.co/datasets/HAERAE-HUB/HAE_RAE_BENCH_1.1) |
111
- | **HRM8K** | 한국어 수학·추론 (GSM8K 한국어판) | [HAERAE-HUB/HRM8K](https://huggingface.co/datasets/HAERAE-HUB/HRM8K) |
112
- | **CLIcK** | 한국어 문화-언어 | [EunsuKim/CLIcK](https://huggingface.co/datasets/EunsuKim/CLIcK) |
113
-
114
- 상기 데이터셋은 HAERAE-HUB와 EunsuKim 등 한국 연구 커뮤니티가 큐레이팅하여 K-AI 허브 평가 표준으로 채택된 공공 자산입니다.
115
 
116
  ---
117
 
118
- ## 🏗 아키텍처 (Architecture)
119
 
120
  | | |
121
  |---|---|
122
  | Class | `Gemma4ForConditionalGeneration` (multimodal: text + image + audio) |
123
- | Parameters | **52B total · 36B active** (MoE, 8 experts, top-2 routing) |
124
  | Layers | 60 |
125
  | Hidden / Intermediate | 5,376 / 21,504 |
126
  | Attention heads / head_dim | 32 / 256 |
@@ -129,29 +175,29 @@ AWAXIS-KR-31B (this model — Darwin-derived)
129
 
130
  ---
131
 
132
- ## 📊 측정 벤치마크 (Measured Benchmarks)
133
 
134
- | 벤치마크 | 설정 | 점수 |
135
  |-----------|---------|-------|
136
- | **한국어 4과목 종합** · n=80, seed=42 | greedy | **80.0%** |
137
- | ↳ KMMLU (지식) | 20Q, greedy | 70.0% |
138
- | ↳ HAERAE-Bench (이해) | 20Q, greedy | 75.0% |
139
- | ↳ HRM8K (수학) | 20Q, greedy | **90.0%** |
140
- | ↳ CLIcK (문화언어) | 20Q, greedy | 85.0% |
141
  | **CLIcK** (n=200) | greedy | **88.0%** |
142
 
143
  ---
144
 
145
- ## 🎯 사용 용도 (Intended Use)
146
 
147
- - 한국어 instruction following
148
- - 지식·문화 QA, 추론·수학
149
- - 일반 한국어 LLM 작업
150
- - 멀티모달 입력(image-text-to-text)은 Gemma-4 베이스 능력 상속
151
 
152
  ---
153
 
154
- ## 🚀 추론 예시 (Inference)
155
 
156
  ```python
157
  from transformers import AutoTokenizer, AutoModelForCausalLM
@@ -175,6 +221,17 @@ print(tok.decode(out[0][inp["input_ids"].shape[-1]:], skip_special_tokens=True))
175
 
176
  ---
177
 
178
- ## 📜 라이선스 (License)
 
 
 
 
 
 
 
 
 
 
 
179
 
180
- 본 모델은 **Gemma-4 계통** 가중치를 포함하며, [Gemma Terms of Use](https://ai.google.dev/gemma/terms)를 준수합니다.
 
18
  - vision-language
19
  - image-text-to-text
20
  - darwin-derived
21
+ - vidraft
22
+ - darwin-crossbreed
23
  - agent
24
  base_model:
25
  - google/gemma-4-31B-it
 
57
 
58
  # AWAXIS-KR-31B
59
 
60
+ ## Overview
61
 
62
+ **AWAXIS-KR-31B**은 **[VIDRAFT](https://huggingface.co/VIDraft) Darwin AI 모델 교배/진화 플랫폼**을 통해 생성된 한국어 특화 MoE 모델입니다. Darwin의 독자적인 **FFN-crossbreed 엔진(V8)**으로 한국어 특화 MoE 베이스(JDONE-Research/AIOne-Agent-52B-A36B-it)에 Opus-distill 추론 시그널(Anserwise/AWAXIS-Think-31B)을 교배 결합하였습니다.
63
 
64
+ Gemma-4 MoE 아키텍처(8 전문가 top-2 라우팅, **52B 총 / 36B 활성** 파라미터, vision/audio 토큰 지원) 기반으로, 한국어 instruction following, 지식/문화 QA, 단계별 추론/수학 작업에 최적화되어 있으며, 한국어 4과목 종합 **80.0%** 성능을 검증했습니다.
65
+
66
+ > AWAXIS-KR-31B is a Korean-focused MoE model (Gemma-4 family, 52B total / 36B active, 8 experts top-2 routing) created through the **[VIDRAFT](https://huggingface.co/VIDraft) Darwin AI Model Breeding/Evolution Platform**. Built via Darwin V8 FFN-crossbreed engine, combining a Korean-specialized MoE base with Opus-distill reasoning signals through automated biological-inspired crossbreeding.
67
 
68
  ---
69
 
70
+ ## VIDRAFT Darwin AI 모델 교배/진화 플랫폼
71
+
72
+ **[VIDRAFT Darwin](https://huggingface.co/VIDraft)**은 AI 모델의 **교배(Crossbreeding)와 진화(Evolution)**를 통해 새로운 고성능 모델을 자동 생성하는 플랫폼입니다. 생물학적 유전 원리에서 영감을 받아, 두 개 이상의 부모 모델에서 각각의 장점을 선택적으로 결합하여 자식 모델을 탄생시킵니다.
73
+
74
+ ### Darwin 교배/진화 핵심 기술
75
+
76
+ | 기술 | 설명 |
77
+ |------|------|
78
+ | **FFN Crossbreed Engine (V8)** | 부모 모델의 Feed-Forward Network(FFN) 레이어를 선택적으로 교차 결합하는 핵심 엔진. 어텐션/임베딩은 어머니(Mother)에서, FFN 시그널은 아버지(Father)에서 추출하여 블렌딩 |
79
+ | **Smart MRI (Model Resonance Imaging)** | 두 모델 간 레이어별 유사도/호환성을 분석하여 최적 교배 비율(alpha)을 자동 탐색하는 기술 |
80
+ | **Alpha Grid Search** | 교배 비율 alpha를 체계적으로 탐색하여 벤치마크 성능이 최대화되는 최적점을 발견 (자연선택 시뮬레이션) |
81
+ | **Multi-Generation Breeding** | 1세대 교배 결과물을 다시 부모로 삼아 2세대, 3세대 교배를 수행하는 다세대 진화 |
82
+
83
+ ### 이 모델의 Darwin 교배 과정
84
+
85
+ **AWAXIS-KR-31B은 2세대(F2) 교배 모델**입니다. 1세대에서 AWAXIS-Think-31B을 생성하고, 이를 다시 아버지로 삼아 한국어 MoE 어머니와 2세대 교배를 수행했습니다.
86
 
87
  ```
88
+ [1세대 교배] AWAXIS-Think-31B 생성
89
+ Mother: TeichAI/gemma-4-31B-it-Claude-Opus-Distill-v2
90
+ Father: google/gemma-4-31B-it
91
+ --> Darwin FFN-crossbreed (alpha=0.1) --> AWAXIS-Think-31B
92
+
93
+ [2세대 교배] AWAXIS-KR-31B 생성 (이 모델)
94
+ Mother: JDONE-Research/AIOne-Agent-52B-A36B-it (한국어 MoE)
95
+ Father: AWAXIS-Think-31B (1세대 교배 결과물)
96
+ --> Darwin FFN-crossbreed --> AWAXIS-KR-31B
 
 
 
 
 
97
  ```
98
 
99
+ 이처럼 Darwin 플랫폼은 **세대를 거듭할수록 능력이 누적 진화**하는 다세대 교배(Multi-Generation Breeding)를 지원합니다.
100
 
101
+ ### 왜 Darwin 교배인가?
 
 
 
102
 
103
+ 기존 모델 합성 방식(단순 가중치 평균, SLERP, TIES 등)과 달리, Darwin 교배는:
104
 
105
+ 1. **생물학적 유전 모방**: 어머니/아버지 역할을 명확히 분리하여 각 부모의 핵심 능력만 선택적으로 상속
106
+ 2. **FFN 선택적 주입**: 어텐션(문맥 이해)은 어머니에서 100% 보존하고, FFN(지식/추론 패턴)만 아버지에서 교차 -> 능력 충돌 최소화
107
+ 3. **벤치마크 기반 자연선택**: alpha grid search로 여러 자식 후보를 생성한 뒤, 실측 벤치마크로 최적 개체를 선택
108
+ 4. **다세대 진화**: 1세대 결과를 부모로 재활용하여 능력 누적 (이 모델 = 2세대)
109
+
110
+ ---
111
+
112
+ ## Model Lineage (모델 족보)
113
+
114
+ ```
115
+ AWAXIS-KR-31B (this model -- 2nd generation Darwin crossbreed)
116
+ |
117
+ +-- Mother (kept full, 100%)
118
+ | JDONE-Research/AIOne-Agent-52B-A36B-it
119
+ | -- Korean-specialized Gemma4 MoE 52B / A36B
120
+ |
121
+ +-- Father (FFN donor)
122
+ Anserwise/AWAXIS-Think-31B (1st generation Darwin crossbreed)
123
+ |
124
+ +-- Grandmother (kept full)
125
+ | TeichAI/gemma-4-31B-it-Claude-Opus-Distill-v2
126
+ | -- Claude Opus reasoning distill base
127
+ |
128
+ +-- Grandfather (FFN donor)
129
+ google/gemma-4-31B-it
130
+ -- Gemma-4 base
131
+ ```
132
+
133
+ ### Direct Parents
134
+
135
+ | Role | Model | Contribution |
136
+ |------|-------|-------------|
137
+ | Mother (kept) | [JDONE-Research/AIOne-Agent-52B-A36B-it](https://huggingface.co/JDONE-Research/AIOne-Agent-52B-A36B-it) | Korean capability, MoE routing, experts, attention, embeddings 100% preserved |
138
+ | Father (FFN donor) | [Anserwise/AWAXIS-Think-31B](https://huggingface.co/Anserwise/AWAXIS-Think-31B) | Opus-distill reasoning signal injected via dense FFN pathway |
139
+
140
+ ### Paternal Grandparents
141
+
142
+ | Role | Model |
143
+ |------|-------|
144
+ | Grandmother | [TeichAI/gemma-4-31B-it-Claude-Opus-Distill-v2](https://huggingface.co/TeichAI/gemma-4-31B-it-Claude-Opus-Distill-v2) |
145
+ | Grandfather | [google/gemma-4-31B-it](https://huggingface.co/google/gemma-4-31B-it) |
146
 
147
+ **Common ancestor**: Google **Gemma-4** architecture.
148
 
149
  ---
150
 
151
+ ## Datasets Used (활용 데이터셋)
152
 
153
  본 모델의 **한국어 능력 평가**에는 **K-AI Hub(NIA AI Hub) / K-AI Leaderboard(aihub.or.kr) 생태계**의 표준 한국어 LLM 벤치마크 데이터셋을 활용했습니다.
154
 
155
+ | Dataset | Domain | Source |
156
+ |---------|--------|--------|
157
+ | **KMMLU** | Korean knowledge (45 subjects) | [HAERAE-HUB/KMMLU](https://huggingface.co/datasets/HAERAE-HUB/KMMLU) |
158
+ | **HAE_RAE_BENCH_1.1** | Korean comprehension/culture (13 subsets) | [HAERAE-HUB/HAE_RAE_BENCH_1.1](https://huggingface.co/datasets/HAERAE-HUB/HAE_RAE_BENCH_1.1) |
159
+ | **HRM8K** | Korean math/reasoning (GSM8K Korean) | [HAERAE-HUB/HRM8K](https://huggingface.co/datasets/HAERAE-HUB/HRM8K) |
160
+ | **CLIcK** | Korean culture-language | [EunsuKim/CLIcK](https://huggingface.co/datasets/EunsuKim/CLIcK) |
 
 
161
 
162
  ---
163
 
164
+ ## Architecture
165
 
166
  | | |
167
  |---|---|
168
  | Class | `Gemma4ForConditionalGeneration` (multimodal: text + image + audio) |
169
+ | Parameters | **52B total / 36B active** (MoE, 8 experts, top-2 routing) |
170
  | Layers | 60 |
171
  | Hidden / Intermediate | 5,376 / 21,504 |
172
  | Attention heads / head_dim | 32 / 256 |
 
175
 
176
  ---
177
 
178
+ ## Measured Benchmarks
179
 
180
+ | Benchmark | Setting | Score |
181
  |-----------|---------|-------|
182
+ | **Korean 4-Subject Composite** (n=80, seed=42) | greedy | **80.0%** |
183
+ | -- KMMLU (knowledge) | 20Q, greedy | 70.0% |
184
+ | -- HAERAE-Bench (comprehension) | 20Q, greedy | 75.0% |
185
+ | -- HRM8K (math) | 20Q, greedy | **90.0%** |
186
+ | -- CLIcK (culture-language) | 20Q, greedy | 85.0% |
187
  | **CLIcK** (n=200) | greedy | **88.0%** |
188
 
189
  ---
190
 
191
+ ## Intended Use
192
 
193
+ - Korean instruction following
194
+ - Knowledge/culture QA, reasoning/math
195
+ - General Korean LLM tasks
196
+ - Multimodal input (image-text-to-text) inherited from Gemma-4 base capability
197
 
198
  ---
199
 
200
+ ## Inference
201
 
202
  ```python
203
  from transformers import AutoTokenizer, AutoModelForCausalLM
 
221
 
222
  ---
223
 
224
+ ## License
225
+
226
+ This model includes **Gemma-4 lineage** weights and complies with the [Gemma Terms of Use](https://ai.google.dev/gemma/terms).
227
+
228
+ ## Acknowledgements
229
+
230
+ - **[VIDRAFT](https://huggingface.co/VIDraft)** -- Darwin AI Model Breeding/Evolution Platform
231
+ - JDONE-Research for the Korean MoE base
232
+ - TeichAI for the Opus-Distill base
233
+ - Google DeepMind for Gemma-4
234
+
235
+ ---
236
 
237
+ *Built with the **VIDRAFT Darwin AI Model Breeding/Evolution Platform** -- FFN-crossbreed V8 engine. This is a 2nd-generation (F2) Darwin crossbreed model, created through automated biological-inspired crossbreeding that selectively combines the strengths of parent models. The Father (AWAXIS-Think-31B) was itself a 1st-generation Darwin crossbreed, demonstrating multi-generation evolution capability. Measured numbers above are exact; nothing inflated.*