Inference Providers
Active filters: grpo
Text Generation
• 3B • Updated • 179k
• 79
oddadmix/Nawah-50M-RAG-Support-2K
Text Generation
• 51.8M • Updated • 627
• 3
Text Generation
• 4B • Updated • 18
• 3
Image-Text-to-Text
• 9B • Updated • 85
• 9
Text Generation
• 8B • Updated • 35
• 2
snap-stanford/humanlm-opinion
Text Generation
• 8B • Updated • 566
• • 12
durgesh-rao/LLaMa-8B-it-GRPO-optimized-weights
Crownelius/Poe-8B-GLM5-Opus4.6-Sonnet4.5-Kimi-Grok-Gemini-3-pro-preview-HERETIC
9B • Updated • 150
• 6
creovateHQ/Qwen2.5-3B-Instruct_BrowserForge_Adapter
Text Generation
• Updated • 4
• 3
Image-Text-to-Text
• 9B • Updated • 17
• 10
Reinforcement Learning
• Updated • 1
gdfhhjk/spectrida-re-gguf
Text Generation
• 8B • Updated • 154
• 4
yuanqianhao/MemSearcher-7B
Text Generation
• 8B • Updated • 54
• 2
vishwr/claim_drafter-merged
Text Generation
• Updated • 1
JiangHoucheng/RePolicy-4B
Text Generation
• 4B • Updated • 437
• 1
riit3sh/qwen2.5-coder-3b-instruct-spider-sft-grpo
Text Generation
• 3B • Updated • 787
• 1
multimodalart/paint-rl-qwen-lora-v2
Text Generation
• 0.1B • Updated • 33
8B • Updated • 13
sergiopaniego/Qwen2-0.5B-GRPO-test
Updated
Novaciano/ESP-NSFW-GRPO-1B-Sin_Censura-GGUF
1B • Updated • 98
• 6
nbd22/Llama-3.1-8B-Instruct-GRPO-gsm8k-ft-lora
Updated
sergiopaniego/Qwen2-0.5B-GRPO
Updated
philschmid/qwen-2.5-3b-r1-countdown
Text Generation
• 3B • Updated • 15
• 8
spinech/qwen-2.5-3b-r1-countdown
Text Generation
• 3B • Updated • 13
Dongwei/Qwen2.5-1.5B-Open-R1-GRPO
Text Generation
• 2B • Updated • 7
• 1
spinech/qwen2.5-3b-r1-rearc-stage1
Text Generation
• 3B • Updated • 8
Dongwei/DeepSeek-R1-Distill-Qwen-7B-GRPO
Text Generation
• 8B • Updated • 12
• 1