Inference Providers
Active filters: GRPO
ykarout/Phi4-ThinkMode-fp16
Text Generation
• 15B • Updated • 6
mradermacher/Phi4-ThinkMode-fp16-GGUF
15B • Updated • 8
Text Generation
• 0.1B • Updated • 5
mradermacher/Nuke_X_Gemma3_1B_Reasoner_Testing-GGUF
1.0B • Updated • 39
• 1
mradermacher/Nuke_X_Gemma3_1B_Reasoner_Testing-i1-GGUF
1.0B • Updated • 141
• 1
Text Generation
• 0.1B • Updated • 8
alonsosilva/SmolGRPO-135M
Text Generation
• 0.1B • Updated • 11
VaidikML0508/Shark-Tank-Offer-Evaluator-llama3.2-3B-Instruct-GRPO-16bits-V1
Text Generation
• 3B • Updated • 8
• 1
mradermacher/Shark-Tank-Offer-Evaluator-llama3.2-3B-Instruct-GRPO-16bits-V1-GGUF
3B • Updated • 53
alfredcs/gemma-3-12b-grpo-firstaid
Updated
Text Generation
• 0.1B • Updated • 7
Thabet/SmolGRPO-135M-learning
Text Generation
• 0.1B • Updated • 8
Text Generation
• 0.1B • Updated • 10
Text Generation
• 0.1B • Updated • 8
yigitkucuk/tint-interact-sft-grpo
Text Generation
• 0.4B • Updated • 4
koochikoo25/SmolGRPO-135M
Text Generation
• 0.1B • Updated • 9
Text Generation
• 0.1B • Updated • 8
TianheWu/VisualQuality-R1-7B
Reinforcement Learning
• 8B • Updated • 2.76k
• 14
pedrocurvo/llama2-grpo-lora
Text Generation
• 7B • Updated • 4
mradermacher/VisualQuality-R1-7B-GGUF
8B • Updated • 73
Text Generation
• 0.1B • Updated • 6
• 1
Ceenen2302/Llama-3.2-1B-Instruct-GRPO-SmartLed
Feature Extraction
• 1B • Updated • 14
alfredcs/torchrun-gemma-3-12b-grpo-icd10pcs-merged
Text Generation
• 8B • Updated • 6
Ceenen2302/Llama-3.2-1B-Instruct-GRPO
Text Generation
• 1B • Updated • 8
alperenyildiz/Mistral-7B-Instruct-v0.3_q8_0_GRPO
Text Generation
• 7B • Updated • 6
hibikigf88/SmolLM-135M-Instruct-smoltldr-GRPO
Text Generation
• 0.1B • Updated • 8
Text Generation
• 0.1B • Updated • 5
Sarahpa/spGRPO-135M-readability
Text Generation
• 0.1B • Updated • 7
alfredcs/gemma-3-27b-grpo-med-merged
Image-Text-to-Text
• Updated • 10
alfredcs/gemma-3-27b-firstaid-icd10-merged
Image-Text-to-Text
• Updated • 6