Other Hugging Face Blog

Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps

GRPOIFStructLFM2.5-350Mstructured output

Structured output — reliably returning valid, parseable output in the requested format — is one of the most common real-world LLM tasks, yet most benchmarks only measure it inside broader reasoning or extraction scores. This guide provides a fully public, inexpensive recipe for making a small model substantially better at such schema compliance. It fine-tunes LFM2.5-350M using Group Relative Policy Optimization (GRPO) with the TRL library and evaluates on the IFStruct benchmark. The entire run consumes only around 500 samples and 100 training steps, small enough to fit on a free-tier Colab or Kaggle GPU, and the notebook is available on GitHub.

The authors note that the pipeline is not the one used to train the RL model described in the IFStruct blog; the goal is to show that task-specific fine-tuning of smaller models can let them match far larger models. The guide splits into two phases: fine-tuning on a GPU and evaluation on a local MacBook (here a MacBook Pro with Apple M5 Max and 36 GB of unified memory) via llama.cpp's OpenAI-compatible server. Prerequisites are uv and llama.cpp installed with Homebrew. As a baseline, they evaluate the base LFM2.5-350M on IFStruct to attempt reproducing its reported score of 21.1%. To do this they serve the BF16 GGUF using llama-server with options to offload all layers (`-ngl 99`), run four parallel requests (`-np 4`), set a 32,768-token context, and execute the 2,000-sample benchmark with `uv run ifstruct-eval`. The benchmark is open-source via Liquid4All/ifstruct, with the public dataset on Hugging Face at LiquidAI/ifstruct-v1.0.

The results show that even a light fine-tuning procedure improves IFStruct performance from 22.6% to 29.7%. This demonstrates a practical, low-cost path to improving schema compliance on small models — often the deciding factor for whether a model can be wired into downstream systems.

Read original →

← Back to home