I had Gemini train its own replacement for $9
The author scrapes Reddit threads about high-end chef's knives to extract every brand, model, and steel mentioned and see what is being bought and argued about. Picking product names out of text is named-entity recognition (NER), a task small models have done for a decade. The author was using Gemini 3.1 Pro as a per-comment paid API, which worked but was overkill: from “picked up a Mazaki in white #2, way better than my old Fibrox” it returned Mazaki as a brand, Fibrox as a model, and white #2 as a steel, and nothing else. Because the scraper pulled every new comment, the bill grew with how much people posted, and the only way to cap it was to skip comments.
The obvious replacement, the open NER model GLiNER run zero-shot, cut the cost to nothing but cut accuracy to about 0.65 F1 against Gemini's answers. The project asked whether Gemini could label 4,290 comments once and teach GLiNER to close that gap. The stated goal was to fine-tune GLiNER large v2.5 (459M) to tag brands, models, and materials in Reddit comments using labels Gemini 3.1 Pro wrote once. The approach was to ask Gemini for strings rather than offsets, compute offsets in code, add comments with no products in them as negatives, and lock a 225-comment validation set before the second run. Five of ten runs produced no usable model: three failed on configuration, and two failed on a tensor called words_mask that the author filled the way an attention mask is filled.
The result was 0.83 F1 against Gemini's labels, with the winning run taking 24 minutes on a Tesla T4. The cost was $9 of labels, about $2.50 of GPU time across all ten runs, and days of debugging. The plan had three steps: have Gemini label a few thousand Reddit comments once, marking every brand, model, and steel; train GLiNER on those labels; then run GLiNER on the author's own machine for every comment after that and stop calling Gemini. Gemini labeled 4,290 comments for $9, or $0.0021 a comment, meaning the trained model pays for itself at roughly comment 4,291, as long as later comments are about the same length and it runs on a GPU the author already owns.
Success was tested on 225 comments the model had never seen: how often does it tag the same words Gemini tagged? One catch applies to every score: nobody checked Gemini's labels by hand, so the model is graded against Gemini, not the truth. Where Gemini was wrong, the model gets marked right for copying the mistake and wrong for fixing it. The labeling run went through OpenRouter at temperature 0 in 25 minutes, and the prompt decision that mattered most was never to ask the model for character offsets. It counts characters badly and returns spans off by two or three positions, so the prompt asks for the exact substring and a label instead.