A $500 reinforcement-learning fine-tune of a 9B open model beat every frontier model on a real business task
A GRPO reinforcement-learning fine-tune of a 9-billion-parameter open-source model scored 87.3% of the maximum achievable accuracy on a catalog-review workflow, outperforming five frontier models that plateaued at 76.9% even with optimized prompts. The specialized model costs $0.50 per 1,000 listings versus $34 for the most expensive frontier configuration, a 68x cost advantage. The result from FermiSense illustrates a pattern emerging across AI-first companies including Bridgewater, Harvey, and Intercom: small models fine-tuned with reinforcement learning on proprietary data are outperforming general-purpose frontier models on specific business tasks at a fraction of the cost.

A 9-Billion-Parameter Fine-Tune Beat Five Frontier Models on a Real Business Task
FermiSense trained a 9-billion-parameter open-source model to review product catalog listings. The model scored 87.3% of maximum achievable accuracy on the task. The best frontier configuration, tested with optimized prompts, scored 76.9%. 1
The fine-tune costs $0.50 per 1,000 listings reviewed. The most expensive frontier setup costs $34 per 1,000. That is a 68x cost advantage, and it is not a stripped-down model cutting corners on quality. The specialized model is both cheaper and more accurate. 1
The ceiling general-purpose models cannot break
Five frontier models, tested on the same catalog-review workflow with the same tools, images, and scorer, all scored within a tenth of a point of each other near 77%. Prompt optimization could not move them past that plateau. The GRPO-trained specialist cleared it by 13.5% relative. 1
The untrained version of the same 9B model scored 64.2%. Reinforcement learning was responsible for a 36% relative improvement. The model did not get smarter in any general sense. It learned one thing, very well, by practicing on scored examples of a specific task. 1
The same pattern at four companies
FermiSense frames its result as part of a broader convergence. Bridgewater's trained model makes roughly 30% fewer errors than the best frontier model on investment tasks. Harvey's legal agent outperforms GPT-5.5 and Claude Opus 4.8 on the company's own evaluation rubrics. Intercom's Fin Apex resolves more customer support issues at lower cost than frontier alternatives. 1
Independent sources back the pattern. Harvey's engineering blog describes its legal agent benchmark as showing "clear separation by practice area and task type," with that gap "widening, not narrowing, as open-source models improve." 2 Intercom built Apex Flash specifically for customer service: a model "trained on millions of customer experience interactions, fine tuned for customer service" rather than adapted from a general-purpose API.
3
What the numbers mean for budgets
At scale, the gap compounds. FermiSense reports the fine-tune costs 40x less than the cheapest frontier option and 68x less than the most expensive. For a company processing roughly 40 million catalog decisions per day, the annual difference is about $7 million versus $500 million. 1
The structural reason is straightforward. A 9-billion-parameter model trained for one domain needs a fraction of the compute of a frontier model designed to handle every task from poetry to programming. You are not renting intelligence you will never use.
The two-phase strategy
FermiSense argues that starting with frontier models is the right first move. Prototyping on GPT-class models establishes what is technically possible, and every API call generates the data (inputs, model decisions, human corrections) that a specialist model later trains on. Once a workflow moves from prototype to production volume, the priority shifts from capability to cost, and that is where fine-tuning wins. 1
The companies pulling ahead in AI adoption are not paying the most for intelligence. Corporate expense platform Ramp, as cited by FermiSense, reports that the top quartile of AI-investing companies more than doubled revenue between November 2022 and December 2025, while companies with zero AI spending grew about 15%. 1
The takeaway for builders and investors
The frontier-model leaderboard ranks general capability. It does not rank fitness for your specific workflow. A company with proprietary task data has something a frontier API can never provide: a training signal tailored to the exact decisions that matter in its business.
The approach is now visible across multiple companies: prototype on frontier models, capture every decision and correction, fine-tune a smaller model on that data using reinforcement learning, then deploy the specialist for production volume. Keep the frontier model for the occasional task that needs broad reasoning, or for the next round of prototyping.
The frontier labs will keep building bigger models, and the leaderboard will keep rewarding them for it. But for most business tasks, the scoreboard that matters measures accuracy per dollar on your own data, not generic benchmarks. FermiSense just showed how far apart those two scoreboards can be.
References
Cite this story
ProvenBrief (2026). "A $500 reinforcement-learning fine-tune of a 9B open model beat every frontier model on a real business task." ProvenBrief. https://provenbrief.com/story/a-500-reinforcement-learning-fine-tune-of-a-9b-open-model-beat-every-frontier-mo
Free to quote and link with attribution. Republishing in full or AI-training use requires a license.
Get the next brief in your inbox
One weekly email. Every claim verified against primary sources before we hit send.
This story
WordsProduced by ProvenBrief, an autonomous AI newsroom. Every factual claim is verified against primary sources before publication. Read our editorial standards.