Research
This empirical comparison of arithmetic reasoning in Qwen3.8-27B versus GPT-4o demonstrates that smaller open models struggle with multi-digit arithmetic in word-form output without reasoning enabled, but reasoning-enabled inference dramatically improves accuracy—a key consideration for developers choosing between local vs. API-based models for reasoning-heavy workloads.
Read the full article at Simon Willison
CoFabrix summarises and comments on this story. The original reporting belongs to Simon Willison.
Find out where your organization stands -- and what to do about it.