Thinking Machines has released Inkling Small, an open-weights reasoning model designed to prioritise efficiency over sheer scale. The model scores 40 on the Intelligence Index, placing it just one point behind its larger sibling, Inkling, while using less than a third of the parameters. Despite the smaller footprint, Inkling Small outperforms the bigger model on specific coding and reasoning benchmarks, including Humanity’s Last Exam and GPQA Diamond. It also demonstrates superior token efficiency, averaging 24K output tokens per task compared to significantly higher counts from competing models like Deepseek V4 Flash and GPT-5.4 mini.
The release signals a strategic shift away from the prevailing trend of increasing model size to achieve better performance. By offering a model that is highly capable yet resource-light, Thinking Machines provides a practical alternative for developers concerned with operational costs and latency. The Apache 2.0 licence allows broad usage, and the inclusion of browser-based fine-tuning tools lowers the barrier for customisation.
* Weights are available on Hugging Face
* Supports text, image, and speech inputs
* Context window reaches 256K tokens




