Thinking Machines Lab has released Inkling-Small, an open weights Mixture-of-Experts model with 276 billion total parameters and 12 billion active. The release sits at roughly a quarter the scale of the main Inkling model, which carries 975 billion total and 41 billion active parameters. Training occurred on NVIDIA GB300 NVL72 systems. The system reasons over text, images and audio. The context window reaches 1 million tokens, and thinking effort is adjustable. Weights ship under Apache 2.0 on Hugging Face.
In this article
Is it deployable
Yes. The quantized checkpoint makes the difference. The BF16 checkpoint requires at least 600 GB of aggregated VRAM. That requirement is met by four NVIDIA B300 units or eight NVIDIA H200 units. The NVFP4 checkpoint drops that floor to 180 GB. It runs W4A4 on a single B300, which requires SM100+, or W4A16 on two H200s. Supported runtimes are SGLang, vLLM, TokenSpeed, Unsloth and Hugging Face.
That single-GPU path moves a 276 billion parameter model out of frontier-lab territory. Startups can self-host on one rented B300 instance. Mid-size enterprises with existing H200 capacity can serve it without new hardware. Regulated sectors gain a private-weights option: financial services, healthcare operations, insurance, telecom and public sector. Applicable workloads include coding agents, terminal automation, and document and chart understanding. Audio widens that to call-center analytics, voice interfaces and meeting summarization.
Architecture
Inkling-Small is a 42-layer decoder-only transformer with a sparse MoE feed-forward backbone. Each token routes to six of 256 experts, plus two shared experts active on every token. Attention is a hybrid of local and global layers. The model is encoder-free and natively multimodal. Images are divided into 40 by 40 pixel patches and transformed using a four-layer hMLP. Audio is represented as dMel spectrograms. Both pass through a lightweight embedding layer and are processed jointly with text tokens. Numerics support covers BF16, MXFP8 and NVFP4. Audio input is WAV at 16 kHz, ideally under two minutes. Output is text only.
Inkling-Small began training after its larger counterpart. That let the research team revise the pre-training data mix and the machine learning recipe. The research team post-trained an earlier checkpoint, Inkling-Small preview, in part using on-policy distillation with Inkling as the teacher. From that checkpoint, it continued scaling agentic coding RL for two weeks.
Benchmark results
The smaller model surpasses its teacher on reasoning and agentic coding. On Humanity’s Last Exam text only Inkling-Small scores 31.6%, ahead of Inkling’s 29.7%. SWE-bench Verified is 80.2% versus 77.6%, using a bash-only harness. Terminal-Bench 2.1 reaches 64.7% at best harness. Toolathlon Verified is 54.4%, against Inkling’s 45.5%. GPQA Diamond is 89.5%, AIME 2026 is 95.5%, and IFBench is 82.2%. ARC-AGI-2 rises to 40.1% from Inkling’s 36.5%.
SimpleQA Verified falls to 20.6% from Inkling’s 43.9%, and the AA Omniscience index drops to -9.0 from 2.1. Tau 3 Banking is 15.5% versus Inkling’s 23.7%. All evaluations ran at effort 0.99 and temperature 1.0, with a 256K max-token trajectory limit on coding evals. External scores are sourced from Artificial Analysis, Scale AI and ARC Prize.
Multimodality, epistemics and safety
Multimodal scores stay close to Inkling at lower cost. MMMU Pro is 74.0%. CharXiv RQ is 77.4%, rising to 81.3% when the model uses Python to crop, zoom and inspect charts programmatically. Audio MC is 54.9%, MMAU is 77.0%, and VoiceBench is 90.1%.
On epistemics, calibration was trained with RL against proper scoring rules on a large corpus of real-world forecasting questions. ForecastBench without search gives a Brier Index of 61.3 ± 0.46, ahead of Inkling’s 60.1 ± 0.54. On safety, StrongREJECT is 98.4%, FORTRESS adversarial is 71.6%, and FORTRESS benign is 96.9%. Thinking Machines Lab concluded the model presents no material uplift beyond the existing open-weight ecosystem. It recommends layering downstream moderation such as Llama Guard on consumer-facing deployments.
Both models are available on Tinker with a limited-time discount. Text, image and audio chat run on Tinker Playground.
Interactive explainer
The embed below breaks the release into four interactive parts. Tab one animates how sparse routing activates eight of 258 experts per layer. Tab two ranks Inkling-Small against comparable open weights models on ten benchmarks. Tab three traces how the three input modalities converge into one decoder. Tab four sizes the hardware each checkpoint format requires.
What it means
Users gain a model that fits on a single high-end GPU, removing the need for massive clusters to run a 276 billion parameter system. The drop in VRAM requirements means startups can host the model themselves without relying on cloud providers. However, the trade-off is a loss in factual recall, with SimpleQA Verified scores falling significantly compared to the larger version. For tasks requiring high accuracy on general knowledge questions, the larger model remains preferable.




