TL;DR

Thinking Machines Lab released its first foundation model, Inkling, with downloadable weights under Apache 2.0 before introducing a closed API. The launch gives organizations more control over deployment, but the model requires costly hardware, omits its training data and trails rivals on several vendor-reported benchmarks.

Thinking Machines Lab, the 17-month-old company founded by former OpenAI technology chief Mira Murati, released the full weights for its first foundation model, Inkling, on July 15 under the Apache 2.0 license. The lab placed the weights on Hugging Face before offering a closed API and acknowledged that Inkling is not the strongest available model, positioning user control and deployment flexibility—not benchmark leadership—as the launch’s main distinction.

Inkling is a 975-billion-parameter Mixture-of-Experts model that activates 41 billion parameters for each token. According to Thinking Machines Lab, it has a 1-million-token context window and was pretrained on 45 trillion tokens spanning text, images, audio and video. Its inputs can include text, images and audio, while its output is text.

The release includes BF16 and NVFP4 checkpoints and day-one support for Transformers, vLLM, SGLang and llama.cpp, among other deployment tools. Apache 2.0 generally permits users to download, modify and commercialize the model. The lab did not publish the underlying training data or full training pipeline, meaning Inkling is an open-weight model rather than a fully open-source training project.

The company also introduced a thinking-effort control ranging from 0.2 to 0.99, allowing operators to trade reasoning work for speed and lower inference cost. The source material reports that Inkling matched Nvidia’s Nemotron 3 Ultra on Terminal-Bench 2.1 while using about one-third of the tokens, but that comparison and the wider benchmark results are vendor-published and await independent replication. Some scores were produced with a prerelease checkpoint.

At a glance
announcementWhen: announced July 15, 2026
The developmentThinking Machines Lab released Inkling’s full model weights on July 15, 2026, making ownership and modification available from launch rather than reserving access for an API.
Top Steam deals right now
Warhammer 40,000: Space Marine 2-70%$17.99
Grand Theft Auto V Enhanced-50%$14.99
Gamble With Your Friends-38%$4.95
Palworld-30%$20.99
PowerWash Simulator 2-25%$18.74
DELTARUNE-20%$19.99
Pathogenic-20%$7.99
Funnel Runners-10%$13.49
Live · Steam store (current discounts)
AI Dispatch · Reality Check · 16 July 2026

The weights came first: what Inkling actually signals

Mira Murati’s lab shipped its first foundation model — and the model isn’t the story. The order of operations is: full weights, Apache 2.0, day one, before any closed API. Plus a rare concession — the lab says it’s not the strongest model available, open or closed.

975B / 41B
total / active · MoE
1M
context window
45T
pretrain tokens
T · I · A
text · image · audio in
Apache 2.0
the licence*
Licence over leaderboard — what’s actually open
Model weightsBF16 + NVFP4 checkpoints on Hugging Face — download, modify, commercialize, keep
Apache 2.0 licenceconfirmed on the model card & HF repo — the real thing, not a source-available lookalike
Day-0 toolingtransformers · vLLM · SGLang · llama.cpp · TokenSpeed · Unsloth
Training data / pipelinenot published — open weights ≠ open source. Industry norm, but say it plainly
Separate use policy?reported: a Model Acceptable Use Policy over parameters & modified versions, barring surveillance, deception & fully automated decisions affecting rights
Unverified — check the model card yourself. If it reads as reported, Apache 2.0 isn’t the whole legal picture, and for ISR / geospatial / public-safety builders that clause is a go/no-go, not a footnote.
▲ Where it’s strong
  • AIME 2026 97.1%
  • GPQA Diamond 87.2%
  • MCP Atlas (Nemotron 44.7%) 74.1%
  • VoiceBench · open-weight audio frontier 91.4%
  • FORTRESS adversarial · best open 78.0%
  • ForecastBench · calibration 61.1
▼ Where it’s behind
  • HLE text-only (GLM-5.2 40.1%) 29.7%
  • SWE-bench Pro (GLM-5.2 62.1%) 54.3%
  • Terminal-Bench 2.1 (GLM-5.2 82.7%) 63.8%
  • SWE-bench Verified (Fable 5 95.0%) 77.6%
  • Design Arena · 2nd open, behind GLM-5.2 ~10th
◆ The dial nobody’s talking about — controllable thinking effort

A 0.2 → 0.99 effort setting trades reasoning tokens against cost & latency, so you get a curve, not a point. On Terminal-Bench 2.1 it reportedly matches Nemotron 3 Ultra at ~⅓ the tokens. Peak score is a vanity metric when you serve millions of calls; the cost curve is what ships. (Bonus: its chain of thought compressed on its own during RL — nobody rewarded it; efficiency did.)

0.2 · fast & cheap 0.99 · max effort
⚑ The China question — & the irony

Pitched as the Western alternative to Chinese open weights (censorship-resistance training is the differentiator). But GLM-5.2 still wins on agentic/reasoning and Kimi K2.6 often on multimodal: best American open model, second in the open field. The irony — post-training was bootstrapped on synthetic data from Kimi K2.5.

⚠ Open weights you probably can’t run

BF16 needs ≥2 TB aggregate VRAM (8× B300 / 16× H200). NVFP4 still needs ≥600 GB. Not a workstation model — a 512 GB fleet falls just short. “Open” ≠ “runnable.” Mitigations: 1-bit GGUFs (~74% acc.), hosted eval routes, and Inkling-Small (12B active) — the release local-first builders actually want.

The take

Open weights used to be a consolation prize. Inkling is a strategic open release — Apache 2.0, natively multimodal, honestly marketed, published complete on day one, optimized for deployment rather than headlines (the model isn’t the product; the fine-tuning platform is). It doesn’t need to win every benchmark for that to matter. The frontier is learning that owning the base beats renting the API — arriving now from the inside. For the sovereignty buyer: ① a real Western hedge against being switched off · ② verify the use policy before you build · ③ check the VRAM, then benchmark vs GLM-5.2 & Kimi K2.6 on your task.

Sources: Thinking Machines Lab (announcement, model card, HF repo, 15 Jul 2026); Hugging Face; VentureBeat, TechCrunch, BenchLM, LinkLoot, XenoSpectrum, NewsCord; Nathan Lambert via X. Benchmarks are vendor-published (some via Artificial Analysis) & await independent replication; some reflect a pre-release checkpoint. The AUP is reported, not verified here.
thorstenmeyerai.com

Model Ownership Takes Priority

Releasing the weights first gives companies, researchers and public institutions the ability to host Inkling on their own infrastructure, modify it and retain customized versions. That can reduce dependence on a vendor-controlled API and limit exposure to price changes, service withdrawals or shifting access rules.

The strategy also marks a different approach from launches that treat downloadable weights as a later or restricted offering. Inkling does not lead every reported test, but its deployment tooling, adjustable reasoning cost and multimodal design may matter more to buyers whose workloads reward predictable operating costs. The release also creates a US-developed open-weight alternative to models including GLM-5.2 and Kimi K2.6, though those Chinese models remain ahead on several reported reasoning, agentic and multimodal measures.

Samsung 55-Inch Class Crystal UHD U8000H Series Samsung Vision AI Smart TV (2026 Model, 55U8000H) Crystal Processor 4K, Endless Free Content, Motion Xcelerator, Color Booster, Alexa Built-in

Samsung 55-Inch Class Crystal UHD U8000H Series Samsung Vision AI Smart TV (2026 Model, 55U8000H) Crystal Processor 4K, Endless Free Content, Motion Xcelerator, Color Booster, Alexa Built-in

  • Crystal Processor 4K: Enhances colors and sharpens details
  • Endless Free Content: Access 2,700+ free streaming options
  • Samsung TV Plus: 750+ free channels included

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Inside Inkling’s Architecture

Inkling uses a 66-layer decoder-only architecture that routes each token to six of 256 experts, alongside two shared experts. Images are divided into 40-by-40-pixel patches, while audio is represented through dMel spectrograms. Thinking Machines Lab says these inputs are projected into a shared space and processed jointly rather than being attached through a separate vision adapter.

The flagship’s scale limits who can operate it directly. The supplied estimates place BF16 deployment at at least 2 terabytes of aggregate VRAM, with NVFP4 still requiring about 600 gigabytes. Thinking Machines Lab has also previewed Inkling-Small, a 276-billion-parameter model with 12 billion active parameters, but its full weights remain pending while testing continues.

“Inkling is not the strongest model available today, closed or open.”

— Thinking Machines Lab, in its Inkling announcement

Benchmarks and Usage Rules Need Checking

Independent evaluators have not yet reproduced Inkling’s main performance and efficiency claims. Vendor figures place it at 97.1% on AIME 2026 and 87.2% on GPQA Diamond, while showing weaker results on Humanity’s Last Exam, SWE-bench Pro and Terminal-Bench 2.1 than some competitors. It is not yet clear how the model performs across real production workloads or whether its adjustable effort control delivers similar savings outside the cited tests.

The supplied reporting also refers to a possible Model Acceptable Use Policy covering the original parameters and modified versions, with restrictions involving surveillance, deception and fully automated decisions affecting rights. That policy was reported but not independently verified in the source material. Prospective users will need to examine the current model card and repository terms because any additional policy could affect uses that appear permitted by Apache 2.0 alone.

Independent Tests and Smaller Weights Awaited

Researchers and prospective adopters will now test Inkling against GLM-5.2, Kimi K2.6 and other open-weight models using their own tasks, hardware and cost limits. Attention will also turn to independent reproduction of the published benchmark scores, clarification of any separate usage restrictions, and the release of Inkling-Small’s full weights. That smaller model may be the more practical measure of the lab’s open strategy for teams without data-center-scale infrastructure.

Key Questions

What did Thinking Machines Lab release?

The company released Inkling’s BF16 and NVFP4 model weights on Hugging Face under Apache 2.0, together with support for several common inference frameworks.

Is Inkling fully open source?

No. The model weights are downloadable, but the lab has not published the training data or complete training pipeline. The release is best described as open weight.

Can Inkling run on a consumer workstation?

The flagship is unlikely to run at useful capacity on ordinary consumer hardware. Supplied estimates call for about 2 terabytes of VRAM for BF16 or at least 600 gigabytes for NVFP4.

Does Inkling outperform every open model?

No. Thinking Machines Lab says Inkling is not the strongest model overall. Its published results are competitive in mathematics, calibration and audio, but it trails other models on several coding, terminal and general reasoning tests.

When will Inkling-Small become available?

The lab has previewed Inkling-Small and said its full weights will follow after testing. The supplied source material gives no confirmed release date.

Source: Thorsten Meyer AI

You May Also Like

China: The Visible Hand

Thorsten Meyer AI’s China entry says Beijing is using state capital and planning for AI and robotics while worker protections remain uneven.

Can AI Elevate Your Work And Gaming Setup? Experts Say Yes

A 10-monitor comparison ranks Dell’s S3425DW first, while MSI, Samsung and LG lead value, productivity and office categories.

How Gaudí’s Crypt paved the way for parametricism

Research shows Gaudí’s crypt influenced the development of parametric design, marking a pivotal moment in architectural history.

The Retail Giant That Supplied Europe’s AI Future Through Private Capital

Schwarz Group is building a subsid-free, 200-megawatt AI data center in Brandenburg designed for up to 100,000 GPUs.