Uploaded July 2026 | Updated September 2026, 2 weeks ago
In this video, I break down Inkling, Thinking Machines’ first open-weight model release under Apache 2.0, trained from scratch as a multimodal decoder-only MoE. I cover the core architecture (near‑trillion parameters, 1M context window, 45T multimodal tokens, 256 experts with 41B active params), how it compares on benchmarks (not SOTA overall but a strong generalist with notable design performance), and why it matters as a major Western open model release. I also walk through the Tinker platform for fine-tuning/post-training, mention the upcoming Inkling Small (276B/12B active), show playground settings like reasoning effort, context limits, and web search/tool calling, demonstrate website generation and a coding example (ISS tracker), and review pricing and options to run weights yourself.
LINK:
thinkingmachines.ai/news/introducing-inkling
My voice to text App: whryte.com
Website: engineerprompt.ai
RAG Beyond Basics Course:
prompt-s-site.thinkific.com/courses/rag
Signup for Newsletter, localgpt:
https://tally.so/r/3y9bb0
Let's Connect:
🦾 Discord: discord.com/invite/t4eYQRUcXB
☕ Buy me a Coffee: ko-fi.com/promptengineering
|🔴 Patreon: patreon.com/PromptEngineering
💼Consulting: calendly.com/engineerprompt/consulting-call
📧 Business Contact: engineerprompt@gmail.com
Become Member: tinyurl.com/y5h28s6h
💻 Pre-configured localGPT VM: bit.ly/localGPT (use Code: PromptEngineering for 50% off).
Signup for Newsletter, localgpt:
https://tally.so/r/3y9bb0
Inkling: Thinking Machines’ First Open-Weight Multimodal MoE Model (Architecture, Playground Demos & Pricing)
00:00 Inkling
00:26 Why This Matters
01:05 Architecture and Scale
01:55 Inkling Small
02:57 Multimodal Strengths
04:06 Reasoning Effort Controls
04:36 Playground Tour and Prompt
06:50 Tool Use and Speed
08:13 Website Design Patterns
09:21 Coding Demo ISS Tracker
In this video, I break down Inkling, Thinking Machines’ first open-weight model release under Apache 2.0, trained from scratch as a multimodal decoder-only MoE. I cover the core architecture (near‑trillion parameters, 1M context window, 45T multimodal tokens, 256 experts with 41B active params), how it compares on benchmarks (not SOTA overall but a strong generalist with notable design performance), and why it matters as a major Western open model release. I also walk through the Tinker platform for fine-tuning/post-training, mention the upcoming Inkling Small (276B/12B active), show playground settings like reasoning effort, context limits, and web search/tool calling, demonstrate website generation and a coding example (ISS tracker), and review pricing and options to run weights yourself.
LINK:
thinkingmachines.ai/news/introducing-inkling
My voice to text App: whryte.com
Website: engineerprompt.ai
RAG Beyond Basics Course:
prompt-s-site.thinkific.com/courses/rag
Signup for Newsletter, localgpt:
https://tally.so/r/3y9bb0
Let's Connect:
🦾 Discord: discord.com/invite/t4eYQRUcXB
☕ Buy me a Coffee: ko-fi.com/promptengineering
|🔴 Patreon: patreon.com/PromptEngineering
💼Consulting: calendly.com/engineerprompt/consulting-call
📧 Business Contact: engineerprompt@gmail.com
Become Member: tinyurl.com/y5h28s6h
💻 Pre-configured localGPT VM: bit.ly/localGPT (use Code: PromptEngineering for 50% off).
Signup for Newsletter, localgpt:
https://tally.so/r/3y9bb0
Inkling: Thinking Machines’ First Open-Weight Multimodal MoE Model (Architecture, Playground Demos & Pricing)
00:00 Inkling
00:26 Why This Matters
01:05 Architecture and Scale
01:55 Inkling Small
02:57 Multimodal Strengths
04:06 Reasoning Effort Controls
04:36 Playground Tour and Prompt
06:50 Tool Use and Speed
08:13 Website Design Patterns
09:21 Coding Demo ISS Tracker










