Back to Directory
Inkling logo

Inkling

New

Work with a full million-token context window across text and images, using free open weights or a managed fine-tuning path through Tinker.

Models
4.5freemium

The verdict on Inkling: Developers who need an entire codebase or long document set to stay in a single context pass, not chunked Inkling's core use case is developers who need an entire codebase or long document set to stay in a single context pass rather than being chunked. Pricing: Free open-weight download (Hugging Face) / paid managed fine-tuning via Tinker. Last reviewed: August 2026.

Best For

Developers who need an entire codebase or long document set to stay in a single context pass, not chunked

Standout Feature

A 975-billion-parameter open-weight model with a genuine one-million-token context window across text and images

TL;DR

Genuinely capable at scale for free, self-hosting the full model requires serious compute and it's a brand-new lab with limited track record.

Alternatives

Overview

Inkling is Thinking Machines Lab's first open-weight model release, a 975-billion-parameter mixture-of-experts model built to handle text, images, and a one-million-token context window in a single system. The large context window means entire codebases, long document sets, or extended conversation histories can stay in a single pass instead of being chunked and stitched back together. Weights are downloadable for free from Hugging Face for teams that want to self-host or fine-tune locally, while Thinking Machines Lab also offers a paid fine-tuning service through its Tinker platform for teams that want a managed path to a custom version without building their own training infrastructure.

A smaller Inkling-Small variant trades some capability for lower compute requirements, useful for teams that want the same architecture on more modest hardware. Inkling fits research teams and technical builders who want an open, multimodal, long-context model they can inspect, fine-tune, and deploy on their own terms.

Our Take

The 975-billion-parameter open-weight model with a genuine one-million-token context window across text and images is the technical centerpiece, and free downloadable weights mean no ongoing API costs. A managed fine-tuning path through Tinker is available if self-hosting isn't the goal. Two constraints apply: full-size self-hosting demands significant compute, and Thinking Machines Lab is a new organization with limited third-party track record. If long-context reasoning on large inputs is the primary requirement, this is one of the few open-weight options that addresses it directly.

Was this useful?

Key Features

  • 975-billion-parameter mixture-of-experts architecture
  • 1-million-token context window
  • Multimodal text and image input
  • Smaller Inkling-Small variant for lighter hardware
  • Managed fine-tuning available via Tinker
Pros
  • Free, downloadable open weights with no usage fees
  • Genuinely long context window for large documents or codebases
  • Managed fine-tuning option for teams without training infrastructure
Cons
  • Full-size model requires significant compute to self-host
  • New lab and release, limited third-party track record so far

Other Models tools builders reach for alongside Inkling.