Inkling
Work with a full million-token context window across text and images, using free open weights or a managed fine-tuning path through Tinker.
Alternatives
Overview
Inkling is Thinking Machines Lab's first open-weight model release, a 975-billion-parameter mixture-of-experts model built to handle text, images, and a one-million-token context window in a single system. The large context window means entire codebases, long document sets, or extended conversation histories can stay in a single pass instead of being chunked and stitched back together. Weights are downloadable for free from Hugging Face for teams that want to self-host or fine-tune locally, while Thinking Machines Lab also offers a paid fine-tuning service through its Tinker platform for teams that want a managed path to a custom version without building their own training infrastructure.
A smaller Inkling-Small variant trades some capability for lower compute requirements, useful for teams that want the same architecture on more modest hardware. Inkling fits research teams and technical builders who want an open, multimodal, long-context model they can inspect, fine-tune, and deploy on their own terms.
Key Features
- 975-billion-parameter mixture-of-experts architecture
- 1-million-token context window
- Multimodal text and image input
- Smaller Inkling-Small variant for lighter hardware
- Managed fine-tuning available via Tinker
- • Free, downloadable open weights with no usage fees
- • Genuinely long context window for large documents or codebases
- • Managed fine-tuning option for teams without training infrastructure
- • Full-size model requires significant compute to self-host
- • New lab and release, limited third-party track record so far
People Also Use
Other Models tools builders reach for alongside Inkling.
Llama 4
Llama 4 Scout and Maverick remain Meta's last open-weight frontier models (April 2025) with up to 10M-token context — Meta paused the open Llama line in 2026 in favor of a new proprietary flagship.
Open WebUI
Run a self-hosted chat interface for local or API-based LLMs like Ollama behind your own login and controls — free at any scale if you keep default branding.
Azure OpenAI Service
Access GPT and other OpenAI models through Azure with enterprise compliance, networking, and regional data controls. Now offers Global, Data Zone, and Regional deployment types.
Qwen 3
Superseded by Qwen3.5 (Feb 2026) and then Qwen3.6 — Alibaba's current flagship line with strong agentic coding, repository-level reasoning, and multimodal understanding.
Poe
Chat with and compare many AI models and community-built bots through one subscription instead of juggling separate accounts. Pricing now transparent against per-model USD/token rates.
HuggingChat
Chat with today's best open-weight AI models for free, with no API key, subscription, or vendor lock-in.