Models AI Tools
Browse and compare the best Models AI tools. Filter by pricing, features, and ratings.
Showing 23 of 23 tools
Llama 4 Scout and Maverick remain Meta's last open-weight frontier models (April 2025) with up to 10M-token context — Meta paused the open Llama line in 2026 in favor of a new proprietary flagship.
Get frontier-level coding and reasoning with a 1M-token context window at a fraction of Western competitor cost. Now on DeepSeek V4 (Flash and Pro tiers) with a permanent 75% price cut locked in May 2026.
Access Mistral Large 3, an open-weight, multilingual, multimodal flagship model at a fraction of the cost of closed competitors — from cloud API to edge deployment.
Run a self-hosted chat interface for local or API-based LLMs like Ollama behind your own login and controls — free at any scale if you keep default branding.
Access GPT and other OpenAI models through Azure with enterprise compliance, networking, and regional data controls. Now offers Global, Data Zone, and Regional deployment types.
Superseded by Qwen3.5 (Feb 2026) and then Qwen3.6 — Alibaba's current flagship line with strong agentic coding, repository-level reasoning, and multimodal understanding.
Chat with and compare many AI models and community-built bots through one subscription instead of juggling separate accounts. Pricing now transparent against per-model USD/token rates.
Chat with today's best open-weight AI models for free, with no API key, subscription, or vendor lock-in.
Train, deploy, and run inference on Gemini and 200+ third-party foundation models, plus build AI agents, billed per token and per compute node-hour.
Download, manage, and run large language models entirely on your own hardware, with a built-in chat interface and an OpenAI-compatible local server.
Get direct API access to generation, embedding, and reranking models built for enterprise search and retrieval, plus dedicated deployment for regulated environments. Now on the Command A family.
Get API access to foundation models from multiple providers, plus fine-tuning and agent tools, without managing infrastructure. New Priority and Flex service levels added alongside On-Demand and Provisioned Throughput.
Experiment with Gemini 2.5 Pro and 1M-token context for free, then ship with the same API key. The fastest path from Gemini prototype to production.
Superseded by Gemma 4 (April 2026) — Gemini-3-derived reasoning and agentic capability in five open sizes from 2B to 31B, running on phones, laptops, or servers with a 256K context window.
Chat with OpenAI, Claude, Gemini, and other models in one self-hosted interface, switching providers mid-conversation without losing context. Free forever once you host it yourself.
Run Llama, Mixtral, and 50+ open-source models at production speed — 3–5x cheaper than OpenAI-equivalent APIs with the same SDK you're already using.
Run vision and multi-step reasoning in a compact 15B model — read documents, ground UI elements, and solve math and science problems on modest hardware. Now Phi-4-reasoning-vision.
Access GLM-5.2 (1M token context, MIT-licensed) plus CogVideoX and CogView3 from one API — frontier-class language, video, and image generation at highly competitive prices.
Deploy optimized AI models as containers on your own GPUs — no inference tuning required. NIM ships every optimization pre-baked so you focus on the application.
Run Llama 70B at 1,800 tokens per second — 20x faster than GPU alternatives. The only inference provider where speed itself is the competitive moat.
Process 256K-token documents faster and cheaper than standard Transformers. Jamba's hybrid architecture is built for long-context enterprise workloads that break other models.
Run frontier-level coding and reasoning tasks on a 2.8-trillion-parameter open-weight model, with the option to self-host once the full weights ship.
Work with a full million-token context window across text and images, using free open weights or a managed fine-tuning path through Tinker.