Explore AI models and compare capabilities, providers, and pricing.
GPT 6.1 Sol
OpenAI
GPT 6 Luna
GPT 6 Sol
DeepSeek V4.1 Flash
DeepSeek
GPT 6 Astra
GLM 5.3 Flash
Zhipu AI
GLM 5.3
Opus 5
Anthropic
Kimi K3
Moonshot AI
GPT 5.6 Luna
GPT 6.1 Sol delivers near-Astra performance at lower cost for complex coding, computer use, professional work, and agentic tool use.
by openai|Reasoning|Tools|Vision
$2.41/M input$1.69/M input30% off|$12.05/M output$8.43/M output
GPT 6 Luna is the fastest GPT-6 variant for high-volume chat, extraction, routing, and lightweight tools.
$0.12/M input$0.084/M input30% off|$0.602/M output$0.422/M output
GPT 6 Sol balances GPT-6 reasoning quality with lower latency and cost for production agent workloads.
DeepSeek V4.1 Flash combines a one-million-token context window with configurable thinking and vision input.
by deepseek|Tools|Vision|Web
$0.361/M input$0.253/M input30% off|$1.45/M output$1.01/M output
GPT 6 Astra is OpenAI's flagship GPT-6 model for complex reasoning, coding, research, and long-running agents.
$12.05/M input$8.43/M input30% off|$60.24/M output$42.17/M output
GLM 5.3 Flash is the faster multimodal GLM variant for coding, screenshots, and tool-driven workflows.
by zhipu-ai|Reasoning|Tools|Vision
$0.181/M input$0.127/M input30% off|$0.602/M output$0.422/M output
GLM 5.3 is Zhipu AI's flagship long-context model for advanced coding and autonomous agents.
by zhipu-ai|Reasoning|Tools|Web
$1.69/M input$1.18/M input30% off|$5.3/M output$3.71/M output
Claude Opus 5 is Anthropic's advanced model for complex agentic coding and enterprise knowledge work.
by anthropic|Reasoning|Tools|Web
$6.02/M input$4.22/M input30% off|$30.12/M output$21.08/M output
Kimi K3 is moonshot's general-purpose language and agent model with a 1,047,000-token context window, vision and tool calling, and API streaming support.
by moonshot-ai|Reasoning|Tools|Vision
$3.61/M input$2.53/M input30% off|$18.07/M output$12.65/M output
GPT 5.6 Luna is the low-latency GPT-5.6 option for high-throughput chat and structured tasks.
$0.241/M input$0.169/M input30% off|$1.45/M output$1.01/M output
GPT 5.6 Sol is the highest-capability GPT-5.6 variant for demanding coding and multimodal reasoning.
$4.82/M input$3.37/M input30% off|$24.1/M output$16.87/M output
GPT 5.6 Terra is the balanced GPT-5.6 option for general-purpose agents, tools, and image understanding.
$2.41/M input$1.69/M input30% off|$14.46/M output$10.12/M output
Claude Sonnet 5 combines fast responses with frontier intelligence for coding and everyday tool-heavy agents.
GLM 5.2 focuses on coding, structured generation, diagrams, and optional reasoning over long contexts.
Claude Fable 5 is Anthropic's long-horizon reasoning model for demanding analysis and autonomous agent workflows.
Nemotron Ultra is NVIDIA's large mixture-of-experts model for difficult reasoning and multi-step agent tasks.
by nvidia|Tools|Web
$0.723/M input$0.506/M input30% off|$2.89/M output$2.02/M output
Gemini 3.5 Flash targets fast, high-throughput multimodal workloads with built-in thinking support.
by google|Tools|Vision|Web
$1.81/M input$1.27/M input30% off|$10.84/M output$7.59/M output
DeepSeek V4 Flash is the legacy Flash identifier now served by DeepSeek V4.1 Flash for fast reasoning.
by deepseek|Tools|Web
DeepSeek V4 Pro is the higher-capability DeepSeek V4 endpoint for long-context reasoning and agentic coding.
$1.59/M input$1.11/M input30% off|$4.77/M output$3.34/M output
GPT 5.5 is openai's general-purpose language and agent model with a 400,000-token context window, vision and tool calling, and API streaming support.
$6.02/M input$4.22/M input30% off|$36.14/M output$25.3/M output
Kimi 2.6 is moonshot's general-purpose language and agent model with a 262,144-token context window, vision and tool calling, and API streaming support.
$1.14/M input$0.801/M input30% off|$4.82/M output$3.37/M output
GPT 5.4 is openai's general-purpose language and agent model with a 400,000-token context window, vision and tool calling, and API streaming support.
$3.01/M input$2.11/M input30% off|$18.07/M output$12.65/M output
GPT 5.4 Thinking is openai's deliberate reasoning and complex analysis model with a 400,000-token context window, vision and tool calling, and API streaming support.
Gemini 3.1 Flash-Lite is optimized for economical high-volume translation, extraction, and simple agent tasks.
$0.301/M input$0.211/M input30% off|$1.81/M output$1.27/M output
Fast Google image generation and editing.
by google|Vision
$0.0807/image$0.0565/image30% off
Mercury 2 is Inception Labs' diffusion language model for fast reasoning, tool use, and multimodal chat.
by inception-labs|Tools|Vision|Web
Gemini 3.1 Pro Preview is Google's advanced multimodal model for agentic workflows, complex reasoning, and coding.
GPT 5.3 Codex is openai's coding and software-engineering model with a 400,000-token context window, vision and tool calling, and API streaming support.
$2.11/M input$1.48/M input30% off|$16.87/M output$11.81/M output
Kimi K2.5 is moonshot's deliberate reasoning and complex analysis model with a 256,000-token context window, vision and tool calling, and API streaming support.
$0.723/M input$0.506/M input30% off|$3.61/M output$2.53/M output
GLM 4.7 is an efficient earlier-generation Zhipu model for coding, structured output, and tool use.
$0.723/M input$0.506/M input30% off|$2.65/M output$1.86/M output
GPT 5.2 Codex is openai's coding and software-engineering model with a 400,000-token context window, vision and tool calling, and API streaming support.
Gemini 3 Flash Preview is Google's legacy Gemini 3 speed-oriented model for multimodal generation.
$0.602/M input$0.422/M input30% off|$3.61/M output$2.53/M output
OpenAI image generation and editing.
by openai|Vision
$0.0602/image$0.0422/image30% off
GPT 5.2 is openai's general-purpose language and agent model with a 400,000-token context window, vision and tool calling, and API streaming support.
GPT 5.2 Thinking is openai's deliberate reasoning and complex analysis model with a 400,000-token context window, vision and tool calling, and API streaming support.
Ministral 3B is Mistral's compact model for local or high-volume classification, routing, and short responses.
by mistral
$0.12/M input$0.084/M input30% off|$0.12/M output$0.084/M output
Mistral Large 3 is an open-weight multimodal mixture-of-experts model for general-purpose enterprise workloads.
$0.602/M input$0.422/M input30% off|$1.81/M output$1.27/M output
Claude Opus 4.5 is Anthropic's high-capability model for demanding coding, research, writing, and agent workflows.
by anthropic|Reasoning|Tools|Vision
Grok 4.1 non-reasoning is a legacy low-latency alias now redirected by xAI to Grok 4.3 with reasoning disabled.
by xai|Reasoning|Tools|Vision
$1.51/M input$1.05/M input30% off|$3.01/M output$2.11/M output
GPT 5.1 is openai's general-purpose language and agent model with a 400,000-token context window, vision and tool calling, and API streaming support.
$1.51/M input$1.05/M input30% off|$12.05/M output$8.43/M output
GPT 5.1 Codex is openai's coding and software-engineering model with a 400,000-token context window, vision and tool calling, and API streaming support.
Kimi K2 Thinking is moonshot's deliberate reasoning and complex analysis model with a 256,000-token context window, vision and tool calling, and API streaming support.
$0.723/M input$0.506/M input30% off|$3.01/M output$2.11/M output
Claude Haiku 4.5 is Anthropic's fastest model for responsive chat, classification, coding, and image analysis.
by anthropic|Tools|Vision|Web
$1.2/M input$0.843/M input30% off|$6.02/M output$4.22/M output
OpenAI video generation.
by openai|video|Vision|Video
$0.12/second$0.0843/second30% off
Claude Sonnet 4.5 balances intelligence and speed for coding, writing, vision, and tool-based agents.
GPT 5 Codex is openai's coding and software-engineering model with a 400,000-token context window, vision and tool calling, and API streaming support.
GPT 5 is openai's general-purpose language and agent model with a 128,000-token context window, vision and tool calling, and API streaming support.
GPT 5 Mini is openai's low-latency and high-volume model with a 400,000-token context window, vision and tool calling, and API streaming support.
$0.301/M input$0.211/M input30% off|$2.41/M output$1.69/M output
Claude Opus 4.1 is a previous-generation Anthropic flagship for deep analysis, coding, and long-form writing.
$18.07/M input$12.65/M input30% off|$90.36/M output$63.25/M output
GPT OSS 120B is openai's general-purpose language and agent model with a 128,000-token context window, tool calling, and API streaming support.
by openai|Tools|Web
GPT OSS 20B is openai's general-purpose language and agent model with a 131,072-token context window, tool calling, and API streaming support.
$0.06/M input$0.042/M input30% off|$0.241/M output$0.169/M output
Gemini 2.5 Flash-Lite is Google's smallest 2.5 model for cost-sensitive classification and data processing.
$0.12/M input$0.084/M input30% off|$0.482/M output$0.337/M output
Gemini 2.5 Flash is Google's hybrid reasoning model for low-latency multimodal applications.
$0.361/M input$0.253/M input30% off|$3.01/M output$2.11/M output
GPT o4 Mini is openai's low-latency and high-volume model with a 128,000-token context window and API streaming support.
by openai|Reasoning|Web
$1.33/M input$0.928/M input30% off|$5.3/M output$3.71/M output
Ashna-X1 is AshnaAI's balanced default for everyday chat, web research, tool use, and image-aware workflows.
by ashnaai|Tools|Vision|Web
GPT 4.1 is openai's general-purpose language and agent model with a 1,047,576-token context window, vision and tool calling, and API streaming support.
by openai|Tools|Vision|Web
$2.41/M input$1.69/M input30% off|$9.64/M output$6.75/M output
GPT 4.1 Mini is openai's low-latency and high-volume model with a 1,047,576-token context window, vision and tool calling, and API streaming support.
$0.482/M input$0.337/M input30% off|$1.93/M output$1.35/M output
GPT 4.1 Nano is openai's low-latency and high-volume model with a 1,047,576-token context window, vision and tool calling, and API streaming support.
Gemini 2.5 Pro is a one-million-context multimodal reasoning model that excels at coding and complex analysis.
Streaming speech-to-text for voice applications.
by deepgram|voice|Voice
$0.00578/minute$0.00405/minute30% off
GPT o3 Mini is openai's low-latency and high-volume model with a 128,000-token context window and API streaming support.
GPT o1 is openai's deliberate reasoning and complex analysis model with a 128,000-token context window and API streaming support.
$18.07/M input$12.65/M input30% off|$72.29/M output$50.6/M output
GPT 4o Mini is openai's low-latency and high-volume model with a 128,000-token context window, vision and tool calling, and API streaming support.
$0.181/M input$0.127/M input30% off|$0.723/M output$0.506/M output
GPT 4o is openai's general-purpose language and agent model with a 128,000-token context window, vision and tool calling, and API streaming support.
$3.01/M input$2.11/M input30% off|$12.05/M output$8.43/M output
Mixtral 8x22B is an open mixture-of-experts model for multilingual text generation, coding, and analysis.
$2.41/M input$1.69/M input30% off|$7.23/M output$5.06/M output
Multilingual text-to-speech model.
by elevenlabs|voice|Voice
$0.00012/character$0.0000843/character30% off
Grok 4.3 is xAI's fast one-million-context model with image input, configurable reasoning, and tool calling.
by xai|Reasoning|Tools|Web
Highest-accuracy Voyage cross-encoder reranker for refining retrieval results.
by voyage-ai
$0.06/M input$0.042/M input30% off|$0/M output$0/M output
Fast and cost-effective Voyage reranker for latency-sensitive retrieval.
$0.024/M input$0.017/M input30% off|$0/M output$0/M output
Code embeddings optimized for coding-agent and repository retrieval.
$0.145/M input$0.101/M input30% off|$0/M output$0/M output
Contextualized chunk embeddings that encode each chunk with the surrounding document context.
Balanced general-purpose and multilingual retrieval embeddings in the shared Voyage 4 vector space.
$0.072/M input$0.051/M input30% off|$0/M output$0/M output
Highest-quality Voyage general-purpose and multilingual retrieval embeddings.
Low-latency, low-cost retrieval embeddings in the shared Voyage 4 vector space.
Interleaved text, image, and video embeddings for rich multimodal retrieval.
by voyage-ai|Vision|Files
$0.145/M text + $0.723/B pixels$0.101/M text + $0.506/B pixels30% off
Previous-generation multilingual reranker with instruction-following support.
Previous-generation low-latency multilingual and instruction-aware reranker.
Previous-generation contextualized chunk embeddings for document-aware retrieval.
$0.217/M input$0.152/M input30% off|$0/M output$0/M output
Previous-generation general-purpose and multilingual retrieval embeddings.
Previous-generation retrieval embeddings optimized for latency and cost.
Previous-generation high-quality general-purpose and multilingual retrieval embeddings.
Previous-generation code retrieval embeddings with flexible dimensions and quantization.
Previous-generation interleaved text and image embeddings for visual retrieval.
General-purpose and multilingual retrieval embeddings with a fixed 1024-dimensional output.
Cost-efficient general-purpose retrieval embeddings with a fixed 512-dimensional output.
Multilingual retrieval and RAG embeddings supporting major global languages.
Domain-specific embeddings optimized for financial retrieval and RAG.
Instruction-tuned embeddings for retrieval, classification, clustering, and similarity.
Domain-specific embeddings optimized for legal and long-context retrieval.
High-quality OpenAI text embeddings.
by openai
$0.157/M input$0.11/M input30% off|$0/M output$0/M output
Low-cost OpenAI text embeddings.
Second-generation semantic code search and code retrieval embeddings.
Legacy general-purpose embeddings balancing retrieval quality, latency, and cost.
$0.12/M input$0.084/M input30% off|$0/M output$0/M output
Legacy general-purpose retrieval model with 1536-dimensional embeddings.