Models

Explore AI model pricing, capabilities, endpoints, and vendor coverage from one production catalog.

Alibaba

Qwen3.8 Max

qwen3.8-max

Qwen3.8-Max is a 2.4T-parameter MoE flagship model designed for advanced reasoning, coding, and enterprise productivity. It can autonomously execute complex, long-running tasks — from software development to professional workflows — delivering production-grade outcomes across domains such as law, finance, and design. Powered by native multimodal intelligence, Qwen3.8-Max understands text, images, and videos across long contexts, supporting deep analysis of extensive documents and long-form content. Its agentic capabilities enable autonomous planning, execution, verification, and continuous iteration throughout complex workflows.

Total Context

1M

Max Output

128K

Released

Aug 3, 2026

Input$2 / M tokens
Output$6 / M tokens
Cache Read$0.25 / M tokens

Moonshot

Kimi K3

kimi-k3

Kimi K3 is a high-performance AI model designed for advanced reasoning, coding, and complex task execution. It delivers strong performance on large-scale programming projects, multi-step problem solving, and knowledge-intensive applications. With broad context handling and multimodal capabilities, K3 is suitable for building powerful AI agents, developer tools, and enterprise-level applications.

Total Context

1M

Max Output

131.1K

Released

Jul 16, 2026

Input$2.8571 / M tokens
Output$14.2857 / M tokens
Cache Read$0.2857 / M tokens

OpenAI

GPT-5.6 Sol

gpt-5.6-sol

GPT-5.6 Sol is the standard version of the Sol series. It is suitable for general Q&A, text processing, content drafting, summarization, lightweight coding assistance, and business automation tasks. Sol is designed for scenarios that need solid general capability while keeping response efficiency and cost in mind. It can be used for product features, chat assistants, content-generation tools, data cleaning, knowledge-base Q&A, and batch-processing tasks. For large-scale usage or rapid prototyping, Sol is a flexible starting point.

Total Context

1.1M

Max Output

128K

Released

Jul 9, 2026

Input$5 / M tokens
Output$30 / M tokens
Cache Read$0.5 / M tokens

OpenAI

GPT-5.6 Sol Pro

gpt-5.6-sol-pro

GPT-5.6 Sol Pro is the Pro version of the Sol series, designed for complex tasks and higher-quality output. It is suitable for advanced writing, precise Q&A, business reasoning, code generation, structured analysis, and multi-turn task execution. Sol Pro is a good fit for tasks that require stronger accuracy, contextual understanding, and instruction completion. It can be used for professional content production, data-analysis assistants, knowledge Q&A systems, complex agent workflows, and internal enterprise AI tools.

Total Context

1.1M

Max Output

128K

Released

Jul 9, 2026

Input$5 / M tokens
Output$30 / M tokens
Cache Read$0.5 / M tokens

OpenAI

GPT-5.6 Terra

gpt-5.6-terra

GPT-5.6 Terra is the general-purpose version of the Terra series. It is suitable for everyday conversation, text generation, summarization, information extraction, lightweight reasoning, and coding assistance. Terra offers a practical balance between general capability, response efficiency, and cost. It can be used for chat assistants, content tools, office automation, knowledge-base Q&A, and batch text-processing workflows. For fast iteration or cost-sensitive applications, Terra is a strong default option within the Terra series.

Total Context

1.1M

Max Output

128K

Released

Jul 9, 2026

Input$2.5 / M tokens
Output$15 / M tokens
Cache Read$0.275 / M tokens

OpenAI

GPT-5.6 Terra Pro

gpt-5.6-terra-pro

GPT-5.6 Terra Pro is the enhanced version of the Terra series, designed for high-quality output across complex tasks. It is suitable for advanced Q&A, in-depth content creation, business analysis, coding assistance, and multi-step reasoning. Compared with the standard version, Terra Pro is better suited for scenarios that require stronger reliability, more complete reasoning, and better instruction following. It can be a good fit for professional writing, complex information synthesis, knowledge-based assistants, automation agents, and enterprise-grade AI applications.

Total Context

1.1M

Max Output

128K

Released

Jul 9, 2026

Input$2.5 / M tokens
Output$15 / M tokens
Cache Read$0.275 / M tokens

OpenAI

GPT-5.6 Luna

gpt-5.6-luna

GPT-5.6 Luna is the general-purpose version of the Luna series, suitable for everyday AI applications. It can be used for chat Q&A, text rewriting, summarization, information extraction, simple reasoning, coding assistance, and office automation. Luna is a practical choice for prototypes, lightweight business features, batch content processing, and common user interaction scenarios. It is a good starting model when speed, simplicity, and cost control are important, while Luna Pro can be considered for tasks that require higher-quality or more stable outputs.

Total Context

1.1M

Max Output

128K

Released

Jul 9, 2026

Input$1 / M tokens
Output$6 / M tokens
Cache Read$0.1 / M tokens

OpenAI

GPT-5.6 Luna Pro

gpt-5.6-luna-pro

GPT-5.6 Luna Pro is the enhanced version of the Luna series, intended for higher-quality answers, more stable instruction following, and stronger task completion. It is suitable for complex conversations, content generation, knowledge Q&A, business analysis, coding assistance, and multi-step automation tasks. Luna Pro is well suited for product experiences where output quality and user experience matter, such as professional writing assistants, enterprise AI assistants, customer support, knowledge-base systems, and agent workflows. It is a good option when the task requires more reliability than the standard Luna model.

Total Context

1.1M

Max Output

128K

Released

Jul 9, 2026

Input$1 / M tokens
Output$6 / M tokens
Cache Read$0.1 / M tokens

DeepSeek

DeepSeek V4 Pro

deepseek-v4-pro

DeepSeek V4 Pro is described as a large-scale Mixture-of-Experts model with 1.6T total parameters and 49B activated parameters, while keeping a 1M-token context window for very large inputs. Its model cards emphasize advanced reasoning, coding, and long-horizon agent workflows rather than simple chat. The Pro variant is the capability-oriented member of the V4 family, making it better suited to full-codebase analysis, large research synthesis, and multi-step automation where depth matters more than the lowest possible latency.

Total Context

1M

Max Output

384K

Released

Apr 24, 2026

Input$1.2857 / M tokens
Output$3.8571 / M tokens
Cache Read$0.0429 / M tokens

Z.ai

GLM-5.2

glm-5.2

GLM-5.2 is Z.ai’s flagship foundation model for long-horizon engineering tasks. It emphasizes a usable 1M-token context window for project-scale code and system context, more stable execution on long tasks, and better adherence to engineering standards. It is positioned for full development workflows, from requirements and repository analysis to implementation, testing, and multi-platform deployment, where large context and sustained agent behavior matter.

Total Context

1M

Max Output

131.1K

Released

Jun 13, 2026

Input$1.1429 / M tokens
Output$4 / M tokens
Cache Read$0.2857 / M tokens

DeepSeek

DeepSeek V4 Flash

deepseek-v4-flash

DeepSeek V4 Flash keeps the V4 family’s 1M-token context window but uses a lighter MoE configuration, commonly described as 284B total parameters with 13B activated parameters. The emphasis is throughput: fast inference, lower cost per call, and production workloads that still need long-context handling. It is the better fit when the task volume is high and the workload benefits from V4-style long-context architecture without always requiring the deepest reasoning tier.

Total Context

1M

Max Output

384K

Released

Apr 24, 2026

Input$0.4286 / M tokens
Output$1.2857 / M tokens
Cache Read$0.0143 / M tokens

Anthropic

Claude Opus 5

claude-opus-5

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis of charts and documents, complex office deliverables, and coordinating parallel subagents. The model maintains strong instruction following and tool use across extended tasks, while remaining effective at lower effort settings for workloads that prioritize latency and token efficiency.

Total Context

1M

Max Output

128K

Released

Jul 24, 2026

Input$5 / M tokens
Output$25 / M tokens
Cache Read$0.5 / M tokens

Anthropic

Claude Opus 5 Fast

claude-opus-5-fast

Fast-mode variant of Opus 5 - identical capabilities with higher output speed at 2x pricing relative to regular Opus 5. Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode

Total Context

1M

Max Output

128K

Released

Jul 24, 2026

Input$10 / M tokens
Output$50 / M tokens
Cache Read$1 / M tokens

xAI

Grok 4.6

grok-4.6

Grok 4.6 is xAI’s advanced reasoning model built for complex programming, long-running agent tasks, and knowledge work. It supports text and image input, a context window of up to 500k tokens, and configurable low/medium/high/xhigh reasoning intensity. With function calling, web and X search, code execution, and structured output, Grok 4.6 is optimized for cross-codebase analysis, sustained multi-step tool use, self-verification, and end-to-end engineering workflows. It is best understood as a general-purpose reasoning and agentic model—not merely a chatbot associated with the X platform.

Total Context

500K

Max Output

500K

Released

Aug 12, 2026

Input$2 / M tokens
Output$6 / M tokens
Cache Read$0.3 / M tokens

Google

Gemini 3.6 Flash

gemini-3.6-flash

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and less hedging, while reducing token use and the number of model calls needed to complete a task.

Total Context

1M

Max Output

65.5K

Released

Jul 21, 2026

Input$1.5 / M tokens
Output$7.5 / M tokens
Cache Read$0.15 / M tokens

xAI

Grok 4.5

grok-4.5

Grok 4.5 is xAI's reasoning model designed for programming, agent tasks, and knowledge work. It supports text and image input, up to 500k token context, and configurable low/medium/high reasoning intensity. Features include function calling, web and X search, code execution, and structured output, making it suitable for long-context analysis, multi-step tool usage, and end-to-end engineering workflows. It should not be seen merely as a chatbot tied to the X platform.

Total Context

500K

Max Output

500K

Released

Jul 8, 2026

Input$2 / M tokens
Output$6 / M tokens
Cache Read$0.3 / M tokens

Alibaba

Qwen3.7 Plus

qwen3.7-plus

Qwen3.7 Plus takes the Qwen3.7 agent-oriented design in a more cost-effective direction. Third-party model cards describe it as supporting text and image input, with stronger vision-language ability and hybrid agent capability for GUI, mobile navigation, and visual-reference tasks. The model is suitable when users need the new Qwen3.7 capability profile without always paying for the Max tier.

Total Context

1M

Max Output

64K

Released

Jun 2, 2026

Input$0.2857 / M tokens
Output$1.1429 / M tokens
Cache Read$0.0571 / M tokens

OpenAI

GPT-4.1

gpt-4.1

GPT-4.1 is an OpenAI model generation focused on improved coding, instruction following, and long-context performance. Official announcements present it as a stronger developer model than GPT-4o for many programming and instruction-heavy tasks. Its catalog description should highlight practical coding reliability and long-context understanding.

Total Context

1M

Max Output

32.8K

Released

Apr 14, 2025

Input$2 / M tokens
Output$8 / M tokens
Cache Read$0.5 / M tokens

OpenAI

GPT-4.1 Mini

gpt-4.1-mini

GPT-4.1 Mini brings the GPT-4.1 family’s coding and instruction-following improvements into a faster, lower-cost form. It is suitable for high-volume developer tools, structured generation, extraction, and product features that do not require the full model. The main distinction is production efficiency while retaining the 4.1 generation’s task discipline.

Total Context

1M

Max Output

32.8K

Released

Apr 14, 2025

Input$0.4 / M tokens
Output$1.6 / M tokens
Cache Read$0.1 / M tokens

OpenAI

GPT-4o

gpt-4o

GPT-4o is OpenAI’s multimodal flagship from the GPT-4o generation, built for text and image input with strong general intelligence. Official docs describe it as a versatile high-intelligence model suitable for a broad range of language and vision tasks. It remains useful where multimodal understanding and natural interaction matter more than the newest reasoning stack.

Total Context

128K

Max Output

16.4K

Released

May 13, 2024

Input$2.5 / M tokens
Output$10 / M tokens
Cache Read$1.25 / M tokens

OpenAI

GPT-4o Mini

gpt-4o-mini

GPT-4o Mini is the fast and affordable small model in the GPT-4o family. OpenAI docs position it for focused tasks with text and image input, structured outputs, fine-tuning, and distillation workflows. It is best introduced as a lightweight multimodal production model rather than a reduced copy of GPT-4o.

Total Context

128K

Max Output

16.4K

Released

Jul 18, 2024

Input$0.15 / M tokens
Output$0.6 / M tokens
Cache Read$0.075 / M tokens

Popular model recommendations

Start with high-signal models from the live catalog, then open a detail page to compare context, endpoints, and effective pricing.

Alibaba

Qwen3.8 Max

Qwen3.8-Max is a 2.4T-parameter MoE flagship model designed for advanced reasoning, coding, and enterprise productivity. It can autonomously execute complex, long-running tasks — from software development to professional workflows — delivering production-grade outcomes across domains such as law, finance, and design. Powered by native multimodal intelligence, Qwen3.8-Max understands text, images, and videos across long contexts, supporting deep analysis of extensive documents and long-form content. Its agentic capabilities enable autonomous planning, execution, verification, and continuous iteration throughout complex workflows.

Total Context

1M

Input Price

$2 / M tokens

View model

Moonshot

Kimi K3

Kimi K3 is a high-performance AI model designed for advanced reasoning, coding, and complex task execution. It delivers strong performance on large-scale programming projects, multi-step problem solving, and knowledge-intensive applications. With broad context handling and multimodal capabilities, K3 is suitable for building powerful AI agents, developer tools, and enterprise-level applications.

Total Context

1M

Input Price

$2.8571 / M tokens

View model

OpenAI

GPT-5.6 Sol

GPT-5.6 Sol is the standard version of the Sol series. It is suitable for general Q&A, text processing, content drafting, summarization, lightweight coding assistance, and business automation tasks. Sol is designed for scenarios that need solid general capability while keeping response efficiency and cost in mind. It can be used for product features, chat assistants, content-generation tools, data cleaning, knowledge-base Q&A, and batch-processing tasks. For large-scale usage or rapid prototyping, Sol is a flexible starting point.

Total Context

1.1M

Input Price

$5 / M tokens

View model

OpenAI

GPT-5.6 Sol Pro

GPT-5.6 Sol Pro is the Pro version of the Sol series, designed for complex tasks and higher-quality output. It is suitable for advanced writing, precise Q&A, business reasoning, code generation, structured analysis, and multi-turn task execution. Sol Pro is a good fit for tasks that require stronger accuracy, contextual understanding, and instruction completion. It can be used for professional content production, data-analysis assistants, knowledge Q&A systems, complex agent workflows, and internal enterprise AI tools.

Total Context

1.1M

Input Price

$5 / M tokens

View model

OpenAI

GPT-5.6 Terra

GPT-5.6 Terra is the general-purpose version of the Terra series. It is suitable for everyday conversation, text generation, summarization, information extraction, lightweight reasoning, and coding assistance. Terra offers a practical balance between general capability, response efficiency, and cost. It can be used for chat assistants, content tools, office automation, knowledge-base Q&A, and batch text-processing workflows. For fast iteration or cost-sensitive applications, Terra is a strong default option within the Terra series.

Total Context

1.1M

Input Price

$2.5 / M tokens

View model

OpenAI

GPT-5.6 Terra Pro

GPT-5.6 Terra Pro is the enhanced version of the Terra series, designed for high-quality output across complex tasks. It is suitable for advanced Q&A, in-depth content creation, business analysis, coding assistance, and multi-step reasoning. Compared with the standard version, Terra Pro is better suited for scenarios that require stronger reliability, more complete reasoning, and better instruction following. It can be a good fit for professional writing, complex information synthesis, knowledge-based assistants, automation agents, and enterprise-grade AI applications.

Total Context

1.1M

Input Price

$2.5 / M tokens

View model

Model Comparison

Quick comparison against selected catalog neighbors.

Model catalog FAQ

A quick guide for choosing, comparing, and using models from the TokenHub catalog.

How should I choose a model from this list?

+

Start with your workload. Use the filters to narrow by provider, tags, endpoint type, and billing group, then compare context size, output limit, modalities, and input or output pricing.

What does effective price mean?

+

Effective price applies the active billing group ratio to the model pricing data. It helps you estimate the real input, output, or per-request cost for the group you are using.

Can I use these models through API endpoints?

+

Yes. Open a model detail page to see the supported endpoint types and documentation links. Availability can differ by model, provider, and current routing configuration.

Why do context window and max output matter?

+

The context window controls how much prompt and conversation history a model can read. Max output controls how much text it can generate in one response, which matters for long-form writing, coding, and document tasks.