Models

Explore AI model pricing, capabilities, endpoints, and vendor coverage from one production catalog.

Anthropic

Claude Fable 5.1

claude-fable-5.1

Claude Fable 5 is Anthropic’s Mythos-class model designed for demanding reasoning, software engineering, knowledge work, visual analysis, and scientific research. Built for long-horizon agentic tasks, it can sustain planning and execution across complex workflows, analyze large codebases and documents, and refine its work over multiple steps. Claude Fable 5 is well suited to production applications that require advanced coding, research, professional analysis, and reliable autonomous task completion.

Total Context

1M

Max Output

128K

Released

Sep 1, 2026

Input$10 / M tokens
Output$50 / M tokens
Cache Read$0.25 / M tokens

Z.ai

GLM-5.3-Flash

glm-5.3-flash

GLM-5.3-Flash is Z.ai's first natively multimodal model in the GLM-5 family, built for fast coding, visual understanding, and agent workflows. Its 320B-parameter MoE architecture activates 18B parameters per token and combines sparse with linear attention to make long-context tasks more efficient. Use the GLM API through TokenHub to evaluate the current GLM-5.3-Flash price, model ID, supported inputs, tool calling, context limits, and endpoints before deploying it for production coding or multimodal agent tasks.

Total Context

1M

Max Output

131.1K

Released

Aug 26, 2026

Input$0.1143 / M tokens
Output$0.4 / M tokens
Cache Read$0.0329 / M tokens

Z.ai

GLM-5.3

glm-5.3

GLM-5.3 is Z.ai’s latest flagship model built for advanced coding, software engineering, and long-horizon agent tasks. With a 1M-token context window, up to 128K output tokens, strong reasoning, and tool-use capabilities, it is ideal for Coding Agents, complex development workflows, and multi-step automation.

Total Context

1M

Max Output

128K

Released

Aug 19, 2026

Input$1.1429 / M tokens
Output$4 / M tokens
Cache Read$0.2857 / M tokens

Tencent

Hy4 preview

hy4-preview

Hy4 preview is Tencent Hunyuan’s next-generation open-source Mixture-of-Experts flagship model, built for real-world productivity across coding, office work, scientific research, long-context reasoning, and complex agent workflows. Also known as Tencent Hy4 preview or Hunyuan 4 preview, it has 770B total parameters, activates 49B per token, and supports a context window of over 1M tokens. It excels at understanding large codebases and documents, multi-step planning, tool use, debugging, validation, and sustained task execution.

Total Context

1M

Max Output

64K

Released

Aug 28, 2026

Input$0.8571 / M tokens
Output$2.5714 / M tokens
Cache Read$0.0429 / M tokens

Alibaba

Qwen3.8 Max

qwen3.8-max

Qwen3.8-Max is a 2.4T-parameter MoE flagship model for advanced reasoning, coding, enterprise productivity, and agent workflows. It is designed to handle complex, long-running tasks from software development to professional workflows across domains such as law, finance, and design. Use the Qwen API through TokenHub to review the current Qwen3.8 Max price, model ID, context and output limits, supported modalities and endpoints, and benchmark results before deploying it in production.Powered by native multimodal intelligence, Qwen3.8-Max understands text, images, and videos across long contexts, supporting deep analysis of extensive documents and long-form content. Its agentic capabilities enable autonomous planning, execution, verification, and continuous iteration throughout complex workflows.

Total Context

1M

Max Output

128K

Released

Aug 3, 2026

Input$2 / M tokens
Output$6 / M tokens
Cache Read$0.25 / M tokens

Moonshot

Kimi K3

kimi-k3

Kimi K3 is Moonshot AI’s flagship model for long-context coding, complex reasoning, and knowledge work. With native vision and a 1M-token context window, the Kimi K3 API is designed for repository analysis, document research, multi-step agent workflows, and production AI applications. Use TokenHub to review Kimi K3 price, API access, model capabilities, and benchmark results before choosing it for your workload.

Total Context

1M

Max Output

131.1K

Released

Jul 16, 2026

Input$2.8571 / M tokens
Output$14.2857 / M tokens
Cache Read$0.2857 / M tokens

DeepSeek

DeepSeek V4 Flash Vision Exp

deepseek-v4-flash-vision-exp

DeepSeek V4 Flash Vision Exp is an experimental vision-enabled variant of DeepSeek V4 Flash 0731, introducing image understanding capabilities while preserving the base model’s text performance, including agent capabilities, reasoning, and world knowledge. It is built on a sparse Mixture-of-Experts (MoE) architecture with 284B total parameters and 13B active parameters.

Total Context

1M

Max Output

384K

Released

Aug 21, 2026

Input$0.4286 / M tokens
Output$1.2857 / M tokens
Cache Read$0.0143 / M tokens

OpenAI

GPT-5.6 Sol

gpt-5.6-sol

GPT-5.6 Sol is the standard version of the Sol series. It is suitable for general Q&A, text processing, content drafting, summarization, lightweight coding assistance, and business automation tasks. Sol is designed for scenarios that need solid general capability while keeping response efficiency and cost in mind. It can be used for product features, chat assistants, content-generation tools, data cleaning, knowledge-base Q&A, and batch-processing tasks. For large-scale usage or rapid prototyping, Sol is a flexible starting point.

Total Context

1.1M

Max Output

128K

Released

Jul 9, 2026

Input$5 / M tokens
Output$30 / M tokens
Cache Read$0.5 / M tokens

OpenAI

GPT-5.6 Sol Pro

gpt-5.6-sol-pro

GPT-5.6 Sol Pro is the Pro version of the Sol series, designed for complex tasks and higher-quality output. It is suitable for advanced writing, precise Q&A, business reasoning, code generation, structured analysis, and multi-turn task execution. Sol Pro is a good fit for tasks that require stronger accuracy, contextual understanding, and instruction completion. It can be used for professional content production, data-analysis assistants, knowledge Q&A systems, complex agent workflows, and internal enterprise AI tools.

Total Context

1.1M

Max Output

128K

Released

Jul 9, 2026

Input$5 / M tokens
Output$30 / M tokens
Cache Read$0.5 / M tokens

OpenAI

GPT-5.6 Terra

gpt-5.6-terra

GPT-5.6 Terra is the general-purpose version of the Terra series. It is suitable for everyday conversation, text generation, summarization, information extraction, lightweight reasoning, and coding assistance. Terra offers a practical balance between general capability, response efficiency, and cost. It can be used for chat assistants, content tools, office automation, knowledge-base Q&A, and batch text-processing workflows. For fast iteration or cost-sensitive applications, Terra is a strong default option within the Terra series.

Total Context

1.1M

Max Output

128K

Released

Jul 9, 2026

Input$2.5 / M tokens
Output$15 / M tokens
Cache Read$0.275 / M tokens

OpenAI

GPT-5.6 Terra Pro

gpt-5.6-terra-pro

GPT-5.6 Terra Pro is the enhanced version of the Terra series, designed for high-quality output across complex tasks. It is suitable for advanced Q&A, in-depth content creation, business analysis, coding assistance, and multi-step reasoning. Compared with the standard version, Terra Pro is better suited for scenarios that require stronger reliability, more complete reasoning, and better instruction following. It can be a good fit for professional writing, complex information synthesis, knowledge-based assistants, automation agents, and enterprise-grade AI applications.

Total Context

1.1M

Max Output

128K

Released

Jul 9, 2026

Input$2.5 / M tokens
Output$15 / M tokens
Cache Read$0.275 / M tokens

OpenAI

GPT-5.6 Luna

gpt-5.6-luna

GPT-5.6 Luna is the general-purpose version of the Luna series, suitable for everyday AI applications. It can be used for chat Q&A, text rewriting, summarization, information extraction, simple reasoning, coding assistance, and office automation. Luna is a practical choice for prototypes, lightweight business features, batch content processing, and common user interaction scenarios. It is a good starting model when speed, simplicity, and cost control are important, while Luna Pro can be considered for tasks that require higher-quality or more stable outputs.

Total Context

1.1M

Max Output

128K

Released

Jul 9, 2026

Input$1 / M tokens
Output$6 / M tokens
Cache Read$0.1 / M tokens

OpenAI

GPT-5.6 Luna Pro

gpt-5.6-luna-pro

GPT-5.6 Luna Pro is the enhanced version of the Luna series, intended for higher-quality answers, more stable instruction following, and stronger task completion. It is suitable for complex conversations, content generation, knowledge Q&A, business analysis, coding assistance, and multi-step automation tasks. Luna Pro is well suited for product experiences where output quality and user experience matter, such as professional writing assistants, enterprise AI assistants, customer support, knowledge-base systems, and agent workflows. It is a good option when the task requires more reliability than the standard Luna model.

Total Context

1.1M

Max Output

128K

Released

Jul 9, 2026

Input$1 / M tokens
Output$6 / M tokens
Cache Read$0.1 / M tokens

DeepSeek

DeepSeek V4 Pro

deepseek-v4-pro

DeepSeek V4 Pro is a DeepSeek model for complex coding, reasoning-intensive workflows, and high-volume text generation. Use the DeepSeek V4 Pro API through TokenHub to review current API pricing, context length, maximum output, available endpoints, and benchmark results before selecting it for production workloads.

Total Context

1M

Max Output

384K

Released

Apr 24, 2026

Input$1.2857 / M tokens
Output$3.8571 / M tokens
Cache Read$0.0429 / M tokens

Z.ai

GLM-5.2

glm-5.2

GLM-5.2 is Z.ai’s flagship foundation model for long-horizon engineering tasks, agent workflows, and complex reasoning. Its usable 1M-token context window supports project-scale codebases, system documentation, and sustained multi-step execution. Use the GLM API through TokenHub to evaluate current pricing, available endpoints, context limits, and benchmark results for workflows spanning requirements analysis, repository understanding, implementation, testing, and deployment.

Total Context

1M

Max Output

131.1K

Released

Jun 13, 2026

Input$1.1429 / M tokens
Output$4 / M tokens
Cache Read$0.2857 / M tokens

DeepSeek

DeepSeek V4 Flash

deepseek-v4-flash

DeepSeek V4 Flash keeps the V4 family’s 1M-token context window but uses a lighter MoE configuration, commonly described as 284B total parameters with 13B activated parameters. The emphasis is throughput: fast inference, lower cost per call, and production workloads that still need long-context handling. It is the better fit when the task volume is high and the workload benefits from V4-style long-context architecture without always requiring the deepest reasoning tier.

Total Context

1M

Max Output

384K

Released

Apr 24, 2026

Input$0.4286 / M tokens
Output$1.2857 / M tokens
Cache Read$0.0143 / M tokens

Anthropic

Claude Opus 5

claude-opus-5

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis of charts and documents, complex office deliverables, and coordinating parallel subagents. The model maintains strong instruction following and tool use across extended tasks, while remaining effective at lower effort settings for workloads that prioritize latency and token efficiency.

Total Context

1M

Max Output

128K

Released

Jul 24, 2026

Input$5 / M tokens
Output$25 / M tokens
Cache Read$0.5 / M tokens

Anthropic

Claude Opus 5 Fast

claude-opus-5-fast

Fast-mode variant of Opus 5 - identical capabilities with higher output speed at 2x pricing relative to regular Opus 5. Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode

Total Context

1M

Max Output

128K

Released

Jul 24, 2026

Input$10 / M tokens
Output$50 / M tokens
Cache Read$1 / M tokens

xAI

Grok 4.6

grok-4.6

Grok 4.6 is xAI’s advanced reasoning model built for complex programming, long-running agent tasks, and knowledge work. It supports text and image input, a context window of up to 500k tokens, and configurable low/medium/high/xhigh reasoning intensity. With function calling, web and X search, code execution, and structured output, Grok 4.6 is optimized for cross-codebase analysis, sustained multi-step tool use, self-verification, and end-to-end engineering workflows. It is best understood as a general-purpose reasoning and agentic model—not merely a chatbot associated with the X platform.

Total Context

500K

Max Output

500K

Released

Aug 12, 2026

Input$2 / M tokens
Output$6 / M tokens
Cache Read$0.3 / M tokens

Google

Gemini 3.6 Flash

gemini-3.6-flash

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and less hedging, while reducing token use and the number of model calls needed to complete a task.

Total Context

1M

Max Output

65.5K

Released

Jul 21, 2026

Input$1.5 / M tokens
Output$7.5 / M tokens
Cache Read$0.15 / M tokens

xAI

Grok 4.5

grok-4.5

Grok 4.5 is xAI's reasoning model designed for programming, agent tasks, and knowledge work. It supports text and image input, up to 500k token context, and configurable low/medium/high reasoning intensity. Features include function calling, web and X search, code execution, and structured output, making it suitable for long-context analysis, multi-step tool usage, and end-to-end engineering workflows. It should not be seen merely as a chatbot tied to the X platform.

Total Context

500K

Max Output

500K

Released

Jul 8, 2026

Input$2 / M tokens
Output$6 / M tokens
Cache Read$0.3 / M tokens

Popular model recommendations

Start with high-signal models from the live catalog, then open a detail page to compare context, endpoints, and effective pricing.

Anthropic

Claude Fable 5.1

Claude Fable 5 is Anthropic’s Mythos-class model designed for demanding reasoning, software engineering, knowledge work, visual analysis, and scientific research. Built for long-horizon agentic tasks, it can sustain planning and execution across complex workflows, analyze large codebases and documents, and refine its work over multiple steps. Claude Fable 5 is well suited to production applications that require advanced coding, research, professional analysis, and reliable autonomous task completion.

Total Context

1M

Input Price

$10 / M tokens

View model

Z.ai

GLM-5.3-Flash

GLM-5.3-Flash is Z.ai's first natively multimodal model in the GLM-5 family, built for fast coding, visual understanding, and agent workflows. Its 320B-parameter MoE architecture activates 18B parameters per token and combines sparse with linear attention to make long-context tasks more efficient. Use the GLM API through TokenHub to evaluate the current GLM-5.3-Flash price, model ID, supported inputs, tool calling, context limits, and endpoints before deploying it for production coding or multimodal agent tasks.

Total Context

1M

Input Price

$0.1143 / M tokens

View model

Z.ai

GLM-5.3

GLM-5.3 is Z.ai’s latest flagship model built for advanced coding, software engineering, and long-horizon agent tasks. With a 1M-token context window, up to 128K output tokens, strong reasoning, and tool-use capabilities, it is ideal for Coding Agents, complex development workflows, and multi-step automation.

Total Context

1M

Input Price

$1.1429 / M tokens

View model

Tencent

Hy4 preview

Hy4 preview is Tencent Hunyuan’s next-generation open-source Mixture-of-Experts flagship model, built for real-world productivity across coding, office work, scientific research, long-context reasoning, and complex agent workflows. Also known as Tencent Hy4 preview or Hunyuan 4 preview, it has 770B total parameters, activates 49B per token, and supports a context window of over 1M tokens. It excels at understanding large codebases and documents, multi-step planning, tool use, debugging, validation, and sustained task execution.

Total Context

1M

Input Price

$0.8571 / M tokens

View model

Alibaba

Qwen3.8 Max

Qwen3.8-Max is a 2.4T-parameter MoE flagship model for advanced reasoning, coding, enterprise productivity, and agent workflows. It is designed to handle complex, long-running tasks from software development to professional workflows across domains such as law, finance, and design. Use the Qwen API through TokenHub to review the current Qwen3.8 Max price, model ID, context and output limits, supported modalities and endpoints, and benchmark results before deploying it in production.Powered by native multimodal intelligence, Qwen3.8-Max understands text, images, and videos across long contexts, supporting deep analysis of extensive documents and long-form content. Its agentic capabilities enable autonomous planning, execution, verification, and continuous iteration throughout complex workflows.

Total Context

1M

Input Price

$2 / M tokens

View model

Moonshot

Kimi K3

Kimi K3 is Moonshot AI’s flagship model for long-context coding, complex reasoning, and knowledge work. With native vision and a 1M-token context window, the Kimi K3 API is designed for repository analysis, document research, multi-step agent workflows, and production AI applications. Use TokenHub to review Kimi K3 price, API access, model capabilities, and benchmark results before choosing it for your workload.

Total Context

1M

Input Price

$2.8571 / M tokens

View model

Model Comparison

Quick comparison against selected catalog neighbors.

Model catalog FAQ

A quick guide for choosing, comparing, and using models from the TokenHub catalog.

How should I choose a model from this list?

+

Start with your workload. Use the filters to narrow by provider, tags, endpoint type, and billing group, then compare context size, output limit, modalities, and input or output pricing.

What does effective price mean?

+

Effective price applies the active billing group ratio to the model pricing data. It helps you estimate the real input, output, or per-request cost for the group you are using.

Can I use these models through API endpoints?

+

Yes. Open a model detail page to see the supported endpoint types and documentation links. Availability can differ by model, provider, and current routing configuration.

Why do context window and max output matter?

+

The context window controls how much prompt and conversation history a model can read. Max output controls how much text it can generate in one response, which matters for long-form writing, coding, and document tasks.