Explore AI model pricing, capabilities, endpoints, and vendor coverage from one production catalog.
AN
Anthropic
Claude Fable 5.1
claude-fable-5.1
ImagePDFText
Text
Claude Fable 5 is Anthropic’s Mythos-class model designed for demanding reasoning, software engineering, knowledge work, visual analysis, and scientific research. Built for long-horizon agentic tasks, it can sustain planning and execution across complex workflows, analyze large codebases and documents, and refine its work over multiple steps. Claude Fable 5 is well suited to production applications that require advanced coding, research, professional analysis, and reliable autonomous task completion.
Total Context
1M
Max Output
128K
Released
Sep 1, 2026
Input$10 / M tokens
Output$50 / M tokens
Cache Read$0.25 / M tokens
Z.
Z.ai
GLM-5.3-Flash
glm-5.3-flash
ImagePDFTextVideo
Text
GLM-5.3-Flash is Z.ai's first natively multimodal model in the GLM-5 family, built for fast coding, visual understanding, and agent workflows. Its 320B-parameter MoE architecture activates 18B parameters per token and combines sparse with linear attention to make long-context tasks more efficient. Use the GLM API through TokenHub to evaluate the current GLM-5.3-Flash price, model ID, supported inputs, tool calling, context limits, and endpoints before deploying it for production coding or multimodal agent tasks.
Total Context
1M
Max Output
131.1K
Released
Aug 26, 2026
Input$0.1143 / M tokens
Output$0.4 / M tokens
Cache Read$0.0329 / M tokens
Z.
Z.ai
GLM-5.3
glm-5.3
Text
Text
GLM-5.3 is Z.ai’s latest flagship model built for advanced coding, software engineering, and long-horizon agent tasks. With a 1M-token context window, up to 128K output tokens, strong reasoning, and tool-use capabilities, it is ideal for Coding Agents, complex development workflows, and multi-step automation.
Total Context
1M
Max Output
128K
Released
Aug 19, 2026
Input$1.1429 / M tokens
Output$4 / M tokens
Cache Read$0.2857 / M tokens
TE
Tencent
Hy4 preview
hy4-preview
Text
Text
Hy4 preview is Tencent Hunyuan’s next-generation open-source Mixture-of-Experts flagship model, built for real-world productivity across coding, office work, scientific research, long-context reasoning, and complex agent workflows. Also known as Tencent Hy4 preview or Hunyuan 4 preview, it has 770B total parameters, activates 49B per token, and supports a context window of over 1M tokens. It excels at understanding large codebases and documents, multi-step planning, tool use, debugging, validation, and sustained task execution.
Total Context
1M
Max Output
64K
Released
Aug 28, 2026
Input$0.8571 / M tokens
Output$2.5714 / M tokens
Cache Read$0.0429 / M tokens
AL
Alibaba
Qwen3.8 Max
qwen3.8-max
ImageTextVideo
Text
Qwen3.8-Max is a 2.4T-parameter MoE flagship model for advanced reasoning, coding, enterprise productivity, and agent workflows. It is designed to handle complex, long-running tasks from software development to professional workflows across domains such as law, finance, and design. Use the Qwen API through TokenHub to review the current Qwen3.8 Max price, model ID, context and output limits, supported modalities and endpoints, and benchmark results before deploying it in production.Powered by native multimodal intelligence, Qwen3.8-Max understands text, images, and videos across long contexts, supporting deep analysis of extensive documents and long-form content. Its agentic capabilities enable autonomous planning, execution, verification, and continuous iteration throughout complex workflows.
Total Context
1M
Max Output
128K
Released
Aug 3, 2026
Input$2 / M tokens
Output$6 / M tokens
Cache Read$0.25 / M tokens
MO
Moonshot
Kimi K3
kimi-k3
ImageText
Text
Kimi K3 is Moonshot AI’s flagship model for long-context coding, complex reasoning, and knowledge work. With native vision and a 1M-token context window, the Kimi K3 API is designed for repository analysis, document research, multi-step agent workflows, and production AI applications. Use TokenHub to review Kimi K3 price, API access, model capabilities, and benchmark results before choosing it for your workload.
Total Context
1M
Max Output
131.1K
Released
Jul 16, 2026
Input$2.8571 / M tokens
Output$14.2857 / M tokens
Cache Read$0.2857 / M tokens
DE
DeepSeek
DeepSeek V4 Flash Vision Exp
deepseek-v4-flash-vision-exp
ImageText
Text
DeepSeek V4 Flash Vision Exp is an experimental vision-enabled variant of DeepSeek V4 Flash 0731, introducing image understanding capabilities while preserving the base model’s text performance, including agent capabilities, reasoning, and world knowledge. It is built on a sparse Mixture-of-Experts (MoE) architecture with 284B total parameters and 13B active parameters.
Total Context
1M
Max Output
384K
Released
Aug 21, 2026
Input$0.4286 / M tokens
Output$1.2857 / M tokens
Cache Read$0.0143 / M tokens
OP
OpenAI
GPT-5.6 Sol
gpt-5.6-sol
ImagePDFText
Text
GPT-5.6 Sol is the standard version of the Sol series. It is suitable for general Q&A, text processing, content drafting, summarization, lightweight coding assistance, and business automation tasks.
Sol is designed for scenarios that need solid general capability while keeping response efficiency and cost in mind. It can be used for product features, chat assistants, content-generation tools, data cleaning, knowledge-base Q&A, and batch-processing tasks. For large-scale usage or rapid prototyping, Sol is a flexible starting point.
Total Context
1.1M
Max Output
128K
Released
Jul 9, 2026
Input$5 / M tokens
Output$30 / M tokens
Cache Read$0.5 / M tokens
OP
OpenAI
GPT-5.6 Sol Pro
gpt-5.6-sol-pro
ImagePDFText
Text
GPT-5.6 Sol Pro is the Pro version of the Sol series, designed for complex tasks and higher-quality output. It is suitable for advanced writing, precise Q&A, business reasoning, code generation, structured analysis, and multi-turn task execution.
Sol Pro is a good fit for tasks that require stronger accuracy, contextual understanding, and instruction completion. It can be used for professional content production, data-analysis assistants, knowledge Q&A systems, complex agent workflows, and internal enterprise AI tools.
Total Context
1.1M
Max Output
128K
Released
Jul 9, 2026
Input$5 / M tokens
Output$30 / M tokens
Cache Read$0.5 / M tokens
OP
OpenAI
GPT-5.6 Terra
gpt-5.6-terra
ImagePDFText
Text
GPT-5.6 Terra is the general-purpose version of the Terra series. It is suitable for everyday conversation, text generation, summarization, information extraction, lightweight reasoning, and coding assistance.
Terra offers a practical balance between general capability, response efficiency, and cost. It can be used for chat assistants, content tools, office automation, knowledge-base Q&A, and batch text-processing workflows. For fast iteration or cost-sensitive applications, Terra is a strong default option within the Terra series.
Total Context
1.1M
Max Output
128K
Released
Jul 9, 2026
Input$2.5 / M tokens
Output$15 / M tokens
Cache Read$0.275 / M tokens
OP
OpenAI
GPT-5.6 Terra Pro
gpt-5.6-terra-pro
ImagePDFText
Text
GPT-5.6 Terra Pro is the enhanced version of the Terra series, designed for high-quality output across complex tasks. It is suitable for advanced Q&A, in-depth content creation, business analysis, coding assistance, and multi-step reasoning. Compared with the standard version, Terra Pro is better suited for scenarios that require stronger reliability, more complete reasoning, and better instruction following. It can be a good fit for professional writing, complex information synthesis, knowledge-based assistants, automation agents, and enterprise-grade AI applications.
Total Context
1.1M
Max Output
128K
Released
Jul 9, 2026
Input$2.5 / M tokens
Output$15 / M tokens
Cache Read$0.275 / M tokens
OP
OpenAI
GPT-5.6 Luna
gpt-5.6-luna
ImagePDFText
Text
GPT-5.6 Luna is the general-purpose version of the Luna series, suitable for everyday AI applications. It can be used for chat Q&A, text rewriting, summarization, information extraction, simple reasoning, coding assistance, and office automation.
Luna is a practical choice for prototypes, lightweight business features, batch content processing, and common user interaction scenarios. It is a good starting model when speed, simplicity, and cost control are important, while Luna Pro can be considered for tasks that require higher-quality or more stable outputs.
Total Context
1.1M
Max Output
128K
Released
Jul 9, 2026
Input$1 / M tokens
Output$6 / M tokens
Cache Read$0.1 / M tokens
OP
OpenAI
GPT-5.6 Luna Pro
gpt-5.6-luna-pro
ImagePDFText
Text
GPT-5.6 Luna Pro is the enhanced version of the Luna series, intended for higher-quality answers, more stable instruction following, and stronger task completion. It is suitable for complex conversations, content generation, knowledge Q&A, business analysis, coding assistance, and multi-step automation tasks.
Luna Pro is well suited for product experiences where output quality and user experience matter, such as professional writing assistants, enterprise AI assistants, customer support, knowledge-base systems, and agent workflows. It is a good option when the task requires more reliability than the standard Luna model.
Total Context
1.1M
Max Output
128K
Released
Jul 9, 2026
Input$1 / M tokens
Output$6 / M tokens
Cache Read$0.1 / M tokens
DE
DeepSeek
DeepSeek V4 Pro
deepseek-v4-pro
Text
Text
DeepSeek V4 Pro is a DeepSeek model for complex coding, reasoning-intensive workflows, and high-volume text generation. Use the DeepSeek V4 Pro API through TokenHub to review current API pricing, context length, maximum output, available endpoints, and benchmark results before selecting it for production workloads.
Total Context
1M
Max Output
384K
Released
Apr 24, 2026
Input$1.2857 / M tokens
Output$3.8571 / M tokens
Cache Read$0.0429 / M tokens
Z.
Z.ai
GLM-5.2
glm-5.2
Text
Text
GLM-5.2 is Z.ai’s flagship foundation model for long-horizon engineering tasks, agent workflows, and complex reasoning. Its usable 1M-token context window supports project-scale codebases, system documentation, and sustained multi-step execution. Use the GLM API through TokenHub to evaluate current pricing, available endpoints, context limits, and benchmark results for workflows spanning requirements analysis, repository understanding, implementation, testing, and deployment.
Total Context
1M
Max Output
131.1K
Released
Jun 13, 2026
Input$1.1429 / M tokens
Output$4 / M tokens
Cache Read$0.2857 / M tokens
DE
DeepSeek
DeepSeek V4 Flash
deepseek-v4-flash
Text
Text
DeepSeek V4 Flash keeps the V4 family’s 1M-token context window but uses a lighter MoE configuration, commonly described as 284B total parameters with 13B activated parameters. The emphasis is throughput: fast inference, lower cost per call, and production workloads that still need long-context handling. It is the better fit when the task volume is high and the workload benefits from V4-style long-context architecture without always requiring the deepest reasoning tier.
Total Context
1M
Max Output
384K
Released
Apr 24, 2026
Input$0.4286 / M tokens
Output$1.2857 / M tokens
Cache Read$0.0143 / M tokens
AN
Anthropic
Claude Opus 5
claude-opus-5
ImagePDFText
Text
Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis of charts and documents, complex office deliverables, and coordinating parallel subagents.
The model maintains strong instruction following and tool use across extended tasks, while remaining effective at lower effort settings for workloads that prioritize latency and token efficiency.
Total Context
1M
Max Output
128K
Released
Jul 24, 2026
Input$5 / M tokens
Output$25 / M tokens
Cache Read$0.5 / M tokens
AN
Anthropic
Claude Opus 5 Fast
claude-opus-5-fast
ImagePDFText
Text
Fast-mode variant of Opus 5 - identical capabilities with higher output speed at 2x pricing relative to regular Opus 5.
Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode
Total Context
1M
Max Output
128K
Released
Jul 24, 2026
Input$10 / M tokens
Output$50 / M tokens
Cache Read$1 / M tokens
xAI
Grok 4.6
grok-4.6
ImageText
Text
Grok 4.6 is xAI’s advanced reasoning model built for complex programming, long-running agent tasks, and knowledge work. It supports text and image input, a context window of up to 500k tokens, and configurable low/medium/high/xhigh reasoning intensity. With function calling, web and X search, code execution, and structured output, Grok 4.6 is optimized for cross-codebase analysis, sustained multi-step tool use, self-verification, and end-to-end engineering workflows. It is best understood as a general-purpose reasoning and agentic model—not merely a chatbot associated with the X platform.
Total Context
500K
Max Output
500K
Released
Aug 12, 2026
Input$2 / M tokens
Output$6 / M tokens
Cache Read$0.3 / M tokens
GO
Google
Gemini 3.6 Flash
gemini-3.6-flash
AudioImagePDFTextVideo
Text
Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and less hedging, while reducing token use and the number of model calls needed to complete a task.
Total Context
1M
Max Output
65.5K
Released
Jul 21, 2026
Input$1.5 / M tokens
Output$7.5 / M tokens
Cache Read$0.15 / M tokens
xAI
Grok 4.5
grok-4.5
ImageText
Text
Grok 4.5 is xAI's reasoning model designed for programming, agent tasks, and knowledge work. It supports text and image input, up to 500k token context, and configurable low/medium/high reasoning intensity. Features include function calling, web and X search, code execution, and structured output, making it suitable for long-context analysis, multi-step tool usage, and end-to-end engineering workflows. It should not be seen merely as a chatbot tied to the X platform.
Total Context
500K
Max Output
500K
Released
Jul 8, 2026
Input$2 / M tokens
Output$6 / M tokens
Cache Read$0.3 / M tokens
Popular model recommendations
Start with high-signal models from the live catalog, then open a detail page to compare context, endpoints, and effective pricing.
A quick guide for choosing, comparing, and using models from the TokenHub catalog.
How should I choose a model from this list?
+
Start with your workload. Use the filters to narrow by provider, tags, endpoint type, and billing group, then compare context size, output limit, modalities, and input or output pricing.
What does effective price mean?
+
Effective price applies the active billing group ratio to the model pricing data. It helps you estimate the real input, output, or per-request cost for the group you are using.
Can I use these models through API endpoints?
+
Yes. Open a model detail page to see the supported endpoint types and documentation links. Availability can differ by model, provider, and current routing configuration.
Why do context window and max output matter?
+
The context window controls how much prompt and conversation history a model can read. Max output controls how much text it can generate in one response, which matters for long-form writing, coding, and document tasks.