Explore AI model pricing, capabilities, endpoints, and vendor coverage from one production catalog.
O
OpenAI
GPT-6.1 Sol
gpt-6.1-sol
ImagePDFText
Text
GPT-6.1 Sol is an advanced reasoning model built for complex problem solving, coding, and agent-based applications. It is designed to deliver strong reasoning performance while maintaining a practical balance between intelligence, speed, and cost.
Unlike traditional chat models that focus mainly on generating responses, GPT-6.1 Sol is optimized for tasks that require deeper thinking, including multi-step reasoning, code generation, debugging, data analysis, and following complex instructions.
GPT-6.1 Sol supports reliable reasoning workflows, structured outputs, and tool-based applications, making it suitable for developers building AI assistants, coding agents, automation systems, and production AI applications.
Total Context
1.1M
Max Output
128K
Released
Sep 29, 2026
Input$2 / M tokens
Output$10 / M tokens
Cache Read$0.1 / M tokens
O
OpenAI
GPT-6.1 Sol Pro
gpt-6.1-sol-pro
ImagePDFText
Text
GPT-6.1 Sol Pro is an enhanced reasoning version of GPT-6.1 Sol. It uses the same underlying model as GPT-6.1 Sol but enables reasoning.mode=pro, allowing the model to spend significantly more computation on each request.
The additional reasoning budget helps GPT-6.1 Sol Pro handle harder problems that require deeper analysis, longer planning, and higher accuracy. It is designed for tasks where getting the best possible answer is more important than response speed or cost efficiency.
Compared with GPT-6.1 Sol, GPT-6.1 Sol Pro consumes significantly more reasoning tokens per request, resulting in higher usage costs and longer response times.
Total Context
1.1M
Max Output
128K
Released
Sep 29, 2026
Input$2 / M tokens
Output$10 / M tokens
Cache Read$0.1 / M tokens
A
Anthropic
Claude Sonnet 5.5
claude-sonnet-5.5
ImagePDFText
Text
Claude Sonnet 5.5 is a next-generation AI model designed for advanced reasoning, coding, analysis, and complex agentic workflows. It delivers strong performance across software development, long-context understanding, multimodal tasks, and professional knowledge work, offering a balanced combination of intelligence, speed, and efficiency. With enhanced reasoning capabilities and improved instruction following, Claude Sonnet 5.5 excels at writing and debugging code, analyzing complex documents, solving multi-step problems, and powering AI agents that require reliable decision-making and tool usage.
Total Context
1M
Max Output
128K
Released
Sep 28, 2026
Input$2 / M tokens
Output$10 / M tokens
Cache Read$0.2 / M tokens
xAI
Grok 4.7
grok-4.7
ImagePDFText
Text
Grok 4.7 is xAI's latest generation model, built for demanding coding tasks, complex reasoning, and knowledge-intensive applications. It improves long-horizon problem solving, self-verification, and workflow execution, making it a strong fit for developers building intelligent assistants and AI-powered products.
Grok 4.7 delivers strong performance across programming, analysis, and general-purpose AI tasks while maintaining competitive efficiency for production deployments.
Total Context
500K
Max Output
500K
Released
Sep 21, 2026
Input$1.6 / M tokens
Output$4.8 / M tokens
Cache Read$0.4 / M tokens
A
Anthropic
Claude Opus 5.5
claude-opus-5.5
ImagePDFText
Text
Claude Opus 5.5 is Anthropic's flagship intelligence model, designed for demanding workloads that require deep reasoning, reliability, and long-running task execution. It delivers strong performance in software engineering, agentic workflows, professional analysis, and knowledge-intensive tasks.
With improved efficiency and enhanced safeguards, Claude Opus 5.5 is optimized for developers building sophisticated AI applications that require consistent performance across complex tasks.
Total Context
1M
Max Output
128K
Released
Sep 22, 2026
Input$4 / M tokens
Output$20 / M tokens
Cache Read$0.2 / M tokens
O
OpenAI
GPT-6 Sol
gpt-6-sol
ImagePDFText
Text
GPT-6 Sol is a high-performance model in OpenAI's GPT-6 series, designed for complex reasoning tasks, software development, and professional AI workflows. It offers enhanced intelligence for multi-step problem solving, advanced coding, data analysis, and agent-based applications.
GPT-6 Sol combines strong reasoning capability with efficient inference, making it suitable for developers who need frontier-level performance for production AI systems.
Total Context
1.1M
Max Output
128K
Released
Sep 22, 2026
Input$2 / M tokens
Output$10 / M tokens
Cache Read$0.2 / M tokens
O
OpenAI
GPT-6 Sol Pro
gpt-6-sol-pro
ImagePDFText
Text
GPT-6 Sol Pro is an enhanced reasoning variant of GPT-6 Sol, designed for tasks that require deeper analysis, higher accuracy, and more reliable multi-step execution.
With increased reasoning capability, GPT-6 Sol Pro delivers stronger performance in software engineering, complex workflows, data analysis, and AI agent applications. It is optimized for developers building advanced AI systems that require consistent results on challenging tasks.
GPT-6 Sol Pro supports long-context processing, tool calling, and structured outputs, making it suitable for production-grade applications that require powerful reasoning and automation capabilities.
Total Context
1.1M
Max Output
128K
Released
Sep 22, 2026
Input$2 / M tokens
Output$10 / M tokens
Cache Read$0.2 / M tokens
O
OpenAI
GPT-6 Luna
gpt-6-luna
ImagePDFText
Text
GPT-6 Luna is a compact, efficient model in OpenAI's GPT-6 family, delivering strong language intelligence at significantly improved cost efficiency. It offers reliable performance for everyday AI applications such as conversational AI, content generation, coding assistance, and automation workflows. With a large context window and optimized inference efficiency, GPT-6 Luna is ideal for developers building scalable AI products that need high throughput, low latency, and predictable operating costs.
Total Context
1.1M
Max Output
128K
Released
Sep 22, 2026
Input$0.1 / M tokens
Output$0.5 / M tokens
Cache Read$0.01 / M tokens
O
OpenAI
GPT-6 Luna Pro
gpt-6-luna-pro
ImagePDFText
Text
GPT-6 Luna Pro is a reasoning-optimized version of GPT-6 Luna, designed to provide stronger problem-solving ability while maintaining the efficiency and affordability of the Luna family.
It improves performance on complex instructions, coding tasks, analytical workflows, and lightweight agent applications compared with standard fast inference models. GPT-6 Luna Pro is ideal for developers who need higher-quality responses while maintaining scalability for production workloads.
With support for long-context processing, function calling, and structured outputs, GPT-6 Luna Pro provides a balanced solution between intelligence, speed, and cost efficiency.
Total Context
1.1M
Max Output
128K
Released
Sep 22, 2026
Input$0.1 / M tokens
Output$0.5 / M tokens
Cache Read$0.01 / M tokens
A
Alibaba
Qwen3.8 Omni Flash
qwen3.8-omni-flash
AudioImageTextVideo
Text
Qwen3.8 Omni Flash is a next-generation omni-modal model from Alibaba’s Qwen family, designed to understand and process text, images, audio, and video within a unified AI system. It combines strong language intelligence with advanced multimodal understanding capabilities for real-world AI applications.
With support for up to 1M-token context windows, Qwen3.8 Omni Flash is optimized for long-document analysis, multimedia understanding, and agentic workflows. It can analyze complex audio and video content, perform reasoning across multiple modalities, and support tool calling for building intelligent AI applications.
Designed for high-efficiency deployment, Qwen3.8 Omni Flash provides a strong balance between multimodal intelligence, scalability, and cost efficiency, making it suitable for developers building next-generation AI assistants, content analysis systems, and automation agents.
Total Context
1M
Max Output
131.1K
Released
Sep 17, 2026
Input$0.1143 / M tokens
Output$0.3857 / M tokens
Cache Read$0.0143 / M tokens
T
TypeSafe
Jev 1.13
jev-1.13
Text
Decisions
Jev 1.13 is TypeSafe AI's first System One decision model, built for structured AI judgment rather than text generation. It analyzes provided context and returns structured choices, scores, and confidence-based decisions. Designed for classification, routing, evaluation, and automation workflows, Jev delivers fast, predictable AI decisions with low latency and efficient token usage.
Total Context
32K
Max Output
N/A
Released
Sep 18, 2026
Input$0.042 / M tokens
Output$0 / M tokens
Z
Z.ai
GLM-5.3-FlashX
glm-5.3-flashx
ImageTextVideo
Text
GLM-5.3-FlashX is a high-speed multimodal model from Z.ai, designed for fast, responsive inference while retaining the core capabilities of GLM-5.3-Flash. It supports text, image, video, and file inputs, with a context window of up to 1 million tokens.
The model performs well across coding, agentic workflows, visual understanding, long-context tasks, and complex reasoning, making it a strong choice for real-time AI applications that require both speed and intelligence.
Total Context
1M
Max Output
131.1K
Released
Sep 18, 2026
Input$0.2857 / M tokens
Output$1 / M tokens
Cache Read$0.0814 / M tokens
D
DeepSeek
DeepSeek V4.1 Flash
deepseek-v4.1-flash
ImageText
Text
DeepSeek-V4.1-Flash is a fast and efficient multimodal AI model with advanced reasoning, coding, and agent capabilities. Supporting up to 1M tokens of context and native vision understanding, it delivers frontier-level intelligence with lower inference costs, making it ideal for AI agents, automation, coding assistants, and scalable applications.
Total Context
1M
Max Output
384K
Released
Sep 10, 2026
Input$0.2857 / M tokens
Output$1.1429 / M tokens
Cache Read$0.0057 / M tokens
O
OpenAI
GPT-6 Astra
gpt-6-astra
ImagePDFText
Text
GPT-6 Astra is OpenAI’s next-generation flagship model, designed for complex reasoning, software engineering, AI agents, research, and demanding end-to-end professional workflows. It excels at multi-step tasks across code, browsers, and professional software, delivering stronger performance in problem solving, tool use, computer use, and long-context understanding.
GPT-6 Astra supports a context window of up to 1.05M tokens and a maximum output of 128K tokens, with multiple reasoning effort levels including Low, Medium, High, XHigh, and Max, allowing developers to balance reasoning capability, efficiency, and cost based on task complexity.
It is well suited for advanced AI agents, complex software development, automated workflows, deep research, long-document processing, data analysis, and enterprise AI applications.
Total Context
1.1M
Max Output
128K
Released
Sep 4, 2026
Input$10 / M tokens
Output$50 / M tokens
Cache Read$1 / M tokens
A
Anthropic
Claude Fable 5.1
claude-fable-5.1
ImagePDFText
Text
Claude Fable 5 is Anthropic’s Mythos-class model designed for demanding reasoning, software engineering, knowledge work, visual analysis, and scientific research. Built for long-horizon agentic tasks, it can sustain planning and execution across complex workflows, analyze large codebases and documents, and refine its work over multiple steps. Claude Fable 5 is well suited to production applications that require advanced coding, research, professional analysis, and reliable autonomous task completion.
Total Context
1M
Max Output
128K
Released
Sep 1, 2026
Input$10 / M tokens
Output$50 / M tokens
Cache Read$0.25 / M tokens
Z
Z.ai
GLM-5.3-Flash
glm-5.3-flash
ImagePDFTextVideo
Text
GLM-5.3-Flash is Z.ai's first natively multimodal model in the GLM-5 family, built for fast coding, visual understanding, and agent workflows. Its 320B-parameter MoE architecture activates 18B parameters per token and combines sparse with linear attention to make long-context tasks more efficient. Use the GLM API through TokenHub to evaluate the current GLM-5.3-Flash price, model ID, supported inputs, tool calling, context limits, and endpoints before deploying it for production coding or multimodal agent tasks.
Total Context
1M
Max Output
131.1K
Released
Aug 26, 2026
Input$0.1143 / M tokens
Output$0.4 / M tokens
Cache Read$0.0329 / M tokens
Z
Z.ai
GLM-5.3
glm-5.3
Text
Text
GLM-5.3 is Z.ai’s latest flagship model built for advanced coding, software engineering, and long-horizon agent tasks. With a 1M-token context window, up to 128K output tokens, strong reasoning, and tool-use capabilities, it is ideal for Coding Agents, complex development workflows, and multi-step automation.
Total Context
1M
Max Output
128K
Released
Aug 19, 2026
Input$1.1429 / M tokens
Output$4 / M tokens
Cache Read$0.2857 / M tokens
T
Tencent
Hy4 preview
hy4-preview
Text
Text
Hy4 preview is Tencent Hunyuan’s next-generation open-source Mixture-of-Experts flagship model, built for real-world productivity across coding, office work, scientific research, long-context reasoning, and complex agent workflows. Also known as Tencent Hy4 preview or Hunyuan 4 preview, it has 770B total parameters, activates 49B per token, and supports a context window of over 1M tokens. It excels at understanding large codebases and documents, multi-step planning, tool use, debugging, validation, and sustained task execution.
Total Context
1M
Max Output
64K
Released
Aug 28, 2026
Input$0.8571 / M tokens
Output$2.5714 / M tokens
Cache Read$0.0429 / M tokens
A
Alibaba
Qwen3.8 Max
qwen3.8-max
ImageTextVideo
Text
Qwen3.8-Max is a 2.4T-parameter MoE flagship model for advanced reasoning, coding, enterprise productivity, and agent workflows. It is designed to handle complex, long-running tasks from software development to professional workflows across domains such as law, finance, and design. Use the Qwen API through TokenHub to review the current Qwen3.8 Max price, model ID, context and output limits, supported modalities and endpoints, and benchmark results before deploying it in production.Powered by native multimodal intelligence, Qwen3.8-Max understands text, images, and videos across long contexts, supporting deep analysis of extensive documents and long-form content. Its agentic capabilities enable autonomous planning, execution, verification, and continuous iteration throughout complex workflows.
Total Context
1M
Max Output
128K
Released
Aug 3, 2026
Input$2 / M tokens
Output$6 / M tokens
Cache Read$0.25 / M tokens
M
Moonshot
Kimi K3
kimi-k3
ImageText
Text
Kimi K3 is Moonshot AI’s flagship model for long-context coding, complex reasoning, and knowledge work. With native vision and a 1M-token context window, the Kimi K3 API is designed for repository analysis, document research, multi-step agent workflows, and production AI applications. Use TokenHub to review Kimi K3 price, API access, model capabilities, and benchmark results before choosing it for your workload.
Total Context
1M
Max Output
131.1K
Released
Jul 16, 2026
Input$2.8571 / M tokens
Output$14.2857 / M tokens
Cache Read$0.2857 / M tokens
D
DeepSeek
DeepSeek V4 Flash Vision Exp
deepseek-v4-flash-vision-exp
ImageText
Text
DeepSeek V4 Flash Vision Exp is an experimental vision-enabled variant of DeepSeek V4 Flash 0731, introducing image understanding capabilities while preserving the base model’s text performance, including agent capabilities, reasoning, and world knowledge. It is built on a sparse Mixture-of-Experts (MoE) architecture with 284B total parameters and 13B active parameters.
Total Context
1M
Max Output
384K
Released
Aug 21, 2026
Input$0.2857 / M tokens
Output$1.1429 / M tokens
Cache Read$0.0057 / M tokens
Popular model recommendations
Start with high-signal models from the live catalog, then open a detail page to compare context, endpoints, and effective pricing.
A quick guide for choosing, comparing, and using models from the TokenHub catalog.
How should I choose a model from this list?
+
Start with your workload. Use the filters to narrow by provider, tags, endpoint type, and billing group, then compare context size, output limit, modalities, and input or output pricing.
What does effective price mean?
+
Effective price applies the active billing group ratio to the model pricing data. It helps you estimate the real input, output, or per-request cost for the group you are using.
Can I use these models through API endpoints?
+
Yes. Open a model detail page to see the supported endpoint types and documentation links. Availability can differ by model, provider, and current routing configuration.
Why do context window and max output matter?
+
The context window controls how much prompt and conversation history a model can read. Max output controls how much text it can generate in one response, which matters for long-form writing, coding, and document tasks.