Models

Explore AI model pricing, capabilities, endpoints, and vendor coverage from one production catalog.

OpenAI

GPT-6.1 Sol

gpt-6.1-sol

GPT-6.1 Sol is an advanced reasoning model built for complex problem solving, coding, and agent-based applications. It is designed to deliver strong reasoning performance while maintaining a practical balance between intelligence, speed, and cost. Unlike traditional chat models that focus mainly on generating responses, GPT-6.1 Sol is optimized for tasks that require deeper thinking, including multi-step reasoning, code generation, debugging, data analysis, and following complex instructions. GPT-6.1 Sol supports reliable reasoning workflows, structured outputs, and tool-based applications, making it suitable for developers building AI assistants, coding agents, automation systems, and production AI applications.

Total Context

1.1M

Max Output

128K

Released

Sep 29, 2026

Input$2 / M tokens
Output$10 / M tokens
Cache Read$0.1 / M tokens

OpenAI

GPT-6.1 Sol Pro

gpt-6.1-sol-pro

GPT-6.1 Sol Pro is an enhanced reasoning version of GPT-6.1 Sol. It uses the same underlying model as GPT-6.1 Sol but enables reasoning.mode=pro, allowing the model to spend significantly more computation on each request. The additional reasoning budget helps GPT-6.1 Sol Pro handle harder problems that require deeper analysis, longer planning, and higher accuracy. It is designed for tasks where getting the best possible answer is more important than response speed or cost efficiency. Compared with GPT-6.1 Sol, GPT-6.1 Sol Pro consumes significantly more reasoning tokens per request, resulting in higher usage costs and longer response times.

Total Context

1.1M

Max Output

128K

Released

Sep 29, 2026

Input$2 / M tokens
Output$10 / M tokens
Cache Read$0.1 / M tokens

Anthropic

Claude Sonnet 5.5

claude-sonnet-5.5

Claude Sonnet 5.5 is a next-generation AI model designed for advanced reasoning, coding, analysis, and complex agentic workflows. It delivers strong performance across software development, long-context understanding, multimodal tasks, and professional knowledge work, offering a balanced combination of intelligence, speed, and efficiency. With enhanced reasoning capabilities and improved instruction following, Claude Sonnet 5.5 excels at writing and debugging code, analyzing complex documents, solving multi-step problems, and powering AI agents that require reliable decision-making and tool usage.

Total Context

1M

Max Output

128K

Released

Sep 28, 2026

Input$2 / M tokens
Output$10 / M tokens
Cache Read$0.2 / M tokens

xAI

Grok 4.7

grok-4.7

Grok 4.7 is xAI's latest generation model, built for demanding coding tasks, complex reasoning, and knowledge-intensive applications. It improves long-horizon problem solving, self-verification, and workflow execution, making it a strong fit for developers building intelligent assistants and AI-powered products. Grok 4.7 delivers strong performance across programming, analysis, and general-purpose AI tasks while maintaining competitive efficiency for production deployments.

Total Context

500K

Max Output

500K

Released

Sep 21, 2026

Input$1.6 / M tokens
Output$4.8 / M tokens
Cache Read$0.4 / M tokens

Anthropic

Claude Opus 5.5

claude-opus-5.5

Claude Opus 5.5 is Anthropic's flagship intelligence model, designed for demanding workloads that require deep reasoning, reliability, and long-running task execution. It delivers strong performance in software engineering, agentic workflows, professional analysis, and knowledge-intensive tasks. With improved efficiency and enhanced safeguards, Claude Opus 5.5 is optimized for developers building sophisticated AI applications that require consistent performance across complex tasks.

Total Context

1M

Max Output

128K

Released

Sep 22, 2026

Input$4 / M tokens
Output$20 / M tokens
Cache Read$0.2 / M tokens

OpenAI

GPT-6 Sol

gpt-6-sol

GPT-6 Sol is a high-performance model in OpenAI's GPT-6 series, designed for complex reasoning tasks, software development, and professional AI workflows. It offers enhanced intelligence for multi-step problem solving, advanced coding, data analysis, and agent-based applications. GPT-6 Sol combines strong reasoning capability with efficient inference, making it suitable for developers who need frontier-level performance for production AI systems.

Total Context

1.1M

Max Output

128K

Released

Sep 22, 2026

Input$2 / M tokens
Output$10 / M tokens
Cache Read$0.2 / M tokens

OpenAI

GPT-6 Sol Pro

gpt-6-sol-pro

GPT-6 Sol Pro is an enhanced reasoning variant of GPT-6 Sol, designed for tasks that require deeper analysis, higher accuracy, and more reliable multi-step execution. With increased reasoning capability, GPT-6 Sol Pro delivers stronger performance in software engineering, complex workflows, data analysis, and AI agent applications. It is optimized for developers building advanced AI systems that require consistent results on challenging tasks. GPT-6 Sol Pro supports long-context processing, tool calling, and structured outputs, making it suitable for production-grade applications that require powerful reasoning and automation capabilities.

Total Context

1.1M

Max Output

128K

Released

Sep 22, 2026

Input$2 / M tokens
Output$10 / M tokens
Cache Read$0.2 / M tokens

OpenAI

GPT-6 Luna

gpt-6-luna

GPT-6 Luna is a compact, efficient model in OpenAI's GPT-6 family, delivering strong language intelligence at significantly improved cost efficiency. It offers reliable performance for everyday AI applications such as conversational AI, content generation, coding assistance, and automation workflows. With a large context window and optimized inference efficiency, GPT-6 Luna is ideal for developers building scalable AI products that need high throughput, low latency, and predictable operating costs.

Total Context

1.1M

Max Output

128K

Released

Sep 22, 2026

Input$0.1 / M tokens
Output$0.5 / M tokens
Cache Read$0.01 / M tokens

OpenAI

GPT-6 Luna Pro

gpt-6-luna-pro

GPT-6 Luna Pro is a reasoning-optimized version of GPT-6 Luna, designed to provide stronger problem-solving ability while maintaining the efficiency and affordability of the Luna family. It improves performance on complex instructions, coding tasks, analytical workflows, and lightweight agent applications compared with standard fast inference models. GPT-6 Luna Pro is ideal for developers who need higher-quality responses while maintaining scalability for production workloads. With support for long-context processing, function calling, and structured outputs, GPT-6 Luna Pro provides a balanced solution between intelligence, speed, and cost efficiency.

Total Context

1.1M

Max Output

128K

Released

Sep 22, 2026

Input$0.1 / M tokens
Output$0.5 / M tokens
Cache Read$0.01 / M tokens

Alibaba

Qwen3.8 Omni Flash

qwen3.8-omni-flash

Qwen3.8 Omni Flash is a next-generation omni-modal model from Alibaba’s Qwen family, designed to understand and process text, images, audio, and video within a unified AI system. It combines strong language intelligence with advanced multimodal understanding capabilities for real-world AI applications. With support for up to 1M-token context windows, Qwen3.8 Omni Flash is optimized for long-document analysis, multimedia understanding, and agentic workflows. It can analyze complex audio and video content, perform reasoning across multiple modalities, and support tool calling for building intelligent AI applications. Designed for high-efficiency deployment, Qwen3.8 Omni Flash provides a strong balance between multimodal intelligence, scalability, and cost efficiency, making it suitable for developers building next-generation AI assistants, content analysis systems, and automation agents.

Total Context

1M

Max Output

131.1K

Released

Sep 17, 2026

Input$0.1143 / M tokens
Output$0.3857 / M tokens
Cache Read$0.0143 / M tokens

TypeSafe

Jev 1.13

jev-1.13

Jev 1.13 is TypeSafe AI's first System One decision model, built for structured AI judgment rather than text generation. It analyzes provided context and returns structured choices, scores, and confidence-based decisions. Designed for classification, routing, evaluation, and automation workflows, Jev delivers fast, predictable AI decisions with low latency and efficient token usage.

Total Context

32K

Max Output

N/A

Released

Sep 18, 2026

Input$0.042 / M tokens
Output$0 / M tokens

Z.ai

GLM-5.3-FlashX

glm-5.3-flashx

GLM-5.3-FlashX is a high-speed multimodal model from Z.ai, designed for fast, responsive inference while retaining the core capabilities of GLM-5.3-Flash. It supports text, image, video, and file inputs, with a context window of up to 1 million tokens. The model performs well across coding, agentic workflows, visual understanding, long-context tasks, and complex reasoning, making it a strong choice for real-time AI applications that require both speed and intelligence.

Total Context

1M

Max Output

131.1K

Released

Sep 18, 2026

Input$0.2857 / M tokens
Output$1 / M tokens
Cache Read$0.0814 / M tokens

DeepSeek

DeepSeek V4.1 Flash

deepseek-v4.1-flash

DeepSeek-V4.1-Flash is a fast and efficient multimodal AI model with advanced reasoning, coding, and agent capabilities. Supporting up to 1M tokens of context and native vision understanding, it delivers frontier-level intelligence with lower inference costs, making it ideal for AI agents, automation, coding assistants, and scalable applications.

Total Context

1M

Max Output

384K

Released

Sep 10, 2026

Input$0.2857 / M tokens
Output$1.1429 / M tokens
Cache Read$0.0057 / M tokens

OpenAI

GPT-6 Astra

gpt-6-astra

GPT-6 Astra is OpenAI’s next-generation flagship model, designed for complex reasoning, software engineering, AI agents, research, and demanding end-to-end professional workflows. It excels at multi-step tasks across code, browsers, and professional software, delivering stronger performance in problem solving, tool use, computer use, and long-context understanding. GPT-6 Astra supports a context window of up to 1.05M tokens and a maximum output of 128K tokens, with multiple reasoning effort levels including Low, Medium, High, XHigh, and Max, allowing developers to balance reasoning capability, efficiency, and cost based on task complexity. It is well suited for advanced AI agents, complex software development, automated workflows, deep research, long-document processing, data analysis, and enterprise AI applications.

Total Context

1.1M

Max Output

128K

Released

Sep 4, 2026

Input$10 / M tokens
Output$50 / M tokens
Cache Read$1 / M tokens

Anthropic

Claude Fable 5.1

claude-fable-5.1

Claude Fable 5 is Anthropic’s Mythos-class model designed for demanding reasoning, software engineering, knowledge work, visual analysis, and scientific research. Built for long-horizon agentic tasks, it can sustain planning and execution across complex workflows, analyze large codebases and documents, and refine its work over multiple steps. Claude Fable 5 is well suited to production applications that require advanced coding, research, professional analysis, and reliable autonomous task completion.

Total Context

1M

Max Output

128K

Released

Sep 1, 2026

Input$10 / M tokens
Output$50 / M tokens
Cache Read$0.25 / M tokens

Z.ai

GLM-5.3-Flash

glm-5.3-flash

GLM-5.3-Flash is Z.ai's first natively multimodal model in the GLM-5 family, built for fast coding, visual understanding, and agent workflows. Its 320B-parameter MoE architecture activates 18B parameters per token and combines sparse with linear attention to make long-context tasks more efficient. Use the GLM API through TokenHub to evaluate the current GLM-5.3-Flash price, model ID, supported inputs, tool calling, context limits, and endpoints before deploying it for production coding or multimodal agent tasks.

Total Context

1M

Max Output

131.1K

Released

Aug 26, 2026

Input$0.1143 / M tokens
Output$0.4 / M tokens
Cache Read$0.0329 / M tokens

Z.ai

GLM-5.3

glm-5.3

GLM-5.3 is Z.ai’s latest flagship model built for advanced coding, software engineering, and long-horizon agent tasks. With a 1M-token context window, up to 128K output tokens, strong reasoning, and tool-use capabilities, it is ideal for Coding Agents, complex development workflows, and multi-step automation.

Total Context

1M

Max Output

128K

Released

Aug 19, 2026

Input$1.1429 / M tokens
Output$4 / M tokens
Cache Read$0.2857 / M tokens

Tencent

Hy4 preview

hy4-preview

Hy4 preview is Tencent Hunyuan’s next-generation open-source Mixture-of-Experts flagship model, built for real-world productivity across coding, office work, scientific research, long-context reasoning, and complex agent workflows. Also known as Tencent Hy4 preview or Hunyuan 4 preview, it has 770B total parameters, activates 49B per token, and supports a context window of over 1M tokens. It excels at understanding large codebases and documents, multi-step planning, tool use, debugging, validation, and sustained task execution.

Total Context

1M

Max Output

64K

Released

Aug 28, 2026

Input$0.8571 / M tokens
Output$2.5714 / M tokens
Cache Read$0.0429 / M tokens

Alibaba

Qwen3.8 Max

qwen3.8-max

Qwen3.8-Max is a 2.4T-parameter MoE flagship model for advanced reasoning, coding, enterprise productivity, and agent workflows. It is designed to handle complex, long-running tasks from software development to professional workflows across domains such as law, finance, and design. Use the Qwen API through TokenHub to review the current Qwen3.8 Max price, model ID, context and output limits, supported modalities and endpoints, and benchmark results before deploying it in production.Powered by native multimodal intelligence, Qwen3.8-Max understands text, images, and videos across long contexts, supporting deep analysis of extensive documents and long-form content. Its agentic capabilities enable autonomous planning, execution, verification, and continuous iteration throughout complex workflows.

Total Context

1M

Max Output

128K

Released

Aug 3, 2026

Input$2 / M tokens
Output$6 / M tokens
Cache Read$0.25 / M tokens

Moonshot

Kimi K3

kimi-k3

Kimi K3 is Moonshot AI’s flagship model for long-context coding, complex reasoning, and knowledge work. With native vision and a 1M-token context window, the Kimi K3 API is designed for repository analysis, document research, multi-step agent workflows, and production AI applications. Use TokenHub to review Kimi K3 price, API access, model capabilities, and benchmark results before choosing it for your workload.

Total Context

1M

Max Output

131.1K

Released

Jul 16, 2026

Input$2.8571 / M tokens
Output$14.2857 / M tokens
Cache Read$0.2857 / M tokens

DeepSeek

DeepSeek V4 Flash Vision Exp

deepseek-v4-flash-vision-exp

DeepSeek V4 Flash Vision Exp is an experimental vision-enabled variant of DeepSeek V4 Flash 0731, introducing image understanding capabilities while preserving the base model’s text performance, including agent capabilities, reasoning, and world knowledge. It is built on a sparse Mixture-of-Experts (MoE) architecture with 284B total parameters and 13B active parameters.

Total Context

1M

Max Output

384K

Released

Aug 21, 2026

Input$0.2857 / M tokens
Output$1.1429 / M tokens
Cache Read$0.0057 / M tokens

Popular model recommendations

Start with high-signal models from the live catalog, then open a detail page to compare context, endpoints, and effective pricing.

OpenAI

GPT-6.1 Sol

GPT-6.1 Sol is an advanced reasoning model built for complex problem solving, coding, and agent-based applications. It is designed to deliver strong reasoning performance while maintaining a practical balance between intelligence, speed, and cost. Unlike traditional chat models that focus mainly on generating responses, GPT-6.1 Sol is optimized for tasks that require deeper thinking, including multi-step reasoning, code generation, debugging, data analysis, and following complex instructions. GPT-6.1 Sol supports reliable reasoning workflows, structured outputs, and tool-based applications, making it suitable for developers building AI assistants, coding agents, automation systems, and production AI applications.

Total Context

1.1M

Input Price

$2 / M tokens

View model

OpenAI

GPT-6.1 Sol Pro

GPT-6.1 Sol Pro is an enhanced reasoning version of GPT-6.1 Sol. It uses the same underlying model as GPT-6.1 Sol but enables reasoning.mode=pro, allowing the model to spend significantly more computation on each request. The additional reasoning budget helps GPT-6.1 Sol Pro handle harder problems that require deeper analysis, longer planning, and higher accuracy. It is designed for tasks where getting the best possible answer is more important than response speed or cost efficiency. Compared with GPT-6.1 Sol, GPT-6.1 Sol Pro consumes significantly more reasoning tokens per request, resulting in higher usage costs and longer response times.

Total Context

1.1M

Input Price

$2 / M tokens

View model

Anthropic

Claude Sonnet 5.5

Claude Sonnet 5.5 is a next-generation AI model designed for advanced reasoning, coding, analysis, and complex agentic workflows. It delivers strong performance across software development, long-context understanding, multimodal tasks, and professional knowledge work, offering a balanced combination of intelligence, speed, and efficiency. With enhanced reasoning capabilities and improved instruction following, Claude Sonnet 5.5 excels at writing and debugging code, analyzing complex documents, solving multi-step problems, and powering AI agents that require reliable decision-making and tool usage.

Total Context

1M

Input Price

$2 / M tokens

View model

xAI

Grok 4.7

Grok 4.7 is xAI's latest generation model, built for demanding coding tasks, complex reasoning, and knowledge-intensive applications. It improves long-horizon problem solving, self-verification, and workflow execution, making it a strong fit for developers building intelligent assistants and AI-powered products. Grok 4.7 delivers strong performance across programming, analysis, and general-purpose AI tasks while maintaining competitive efficiency for production deployments.

Total Context

500K

Input Price

$1.6 / M tokens

View model

Anthropic

Claude Opus 5.5

Claude Opus 5.5 is Anthropic's flagship intelligence model, designed for demanding workloads that require deep reasoning, reliability, and long-running task execution. It delivers strong performance in software engineering, agentic workflows, professional analysis, and knowledge-intensive tasks. With improved efficiency and enhanced safeguards, Claude Opus 5.5 is optimized for developers building sophisticated AI applications that require consistent performance across complex tasks.

Total Context

1M

Input Price

$4 / M tokens

View model

OpenAI

GPT-6 Sol

GPT-6 Sol is a high-performance model in OpenAI's GPT-6 series, designed for complex reasoning tasks, software development, and professional AI workflows. It offers enhanced intelligence for multi-step problem solving, advanced coding, data analysis, and agent-based applications. GPT-6 Sol combines strong reasoning capability with efficient inference, making it suitable for developers who need frontier-level performance for production AI systems.

Total Context

1.1M

Input Price

$2 / M tokens

View model

Model Comparison

Quick comparison against selected catalog neighbors.

Model catalog FAQ

A quick guide for choosing, comparing, and using models from the TokenHub catalog.

How should I choose a model from this list?

+

Start with your workload. Use the filters to narrow by provider, tags, endpoint type, and billing group, then compare context size, output limit, modalities, and input or output pricing.

What does effective price mean?

+

Effective price applies the active billing group ratio to the model pricing data. It helps you estimate the real input, output, or per-request cost for the group you are using.

Can I use these models through API endpoints?

+

Yes. Open a model detail page to see the supported endpoint types and documentation links. Availability can differ by model, provider, and current routing configuration.

Why do context window and max output matter?

+

The context window controls how much prompt and conversation history a model can read. Max output controls how much text it can generate in one response, which matters for long-form writing, coding, and document tasks.