Supported Models

Overview

ModelStack provides access to 50+ models from leading AI providers through a single OpenAI-compatible API. Switch between any model by changing one parameter.

info

Versioned model variants are also supported. Any model with a supported prefix (claude-, gpt-, gemini-, deepseek-, qwen-, glm-) is accepted.

Anthropic

Model IDDisplay NameContextBest For
claude-opus-4-6Claude Opus 4.61MAgents, complex reasoning, long-context
claude-opus-4-7Claude Opus 4.71MComplex coding, agentic tasks (legacy)
claude-sonnet-4-6Claude Sonnet 4.61MBest value — near-Opus at Sonnet price
claude-haiku-4-5Claude Haiku 4.5200KHigh-volume, low-latency, sub-agents

OpenAI

Model IDDisplay NameContextBest For
gpt-5.5GPT-5.51.05MFrontier reasoning, complex problem-solving
gpt-5.4GPT-5.41.05MExtended reasoning, professional knowledge work
gpt-5GPT-5400KStrong general performance
gpt-4.1-miniGPT-4.1 Mini1.05MCost-sensitive coding, vision
gpt-4o-miniGPT-4o Mini128KHigh-volume, cost-optimized tasks

Google

Model IDDisplay NameContextBest For
gemini-3.1-proGemini 3.1 Pro Preview1.05MAgentic coding, software engineering
gemini-3.1-flash-liteGemini 3.1 Flash Lite1.05MHigh-volume, low-latency multimodal
gemini-3-flashGemini 3 Flash Preview1.05MMultimodal, computer use, vibe coding
gemini-2.5-proGemini 2.5 Pro1.05MComplex reasoning, math, STEM
gemini-2.5-flashGemini 2.5 Flash1.05MPrice-performance, large-scale processing
gemma-4-31b-itGemma 4 31B262KOpen-weight reasoning and coding

DeepSeek

Model IDDisplay NameContextBest For
deepseek-v3.2DeepSeek V3.2131KReasoning, coding, math at low cost
deepseek-v4-proDeepSeek V4 Pro1.05MAgentic workflows, cost-effective intelligence
deepseek-v4-flashDeepSeek V4 Flash1.05MFast reasoning at low cost

MiniMax

Model IDDisplay NameContextBest For
minimax-m2.7MiniMax M2.7205KAgentic workflows, live debugging, document generation
minimax-m2.5MiniMax M2.5205KCoding (SWE-Bench 80.2%), office productivity
minimax-m2.1MiniMax M2.1205KCoding, agentic workflows

ZAI (GLM)

Model IDDisplay NameContextBest For
glm-5GLM 5203KGeneral chat and vision
glm-4.7GLM 4.7203KGeneral chat and vision

Qwen (Alibaba)

Model IDDisplay NameContextBest For
qwen3-coder-nextQwen3 Coder Next262KCoding agents, tool use, failure recovery
qwen3-coder-flashQwen3 Coder Flash1MMulti-file codebases, repository-scale operations

Model Selection Guide

emoji_events

Best Quality

gpt-5.5 or claude-opus-4-6 — Frontier models for the most complex reasoning, agents, and long-context work.

balance

Best Balance

claude-sonnet-4-6 or gemini-3.1-pro — Near-flagship performance at significantly lower cost.

bolt

Best Speed

claude-haiku-4-5 or gemini-3.1-flash-lite — Sub-10ms latency for high-volume, latency-sensitive tasks.

paid

Best Value

deepseek-v3.2 or gpt-4o-mini — Lowest cost per token for budget-conscious applications.