Models#
gptme is model-agnostic: it works with any LLM through a single --model flag, and
you can pick a different model for every task. Use a small, fast model for quick
questions and a powerful reasoning model for complex code — without changing tools,
formats, or workflow.
This page helps you pick a model. To set up access to one — API keys, subscriptions, local servers — see Providers.
Recommended models#
For API-key use we recommend Claude Sonnet 5.5 (anthropic/claude-sonnet-5-5, or openrouter/anthropic/claude-sonnet-5.5) as the everyday model, and Claude Opus 5.5 (anthropic/claude-opus-5-5, or openrouter/anthropic/claude-opus-5.5) for the hardest tasks. Both IDs work today; gptme currently inherits context-length and pricing metadata from the closest 4.x catalog entry until 5.5-specific entries land. Both offer:
Strong agentic capabilities
Strong coder capabilities
Strong performance across all tool types and formats
Reasoning capabilities
Vision & computer use capabilities
Claude Sonnet 4.6 (anthropic/claude-sonnet-4-6), the previous recommendation, was a strong and dependable workhorse and still works well. The 5.5 pair is now the easier recommendation over Sonnet 4.6 and the intermediate releases in between.
If you already pay for a frontier subscription, use it instead of an API key (see Subscriptions): GPT-6 Astra (openai-subscription/gpt-6-astra, OpenAI’s frontier flagship), GPT-6.1 Sol (openai-subscription/gpt-6.1-sol, the newer workhorse tier, near-Astra performance at a lower cost) and Grok 4.6 via SuperGrok (grok-subscription/grok-4.6) are all frontier-class and cost nothing per token. Prefer GPT-6.1 Sol over GPT-5.6 Sol for everyday work, and reach for Astra on the hardest tasks.
For high-volume or cost-sensitive work, two open-weight “flash” models hold up well in agentic use for a small fraction of the price:
DeepSeek V4.1 Flash (
openrouter/deepseek/deepseek-v4.1-flash, ordeepseek/deepseek-flashon DeepSeek’s own API). It is the default model on gptme.ai and what-m openrouterresolves to.GLM 5.3 Flash (
openrouter/z-ai/glm-5.3-flash)
DeepSeek V4.1 Flash is not limited to DeepSeek’s own API: OpenRouter routes it to around 30 third-party hosts as well (Together, Fireworks, DeepInfra, and others). Pin one host, or an ordered list that falls back within itself, by appending @ and OpenRouter provider slugs to the model:
gptme -m "openrouter/deepseek/deepseek-v4.1-flash@together"
gptme -m "openrouter/deepseek/deepseek-v4.1-flash@together,fireworks,deepinfra"
Without a pin, OpenRouter picks among all hosts that pass gptme’s privacy defaults (see Data policy). Hosts differ in reliability, speed, cache pricing, and data policy; see Choosing a subprovider before picking one. The earlier deepseek-v4-flash-0731 is still served by third-party hosts on OpenRouter, but DeepSeek’s own API has retired it.
Decent alternatives include:
GPT-6.1 Sol (
openai/gpt-6.1-sol,openai-subscription/gpt-6.1-sol): the current OpenAI workhorse, preferred over GPT-5.6 Sol ($5/$30 per 1M tokens) at a lower API price ($2/$10 per 1M tokens); GPT-6 Astra (openai/gpt-6-astra) is the pricier flagship above it ($10/$50 per 1M tokens)GPT-5.6 Sol / Terra / Luna (
openai/gpt-5.6-sol,openai-subscription/gpt-5.6-sol), the previous generationGemini 3.1 Pro (
gemini/gemini-3.1-pro-preview,openrouter/google/gemini-3.1-pro-preview)Grok 4.6 via the API (
xai/grok-4.6,openrouter/x-ai/grok-4.6)DeepSeek V4 Pro (
openrouter/deepseek/deepseek-v4-pro-0813; DeepSeek’s own API now servesdeepseek-v4-prorequests with V4.1 Flash)Kimi K3 / K2.6 (
moonshot/kimi-k3,openrouter/moonshotai/kimi-k2.6)Qwen3 Max (
openrouter/qwen/qwen3-max)MiniMax M2 (
openrouter/minimax/minimax-m2)
Some models perform better or worse with different --tool-format options (markdown, xml, or tool for native tool-calling); see Tool Formats.
To see how models actually perform on gptme’s eval suites, check the model leaderboard. For an overview of model usage in the wild, see the OpenRouter app analytics for gptme. When you pass only a provider name (-m anthropic), gptme uses that provider’s default model (currently Sonnet 4.6 — not the recommended 5.5). Use the full model ID to target the recommended version.
Pick a model per session#
Pass --model (-m) as <provider>/<model> to choose the model for a single run:
# Quick question — small, cheap, fast
gptme "what does this regex match?" -m openrouter/qwen/qwen3-max
# Complex coding — powerful reasoning model
gptme "refactor this module for testability" -m anthropic/claude-sonnet-5-5
# Use a provider default (no model specified)
gptme "hello" -m anthropic
List the models gptme knows about at any time:
gptme '/models' - '/exit'
The rule of thumb: match the model to the job. Triage, summarization, and quick lookups run fine on small models; multi-step coding and reasoning benefit from a frontier model. Picking per task keeps cost down without capping capability.
One key, many models: OpenRouter#
OpenRouter is the easiest way to reach many models
without managing a separate API key for each provider. With one
OPENROUTER_API_KEY you can route to 100+ models from Anthropic, OpenAI, Google,
DeepSeek, xAI, and more:
gptme "hello" -m openrouter/anthropic/claude-sonnet-5.5
gptme "hello" -m openrouter/deepseek/deepseek-v4.1-flash
gptme "hello" -m openrouter/x-ai/grok-4.6
gptme applies privacy-first defaults for OpenRouter (data collection denied, provider routing requires full parameter support). See OpenRouter for configuration details, quantization controls, and provider pinning.
Data policy#
On gptme.ai, prompts are never routed to a provider that trains on them. The
gptme.ai gateway forces OpenRouter’s no-training filter
(data_collection: "deny") on every request, also excludes every provider
OpenRouter lists as training on prompts, including DeepSeek’s first-party API,
and serves the default DeepSeek V4.1 Flash only from a vetted allowlist of
no-training, zero-data-retention hosts. A request pinned to a training provider
is rejected rather than silently rerouted.
When self-hosting, gptme already sends data_collection: "deny" to
OpenRouter by default (OPENROUTER_DATA_COLLECTION), which skips hosts that
train on prompts, DeepSeek’s official endpoint among them. To go further:
Enforce it account-wide in OpenRouter’s privacy settings, so it holds for every client using your key: turn off providers that may train on inputs, optionally require Zero Data Retention endpoints, and add providers to the account-wide ignore list.
Restrict gptme to hosts you have vetted with an
@pin, or setOPENROUTER_PROVIDER_ORDERto apply the same allowlist to every request (see OpenRouter).Keep in mind that a direct provider such as
deepseek/...bypasses OpenRouter, so only that provider’s own data policy applies; DeepSeek’s allows training on API inputs.
Set a default model#
If you mostly use one model, set it once in your global config
(~/.config/gptme/config.toml) instead of passing --model every time:
[models]
default = "openrouter/qwen/qwen3-max"
With a default configured, gptme "query" uses that model, and --model still
overrides it per run when you need something stronger or cheaper.
See Configuration for the full config reference.
Per-agent models#
When you run multiple agents — for example a team of agents each handling a different role — each can have its own model. Set the model in the agent’s own config so a fast routing agent and a reasoning-heavy coding agent can coexist without per-call flags:
# router-agent/gptme.toml — cheap, fast, handles triage and dispatch
[env]
MODEL = "openrouter/qwen/qwen3-max"
# coder-agent/gptme.toml — frontier model for complex implementation
[env]
MODEL = "anthropic/claude-sonnet-5-5"
A model set this way wins over a global [models].default: the project’s
gptme.toml is the more specific layer. To override it for a single run, pass
--model in the command that runs the agent, or export MODEL for that
command. See How model selection works for the full order.
This is how an agent “brain” pins its default model: configure it once in the agent’s config, override per session only when a specific task needs a different model. No vendor lock-in, no format changes.
See also#
Providers — set up access: credentials, subscriptions, default models
Supported Providers — setup details for each built-in provider
Custom and Local Providers — local and OpenAI-compatible servers
Tool Formats — which tool format suits a model
Evals — how models perform on gptme’s benchmark suite