Provider Integration Guide#
This guide explains how to connect a new LLM provider to gptme — whether as a standalone pip-installable plugin, a user-level custom provider config, or a built-in provider PR.
Choosing the Right Integration Path#
Scenario |
Recommended path |
Review required? |
|---|---|---|
Personal/team API endpoint |
No |
|
Shareable OpenAI-compatible | Option 2: Provider Plugin Package provider (e.g. MiniMax, Chutes)| |
Only your package |
|
Widely-used provider that many gptme users will want |
gptme core PR |
|
Start with a plugin package unless the provider has broad community demand — it lets you iterate quickly without waiting for a core PR review.
Option 1: Custom Provider Config#
For a personal or team endpoint (Ollama, vLLM, a private proxy), add a
[[providers]] block to ~/.config/gptme/config.toml:
[[providers]]
name = "mycompany"
base_url = "https://llm.mycompany.internal/v1"
api_key_env = "MYCOMPANY_API_KEY" # env var holding the key
default_model = "llama3.1-8b"
Then use it with gptme 'hello' -m mycompany/llama3.1-8b.
See Custom and Local Providers for the full config reference.
Option 2: Provider Plugin Package#
A provider plugin is a Python package you publish to PyPI. Once installed, users can use your provider without any config file changes.
Note
The gptme-provider-template repository provides a ready-to-fork starting point with a minimal OpenAI-compatible example and an OAuth flow example. Clone it instead of starting from scratch.
Minimum working example#
Create a package (e.g. gptme-provider-acme) with a single module:
# src/gptme_provider_acme/__init__.py
from gptme.llm.models import ModelMeta, ProviderPlugin
provider = ProviderPlugin(
name="acme",
api_key_env="ACME_API_KEY",
base_url="https://api.acme.ai/v1",
models=[
ModelMeta(
provider="unknown",
model="acme/turbo-v1",
context=128_000,
max_output=4_096,
price_input=0.10, # USD per 1M input tokens
price_output=0.40, # USD per 1M output tokens
supports_vision=True,
),
ModelMeta(
provider="unknown",
model="acme/mini-v1",
context=32_000,
max_output=2_048,
price_input=0.02,
price_output=0.08,
),
],
)
Declare the entry point in pyproject.toml:
[project.entry-points."gptme.providers"]
acme = "gptme_provider_acme:provider"
After pip install gptme-provider-acme:
ACME_API_KEY="sk-..." gptme 'hello' -m acme/turbo-v1
Required ProviderPlugin fields#
Field |
Required |
Description |
|---|---|---|
|
Yes |
Unique provider name used as model prefix |
|
Yes |
Environment variable that holds the API key |
|
Yes |
Base URL for the OpenAI-compatible endpoint |
|
No |
List of |
|
No |
|
If init is omitted, gptme auto-initialises an OpenAI client using
base_url and the key from api_key_env.
ModelMeta capability fields#
When declaring models, set these flags to match your provider’s capabilities. gptme uses them to route requests correctly and render UI hints.
Field |
Default |
When to set |
|---|---|---|
|
Required. Max context window (tokens) |
|
|
|
Max tokens the model returns (if capped) |
|
False |
Model accepts image inputs |
|
False |
Model has native chain-of-thought reasoning |
|
True |
Model streams tokens (nearly universal) |
|
False |
Model can emit multiple tool calls at once |
|
False |
Supports OpenAI |
|
False |
Compatible with the OpenAI Responses API |
|
per_token |
|
|
0 |
USD per 1M tokens (0 = free/subscription) |
Custom auth (subscription-backed providers)#
If your provider uses OAuth or a non-API-key auth flow, supply a custom
init function. It must register an OpenAI client before returning:
def _init(config):
token = load_my_token() # read from ~/.myapp/auth.json, etc.
from gptme.llm.llm_openai import _init_openai_client
_init_openai_client(
"myprovider",
api_key=token,
base_url="https://api.myprovider.ai/v1",
)
provider = ProviderPlugin(
name="myprovider",
api_key_env="MYPROVIDER_API_KEY", # keep as fallback even with custom init
base_url="https://api.myprovider.ai/v1",
models=[...],
init=_init,
)
If init runs but does not register a client, gptme raises RuntimeError.
See gptme/llm/llm_grok_subscription.py for a full OAuth PKCE example.
Testing Your Plugin#
Copy this test scaffold into your package’s tests/ directory and adapt it:
"""Test scaffold for gptme provider plugins.
Copy this file into your plugin package's ``tests/`` directory and adapt it:
cp provider_test_scaffold.py tests/test_my_provider.py
Then replace:
- ``acme`` → your provider name
- ``ACME_API_KEY`` → your API key env var
- ``https://api.acme.ai/v1`` → your endpoint
- ``acme/turbo-v1`` → your model IDs
All tests here run *without* hitting the real API — they mock the OpenAI client
layer. Before publishing, also run a live smoke test::
ACME_API_KEY="sk-..." GPTME_RUN_LIVE_TESTS=1 pytest tests/ -m live -v
Usage with pytest::
pytest tests/test_my_provider.py -v
GPTME_RUN_LIVE_TESTS=1 pytest tests/test_my_provider.py -m live -v # requires real API key
"""
from __future__ import annotations
import importlib.metadata
from unittest.mock import MagicMock, patch
import pytest
# ── Adapt these to your provider ─────────────────────────────────────────────
PROVIDER_NAME = "acme"
API_KEY_ENV = "ACME_API_KEY"
BASE_URL = "https://api.acme.ai/v1"
MODEL_IDS = ["acme/turbo-v1", "acme/mini-v1"]
# Import your plugin's ProviderPlugin instance here:
# from gptme_provider_acme import provider as PROVIDER
# For this scaffold we build a synthetic one so it runs without installation:
from gptme.llm.models.types import ModelMeta, ProviderPlugin
PROVIDER = ProviderPlugin(
name=PROVIDER_NAME,
api_key_env=API_KEY_ENV,
base_url=BASE_URL,
models=[
ModelMeta(
provider="unknown",
model=MODEL_IDS[0],
context=128_000,
max_output=4_096,
price_input=0.10,
price_output=0.40,
supports_vision=True,
),
ModelMeta(
provider="unknown",
model=MODEL_IDS[1],
context=32_000,
max_output=2_048,
price_input=0.02,
price_output=0.08,
),
],
)
# ─────────────────────────────────────────────────────────────────────────────
def _make_entry_point(plugin: ProviderPlugin) -> MagicMock:
"""Return a mock entry point that loads *plugin*.
Note: this validates the in-memory plugin wiring only. It does NOT verify
that the entry-point name or import target declared in your pyproject.toml
match the installed package. Test that separately by installing the package
and running ``importlib.metadata.entry_points(group="gptme.providers")``.
"""
ep = MagicMock(spec=importlib.metadata.EntryPoint)
ep.name = plugin.name
ep.load.return_value = plugin
return ep
@pytest.fixture(autouse=True)
def _reset_plugin_cache():
from gptme.llm.provider_plugins import clear_plugin_cache
clear_plugin_cache()
yield
clear_plugin_cache()
@pytest.fixture
def installed_provider():
"""Patch entry-point discovery so tests see your plugin as if it were installed."""
ep = _make_entry_point(PROVIDER)
with patch("importlib.metadata.entry_points", return_value=[ep]):
yield PROVIDER
# ── 1. Plugin metadata ────────────────────────────────────────────────────────
class TestPluginMetadata:
"""Verify the ProviderPlugin declaration is correct."""
def test_name(self):
assert PROVIDER.name == PROVIDER_NAME
def test_api_key_env(self):
assert PROVIDER.api_key_env == API_KEY_ENV
def test_base_url(self):
assert PROVIDER.base_url == BASE_URL
# Must be a valid URL (no trailing slash is fine; leading scheme required)
assert PROVIDER.base_url.startswith(("https://", "http://"))
def test_models_not_empty(self):
assert len(PROVIDER.models) > 0, (
"Provide at least one ModelMeta so users know what models exist"
)
def test_model_ids_fully_qualified(self):
for m in PROVIDER.models:
assert m.model.startswith(f"{PROVIDER_NAME}/"), (
f"Model '{m.model}' must start with '{PROVIDER_NAME}/' "
f"(e.g. '{PROVIDER_NAME}/my-model')"
)
def test_model_context_positive(self):
for m in PROVIDER.models:
assert m.context > 0, f"Model '{m.model}' context must be > 0"
def test_pricing_non_negative(self):
for m in PROVIDER.models:
assert m.price_input >= 0, f"price_input for '{m.model}' must be >= 0"
assert m.price_output >= 0, f"price_output for '{m.model}' must be >= 0"
# ── 2. Plugin discovery ───────────────────────────────────────────────────────
class TestPluginDiscovery:
"""Verify gptme can discover and load your plugin."""
def test_discovered_by_name(self, installed_provider):
from gptme.llm.provider_plugins import get_provider_plugin
result = get_provider_plugin(PROVIDER_NAME)
assert result is not None
assert result.name == PROVIDER_NAME
def test_is_plugin_provider(self, installed_provider):
from gptme.llm.provider_plugins import is_plugin_provider
assert is_plugin_provider(PROVIDER_NAME) is True
def test_not_mistaken_for_builtin(self, installed_provider):
from gptme.llm.models.types import PROVIDERS
assert PROVIDER_NAME not in PROVIDERS, (
f"'{PROVIDER_NAME}' is registered as a built-in provider. "
"Plugin providers must use a name not already in PROVIDERS."
)
# ── 3. Model resolution ───────────────────────────────────────────────────────
class TestModelResolution:
"""Verify model IDs resolve to the right ModelMeta."""
@pytest.mark.parametrize("model_id", MODEL_IDS)
def test_get_model_resolves(self, installed_provider, model_id):
from gptme.llm.models.resolution import get_model
result = get_model(model_id)
assert result.model == model_id
assert result.context > 0
def test_unknown_model_falls_back(self, installed_provider):
from gptme.llm.models.resolution import get_model
result = get_model(f"{PROVIDER_NAME}/nonexistent-model-xyz")
# Should fall back to 128k default, not raise
assert result.context == 128_000
def test_models_appear_in_listing(self, installed_provider):
from gptme.llm.models.listing import get_model_list
models = get_model_list(dynamic_fetch=False)
listed = [m.model for m in models]
for mid in MODEL_IDS:
assert mid in listed, f"Model '{mid}' not in get_model_list() output"
# ── 4. Request routing ────────────────────────────────────────────────────────
class TestRequestRouting:
"""Verify that chat/stream calls are routed to the OpenAI client path."""
def test_chat_routed_through_openai(self, installed_provider):
from gptme.llm import _chat_complete
from gptme.message import Message
messages = [Message("user", "ping")]
fake_result = ("pong", {"model": MODEL_IDS[0]})
with patch("gptme.llm.llm_openai.chat", return_value=fake_result) as mock:
result = _chat_complete(messages, MODEL_IDS[0], None)
assert result == fake_result
mock.assert_called_once()
call_args = mock.call_args
assert call_args.args[1] == MODEL_IDS[0]
def test_stream_routed_through_openai(self, installed_provider):
from gptme.llm import _stream
from gptme.message import Message
messages = [Message("user", "ping")]
def _fake_stream(*args, **kwargs):
yield "pong"
return {"model": MODEL_IDS[0]}
with patch("gptme.llm.llm_openai.stream", side_effect=_fake_stream) as mock:
gen = _stream(messages, MODEL_IDS[0], None)
chunks = list(gen)
assert chunks == ["pong"]
mock.assert_called_once()
# ── 5. Init function (only if you supply a custom init) ──────────────────────
#
# Uncomment and adapt if PROVIDER.init is not None.
#
# class TestCustomInit:
# def test_init_registers_client(self, installed_provider):
# from gptme.llm import init_llm
# from gptme.llm.models import CustomProvider
#
# client_registered = False
#
# def _patched_init_openai_client(provider, **kwargs):
# nonlocal client_registered
# client_registered = True
#
# with (
# patch("gptme.llm.llm_openai._init_openai_client", side_effect=_patched_init_openai_client),
# patch("gptme.llm.llm_openai.has_client", side_effect=lambda _: client_registered),
# patch.dict("os.environ", {API_KEY_ENV: "test-key"}),
# ):
# init_llm(CustomProvider(PROVIDER_NAME))
#
# assert client_registered, "init() must register an OpenAI-compatible client"
# ── 6. Live smoke test (skipped unless --live flag / env var set) ─────────────
@pytest.mark.live
@pytest.mark.skipif(
not (
__import__("os").getenv(API_KEY_ENV)
and __import__("os").getenv("GPTME_RUN_LIVE_TESTS")
),
reason=f"Set {API_KEY_ENV} and GPTME_RUN_LIVE_TESTS=1 to run live API tests",
)
class TestLiveAPI:
"""Optional smoke test that hits the real API.
Requires BOTH the API key AND explicit opt-in to avoid unexpected charges::
{API_KEY_ENV}="sk-..." GPTME_RUN_LIVE_TESTS=1 pytest tests/ -m live -v
These tests are skipped automatically unless both env vars are set.
"""
def test_live_completion(self):
"""A real completion returns non-empty content."""
import gptme.llm as llm
from gptme.llm.models import CustomProvider
from gptme.message import Message
model = MODEL_IDS[0]
llm.init_llm(CustomProvider(PROVIDER_NAME))
messages = [Message("user", 'Reply with the single word "pong".')]
# _chat_complete returns (response_str, metadata_dict)
result = llm._chat_complete(messages, model, None)
assert str(result[0]).strip(), "Expected non-empty response from live API"
Run with:
pytest tests/test_my_provider.py -v
All tests in the scaffold run without hitting the real API — they mock the OpenAI client layer. Before publishing, also run a live smoke test. Two env vars are required to prevent accidental API charges:
ACME_API_KEY="sk-..." GPTME_RUN_LIVE_TESTS=1 pytest tests/ -m live -v
Unified Plugin System#
If your package also provides gptme tools, hooks,
or commands, register via the unified gptme.plugins
entry-point group instead, which lets a single package provide everything:
# src/gptme_plugin_acme/plugin.py
from gptme.plugins.plugin import GptmePlugin
from .provider import provider
plugin = GptmePlugin(
name="acme",
provider=provider, # the ProviderPlugin from above
tools=[...], # optional
)
[project.entry-points."gptme.plugins"]
acme = "gptme_plugin_acme.plugin:plugin"
See Plugin System for the full unified plugin reference.
Package naming#
We recommend the convention:
Provider only:
gptme-provider-<name>(e.g.gptme-provider-minimax)Full plugin:
gptme-plugin-<name>
Option 3: Built-in Provider PR#
Add the provider to gptme core when:
It is widely requested by the community (see open issues/discussions)
It requires special auth that benefits from gptme’s built-in
gptme authflow (e.g. OAuth PKCE, subscription tokens)Its model metadata (capabilities, pricing) is stable and publicly documented
Steps#
Create
gptme/llm/llm_<name>.pyfollowing the pattern inllm_openai.py(for standard API keys) orllm_grok_subscription.py(for OAuth/subscription flows).Register the provider name in
gptme/llm/models/types.py:BuiltinProvider = Literal[ ..., "myprovider", # ← add here ]
If it uses the OpenAI client path, also add to
PROVIDERS_OPENAI.Add models to
gptme/llm/models/openai_models.py(for OpenAI-compat providers) or create a dedicatedllm_<name>_models.py.Wire init — in
gptme/llm/__init__.py, add a branch ininit_llm()to callllm_<name>.init(config).Add auth support — if the provider needs
gptme auth <name>, add a branch ingptme/cli/auth.py.Write tests — add
tests/test_llm_<name>.pycovering at minimum:Auth/init with a mocked HTTP response
Model listing resolves correctly
A streaming response produces expected message chunks
Update docs — add the provider to
docs/providers.rst.Open a PR with the title
feat(providers): add <name> provider.
Before opening the PR, run:
make typecheck
make test
make lint
Common Pitfalls#
Model IDs#
Use the provider-prefixed form in ModelMeta.model — not bare model names:
# ✅ Correct
ModelMeta(provider="unknown", model="acme/turbo-v1", context=128_000)
# ❌ Wrong — provider prefix missing
ModelMeta(provider="unknown", model="turbo-v1", context=128_000)
The prefix must match your ProviderPlugin.name so that
get_model("acme/turbo-v1") can look up the right plugin.
Custom init must register a client#
If you provide init, it must call _init_openai_client() or an
equivalent before returning. Failing to do so raises RuntimeError at
startup and blocks all requests to your provider.
Streaming#
All built-in gptme providers stream by default. The supports_streaming
field in ModelMeta is a metadata-only flag — it affects UI hints and
capability-check display, but does not change how the underlying OpenAI
client path issues requests. The transport currently sends stream=True
unconditionally.
If your endpoint does not support streaming and rejects stream=True:
Configure a proxy layer (e.g. Nginx, Caddy, or LiteLLM) that converts streaming calls to blocking responses; or
Open an issue to request native non-streaming support.
Setting supports_streaming=False in ModelMeta documents the limitation
and suppresses misleading UI streaming indicators, but will not prevent the
stream=True parameter from being sent to your endpoint.
Vision inputs#
If your provider supports image inputs (supports_vision=True), it must
accept the standard OpenAI image_url content-part format — gptme passes
images through without transcoding.
Frequently Asked Questions#
How do I choose a tokenizer?#
Provider plugins do not currently declare a tokenizer. gptme passes the
fully-qualified model ID to its shared token-counting helper, strips known
provider prefixes where supported, and asks tiktoken for that model’s
encoding. Unknown models fall back to cl100k_base; token counts are
therefore estimates for models with a different tokenizer.
There is no tokenizer field on ModelMeta yet, so a plugin cannot select
cl100k_base or o200k_base directly. Keep the model’s canonical upstream
name when possible so tiktoken can recognize it. If exact accounting is
required, open an issue describing the provider and tokenizer family rather
than adding an undocumented field to the plugin.
Which capability flags should I set?#
Capabilities are declared per model in ModelMeta, not on
ProviderPlugin. Set only capabilities the endpoint actually supports:
supports_vision, supports_reasoning,
supports_parallel_tool_calls, supports_strict_tools, and
supports_responses_api. supports_streaming defaults to True but is
currently metadata-only, as described above.
These flags are used in different places: model listing and setup surfaces show
them, vision handling checks supports_vision, and streamed generation uses
supports_parallel_tool_calls to decide whether to stop after the first tool
call. A flag is not a generic transport adapter — the provider must still
implement the corresponding OpenAI-compatible request and response behavior.
Which entry-point group should I use?#
Use gptme.providers when the package exports a standalone
ProviderPlugin. Use gptme.plugins only when the package exports a
GptmePlugin that bundles a provider with tools, hooks, or commands. The two
entry-point groups have different exported object types; they are not
interchangeable aliases.
What goes in ModelMeta.model?#
For a third-party provider, set provider="unknown" and use the fully
qualified runtime ID in model, for example "acme/turbo-v1". The value
is an identifier, not a display label, and its prefix must match
ProviderPlugin.name. Users pass that same ID to --model.
What units do the price fields use?#
price_input and price_output are USD per one million tokens. Enter
the provider’s published per-million values directly — for example, a price of
$0.10 per million input tokens is price_input=0.10, not 0.0001.