Providers and models

Connect a built-in or plugin-provided model service, select a model, and configure provider-specific behavior.

Built-in providers

Built-in provider IDs are anthropic, google, openai, openai-chatgpt, openrouter, github-copilot, xai, deepseek, alibaba-cloud, minimax, minimax-coding-plan, opencode-go, and typesafe. claude is an alias for anthropic, gemini is an alias for google, openai-api is an alias for openai, chatgpt is an alias for openai-chatgpt, copilot is an alias for github-copilot, grok is an alias for xai, and dashscope is an alias for alibaba-cloud.

The only built-in UI ID is tui. Plugins may register more providers, aliases, and UIs.

Each provider connection is stored as a named profile. Profile names are globally unique and case-insensitive, while an internal immutable ID keeps sessions, background workers, token refreshes, and caches bound to the same account after a rename.

The profile behind the selected model is stored as profile alongside provider and model in Configuration. For a one-off headless run, use xal run --connection <profile>. If --provider identifies a provider with multiple profiles and no selected profile resolves the ambiguity, Xal requires --connection.

Model discovery

The active profile's catalog is loaded into the process cache when a session starts. When the interactive UI launches, Xal also refreshes every connected text-generation profile's catalog in the background: each profile's stored catalog is served immediately, live discovery runs asynchronously, and nothing waits on the network — while a refresh is pending, readers get the last resolved catalog. /model reuses the process cache and requests each other connected profile's non-refresh catalog at most once, so reopening the picker does not reload successful or failed catalogs. A provider may perform initial live discovery when it has no persistent or bundled catalog. /model refresh and xal models explicitly refresh every connected text-generation profile. A provider that fails or returns an invalid catalog is reported without hiding models from the other providers or preventing the session from starting. If a refresh fails after that profile supplied a valid catalog, Xal keeps the previous in-process catalog available. Catalogs supply the model picker, context-window tracking, input modalities, and the choices shown by /thinking.

Automatic context compaction starts at 80% of the model's active context window by default. A limit saved with /compaction-limit for the canonical provider/model takes precedence over provider metadata. Without a saved value, Xal honors a lower provider-advertised limit. Saved, stale, and provider-advertised values are all capped at the 80% safety ceiling. The command is available only when the active model has a known context window. The full request is checked again after compaction, and Xal never samples a normal request at or beyond the model's hard context window.

The OpenAI provider discovers models from the API key's /models endpoint, keeps GPT-4o and later, o-series, and Codex models that use the Responses API, and stores the last successful result in <app-home>/cache/openai-models-<profile-id>.json. The endpoint does not report context windows, input modalities, or reasoning controls, so Xal applies the configured context cap, lowers it for families with smaller documented windows, and marks the discovered agent models as image-capable. It layers model-family reasoning controls over the result, including the full none through max range for GPT-5.6 and the narrower ranges accepted by earlier GPT-5 models. If live discovery is unavailable, Xal uses that profile's cache or fails when no cache exists.

The ChatGPT provider discovers the account-visible catalog from the authenticated Codex service and stores the last successful result in <app-home>/cache/openai-chatgpt-models-<profile-id>.json, including an optional provider-declared auto-compaction limit. The limit is clamped after Xal applies its context-window cap. If live discovery is unavailable, Xal reports the failure and uses that profile's cache, then its bundled catalog.

GitHub Copilot discovers the models enabled for each connected subscription and stores the compatible subset in <app-home>/cache/github-copilot-models-<profile-id>.json, bound to the token and GitHub domain that produced it. It exposes tool-capable models that advertise /chat/completions, /responses, or omit endpoint metadata, and routes each model through its advertised protocol. Models that explicitly advertise only Anthropic Messages remain hidden. Personal catalogs include both picker-visible and policy-enabled compatible models, which preserves enabled models whose picker flag is unset; if all visibility metadata is absent, every compatible model is used. Enterprise endpoints keep strict picker visibility. If live discovery is unavailable, only that profile's matching validated cache is used; without one, model discovery fails.

Anthropic discovers models from its authenticated /models endpoint and layers bundled context windows, output limits, and thinking options over the result because that endpoint reports none of them. Google Gemini discovers models from /models, keeps the ones that support generateContent, and reads each context window from the reported input token limit. OpenRouter discovers its full catalog from /models and reads context window, input modalities, and reasoning support directly from the response, so no bundled metadata is layered over it. Each of the three falls back to a small bundled catalog and reports the failure when live discovery is unavailable.

xAI discovers models from its authenticated /models endpoint, hides the image, speech, and voice models that the chat endpoint rejects, and layers bundled context windows and thinking options over the result because that endpoint reports neither. The account's credential decides what the endpoint returns, so a Grok subscription and an API key each see their own catalog. DeepSeek discovers models from its authenticated /models endpoint and reports when it must use bundled model metadata. Alibaba Cloud uses a bundled catalog of Qwen models shared by Model Studio and Coding Plan. MiniMax and MiniMax Coding Plan use the bundled minimax.io catalog. OpenCode Go discovers the account-visible catalog from its /models endpoint, which reports only IDs, so bundled metadata supplies names, context windows, input modalities, thinking controls, and each model's wire protocol; unknown IDs are served over Chat Completions with conservative defaults. It falls back to its bundled catalog and reports the failure when live discovery is unavailable.

Anthropic

pluginConfig.anthropic.clientName is a non-empty string used in the provider request user agent. It defaults to the package application name. Anthropic currently has no other configuration options.

Connect with an API key from the Anthropic Console. Reasoning controls depend on the model generation, because the two families accept different request shapes. Claude 4.6 and later take adaptive thinking with summarized reasoning plus an effort level; Claude 4.5 and earlier reject both and take an explicit thinking token budget instead, so Xal maps the selected effort onto a budget for them and omits the effort field. none disables thinking on either family. Xal never sends sampling parameters, because current Claude models reject them. Reasoning blocks are replayed to the same model with their signatures intact, so a thinking turn stays valid across later requests, and a turn that stops at the model output limit is reported as an error rather than returned as a short answer.

Google Gemini

pluginConfig.google.clientName is a non-empty string used in the provider request user agent. It defaults to the package application name. Google Gemini currently has no other configuration options.

Connect with an API key from Google AI Studio. Effort maps onto the reasoning control each model family accepts: Gemini 3 Pro takes only the low and high thinking levels, other Gemini 3 models take the full minimal-to-high range, and Gemini 2.x takes a thinking token budget instead of a level. none requests the lowest level each family allows, with thought summaries hidden, because Gemini 3 cannot disable thinking outright; Gemini 2.x disables it with a zero budget. Thought signatures are carried across streamed parts and replayed to the same model, so multi-step tool use keeps its reasoning context, and a response that stops for any reason other than normal completion is reported as an error.

OpenRouter

pluginConfig.openrouter.clientName is a non-empty string used in the provider request user agent. It defaults to the package application name. OpenRouter currently has no other configuration options.

Connect with an API key from openrouter.ai. Models are addressed by their OpenRouter IDs, such as anthropic/claude-opus-5. Efforts above high are clamped to high, which is the highest level OpenRouter accepts.

Alibaba Cloud

Configure options under pluginConfig.alibaba-cloud:

OptionTypeDefaultDescription
baseUrlstringhttps://dashscope-intl.aliyuncs.com/compatible-mode/v1HTTPS OpenAI-compatible endpoint for the API key's region, workspace, or Coding Plan.
clientNamestringPackage application nameClient name used in the provider request user agent.

Alibaba Cloud Model Studio API keys are region-specific. Set baseUrl to the OpenAI-compatible API Host shown when the key is created. Coding Plan keys use https://coding-intl.dashscope.aliyuncs.com/v1. /connect stores the key without making a billable model request; the first turn validates that the key, endpoint, and selected model are compatible.

MiniMax

pluginConfig.minimax.clientName is a non-empty string used in the provider request user agent. It defaults to the package application name. MiniMax currently has no other configuration options.

Use minimax for standard minimax.io API billing and minimax-coding-plan for a minimax.io Coding Plan. Both providers use the official Anthropic-compatible endpoint at https://api.minimax.io/anthropic/v1 and have separate connection profiles. Run xal connect minimax or xal connect minimax-coding-plan, then paste the matching API key. Connection stores the key without making a billable model request.

The bundled catalog includes MiniMax M3 and the M2 family, including the high-speed M2.5 and M2.7 variants. M3 supports image input and exposes a thinking on/off control. Its Anthropic-compatible interface defaults thinking off, so Xal explicitly requests adaptive thinking unless none is selected. M2 models reason natively without an effort dial and use MiniMax's recommended sampling values.

GitHub Copilot

Configure options under pluginConfig.github-copilot:

OptionTypeDefaultDescription
enterpriseDomainstringgithub.comGitHub Enterprise domain or HTTPS URL used for device login.
clientNamestringPackage application nameClient name used in the provider request user agent.

Run xal connect copilot, open the displayed GitHub device-login URL, and enter its one-time code. Xal uses the resulting GitHub OAuth token directly with the Copilot API and validates that the account returns at least one compatible agent model before storing the token. Personal catalogs that omit endpoint, picker, or policy metadata are accepted unless a model is explicitly incompatible or disabled. Models advertising Responses use that protocol, including newer GPT families, while chat-compatible Claude and other models continue using Chat Completions. Image input is enabled per model from the vision capability advertised by Copilot, so vision-capable models accept PNG and JPEG attachments while text-only models continue to reject them. For GitHub Enterprise, configure enterpriseDomain before connecting.

xAI

Configure options under pluginConfig.xai:

OptionTypeDefaultDescription
baseUrlstringhttps://api.x.ai/v1HTTPS OpenAI-compatible endpoint used for inference and discovery.
clientNamestringPackage application nameClient name used in the provider request user agent.

Run xal connect xai and choose how to authenticate:

Both credential types stream over the OpenAI Responses API, where Grok models expose low, medium, high, and xhigh thinking effort. max maps to xhigh, the highest level xAI accepts. The model catalog is the single source of truth for that dial, so /thinking and the wire never disagree. A few Grok reasoners — the grok-build and grok-4.20-0309 families and grok-composer — think natively but reject the effort parameter, so /thinking does not offer it for them and no effort is sent.

OpenAI

The openai plugin registers both OpenAI providers: openai for OpenAI Platform API keys and openai-chatgpt for ChatGPT subscriptions. They authenticate and bill separately, but stream over the same OpenAI Responses API and share options under pluginConfig.openai:

OptionTypeDefaultDescription
contextWindowPositive integer260000Context-window cap for ChatGPT and assumed context window for OpenAI models.
clientNamestringcodex_cli_rsClient name used in both providers' request user agent.

GPT-5.6 Sol, Terra, and Luna support configurable context windows. Select one of those models, then run /context-window to choose its default 2xxK window, 400K, 600K, 800K, or the provider-advertised maximum. The maximum is 872K for ChatGPT profiles and 1M for OpenAI API profiles. Xal remembers the choice separately for each provider and model and uses it for context tracking and hard request admission. Automatic compaction starts at 80% unless /compaction-limit saves an earlier limit or the provider advertises a lower limit; a saved limit takes precedence over provider metadata. Other providers can expose the context-window command by advertising multiple options for a model.

OpenAI API

Run xal connect openai, name the profile, and paste an API key created in the OpenAI Platform. Xal validates the key against https://api.openai.com/v1/models before storing it. Requests stream through https://api.openai.com/v1/responses with response storage disabled. API profiles are independent from ChatGPT subscription profiles and use the openai provider ID in configuration, thinking preferences, and replay data.

OpenAI ChatGPT

Run xal connect chatgpt and choose browser login, pasted callback, or headless device login for a ChatGPT Pro or Plus subscription. ChatGPT subscription requests use the authenticated Codex service and remain separate from OpenAI API billing and API keys.

DeepSeek

pluginConfig.deepseek.clientName is a non-empty string used in the provider request user agent. It defaults to the package application name. DeepSeek currently has no other configuration options.

OpenCode Go

pluginConfig.opencode-go.clientName is a non-empty string used in the provider request user agent. It defaults to the package application name. OpenCode Go currently has no other configuration options.

All OpenCode Go model requests send x-opencode-session with Xal's conversation ID across Chat Completions, Responses, and Messages. The ID stays the same across turns, retries, and resumed sessions; new or forked conversations get a new ID. No configuration is required.

OpenCode Go is opencode's low-cost subscription for popular open coding models, served from https://opencode.ai/zen/go/v1. Run xal connect opencode-go, then paste the API key from opencode.ai/auth. Connection stores the key without making a billable model request; the first turn validates that the key and subscription cover the selected model.

Each model streams over the protocol its family advertises: Grok 4.5, GPT-5.6 Luna, and Muse Spark use OpenAI Responses; MiniMax M3/M2.x and Qwen3.x use an Anthropic-compatible endpoint; GLM, Kimi, MiMo, Hy3, DeepSeek, and Ox Alpha Free use Chat Completions. GPT-5.6 Luna exposes the full none through max effort range and Grok 4.5 low through xhigh; MiniMax M3 offers a thinking on/off control that Xal maps onto adaptive or disabled thinking. Other models reason natively without a dial.

TypeSafe AI

TypeSafe's Jev is a decision model, not a text generator. It evaluates shared state against typed Noul (yes/no probability), Choice, and Score questions using POST https://api.typesafe.ai/v1/systemone. It cannot run the coding harness and never appears in /model, /models, or xal models.

Get an API key from TypeSafe settings, then run /connect and choose TypeSafe AI, or run xal connect typesafe <profile>. Xal validates the bearer token with GET /v1/models before storing it in the existing protected credentials file. Connecting does not select Jev or replace your harness profile. /profiles and /logout manage TypeSafe profiles normally. Xal does not automatically connect using an environment variable.

Decision requests have a 60-second timeout, support cancellation, and retry transient HTTP/network failures up to twice with backoff and Retry-After. Authentication and malformed responses fail without retry. Model discovery has a 15-second timeout. Plugins use the shared decision service, not another plugin's implementation or credentials.

After connecting, open /config, choose Use TypeSafe AI, and select On. This is the single switch for Jev compaction, Jev read-ahead, and the general-purpose classify tool. Off uses ordinary summary compaction, prefetches nothing, and blocks TypeSafe inference. Search and thinking effort are unchanged by this setting. Connecting alone does not enable compaction. See TypeSafe AI configuration for privacy, profile selection, migration, and fallback details.

API contract: TypeSafe API reference, models, and structured question descriptions.