# LLM providers

Two options decide which models your deployment runs on. The `providers` option holds the accounts and endpoints that CKEditor AI On-Premises calls. The `models` option holds the models the service offers to end users and says which feature each model serves.

Set both. The service does not start until every conversation, review, and action feature has a model.

This is the configuration behind [Model providers](../../guides/ckeditor-ai/extensions/model-providers.md). Read that page first if you are still deciding whether to run CKEditor AI on your own models. For secrets, databases, and storage, see [Required configuration](configuration.md).

<a id="providers">

## Providers

The `providers` option is required. The service does not start without at least one provider, and a model that names a provider you did not define stops startup too.

The service supports the following LLM providers:

* [OpenAI](#openai-anthropic-and-google)
* [Anthropic](#openai-anthropic-and-google)
* [Google](#openai-anthropic-and-google) (Gemini models)
* [Azure OpenAI](#azure-openai)
* [Amazon Bedrock](#amazon-bedrock)
* [Google Vertex AI](#google-vertex-ai) (Gemini and Anthropic models)
* [Custom providers](#custom-providers) compatible with the [OpenAI Chat Completions API](https://platform.openai.com/docs/api-reference/chat)

`google` and `vertex` are separate provider types. `google` calls the Gemini API with an API key, and `vertex` calls Vertex AI in your Google Cloud project.

Declare each provider in the `providers` option. You choose the key, and a model names that key in its own `provider` option.

**JSON**

```json
{
	"providers": {
		"anthropic": {
			"type": "anthropic",
			"name": "Anthropic",
			"apiKeys": ["api_key1", "api_key2", "api_key3"]
		},
		"google": {
			"type": "google",
			"name": "Google",
			"apiKeys": ["api_key1", "api_key2", "api_key3"]
		},
		"openai": {
			"type": "openai",
			"name": "OpenAI",
			"apiKeys": ["api_key1", "api_key2", "api_key3"]
		},
		"azure": {
			"type": "azure",
			"name": "Azure",
			"resourceName": "azure-resource-name",
			"apiKeys": ["api_key1", "api_key2", "api_key3"]
		},
		"bedrock": {
			"type": "bedrock",
			"name": "Bedrock",
			"region": "us-east-1",
			"apiKeys": ["api_key1", "api_key2", "api_key3"]
		},
		"vertex": {
			"type": "vertex",
			"name": "Vertex",
			"project": "google-project-id",
			"location": "global",
			"apiKeys": ["api_key1", "api_key2", "api_key3"]
		},
		"your-custom-provider": {
			"type": "openai-compatible",
			"name": "Custom provider",
			"baseUrl": "https://your-custom-provider.com",
			"headers": {
				"Authorization": "Bearer token",
				"X-Custom-Header": "custom_value"
			}
		}
	}
}
```

**Environment variable**

The value should be provided as a stringified JSON object:

```bash
PROVIDERS='{
	"anthropic": {
		"type": "anthropic",
		"name": "Anthropic",
		"apiKeys": ["api_key1", "api_key2", "api_key3"]
	},
	"google": {
		"type": "google",
		"name": "Google",
		"apiKeys": ["api_key1", "api_key2", "api_key3"]
	},
	"openai": {
		"type": "openai",
		"name": "OpenAI",
		"apiKeys": ["api_key1", "api_key2", "api_key3"]
	},
	"azure": {
		"type": "azure",
		"name": "Azure",
		"resourceName": "azure-resource-name",
		"apiKeys": ["api_key1", "api_key2", "api_key3"]
	},
	"bedrock": {
		"type": "bedrock",
		"name": "Bedrock",
		"region": "us-east-1",
		"apiKeys": ["api_key1", "api_key2", "api_key3"]
	},
	"vertex": {
		"type": "vertex",
		"name": "Vertex",
		"project": "google-project-id",
		"location": "global",
		"apiKeys": ["api_key1", "api_key2", "api_key3"]
	},
	"your-custom-provider": {
		"type": "openai-compatible",
		"name": "Custom provider",
		"baseUrl": "https://your-custom-provider.com",
		"headers": {
			"Authorization": "Bearer token",
			"X-Custom-Header": "custom_value"
		}
	}
}'
```

Every provider accepts the following options:

* `type` (required) – the provider type: `openai`, `anthropic`, `google`, `azure`, `bedrock`, `vertex`, or `openai-compatible`.
* `name` (optional) – the name displayed in the models list. Defaults to the provider key.
* `apiKeys` (required for `openai`, `anthropic`, and `google`; optional for the remaining types) – the API keys for the provider. The service uses the first key. If the provider rejects that key as invalid, the service uses the next one.
* `baseUrl` (required for `openai-compatible`, optional for the other types) – the base URL of the provider. For an `openai-compatible` provider, every request goes to `{baseUrl}/chat/completions`. For the other types, `baseUrl` replaces the default API URL.
* `headers` (optional) – additional headers sent with every request to the provider.

Azure OpenAI needs either `apiKeys` or the `authentication` option, see [Azure OpenAI](#azure-openai). Amazon Bedrock and Google Vertex AI take their native credentials instead: an access key or a service account key, a role the service assumes, or the credentials of the host it runs on. An `openai-compatible` provider needs no keys when its endpoint is open or authenticates through `headers`.

<a id="openai-anthropic-and-google">

### OpenAI, Anthropic, and Google

These three need only an API key. Set `type` and `apiKeys`, and add `baseUrl` or `headers` if your traffic goes through a proxy. The `google` type calls the Gemini API. For Gemini or Claude models served from your own Google Cloud project, use [Google Vertex AI](#google-vertex-ai) instead.

```json
{
	"providers": {
		"openai": {
			"type": "openai",
			"apiKeys": ["your-openai-api-key"]
		},
		"anthropic": {
			"type": "anthropic",
			"apiKeys": ["your-anthropic-api-key"]
		},
		"google": {
			"type": "google",
			"apiKeys": ["your-gemini-api-key"]
		}
	}
}
```

<a id="azure-openai">

### Azure OpenAI

The Azure OpenAI provider also accepts:

* `resourceName` (optional) – the Azure resource name. You can point to the resource with `baseUrl` instead.

* `apiVersion` (optional, default: `v1`) – the API version, sent as the `api-version` query parameter.

* `useDeploymentBasedUrls` (optional, default: `false`) – use the legacy deployment-based URLs, `{baseUrl}/deployments/{model id}/chat/completions`, instead of the default Responses API URLs under `{baseUrl}/v1`. Set it to `true` when your deployment exposes only the deployment-based Chat Completions endpoint.

* `authentication` (optional) – token-based authentication. Set either `apiKeys` or `authentication`, not both.

  * `type` (required) – either `bearer-token` or `entra-id`.
  * `tokenSecret` (required for `bearer-token`) – the token sent with every request to the Azure endpoint.
  * `scope` (required for `entra-id`) – the scope requested from Microsoft Entra ID, for example `https://cognitiveservices.azure.com/.default`.

With `bearer-token`, the service sends the token you provide. With `entra-id`, the service obtains tokens from the Azure default credential chain. The identity comes from the host: environment variables, a managed identity, or a workload identity. The service refreshes the tokens itself.

The `id` of every model this provider serves must be your Azure deployment name, and the name must start with the id of the OpenAI model behind it. See `id` under [Custom models](#custom-models).

```json
{
	"providers": {
		"azureBearer": {
			"type": "azure",
			"name": "Azure OpenAI with bearer token",
			"baseUrl": "https://your-resource.openai.azure.com",
			"authentication": {
				"type": "bearer-token",
				"tokenSecret": "your_token"
			}
		},
		"azureEntra": {
			"type": "azure",
			"name": "Azure OpenAI with Microsoft Entra ID",
			"baseUrl": "https://your-resource.openai.azure.com",
			"authentication": {
				"type": "entra-id",
				"scope": "https://cognitiveservices.azure.com/.default"
			}
		}
	}
}
```

<a id="amazon-bedrock">

### Amazon Bedrock

The Amazon Bedrock provider also accepts:

* `region` (required) – the AWS region that serves the models, for example `us-east-1`.

* `credentials` (optional) – IAM credentials, an alternative to `apiKeys`.

  * `accessKeyId` (required) – the AWS access key ID.
  * `secretAccessKey` (required) – the AWS secret access key.
  * `sessionToken` (optional) – the AWS session token.

* `assumedRole` (optional) – an IAM role the service assumes through AWS STS before it calls Bedrock.

  * `roleArn` (required) – the ARN of the role to assume.
  * `externalId` (optional) – the external ID that the trust policy of the role requires.
  * `sessionDurationSeconds` (optional) – the lifetime of the temporary credentials, from 900 to 43200 seconds, capped by the maximum session duration of the role. Left unset, the AWS default applies.
  * `roleSessionName` (optional, default: `tiugo-ai-service`) – the session name recorded in AWS CloudTrail. It has to match `^[\w+=,.@-]{2,64}$`.

The service authenticates with the first of these you configure:

1. **`assumedRole`:** the service calls `sts:AssumeRole` and signs Bedrock requests with the temporary credentials that come back, renewing them before they expire. The `AssumeRole` call itself uses `credentials` when you set both options, and the credentials of the host when `assumedRole` stands alone. The role carries the Bedrock permissions, so this is the cross-account setup.
2. **`credentials`:** the service signs requests with the access key you provide. Temporary keys also need `sessionToken`, and the service does not renew them.
3. **`apiKeys`:** the service sends the Bedrock API key as a bearer token, with no request signing.
4. **The credentials of the host:** with none of the three options set, or with `apiKeys` resolving to an empty string, the service takes the identity of the machine it runs on. The AWS SDK looks for it in its own order: environment variables, the shared AWS configuration files, a web identity token file, the container credentials endpoint, and the instance metadata service. A service on EC2, ECS, or EKS reaches Bedrock with no secret in `providers` at all.

```json
{
	"providers": {
		"bedrockAssumedRole": {
			"type": "bedrock",
			"name": "Bedrock with an assumed role",
			"region": "us-east-1",
			"assumedRole": {
				"roleArn": "arn:aws:iam::123456789012:role/ckeditor-ai-bedrock",
				"externalId": "your-external-id",
				"sessionDurationSeconds": 3600
			}
		},
		"bedrockHostCredentials": {
			"type": "bedrock",
			"name": "Bedrock with the credentials of the host",
			"region": "us-east-1"
		}
	}
}
```

> **Warning**
>
> The AWS SDK reads the `AWS_BEARER_TOKEN_BEDROCK` environment variable as a Bedrock API key, and it wins over every signed path above, `assumedRole` and `credentials` included. Leave it unset unless it is the identity you want.

<a id="google-vertex-ai">

### Google Vertex AI

The Google Vertex AI provider also accepts:

* `project` (required) – the Google Cloud project ID. You can set it here or in the `GOOGLE_VERTEX_PROJECT` environment variable.

* `location` (required) – the Google Cloud location, for example `us-central1`. You can set it here or in the `GOOGLE_VERTEX_LOCATION` environment variable.

* `credentials` (optional) – service account credentials, an alternative to `apiKeys`.

  * `clientEmail` (required) – the service account email.
  * `privateKey` (required) – the service account private key.

* `impersonation` (optional) – a service account the service impersonates before it calls Vertex AI.

  * `targetPrincipal` (required) – the email of the service account to impersonate.
  * `lifetimeSeconds` (optional, default: 3600, maximum: 43200) – the lifetime of the access tokens issued for that service account.

As with Bedrock, the service authenticates with the first of these you configure:

1. **`impersonation`:** the service asks the IAM Credentials API for a short-lived access token for `targetPrincipal`, scoped to `https://www.googleapis.com/auth/cloud-platform`, and calls Vertex AI with that token, renewing it as it expires. The token request uses `credentials` when you set both options, and the credentials of the host when `impersonation` stands alone. Grant the source identity the `roles/iam.serviceAccountTokenCreator` role on the target service account.
2. **`credentials`:** the service signs a JWT with the service account key you provide and exchanges it for an access token.
3. **`apiKeys`:** the service calls Vertex AI in express mode with the API key.
4. **The credentials of the host:** with none of the three options set, the service takes the application default credentials of the machine it runs on. The Google auth library looks for them in its own order: the file that `GOOGLE_APPLICATION_CREDENTIALS` points to, the file that `gcloud auth application-default login` writes, and the service account attached to the host. A service on Google Cloud reaches Vertex AI with no secret in `providers` at all.

```json
{
	"providers": {
		"vertexImpersonation": {
			"type": "vertex",
			"name": "Vertex AI with an impersonated service account",
			"project": "your-google-project-id",
			"location": "us-central1",
			"impersonation": {
				"targetPrincipal": "ckeditor-ai@your-google-project-id.iam.gserviceaccount.com",
				"lifetimeSeconds": 3600
			}
		},
		"vertexHostCredentials": {
			"type": "vertex",
			"name": "Vertex AI with the credentials of the host",
			"project": "your-google-project-id",
			"location": "us-central1"
		}
	}
}
```

> **Warning**
>
> An API key authenticates Gemini models only, because express mode is a Gemini feature. Requests to Claude models on Vertex AI always carry a Google OAuth token, so a provider that has only `apiKeys` falls back to the credentials of the host for them. Give a provider that serves Claude models `credentials` or `impersonation`, or run it on a host that carries application default credentials.

<a id="custom-providers">

### Custom providers

Any service that implements the [OpenAI Chat Completions API](https://platform.openai.com/docs/api-reference/chat) can be a provider. Examples: a LiteLLM proxy in front of your own models, a hosted gateway such as Groq, or an Ollama instance on your own hardware. Declare it with `type: "openai-compatible"`. Every request goes to `{baseUrl}/chat/completions`, so include any prefix the routes sit under, such as the `/v1` in the examples below.

**JSON**

```json
{
	"providers": {
		"litellm": {
			"type": "openai-compatible",
			"name": "LiteLLM",
			"baseUrl": "http://litellm.internal:4000/v1",
			"apiKeys": ["sk-litellm-master-key"]
		},
		"groq": {
			"type": "openai-compatible",
			"name": "Groq",
			"baseUrl": "https://api.groq.com/openai/v1",
			"apiKeys": ["gsk_your_groq_key"]
		},
		"ollama": {
			"type": "openai-compatible",
			"name": "Ollama",
			"baseUrl": "http://ollama.internal:11434/v1"
		}
	}
}
```

**Environment variable**

The value should be provided as a stringified JSON object:

```bash
PROVIDERS='{
	"litellm": {
		"type": "openai-compatible",
		"name": "LiteLLM",
		"baseUrl": "http://litellm.internal:4000/v1",
		"apiKeys": ["sk-litellm-master-key"]
	},
	"groq": {
		"type": "openai-compatible",
		"name": "Groq",
		"baseUrl": "https://api.groq.com/openai/v1",
		"apiKeys": ["gsk_your_groq_key"]
	},
	"ollama": {
		"type": "openai-compatible",
		"name": "Ollama",
		"baseUrl": "http://ollama.internal:11434/v1"
	}
}'
```

The Ollama entry above has no `apiKeys`, because a local Ollama endpoint requires no authentication. If the endpoint expects something other than a bearer token, send it in `headers`, as in the `your-custom-provider` example at the top of the page.

An `openai-compatible` provider has no default model list. Declare every model it serves in the `models` option, with `provider` set to the provider key. See [Custom models](#custom-models).

<a id="custom-models">

## Custom models

The OpenAI, Anthropic, and Google providers come with a default list of models. The `models` option replaces the default list with your own. It is also where you declare the models served by Azure OpenAI, Amazon Bedrock, Google Vertex AI, and `openai-compatible` providers, which have no defaults.

**JSON**

```json
{
	"models": [
		{
			"id": "model1",
			"name": "Model 1",
			"description": "Model 1 description",
			"provider": "your-custom-provider",
			"recommended": true,
			"capabilities": {
				"webSearch": true,
				"reasoning": false
			},
			"features": ["conversations", "reviews", "actions"]
		},
		{
			"id": "model2",
			"name": "Model 2",
			"description": "Model 2 description",
			"provider": "your-custom-provider",
			"recommended": true,
			"capabilities": {
				"webSearch": true,
				"reasoning": false
			},
			"features": ["conversations", "reviews", "actions"]
		}
	]
}
```

**Environment variable**

The value should be provided as a stringified JSON array:

```bash
MODELS='[
	{
		"id": "model1",
		"name": "Model 1",
		"description": "Model 1 description",
		"provider": "your-custom-provider",
		"recommended": true,
		"capabilities": {
			"webSearch": true,
			"reasoning": false
		},
		"features": ["conversations", "reviews", "actions"]
	},
	{
		"id": "model2",
		"name": "Model 2",
		"description": "Model 2 description",
		"provider": "your-custom-provider",
		"recommended": true,
		"capabilities": {
			"webSearch": true,
			"reasoning": false
		},
		"features": ["conversations", "reviews", "actions"]
	}
]'
```

> **Note**
>
> For the best results, use models of the same class as the latest major releases from Anthropic, Google, or OpenAI. Older models and smaller models produce weaker results.

For each model you can set the following options:

* `id` (required) – the model identifier used when calling the provider. It must be unique across all models.

  > **Note**
  >
  > For the Azure OpenAI provider, the model `id` must be your Azure deployment name. Deployment names are free-form, so start them with the underlying OpenAI model id. The service reads that prefix to recognize the model family and apply the right request options.

* `provider` (required) – the key of the provider that serves the model, as defined in `providers`. The match is case-insensitive.

* `description` (required) – the description displayed in the models list.

* `type` (optional) – either `standard` (the default) or `agent`, described in the [Agent models](#agent-models) section below.

* `name` (optional) – the name displayed in the models list. Defaults to the model `id`.

* `recommended` (optional) – whether the model belongs to the recommended list. CKEditor 5 selects a recommended model by default.

* `capabilities` (optional) – what the model can do, as an object with the following keys:

  * `webSearch` (optional, default: `false`) – whether the model can use the web search feature.
  * `reasoning` (optional, default: `false`) – whether the model can use the reasoning feature.

* `contextLimits` (optional) – the limits applied to a single conversation, as an object with the following keys:

  * `maxContextLength` (optional, default: 256000) – the maximum context length in characters.
  * `maxFiles` (optional, default: 100) – the maximum number of files in a context.
  * `maxFileSize` (optional, default: 25 MB) – the maximum size of a single file, in bytes.
  * `maxTotalFileSize` (optional, default: 30 MB) – the maximum total size of all files in a context, in bytes.
  * `maxTotalPdfFilePages` (optional, default: 100) – the maximum total number of pages across all PDF files in a context.

  > **Note**
  >
  > `contextLimits` only caps what the model accepts. If a model does not accept file input at all, turn off the file upload permission for the users who can select it.

* `features` (optional) – the features the model serves. System reviews, system actions, and conversation title generation run on a model that the service picks. This option is how you decide which one.

  > **Note**
  >
  > The service fails to start unless every conversation, review, and action feature has a model. List `conversations`, `reviews`, and `actions` on at least one model each, or cover their sub-features one by one. Some features also need a model that supports structured output.

<a id="available-features">

### Available features

* `conversations` – all conversation features.
* `conversations.titleGeneration` – conversation title generation.
* `reviews` – all reviews.
* `reviews.correctness` – the correctness review.
* `reviews.clarity` – the clarity review.
* `reviews.readability` – the readability review.
* `reviews.make-longer` – the review that makes text longer.
* `reviews.make-shorter` – the review that makes text shorter.
* `reviews.make-tone-casual` – the review that makes the tone casual.
* `reviews.make-tone-direct` – the review that makes the tone direct.
* `reviews.make-tone-friendly` – the review that makes the tone friendly.
* `reviews.make-tone-confident` – the review that makes the tone confident.
* `reviews.make-tone-professional` – the review that makes the tone professional.
* `reviews.translate` – the translation review.
* `actions` – all actions.
* `actions.make-longer` – the action that makes text longer.
* `actions.make-shorter` – the action that makes text shorter.
* `actions.make-tone-casual` – the action that makes the tone casual.
* `actions.make-tone-direct` – the action that makes the tone direct.
* `actions.make-tone-friendly` – the action that makes the tone friendly.
* `actions.make-tone-confident` – the action that makes the tone confident.
* `actions.make-tone-professional` – the action that makes the tone professional.
* `actions.translate` – the translation action.
* `actions.continue` – the continue action.
* `actions.fix-grammar` – the action that fixes grammar.
* `actions.improve-writing` – the action that improves writing.

<a id="model-resolution-order">

### Model resolution order

Feature names form a dot-separated hierarchy. In `reviews.correctness`, the parent feature is `reviews`. When several models declare features at different levels, a request goes to the model with the most specific match:

* A request for `reviews.correctness` goes to a model that declares `reviews.correctness`, even if a model that declares `reviews` comes first in the configuration.
* A feature with no exact match falls back to its parent. With only `reviews` and `reviews.correctness` configured, a `reviews.clarity` request goes to the model that declares `reviews`.
* When several models declare the same feature, the order in the configuration ranks them. The first model serves the request. If that model is unavailable, the next model serves it.

For example, take the following configuration, in this order:

1. **Claude Haiku 4.5** with `["reviews.correctness"]`
2. **GPT 5 Mini** with `["reviews.correctness"]`
3. **Gemini 3 Flash** with `["reviews"]`

| Request               | Routed to        | Fallback   | Reason                                                                     |
| --------------------- | ---------------- | ---------- | -------------------------------------------------------------------------- |
| `reviews.correctness` | Claude Haiku 4.5 | GPT 5 Mini | exact match, and Claude Haiku 4.5 comes first in the configuration         |
| `reviews.clarity`     | Gemini 3 Flash   | –          | no exact match, so the request falls back to the parent feature, `reviews` |

<a id="agent-models">

## Agent models

An **agent model** is a virtual model. It calls no provider of its own. It lists standard models in `fallbackOrder`, and each request goes to the first model on that list that is available. If none is available, the request fails.

To the end user, an agent model looks like any other model. The service exposes only the `id`, `name`, and `description` of the agent model. It never sends the fallback list or the provider names to the client.

The `type` property of a `models` entry sets the kind: `standard` (the default, described in [Custom models](#custom-models)) or `agent`.

**JSON**

```json
{
	"models": [
		{
			"type": "standard",
			"id": "claude-major-model",
			"name": "Claude major model",
			"description": "Most powerful model in Claude family",
			"provider": "anthropic",
			"capabilities": {
				"webSearch": true,
				"reasoning": true
			}
		},
		{
			"id": "gpt-newest-model",
			"name": "GPT newest model",
			"description": "Newest model in GPT family",
			"provider": "openai",
			"capabilities": {
				"webSearch": true,
				"reasoning": true
			}
		},
		{
			"type": "agent",
			"id": "default-agent",
			"name": "CKEditor AI Agent",
			"description": "Automatically selects the best available model.",
			"recommended": true,
			"capabilities": {
				"webSearch": true,
				"reasoning": true
			},
			"features": ["conversations", "reviews", "actions"],
			"fallbackOrder": ["claude-major-model", "gpt-newest-model"]
		}
	]
}
```

**Environment variable**

The value should be provided as a stringified JSON array:

```bash
MODELS='[
	{
		"type": "standard",
		"id": "claude-major-model",
		"name": "Claude major model",
		"description": "Most powerful model in Claude family",
		"provider": "anthropic",
		"capabilities": {
			"webSearch": true,
			"reasoning": true
		}
	},
	{
		"id": "gpt-newest-model",
		"name": "GPT newest model",
		"description": "Newest model in GPT family",
		"provider": "openai",
		"capabilities": {
			"webSearch": true,
			"reasoning": true
		}
	},
	{
		"type": "agent",
		"id": "default-agent",
		"name": "CKEditor AI Agent",
		"description": "Automatically selects the best available model.",
		"recommended": true,
		"capabilities": {
			"webSearch": true,
			"reasoning": true
		},
		"features": ["conversations", "reviews", "actions"],
		"fallbackOrder": ["claude-major-model", "gpt-newest-model"]
	}
]'
```

An agent model takes the same options as a standard model (`id`, `name`, `description`, `recommended`, `capabilities`, `contextLimits`, and `features`), with the following differences:

* `type` (required) – must be set to `agent`.
* `fallbackOrder` (required) – a non-empty, ordered array of `id`s of standard models defined in the same `models` array. The service uses the first available one.
* `provider` – ignored. Each request runs on the standard model that serves it and uses the provider of that model.
* `id` – must be unique across all models, as for standard models. No model of either type may use the reserved `agent-` prefix.

> **Note**
>
> If an agent model declares `webSearch` or `reasoning` as `true`, every model in its `fallbackOrder` must declare it too. Otherwise the service fails to start with a validation error.

<a id="when-to-use-agent-models">

### When to use agent models

Agent models separate the model the end user picks from the model that runs. Use them to:

* **Change models without touching clients** – reorder or replace the entries in `fallbackOrder` to roll out a newer model, cut costs, or replace a retired one. End users keep the same agent model, and editor configuration and permissions stay as they are.
* **Keep your model choices private** – end users and your integration see the agent model only, not the models and providers behind it.
* **Keep answering during provider outages** – the service moves to the next available model, so the agent model still answers while one provider is down or throttled.

<a id="model-availability">

## Model availability

The service tracks the health of every model. It takes an unhealthy model out of rotation for 60 seconds. All instances share this state through Redis, so a model that is out of rotation is out everywhere.

A model goes out of rotation after three failures, or three responses that took more than 10 seconds to produce their first token, within 5 minutes. The service counts the failures and the slow responses separately. One healthy response resets both counts. After 60 seconds, the service sends one request to the model as a probe. If the request succeeds, the model returns to rotation. If it fails, the model waits another 60 seconds.

The service picks the model before it sends the request, and the choice does not change mid-request. A request to a standard model that is out of rotation fails, as there is nothing to fall back to. An agent model takes the first model in its `fallbackOrder` that is in rotation and fails only when none is.

A standard model offered directly to end users has no fallback, so an outage at its provider reaches them as failed requests. To avoid that, offer an agent model instead and list models from two providers in its `fallbackOrder`.

---

Full index of the Cloud Services documentation: [llms.txt](../../../llms.txt)
