Skip to guide
Deployment/Configure
Self-hosting reference

Chat & embeddings

Choose one chat endpoint, define model routes, and configure the embedding provider independently.

Chat endpoint and model routes

CHAT_BASE_URL and CHAT_API_KEY select one OpenAI-compatible chat endpoint. apps/web/config/chat-models.yaml assigns the default, fast, reasoning, and vision routes. The checked-in file uses vendor-prefixed IDs; those IDs must be available at the endpoint you select.

For a direct Gemini endpoint, use bare model IDs with a matching registered preset or explicit behavior. The example below illustrates the current schema with an existing preset; confirm the model is available on your provider account.

.env
CHAT_BASE_URL=https://generativelanguage.googleapis.com/v1beta/openai
CHAT_API_KEY=replace-with-your-gemini-key
CHAT_MODELS_CONFIG=config/chat-models.yaml

An explicit model configuration

This minimal example uses one model for every route. Add distinct models when you want a separate fast or reasoning tier. Unknown models require a complete behavior definition; the app does not infer capabilities from the model name.

apps/web/config/chat-models.yaml
version: 1
models:
  primary:
    id: gemini-2.5-flash
    preset: google/gemini-2.5-flash
routes:
  default: primary
  fast: primary
  reasoning: primary
  vision: primary

Compose mounts this file read-only into both app and worker. After a model edit, restart both processes. For host development, use an absolute CHAT_MODELS_CONFIG path if their working directories differ.

Embeddings are a separate configuration

Chat credentials do not configure embeddings. Supply EMBEDDING_API_BASE_URL and EMBEDDING_API_KEY together, plus a model compatible with the selected index. The default index, legacy-openai-1536, stores 1536-dimensional vectors.

For a new, empty workspace using the registered Gemini index, use the following settings. Do not apply this switch to an existing corpus without a reindex plan: stored vectors belong to a specific model and dimension.

.env · new Gemini corpus only
EMBEDDING_INDEX=gemini-embedding-768
EMBEDDING_API_BASE_URL=https://generativelanguage.googleapis.com/v1beta/openai
EMBEDDING_API_KEY=replace-with-your-gemini-key
EMBEDDING_MODEL=gemini-embedding-001

Supporting capabilities and workspace settings

AI_BASE_URL + AI_API_KEY provide a global endpoint pair for non-chat capabilities. Per-capability EMBEDDING_*, RERANK_*, NER_*, and TRANSCRIPTION_* API pairs can override it. Keep each key paired with its intended endpoint.

Workspace Settings → Models shows the resolved chat models and routes as a read-only view. Settings → Processing manages workspace embedding settings. EMBEDDING_SECRETS_KEY must contain 32 random bytes encoded as base64 to encrypt workspace provider credentials. Preserve that key across restarts and restores.

GOOGLE_MODEL is not the chat model selector. Use the mounted YAML. Do not replace the embedding index just to change the chat model.

Deployment Guide — Self-Host Launchstack | Launchstack