Chat & embeddings
Choose one chat endpoint, define model routes, and configure the embedding provider independently.
Chat endpoint and model routes
CHAT_BASE_URL and CHAT_API_KEY select one OpenAI-compatible chat endpoint. apps/web/config/chat-models.yaml assigns the default, fast, reasoning, and vision routes. The checked-in file uses vendor-prefixed IDs; those IDs must be available at the endpoint you select.
For a direct Gemini endpoint, use bare model IDs with a matching registered preset or explicit behavior. The example below illustrates the current schema with an existing preset; confirm the model is available on your provider account.
CHAT_BASE_URL=https://generativelanguage.googleapis.com/v1beta/openai
CHAT_API_KEY=replace-with-your-gemini-key
CHAT_MODELS_CONFIG=config/chat-models.yamlAn explicit model configuration
This minimal example uses one model for every route. Add distinct models when you want a separate fast or reasoning tier. Unknown models require a complete behavior definition; the app does not infer capabilities from the model name.
version: 1
models:
primary:
id: gemini-2.5-flash
preset: google/gemini-2.5-flash
routes:
default: primary
fast: primary
reasoning: primary
vision: primaryCompose mounts this file read-only into both app and worker. After a model edit, restart both processes. For host development, use an absolute CHAT_MODELS_CONFIG path if their working directories differ.
Embeddings are a separate configuration
Chat credentials do not configure embeddings. Supply EMBEDDING_API_BASE_URL and EMBEDDING_API_KEY together, plus a model compatible with the selected index. The default index, legacy-openai-1536, stores 1536-dimensional vectors.
For a new, empty workspace using the registered Gemini index, use the following settings. Do not apply this switch to an existing corpus without a reindex plan: stored vectors belong to a specific model and dimension.
EMBEDDING_INDEX=gemini-embedding-768
EMBEDDING_API_BASE_URL=https://generativelanguage.googleapis.com/v1beta/openai
EMBEDDING_API_KEY=replace-with-your-gemini-key
EMBEDDING_MODEL=gemini-embedding-001Supporting capabilities and workspace settings
AI_BASE_URL + AI_API_KEY provide a global endpoint pair for non-chat capabilities. Per-capability EMBEDDING_*, RERANK_*, NER_*, and TRANSCRIPTION_* API pairs can override it. Keep each key paired with its intended endpoint.
Workspace Settings → Models shows the resolved chat models and routes as a read-only view. Settings → Processing manages workspace embedding settings. EMBEDDING_SECRETS_KEY must contain 32 random bytes encoded as base64 to encrypt workspace provider credentials. Preserve that key across restarts and restores.
GOOGLE_MODEL is not the chat model selector. Use the mounted YAML. Do not replace the embedding index just to change the chat model.