β¨ AI Chat & Models
Chat already streams through FastAPI to OpenAI or Anthropic. Change identity, the system prompt, and which models each plan may use.
Chat already streams through your API. Change the voice, the models, and what the product is for.
The browser never calls OpenAI or Anthropic. It talks to your API. The API streams tokens back, stores the conversation, and tracks usage.
User
β
Next.js (`apps/web`)
β
FastAPI (`apps/api`)
β
OpenAI / Anthropic
β
streamed responseHow do I turn the demo AI into my product?
Do these first. You do not need to read the whole llm/ package.
| Order | Change | Where |
|---|---|---|
| 1 | Name, branding, which AI features exist | starter.config.json |
| 2 | Who may use which models and quotas | apps/api/app/plans/config.py |
| 3 | Product voice (system prompt) | Default in chat_service.py or use-chat-stream.ts β see below |
| 4 | When to call tools | apps/api/app/ai/tools/prompts.py β see Tools |
| 5 | Turn off demo tools you do not ship | features.tools.* |
| 6 | Provider keys | apps/api/.env |
Chat empty-state copy lives in the frontend (empty-state.tsx). That is not the model prompt.
Providers
Configure at least one chat key:
Defaults: DEFAULT_MODEL=gpt-4o, ANTHROPIC_DEFAULT_MODEL=claude-opus-4-8.
There is no Gemini provider in this release.
Users pick a provider/model in Settings β AI. Resolution order: the request, then saved preferences, then app defaults. Plans still restrict allowed_models.
The system prompt
There is no single SYSTEM_PROMPT file. Each conversation can store one (conversations.system_prompt). It is injected first, then tools, memory, file context, and language.
The shipped chat UI currently sends system_prompt: null (apps/web/src/hooks/use-chat-stream.ts). Until you set one, the model only gets the built-in tool/memory/file instructions.
To give every chat your product voice, pick one:
- Set a default string in
_build_llm_messagesinapps/api/app/services/chat_service.pywhenconversation.system_promptis empty - Send it from
use-chat-stream.tson each request (the API stores it on the conversation the first time) PATCH /api/v1/conversations/{id}withsystem_promptfor a single thread
Keep it under MAX_SYSTEM_PROMPT_CHARS (default 8000).
Tool routing copy is separate: apps/api/app/ai/tools/prompts.py. Planner copy is apps/api/app/ai/planner/prompts.py. See Tools and Agents.
Streaming and history
| Piece | Path |
|---|---|
| Browser | apps/web/src/hooks/use-chat-stream.ts |
| Route | POST /api/v1/chat/completions/stream |
| Service | apps/api/app/services/chat_service.py |
| Providers | apps/api/app/llm/ |
Conversations and messages persist in Postgres. The first turn can generate a title. Failed hard errors refund the chat-message quota.
Usage
Each send consumes a chat_message against the plan. Token usage is recorded from the provider response. Settings β Usage shows the summary (GET /api/v1/usage).
Tune limits in apps/api/app/plans/config.py, not in the React tree.
Settings that matter
| Setting | Where |
|---|---|
| Feature flags (files, memory, Deep Research, tools) | starter.config.json |
TOOLS_ENABLED / TOOL_MAX_STEPS | apps/api/.env |
| User memory toggle | Settings β AI |
| Deep Research in the UI | Plan entitlement + features.deepResearch |
Next: give the model something to do besides talk.