Skip to content

Models

obsei talks to any OpenAI-compatible chat endpoint with structured JSON output. Endpoints are named under llms in obsei.yaml; plugins refer to them by name (default unless set):

llms:
default:
base_url: http://localhost:11434/v1
model: qwen3:8b
max_requests: 500 # optional; also api_key_env, api_key_header, timeout
Where base_url
Ollama http://localhost:11434/v1
vLLM, llama.cpp, LM Studio http://your-host:8000/v1
Azure OpenAI https://<resource>.openai.azure.com/openai/v1 with api_key_header: api-key
OpenAI, Mistral, Groq, OpenRouter the provider’s /v1 URL
AWS Bedrock, Google Vertex AI, Anthropic a LiteLLM proxy in your network
Mode Model and sink endpoints allowed
air_gapped (default) localhost, private and link-local IPs, .localhost, .internal, .local, single-label hosts
private the above plus OBSEI_EGRESS_ALLOW (comma-separated hosts and their subdomains)
hybrid any host

Set the mode with OBSEI_EGRESS_MODE, or with egress: {mode: private, allowed_hosts: [...]} in obsei.yaml, which takes precedence over the environment. A public endpoint is refused before any request is made.

max_requests caps requests per endpoint client, so a run cannot exceed it. With fallback_llm on the classify enricher, the llm model labels everything and only results below threshold confidence (default 0.7) go to the larger model.

Feedback is sent to the model as JSON data, and the system prompt tells the model never to follow instructions inside it.