The assistant and model endpoints
How it works
The assistant is a short loop in xpsflow/agent/orchestrator.py. It sends the conversation and
the list of tool schemas to a chat-completions endpoint, runs whatever tools the model asks for
against the session, returns each result as compact JSON, and stops when the model answers in
prose or a step limit is reached. Every exchange is appended to events.jsonl in the session
directory.
The division of labor is strict: tools compute, the model orchestrates. The model chooses the order of calls and explains the results; it never produces a number of its own. The tools are the same functions the command line, the web workbench and the MCP server use, so a value quoted in the chat is the value in the report. The system prompt tells the model to never invent a number, to read the audit before quoting a fit, and to ask one question when something is ambiguous.
The assistant needs a model that supports tool calling (sometimes called function calling). Models of roughly 7B parameters and up handle multi-step requests reliably; smaller ones get single calls right but tend to stop partway through a long request.
Pointing it at an endpoint
Any server that implements the /chat/completions protocol works. Four environment variables
configure it; a preset supplies defaults and an explicit variable always overrides the preset:
| Variable | Meaning |
|---|---|
XPSFLOW_LLM_PRESET |
ollama (default), cborg, openai or custom |
XPSFLOW_LLM_BASE_URL |
Base URL of the endpoint (overrides the preset) |
XPSFLOW_LLM_API_KEY |
Sent as Authorization: Bearer <key>; none for local servers |
XPSFLOW_LLM_MODEL |
Model name (overrides the preset's default) |
| Preset | Base URL | Default model | Notes |
|---|---|---|---|
ollama |
http://localhost:11434/v1 |
qwen2.5:7b |
Local. ollama pull qwen2.5:7b first. No key. |
cborg |
https://api.cborg.lbl.gov |
lbl/cborg-coder |
An OpenAI-compatible LiteLLM proxy with per-user budgets. |
openai |
https://api.openai.com/v1 |
gpt-4.1 |
Needs an OpenAI key. |
custom |
(yours) | (yours) | Any gateway: vLLM, llama.cpp, LM Studio, a hosted proxy. |
xpsflow models lists what the configured endpoint serves. It queries GET /models and, when
the endpoint is a LiteLLM proxy, merges GET /model/info to show context sizes and whether each
model can call tools. Models that cannot call tools are marked, and the command warns if the
configured model is one of them.
xpsflow models # the configured endpoint
xpsflow models --preset cborg --api-key … # another endpoint, one-off
xpsflow models --json # machine-readable
In the web workbench, Model settings has the same presets: choosing one fills the base URL and model, and Test connection lists the endpoint's models and says whether the chosen model can call tools. Settings typed in the browser stay in the browser; when a custom base URL is set, the server probes it with the key typed there and never with its own.
Ollama (local)
ollama pull qwen2.5:7b
export XPSFLOW_LLM_PRESET=ollama
xpsflow chat examples/demo.vms
CBORG, step by step
- Request a key at https://computeview.lbl.gov/cborg/budget. Keys carry a personal budget;
lab-hosted models such as
lbl/cborg-coderare free against it, while commercial models (anthropic/claude-sonnet,openai/gpt-4.1, …) draw on it. - Configure the endpoint. On the lab network
https://api-local.cborg.lbl.govalso works and avoids the public route.
bash
export XPSFLOW_LLM_PRESET=cborg
export XPSFLOW_LLM_API_KEY=your-cborg-key
- See what is available and pick a tool-capable model:
bash
xpsflow models
export XPSFLOW_LLM_MODEL=lbl/cborg-coder # or anthropic/claude-sonnet, openai/gpt-4.1, …
- Talk to it, from the terminal or the workbench:
bash
xpsflow chat data/sample.vms
xpsflow serve # then open http://127.0.0.1:8765
Any OpenAI-compatible endpoint
export XPSFLOW_LLM_PRESET=custom
export XPSFLOW_LLM_BASE_URL=https://llm.example.org/v1
export XPSFLOW_LLM_API_KEY=your-key
export XPSFLOW_LLM_MODEL=your-model
xpsflow models && xpsflow chat data/sample.vms
xpsflow chat --base-url … --model … overrides the environment for one session.
What the model sees
Only compact tool results go to the model: spectrum inventories, quality grades, a down-sampled sketch of at most 40 points on request, component tables, audit findings, compositions and figure paths. Raw intensity arrays are stripped before serialization and large results are truncated. Instrument files never leave the machine that runs xpsflow; the only outbound connection the assistant makes is to the configured model endpoint, and the only secret it sends is the key you configured for that endpoint. In the workbench, a base URL chosen in the browser is never contacted with the server's key.