# The assistant and model endpoints

## How it works

The assistant is a short loop in `xpsflow/agent/orchestrator.py`. It sends the conversation and
the list of tool schemas to a chat-completions endpoint, runs whatever tools the model asks for
against the session, returns each result as compact JSON, and stops when the model answers in
prose or a step limit is reached. Every exchange is appended to `events.jsonl` in the session
directory.

The division of labor is strict: **tools compute, the model orchestrates.** The model chooses the
order of calls and explains the results; it never produces a number of its own. The tools are the
same functions the command line, the web workbench and the [MCP server](mcp.md) use, so a value
quoted in the chat is the value in the report. The system prompt tells the model to never invent
a number, to read the audit before quoting a fit, and to ask one question when something is
ambiguous.

The assistant needs a model that supports **tool calling** (sometimes called function calling).
Models of roughly 7B parameters and up handle multi-step requests reliably; smaller ones get single
calls right but tend to stop partway through a long request.

## Pointing it at an endpoint

Any server that implements the `/chat/completions` protocol works. Four environment variables
configure it; a preset supplies defaults and an explicit variable always overrides the preset:

| Variable | Meaning |
|---|---|
| `XPSFLOW_LLM_PRESET` | `ollama` (default), `cborg`, `openai` or `custom` |
| `XPSFLOW_LLM_BASE_URL` | Base URL of the endpoint (overrides the preset) |
| `XPSFLOW_LLM_API_KEY` | Sent as `Authorization: Bearer <key>`; `none` for local servers |
| `XPSFLOW_LLM_MODEL` | Model name (overrides the preset's default) |

| Preset | Base URL | Default model | Notes |
|---|---|---|---|
| `ollama` | `http://localhost:11434/v1` | `qwen2.5:7b` | Local. `ollama pull qwen2.5:7b` first. No key. |
| `cborg` | `https://api.cborg.lbl.gov` | `lbl/cborg-coder` | An OpenAI-compatible LiteLLM proxy with per-user budgets. |
| `openai` | `https://api.openai.com/v1` | `gpt-4.1` | Needs an OpenAI key. |
| `custom` | (yours) | (yours) | Any gateway: vLLM, llama.cpp, LM Studio, a hosted proxy. |

`xpsflow models` lists what the configured endpoint serves. It queries `GET /models` and, when
the endpoint is a LiteLLM proxy, merges `GET /model/info` to show context sizes and whether each
model can call tools. Models that cannot call tools are marked, and the command warns if the
configured model is one of them.

```bash
xpsflow models                              # the configured endpoint
xpsflow models --preset cborg --api-key …   # another endpoint, one-off
xpsflow models --json                       # machine-readable
```

In the web workbench, **Model settings** has the same presets: choosing one fills the base URL
and model, and **Test connection** lists the endpoint's models and says whether the chosen model
can call tools. Settings typed in the browser stay in the browser; when a custom base URL is set,
the server probes it with the key typed there and never with its own.

### Ollama (local)

```bash
ollama pull qwen2.5:7b
export XPSFLOW_LLM_PRESET=ollama
xpsflow chat examples/demo.vms
```

### CBORG, step by step

1. Request a key at <https://computeview.lbl.gov/cborg/budget>. Keys carry a personal budget;
   lab-hosted models such as `lbl/cborg-coder` are free against it, while commercial models
   (`anthropic/claude-sonnet`, `openai/gpt-4.1`, …) draw on it.
2. Configure the endpoint. On the lab network `https://api-local.cborg.lbl.gov` also works and
   avoids the public route.

   ```bash
   export XPSFLOW_LLM_PRESET=cborg
   export XPSFLOW_LLM_API_KEY=your-cborg-key
   ```

3. See what is available and pick a tool-capable model:

   ```bash
   xpsflow models
   export XPSFLOW_LLM_MODEL=lbl/cborg-coder     # or anthropic/claude-sonnet, openai/gpt-4.1, …
   ```

4. Talk to it, from the terminal or the workbench:

   ```bash
   xpsflow chat data/sample.vms
   xpsflow serve        # then open http://127.0.0.1:8765
   ```

### Any OpenAI-compatible endpoint

```bash
export XPSFLOW_LLM_PRESET=custom
export XPSFLOW_LLM_BASE_URL=https://llm.example.org/v1
export XPSFLOW_LLM_API_KEY=your-key
export XPSFLOW_LLM_MODEL=your-model
xpsflow models && xpsflow chat data/sample.vms
```

`xpsflow chat --base-url … --model …` overrides the environment for one session.

## What the model sees

Only compact tool results go to the model: spectrum inventories, quality grades, a down-sampled
sketch of at most 40 points on request, component tables, audit findings, compositions and
figure paths. Raw intensity arrays are stripped before serialization and large results are
truncated. Instrument files never leave the machine that runs xpsflow; the only outbound
connection the assistant makes is to the configured model endpoint, and the only secret it sends
is the key you configured for that endpoint. In the workbench, a base URL chosen in the browser
is never contacted with the server's key.
