# Security

xpsflow is a local analysis tool. This page says what that means in practice:
what the threat model is, what leaves the machine and when, which limits and
headers the workbench enforces, how to report a problem, and what the stress
suite checks on every commit.

## Threat model

**What it is.** `xpsflow serve` starts a FastAPI server on `127.0.0.1:8765`
and opens a single-page workbench in your browser. The CLI, the workbench and
the chat assistant all call the same deterministic tool layer
(`agent/tools.py`). There is no login, no user database and no multi-tenant
intent: the server trusts whoever can reach the port, which by default is only
a process on the same machine.

**What it is not.** It is not an internet-facing service. Binding to
`0.0.0.0` or any non-loopback address prints a warning, because every client
that can reach the port can upload files, run every tool and read every
session. If you need remote access, keep the loopback bind and use an SSH
tunnel, or put the app behind a reverse proxy that performs authentication.

**Sessions.** Each browser tab gets a server-issued session id (12 hex
characters). Only ids of that exact form are accepted, so a client cannot
point a session at another directory. A session's files live under
`$XPSFLOW_RUNS_DIR/<sid>/` (default `runs/`): `uploads/` for the instrument
files, `figures/` for PNGs, `events.jsonl` for the tool log and the reports.
Every path a tool receives in the web app is resolved against the session's
`uploads/` folder and refused when it lands outside it, so neither a browser
client nor a language model can read other files on the machine. The CLI and
the Python API run without this sandbox because they act on behalf of the
local user. The session store is bounded (`XPSFLOW_MAX_SESSIONS`, default
64); the least recently used session is dropped first.

**Trust boundaries.**

| Input | Who controls it | What the code assumes |
|---|---|---|
| Instrument files | the user, or anyone who gave them a file | Nothing. Readers are fuzzed with random bytes, truncation and byte flips and may only raise `ValueError`. Uploads that no reader accepts are deleted and reported by name and error type, never by content. |
| Templates (YAML) | the user, or a model via `propose_fit` | Parsed with `yaml.safe_load` and validated by pydantic. `expr` constraints are restricted to `{Component.param}` references, numbers, `+ - * / **` with small constant exponents, parentheses and `abs/min/max/sqrt/exp/log` before lmfit's interpreter ever sees them. In the web app, `use_template` accepts built-in template names only, never file paths. |
| Tool arguments | the browser or the language model | Coerced to the schema derived from the function signature; unknown keys are dropped, enumerations such as lineshape and background are validated by pydantic, paths go through the sandbox. |
| Chat endpoint settings | the browser | A client-chosen `base_url` never receives the server's own API key. |

**What leaves the machine, and when.**

- Nothing, unless you use the assistant. Opening the workbench loads Plotly
  and the fonts from the package itself (`static/vendor/`, `static/fonts/`),
  the HTML report inlines its fonts and figures, and the Content-Security-Policy
  forbids the page from connecting anywhere else.
- When you send a chat message, the orchestrator posts the conversation, the
  tool schemas and compact tool results to the configured chat-completions
  endpoint (`XPSFLOW_LLM_BASE_URL`). Raw spectra and curves are stripped before
  serialization (`for_model` drops arrays, sketches and figures), but file
  names, labels, acquisition metadata, fitted numbers, audit findings and your
  own notes are part of the conversation. Point the endpoint at a local model
  if that must not leave the machine.
- `xpsflow` never phones home, fetches reference data or checks for updates.

## Headers and limits

Every response from the server carries:

| header | value |
|---|---|
| `Content-Security-Policy` | `default-src 'self'; script-src 'self'; style-src 'self' 'unsafe-inline'; img-src 'self' data: blob:; connect-src 'self'; font-src 'self' data:; frame-ancestors 'none'; base-uri 'self'; form-action 'self'; object-src 'none'` |
| `X-Content-Type-Options` | `nosniff` |
| `Referrer-Policy` | `no-referrer` |
| `X-Frame-Options` | `DENY` |
| `Permissions-Policy` | `camera=(), microphone=(), geolocation=()` |
| `Cache-Control` (on `/api/*`) | `no-store` |

`style-src` allows inline styles because Plotly sets them on the elements it
creates; scripts are never inline. `font-src` and `img-src` allow `data:` so
that the self-contained HTML report renders under the same policy.

Cross-site writes are refused: a `POST` with an `Origin` header that does not
match the `Host` gets a 403, so a page on another origin cannot drive the
server through your browser.

Limits (all configurable through environment variables):

| limit | default | variable |
|---|---|---|
| size of one uploaded file | 50 MB | `XPSFLOW_MAX_UPLOAD_MB` |
| files per upload request | 32 | `XPSFLOW_MAX_UPLOAD_FILES` |
| total upload volume per session | 500 MB | `XPSFLOW_MAX_SESSION_UPLOAD_MB` |
| JSON body on `/api/chat` and `/api/tool/*` | 256 KB | `XPSFLOW_MAX_JSON_KB` |
| chat requests per session per minute | 30 | `XPSFLOW_CHAT_RATE_LIMIT` |
| sessions kept in memory | 64 | `XPSFLOW_MAX_SESSIONS` |

A body without `Content-Length` on the capped endpoints is refused (411), an
oversized one gets 413, and the rate limit answers 429 with `Retry-After`.
Upload file names are reduced to a bare name and refused when they are `.`,
`..` or contain path separators or control characters. When any file in an
upload fails (too large, over quota, unreadable) every file written by that
request is deleted.

## Static analysis and dependency audit

CI runs `bandit` over `src` and `scripts` and fails on high-severity findings,
and `pip-audit` over the installed environment and fails on any known
vulnerability. The bandit configuration is in `pyproject.toml`
(`[tool.bandit]`); the skipped checks are listed there with a reason each.
Run both locally with:

```bash
bandit -c pyproject.toml -r src scripts --severity-level high
pip-audit --skip-editable
```

## Reporting a vulnerability

Please do not open a public issue for a security problem. Use GitHub's
private vulnerability reporting on the repository
(**Security → Report a vulnerability**), as described in
[`SECURITY.md`](../SECURITY.md). Include the version or commit, the steps to
reproduce and what you think the impact is.

## What the stress suite covers

`tests/test_stress.py` runs on every commit (about 30 s; the memory test is
marked `slow` and CI runs it too). It checks, with `hypothesis` and
`pytest-timeout`:

- **Readers.** Random datasets written by `write_vamas` round-trip through
  `read_vamas` with energies, intensities and metadata intact. Random ASCII
  tables with mixed separators, decimal commas, comment lines, headers,
  chopped rows and every line ending either parse into finite, sorted spectra
  or raise `ValueError`. Random bytes and truncated or byte-flipped copies of
  the Kratos and VAMAS fixtures never raise anything but `ValueError`.
- **Fitting.** Flat, all-zero, NaN, infinite, single-point, duplicated,
  reversed, negative, enormous, tiny, out-of-window, constant-axis and
  100 000-point spectra either fit with no NaN anywhere in the result or fail
  with a clear error, under a 60 s timeout, through both the engine and model
  selection. The quality gate never raises on any of them.
- **Pipeline.** Ten synthetic samples analyzed in one process (quality,
  calibration, survey, fit, quantification) must not grow peak RSS by more
  than 200 MB. A session that re-uploads the same file twenty times keeps a
  consistent inventory, no stale results and a single file on disk.
- **Concurrency.** Eight threads on separate sessions upload different files,
  run tools and read spectra at the same time; each only ever sees its own
  data, and each session's folder holds only its own upload.
