Security
xpsflow is a local analysis tool. This page says what that means in practice: what the threat model is, what leaves the machine and when, which limits and headers the workbench enforces, how to report a problem, and what the stress suite checks on every commit.
Threat model
What it is. xpsflow serve starts a FastAPI server on 127.0.0.1:8765
and opens a single-page workbench in your browser. The CLI, the workbench and
the chat assistant all call the same deterministic tool layer
(agent/tools.py). There is no login, no user database and no multi-tenant
intent: the server trusts whoever can reach the port, which by default is only
a process on the same machine.
What it is not. It is not an internet-facing service. Binding to
0.0.0.0 or any non-loopback address prints a warning, because every client
that can reach the port can upload files, run every tool and read every
session. If you need remote access, keep the loopback bind and use an SSH
tunnel, or put the app behind a reverse proxy that performs authentication.
Sessions. Each browser tab gets a server-issued session id (12 hex
characters). Only ids of that exact form are accepted, so a client cannot
point a session at another directory. A session's files live under
$XPSFLOW_RUNS_DIR/<sid>/ (default runs/): uploads/ for the instrument
files, figures/ for PNGs, events.jsonl for the tool log and the reports.
Every path a tool receives in the web app is resolved against the session's
uploads/ folder and refused when it lands outside it, so neither a browser
client nor a language model can read other files on the machine. The CLI and
the Python API run without this sandbox because they act on behalf of the
local user. The session store is bounded (XPSFLOW_MAX_SESSIONS, default
64); the least recently used session is dropped first.
Trust boundaries.
| Input | Who controls it | What the code assumes |
|---|---|---|
| Instrument files | the user, or anyone who gave them a file | Nothing. Readers are fuzzed with random bytes, truncation and byte flips and may only raise ValueError. Uploads that no reader accepts are deleted and reported by name and error type, never by content. |
| Templates (YAML) | the user, or a model via propose_fit |
Parsed with yaml.safe_load and validated by pydantic. expr constraints are restricted to {Component.param} references, numbers, + - * / ** with small constant exponents, parentheses and abs/min/max/sqrt/exp/log before lmfit's interpreter ever sees them. In the web app, use_template accepts built-in template names only, never file paths. |
| Tool arguments | the browser or the language model | Coerced to the schema derived from the function signature; unknown keys are dropped, enumerations such as lineshape and background are validated by pydantic, paths go through the sandbox. |
| Chat endpoint settings | the browser | A client-chosen base_url never receives the server's own API key. |
What leaves the machine, and when.
- Nothing, unless you use the assistant. Opening the workbench loads Plotly
and the fonts from the package itself (
static/vendor/,static/fonts/), the HTML report inlines its fonts and figures, and the Content-Security-Policy forbids the page from connecting anywhere else. - When you send a chat message, the orchestrator posts the conversation, the
tool schemas and compact tool results to the configured chat-completions
endpoint (
XPSFLOW_LLM_BASE_URL). Raw spectra and curves are stripped before serialization (for_modeldrops arrays, sketches and figures), but file names, labels, acquisition metadata, fitted numbers, audit findings and your own notes are part of the conversation. Point the endpoint at a local model if that must not leave the machine. xpsflownever phones home, fetches reference data or checks for updates.
Headers and limits
Every response from the server carries:
| header | value |
|---|---|
Content-Security-Policy |
default-src 'self'; script-src 'self'; style-src 'self' 'unsafe-inline'; img-src 'self' data: blob:; connect-src 'self'; font-src 'self' data:; frame-ancestors 'none'; base-uri 'self'; form-action 'self'; object-src 'none' |
X-Content-Type-Options |
nosniff |
Referrer-Policy |
no-referrer |
X-Frame-Options |
DENY |
Permissions-Policy |
camera=(), microphone=(), geolocation=() |
Cache-Control (on /api/*) |
no-store |
style-src allows inline styles because Plotly sets them on the elements it
creates; scripts are never inline. font-src and img-src allow data: so
that the self-contained HTML report renders under the same policy.
Cross-site writes are refused: a POST with an Origin header that does not
match the Host gets a 403, so a page on another origin cannot drive the
server through your browser.
Limits (all configurable through environment variables):
| limit | default | variable |
|---|---|---|
| size of one uploaded file | 50 MB | XPSFLOW_MAX_UPLOAD_MB |
| files per upload request | 32 | XPSFLOW_MAX_UPLOAD_FILES |
| total upload volume per session | 500 MB | XPSFLOW_MAX_SESSION_UPLOAD_MB |
JSON body on /api/chat and /api/tool/* |
256 KB | XPSFLOW_MAX_JSON_KB |
| chat requests per session per minute | 30 | XPSFLOW_CHAT_RATE_LIMIT |
| sessions kept in memory | 64 | XPSFLOW_MAX_SESSIONS |
A body without Content-Length on the capped endpoints is refused (411), an
oversized one gets 413, and the rate limit answers 429 with Retry-After.
Upload file names are reduced to a bare name and refused when they are .,
.. or contain path separators or control characters. When any file in an
upload fails (too large, over quota, unreadable) every file written by that
request is deleted.
Static analysis and dependency audit
CI runs bandit over src and scripts and fails on high-severity findings,
and pip-audit over the installed environment and fails on any known
vulnerability. The bandit configuration is in pyproject.toml
([tool.bandit]); the skipped checks are listed there with a reason each.
Run both locally with:
bandit -c pyproject.toml -r src scripts --severity-level high
pip-audit --skip-editable
Reporting a vulnerability
Please do not open a public issue for a security problem. Use GitHub's
private vulnerability reporting on the repository
(Security → Report a vulnerability), as described in
SECURITY.md. Include the version or commit, the steps to
reproduce and what you think the impact is.
What the stress suite covers
tests/test_stress.py runs on every commit (about 30 s; the memory test is
marked slow and CI runs it too). It checks, with hypothesis and
pytest-timeout:
- Readers. Random datasets written by
write_vamasround-trip throughread_vamaswith energies, intensities and metadata intact. Random ASCII tables with mixed separators, decimal commas, comment lines, headers, chopped rows and every line ending either parse into finite, sorted spectra or raiseValueError. Random bytes and truncated or byte-flipped copies of the Kratos and VAMAS fixtures never raise anything butValueError. - Fitting. Flat, all-zero, NaN, infinite, single-point, duplicated, reversed, negative, enormous, tiny, out-of-window, constant-axis and 100 000-point spectra either fit with no NaN anywhere in the result or fail with a clear error, under a 60 s timeout, through both the engine and model selection. The quality gate never raises on any of them.
- Pipeline. Ten synthetic samples analyzed in one process (quality, calibration, survey, fit, quantification) must not grow peak RSS by more than 200 MB. A session that re-uploads the same file twenty times keeps a consistent inventory, no stale results and a single file on disk.
- Concurrency. Eight threads on separate sessions upload different files, run tools and read spectra at the same time; each only ever sees its own data, and each session's folder holds only its own upload.