Architecture and design decisions
xpsflow is built around one rule: deterministic tools compute; everything else only calls them.
The CLI, the web workbench and the conversational agent all sit on the same tool layer, so a
number seen in a chat is the number in the PDF, and either can be reproduced from results.json
and the saved templates.
┌──────────────┐ ┌──────────────┐ ┌──────────────────────┐
people → │ CLI (typer) │ │ web (FastAPI)│ │ agent (tool calling) │
└──────┬───────┘ └──────┬───────┘ └──────────┬───────────┘
└──────────────────┼──────────────────────┘
▼
pipeline.py / agent/tools.py ← one tool layer
▼
io → physics.quality → physics.calibration → survey → fitting.strategy
├─ engine (lmfit)
├─ auto (proposals)
└─ audit
▼
quantification → report (PDF · HTML · JSON)
Data model
core.Spectrum holds numpy arrays on the binding-energy axis plus a metadata dictionary with a
small set of normalized keys (source_file, element, transition, pass_energy_ev,
dwell_time_s, sweeps, source_label, …). Every reader produces it and every tool consumes it.
Spectrum.key ("<file stem>/<label>") is the identifier used everywhere, from the agent and the
API to the report.
Readers
- Kratos
.kal. A tree parser for the<id> <name> = <value>format with nested{}blocks. Spectra areF_SPECTRUMobjects. The exported kinetic energies already include the instrument work function, so binding energy ishν − KE. That was verified against the CasaXPS VAMAS export of the same file to the last digit. The sparse transmission-function table is interpolated, and linearly extrapolated, onto each spectrum. - VAMAS. A complete ISO 14976 reader: every experiment and scan mode, header inclusion flags,
experimental variables, interleaved corresponding variables (intensity and transmission) and the
1e37"not specified" sentinel. A writer for round trips is included.
Quality before fitting
Signal-to-noise is measured against the outer tenth of each window. Spikes, saturation, coarse sampling and missing coverage become explicit warnings and a grade. Smoothing exists but is never applied to data that are fitted.
Calibration
A single shift, found from a reference line (adventitious C 1s, ISO 15472 metals, a Fermi-edge fit or a user peak), is applied to every XPS spectrum, while UPS spectra are left alone. Shifts beyond a limit are refused rather than applied.
Survey identification
The identifier scores every element against the observed peak list. It only expects lines that lie inside the measured range. It accepts positions anywhere within a literature chemical-state range, rewards doublets whose observed splitting matches the reference, penalizes resolved doublets with a missing partner, single-feature matches and Auger coincidences, and applies a coarse prior against noble gases, radioactive elements and most lanthanides. Matches are claimed greedily so that one peak cannot justify two elements.
Fitting: templates, then selection, then audit
- Templates (
fitting/template.py) are pydantic-validated YAML. Doublets expand into tied pairs,{Component.param}expressions become lmfit constraints, and region-wide width tolerances tie widths to a reference component. - The engine (
fitting/engine.py) builds a composite lmfit model. It uses trust-region least squares with Poisson weights and a Levenberg–Marquardt fallback, and reports uncertainties only when they are meaningful. - Selection (
fitting/strategy.py) fits a small family of models: the template as written, the template rigidly registered to absorb charging (not applied to the calibration reference), a data-driven proposal, and residual-driven additions placed away from existing components. The lowest BIC wins, and an added component must improve BIC by at least 10. Background variants are fitted as a sensitivity check: the spread of region areas across Shirley, active Shirley and Tougaard is reported, not optimized away. - The audit (
fitting/audit.py) checks widths and positions against the curated literature, bound pinning, residual autocorrelation and runs, doublet integrity, and peak necessity by warm-started leave-one-out refits.
Literature labels are only offered when the elements a state implies were detected on the sample. "Li2O" is never suggested for a lithium-free film.
Quantification
Region areas are divided by acquisition time (dwell × sweeps), by the transmission function the
instrument wrote into the file, and by the whole-doublet sensitivity factor. physics/intensity.py
resolves the intensity scale once per run and records it (intensity_scale in results.json):
file transmission when present; a flagged pass-energy fallback when regions differ in pass
energy and no transmission exists; nothing when the sensitivity-factor library already includes
the instrument response. Per-line corrections derived from a reference sample
(physics/reference.py) and the NPL angular-asymmetry factor are optional. The composition is
reported in three labelled views (as measured, carbon excluded, overlayer corrected) and
cross-checked against the survey. Chemical states are partitioned in proportion to fitted
component areas. See docs/quantification.md for the formulas and the evidence behind each rule.
Reports
report/narrative.py turns the numbers into sentences deterministically, never with a language
model: a headline finding, a summary, three things to know, and a checklist. One Jinja template
renders both the HTML document and the PDF (WeasyPrint, with bundled Inter and JetBrains Mono). The
last page is a sign-off sheet, because a fit becomes a result only when a person who knows the
sample accepts it.
The agent
agent/tools.py turns plain Python functions into tool schemas from their signatures and
docstrings. agent/orchestrator.py is a short loop: ask the model, run the requested tools, return
compact JSON, and stop when the model answers in prose or a step limit is reached. The client speaks
the standard chat-completions protocol and also accepts models that emit a JSON tool call as text.
Every exchange is appended to events.jsonl in the session directory.
Security
The web app is a local tool without authentication: it binds to the loopback interface, issues
its own session ids, resolves every path a tool receives inside the session's upload folder, serves
scripts and fonts from the package under a strict Content-Security-Policy, and caps request sizes,
upload volume and chat rate. Template expressions are whitelisted before lmfit evaluates them.
docs/security.md has the threat model, the limits and what the stress suite exercises.
agent/llm.py resolves the endpoint from a preset (ollama, cborg, openai, custom) and
the environment, and can list an endpoint's model catalog, merging LiteLLM's /model/info when
present so tool-calling support is known before a conversation starts. See docs/assistant.md.
The MCP layer
mcp_server.py publishes the same tool registry over the Model Context Protocol, so an external
AI client can drive the analysis with its own model. Each tool's MCP schema is generated from the
registry's ToolSpec, which keeps the three front ends (CLI, workbench, agent) and the MCP server
in lockstep. Every MCP connection gets its own sandboxed Session, exactly as a browser session
does: tools may only read files inside the session's data directory, and results are the compact
JSON the built-in assistant receives. The built-in templates and the curated chemical-state table
are exposed as resources, and an analyze-sample prompt encodes the standard workflow. Two
transports are offered, stdio (xpsflow mcp) and streamable HTTP (xpsflow mcp --http, also
mounted at /mcp on the web application). See docs/mcp.md.