xpsflow

Architecture and design decisions

xpsflow is built around one rule: deterministic tools compute; everything else only calls them. The CLI, the web workbench and the conversational agent all sit on the same tool layer, so a number seen in a chat is the number in the PDF, and either can be reproduced from results.json and the saved templates.

            ┌──────────────┐   ┌──────────────┐   ┌──────────────────────┐
  people →  │  CLI (typer) │   │ web (FastAPI)│   │ agent (tool calling) │
            └──────┬───────┘   └──────┬───────┘   └──────────┬───────────┘
                   └──────────────────┼──────────────────────┘
                                      ▼
                       pipeline.py / agent/tools.py   ← one tool layer
                                      ▼
   io → physics.quality → physics.calibration → survey → fitting.strategy
                                                            ├─ engine (lmfit)
                                                            ├─ auto (proposals)
                                                            └─ audit
                                      ▼
                      quantification → report (PDF · HTML · JSON)

Data model

core.Spectrum holds numpy arrays on the binding-energy axis plus a metadata dictionary with a small set of normalized keys (source_file, element, transition, pass_energy_ev, dwell_time_s, sweeps, source_label, …). Every reader produces it and every tool consumes it. Spectrum.key ("<file stem>/<label>") is the identifier used everywhere, from the agent and the API to the report.

Readers

Quality before fitting

Signal-to-noise is measured against the outer tenth of each window. Spikes, saturation, coarse sampling and missing coverage become explicit warnings and a grade. Smoothing exists but is never applied to data that are fitted.

Calibration

A single shift, found from a reference line (adventitious C 1s, ISO 15472 metals, a Fermi-edge fit or a user peak), is applied to every XPS spectrum, while UPS spectra are left alone. Shifts beyond a limit are refused rather than applied.

Survey identification

The identifier scores every element against the observed peak list. It only expects lines that lie inside the measured range. It accepts positions anywhere within a literature chemical-state range, rewards doublets whose observed splitting matches the reference, penalizes resolved doublets with a missing partner, single-feature matches and Auger coincidences, and applies a coarse prior against noble gases, radioactive elements and most lanthanides. Matches are claimed greedily so that one peak cannot justify two elements.

Fitting: templates, then selection, then audit

  1. Templates (fitting/template.py) are pydantic-validated YAML. Doublets expand into tied pairs, {Component.param} expressions become lmfit constraints, and region-wide width tolerances tie widths to a reference component.
  2. The engine (fitting/engine.py) builds a composite lmfit model. It uses trust-region least squares with Poisson weights and a Levenberg–Marquardt fallback, and reports uncertainties only when they are meaningful.
  3. Selection (fitting/strategy.py) fits a small family of models: the template as written, the template rigidly registered to absorb charging (not applied to the calibration reference), a data-driven proposal, and residual-driven additions placed away from existing components. The lowest BIC wins, and an added component must improve BIC by at least 10. Background variants are fitted as a sensitivity check: the spread of region areas across Shirley, active Shirley and Tougaard is reported, not optimized away.
  4. The audit (fitting/audit.py) checks widths and positions against the curated literature, bound pinning, residual autocorrelation and runs, doublet integrity, and peak necessity by warm-started leave-one-out refits.

Literature labels are only offered when the elements a state implies were detected on the sample. "Li2O" is never suggested for a lithium-free film.

Quantification

Region areas are divided by acquisition time (dwell × sweeps), by the transmission function the instrument wrote into the file, and by the whole-doublet sensitivity factor. physics/intensity.py resolves the intensity scale once per run and records it (intensity_scale in results.json): file transmission when present; a flagged pass-energy fallback when regions differ in pass energy and no transmission exists; nothing when the sensitivity-factor library already includes the instrument response. Per-line corrections derived from a reference sample (physics/reference.py) and the NPL angular-asymmetry factor are optional. The composition is reported in three labelled views (as measured, carbon excluded, overlayer corrected) and cross-checked against the survey. Chemical states are partitioned in proportion to fitted component areas. See docs/quantification.md for the formulas and the evidence behind each rule.

Reports

report/narrative.py turns the numbers into sentences deterministically, never with a language model: a headline finding, a summary, three things to know, and a checklist. One Jinja template renders both the HTML document and the PDF (WeasyPrint, with bundled Inter and JetBrains Mono). The last page is a sign-off sheet, because a fit becomes a result only when a person who knows the sample accepts it.

The agent

agent/tools.py turns plain Python functions into tool schemas from their signatures and docstrings. agent/orchestrator.py is a short loop: ask the model, run the requested tools, return compact JSON, and stop when the model answers in prose or a step limit is reached. The client speaks the standard chat-completions protocol and also accepts models that emit a JSON tool call as text. Every exchange is appended to events.jsonl in the session directory.

Security

The web app is a local tool without authentication: it binds to the loopback interface, issues its own session ids, resolves every path a tool receives inside the session's upload folder, serves scripts and fonts from the package under a strict Content-Security-Policy, and caps request sizes, upload volume and chat rate. Template expressions are whitelisted before lmfit evaluates them. docs/security.md has the threat model, the limits and what the stress suite exercises. agent/llm.py resolves the endpoint from a preset (ollama, cborg, openai, custom) and the environment, and can list an endpoint's model catalog, merging LiteLLM's /model/info when present so tool-calling support is known before a conversation starts. See docs/assistant.md.

The MCP layer

mcp_server.py publishes the same tool registry over the Model Context Protocol, so an external AI client can drive the analysis with its own model. Each tool's MCP schema is generated from the registry's ToolSpec, which keeps the three front ends (CLI, workbench, agent) and the MCP server in lockstep. Every MCP connection gets its own sandboxed Session, exactly as a browser session does: tools may only read files inside the session's data directory, and results are the compact JSON the built-in assistant receives. The built-in templates and the curated chemical-state table are exposed as resources, and an analyze-sample prompt encodes the standard workflow. Two transports are offered, stdio (xpsflow mcp) and streamable HTTP (xpsflow mcp --http, also mounted at /mcp on the web application). See docs/mcp.md.