# Constraints, provenance and credibility

A curve fit in XPS is a hypothesis dressed in numbers. The lineshape, the position bounds, the
width bounds and the spin–orbit ties decide what the fit *can* say, so they must be visible,
sourced and reviewable. This page describes how xpsflow records where every constraint comes
from, how the reference database is built, and what the report prints so a reader can judge
the result without reading code.

## 1. Every constraint has a provenance

Each component in a template may carry a `provenance` block:

```yaml
- name: Te-M
  chemical_state: telluride / metallic Te
  provenance: {source: bahl1978, method: literature}
  center: {value: 572.8, min: 571.8, max: 573.6}
  fwhm:   {value: 1.0, min: 0.6, max: 1.6}
  doublet: {splitting: 10.39, ratio: 0.6667}
```

| field | meaning |
|---|---|
| `source` | a key in the bibliography (`src/xpsflow/elements/data/bibliography.yaml`) or a free-text citation |
| `method` | `literature` (tabulated values), `instrument` (measured on this instrument), `user` (set by the analyst), `data` (proposed from the spectrum), `default` (package default, no specific source) |
| `note` | anything a reviewer should know |

Provenance is inherited: a component without its own block takes the region's, then the
template's first entry in `references`. Components that xpsflow adds itself are always marked
`data`, with the literature state they were matched to (if any) and the residual feature that
motivated them. The shipped templates cite a source for every component; a test
(`tests/test_provenance.py`) fails if a template or a database entry cites a key that is not in
the bibliography.

## 2. What the report prints

The page "Where the constraints come from" lists, per region and per component, the position
bounds, width bounds, doublet ties, lineshape and origin, followed by a reporting record in the
spirit of ISO 19830 (minimum reporting requirements for peak fitting): instrument and source,
pass energy, step, dwell and sweeps, charge neutralization, energy referencing method and
shift, background type and window, lineshapes, widths, R² and reduced χ², and the software
version. The same record is in `results.json` under `provenance`, and the bibliography of
every source actually used in the run is printed at the end.

## 3. How the reference database is built

`src/xpsflow/elements/data/chemical_states.yaml` holds, for 53 elements, the principal core
levels with their spin–orbit splitting and ratio, typical widths, a conservative fit window and
the chemical states with binding energies, ranges and a `source` key. Rules:

1. **Every state cites a source.** The schema rejects a state whose `source` is not in the
   bibliography. Prefer primary literature (Biesinger's transition-metal series, Beamson and
   Briggs for polymers, Moulder's handbook for elements) and cite NIST SRD 20 per entry.
2. **Ranges, not points.** `be_range_ev` records the spread across the cited sources and across
   instruments; the audit uses the range, the central value is only a starting point.
3. **Charge referencing is stated.** Values are on the adventitious C 1s = 284.8 eV scale (or the
   Fermi level for conductors). That scale is a convention, not a physical constant: on
   conductors the adventitious carbon peak moves by more than 1 eV between substrates
   (Greczynski and Hultman 2020). xpsflow records the referencing method and shift in every
   run, supports Fermi-edge and ISO 15472 metallic references, and flags shifts beyond a limit.
4. **NIST SRD 20 is linked, not copied.** The database is a copyrighted Standard Reference
   Database with no bulk export. Entries cite it; nothing is redistributed from it.
5. **Context matters.** A literature state is only offered as a label when the elements it
   implies were detected on the sample (no "Li₂O" on a lithium-free film).

To add an entry: add the citation to `bibliography.yaml`, add the state with its range under the
element's core level, run the tests, and open a pull request that quotes the source.

## 4. Fitting practice encoded in the audit

The audit (`fitting/audit.py`) checks each fit against the guidance in Major et al. 2020: widths
inside the typical range for the core level, no parameter pinned on a bound, residuals without
structure, doublets intact, and every component necessary (removing it must worsen BIC by at
least 10). Reduced χ² is only meaningful because the noise scale is calibrated from the data
(`docs/quantification.md`, section 8).

## 5. What is still a judgement call

Templates encode one view of a material system. Differential charging, unexpected species,
and multiplet-split transition-metal lines can all defeat a template; the model selection and
the audit are designed to make that visible, not to hide it. The last page of every report is a
sign-off because a fit becomes a result only when someone who knows the sample accepts it.
