# Quantification: how areas become a composition, and how far to trust it

This page explains every step between a fitted peak area and an atomic percentage, the
assumptions each step makes, and what xpsflow records so the result can be audited.

## 1. The formula

For each element one region is used (the one with the largest count rate when several were
fitted). Its corrected intensity is

```
I_i = A_i / (t_i · T_i · S_i · k_i · G_i)
```

| symbol | meaning | where it comes from |
|---|---|---|
| `A_i` | fitted area of the whole region (both spin-orbit partners) | `fitting/engine.py` |
| `t_i` | acquisition time per channel, dwell × sweeps | file metadata (`physics/intensity.py`) |
| `T_i` | analyzer transmission at the main line | file transmission function, or a flagged fallback |
| `S_i` | sensitivity factor of the whole doublet | Scofield cross sections (`elements/data/lines.json`) or an instrument library |
| `k_i` | optional per-line correction from a reference sample | `physics/reference.py` |
| `G_i` | optional angular-asymmetry factor | NPL form, `physics/quantify.py` |

and the atomic percentage is `x_i = 100 · I_i / Σ_j I_j`.

## 2. Count rate, and why "normalize to 100 ms" is not a separate step

Exports hold total counts per channel accumulated over `sweeps` scans of `dwell_time_s` each.
Regions with different dwell times or sweep counts are only comparable as *rates*, so every
area is divided by `dwell × sweeps`. Files that already contain counts per second are left
alone. Normalizing to 100 ms instead of 1 s multiplies every region by the same constant and
cannot change a composition; it is available as a display unit for plots and exports
(`xpsflow.physics.intensity.display_scale`).

Before this normalization existed, a test sample with mixed 6- and 8-sweep regions came out
about a factor of two wrong for some elements. The regression test for that is
`tests/test_validation_demo.py`.

## 3. Pass energy and the transmission function

In constant-analyzer-energy mode the analyzer passes more electrons at a higher pass energy.
The first-order expectation is *area ∝ pass energy*, but the real response also depends on
kinetic energy, lens mode and apertures. On the Kratos Axis files used to develop xpsflow the
transmission function stored by the instrument gives, for pass energy 40 relative to 20:

| kinetic energy (eV) | 300 | 600 | 900 | 1200 | 1400 |
|---|---|---|---|---|---|
| T(40 eV) / T(20 eV) | 2.1 | 2.4 | 2.6 | 3.0 | 3.3 |

So assuming a factor of 2 between pass energies 20 and 40 can be wrong by more than 60 %.
A direct check on the same files, comparing each narrow scan with the same window in the
survey (pass energy 160), gave survey-to-narrow area ratios of 8 to 10 for O 1s where the
pass-energy ratio alone predicts 4 and the transmission function predicts 7.3.

xpsflow therefore applies, in this order:

1. **The transmission function from the file**, per spectrum. Kratos Vision and ESCApe write
   one per operating mode; CasaXPS exports it as a corresponding variable in VAMAS.
2. **A pass-energy fallback** (`area ∝ pass energy`) only when no transmission function exists
   and regions differ in pass energy. The run records `transmission_source =
   "pass_energy_fallback"` and warns that errors of tens of percent are expected.
3. **Nothing** when the sensitivity-factor library already contains the instrument response
   (Wagner-type or vendor tables), so the correction is never applied twice.

What was done is stored in `results.json` under `intensity_scale` and printed in the report's
"Method in one table".

## 4. Sensitivity factors and reference samples

Scofield cross sections are theoretical; they are right only when the transmission is
right. The accuracy ceiling for routine XPS with standard factors is about 10 % relative.
To do better, measure a flat, clean sample of known stoichiometry in the same session and
derive per-line corrections:

```
xpsflow run ptfe.vms --out runs/ptfe
xpsflow reference runs/ptfe/results.json ptfe --out ptfe-factors.yaml
xpsflow run sample.vms --line-corrections ptfe-factors.yaml
```

The YAML carries the material, the date, the pass energies and the measured and expected
compositions, because an instrument factor is only valid for the operating mode it was measured
in (ISO 5861, ISO 18118). Built-in materials: PTFE, PET, PMMA, LiF and SiO₂
(`physics/reference.py`). ISO 15472 energy-scale checks against Au 4f7/2, Ag 3d5/2 and
Cu 2p3/2 are in the same module.

## 5. Three readings of one composition

Air-exposed surfaces carry one to two nanometers of adventitious hydrocarbon, which can
dominate the equivalent homogeneous composition. No standard prescribes a single correction, so
xpsflow reports three labelled views:

| view | what it assumes |
|---|---|
| as measured | nothing: every detected element, the surface as found |
| carbon excluded | all carbon is contamination; the rest is renormalized to 100 % |
| overlayer corrected | carbon and an O/C ≈ 0.11 share of the oxygen belong to the overlayer (ISO 22581 spirit); an estimate with > 20 % relative uncertainty |

Quote the as-measured view for the surface and the carbon-excluded view for the material, and
say which one you used.

## 6. The survey as a cross-check

The survey is quantified too (one line per element, local Shirley background, Scofield
factors, the survey's own transmission function). Its coarse step, lower resolution and crude
background endpoints make it less precise than the narrow scans, but it is acquired at a high
pass energy where scattering matters less, and it covers every element. The pipeline compares
the two compositions over the elements they share and warns when they disagree by more than
the configured tolerance; the usual causes are a hidden overlap in the survey, a sample that
changed between scans, or an intensity-scale problem.

## 7. Angular asymmetry

The Kratos Axis geometry places the X-ray source 60° from the analyzer axis, not at the magic
angle, so the dipole asymmetry of each subshell changes the measured intensities slightly.
xpsflow can divide by the NPL factor

```
G = 1 + ½ (0.69 β) (3/2 sin²γ − 1)
```

with β = 2 for s subshells (exact in the dipole approximation) and user-supplied β for
others. It is off by default and is only applied when a parameter exists for every line used;
with s levels only the factor is 1.086 at 60°.

## 8. What makes BIC comparable

Model selection relies on the Bayesian information criterion, which needs the noise model to
be right. Raw counts follow Poisson statistics, but averaged sweeps, rate data or rescaled
exports do not have σ = √y. The engine estimates one scale factor from the second differences
of the data (`fitting/engine.py: estimate_noise_scale`) and weights the fit with σ = s·√y, so
χ² and BIC keep their meaning on any export. Each candidate model is reported with a
probability (BIC weight) so a near tie is visible instead of hidden behind a single winner.

## References

- A. G. Shard, "Practical guides for x-ray photoelectron spectroscopy: Quantitative XPS",
  J. Vac. Sci. Technol. A 38, 041201 (2020).
- B. P. Reed et al., "Versailles Project on Advanced Materials and Standards interlaboratory
  study on intensity calibration for x-ray photoelectron spectroscopy instruments using
  low-density polyethylene", J. Vac. Sci. Technol. A 38, 063208 (2020); ISO 5861:2024.
- G. H. Major et al., "Practical guide for curve fitting in x-ray photoelectron spectroscopy",
  J. Vac. Sci. Technol. A 38, 061203 (2020).
- M. P. Seah, I. S. Gilmore and S. J. Spencer, NPL average-matrix relative sensitivity factors.
- G. Greczynski and L. Hultman, "X-ray photoelectron spectroscopy: Towards reliable binding
  energy referencing", Prog. Mater. Sci. 107, 100591 (2020).
- ISO 15472:2010, ISO 18118:2024, ISO 19830:2015, ISO 22581:2021.
