xpsflow

Quantification: how areas become a composition, and how far to trust it

This page explains every step between a fitted peak area and an atomic percentage, the assumptions each step makes, and what xpsflow records so the result can be audited.

1. The formula

For each element one region is used (the one with the largest count rate when several were fitted). Its corrected intensity is

I_i = A_i / (t_i · T_i · S_i · k_i · G_i)
symbol meaning where it comes from
A_i fitted area of the whole region (both spin-orbit partners) fitting/engine.py
t_i acquisition time per channel, dwell × sweeps file metadata (physics/intensity.py)
T_i analyzer transmission at the main line file transmission function, or a flagged fallback
S_i sensitivity factor of the whole doublet Scofield cross sections (elements/data/lines.json) or an instrument library
k_i optional per-line correction from a reference sample physics/reference.py
G_i optional angular-asymmetry factor NPL form, physics/quantify.py

and the atomic percentage is x_i = 100 · I_i / Σ_j I_j.

2. Count rate, and why "normalize to 100 ms" is not a separate step

Exports hold total counts per channel accumulated over sweeps scans of dwell_time_s each. Regions with different dwell times or sweep counts are only comparable as rates, so every area is divided by dwell × sweeps. Files that already contain counts per second are left alone. Normalizing to 100 ms instead of 1 s multiplies every region by the same constant and cannot change a composition; it is available as a display unit for plots and exports (xpsflow.physics.intensity.display_scale).

Before this normalization existed, a test sample with mixed 6- and 8-sweep regions came out about a factor of two wrong for some elements. The regression test for that is tests/test_validation_demo.py.

3. Pass energy and the transmission function

In constant-analyzer-energy mode the analyzer passes more electrons at a higher pass energy. The first-order expectation is area ∝ pass energy, but the real response also depends on kinetic energy, lens mode and apertures. On the Kratos Axis files used to develop xpsflow the transmission function stored by the instrument gives, for pass energy 40 relative to 20:

kinetic energy (eV) 300 600 900 1200 1400
T(40 eV) / T(20 eV) 2.1 2.4 2.6 3.0 3.3

So assuming a factor of 2 between pass energies 20 and 40 can be wrong by more than 60 %. A direct check on the same files, comparing each narrow scan with the same window in the survey (pass energy 160), gave survey-to-narrow area ratios of 8 to 10 for O 1s where the pass-energy ratio alone predicts 4 and the transmission function predicts 7.3.

xpsflow therefore applies, in this order:

  1. The transmission function from the file, per spectrum. Kratos Vision and ESCApe write one per operating mode; CasaXPS exports it as a corresponding variable in VAMAS.
  2. A pass-energy fallback (area ∝ pass energy) only when no transmission function exists and regions differ in pass energy. The run records transmission_source = "pass_energy_fallback" and warns that errors of tens of percent are expected.
  3. Nothing when the sensitivity-factor library already contains the instrument response (Wagner-type or vendor tables), so the correction is never applied twice.

What was done is stored in results.json under intensity_scale and printed in the report's "Method in one table".

4. Sensitivity factors and reference samples

Scofield cross sections are theoretical; they are right only when the transmission is right. The accuracy ceiling for routine XPS with standard factors is about 10 % relative. To do better, measure a flat, clean sample of known stoichiometry in the same session and derive per-line corrections:

xpsflow run ptfe.vms --out runs/ptfe
xpsflow reference runs/ptfe/results.json ptfe --out ptfe-factors.yaml
xpsflow run sample.vms --line-corrections ptfe-factors.yaml

The YAML carries the material, the date, the pass energies and the measured and expected compositions, because an instrument factor is only valid for the operating mode it was measured in (ISO 5861, ISO 18118). Built-in materials: PTFE, PET, PMMA, LiF and SiO₂ (physics/reference.py). ISO 15472 energy-scale checks against Au 4f7/2, Ag 3d5/2 and Cu 2p3/2 are in the same module.

5. Three readings of one composition

Air-exposed surfaces carry one to two nanometers of adventitious hydrocarbon, which can dominate the equivalent homogeneous composition. No standard prescribes a single correction, so xpsflow reports three labelled views:

view what it assumes
as measured nothing: every detected element, the surface as found
carbon excluded all carbon is contamination; the rest is renormalized to 100 %
overlayer corrected carbon and an O/C ≈ 0.11 share of the oxygen belong to the overlayer (ISO 22581 spirit); an estimate with > 20 % relative uncertainty

Quote the as-measured view for the surface and the carbon-excluded view for the material, and say which one you used.

6. The survey as a cross-check

The survey is quantified too (one line per element, local Shirley background, Scofield factors, the survey's own transmission function). Its coarse step, lower resolution and crude background endpoints make it less precise than the narrow scans, but it is acquired at a high pass energy where scattering matters less, and it covers every element. The pipeline compares the two compositions over the elements they share and warns when they disagree by more than the configured tolerance; the usual causes are a hidden overlap in the survey, a sample that changed between scans, or an intensity-scale problem.

7. Angular asymmetry

The Kratos Axis geometry places the X-ray source 60° from the analyzer axis, not at the magic angle, so the dipole asymmetry of each subshell changes the measured intensities slightly. xpsflow can divide by the NPL factor

G = 1 + ½ (0.69 β) (3/2 sin²γ − 1)

with β = 2 for s subshells (exact in the dipole approximation) and user-supplied β for others. It is off by default and is only applied when a parameter exists for every line used; with s levels only the factor is 1.086 at 60°.

8. What makes BIC comparable

Model selection relies on the Bayesian information criterion, which needs the noise model to be right. Raw counts follow Poisson statistics, but averaged sweeps, rate data or rescaled exports do not have σ = √y. The engine estimates one scale factor from the second differences of the data (fitting/engine.py: estimate_noise_scale) and weights the fit with σ = s·√y, so χ² and BIC keep their meaning on any export. Each candidate model is reported with a probability (BIC weight) so a near tie is visible instead of hidden behind a single winner.

References