Replications

Thin-film stacks read out of the open-access literature by Reviewer3 and recomputed here with a transfer-matrix solver. Each record opens in the calculator with the stack preloaded. Spotted a score that looks wrong, or a stack we have read incorrectly? Send us the details and we will re-check it.

Method, and how to read a record

How a record is built. Reviewer3’s extraction pipeline reads each paper’s PDF and records the stack, the claim window, and any optical constants stated in the text. We recompute the spectrum from that reading with a transfer-matrix solver and compare it against a calculated R/T/A curve digitized from one of the paper’s own figures. Measured curves appear for context only. Where the published inputs support no comparable curve, the paper is listed in the appendix at the foot of this page.

What the fit score measures. The residual carries both the reading and the optical data, so a low score can come from a misread layer as easily as from the physics. Publisher figures appear only where the paper’s licence permits display. These records describe reproducibility, and say nothing about the authors.

A continuous 0–1 score compares the recomputed spectrum with the digitized calculated curve. It uses span-normalized RMSE with a maximum-deviation guard.

Every curve the paper published is scored, and they pool into one number. Reflectance, transmittance and absorptance are not three verdicts about a stack — they are three statements about the same one, and R + T + A = 1, so matching reflectance while transmittance floats free pins nothing down. The per-curve residuals are combined in quadrature and the maximum deviation is taken across all of them, so one badly missed curve still caps the score instead of being averaged away, and a paper that published three curves is not scored more harshly than one that published a single curve. Where more than one curve was digitized, the record lists each one and the spread between them.

Hatched: one or more optical inputs are substituted, so the fit cannot be attributed to the paper. Whisker: the worst-to-best fit across covering literature datasets for a matched material. Spread: where a paper published several curves, how far the per-curve fits disagree. A wide spread means the reproduction contradicts itself rather than being uniformly approximate, and is worth more attention than the headline alone.

Where the optical constants behind the curve came from. This runs independently of the fit, because a close match built on a substituted index reflects the proxy as much as the paper’s own recipe.

Documented Every input traces to a dataset the paper itself names.
Matched The paper names the material but not a dataset. Every covering same-material dataset is evaluated and the worst-to-best range is published beside the score.
Substituted No usable published data exists for at least one of the paper’s materials, so a named stand-in was used in its place: for instance, measured TiO2 for a non-stoichiometric TiOx. The fit is hatched, because part of what it measures is the stand-in.

The corpus at a glance

Loading summary…

Literature records

Not computable from the published inputs Loading records...

These records lack a comparable calculated curve, so they carry no fit score. A missing input says nothing about whether the published calculation is correct.

DOIPaperReason