The shade library
50,000 samples, and a plan for every one of them.
A shade engine inherits the biases of whatever it learned from. This page is about how ours was assembled, because that decision determines who the product works for.
Why the sampling plan is the product
Convenience data fails in a predictable direction.
If you build a skin dataset from whatever imagery is easiest to obtain, you get something dense at the light end of the range and thin at the deep end. Not through anyone's intent — that is simply the distribution of what is abundant and easy to scrape.
The consequence is not evenly spread either. A model trained that way performs respectably on average and degrades exactly where it has least data, which is also exactly where sensor noise and channel clipping are worst, and where customers have the longest history of being let down. The failures compound instead of cancelling.
So we treat coverage as a specification rather than an outcome. Targets are set per Fitzpatrick type and per undertone family within each type, and a collection session that would skew the distribution is rejected rather than merged in because it was already paid for.
The honest version: this is slower and more expensive than scraping, and it is the single largest cost in building the engine. It is also the only part that cannot be retrofitted later, which is why it came first.
Balanced by construction
No band below 15% or above 18%. The small variation is real — recruitment is never perfectly even — but it is monitored against target rather than discovered afterwards.
How a session works
Controlled conditions, informed people, paid time.
Every sample in the library came from a session that ran like this. None came from scraped imagery, and none came from a brand's customers.
Before anything is recorded
Contributors get a written information sheet explaining what is being measured, what it will be used for, how long it is kept and how to withdraw. Consent is specific and recorded, not buried in a terms-of-use acceptance. Anyone can stop mid-session and still be paid.
Calibration first
A physical calibration target is in frame for every capture, and the session is validated against it before any of its data is accepted. If the target readings drift outside tolerance, the whole session is discarded — not corrected after the fact.
Contact measurement alongside imaging
A spectrophotometer takes direct readings at the same sites as the image captures. This is what makes the library a set of measurements rather than a set of inferences, and it is what every accuracy number we publish is ultimately checked against.
Multiple illuminants per contributor
The same skin is captured under several modelled light sources in one sitting. Without that, a model can learn to match faces rather than learning how a given skin behaves as the light changes — which is the entire point.
Separation and storage
Samples are stored under a pseudonymous identifier. Contact details live somewhere else entirely, linked only well enough to honour a withdrawal request. Nobody browsing the library can tell who anyone is, and there is no reason for anyone to try.
What the library is used for
- Building and validating the reflectance-recovery model
- Constructing benchmark sets that are balanced by band and undertone
- Catching regressions before a release, per band rather than on average
- Deciding where the engine should decline rather than answer
What it is never used for
- Identifying anyone, at any point, for any reason
- Training facial recognition or verification of any kind
- Inferring ethnicity, health, age or any protected characteristic
- Selling, licensing or sharing — the library does not leave us
- Any dermatological or medical assessment
Withdrawal actually works
Ask and your samples come out of the library and out of future training runs. We cannot un-train a model that already shipped, and we say so rather than implying otherwise — but nothing you contributed will inform any subsequent version.
Retention has an end date
Samples are held until withdrawal or five years from collection, whichever comes first. After that they are deleted rather than quietly archived, and the deletion is on a schedule rather than waiting for someone to remember.
Reported per band, always
Every accuracy figure we publish or hand a customer is broken out by Fitzpatrick band and undertone family. A single blended average would let us hide the exact failure this library exists to prevent, so we do not produce one.
A caveat we take seriously
Fitzpatrick is a coarse instrument.
It was devised to describe how skin responds to ultraviolet light, not to categorise people, and it flattens an enormous amount of real variation into six bins.
We use it anyway, for one reason: it is the shared vocabulary the industry already has, so reporting against it makes our claims checkable by people outside this company. A private taxonomy of our own would be more precise and completely unauditable.
What we do to compensate is stratify within each band. Undertone families — warm, neutral, olive, cool — are sampled separately inside every depth band, so the library carries far more structure than six categories suggest. Olive skin in particular sits badly on a single depth axis, and treating it as a first-class family rather than an afterthought is one of the clearer wins in the data.
The stratification, in short
Six depth bands, four undertone families within each, sampled toward a target rather than filled opportunistically.
Depth bands covered per undertone family. All twenty-four cells carry samples; none is filled by extrapolation from a neighbour.
Questions about the data?
Data-protection officers, science teams and anyone who has been burned by a shade tool before — ask us directly. We would rather answer awkward questions now than have them surface during a security review.