← Writing · AI Automation
Flux Working Paper No. 41

The Word "Personalized" Is Doing Two Different Jobs

Ken Ruto · Flux (FluxImpact) · September 2026 · 31 min
↩ Read as essay
BibTeX · RIS
Abstract

Skincare has sold "personalization" for a decade — quizzes, skin-type sliders, AI-matched routines — and most of what it delivers is cosmetic segmentation: oily or dry, budget tier, fragrance preference. Almost none of it is the specific personalization that melanin-rich skin actually needs, which is a routine matched against a documented, disproportionate risk: post-inflammatory hyperpigmentation (PIH), a complication that is measurably more common, more severe, and slower to resolve in darker Fitzpatrick phototypes. This paper traces that gap to its source. Dermatology's own trial base, textbook imagery, and — now — the image datasets training AI dermatology tools all underrepresent skin of color by a repeatedly measured margin. We argue that Zamia, Flux's AI skincare consultant, is a real attempt at the second, harder kind of personalization, deliberately "boxed" so a language model narrates while a fixed, vetted product-and-evidence engine decides. We do not treat that as self-evidently true. We hold Flux's own commercial stake in this conclusion up for scrutiny, take seriously the objection that "personalized skincare" is mostly marketing language borrowing the authority of science, and set out the regulatory line — cosmetic advice, not diagnosis — that a tool like Zamia has to stay on the right side of. We close with falsification conditions and an honest accounting of what the evidence does not yet show.

Keywords: personalization, post-inflammatory hyperpigmentation, skin of color, clinical trial representation, cosmetic regulation, AI skincare, melanin-rich skin, Zamia

Type "personalized skincare" into any app store and you'll get the same six questions back, in some order: oily, dry, or combination. Sensitive, yes or no. Budget: under $30, $30–60, whatever's next. Answer them and an algorithm sorts you into a bucket that maybe forty thousand other people are also sitting in, and calls the bucket a routine made for you.

That's not nothing. Matching a moisturizer to an oil level is a real problem and a real fix. But it is not the personalization that the highest-stakes gap in skincare actually calls for, and calling it that has let the industry claim credit for solving a problem it never measured.

WP21 made the diagnostic case for why melanin-rich skin needs something built for it specifically — glow is not a purchasable finish, healthy skin reacts to irritation differently in deep pigmentation, and the market's historical answer to that difference was to sell bleaching instead of care.1 This paper is about the evidence underneath that case, and about what it takes for an AI consultant to actually close the gap instead of relabeling it.

Here is the distinction the rest of this paper turns on. Call the first kind segmentation-personalization: sorting people by attributes that mostly don't change how urgently a product needs to be right — skin type, scent preference, price point. Call the second kind risk-personalization: matching recommendations against a complication that is unevenly distributed by skin type in a way that has real, measured clinical consequences if you get it wrong. Post-inflammatory hyperpigmentation is the clearest case of the second kind in skincare, and it is disproportionately a risk for deeper Fitzpatrick phototypes. The industry has spent a decade getting very good at the first kind of personalization and calling it the second.

KANAIRO://WP41 — TWO JOBS, ONE WORD FITZPATRICK IV PIH RISK AND SEVERITY TRIAL & TEXTBOOK COVERAGE THE SEGMENT PERSONALIZATION DOESN'T PRICE IN LIGHT SKIN DEEP SKIN THE RISK CLIMBS EXACTLY WHERE THE EVIDENCE BASE THINS OUT. The risk climbs exactly where the evidence base thins out.

1. Scope, method, and standing

This paper argues from four kinds of source, and only these: published dermatology and cosmetic-science literature on skin-of-color representation in clinical trials, textbooks, and imaging datasets; the clinical literature on post-inflammatory hyperpigmentation specifically; public health and regulatory documents on skin-lightening products and on the regulation of cosmetic versus medical claims; and one existing systematic review of AI-driven personalized skincare research. Every empirical claim below is footnoted to a source we actually retrieved and read; where we could not verify a number to our own satisfaction, it is marked CITATION-NEEDED in Appendix B rather than stated as fact. Press-reported market figures are labeled as order-of-magnitude estimates, not precise rates, because that is what they are.

This paper does not argue from Zamia's own usage data. We have not run a controlled comparison of Zamia against a generic personalization app, we have no outcome data on PIH incidence among Zamia users, and nothing here should be read as evidence that Zamia works better than any specific competitor. That comparison would require a trial we have not run. What we can establish is narrower and, we think, still worth establishing: that the representation gap in the underlying clinical and testing literature is real and measured, that PIH is a real and disproportionate risk for darker phototypes, and that an AI consultant's architecture — whether it is genuinely constrained to a vetted evidence engine or is free to improvise — determines whether it can even in principle close that gap rather than just repackage it.

We are also not neutral. Flux builds and sells Zamia. That is a direct commercial interest in the conclusion that risk-personalization matters and that Zamia is a serious attempt at it, and the reader should weigh everything that follows with that in mind. We address the objection directly in §7 rather than disclosing it once and moving on.

What would refute this paper's thesis, stated plainly: if the PIH literature turned out to be weaker or more contested than it presents as being — if darker phototypes did not, in fact, show measurably worse and more persistent post-inflammatory pigmentation — the central premise collapses. If the representation gaps in trials, textbooks, and datasets turned out to be closing fast enough that they no longer describe today's evidence base, the "the base was never built for this" framing goes stale. And if it turned out that mainstream personalization tools already do quietly price in pigmentary risk — that the quiz behind a $40 serum subscription is already asking about scarring history and adjusting active-ingredient concentration accordingly — then the "two jobs" distinction this paper rests on is a distinction without a difference, and the paper's contribution is much smaller than claimed. We have looked for that evidence and not found it; we say more about the search in §7 and §8.

2. Segmentation is not the personalization at stake

Every personalization quiz asks about skin type because skin type is easy to self-report, easy to act on with existing product lines, and low-risk to get slightly wrong. Tell an app you're "combination" when you're actually "oily," and worst case you get a slightly heavier moisturizer. That's a soft failure.

PIH-aware personalization is a different kind of problem. Get it wrong — recommend an aggressive retinoid concentration, an unbuffered acid exfoliant, a treatment protocol lifted from a study population that was mostly Fitzpatrick I–III — and the failure mode for someone with deeper pigmentation isn't mild irritation that resolves in a day. It's a pigmented mark that can sit on the skin for months to years after the original irritation has healed.2 That is not a symmetrical risk across skin types, and treating it as though a skin-type-and-budget quiz already covers it is the error this paper is naming.

The reason mainstream "personalization" doesn't ask about PIH history, scarring tendency, or prior hyperpigmentation the way it asks about oiliness is not malice. It's that the underlying product-and-evidence base most skincare brands build from was assembled, tested, and photographed on a study population that skews light, and you can't personalize against a risk your evidence base didn't systematically measure. That's the claim the next section grounds.

3. The measured gap: who dermatology's evidence was built on

Start with clinical trials. A 2023 analysis of three core dermatology surgery textbooks found that of 1,501 clinical images, only 5.6% depicted Fitzpatrick phototypes IV–VI — the phototypes most associated with visible skin of color — and 37.9% of surgical topics covered included no skin-of-color images at all.3 That's not a niche finding. It's echoed by an earlier, separately conducted analysis of six major dermatology textbooks, which found the share of images depicting dark skin ranging from 4% to 18% across texts, and — more damning for any argument that this is old news already fixed — found that only one of the six textbooks showed more than a one-percentage-point increase in Fitzpatrick V/VI representation between 2006 and 2020.4 Fourteen years, and the needle barely moved in five of six books.

This is not confined to textbooks. A 2025 review of melasma clinical trials — melasma being a pigmentary condition that disproportionately affects people with more melanin — found that Fitzpatrick skin types III and IV together made up more than 75% of trial participants.5 Melasma is a condition where deeper phototypes are, if anything, overrepresented among the people who actually have it; even here, the trial population clusters toward the lighter end of the range it should be studying.

The pattern repeats again, and more starkly, in the datasets now training AI dermatology tools — the same kind of imaging data that would underlie any AI skincare consultant's visual assessment layer, if it had one. A 2022 systematic review of publicly available skin-cancer image datasets found that ethnicity data existed for only 1.3% of images and Fitzpatrick skin type for only 2.1%.6 Of the 2,436 images across the three datasets that did report skin type, ten were Fitzpatrick V and exactly one was Fitzpatrick VI.6 Eleven images, out of thousands, representing the deepest end of the phototype scale, in the data that AI tools are trained and validated against.

None of these numbers describe malice by any single actor. They describe an evidence base built, image by image and trial by trial, on a study population that was easiest and cheapest to recruit and photograph in the institutions where dermatology research has historically concentrated. But an evidence base with that shape cannot, by construction, tell you as much about how a treatment performs on deeper skin as it tells you about lighter skin — and any personalization claim built on top of it inherits the same blind spot, whether or not the app admits it.

4. Why post-inflammatory hyperpigmentation specifically deserves this attention

PIH is not a rare or marginal complication for people with deeper skin. In one of the foundational reviews of the condition, pigmentary disorders (excluding vitiligo) ranked as the third most common diagnosis among Black American dermatology patients in an early-1980s study, at roughly 9% of visits, versus seventh and roughly 1.7% among white patients in the same analysis; a later, 2007 study again found dyschromias as the second most common diagnosis among Black patients while the category failed to place in the top ten diagnoses at all for white patients.2 The same review reports that among patients who developed acne-related PIH, 65.3% were Black, compared with 52.7% Hispanic and 47.4% Asian patients in the cohorts studied — again, not a uniform risk.2

The mechanism is now reasonably well characterized. Inflammation — from acne, eczema, an aggressive product, a cosmetic procedure — releases cytokines and reactive oxygen species that stimulate melanocyte activity and melanin synthesis; in skin with a higher baseline melanocyte and tyrosinase activity, that inflammatory signal produces more pigment, transferred more efficiently into surrounding skin.7 The same review notes that dark skin also shows measurable differences in barrier structure — reduced terminal keratinocyte differentiation, weaker basement-membrane markers at the dermal-epidermal junction — that compound the risk once inflammation starts.7 This is not a claim that darker skin is more fragile in some vague sense. It's a specific, mechanistic account of why the same irritant event produces a different, more visible, and more durable outcome depending on phototype.

And durable is the operative word. Epidermal PIH can take months to resolve untreated; dermal PIH — pigment that has migrated deeper — can persist for a protracted period or, in some cases, be effectively permanent.2 A skincare recommendation that triggers PIH in someone with deeper pigmentation is not producing a cosmetic inconvenience that fades by the next photo. It is producing a mark that can outlast the product that caused it by a year or more, and for the person wearing it, that is not incidental to how well a routine "worked."

This is the risk that segmentation-personalization does not ask about, and it is the risk that a routine built without reference to skin-of-color-specific evidence is structurally unable to price in — not because any individual formulator wants that outcome, but because the trial and textbook base most formulations, marketing claims, and even AI training sets were built against did not systematically include the population for whom this risk is highest.

5. The bad answer the industry already tried, and why it wasn't personalization either

For decades, the market's dominant response to "skin of color has different needs" was not personalization. It was brightening — a euphemism, in a meaningful share of cases, for lightening or bleaching. This is the history WP21 named. It's worth being precise here about what the evidence actually says, because this is a place where the paper's thesis could collapse into moralizing if it isn't grounded.

The clearest, hardest evidence is about mercury. The World Health Organization's 2019 information note on mercury in skin-lightening products states plainly that mercury-containing lightening products are hazardous to health and have been banned in numerous countries, and that despite those bans, such products continue to circulate through internet sales and informal trade — which is precisely why the WHO frames this as requiring coordinated customs enforcement and public health messaging under the Minamata Convention, not just a labeling rule.8 The U.S. FDA has separately and repeatedly warned consumers about skin products containing mercury or hydroquinone specifically, flagging the combination as a recurring health-fraud pattern rather than an isolated bad actor.9 The European Union banned hydroquinone from cosmetic formulations outright — it sits in Annex II of the EU cosmetics regulation as a prohibited substance, with a narrow, unrelated exception for artificial nail systems that does not extend to skin care.10 Long-term or improper hydroquinone use is associated with exogenous ochronosis, a treatment-resistant blue-black discoloration — which is to say, the "fix" for pigmentation risk can itself produce a pigmentation disorder.10

Enforcement testing gives some sense of how much of this persists in the market despite the bans, though these figures are testing-program findings on products flagged for scrutiny, not random samples of the global market, and should be read as order-of-magnitude indicators rather than population rates: in one European testing program, roughly 18% of products tested were non-compliant specifically due to the presence of prohibited mercury, hydroquinone, or corticosteroids.11 The market these products sit in is not small — industry market-research estimates put the global skin-lightening products market at roughly $11 billion in 2023, projected toward the mid-teens of billions by 2030;12 we flag this figure explicitly as a press-reported, order-of-magnitude commercial estimate, not a scientific measurement, because that is the kind of source it is.

The reason this matters for the personalization argument, rather than being a separate topic, is this: bleaching was never personalization, and it was never trying to be. It was a single, universal intervention — reduce the melanin — applied regardless of what any individual's skin actually needed, marketed as care. The correct response to "the mainstream market underserves darker skin" is not "sell darker skin a different universal product." It's the harder, more specific thing: build recommendations that account for how darker skin actually responds to irritation, without ever treating the pigmentation itself as the thing to be corrected. That is the line WP21 draws as "radiance, never lightness," and this paper's job is to show that the line has to be enforced not just as a brand promise but as a structural constraint on what a recommendation engine is allowed to output.

6. What "boxed" has to buy, and the regulatory line it has to sit inside

Zamia's architecture, per WP21, separates two jobs: a language model that carries the conversation — empathetic, plain-spoken, reassuring — and a fixed, vetted product-and-evidence engine underneath that actually decides what gets recommended. The model narrates; it does not invent the recommendation.1 The point of that separation, for this paper's purposes, is narrower than "it's safer" in the abstract. It's that the separation is the only architecture under which the evidence gaps described in §§3–4 can be corrected deliberately, on a known and auditable basis, rather than reproduced silently.

An unconstrained model asked "what should I use for these dark marks after a breakout" will answer from whatever it absorbed in training, which — per the dataset findings in §3 — is disproportionately drawn from a literature and image base that underrepresents exactly the skin type asking the question. A model that is boxed against a specific, chosen evidence engine can, in principle, be pointed deliberately at the skin-of-color-specific literature this paper cites, rather than at the average of everything on the internet about skincare. Whether Zamia's engine is actually built that way, curated against sources like the ones footnoted here rather than against generic skincare content, is an empirical claim about Flux's own product that we have not independently audited for this paper, and we say so plainly rather than asserting it.

There is also a regulatory line that determines how far "boxed" can go before it becomes something else entirely, and it's worth being precise about where that line sits rather than treating "not medical advice" as a throwaway disclaimer. The U.S. FDA's general wellness policy exempts software intended for maintaining or encouraging a healthy lifestyle — and unrelated to diagnosing, curing, mitigating, preventing, or treating a disease or condition — from regulation as a medical device.13 That is the carve-out a cosmetic AI consultant has to stay inside: routine and product recommendations, yes; diagnosing a skin condition, adjudicating whether a lesion is cancerous, or making a disease-treatment claim, no. The moment an AI skincare tool starts telling someone what a mark on their skin is, rather than how to care for skin generally, it has crossed from cosmetic advice into something the FDA's own framework treats as a different, higher-stakes category, requiring a different kind of validation than this paper's evidence base can speak to.

The EU's cosmetics claims regime draws a parallel but distinct line: Article 20 of the EU cosmetics regulation prohibits claims implying characteristics a product doesn't have, and the "evidential support" criterion under Commission Regulation 655/2013 requires that any claim — explicit or implicit — be backed by adequate, verifiable evidence, held in a documented product information file.14 An AI consultant that tells a user "this will fix your hyperpigmentation" is making exactly the kind of claim that regime is built to test, and "the AI said so" is not itself evidence under that standard; the evidence has to sit underneath the claim, in the engine, before the model is allowed to say it.

And in the U.S., the FTC's 2024 "Operation AI Comply" enforcement sweep put beauty and personal-care AI claims on explicit notice: AI-generated claims are held to the same substantiation standard as any other advertising claim, and the FTC has signaled it will treat unsubstantiated accuracy or personalization claims from an AI system as no different from an unsubstantiated claim made by a human copywriter.15 For a company like Flux, that means "boxed" is not just a safety design — it's the design that makes a personalization claim defensible at all, because it lets Flux point to what the engine's recommendations are actually based on, rather than to what a language model felt like generating.

7. Objections

Objection 1: Flux has a commercial interest in this exact conclusion, and this paper was written by people who build and sell Zamia.

This is true, and it is the single strongest reason to discount everything else in this paper somewhat. We are not a disinterested academic lab; we are a company arguing that the problem we identified is real and that our product is a serious attempt at solving it. The honest response is not to protest neutrality but to separate the two kinds of claim this paper makes and note that they carry different risk of bias. The claims about representation gaps in dermatology trials, textbooks, and AI training datasets (§3) and about PIH mechanism and persistence (§4) are drawn entirely from independent, peer-reviewed literature that exists whether or not Zamia exists, and would be true or false regardless of Flux's commercial fortunes. The claim that Zamia's architecture is a serious attempt at addressing that gap (§6) is the one claim in this paper that is Flux's to make and the reader's to doubt; we have deliberately not asked the reader to take our word for how well Zamia's evidence engine is actually curated, because we have not audited it independently for this piece. Readers should weight §6 accordingly, and weight §§3–4 on their own scientific merits.

Objection 2: "Personalized skincare" is industry marketing language, not a scientific category — and this paper is dressing Zamia's marketing claim up in the vocabulary of a research paper.

This objection has real force, and it's worth stating the strongest version of it: a 2025 systematic review of AI- and genomics-driven personalized skincare research — cosmetogenomics, in the field's own term — found that most of the research base behind this entire category "remains at early proof-of-concept stages" and lacks independent validation, that the studies that do exist show "limited geographic diversity and underrepresentation of darker skin phototypes," and that the reviewers concluded the evidence base is, in their words, "modest, geographically limited, and often exploratory."16 That finding cuts directly against this paper's argument, because it means the AI-personalization research literature — the field Zamia sits in — has exactly the same representation problem this paper accuses the older dermatology literature of having. We are not exempting AI-driven personalization from the critique we're making of the market it's replacing; if anything, this review shows the critique applies to the AI-personalization research base too, and Zamia's evidence engine is only as good as its willingness to correct for that rather than inherit it. We think this is a reason to be more specific and more skeptical about any personalization claim — including ours — not a reason to abandon the distinction between segmentation and risk-personalization. But readers who conclude that "personalization" as a category is currently more marketing than science, across the whole industry including AI-driven entrants, are reading the same evidence we are and reaching a defensible conclusion.

Objection 3: "Boxed" and "never invents" don't solve the problem if the underlying engine is wrong, and PIH is exactly the risk where getting it wrong costs the most.

This is the objection that worries us most, because it doesn't have a clean rebuttal — it has an operational answer that depends on facts about Zamia's engine we have not independently verified. The architecture in §6 is only as good as the evidence curated inside the box. If Zamia's engine were trained or validated against the same lopsided image and trial base described in §3 — the same one where AI dermatology tools have been shown, in at least one dataset review, to have almost no Fitzpatrick V/VI representation to learn from6 — then "the model doesn't invent" is a true but insufficient claim; a fixed engine built on biased inputs still produces biased outputs, just consistently instead of randomly. We looked for a direct, published measurement of how much diagnostic or recommendation accuracy specifically degrades across Fitzpatrick skin types for AI dermatology tools, and found that the research question is being actively studied — a 2025 letter in the International Journal of Dermatology specifically evaluated multimodal LLM diagnostic accuracy across skin types17 — but we could not independently verify the letter's specific quantitative findings within the scope of this paper, and we are not going to state a number we haven't confirmed. We flag that as CITATION-NEEDED in Appendix B rather than filling the gap with an estimate. What we can say is that this objection identifies the actual, non-rhetorical stakes of Zamia's engineering: the "boxed" design only earns the trust this paper is extending to it if the box's contents are demonstrably, auditably built against skin-of-color-specific evidence, and that is a standard Flux should expect to be held to rather than one this paper can certify.

8. Falsification conditions

In decreasing order of how much damage each would do to this paper's thesis:

  1. The PIH literature itself is overstated or contested. If subsequent, larger, or better-controlled studies show that post-inflammatory hyperpigmentation is not, in fact, measurably more severe, more common, or more persistent in deeper Fitzpatrick phototypes than in lighter ones — if the mechanistic account in §4 turns out to rest on a small or non-representative evidence base of its own — the entire premise that there is a specific, evidenced risk to personalize against collapses, and this paper's thesis fails outright.

  2. The representation gap is closing fast enough to be stale. If new, large-scale audits of dermatology trials, textbooks, and — especially — AI training datasets show that Fitzpatrick IV–VI representation has risen sharply and recently, closing the gap documented in §3 to a point where "the evidence base wasn't built for this population" no longer describes the current literature, the historical framing this paper leans on would need to be substantially revised, even if the PIH science in §4 still held.

  3. Mainstream segmentation-personalization already prices in pigmentary risk. If it turned out that leading skin-type-and-budget personalization tools already collect and act on PIH-relevant history — prior scarring, phototype, reaction history — as a matter of course, the "two jobs" distinction this paper is built on would be a distinction without a difference, and the paper's central contribution would need to be withdrawn or narrowed to a much smaller claim about labeling rather than substance.

  4. Zamia's own evidence engine is shown to be built on the same skewed base it critiques. If an independent audit of Zamia's underlying product-and-evidence engine found that it draws primarily on the same underrepresented dermatology literature and imaging data described in §3, rather than on skin-of-color-specific sources, the architecture described in §6 would not deliver what this paper claims it's built to deliver, even though the "boxed" design itself might still be structurally sound. This would not falsify the paper's account of the market-wide gap, but it would falsify the claim that Zamia specifically closes it.

  5. AI personalization tools measurably underperform on darker skin types by a wide margin. If the accuracy question flagged as CITATION-NEEDED in §7 and Appendix B is resolved with a clear, well-replicated finding that multimodal AI tools perform substantially worse on Fitzpatrick V/VI than on lighter phototypes, that would not falsify the thesis that risk-personalization is the right goal, but it would sharply raise the bar for what "boxed" has to guarantee before any AI consultant, Zamia included, should be trusted with this specific risk.

9. Limitations

This paper is not medical advice, and nothing in it should be read as instructing anyone on how to treat a specific skin condition. We do not claim that Zamia, or any AI system referenced here, can diagnose a skin condition, distinguish PIH from another pigmentary disorder in an individual case, or substitute for a dermatologist's clinical judgment. Where the FDA's general wellness framework draws a line between cosmetic guidance and disease diagnosis or treatment (§6), this paper sits, and asks Zamia to sit, on the cosmetic-guidance side of that line, and everything we've argued should be read within that boundary.

We do not present new primary research. Every empirical claim here is drawn from existing published literature, and this paper's contribution is the argument connecting that literature to a specific product-design question, not new data. We have not audited Zamia's underlying evidence engine independently for this paper, and §7's third objection and §8's fourth falsification condition should be read as open questions about Flux's own product, not as settled facts in Zamia's favor. We have not run or cited any controlled study comparing Zamia's outcomes to any competitor's, and this paper makes no claim that Zamia performs better than any named or unnamed alternative. Finally, several of the figures cited here — market size estimates, testing-program non-compliance rates — are explicitly order-of-magnitude, press-reported, or program-specific figures rather than population-level scientific measurements, and we have tried to flag each one as such at the point of citation rather than only in this section.


Appendix A: Evidence table

Claim Measurement needed Status Source
Dermatology surgical textbooks underrepresent Fitzpatrick IV–VI Share of clinical images depicting FST IV–VI in core surgical textbooks Sourced — 5.6% of 1,501 images; 37.9% of topics with zero skin-of-color images Porras Fimbres et al., Arch Dermatol Res (2023)3
Dermatology textbooks show little improvement over time Change in FST V/VI image share, 2006–2020, across major textbooks Sourced — only 1 of 6 textbooks showed >1 percentage-point increase Adelekun, Onyekaba & Lipoff, J Am Acad Dermatol (2020/2021)4
Melasma trials underrepresent the phototypes most affected FST distribution among melasma trial participants Sourced — FST III/IV >75% of participants Wang et al., J Drugs Dermatol (2025)5
AI dermatology training data underrepresents dark skin severely Share of skin-cancer image datasets reporting FST; share of those images that are FST V/VI Sourced — FST reported for 2.1% of images; 11 of 2,436 typed images were FST V/VI Wen et al., Lancet Digit Health (2022)6
PIH is a leading, disproportionate complaint among Black dermatology patients Diagnosis ranking and prevalence of dyschromia/PIH by patient race Sourced — ranked 2nd–3rd most common diagnosis in Black patients across cited studies; not in top 10 for white patients Davis & Callender, J Clin Aesthet Dermatol (2010)2
PIH persists longer than typical irritation Duration of epidermal vs. dermal PIH resolution Sourced — epidermal PIH: months to years untreated; dermal PIH: can be permanent Davis & Callender (2010)2
Mechanistic basis for greater PIH severity in dark skin Cytokine/melanocyte/barrier differences by phototype Sourced — inflammatory mediator and barrier-structure differences described Markiewicz et al., Clin Cosmet Investig Dermatol (2022)7
Mercury-containing lightening products remain in circulation despite bans Scale of ongoing global trade in banned lightening products Sourced qualitatively (persistence via internet/informal trade); no reliable global volume estimate found WHO information note (2019)8
Non-compliant (mercury/hydroquinone-containing) lightening products in circulation Share of tested market products violating bans Sourced, testing-program figure only — ~18% non-compliant in one EU testing program EDQM report11
Global skin-lightening market size Market revenue estimate Sourced, order-of-magnitude industry estimate only — ~$11B (2023) Grand View Research12
AI diagnostic accuracy differs by Fitzpatrick skin type Quantified accuracy gap between lighter and darker phototypes for multimodal AI dermatology tools Not independently verified in this paper — study exists but specific figures unconfirmed
Consumer trust in "dermatologist recommended/tested" labeling Survey-based measure of how such labeling affects purchase or trust decisions Not sourced to a verifiable primary study within this paper's research pass
Zamia's evidence engine's actual source composition Share of Zamia's underlying product/evidence base drawn from skin-of-color-specific literature vs. general dermatology sources Not audited — no independent measurement available
Incidence rate of PIH among skincare-app users specifically Population-level PIH incidence tied to app-driven routine choices Not measured anywhere found — no study connects personalization-app usage to PIH incidence

Appendix B: Citation register (CITATION-NEEDED)

  • Quantified AI diagnostic/recommendation accuracy gap by Fitzpatrick skin type. A relevant study exists — Kim GH, Kim NK, Moon IJ, "Evaluating Dermatological Diagnostic Accuracy and Consistency in Multimodal Large Language Models Across Skin Types," International Journal of Dermatology (2025) — but this paper could not independently verify its specific quantitative findings within the current research pass. Any future revision citing a specific accuracy differential from this or a comparable study should re-verify the figure directly from the primary source before publication.
  • Exact prevalence rate of PIH among users of over-the-counter skincare specifically (as opposed to diagnosis-ranking data among dermatology patients generally). The Davis & Callender (2010) review establishes PIH's rank among diagnoses and its persistence, but does not give a population-wide incidence rate tied to product use, and no such rate should be inferred from it.
  • A direct empirical link between mainstream segmentation-personalization tools and PIH incidence or outcomes. No study was found connecting the design of existing skin-type-and-budget personalization quizzes to measured PIH outcomes in their users, positive or negative. This paper's claim that such tools do not price in PIH risk rests on their design (the questions they ask), not on outcome data, and should not be overstated as an outcomes claim.
  • Survey data behind "dermatologist recommended/tested" consumer trust figures occasionally cited in industry commentary (e.g., a frequently repeated figure that a large share of consumers trust such labeling as a guarantee of efficacy). This paper found the claim repeated in secondary commentary but could not trace it to a primary, citable survey within this research pass, and has therefore omitted any such figure from the body text.
  • Zamia's internal evidence-engine composition and curation process. This paper asserts that Zamia is designed to draw on skin-of-color-specific evidence, per Flux's own architecture description in WP21, but has not independently audited the engine's actual source composition. Any future paper claiming a specific, audited answer to this question should be treated as a distinct, verifiable claim, not an extension of this one.

Footnote definitions

  1. Ruto K. "Personalized Is the New Glow." Flux Working Papers, fluximpact.org/blog/012-personalized-is-the-new-glow/. Cited for the argument structure (glow as marketed myth, the trial-and-error tax, Zamia's "boxed" architecture, and the "radiance, never lightness" principle) that this paper builds on rather than repeats.

  2. Davis EC, Callender VD. "Postinflammatory Hyperpigmentation: A Review of the Epidemiology, Clinical Features, and Treatment Options in Skin of Color." Journal of Clinical and Aesthetic Dermatology. 2010;3(7):20–31.

  3. Porras Fimbres D, et al. "Representation of Fitzpatrick skin phototype in dermatology surgical textbooks." Archives of Dermatological Research (2023). Analysis of 1,501 images across three core dermatology surgery textbooks.

  4. Adelekun A, Onyekaba G, Lipoff JS. "Skin color in dermatology textbooks: An updated evaluation and analysis." Journal of the American Academy of Dermatology (2020/2021).

  5. Wang JY, et al. "Gender, Racial, and Fitzpatrick Skin Type Representation in Melasma Clinical Trials." Journal of Drugs in Dermatology. 2025;24(1):19–22.

  6. Wen D, Khan SM, Xu AJ, Ibrahim H, et al. "Characteristics of publicly available skin cancer image datasets: a systematic review." The Lancet Digital Health. 2022;4(1):e64–e74.

  7. Markiewicz E, Karaman-Jurukovska N, Mammone T, Idowu OC. "Post-Inflammatory Hyperpigmentation in Dark Skin: Molecular Mechanism and Skincare Implications." Clinical, Cosmetic and Investigational Dermatology (2022).

  8. World Health Organization. "Mercury in skin lightening products." Information note WHO/CED/PHE/EPE/19.13 (2019).

  9. U.S. Food and Drug Administration. "FDA Warns Consumers of Skin Products Containing Mercury and/or Hydroquinone." FDA consumer update, fda.gov.

  10. European Union. Regulation (EC) No 1223/2009 on cosmetic products, Annex II (prohibited substances), entry covering hydroquinone; related reporting on associated risk of exogenous ochronosis.

  11. European Directorate for the Quality of Medicines & HealthCare (EDQM). Report on the presence of banned substances (mercury, hydroquinone, corticosteroids) in tested skin-whitening products. Figure cited (~18% non-compliance) is a targeted testing-program result, not a random population sample, and is treated here as order-of-magnitude.

  12. Grand View Research. "Skin Lightening Products Market Size, Share & Trends Analysis Report" (2024–2030 edition). Market estimate (~$11B in 2023, projected toward the mid-teens of billions by 2030) is a press-reported industry estimate, treated here as order-of-magnitude, not a scientific measurement.

  13. U.S. Food and Drug Administration. "General Wellness: Policy for Low Risk Devices." Guidance for Industry and FDA Staff (revised 2026; originally issued 2016, updated 2019).

  14. European Commission. Regulation (EC) No 1223/2009, Article 20; Commission Regulation (EU) No 655/2013 laying down common criteria for the justification of claims used in relation to cosmetic products (the "evidential support" criterion).

  15. U.S. Federal Trade Commission. "FTC Announces Crackdown on Deceptive AI Claims and Schemes" [Operation AI Comply]. Press release, September 25, 2024, ftc.gov.

  16. Haykal D, Flament F, Amar D, Cartier H, Kourosh AS, Lee DH, Rowland-Payne C. "Cosmetogenomics unveiled: a systematic review of AI, genomics, and the future of personalized skincare." Frontiers in Artificial Intelligence (2025).

  17. Kim GH, Kim NK, Moon IJ. "Evaluating Dermatological Diagnostic Accuracy and Consistency in Multimodal Large Language Models Across Skin Types." International Journal of Dermatology (2025). Cited for the existence of the study; specific quantitative findings not independently verified for this paper — see Appendix B.

Provenance
Flux Working Paper No. 41 · Ken Ruto, Flux (FluxImpact)
Published 24 Sep 2026
Content hash (SHA-256): 692df205295dce23… · build d2a1837
DOI: pending deposit
Ken Ruto
About the author
Ken Ruto

Founder of Flux. Building vertical AI-powered SaaS for Africa's institutions — and writing the thesis behind every bet. kenruto.fluximpact.org →

Share X LinkedIn WhatsApp
Did this land?
Was it useful?

Comments

No comments yet — be the first.

Replying to · cancel
Get new essays

No spam — just the next piece when it's out.

Think I got something wrong? Highlight any sentence to push back on it — or It comes straight to me, never shown publicly.

Push back
Related writing
35 min
The Switching Cost Is Why Nothing Sticks
WP26 said the useful personal AI would live in your messages, not in an app you open. This one asks why that actually works — and finds the mechanism in a decades-old finding about what happens to a goal while you're off looking for the app.
9 min
The Second Question: What Changes When an Explainer Can Be Asked Again
An inline explainer that answers once is a dictionary. The follow-up carries the value — and it is why the thing has to go standalone.
9 min
The Panel Is Also a Tab
Every AI answer surface solves the new-tab reflex by moving you somewhere smaller. The cost was never the tab — it was getting back.