A validity and fairness framework for multilingual beauty-licensing examinations
Prepared September 14, 2026 Prepared by Amy Comp for Di Tran University Operating disclosure: Amy Comp is operated by Di Tran Enterprise (Di Tran LLC) under human oversight. Research status: Independent educational and policy analysis. This paper is not legal advice, does not reproduce secure examination content, and does not allege intentional misconduct by PSI, ETS, a licensing board, or any individual.
Abstract
Beauty professionals work near skin, eyes, nails, hair, chemicals, sharp tools, heat, and surfaces capable of transmitting infection. Licensing examinations therefore have a legitimate public-protection purpose: identify whether a candidate has the minimum knowledge and judgment needed to practice safely.
Louisville Beauty Academy has heard recurring candidate descriptions of some licensing questions as “tricky.” Those reports deserve attention, but they are not enough to prove deliberate trickery, discrimination, or an invalid examination. The correct research question is narrower and more useful: Do any test items introduce reading, wording, or translation demands that are not necessary to measure safe occupational practice?
The governing Kentucky materials, PSI’s own published test-development principles, ETS’s Standards for Quality and Fairness, the joint AERA/APA/NCME testing standards, the International Test Commission’s test-adaptation guidelines, and peer-reviewed research converge on a balanced answer. Applied scenarios can validly test professional judgment. However, ambiguous stems, unnecessary conditions, multiple negatives, obscure nontechnical vocabulary, cultural assumptions, or non-equivalent translations can create construct-irrelevant variance—score differences caused by something other than the safety competence the examination is supposed to measure.
The concern is therefore fair and evidence-worthy, but not yet proven at the PSI program level. A responsible resolution requires secure item review, cognitive laboratories, readability and linguistic-demand review, bilingual adaptation panels, differential item functioning analysis where sample sizes permit, documented challenge procedures, and privacy-protected outcome reporting. The standard should be rigorous and simple: test the judgment that protects the public; remove language difficulty that does not.
Executive finding
What the evidence supports
- Safety and infection control are central constructs. Kentucky’s current PSI guides allocate 30% of the cosmetology theory outline, 40% of esthetics, 50% of nail technology, and 40% of shampoo styling to safety and infection control.
- Applied judgment is legitimate. Safe practice is not only memorizing definitions. A professional may need to identify contamination, choose the correct response to blood exposure, interpret an SDS, or decide whether a service should stop. A concise, job-authentic scenario can measure that judgment.
- Language must remain proportionate to the construct. PSI says credentialing items should directly assess intended knowledge without extraneous factors and should be clear and unambiguous. ETS standards say linguistic and reading demands should be no greater than necessary for the test’s purpose.
- Negative and convoluted wording can threaten validity. Peer-reviewed research identifies negatively worded multiple-choice questions as an avoidable source of construct-irrelevant variance, especially when candidates respond as though a negative stem were positive.
- Translation is an adaptation and validation problem. PSI says translated items should be equivalent to the English form and independently verified. ETS and International Test Commission standards call for documented adaptation and empirical evaluation of comparability.
- Candidate reports are signals, not verdicts. Without secure item access, response data, translation records, or psychometric analyses, the public cannot determine whether a reported difficulty reflects necessary safety judgment, weak preparation, one flawed item, a translation defect, or a broader pattern.
What the evidence does not support
- It does not prove that PSI intentionally writes “trick questions.”
- It does not prove that current PSI National beauty examinations are invalid or discriminatory.
- It does not show that every scenario, negative stem, or conditional question is improper.
- It does not establish psychometric equivalence or inequivalence across Kentucky’s six available languages; that requires data not presently public.
- It does not authorize disclosure, collection, or reconstruction of secure examination items.
1. The question that matters
Calling a question “tricky” describes a test taker’s experience, not a technical finding. A fairness review must translate that experience into testable hypotheses:
- Is the item aligned to the published content outline and required curriculum?
- Is the knowledge or judgment actually required for safe entry-level practice?
- Could a qualified practitioner answer correctly without solving an unnecessary language puzzle?
- Are the stem and options unambiguous to representative subject-matter experts?
- Does the item behave unexpectedly across language groups after relevant proficiency is controlled, where data permit?
- Is a translated item functionally equivalent in meaning, difficulty, and intended cognitive demand?
- Is there a secure, documented route to report and investigate a suspected ambiguity, mistranslation, outdated reference, or miskey?
This framing protects both sides. It preserves rigorous testing and public safety while refusing to treat irrelevant reading difficulty as evidence of occupational competence.
2. Kentucky’s legal and examination context
Kentucky Revised Statute 317A.120 authorizes examinations that cover qualifications for the license, including skill, technique, scientific knowledge, and other knowledge. It permits use of a national examination approved by the Board. Kentucky regulation 201 KAR 12:030 provides that theory and practical examinations are drawn from the curriculum requirements in 201 KAR 12:082 and establishes passing standards.
The current Kentucky PSI candidate guides identify PSI National theory examinations for cosmetology, esthetics, nail technology, and shampoo styling. They state that theory forms are available in English, Korean, Portuguese, Simplified Chinese, Spanish, and Vietnamese. They also identify unscored experimental questions used in developing future tests.
These authorities establish a legitimate domain—fitness to practice under the required curriculum. They do not make unnecessary linguistic complexity part of that domain. Nor do they establish that every difficult question is defective. The controlling distinction is whether the difficulty belongs to the work.
3. Direct knowledge and applied judgment are not opposites
The proposition “safety should be direct” is substantially right, but it needs one qualification. Safety rules should be teachable, memorable, and executable. Yet a licensing examination may also need to determine whether a candidate can apply those rules when facts change.
Consider two abstract examples that do not reproduce any secure item:
- Direct knowledge: Which class of disinfectant is appropriate under a specified sanitation rule?
- Applied judgment: After a tool contacts blood, which sequence of actions protects the client and practitioner?
The second item asks for reasoning, but the reasoning is occupationally relevant. The U.S. Office of Personnel Management describes situational-judgment tests as work-related scenarios built from critical incidents and subject-matter expert judgments; their strength is content relevance to actual tasks. PSI likewise states that the construct should drive the item type and that cosmetology exams should measure applied understanding of infection control, hygiene, and tool safety.
The problem begins when the scenario adds irrelevant narrative, syntactic complexity, hidden exceptions, or vocabulary that changes the task from “choose the safe action” to “decode the examiner’s prose.” Rigor is not the same as obscurity.
4. Construct-irrelevant variance: the core technical concern
In assessment science, a score is valid only in relation to the interpretation and use claimed for it. A beauty-licensing score is intended to support an inference about minimum safe practice—not literary analysis, generalized academic English, or test-taking cleverness.
Construct-irrelevant variance occurs when scores vary because of factors outside the intended construct. Potential sources include:
- reading load beyond what the job requires;
- multiple negatives or exception formats;
- avoidable conditional clauses;
- ambiguous pronouns or time order;
- distractors that depend on verbal tricks rather than plausible unsafe practices;
- uncommon nontechnical vocabulary;
- cultural context unrelated to beauty work;
- translation shifts in meaning, difficulty, or emphasis; and
- computer-interface burdens unrelated to the licensed task.
ETS Standard 5.1 calls for linguistic or reading demands no greater than necessary to achieve the purpose of the assessment. It also identifies judgmental and empirical fairness reviews, including differential item functioning. PSI’s published item-writing principles similarly emphasize relevance, clarity, fairness, representative experts, and multiple layers of review.
Those are strong standards. The present question is not whether the standards exist; it is whether the program can demonstrate how they are operationalized for multilingual beauty examinations.
5. Why negative wording deserves special scrutiny
Words such as not, except, least, or layered negatives can sometimes be necessary—for example, if safe practice genuinely requires recognizing a prohibited action. But unnecessary negative construction can make an item measure whether the candidate noticed the linguistic reversal.
Chiavaroli’s peer-reviewed review of negatively worded multiple-choice questions reports that even typographic emphasis may not prevent candidates from responding as though a negative item were positive. The author characterizes this as an avoidable threat to validity. A broader medical-education literature likewise warns that flawed items can inflate or deflate scores through systematic measurement error.
The practical policy is not “ban every negative word.” It is:
- use positive, direct stems by default;
- permit negative construction only when the negative distinction is part of safe practice;
- prominently signal the negative term;
- require an independent reviewer to justify its necessity;
- monitor its item statistics and distractor behavior; and
- retire or revise items that show anomalous or group-specific behavior that cannot be explained by the intended construct.
6. Multilingual testing requires adaptation, not substitution
Kentucky’s availability of six theory languages is meaningful access. Availability alone, however, does not establish equivalence.
Literal word-for-word translation can change:
- the difficulty of vocabulary;
- the location or salience of a negative;
- grammatical time and sequence;
- the plausibility of distractors;
- professional terminology;
- cultural familiarity; or
- the amount of reading needed to reach the same meaning.
The International Test Commission organizes test adaptation around preconditions, development, empirical confirmation, administration, scoring and interpretation, and documentation. PSI’s own translation guidance calls for equivalence to the English form, alignment to the same specifications, and verification after initial translation. ETS Standard 5.7 calls for the adaptation process and comparability outcomes to be described and evaluated.
Research by Allalouf found that translated tests are not automatically psychometrically equivalent and that differential item functioning analysis can identify items affected by word difficulty, format, content, or cultural relevance. Revision reduced DIF in that study. DIF is not by itself proof of bias; it is a flag requiring substantive investigation.
A strong multilingual program therefore combines:
- two or more qualified translators with domain knowledge;
- independent review or reconciliation;
- representative licensed practitioners and educators;
- cognitive interviews with target-language candidates;
- preservation of technical terms that are genuinely required on the job;
- item-level statistical monitoring by language when sample sizes permit;
- qualitative review when samples are too small for stable DIF estimates; and
- documented revision, re-equating, and communication controls.
7. Disability law and language access: related but distinct
The Americans with Disabilities Act expressly covers examinations used for trade licensing, including cosmetology. U.S. Department of Justice guidance explains that covered tests must be accessible to candidates with disabilities and, when accommodation is required, should reflect the aptitude or skill the examination is intended to measure rather than the impairment—except when the impaired skill is itself the construct.
That principle is highly relevant by analogy to construct clarity, but English-language-learner status is not itself a disability. Language access may implicate different federal or state authorities depending on the entity, funding, and facts. This paper therefore does not claim that multilingual differences automatically establish an ADA or Title VI violation. Any candidate-specific legal claim should be evaluated by qualified counsel.
8. Is the concern fair?
Yes—as a disciplined inquiry
The concern is fair because:
- the stakes include entry into a licensed occupation;
- safety content comprises a major share of the published outlines;
- Kentucky serves adult and multilingual learners;
- established testing standards recognize unnecessary language demand and translation inequivalence as validity threats;
- negative wording has documented risks; and
- the public guides reviewed do not explain a Kentucky-specific, item-security-safe challenge and disposition process.
No—as a settled accusation
It would not be fair, on present evidence, to say:
- PSI deliberately deceives test takers;
- a particular candidate failed because of a defective item;
- one language form is harder than another;
- the passing standard is invalid; or
- the entire examination measures English instead of safety.
Those conclusions require evidence not available in public materials: secure item review, form assembly records, cognitive-lab results, item statistics, equating evidence, language-level performance, challenge outcomes, and expert analysis.
9. A rigorous, privacy-protective study design
Di Tran University proposes a research program that does not collect or publish secure test content.
Phase 1 — Candidate experience instrument
Collect de-identified reports using categories rather than recalled wording:
- program and test language;
- first or repeat attempt;
- content domain;
- perceived issue: negative stem, time sequence, vocabulary, translation, outdated terminology, interface, or preparation gap;
- confidence that the concern affected comprehension;
- whether the candidate used the official reporting route; and
- result of any response received.
Do not ask candidates to reproduce questions, answer choices, screenshots, or confidential testing material.
Phase 2 — Public-document alignment review
Map published content outlines to Kentucky curriculum requirements and cited references. Identify where precise technical vocabulary is necessary and where plain-language alternatives would preserve the construct.
Phase 3 — Independent secure audit
Under appropriate nondisclosure and test-security controls, qualified psychometricians, beauty SMEs, infection-control educators, and bilingual reviewers should examine flagged categories. Reviewers should assess content alignment, linguistic necessity, ambiguity, translation equivalence, key accuracy, distractor quality, and entry-level relevance.
Phase 4 — Empirical analysis
Where sample sizes support responsible analysis:
- classical item difficulty and discrimination;
- distractor analysis;
- DIF by language and other legally/ethically appropriate groups;
- form comparability and equating;
- first-attempt versus repeat patterns;
- response-time anomalies; and
- pre/post-revision performance.
Small cells should be suppressed, and DIF flags should trigger expert review rather than automatic conclusions of bias.
Phase 5 — Correction and transparency
Publish process-level findings: number of items reviewed, categories of revisions, adaptation methodology, frequency and disposition of challenges, and aggregate outcomes. Never publish secure items or information that enables memorization.
10. Recommendations
For PSI and ETS
- Publish a plain-language statement of the intended construct for each beauty examination and whether general English proficiency is part of it.
- Document controls for negative wording, conditional complexity, reading demand, cultural fairness, and translation adaptation.
- Provide a secure item-concern process with deadlines, review roles, disposition categories, and score-remedy rules.
- Use cognitive labs with representative adult and multilingual test takers before operationalizing materially revised forms.
- Conduct DIF and response-process studies where sample sizes permit; use structured qualitative review where they do not.
- Publish the composition and selection principles—not secure deliberations—of relevant advisory panels.
- Include working licensees, safety educators, small-school faculty, multilingual professionals, and recent candidates in a balanced expert pool.
For schools and educators
- Teach the safety rule, the reason for it, and the action sequence.
- Practice concise, authentic scenarios without teaching “tricks.”
- Separate required technical vocabulary from avoidable academic complexity.
- Train learners to identify negative terms while advocating for positive stems by default.
- Preserve de-identified trend data and use the official secure reporting process.
- Never solicit or circulate recalled exam items.
For researchers and policymakers
- Treat candidate narratives as early-warning data, not proof.
- Require outcome definitions and denominators before comparing groups.
- Separate language access, disability accommodation, curriculum alignment, and psychometric validity; they overlap but are not identical legal questions.
- Reward transparent correction systems. A revised item is evidence of quality control, not institutional weakness.
11. The human standard
The ethical purpose of a licensing test is not to demonstrate the cleverness of the test writer. It is to protect a person in the chair, the practitioner serving that person, and the public that trusts the license.
Humanization does not mean lowering the standard. It means locating rigor in the right place: infection control, chemical safety, blood-exposure response, tool handling, contraindications, and professional judgment. A fair test can be difficult because safe practice is demanding. It should not be difficult because meaning is unnecessarily hidden.
The governing principle is simple: require every decision safety requires; remove every obstacle safety does not.
12. Limitations and next evidence needed
This paper is based on public law, regulations, candidate guides, provider standards, professional standards, and published research. It does not include secure PSI items, proprietary psychometric data, Kentucky form-level statistics, candidate records, or PSI/ETS’s response to the September 14 governance inquiry. It cannot estimate the prevalence of defective wording or translation effects.
The next evidence needed is:
- PSI/ETS’s documented response about its operational controls;
- the current advisory-board roster and representation criteria;
- an item-security-safe challenge policy;
- de-identified aggregate outcomes by program and language with small-cell protection;
- cognitive-lab and adaptation documentation; and
- independent psychometric review under controlled access.
References and authority register
Controlling Kentucky authority
- Kentucky General Assembly. KRS 317A.120 — Examinations by board. Current official text, accessed September 14, 2026. https://apps.legislature.ky.gov/law/statutes/statute.aspx?id=56213
- Kentucky Legislative Research Commission. 201 KAR 12:030 — Licensing and examinations. Current official text, accessed September 14, 2026. https://apps.legislature.ky.gov/law/kar/titles/201/012/030/
- Kentucky Legislative Research Commission. 201 KAR 12:082 — Education requirements and school administration. Current official text, accessed September 14, 2026. https://apps.legislature.ky.gov/law/kar/titles/201/012/082/
- PSI Services. Kentucky candidate guides, bulletin IDs 8564–8568. Effective September 8, 2026; downloaded September 14, 2026 from the KBC-linked PSI portal. Example: https://test-takers.psiexams.com/api/content/bulletin/8565
Provider and professional standards
- Gasperson, S. Item writing and exam assembly in credentialing: Importance and best practices. PSI Services, May 15, 2024. https://www.psiexams.com/knowledge-hub/item-writing-and-exam-assembly-in-credentialing-importance-and-best-practices/
- PSI Services. Greater equity, better experience: How language translations can make more sense for your test takers. https://www.psiexams.com/knowledge-hub/translating-tests-into-additional-languages/
- Conder, S. Public health and safety-first cosmetology state board exams in a multi-modal era. PSI Services, March 4, 2026. https://www.psiexams.com/knowledge-hub/public-health-and-safety-first-cosmetology-state-board-exams-in-a-multi-modal-era/
- Educational Testing Service. ETS Standards for Quality and Fairness. 2014. https://www.ets.org/content/dam/ets-org/pdfs/about/standards-quality-fairness.pdf
- American Educational Research Association, American Psychological Association, and National Council on Measurement in Education. Standards for Educational and Psychological Testing. 2014; open-access portal accessed September 14, 2026. https://www.apa.org/science/programs/testing/standards
- International Test Commission. The ITC Guidelines for Translating and Adapting Tests, Second Edition. 2017. https://www.intestcom.org/files/guideline_test_adaptation_2ed.pdf
- U.S. Department of Justice. ADA Requirements: Testing Accommodations. Updated February 28, 2020. https://www.ada.gov/resources/testing-accommodations/
- U.S. Office of Personnel Management. Situational Judgment Tests. Accessed September 14, 2026. https://www.opm.gov/policy-data-oversight/assessment-and-selection/other-assessment-methods/situational-judgment-tests/
Research literature
- Chiavaroli, N. (2017). Negatively-Worded Multiple Choice Questions: An Avoidable Threat to Validity. Practical Assessment, Research & Evaluation, 22(3). https://doi.org/10.7275/ca7y-mm27
- Allalouf, A. (2003). Revising Translated Differential Item Functioning Items as a Tool for Improving Cross-Lingual Assessment. Applied Measurement in Education, 16(1), 55–73. https://doi.org/10.1207/S15324818AME1601_3
- Downing, S. M. (2002). Threats to the validity of locally developed multiple-choice tests in medical education: construct-irrelevant variance and construct underrepresentation. Advances in Health Sciences Education, 7, 235–241. https://pubmed.ncbi.nlm.nih.gov/12510145/
Claim-control note
“Trickery” remains a reported perception, not an established PSI practice. Publications derived from this paper must preserve that distinction, avoid candidate-identifying information and secure exam content, and update the analysis if PSI/ETS supplies responsive evidence.