Methodology
How each track is built
Published in full, because an assessment whose method is not inspectable is not an assessment. It is an opinion with a number attached.
Five stages, three of them before anyone can take it
A track is not a variant of one master questionnaire with the nouns swapped. Each is researched, written and validated against the standards governing its own profession, which is why they reach publication at different times.
Framework research
Pre-publishingEach track is anchored to the regulatory and professional body frameworks governing its own target population. No cross-sector borrowing.
- Review of the professional body AI guidance specific to the track, such as ABA Formal Opinion 512 for legal, FSB and IOSCO for financial services, NACD and OECD for board leadership, or C2PA for creative and media.
- Competency domains derived from those standards rather than from a generic template.
- Target population and exclusion criteria defined before any item is written.
- Literature review of existing measurement instruments in the field.
- The theoretical construct documented and agreed before drafting begins.
Expert item development
Pre-publishingSubject matter experts draft 40 items per track, calibrated to observable professional behavior rather than to abstract knowledge.
- 25 behaviorally anchored rating-scale items distributed across the competency domains.
- 15 scenario-based judgment items concentrated in the highest-stakes areas.
- Three proficiency levels per item: Acquire as the safe minimum, Deepen as competent applied practice, Create at expert and governance level.
- Anchors describe what practitioners at each level actually do, not what they know in the abstract.
- Reviewed by the track subject matter lead before it reaches the expert panel.
Content Validity Index review
Pre-publishingAn independent panel scores every item on relevance and clarity. Nothing enters a published instrument without passing this gate.
- A minimum of six subject matter expert judges per track, independent of the item authors, following Lynn (1986) and Polit and Beck (2006).
- Every item rated separately on two criteria: relevance to the construct, and clarity of wording.
- Item-level (I-CVI) and scale-level (S-CVI/Ave) indices computed on the four-point ordinal scale established by Lynn (1986).
- Acceptance threshold for the full instrument: S-CVI/Ave at or above 0.90.
- Items below threshold on either criterion are revised, not removed, and resubmitted for a second independent round.
Track publishing
LaunchThe validated instrument is integrated into the platform and the track opens to organizations.
- Scoring rubric and level thresholds locked. No post-publishing item change without re-validation.
- The track becomes available to organizations holding an applicable access code.
- Response data begins accumulating from licensed professional respondents.
- Volume monitored in preparation for the psychometric stage.
Psychometric validation
OngoingAs responses accumulate, the instrument is tested formally for factor structure and internal consistency.
- Exploratory Factor Analysis once 150 or more responses per track are available.
- Confirmatory Factor Analysis at 300 or more responses, to confirm the domain factor structure.
- Internal consistency measured against a Cronbach's alpha target of 0.80 or above per domain subscale.
- Item performance and score drift monitored between formal cycles.
- Periodic refinement as the professional and regulatory landscape moves.
What content validity does and does not establish
It establishes that the items represent the construct the instrument claims to measure, as judged by experts in the field. It does not establish reliability, construct validity, or responsiveness. Those require accumulated response data, which is what stage five is for.
What we will not claim
On self-report
There is a fair criticism of self-report instruments: they can capture confidence rather than ability, and confidence is a poor proxy. Two things address it here.
- Behavioral anchoring. A respondent chooses between described practices rather than rating themselves on a scale of one to five, and the instrument asks explicitly for current practice rather than intent.
- Scenario items. Fifteen of the forty describe a situation and ask for a course of action, concentrated in the areas where a wrong answer costs the most.
Used this way the instrument is a baseline and placement tool, which is what it is good at, rather than a measurement of demonstrated capability, which it is not and does not claim to be.
Framework alignment
Results report against the UNESCO AI Competency Frameworks, the UNESCO Recommendation on the Ethics of Artificial Intelligence, the AI Fluency 4D Framework, and the U.S. Department of Labor AI Literacy Framework, plus the professional body guidance specific to each track. The frameworks are the measurement and reporting layer. They are not the syllabus, and courses are not organized around them.
Not affiliated with UNESCO
Sources
- Lynn, Nursing Research, 1986Determination and quantification of content validity. The origin of the four-point ordinal CVI scale and its acceptance thresholds.
- Polit and Beck, Research in Nursing & Health, 2006The content validity index: are you sure you know what's being reported? Critique and recommendations.
- UNESCO AI Competency Framework for TeachersUNESCO, September 2024. 15 competencies across five dimensions and three progression levels.
- UNESCO Recommendation on the Ethics of Artificial IntelligenceUNESCO, adopted November 2021 by all member states. The first global standard on AI ethics.
- AI Fluency FrameworkThe 4D framework: Delegation, Description, Discernment, Diligence. Developed by Anthropic with Rick Dakan and Joseph Feller.
- U.S. Department of Labor AI Literacy FrameworkEmployment and Training Administration. Five foundational content areas and seven delivery principles.
- ABA Formal Opinion 512American Bar Association Standing Committee on Ethics and Professional Responsibility, July 2024. Generative artificial intelligence tools.
- FSB, The Financial Stability Implications of Artificial IntelligenceFinancial Stability Board, November 2024.
- IOSCO, Artificial Intelligence in Capital MarketsInternational Organization of Securities Commissions, March 2025. Use cases, risks and challenges.
- NACD AI governance guidanceNational Association of Corporate Directors. Board oversight of artificial intelligence.
- OECD AI PrinciplesOECD, adopted 2019 and updated 2024. The first intergovernmental standard on AI, with 47 adherents.
- C2PACoalition for Content Provenance and Authenticity, a Linux Foundation project. Content Credentials provenance standard.
The expert panel
Panels are recruited openly. Researchers, academics and credentialed professionals in a track area can apply to review; a cycle runs roughly 40 to 60 minutes and scores items on relevance and clarity.
