Assessment
When the AI sounds certain, does your team check?
Generative AI answers with structure and confidence whether it is right or wrong. The AI Critical Thinking Assessment measures whether a person actually verifies, interrogates, and corrects AI output under working conditions, not whether they say they would.
What it is
A two-part instrument. Part one is a validated thirteen-item scale of how a person reports using generative AI. Part two is a workplace simulation: the participant advises on a realistic business decision, consults an AI advisor that has been deliberately seeded with flawed reasoning, and writes two short memos that are scored against a published behavioral rubric.
- Structure
- A 13-item self-report scale, then an interactive simulation with two written memos.
- Time
- 35 to 50 minutes, in one sitting. Progress is saved on the device.
- Scoring
- Eleven behavioral indicators across three dimensions, plus three self-report subscales.
- Result
- A disposition profile, a demonstrated-competence index, and the gap between them.
- Access
- By organization code. There is no individual purchase today.
The problem it measures
The failure mode it targets is automation bias: accepting a well-structured, authoritative answer without checking it. In the simulation, the AI advisor is confidently wrong in specific, planted ways: a dismissed labor risk, a fabricated regulatory certainty, a security guarantee that conflates uptime with safety, a statistical bias reframed as natural volatility. The score reflects which of those the participant caught, challenged, and corrected in writing.
Because part one records what people say and part two records what they do, the assessment also surfaces the intention–behavior gap: the people who report strong verification habits but did not verify when it counted. That group is invisible to survey-only instruments, and it is exactly the group targeted training helps most.
How the simulation works
Part 1
Self-reported disposition
Thirteen statements about verification habits, motivation to understand how AI produces its answers, and reflection on responsible use. Scored on three subscales from the published Critical Thinking in AI Use scale (Lau et al., 2026).
Part 2
Live consultation and memos
The participant questions an AI advisor about a strategic decision, then submits a board memo and a critique of the AI's output. Chat behavior and both memos are scored against the eleven-indicator Human-AI Interaction rubric (Li et al., 2025).
What the scores cover
Dimension 1
Retrieval and analysis
How the participant works the advisor: question depth, coverage of distinct risk perspectives, follow-up chains, and how prompts are refined as the picture develops.
Dimension 2
Evaluation
Whether flawed claims are challenged at all, and whether the challenge is substantiated: evidence demanded, logic tested, errors named in writing.
Dimension 3
Integration
What the memos are built from: screened and filtered evidence rather than pasted output, multiple perspectives weighed, and a logically consistent argument to a clear recommendation.
What an organization receives
- Per-participant results: the disposition profile, the demonstrated-competence index, and every indicator with its grading rationale.
- An intention–behavior map of the cohort, plotting reported habits against demonstrated performance to show who needs which intervention.
- Excel and PDF exports, and optional emailed reports to participants (none, a summary, or the full report), set per access code.
The simulation is standardized, deliberately
Methodological grounding
Part one applies the Critical Thinking in AI Use scale of Lau et al. (2026), published in Computers in Human Behavior Reports. Part two applies the performance-based Human-AI Interaction competence rubric of Li et al. (2025), published in the Journal of Intelligence, scored on its eleven indicators across the three dimensions above. Interpretation bands are provisional pending standard-setting, and are labeled that way wherever they appear.
If you were sent here without a code
Read further
Assessment
The AI Competency Assessment
The companion instrument: a content-validated measure of AI competency by domain, in nine sector tracks. A different question: what someone knows and applies, rather than how they think when the AI answers back.
Future Skills
Critical thinking as a future skill
Why employers rank critical thinking where they do, and what the evidence says about it: the skill this assessment measures in the specific context of AI-assisted work.
