Skip to content
American Institute of Business & TechnologyAmerican Institute ofBusiness & TechnologyAmerican Institute of Business & Technology home

Assessment

When the AI sounds certain, does your team check?

Generative AI answers with structure and confidence whether it is right or wrong. The AI Critical Thinking Assessment measures whether a person actually verifies, interrogates, and corrects AI output under working conditions, not whether they say they would.

What it is

A two-part instrument. Part one is a validated thirteen-item scale of how a person reports using generative AI. Part two is a workplace simulation: the participant advises on a realistic business decision, consults an AI advisor that has been deliberately seeded with flawed reasoning, and writes two short memos that are scored against a published behavioral rubric.

Structure
A 13-item self-report scale, then an interactive simulation with two written memos.
Time
35 to 50 minutes, in one sitting. Progress is saved on the device.
Scoring
Eleven behavioral indicators across three dimensions, plus three self-report subscales.
Result
A disposition profile, a demonstrated-competence index, and the gap between them. Preliminary at completion, final after human review.
Access
By organization code. There is no individual purchase today.

The problem it measures

The failure mode it targets is automation bias: accepting a well-structured, authoritative answer without checking it. In the simulation, the AI advisor is confidently wrong in specific, planted ways: a dismissed labor risk, a fabricated regulatory certainty, a security guarantee that conflates uptime with safety, a statistical bias reframed as natural volatility. The score reflects which of those the participant caught, challenged, and corrected in writing.

Because part one records what people say and part two records what they do, the assessment also surfaces the intention–behavior gap: the people who report strong verification habits but did not verify when it counted. That group is invisible to survey-only instruments, and it is exactly the group targeted training helps most.

How the simulation works

Part 1

Self-reported disposition

Thirteen statements about verification habits, motivation to understand how AI produces its answers, and reflection on responsible use. Scored on three subscales from the published Critical Thinking in AI Use scale (Lau et al., 2026).

Part 2

Live consultation and memos

The participant questions an AI advisor about a strategic decision, then submits a board memo and a critique of the AI's output. Chat behavior and both memos are scored against the eleven-indicator Human-AI Interaction rubric (Li et al., 2025).

What the scores cover

Dimension 1

Retrieval and analysis

How the participant works the advisor: question depth, coverage of distinct risk perspectives, follow-up chains, and how prompts are refined as the picture develops.

Dimension 2

Evaluation

Whether flawed claims are challenged at all, and whether the challenge is substantiated: evidence demanded, logic tested, errors named in writing.

Dimension 3

Integration

What the memos are built from: screened and filtered evidence rather than pasted output, multiple perspectives weighed, and a logically consistent argument to a clear recommendation.

What an organization receives

The simulation is standardized, deliberately

Every participant faces the same advisor with the same planted flaws. That is what makes scores comparable across a cohort and defensible to the people being assessed. The advisor is scripted, not a live model; the grading of memos and consultation behavior is where careful automation is applied, and every automated grade carries a stated reason a reviewer can audit.

Two stages: preliminary, then final

Scores are issued twice, and say so. At completion the participant receives a preliminary report, produced by automated grading and labeled as preliminary on every surface. A reviewer then audits the automated grades, re-scores once, and the final report goes out, stamped with both the submission and the review date. No participant is ever left holding a number the institution has not stood behind.

The reports, at a glance

Each participant receives a personal report: what they say, what they did, the gap between the two stated plainly, and concrete development recommendations. It arrives first as a clearly labeled preliminary report, then as the final version once a human review is complete.

Sample participant report: disposition and competence headline scores, subscale and dimension breakdowns, and development recommendations
The participant report. Sample data for illustration.

The organization sees how every score was produced: each behavioral indicator with its score, its source, and the grader's stated reason, above the full consultation transcript and both memos. Every grade can be audited and re-scored.

Sample administrator view: score summary cards and the indicator grading table with per-indicator grading reasons
The administrator's grading audit view. Sample data for illustration.

And the cohort map plots reported habits against demonstrated performance. The upper-left group, people who report strong verification habits but did not verify in the simulation, is invisible to survey-only instruments and is exactly where targeted development pays off.

Sample cohort intention-behavior map: a scatter of participants by self-reported disposition versus demonstrated competence, with the gap group flagged
The cohort intention-behavior map. Sample data for illustration.

Methodological grounding

Part one applies the Critical Thinking in AI Use scale of Lau et al. (2026), published in Computers in Human Behavior Reports. Part two applies the performance-based Human-AI Interaction competence rubric of Li et al. (2025), published in the Journal of Intelligence, scored on its eleven indicators across the three dimensions above. Interpretation bands are provisional pending standard-setting, and are labeled that way wherever they appear.

If you were sent here without a code

You need a code from your organization to begin, and there is currently no way to buy one on your own. Registering interest is the whole of what you can do today, and it takes a minute.

Read further