Frequently asked questions
Assessment Analytics Mistakes FAQs
Review common questions about assessment averages, benchmarks,
percentiles, sample size, incomplete attempts, question analytics,
validity, fairness, and dashboard interpretation.
What is the most common mistake in assessment analytics?
One of the most common mistakes is using a single summary metric,
such as the average score, without reviewing the score
distribution, sample size, candidate statuses, assessment version,
or differences between candidate groups.
Why can an average assessment score be misleading?
An average can hide variation, extreme results, multiple candidate
clusters, incomplete attempts, and major differences between
subgroups. It should be reviewed with the median, range,
distribution, sample size, and assessment context.
Can scores from two different assessments be compared?
Direct comparison may be inappropriate when the assessments differ
in content, difficulty, time limit, scoring method, delivery
conditions, language, or candidate population. Comparability should
be established before drawing conclusions.
Why should incomplete assessment attempts be analysed?
Incomplete attempts may reveal technical failures, accessibility
barriers, unclear instructions, excessive duration, candidate
withdrawal, or assessment design problems. Excluding them can make
completion and performance results appear stronger than they are.
What is the difference between a percentile and a percentage score?
A percentage score generally describes the proportion of available
marks earned. A percentile describes a candidate’s relative
position within a reference group. The benchmark population should
be explained when reporting percentiles.
How should small assessment samples be reported?
Display the actual candidate count, avoid excessive decimal
precision, describe the result as preliminary when appropriate, and
avoid broad conclusions that the limited sample cannot reliably
support.
What question-level analytics should be reviewed?
Review question difficulty, response distribution, skipped
responses, time spent, discrimination, candidate complaints,
content relevance, technical display, answer-key accuracy, and
behaviour across relevant groups.
Does a reliable assessment automatically have strong validity?
No. Reliability concerns consistency, while validity concerns
whether the assessment supports the intended interpretation and
decision. An assessment can be consistently scored while measuring
content that is not sufficiently relevant to the role.
How can assessment analytics support fairness review?
Teams can compare completion, technical experience, score
distribution, progression, question behaviour, and other relevant
outcomes across appropriately defined candidate groups while
protecting privacy and avoiding unreliable conclusions from very
small samples.
How can CloudTest support assessment analytics?
CloudTest can support structured online assessment workflows,
candidate attempt tracking, score reporting, question-level review,
and more consistent assessment administration. Available analytics
and platform capabilities may vary by plan and implementation.