The science library

CHAPTER 05 / 4 MIN READ

Consistency, without false precision

Internal consistency, repeated measurement, and the uncertainty behind a score.

Reliability is a question about a procedure

Reliability concerns how consistently a measurement procedure distinguishes among people under specified conditions. Those conditions matter. Repeating the same questionnaire next week, using a different item set, and asking a different observer are different questions. A reliability coefficient needs a sample, a design, and an interpretation. It is not a permanent property that can be attached to a name and carried into every language, population, or purpose.

For our inventory, useful evidence would include how each domain’s items work together and how scores behave when the same adults answer again after a defined interval. A report should also communicate the limits of precision. Before those studies exist, displaying a very exact number cannot make its psychological interpretation equally exact. We therefore treat current scores as descriptive indices and avoid presenting tiny differences as established differences in a person’s underlying preferences.

Why alpha is not the whole answer

Sijtsma’s examination of coefficient alpha explains why it cannot, by itself, establish that a scale measures one dimension. Alpha depends on the pattern of item relationships and on the number of items. A collection of overlapping questions can produce an impressive value while covering a narrow slice of behavior. Structural evidence is needed to understand what the items share. A numerical threshold should not replace that investigation.

Dunn, Baguley, and Brunsden discuss coefficient omega as a practical alternative for internal consistency estimation and emphasize interval estimates. Omega also involves assumptions and a measurement model; choosing a more sophisticated statistic does not remove the need to check them. Our proposed technical work should report the estimator, model, sample, and uncertainty for each domain. We should examine whether the model is defensible before treating a coefficient as support for interpretation.

Repeating a score asks another question

Koo and Li show that intraclass correlation coefficients come in different forms, with different assumptions. In particular, consistency and absolute agreement address different questions. Two sets of scores can preserve people’s ordering while all scores shift upward. A study interested in agreement must not mistake that pattern for identical measurement. The exact coefficient and its confidence interval should be reported, together with the design that justifies it.

For a proposed retest study, we would specify the interval, participation conditions, and any major events between administrations. An interval that is too short may involve memory of answers; a longer interval gives more opportunity for genuine changes in circumstances. Our application-specific question is whether the inventory produces a sufficiently stable summary for reflective use while remaining sensitive to the fact that people answer from a particular context. No retest result is claimed for the present release.

What a careful reader should do

Suppose a participant receives two neighboring domain indices, 64 and 66. Without an estimate of measurement error, there is no justified basis for saying that this two-point difference marks a clear psychological boundary. Even with uncertainty estimates, interpretation would depend on the purpose and the evidence. In a developmental discussion, it is often more productive to explore both preferences and their examples than to elevate one number into a definitive identity.

Our report should make room for that ambiguity. Close scores invite a blended reading; a repeated assessment invites a conversation about circumstances before a claim of personal change. Current outputs must not contain invented reliability values, confidence bands, or a claim that 100 questions guarantees precision. The planned research can estimate uncertainty. Until then, transparent limits are part of a useful explanation of the results, not an optional technical footnote.