Start with a precise claim
The Standards for Educational and Psychological Testing frame validity around evidence and theory supporting proposed score interpretations and uses. Different evidence can address item content, response processes, internal structure, and relationships with other variables. The practical consequence is that ‘validated’ is incomplete unless we know what interpretation is supported, for whom, and for what purpose. An attractive report or a long bibliography is not a substitute for evidence about the actual instrument.
Our current claim is deliberately bounded: the questionnaire summarizes how an adult describes their behavioral preferences across four intended domains. It is a research-informed developmental inventory. We have not established its factor structure, norms, criterion validity, or suitability for consequential decisions. That statement applies to the current item set and scoring approach. It should remain visible wherever a purchaser or participant might otherwise assume that published research on personality automatically validates this product.
Relationships should be meaningful and selective
Campbell and Fiske’s multitrait–multimethod framework distinguishes convergence from discrimination. A proposed measure should relate to other ways of assessing a similar construct while retaining meaningful separation from different constructs. It also encourages attention to shared methods: two self-report questionnaires may resemble one another partly because of how information is collected. A strong validation programme therefore looks beyond a single correlation and considers multiple sources of evidence.
Applied to this inventory, we might preregister expectations about relationships between directness items and independent observations of speaking up in a structured group task. We would also specify constructs those items should not strongly represent. The task and comparison measures would need their own justification. This is an example of a possible study, not a claim that the research has been conducted or that speaking often is necessarily an effective behavior.
Uses introduce additional assumptions
Kane’s argument-based approach asks developers to set out the inferences and assumptions connecting responses to conclusions and uses. More ambitious claims require more support. Evidence for an interpretation does not automatically justify every decision that could be made from it. This distinction is especially important when a report designed for reflection is repurposed as a ranking system. The evidence needed changes when the consequence changes.
For Machiavelli Assessments, helping someone prepare a conversation about their preferred working pace is different from predicting job success, diagnosing a condition, or deciding who should receive an opportunity. The present inventory is not offered for those latter purposes. A purchaser should not convert a development profile into a pass/fail threshold, a candidate shortlist, or a claim about someone’s trustworthiness. Such conclusions extend beyond both the questions and the available evidence.
A useful research result can be inconvenient
Imagine a study in which two intended domains cannot be distinguished clearly. The responsible response would be to inspect the definitions, items, and model rather than conceal the overlap. Or suppose the profile prompts engaging conversations but does not predict an outcome that was proposed in advance. The conversational experience can still be studied on its own terms, while the unsupported prediction is withdrawn. A coherent scientific position needs room for both possibilities.
Our editorial principle is that the strength of the language should follow the strength of the evidence. ‘Designed to explore’ describes an intention. ‘Associated with’ requires observed data and context. ‘Predicts’ requires a suitable predictive study and an account of uncertainty. ‘Improves’ requires evidence about an intervention and competing explanations. Keeping those verbs distinct helps readers understand what is known today and what remains a research question.
SOURCES & FURTHER READING
AERA, APA, & NCME (2014). Standards for Educational and Psychological Testing. Chapters 1–3. ↗Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait–multimethod matrix. Psychological Bulletin, 56, 81–105. ↗Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50, 1–73. ↗