The science library

CHAPTER 04 / 4 MIN READ

The work behind every question

Why a 100-item inventory needs a blueprint, participant feedback, and empirical revision.

Coverage comes before length

One hundred questions is a design specification, not a quality certificate. Clark and Watson’s account of scale development begins with a clear construct and a broad initial item pool. Their argument emphasizes content coverage, careful wording, and testing with a sample that represents the intended range. Repeating similar statements can make an inventory longer without making its interpretation richer. A development process must decide what each question adds and what relevant behavior remains unrepresented.

Our blueprint assigns 25 original items to each of four intended domains. The purpose is balanced coverage for a developmental profile. Equal item counts do not imply equal reliability or establish that all questions work equally well. Some items may eventually need revision or removal. The commitment is to the quality of the final interpretation, not to preserving a particular sentence merely because it appeared in the first release.

Writing is only the first stage

Boateng and colleagues describe a development sequence that includes defining the domain, generating items, assessing content, pretesting, examining structure, and evaluating reliability and validity. Their framework also distinguishes expert review from feedback by intended respondents. Experts can assess whether an item belongs in a domain; participants can reveal how its language is actually understood. These forms of evidence serve different purposes and are both more informative than assuming that a plausible sentence is a good measurement item.

Our proposed review asks one question at a time. Does this statement describe a single behavior? Can someone answer without managing a team or working in a particular occupation? Does agreement sound like a moral achievement? Does the wording depend on an idiom? Could an answer change simply because the participant imagines a different time period? These are concrete design checks to apply before statistical evaluation, followed by interviews that test whether the intended meaning survives actual use.

Response options shape the task

The response scale is part of the instrument. In a randomized study of personality questionnaire formats, Simms and colleagues compared different numbers of response options and found that precision and criterion relationships did not follow exactly the same pattern. That study does not establish a universally best format for our questions. It shows why the number and meaning of response options should be examined empirically instead of treated as decoration added after item writing.

For this inventory, participants use a consistent five-point agreement scale. Keeping labels visible helps them apply the same frame throughout the questionnaire. The middle option should not be portrayed as a poor response: sometimes a statement fits only in certain settings. Similarly, an extreme response is not automatically a stronger personality. It is simply stronger endorsement of a statement under the questionnaire’s instructions. The development study must examine how people use each category.

A better item earns its place

Imagine a hypothetical statement that combines speed and confidence: ‘I decide quickly and rarely doubt myself.’ A participant who decides slowly but confidently cannot answer it cleanly. Splitting the behaviors may produce more interpretable responses. A different statement might mention managing employees and therefore miss people who demonstrate initiative in informal groups. These examples show why item editing is a substantive part of assessment science rather than a final copyediting pass.

The current questions remain candidates within an original developmental inventory. Cognitive interviews, item distributions, response-category analyses, and structural studies have not yet been completed for them. Future changes should be documented by version so that a report can always be traced to its questionnaire and scoring rules. A scientifically serious product should be willing to improve its items when evidence warrants it, while explaining what changed and whether previous scores remain comparable.