{"id":34162,"date":"2017-03-22T01:56:00","date_gmt":"2017-03-22T01:56:00","guid":{"rendered":"https:\/\/hoozoocomm.co.za\/cognadev\/?p=34162"},"modified":"2026-06-29T09:53:18","modified_gmt":"2026-06-29T09:53:18","slug":"assessing-the-validity-of-a-psychological-assessment","status":"publish","type":"post","link":"https:\/\/hoozoocomm.co.za\/cognadev\/2017\/03\/22\/assessing-the-validity-of-a-psychological-assessment\/","title":{"rendered":"Assessing the Validity of a Psychological Assessment"},"content":{"rendered":"<div class=\"wpb-content-wrapper\"><p>[vc_row][vc_column][vc_empty_space height=&#8221;20px&#8221;][vc_custom_heading text=&#8221;Assessing the Validity of a Psychological Assessment&#8221; font_container=&#8221;tag:h1|font_size:35px|text_align:left|color:%23049EB3&#8243; google_fonts=&#8221;font_family:Arvo%3Aregular%2Citalic%2C700%2C700italic|font_style:700%20bold%20regular%3A700%3Anormal&#8221; css=&#8221;&#8221;][vc_empty_space height=&#8221;10px&#8221;][vc_custom_heading text=&#8221;By Paul Barrett&#8221; font_container=&#8221;tag:p|font_size:20px|text_align:left|color:%23000000&#8243; google_fonts=&#8221;font_family:Lato%3A100%2C100italic%2C300%2C300italic%2Cregular%2Citalic%2C700%2C700italic%2C900%2C900italic|font_style:700%20bold%20regular%3A700%3Anormal&#8221; css=&#8221;&#8221;][\/vc_column][\/vc_row][vc_row][vc_column][vc_empty_space height=&#8221;40px&#8221;][vc_column_text css=&#8221;&#8221;]Within test-publisher, psychological assessment training, or assessment\/test review \u2018guidelines\u2019, the terms: construct, content, face, predictive, concurrent, and ecological validity are presented as the definitive \u2018types of validity\u2019.<\/p>\n<p>However, these \u2018types\u2019 were introduced in the mid-20th Century and should have been quietly retired in the late 1990s-early 2000s with the near-simultaneous but quite independent stream of publications from three philosophers of science and measurement who revolutionised how we should think about validity and the use of that word \u2018quantitative\u2019 when applied to psychological attribute assessment.<\/p>\n<h4><strong><b>So who are these \u201cBig Three\u201d of Psychological Measurement &amp; Validity?<\/b><\/strong><\/h4>\n<p>First came Mike Maraun (1997), who demonstrated so clearly the huge logical mistake made by Cronbach and Meehl when they initially proposed their definition of construct validity.<\/p>\n<p>Then came Joel Michell (1998) who laid out the facts and axioms which constitute the definition of a quantity and quantitative measurement, as embodied in all sciences for two or more centuries.<\/p>\n<p>Finally, it was Denny Borsboom (and colleagues) in 2004 who provided a simple definition of validity which clarified the distinction between the \u2018scientific\u2019 perspective and the pragmatic-practical perspective. Note how this\u00a0<strong><b>scientific<\/b><\/strong>\u00a0definition is focused on measurement:<\/p>\n<p>\u201cValidity is not complex, faceted, or dependent on nomological networks and social consequences of testing. It is a very basic concept and was correctly formulated, for instance, by Kelley (1927, p. 14) when he stated that a test is valid if it measures what it purports to measure<\/p>\n<p>A test is valid for measuring an attribute if and only if (a) the attribute exists and (b) variations in the attribute causally produce variations in the outcomes of the measurement procedure.\u201d (p. 1061)<\/p>\n<h4><strong><b>Investigating Validity from a Scientific Perspective<\/b><\/strong><\/h4>\n<p>It\u2019s all about phenomenon detection and\/or an initial theory-claim proposing an attribute\u2019s \u2018existence\u2019 and qualities, then proposing and\u00a0testing\u00a0the \u2018rules\u2019 by which the attribute variations are claimed to be measurable. Those \u2018rules\u2019 embody particular axiomatic equalities\/inequalities which must hold if a rule is to be adjudged valid. This investigative work requires careful experimentation and a \u2018<a href=\"https:\/\/en.wikipedia.org\/wiki\/Metrology\">metrology<\/a>\u2019 mind-set in contrast to the iterations of descriptive correlational workups and covariance analyses by psychologists; as these are already predicated upon an\u00a0assumption\u00a0of quantity.<\/p>\n<p>And therein lies the problem for all psychological attribute measurement which attempts to claim it is making \u2018quantitative\u2019 measurement of an attribute.\u00a0<strong><b>There is no evidence, to date, that any psychological attribute varies as a quantity<\/b><\/strong>. For physical\u00a0<a href=\"https:\/\/physics.nist.gov\/cuu\/Units\/units.html\">quantity examples<\/a>, think length, mass, electrical current;<\/p>\n<h4><strong><b>Investigating Validity from a Pragmatic Perspective<\/b><\/strong><\/h4>\n<p>What this actually boils down to is \u2018validation\u2019; as Borsboom and colleagues (2004) put it:<\/p>\n<p>\u201cValidity is a\u00a0property {of tests}, whereas validation is an activity. In particular, validation is the kind of activity researchers undertake to find out whether a test has the property of validity. Validity is a concept like truth: It represents an ideal or desirable situation. Validation is more like theory testing: the muddling around in the data to find out which way to go.\u201d p. 1063.<\/p>\n<p>In applied practice, the answers to certain obvious, relevant-to-deployment questions are what test publishers\/authors need to convey to those wishing to use their assessments. Why? Because these answers form the \u2018evidence-base\u2019 which will be used by others to form a\u00a0validation\u00a0judgement i.e. an informed judgement as to whether the assessment will be fit-for-purpose, a\u00a0priori\u00a0justifiable, and ultimately legally defensible.<\/p>\n<p>So, what are some of these questions?<\/p>\n<h4><strong><b>1.Do those who use this test find it of benefit; if so, what is\/are these benefits?<\/b><\/strong><\/h4>\n<p>Such information has to be acquired via a few standard but open-ended survey questions asked of assessment users\u00a0(whether phone-call, personal meeting, or on-line survey). That qualitative information can be formally categorized and summarised, then written up as a simple one-page infographic. If the previous deployments of an assessment are adjudged favourably by users for the various reasons they provide, that\u2019s \u2018good enough\u2019\u00a0preliminary\u00a0evidence of pragmatic validity. Why? Because if the assessment was producing random results which made no coherent or consistent sense, no user would give it a positive rating.<\/p>\n<p>&nbsp;<\/p>\n<h4><strong><b>2. Does it assess what it is claimed it assesses?<\/b><\/strong><\/h4>\n<p>This is all about presenting information which justifies a claim that an assessment assesses magnitudes of a particular attribute, or class-category types. In practice this is more about developing a line of plausible reasoning based upon some empirical and logical analytical workups rather than referring to some abstract notion of \u2018concurrent validity\u2019. For example, we already know that in personality research, there is only moderate agreement between assessments claiming to assess the same-named \u2018constructs\u2019 (Pace and Brannick, 2010). And, as we know from Mike Maraun\u2019s expositions, given we have no \u2018technical\u2019 definition of any attribute and no evidence that any psychological attribute varies as a quantity, we are simply looking for \u2018good enough\u2019 justifications here. This is not physics or chemistry, no matter how many wish it to be so.<\/p>\n<p>&nbsp;<\/p>\n<h4><strong><b>3. If I give the same assessment tomorrow or next week to the same candidate, will they attain more-or-less the same results?<\/b><\/strong><\/h4>\n<p>I know, I can hear you say \u201cbut surely this about reliability!\u201d And so it is. But when forming a judgement about whether an assessment is appropriately justified\/ validated for your particular deployment, you need to know the answer to this question. Obviously, if what you propose to assess is something you expect to change dramatically on a day-to-day basis, this question is irrelevant. But for the vast majority of assessment applications used in the workplace, we are looking at attributes which comprise a stable feature of individuals over short periods of time.<\/p>\n<p>&nbsp;<\/p>\n<h4><strong><b>4. If an assessment is being used on the basis that its \u2018scores\u2019 or \u2018indicators\u2019 predict certain outcomes, do they actually do so? i.e. What is its predictive accuracy?<\/b><\/strong><\/h4>\n<p>The usual Pearson-correlations-as-validity-coefficients used by many to answer this question are assumption-laden (yes, that quantity assumption again!) estimates of monotonic agreement in which all magnitude information has been carefully removed by the computations forming the estimate. In short, mildly-amusing but not what a user really wants to know here. Indexing predictive accuracy requires analyses conducted in the metric of the observations themselves, looking at observed vs predicted magnitude discrepancies, counting success\/failures, producing mis-classification tables and rates, and above all, using V-fold or holdout-sample cross-validation of any model which claims to be \u2018predictive\u2019 of an outcome.<\/p>\n<p>&nbsp;<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.cognadev.com\/assets\/images\/blog\/Botttom_line.jpg\" \/><\/p>\n<p>Practitioners do not set out to evaluate the property of validity of a test or assessment by looking for evidence justifying: \u201cvariations in the attribute causally produce variations in the outcomes of the measurement procedure\u201d\u00a0(the scientific question). Rather, they want to make an informed personal judgement based upon the \u2018validation\u2019 information associated with an assessment that is directly relevant to its deployment in the workplace. To that end they need clear answers to some basic questions about\u00a0if and how\u00a0an assessment has been validated for\u00a0productive\u00a0use in the workplace. Put more simply, does it do what it says on the box?<\/p>\n<p>&nbsp;<\/p>\n<h4>References<\/h4>\n<p>Borsboom, D., Mellenbergh, G.J., &amp; Van Heerden, J. (2004). The concept of validity.\u00a0Psychological Review, 111, 4, 1061-1071.<\/p>\n<p>Borsboom, D., Cramer, A.O.J., Kievit, R.A., Scholten, A.Z., &amp; Franic, S. (2009). The end of construct validity. In Lissitz, R.W. (Eds.).\u00a0The Concept of Validity: Revisions, New Directions, and Applications\u00a0(Chapter 7, pp. 135-170). Charlotte: Information Age Publishing.<\/p>\n<p>Maraun, M.D. (1998). Measurement as a Normative Practice: Implications of Wittgenstein\u2019s Philosophy for Measurement in Psychology.\u00a0Theory &amp; Psychology, 8, 4, 435-461.<\/p>\n<p>Michell, J. (1997). Quantitative science and the definition of measurement in Psychology.\u00a0British Journal of Psychology, 88, 3, 355-383.<\/p>\n<p>Michell, J. (2009). Invalidity in Validity. In Lissitz, R.W. (Eds.).\u00a0The Concept of Validity: Revisions, New Directions, and Applications\u00a0(Chapter 6, pp. 111-133). Charlotte: Information Age Publishing.<\/p>\n<p>Pace, V.L., &amp; Brannick, M.T. (2010). How similar are personality scales of the \u201csame\u201d construct? A meta-analytic investigation.\u00a0Personality and Individual Differences, 49, 7, 669-676.[\/vc_column_text][\/vc_column][\/vc_row]<\/p>\n<\/div>","protected":false},"excerpt":{"rendered":"<p>[vc_row][vc_column][vc_empty_space height=&#8221;20px&#8221;][vc_custom_heading text=&#8221;Assessing the Validity of a Psychological Assessment&#8221; font_container=&#8221;tag:h1|font_size:35px|text_align:left|color:%23049EB3&#8243; google_fonts=&#8221;font_family:Arvo%3Aregular%2Citalic%2C700%2C700italic|font_style:700%20bold%20regular%3A700%3Anormal&#8221; css=&#8221;&#8221;][vc_empty_space height=&#8221;10px&#8221;][vc_custom_heading text=&#8221;By Paul Barrett&#8221; font_container=&#8221;tag:p|font_size:20px|text_align:left|color:%23000000&#8243; google_fonts=&#8221;font_family:Lato%3A100%2C100italic%2C300%2C300italic%2Cregular%2Citalic%2C700%2C700italic%2C900%2C900italic|font_style:700%20bold%20regular%3A700%3Anormal&#8221; css=&#8221;&#8221;][\/vc_column][\/vc_row][vc_row][vc_column][vc_empty_space height=&#8221;40px&#8221;][vc_column_text<\/p>\n","protected":false},"author":1,"featured_media":34161,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[44],"tags":[],"class_list":["post-34162","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-assessment-issues"],"_links":{"self":[{"href":"https:\/\/hoozoocomm.co.za\/cognadev\/wp-json\/wp\/v2\/posts\/34162","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/hoozoocomm.co.za\/cognadev\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/hoozoocomm.co.za\/cognadev\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/hoozoocomm.co.za\/cognadev\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/hoozoocomm.co.za\/cognadev\/wp-json\/wp\/v2\/comments?post=34162"}],"version-history":[{"count":1,"href":"https:\/\/hoozoocomm.co.za\/cognadev\/wp-json\/wp\/v2\/posts\/34162\/revisions"}],"predecessor-version":[{"id":34163,"href":"https:\/\/hoozoocomm.co.za\/cognadev\/wp-json\/wp\/v2\/posts\/34162\/revisions\/34163"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/hoozoocomm.co.za\/cognadev\/wp-json\/wp\/v2\/media\/34161"}],"wp:attachment":[{"href":"https:\/\/hoozoocomm.co.za\/cognadev\/wp-json\/wp\/v2\/media?parent=34162"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/hoozoocomm.co.za\/cognadev\/wp-json\/wp\/v2\/categories?post=34162"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/hoozoocomm.co.za\/cognadev\/wp-json\/wp\/v2\/tags?post=34162"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}