AI Series

The AI Series 1.2

Do the new methods of “AI” analysis generate more accurate measures or predictions?

By Paul Barrett

Predicting the scores of existing self-report personality assessments

I’m looking at a selection of studies from researchers who have deployed a variety of Machine Learning (ML) algorithms with the goal of forming machine-generated measures of personality attributes. I’m specifically reporting studies that provide correlations between self-report vs machine-generated personality scores. The list is not intended to be exhaustive; I’m just summarising those studies which provide clear validity data where the author/s are attempting to recommend supplanting self-report or existing psychometric scores with machine-generated versions. And I’m reporting actual comparative-validity numerical magnitudes rather than the usual inaccurate verbal descriptions of them.

# Author/s Year Research Theme Results
1 Kosinski et al 2013 Personality traits are predictable from Facebook Likes

Correlations between machine generated personality scores and a 20-item Big Five self-report questionnaire scores (Figure 3, p. 5804):

.43 → Openness
.29 → Conscientiousness
.40 → Extraversion
.30 → Agreeableness
.30 → Emotional Stability (Neuroticism)

2 Park et al 2015 Automatic Big Five personality assessment using Facebook personal status message content analysis Correlations between machine generated personality scores and those from a 100-item Big Five self-report questionnaire (Table 1, p. 940):

.46 → Openness
.38 → Conscientiousness
.41 → Extraversion
.40 → Agreeableness
.39 → Neuroticism

3 Janz (Humantic.ai) 2019 This study reported the relationship between the author’s subjective professional evaluation on 16 personal characteristics of 120 members of his professional network and machine-derived scores on the same factors.

Correlations between the machine generated Big 5 personality scores and the subjective ratings from Tom Janz (p. 4, column 1 of the table; Neuroticism wasn’t rated):

.53 → Openness
.37 → Conscientiousness
.39 → Extraversion
.43 → Agreeableness

4 Stachl et al 2020 Predicting scores on the 60-item German version of the Big Five Structure Inventory using analysis of smartphone use/activity over 30 days.

Correlations between machine generated personality scores and the Big Five scale scores (Table S4, p. 11, Supplemental Material):

.29 → Openness
.31 → Conscientiousness
.37 → Extraversion
.05 → Agreeableness
.11 → Neuroticism

5 Kachur et al 2020 Explored the relationship between participant self-report Big Five personality traits and their real-life static facial images Correlations between machine generated personality scores and a modified Russian version of the 5PFQ questionnaire, which is a 75-item measure of the Big Five model (Table 2, p. 4). The results are reported separately for males and females:

Males (n=505):

.19 → Openness
.36 → Conscientiousness
.19 → Extraversion
.21 → Agreeableness
.21 → Neuroticism

Females (n=740):

.14 → Openness
.34 → Conscientiousness
.27 → Extraversion
.24 → Agreeableness
.28 → Neuroticism

 

6 Marengo et al 2023 A meta-analysis of 21 published studies assessing the association between smartphone use/activity data and Big Five traits.

Meta-analytic correlations between machine generated personality scores and the Big Five scale scores (Table S3, p. 8, Supplemental Material; after removal of outliers):

.22 → Openness
.23 → Conscientiousness
.34 → Extraversion
.22 → Agreeableness
.22 → Neuroticism

7 Fan et al 2023 The study reports the results from exploring the plausibility of measuring Big Five self-report personality scores indirectly from data acquired from individuals interacting between 20-30 minutes with an artificial intelligence (AI) chatbot.

Correlations between machine generated personality scores and the Big Five IPIP-300-item questionnaire scale scores (Table 6, p. 1291):

.58 → Openness
.46 → Conscientiousness
.45 → Extraversion
.42 → Agreeableness
.40 → Neuroticism

The results are hardly encouraging. Mostly inaccurate and just offering ‘somewhere in the ball-park’ information. But this kind of study assumes self-report assessed personality is the ‘gold standard’ measurement against which alternative machine-generated measures must be compared. The problem of course is that self-report questionnaire scores may not be the kind of accurate criterion measurement required to establish ‘equivalence of measurement’, especially since several articles have previously shown how even personality scales with the same name differ substantially (in score comparability/meaning) from one another, or the measurement assumptions upon which the scales rely for their validity are simply untenable:

Let’s face it, we are not calibrating an alternative measure of length using a previously calibrated measuring instrument (e.g. a ruler or steel tape). But then length is a quantitively structured base-unit variable; a personality attribute is not. So, attempting to predict personality scores with sufficient accuracy and generalizability that they could be replaced by machine-generated scores using other kinds of observed behavioural or “digital-footprint” data was never going to work – from first principles of measurement let alone the conceptual and known semantic ‘haze’ of scale content.

Articles listing or referencing other kinds of “AI” studies

  • Ihsan, Z., & Furnham, A. (2018). The new technologies in personality assessment: A review. Consulting Psychology Journal: Practice and Research, 70, 2, 147-166. http://dx.doi.org/10.1037/cpb0000106  [Paywall]

listed 39 studies which have attempted to relate ‘machine-generated’ assessments of personality with external behavioural criteria and sometimes actual self-report questionnaire scores. On the basis of the reported evidence, they concluded:

“At this stage there is more absence of evidence of the psychometric properties of these new approaches than evidence of absence of their validity.” (from the article abstract)
and concluded the article with this final paragraph:

“It has been said about some psychological tests that they are techniques in search of a theory or solutions to non-existent problems. There is always the danger when exploiting the possibilities of new technology that insufficient evidence is collected and assessed to show incremental validity over existing and established methods. The history of psychology is littered with examples of how the primary technology of the time shapes not only theories and technology but also how many promises were never delivered.

More recently, the research strategy has focused on using AI algorithms to create predictive models of relevant variables from a wide variety of ‘digital-footprint’ and natural language processing (NLP) of text analysis of application material (e.g. application letters, resumes, open-ended text responses to questions). No longer concerned with predicting self-report questionnaire scale scores but now developing new ‘summary variables’ which possess some predictive facility with relevant outcome variables. Examples of these studies are provided in two recent “editorials to a special issue”, which summarize the research papers in the respective issues.

  • Campion, M.A., & Campion, E.D. (2023). Machine learning applications to personnel selection: Current illustrations, lessons learned, and future research. Personnel Psychology, 76, 4, 993-1009. https://doi.org/10.1111/peps.12621  [Open Access]
  • Woo, S.E., Tay, L., & Oswald, F. (2024). Artificial intelligence, machine learning, and big data: Improvements to the science of people at work and applications to practice. Personnel Psychology, Online First, 1-16. https://doi.org/10.1111/peps.12643  [Paywall]

Of interest perhaps is the “Lessons Learned” section in Campion and Campion (2023; Table 3, pp. 1001-1002), abbreviated here:

“1. Relevance of estimating the reliability of ML models. ML researchers do not always analyze reliabilities, but they should for all the same reasons we do otherwise.

2. Influence of reliability on prediction. If models are trained against a criterion, then the correlation with that criterion is influenced by the reliability of the criterion.

3. Criterion-related validity of ML algorithms in employment applications. Evidence currently suggests that we may be able to build ML models that demonstrate equivalent or better criterion-related validity. Criterion-related validity is essential to supporting the utility of an employment assessment.

5. Improvement in prediction. When examining ML algorithms compared to traditional methods with large samples, the improvement in prediction is probably not going to be large in most instances.

8. Scoring new types of data is the greatest opportunity. We believe the greatest opportunity afforded by ML is the ability to score data that have been relatively neglected—text data, as well as other unstructured responses—as opposed to improving prediction from traditional data such as the dominant multiple-choice and rating scale responses or other structured numeric responses in current assessments.“

With reference to Point #5 above: “When examining ML algorithms compared to traditional methods with large samples, the improvement in prediction is probably not going to be large in most instances”, this is exemplified in a more recent article: Campion, E., Campion, M., Johnson, J., Carretta, T., Romay, S., Dirr, B., Deregla, A., & Mouton, A. (2024). Using natural language processing to increase prediction and reduce subgroup differences in personnel selection decisions. Journal of Applied Psychology, 109, 3, 307-338. https://doi.org/10.1037/apl0001144  [Paywall] , where incremental prediction of US Air Force Officer Training Assessment Board applicant scores using AI-created variables resulted in regression-model deviation R-squares of between 0.03 and 0.09 over existing variables.

The latest British Psychological Society (BPS) article on these issues (Spring issue, ADM, 2024)

Some readers of this blog who are members of the British Psychological Society I/O section, might be aware of this article appearing in the 2024 Spring issue of Assessment and Development Matters, entitled:

“In what ways will AI enhance psychometric testing in the workplace?”

https://explore.bps.org.uk/content/bpsadm/16/1/24

 

Abstract ” It has been said about some psychological tests that they are techniques in search of a theory or solutions to non-existent problems. There is always the danger when exploiting the possibilities of new technology that insufficient evidence is collected and assessed to show incremental validity over existing and established methods. The history of psychology is littered with examples of how the primary technology of the time shapes not only theories and technology but also how many promises were never delivered. “

It was based around six article references from a variety of APA and other journals, none of which I’d come across. So, I tried to access each reference for my own records. All are non-existent. The author details at the end of the article includes a sentence stating that the article was written entirely by ChatGPT4, but it was published without any warning that it was a complete fabrication (as the journal editorial team thought it would be self-obvious to readers).

 

But those 6 references look entirely credible at first glance e.g.

  • Garcia, F., et al. (2020). AI-powered analytics in psychometric testing: Improving talent prediction accuracy. Journal of Human Resources Management, 35(4), 201–218.
  • Johnson, A. & Thompson, B. (2018). The impact of artificial intelligence on psychometric testing. Journal of Applied Psychology, 123(2), 86–94.
  • Li, T., et al. (2019). Application of artificial intelligence in psychometric testing: A comparative analysis. Journal of Organizational Behavior, 40(3), 115–131.
  • Li, M., Wang, Q. & Chen, H. (2020). The predictive validity of AI-driven psychometric testing in job performance. Journal of Organizational Behavior, 67(3), 210–225.
  • Smith, A. & Johnson, L. (2019). Enhancing personality assessment accuracy through AI algorithms Journal of Applied Psychology, 45(2), 123–140.
  • Wang, H. & Chen, S. (2021). Reducing bias in personality assessments using artificial intelligence: A comparative study. Journal of Applied Social Psychology, 45(2), 76–89.

 

So, some readers of the BPS article might well have wondered why I did not quote any of these references or even mention the article. Now they know! A recent editorial in Nature magazine is perhaps relevant here:

Editorial (online, 6th March) (2024). Why scientists trust AI too much – and what to do about it. Nature, 627, 8002, https://doi.org/10.1038/d41586-024-00639-y  [Open Access] summarizing the full article in the same issue:

Messeri, L., & Crockett, M.J. (2024). Artificial intelligence and illusions of understanding in scientific research. Nature, 627, 8002, 49-58. https://doi.org/10.1038/s41586-024-07146-0  [Paywall]

In Conclusion

Well, you’ve now seen some of what I see, while being told by a guru, academic, or consultant that if I’m not using AI or redeveloping existing assessments to incorporate AI, then it’s “game over” – a sentiment similarly observed in a recent humorous but thought-provoking “Rolling Stone” article:

Evans, R. (2024). The cult of AI. Rolling Stone, January, 27th, , 1-10. https://www.rollingstone.com/culture/culture-features/ai-companies-advocates-cult-1234954528/  [Open Access] summarizing the full article in the same issue:

And reflected in a Financial Times article published at the end of 2023:
The AI revolution’s first year: has anything changed?
The launch of ChatGPT was heralded as the dawn of a new age. But companies are wondering how useful generative AI really is https://www.ft.com/content/f84bd56f-484d-4393-bafc-4428da6a7873  [Paywall]

or just a week or so ago:
Beware AI euphoria: Like all great bubble stories, the latest tech narrative conveys a sense of inevitability. https://www.ft.com/content/599a5c5b-dc59-4724-8248-2d4132ffdb7f  [Paywall]

However, there are applications that make clear financial and logistical sense as the article I referenced above explains (Campion et all, 2024, p. 334, 2nd column, “Practical Implications”) with an estimated saving of one-third the current cost ($169,334.40 to $310,382.40 annually) of the selection decisions, if using the NLP model to replace just one selection Board member. That’s significant and a clear guide as to where and how AI will make a difference in organizations i.e. Decision-Support.

For example, in the legal profession, this recent article shows just how much cost-saving could be made by replacing certain ‘junior’ legal services with ChatGPT:
Martin, L., Whitehouse, N., Yiu, S., Catterson, L., & Perera, R. (2024). Better call GPT, comparing large language models against lawyers. arXiv Preprint , arXiv:2401.16212v1 [cs.CY] 24 Jan 2024, 1-16. https://arxiv.org/html/2401.16212v1  [Open Access]

 

Abstract This paper presents a groundbreaking comparison between Large Language Models (LLMs) and traditional legal contract reviewers—Junior Lawyers and Legal Process Outsourcers (LPOs). We dissect whether LLMs can outperform humans in accuracy, speed, and cost-efficiency during contract review. Our empirical analysis benchmarks LLMs against a ground truth set by Senior Lawyers, uncovering that advanced models match or exceed human accuracy in determining legal issues. In speed, LLMs complete reviews in mere seconds, eclipsing the hours required by their human counterparts. Cost-wise, LLMs operate at a fraction of the price, offering a staggering 99.97 percent reduction in cost over traditional methods. These results are not just statistics—they signal a seismic shift in legal practice. LLMs stand poised to disrupt the legal industry, enhancing accessibility and efficiency of legal services. Our research asserts that the era of LLM dominance in legal contract review is upon us, challenging the status quo and calling for a reimagined future of legal workflows.

 

This was reported in New Scientist Business Insights newsletter (21st February, 2024):

 

” While junior lawyers would pore over a document for 56 minutes on average and the legal outsourcers for 3 hours 21 minutes, GPT-4 could spend as little as 2 minutes to get largely the same results. Other LLMs, including Claude and Google’s PaLM 2 were as quick or even quicker: PaLM 2 took around 45 seconds. And while junior lawyers cost around $75 per document analyzed, LLMs could cost mere cents. It all adds up to a worrying future for those wanting to make a living in the field of law “.

 

But in answer to my own question, the title of this blog: Do the new methods of “AI” analysis generate more accurate measures or predictions? I have to conclude ‘not really’. We can do a bit of metaphorical hand-waving and say: “this is new technology and will take time to fully leverage’, but the real issue is “us” humans – complex cognitive open systems on two legs. All of which is explained in detail in the forthcoming blog #3 of this series!

References

Table References

#1: Kosinski, M.D., Stillwell, D., & Graepel, T. (2013). Private traits and attributes are predictable from digital records of human behavior. Proceedings of the National Academy of Sciences, 110, 15, 5802-5805. https://doi.org/10.1073/pnas.1218772110  [Open Access]

#2: Park, G., Schwartz, H.A., Eichstaedt, J.C., Kern, M.L., Kosinski, M., Stillwell, D.J., Ungar, L.H., & Seligman, M.E.P. (2015). Automatic personality assessment through social media language. Journal of Personality and Social Psychology, 108, 6, 934-952. https://doi.org/10.1037/pspp0000020  [Paywall]

#3: Janz, T. (2019). The Accuracy of Instant Talent Analytics: A Personal Network Validation. HRExaminer, https://www.hrexaminer.com/instant-talent-assessment-appeals-but-does-it-work/ , v10.24, , 1-7. https://www.hrexaminer.com/wp-content/uploads/2019/06/The-Accuracy-of-Instant-Talent-Analytics.pdf  [Open Access]

#4: Stachl, C., Au, Q., Schoedel, R., Gosling, S.D., Harari, G., Buschek, D., Volkel, S.T., Schuwerk, T., Oldemeier, M., Ullmann, T., Hussmann, H., Bischi, B., & Bühner, M. (2020). Predicting personality from patterns of behavior collected with smartphones. Proceedings of the National Academy of Sciences, 117, 30, 17680-17687. https://doi.org/10.1073/pnas.1920484117  [Open Access]

#5: Kachur, A., Osin, E., Davydov, D., Shutilov, K., & Novokshonov, A. (2020). Assessing the Big Five personality traits using real-life static facial images. Nature: Scientific Reports, 10, 1, #8487, 1-11. https://doi.org/10.1038/s41598-020-65358-6  [Open Access]

#6: Marengo, D., Elhai, J.D., & Montag, C. (2023). Predicting Big Five personality traits from smartphone data: a meta-analysis on the potential of digital phenotyping. Journal of Personality, 91, 6, 1410-1424. https://doi.org/10.1111/jopy.12817  [Open Access]

#7: Fan, J., Sun, T., Liu, J., Zhao, T., Zhang, B., Chen, Z., Glorioso, M., & Hack, E. (2023). How well can an AI chatbot infer personality? Examining psychometric properties of machine-inferred personality scores. Journal of Applied Psychology, 108, 8, 1277-1299. https://doi.org/10.1073/pnas.1218772110  [Paywall]