Using Large Language Models in Selection scenarios: What has been shown to work, why has it worked, and what this suggests for future low and high-stakes application deployments.
By Paul Barrett
What has been published or made available publicly
This kind of information is limited as many organizations may well be experimenting ‘in house’ but, for commercial/financial reasons, will not be publicising their activities/results. So, what’s readily available is from academics undertaking research in the area, the odd news report/analysis, and a LinkedIn subscriber developing a “competencies for various job-roles’ niche LLM application.
The Legal and Contract-Management Domains
From AI-Series Blog #2; this example from Whitehouse, Yiu, Catterson, & Perera (2024) shows just where a LLM can have a huge impact on logistical efficiencies, selection/job-roles, and ROI.
Abstract
This paper presents a groundbreaking comparison between Large Language Models (LLMs) and traditional legal contract reviewers—Junior Lawyers and Legal Process Outsourcers (LPOs). We dissect whether LLMs can outperform humans in accuracy, speed, and cost-efficiency during contract review. Our empirical analysis benchmarks LLMs against a ground truth set by Senior Lawyers, uncovering that advanced models match or exceed human accuracy in determining legal issues. In speed, LLMs complete reviews in mere seconds, eclipsing the hours required by their human counterparts. Cost-wise, LLMs operate at a fraction of the price, offering a staggering 99.97 percent reduction in cost over traditional methods. These results are not just statistics—they signal a seismic shift in legal practice. LLMs stand poised to disrupt the legal industry, enhancing accessibility and efficiency of legal services. Our research asserts that the era of LLM dominance in legal contract review is upon us, challenging the status quo and calling for a reimagined future of legal workflows.
This was reported in New Scientist Business Insights newsletter (21st February, 2024):
“While junior lawyers would pore over a document for 56 minutes on average and the legal outsourcers for 3 hours 21 minutes, GPT-4 could spend as little as 2 minutes to get largely the same results. Other LLMs, including Claude and Google’s PaLM 2 were as quick or even quicker: PaLM 2 took around 45 seconds. And while junior lawyers cost around $75 per document analyzed, LLMs could cost mere cents. It all adds up to a worrying future for those wanting to make a living in the field of law”.
Then there is the October, 2023 report in the Financial Times (FT) using four case-studies using ChatGPT, entitled: Legal tech teams turn to AI to advance business goals: Our case studies highlight how the latest tools are helping speed up legal work. [Open Access] .
At the same time, LexisNexis, the online application for searching legal documents / case-law also adopted a LLM-supported strategy for the more efficient searching and preparation of documentary evidence for lawyers, resulting in efficiencies which again mean fewer junior lawyers required for such ‘support’ work.
Then in the FT again, July 4th, 2024: Generative AI turns spotlight on contract management: The prospect of applying tech breakthroughs to handling data-rich digital documents has prompted a flurry of deals. [Open Access]
This is clearly one workplace area where the use of LLMs is having a significant impact in terms of role requirements and ultimately applicant selection.
The Military Domain
In a recent article, also described in AI-Series Blog #2; , Campion, Campion, Johnson, Carretta, Romay, Dirr, Deregla, & Mouton (2024). where incremental prediction of US Air Force Officer Training Assessment Board applicant scores using AI-created variables resulted in regression-model deviation R-squares of between 0.03 and 0.09 over existing variables. The abstract to the article explains the design and application logic:
The purpose of this research is to demonstrate how using natural language processing (NLP) on narrative application data can improve prediction and reduce racial subgroup differences in scores used for selection decisions compared to mental ability test scores and numeric application data. We posit there is uncaptured and job-related constructs that can be gleaned from applicant text data using NLP. We test our hypotheses in an operational context across four samples (total N = 1,828) to predict selection into Officer Training School in the U.S. Air Force. Boards of three senior officers make selection decisions using a highly structured rating process based on mental ability tests, numeric application information (e.g., number of past jobs, college grades), and narrative application information (e.g., past job duties, achievements, interests, statements of objectives). Results showed that NLP scores of the narrative application generally (a) predict Board scores when combined with test scores and numeric application information at a level of correlation equivalent to the correlation between human raters (.60), (b) add incremental prediction of Board scores beyond mental ability tests and numeric application information, and (c) reduce subgroup differences between racial minorities and non-racial minorities in Board scores compared to mental ability tests and numeric application information. Moreover, NLP scores predict (a) job (training) performance, (b) job (training) performance beyond mental ability tests and numeric application information, and (c) even job (training) performance beyond Board scores. Scoring of narrative application data using NLP shows promise in addressing the validity-adverse impact dilemma in selection.
Campion et al, 2024, p. 334, 2nd column go on to show the “Practical Implications” of this work with an estimated saving of one-third the current cost ($169,334.40 to $310,382.40 annually) of the selection decisions, if using the NLP model to replace just one selection Board member. That’s significant and a clear guide as to where and how AI can make a difference in some organizations i.e. Decision-Support. But, it’s a niche application with very specific contextual features.
Exploratory Academic Research on LLMs and NLP
In an earlier editorial to a special issue on machine learning applications to personnel selection, Campion & Campion (2023) provided a summary of the various studies contributing to the special issue, and their results (table 2, pp. 999-1000):
| Authors | Title | Key Findings |
|---|---|---|
| Hernandez and Nie (2023) | The AI-IP: Minimizing the Guesswork of Personality Scale Item Development Through Artificial Intelligence | ML models can be used to efficiently create personality item pools with reliability and construct validity similar to traditional methods. |
| Landers et al. (2023) | A Simulation of the Impacts of Machine Learning to Combine Psychometric Employee Selection System Predictors on Performance Prediction, Adverse Impact, and Number of Dropped Predictors | In a large-scale set of simulations, ML does not greatly predict beyond traditional methods like regression unless samples are small relative to parameters (i.e., an n-to-k ratio of less than 3 for scales or 14 for items), but there are many nuanced findings where ML may be better such as when item-level models are used. |
| Composite Article 1: Improving measurement and prediction in personnel selection through the application of machine learning (Koenig et al., 2023) | ||
| Study 1 (Koenig et al., 2023) | Algorithmic Construct Generalizability: Scoring Novel Open-Ended Prompts with Deep Learning Trained on Alternative Prompts | ML algorithms can generalize to scoring responses from novel prompts, especially when the assessment is the same, when content is similar, and when training data are seeded. |
| Study 2 (Yankov and Speer) | Comparing Three Machine Learning Algorithms for Scoring Assessment Center Text Data | ML can score constructed responses to assessment center exercises with as much reliability and criterion-related validity as humans or better, and there are some differences by ML methods. |
| Study 3 (Hardy et al.) | Using Artificial Intelligence to Make Better Pre-Hire Assessments | ML can complement existing assessments by scoring open-ended questions as well as humans, but more efficiently, with slight criterion-related validity gains and also only slight adverse impact |
| Study 4 (Liu et al.) | Developing and Validating Automated Scoring for an Audio Constructed Response Simulation | ML can score audio constructed responses to a simulation assessment with as much reliability and criterion-related validity as humans, and incremental validity beyond existing assessments. |
| Study 5 (Sun et al.) | Practical ML Algorithms for Selection Assessment Scoring: A Use Case Report on Multi-Outcome Prediction | ML can be used to predict multiple outcomes simultaneously (e.g., productivity and turnover), but the gains over traditional methods may only be marginal with highly structured data (e.g., multiple-choice). |
| Composite article 2: Reducing subgroup differences in personnel selection through the application of machine learning (Zhang et al., 2023) | ||
| Study 1 (Zhang et al., 2023) | Are Fairness-Aware ML Algorithms Really Fair? Predictive Bias of Using ML in Personnel Selection | Fairness-aware ML algorithms that statistically eliminate subgroup differences must create predictive bias mathematically, which may reduce validity and penalize high-scoring racial minorities. |
| Study 2 (Hickman et al.) | Oversampling Higher-Performing Minorities During Machine Learning Model Training Reduces Adverse Impact Slightly but Also Reduces Model Accuracy | Statistically removing subgroup differences in the training data only slightly reduces adverse impact ratios of the resulting ML model but also slightly reduces model accuracy (convergent validity in this study). |
| Study 3 (Song et al.) | Multi-Objective Optimization forPersonnel Selection: A Guide, Tutorial, and User-Friendly Tool | Presents a tool for achieving optimization (Pareto optimal) for up to three objectives, which has many applications in selection. |
Likewise, the article from Canagasuriam & Lukacik, E. (2024) entitled: ChatGPT, can you take my job interview? Examining artificial intelligence cheating in the asynchronous video interview. The abstract reveals the same ‘none of this is usable’ messaging as so much of this work:
Artificial intelligence (AI) chatbots, such as Chat Generative Pre‐trained Transformer (ChatGPT), may threaten the validity of selection processes. This study provides the first examination of how AI cheating in the asynchronous video interview (AVI) may impact interview performance and applicant reactions. In a preregistered experiment, Prolific respondents (N = 245) completed an AVI after being randomly assigned to a non‐ChatGPT, ChatGPT‐Verbatim (read AI‐generated responses word‐for‐word), or ChatGPT‐Personalized condition (provided their résumé/contextual instructions to ChatGPT and modified the AI‐generated responses). The ChatGPT conditions received considerably higher scores on overall performance and content than the non‐ChatGPT condition. However, response delivery ratings did not differ between conditions and the ChatGPT conditions received lower honesty ratings. Both ChatGPT conditions rated the AVI as lower on procedural justice than the non‐ChatGPT condition.
Then there is the experimental work using Natural Language Processing (NLP). E.g. Fyffe, Lee, & Kaplan (2024). “Transforming” personality scale development: Illustrating the potential of state-of-the-art natural language processing. The abstract again shows there really is nothing more here than recreating ‘same-old’ – but now using LLMs and NLP.
Natural language processing (NLP) techniques are becoming increasingly popular in industrial and organizational psychology. One promising area for NLP-based applications is scale development; yet, while many possibilities exist, so far these applications have been restricted—mainly focusing on automated item generation. The current research expands this potential by illustrating an NLP-based approach to content analysis, which manually categorizes scale items by their measured constructs. In NLP, content analysis is performed as a text classification task whereby a model is trained to automatically assign scale items to the construct that they measure. Here, we present an approach to text classification—using state-of-the-art transformer models—that builds upon past approaches. We begin by introducing transformer models and their advantages over alternative methods. Next, we illustrate how to train a transformer to content analyze Big Five personality items. Then, we compare the models trained to human raters, finding that transformer models outperform human raters and several alternative models. Finally, we present practical considerations, limitations, and future research directions.
Finally, there is the editorial article from Woo, Tay, & Oswald, F. (2024) entitled: Artificial intelligence, machine learning, and big data: Improvements to the science of people at work and applications to practice. Which is just the usual handwaving about the ‘potential’ for AI to impact selection and assessment. The abstract is:
Currently, in the organizational research community, artificial intelligence (AI), machine learning (ML), and big data techniques are being vigorously explored as a set of modern-day approaches contributing to a multidisciplinary science of people at work. This paper discusses more specifically how these sophisticated technologies, methods, and data might together advance the science of people at work through various routes, including improving theory and knowledge, construct measurements, and predicting real-world outcomes. Inspired by the four articles in the current special issue highlighting several of these aspects in essential ways, we also share other possibilities for future organizational research. In addition, we indicate many key practical, ethical, and institutional challenges with research involving AI/ML and big data (i.e., data accessibility, methodological skill gaps, data transparency, privacy, reproducibility, generalizability, and interpretability). Taken together, the opportunities and challenges that lie ahead in the areas of AI and ML promise to reshape organizational research and practice in many exciting and impactful ways.
There are other articles published but none of them showcase anything that can be used ‘in practice’ as of now. It is quite literally ‘academic’ research; exploratory and usually more concerned with exploring a variety of “AI” algorithms rather than attempting to do something which has direct, practical, logistic or financial value.
Job-Role Competencies
This is an interesting development from Richelle Arugay ( https://www.linkedin.com/in/richelle-arugay-ph-d-1a6bb83/) who recently announced (June, 2024):
I’m thrilled to announce that I’ve developed JobDissect, a GPT model designed to streamline and enhance HR processes. This AI-powered tool customizes job competencies with precision, making strategic decision-making easier for HR professionals.
BONUS: JobDissect offers the corresponding competency-based interviewing questions… Check it out.
https://chatgpt.com/g/g-KtwURd4fa-job-dissect
Which states on the ChatGPT site:
“This is a job competency customization tool. Please provide the job title, responsibilities, or job description of the role specific to the organization to assess and tailor the competencies required for success in this position. Let me know if you need competency-based interview questions as well”
As someone deeply passionate about the intersection of HR and technology, I’m excited about the potential of JobDissect to revolutionize how we approach HR tasks and contribute to the evolving landscape of human capital management.
Stay tuned for more updates and innovations in the HR tech space! Follow me on LinkedIn for more insights and updates. Let’s connect and explore the future of HR together!
I was somewhat dubious about the utility of such an ‘innovation’, posting a response:
Apart from the saving in clerical time (having an LLM assemble a list of competencies and questions from what’s already “out there” in digital traces) – which is OK as far as it goes, what improvement beyond the solely clerical is being proposed? i.e. how do we assess someone’s listed ‘competency’ any more accurately than before?
What if you already use a searchable competency system (SHL, DDI, O*NET, Cognadev’s Contextualised Competency Mapping (CCM) system etc. with ‘established’ assessments tied to those competencies)? Why bother with a ChatGPT-mediated clerical list?
Richelle replied:
The GPT model Job Dissect dynamically customizes job competencies based on specific job descriptions unique to an organization, rather than generalizing from digital footprints. This ensures the assessments are highly relevant and specific.
Unlike standardized tools like O*NET, Job Dissect adapts in real-time to market trends and organizational needs, enhancing HR processes such as recruitment, training, and performance management. I acknowledge the value of established competency assessment tools and aim to complement them. With further enhancement, Job Dissect can integrate with preferred systems such as SHL or DDI to improve talent assessment, providing deeper insights and more precise competency profiles.
Whether we agree or not with Richelle, what’s important is that this is an example of what might evolve over time with a customised LLM approach. Those already using established commercial competency models might wish to explore Job Dissect and evaluate its potential to supplant or augment their current approaches.
What inferences might we derive from these examples?
1. Niche role-areas involving the use of the use of factual information, procedural knowledge, and reference material will render some selection assessments and job-roles obsolete (legal). This also applies to software and other ‘tech’ roles.
2. Prior and acquired ‘private’ organizational data of sufficient quantity can be used to construct evidence-bases for new, optimised logistical processes and create significant potential financial gain through the instantiation of those new logistical efficiencies (the US Army example). But the data remain ‘private’ and specific to the organization.
3. Existing organizationally-specific psychological assessment information (assessment scores, whatever) might now become the ‘input-data’ into any organizational-specific model, where prediction modelling is conducted by an LLM ‘fed’ both assessment data and outcomes, Organizations can create their own selection/placement filters, job-role-mappings, and relevant employee profiles based upon LLM analytics. This is potentially a real ‘game-changer’ as organizations no longer have to reply upon a test publisher or commercial application to create job-role profiles or prediction models using their own data on their own workforce. However, it assumes an organization has someone ‘in house’ who is charged with working with an LLM like this. Not necessarily a data scientist, but someone who is charged with data uploads and probing an LLM with the kinds of workforce questions of interest to HR.
It also requires a user to have at least a personal paid subscription to ChatGPT4-Plus (a ‘Professional’ subscription at US$19.99 a month) or be a member of an organization which has an Enterprise subscription, and in which they are a recognized user. Data upload and analytics can’t be achieved using a free account.
I’ve created a video which shows how to upload an Excel assessment data file to Chatgpt, and engage with it with prompts / questions so as to produce predictive models of outcome performance categories (Poor, Acceptable, Excellent), and various graphics highlighting the most important predictive variables. For 2,000 ‘employees’, the data file contains simulated CPP scores and ranked styles text-data as predictors, and performance classifications as the outcomes to be predicted, with sufficient random ‘variation’ around the predictor variables to make things semi-realistic. It’s an eye-opener for those who haven’t seen just how easy this is to do.
But, as this article indicates, this is still ‘work in progress’ .. In Medium magazine, by Yu Dong, 20th July 2024 Evaluating ChatGPT’s Data Analysis Improvements: Interactive Tables and Charts: Is ChatGPT becoming a BI tool? [Paywall]
It begins:
“In May 2024, alongside the exciting release of the GPT-4o, OpenAI announced its improvements to data analysis in ChatGPT, featuring interactive tables and charts, and integration with Google Drive and Microsoft OneDrive.”
For #2 and #3, the key principle is that an organization uses its own data to form its own analytics The data is ‘private’ and not allowed to be used by any open-access LLM. That requires an “Enterprise” paid account with ChapGPT, for which only organizations with a minimum of 150 ‘users’ can apply.
4. It is possible to construct competency or other job-role information from existing “open-access’ information, using a customised LLM. And interestingly, there is a move now to creating smaller, ‘niche-area/specific-task’ LLMs which are faster and cheaper to run than the larger ones. See the recent article in the Wall Street Journal (6th July, 2024): For AI Giants, Smaller Is Sometimes Better: Companies are turning their attention to less powerful models, hoping lower costs and solid performance will win more customers. [Paywall]
5. Other than that. For high volume, low-stakes candidate handling, it’s mainly about enhancing the clerical logistics – in a sense competing with existing ATS systems – and handling large volumes of candidates efficiently with automated rules, communications, and ‘candidate tracking’ logistics. Not very exciting or even “AI” – but likely to be logistically financially rewarding with a corresponding reduction in clerical work.
6. For “high Stakes” managerial-leadership assessment – nothing significant changes as this kind of assessment requires careful evaluative judgments based upon a complex mix of information. And that information will be partly unique to any single candidate. An LLM might be employed for acquiring background or ancillary information that adds context for any job-role, but it can’t do more than provide ‘probe-targeted’ information rather than be used for assessment or even ‘prediction’. The role-specificity for such roles contraindicates the use of any LLM or ML strategy.
Isn’t this is all a bit “underwhelming”?
Yes, because the reality is that LLMs can’t do much more in this area than the examples shown. It can be seriously significant for some niche areas and for large organizations who can draw upon their own existing assessment information and outcomes in order to form optimised selection strategies. But it requires a serious strategy, time, effort, a particular kind of expertise, and a financial commitment that may not convince a CFO or CEO of its ROI.
It’s why more are more articles are appearing that are beginning to question the utility and financial ROI of “AI” to date. A sample of recently published articles is provided below.
In the Wall Street Journal, May 9th, 2024:
Where Is the AI Boom Taking Us? Business Leaders Disagree on Outlook: The technology’s promises and perils were a central focus of the WSJ’s CEO Council Summit. [Paywall]
In the CIO newsletter, 22nd May, 2024:
Where’s the ROI for AI? CIOs struggle to find it: Nearly half of all AI leaders question how to estimate or demonstrate the value of AI-related technologies — and for good reason, based on early implementations at many companies. [Open Access]
In Medium magazine, June 13th, 2024:
The Tech Industry Has Stopped Building Things Customers Want. Consumers and business don’t want new technology, they want the benefits of new technology. [Paywall]
In the Wall Street Journal, June 21st, 2024:
Can AI Startups Outrun Dot-Com Bubble Comparisons? Investors Aren’t So Sure. Venture capitalists at this week’s Collision tech event in Toronto approached the next wave of artificial-intelligence startups with increasing skepticism. [Paywall]
In the Wall Street CIO Journal, June 25th, 2024:
AI Work Assistants Need a Lot of Handholding: Getting full value out of AI workplace assistants is turning out to require a heavy lift from enterprises. ‘It has been more work than anticipated,’ says one CIO. [Open Access]
In Medium magazine, June 27th, 2024:
Why I Believe AI Is the Biggest Lie Ever and We’re Buying It: During a gold rush, sell shovels. [Paywall]
In the Economist magazine, July 2nd, 2024:
What happened to the artificial-intelligence revolution? So far the technology has had almost no economic impact. [Paywall]
In the New York Intelligencer newspaper, July 10th, 2024:
AI Investors Are Starting to Wonder: Is This Just a Bubble? [Open Access]
In the FT, July 12th, 2024:
AI bubble set to inflate further. It will take time for the technology to be put to productive use by customers [Paywall]
In the Atlantic magazine, July 12th, 2024:
AI Has Become a Technology of Faith: Sam Altman and Arianna Huffington told me that they believe generative AI can help millions of suffering people. [Paywall]
It begins with a truly meaningful paragraph:
“An important thing to realize about the grandest conversations surrounding AI is that, most of the time, everyone is making things up. This isn’t to say that people have no idea what they’re talking about or that leaders are lying. But the bulk of the conversation about AI’s greatest capabilities is premised on a vision of a theoretical future. It is a sales pitch, one in which the problems of today are brushed aside or softened as issues of now, which surely, leaders in the field insist, will be solved as the technology gets better. What we see today is merely a shadow of what is coming. We just have to trust them.”
While these articles are all concerned with ‘commercial’ implementations and ROI, the same question of “where is the ‘actual benefit’ rather than ‘promised’ benefit” might be targeted at the psychological assessment domain.
We also need to be aware that LLMs really do make things up, as this July 17th, 2024, article in Scientific American explains: ChatGPT Isn’t ‘Hallucinating’—It’s Bullshitting! [Open Access]
The above assessment-oriented examples perhaps indicate where LLMs might have a substantive impact, but these are found in niche areas. There is no ‘it will disrupt the entire psychological assessment market message’ – for reasons in AI-Series Blog #3 , let alone the information above.
But, I can see where organizations might be able to benefit in various ways from analyzing their own assessment and workforce data using a LLM; likewise when considering automating some of the clerical tasks associated with dealing with a high-volume of applicants if an organization is not already using an ATS (Applicant Tracking System).
Perhaps consideration of the Gartner Hype cycle is relevant here?
“The Gartner hype cycle is a graphical presentation developed, used and branded by the American research, advisory and information technology firm Gartner to represent the maturity, adoption, and social application of specific technologies. The hype cycle claims to provide a graphical and conceptual presentation of the maturity of emerging technologies through five phases” (Wikipedia)

Although I am tempted to modify this as:

Which reflects what may be already happening; gradually fading interest in AI/LLMs as they haven’t had the huge impact many were led to believe except in niche areas. Don’t get me wrong, I find ChatGPT absolutely invaluable for quickly finding algorithm code-solutions for some problems I work on, as well as a quick ‘summariser’ of factual information. And, as my video with this blog shows, it can provide some substantive and valuable insights on your own employee assessment data and job-performance prediction. That’s pretty impressive but is it just a provider of ‘powerpoint’ graphics for a presentation, or something that can revolutionise your HR selection and workforce-planning practices?
In Conclusion
There are ‘opportunities’ with current LLMs, NLP, and ML models which have been shown to be game-changing’ in niche areas where clerical, logistical, and ‘informatics-content’ will show a significant ROI. Otherwise, there isn’t much an LLM can do for you other than provide information support in high-stakes assessment. The Fear of Missing Out (FOMO) induced by the hype from “LinkedIn” gurus is just a that, an irrational fear promoted by those looking for their next consultancy contract.
To make intelligent use of LLMs, think about the examples above, the inferences I’ve drawn, the analytics video, and whether you can see a way of optimising something in your work-area that would be worth doing – something that would convince a CFO or CEO of its potential to ‘make a difference’ in their organization.
And, you are not alone in trying to answer the question: “What is AI?”. In MIT Technology Review, July 10th, 2024, Will Douglas Heaven answered this question with a deep historical overview of where and when the term was initially proposed and how the concept has evolved over time to the present day: “Everyone thinks they know but no one can agree. And that’s a problem.” [Open Access]
His piece ends with:
“AI is many things. But I don’t think it’s humanlike. I don’t think it’s the solution to all (or even most) of our problems. It isn’t ChatGPT or Gemini or Copilot. It isn’t neural networks. It’s an idea, a vision, a kind of wish fulfilment. And ideas get shaped by other ideas, by morals, by quasi-religious convictions, by worldviews, by politics, and by gut instinct. “Artificial intelligence” is a helpful shorthand to describe a raft of different technologies. But AI is not one thing; it never has been, no matter how often the branding gets seared into the outside of the box.”
References
Campion, M.A., & Campion, E.D. (2023). Machine learning applications to personnel selection: Current illustrations, lessons learned, and future research. Personnel Psychology, 76, 4, 993-1009. https://doi.org/10.1111/peps.12621 [Open Access] .
Campion, E., Campion, M., Johnson, J., Carretta, T., Romay, S., Dirr, B., Deregla, A., & Mouton, A. (2024). Using natural language processing to increase prediction and reduce subgroup differences in personnel selection decisions. Journal of Applied Psychology, 109, 3, 307-338. https://doi.org/10.1037/apl0001144 . [Paywall] .
Canagasuriam, D., & Lukacik, E. (2024). ChatGPT, can you take my job interview? Examining artificial intelligence cheating in the asynchronous video interview. International Journal of Selection and Assessment, Early View, 1-16. https://doi.org/10.1111/ijsa.12491 [Open Access] .
Fyffe, S., Lee, P., & Kaplan, S. (2024). “Transforming” personality scale development: Illustrating the potential of state-of-the-art natural language processing. Organizational Research Methods, 27, 2, 265-300. https://doi.org/10.1177/10944281231155771 . [Paywall]
Hernandez, I., & Nie,W. (2023). The AI-IP: Minimizing the guesswork of personality scale item development through artificial intelligence. Personnel Psychology, 76, 4, 1011-1035. https://doi.org/10.1111/peps.12543 . [Open Access]
Koenig, N., Tonidandel, S., Thompson, I., Albritton, B., Koohifar, F., Yankov, G., Speer, A., Hardy III, J.H., Gibson, C., Frost, C., Liu, M., McNeney, D., Capman, J., Lowery, S., Kitching, M., & … Newton, C. (2023). Improving measurement and prediction in personnel selection through the application of machine learning. Personnel Psychology, 76, 4, 1061-1123. https://doi.org/10.1111/peps.12608. [Open Access]
Landers, R.L., Auer, E.M., Dunk, L., Langer, M., & Tran, K.N. (2023). A simulation of the impacts of machine learning to combine psychometric employee selection system predictors on performance prediction, adverse impact, and number of dropped predictors. Personnel Psychology, 76, 4, 1037-1060. https://doi.org/10.1111/peps.12587. [Open Access]
Martin, L., Whitehouse, N., Yiu, S., Catterson, L., & Perera, R. (2024). Better call GPT, comparing large language models against lawyers. arXiv Preprint , arXiv:2401.16212v1 [cs.CY] 24 Jan 2024, 1-16. https://arxiv.org/html/2401.16212v1. [Open Access]
Stevenor, B.A., Hickman, L., Zickar, M.J., Wimbush, F., & Beck, W. (2024). Validity evidence for personality scores from algorithms trained on low-stakes verbal data and applied to high-stakes interviews. International Journal of Selection and Assessment, Early View, 1-17. https://doi.org/10.1111/ijsa.12480 . [Paywall]
Woo, S.E., Tay, L., & Oswald, F. (2024). Artificial intelligence, machine learning, and big data: Improvements to the science of people at work and applications to practice. Personnel Psychology, Online First, , 1-16. https://doi.org/10.1111/peps.12643 . [Paywall]
Zhang, N., Wang, M., Xu, H., Koenig, N., Hickman, L., Kuruzovich, J., Ng, V., Arhin, K., Wilson, D., Song, Q.C., Tang, C., Alexander III, L., & Kim, Y. (2023). Reducing subgroup differences in personnel selection through the application of machine learning. Personnel Psychology, 76, 4, 1125-1159. https://doi.org/10.1111/peps.12593. [Open Access]