How to Tell if a Screening Tool Actually Works: Two Questions From the Man Who Used to Certify Them
Wednesday 26th August
If you're buying a tool to sift high volumes of applications, the hard part isn't the shortlist of vendors. It's working out which of them can prove their tool does what they say it does.
That problem sits at the centre of this episode of the TA Disruptors podcast. Host Robert Newry sits down with Dr Charlie Eyre, Director at Workspheres Ltd, an occupational psychologist who spent years as Senior Editor at the British Psychological Society. In that role, his job was reviewing psychometric assessments against the BPS standards and deciding whether they met them.
Randstad estimated that 1,200 new AI tools for TA teams launched in the last year in the US alone. Almost none of them will ever be reviewed the way Charlie used to review assessments, because the process is voluntary. So this conversation works through what that review actually involves, and how a TA leader with no psychometrics background can run a version of it themselves.
Why is it so hard to tell whether a screening tool works?
Screening tools are sold on the things a demo can show. Sift volume, candidate experience, how quickly it talks to your ATS. All of it observable in 40 minutes.
What a demo can't show you is whether the tool picks the right people. That evidence arrives much later, and usually from the wrong direction. In The State of Volume Hiring 2026, the number one reason hiring managers lose faith in a hiring process is a candidate who interviewed well and then couldn't do the job. 39% named it, ahead of exaggerated CVs and ahead of poor-fit applications.
The market pressure behind this is real. Application volumes have risen sharply, TA teams haven't grown to match, and the obvious response is a tool that cuts thousands of applications down to a number a human can engage with. Charlie's point is that the discipline for evaluating those tools already exists, and it predates the current wave by about 40 years.
The gap is regulatory. Charlie uses driverless taxis as the comparison: legislation, trials, accreditation of each vehicle and each provider, and a safety case established before one carries a passenger. "The difference between the taxi analogy and this is that the regulatory context is far less. It's primarily voluntary for providers and developers of new tools."
What makes high volume hiring in a betting shop so complex?What are the three principles behind any good selection tool?
Charlie starts with three questions that have been in use in the assessment profession for decades. He translates them into plain terms early in the conversation.
Is the tool measuring what it claims to measure? That's validity. Does it measure consistently? That's reliability, sometimes called precision. And does it measure in a way that doesn't disadvantage any particular group in the candidate pool? That's fairness.
Two things make those questions harder to answer in 2026 than they were 10 years ago. The first is that many of the new tools aren't psychometric at all. They're matching engines, working from a job description rather than a measured construct. Charlie is explicit that he comes to this agnostic on methodology, and the standards apply regardless: if a selection decision is being made, the same three questions hold.
The second is that standards exist without a regulator to enforce them. The European Federation of Psychological Associations publishes standards for psychological and educational tests, recently updated and free to read. The BPS Psychological Testing Centre publishes guidance too. There's no badge a vendor legally has to hold before going to market, which Charlie describes as an ethical obligation rather than a legal one.
Does the tool measure what it claims to measure?
This is where most of the conversation lives, and where the thought experiment in the middle of the episode does its work. Robert plays the founder of a new AI matching tool: it reads your job description, extracts the frequently used words, then returns the 10 best-matching CVs from the internet.
Charlie's first response is to go upstream of the tool entirely. He won't assess the matching until he's satisfied the job description reflects the attributes the role actually needs. Robert's summary of the risk: "you've got rubbish in, rubbish out." Get that wrong and everything downstream inherits it, and the risk grows the harder the sift gets.
Then the measurement question. A keyword matcher establishes how closely a CV resembles a job description. Whether the 10 it surfaces would outperform the hundreds it discarded is a separate claim, and one that needs evidence: "what's the evidence that they are ultimately the best people that would meet the job requirements we have here?"
Charlie also notes what's changed underneath this. Candidates are now investing significant time and AI assistance in honing their CVs, and our research puts numbers on it: 42% of candidates have used or would use AI to write their CV. Rank applicants on CV wording and an increasing share of what you're measuring is the quality of the help they had.
His answer to what evidence looks like is specific. Criterion validity, meaning the extent to which scores on a measure relate to a real-world outcome. A predictive validity study, checking scores against how those same people went on to perform in the role. Or convergent validity, running a new measure alongside an established one with a common group and comparing the results.
He's also honest about the limits, which is worth quoting in full: "Validity is about piecing together a jigsaw, really. It's not necessarily that there is one universally applied measure that will tell you whether something is valid. Likewise, in all my time working in this area, you never see a fully valid measure. But it is a case of seeing some evidence."
One vendor response he doesn't accept is that a new tool can't have evidence yet. "The reality is no new tool is going to be born with a complete evidence base behind it," he says, and then explains why that's a starting point rather than an excuse. Research done properly begins in a test environment, and the question stays the same: what's there, and what does it tell a customer about the outcomes to expect?
Does it measure consistently, and why does that matter more now?
Reliability gets asked about less often, and Charlie's analogy is deliberately mundane. A tape measure in the garden should do the same job every time you use it. Any measurement instrument should.
In practice, that means test-retest reliability. Administer a tool to 100 candidates, administer it to the same 100 two weeks later, and you should see a similarity in the outputs. If you don't, something other than the attribute you meant to measure is influencing the result.
This has become a harder standard to meet. Older assessments were static and identical for every candidate, so consistency was simple to demonstrate. Charlie's observation is that as technology evolves, each candidate is more likely to get a slightly different administration, which raises the question of whether different forms of the same tool are still measuring the same attribute in the same way.
For a TA leader, this is the check that stops a score from being an accident of timing. Without it, a ranking can shift for reasons that have nothing to do with the candidates.
Why "ethical AI" is not evidence of fairness
Fairness is the principle most vendors already have language for, and Charlie is direct about the limits of that language. Good intentions and bad outputs coexist comfortably: "Tools can be used with the best of intentions but have an undesirable output from them."
What gives you confidence is analysis. Specifically, evidence that someone has looked at how a tool performs across protected characteristics in a large candidate pool. Sex differentials, differentials by ethnicity, and so on. "It's only through looking at those hard numbers that you get a better sense of what the risk is."
There's a volume-hiring detail here that deserves more attention than it gets. The harder the sift, the higher the stakes. As Charlie puts it, the lower the selection ratio and the larger the proportion of candidates deselected, the greater the risk of seeing bias between groups. Sifting 5,000 down to 50 carries more fairness risk than sifting 500 down to 50, which is exactly the direction application volumes have been moving.
He also makes the point that matters most to anyone using an early-stage sift: you'll never meet the people it removes. Which is why the evidence has to exist before the tool goes live, not after.
What should you ask a vendor before you buy?
Charlie's practical answer is a document. Ask for the technical manual.
"In effect it's like looking under the bonnet of the car and just seeing what have we got here," he explains. What's the stated purpose, what's the theoretical basis the instrument was built on, what trials were run before it went live, and what does the evidence say on validity, reliability and fairness. For a psychometric going through BPS review, a considerable technical manual is a precondition of being considered at all.
The practical value for a TA leader is comparative. Ask every vendor on your shortlist the same question and the answers themselves become evidence. Some will send a document. Some will send a case study. Some will explain why they don't have one. That spread tells you where the risk sits, and it lets you decide how much of it you want to carry.
Two follow-up questions make it sharper. What were the scores validated against, meaning real performance in a real role rather than a convenient proxy. And, before you ask anyone anything, read what good looks like: both the BPS Psychological Testing Centre and the EFPA standards are published and free. Charlie's point about the EFPA document is that it's useful even if you skip the technical detail. It tells you the standards professional developers work to, which tells you what to ask.
Why defensibility is now a regulatory question
The last part of the conversation moves from evidence to consequence, and it's the part that has changed most since Charlie was reviewing assessments at the BPS.
His test is a candidate-facing one. Can you explain to someone who was rejected what the process assessed and why the outcome went the way it did? And if that explanation were challenged, at an employment tribunal or elsewhere, would it hold? He points to US case law where validation evidence was the deciding element in an organisation justifying its use of a measure.
That question now has a regulatory frame in the UK. The ICO published draft guidance on automated decision-making in recruitment on 31 March 2026, underpinned by the Data (Use and Access) Act 2025. Where automated decision-making was previously restricted, it's now permitted provided the right safeguards are in place. And the guidance weights validity alongside explainability, rewarding processes that can show why a decision was made on the basis of evidence rather than a black box.
Which puts Charlie's 40-year-old questions in a new position. The evidence a technical manual contains has moved from good practice towards something closer to a defence file.
Charlie's closing point is about where that expertise sits. Chartered psychologists work under a statutory regulator and to a set of published ethical standards, and their training is grounded in evidence-based practice. His argument, and Robert's, is that more of that discipline belongs inside the recruitment process than currently sits there.
Key takeaways
- Ask what the tool is measuring before you ask how accurate it is. A matching engine establishes similarity between a CV and a job description. Whether that similarity predicts performance in the role is a separate claim that needs its own evidence.
- Ask for the technical manual, then compare what comes back. Theoretical basis, trials run before launch, and the evidence on validity, reliability and fairness. The differences between vendors' answers tell you where your risk is.
- "Too new to have evidence" is a starting point, not an answer. No tool launches with a complete evidence base. But research done properly begins in a test environment, and there should be something to show.
- Check the job description before you blame the tool. If the input doesn't reflect what the role actually needs, nothing downstream will fix it. And the narrower your sift, the more that flaw is amplified.
- The harder the sift, the higher the fairness risk. A lower selection ratio means a larger share of candidates deselected, and more scope for differences between groups to appear. Rising volumes make this more pressing, not less.
- Read the free standards before you talk to anyone. The BPS Psychological Testing Centre and the European Federation of Psychological Associations both publish what good looks like, at no cost.
Listen now 👇
Charlie's full conversation with Robert Newry covers more than we could fit here, including how to decide what to measure when competencies, strengths, personality, motivation and skills all overlap, why work samples are so effective and so expensive, and the difference between self-report and norm-based comparison when you're trying to justify why 10 candidates progressed and 490 didn't.
Listen to the full episode below.
If you're working out how to evidence the tools in your hiring process, explore how Arctic Shores can help.
Transcript:
Robert Newry - Arctic Shores Co-Founder and Chief Explorer
Dr Charlie Eyre - Director, Workspheres Ltd
Robert: Welcome to the TA Disruptors podcast. I'm Robert Newry, co-founder and Chief Explorer at Arctic Shores, the task-based psychometric company that uncovers potential and helps organisations see more in people. In today's episode, we're going to explore how you decide if a screening tool for high-volume sifting is safe to use. The science of screening, as it were.
I'm very excited to pick the brains of a true expert in this area, Dr Charlie Eyre, Director at Workspheres Ltd. Charlie has a long and distinguished career in occupational psychology and in evaluating psychometric assessments. He was Principal Psychologist at the College of Policing for many years before branching out to run his own consultancy, which focuses on the effective use of psychology in the workplace.
During that time he was also Senior Editor at the British Psychological Society, where he oversaw the process and certification of psychometric assessments. The BPS uniquely sets out the standards and requirements for a psychometric assessment to meet, and then offers certification against those standards.
It was in this capacity that I first met Charlie almost 10 years ago, when I wanted to understand what Arctic Shores needed to do to meet those standards. He was so helpful to me then, and has continued to be since, that it was only natural for me to want to ask him whether and how those standards and best practices might apply in the era of AI, and how they might apply to the plethora of new tools that now exist to help with screening.
Welcome to the TA Disruptors podcast, Charlie.
Charlie: Thank you very much, Robert, and thank you for that generous introduction. It's an absolute pleasure to be here to talk to you about these subjects.
Robert: It's brilliant to have you here, and I'm very much looking forward to this discussion on the science of screening. It's so important now to be looking at this area, because all the time we are hearing more and more stories about the high volumes of applications that TA teams are having to deal with.
The volumes just seem to be going through the roof. I'm hearing crazy stories of thousands of applications for one role. There was something on LinkedIn the other day where somebody said they'd had 18,000 applications for one role and wondered if that was a new record.
So it's no surprise that TA teams who haven't been given more resources are now running around trying to find tools that will help them deal with this huge volume, and screen these applications down to a more manageable number that they can then engage with on a more human, face-to-face level.
Interestingly, Randstad, I heard on a webinar recently, did some research and in the last 12 months they found 1,200 new tools that have been launched to help in talent acquisition with this screening problem. 1,200. So it's a big problem, and there is a wave of new tools coming out to help people try and address it.
But screening tools are nothing new for the world that you and I come from, Charlie, and the psychometric sector has been developing standards and best practice about how you do this in a good way. I think now there's a lot that we can learn from that, in how we can help TA leaders think about what is a good tool to use and what is not.
Perhaps we can kick this off with some thoughts from you on what those first principles are around what makes a good selection tool, and why it's important to have standards around this.
Charlie: First of all, in terms of what the basic principles are that we need to be looking at, these have been around for some time, but they do ask some very basic questions of the methodology and the tools that we're looking at.
Are they actually measuring what they claim to measure, which is the concept of validity? There's the point around consistency, which is: does the tool actually measure on a consistent basis? That's often referred to as reliability, or precision. But there's also the point around whether these particular tools measure in a way that's fair and doesn't disadvantage any particular group within the overall candidate pool.
So those three principles of validity, reliability and fairness have been around for a long time. Psychometricians have been applying them for decades in the selection space. It's also thinking about how those might apply within the new and very different context that we're seeing now in terms of emerging technology.
Robert: It makes a lot of sense to me, and I can see how those are three fundamental questions. I really like the way you put that into general TA language: do they do what they claim to do, do they do that consistently, and do they do that fairly?
Are there a set of standards that we can all look to around this, or is this something that in our space doesn't have that regulatory oversight?
Charlie: I suppose there's the point of the standards, and then there's the point of the regulation of those standards.
If we look at an analogy in terms of innovative technologies which are having a significant effect on people's day-to-day lives, we might look at something like the driverless taxis being introduced in London in the coming months. What we see within that is a systematic testing process going on as we speak: new legislation, trials, accreditation of different vehicles, accreditation of providers, and ultimately a fairly rigorous process of establishing whether this is fit for purpose and whether it's going to be safe, before we actually see those taxis on the road.
That's the regulatory component. In psychometrics particularly, but in some ways in the broader assessment space, there are standards which are set out. For example, the European Federation of Psychological Associations has a clear set of standards for psychological and educational tests. They've recently been updated and are freely available to the public. Those provide a good set of guidance principles for what you should expect to see within a psychological test, and they do cover off these three components of fairness, reliability and validity in that context.
The difference between the taxi analogy and this is that the regulatory context is far less. It's primarily voluntary for providers and developers of new tools. There is no unified, legally required badge you have to get to bring something onto the market. So there is in some ways an ethical obligation to look at these issues, but it's not a legal obligation in the same way it might be in other analogous situations.
Robert: Yes. So we can have a separate debate as to whether there should be a legal requirement for it. But there isn't. And therefore that puts the onus on organisations to understand those standards and determine if they are being applied to the tools that they're either using or considering using.
So let's dig into some of those three things you talked about. What does it mean, from a standards point of view, and how can we understand whether a tool does what it claims to do?
And we're not just talking about a psychometric tool here. This could be an AI tool that is doing matching in some form in the recruitment process. There is a selection activity going on, and therefore the standards that are applied to that selection activity, whether it's matching or measuring, are really important.
How, in your experience, do we break down that critical piece of: does it do what it says it does on the tin?
Charlie: Often the first thing we need to go to is what research is available on the particular tool that we're talking about.
In the space of validity, we often look at what's known as criterion validity, which is simply the extent to which scores on this particular measure equate to a real-world outcome. So for example, if there is a particular tool which is claiming to measure a particular dimension of an individual's suitability, then how much do the scores on the tool relate to how the individual is actually performing in the workplace later on down the line? That would be a predictive validity study. It gives some level of reassurance that rather than it being speculative, there is actually research evidence that there is a relationship between the outcomes of the measure and what you would expect to see in the workplace later down the line.
There are other ways in which this can be established from a research basis. For example, convergent validity: if you have a particular new measure that you want to assess, is this measuring what we expect it to? There may be existing tools on the market that measure aligned constructs. In those situations there is an opportunity to administer both side by side with a common group of people, and to see whether the scores relate to one another.
Validity is about piecing together a jigsaw, really. It's not necessarily about there being one universally applied measure that will tell you whether or not something is valid. Likewise, in all my time working in this area, you never see a fully valid measure. But it is a case of seeing some evidence.
Robert: You want to see some evidence.
Charlie: The more objective that is, derived from research methodology which is universally accepted, or widely accepted, then that's going to give you that level of reassurance that it is actually doing what it is supposed to do. Which ultimately is about the quality of the hire. And it leads to all sorts of things like retention and performance in the workplace.
Robert: So those are useful metrics that you'd want to see in the research: that when somebody claims this is going to find and select better candidates than whatever the current process an organisation uses, it's got some evidence to support that claim from the research they've done. Whether that be attrition, whether it be performance.
And I think the interesting thing is that when a new tool comes out, they'll say, well, we couldn't possibly have done that because it's a new tool. I think that's a little bit of hiding behind a story to get an out-of-jail card on this one, rather than: well, if you've done your research properly, you would have spent some time introducing this tool into a testing environment where you can pick up some of that data.
Charlie: The reality is no new tool is going to be born with a complete evidence base behind it. It's a self-evident thing to say, but it's an important recognition that you've got to start somewhere. It's difficult to enter into a competitive marketplace with this.
But I suppose it still comes down to the bottom line of what is there in terms of the evidence base behind the tool. How effective does that stack up in terms of demonstrating all of these areas, whether it be validity, reliability, or fairness? And in simple terms, what does that then say to a potential customer organisation in terms of the outcomes that they can expect from using that tool in a live context?
Robert: Really, really useful. I think validity is a term that we should want to see as table stakes when somebody is talking about taking on a new tool: what's the validity behind it?
But that's only one piece of the puzzle. We then have reliability, or consistency. Why is consistency just as important?
Charlie: All of this is going to be based upon what an organisation has identified as being important in terms of the attributes which are needed for a particular role.
If you are using any kind of measurement, whether it be a tape measure in the garden or anything like that, you expect it to do the same job each time you use it. You don't expect it to vary in terms of how it's applied and the readout that it's giving you.
In terms of any kind of measurement instrument such as this, it's no different, in that you would expect to see a level of consistency. So that if you were to administer it to 100 candidates, and then two weeks later administer it to the same 100 candidates, you would expect to see a similarity in terms of the outputs that it produces. So you can have that reassurance that there aren't external factors that might be influencing the results beyond the particular attributes that you want to be measuring through this tool.
Robert: That is something that makes sense, and I think a lot of people would understand it, but it's interesting that it doesn't always get asked. And like all these things: what does consistency mean?
I assume that's why you go to the European standards, or the British Psychological Society, that will give you a sense of those hundred who did it at time point one and then when they did it at time point two, and what level of consistency you'd want. There are standards, I'm assuming, that enable people to go and look at and say, okay, this is what good looks like.
Charlie: Absolutely. And there are standards within the EFPA criteria which I mentioned earlier on. What I just mentioned there is referred to as test-retest reliability, where you administer on two occasions.
If we're looking back at the history of some of these tools, we used to see very static, fixed measures, so you could be confident that it was exactly the same administration each time. As technology evolves, there is more likelihood that each candidate will get a slightly different administration of that test depending on what happens, which ups the need to be able to say: well, is it measuring this attribute in a consistent way, regardless of when you take the test? So again, in those situations, in terms of the different forms of the test that might be produced, are they still measuring in that consistent way?
Robert: Brilliant. So we've got validity, we've got reliability. And then this is the big worry for everybody: fairness.
One of the things that people are worried about, particularly with all the new AI tools that come out, is: are we just industrialising bias? And the other thing I often hear as an overlay on this is, well, humans have bias, so they're not particularly fair, so do we need to worry about whether the screening tools are fair or not?
I think that kind of misses the point a bit. The fact that humans are biased doesn't necessarily mean we should accept that the tools that humans develop can also be biased, because then everything is flawed, as opposed to: could we actually make it better?
Charlie: Yes. And there are a few factors about this.
There's the simple question of who are the people who you want to bring into your organisation? Because if you are possibly never going to see many of these people, because they are deselected at a very early stage in the recruitment process, then you want to have the confidence that there isn't any unfair differentiation taking place right at the outset, which might disadvantage particular groups.
There is that point of fairness, and there are also legislative requirements in terms of fairness, which we all need to be aware of in terms of hiring practices as well. So it all feeds into that.
The one way in which we can have more confidence about the fairness of a tool is simply through the analysis which has been completed on it, in terms of any administrations and experimental research which has been done to demonstrate the extent to which, if you are looking at a large pool of candidates, analysis has been completed across different protected characteristics. Might there be sex differentials, or might there be differentials in terms of ethnicity, and so on? It's only through looking at those hard numbers that you can get that better sense of what the risk is with this.
One of the factors with this as well, given what you said in the intro about the huge numbers of applicants coming in: the harder the deselection, or the lower the selection ratio, and the larger the proportion of candidates being deselected, the greater the risk that you are going to see some bias in there between groups. So it ramps up the pressure to be able to look at the data and get a better understanding of how a particular tool is working.
Robert: And what's clear from everything that you've said too is that it needs to be evidence-based. Because I hear a lot of things about our tool uses ethical AI, as if that term in its own definition implies fairness.
What you've shared there is that if we're really going to determine whether fairness exists and is built into the way a particular tool operates, then we need to see that there's some evidence around that. Just using the term ethical design doesn't make it fair per se, because we can all have approached it ethically and still come out with an unethical outcome.
Charlie: Potentially, yes. Tools can be used with the best of intentions but have an undesirable output from them.
So whether it's ethical AI or anything, I suppose it's about better understanding what that actually means, and what is involved in that. But ultimately, the numbers are what's needed to be able to provide that objective evaluation in terms of how this is stacking up in the real world when you use it in a live context. What's it showing?
Robert: That's been incredibly helpful, and I think it can be really useful for all the listeners on this.
So now that we understand the three principles, let's go down a thought experiment, if you're happy with doing that.
A lot of the tools that are coming out now to deal with this high volume of applications are matching-based rather than psychometric measuring-based. And I know you're not an AI matching expert, but in some ways that's kind of good for this thought experiment.
I'm going to pretend that I'm the founder of a brand new AI matching tool. And my pitch to you is that I've got a job description that I've loaded up into my platform. It's analysed it, it's picked out some keywords from it, and I'm going to be able to bring to you the top 10 best candidates that match to that job description, rather than you having to wade through hundreds or thousands of CVs.
So if that was my pitch to you, what would you want to understand from me based on those three principles?
Charlie: I think some of the questions I would want to ask are very simple in some ways. It's just opening up what's going on between the starting point, where you have some kind of input, and it might be CV matching, it might be performance on a measure, whatever it is, and the output, which ultimately is a binary outcome of do you pass or are you unsuccessful. Knowing what happens in the middle there is where I would be interested.
Before we get to any of that, it's really truly understanding that the job description is reflecting the right attributes which are actually needed to perform effectively in the role.
Robert: That's quite right actually. Because a lot of the time when we think about the job description, people are bringing it out from a drawer, metaphorically as it were, and then saying, here's the job description, as if the job description of old was still appropriate for the job role of the future.
This is actually a really important document, because if that's the input into something that is going to determine if somebody gets brought forward or not, getting that job description right seems to me an essential part of all this. Because otherwise you've got rubbish in, rubbish out.
Charlie: Absolutely. And I suppose the risk of that, again, is heightened if you do have that very strong narrowing effect. For those 10 candidates that come out, there might be 500 going in at the start, however many it may be.
So I suppose within that, it's understanding: given these attributes that have been identified in the job description, first of all, what's being measured and how is it being measured? And is there evidence of how this actually predicts future performance in the role?
It may be that for this measure there is some proxy evidence that the provider is drawing upon to be able to exemplify that this is going to do what it's supposed to do. But ultimately, is there some kind of research-based evidence to show that this will predict future performance in this area? And that in using this methodology, we will be able to confidently say that the 10 that come out at the end are the best. And that can be, I suppose, both morally and legally defended in that context.
Robert: So let me throw back a bit at your good questions. My response to you on that is that you've given me the job description. So the onus of responsibility is on you for that job description.
Charlie: On me. Absolutely, yes.
Robert: So I take that job description that you've given me. I'm not the expert in your company or that role, so I just have to take it for granted that the job description you've given me is fit for purpose. I then scan that job description and find the words that you've most frequently used, levels of experience that might be required. And then I go out and search the internet, LinkedIn, CV Library, Indeed, wherever, to find people with CVs or profiles that match the words most frequently applied in your job description.
So I don't really need to give any evidence, do I, around that? Because I'm just giving you a word match. Or are there other things that you'd want to understand?
Charlie: I suppose it would be that in using that methodology, you are narrowing down to a particular pool of the people who happen to meet these criteria which have been set. And again, it still comes back to: well, what's the evidence that they are ultimately the best people that would meet the job requirements which we have here?
As you mentioned, I'm not an expert in AI, but I do know there are a lot of people out there who are using a lot of time and effort and AI in terms of honing their CVs. So I guess the point of this is: how does this actually predict future performance in the role? And is there an exclusion of some high performers who might not be meeting these particular criteria in using that approach?
Again, these are difficult questions to answer, because the numbers which are being applied to organisations mean that there are really challenging problems that recruiters are grappling with. The question on this is having the confidence that whatever tool you do choose is performing in a way that actually has that fairness, and the confidence that it will predict future performance as well.
Robert: That's right. And I think at the end of the day, it's got to be some evidence around that.
We talked before we started the podcast about the difference between social science and data science. In the world that we come from, of social science, the thing that we're looking for around all of this is not just a correlation. Does Robert on his CV have the word project management mentioned 10 times? Oh look, project management was mentioned several times in the job description, therefore Robert is a good match. We have a correlation there.
There has to be a bit more thought around this. What about if somebody has been doing a task or job that has project-management-like tasks, but it's not mentioned there? How do we bring in some of those factors? How do we make sure that we've got social mobility around this too, that somebody could be perfectly capable of doing project management tasks, but just not being in a role where they've been able to express that? You could be working in the retail sector in logistics and have amazing project management capabilities, but you just don't use those words to describe the things that you've been doing.
So the way that we look at the world through social science, trying to look at other factors in there, making sure that this is fair to everybody and looking for the evidence on this, not just a correlation, is an important discipline that we need to think about when we're looking at whether we are using the right tools to screen people, and whether we're not creating something that either is ineffective or unfair.
Charlie: Absolutely. And I suppose what you've described there is looking at the aptitude to be able to perform a role, as opposed to having the ticket already.
Now it obviously depends on the level of seniority you're looking at in terms of the hiring process. But ultimately, some of those underpinning skills that you've described in that logistics role: how are they being analysed to get that better understanding of, okay, these are the underpinning ways of thinking, the skills that we would expect to see from an effective hire in this situation? How are they captured in the job description? How might we measure those in a recruitment process? And then it can work out from that, rather than purely the assumption that there is a mention of this in a CV on a number of occasions.
And again, I'm coming to this agnostic on the particular methodologies. But I suppose there is that question as to how, does how one writes a CV naturally lend itself to future performance in the role? And then that's where the evidence is hopefully pointing in the direction for any individual in terms of the effectiveness of a particular tool.
Robert: Great point. And at the end of the day, this has to be not about whether you are in favour of AI or not in favour of AI. It's got to be about the process and the methodology that you've applied to evidence that what you claim to be a great way to identify the most appropriate candidates is indeed a great way.
Charlie: Absolutely.
Robert: Let's go back. You brought up a really good point about skills and what is predictive of success in a role. And we move from the world of matching to measuring a bit. There's a little bit of a grey area between the two, which we can come back to. But let's just focus on this area of measuring.
There seem to be a lot of different terms that are thrown around in the recruitment world around things that we can measure in order to find the right people. We've got competencies, we've got strengths, we've got personality traits, we've got intelligence, motivation, values, and now everybody's talking about skills.
Do we need a tool that measures all of those things, some of those things? What kind of advice and thought do you have on somebody that's trying to think about, I need to measure something, but how do I determine what I should be measuring?
Charlie: It is a challenge, and human performance is a complex thing to categorise and then to measure subsequently.
I suppose if we were to look at all of those areas in a very crude sense, there's a bit of an overlapping Venn diagram, if you like, of all of those different elements. It might be, for example, that someone has identified that error identification is something which is really important for a particular job. There's a particular skill within that, but that in turn links to a certain technical competency of data analysis, that may overlap with an area of behavioural competency around decision making. And in turn, that might be influenced by particular personality facets such as risk tolerance.
So in some ways there is an interlinkage between all of these different areas. That initial analysis of what's important is going to be a critical area with this, really. What's going to be realistic to assess in terms of the volumes that might be coming through in a TA context with high volumes of applicants coming through?
But I suppose within that there always has to be the decision of: if we are screening people out on these particular criteria, does it ultimately leave us with the best people?
Robert: Yes. And it's a very simple point, but it's an incredibly important point with that.
Charlie: Yes.
Robert: And what makes up the best people will be, as you say, a number of different elements to this. There will be: do they have the capability and the aptitude to do this? Then there's the whole separate piece of, are they motivated to do it, and to be that within that organisation? And increasingly they're working in teams and groups, so do they have the right values that align with that organisation?
So would you measure those things separately, or would you try and do them all within one tool? What would be your guidance to somebody thinking, I need to capture all those things, and how would I go about it?
Charlie: If you're wanting to capture all of those things, then there's a huge range of complexity within that.
For example, in my history, in my career, a lot of it has revolved around assessment centres, which are immersive work-sample-type exercises where people will come for possibly a day or two days and are put into those simulated work sample situations. Those are time consuming. They are expensive to run. Clearly there's going to be a finite number of candidates you can put in that situation. But I suppose the positive of it is there are a number of dimensions that you can take from that, in terms of truly putting people into that immersive environment.
There needs to be the clarity in terms of: what are the requirements for the role? What's realistic to measure? What tools do we have available to us to do that? What's the evidence base behind them? And then obviously there is the point of sequencing within a recruitment process, in terms of what might come where, and what tools do you want to expose people to at that high volume stage as well?
Robert: And I really like your point about work samples. I think we'll see more focus on work sample usage, whether that be in an interview or in high volume situations, where often it is an assessment centre. That can be highly predictive, because then you can see, and they can demonstrate at the same time, their true capabilities and motivations and values when they are in the work sample environment.
Prior to that, you just need to know, do they have the capability and the aptitude to do that? Because as you say, it's expensive. It requires a lot of careful observation as well, which you can do in an assessment centre, but you can't do that at scale.
So I often talk about the golden thread around this, and how important it is: if this is the output, how do we go back from that and look at each stage along the lines that you were saying, to say, okay, what tools have we got available? How do they link to each other? How are they building up that picture of a candidate's full set of requirements to be successful in the role? All those elements and stages should be interlinked and thought through in terms of their input and output, to get us to that final decision.
Charlie: Absolutely. And before people get to that expensive stage, whatever that may look like, when you are interacting more closely with applicants, that's where the validity and the reliability evidence becomes so important, to give the recruiter that confidence that the people we're investing the money in, even to assess at that late stage, are the right people, and there is evidence to back up that fact.
Robert: One thing we haven't talked about, that I think is important but also a little bit confusing for people, is that when you have a measurement tool, and to some extent even a matching tool, we are trying to determine individual differences. As you said, if we start with 500, we've got to get down to 10. So we need to be able to justify why those 10 are not the other 490.
There seem, in the market anyway, to be two ways that we can go about this. We can ask people to self-report on what they think their individual differences are, or we can compare them to a general population group, sometimes called a norm group.
In your opinion, what's the importance of understanding the difference between those two, and what are the pros and cons of how they might be applied in a recruiting setting?
Charlie: I suppose if we look back at the psychometric context, there are tools which are referred to as ipsative, where ultimately candidates are given a number of options where they are forced to choose between them. But it's self-reported at that stage. Within that, largely the test is self-contained in that respect, so it will give more of an indication in terms of does a candidate lean towards this particular disposition or this particular attribute, or less so from that one, but it's all comparative internally for that particular candidate.
That can have its place in terms of development. There is an argument that it helps to minimise levels of socially desirable responses, some of the criticisms which are levelled at other types of psychometrics, and that's where people are managing their image and responding in a particular way that they feel that the recruiter might find desirable. So it controls for that, but it's possibly less effective in that broader context of comparison.
And you mentioned norm-based measures. A norm is simply a means of comparing one person, within a particular group who has taken a particular measure, with a larger group of people.
Robert: And that doesn't necessarily mean normal people.
Charlie: It's not about normal people as such. People may have come across the normal distribution, that bell-shaped curve. And really, this is about giving some context to a score. Because I could tell you I've got a raw score of 36 on this measure. In isolation, that means absolutely nothing. It needs a context within which to interpret that. The purpose of norms within psychometrics is simply to provide some context which gives more meaning to that particular score.
Robert: And so a good norm group would be one that was reflective of the general population, and also potentially specific to that kind of area, so early careers, or a region, or something like that.
Charlie: Absolutely. And if we look back at the way in which psychometrics are evaluated, part of it is: if they are using norm groups, how are they constructed? Who's been selected to take part in those? Are they representative of the people with whom this test or particular measure will ultimately be used?
And you can break that down. There might be a global norm for a particular measure, but it might be that actually we need a particular region for this, because it would be unfair, or actually counterproductive, to compare one individual with someone else from a completely different culture. So it depends on what it is we're measuring in respect of this.
As you mentioned, it might be particular job clusters. So it might be that there is perhaps more of a clerical kind of level of norm, or there might be a senior managerial norm. So again, it's putting the context behind what this score means in the broader context of people who have taken this measure previously. That is an incredibly important area to understand and make sure of, because as you said, if you are getting a score and comparing it to the wrong contextual group, then it's not really reflective, or may not be as reflective, of that individual's capabilities within that area.
Robert: So understanding if somebody is using a comparison group, and it's important to do that to see individual differences, that they have put together the right kind of comparison group, and that that comparison group has been carefully curated to make sure that it is not in any way biased or weighted to one group over another.
Charlie: Absolutely. And I suppose within a recruitment context, often there is almost a de facto norm in some ways, because you have a particular number of applicants who might take a particular kind of tool, it scores them in a particular way, and the decision for the recruiter is which of these individuals do we progress on to the next stage? So if you've got 50 applicants and 10 people progressing on to the next stage, in some ways you have an informal norm group there in and of itself, simply because there is comparison between them.
Robert: We're having to make a comparison.
Charlie: And you are taking, okay, this is an important criterion, we've identified it's a critical attribute for the job. These are the people that have been shown through this valid and reliable measure to perform best on this. Therefore we're going to select those people as the people to progress to the next stage.
Robert: I think that's a great point: that in the world of selection, we're always making a comparison. We're always having to provide a difference between people, because how else do we end up, as you say, with the group that we're going to move forward and the group that we're going to decline?
So there has to be a ranking mechanism in there. And as soon as there's a ranking mechanism, there's a comparison and there's a measurement piece. So whatever process somebody uses, what's important is that whatever mechanism they're using to make that comparison is done in a way that's fair, consistent and valid to the principles that you shared at the beginning.
Charlie: Absolutely.
Robert: So we've covered quite a lot of things over this podcast so far, a lot of different information about how you might consider screening tools and think about what the differences are between them, and whether they're safe to use or not.
And so now, if somebody having listened to this goes, okay, some really good and meaty things have been covered here, where do I go to get some external reading, advice, standards on this? What would you recommend for people to go and look at?
Charlie: As a starting point, the British Psychological Society have a number of documents on their website, the Psychological Testing Centre, which cover off a number of these concepts and explain them, and talk about the evidence that you might want to look for in terms of evaluating a particular measure.
If you want more detail of this, then I mentioned earlier on the standards for psychological and educational tests which are published by the European Federation of Psychological Associations. Again, they're freely available out there. They're perhaps more of a technical read, but in looking at those, it's not necessarily the case of having to understand every single word within it, or all of the concepts. It does give a signpost as well: these are the types of standards which professional test developers are working to. And in that context, it may give some useful pointers in terms of what would be the types of question I should be asking of a potential vendor who has a tool, in terms of the process that they've gone through to get this to market, and on what evidence are they judging its validity, its reliability and its fairness.
Robert: Because you can ask the questions, you can get the answers, and then how do I know what good looks like? That's where those standards are incredibly helpful.
What about a technical manual? In the psychometric world we have been used to that idea of, right, here's a technical manual to show the evidence and the research and the testing that we've done, in order to demonstrate to a subject matter expert such as yourself that the tool meets the standards, and you can go and look at that evidence.
What would you expect to be in a technical manual, and do you think they are, and should continue to be, useful documents that you would want to evaluate?
Charlie: The technical manual, in reviewing any kind of standardised assessment process, has always been the go-to place. In effect it's like looking under the bonnet of the car and just seeing what have we got here. What's the supposed purpose of this? But also, what are the steps which have been taken? First of all, the theoretical basis on which this particular instrument has been developed. What were the steps that were taken to trial it before it's used in a live setting? But also, what is the evidence around all of those attributes, the reliability, the validity, and the fairness?
If, for example, a psychometric were to be approved by the British Psychological Society, you would expect a fairly considerable technical manual to be available for it to be considered. To get certified, it is going to have to have the necessary characteristics and evidence base underpinning it, in order to give anyone who uses it, whether it be the recruiter or the candidate who is actually going through it, confidence that it actually does what it claims to do.
Robert: So it's an important document. And it provides a great way for TA leaders when they're thinking about different tools: they can ask, do you have a technical manual? Then you can compare vendors on the degree and level of evidence that they provide behind the tool. And that in itself can help you in making the right decision, and the level of risk that you want to take when you are adopting a new tool or not.
Charlie: Absolutely, and the defensibility of the tool is critical here.
There is the point around: can you look at a candidate who's been successful or unsuccessful and actually explain to them what the process has been through, and why they have been unsuccessful or successful in this situation?
It can also be, and it's not to look at it in too defensive a setting, the evidence that if the process were challenged, it could be defended, whether it be through an employment tribunal or anything like that. The evidence that comes out of validation studies is critical in those areas. And there's various case law, particularly in the US, where validation evidence has been the central element of an organisation justifying why they've used a particular measure. So there are real-world practical reasons for that, as well as the simple point that it helps evidence that you're getting the right people into the organisation.
Robert: Such a good point about the defensibility. And we're living in a world where there's more regulation coming out, whether it's the EU AI Act, whether it's GDPR, European or UK, there's the new Data (Use and Access) Act. So that defensibility is going to be really important.
You alluded to something there, and this would be the final point: the importance not just of the validity behind the research and the evidence that went into creating the tool, but also how you practically implement it.
So what's your advice to TA leaders thinking about implementation? Because I assume you shouldn't just be given the keys to the car without any expert guidance or instruction, and just told, here are the keys to the car, go drive.
Charlie: I suppose within all of this, there is a point about considering the broader context of the recruitment process overall. And the point I made earlier on about the foundations on which this is built: how well do we truly understand what it is we're measuring here, before we even get to the point of going to market to see what's actually out there?
But it would be important to have that point around who within the organisation, or who is brought into the organisation, to support in terms of the implementation and the analysis of how things are going. The analytics are really important with this, in terms of just being able to understand what the data is saying in terms of how the measure is landing in respect of this. So there are various functions which the organisation can serve, and potentially others might be able to assist with as well.
Robert: And I think that's where your profession, as an occupational psychologist, has such an important part to play in this. Because the discipline, the standards, how you do this well, how you make sure that when you're setting this up at the beginning you're giving yourself the best chance for success, that you're ensuring the defensibility is in there. It's much, I think, underrated in many ways.
And I really would like to see the increasing use of business psychologists within the recruitment process, because there is a level of expertise and discipline and training that I think you have gone through that is incredibly important and valuable when you're introducing a matching or a measuring tool of some form. Though naturally you'd probably agree with that.
Charlie: I don't know how reliable and impartial an answer I can give to that. But in seriousness, I suppose where we can potentially add value as a discipline is in a couple of areas.
One is that for a fully qualified psychologist, for example someone who has chartership with the BPS, they will have been through many years of both study and practice to get to where they are. And a lot of that is grounded in the idea of evidence-based practice. Within looking at any tools like this, it's about looking at where the objective evidence lies, and what research is needed to support that. So that's part of the discipline, and anyone in this discipline would hopefully be able to assist with that.
The other factor is that within our discipline, for chartered psychologists, we are overseen by healthcare profession standards. So there is a statutory regulator on us as well. As many professions have, there are ethical standards that we need to comply with. There will be professional bodies that provide these, but there are also clear demarcations in terms of the standards of ethics that we should be working to as well.
Robert: Charlie, as always, it's been fascinating talking to you, and it's been so interesting and helpful to hear about those three principles, and why that science of the world of occupational psychology is as important now as it's always been.
I am sure lots of people who've been listening will have taken a lot of great takeaways and insights and useful tips from you on this. So thank you so much for coming on the podcast.
Charlie: Thank you very much, Robert. It's been a pleasure.
Read Next
Sign up for our newsletter to be notified as soon as our next research piece drops.
Join over 2,000 disruptive TA leaders and get insights into the latest trends turning TA on its head in your inbox, every week
Sign up for our newsletter to be notified as soon as our next research piece drops.
Join over 2,000 disruptive TA leaders and get insights into the latest trends turning TA on its head in your inbox, every week