Faking Personality Assessment
Contents
- Personality Assessment Relies on Honest Responses
- How Generative AI Is Transforming Faking in Personality Assessment
- What We Know So Far: Generative AI and Its Ability to Fake Personality Assessments
- Responding to AI-Assisted Faking
- Prevention of Faking Behavior
- Rethinking the Role of AI in Assessment
- Rethinking the Role of Recommendations
- Recommendations
- PDF download
- Author
- References
Personality Assessment Relies on Honest Responses
Personality assessments have become a widely used tool in modern hiring and selection processes.
Employers increasingly recognize that success in many roles depends not only on a candidate’s technical skills but also on how they typically behave – whether they are dependable, adaptable, cooperative, or resilient, for example. To make informed hiring decisions, organizations try to match candidates to job profiles based not only on what they can do, but also on who they are.
This approach assumes a predictive relationship between test scores and future workplace behavior: If we know how someone tends to behave across time and in different situations, we can make educated guesses about how they will perform in a given job context (e.g., Tett et al., 1991). But this predictive power hinges on a critical condition – that the test scores accurately reflect the candidate’s true personality traits.
Since the early days of personality testing, however, candidates have found ways to intentionally distort their responses. This behavior, known as faking, has been a long-standing concern in personnel psychology. Faking is broadly defined as a deliberate, goal-directed behavior where individuals consciously shape their responses to create a favorable impression – often based on what they believe the employer wants to see (Ziegler et al., 2011). In this sense, faking is not a personality trait in itself, but a motivated behavior.
The consequence of faking is straightforward but serious: If a person’s answers are not an honest reflection of their traits, then any predictions made based on those scores – such as their likelihood to be reliable, collaborative, or stress-resilient – become less valid. Research confirms that faked scores undermine the validity of personality assessments, making it harder for organizations to identify the right people for the job (MacCann et al., 2011; Tett & Simonet, 2021; Ziegler et al., 2011).
Importantly, not everyone fakes, and not everyone fakes to the same extent. Whether someone engages in this behavior depends on the interaction of personal and situational factors (Heggestad, 2011). For instance, high-stakes situations such as competitive job applications create a strong incentive to fake, especially if the candidate understands the role’s requirements and believes they can game the system. Yet not everyone will act on that motivation. A person who is anxious about being dishonest – or fears getting caught – might still respond truthfully, even in a high-pressure context (MacCann et al., 2011).
Distinguishing between faked and honest responses is not always straightforward. An unusually high or low score might raise suspicion, but it could just as well reflect an authentic response from an unusual person. There are methods that allow us to evaluate whether a person has faked: Validity items can provide information about the test taker’s attention level during the assessment (e.g., Meade & Craig, 2012). Impression management scales can offer insight into how much a person tends to alter their answers to reflect more socially desirable traits and behaviors (Crowne & Marlowe, 1960). Bogus items can be used in the overclaiming technique to gauge how much people inflate the presentation of their knowledge in certain areas (Paulhus et al., 2003). As a test publisher, Hogrefe has long relied on these and other psychometric safeguards to support the responsible use of personality assessments. All of these methods work to some degree, which will also be discussed further in a later part of this white paper. However, they never allow for complete certainty if a person has faked their re - sponses, only whether responses should be treated with caution. This inherent ambiguity makes it difficult to confidently identify faking, even when there is reason for concern (MacCann et al., 2011).
The emergence of generative artificial intelligence (AI) tools further complicates this picture. For the first time, individuals have access to powerful systems that can generate highly plausible, tailored responses to personality questionnaire items (Phillips & Robie, 2024a, 2024b; Robie et al., 2025), making it even harder to tell whether a test result reflects the person or the prompt they gave to an AI. Importantly, traditional methods that have proven useful against human faking may not work as expected when applied to responses that were generated or assisted by AI (Robie et al., 2025). This underscores the need for specific guidelines and new methods to detect and prevent faking in the age of generative AI.
How Generative AI Is Transforming Faking in Personality Assessment
Faking is a welldocumented phenomenon in the context of personality assessment.
Empirical data indicate that approximately 30–50% of applicants have, at some point, misrepresented themselves during a hiring or selection process (Griffith & Converse, 2011). More specifically, Griffith and Converse (2011) found that 32% of respondents admitted to exaggerating positive traits on questionnaires, and 15% acknowledged providing outright false responses. With the increasing accessibility of generative AI, this behavior appears to be evolving. Recent surveys suggest that between 29% and 38% of respondents report being at least somewhat likely to use AI to assist them in completing recruitment assessments (Westfall, 2024; Thresher, 2024).
Until recently, individuals who sought to fake their responses faced a number of cognitive and strategic challenges. Effective faking required an understanding of the target role, insight into what the employer was seeking, and the ability to map those expectations onto specific questionnaire items. The success of faking was, in part, limited by an individual’s ability to strategically manage impressions – something that not all candidates were equally capable of doing (e.g., McFarland & Ryan, 2000). In some cases, this behavior may even have been interpreted as a form of test-taking skill or situational intelligence.
The introduction of generative AI technologies has changed this dynamic. Tools such as large language models (LLMs) – for example, ChatGPT – are now capable of producing coherent and contextually appropriate responses to simple user prompts (Budhwar et al., 2023). By lowering the effort required to fake effectively, these tools may allow a much broader range of applicants to generate responses that align closely with role expectations. Initial evidence suggests that generative AI can, in some cases, outperform human fakers (Phillips & Robie, 2024a, 2024b), and that traditional detection methods may not be well equipped to identify such responses (Robie et al., 2025). In the next section, we review recent studies that directly examine the capacity of LLMs to fake personality assessments.
What We Know So Far: Generative AI and Its Ability to Fake Personality Assessments
AI Can Outperform Humans – With Some Limitations
Although LLMs such as ChatGPT are already widely used across a range of tasks, their role in personality assessment – particularly in the context of faking – has only recently begun to attract scientific attention. Given the easy accessibility of these tools, understanding how they might be employed by candidates during high-stakes situations is essential. It is now entirely feasible for individuals to copy and paste job descriptions and assessment items into a publicly available AI model to generate tailored responses that align with employer expectations (Phillips & Robie, 2024a).
Two recent studies have systematically investigated whether LLMs are capable of faking personality assessments equally well as, or even better than, humans (Phillips & Robie, 2024a, 2024b). The results suggest that, at least under certain conditions, the answer is yes, LLMs can be more capable of faking ideal personality profiles than humans are. Another key finding from both studies is that not all test formats are equally vulnerable to faking, and that, for example, forced-choice formats show greater resistance (for more details, see A. LLMs outperform humans in faking ideal personality profiles).
A) LLMs outperform humans in faking ideal personality profiles
Phillips & Robie (2024a, 2024b) compared AI-generated responses to those produced by human participants instructed to fake. The LLMs were given job descriptions and were prompted to generate ideal answers to personality items designed to measure traits relevant to those jobs. LLMs consistently produced personality profiles that were more socially desirable and better aligned with job requirements than those of human participants. GPT-based models, in particular, demonstrated a strong ability to inflate personality scores. In several cases, AI-generated responses outperformed student participants on assessments designed to measure traits such as conscientiousness – a trait frequently associated with job success (e.g., Sackett et al., 2022; Wilmot & Ones, 2019).
However, the LLMs were not uniformly successful. They tended to struggle with more abstract or less frequently discussed constructs, such as integrity, likely due to the nature and limitations of their training data (Phillips & Robie, 2024b). In addition, when asked to respond to each item after resetting the model – that is, relying solely on their pretraining knowledge – models generated responses that were less precisely tailored, indicating that access to contextual data influences model performance (Phillips & Robie, 2024a).
One key finding from these studies is that not all test formats are equally vulnerable. Assessments using forced-choice formats – where individuals must compare and choose between equally desirable statements – were generally more difficult for both humans and AI to fake effectively. This effect was particularly evident in phrase-based forced-choice measures, which appeared to disrupt the straightforward generation of “ideal” answers. In phrase-based forced-choice assessments, participants are presented with a small set of statements at the same time and must indicate which statement is “most like me” and which is “least like me.” For example, to measure conscientiousness, an assessment might include the statements “I want to get better marks than my fellow students” (positively keyed) and “I put just enough effort into my studies to get the marks that I want” (negatively keyed; Phillips & Robie, 2024b). Based on the pattern of “most” and “least” choices across many such sets, personality trait scores can be derived. These findings point to promising design strategies for increasing resistance to faking, although further testing is needed to validate these approaches across different platforms and populations.
Implications for Practitioners and Test Publishers
These initial studies by Phillips and Robie (2024a, 2024b) provide a valuable starting point, but several open questions remain. The current research sheds little light on how to detect AI-generated responses once they are submitted. It is difficult to draw conclusions from patterns in aggregated data to make decisions about individual profiles. More importantly, the models themselves are evolving rapidly. Future versions may be better able to simulate traits like integrity, decode more complex question formats, and tailor answers with even greater nuance.
This highlights the need for ongoing research and vigilance. At Hogrefe, we are actively working to investigate and monitor AI-assisted faking behaviors, and to explore mechanisms for identifying such patterns in assessment data. This includes both psychometric and technological solutions.
In the meantime, what can organizations and practitioners do now to minimize the risk of AI-assisted faking? To answer this question, we review traditional strategies for detection and prevention as well as the avenues opened by the introduction of generative AI, in the next section.
Responding to AI-Assisted Faking
The Challenge of Detecting Faking Behavior
Detecting faking behavior in personality assessments has long presented a methodological and ethical challenge. For a detection method to be practically applicable, it must demonstrate both a high detection rate and a low false-positive rate (Kuncel et al., 2011; Stark et al., 2011; Zickar & Sliter, 2011). Incorrectly accusing a candidate of faking can have serious consequences – both internally, in terms of flawed selection decisions, and externally, with potential legal or reputational implications (MacCann et al., 2011).
Over the years, a range of psychometric techniques has been developed to address this problem, such as validity items, impression management scales, the overclaiming technique, and analyses of response patterns (MacCann et al., 2011; Paulhus et al., 2003). While each of these approaches provides useful information, none allows for complete certainty about whether the test taker has faked their responses. At best, they indicate that the results should be treated with caution. This general ambiguity means that detection has always been limited, even under traditional conditions (see B. Examples of traditional psychometric techniques for faking-detection for more details).
The critical question today is whether these traditional tools can detect AI-assisted faking with the same effectiveness, or whether generative AI makes existing safeguards less reliable. In the following, we will review recent evidence that directly examines this issue.
B) Examples of traditional psychometric techniques for faking-detection
One widely used approach is the use of social desirability scales or impression management scales, which are intended to measure the extent to which individuals tailor their responses to align with perceived expectations (e.g., MacCann et al., 2011). While this concept appears promising at first glance, several issues quickly emerge (MacCann et al., 2011). Notably, these scales have been shown to correlate with stable personality traits, particularly Emotional Stability, Conscientiousness, and Agreeableness (Li & Bagger, 2006; Ones et al., 1996). This raises concerns about construct overlap rather than genuine detection of faking tendencies. Moreover, efforts to adjust or “correct” personality scores using social desirability measures have been shown to be ineffective or even counterproductive. Empirical studies report that such corrections often result in no meaningful improvement or even a reduction in the predictive validity of personality scores for job performance and other criteria (Reeder & Ryan, 2011; Dilchert & Ones, 2011; Stark et al., 2011).
Furthermore, social desirability items may not be interpreted consistently by test takers. For example, agreement with statements such as “I have never stolen anything” may reflect a respondent’s genuine self-perception or an endorsement of social norms, rather than a deliberate attempt to deceive (Kuncel et al., 2011). This interpretation is supported by consistent correlations between social desirability scores and traits such as Conscientiousness and Agreeableness.
Another technique that has received attention is the overclaiming method (Paulhus et al., 2003). This approach measures self-enhancement by asking respondents to indicate familiarity with both real and nonexistent items (e.g., fictional people, places, or things). Overclaiming bias is derived from the tendency to falsely recognize illegitimate items, while accuracy scores reflect actual knowledge. Though promising, this method also has notable limitations for applied contexts. Overclaiming bias scores show moderate correlations with other self-enhancement measures, but they also display positive associations with cognitive ability – raising the concern that overclaiming scales might inadvertently penalize more cognitively able candidates, potentially leading to adverse impact or lower-quality hiring decisions.
These examples illustrate broader shortcomings in traditional fakingdetection methods. Other approaches – including the Bayesian truth serum (Prelec, 2004) and linguistic analysis (e.g., Ventura, 2011) – have likewise failed to produce satisfactory results (MacCann et al., 2011). These methods tend to suffer from either low sensitivity (missing actual cases of faking) or high false-positive rates, meaning they incorrectly flag honest respondents. Any method that penalizes a meaningful proportion of truthful individuals is, by definition, unfit for practical implementation.
Recent Evidence: Can Traditional Methods Detect Faking by AI?
A recent study by Robie et al. (2025) directly examined the effectiveness of traditional faking-detection techniques when applied to AI-assisted faking. The researchers conducted a series of experiments comparing responses generated by ChatGPT and human participants, both instructed to present themselves as favorably as possible for a specific job, based on a provided job description.
The study assessed whether ChatGPT could outperform humans in faking personality assessments, while avoiding detection by two commonly used faking-detection methods: impression management scales and the overclaiming technique (see B. Examples of traditional psychometric techniques for faking-detection). Additionally, they tested the effect of “coaching,” by providing participants (both human and AI) with information about the presence of faking-detection tools (see C. How effectively can LLMs avoid detection when faking? for study details).
C) How effectively can LLMs avoid detection when faking?
The research of Robie et al. (2025) revealed that ChatGPT demonstrated some ability to manipulate job-relevant trait scores, but it did not consistently outperform human respondents. This stands in contrast to the two earlier studies by Phillips and Robie (2024a, 2024b), as described in Box 1, where LLMs generally achieved higher desired trait scores than humans. Taken together, the results indicate that LLMs are not universally more effective fakers than humans: Their performance appears to vary by trait and testing context. At the same time, the continuing development and improvement of LLMs suggests that this might change in the future.
With regard to detection methods, the results were also mixed. Coaching had only a minimal effect on impression management scores, but it did lead to a notable reduction in overclaiming bias for ChatGPT. In fact, the reduction was large enough that, post-coaching, there was no longer a meaningful difference between the AI and human faking groups in overclaiming scores (faking-detection scores). This suggests that current detection tools are not uniformly reliable when applied to AI-generated responses, and their effectiveness may be further undermined as LLMs evolve.
The results of the study by Robie et al. (2025) offer two key takeaways: First, LLMs are not consistently superior at faking, but they are already capable of producing responses that rival or exceed those of human fakers in some contexts. Second, traditional detection methods show important limitations, underscoring the need for refinement and revalidation as AI tools continue to develop.
Using Digital Traces and Metadata to Detect AI-Assisted Faking
Paradoxically, while generative AI presents a new avenue for faking, it may also offer new opportunities for faking-detection. To use an LLM to assist in faking, a candidate must typically provide the model with a job profile or summary, as well as the actual test items or closely paraphrased prompts. This interaction can leave detectable behavioral traces.
For example, system metadata such as response times, tab-switching events, and browser activity patterns can be used to identify anomalous test-taking behavior. Candidates who copy test items into an AI tool may display atypical pausing, refocusing, or tab navigation behaviors that suggest external tool usage. These so-called page-focus events can be recorded in online assessment tools, such as the Hogrefe Test System (HTS), which provides a report on test-taking behavior after a test session has been completed. This report also includes response times, which could be used to evaluate the likelihood of AI-assisted faking. However, the mentioned indicators are not definitive on their own, and it can never be guaranteed that a person has actually faked their responses. However, they could inform the development of risk indices to help flag potentially AI-assisted responses.
The usefulness of metadata to detect AI-assisted faking is complicated even more by the emergence of AI agents (e.g., Open AI, 2025). This new feature of some AIs lets users give an LLM a task which it then carries out on its own – for example, to book a table at a restaurant via online booking services. This kind of AI behavior still has to be investigated and might produce patterns of metadata that are different from candidates interacting with an AI during an assessment.
However, it is important to emphasize that no faking-detection system can guarantee perfect accuracy. False positives remain a concern, and the current evidence base is too limited to justify widespread adoption of such tools. More research is needed to explore how and under what conditions AI-assisted faking can be identified. At Hogrefe, we are actively contributing to this research and are committed to developing evidence-based solutions. Until more robust and validated detection methods are available, we recommend that practitioners exercise caution when interpreting personality test results that appear inconsistent, implausibly extreme, or unusually well-aligned with job requirements. In such cases, supplementary assessment methods – such as retesting, structured interviews, or reference checks – may help clarify remaining uncertainties. However, a sole focus on detection is unlikely to provide a complete solution. It is equally important to consider preventive strategies that reduce the likelihood of faking before it occurs. Prevention holds a distinct advantage: It eliminates the need for difficult, and often legally sensitive, decisions regarding how to treat suspected cases of faking (MacCann et al., 2011). In contrast to detection-based approaches, preventive measures mitigate risk without requiring post hoc judgment about a candidate’s honesty.
In the following section, we outline a range of prevention strategies that can be integrated into assessment design and administration to help safeguard the integrity of personality assessments – particularly in the age of AI.
Prevention of Faking Behavior
Selecting Assessments That Are More Resistant to Faking
One strategy to mitigate both traditional and AI-based faking is to use assessment formats that are inherently more difficult to manipulate. For example, asking candidates to provide verifiable information may increase accountability and discourage exaggeration. A question such as “How many hours of overtime did you work last month?” is more likely to elicit a truthful answer than a general statement like “I take on extra work,” because it is more concrete and potentially subject to verification (MacCann et al., 2011).
Situational judgment tests (SJTs) offer another approach. These assessments present test takers with realistic workplace scenarios and ask them to evaluate or select from possible responses. Evidence suggests that SJTs are less vulnerable to faking than traditional Likert-type rating scales, potentially because evaluating specific behavioral responses in context demands more cognitive effort than endorsing generalized statements (Nguyen et al., 2005), and the “correct” answer is harder to determine. One example is Hogrefe’s Leadership Judgment Indicator-2, which measures leadership behavior through hypothetical situations in which candidates must choose how they would act, from four possible options. By grounding responses in context, such instruments make it more difficult to simply select the “ideal” answer. However, even with SJTs, faking remains possible for humans (Peeters & Lievens, 2005), and it is not yet known how effectively AI models can simulate desirable responses in SJT formats.
Forced-choice formats have also been shown to be less fakable, particularly in recent studies discussed above (see the section AI Can Outperform Humans – With Some Limitations; Phillips & Robie, 2024a, 2024b). In this format, test takers are required to choose between equally socially desirable statements, limiting their ability to endorse all positive traits at once. While effective in curbing faking, forced-choice formats result in ipsative or partially ipsative scores, which complicates comparisons across candidates – an important limitation for use in selection contexts.
Reducing Motivation to Fake, Through Warnings
Another strategy to prevent faking involves the use of warnings designed to reduce the test takers’ motivation to engage in deception. These warnings can communicate that dishonest responses may be detected, that misrepresentation could lead to disqualification, or that faking may result in a poor person–job fit (for an example see D. Reducing faking through warnings – An example). Warnings may also appeal to moral norms by emphasizing the value of honesty in the selection process (MacCann et al., 2011).
A recent meta-analysis by Moon et al. (2025) examined the effectiveness of various types of faking warnings. Overall, warnings were found to have a moderate effect in reducing applicant faking. The most effective warnings were those that targeted the ability, motivation, and opportunity to fake. However, the authors also noted substantial variability in effectiveness across studies and contexts, suggesting that no single warning strategy will be universally effective.
D) Reducing faking through warnings – An example
One example of a warning that targets the above-mentioned domains (motivation to engage in deception, moral norms, value of honesty) would be the following: “Please answer all questions honestly. Personality assessments are designed to identify how well your personal traits fit the requirements of the role. Trying to present yourself in a way you think the employer wants, including by using external tools such as AI chatbots, may reduce your chances of success. The assessment includes measures that can detect unusual or inconsistent response patterns, and dishonest responses can result in an inaccurate profile. Even if not detected directly, providing responses generated by AI or otherwise faked responses can lead to poor fit with the role, which may result in lower job satisfaction and performance. The best strategy is to respond truthfully and consistently, as this gives both you and the employer the most accurate picture of whether the role is a good match.”
Additionally, the use of warnings raises important ethical and legal considerations. If test takers are warned that deception will be detected and punished, but no valid detection method is in place, a warning may be considered misleading. This could create legal exposure for organizations and reduce trust in the assessment process (MacCann et al., 2011). In addition, warnings may have unintended consequences for honest test takers, particularly those who are anxious. Some individuals may respond more conservatively or alter their answers unnecessarily out of fear of being falsely flagged as deceptive.
For these reasons, warnings must be used thoughtfully. Their wording should reflect the actual capacities of the testing system, and they should be designed to reduce faking without introducing undue stress or confusion for genuine candidates. In the case of generative AI, test takers may benefit from clear communication explaining that the use of AI tools to influence test results could misrepresent their suitability for the role – and may ultimately work against their best interests.
Technological Approaches to Preventing AI-Assisted Faking
As generative AI becomes more accessible, technological safeguards are gaining importance as a means of preventing its misuse during assessment. Although empirical data on their effectiveness is still limited, several practical interventions have emerged as promising options.
Proctored testing environments, either in person or online, offer one of the most reliable means of deterring AI use. By restricting access to electronic devices and monitoring candidate behavior, these settings can significantly limit opportunities for external assistance. When in-person testing is not feasible, remote assessments can be administered through digital platforms that incorporate browser restrictions – such as blocking copy–paste functionality, disabling screenshots, or preventing candidates from leaving the testing tab. Using assessments with time limits reduces the opportunity for individuals to consult AI tools during the test session.
Webcam-based proctoring may serve as an additional safeguard, although this approach raises privacy concerns and may not be appropriate in all contexts. In such cases, it is essential that candidates are informed in advance about the use of webcam monitoring, that they provide explicit consent, and that they are given clear information on how their data will be handled, stored, and protected. Regardless of the method, it is important to maintain transparency: Test takers should be informed about any monitoring tools or restrictions that are in place. This not only ensures compliance with data protection regulations but also helps support candidate trust and reduce dissatisfaction with the testing process.
It is important to acknowledge that no single method can fully prevent faking, especially as AI models continue to improve. However, a combination of assessment selection, warnings, and technological controls can substantially reduce the likelihood of dishonest responding and help preserve the validity of test results.
Rethinking the Role of AI in Assessment
Finally, some researchers have begun to propose a shift in perspective. Rather than treating AI use solely as a threat to test validity, it may be appropriate in certain contexts to consider AI use as a relevant skill in its own right. Lievens and Dunlop (2025) argue that if the ability to use AI is essential to the job being assessed, then AI-assisted performance on assessments might reflect job-relevant behavior rather than response distortion. In such cases, AI use could be incorporated into the design of the assessment itself, transforming what would otherwise be considered error variance into a construct of interest.
While this perspective is not yet widely adopted, it underscores the importance of context in assessment design. Not all uses of AI are inherently deceptive or undesirable, and future practices may benefit from a more nuanced understanding of when and how such tools are used.
Rethinking the Role of Personality Tests in Assessment
In selection contexts, the primary role of personality assessment is often to identify candidates who best fit the requirements of the position. In this sense, personality tests are commonly used to “select-in” the strongest applicants. However, research suggests that faking poses less of a threat when personality tests are used differently – namely, to “select-out” candidates who clearly do not fit the position (MacCann et al., 2011).
The reasoning is straightforward: Faking primarily affects rank order at the upper end of the distribution, where individuals may displace one another by faking higher scores. By contrast, those who engage in faking are not necessarily drawn from the very bottom of the distribution. As a result, using personality tests to exclude the lowest-scoring group of applicants reduces the disruptive impact of faking on the overall selection outcome.
Supporting this view, Griffith et al. (2007) looked at cases where the selection ratio is high – specifically, when more than 50% of candidates are selected to continue in the selection process. In such cases, even if some candidates inflate their scores, the relative rank among top scorers is less important, because the majority of them are advancing anyway. What matters more is that those with genuinely low scores are correctly identified and excluded. Put differently, personality assessments may be most effective when used as a screen-out tool to remove the lowest 10–20% of candidates, rather than to identify only the very top performers.
Taking these considerations into account, the next section of this paper will present a set of recommendations designed to help practitioners respond effectively to the risks of AI-assisted faking in personality assessment.
Recommendations
1 Focus on prevention rather than detection alone.
Consider using assessment designs that are less vulnerable to manipulation, such as:
- Situational judgment tests, which demand contextual reasoning.
- Forced-choice formats, which limit the ability to universally endorse desirable traits.
- Verifiable response formats, which ask for concrete behaviors rather than self-perception.
Note that each of these formats has advantages and trade-offs. Practitioners should choose the format that best aligns with the specific selection requirements and role in question. For advice on this matter, Hogrefe Consulting offers expert consultation and support.
If implementation of proctored testing is possible, this option has the highest chance of eliminating AI-assisted faking. However, it does not eliminate the possibility of traditional faking, so other measures still need to be considered.
If unproctored testing is necessary, consider using secure digital platforms such as the HTS that restrict or report test-taking behaviors such as copying text, leaving the test window, reaction times, or taking screenshots. Shortening response time windows can also reduce opportunities for AI-assisted faking. For high-stakes decisions, remote proctoring via webcam or in-person testing may be warranted.
2 Use personality assessment as a “screen-out” tool.
Using personality assessment to exclude the candidates with inappropriately low levels of desired attributes will give “fakers” a smaller chance to displace desirable candidates in the selection procedure.
3 Communicate clearly with candidates.
Transparent communication about the purpose of the assessment, the importance of honest responding, and any technical safeguards in place can increase test taker compliance, trust, and engagement. Consider also including a note that AI-assisted responses may not reflect a candidate's actual fit for the role and could work against their interests. At Hogrefe Consulting we provide practical advice on how to communicate the implementation of assessments and the measures against faking in place.
4 Use warnings thoughtfully and transparently.
Warnings about the detection and consequences of faking reduce deceptive behavior but must be used with care. Avoid overstating the capabilities of detection systems and be mindful of the potential for increasing anxiety among honest respondents. Where possible, frame warnings in an educational way, emphasizing the value of honest responses for achieving good person–job fit, which is in the best interest of both the candidate and the employer.
5 Interpret test scores with caution – especially when results are atypical.
When existing detection methods suggest that faking has occurred, we advise practitioners to retest or interpret scores cautiously rather than to exclude the identified fakers or try to correct their scores. Supplementary assessment methods, such as structured interviews or reference checks, may help clarify whether results reflect genuine self-presentation and should be applied before making final decisions.
6 Stay informed and monitor emerging developments.
Generative AI technology is evolving rapidly. Assessment providers and HR professionals should continue to monitor new research, updates to AI capabilities, and developments in detection and prevention methods. Hogrefe remains committed to contributing to this research and providing clients with timely guidance.
7 Consider contextual flexibility.
Finally, it is important to remember that not all uses of AI are inappropriate. In some cases, particularly where AI use is part of the job, applicant’s ability to leverage these tools may represent a job-relevant skill rather than deception. Assessment strategies should remain flexible and context-sensitive.
© 2025 Hogrefe Publishing Group
PDF download
Author
R&D Dept
Aline Schwarz
- Send email
- R&D@hogrefe.com
Hogrefe Publishing Group
Merkelstr. 3
37085 GöttingenGermany
References
Budhwar, P., Chowdhury, S., Wood, G., et al., (2023). Human resource management in the age of generative artificial intelligence: Perspectives and research directions on ChatGPT. Human Resource Management Journal, 33(3), 606-659. DOI: 10.1111/1748-8583.12524
Crowne, D. P., & Marlowe, D. (1960). A new scale of social desirability independent of psychopathology. Journal of Consulting Psychology, 24(4), 349–354. doi.org/10.1037/h0047358
Dilchert, S., & Ones, D. S. (2011). Application of preventive strategies. In M. Ziegler, C. MacCann, & R. Roberts (Eds.), New perspectives on faking in personality assessment (pp. 177–200). Oxford University Press. 10.1093/acprof:oso/9780195387476.003.0054
Griffith, R. L., Chmielowski, T., & Yoshita, Y. (2007). Do applicants fake? An examination of the frequency of applicant faking behavior. Personnel Review, 36(3), 341–355. doi.org/10.1108/00483480710731310
Griffith, R. L., & Converse, P. D. (2011). The rules of evidence and the prevalence of applicant faking. In M. Ziegler, C. MacCann, & R. Roberts (Eds.), New perspectives on faking in personality assessment (pp. 34–52). Oxford University Press. 10.1093/acprof:oso/9780195387476.003.0018
Heggestad, E. D. (2011). A conceptual representation of faking: Putting the horse back in front of the cart. In M. Ziegler, C. MacCann, & R. Roberts (Eds.), New perspectives on faking in personality assessment (pp. 87–101). Oxford University Press. 10.1093/acprof:oso/9780195387476.003.0033
Westfall, B. (2024). Nearly Half of Job Seekers Are Using AI To Cheat – Here’s How Recruiters Can Fight Back. Capterra. www.capterra.com/resources/how-to-stop-jobapplication-ai-cheating/
Kuncel, N. R., Borneman, M., & Kiger, T. (2011). Innovative item response process and Bayesian faking detection methods: More questions than answers. In M. Ziegler, C. Maccann, & R. Roberts (Eds.), New perspectives on faking in personality assessment (pp. 102–112). Oxford University Press. 10.1093/acprof:oso/9780195387476.003.0036
Li, A., & Bagger, J. (2006). Using the BIDR to distinguish the effects of impression management and self-deception on the criterion validity of personality measures: A meta-analysis. International Journal of Selection and Assessment, 14(2), 131–141. doi.org/10.1111/j.1468-2389.2006.00339.x
Lievens, F., & Dunlop, P. D. (2025). Effects of applicants’ use of generative AI in personnel selection: Towards a more nuanced view? International Journal of Selection and Assessment, 33(1), e12516. doi.org/10.1111/ijsa.12516
Maccann, C., Ziegler, M., & Roberts, R. (2011). Faking in personality assessment: Reflections and recommendations. In M. Ziegler, C. MacCann, & R. Roberts (Eds.), New perspectives on faking in personality assessment (pp. 309–329).
doi.org/10.1093/acprof:oso/
9780195387476.003.0087
McFarland, L. A., & Ryan, A. M. (2000). Variance in faking across noncognitive measures. Journal of Applied Psychology, 85(5), 812–821. doi.org/10.1037/0021-9010.85.5.812
Meade, A. W., & Craig, S. B. (2012). Identifying careless responses in survey data. Psychological Methods, 17(3), 437–455. doi.org/10.1037/a0028085
Moon, B., Bourdage, J.S. & Roulin, N. (2025). Investigating Deceptive Impression Management in Behavioral Description Interviews Through a Cognitive Load Perspective. J Bus Psychol. doi.org/10.1007/s10869-025-10042-7
Nguyen, N. T., Biderman, M. D., & McDaniel, M. A. (2005). Effects of response instructions on faking a situational judgment test. International Journal of Selection and Assessment, 13(4), 250–260. doi.org/10.1111/j.1468-2389.2005.00322.x
Ones, D. S., Viswesvaran, C., & Reiss, A. D. (1996). Role of social desirability in personality testing for personnel selection: The red herring. Journal of Applied Psychology, 81(6), 660–679. doi.org/10.1037/0021-9010.81.6.660
Open AI. (2025). ChatGPT agent. chatgpt.com/features/agent
Paulhus, D. L., Harms, P. D., Bruce, M. N., & Lysy, D. C. (2003). The over-claiming technique: Measuring self-enhancement independent of ability. Journal of Personality and Social Psychology, 84(4), 890–904. doi.org/10.1037/0022-3514.84.4.890
Peeters, H., & Lievens, F. (2005). Situational judgment tests and their predictiveness of college students’ success: The influence of faking. Educational and Psychological Measurement, 65(1), 70–89. doi.org/10.1177/0013164404268672
Phillips, J., & Robie, C. (2024a). Can a computer outfake a human? Personality and Individual Differences, 217, 112434. doi.org/10.1016/j.paid.2023.112434
Phillips, J., & Robie, C. (2024b). Hacking the perfect score on high-stakes personality assessments with generative AI. Personality and Individual Differences, 231, 112840. doi.org/10.1016/j.paid.2024.112840
Prelec, D. (2004). A Bayesian truth serum for subjective data. Science, 306(5695), 462–466. doi.org/10.1126/science.1102081
Reeder, M. C., & Ryan, A. M. (2011). Methods for correcting for faking. In M. Ziegler, C. MacCann, & R. Roberts (Eds.), New perspectives on faking in personality assessment (pp. 131–150). Oxford University Press. 10.1093/acprof:oso/9780195387476.003.0041
Robie, C., Phillips, J., Bourdage, J. S., Christiansen, N. D., Dunlop, P. D., Risavy, S. D., & Speer, A. B. (2025). Can ChatGPT outperform humans in faking a personality assessment while avoiding detection? International Journal of Selection and Assessment, 33(3), e70015. doi.org/10.1111/ijsa.70015
Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068. doi.org/10.1037/apl0000994
Stark, S., Chernyshenko, O. S., & Drasgow, F. (2011). Constructing fake-resistant personality tests using item response theory: High-stakes personality testing with multidimensional pairwise preferences. In M. Ziegler, C. MacCann, & R. Roberts (Eds.), New perspectives on faking in personality assessment (pp. 214–239). Oxford University Press. 10.1093/acprof:oso/9780195387476.003.0061
Tett, R. P., Jackson, D. N., & Rothstein, M. (1991). Personality measures as predictors of job performance: A meta-analytic review. Personnel Psychology, 44(4), 703–742. doi.org/10.1111/j.1744-6570.1991.tb00696.x
Tett, R. P., & Simonet, D. V. (2021). Applicant faking on personality tests: Good or bad and why should we care? Personnel Assessment and Decisions, 7. scholarworks.bgsu.edu/pad/vol7/iss1/2/
Thresher, S. (2024, December 17). Navigating AI and cheating in early talent hiring. Talogy. talogy.com/en/blog/preparing-for-the-future-aiand-cheating-in-early-talent-hiring/
Ventura, M. (2011). The detection of faking through word use. In M. Ziegler, C. MacCann, & R. Roberts (Eds.), New perspectives on faking in personality assessment (pp. 165–174). Oxford University Press. 10.1093/acprof:oso/9780195387476.003.0049
Wilmot, M. P., & Ones, D. S. (2019). A century of research on conscientiousness at work. Proceedings of the National Academy of Sciences, 116(46), 23004–23010. doi.org/10.1073/pnas.1908430116
Zickar, M. J., & Sliter, K. A. (2011). Searching for unicorns: Item response theory-based solutions to the faking problem. In M. Ziegler, C. MacCann, & R. Roberts (Eds.), New perspectives on faking in personality assessment (pp. 113–130). Oxford University Press. doi.org/10.1093/acprof:oso/9780195387476.003.0038
Ziegler, M., MacCann, C., Roberts, R. D., Ziegler, M., MacCann, C., & Roberts, R. (Eds.). (2011). Faking: Knowns, unknowns, and points of contention. In M. Ziegler, C. MacCann, & R. Roberts (Eds.), New perspectives on faking in personality assessment (pp. 3–16). Oxford University Press. doi.org/10.1093/acprof:oso/9780195387476.003.0011