The Screening Assessment Decoder, One Row per Test Type
You can name each test in a screening link, state what it measures and how it is scored, and decide per module whether to prepare, answer straight, or decline.
A link arrives after you apply: complete these online assessments before your next interview. It is rarely one test. It is a stack of modules, each measuring a different thing and scored by a different rule, and the invite almost never tells you which is which. This guide is a lookup document for that stack. Jump to the row for the module in front of you, read what it proves, how it is scored, how it misleads, and whether preparation, a straight answer, or declining is the right move.
The automated psychometric and skills-screening layer is now the default gate. In the aggregate, employer surveys put pre-hire assessment use somewhere between 65% and 82% of companies, with the spread driven by survey year, company size, and definition rather than by disagreement about the trend. The mechanism is simpler than the number: high applicant volume plus a link that scores itself. Your job is to read the stack correctly before the timer starts.
What the six assessment types measure
There are four core families and three common sub-types. Cognitive ability, personality, situational judgement, and skills are the families; coding challenges, integrity, and typing or data-entry are sub-types that show up inside a link. Each measures a different construct, and reading the construct is the whole game, because it tells you what a score can and cannot move.
- Cognitive ability measures reasoning capacity: the mental processes involved in thinking, including reasoning, perception, memory, verbal and mathematical ability, and problem solving. It estimates your potential to solve work-related problems and acquire new job knowledge.
- Situational judgement (SJT) measures your choices in work scenarios against what experts or high performers would do. It probes judgement, not knowledge.
- Personality measures stable traits, typically along dimensions such as conscientiousness and emotional stability, and carries built-in checks for inconsistent or overly virtuous answering.
- Integrity is a personality sub-type aimed at counterproductive-work-behavior risk, and it leans heavily on the same consistency and social-desirability machinery.
- Coding challenge is a timed skills test that filters for solving structured problems under time pressure in a browser or take-home environment.
- Skills and job-simulation tests measure acquired expertise: job-knowledge or achievement tests for a specialized area, and timed hard-skill checks like typing or SQL.
A word on stable versus trainable, because it decides your prep budget. Cognitive ability is a stable construct: it measures learning ability, which remains relatively stable across age groups, so drilling improves your speed with the format more than your underlying score. Coding is the opposite. It rewards trained pattern recognition, so practice converts directly to pass rate.
The assessment stack, outermost to innermost
- Cognitive gateTimed reasoning test, often first and short (e.g. a 12-minute general-ability assessment)
- Judgement layerSituational judgement scenarios keyed to a norm group
- Trait layerPersonality or integrity items with consistency and social-desirability checks
- Skills layerCoding, SQL, typing, or job-knowledge tests measuring acquired expertise
The decoder table, one row per test type
This is the reference. For each type, the row gives what it measures, how it is scored, and the single move that pays. Read the row for the module you are looking at and skip the rest.
| Test type | What it measures | How it is scored | The move |
|---|---|---|---|
| Cognitive ability | Reasoning capacity and learning ability | Job-relative, norm-referenced; no universal pass mark | Practice format for speed; do not expect large score jumps |
| Situational judgement | Judgement in work scenarios | Expert or empirical key, reported as percentile or band | Aim for the norm group, not a high raw count |
| Personality | Stable traits (conscientiousness, stability) | Trait profile plus consistency and social-desirability flags | Answer consistently and honestly |
| Integrity | Counterproductive-behavior risk | Same flag machinery as personality | Do not over-endorse virtuous items |
| Coding challenge | Structured problem-solving under time | Test cases pass or fail, timed | Practice under timed, no-autocomplete conditions |
| Skills / typing | Acquired hard or soft skill | Right/wrong or measured throughput (words per minute) | Rehearse in the exact tool |
Two scoring distinctions do most of the work here. Norm-referenced scoring compares you to other people and reports where you rank, which is how SJTs and most cognitive tests behave. Keyed scoring marks answers right or wrong against a fixed key or test cases, which is how coding and job-knowledge tests behave. Knowing which one you face tells you whether beating the average or hitting a correct answer is the target.
How each type is scored, and the threshold that matters
Scoring model, not effort, decides most outcomes. The clearest example is the SJT: results are compared against others to place you in a percentile group, and a good score typically ranges from the 70th to 80th percentile. Band scores work the same way, calculated as the percentage of your answers that match employer-preferred answers, with Band 1 or Band 2 most likely to see an application proceed.
That relativity is a trap. Other applicants also understand what an effective response looks like. If they answer 27 of 30 correctly on average, your 26 of 30 "good result" is actually below average, because it ranks lower than people in similar roles. The raw count feels strong and the ranking is weak. There is no fixing this by trying harder on a near-consensus item; the fix is understanding that the target is the norm group.
Cognitive tests are scored job-relative rather than on a universal scale. Unlike an IQ test, interpretation is job-specific and relative to the skills the role requires, so there is no publicly established pass mark you can aim at. Personality and integrity tests layer trait scores under consistency and social-desirability checks, described in their own section below.
Predictive validity is worth knowing because it explains why cognitive tests keep showing up first. But the published coefficients diverge, and a single confident number would hide that.
| Source | Coefficient or claim |
|---|---|
| Schmidt 2016 (via SHRM) | 0.54 |
| x0pa citing research | 0.62 |
| NBER (via taggd.in) | ~25% of performance variance |
Read that spread as a caveat, not a contradiction. Cognitive ability carries the highest reported predictive validity among assessment types, which is why employers front-load it, but "highest" still leaves most of job performance unexplained.
The scoring model, not the effort, decides the outcome; a 26 of 30 fails when the room averages 27.
The step-by-step: decode a link before you start
Run this in order. It takes under an hour of desk work before any timer starts, and it prevents the two costly errors: burning a coding attempt unprepared, and answering a norm-referenced test as if effort alone carried it.
Decode and clear a screening link
- Inventory the linkOpen the invite and list every sub-test with its stated time limit before starting. Done when each module is named and you know the total time commitment.
- Classify each module by constructTag each as cognitive, SJT, personality, integrity, coding, skills, or typing. Done when every module has a label for what it measures.
- Determine the scoring modelFor each module, decide whether it is percentile, band, expert-keyed, or a consistency and social-desirability check. Done when you know norm-referenced from keyed.
- Decide prepare versus declinePrepare where practice moves the score, answer straight where it does not, reserve declining for real grounds. Practice helps coding and format; the cognitive construct is stable.
- Prepare where it paysFor coding, do 3 to 5 problems per day timed, autocomplete off, tabs closed. Budget 2 to 4 weeks with solid fundamentals, 4 to 6 weeks if rusty.
- Answer personality and integrity straightRespond consistently and honestly so same-trait items match and you do not over-endorse virtuous statements. Done when responses are self-consistent.
- Check for legal red flagsFlag any pre-offer medical, genetic, age, or religious item. Done when each module is confirmed job-related or the item is documented.
- Complete under controlled conditionsSit it in a quiet room with a stable connection, since a coding assessment can run 60 to 90 minutes. Done when every module is submitted cleanly.
Sources disagree on where declining belongs in the sequence. Some frame it as always process-ending, others note the narrow ADA and religious grounds. I place the prepare-versus-decline decision after classification because you cannot judge whether a module is worth preparing for, or worth refusing, until you know what it measures.
Tailoring the surrounding application to the posting is a separate job from clearing the link, but it feeds the same interview. Refolk writes your resume from your own history, tailors it to each posting, and scores how well you fit, which frees the hours you would otherwise spend on wording to spend on timed coding practice instead.
Prepare, answer straight, or decline: the decision per type
The move depends on two axes: whether practice moves the score, and whether the module is norm-referenced or keyed. Plot a module on those axes and the decision falls out.
What to do with a module
For coding specifically, preparation is not optional if you want to pass. These assessments are not designed to find genius-level programmers; they filter for people who can solve structured problems under time pressure, which is a trainable skill. Get comfortable with the environment before the timer, run a "hello world," and understand how code is submitted. Then practice the constraint, not just the concept.
Daily (60 to 90 min): - Set a hard timer per problem (match the real limit if known) - Close all other tabs; disable IDE autocomplete - Solve 3 to 5 problems: 1 easy warm-up, 2 to 3 medium - Write out complexity before coding; state it aloud Weekly: - One full-length mock in the actual platform (e.g. a 60 to 90 min sitting) - Log every problem you failed and why: syntax, edge case, or approach - Re-solve last week's failures cold, no notes Before the real test: - Run a "hello world" in the real environment first - Confirm how code is submitted and whether tests run on submit
Compress to your rustiness; 2 to 4 weeks with solid fundamentals, 4 to 6 weeks if rusty.
On declining: you can refuse, but the practical consequence is usually removal from the process. Documented legal protection to refuse is narrow. There are situations where you may legally refuse, for example if you have an ADA condition and the test would conflict with the reasonable-accommodations provision, or on genuine religious grounds. If you have such a ground, request an accommodation in writing rather than simply not completing the link.
How each type misleads: false positives and failure modes
This is the section that saves you. Every assessment type has a way of producing a confident wrong read, either by you about your result or by the test about you. Here is what each looks like when it lies.
- Personality, gamed. You think you passed but tripped a consistency flag. The same trait is measured by several differently worded items, and contradictory answers flag inconsistency. The tell that you did this: you tried to "read" what each item wanted instead of answering the same way every time.
- Integrity, over-endorsed. Agreeing to items like "I have never told a lie" flags impression management. A perfect virtue score is itself the red flag.
- SJT, misread raw. A high raw count reads as a win but sits below the norm group. Check whether scoring is percentile or band, not raw, before you celebrate a 26 of 30.
- Cognitive, prep illusion. Drilling improves format speed but not the stable underlying construct, so do not expect a large score jump from practice. Practice to remove format friction, not to raise the ceiling.
- Coding, burned attempt. Starting unprepared, then hitting the cooldown, then being locked out for months. The fail is the environment and the timer, not the algorithm.
- Decline, assumed safe. Refusing a lawful, job-related test usually ends candidacy with no recourse; the narrow exceptions are ADA and religious grounds.
- Skills test, wrongly reported. Treating a lawful skills check as reportable. Only pre-offer medical, genetic, or protected-class items are reportable, not job-related skills.
- Faking flags, over-trusted. Social-desirability scales are a contested faking indicator; there is evidence they are a poor one, and correcting scores on that basis does not improve validity. Do not assume a low social-desirability score is a clean bill, or a high one a conviction.
The through-line: honesty and consistency beat gaming on the trait modules, and preparation beats talent on the coding module. The people who lose do the opposite, gaming the trait tests and winging the coding.
The legal red flags worth checking
Before you complete any module, confirm each item is job-related. A narrow set of questions is genuinely off-limits or reportable, and the boundary is pre-offer medical and genetic inquiry, plus protected-class discrimination, not skills testing.
- Under the ADA, an employer may not ask a job applicant medical questions or require a medical exam before making a job offer.
- Asking about genetic information is specifically prohibited under the Genetic Information Nondiscrimination Act of 2008 (GINA).
- Title VII prohibits discrimination on the basis of protected criteria such as race, gender, religion, and national origin, so items screening on those grounds are improper.
A job-related aptitude, personality, or coding test is lawful even though it is demanding. What crosses the line is a pre-offer health question, a genetic inquiry, or an item that sorts on a protected class. If you see one, document it and raise it with the employer or the relevant agency rather than assuming the whole assessment is invalid.
Who builds these tests, and who takes them
The people who design assessments are heavily concentrated in one country, which shapes the norms you are measured against. In Refolk's index of professional profiles, there are 472 US profiles with psychometrician or industrial-organizational psychologist titles versus 12 in the UK, roughly a 39 times gap. Most candidates worldwide, then, face tests built to US I-O norms even when the role is elsewhere.
| Country | Psychometrician / I-O titles | Ratio vs UK |
|---|---|---|
| US | 472 | 39x |
| UK | 12 | 1x |
The candidate side of the coding pipeline tilts the other way. In Refolk's index, there are 349,827 US "Software Engineer" profiles versus 559,906 in India, so India is 62% of that two-country pool. That volume is part of why coding-screen automation scaled first for high-volume engineering pipelines: a link that scores itself is how you triage hundreds of thousands of applicants. Understanding that mechanism is more useful than any single prevalence percentage.
| Country | Software Engineer profiles | Share of the two-country pool |
|---|---|---|
| India | 559,906 | 62% |
| US | 349,827 | 38% |
If you want to see who actually builds and runs these screens for a given sector, you can search the same index by role and industry.
Before you submit: the clearing checklist
Run this last, once per link, right before you start the first timed module. It converts everything above into a single pass.
Screening-link readiness
- Every module is named and tagged by construct (cognitive, SJT, personality, integrity, coding, skills, typing)
- You know each module's scoring model: norm-referenced or keyed
- You have decided prepare, answer straight, or decline for each module
- Coding modules are rehearsed under timed, no-autocomplete conditions and you have run a "hello world" in the real environment
- Personality and integrity answers will be consistent across repeated items and free of over-endorsed virtue statements
- No module contains a pre-offer medical, genetic, or protected-class item; any that does is documented
- You are in a quiet room with a stable connection for a 60 to 90 minute sitting
Keeping this current
The construct rows do not age: cognitive tests still measure stable reasoning, SJTs are still norm-referenced, coding is still trainable under time. What moves is the prevalence numbers and the platform mechanics. Treat any single "percent of employers using tests" figure with suspicion, because the published range runs from 65% to 82% purely on differences of year, company size, and definition. If you need a current number, note which survey, which year, and which company-size band produced it before you trust it.
Two mechanics worth re-checking when a link arrives: the cooldown window after a failed attempt, which commonly runs 6 to 12 months but varies by employer, and the exact environment a coding module runs in, since submission behavior and available tooling change between platforms. Confirm both from the invite itself rather than from memory. The rest of the decoder holds, one row per type, whenever the next link lands.
Put this to work
Paste your career in once. Every application after that is written for you.
Drop a resume or a LinkedIn URL. I rank the live openings against it, rewrite the resume and write a cover letter for the best of them, and fill in the employer's form when you press the button. You read, you decide what goes out.
01Drop your resume
A PDF or a LinkedIn URL. About a minute, once.
02I rank the openings
Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.
03Each one is written up
Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.
- New matches ranked and written before you are up.
- Every bullet stays inside what your history supports. Nothing invented.
- Queued, submitted, interviewing, offer: one screen, not a spreadsheet.
500 free credits on sign-up. No card. Nothing is sent until you say so.