Blog Post

When Resume Scores Keep Changing: What ATS Algorithms Teach Us About Fairness at Work

Khaled Editor · 2026-06-29 17:31

When Resume Scores Keep Changing: What ATS Algorithms Teach Us About Fairness at Work

A recent online debate, including discussion around HackerRank open-sourcing an applicant tracking system, focused on a complaint many job seekers already know well: resume scores do not always stay still. Applicants and developers described cases where similar resumes, or even the same resume in slightly different formats, received different evaluations. These are reported experiences and demonstrations, not proof that every ATS works this way. But the issue matters because automated screening is now common in hiring, and a shaky score can shape who gets seen by a recruiter and who disappears from the process.

The real debate is bigger than transparency. It is about whether systems that help decide access to work are stable, understandable, and open to challenge. A company can publish code, show an AI-generated score, or say a human is still involved, and still run a process that feels arbitrary. If a resume score changes for reasons an applicant cannot see or contest, then the fairness problem is not only opacity. It is unreliable decision-making.

A moving score is a warning sign

In hiring, a score that cannot stay reasonably stable should not be used as a gatekeeper.

That is the clearest lesson from the recent discussion. Fairness is not only about treating groups equally in the aggregate. It is also about procedural consistency. Similar candidates should be judged by similar standards. If the same candidate can look strong on Monday and weak on Tuesday because a parser handled a PDF differently, a model version changed, or a job description was rewritten without notice, the process is not ready to make high-stakes decisions.

This does not mean every score change is automatically unfair. Context matters. A candidate may score differently for different jobs, and that is normal. A resume may rank differently when compared against a new applicant pool, and that can also be legitimate. But if employers present the output as an objective measure of merit, they create a false sense of precision. A number on a screen can look solid even when the system behind it is not.

There are many ordinary reasons scores can shift:

  • A resume parser may read a two-column layout badly and miss key experience.
  • A tool may weight exact phrases like “project management” more heavily than close alternatives.
  • A system may be updated between hiring cycles, or even during one.
  • Different recruiters may set different thresholds for the same role.
  • A score may be relative to other applicants rather than an absolute measure.

None of these issues is exotic. That is exactly why they matter. Applicants are often told that ATS tools are neutral, consistent, and efficient. In practice, many systems are highly sensitive to formatting choices, keyword choices, and internal configuration choices that applicants never see.

Why transparency does not solve the problem

There is value in more openness. If a company or vendor reveals how a system works, researchers, lawyers, journalists, and customers have more to inspect. Open-sourcing software can help uncover bugs, bad assumptions, and weak evaluation methods. That is better than pure secrecy.

But transparency alone is not enough. A public code repository does not tell applicants how a specific employer has configured the system, what data it relies on, what thresholds trigger rejection, or whether a recruiter ignores the score half the time. It does not reveal how the parser handles nonstandard resumes, how often the model is updated, or whether the tool has been tested on international CVs, career changers, or people returning after a gap.

Even experts cannot judge fairness from code alone. They need to know the surrounding process. What counts as a minimum qualification? Which features are allowed to influence the score? Is there human review for borderline cases? Can a candidate correct bad extraction or missing information? Are outcomes audited across gender, disability, race, age, language background, and school prestige? Fairness lives in these operational details.

There is also a practical point. Most applicants cannot audit an algorithm. Telling them a tool is transparent may sound reassuring, but it often changes very little about their power in the process. A fair system needs more than visibility. It needs accountability, documentation, and a way to fix mistakes.

The case for ATS tools, and where it holds up

Employers do have a real problem to solve. Many roles attract hundreds or thousands of applications. Small recruiting teams cannot review every file in depth. Search, deduplication, scheduling, and basic qualification checks save time. Structured screening can also reduce some human inconsistency. A tired recruiter with 600 resumes to scan is not a gold standard of fairness either.

That counterpoint matters. The choice is not between perfect humans and flawed software. Human hiring already contains bias, noise, and snap judgment. In some cases, structured tools can improve consistency. A clear requirement like a license, language certification, or years of specific experience can be easier to apply uniformly with software than by rushed manual review.

But this defense has limits. The fact that hiring is hard does not justify using unstable scoring as a hidden filter. Automation can help with logistics and broad organization. It should be treated much more carefully when it starts ranking, excluding, or downgrading people on fuzzy signals. A tool that saves recruiters time but quietly blocks qualified candidates is not a neutral efficiency gain. It shifts the burden of error onto applicants.

There is another reason to be cautious: algorithmic inconsistency scales fast. One recruiter’s bad afternoon can harm a few applicants. A weak screening system can affect thousands, and do it with the appearance of objectivity. That is why “humans are biased too” is a true point, but not a full defense.

What fairness should mean in automated hiring

If employers want to use AI-assisted screening responsibly, they need a stricter standard than “the model generally works.” Hiring is not a product recommendation problem. It is a gateway to income, dignity, and opportunity. The system should be designed around job relevance, consistency, and reviewability.

At a minimum, fairer use would look like this:

  • Freeze the rules for each hiring round. If scoring criteria or model versions change, document it and avoid switching standards mid-process.
  • Test for stability. The same resume should not swing wildly because of file format, layout, minor wording changes, or synonyms.
  • Separate hard requirements from soft scoring. If a role legally requires a license, say that clearly. Do not hide soft preference judgments inside an opaque numeric score.
  • Give meaningful explanations. “You were not selected” is not enough when automation played a role. Candidates should know whether the issue was missing requirements, poor parsing, or competitive ranking.
  • Offer a human review path. Especially for near-threshold cases, career changers, and applicants who believe the system extracted their resume incorrectly.
  • Audit outcomes, not just code. Employers should test how the system treats different formats, schools, languages, employment gaps, and nontraditional backgrounds.
  • Limit what the tool is allowed to decide. Organizational convenience is not a good reason to give a probabilistic score the power to reject people automatically.

These steps are not radical. They are basic quality control for a tool used in a high-stakes setting.

The keyword game is not a fairness strategy

One bad response to ATS uncertainty is to turn hiring into a game of optimization. Job seekers are told to copy phrases from job descriptions, strip personality from their resumes, and write for a parser before they write for a human. Some of that advice is practical. Simple formatting helps. Clear section headings help. Matching the employer’s language, when truthful, helps.

But the bigger trend is unhealthy. When career guidance becomes “beat the machine,” the burden shifts away from the employer and onto the applicant. Students, recent graduates, and non-native English speakers are especially affected. They may spend hours guessing which words a system prefers instead of showing what they can actually do.

Universities and career centers should teach the basics of machine-readable resumes, but they should stop there. Their more important role is to push employers for better practice. They can ask whether automated scoring is used, whether candidates can request a human review, and how the system handles international experience, portfolio-based work, and unconventional career paths. That is more useful than endless workshops on keyword stuffing.

What applicants should remember

For job seekers, the practical lesson is simple. Use a clean format. Avoid putting key information in text boxes, tables, or images. Tailor your resume to the role with honest language that matches the job description. Keep a plain-text or simple-layout version ready. Those steps improve the odds that a system will read your file correctly.

But applicants should also resist a damaging conclusion: a low or changing score is not a final verdict on ability. It may reflect a tool’s limits, a bad parser, shifting criteria, or a ranking system that was never designed to capture your strengths well. The problem is often structural, not personal.

The standard that matters

The best test of fairness is not whether an employer can say it uses AI responsibly. It is whether a qualified person has a real chance to be assessed on relevant evidence under stable rules. That standard is harder to meet than posting an explanation page or publishing code, but it is the standard that work deserves.

When resume scores keep changing, the burden should not fall on applicants to decode the system. It should fall on employers and vendors to prove the process is reliable, limited, and open to correction. In hiring, consistency is not a luxury feature. It is part of equal access to work.

← Back to Blog