1. Set the criteria and what they are worth

Pick a role to load a starting set, then adjust the weights to your situation. Do this before you score anyone, so the standard is fixed while you are still objective.

Shared remote criteria

Inbox and calendar judgment

Sorts urgent from noisy without being told, and knows what to escalate rather than sitting on it.

Accuracy and attention to detail

Data entry, filing, and scheduling done right the first time. Check the test task for silent errors.

Follows a brief

Did the thing that was asked, and flagged assumptions rather than guessing quietly.

Tool fluency

Comfortable in your stack, or shows a track record of picking up new tools fast.

Works without hand-holding

Evidence of owning recurring work end to end, not waiting to be assigned each task.

Written communication

Clear, well organised writing in their application, emails, and test task. This is most of the job in a remote role.

Remote reliability and home setup

Named internet speed, a mobile-data fallback, and a power plan such as an inverter, UPS, or generator for load-shedding.

2. Score every candidate against the same standard

Score one criterion across all candidates before moving to the next, rather than finishing one person at a time. It is much harder to drift into a favourite that way. Use reference codes or initials rather than full names, so this scorecard is safe to share internally.

Inbox and calendar judgment (weight 5, must-have)

Candidate A

Not scored

Candidate B

Not scored

Candidate C

Not scored

Accuracy and attention to detail (weight 4)

Candidate A

Not scored

Candidate B

Not scored

Candidate C

Not scored

Follows a brief (weight 4, must-have)

Candidate A

Not scored

Candidate B

Not scored

Candidate C

Not scored

Tool fluency (weight 3)

Candidate A

Not scored

Candidate B

Not scored

Candidate C

Not scored

Works without hand-holding (weight 4)

Candidate A

Not scored

Candidate B

Not scored

Candidate C

Not scored

Written communication (weight 4, must-have)

Candidate A

Not scored

Candidate B

Not scored

Candidate C

Not scored

Remote reliability and home setup (weight 4, must-have)

Candidate A

Not scored

Candidate B

Not scored

Candidate C

Not scored

What the numbers mean

  • 1 - Well below bar. Clear evidence they cannot do this part of the job yet
  • 2 - Below bar. Some evidence, but noticeably weaker than what the role needs
  • 3 - Meets bar. Does what the role needs, no concerns and no standout
  • 4 - Above bar. Better than the role needs, with concrete evidence behind it
  • 5 - Exceptional. Best you have seen in this pool, and you can point to why

3. Read the result

Each score is multiplied by its weight, then shown as a percentage of the weighted points available on the criteria you have actually scored.

21 of 21 boxes are still unscored. The ranking below only counts the criteria you have scored, so treat it as provisional until the grid is complete.

1. Candidate A

Not scored

0 of 0 weighted points across 0 scored criteria, 7 still unscored.

2. Candidate B

Not scored

0 of 0 weighted points across 0 scored criteria, 7 still unscored.

3. Candidate C

Not scored

0 of 0 weighted points across 0 scored criteria, 7 still unscored.

Summary for your hiring file

Candidate scorecard: General admin virtual assistant

Criteria and weights
- Inbox and calendar judgment (weight 5, must-have)
- Accuracy and attention to detail (weight 4)
- Follows a brief (weight 4, must-have)
- Tool fluency (weight 3)
- Works without hand-holding (weight 4)
- Written communication (weight 4, must-have)
- Remote reliability and home setup (weight 4, must-have)

Results
1. Candidate A: 0% | 7 criteria not yet scored
2. Candidate B: 0% | 7 criteria not yet scored
3. Candidate C: 0% | 7 criteria not yet scored

Scores by criterion
- Inbox and calendar judgment: Candidate A not scored, Candidate B not scored, Candidate C not scored
- Accuracy and attention to detail: Candidate A not scored, Candidate B not scored, Candidate C not scored
- Follows a brief: Candidate A not scored, Candidate B not scored, Candidate C not scored
- Tool fluency: Candidate A not scored, Candidate B not scored, Candidate C not scored
- Works without hand-holding: Candidate A not scored, Candidate B not scored, Candidate C not scored
- Written communication: Candidate A not scored, Candidate B not scored, Candidate C not scored
- Remote reliability and home setup: Candidate A not scored, Candidate B not scored, Candidate C not scored

Scored with the HireSava virtual assistant candidate scorecard: https://hiresava.com/virtual-assistant-candidate-scorecard

Got a winner? Post your role and shortlist vetted South African professionals directly, with video introductions and no middleman. Start hiring on HireSava.

Why the last interview usually wins, and why that is a problem

Ask any small business owner how they chose their assistant and you will usually get some version of the same answer: they just clicked with one of them. It sounds like intuition, and sometimes it is. More often it is a handful of well-documented distortions doing their work quietly. Recency means the person you spoke to yesterday is vivid while the equally good candidate from ten days ago has faded to a few adjectives in your notes. Similarity means the person who reminds you of yourself reads as competent. And the halo effect means one impressive answer, or simply a warm and confident manner, spreads outward until it colours your view of skills that person never actually demonstrated.

The uncomfortable part is how confident this feels from the inside. In a much-cited experiment, Dana, Dawes and Peterson had participants interview subjects who, in one condition, were secretly instructed to answer questions randomly. Interviewers formed confident impressions anyway, made sense of answers that had no meaning in them, and in a later study said they would rather conduct a random interview than none at all. Worse, the impressions they formed diluted the value of the genuinely predictive information they already held. Their paper is titled Belief in the unstructured interview: The persistence of an illusion, which is about as blunt as academic titles get.

A scorecard does not make you objective. Nothing does. What it does is much narrower and much more useful: it makes you commit to a standard while you are still capable of being neutral, and then it holds you to that standard after you have started having feelings about individual people. You decide what the job needs on Monday, before you have met anyone. On Friday, when the charming candidate is fresh in your mind, the weights you set on Monday are still sitting there, unmoved, quietly asking whether charm was on the list.

What the research actually says about structure

For twenty-five years the reference point for this question was Schmidt and Hunter's 1998 meta-analysis, which put cognitive ability at the top of the list of predictors of job performance. In 2022, Paul Sackett, Charlene Zhang, Christopher Berry and Filip Lievens published a re-analysis in the Journal of Applied Psychology showing that decades of these estimates had been inflated by a systematic overcorrection for range restriction. Their paper, Revisiting meta-analytic estimates of validity in personnel selection, reordered the field. The revised figures put structured interviews first, ahead of cognitive ability tests.

PredictorRevised validityWhat it is
Structured interviews.42Same questions, same order, scored against a fixed rubric. The strongest common predictor in the revised set.
Job knowledge tests.40Direct tests of what the person knows about the work itself.
Empirically keyed biodata.38Background and history questions, scored against patterns validated on past hires.
Work sample tests.33A short piece of the real job, done and scored before you hire.
Cognitive ability tests.31General mental ability, long assumed to top this list and no longer at the top of it.

Two things are worth pulling out of that table. The first is that the word doing the work in structured interview is structured, not interview. An unstructured interview is a different activity with a different and much weaker track record. The structure is the intervention: the same questions in the same order, scored against a fixed rubric, for every candidate. The second is that the top two entries are things a small business can genuinely do. You are unlikely to build a validated biodata key for a one-person hire. You can absolutely ask five people the same six questions and score the answers the same way, and you can absolutely give them all the same short work-sample task.

Do not over-read the decimals either. A validity of .42 is a real and useful edge across many hires; it is not a promise about your next one. What the numbers justify is the modest claim this page is making: that structure beats impression reliably enough to be worth the twenty minutes it costs you.

How to build a scorecard that is worth filling in

Keep it to three to seven criteria. This is the single most common mistake, and it runs in one direction. A scorecard with fifteen rows feels rigorous and behaves like noise: every criterion is diluted, the weights stop separating anything, and candidates converge on the same middling total because everyone is good at something. Ask yourself which handful of things this hire genuinely lives or dies on. Usually there are four.

Write criteria you can actually observe. Culture fit is not a criterion, it is a container for whatever you already felt. Replace it with the specific behaviour you meant: responds to feedback without defensiveness, or raises a problem the same day rather than at the deadline. The test is whether you could point at a piece of evidence and say that is a four. If you cannot, the row will quietly become a place to record your overall impression twice.

Set the weights before you screen anyone. Weights set afterwards are not weights, they are a justification for a decision you have already made. This is why the tool asks you to build the criteria list first and only then start scoring. If more than one person has a vote, agree the weights together in that first conversation, because arguing about what matters is much more productive before you have candidates than after.

Use must-haves sparingly and mean them. A must-have is a criterion where a below-the-bar score should stop the hire no matter how good the total is. Reconciliation accuracy for a bookkeeper is a real must-have: someone who forces a balance rather than investigating a difference is not a slightly weaker bookkeeper, they are a liability. Familiarity with your project management tool is not, however convenient it would be. Mark two or three, not eight, because a scorecard where everything is critical has no weights at all.

Score across, not down. Take one criterion and score every candidate on it before you move to the next. Scoring one person completely, then the next, invites a running overall impression to leak into every row, which is exactly the halo effect the scorecard exists to interrupt. Scoring across also gives you a live sense of the spread on each criterion, which is often more informative than the totals.

Score from evidence, and score immediately. Fill in the row within a few minutes of the call or the test submission, while you can still remember what they actually said rather than how you felt about it. Where you have no evidence for a criterion yet, leave it blank. Guessing to fill the grid is worse than an incomplete grid, which is why the tool counts unscored boxes and marks the ranking as provisional until they are filled.

What the scores mean, and what they do not

The arithmetic here is deliberately simple, and you should be able to reproduce it on paper. Each score from one to five is multiplied by that criterion's weight from one to five. Those weighted points are added up and divided by the maximum available on the criteria you have actually scored, so a candidate you have half-assessed is not silently punished for the empty rows. The result is a percentage, and the percentage exists only to make comparison easy. It is not a probability of success and it is not a grade.

The most important output on the results panel is not the ranking, it is the too-close-to-call flag. When your top two are within five points, the scorecard is telling you it cannot separate them, and the correct response is to go and get more evidence rather than to trust the third decimal place of a judgment you made in fifteen minutes. Send both finalists a short paid trial task. Run a reference check aimed squarely at whichever criterion you feel least certain about. Ask each of them to walk you through their work-sample submission and watch which one gets more impressive under gentle questioning. Any of these produces real new information. Staring at 78 versus 76 does not.

The must-have flag works in the opposite direction, and it is there to stop a specific failure. A candidate can be delightful, fast, well-organised, and strong on four of your five criteria, and still be below the bar on the one thing that would make the role work. The weighted total will happily bury that under everything they are good at. When the tool tells you someone failed a must-have, the total is not the answer to it. Either you decide, deliberately and in writing, that you are willing to hire around that gap and how you will cover it, or you do not hire.

One last caution. A scorecard makes a decision legible, which makes it easier to defend, and that can shade into using it as cover for a call you made on other grounds. If you find yourself adjusting weights after the scores are in, stop. The output has stopped being a decision aid and become a document justifying your gut. That is a worse position than never having built the scorecard, because now you believe the number.

Where the scorecard fits in a hiring process

The scorecard is not a step. It is the spine that the steps hang off, and it is open from the moment you write the role to the moment you send the offer. Here is the shape of a process that works for a remote assistant hire, with the scoring behaviour that belongs at each stage.

StageWhat you do with the scorecard
Before you postWrite the criteria and weights. Three to seven, agreed with anyone else who gets a vote.
ScreeningScore only the criteria an application can evidence. Leave the rest blank rather than guessing.
First interviewSame core questions for everyone. Score immediately afterwards, before the next call starts.
Work sample taskSame task, same time limit, paid. Score submissions with the names hidden if you can.
Final conversationAsk each finalist to walk you through their submission. Update scores where you learn something new.
DecisionRead the weighted totals, check the must-haves, and take a gap under five points as a tie.

Each of those stages has a tool on this site if you want the drafting done for you. The job description generator writes the posting, and the criteria you set here should map onto what it advertises. The interview questions generator gives you the same question set to run with every candidate, which is the structure the research is pointing at. The skills test generator builds the work-sample task and its rubric. The reference check questions generator is what you reach for when the top two are within five points. And once you have decided, the offer letter generator and the candidate rejection email generator close the loop with the people who did not get it, which matters more than most employers think.

Scoring South African candidates specifically

Most of what is on this page applies to any remote hire anywhere. Two criteria are worth building in deliberately when you are hiring from South Africa, and the tool adds both by default.

The first is written communication, which is less about language and more about how much of the job it represents. English is the language of business in South Africa and is spoken at a native or near-native level in professional settings, so you are judging writing against the same bar you would use at home rather than making allowances. In a remote role, that writing is most of the working relationship. Score it from something they actually wrote to you, not from how articulate they were on a call, because those are different skills and only one of them is what you will live with daily.

The second is the home setup, and it is the single most useful filter you have. South Africa has scheduled power cuts, known locally as load-shedding, and the difference between a reliable hire and an unreliable one is usually just whether they have prepared for it. Strong candidates answer this without hesitation and in specifics: their fibre line and its speed, a mobile-data fallback with a data budget, and a power plan such as an inverter, a UPS, or a small generator, along with roughly how many hours it carries them. Candidates who have not solved it tend to answer in reassurances rather than equipment. This is not a reason to hesitate about South African talent; the good candidates solved this years ago, and asking is how you find them.

Time zone overlap is worth scoring only if your role genuinely needs live hours. South Africa sits at UTC+2 with no daylight saving, which gives near-total overlap with the UK and Europe and a solid morning overlap with the US east coast, so for most roles it is a non-issue rather than a criterion. If yours depends on a specific standup or a support window, add it as a criterion and check the actual hours with the time zone overlap calculator before you make it a must-have. And when you get to the offer, the salary calculator shows what a strong local package looks like, which is how hiring directly with no agency in the middle lets you pay well locally and still save up to 80% against a comparable hire at home.

Fairness, records, and candidate data

A scorecard is a selection procedure, and that has consequences worth knowing about even for a one-person hire. Under the US Uniform Guidelines on Employee Selection Procedures, federal enforcement agencies generally treat a selection rate for any race, sex, or ethnic group below four-fifths of the highest group's rate as evidence of adverse impact. Small employers hiring one assistant are not running the statistics that rule contemplates, but the underlying discipline is exactly what a scorecard gives you: consistent criteria applied the same way to everyone, recorded at the time, so that if you ever have to explain a decision you have something better than a memory of liking someone.

If your process includes any formal psychometric or aptitude testing of South African candidates, note that section 8 of the Employment Equity Act sets a real bar: such testing is prohibited unless the instrument has been scientifically shown to be valid and reliable, can be applied fairly to all employees, and is not biased against any employee or group. That is a good reason for most small employers to stick to structured interviews and work-sample tasks, which test the job directly, rather than buying a personality instrument whose validity for your role you cannot actually establish.

On the data itself, keep the scorecard lean. Nothing you type into the tool on this page leaves your browser, but the summary you copy out will end up in an email or a document somewhere, so use reference codes or initials rather than full names and do not paste in CVs, ID numbers, or anything else you would not want forwarded. South Africa's Protection of Personal Information Act requires that records of personal information are not kept longer than necessary for the purpose they were collected for, which in practice means unsuccessful candidates' details should be deleted once the role is filled unless they have agreed to you keeping them on file for future openings. Deciding that in advance is a two-minute conversation. Discovering it afterwards is not.

Candidate scorecard FAQs

What is a hiring scorecard?

A hiring scorecard is a short list of the things that will actually decide a role, each given a weight, which every candidate is then scored against using the same anchored scale. Instead of comparing people on overall impression, you compare them criterion by criterion and let the arithmetic hold your earlier judgment in place. The point is not that the number is objective. The point is that you fixed the standard before you met anyone, so a late charming interview cannot quietly rewrite what the job needed.

How do I compare virtual assistant candidates fairly?

Agree the criteria and their weights before you screen anyone, ask every candidate the same core questions, give them all the same work-sample task, and score one criterion across the whole shortlist at a time rather than finishing one person before starting the next. Score from evidence you can point to, not from recall. The tool on this page does the weighting and ranking, flags any must-have a candidate falls below, and tells you when the top two are too close for the total to be meaningful.

What criteria should I use to evaluate a virtual assistant?

Three to seven criteria is the working range. Fewer and you are not really discriminating between candidates, more and everything is diluted to the point where the weights stop mattering. For almost any remote assistant role, written communication and a real home-office setup with a power and internet backup plan belong on the list. Beyond that, pick the two or three things this specific role lives or dies on, such as reconciliation accuracy for a bookkeeper or tone with an upset customer for a support hire. The scorecard loads a starting set for nine common roles and lets you edit every weight.

Should I weight hiring criteria differently?

Yes, because they genuinely are not equal. If an assistant will spend most of their week in your inbox, judgment about what to escalate matters far more than familiarity with any particular tool, which they can learn in a fortnight. Weighting forces you to say that out loud and in advance. The practical benefit shows up later, when a candidate is strong on something that turns out to be worth a two and average on the thing worth a five, and the weighted total notices what your gut would have missed.

What if two candidates score almost the same?

Treat it as a genuine tie rather than reading precision into the total. A scorecard is a structured judgment, not a measuring instrument, and a three point gap is well inside the noise. Break the tie with new evidence instead: a short paid trial task, a reference check aimed at your least certain criterion, or a conversation where you ask each finalist to walk you through their test submission. The tool flags any result where the top two are within five points so you are not tempted to over-read it.

Does a scorecard actually improve hiring decisions?

The structure is the part that helps. In the largest recent re-analysis of selection research, Sackett and colleagues (2022) found structured interviews to be the strongest common predictor of job performance, ahead of cognitive ability tests, and structure is precisely what a scorecard adds to a process that would otherwise run on impression. There is also good evidence that unstructured interviews actively hurt: Dana, Dawes and Peterson (2013) found interviewers form confident impressions even from deliberately random answers, and that those impressions crowd out the valid information they already had.

Is the candidate scorecard tool free?

Yes, it is free and needs no signup. Everything runs in your browser, nothing you type is sent anywhere or stored, and you can copy the summary into your own hiring file. Use reference codes or initials rather than full names so the scorecard stays safe to pass around your team.

Sources

  • Sackett, P. R., Zhang, C., Berry, C. M., and Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040-2068. Record on PubMed.
  • Dana, J., Dawes, R., and Peterson, N. (2013). Belief in the unstructured interview: The persistence of an illusion. Judgment and Decision Making, 8(5), 512-520. Full text.
  • Uniform Guidelines on Employee Selection Procedures (1978), 29 CFR Part 1607, section 1607.4. Text at Cornell LII.
  • Employment Equity Act 55 of 1998 (South Africa), section 8, psychometric testing. Department of Employment and Labour summary.
  • Protection of Personal Information Act 4 of 2013 (South Africa), section 14, retention and restriction of records. Section text.

Shortlist people worth scoring

Post a role and compare vetted South African professionals directly, with video introductions and no middleman.

SAVA mobile app