How antivirus detection works, term by term
When security software blocks a file, it has made a judgement using one or more of a small number of techniques. Knowing which technique is involved explains why detection is fast for some threats and uncertain for others, and why mistakes happen in both directions.
Quick answer. Antivirus products recognise known malware by its fingerprint (signatures), flag unfamiliar files that look or act suspicious (heuristics and behaviour monitoring), and check unknown items against the vendor's online records (cloud lookups). Suspicious files are usually quarantined rather than deleted, so that a mistaken detection can be reversed.
Signatures: recognising what is already known
The oldest technique, and still the fastest, is matching files against a list of identifying patterns taken from malware that has already been analysed. A signature may be a hash of an entire file or a short sequence of bytes that appears in every copy of a particular malware family.
Signature matching is precise. When a file matches, the product can name the malware with confidence, and the chance of a mistake is low. Its weakness is equally clear: a file that has never been seen before has no signature. Malware authors exploit this by producing many slightly different versions of the same program, each with a new fingerprint. That is why every product also uses the techniques below, and why signature lists are updated many times a day.
Heuristics: judging by characteristics
Heuristic analysis looks at a file's structure and contents for features that are common in malware and rare in legitimate software. Examples include code that is deliberately scrambled to resist inspection, instructions to disable security tools, and a file pretending to be a document while actually being a program.
Each feature adds to a score. Above a threshold, the file is flagged. Vendors tune the threshold: set it low and more new malware is caught, along with more innocent files; set it high and false alarms drop, but more threats slip through. Some products let you choose a sensitivity level. The default is usually a reasonable balance for home users.
Behaviour monitoring: judging by actions
Rather than examining a file before it runs, behaviour monitoring watches what programs do while they are running. Certain sequences of actions are strong indicators of harm whatever the file looks like: rapidly reading and rewriting large numbers of documents, attempting to disable backups, injecting code into other running programs, or recording keystrokes.
Behaviour monitoring is particularly relevant to ransomware, which can be stopped partway through once its pattern of activity becomes clear. Some products pair it with the ability to roll back the changes made before the process was stopped. Whether a given product does this is a question for its documentation; the feature varies between vendors and editions.
Sandboxing: watching in isolation
A sandbox runs an unknown program in a sealed environment and observes it. If it tries something harmful, it does so harmlessly, and the verdict can be applied to the real file. Sandboxing is resource-heavy, so it tends to be used selectively, for files that are both unknown and suspicious, and in some products it happens on the vendor's servers rather than on your device.
Cloud lookups and reputation
Many products send a file's hash, or a website address, to the vendor's servers and ask whether it has been seen before and with what result. A file seen on millions of machines for years with no problems is very likely safe; a file first seen an hour ago on a handful of machines deserves more caution. This is called reputation.
Cloud lookups let a product respond to new threats within minutes, without waiting for a full update. They also mean that information about the files and sites on your device is sent to the vendor. Reputable vendors describe what is sent and how it is handled in their privacy documentation, which is worth reading if this matters to you. New Zealand's Office of the Privacy Commissioner explains your rights over personal information held by organisations.
"Machine learning" and "AI" detection
Vendors increasingly describe their detection as using machine learning or artificial intelligence. In this context it usually means that the scoring used in heuristic or behaviour analysis has been produced by training a statistical model on very large collections of known-good and known-bad files, rather than hand-written. It is a refinement of the techniques above rather than a separate category, and the label by itself tells you nothing about how well a product performs. Results from independent testing laboratories are a better guide than the terminology.
Real-time and scheduled scanning
All of these techniques can be applied in two modes. Real-time protection checks items as they are downloaded, opened or run, and is where nearly all detections happen in practice. On-demand scans examine files already stored on the device, either when you ask or on a schedule. A full scan is useful after removing a threat, or when you first install a product, to check that nothing was already present.
Quarantine, removal and repair
When a product decides a file is harmful, it usually moves it to quarantine, an isolated area where it cannot run. From there you can delete it permanently or restore it if the detection was a mistake. Where malware has changed system settings, some products also attempt to reverse those changes. Products typically keep quarantined items for a set period before deleting them automatically, so a quarantined file you need should be dealt with promptly.
When detection gets it wrong
Two kinds of error are possible. A false positive flags something harmless; a false negative misses something harmful. Every product makes both kinds, because every technique beyond exact signature matching involves a judgement. Small, specialised or newly released legitimate programs are the ones most often caught by false positives, because they have little reputation and may share features with malware.
If you believe a detection is mistaken
- Leave the file in quarantine for now. Quarantine is reversible. Restoring the file before you have checked it removes the safety margin.
- Note the name of the detection. The product will show a detection name, often a family name and a variant code. Write it down exactly.
- Consider where the file came from. A file downloaded directly from the maker's own website, or installed from an official app store, is a much stronger candidate for a false positive than one from a forum, a file-sharing site or an email.
- Check with the software's maker. Developers of legitimate software that is wrongly flagged often publish a note about it, and can confirm the correct file details.
- Report it to the security vendor. Most vendors have a submission process for suspected false positives, described in their help pages. A confirmed false positive is usually corrected for all users.
- Restore only once you are satisfied. If you do need to restore the file, add an exclusion for that specific file only, not a whole folder or drive, and remove the exclusion when it is no longer needed.
What to watch out for
- Excluding an entire drive or the Downloads folder to stop repeated alerts. This switches off protection where it is most needed.
- Running two real-time antivirus products at the same time. They can conflict with each other; operating systems generally expect one active product.
- Treating a clean scan result as proof that a device is safe. It means only that nothing recognisable was found.
Reading independent test results
Several independent organisations test security products against collections of real malware and publish their results. When reading such a report, check three things: the date of the test, because products change quickly; whether false positive results are reported alongside detection rates, because a product that flags everything would score perfectly on detection alone; and whether the edition tested is the one you are considering. We do not reproduce or summarise test scores on this site, because we cannot keep them current, and an out-of-date score is a misleading one.