Updated October 6, 2026
Every card scanner in this category advertises a percentage. Almost none of them say what was measured, on how many cards, or what counts as a pass — which makes the number unfalsifiable and therefore worth nothing.
So here are ours. Same corpus every time, checked into the repo, run on every change. The figures on this page are injected from the benchmark result files at build time, so this page cannot quietly drift away from what the shipped app actually scores.
Three different things get called "accuracy", and they are not interchangeable. We measure them separately.
The question for bulk capture: you photograph a binder page or a laid-out grid, and we have to pull out every card in the frame.
Corpus: 50 photos, 373 cards. Every photo's true card count is in
its filename, so the scoring is not a judgement call. Model 25222858.
| Cards present | Cards found | ||
|---|---|---|---|
| Flat layouts — binder pages, laid-out grids, shoebox spreads | 313 | 311 | 99% |
| Overlapping layouts — fanned, piled, stacked | 60 | 33 | 55% |
| All photos | 373 | 344 |
40 of the 50 photos return exactly the right count.
The split is the whole story, and it is more useful than the blended number. Lay cards flat and separated and detection is essentially solved: the only flat photos that lose anything are bright_5 (4 of 5), tilted_7 (6 of 7). Fan or pile them and it falls apart — a card that is 60% hidden behind another card has no edges to find.
That is a real limit, not a tuning problem, and it is why the app asks you to lay cards out rather than promising to read a stack.
| Benchmark photo | Cards present | Cards found |
|---|---|---|
overlap_heavy_8 |
8 | 1 |
pile_15 |
15 | 8 |
fanned_10 |
10 | 5 |
test_pile_9 |
9 | 5 |
bright_5 |
5 | 4 |
overlap_light_6 |
6 | 5 |
overlap_two_2 |
2 | 1 |
test_fan_7 |
7 | 6 |
test_overlap_3 |
3 | 2 |
tilted_7 |
7 | 6 |
Finding a card is not the same as cutting it out cleanly. A crop that is tilted, trapezoidal or carrying the table around the edges still counts as "detected", and it still ruins the identification downstream.
We score crops against hand-clicked corners — a human clicked the four corners of the real
card — using IoU, with 0.7 as the pass floor. Model d0a3bca7.
| Section | Cards | Passed | Median IoU |
|---|---|---|---|
| Close-up capture (hand-labelled) | 39 | 39 | 0.8952 |
| Tilted phone angles, holdout — never trained on | 14 | 14 | 0.8586 |
Only hand-labelled sections appear here. Our full crop corpus is larger, but part of it was labelled by an earlier model — those sections are useful as a regression check and would be dishonest as an accuracy claim, so they are not on this page and there is no combined total.
The holdout row is the one that matters. It is tilted, hand-held, real-angle photography that the model never saw during training, which is the only condition under which a number means anything.
The one that actually matters to you. Detection and crop are both just proxies for this.
A card counts as correct only if the identity we stored matches the card in the photo, scored by comparing the crop against the official art for the card we claim it is. Not "we returned something" — "we returned the right thing".
114 of 192 cards: 59%.
| Cards | Correct | ||
|---|---|---|---|
| Page capture | 152 | 94 | 62% |
| Single-card capture | 40 | 20 | 50% |
Those are different corpora, not the same cards shot both ways, so this is not evidence that page capture beats single capture. What it does show is that bulk capture is not the accuracy compromise people assume it is.
Of the 120 cards the pipeline was confident about, 15 were the wrong card — 12%.
That is the failure mode that actually costs you something. A miss is visible: the card isn't there, you scan it again. A confident wrong answer is invisible — it files Dak Prescott into your collection and you find out two years later when you go to sell.
It is also why bulk imports land in a review queue instead of going straight into your collection. We would rather show you 120 cards to glance at than quietly file 15 wrong ones.
Two reasons, and neither is modesty.
A percentage with no failure list is not a measurement, it's a marketing asset. If we only showed you 99% we would be doing exactly what we are complaining about.
And the failures are the useful part. "55% on overlapping layouts" tells you to lay your cards flat, which takes ten seconds and roughly doubles what you get back. A number that only goes up tells you nothing you can act on.
Four questions. They work on us too.
It depends entirely on which stage you mean. On our own fixed benchmark of 373 cards across 50 photos, detection finds 99% of cards in flat layouts and 55% in overlapping ones, and end-to-end identification — the stored card actually being the card photographed — is 59%. Most published accuracy claims in this category are detection-stage numbers on clean cards, which is the most flattering of the three.
114 of 192 cards end-to-end (59%) on our benchmark, split 62% for page capture and 50% for single-card capture. Detection alone is 311 of 313 cards in flat layouts. All of these come from benchmark files in our repository that are re-run on every change.
Overlap, almost exclusively. In our 373-card benchmark, flat layouts returned 311 of 313 cards while fanned, piled and overlapping layouts returned only 33 of 60. A card that is mostly covered by another card has no visible edges to detect. Laying cards out flat and separated is the single biggest thing you control.
Yes, and it is the failure that matters. On our benchmark the pipeline was confident about 120 cards and 15 of those were wrong — 12%. A missed card is obvious and you rescan it; a confidently wrong card gets filed and never re-checked. This is why bulk imports go through a review queue rather than straight into your collection.
There isn't a single one, because "accuracy" is three different measurements. Ask instead what was measured, on how many cards, what counted as correct, and where it failed. Any claim that cannot answer those four questions is unverifiable, however high the percentage is.
We run the shipped pipeline over a fixed, checked-in corpus on every change and compare against stored baselines. Detection scores found-vs-present card counts from 50 photos whose true counts are in their filenames. Crop scores the cut card against hand-clicked corners by IoU, with a 0.7 pass floor. End-to-end scores the stored identity against the official art for the card we claimed. Same corpus every time, so results are comparable across changes rather than across press releases.
Try SnapMyCards — it runs in your browser