How Accurate Is AI at Identifying Rocks?
By The Any Rock Identifier Team · Published 26 June 2026 · Updated 31 August 2026
Short answer: it depends on the specimen, and the confidence score is the part worth reading. In our most recent test — 76 specimens run through the live tool — every identification it labelled "Confident" was correct, all 22 of them. Across all 76, including specimens picked specifically because they are easy to mix up, it was right about two thirds of the time. On the rest it lowered its confidence and offered alternatives rather than guessing.
We did not want to wave our hands at this, so we wrote the numbers down. Below are the results by confidence band, how we tested, and the one claim we published earlier that did not survive a harder test.
The numbers, by confidence band
In July 2026 we sent 76 specimens with known correct answers through anyrockidentifier.com itself — the live tool, not a separate test copy — and scored what came back. Thirty-five were our original blind-test specimens. The other 41 were picked on purpose to be hard: quartz against calcite against selenite, jasper against agate, moonstone against opal against labradorite, granite against diorite against gabbro.
Sorting the results by the confidence score the tool showed:
The pattern is the point. The confidence number is honest — high scores really were reliable, low scores really were shaky. Every wrong answer that still looked confident sat between 78 and 82: a hematite read as magnetite, a green fluorite read as emerald, a variscite read as malachite.
That mattered because our "Confident" badge used to start at 75, so it was being given to a band that was only about 81% right. We raised the badge to 85 and above. Fewer results get called Confident now, and the ones that do have earned it. We also replaced four photos in our own field guide, including quartz and granite, after the test showed they were unrepresentative enough to fool the tool.
One honest limit: 76 specimens is a real test, not a lab study, and 22 out of 22 from 22 tries is a clean result rather than a promise. The model also samples, so runs vary — we ran the same 35 specimens twice an hour apart and got slightly different mistakes each time. Read the bands as the signal and any single result as an anecdote.
- 85–100 ("Confident"): 22 specimens, 22 correct.
- 75–84: 21 specimens, 17 correct — about 81%.
- 65–74: 13 specimens, 6 correct — about 46%.
- 40–64: 20 specimens, 7 correct — about 35%.
What we published first, and why we withdrew it
Our first write-up, in June 2026, came from a blind test of 35 specimens. It reported 82.9% strictly correct on the exact mineral or rock name, rising to 88.6% when we allowed the right mineral family or an accepted synonym — calling massive purple quartz "amethyst" rather than "quartz," for example. On common everyday specimens it scored about 91.7%, which we summarised as roughly nine in ten. We also reported that it never made a confidently wrong call.
Both of those claims are now withdrawn, for two separate reasons.
First, that June run went through a separate test harness rather than the live site, so it never measured what a visitor actually gets. Its numbers are not comparable to the July ones and we no longer quote them as current.
Second, the harder July set broke both claims outright. Confidently wrong calls did happen once look-alikes were in the mix. And on the common, everyday specimens — the ones "nine in ten" was about — the re-run scored 9 out of 12, which does not support the claim. So we stopped making it.
We would rather publish this than quietly leave the old number up. The numbers in the section above are the ones we stand behind.
Want an honest, confidence-scored ID of your own specimen? Try the rock identifier
How we tested it
Accuracy claims are cheap, so we built the test to be hard to game. One step assembled 35 photos of real specimens from authoritative sources, saved each under a scrambled, meaningless filename, and sealed the list of correct answers away. A separate step identified every photo cold — no peeking at the answer key — and wrote those predictions to disk. Only then was the key opened and the predictions scored against it. The model never saw what it was supposed to say while it was saying it.
We did not stack the deck with easy wins, either. The set ran from iconic, beginner-friendly specimens down to deliberately nasty cases: minerals that look almost identical to one another, and a handful of dim, cluttered, real-world photos. We also recorded a confidence number for each call, so we could check not just whether the answer was right, but whether the model knew how sure it should have been.
What the numbers actually mean
A single accuracy percentage hides the useful part. Three things do most of the work:
- The badge is the number that matters. "Confident" is not decoration — it marks the band where the tool was 22 for 22. "Most likely" marks the band where it was right about four times in five. Treat them as different kinds of answer, because they are.
- Where the misses cluster. Near-twins and poor photos, almost every time: the mineral pairs that need a streak plate or a hardness test to separate, and shots that are blurry, dim, or full of other rocks. That is where a careful human slows down too.
- Strict vs. lenient naming. Some "wrong" answers are naming, not blindness — calling a stone "quartz" when the precise answer is "amethyst" reads as an error in a strict count. Our June test measured that gap at about six points. We have not re-measured it on the larger set, so treat it as a rough sense of scale rather than a current figure.
- A small sample. Seventy-six specimens is honest and directional, not a guarantee stamped on every future photo. A clear shot of a common stone is the easy case; one gray rock among many is the hard one.
Why calibration beats raw accuracy
Here is the uncomfortable truth about identification tools: being right most of the time is not enough. What ruins trust is being wrong while sounding certain. A tool that says "this is malachite, 98%" about a dyed howlite has done real damage — someone overpays, or mislabels a piece, or stops looking. A tool that says "possibly malachite, but I am not certain — also consider chrysocolla, and check the hardness" has done its job even when it does not nail the name.
That property has a name: calibration. A well-calibrated model's confidence tracks reality — high confidence is almost always right, and low confidence is the model honestly flagging a coin-flip. In our test the calibration was strong. Every error landed in the low-confidence band; the high-confidence calls were right across the board. The model knew when it did not know, which is the single hardest and most valuable thing a vision model can do here.
So we built the rock identifier around that instead of hiding it. You get a real confidence score, not a fake one. When the model is unsure, it shows you the runner-up candidates rather than forcing a single guess. And every result points you to a simple at-home test to confirm it yourself. We would rather tell you "I think it is X, here is how to be sure" than impress you with a number we cannot stand behind.
Where AI struggles — and how to still get a reliable ID
It is worth understanding why this is genuinely hard, because it explains both the misses and the fix. Geologists do not identify minerals from looks alone. They scratch them to gauge hardness, drag them across a plate to see the streak (the powder color), weigh them for density, drip acid on them, test them with a magnet. A photo throws all of that away. You are left with color, shape, and luster — and plenty of different minerals share those. Some look-alikes are close to impossible to separate from a picture, no matter how good the model is. This is just as true for a crystal identifier as for any rock — two clear, glassy crystals can be different minerals entirely.
Two situations cause most of the trouble. The first is genuine twins: labradorite, for instance, is unmistakable when its blue-green flash catches the light, but photographed flat with no flash showing it is just a gray feldspar that any system will hedge on. See labradorite for what that flash should look like. The second is bad inputs — dim light, a busy background, no sense of scale, a thumb in the frame. Garbage in, hedged answer out, which is the honest behavior but not the satisfying one.
The good news is you can do a lot to swing the odds. Better photos alone close most of the gap, and a single physical test usually settles whatever the photo could not.
When you do this, the AI stops being a final verdict and becomes what it is good at: a fast, well-read first opinion that narrows a mystery rock down to a short, testable shortlist. Pair it with a five-minute test and your real-world accuracy climbs well past any single photo on its own.
- Shoot in bright, even daylight — near a window beats overhead light. Avoid harsh glare and deep shadow.
- Fill the frame with the specimen against a plain background, and keep it in sharp focus.
- Include a few angles, and wet the stone or show a freshly broken surface so the true color and luster come through.
- Confirm with a streak test — drag the specimen across unglazed porcelain; the powder color cuts through surface staining and separates a lot of look-alikes.
- Check hardness against the Mohs hardness scale — whether a fingernail, coin, knife, or glass scratches it (or it scratches them) is one of the most decisive tests there is. It is how you tell heavy, brassy pyrite from soft real gold, or hard quartz from soft calcite.
What we will not identify
One deliberate limit, because it is part of being honest about accuracy. We identify rocks, crystals, minerals, gemstones, and fossils — and nothing else. We do not identify mushrooms, plants, berries, or wildlife. A misidentified rock costs you a label. A misidentified mushroom or snake can cost someone their life, and no confidence score is good enough to take that risk. Staying in the safe lane is a feature, not a gap.
Frequently asked questions
How accurate is AI at identifying rocks?
In our most recent test — 76 specimens sent through the live tool, 41 of them chosen because they are easy to confuse — every identification labelled "Confident" was correct, 22 out of 22. Across all 76 it was right about two thirds of the time, and it lowered its confidence and offered alternatives on the rest. Accuracy drops on look-alike minerals and on poorly-lit or cluttered photos, which is why a confirming physical test still matters. An earlier figure we published — about 91.7% on common specimens, from a smaller June test that did not run against the live site — has been withdrawn.
Can AI tell when it is unsure, or does it just guess?
A well-built one tells you, and how well it does that is measurable. In our original 35-specimen test every mistake happened on a specimen the model had already flagged with low confidence. Our July re-test on a harder 76-specimen set found that this no longer held at the threshold we were using: a few wrong answers were arriving with a "Confident" badge, all of them scoring between 78 and 82. Everything at 85 or above was correct, so we raised the badge to 85. That property is called calibration, and it matters more than raw accuracy — a tool that honestly says "I am not sure" is safer than one that is wrong while sounding certain.
Why can't AI identify a mineral as well as an expert?
Because experts do not rely on looks alone. They test hardness, streak (powder color), density, acid reaction, and magnetism — none of which a photo captures. Color, shape, and luster are often shared by several different minerals, so some look-alikes are nearly impossible to separate from a picture. AI is a strong first opinion, but a quick streak or hardness test is what confirms it.
How can I make AI rock identification more accurate?
Photograph the specimen in bright, even daylight against a plain background, in sharp focus, from a few angles, and wet it or show a fresh broken surface so the true color shows. Then confirm the AI's top guess with a streak test or a Mohs hardness scratch test. Better photos plus one physical test push real-world accuracy well past any single photo on its own.
Will AI identify mushrooms, plants, or other things too?
We deliberately do not. We identify only rocks, crystals, minerals, gemstones, and fossils. Misidentifying a mushroom, plant, berry, or animal can be dangerous or fatal, and no confidence score justifies that risk. Staying strictly in the rock-and-mineral lane is an intentional safety decision.
Got a rock or crystal to identify?
Snap a photo and get an instant identification with an honest confidence score — free to start.
Identify yours freeMentioned in this article
Keep reading
Educational content — confirm important identifications with the diagnostic tests described or a qualified expert before relying on them.