I have a black leather sectional in my living room. I am sitting on it right now.
It is an IKEA FINNALA with a chaise, and I know that with certainty because I bought it. The brand is not visible in any photograph of it. There is no tag in frame, no logo, nothing.
So I took three photos and asked five AI assistants the same question: what is this worth?
They said $250, $275, $450, $450, and $550.
The setup
Same three photos to all five. One wide shot, one alternate angle, one close-up of the arm. Identical prompt asking for a price, a range, a condition, dimensions, weight, and what the item is.
First answers only. No follow-ups, no corrections, no telling any of them what the others said.
The five: ChatGPT, Claude, Gemini, Grok, and ClearList, which is the tool I build. I will get to how mine did, and it is not a clean win.
What we can actually check: it is a FINNALA in Grann/Bomstad black, meaning real grain leather on the contact surfaces and coated fabric elsewhere. The chaise lid lifts and there is storage underneath. It measures 98 inches wide. IKEA lists the three-seat at $1,299 today, with the chaise section priced on top. The left arm has cat scratches.
It has not sold, so there is no "correct" price. That turns out not to matter, because the interesting failures were not about money.
Round one: the numbers
| Price | Range | Width guess | Model | |
|---|---|---|---|---|
| Gemini | $250 | $150–350 | 110–120" | KIVIK |
| Claude | $275 | $200–375 | 96" | none |
| ChatGPT | $450 | $350–550 | 110–115" | "VIMLE/FINNALA-style" |
| Grok | $450 | $350–600 | 105–115" | none |
| ClearList | $550 | $450–650 | not shown | told by me |
The spread is 2.2x from bottom to top. It is also not evenly scattered. Two clustered near $250 and two landed on exactly $450 independently, which is a strange thing to see from different models looking at the same pictures.
Gemini's ceiling of $350 is the floor of the other two. If you had asked only Gemini and only ChatGPT, you would have received two answers with almost no overlap.
The dimension problem
This is the finding I would actually act on.
The sofa is 98 inches wide. Three of the four vision guesses came in between 105 and 120 inches. Gemini's high end was off by 22 inches, which is nearly two feet.
Photos give a model very little to work with on scale. There is no reference object, and rooms distort. So the number a buyer most needs, the one that determines whether the thing fits through their door and into their vehicle, is the number these tools are worst at.
Check your dimensions with a tape measure. It takes thirty seconds and no model can do it for you.
Round two: I told them the model
Then I went back to each one and said: that is a FINNALA.
What happened next was the most interesting part of the whole exercise.
Grok transformed. Given only the name, it came back with 99¼ inches wide against my measured 98. It gave 38⅝ inches deep and 33½ high, both matching IKEA's published spec exactly. It named the Grann/Bomstad leather. And it mentioned the storage under the chaise.
It knew all of that the entire time. What it could not do was look at the photograph and connect it to what it knew.
That reframes the whole test. This is not a knowledge problem. It is a seeing problem.
Gemini announced a correction it did not make. It thanked me for the catch, said the footprint was "slightly more compact than my initial guess," and then gave 111 inches. Its initial guess had been 110 to 120. That is inside the original range, not more compact, and still 13 inches over the real number. It kept its $250.
ChatGPT held at $450 and confirmed the model without much change. It had been closest to right in round one anyway.
The pet hair that was not pet hair
Gemini's original answer said there was "lingering pet hair on the armrests," citing the specific photo file. It then used that in its reasoning: the need for "a deep cleaning to clear away the pet hair" put the sofa in the lower-mid tier.
There is no pet hair. There are cat scratches.
I want to be fair about this, because it is more interesting than a simple mistake. Gemini correctly inferred that a cat had been involved. What it got wrong was the kind of damage, and the kind of damage is the whole point. Hair comes off with a lint roller. Scratches in leather do not come off at all.
So it recommended a remedy that would not have worked, and then discounted the price for a problem it had misdiagnosed. It was directionally right and practically wrong, which is a failure mode worth knowing about.
Nobody found the storage
Zero out of five spotted that the chaise opens.
Not from the photos, which is fair enough, though the lid seam is visible if you know to look. But after I supplied the model name, only Grok mentioned it. My own tool, knowing it was a FINNALA, still did not.
That is a genuine feature and a real selling point. Five AI systems looked at this sofa and none of them told me to mention it in the listing.
How mine did, honestly
ClearList priced it highest at $550, with a range of $450 to $650 and a retail anchor of $1,500 to $1,800 new, which is about right for this configuration.

Once it had the model, its physical specs were the best of the group: 99 by 65 by 34 inches and 230 pounds, pulled from manufacturer data. It was also the only one that produced structured logistics as actual fields rather than prose, flagging that the item needs a truck and a second person.
Now the parts I am less pleased about.
It did not identify the sofa. Its own description says "seller states this is an IKEA FINNALA." I told it. The other four had to work from photos, and it did not. The good specs came after the name, same as Grok.
It missed the storage too, even holding the model name.
And it undersold the damage. It described "normal creasing and light wear from use." Gemini, ChatGPT, Grok and Claude all flagged the arm scratching in some form. Mine called it light wear.
That last one bothers me most, because it errs in the direction that costs a seller a real sale. A buyer who drives across town expecting light wear and finds cat scratches does not buy the sofa. They leave, and you have burned an afternoon and a lead.
Being the highest price and the softest description at the same time is not a good combination. That is on my roadmap now, and it is the useful kind of result to get from a test you ran on yourself.
What I would tell you to do
Use these tools. They are genuinely fast and every one of them produced an answer I could work from in under a minute.
Then check three things yourself, because they are things you can verify and the models cannot.
Measure it. This is the big one. Get the width and the depth with a tape.
Look at the damage properly and name it accurately. Not "light wear." Say scratches if it is scratches. Honest descriptions bring buyers who show up.
List the features the AI did not notice. If your sofa has storage under the chaise, that belongs in the listing, and apparently you will have to be the one to put it there.
AI can be accurate, and mine gets things wrong too. That is not a reason to skip it. It is a reason to spend the thirty seconds it saved you on checking its work.
I am still sitting on this couch, by the way. Five different opinions on what it is worth and not one buyer yet, which is its own kind of answer.
Related reading: how to price anything in your house in under a minute and how much should I sell my couch for.