Preloader
Others
  • Estimated reading time: 5 Minutes

Every Secondhand Listing Is a Data Entry Problem: How AI Turns Product Photos into Structured Records

Every Secondhand Listing Is a Data Entry Problem: How AI Turns Product Photos into Structured Records

A friend of mine flips power tools on weekends. He buys them in lots at estate sales, tests each one in his garage, and sells them on eBay and Facebook Marketplace. The buying is the fun part. The listing is where everything stalls: fifteen or twenty minutes per item of squinting at a half-peeled label, hunting for the right category, filling in voltage and chuck size and whether a battery is included. Last time I was over there he had a shelf of tested, working drills and maybe a third of them listed.

Watching him do it, I kept thinking the same thing. This is data entry.

The input is a physical object, documented by a few phone photos. The output is a structured record with required fields, enumerated values and a character limit on the title. Multimodal models make automating that workable now, but the distance between a demo that captions a photo and a listing a seller can actually post is bigger than it looks.

What the Record Actually Looks Like

Start with the target, not the photos. A marketplace listing is closer to a database row than a blog post. On eBay, the title tops out at 80 characters, the category comes from a deep tree, and each category carries its own item specifics, some required, many restricted to a fixed list of values. Other marketplaces differ in detail, but the shape is roughly the same:

{
  "title": "DeWalt DCD771 20V Max Cordless Drill Driver, Tool Only",
  "category_id": "...",
  "item_specifics": { "Brand": "DeWalt", "Model": "DCD771", "Voltage": "20 V", "Power Source": "Battery" },
  "condition": "used",
  "condition_notes": "Light scuffing on housing. No battery or charger included."
}

The title is the easy bit to improvise. The item specifics are what filters run against, and they're exactly what casual sellers skip. An empty "Voltage" field can drop a drill out of the results the moment a buyer filters for 20V tools.

Extract Across Photos, Not Per Photo

The first version most people build captions each photo separately and merges the text. It breaks fast. Sellers shoot an item the way they walk around it, not in an order that suits a parser. On that drill, the brand is on the side, the model number is on a sticker under the battery mount, and the voltage might only be printed on the battery. Three separate captions give you three partial guesses that disagree.

Send all the photos in one request, with one instruction: these images show a single item. The model can then reconcile evidence across shots, taking the logo from the first photo and the part number from the third, and it stops treating a close-up of a scuff as a second object. A few things help once that's working:

  • Ask for the source of each value (which photo, which label). Wrong answers get much easier to trace.
  • Trust printed model numbers and barcodes over visual resemblance. Two generations of the same drill can look nearly identical.
  • Allow "unknown." An unbranded ceramic vase is a valid answer. An invented maker is not.

Fetch the Schema Before You Generate

This step separates prototypes from listings that actually post. Ask a vision model to "fill in the item specifics" with no schema in context and it will write attributes that sound plausible and don't exist: a value outside the allowed list, a field that category simply doesn't have, a category path that was never in the tree. The marketplace rejects some of that outright. The rest gets accepted and ignored, which is worse, because nobody notices.

The fix is ordering. Classify the item first, look up that category's field list and allowed values, then generate against that schema with structured output. Required fields get filled or explicitly flagged. Enumerated fields can only take values the marketplace offers.

Listing Tool implementations that work this way, checking the category tree and field list before writing anything, tend to produce output that survives the marketplace's own validation on the first pass. A seller hands over three or four photos and, a minute or so later, get back a title, description, category, item specifics, condition and a suggested price range, all editable before anything goes live. How far a seller trusts that output depends mostly on the next part.

Condition Comes From the Photos, and the Seller Gets the Last Word

Condition is where sellers and models are both tempted to be generous. Asked to describe condition, a model will happily write "excellent pre-owned condition" for a jacket with a pulled seam, because that's what countless listings say. On eBay, that's how you earn "not as described" returns.

So ground condition in the pixels. A mark on the hem of a dress gets named, with its location, near the top of the description where the buyer will actually read it.

Photos only go so far, though. They can't show that a zipper sticks, or that the charger in the shot belongs to a different tool. The seller knows those things. Every generated field has to stay editable, and the review step is the part of the pipeline with the most information, not a formality.

Cut the Background, Never Repaint the Item

Photo cleanup is the tempting place to reach for a generative model. Why not turn a garage-floor photo into a studio shot? Because the buyer is paying for that exact object, scratches included. Any step that redraws the item's pixels, whether it's generative fill or an upscaler inventing texture, turns a record into an illustration.

Segmentation is the right tool. Cut the item out, put it on plain white or light grey (or leave it as a transparent PNG), rotate the shot level, crop it square, fix the exposure. The item's pixels stay the item's pixels. It's an easy rule to break by accident when an image library defaults to generative fill, so it's worth a test that flags any local edits inside the product mask.

Conclusion

Take the reselling context away and this is a general pattern for turning images into records: read all the evidence together, constrain output to the schema of wherever it's going, report what is visible rather than what sounds good, and keep generative steps away from anything the user treats as proof. Marketplace listings just happen to be an unforgiving place to test it, because buyers check.

My friend still lists some tools by hand, mostly out of habit. The shelf in his garage is shorter than it used to be, though.

Frequently Asked Questions

Can a vision model write a usable listing from photos alone?

It can write text that reads well. Without the category's field list and allowed values in context, it will invent attributes and guess at categories, and those listings get rejected or never show up in filtered search.

Should a photo-to-listing pipeline publish automatically?

Keep a review step. Functional faults, missing accessories and some condition details aren't visible in photos, and only the seller knows them.

Related articles
Where to Buy YouTube Like Packages Safely
7 Oct, 2026
  • Estimated reading time: 9 Minutes
How Developers Track AI Search Visibility for Their Product
7 Oct, 2026
  • Estimated reading time: 6 Minutes
How to Optimize Complex 3D Anatomy Models for Browser Performance
7 Oct, 2026
  • Estimated reading time: 5 Minutes
The Instagram Scam That Looked Like a Better Exchange Rate
7 Oct, 2026
  • Estimated reading time: 5 Minutes
Weekly trending
Our Sponsors

Our blog is proudly supported by industry-leading sponsors.