Width.ai

AI in Fashion Retail: The 2026 Technical Guide to Models, Pipelines, and Real Production Numbers

Matt Payne
·
October 6, 2026

The demo always “works”.

Somebody points a phone at a shelf and the products come back named. Somebody uploads a photo of a dress and the catalog returns three near-identical ones. It is a good meeting, and it is usually where the budget gets approved.

Then it meets the real catalog. This is the most common way an AI in fashion retail project goes wrong, and it rarely announces itself. The products it names confidently are the ones that look nothing like each other. The near-identical dresses turn out to be near-identical in color and in nothing else. Whatever the demo was measuring, it was not this.

The gap happens due to real complexity in the fashion space at scale. The use cases are fairly easy to get to something usable, but very complex to get to production level accuracy. Most of the customers coming to us have built something, but its not working at a good enough level to roll out. We see it all the time:

In a use case focused on recognizing matching products across a catalog the off the shelf model only got the correct product 41% of the time. After switching to a “fashion” focused model the accuracy was still only 50%. We customized the approach deeply to get to 89%. The issues were all related to real complexities in fashion product records.

So this is written for the person who has to build an ai system for fashion retail. Four applications get a full breakdown: what the model actually is, what the evaluation number actually measured, and what breaks in production. Where we have shipped the system, the numbers are ours and they are stated with the dataset they came from.

  
    ✂️ Definition    

What is AI in fashion retail?

    

AI in fashion retail refers to machine learning systems applied across the apparel commerce funnel: visual search and product similarity, shelf and SKU recognition, automated product content for search and answer engines, virtual try-on, demand forecasting, and personalization. The technically demanding applications are the vision ones, because fashion products differ along attributes that general-purpose models handle poorly: silhouette, fabric, neckline, sleeve length, and pattern. General image embedding models perform substantially worse on apparel than models fine-tuned on fashion data, which is why production systems in this category are almost always domain-tuned rather than used off the shelf.

  

‍

Key Focus Areas

  • Fashion products differ with specific features general vision models were never trained to separate, which is why off-the-shelf embeddings underperform badly on apparel catalogs.
  • On retail shelf SKU matching, base CLIP reached 41% top-1 and a fashion-domain tune reached 50%. A task-tuned model reached 89% on the same evaluation.
  • Product similarity at catalog scale is a retrieval problem, not a classification problem, and it is evaluated differently.
  • Virtual try-on stopped being a feature you build. Google moved it into Search in April 2026, which makes your product imagery the input to someone else's model.
  • The build decision is rarely build versus buy. It is off-the-shelf versus fine-tune, and catalog size plus attribute specificity decides it.

‍

Diagram of the AI stack in fashion retail showing discovery, search, conversion and operations layers with the model class and evaluation metric for each

‍

The State of AI in Fashion Retail in 2026

Fashion retail is past the question of whether AI belongs in the stack. Large catalogs already run it in search, recommendations, sizing, product content and store operations, and the budget conversation has moved from whether to fund a pilot to which systems get owned internally. 

‍

Two business changes over the last eighteen months matter more than the capabilities themself, because they change what a retailer is actually deciding when a project gets scoped. 

‍

Google retired its standalone Doppl try-on app in April 2026 and moved the capability directly into Search and Shopping, where a shopper taps a product listing, taps “Try It On”, and sees the garment rendered on themselves. Since December 2025 it works from a single selfie. Zalando, Zara and L'Agence have run try-on campaigns on it. Macy's launched a Gemini-powered stylist in March 2026 and reports that shoppers who engage with it spend several times more per visit than those who do not.

‍

This is more interesting than any vendor announcement of a new model. They changed how the ai tech is used to be more of an integrated feature than standalone. Lots of products are moving towards a focus of being ai native, instead of standalone.  A retailer evaluating whether to build try-on ai in 2024 was deciding whether to add a new feature to their product page. A retailer evaluating it in 2026 is deciding how to integrate ai deeper into the shopping experience, and own the data. 

‍

The second change is less visible and it affects a few major applications: catalog turnover got faster while catalog data got thinner. Drops arrive more often, from more suppliers, with less accompanying information than a decade ago, which means the volume of product data work per season went up at the same time as the tolerance for slow onboarding went down. That is the pressure behind most fashion AI budgets, whatever the deck says. 

‍

Neither change addresses the ai project failure rate, which is the part worth planning around. Fashion AI pilots do not usually fail because the technology does not work. They fail because the system that was bought was evaluated on curated examples and deployed against a catalog that looks nothing like them, and nobody established up front what accuracy the merchandising team would actually accept. 

‍

This produces a specific and avoidable sequence: a successful proof of concept on a few hundred hand-picked products, a rollout that exposes the long tail, a quiet drop in the accuracy nobody agreed to measure, and a merchandising team that stops trusting the output and goes back to doing it by hand while the license keeps renewing.

‍

Visual Search and Product Similarity

Drive more conversions, visual search for customers

‍

A shopper sees a jacket on someone in a coffee shop, or in a video that scrolled past, and wants to find it. What they have is a photograph and no vocabulary. No logo is visible, they do not know the brand, and "olive green jacket with a collar" returns four thousand products. Shoppers are not patient with this: 47% give up after a single failed search and only 23% try three or more. An image is the shortest path from wanting the thing to finding it.

What Shoppers Actually Do With an Image

Three different requests hide behind one feature, and they are not the same system.

Find this exact product. The shopper has a photo of a specific item and wants that item, in their size, in stock. There is a right answer and the system either returns it or fails.

Find things like this. No specific product in mind, just an aesthetic they are chasing. This is product similarity, and it is one of the three jobs rather than the whole feature. There is no single right answer, which means it is judged on the quality of a ranking instead of a hit or a miss.

Find the product inside this photo. The image is an outfit, a room, a street. Pinterest found that shoppers usually care about one object in a busy picture, which means the system has to detect and crop the right object before any search runs. A failure at that stage means the other two never get their chance.

The same machinery runs internally, pointed the other way. A marketplace where forty sellers listed the same product with forty different photographs needs near-duplicate detection across its own catalog, which is the exact-match job with no shopper attached to it.

‍

Why This Stopped Being a Novelty

More product discovery now happens on surfaces that carry no product text at all. An outfit in a short-form video has no title, no attributes and no SKU. The only thing a shopper can carry from that surface into your store is a screenshot, and if your store cannot read one, the trip ends at your search bar.

What it moves is measurable on the same dashboard you already run. Conversion rate and click-through improve because the shopper reaches the product with less friction. Basket size follows, because a shopper who found the thing they were picturing keeps shopping instead of abandoning. Bounce rate is the early signal, since it tracks how well the results matched what was actually in the picture. There is a merchandising payoff too: the stream of query images is demand data, and what people photograph is a read on what to buy next season.

Why Apparel Is the Hard Case

The architecture follows from the job. A classifier maps an image to one of N known labels and needs retraining whenever the catalog changes, which a seasonal catalog does constantly. Retrieval embeds every product image into a vector space once and answers a query by finding nearest neighbors, so new products get indexed without retraining. That is the only version that survives a catalog that turns over.

Which makes the model's quality a property of the space it produces. Two photographs of the same jacket under different lighting should land near each other; a similar jacket in a different fabric should land further away than one in the same fabric. General models are bad at this on clothing specifically. CLIP and its descendants learned from broad web image-text pairs, which separates a jacket from a lamp extremely well and a bomber from a harrington poorly, because neckline, sleeve construction, closure type and fabric weight are treated as noise by broad pretraining.

The failure modes are recognizable once you have seen them. Color dominance, where everything returned is the right shade and the wrong silhouette. Background dominance, where an outdoor photo retrieves other outdoor photos. Brand-logo bias, where a visible logo drags retrieval toward that brand regardless of the garment. And the asymmetry that matters most in production: a clean seller photograph on white versus a customer photograph taken at an angle in bad light, which are precisely the two things the system is being asked to match.

What Fine-Tuning Changes

Domain specific fine-tuning closes the gap. Fashion CLIP is CLIP further trained on fashion imagery, and on apparel it separates garment types that base CLIP collapses together. We wrote up where it improves on base CLIP and where it still falls short, because the second half is what decides whether you need a custom model.

Use case specific tuning closes the rest. On a product similarity pipeline evaluated against a catalog of more than 10 million products, our fine-tuned pipeline reached 92.44% top-1 and 99.3% top-3 accuracy. Top-1 is the strict measure, where the correct product is the single first result. Top-3 matters commercially because visual search interfaces show a row of results, and a correct match in position three still converts. That same approach indexed 3.2 million unique SKUs from one example image per SKU, which is the realistic condition for a catalog: you have a product photograph, not a labelled training set per item.

Two Questions Before You Buy

The index is the other half of the system and it is where scale bites. Embeddings have to be stored and sharded, queries have to return in a couple of hundred milliseconds rather than a couple of seconds, and sold-through styles have to leave the index or the system will confidently recommend things nobody can buy. That last one gets underestimated constantly, and designing for it upfront is far cheaper than retrofitting it.

Then ask the vendor two things. What was the evaluation catalog size, because top-1 accuracy against a thousand products and against ten million are not comparable numbers. And were both sides of the evaluation clean catalog images, or were the queries real customer photographs, because that difference is worth tens of points.

‍

‍

  
    

Building Visual Search on a Real Catalog?

    

Width.ai builds the retrieval systems behind exact product matching, similarity and scene-to-product search. Our fine-tuned pipeline reached 92.44% top-1 and 99.3% top-3 against a catalog of more than 10 million products, indexed from a single example image per SKU.

    

Tell us your catalog size, where your query images come from, and the accuracy your merchandising team needs to sign off on.

    Talk to us about visual search →  

Shelf and SKU Recognition

A field rep photographs a shelf and the system returns every product on it, matched to the exact SKU. It is used for planogram compliance, out-of-stock detection and competitive audits, and it is the hardest of the four problems here because the products are adjacent, partially occluded, photographed at an angle, and differ from each other by a flavor variant or a pack size.

What Three Models Score on the Same Shelf Photographs

The evaluation is RP2K, a public dataset of more than 500,000 shelf photographs taken in real stores rather than staged for a catalog. That distinction carries most of the weight, because accuracy measured on clean product imagery does not survive contact with a phone photograph taken at an angle under fluorescent light.

Three models, the same photographs, top-1 accuracy. Base CLIP, a general-purpose image model, reached 41%. Fashion CLIP, which is CLIP further trained on fashion imagery, reached 50%. A model we trained on this specific job reached 89%.

The middle number is the one worth sitting with. A model tuned on the right domain gained nine points over the general one and was still unusable, because it had been tuned for a different task inside that domain. Fashion CLIP learned to tell garments apart. Shelf recognition asks a model to tell packages apart. Domain tuning and task tuning are separate interventions with separate costs, and they get talked about as though they were one thing.

Which means a vendor demonstrating a fashion-tuned model against a shelf recognition problem is showing you 50%. Whether they can reach 89% is a question about their training data and their evaluation discipline, not about which base model they started from.

Why a Fashion-Tuned Model Stalls at 50%

Fine-grained distinction is the whole problem. Two SKUs from the same product line in different pack sizes carry nearly identical packaging separated by one number. The model has to be sensitive to that difference and at the same time invariant to lighting, angle, partial occlusion by a shelf edge, and motion blur from a phone camera.

Those two requirements pull against each other, which is why generic models plateau. Getting past the plateau means training on the actual failure cases: photographs at the angles your reps take, with the occlusion patterns your shelves produce, on the SKUs whose packaging differs least.

Detection and Recognition Are Two Different Models

A shelf photograph is not one prediction, it is dozens. The system has to first find each product instance in the image, then identify which SKU each one is. Those are separate problems with separate failure modes, and conflating them is a common source of disappointing results.

Detection struggles with products that are flush against each other, where the boundary between two identical-looking packages is ambiguous even to a person, and with items at the edges of the frame or behind a price rail. Recognition struggles with the fine-grained distinctions described above. A system can have excellent detection and poor recognition, which shows you the right number of products with the wrong names, and that failure looks very different in a report than low detection with good recognition, which silently undercounts.

Worth asking any vendor which of the two their number describes. An end-to-end accuracy figure that does not decompose into detection and recognition is hiding which half is weak.

What This Is Actually Used For

Three jobs justify the spend. Planogram compliance, meaning whether the shelf matches the plan the brand paid for, which is a real dispute between brands and retailers with money attached. Out-of-stock detection, where the value is the speed of the alert rather than the precision of the count. And competitive audit, where a rep walks a competitor aisle and the system captures assortment and facings.

Those three tolerate different error rates, which is the useful thing to notice when setting a target. Out-of-stock detection can live with a few percent of false positives because the cost of checking is low. Planogram compliance backing a commercial claim cannot, because the number is going into a conversation with a trading partner.

‍

Product Content for Search and Answer Engines

Fashion catalogs turn over faster than any other retail category and arrive with the thinnest data. A seasonal drop lands as a vendor spreadsheet with a style code, a colorway, a price and two sentences written by someone who has never seen the garment. That has to become a product page that ranks, converts, and can be cited by an answer engine.

What the Demo Shows and What Production Requires

The demo takes a product name and returns a description. It is fluent and it is frequently wrong, because a language model given a sparse record has nothing to work from except what it can infer from the product name, and inference on apparel attributes produces confident fabrication: the wrong fibre content, a neckline that does not exist, a fit description contradicted by the size chart.

Production requires the generation step to run second. First the system retrieves real data about that specific product from sources you nominate, which for fashion usually means the brand's own product page or the line sheet the vendor sent. Then it confirms the retrieved data describes that exact style and colorway, which matters more in apparel than anywhere else because a style code with six colorways produces six nearly identical records. Only then does generation run, against your attribute schema and your content rules.

Categorization Is the Part That Decides Discoverability

Assigning each product to the right node of your taxonomy determines which category pages it appears on, which filters surface it, and what your channel feeds submit. Fashion taxonomies run deep, and the bottom levels are where accuracy collapses: separating a shift dress from a shirt dress is a harder call than separating dresses from footwear.

Production systems we have built reach 97% accuracy on a five level deep taxonomy for an ecommerce solutions company, 92% at the lowest level of a 5,585 category multilingual taxonomy for a wholesale marketplace, and 97.62% at 50 million products per month for a growing marketplace. The floor is contractual at 90% on the lowest level.

The reason those numbers are quoted per-taxonomy rather than as a single figure is that depth and language change the problem. A top-level accuracy number on a shallow tree tells you very little about what happens at the leaf where filters actually operate.

Writing for Answer Engines Is the Same Work, Measured Differently

The requirements for being cited by an answer engine and for ranking in search have converged more than the separate vocabulary suggests. Both reward the same thing: specific, attributable, structured facts about the product. A description that states fibre content, measured length, fit relative to standard sizing and care requirements gives a model something to quote. A description built from adjectives gives it nothing.

Where they diverge is in what gets rewarded at the margin. Search still rewards the page as a whole, so internal linking and category structure matter. Answer engines reward the passage, which means a product page that answers a specific question in a self-contained paragraph is more citable than one where the answer has to be assembled from four places on the page.

Where Generated Content Goes Wrong in Fashion Specifically

Two failure modes recur. The first is colorway confusion, where a system retrieves the correct style but the wrong color and writes a description that contradicts the product image. The second is fit language, which is the most commercially dangerous thing to fabricate, because a description that says a garment runs true to size when it does not converts the sale and then produces the return.

Both are avoidable with the same discipline: validate that retrieved data matches the exact variant before generation, and treat fit and sizing as fields that may only be populated from a source, never inferred.

Virtual Try-On Is Now a Distribution Channel

The strategic position of try-on changed in 2026 and most planning has not caught up.

When Google moved try-on into Search and Shopping, the render started happening before the shopper reaches your site. Your product images and attributes became the input to a model you do not control, on a surface you do not own, producing an image of your garment that you never approved. The build question changed from whether to add try-on to your product page into whether your imagery is good enough for someone else's model to render your clothes correctly.

If You Are Feeding Someone Else's Model

Then the work is upstream and it is unglamorous. Consistent garment photography, flat or on-model, shot the same way across the catalog. Accurate colorway data, because a render in the wrong shade is worse than no render. Complete size and fit attributes. Clean backgrounds on the garment images the model will reference. This is product data quality work wearing a different hat, and it is the highest-leverage thing most retailers can do about try-on this year.

If You Are Building Your Own

The architecture moved. Segmentation-first pipelines, which mask the original garment and warp a new one onto the body, were the previous generation. Current work is diffusion-based and largely end to end. CatVTON concatenates the person and garment images spatially and runs self-attention inside a single UNet with no high-level semantic reference. IDM-VTON uses a frozen UNet for low-level garment detail alongside IP-Adapter for semantics, with pose as an additional control signal. Leffa is a dual-UNet that outperforms CatVTON at considerably higher compute. Re-CatVTON now reports beating Leffa on FID, KID and LPIPS on VITON-HD from a single UNet, at a fraction of the memory.

The evaluation vocabulary is worth knowing before a vendor conversation. FID and KID measure distributional realism, LPIPS measures perceptual similarity, SSIM measures structural similarity, and VITON-HD is the standard benchmark. A vendor who cannot state their numbers on a named benchmark is showing you selected outputs.

The Data Nobody Budgets For

Training or fine-tuning a try-on model needs paired data: the garment photographed flat or on a hanger, and the same garment worn, ideally by several body types at several angles. Most catalogs have one of those two. Assembling the other is a photography project with a real cost, and it is the line item that gets discovered after the model selection rather than before it.

The edge cases are where the requirement gets expensive. Loose and draped garments behave differently from fitted ones and need more examples. Patterns have to align across seams or the render reads as obviously fake. Sheer and textured fabrics are the hardest category in the field, because the model has to represent something partially transparent over a body it is also generating.

The Failure Mode That Increases Returns

Try-on is sold on return reduction, and the reported effects are real. But the mechanism is honesty rather than imagery, and that cuts both ways. A system that renders every garment as flattering, that quietly smooths fit on a loose cut or drapes a stiff fabric like jersey, sets an expectation the physical garment cannot meet.

That produces a shopper who was more confident at checkout and more disappointed at delivery, which is a worse return than one caused by uncertainty. The systems that reduce returns are the ones willing to render a garment as slightly unflattering when it is.

For background on the segmentation approach that preceded the diffusion generation, we have written an introduction to the open-source Clothseg repository, which remains a reasonable entry point for understanding the masking step even though current systems handle it differently.

The Rest of the Landscape, Briefly

Seven more applications come up in most fashion AI conversations. Each gets a sentence on what it is and a sentence on the part that decides whether it works, because in every case the second sentence is where the project lives or dies.

  • On-model product imagery. Generating a garment on a model rather than photographing it. The quality ceiling is set by how well the underlying embedding preserves fabric behavior, which is why fine-tuning matters more than prompt engineering.
  • Generative design and sketch-to-image. Diffusion with ControlNet for structure and LoRA adapters trained on a house aesthetic. Useful for ideation, and the output is a mood board rather than a tech pack.
  • Trend forecasting from social imagery. Vision-language models over runway and street-style feeds. The hard part is attribution rather than detection: knowing a silhouette is trending is easy, knowing whether it converts in your price band is not.
  • Demand forecasting and allocation. Attribute-aware models that forecast at the style-color-size level rather than the style level, with weather and calendar features. Size curve accuracy is where the money is and where most models are weakest.
  • Dynamic pricing and markdown optimization. Elasticity modeling per style, informed by catalog embeddings so that similar unsold products inform each other. Fashion's constraint is that the season ends whether the model was right or not.
  • Conversational and agentic shopping. Retrieval over the product catalog with tool calls into inventory and checkout. Retrieval quality is the whole system, and retrieval quality is a product data problem.

A pattern runs through all seven. The modeling in each is well understood and mostly available off the shelf, and the differentiator in each is the quality and structure of the product data underneath it. A demand forecast is only as good as the attribute data it forecasts on. A conversational agent is only as good as what retrieval returns. Trend detection is only useful if it maps onto categories your catalog actually uses.

That is the unglamorous conclusion of most fashion AI evaluations, and it is why the catalog work in the earlier sections tends to gate the rest. Teams looking for the highest-leverage first project usually find it upstream of whichever application they came in asking about.

Email and lifecycle optimization. Send-time and next-purchase prediction from transaction history, typically gradient boosting over customer features with monthly retraining. We have written up the model and pipeline we use for this, and it is the one application on this list where the fashion specificity is low and the transferable ROI is high.

Off the Shelf, Fine-Tune, or Custom

Almost nobody needs to build a vision model from scratch, so the real decision is narrower than build versus buy. It is whether a general model is good enough, whether a domain tune closes the gap, or whether the task needs its own training.

Off the Shelf Is Right When

Your catalog is small enough that errors are individually correctable, your categories are standard, your attribute vocabulary matches common usage, and nobody has committed to an accuracy number. A few thousand SKUs in mainstream apparel categories, with a merchandiser who can fix the misses, is a good fit for a general model. Spending six figures on a custom pipeline for that catalog is how budget gets wasted.

Fine-Tune When

Your catalog is mid-size, your attributes carry meaning general models do not capture, and someone downstream depends on a stated accuracy. The RP2K progression is the argument here: a domain tune moved base performance from 41% to 50%, which is real improvement and still not a system anyone would deploy. Fine-tuning is the right call when you can name the specific distinctions the general model gets wrong, because those are what you train on.

Build Custom When

Your attributes are proprietary, your catalog runs into the millions, your vertical is unusual enough that no public dataset resembles it, or your accuracy requirement is contractual. Jewelry, technical outerwear, footwear sizing and textiles all sit here, because the distinctions that matter commercially are ones no general pretraining corpus contains.

What Fine-Tuning Actually Costs

The compute is rarely the expensive part, which surprises people. Fine-tuning an embedding model on a fashion catalog is measured in hours on a single machine for most catalog sizes, and the cloud bill is not what decides the project.

The expensive parts are labelled data and evaluation. Someone has to assemble examples covering the distinctions that matter, weighted toward the hard ones, and someone has to build an evaluation set that reflects real query conditions rather than clean catalog images. Both require merchandising knowledge rather than ML knowledge, which means they land on the team least able to absorb extra work.

Budget the evaluation set first. A team with a good evaluation set and a mediocre model can improve. A team with a good model and no evaluation set cannot tell whether they are improving, which is a worse position to be in and a much more common one.

Three Questions Worth Asking Any Vendor

What catalog size was the accuracy number measured against, because top-1 on a thousand products and on ten million are different claims. What did the query images look like, because clean catalog photographs on both sides will flatter a system that collapses on customer photographs. And what happens to a product the model is uncertain about, because a system with no confidence threshold and no review queue will publish its worst guesses alongside its best ones and give you no way to tell them apart.

Those three answers separate vendors who have run something in production from vendors who have run a demo. Neither answer is disqualifying on its own. Not having an answer is.

Conclusion

Every application above works. That was not true five years ago and it is the reason the budgets exist. What has not changed is that the version that works in a demo and the version that survives a production catalog are separated by model choice, training data, evaluation discipline and a set of failure modes that only appear at scale.

If you are scoping one of these projects, the useful first move is to write down the accuracy your team would actually accept and the catalog you would measure it against, before talking to anyone. Most of the disappointing outcomes here trace back to nobody having written that down, because without it every demo looks impressive and no two vendor claims are comparable.

The second move is to decide what happens to the products the model is unsure about. Every system described above will be uncertain about some share of any real catalog, and the difference between a system a merchandising team trusts and one they quietly abandon is almost always whether the uncertain cases were surfaced for review or published alongside the confident ones.

Get those two decisions right and most of the technical choices become tractable. Get them wrong and the best model available will still produce a pilot nobody renews.

‍

Scoping a Fashion AI Build?

Width.ai builds the production systems behind visual search, product similarity, SKU recognition and catalog content at scale. We have published our numbers and the datasets they came from, including 92.44% top-1 product similarity against a 10M product catalog and 89% top-1 SKU classification on RP2K.

Tell us your catalog size, your attribute schema, and the accuracy your team needs to sign off on.

Scope your first AI project →

‍

Frequently Asked Questions

What is AI in fashion retail?

Machine learning applied across apparel commerce: visual search and product similarity, shelf and SKU recognition, automated product content, virtual try-on, demand forecasting and personalization. The vision applications are the technically demanding ones, because apparel differs along attributes general models handle poorly.

How is AI used in retail stores?

In physical retail the main applications are shelf and planogram compliance from photographs, out-of-stock detection, and competitive price and assortment audits. All three are the same underlying problem: identifying exact SKUs in a photograph where products are adjacent, angled and partially occluded.

What are the biggest AI use cases in fashion retail?

By deployment volume, product content generation and categorization, because every retailer has a catalog and every catalog has gaps. By technical difficulty, shelf-level SKU recognition. By strategic importance right now, virtual try-on, because the surface it runs on moved outside the retailer's control this year.

How does virtual try-on actually work?

Current systems are diffusion-based. Rather than masking the original garment and warping a new one onto the body, models like CatVTON and IDM-VTON condition a diffusion process on the person image, the garment image and usually a pose signal, generating the result end to end. They are benchmarked on VITON-HD using FID, KID, LPIPS and SSIM.

What is Fashion CLIP and why does it matter?

It is CLIP further trained on fashion imagery, so it separates apparel categories that base CLIP treats as similar. It matters as a demonstration of both the value and the limit of domain tuning: on our shelf recognition evaluation it improved top-1 accuracy from 41% to 50%, which is a real gain and still far from deployable.

Should we buy an off-the-shelf tool or fine-tune our own?

Off the shelf when the catalog is small, the categories are standard and no accuracy number is committed. Fine-tune when you can name the specific distinctions a general model gets wrong and something downstream depends on getting them right. Custom when attributes are proprietary, the catalog runs to millions, or accuracy is contractual.

What data do we need to start?

For retrieval, one clean image per SKU is often enough, which is what makes visual search more approachable than teams expect. For classification and categorization you need labelled examples covering the distinctions you care about, weighted toward the ones the model will find hardest. For try-on, paired garment and on-model photography with consistent framing.

How long does a fashion AI project take to reach production?

For retrieval against an existing catalog, weeks rather than months, because the data requirement is light. For categorization against a deep custom taxonomy, longer, and the time goes into labelling and evaluation rather than modeling. For try-on built in-house, budget a photography programme before the model work, which usually makes it the longest of the three by some margin.

Keep Reading

  • Product Similarity Search with Fashion CLIP: where domain tuning helps and where it stops.
  • Product Image Recognition for Retail Shelves: the full RP2K case study behind the 89% number.
  • Image Embedding Models: choosing an embedding model and when fine-tuning pays for itself.
  • Difference Between SEO and AEO: what product content has to do to be cited rather than just ranked.

‍