Ever seen a beautiful perfume bottle online (on TikTok or Instagram) through a review by an influencer or in a shop, but couldn't get a sense of its scent? As a perfume enthusiast, I've been there! That challenge inspired my latest project: a computer vision pipeline that provides deep insights without a single sniff. I wanted to build a system that goes beyond simple recognition. In the messy reality of social media posts and retail displays, perfume bottles are often surrounded by clutter. My pipeline aims to mimic human perception: first, finding the bottle; second, recognizing it; and finally, inferring its core characteristics like its fragrance family (or notes). Here's a look at how this end-to-end system works: Precise Object Detection: I leveraged YOLOv5 to train a model that expertly locates perfume bottles in diverse, real-world images, outputting exact bounding boxes. Fine-Grained Product Identification: The cropped bottle images are then fed into a fine-tuned ResNet50 classifier, capable of distinguishing specific products (e.g., "Chanel No. 5," "YSL Black Opium") even among similar-looking bottles. Fragrance Family Classification: The identified product is then mapped to its corresponding fragrance family (e.g., Floral, Woody, Oriental) using a second-stage ResNet classifier trained on curated scent metadata. To ensure this sophisticated backend was easily accessible, I architected the deployment in two key parts: The core inference logic is exposed via a Flask API, providing a robust and scalable way to integrate the computer vision models. The entire user experience, from image upload to results display, is powered by a Streamlit app, seamlessly hosted on Streamlit Cloud for global accessibility and ease of use. This project showcases how we can bridge the gap between clean datasets and real-world image complexity, leveraging modern deployment practices to deliver a tangible solution. It excites me because it touches on real-world problems in visual search and AI-powered retail, all while helping fellow perfume lovers explore new scents! PS: Still in demo phase so let me know about the bugs, what you think or if you have any feedback? Explore the live demo here: https://lnkd.in/gxp53uCK
AI-Driven Personalization In E-Commerce
Explore top LinkedIn content from expert professionals.
-
-
Exciting Research Alert: Multimodal Semantic Retrieval Revolutionizing E-commerce Product Search! Just came across a fascinating paper from Amazon researchers that tackles a crucial challenge in e-commerce search - integrating both text and image data for better product discovery. >> Key Innovations The researchers developed two groundbreaking architectures: - A 4-tower multimodal model combining BERT and CLIP for processing both text and images - A streamlined 3-tower model that achieves comparable performance with reduced complexity >> Technical Deep Dive The system leverages dual-encoder architecture with some impressive components: - Bi-encoder BERT model for processing text queries and product descriptions - Visual transformers from CLIP for image processing - Advanced fusion techniques including concatenation and MLP-based approaches - Cosine similarity scoring for efficient large-scale retrieval >> Real-world Impact The results are remarkable: - Up to 78.6% recall@100 for product retrieval - Over 50% exact match precision - Significant reduction in irrelevant results to just 11.9% >> Industry Applications This research has major implications for: - E-commerce search optimization - Visual product discovery - Large-scale retrieval systems - Cross-modal product recommendations What's particularly impressive is how the system handles millions of products while maintaining computational efficiency through smart architectural choices. This work represents a significant step forward in making online shopping more intuitive and accurate. The researchers from Amazon have demonstrated that combining visual and textual information can dramatically improve search relevance while maintaining scalability.
-
Google just announced the visual fan-out technique‼️ Visual Search Fan-Out marks a significant leap in how AI interprets and interacts with images in search. Instead of simply identifying objects, the system now understands images in full context, combining visual details with natural language to deliver richer and more relevant results. ✅ Interactive visual exploration You can describe what you’re looking for in plain language, and the AI turns vague ideas into clear visual results. It learns from your follow-up questions to improve and refine the images it shows. ✅ Context-aware shopping Finding products is easier. Instead of using complicated filters, you just describe what you want—like style, fit, or color—and the AI shows matching items using up-to-date product data. ✅ Advanced image understanding The AI looks at the main subject of an image as well as smaller details and background elements. It combines all this information to give more accurate and relevant results. ✅ Flexible, multimodal input You can start your search with text, an image, or a photo. The AI blends these inputs to guide you to the most useful results.
-
Search isn’t about typing anymore. It’s about snapping. 📸 Visual search is transforming how people discover products online. Google Lens handles 12 billion searches per month, and Pinterest Lens sees 600 million. Younger generations, especially Gen Z, prefer visual-first discovery over text. Smartphone cameras + AI = instant product recognition. Visual search drives faster discovery, better decisions, and higher confidence. Google Lens and Pinterest Lens are shaping the future of SEO beyond keywords. Brands with weak visuals or missing structured data risk being invisible. Optimized images, alt text, and product context can boost e-commerce discovery 30%+. Retailers using visual search see 48% faster discovery and 25% higher conversion rates. By 2028, 50% of searches will be visual or voice-driven. Early optimization means brands dominate tomorrow’s AI-powered shopping landscape. Are your images ready for the future of search?
-
Your thumbnail might legitimately outperform your title in AI search results. While everyone obsesses over text optimization, we discovered something unexpected: clients with strong visual assets are getting cited more often in multimodal AI responses. Google I/O 2025 confirmed what we've been seeing: Enhanced multimodal capabilities are making visual search a primary discovery channel. The shift makes sense when you think about user behavior. People are uploading screenshots to ChatGPT, asking Claude to analyze images, and using Google Lens for everything from product identification to problem-solving. But most companies are completely unprepared for visual AI optimization. They're still thinking about images as decoration instead of discoverable content that AI systems can parse and cite. What's actually driving visual AI citations: • Images that directly answer queries at a glance work best. Structure visual content to solve specific problems or demonstrate clear outcomes rather than generic stock photos or logos. • Proper image schema markup using ImageObject schema with detailed alt text, captions, and structured data helps LLMs understand and cite visual content accurately. • Consistent visual authority through unified branding and professional quality across all visual assets. AI systems recognize and favor brands with coherent visual identity. • Context-rich visuals that work standalone while supporting surrounding text. LLMs prefer content that provides clear, actionable information whether viewed independently or with accompanying text. • Systematic visual performance tracking to monitor how images appear in AI responses and search features, then optimize based on actual citation patterns. The opportunity is massive because so few companies are thinking about visual AI optimization yet. The brands that nail this early will dominate multimodal discovery in their categories. Visual content optimized for AI comprehension dramatically increases citation chances in multimodal search results. How are you thinking about visual content for AI discovery? Are you seeing any of your images get referenced in LLM responses yet?
-
Visual search is quietly killing keywords. $41.7B in 2024 → $151.6B by 2032. 12 billion+ searches on Google Lens every month. The next SEO revolution is image understanding. Pinterest Lens hit 600M monthly searches. In 2025, the brands winning are the ones who realized: Google Lens understands "leather handbag in burgundy" better than keywords ever could. Pinterest's multimodal AI generates the right keywords from images automatically. Your customer doesn't want to type "green sofa." They want to grab a photo of a sofa they like and say, "Find me this vibe." When I audit brand websites for visual search readiness, I find: 84% have product images named IMG123.jpg (AI can't understand them) 71% have zero image sitemaps 0% have structured data for visual search Not because they're lazy. Because nobody told them it mattered. But here's what kills me: these same brands are spending $50K/month on Google Ads trying to rank for keywords that matter less every quarter. Meanwhile, visual discovery is just sitting there. WHAT I'D DO DIFFERENTLY: Week 1: Audit your product photography. Are images high-resolution? Do file names use keywords? Do you have image sitemaps? (Most brands: 0/3) Week 2: Rename your product images. Sounds basic. But black-leather-handbag.jpg beats IMG123.jpg by 340% in visual search visibility. Week 3: Build image sitemaps. Takes 2 hours. ROI? 30% engagement boost within 3 months based on early data. Week 4: Implement AR try-on for your top 5 SKUs. Expected ROI: 200% conversion increase. (I'm not exaggerating, PerfectCorp data backs this) Week 5: Set up Pinterest Lens optimization. 600M monthly searches. 36% of your audience starts there. If you're not there, you're dead. By Month 2, you'll start seeing visual search traffic. By Month 3, you won't go back to keyword-only strategies. And AI is getting so good at it, text-based search feels like a flip phone in a smartphone world. IF YOU TAKE ONE THING FROM THIS: Stop asking, "How do we rank for keywords?" Start asking, "How do we get discovered when customers snap a photo of what they want?" The answer is visual search. And if you're not ready, your competitor will be. #Marketingexpert #visuals #GTMLeader
-
Typing for product search is no longer required. For years, Google Lens has been the go-to tool for visual search. You point your camera at something, and Google shows you similar images or links across the web. But here’s where Amazon Lens Live changes the game: it skips the middle step. Instead of sending customers to multiple websites, blogs, or stores, Amazon shows real-time product matches directly inside the app, alongside prices, reviews, and add-to-cart options. It’s like Google Lens plus checkout in one tap. What’s changing right now: ➡️ Visual matching matters more → If your hero images don’t align with what Lens detects, you might not appear at all ➡️ Reviews are front and center → Rufus highlights customer sentiment in real time ➡️ Structured data drives visibility → Clean attributes help Rufus compare products accurately ➡️ Competition gets sharper → Customers will see your product right next to your competitors’, instantly So one thing is for sure, Amazon is compressing the path from intent → discovery → purchase into a single interface. That changes where and how visibility is won: • Products without clean data, optimized attributes, and strong imagery will get bypassed instantly. • Reviews and ratings will matter even more because Rufus summarizes them directly in the shopping flow. • Visual alignment becomes a ranking factor. If your hero images don’t match what Lens detects, your products might not even show up. This isn’t just a new feature. It’s an early signal of where Amazon is taking shopping: AI agents curating matches, recommending products, and deciding visibility based on data quality. Amazon isn’t just making visual search easier, it’s moving shopping closer to instant decision-making. Imagine spotting a bag on TikTok, scanning it with Lens Live, and buying the exact match without leaving Amazon. That’s where we’re heading. The shift here is subtle but big: it’s no longer just about being found. It’s about being matched, by AI, in real time. If Google Lens taught us how to search visually, Amazon Lens Live is teaching customers how to shop visually. #Amazon #AIShopping #ProductDiscovery #EcommerceStrategy
-
We’ve all been there during the holidays: trying to find the perfect gift based on a vague idea, like "something with a vintage vibe" or a "cozy sweater in a trending color,” but the right keywords just don't exist. With a major new update to AI Mode in Google Search, you can now simply show or tell Google what you’re thinking and get rich, visual results that you can instantly shop. This is a game changer for everything from designing a room, to completing an outfit, to figuring out the right gift for that person who’s hard to shop for. Just in time for the holiday shopping season. This breakthrough visual search experience is rooted in combining our visual understanding (Lens and Image search) with Gemini 2.5’s advanced multimodal capabilities. For advertisers, this means two critical shifts: Richer Intent Signals: Consumers are moving from vague keywords to detailed, natural language descriptions of their intent ("I want more ankle length"). This gives marketers significantly richer data signals to optimize campaigns and ensure their ads are perfectly relevant. Visual Content is Your New Keyword: Your product images, high-quality visuals, and associated metadata are now more important than ever. The new "visual search fan-out" technique allows AI Mode to understand subtle details within an image, meaning brands must prioritize comprehensive, structured content to ensure their products are discoverable when the customer is searching by image or "vibe." AI is making search more intuitive and shoppable than ever. Read the full announcement here: https://bit.ly/46RlgAH #GoogleSearch #AIMode #HolidayShopping
-
Google’s latest AI shopping update adds a visual “fan-out” mode that surfaces product images directly within search. It’s yet another clear sign of where shopping is headed. All signs point away from keyword-based towards context-based discovery. The KEY for consumer brands: these new visual experiences only work when your underlying product data is clean, structured, and trusted. AI can only match “that striped blue shirt with organic cotton” if the product feed actually includes structured attributes for color, pattern, and material, tied to verified sources. Without that metadata, the model doesn’t surface you. At Novi, this is exactly what we’re solving: helping brands and retailers turn product attributes, claims, and certifications into verified, machine-readable data that AI systems (like Google’s) can interpret correctly. It’s the key to preventing hallucinated search results that erode consumer trust fast. Visual search makes product storytelling feel effortless and behind every effortless experience is hard data work. First mover advantage folks! The brands investing in structured, trustworthy, authoritative product data now will be the ones whose products surface first in AI-powered discovery.
-
🚀 Amazon has just taken visual search to another level. They’ve launched Lens Live - an AI-powered tool that lets shoppers point their phone at an item and instantly see real-time matches in a swipeable carousel. No more snapping, uploading, or scanning barcodes. What makes this bigger is the integration with Rufus, Amazon’s generative AI shopping assistant. That means shoppers don’t just see “similar items” - they get product summaries, comparisons, and answers to their questions in the moment. 🔎 Think about what this means: - Faster product discovery - Less friction from inspiration to purchase - AI shaping the buying decision inside Amazon’s ecosystem For brands, this isn’t just a feature update - it’s a reminder that Amazon is redefining how shoppers discover products. If your content, imagery, and product data aren’t optimised, you’re already behind.