🤖 AI Summary
This study addresses the challenge of semantic modeling in unstructured, noisy, and label-scarce consumer-to-consumer (C2C) e-commerce data—specifically, over 1.2 million multimodal listings (images + text) from OfferUp and Craigslist for used automotive parts. Method: We pioneer the application of OpenFlamingo, a vision-language foundation model, to generate joint image-text embeddings, followed by unsupervised k-means clustering to discover cross-modal semantic patterns without human annotation. Contribution/Results: Experiments demonstrate that OpenFlamingo effectively identifies coherent, semantically consistent clusters corresponding to major automotive part categories—validating its transferability to real-world, non-standard e-commerce settings. However, several disordered clusters emerge, exposing architectural limitations in fine-grained part representation and precise image–text alignment. This work provides empirical evidence and actionable insights for adapting multimodal foundation models to vertical-domain, unstructured transactional data, bridging a critical gap between general-purpose vision-language models and practical C2C marketplace applications.
📝 Abstract
In this paper, we aim to investigate the capabilities of multimodal machine learning models, particularly the OpenFlamingo model, in processing a large-scale dataset of consumer-to-consumer (C2C) online posts related to car parts. We have collected data from two platforms, OfferUp and Craigslist, resulting in a dataset of over 1.2 million posts with their corresponding images. The OpenFlamingo model was used to extract embeddings for the text and image of each post. We used $k$-means clustering on the joint embeddings to identify underlying patterns and commonalities among the posts. We have found that most clusters contain a pattern, but some clusters showed no internal patterns. The results provide insight into the fact that OpenFlamingo can be used for finding patterns in large datasets but needs some modification in the architecture according to the dataset.