Billions of embeddings make vector search a memory problem. Pinterest's Manas retrieval platform tackles it by changing how vectors are represented and…
// the short version
Billions of embeddings make vector search a memory problem. Pinterest's Manas retrieval platform tackles it by changing how vectors are represented and where the search index lives.
// what to take away
- Scalar and product quantization trade vector precision for smaller indexes and cheaper comparisons.
- SPANN keeps routing centroids in memory while fetching promising vector groups from SSD.
- Memory and throughput improvements must be evaluated alongside recall, not in isolation.
// transcript
Pinterest's search data was outgrowing expensive memory. Each Pin becomes an embedding: a vector of numbers learned by a model. Similar content gets similar vectors, so retrieval looks for vectors near the query's vector. But keeping billions of them in memory makes every extra byte expensive. Pinterest shrank those vectors with quantization. Scalar quantization maps each coordinate onto a smaller set of integer values, trading precision for fewer bits. In one benchmark, that shrank a search index from a hundred and twenty-one gigabytes to fifty. More aggressive compression cut it to thirty-two gigabytes, but missed more of the nearest matches. That's the trade: a smaller index is only useful if it still finds relevant Pins. After testing user engagement, Pinterest rolled out quantization across major uses, cutting serving costs by twenty to thirty percent. Smaller numbers also let the processor compare several coordinates in one instruction. Pinterest's scaling approach avoids decoding them back to floating point first, reducing compute per query by ten to fifteen percent. But a smaller index still takes up memory, so they also experimented with moving most of it to solid state drives. The catch is that disk reads take longer, making scattered reads expensive. Their disk search groups similar vectors together, keeping representative vectors, called centroids, in memory. A query searches those centroids first, then fetches promising groups from disk. Balancing the groups keeps one oversized list from holding up a query. To shrink those disk reads, product quantization splits each vector into chunks. Each chunk becomes the code for its nearest representative in a learned codebook. The disk data gets compressed, while the centroids used to choose groups stay at full precision. In their benchmark, that combination handled about four and a half times as many queries per second as uncompressed disk search, with slightly lower recall. Early disk serving experiments used a tenth of the memory compared with in-memory serving. Those disk results are experimental; quantization is already deployed. Keep the routing precise, compress the collection, and fetch only the promising parts.
// source
This explainer is based on Evolving Pinterest's Embedding Retrieval Platform by Pinterest ↗. The original reporting and technical work belong to its publisher.