Speeding Up Embedding Comparisons
As I was reading through the Petar Maymounkov and David Mazières paper "Kademlia: A Peer-to-peer Information System Based on the XOR Metric", I was interested in the application of this XOR metric on work I had previously done on clustering embeddings. When implementing the Infer API for Clarity Hub, we focused on using the distance between vectors as a way to cluster utterances together into a topic. We did this via Euclidean distance between vectors, which we had to compare against all other vectors within the space.



