🏠 Back to Blog
a vector database is a thing that stores vectors in a database (which has indices and other database semantics)
a vector is a numerical representation of data points (as an array)
vectors place data in a multi-dimensional space
similar things are close together in that space (they have high spatial locality)
vectors allow us to search by meaning instead of exact words
Example: “Animal that is a bird but cannot fly”
Embedding is the act of inserting data into a vector database. Data is ‘chunked’ into fixed size pieces that fit within the embedding model’s context window.
vector databases store embeddings
embeddings are arrays of floating point numbers that represent data (image, text, etc.). These numbers carry meaning.
These numbers are generated by machine learning models
they capture the semantic meaning or features of the original data
An embedding model converts the data into lists of numbers and each list is a vector
words with similar meaning get similar numerical representation
Example:
Dog: [0.8, 0.9, 0.2]
Cat: [0.7, 0.3, 0.9]
Apple: [0.1, 0.3, 0.2]
In the above examples, you can see that dog and cat have closer values than dog and apple or cat and apple
Your query gets converted to a vector
The semantic space is then searched using your vector and matches are returned