Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

🏠 Back to Blog

What is a vector database

  • a vector database is a thing that stores vectors in a database (which has indices and other database semantics)

What is a vector

  • a vector is a numerical representation of data points (as an array)
  • vectors place data in a multi-dimensional space
    • similar things are close together in that space (they have high spatial locality)
  • vectors allow us to search by meaning instead of exact words
    • Example: “Animal that is a bird but cannot fly”

What is embedding?

  • Embedding is the act of inserting data into a vector database. Data is ‘chunked’ into fixed size pieces that fit within the embedding model’s context window.

What does a vector database store?

  • vector databases store embeddings
    • embeddings are arrays of floating point numbers that represent data (image, text, etc.). These numbers carry meaning.
    • These numbers are generated by machine learning models
    • they capture the semantic meaning or features of the original data

From data to numbers

  • An embedding model converts the data into lists of numbers and each list is a vector
  • words with similar meaning get similar numerical representation
  • Example:
    • Dog: [0.8, 0.9, 0.2]
    • Cat: [0.7, 0.3, 0.9]
    • Apple: [0.1, 0.3, 0.2]
    • In the above examples, you can see that dog and cat have closer values than dog and apple or cat and apple

Querying a vector database

  • Your query gets converted to a vector
  • The semantic space is then searched using your vector and matches are returned