JOURNAL

English Lesson — Tuesday Morning: Embeddings & Vector Search Vocabulary

Daily English practice for tech professionals: embed, vector database, and nearest neighbor search vocabulary with pronunciation and exercises.

Diagram: a text query is embedded into a vector, compared against vectors stored in a vector database, and ranked by nearest neighbor search to return the top matching results

Read with AI

Choose content to copy and paste into your AI assistant. Nothing is sent automatically. CMS content is converted to Markdown; original Markdown is used when available.

Good morning! Today’s AI & Machine Learning session is about embeddings and vector search — the vocabulary behind how a search bar finds the right answer even when you don’t type the exact words that are in the document. About 20 minutes, start to finish.

Word of the Day: Embed

Word: embed (verb)
IPA: /ɪmˈbɛd/
Vietnamese meaning: nhúng — chuyển một đoạn văn bản, hình ảnh hoặc dữ liệu khác thành một vector số, để máy tính có thể so sánh chúng theo ý nghĩa thay vì theo từ khóa chính xác.

Example sentences:

  • “We embed every support ticket into a 1,536-dimension vector before storing it.”
  • “If you embed the query with one model and the documents with another, the similarity scores stop making sense.”
  • “She spent the afternoon embedding the whole knowledge base so search could finally understand paraphrases, not just exact keywords.”

Pronunciation and dictionary links: Cambridge Dictionary — embed · YouGlish — hear “embed” in real videos

How It Fits Together

Diagram: a text query is embedded into a vector, compared against vectors stored in a vector database, and ranked by nearest neighbor search to return the top matching results
A query gets embedded into a vector, then compared against a vector database to find its nearest neighbors. Click the image to zoom in.

In plain English: a sentence goes in, a list of numbers comes out. Similar meanings end up as nearby numbers. “Reset my password” and “I forgot my login” land close together in that number space even though they share almost no words — that’s the whole trick behind semantic search.

Vocabulary Table

PhraseVietnameseExample
vector databasecơ sở dữ liệu vector“We store all 50,000 product embeddings in a vector database so similarity search stays fast.”
cosine similarityđộ tương tự cosine“Cosine similarity between the query and the top result was 0.91, so we trust the match.”
nearest neighborláng giềng gần nhất“The algorithm returns the five nearest neighbors to the query vector.”
semantic searchtìm kiếm theo ngữ nghĩa“Semantic search found the right article even though the user typed completely different words.”
dimensionalitysố chiều“Reducing the dimensionality from 1,536 to 256 cut our storage cost by 80%.”

Pronunciation Guide

embed — /ɪmˈbɛd/ — two syllables, stress on the second one: “im-BED”, not “EM-bed”. The ending rhymes with “bed”, not with “bead”. Say it out loud a few times: im-BED, im-BED, im-BED.

Practice sentence (read aloud 3 times):

“We embed the query, then compare it against every vector in the database.”

Watch your stress on “em-BED” and keep “database” smooth: /ˈdeɪ.tə.beɪs/, three even beats — DAY-ta-base.

Exercise 1: Fill in the Blank

  1. Before we can search by meaning, we need to _______ every document into a vector.
  2. The two closest results had a _______ of 0.87, so we treated them as near-duplicates.
  3. Our _______ stores two million product vectors and answers queries in under 50 milliseconds.
  4. We cut storage costs by lowering the _______ of each vector from 1,536 to 384.
Show answers

1. embed  2. cosine similarity  3. vector database  4. dimensionality

Exercise 2: Translate

Translate these into English using today’s vocabulary:

  1. Chúng tôi nhúng mỗi câu hỏi của khách hàng thành một vector trước khi lưu trữ.
  2. Tìm kiếm theo ngữ nghĩa đã tìm đúng bài viết dù người dùng gõ từ hoàn toàn khác.
  3. Cơ sở dữ liệu vector của chúng tôi trả về năm láng giềng gần nhất trong vài mili giây.
Show answers

1. “We embed every customer question into a vector before storing it.”
2. “Semantic search found the right article even though the user typed completely different words.”
3. “Our vector database returns the five nearest neighbors in a few milliseconds.”

Idiom of the Day: “Needle in a Haystack”

Vietnamese meaning: mò kim đáy bể — tìm một thứ rất nhỏ giữa một khối lượng khổng lồ, gần như bất khả thi nếu làm thủ công.

  • “Finding the one relevant paragraph in ten thousand documents used to be a needle in a haystack — now the embedding model does it in milliseconds.”
  • “Without an index, searching raw text for an exact phrase across millions of rows is like looking for a needle in a haystack.”

Recommended Watching

Today’s Challenge

Pick one real sentence from your own work — a commit message, a ticket, a Slack message — and say it out loud three times using “embed” and “vector database” correctly. If you can explain to a coworker what embedding does in one sentence today, you’ve got it.

Discussion

Comments are reviewed before publication. Your email is kept private.

← Back to allĐọc tiếng Việt