### **1. Libraries & Tools** - **NumPy**: Fundamental for numerical operations, used to handle vectorized computations efficiently. - **scikit-learn**: Provides tools like `TfidfVectorizer` for text data and `StandardScaler` for numerical data. - **TensorFlow / PyTorch**: Used for deep learning models (e.g., BERT, Word2Vec) to generate embeddings. - **Pandas**: For data loading and preprocessing (e.g., CSV, JSON). ### **2. Formats** - **JSON**: Use `json` module or `pandas.read_json()` to load data. - **CSV**: Use `pandas.read_csv()` or `numpy.loadtxt()` for numerical data. - **Pickle**: Use `pickle.dumps()`/`pickle.load()` for serializing/deserializing vectorized data. ### **3. Algorithms** - **TF-IDF**: Converts text documents into term frequency-inverse document frequency vectors (scikit-learn). - **Word2Vec**: Neural network-based word embeddings (using `gensim`). - **BERT**: Contextualized embeddings via transformer models (Hugging Face `transformers` library). ### **4. Example Code** #### **TF-IDF (scikit-learn)** ```python from sklearn.feature_extraction.text import TfidfVectorizer import pandas as pd # Load CSV data df = pd.read_csv('data.csv') texts = df['text_column'].tolist() # Vectorize text vectorizer = TfidfVectorizer() tfidf_matrix = vectorizer.fit_transform(texts) ``` #### **Word2Vec (gensim)** ```python from gensim.models import KeyedVectors from sklearn.decomposition import TruncatedSVD # Load pre-trained Word2Vec model word_vectors = KeyedVectors.load_word2vec_format('word2vec.bin', binary=True) # Average word vectors into document vectors doc_vectors = [np.mean([word_vectors[w] for w in doc.split()], axis=0) for doc in texts] ``` #### **BERT (Hugging Face transformers)** ```python from transformers import BertTokenizer, TFBertModel import numpy as np # Load pre-trained BERT model and tokenizer tokenizer = BertTokenizer.from_pretrained('bert-base-uncased') model = TFBertModel.from_pretrained('bert-base-uncased') # Convert text to BERT embeddings def bert_encode(text): input_ids = tokenizer(text, return_tensors='tf')['input_ids'] outputs = model(input_ids) return np.array(outputs.last_hidden_state).mean(axis=1) embeddings = [bert_encode(doc) for doc in texts] ``` ### **5. Notes** - Choose TF-IDF for traditional text vectorization. - Use Word2Vec for semantic word-level representations. - BERT provides contextualized, transformer-based embeddings for complex NLP tasks.