Files
infrastructure/data_vectorization_tools.md
T

63 lines
2.4 KiB
Markdown

### **1. Libraries & Tools**
- **NumPy**: Fundamental for numerical operations, used to handle vectorized computations efficiently.
- **scikit-learn**: Provides tools like `TfidfVectorizer` for text data and `StandardScaler` for numerical data.
- **TensorFlow / PyTorch**: Used for deep learning models (e.g., BERT, Word2Vec) to generate embeddings.
- **Pandas**: For data loading and preprocessing (e.g., CSV, JSON).
### **2. Formats**
- **JSON**: Use `json` module or `pandas.read_json()` to load data.
- **CSV**: Use `pandas.read_csv()` or `numpy.loadtxt()` for numerical data.
- **Pickle**: Use `pickle.dumps()`/`pickle.load()` for serializing/deserializing vectorized data.
### **3. Algorithms**
- **TF-IDF**: Converts text documents into term frequency-inverse document frequency vectors (scikit-learn).
- **Word2Vec**: Neural network-based word embeddings (using `gensim`).
- **BERT**: Contextualized embeddings via transformer models (Hugging Face `transformers` library).
### **4. Example Code**
#### **TF-IDF (scikit-learn)**
```python
from sklearn.feature_extraction.text import TfidfVectorizer
import pandas as pd
# Load CSV data
df = pd.read_csv('data.csv')
texts = df['text_column'].tolist()
# Vectorize text
vectorizer = TfidfVectorizer()
tfidf_matrix = vectorizer.fit_transform(texts)
```
#### **Word2Vec (gensim)**
```python
from gensim.models import KeyedVectors
from sklearn.decomposition import TruncatedSVD
# Load pre-trained Word2Vec model
word_vectors = KeyedVectors.load_word2vec_format('word2vec.bin', binary=True)
# Average word vectors into document vectors
doc_vectors = [np.mean([word_vectors[w] for w in doc.split()], axis=0) for doc in texts]
```
#### **BERT (Hugging Face transformers)**
```python
from transformers import BertTokenizer, TFBertModel
import numpy as np
# Load pre-trained BERT model and tokenizer
tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')
model = TFBertModel.from_pretrained('bert-base-uncased')
# Convert text to BERT embeddings
def bert_encode(text):
input_ids = tokenizer(text, return_tensors='tf')['input_ids']
outputs = model(input_ids)
return np.array(outputs.last_hidden_state).mean(axis=1)
embeddings = [bert_encode(doc) for doc in texts]
```
### **5. Notes**
- Choose TF-IDF for traditional text vectorization.
- Use Word2Vec for semantic word-level representations.
- BERT provides contextualized, transformer-based embeddings for complex NLP tasks.