mirror of
https://gitverse.ru/kpa39l/infrastructure.git
synced 2026-09-29 18:15:02 +00:00
63 lines
2.4 KiB
Markdown
63 lines
2.4 KiB
Markdown
### **1. Libraries & Tools**
|
|
- **NumPy**: Fundamental for numerical operations, used to handle vectorized computations efficiently.
|
|
- **scikit-learn**: Provides tools like `TfidfVectorizer` for text data and `StandardScaler` for numerical data.
|
|
- **TensorFlow / PyTorch**: Used for deep learning models (e.g., BERT, Word2Vec) to generate embeddings.
|
|
- **Pandas**: For data loading and preprocessing (e.g., CSV, JSON).
|
|
|
|
### **2. Formats**
|
|
- **JSON**: Use `json` module or `pandas.read_json()` to load data.
|
|
- **CSV**: Use `pandas.read_csv()` or `numpy.loadtxt()` for numerical data.
|
|
- **Pickle**: Use `pickle.dumps()`/`pickle.load()` for serializing/deserializing vectorized data.
|
|
|
|
### **3. Algorithms**
|
|
- **TF-IDF**: Converts text documents into term frequency-inverse document frequency vectors (scikit-learn).
|
|
- **Word2Vec**: Neural network-based word embeddings (using `gensim`).
|
|
- **BERT**: Contextualized embeddings via transformer models (Hugging Face `transformers` library).
|
|
|
|
### **4. Example Code**
|
|
#### **TF-IDF (scikit-learn)**
|
|
```python
|
|
from sklearn.feature_extraction.text import TfidfVectorizer
|
|
import pandas as pd
|
|
|
|
# Load CSV data
|
|
df = pd.read_csv('data.csv')
|
|
texts = df['text_column'].tolist()
|
|
|
|
# Vectorize text
|
|
vectorizer = TfidfVectorizer()
|
|
tfidf_matrix = vectorizer.fit_transform(texts)
|
|
```
|
|
#### **Word2Vec (gensim)**
|
|
```python
|
|
from gensim.models import KeyedVectors
|
|
from sklearn.decomposition import TruncatedSVD
|
|
|
|
# Load pre-trained Word2Vec model
|
|
word_vectors = KeyedVectors.load_word2vec_format('word2vec.bin', binary=True)
|
|
|
|
# Average word vectors into document vectors
|
|
doc_vectors = [np.mean([word_vectors[w] for w in doc.split()], axis=0) for doc in texts]
|
|
```
|
|
#### **BERT (Hugging Face transformers)**
|
|
```python
|
|
from transformers import BertTokenizer, TFBertModel
|
|
import numpy as np
|
|
|
|
# Load pre-trained BERT model and tokenizer
|
|
tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')
|
|
model = TFBertModel.from_pretrained('bert-base-uncased')
|
|
|
|
# Convert text to BERT embeddings
|
|
def bert_encode(text):
|
|
input_ids = tokenizer(text, return_tensors='tf')['input_ids']
|
|
outputs = model(input_ids)
|
|
return np.array(outputs.last_hidden_state).mean(axis=1)
|
|
|
|
embeddings = [bert_encode(doc) for doc in texts]
|
|
```
|
|
|
|
### **5. Notes**
|
|
- Choose TF-IDF for traditional text vectorization.
|
|
- Use Word2Vec for semantic word-level representations.
|
|
- BERT provides contextualized, transformer-based embeddings for complex NLP tasks. |