mirror of
https://gitverse.ru/kpa39l/infrastructure.git
synced 2026-09-29 01:55:03 +00:00
2.4 KiB
2.4 KiB
1. Libraries & Tools
- NumPy: Fundamental for numerical operations, used to handle vectorized computations efficiently.
- scikit-learn: Provides tools like
TfidfVectorizerfor text data andStandardScalerfor numerical data. - TensorFlow / PyTorch: Used for deep learning models (e.g., BERT, Word2Vec) to generate embeddings.
- Pandas: For data loading and preprocessing (e.g., CSV, JSON).
2. Formats
- JSON: Use
jsonmodule orpandas.read_json()to load data. - CSV: Use
pandas.read_csv()ornumpy.loadtxt()for numerical data. - Pickle: Use
pickle.dumps()/pickle.load()for serializing/deserializing vectorized data.
3. Algorithms
- TF-IDF: Converts text documents into term frequency-inverse document frequency vectors (scikit-learn).
- Word2Vec: Neural network-based word embeddings (using
gensim). - BERT: Contextualized embeddings via transformer models (Hugging Face
transformerslibrary).
4. Example Code
TF-IDF (scikit-learn)
from sklearn.feature_extraction.text import TfidfVectorizer
import pandas as pd
# Load CSV data
df = pd.read_csv('data.csv')
texts = df['text_column'].tolist()
# Vectorize text
vectorizer = TfidfVectorizer()
tfidf_matrix = vectorizer.fit_transform(texts)
Word2Vec (gensim)
from gensim.models import KeyedVectors
from sklearn.decomposition import TruncatedSVD
# Load pre-trained Word2Vec model
word_vectors = KeyedVectors.load_word2vec_format('word2vec.bin', binary=True)
# Average word vectors into document vectors
doc_vectors = [np.mean([word_vectors[w] for w in doc.split()], axis=0) for doc in texts]
BERT (Hugging Face transformers)
from transformers import BertTokenizer, TFBertModel
import numpy as np
# Load pre-trained BERT model and tokenizer
tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')
model = TFBertModel.from_pretrained('bert-base-uncased')
# Convert text to BERT embeddings
def bert_encode(text):
input_ids = tokenizer(text, return_tensors='tf')['input_ids']
outputs = model(input_ids)
return np.array(outputs.last_hidden_state).mean(axis=1)
embeddings = [bert_encode(doc) for doc in texts]
5. Notes
- Choose TF-IDF for traditional text vectorization.
- Use Word2Vec for semantic word-level representations.
- BERT provides contextualized, transformer-based embeddings for complex NLP tasks.